Method for improving the overall performance of cas9-based genetic editors
Patent Information
- Application Number
- PCT/US2025/017868
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-28
- Filing Date
- 2025-02-28
- Publication Date
- 2025-10-30
AI Technical Summary
Existing Cas9-based genetic editors face challenges in enhancing specificity and accuracy due to the lack of experimental evidence supporting the role of the REC lobe in Cas9 evolution, leading to off-target cleavage and inefficiencies in gene editing, particularly with large cargo delivery.
Incorporation of a Streptococcus thermophilus Cas9 (StlCas9) recognition (REC) domain into the N-terminus of Streptococcus pyogenes Cas9 (SpCas9), either as a nickase or with a linker, to expand the Cas9 scaffold and improve binding specificity.
Enhances the precision and accuracy of Cas9-based gene editing by reducing off-target effects and enabling more efficient delivery of larger cargo through lipid nanoparticle methods.
Smart Images

Figure US2025017868_30102025_PF_FP_ABST
Abstract
Description
[0001] METHOD FOR IMPROVING THE OVERALL PERFORMANCE OF CAS9-BASED GENETIC EDITORS
[0002] Related Application
[0003] This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 558,998, filed February 28, 2024; the entire contents of which is hereby incorporated by reference.
[0004] Statement of Rights
[0005] This invention was made with government support under grant UG3TR002636 awarded by the National Institutes of Health. The government has certain rights in the invention.
[0006] Background of the Invention
[0007] The Cas9-diversification in size provides rich materials for studying the evolution (Altae-Tran, Kannan et al. 2021, Altae-Tran, Kannan et al. 2023). The hypothesis on Cas9 evolution suggests that Cas9 originated from an IscB ancestor and underwent a complex evolution from small to large sizes, involving the acquisition of various domains. Among these domains, the recognition (REC) lobe is a critical component. Moreover, the REC lobe exhibits a high degree of conservation and uniqueness across various RNA-guided endonuclease proteins (Nishimasu, Ran et al, 2014). Acquisition of the REC lobe is predicted to play a crucial role in enhancing the specificity of the parent, contributing to the overall evolution of Cas9 (Kapitonov, Makarova et al. 2016, Makarova, Wolf et al, 2020, Altae-Tran, Kannan et al. 2021). A pair of mutual comparisons have not been found available for researching this question. Despite a lot of effort over the past decade to minimize Cas9 off-target cleavage through introducing amino acid substitutions (Kleinstiver, Pattanayak et al, 2016, Slaymaker, Gao et al. 2016, Chen, Dagdas et al, 2017, Casini, Olivieri et al, 2018, Lee, Jeong et al, 2018, Kim, Kim et al. 2023). there is currently no direct experimental evidence supporting the hypothesis that REC acquisition plays a critical role in enhancing specificity (FIG. 1A). Meanwhile, scientists paid excessive attention to exploring and pursuing more hypercompact editors including miniCas9 variants to overcome the size limitations associated with viral-based delivery in gene editing (Wang and Doudna 2023, Wu, Liu et al. 2023). This trend has increasingly hindered understanding of Cas9 protein expansion (FIG. 8A). Encouragingly, as a synthetic materialbased delivery method, lipid nanoparticle delivery has emerged as a safer and more efficient solution for the delivery of large cargo (Gillmore, Gane et al, 2021, Musunuru, Chadwick et ah 2021, Madigan, Zhang et al. 2023, Wang and Doudna 2023). Exploring REC lobe expansion not only can aid in refining the Cas9 evolutionary hypothesis and deep understanding of Cas9 structure, but also contributes significant guidance for optimizing Cas9-based gene editors to improve accuracy and precision.
[0008] Summary of the Invention
[0009] In some aspects, provided herein is a Cas9 variant comprising a Streptococcus pyogenes Cas9 (SpCas9) that comprises an insertion of a Streptococcus thermophilus Cas9 (StlCas9) recognition (REC) domain.
[0010] In some embodiments, the StlCas9 REC domain is inserted at the N-terminus of the SpCas9 REC domain. In some embodiments, the StlCas9 REC domain is inserted between the SpCas9 Bridge helix (BH) domain and the SpCas9 REC domain. In some embodiments, the StlCas9 REC domain is inserted between residues 93D and 94D of the SpCas9. In some embodiments, the SpCas9 is a wild-type SpCas9. In some embodiments, the SpCas9 comprises an amino acid sequence of SEQ ID NO: 1. In some embodiments, the SpCas9 is a SpCas9 nickase (nSpCas9). In some embodiments, the SpCas9 nickase is nSpCas9 (D10A).
[0011] In some embodiments, the StlCas9 REC domain comprises an RECI domain and an REC2 domain. In some embodiments, the StlCas9 REC domain comprises residues 74-466 of StlCas9. In some embodiments, the StlCas9 REC domain comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the sequence of SEQ ID NO: 19. In some embodiments, the StlCas9 REC domain comprises the amino acid of SEQ ID NO: 19. In some embodiments, the Cas9 variant does not comprise a BH domain from StlCas9.
[0012] In some embodiments, the StlCas9 REC domain is connected to the N-terminus of the SpCas9 REC domain via a linker. In some embodiments, the linker comprises a linker from Staphylococcus aureus (SaCas9). In some embodiments, the linker from SaCas9 comprises residues T205 to D223 of SaCas9. In some embodiments, the linker from SaCas9 comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the sequence of SEQ ID NO: 20. In some embodiments, the linker from SaCas9 comprises the amino acid sequence of SEQ ID NO: 20.
[0013] In some embodiments, the SpCas9 comprises an insertion of the amino acid sequence of SEQ ID NO: 8. In some embodiments, the Cas9 variant comprises the amino acid sequence of SEQ ID NO: 21. In some aspects, provided herein is a nucleic acid encoding the Cas9 variant described herein. In some embodiments, the nucleic acid is a DNA or RNA.
[0014] In some aspects, provided herein is a vector comprising a nucleic acid described herein.
[0015] In some aspects, provided herein is a composition or kit, comprising:
[0016] (a) a Cas9 variant described herein, a nucleic acid described herein, or a vector described herein; and
[0017] (b) a Cas9 guide RNA (gRNA) .
[0018] In some embodiments, the Cas9 gRNA is a spCas9 gRNA.
[0019] In some aspects, provided herein is a composition or kit, comprising:
[0020] (a) a Cas9 variant described herein, a nucleic acid described herein, or a vector described herein;
[0021] (b) a Cas9 crRNA; and
[0022] (c) a Cas9 tracrRNA.
[0023] In some embodiments, the Cas9 crRNA is a spCas9 crRNA, and the Cas9 tracrRNA is a spCas9 tracrRNA.
[0024] In some aspects, provided herein is a method of cleaving or modifying a target DNA, comprising: contacting the target DNA or a cell comprising the target DNA with:
[0025] (a) a Cas9 variant described herein, a nucleic acid described herein, or a vector described herein; and
[0026] (b) a Cas9 guide RNA (gRNA) ; wherein the Cas9 variant and the Cas9 gRNA form a complex that cleaves or modifies the target DNA.
[0027] In some embodiments, the Cas9 gRNA is a spCas9 gRNA.
[0028] In some aspects, provided herein is a method of cleaving or modifying a target DNA, comprising: contacting the target DNA or a cell comprising the target DNA with:
[0029] (a) a Cas9 variant described herein, a nucleic acid described herein, or a vector described herein;
[0030] (b) a Cas9 crRNA; and
[0031] (c) a Cas9 tracrRNA; wherein the Cas9 variant, the Cas9 crRNA, and the Cas9 tracrRNA form a complex that cleaves or modifies the target DNA. In some embodiments, the Cas9 crRNA is a spCas9 crRNA, and the Cas9 tracrRNA is a spCas9 tracrRNA.
[0032] In some aspects, provided herein is a method of cleaving or modifying a target DNA, comprising contacting the target DNA or a cell comprising the target DNA with a composition or kit described herein.
[0033] In some aspects, provided herein is a Cas9-guided base editor comprising a base editor linked to a Cas9 variant, wherein the Cas9 variant comprises a SpCas9 nickase (nSpCas9) comprising an insertion of a StlCas9 REC domain.
[0034] In some embodiments, the StlCas9 REC domain is inserted at the N-terminus of the nSpCas9 REC domain. In some embodiments, the StlCas9 REC domain is inserted between the BH domain and the REC domain of the nSpCas9. In some embodiments, the StlCas9 REC domain is inserted between residues 93D and 94D of the nSpCas9. In some embodiments, the nSpCas9 is nSpCas9 (D10A). In some embodiments, nSpCas9 comprises an amino acid sequence of SEQ ID NO: 22. In some embodiments, the StlCas9 REC domain comprises an RECI domain and an REC2 domain. In some embodiments, the StlCas9 REC domain comprises residues 74-466 of StlCas9. In some embodiments, the StlCas9 REC domain comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the sequence of SEQ ID NO: 19. In some embodiments, the StlCas9 REC domain comprises the amino acid of SEQ ID NO: 19.
[0035] In some embodiments, the StlCas9 REC domain is connected to the N-terminus of the SpCas9 REC domain via a linker. In some embodiments, the linker comprises a linker from Staphylococcus aureus (SaCas9). In some embodiments, the linker from SaCas9 comprises residues T205 to D223 of SaCas9. In some embodiments, the linker from SaCas9 comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the sequence of SEQ ID NO: 20. In some embodiments, the linker from SaCas9 comprises the amino acid sequence of SEQ ID NO: 20. In some embodiments, the Cas9 variant comprising half (0.5x), single (lx), or two (2x) BH domains immediately upstream of the StlCas9 REC domain. In some embodiments, the Cas9 variant comprising a single BH domain immediately upstream of the StlCas9 REC domain.
[0036] In some embodiments, the nSpCas9 comprises an insertion of an amino acid sequence selected from SEQ ID NOs: 3, 7 and 8. In some embodiments, the nSpCas9 comprises an insertion of the amino acid sequence of SEQ ID NOs: 8. In some embodiments, the cas9 variant comprises an amino acid sequence selected from SEQ ID NOs: 23-25. In some embodiments, the cas9 variant comprises the amino acid sequence selected from SEQ ID NOs: 24.
[0037] In some embodiments, the base editor is linked to N-terminus of the Cas9 variant. In some embodiments, the base editor is an adenine deaminase. In some embodiments, the adenine deaminase is selected from TadA8e, TadA9, TadA8e(N46L), and TadA8e(VI06W).
[0038] In some aspects, provided herein is a nucleic acid encoding the Cas9-guided base editor described herein. In some embodiments, the nucleic acid is a DNA or RNA.
[0039] In some aspects, provided herein is a vector comprising a nucleic acid described herein.
[0040] In some aspects, provided herein is a composition or kit, comprising:
[0041] (a) a Cas9-guided base editor described herein, a nucleic acid described herein, or a vector described herein; and
[0042] (b) a Cas9 guide RNA (gRNA).
[0043] In some embodiments, the Cas9 gRNA comprises a gRNA-scaffold selected from SpCas9 gRNA scaffold, SaCas9 gRNA-scaffold, StlCas9 gRNA-scaffold, and FnCas9 gRNA-scaffold. In some embodiments, the Cas9 gRNA comprises a guide sequence of 8 nucleotides (nt), lOnt, 12nt, 20nt, 35nt, or 40 nt in length.
[0044] In some aspects, provided herein is a composition or kit, comprising:
[0045] (a) a Cas9-guided base editor described herein, a nucleic acid described herein, or a vector described herein;
[0046] (b) a Cas9 crRNA; and
[0047] (c) a Cas9 tracrRNA.
[0048] In some embodiments, the Cas9 crRNA is selected from SpCas9 crRNA, SaCas9 crRNA, StlCas9 crRNA, and FnCas9 crRNA, and wherein the Cas9 tracrRNA is selected from the same species as the Cas9 crRNA. In some embodiments, the Cas9 crRNA comprises a guide sequence of 8 nt, lOnt, 12nt, 20nt, 35nt, or 40 nt in length.
[0049] In some aspects, provided herein is a method of modifying a target DNA or RNA comprising: contacting the target DNA or RNA or a cell comprising the target DNA or RNA with:
[0050] (a) a Cas9-guided base editor described herein, a nucleic acid described herein, or a vector described herein; and
[0051] (b) a Cas9 guide RNA (gRNA) ; wherein the Cas9-guided base editor and the Cas9 gRNA form a complex that modifies the target DNA or RNA.
[0052] In some embodiments, the Cas9 gRNA comprises a gRNA-scaffold is selected from SpCas9 gRNA-scaffold, SaCas9 gRNA-scaffold, StlCas9 gRNA-scaffold, and FnCas9- gRNA-scaffold. In some embodiments, the Cas9 gRNA comprises a guide sequence of 8 nt, lOnt, 12nt, 20nt, 35nt, or 40 nt in length.
[0053] In some aspects, provided herein is a method of modifying a target DNA or RNA comprising: contacting the target DNA or a cell comprising the target DNA or RNA with:
[0054] (a) a Cas9-guided base editor described herein, a nucleic acid described herein, or a vector described herein; and
[0055] (b) a Cas9 crRNA; and
[0056] (c) a Cas9 tracrRNA; wherein the Cas9-guided base editor, the Cas9 crRNA, and the Cas9 tracrRNA form a complex that modifies the target DNA or RNA.
[0057] In some embodiments, the Cas9 crRNA is selected from SpCas9 crRNA, SaCas9 crRNA, StlCas9 crRNA, and FnCas9 crRNA, and wherein the Cas9 tracrRNA is selected from the same species as the Cas9 crRNA. In some embodiments, the Cas9 crRNA comprises a guide sequence of 8 nt, lOnt, 12nt, 20nt, 35nt, or 40 nt in length.
[0058] In some aspects, provided herein is a method of modifying a target DNA or RNA comprising contacting the target DNA or a cell comprising the target DNA or RNA with a composition or kit described herein.
[0059] In some embodiments, modifying the target DNA or RNA comprises inducing a base change of the target DNA or RNA. In some embodiments, the base change is a A-to-G or A- to-I change. In some embodiments, the cell is a mammalian cell.
[0060] In some aspects, provided herein is a Cas9 variant comprising an enlarged SpCas9 that comprises an insertion of one or more domains selected from a non-catalytic BH domain and a REC domain. In some embodiments, the enlarged SpCas9 comprises an insertion of one or more domains selected from a non-catalytic StlCas9 BH domain and a StlCas9 REC domain. In some embodiments, the enlarged SpCas9 comprises an insertion of a StlCas9 REC domain.
[0061] In some aspects, provided herein is a method of producing a Cas9 variant described herein. In some aspects, provided herein is a method of producing a Cas9-guided base editor described herein.
[0062] In some aspects, provided herein is a method of regulating activity and / or specificity of a catalytic domain linked to a Cas9 scaffold comprising expanding the Cas9 scaffold by inserting one or more domains selected from non-catalytic BH domain and REC domain. In some embodiments, the catalytic domain is an N-terminal catalytic domain. In some embodiments, the catalytic domain is a base editor. In some embodiments, the catalytic domain has reduced byproducts. In some embodiments, the Cas9 scaffold is a SpCas9. In some embodiments, the method comprises expanding the Cas9 scaffold by inserting one or more domains selected from a non-catalytic BH domain and a REC domain from StlCas9.
[0063] In some aspects, provided herein is a method of repositioning an N-terminal catalytic domain linked to Cas9, comprising adjusting the length of a BH domain. In some embodiments, the BH domain is a half (0.5x), single (lx), or two (2x) BH domains.
[0064] In some aspects, provided herein is a method of reducing bystander editing of a deaminase-based base editor, comprising modifying a Cas9 scaffold. In some embodiments, modifying the Cas9 scaffold comprises inserting one or more domains selected from a non- catalytic BH domain and a REC domain into the Cas9 scaffold. In some embodiments, modifying the Cas9 scaffold comprises inserting one or more domains selected from a non- catalytic StlCas9 BH domain and a StlCas9 REC domain into the Cas9 scaffold. In some embodiments, the Cas9 scaffold is a nSpCas9.
[0065] In some aspects, provided herein is a method of improving Cas9 binding specificity, comprising enlarging Cas9 by inserting one or more domains selected from a non-catalytic BH domain and a REC domain. In some embodiments, the method described herein comprises inserting a StlCas9 REC domain to the Cas9.
[0066] In some aspects, provided herein is a Cas9 variant, comprising nSpCas9 (D10A) that comprises an insertion of BH and REC domains of StlCas9. In some embodiments, the nSpCas9 (D10A) comprises an insertion of the amino acid sequence of SEQ ID NO: 3. In some embodiments, the nSpCas9 (D10A) comprises an insertion of the amino acid sequence of SEQ ID NO: 3, 7, or 8 between residues 93D and 94D of the nSpCas9 (D10A). In some embodiments, the Cas9 variant comprises the amino acid sequence selected from SEQ ID NOs: 23-25. In some embodiments, the Cas9 variant is used as a binding scaffold to generate a base editor.
[0067] Brief Description of the Drawings FIGS. 1A-1M show comparative analysis of various Cas9 sequences and design artificial REC expansion of SpCas9. (A) Schematic illustration of REC expansion in the Cas9- evolution-hypothesis. (B) Domain size of IscB and Cas9s. Crystal structures of all the protein have been identified. REC, recognition (REC) lobe; BH, bridge helix; PAM, protospacer- adjacent motifs; PI, PAM-Interacting Domain. (C) Violin plot illustrating distribution of Cas9- sizes in various species. One pot means one sequence, n, numbers of sequences in each group. (D) Highlighting evolution of Cas9 proteins in different species. F, Francisella,' Sp, Streptococcus pyogenes,' Stl, Streptococcus thermophilus,' Sa, Staphylococcus aureus; Nme, Neisseria meningitidis and Cj, Campylobacter jejuni. (E) A crystal structure of StlCas9 (protein data bank ID 6m0v). (F) A crystal structure of SaCas9 (protein data bank ID 5axw). (G) Schematic of REC expansion from SpCas9. Insert positions are shown in cells. (H) workflow for testing Cas9 variants activity in HEK293T cells. (I) Activities induced by SpCas9 and variants. Data are presented as mean ±s.d. (n=3). P value was determined by two-way ANOVA. (J) Testing influences from BHs with different lengths. Data are presented as mean ±s.d.(n=3). P value was determined by two-way ANOVA. (K) Variant Left-REC12 mediated EGFP disruption at different plasmid dosage. (L) Disruption activities of SpCas9 and Left- REC12 on endogenous GFP site in HEK293-deGFP cells. Data are presented as mean ±s.d.(n=3). (M) Schematic illustrating the SpCas9 and hypothetical giant SpCas9 (GS-Cas9).
[0068] FIGS. 2A-2H show investigating GS-Cas9 performance on human genome. (A) Workflow for systemic testing GS-Cas9 activity in HEK293T cells. (B) Indel rates generated by SpCas9 and GS-Cas9 at two loci in HEK293T cells. Data are presented as mean ±s.d. (SpCas9, n=2; GS-Cas9, n=3). (C, D) Representative sequences of the human EMX1 and P205 loci targeted by GS-Cas9. sgRNA target site and PAM are indicated by blue and purple respectively. Below, selected sequences showing representative modified reads. (E, F) Average indel length at EMX1 and P205 sites from (B). (G) Targeted deep-sequencing analysis of off- target sites for EMX1 site Data are presented as mean ±s.d. (SpCas9, n=2; GS-Cas9, n=3) (H) Modified levels with guide sequences containing double-base mismatches. Data are presented as mean ±s.d. (n=3).
[0069] FIG. 3 shows a crystal structure of SpCas9 based adenosine base editor (protein data bank ID 6vpc). ABE8e contains an engineered deoxyadenosine deaminase TadA from Escherichia coli and a SpCas9 (D10A) nickase.
[0070] FIGS. 4A-4L show REC expansion could mediate catalytic domain fused to N-end of SpCas9. (A) Schematic illustrating the ABE8e and hypothetical GS-ABE8e). (B) Illustrating work-process of cpREADAR assay. (C) Representative sequences of reporter plasmid pCMV- EGFP(W58stop) targeted by SpCas9 or GS-Cas9. Below, guide sequences with different lengths. (C, D) Activation of EGFP mediated by various base editors. ABE8e / LBR12, BH- REC12 domain was inserted into N-end of REC domain in nSpCas9(D10A); P102, BH from StlCas9 was removed from ABE8e / LBR12; P103, half of BH from StlCas9 was removed from ABE8e / LBR12. Data are presented as mean ±s.d. (n=3). (E, F) Disrupting EGFP with DNA editing mediated by various base editors. Data are presented as mean ±s.d. (n=3). (G) REC expansion could be expanded to different base editors. Data are presented as mean ±s.d. (n=3). (H) Activation of EGFP mediated by various base editors loaded with gRNAs carrying guide sequences in various lengths. Data are presented as mean ±s.d. (n=3), P value was determined by two-way ANOVA. (I, H) Cumulative A«T to G*C base editing efficiencies and indel formation at human genomic EMX1 site in HEK293T cells treated with various base editors loaded with 20nt or with 35nt guide RNA. Data are presented as mean ±s.d. (n=3). (K) Average A-to-I RNA editing frequencies by nSpCas9, ABE8e and GS-ABE8e mutants among 77 adenosines in IP90 mRNA transcripts. Data are presented as mean of three repeats. (L) MFI disruption activities with SpCas9, ABE8e, ABE9, and GS-ABE9.
[0071] FIGS. 5A-5D show EGFP-disrupting activities of GS-ABE8e loaded with various gRNAs in human HEK293T cells. (A) Targeting ABE8e and GS-ABE8e to a representative EGFP locus in HEK293FT cells with gRNAs containing guides of various lengths. Data are presented as mean ±s.d. (n=3). (B) EGFP disruption with guide sequences containing consecutive transversion mismatches. Data are presented as mean ±s.d. (n=3). (C) Investigating base editors’ activities on out of protospacer A-to-G editing. lOng of EGFP(Q81 stop) plasmid was used for each well. (D) Plasmid based orthogonal R-loop assay. Here, EGFPstop was activated by Cas9-independent off-target DNA editing. Data are presented as mean ±s.d. (n=3). P value was determined by two-way ANOVA.
[0072] FIGS. 6A-6R show systematically comparing ABE8e and GS-ABE8e gene editing performance in HEK293T cells. (A- 1) Evaluation of the A-to-G editing efficiencies of ABE8e and GS-ABE8e at nine representative endogenous genomic sites in HEK293T cells. Data are presented as mean ±s.d. (n=3). (J) Indel formation in HEK293T cells treated as described in A- I and FIG. 24. Data are presented as mean ±s.d. (n=3). (K) Comparison of indels induced by ABE8e and GS-ABE8e at 11 target sites from (J). Each data point represents the average indel frequency at each site calculated from 3 independent experiments. P value was determined by two-tailed Student’s t-test. (L) Frequencies of A-to-G by ABE8e or GS-ABE8e editor across the protospacer positions 1-20 from (A-I) and FIG. 34. (M) Cas9-dependent DNA off-target analysis comparing ABE8e and GS-ABE8e at site P206. (N) Indel formation at two off-target sites in (M). (0) Cas9-dependent DNA off-target analysis comparing ABE8e and GS-ABE8e at site EMX1. (P) Indel formation at off-target sites in (O). (Q, R) On-target:off-target editing ratios for two sites in (M) and (O). For all plots, bars represent mean values, and error bars represent the s.d. of three independent biological replicates.
[0073] FIGS. 7A-7C show Cas9-independent DNA off-target analysis of the cumulative adenosine editing induced by ABE8e and GS-ABE8e. (A) Orthogonal R-loop assay overview. (B) Cas9-independent off-target A*T-to-G*C editing frequencies detected by the orthogonal R- loop assay at each R-loop site with dSaCas9 and a SaCas9 sgRNA. Each R-loop was performed by co-transfection of ABE8e or GS-ABE8e, and an SpCas9 sgRNA targeting site P204 with dSaCas9 and a SaCas9 sgRNA targeting R-loops 1-5, respectively. Amplicons of site P204 were analyzed by high-throughput sequencing. For all plots, bars represent mean values, and error bars represent the s.d. of three independent biological replicates. (C) On-target base editing efficiencies for ABE8e and GS-ABE8e in HEK293T cells at site P204 in HEK293T cells for the orthogonal R-loop assay. Amplicons of site P204 were analyzed by Sanger sequencing. For all plots, dots represent individual biological replicates and bars represent mean ±s.d. of three independent biological replicates.
[0074] FIGS. 8A-8D show an overview of research status of RNA- guided DNA nucleases . (A) Representative RNA-guided endonucleases which characteristics and crystal structures have been experimentally identified. There are no RNA-guided endonuclease with lengths larger than 1700aa identified to date. (B) percentage of proteins with various lengths. Original data were quoted from previous literature (Altae-Tran, Kannan et al, 2021). These data contain all IsrB, IscB and Cas9. (C) Characteristics of crystally identified IscB and Cas9s. FnCas9 (PDB:5B2O), SpCas9 (PDB:4OO8), StlCas9 (PDB:6M0V), SaCas9 (PDB:5AXW), NmeCas9 (PDB:6JDQ), CjCas9 (PDB: 6JOO), OgeuIscB (PDB:7UTN). (D) Classifications and characteristics of identified Cas9s.
[0075] FIG. 9 shows an analysis of sequences in Table 1. The above, alignment analysis of sequences large than 1629aa with FnCas9. Below, illustrating sizes of Cas9 proteins and predicted-REC domains.
[0076] FIG. 10 shows an alignment analysis of Cas9 sequences from Francisella organism. Clustal Omega 1.2.2 algorithm was used. Purple frame means predicted REC expansion regions. Note: KN046796.1, length 749aa; JOOV01000026.1, length U23aa, REC 794aa; AMPPO 1000018.1, length 1125aa, REC 796aa; CP002557.1 FnCas9, length 1629aa, REC 776aa; DS264128.1, length 1629aa, REC 776aa; CP018093.1, length 1639aa, REC 785aa; CP002558.1, length 1646aa, REC 785aa; CP041030.1, length 1695aa, REC 841aa. FIG. 11 shows an alignment analysis of Cas9 sequences from Streptococcus pyogenes organism. Clustal Omega 1.2.2 algorithm was used. Note: LVYNO 1000125.1, length 509aa; APMZO 1000018.1, length 986aa; AURU01000042.1, length 1029aa; AAFV01000057.1, length 1052aa; NGQ JO 1000001.1, length 1179aa; JHTP01000061. 1, length 1188aa; CABFEC010000007.1, length 1218aa; CAAIGD010000001.1, length 1240aa; CABEWD0 10000001.1, length 1245aa; C AAHQ JOI 0000001. 1, length 1263aa; CABFBN010000008.1, length 1296aa; CABEXB010000001.1, length 1336aa, rec 624aa; CABFFS010000018.1, length 1341aa, rec 624aa; LS483321.1, length 1346aa, rec 624aa; CABFHG010000001.1, length 1349aa, rec 624aa; CABFEUO 10000001.1, length 1359aa, rec 624aa; Cas9-SpCas9, length 1368aa, rec 624aa.
[0077] FIG. 12 shows an alignment analysis of Cas9 sequences from Streptococcus thermophilus organism. Clustal Omega 1.2.2 algorithm was used. Two boxes indicated REC expansion regions. Note: CM003139.1, length 565aa; AZTM01000060.1, length 584aa; AGFN01000315.1, length 595aa; CM003139.1, length 629aa; LR822031.1, length 718aa; CP047191.1, length 743aa; LVWW01000001.1, length 766aa; CM002370.1, length 940aa; ALIL01000052.1, length 986aa, REC 393aa; Cas9-StlCas9, length 1122aa, REC 393aa; LR822041.1, length 1128aa, REC 399aa; LR822017.1, length 1131aa, REC 631aa; BJMZ01000002.1, length 1188aa, REC 629aa; LR822042.1, length 1299aa, REC 628aa ; CM003138.1, length 1332aa, REC 629aa; VBTK01000005.1, length 1389aa, REC 629aa; LR822033.1, length 1409aa, REC 629aa.
[0078] FIG. 13 shows an alignment analysis of Cas9 sequences from Staphylococcus aureus organism. Clustal Omega 1.2.2 algorithm was used. Two boxes indicate REC expansion regions. Note: AIW001000003.1, 749aa; SaCas9, length 1053aa, REC 353aa; UHCS01000002.1, length 1371 aa, REC 621aa; WWFR01000017.1, 1764aa, REC 621aa.
[0079] FIG. 14 shows an alignment analysis of Cas9 sequences from Neisseria meningitidis organism. Clustal Omega 1.2.2 algorithm was used. Two boxes indicate predicated REC expansion regions. Note: CMZM01000029.2, length 500aa; OALV01000020.1, length 516aa; NWYP01000005.1, length 553aa; QQFM01000006.1, length 664aa; GAJT01000156.1, length 839aa; OAF V01000026.1, length 976aa; JWOBO 1000019.1, length 997aa; JWNZ01000005.1, length 1056aa, REC 365aa; OADFO 1000033.1, length 1081aa, REC 365aa; C9X1G5 A- CAS9_NEIM8 pdb 6JDQ, length 1082aa, REC 365aa; OAGL01000133.1, length 1120aa, REC 384aa; OAG001000026.1, length 1130aa, REC 387aa. FIG. 15 shows an alignment analysis of Cas9 sequences from Campylobacter jejuni organism. Clustal Omega 1.2.2 algorithm was used. The box indicates predicated REC expansion regions. Note: CjCas9, length 984aa, REC, 351aa; AANLCG01, length lOOOaa, REC 360aa; NZ_NFNE01, length 1002aa, REC 360; AAJFMH01, length 1003aa, REC.370; NZ_NFOE01, length 1004aa, REC 360aa; AACDNH01, length 1283aa, REC 351aa.
[0080] FIG. 16 shows an alignment analysis of SpCas9 and UHCSO 1000002.1 from FIG. 13. Clustal Omega 1.2.2 algorithm was used.
[0081] FIG. 17 shows an alignment analysis of SpCas9 and sequences from FIG. 13 with length larger than 1188aa. Clustal Omega 1.2.2 algorithm was used.
[0082] FIG. 18 shows an alignment analysis of three Cas9 proteins. Clustal Omega 1.2.2 algorithm was used. For the consensus, black background, 100% similar; gray background, 60 to 80% similar; white background, less than 60% similar.
[0083] FIG. 19 shows a schematic illustrating hypothetical variants of SpCas9.
[0084] FIG. 20 shows a schematic showing the design of RNA editing A-to-I mediated by base editor. EGFP(W58stop) mRNA contains a stop codon that prevents translation. Adenine deaminase of base editor could mediate A-to-I editing and convert the UAG stop codon to UGG Trp codon, switching on translation of EGFP protein.
[0085] FIG. 21 shows comparison among different EGFP variants used in reporter assay. Data are presented as mean ±s.d.(n=3). P values were determined by two-way ANOVA sidak’s multiple comparisons test.
[0086] FIGS. 22A-22I show optimizing conditions for reporter assay. (A) Architectures of ABE8e, ABE9, and dimer TadA deaminase nls-TadA-TadA*. (B) EGFP activating activities induced by different dosages of ABE8e. (C) EGFP activating activities induced by ABE8e, ABE9, and nls-TadA-TadA* with 200ng plasmid dosage. (D) Stop_mRNA was used for testing activating activities mediated by ABE8e and ABE9 with 200ng plasmid dosage. (E) schematic showing of EGFP DNA editing design. MMR, mismatch repair. Cells are transformed with a plasmid containing an EGFP gene, which is expressed under the control of the CMV promoter. In the presence of MMR, the repair of the bottom strand results in the generation of the mutated EGFP gene, and fluorescence is disrupted in the cells (Ito, Shiraishi et al. 2018). (F) Percentage of EGFP+ cells from RNA A-to-I assay. (G) MFI of EGFP from RNA A-to-I assay. (H) Percentage of EGFP+ cells from DNA editing assay. (I) MFI of EGFP from DNA editing assay. Data are presented as mean ±s.d., (n=3). P value was determined by two-way ANOVA sidak’s multiple comparisons test, ns, not significant. FIGS. 23A and 23B show RNA-editing and DNA editing assay with EGFPstop or EGFP reporter system. Data are presented as mean ±s.d. (n=3). P value was determined by two-way ANOVA.
[0087] FIG. 24 shows percentage and MFI of EGFP+ cells from RNA A-to-I assay.
[0088] FIG. 25 shows MFI of FIG. 3G.
[0089] FIG. 26 shows A-to-G editing rate at EMX1 site. Editing frequencies of three independent replicates (n = 3) at each A base are displayed. Here shows the details of FIG. 41.
[0090] FIG. 27 shows A-to-G editing rate at EMX1 site. Editing frequencies of three independent replicates (n = 3) at each A base are displayed. Here shows the details of FIG. 4J.
[0091] FIGS. 28A-28E show EGFP-disrupting activities of GS-ABE8e and ABE8e loaded with various gRNAs in human HEK293T cells. (A) alignment analysis of gRNA-scaffold for SpCas9, SaCas9, StlCas9 and FnCas9. Black background, 100% similar; gray background, 60 to 80% similar; white background, less than 60% similar. (B) representative sequences of reporter plasmid pCMV-GFP targeted by base editor. Guide sequence and PAM were highlighted. Below, architectures of various sgRNAs. (C) ABE8e and GS-ABE8e were assessed using human cell EGFP disruption assay when programmed with sgRNAs carrying different gRNA- scaffolds. Reporter plasmid used from (B). Data are presented as mean ±s.d. (n=3). (D), Targeting site bearing a promiscuous PAM for SaCas9 and SpCas9. (E) EGFP disruption assay using plasmid from (D) mediated by ABE8e and GS-ABE8e. Data are presented as mean ±s.d. (n=3).
[0092] FIG. 29 shows investigating base editors’ activities on A-to-G editing outside of protospacer. lOng of EGFP(Q81stop) plasmid was used for each well. SpCas9 could target TGAA PAM (Walton, Christie et al. 2020). Data are presented as mean ±s.d. (n=2 or 3).
[0093] FIG. 30 shows GS-ABE8e showed higher editing resolution on low concentration of EGFP plasmid. Ing of pCMV-GFP plasmid was used for each transfection. P values were determined by two-way ANOVA sidak’s multiple comparisons test, ns, not significant.
[0094] FIGS. 31A-31B show investigating anti-Cas9 protein AcrIIA4 effects on ABE8e and GS-ABE8e. (A) Workflow of evaluating effects of AcrIIA4 on ABE8e and GS-ABE8e. (B) EGFP disruptions mediated by ABE8e and GS-ABE8e in presence of AcrIIA4 protein. P values were determined by two-way ANOVA sidak’s multiple comparisons test, ns, not significant.
[0095] FIG. 32 shows A-to-G editing efficiencies of ABE8e and GS-ABE8e at site P208 and site P255.
[0096] FIGS. 33A-33C show comparing editing performances of ABE9 and GS-ABE9. FIG. 34 shows comparing compatibilities of GS-nCas9(D10A) with TadA8e(N46L) mutant.
[0097] FIGS. 35A-35L show comparative analysis of various Cas9 sequences and investigating REC expansion of SpCas9. (A) A Schematic illustration of REC expansion in the Cas9-evolution-hypothesis. (B) Domain size of IscB and Cas9s. Crystal structures of all the protein have been identified. REC, recognition (REC) lobe; BH, bridge helix; PAM, protospacer- adjacent motifs; PI, PAM-Interacting Domain. FnCas9 (PDB:5B2O), SpCas9 (PDB:4OO8), StlCas9 (PDB:6M0V), SaCas9 (PDB:5AXW), NmeCas9 (PDB:6JDQ), CjCas9 (PDB: 6JOO), OgeuIscB (PDB:7UTN). (C) Violin plot illustrating distribution of Cas9-sizes in various species. One pot means one sequence, n, numbers of sequences in each group. (D) Highlighting evolution of Cas9 proteins in different species. F, Francisella; Sp, Streptococcus pyogenes,' Stl, Streptococcus thermophilus,' Sa, Staphylococcus aureus; Nme, Neisseria meningitidis and Cj, Campylobacter jejuni. (E) A crystal structure of StlCas9 (PDB:6M0V). (F) A crystal structure of SaCas9 (PDB:5AXW). (G) Schematic of REC expansion from SpCas9. Insert positions are shown in cells. (H) Workflow for testing Cas9 variants activity in HEK293T cells. Episomal EGFP plasmid was co-transfected with Cas9 and gRNA plasmids to monitor Cas9 activity. (I) Activities induced by SpCas9 and variants. Data are presented as mean ±s.d. (n=3). P values were determined by two-way ANOVA sidak's multiple comparisons test. (J) Testing influences of BHs with different lengths on activities. Data are presented as mean ±s.d.(n=3). P values were determined by two-way ANOVA sidak's multiple comparisons test. (K) Variant Left-REC12 mediated EGFP disruption at different plasmid dosage. Low dosage 1, 2 ng of reporter plasmid; Low dosage 2, 4 ng of reporter plasmid. (L) Disruption activities of SpCas9 and GS-Cas9 on endogenous GFP site in HEK293-deGFP cells. Data are presented as mean ±s.d.(n=3).
[0098] FIGS. 36A-36H show investigating GS-Cas9 performance on human genome. (A) Domain organization of SpCas9 and GS-Cas9. (B) Workflow for systemically testing GS-Cas9 activity in HEK293T cells. (C) Edits rates generated by SpCas9 and GS-Cas9 at seven loci in HEK293T cells. Data are presented as mean ±s.d. (SpCas9, n=2; GS-Cas9, n=3). (D) Normalized editing frequencies for 7 target sites for SpCas9 and GS-Cas9. Each dot represents a different guide. (E) Representative sequences of the human EMX1 site targeted by GS-Cas9. sgRNA target site and PAM are indicated by blue and purple respectively. (F) Average indel length at EMX1 site from (C). (G) Targeted deep-sequencing analysis of off-target sites for the EMX1 site. Data are presented as mean ±s.d. (SpCas9, n=2; GS-Cas9, n=3). (H) Modified levels with guide sequences containing double-base mismatches at VEGFA site. Data are presented as mean ±s.d. (n=3).
[0099] FIGS. 37A-37B show that ABE8e was chosen for investigating the influences of REC expansion on N-end catalytic domain. (A) Protein structure. Left, FnCas9 (PDB 5b2o); right, ABE8e (PDB 6vpc) contains an engineered deoxy adenosine deaminase TadA from Escherichia coli and a nickase SpCas9 (D10A). (B) Hypothetical model of REC expansion for Cas9 fusion. In this proposed model, the insertion may lead to a rearrangement of the original REC domain and shorten the distance between REC and TadA by enlarging the coverage of the engineered REC lobe, thereby generating additional interactions.
[0100] FIGS. 38A-38J show REC expansion could regulate the catalytic domain fused to blend of SpCas9. (A) An illustrating work-process illustration of the pbREADAR assay. (B) A schematic illustration of the ABE8e and hypothetical GS-ABE8e variants. (C) Representative sequences of reporter plasmid pCMV-EGFP(W58stop) targeted by SpCas9 or GS-Cas9. Below, guide sequences with 20-nt length. (D, E) Activation of EGFP mediated by various base editors. Data are presented as mean ±s.d. (n=3). (F, G) Disrupting EGFP with DNA editing mediated by various base editors. Data are presented as mean ±s.d. (n=3). (H) Average A-to-I RNA editing frequencies by nSpCas9, ABE8e and GS-ABE8e mutants among 77 adenosines in IP90 mRNA transcripts. Data are presented as mean ±s.d. (n=3). (I, J) Cumulative A»T to G’C base editing efficiencies and indel formation at the human genomic EMX1 site in HEK293T cells treated with various base editors loaded with 20nt or with 35nt guide RNA. Data are presented as mean ±s.d. (n=3). P values were determined by two-way ANOVA sidak’s multiple comparisons test, ns, not significant.
[0101] FIGS. 39A-39G show investigating performance of GS-ABE8e using a plasmid-based reporter system. (A) Targeting ABE8e and GS-ABE8e to a representative EGFP locus in HEK293T cells with gRNAs containing guides in various lengths. Data are presented as mean ±s.d. (n=3). (B) EGFP disruption with guide sequences containing consecutive transversion mismatches. Data are presented as mean ±s.d. (n=3). (C) Investigating base editors’ activities on out of protospacer A-to-G editing. lOng of EGFP(Q8 Istop) plasmid was used for each well. (D) Plasmid based orthogonal R-loop assay. Here, EGFPstop was activated by Cas9- independent off-target DNA editing. Data are presented as mean ±s.d. (n=3). P value was determined by two-way ANOVA sidak’s multiple comparisons test. NS, not significant. (E) Orthogonal R-loop assay overview on human genomic loci. (F) Cas9-independent off-target A«T-to-G*C editing frequencies detected by the orthogonal R-loop assay at each R-loop site with dSaCas9 and a SaCas9 sgRNA. Each R-loop was performed by co-transfection of ABE8e or GS-ABE8e, and an SpCas9 sgRNA targeting site ABE sitel with dSaCas9 and a SaCas9 sgRNA targeting R-loops 1-5, respectively. Amplicons of R-loops sites were analyzed by high- throughput sequencing. For all plots, bars represent mean ±s.d. of three independent biological replicates. (G) On-target base editing efficiencies for ABE8e and GS-ABE8e in HEK293T cells at ABE site 1 in HEK293T cells for the orthogonal R-loop assay. Amplicons of ABE site 1 were analyzed by Sanger sequencing. For all plots, dots represent individual biological replicates and bars represent mean ±s.d. of three independent biological replicates.
[0102] FIGS. 40A-40K show systematically comparing ABE8e and GS-ABE8e gene editing performance on human genomic loci in HEK293T cells. (A) Workflow for testing GS-ABE8e in HEK293T cells. (B) Evaluation of the A-to-G editing efficiencies of ABE8e and GS-ABE8e at 11 representative endogenous genomic sites in HEK293T cells. Data are presented as mean ±s.d. (n=3). (C) Indel formation in HEK293T cells treated as described in (B). Data are presented as mean ±s.d. (n=3). (D) Comparison of indels induced by ABE8e and GS-ABE8e at 12 target sites including 11 sites in (C) and EMX1 site in FIG. 381. Each data point represents the average indel frequency at each site calculated from 3 independent experiments. P values were determined by two-tailed student’s t-test. (E) Frequencies of A-to-G editing by ABE8e or GS-ABE8e editor across the protospacer positions 1-20 from (B). (F) Cas9-dependent DNA off-target analysis comparing ABE8e and GS-ABE8e at ABE site 2. (G) Cas9-dependent DNA off-target analysis comparing ABE8e and GS-ABE8e at EMX1 site. (H, I) On-target:off-target editing ratios for two sites in (F) and (G). For all plots, bars represent mean values, and error bars represent the s.d. of three independent biological replicates. (J) Indel formation at two off- target sites in (J). (K) Indel formation at five off-target sites in (G).
[0103] FIGS. 41A-41C show schematic-illustrating hypothetical models and expressions of SpCas9 variants. (A) Excessive length of BH domain reduced SpCas9 cleavage activities. Cells were analyzed 2 days post transfections. Data are presented as mean ±s.d., (n=3). P values were determined by two-way ANOVA sidak’s multiple comparisons test. (B) Hypothetical cartoon models of Cas9 variants. (C) Western blotting analysis of Cas9 variants using anti-Flag antibody. Different lysates of HEK293T cells transfected with PBS, SpCas9, SpCas9-2BH and GS-Cas9 were analyzed.
[0104] FIG. 42 shows Alleles_frequency and indel length. Left, representative sequences of 6 sites targeted by GS-Cas9. Right, average indel length at 6 sites mediated by SpCas9 and GS- Cas9 in FIG.36C. FIG. 43A-43B show optimizing conditions for reporter assay. (A) Architectures of ABE8e, ABE9, and dimer TadA deaminase nls-TadA-TadA*. (B) EGFP activating activities induced by different dosages of ABE8e.
[0105] FIGS. 44A-44D show evaluating various ABE variants using pbREADER method. (A) EGFP activations mediated by various ABE editors. (B) EGFP disruptions mediated by various ABE editors. (C) EGFP disruptions mediated by ABE8e and ABE8e-2BH. Data are presented as mean ±s.d. (n=3). P values were determined by two-way ANOVA sidak’s multiple comparisons test, ns, not significant. (D) Different lysates of HEK293T cells transfected with PBS, ABE8e, GS-ABE8e, GS-ABE8e-1.5BH, and GS-ABE8e-2BH were analyzed. A Flag tag was fused to c-end of each editor, which was used for western blotting analysis.
[0106] FIG. 45 shows A-to-G editing rate at EMX1 site. Editing frequencies of three independent replicates (n = 3) at each A base are displayed. Here shows the details of Fig. 4i.
[0107] FIG. 46 shows A-to-G editing rate at EMX1 site. Editing frequencies of three independent replicates (n = 3) at each A base are displayed. Here show the details of FIG. 38J.
[0108] FIG. 47 shows edits outside of the protospacer sequence mediated by ABE8e and GS- ABE8e in FIG. 40.
[0109] FIG. 48 shows investigating ABE9 and GS-ABE9 activities using plasmid reporter system. Data are presented as mean ±s.d. (n=3) .
[0110] FIG. 49A-49C show comparison of editing performances of ABE9 and GS-ABE9. (A) A-to-G editing efficiencies. (B) Indel formation in HEK293T cells treated as described in A. Data are presented as mean ±s.d. (n=3). (C) Frequencies of A-to-G by ABE9 or GS-ABE9 editor across the protospacer positions 1-20 from (A).
[0111] FIGS. 50A and 50B show comparison of compatibilities of GS-nCas9(D10A) with TadA8e(N46L) and TadA8e (V106W) mutant. (A) Architectures of tested base editors’ variants. (B) A-to-G editing efficiencies. Data are presented as mean ±s.d. (n=3) .
[0112] FIG. 51 shows activation of EGFP mediated by ABE8e (V106W) and GS-ABE8e (V106W) base editors. Data are presented as mean ±s.d. (n=3). P values were determined by two-way ANOVA sidak’s multiple comparisons test, ns, not significant.
[0113] FIG. 52 shows size distributions of domains in representative BEs, Cas9s and IscB.
[0114] Detailed Description of the Invention Domain expansion contributes to diversification of RNA-guided-endonucleases including Cas9. However, it remains unclear how REC domain expansion could benefit Cas9. In this study, we identified a spot that is compatible with large REC insertion and succeeded in enlarging the non-catalytic REC domain of Streptococcus pyogenes Cas9. The naturalevolution-like giant SpCas9 (GS-Cas9) was created and showed a substantial improvement of editing precision. We further discovered that enlarging the REC domain could enable regulation of the N-terminal adenine deaminase TadA8e tethered to the Cas9 scaffold, which contributes to substantially reducing unexpected editing and improving the precision of the adenine base editor ABE8e.
[0115] Definitions
[0116] For convenience, certain terms employed in the specification, examples, and appended claims are collected here.
[0117] The articles “a” and “an” are used herein to refer to one or to more than one (e.g., to at least one) of the grammatical object of the article. By way of example, “an element” means one element or more than one element.
[0118] As used herein, two nucleic acid sequences “complement” one another or are “complementary” to one another if they base pair one another at each position.
[0119] The term “polynucleotide” and “nucleic acid” refer to a polymeric form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogs thereof. It may be composed of four standard nucleotides, each with a different nucleobase: adenine (A), thymine (T) / Uridine (U), guanine (G), and cytosine (C). It may contain non-conventional nucleobases, such as the Z-base. Polynucleotides may have any three-dimensional structure, and may perform any function, known or unknown. The following are non-limiting examples of polynucleotides: coding or non-coding regions of a gene or gene fragment, loci (locus) defined from linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, ribozymes, cDNA, synthetic polynucleotides, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. A polynucleotide may comprise modified nucleotides, such as methylated nucleotides and nucleotide analogs. If present, modifications to the nucleotide structure may be imparted before or after assembly of the polymer. The sequence of nucleotides may be interrupted by non-nucleotide components. A polynucleotide may be further modified, such as by conjugation with a labeling component. The term “modifying” a target DNA refers to changing the sequence of the target DNA, for example, by introducing a deletion, an insertion, and / or a substitution of the target DNA sequence. In some embodiments, “modifying” a target DNA refers to inducing an indel in the target DNA. In some embodiments, “modifying” a target DNA refers to inducing a base change, e.g., a A-to-G base change in the target DNA.
[0120] “Guide RNA,” “gRNA,” and simply “guide” are used herein interchangeably to refer to either a guide that comprises a guide sequence, e.g. either a crRNA (also known as CRISPR RNA), or the combination of a crRNA and a trRNA (also known as tracrRNA crRNA (also known as CRISPR RNA), or the combination of a crRNA and a trRNA (also known as tracrRNA). The crRNA and trRNA may be associated as a single RNA molecule (single guide RNA, sgRNA) or, for example, in two separate RNA molecules (dual guide RNA, dgRNA). “Guide RNA” or “gRNA” refers to each type. The trRNA may be a naturally-occurring sequence, or a trRNA sequence with modifications or variations compared to naturally-occurring sequences. Guide RNAs, such as sgRNAs or dgRNAs, can include modified RNAs as described herein.
[0121] As used herein, a “guide sequence” refers to a sequence within a guide RNA that is complementary to a target sequence and functions to direct a guide RNA to a target sequence for binding or modification (e.g., cleavage) by a Cas9 variant descried herein. A “guide sequence” may also be referred to as a “targeting sequence,” or a “spacer sequence.” A guide sequence can be 20 base pairs in length, e.g., in the case of Streptococcus pyogenes (i.e., Spy Cas9) and related Cas9 homologs / orthologs. Shorter or longer sequences can also be used as guides, e.g., 15-, 16-, 17-, 18-, 19-, 21-, 22-, 23-, 24-, or 25-nucleotides in length. In some embodiments, the target sequence is in a gene or on a chromosome, for example, and is complementary to the guide sequence. In some embodiments, the degree of complementarity or identity between a guide sequence and its corresponding target sequence may be about 75%, 80%, 85%, 90%, 95%, or 100%. In some embodiments, the guide sequence and the target region may be 100% complementary or identical. In other embodiments, the guide sequence and the target region may contain at least one mismatch. For example, the guide sequence and the target sequence may contain 1, 2, 3, or 4 mismatches, where the total length of the target sequence is at least 15, 16, 17, 18, 19, 20 or more base pairs. In some embodiments, the guide sequence and the target region may contain 1-4 mismatches where the guide sequence comprises at least 15, 16, 17, 18, 19, 20 or more nucleotides. In some embodiments, the guide sequence and the target region may contain 1, 2, 3, or 4 mismatches where the guide sequence comprises 20 nucleotides. Target sequences for Cas9 include both the positive and negative strands of genomic DNA i.e., the sequence given and the sequence’s reverse complement), as a nucleic acid substrate for Cas9is a double stranded nucleic acid. Accordingly, where a guide sequence is said to be “complementary to a target sequence,” it is to be understood that the guide sequence may direct a guide RNA to bind to the sense or antisense strand (e.g. reverse complement) of a target sequence. Thus, in some embodiments, where the guide sequence binds the reverse complement of a target sequence, the guide sequence is identical to certain nucleotides of the target sequence (e.g. , the target sequence not including the PAM) except for the substitution of U for T in the guide sequence.
[0122] As used herein, a first sequence is considered to “comprise a sequence with at least X% identity to” a second sequence if an alignment of the first sequence to the second sequence shows that X% or more of the positions of the second sequence in its entirety are matched by the first sequence. For example, the sequence AAGA comprises a sequence with 100% identity to the sequence AAG because an alignment would give 100% identity in that there are matches to all three positions of the second sequence. The differences between RNA and DNA (generally the exchange of uridine for thymidine or vice versa) and the presence of nucleoside analogs such as modified uridines do not contribute to differences in identity or complementarity among polynucleotides as long as the relevant nucleotides (such as thymidine, uridine, or modified uridine) have the same complement (e.g., adenosine for all of thymidine, uridine, or modified uridine; another example is cytosine and 5-methylcytosine, both of which have guanosine or modified guanosine as a complement). Thus, for example, the sequence 5’-AXG where X is any modified uridine, such as pseudouridine, N1 -methyl pseudouridine, or 5-methoxyuridine, is considered 100% identical to AUG in that both are perfectly complementary to the same sequence (5’-CAU). Exemplary alignment algorithms are the Smith-Waterman and Needleman-Wunsch algorithms, which are well-known in the art. One skilled in the art will understand what choice of algorithm and parameter settings are appropriate for a given pair of sequences to be aligned; for sequences of generally similar length and expected identity >50% for amino acids or >75% for nucleotides, the Needleman- Wunsch algorithm with default settings of the Needleman-Wunsch algorithm inteace provided by the EBI at the www.ebi.ac.uk web server is generally appropriate.
[0123] As used herein, a first sequence is considered to be “X% complementary to” a second sequence if X% of the bases of the first sequence base pairs with the second sequence. For example, a first sequence 5’AAGA3’ is 100% complementary to a second sequence 3’TTCT5’, and the second sequence is 100% complementary to the first sequence. In some embodiments, a first sequence 5’AAGA3’ is 100% complementary to a second sequence 3’TTCTGTGA5’, whereas the second sequence is 50% complementary to the first sequence.
[0124] As used herein, “indels” refer to insertion / deletion mutations consisting of a number of nucleotides that are either inserted or deleted at the site of double-stranded breaks (DSBs) in a target nucleic acid.
[0125] As used herein, a “target sequence” refers to a sequence of nucleic acid in a target gene that has complementarity to the guide sequence of the gRNA. The interaction of the target sequence and the guide sequence directs a Cas9 variant described herein to bind, and potentially nick or cleave (depending on the activity of the agent), within the target sequence.
[0126] Cas9 variants
[0127] In some aspects, provided herein is a Cas9 variant comprising an enlarged Streptococcus pyogenes Cas9 (SpCas9) that comprises an insertion of one or more domains selected from a non-catalytic Bridge helix (BH) domain and a recognition (REC) domain. In some embodiments, the enlarged SpCas9 comprises an insertion of a BH domain. In some embodiments, the enlarged SpCas9 comprises an insertion of a REC domain. In some embodiments, the enlarged SpCas9 comprises an insertion of a BH and a REC domain.
[0128] The non-catalytic Bridge helix (BH) domain and / or the recognition (REC) domain can be derived from a Cas9 nuclease from any species other than Streptococcus pyogenes. In some embodiments, the enlarged SpCas9 comprises an insertion of one or more domains selected from a non-catalytic StlCas9 BH domain and a StlCas9 REC domain. In some embodiments, the enlarged SpCas9 comprises an insertion of a non-catalytic StlCas9 BH domain and a StlCas9 REC domain. In some embodiments, the enlarged SpCas9 comprises an insertion of a non-catalytic StlCas9 BH domain. In some embodiments, the enlarged SpCas9 comprises an insertion of a StlCas9 REC domain.
[0129] In some aspects, provided herein is a Cas9 variant comprising a Streptococcus pyogenes Cas9 (SpCas9) that comprises an insertion of a Streptococcus thermophilus Cas9 (StlCas9) REC domain. In some aspects, provided herein is a Cas9 variant comprising a Streptococcus pyogenes Cas9 (SpCas9) that comprises an insertion of a StlCas9 BH domain. In some aspects, provided herein is a Cas9 variant comprising a Streptococcus pyogenes Cas9 (SpCas9) that comprises an insertion of a StlCas9 BH domain and a StlCas9 REC domain.
[0130] In some embodiments, the StlCas9 REC domain is inserted at the N-terminus of the SpCas9 REC domain. In some embodiments, the StlCas9 REC domain is inserted between the BH domain and the REC domain of the SpCas9. In some embodiments, the StlCas9 REC domain is inserted between residues 93D and 94D of the SpCas9. In some embodiments, the StlCas9 REC domain is inserted between residues 93D and 94D of SEQ ID NO: 1. In some embodiments, the StlCas9 REC domain is inserted between residues 93D and 94D of SEQ ID NO: 22.
[0131] In some embodiments, the Cas9 variant is derived from a wild-type SpCas9. In some embodiments, the Cas9 variant is derived from a SpCas9 that comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the sequence of the wild-type SpCas9. In some embodiments, the Cas9 variant is derived from a SpCas9 that comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the sequence of SEQ ID NO: 1. In some embodiments, the Cas9 variant is derived from a SpCas9 that comprises the amino acid sequence of SEQ ID NO: 1.
[0132] In some embodiments, the Cas9 variant is derived from a SpCas9 nickase (nSpCas9). In some embodiments, the Cas9 variant is derived from a nSpCas9 (D10A). In some embodiments, the Cas9 variant is derived from a nSpCas9 that comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the sequence of SEQ ID NO: 22. In some embodiments, the Cas9 variant is derived from a nSpCas9 that comprises the amino acid sequence of SEQ ID NO: 22.
[0133] In some embodiments, the StlCas9 REC domain comprises an RECI domain and / or an REC2 domain. In some embodiments, the StlCas9 REC domain comprises an RECI domain and an REC2 domain (e.g., designated as “REC12” herein). In some embodiments, the StlCas9 REC domain comprises residues 74-466 of StlCas9. In some embodiments, the StlCas9 REC domain comprises the amino acid of SEQ ID NO: 19. In some embodiments, the StlCas9 REC domain comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the sequence of SEQ ID NO: 19.
[0134] In some embodiments, the Cas9 variant does not comprise a BH domain from StlCas9.
[0135] In some embodiments, the StlCas9 REC domain is connected to the N-terminus of the SpCas9 REC domain via a linker. The linker can be a chemical linker, or a peptide. In some embodiments, the linker is a peptide fragment from a Cas9 nuclease. In some embodiments, the linker is a peptide fragment from a Cas9 nuclease from a species other than Streptococcus pyogenes. In some embodiments, the linker comprises a linker from Staphylococcus aureus (SaCas9). In some embodiments, the linker from SaCas9 comprises residues T205 to D223 of SaCas9. In some embodiments, the linker from SaCas9 comprises the amino acid sequence of SEQ ID NO: 20. In some embodiments, the linker from SaCas9 comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the sequence of SEQ ID NO: 20.
[0136] In some embodiments, the SpCas9 of the Cas9 variant comprises an insertion of the amino acid sequence of SEQ ID NO: 8. In some embodiments, the SpCas9 of the Cas9 variant comprises an insertion of an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the sequence of SEQ ID NO: 8. In some embodiments, the Cas9 variant comprises the amino acid sequence of SEQ ID NO: 21. In some embodiments, the Cas9 variant comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the sequence of SEQ ID NO: 21.
[0137] In some aspects, provided herein is a Cas9 variant, comprising a nSpCas9 (e.g., nSpCas9 (D10A)) that comprises an insertion of BH and / or REC domains. In some aspects, provided herein is a Cas9 variant, comprising a nSpCas9 (e.g., nSpCas9 (D10A)) that comprises an insertion of BH and / or REC domains of StlCas9. In some embodiments, the BH and / or REC domains of StlCas9 is inserted at the N-terminus of the nSpCas9 (e.g., nSpCas9 (D10A)) REC domain. In some embodiments, the BH and / or REC domains of StlCas9 is inserted between the BH domain and the REC domain of the nSpCas9 (e.g., nSpCas9 (D10A)). In some embodiments, the BH and REC domains of StlCas9 is inserted between the BH domain and the REC domain of the nSpCas9 (e.g., nSpCas9 (D10A)), for example, from N-terminus to C-terminus, in the order of SpCas9 BH domain, StlCas9 BH domain, StlCas9 REC domain, SpCas9 REC domain. In some embodiments, the BH and / or REC domains of StlCas9 is inserted between residues 93D and 94D of the nSpCas9 (e.g., nSpCas9 (D10A)). In some embodiments, , the BH and / or REC domains of StlCas9 is inserted between residues 93D and 94D of SEQ ID NO: 22.
[0138] In some embodiments, the nSpCas9 (e.g., nSpCas9 (D10A)) comprises an insertion of the amino acid sequence of SEQ ID NO: 3, 7, or 8. In some embodiments, the nSpCas9 (e.g., nSpCas9 (D10A)) comprises an insertion of an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the amino acid sequence of SEQ ID NO: 3, 7, or 8. In some embodiments, the nSpCas9 (D10A) comprises an insertion of the amino acid sequence of SEQ ID NO: 3, 7, or 8 between residues 93D and 94D of the nSpCas9 (e.g., nSpCas9 (D10A)). In some embodiments, the Cas9 variant comprises an amino acid sequence selected from SEQ ID NOs: 23-25. In some embodiments, the Cas9 variant is used as a binding scaffold to generate a base editor.
[0139] As used herein, “Cas9 variant” refers to a variant e.g., mutant, fragment, fusion, or combinations thereof) of a wild-type (i.e., naturally occurring) Cas9. A Cas9 variant may possess at least or about 5%, 10%, 15%, 20%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% functional activity of the wild-type Cas9. In some embodiments, the Cas9 variant is at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the sequence of the wild-type Cas9. In some embodiments, a Cas9 variant may be a hyperactive variant. In certain instances, the Cas9 variant possesses between about 80% and about 120%, 140%, 160%, 180%, 200% of the functional activity of the wild-type Cas9. In some embodiments, the Cas9 variant is an enlarged SpCas9, i.e., comprising a longer amino acid sequence than the wildtype SpCas9. In some embodiments, the Cas9 variant comprises a larger REC domain compared to the wild-type SpCas9.
[0140] In some embodiments, the Cas9 variant reduces off-target activity of the wild-type Cas9. In certain instances, the Cas9 variant reduces at least or about 5%, 10%, 15%, 20%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% off-target activity of the wild-type Cas9. The off-target activity comprises the percentage of indels induced at the one off-target site and / or the total number of off-target sites that have detectable indels. That is, in some embodiments, the Cas9 variant reduces at least or about 5%, 10%, 15%, 20%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% of the indels induced at an off-target site compared to the wild-type Cas9. In some embodiments, the Cas9 variant reduces at least or about 5%, 10%, 15%, 20%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% of the total number of off-target sites that have detectable indels compared to the wild-type Cas9.
[0141] In some embodiments, the Cas9 variant, has cleavase activity, which can also be referred to as double-strand endonuclease activity. In some embodiments, the Cas9 variant, has nickase activity, which can also be referred to as single-strand endonuclease activity.
[0142] In some embodiments, the Cas9 variant is a variant of a Cas9 nuclease selected from those of the type II CRISPR systems of 5. pyogenes, S. aureus, and other prokaryotes. In some embodiments, the Cas9 variant is a variant of a Cas9 nuclease from Streptococcus pyogenes. In some embodiments, the Cas9 variant is a variant of a Cas9 nuclease from Streptococcus thermophilus. In some embodiments, the Cas9 variant is a variant of a Cas9 nuclease from Neisseria meningitidis. In some embodiments, the Cas9 variant is a variant of a Cas9 nuclease from Staphylococcus aureus.
[0143] Wild type Cas9 has two nuclease domains: RuvC and HNH. The RuvC domain cleaves the non-target DNA strand, and the HNH domain cleaves the target strand of DNA. In some embodiments, the Cas9 protein comprises more than one RuvC domain or more than one HNH domain. In some embodiments, the Cas9 variant comprises an enlarged wild type Cas9. In each of the composition, use, and method embodiments, the Cas9 variant induces a double strand break in target DNA.
[0144] In some embodiments, the Cas9 variant is derived from a chimeric Cas9 nuclease, where one domain or region of the Cas9 is replaced by a portion of a different protein. In some embodiments, a Cas9 nuclease domain may be replaced with a domain from a different nuclease such as Fokl. In some embodiments, a Cas9 nuclease may be a modified nuclease.
[0145] In some embodiments, the Cas9 variant has single-strand nickase activity, i.e., can cut one DNA strand to produce a single-strand break, also known as a “nick.” In some embodiments, the Cas9 variant comprises a Cas9 nickase. A nickase is an enzyme that creates a nick in dsDNA, i.e., cuts one strand but not the other of the DNA double helix. In some embodiments, a Cas9 nickase is a version of a Cas9 nuclease (e.g., a Cas9 nuclease discussed above) in which an endonucleolytic active site is inactivated, e.g., by one or more alterations e.g., point mutations) in a catalytic domain. In some embodiments, a Cas9 nickase such as a spCas9 nickase has an inactivated RuvC or HNH domain.
[0146] In some embodiments, the Cas9 variant is modified to contain only one functional nuclease domain. For example, the Cas9 variant may be modified such that one of the nuclease domains is mutated or fully or partially deleted to reduce its nucleic acid cleavage activity. In some embodiments, a nickase is used having a RuvC domain with reduced activity. In some embodiments, a nickase is used having an inactive RuvC domain. In some embodiments, a nickase is used having an HNH domain with reduced activity. In some embodiments, a nickase is used having an inactive HNH domain.
[0147] In some embodiments, a conserved amino acid within a Cas9 protein nuclease domain is substituted to reduce or alter nuclease activity. In some embodiments, a Cas9 nuclease may comprise an amino acid substitution in the RuvC or RuvC-like nuclease domain. Exemplary amino acid substitutions in the RuvC or RuvC-like nuclease domain include D10A (based on the S. pyogenes Cas9 protein). See, e.g., Zetsche et al. (2015) Cell Oct 22:163(3): 759-771. In some embodiments, the Cas9 nuclease may comprise an amino acid substitution in the HNH or HNH-like nuclease domain. Exemplary amino acid substitutions in the HNH or HNH-like nuclease domain include E762A, H840A, N863A, H983A, and D986A (based on the S. pyogenes Cas9 protein). See, e.g., Zetsche et al. (2015).
[0148] In some embodiments, the Cas9 variant comprises one or more heterologous functional domains (e.g., is or comprises a fusion polypeptide).
[0149] In some embodiments, the heterologous functional domain may facilitate transport of the Cas9 variant into the nucleus of a cell. For example, the heterologous functional domain may be a nuclear localization signal (NLS). In some embodiments, the Cas9 variant may be fused with 1-10 NLS(s). In some embodiments, the Cas9 variant may be fused with 1-5 NLS(s). In some embodiments, the Cas9 variant may be fused with one NLS. Where one NLS is used, the NLS may be linked at the N-terminus or the C-terminus of the Cas9 variant. It may also be inserted within the Cas9 variant. In other embodiments, the Cas9 variant may be fused with more than one NLS. In some embodiments, the Cas9 variant may be fused with 2, 3, 4, or 5 NLSs. In some embodiments, the Cas9 variant may be fused with two NLSs. In certain circumstances, the two NLSs may be the same (e.g., two SV40 NLSs) or different. In some embodiments, the Cas9 variant is fused to two SV40 NLS sequences linked at the carboxy terminus. In some embodiments, the Cas9 variant may be fused with two NLSs, one linked at the N-terminus and one at the C-terminus. In some embodiments, the Cas9 variant may be fused with 3 NLSs. In some embodiments, the Cas9 variant may be fused with no NLS. In some embodiments, the NLS may be a monopartite sequence, such as, e.g., the SV40 NLS, PKKKRKV or PKKKRRV. In some embodiments, the NLS may be a bipartite sequence, such as the NLS of nucleoplasmin, KRPAATKKAGQAKKKK. In a specific embodiment, a single PKKKRKV NLS may be linked at the C-terminus of the Cas9 variant. One or more linkers are optionally included at the fusion site.
[0150] In some aspects, provided herein is a nucleic acid encoding the Cas9 variant described herein. In some embodiments, the nucleic acid is a DNA or RNA.
[0151] In some aspects, provided herein is a vector comprising a nucleic acid described herein. In some embodiments, the vector is a plasmid. In some embodiments, the vector is a viral vector, such as an AAV vector or a lentiviral vector.
[0152] Base Editors In some aspects, provided herein is a Cas9-guided base editor comprising a base editor linked to a Cas9 variant described herein. For example, in some aspects, provided herein is a Cas9-guided base editor comprising a base editor linked to a Cas9 variant, wherein the Cas9 variant comprises a SpCas9 nickase (nSpCas9) comprising an insertion of a REC domain. In some embodiments, provided herein is a Cas9-guided base editor comprising a base editor linked to a Cas9 variant, wherein the Cas9 variant comprises a SpCas9 nickase (nSpCas9) comprising an insertion of a BH domain and a REC domain. The non-catalytic Bridge helix (BH) domain and / or the recognition (REC) domain can be derived from a Cas9 nuclease from any species other than Streptococcus pyogenes. In some embodiments, the enlarged SpCas9 comprises an insertion of one or more domains selected from a non- catalytic StlCas9 BH domain and a StlCas9 REC domain.
[0153] In some embodiments, provided herein is a Cas9-guided base editor comprising a base editor linked to a Cas9 variant, wherein the Cas9 variant comprises a SpCas9 nickase (e.g., nSpCas9 (D10A)) comprising an insertion of a StlCas9 REC domain. In some embodiments, provided herein is a Cas9-guided base editor comprising a base editor linked to a Cas9 variant, wherein the Cas9 variant comprises a SpCas9 nickase (e.g., nSpCas9 (D10A)) comprising an insertion of a BH domain and a REC domain from StlCas9.
[0154] In some embodiments, the BH and / or REC domains of StlCas9 is inserted at the N- terminus of the nSpCas9 (e.g., nSpCas9 (D10A)) REC domain. In some embodiments, the BH and / or REC domains of StlCas9 is inserted between the BH domain and the REC domain of the nSpCas9 (e.g., nSpCas9 (D10A)). In some embodiments, the BH and REC domains of StlCas9 is inserted between the BH domain and the REC domain of the nSpCas9 (e.g., nSpCas9 (D10A)), for example, from N-terminus to C-terminus, in the order of SpCas9 BH domain, StlCas9 BH domain, StlCas9 REC domain, SpCas9 REC domain. In some embodiments, the BH and / or REC domains of StlCas9 is inserted between residues 93D and 94D of the nSpCas9 (e.g., nSpCas9 (D10A)). In some embodiments, the BH and / or REC domains of StlCas9 is inserted between residues 93D and 94D of SEQ ID NO: 22.
[0155] In some embodiments, the StlCas9 REC domain comprises an RECI domain and / or an REC2 domain. In some embodiments, the StlCas9 REC domain comprises an RECI domain and an REC2 domain (e.g., designated as “REC12” herein). In some embodiments, the StlCas9 REC domain comprises residues 74-466 of StlCas9. In some embodiments, the StlCas9 REC domain comprises the amino acid of SEQ ID NO: 19. In some embodiments, the StlCas9 REC domain comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the sequence of SEQ ID NO: 19.
[0156] In some embodiments, the StlCas9 REC domain is connected to the N-terminus of the nSpCas9 (e.g., nSpCas9 (D10A)) REC domain via a linker. The linker can be a chemical linker, or a peptide. In some embodiments, the linker is a peptide fragment from a Cas9 nuclease. In some embodiments, the linker is a peptide fragment from a Cas9 nuclease from a species other than Streptococcus pyogenes. In some embodiments, the linker comprises a linker from Staphylococcus aureus (SaCas9). In some embodiments, the linker from SaCas9 comprises residues T205 to D223 of SaCas9. In some embodiments, the linker from SaCas9 comprises the amino acid sequence of SEQ ID NO: 20. In some embodiments, the linker from SaCas9 comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the sequence of SEQ ID NO: 20.
[0157] In some embodiments, the Cas9 variant comprising half (0.5x), single (lx), or two (2x) BH domains immediately upstream of the StlCas9 REC domain. In some embodiments, the Cas9 variant comprising a single BH domain immediately upstream of the StlCas9 REC domain.
[0158] In some embodiments, the nSpCas9 comprises an insertion of an amino acid sequence selected from SEQ ID NOs: 3, 7 and 8. In some embodiments, the nSpCas9 comprises an insertion of the amino acid sequence of SEQ ID NOs: 3. In some embodiments, the nSpCas9 of the Cas9 variant comprises an insertion of an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the sequence of SEQ ID NO: 3. In some embodiments, the nSpCas9 comprises an insertion of the amino acid sequence of SEQ ID NOs: 7. In some embodiments, the nSpCas9 of the Cas9 variant comprises an insertion of an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the sequence of SEQ ID NO: 7. In some embodiments, the nSpCas9 comprises an insertion of the amino acid sequence of SEQ ID NOs: 8. In some embodiments, the nSpCas9 of the Cas9 variant comprises an insertion of an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the sequence of SEQ ID NO: 8.
[0159] In some embodiments, the cas9 variant comprises an amino acid sequence selected from SEQ ID NOs: 23-25. In some embodiments, the cas9 variant comprises the amino acid sequence of SEQ ID NOs: 23. In some embodiments, the cas9 variant comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the sequence SEQ ID NOs: 23. In some embodiments, the cas9 variant comprises the amino acid sequence of SEQ ID NOs: 24. In some embodiments, the cas9 variant comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the sequence SEQ ID NOs: 24. In some embodiments, the cas9 variant comprises the amino acid sequence of SEQ ID NOs: 25. In some embodiments, the cas9 variant comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the sequence SEQ ID NOs: 25.
[0160] In some embodiments, the base editor is linked to N-terminus of the Cas9 variant. In some embodiments, the base editor is an adenine deaminase. In some embodiments, the adenine deaminase is selected from TadA8e, TadA9, TadA8e(N46L), and TadA8e(V106W), or a fragment or variant thereof. In some embodiments, the adenine deaminase is selected from TadA8e, or a fragment or variant thereof.
[0161] In some aspects, provided herein is a nucleic acid encoding the Cas9-guided base editor described herein. In some embodiments, the nucleic acid is a DNA or RNA.
[0162] In some aspects, provided herein is a vector comprising a nucleic acid described herein. In some embodiments, the vector is a plasmid. In some embodiments, the vector is a viral vector, such as an AAV vector or a lentiviral vector.
[0163] Compositions and Kits
[0164] In some aspects, provided herein is a composition or kit, comprising: (a) a Cas9 variant described herein, a nucleic acid described herein, or a vector described herein; and (b) a Cas9 guide RNA (gRNA). In some embodiments, the Cas9 gRNA comprises a gRNA- scaffold selected from SpCas9 gRNA scaffold, SaCas9 gRNA-scaffold, StlCas9 gRNA- scaffold, and FnCas9 gRNA-scaffold. In some embodiments, the Cas9 gRNA comprises a guide sequence of 8-40 nt in length, e.g., 8 nucleotides (nt), lOnt, 12nt, 20nt, 35nt, or 40 nt in length. In some embodiments, the Cas9 gRNA is a spCas9 gRNA.
[0165] In some aspects, provided herein is a composition or kit, comprising: (a) a Cas9 variant described herein, a nucleic acid described herein, or a vector described herein; (b) a Cas9 crRNA; and (c) a Cas9 tracrRNA. In some embodiments, the Cas9 crRNA is selected from SpCas9 crRNA, SaCas9 crRNA, StlCas9 crRNA, and FnCas9 crRNA, and wherein the Cas9 tracrRNA is selected from the same species as the Cas9 crRNA. In some embodiments, the Cas9 crRNA comprises a guide sequence of 8-40 nt in length, e.g., 8 nt, lOnt, 12nt, 20nt, 35nt, or 40 nt in length.
[0166] In some aspects, provided herein is a composition or kit, comprising: (a) a Cas9- guided base editor described herein, a nucleic acid described herein, or a vector described herein; and (b) a Cas9 guide RNA (gRNA). In some embodiments, the Cas9 gRNA comprises a gRNA-scaffold selected from SpCas9 gRNA scaffold, SaCas9 gRNA-scaffold, StlCas9 gRNA-scaffold, and FnCas9 gRNA-scaffold. In some embodiments, the Cas9 gRNA comprises a guide sequence of 8-40 nt in length, e.g., 8 nucleotides (nt), lOnt, 12nt, 20nt, 35nt, or 40 nt in length. In some embodiments, the Cas9 gRNA is a spCas9 gRNA.
[0167] In some aspects, provided herein is a composition or kit, comprising: (a) a Cas9- guided base editor described herein, a nucleic acid described herein, or a vector described herein; (b) a Cas9 crRNA; and (c) a Cas9 tracrRNA. In some embodiments, the Cas9 crRNA is selected from SpCas9 crRNA, SaCas9 crRNA, StlCas9 crRNA, and FnCas9 crRNA, and wherein the Cas9 tracrRNA is selected from the same species as the Cas9 crRNA. In some embodiments, the Cas9 crRNA comprises a guide sequence of 8-40 nt in length, e.g., 8 nt, lOnt, 12nt, 20nt, 35nt, or 40 nt in length.
[0168] In some embodiments, the guide sequences may further comprise a SpCas9 sgRNA sequence. An example of a SpCas9 sgRNA sequence is: 5’- GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUU GAAAAAGUGGCACCGAGUCGGUGC-3’, included at the 3’ end of the guide sequence. In some embodiments, the sgRNA scaffold for SaCas9 comprises the amino acid sequence of SEQ ID NO: 16. In some embodiments, the sgRNA scaffold for StlCas9 comprises the amino acid sequence of SEQ ID NO: 17. In some embodiments, the sgRNA scaffold for FnCas9 comprises the amino acid sequence of SEQ ID NO: 18.
[0169] In some embodiments, the example of a nucleotide sequence of SpCas9 sgRNA listed above may serve as a template sequence for specific chemical modifications, sequence substitutions and truncations.
[0170] In certain embodiments, the gRNA is an sgRNA or a dgRNA, for example, and it optionally comprises a chemical modification. In some embodiments, the modified sgRNA comprises a guide sequence and a SpCas9 sgRNA sequence, e.g., exemplary SpyCas9 sgRNA shown above. A gRNA, such as an sgRNA, may include modifications on the 5’ end of the guide sequence and / or on the 3’ end of the SpCas9 sgRNA sequence, such as, e.g., at one or more of the terminal nucleotides, e.g., at 1, 2, 3, or 4 of the nucleotides at the 3’ end or at the 5’ end. In certain embodiments, the modified nucleotide is selected from a 2’-O- methyl (2’-0Me) modified nucleotide, a 2’-O-(2-methoxyethyl) (2’-O-moe) modified nucleotide, a 2’-fluoro (2’-F) modified nucleotide, a phosphorothioate (PS) linkage between nucleotides, an inverted abasic modified nucleotide, or a combination thereof. In certain embodiments, the modified nucleotide includes a 2’-0Me modified nucleotide. In certain embodiments, the modified nucleotide includes a PS linkage. In certain embodiments, the modified nucleotide includes a 2’-0Me modified nucleotide and a PS linkage.
[0171] In each composition and method embodiment described herein, the crRNA and trRNA may be associated as a single RNA (sgRNA) or may be on separate RNAs (dgRNA). In the context of sgRNAs, the crRNA and trRNA components may be covalently linked, e.g., via a phosphodiester bond or other covalent bond. In some embodiments, the sgRNA comprises one or more linkages between nucleotides that is not a phosphodiester linkage
[0172] In each of the composition, use, and method embodiments described herein, the guide RNA may comprise a single RNA molecule as a "single guide RNA” or “sgRNA” or “gRNA”.
[0173] In some embodiments, the crRNA and the trRNA (also referred to as “tracrRNA”) are covalently linked via a linker. In some embodiments, the sgRNA forms a stem- loop structure via the base pairing between portions of the crRNA and the trRNA. In some embodiments, the crRNA and the trRNA are covalently linked via one or more bonds that are not a phosphodiester bond.
[0174] In some embodiments, the trRNA may comprise all or a portion of a trRNA sequence derived from a naturally occurring CRISPR / Cas system. In some embodiments, the trRNA comprises a truncated or modified wild type trRNA. The length of the trRNA depends on the CRISPR / Cas system used. In some embodiments, the trRNA comprises or consists of 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, or more than 100 nucleotides. In some embodiments, the trRNA may comprise certain secondary structures, such as, for example, one or more hairpin or stem-loop structures, or one or more bulge structures.
[0175] It will be appreciated that for methods that use the guide RNAs for a Cas nuclease, such as a Cas9 nuclease disclosed herein, the methods include the use of the CRISPR / Cas system.
[0176] Methods of Use
[0177] In some aspects, provided herein is a method of cleaving or modifying a target DNA, comprising: contacting the target DNA or a cell comprising the target DNA with: (a) a Cas9 variant described herein, a nucleic acid described herein, or a vector described herein; and (b) a Cas9 guide RNA (gRNA), wherein the Cas9 variant and the Cas9 gRNA form a complex that cleaves or modifies the target DNA. In some embodiments, the Cas9 gRNA comprises a gRNA-scaffold selected from SpCas9 gRNA scaffold, SaCas9 gRNA-scaffold, StlCas9 gRNA-scaffold, and FnCas9 gRNA-scaffold. In some embodiments, the Cas9 gRNA comprises a guide sequence of 8-40 nt in length, e.g., 8 nucleotides (nt), lOnt, 12nt, 20nt, 35nt, or 40 nt in length. In some embodiments, the Cas9 gRNA is a spCas9 gRNA.
[0178] In some aspects, provided herein is a method of cleaving or modifying a target DNA, comprising: contacting the target DNA or a cell comprising the target DNA with: (a) a Cas9 variant described herein, a nucleic acid described herein, or a vector described herein; (b) a Cas9 crRNA; and (c) a Cas9 tracrRNA; wherein the Cas9 variant, the Cas9 crRNA, and the Cas9 tracrRNA form a complex that cleaves or modifies the target DNA. In some embodiments, the Cas9 crRNA is selected from SpCas9 crRNA, SaCas9 crRNA, StlCas9 crRNA, and FnCas9 crRNA, and wherein the Cas9 tracrRNA is selected from the same species as the Cas9 crRNA. In some embodiments, the Cas9 crRNA comprises a guide sequence of 8-40 nt in length, e.g., 8 nt, lOnt, 12nt, 20nt, 35nt, or 40 nt in length.
[0179] In some aspects, provided herein is a method of modifying a target DNA or RNA comprising: contacting the target DNA or RNA or a cell comprising the target DNA or RNA with: (a) a Cas9-guided base editor described herein, a nucleic acid described herein, or a vector described herein; and (b) a Cas9 guide RNA (gRNA), wherein the Cas9-guided base editor and the Cas9 gRNA form a complex that modifies the target DNA or RNA. In some embodiments, the Cas9 gRNA comprises a gRNA-scaffold selected from SpCas9 gRNA scaffold, SaCas9 gRNA-scaffold, StlCas9 gRNA-scaffold, and FnCas9 gRNA-scaffold. In some embodiments, the Cas9 gRNA comprises a guide sequence of 8-40 nt in length, e.g., 8 nucleotides (nt), lOnt, 12nt, 20nt, 35nt, or 40 nt in length. In some embodiments, the Cas9 gRNA is a spCas9 gRNA.
[0180] In some aspects, provided herein is a method of modifying a target DNA or RNA comprising: contacting the target DNA or a cell comprising the target DNA or RNA with: (a) a Cas9-guided base editor described herein, a nucleic acid described herein, or a vector described herein; and (b) a Cas9 crRNA; and (c) a Cas9 tracrRNA; wherein the Cas9-guided base editor, the Cas9 crRNA, and the Cas9 tracrRNA form a complex that modifies the target DNA or RNA. In some embodiments, the Cas9 crRNA is selected from SpCas9 crRNA, SaCas9 crRNA, StlCas9 crRNA, and FnCas9 crRNA, and wherein the Cas9 tracrRNA is selected from the same species as the Cas9 crRNA. In some embodiments, the Cas9 crRNA comprises a guide sequence of 8-40 nt in length, e.g., 8 nt, lOnt, 12nt, 20nt, 35nt, or 40 nt in length.
[0181] In some aspects, provided herein is a method of modifying a target DNA or RNA comprising contacting the target DNA or a cell comprising the target DNA or RNA with a composition or kit described herein.
[0182] In some embodiments, modifying the target DNA or RNA comprises inducing an insertion, a deletion, and / or a mutation of the target DNA or RNA. In some embodiments, modifying the target DNA or RNA comprises correcting a mutation of the target DNA or RNA.
[0183] In some embodiments, modifying the target DNA or RNA comprises inducing a base change of the target DNA or RNA. In some embodiments, the base change is a A-to-G or A- to-I change. In some embodiments, the cell is a mammalian cell.
[0184] In some aspects, provided herein is a method of regulating activity and / or specificity of a catalytic domain linked to a Cas9 scaffold comprising expanding the Cas9 scaffold by inserting one or more domains selected from non-catalytic BH domain and REC domain. In some embodiments, the Cas9 scaffold is a SpCas9. In some embodiments, the non-catalytic BH domain and REC domain are from StlCas9.
[0185] In some embodiments, the catalytic domain is an N-terminal catalytic domain. In some embodiments, the catalytic domain is a base editor, such as those described herein.
[0186] In some embodiments, the catalytic domain has reduced (e.g., 5%, 10%, 15%, 20%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% reduced) byproducts compared to a catalytic domain that is linked to Cas9 or Cas9 nickase scaffold without the insertion of the BH and / or REC domains.
[0187] In some embodiments, the method comprises expanding the Cas9 scaffold by inserting one or more domains selected from a non-catalytic BH domain and a REC domain from StlCas9.
[0188] In some aspects, provided herein is a method of repositioning an N-terminal catalytic domain linked to Cas9, comprising adjusting the length of a BH domain. In some embodiments, the BH domain is a half (0.5x), single (lx), or two (2x) BH domains.
[0189] In some aspects, provided herein is a method of reducing bystander editing of a deaminase-based base editor, comprising modifying a Cas9 scaffold. In some embodiments, modifying the Cas9 scaffold comprises inserting one or more domains selected from a non- catalytic BH domain and a REC domain into the Cas9 scaffold. In some embodiments, modifying the Cas9 scaffold comprises inserting one or more domains selected from a non- catalytic StlCas9 BH domain and a StlCas9 REC domain into the Cas9 scaffold. In some embodiments, the Cas9 scaffold is a nSpCas9.
[0190] In some aspects, provided herein is a method of improving Cas9 binding specificity, comprising enlarging Cas9 by inserting one or more domains selected from a non-catalytic BH domain and a REC domain. In some embodiments, the method described herein comprises inserting a StlCas9 REC domain to the Cas9.
[0191] Methods of using a Cas nuclease, e.g., Cas9, are also well known in the art. It will be appreciated that, depending on the context, the Cas9 variant can be provided as a nucleic acid (e.g. , DNA or mRNA) or as a protein. In some embodiments, the present method can be practiced in a host cell that already expresses an Cas9 variant. In some embodiments, methods described herein modifies the target DNA by introducing an insertion, a deletion, and / or a substitution of the target DNA sequence.
[0192] The guide RNA, Cas9 variant, and an optional nucleic acid construct disclosed herein can be delivered to a host cell or subject, in vivo or ex vivo, using various known and suitable methods available in the art. The guide RNA Cas9 variant, and nucleic acid construct can be delivered individually or together in any combination, using the same or different delivery methods as appropriate.
[0193] Conventional viral and non- viral based gene delivery methods can be used to introduce the guide RNA disclosed herein as well as the Cas9 variant and / or donor construct in cells e.g., mammalian cells) and target tissues. As further provided herein, non-viral vector delivery systems nucleic acids such as non-viral vectors, plasmid vectors, and, e.g., naked nucleic acid, and nucleic acid complexed with a delivery vehicle such as a liposome, lipid nanoparticle (LNP), or poloxamer. Viral vector delivery systems include DNA and RNA viruses.
[0194] Methods and compositions for non-viral delivery of nucleic acids include electroporation, lipofection, microinjection, biolistics, virosomes, liposomes, immunoliposomes, LNPs, polycation or lipidmucleic acid conjugates, naked nucleic acid (e.g., naked DNA / RNA), artificial virions, and agent-enhanced uptake of DNA. Sonoporation using, e.g., the Sonitron 2000 system (Rich-Mar) can also be used for delivery of nucleic acids.
[0195] Additional exemplary nucleic acid delivery systems include those provided by AmaxaBiosystems (Cologne, Germany), Maxcyte, Inc. (Rockville, Md.), BTX Molecular Delivery Systems (Holliston, Ma.) and Copernicus Therapeutics Inc., (see for example U.S. Pat. No. 6,008,336). Lipofection is described in e.g., U.S. Pat. Nos. 5,049,386; 4,946,787; and 4,897,355) and lipofection reagents are sold commercially (e.g., Transfectam™ and Lipofectin™). The preparation of lipid: nucleic acid complexes, including targeted liposomes such as immunolipid complexes, is well known in the art, and as described herein.
[0196] Various delivery systems (e.g., vectors, liposomes, LNPs) containing the guide RNAs, RNA-guided DNA binding agent, and donor construct, singly or in combination, can also be administered to an organism for delivery to cells in vivo or administered to a cell or cell culture ex vivo. Administration is by any of the routes normally used for introducing a molecule into ultimate contact with blood, fluid, or cells including, but not limited to, injection, infusion, topical application and electroporation. Suitable methods of administering such nucleic acids are available and well known to those of skill in the art.
[0197] Additional Embodiments:
[0198] Embodiment 1. A Cas9 variant comprising a SpCas9 with an insertion of a StlCas9 recognition (REC) domain between residues 93D and 94D of the SpCas9.
[0199] Embodiment 2. The cas9 variant of embodiment 1, wherein the SpCas9 is a wild-type SpCas9 or a SpCas9 nickase (nSpCas9).
[0200] Embodiment 3. The cas9 variant of embodiment 2, wherein the SpCas9 nickase is nSPCas9 (D10A).
[0201] Embodiment 4. The cas9 variant of embodiment 2, wherein SpCas9 comprises an amino acid sequence of SEQ ID NO: 1.
[0202] Embodiment 5. The cas9 variant of any one of embodiments 1-4, wherein the StlCas9 REC domain comprises residues 74-466 of StlCas9.
[0203] Embodiment 6. The cas9 variant of any one of embodiments 1-5, comprising a linker from SaCas9 immediately downstream of the StlCas9 REC domain.
[0204] Embodiment 7. The cas9 variant of embodiment 6, wherein the linker from SaCas9 comprises residues T205 to D223 of SaCas9.
[0205] Embodiment 8. The cas9 variant of any one of embodiments 1-7, comprising a SpCas9 with an insertion of an amino acid sequence of SEQ ID NO: 8 between residues 93D and 94D of the SpCas9.
[0206] Embodiment 9. The cas9 variant of embodiment 9, comprising a SpCas9 with an insertion of an amino acid sequence of SEQ ID NO: 8 between residues 93D and 94D of SEQ ID NO: 1.
[0207] Embodiment 10. A nucleic acid encoding the Cas9 variant of any one of embodiments 1-9. Embodiment 11. A composition or kit, comprising:
[0208] (a) a Cas9 variant of any one of embodiments 1-9, or a nucleic acid of embodiment 10; and
[0209] (b) a Cas9 guide RNA (gRNA).
[0210] Embodiment 12. A composition or kit, comprising:
[0211] (a) a Cas9 variant of any one of embodiments 1-9, or a nucleic acid of embodiment 10;
[0212] (b) a Cas9 crRNA; and
[0213] (c) a Cas9 tracrRNA.
[0214] Embodiment 13. A method of cleaving or modifying a target DNA, comprising: contacting the target DNA or a cell comprising the target DNA with:
[0215] (a) a Cas9 variant of any one of embodiments 1-9, or a nucleic acid of embodiment 10; and
[0216] (b) a Cas9 guide RNA (gRNA); wherein the Cas9 variant and the Cas9 gRNA form a complex that cleaves or modifies the target DNA.
[0217] Embodiment 14. A method of cleaving or modifying a target DNA, comprising: contacting the target DNA or a cell comprising the target DNA with:
[0218] (a) a Cas9 variant of any one of embodiments 1-9, or a nucleic acid of embodiment 10;
[0219] (b) a Cas9 crRNA; and
[0220] (c) a Cas9 tracrRNA; wherein the Cas9 variant, the Cas9 crRNA, and the Cas9 tracrRNA form a complex that cleaves or modifies the target DNA.
[0221] Embodiment 15. A method of cleaving or modifying a target DNA, comprising contacting the target DNA or a cell comprising the target DNA with a composition or kit of embodiment 11 or 12.
[0222] Embodiment 16. A Cas9-guided base editor comprising a base editor linked to a Cas9 variant, wherein the Cas9 variant comprising a SpCas9 nickase (nSpCas9) with an insertion of a StlCas9 REC domain between residues 93D and 94D of the nSpCas9.
[0223] Embodiment 17. The Cas9-guided base editor of embodiment 16, wherein the base editor is linked to N-terminus of the Cas9 variant.
[0224] Embodiment 18. The Cas9-guided base editor of embodiment 16 or 17, wherein the SpCas9 nickase is nSpCas9 (D10A). Embodiment 19. The Cas9-guided base editor of any one of embodiments 16-18, wherein the StlCas9 REC domain comprises residues 74-466 of StlCas9.
[0225] Embodiment 20. The Cas9-guided base editor of any one of embodiments 16-19, comprising a linker from SaCas9 immediately downstream of the StlCas9 REC domain.
[0226] Embodiment 21. The Cas9-guided base editor of embodiment 20, wherein the linker from SaCas9 comprises residues T205 to D223 of SaCas9.
[0227] Embodiment 22. The Cas9-guided base editor of any one of embodiment 16-21, wherein the Cas9 variant comprising half (0.5x), single (lx), or two (2x) Bridge helix (BH) domains immediately upstream of the StlCas9 REC domain.
[0228] Embodiment 23. The Cas9-guided base editor of embodiment 22, wherein the Cas9 variant comprising single BH domain immediately upstream of the StlCas9 REC domain.
[0229] Embodiment 24. The Cas9-guided base editor of any one of embodiment 16-23, wherein the Cas9 variant comprises a nSpCas9 with an insertion of an amino acid sequence of SEQ ID NO: 3, 7, or 8 between residues 93D and 94D of the nSpCas9.
[0230] Embodiment 25. The Cas9-guided base editor of embodiment 24, wherein the base editor is an adenine deaminase.
[0231] Embodiment 26. The Cas9-guided base editor of embodiment 25, wherein the adenine deaminase is TadA8e.
[0232] Embodiment 27. The Cas9-guided base editor of embodiment 25, wherein the adenine deaminase is TadA9.
[0233] Embodiment 28. The Cas9-guided base editor of embodiment 25, wherein the adenine deaminase is TadA8e(N46L).
[0234] Embodiment 29. A nucleic acid encoding the Cas9-guided base editor of any one of embodiments 16-28.
[0235] Embodiment 30. A composition or kit, comprising:
[0236] (a) a Cas9-guided base editor of any one of embodiments 16-28, or a nucleic acid of embodiment 29; and
[0237] (b) a Cas9 guide RNA (gRNA).
[0238] Embodiment 31. The composition or kit of embodiment 30, wherein the Cas9 gRNA scaffold is selected from SpCas9 gRNA-scaffold, SaCas9-gRNA-scaffold, StlCas9-gRNA- scaffold, and FnCas9-gRNA-scaffold.
[0239] Embodiment 32. The composition or kit of embodiment 30 or 31, wherein the guide RNA contains 8nt, lOnt, 12nt, 20nt, 35nt, or 40 nt guide sequence.
[0240] Embodiment 33. A composition or kit, comprising: (a) a Cas9-guided base editor of any one of embodiments 16-28, or a nucleic acid of embodiment 29;
[0241] (b) a Cas9 crRNA; and
[0242] (c) a Cas9 tracrRNA.
[0243] Embodiment 34. A method of modifying a target DNA or RNA comprising: contacting the target DNA or RNA or a cell comprising the target DNA or RNA with:
[0244] (a) a Cas9-guided base editor of any one of embodiments 16-28, or a nucleic acid of embodiment 29; and
[0245] (b) a Cas9 guide RNA (gRNA); wherein the Cas9-guided base editor and the Cas9 gRNA form a complex that induces a base change of the target DNA or RNA.
[0246] Embodiment 35. The method of embodiment 34, wherein the Cas9 gRNA scaffold is selected from SpCas9 gRNA- scaffold, SaCas9-gRNA-scaffold, StlCas9-gRNA-scaffold, FnCas9-gRNA-scaffold.
[0247] Embodiment 36. The method of embodiment 34 and 35, wherein the guide RNA contains 8nt, lOnt, 12nt, 20nt, 35nt, or 40 nt guide sequence.
[0248] Embodiment 37. A method of modifying a target DNA or RNA comprising: contacting the target DNA or a cell comprising the target DNA or RNA with:
[0249] (a) a Cas9-guided base editor of any one of embodiments 16-28, or a nucleic acid of embodiment 29; and
[0250] (b) a Cas9 crRNA; and
[0251] (c) a Cas9 tracrRNA; wherein the Cas9-guided base editor, the Cas9 crRNA, and the Cas9 tracrRNA form a complex that induces a base change of the target DNA or RNA.
[0252] Embodiment 38. A method of modifying a target DNA or RNA comprising contacting the target DNA or a cell comprising the target DNA or RNA with a composition or kit of any one of embodiments 30-33.
[0253] Embodiment 39. The method of any one of embodiments 34-38, wherein the base change is a A-to-G or A-to-I change.
[0254] Embodiment 40. The method of any one of embodiments 13-15 and 34-38, wherein the cell is a mammalian cell.
[0255] Embodiment 41. A Cas9 variant comprising an enlarged SpCas9 with an insertion of one or more domains selected from non-catalytic BH domain and REC domain. Embodiment 42. The Cas9 variant of embodiment 41, comprising an enlarged SpCas9 with an insertion of one or more domains from non-catalytic StlCas9 BH domain and StlCas9 REC domain.
[0256] Embodiment 43. The Cas9 variant of embodiment 41 or 42, comprising an enlarged SpCas9 with an insertion of a StlCas9 REC domain.
[0257] Embodiment 44. A method of producing an enlarged SpCas9 of any one of embodiments 41-43.
[0258] Embodiment 45. A method of regulating activity and / or specificity of a catalytic domain linked to a Cas9 scaffold comprising expanding the Cas9 scaffold by inserting one or more domains selected from non-catalytic BH domain and REC domain.
[0259] Embodiment 46. The method of embodiment 45, wherein the catalytic domain is an N-terminal catalytic domain.
[0260] Embodiment 47. The method of embodiment 46, wherein the N-terminal catalytic domain has reduced byproducts.
[0261] Embodiment 48. A method of repositioning an N-terminal catalytic domain linked to Cas9, comprising adjusting the length of the BH domain.
[0262] Embodiment 49. A method of reducing bystander editing of a deaminase-based base editor, comprising modifying Cas9 scaffold.
[0263] Embodiment 50. The method of embodiment 49, wherein one or more domains selected from non-catalytic BH domain and REC domain is inserted into Cas9 scaffold.
[0264] Embodiment 51. The method of embodiment 49 or 50, wherein one or more domains selected from non-catalytic StlCas9 BH domain and StlCas9 REC domain is inserted into Cas9.
[0265] Embodiment 52. A method of improving Cas9 binding specificity, comprising enlarging Cas9 by inserting one or more domains selected from non-catalytic BH domain and REC domain.
[0266] Embodiment 53. The method of embodiment 52, comprising inserting a StlCas9 REC domain to the Cas9.
[0267] Embodiment 54. A Cas9 variant, comprising SpCas9 (D10A) with an insertion of BH and REC domains of StlCas9.
[0268] Embodiment 55. The Cas9 variant of embodiment 54, comprising SpCas9 (D10A) with an insertion of an amino acid sequence of SEQ ID NO: 3. Embodiment 56. The Cas9 variant of embodiment 54 or 55, comprising SpCas9 (D10A) with an insertion of an amino acid sequence of SEQ ID NO: 3 between residues 93D and 94D of the SpCas9 (D10A).
[0269] Embodiment 57. The Cas9 variant of any one of embodiments 54-56, wherein the Cas9 variant is used as a binding scaffold to generate a based editor.
[0270] In some aspects, provided herein is a Cas9 variant comprising a SpCas9 with an insertion of a StlCas9 recognition (REC) domain between residues 93D and 94D of the SpCas9.
[0271] In some embodiments, the SpCas9 is a wild-type SpCas9 or a SpCas9 nickase (nSpCas9). In some embodiments, the SpCas9 nickase is nSpCas9 (D10A). In some embodiments, SpCas9 comprises an amino acid sequence of SEQ ID NO: 1 . In some embodiments, the StlCas9 REC domain comprises residues 74-466 of StlCas9. In some embodiments, the cas9 variant comprises a linker from SaCas9 immediately downstream of the StlCas9 REC domain. In some embodiments, the linker from SaCas9 comprises residues T205 to D223 of SaCas9. In some embodiments, the cas9 variant comprises a SpCas9 with an insertion of an amino acid sequence of SEQ ID NO: 8 between residues 93D and 94D of the SpCas9. In some embodiments, the cas9 variant comprises a SpCas9 with an insertion of an amino acid sequence of SEQ ID NO: 8 between residues 93D and 94D of SEQ ID NO: 1.
[0272] In some aspects, provided herein is a nucleic acid encoding the Cas9 variant described herein.
[0273] In some aspects, provided herein is a composition or kit, comprising:
[0274] (a) a Cas9 variant described herein, or a nucleic acid described herein; and
[0275] (b) a Cas9 guide RNA (gRNA).
[0276] In some aspects, provided herein is a composition or kit, comprising:
[0277] (a) a Cas9 variant of described herein, or a nucleic acid described herein;
[0278] (b) a Cas9 crRNA; and
[0279] (c) a Cas9 tracrRNA.
[0280] In some aspects, provided herein is a method of cleaving or modifying a target DNA, comprising: contacting the target DNA or a cell comprising the target DNA with:
[0281] (a) a Cas9 variant described herein, or a nucleic acid described herein; and
[0282] (b) a Cas9 guide RNA (gRNA); wherein the Cas9 variant and the Cas9 gRNA form a complex that cleaves or modifies the target DNA. In some aspects, provided herein is a method of cleaving or modifying a target DNA, comprising: contacting the target DNA or a cell comprising the target DNA with:
[0283] (a) a Cas9 variant described herein, or a nucleic acid described herein;
[0284] (b) a Cas9 crRNA; and
[0285] (c) a Cas9 tracrRNA; wherein the Cas9 variant, the Cas9 crRNA, and the Cas9 tracrRNA form a complex that cleaves or modifies the target DNA.
[0286] In some aspects, provided herein is a method of cleaving or modifying a target DNA, comprising contacting the target DNA or a cell comprising the target DNA with a composition or kit described herein.
[0287] In some aspects, provided herein is a Cas9-guided base editor comprising a base editor linked to a Cas9 variant, wherein the Cas9 variant comprises a SpCas9 nickase (nSpCas9) with an insertion of a StlCas9 REC domain between residues 93D and 94D of the nSpCas9.
[0288] In some embodiments, the base editor is linked to N-terminus of the Cas9 variant. In some embodiments, the SpCas9 nickase is nSpCas9 (D10A). In some embodiments, the StlCas9 REC domain comprises residues 74-466 of StlCas9. In some embodiments, the Cas9-guided base editor comprises a linker from SaCas9 immediately downstream of the StlCas9 REC domain. In some embodiments, the linker from SaCas9 comprises residues T205 to D223 of SaCas9. In some embodiments, the Cas9 variant comprises half (0.5x), single (lx), or two (2x) Bridge helix (BH) domains immediately upstream of the StlCas9 REC domain. In some embodiments, the Cas9 variant comprises single BH domain immediately upstream of the StlCas9 REC domain. In some embodiments, the Cas9 variant comprises a nSpCas9 with an insertion of an amino acid sequence of SEQ ID NO: 3, 7, or 8 between residues 93D and 94D of the nSpCas9. In some embodiments, the base editor is an adenine deaminase. In some embodiments, the adenine deaminase is TadA8e. In some embodiments, the adenine deaminase is TadA9. In some embodiments, the adenine deaminase is TadA8e(N46L).
[0289] In some aspects, provided herein is a nucleic acid encoding the Cas9-guided base editor described herein.
[0290] In some aspects, provided herein is a composition or kit, comprising:
[0291] (a) a Cas9-guided base editor described herein or a nucleic acid described herein; and
[0292] (b) a Cas9 guide RNA (gRNA). In some embodiments, the Cas9 gRNA scaffold is selected from SpCas9 gRNA- scaffold, SaCas9-gRNA-scaffold, StlCas9-gRNA-scaffold, and FnCas9-gRNA-scaffold. In some embodiments, the guide RNA contains 8nt, lOnt, 12nt, 20nt, 35nt, or 40 nt guide sequence.
[0293] In some aspects, provided herein is a composition or kit, comprising:
[0294] (a) a Cas9-guided base editor described herein, or a nucleic acid described herein;
[0295] (b) a Cas9 crRNA; and
[0296] (c) a Cas9 tracrRNA.
[0297] In some aspects, provided herein is a method of modifying a target DNA or RNA comprising: contacting the target DNA or RNA or a cell comprising the target DNA or RNA with:
[0298] (a) a Cas9-guided base editor described herein or a nucleic acid described herein; and
[0299] (b) a Cas9 guide RNA (gRNA); wherein the Cas9-guided base editor and the Cas9 gRNA form a complex that induces a base change of the target DNA or RNA.
[0300] In some embodiments, the Cas9 gRNA scaffold is selected from SpCas9 gRNA- scaffold, SaCas9-gRNA-scaffold, StlCas9-gRNA-scaffold, FnCas9-gRNA-scaffold. In some embodiments, the guide RNA contains 8nt, lOnt, 12nt, 20nt, 35nt, or 40 nt guide sequence.
[0301] In some aspects, provided herein is a method of modifying a target DNA or RNA comprising: contacting the target DNA or a cell comprising the target DNA or RNA with:
[0302] (a) a Cas9-guided base editor described herein, or a nucleic acid described herein; and
[0303] (b) a Cas9 crRNA; and
[0304] (c) a Cas9 tracrRNA; wherein the Cas9-guided base editor, the Cas9 crRNA, and the Cas9 tracrRNA form a complex that induces a base change of the target DNA or RNA.
[0305] In some aspects, provided herein is a method of modifying a target DNA or RNA comprising contacting the target DNA or a cell comprising the target DNA or RNA with a composition or kit described herein. In some embodiments, the base change is a A-to-G or A- to-I change. In some embodiments, the cell is a mammalian cell.
[0306] In some aspects, provided herein is a Cas9 variant comprising an enlarged SpCas9 with an insertion of one or more domains selected from non-catalytic BH domain and REC domain. In some embodiments, the Cas9 variant comprises an enlarged SpCas9 with an insertion of one or more domains from non-catalytic StlCas9 BH domain and StlCas9 REC domain. In some embodiments, the Cas9 variant comprises an enlarged SpCas9 with an insertion of a StlCas9 REC domain.
[0307] In some aspects, provided herein is a method of producing an enlarged SpCas9 described herein.
[0308] In some aspects, provided herein is a method of regulating activity and / or specificity of a catalytic domain linked to a Cas9 scaffold comprising expanding the Cas9 scaffold by inserting one or more domains selected from non-catalytic BH domain and REC domain. In some embodiments, the catalytic domain is an N-terminal catalytic domain. In some embodiments, the N-terminal catalytic domain has reduced byproducts.
[0309] In some aspects, provided herein is a method of repositioning N-terminal catalytic domain linked to Cas9 comprising adjusting length of BH domain.
[0310] In some aspects, provided herein is a method of reducing bystander editing of a deaminase-based base editor comprising modifying Cas9 scaffold. In some embodiments, one or more domains selected from non-catalytic BH domain and REC domain is inserted into Cas9 scaffold. In some embodiments, one or more domains selected from non-catalytic StlCas9 BH domain and StlCas9 REC domain is inserted into Cas9.
[0311] In some aspects, provided herein is a method of improving Cas9 binding specificity comprising enlarging Cas9 by inserting one or more domains selected from non-catalytic BH domain and REC domain. In some embodiments, the method comprises inserting a StlCas9 REC domain to the Cas9.
[0312] In some aspects, provided herein is a Cas9 variant comprising SpCas9 (D10A) with an insertion of BH and REC domains of StlCas9. In some embodiments, the Cas9 variant comprises SpCas9 (D10A) with an insertion of an amino acid sequence of SEQ ID NO: 3. In some embodiments, the Cas9 variant comprises SpCas9 (D10A) with an insertion of an amino acid sequence of SEQ ID NO: 3 between residues 93D and 94D of the SpCas9 (D10A). In some embodiments, the Cas9 variant is used as a binding scaffold to generate a based editor.
[0313] EXAMPLES
[0314] Example 1: Materials and methods for Examples 2-8
[0315] Analysis of sequences Alignment analysis and annotation analysis were manually performed with Geneious Prime 11.0.18. Filtration of sequences was manually performed using Microsoft Excel. NCBI blastp (protein-protein BLAST) was used to double check candidate-sequences.
[0316] Construction of Plasmid
[0317] We used general cloning methods including Gibson assembly with GeneArt Gibson Assembly HiFi Master Mix (Thermo Fisher, A46628) and Quick ligation kit (New England Biolabs, M2200S) with variety of type IIS restriction enzymes. One Shot™ MAX Efficiency™ DH5ot-TlR Competent Cells (Thermo Fisher, 12297016) was used for DNA cloning. The sequences of cloned constructs were confirmed by sanger sequencing following extraction with QIAprep kit (QIAGEN, 27106). DNA sequences of plasmids including vectors purchased from Addgene used in this study can be found in Example 8. sgRNA target sites are available in Table 3 and Tables 4A-4C. Oligonucleotides used in this research can be found in Table 5. EGFP expression plasmids containing amino acid substitutions were generated by standard PCR with Q5 Site-Directed Mutagenesis Kit (New England Biolabs, E0554S). Homan-codon- optimized fragments used for REC expansions, oligonucleotides and sgRNA expression plasmids were synthesized by IDT (Integrated DNA Technologies, USA). Unless otherwise indicated, all sgRNAs were designed to target sites containing a 5’ guanine nucleotide.
[0318] Cell Culture and Transfection
[0319] HEK293T cells and HEK293-deGFP cells (Andreatta, Nahreini et al. 2001) (a gift from Prof. Eben Alsberg, University of Illinois at Chicago) were cultured in DMEM (Sigma, D6429) supplemented with 10% FBS and 1% Gibco™ Penicillin-Streptomycin (10,000 U / mL) at 37 °C and 5% CO2. Cells were seeded one day prior to transfection in 24-well plates. For cpREADER assay, cells were plated at a density of approximately 70,000 cells per well, cells were transfected with 200ng of nuclease or base editor expression plasmid DNA, 200 ng of sgRNA plasmid, and 200 ng or 10 ng of reporter plasmid (200 ng of EGFPstop plasmid, 10 ng of pCMV-GFP plasmid). Unless otherwise noted, cells were plated at a density of approximately 50,000 cells per well for in vitro cell genome editing, cells were transfected with 400ng of nuclease or base editor expression plasmid DNA and 200 ng of sgRNA plasmid per well. Transfections were performed with Lipofectamine 2000 (Invitrogen, 11668027) and Lipofectamine 3000 (Thermo Fisher, L3000008) according to the manufacturer’s recommended protocol in cpREADER assay and genome editing, respectively. Of note, the ratio of lipofectamine to DNA was set at 2.5:1.
[0320] EGFPstop mRNA activation assay was performed as the following protocol. mRNA carrying a stop codon UAG in place of UGG at the 58-position. mRNA was prepared as our previous study (Gao, Guan et al. 2024). HEK293T cells were plated at a density of approximately 70,000 cells per well in 24-well plates. Cells were transfected with 200 ng of ABE8e plasmid or ABE9 plasmid using Lipofectamine 3000. After 6 hours, cells were transfected with 400 ng of mRNA using Lipofectamine 2000. Then 48 hours, cells were analyzed using BD Accuri™ C6 Flow Cytometer.
[0321] Genomic DNA isolation
[0322] Cells were harvested 3 days post transfection. Genomic DNA was extracted from transfected cells using DNeasy Blood & Tissue Kit (Qiagen, 69504) following the manufacture’s protocol. Extracted DNA was normalized to a final concentration of 50 ng or 100 ng per pl with ddtFO.
[0323] Flow Cytometry Detection
[0324] Adherent cells were treated with 100 pl of 0.25% Trypsin-EDTA (Gibco, 25200056) incubated for 5 min at 37 °C to completely detach cells. 400 pl of DMEM was used to stop trypsin digestion. Samples were applied to the BD Accuri™ C6 Flow Cytometer (BD Biosciences) directly, and GFP fluorescence was measured. Data analysis was performed with BD Accuri C6 Software and Microsoft Excel. Prism (GraphPad) was used to generate column graph and for P value calculations. EGFP disruption or EGFPstop activation experiments in cpREADER assays were performed as previously described (Gao, Guan et al. 2024). Briefly, transfected cells were analyzed 48h after transfection for loss or restoration of EGFP fluorescence. The background was determined by gating a negative control transfection.
[0325] RNA isolation from mammalian cells
[0326] HEK293T cells were transfected with nSpCas9-EGFP, ABE8e-EGFP and GS-ABE8e- EGFP plasmid. 48 hours post transfection, approximately 300,000 cells were harvested for mRNA isolation. RNA isolation was performed with the RNeasy Plus Mini Kit (QIAGEN, 74134) according to the manufacturer’s instructions. In short, RNA isolation began with removal of the culture medium and washing of the cells with lx PBS (Thermo Fisher Scientific). 350 pl of RLT Plus Buffer was added into each well; cells were homogenized by pipetting and transferred into a DNA eliminator column, and the subsequent binding and washing steps for RNA isolation using the RNeasy columns were performed as recommended by the manufacturer. Upon elution of RNA from the RNeasy column with 45 pl of RNase (ribonuclease) free water (QIAGEN), 2 pl of RNase inhibitor (New England Biolabs, M0314S) was added to prevent RNA degradation, and RNA was stored at -80°C. cDNA Synthesis cDNA synthesis was performed using ProtoScript II First Strand cDNA Synthesis Kit (New England Biolabs, E6560) with 1 pg ofRNA in a total reaction volume of 20 pl. Reactions were incubated for 2 hr at 42 °C. 2 pl cDNA was input into 50 pl NEBNext High-Fidelity 2X PCR Master Mix (New England Biolabs, M0541S) containing specific primers. The purification of target products was performed with GeneJET PCR Purification Kit (Thermo Fisher, K0702).
[0327] Preparation of genomic DNA amplicons for deep sequencing
[0328] Targeted regions flanking the on-target or off-targeted sites were amplified with specific primers and Q5 Hot Start High-Fidelity 2X Master Mix (NEB, M0494S) under the following thermal cycling conditions: one cycle, 98°C, 1 min; one cycle, 98°C, 30 s; 35 cycles, 98°C, 10 s, 65°C, 30 s, 72°C, 15s; one cycle, 72°C, 2 min; 4°C hold. 100 ng of isolated genomic DNA was input into each 50 pl of PCR. PCR products were analyzed on an agarose gel electrophoresis system to verify both size and purity. PCR products were purified with GeneJET PCR Purification Kit (Thermo Fisher, K0702) followed by normalizing to a final concentration of 20 ng per pl with ddH2O.
[0329] Analysis of HTS data for targeted amplicon sequencing
[0330] Targeted amplicon sequencing was carried out by Genwiz (Azenta, South Plainfield, NJ, US) with Amplicon-EZ protocol. Batch analysis with CRISPResso2 pipeline was used for targeted amplicon sequencing. For DNA analysis, a 30-bp window was used to quantify indels around the DNA nick site. Otherwise, the default parameters were used for analysis. The output file “NUCLEOTIDE_PERCENTAGE_TABLE.txt” was imported into Microsoft Excel for quantification of editing frequencies and “Indel_histogram.txt” for quantification of indel frequencies. Indel percentage was re-checked with BEanalyzer if requested. For analysis of RNA amplicon editing, no sgRNA flag was used. Instead, the output file “NUCLEOTIDE_PERCENTAGE_TABLE.txt” was imported into Microsoft Excel for analysis of A-to-G editing rates associated with each sample (inosine in RNA is read as a guanosine by polymerases).
[0331] Prism (GraphPad) was used to generate dot plots and bar plots of these data. For instances in the text where means have been calculated across multiple genomic or transcriptomic loci, the SDs reported represent the SD of the mean for all biological replicates.
[0332] Orthogonal R-loop assay
[0333] Orthogonal R-loop assays were performed to measure Cas9-independent off-target editing as described previously, with minor modifications. Under standard conditions, 200 ng of base editor plasmid, 300 ng of dSaCas9 plasmid, 100 ng of SpCas9 sgRNA plasmid and 100 ng of SaCas9 sgRNA plasmid were co-transfected into HEK293T cells using 1.5 pl of Lipofectamine 3000. Cells were cultured for 3 days after treatment followed by genomic DNA isolation. mRNA production for ABE editors (Gaudelli, Lam et al. 2020, Neugebauer, Hsu et al. 2023)
[0334] Base editor mRNA was generated from PCR product amplified from a plasmid template. Plasmid contains a dead T7 promoter followed by a 5’ untranslated region (UTR), Kozak sequence, open reading and 3’ UTR. The dead T7 promoter carries an inactivating point mutation within the T7 promoter that prevents transcription from circular plasmid in E. coli. PCR reaction was performed with Q5 Hot Start 2X Master Mix, in which the forward primer G200 corrected the SNP within the T7 promoter and the reverse primer G201 appended a poly A tail to the 3’UTR. 20 ng of plasmid was input into each of 50 pl PCR volume. PCR products were verified with gel detection followed by purification with GeneJET PCR Purification Kit. DNA was eluted from column with nuclease free ddH2O. IVT reactions were performed using HiScribe T7 High-Yield RNA Synthesis Kit (New England Biolabs, E2040S) in 40 pl volume according to the manufacturer’s protocol for Co-transcriptional capping using CleanCap® Reagent AG from TriLink but with full substitution of N^methyl-pseudouridine (TriLink BioTechnologies, N-1081) in place of uridine and co-transcriptional capping with CleanCap reagent AG (TriLink BioTechnologies, N-7113). mRNA isolation was performed with Monarch RNA Cleanup Kit (New England Biolabs, T20240S). mRNA was eluted from column by 30 pl RNase free ddH2O. An aliquot of the elution was diluted fivefold for quantification by NanoDrop. Samples were normalized to 2 pg pl1and stored at -80 °C.
[0335] Statistical Analysis and Graphical Illustrations.
[0336] Curve plotting and statistical analysis were performed using Prism 8 (GraphPad, La Jolla, CA). Data are shown as means ± standard error of the mean for groups of two or more replicates or as individual values with the mean indicated. Graphical illustrations were created using BioRender (https: / / biorender.com / ) and Office Power point.
[0337] Example 2: Comparative analysis of various Cas9 sequences from public database
[0338] It was unknow whether a pair of mutual comparisons of Cas9-expansion existed in nature. We first tried to search pairs candidate sequences assessable for side-by-side comparing REC expansion. As a pair of mutual comparisons, we proposed that they should differ only in protein size and the REC lobe domain. Han and colleagues predicted 161,859 IsrB_IscB_Cas9 sequences from the NCBI database (FIG.8B), among which there are 138,334 Cas9 sequences. However, only a minuscule fraction of these sequences has been experimentally characterized and structurally resolved (Nishimasu, Ran et al. 2014, Nishimasu, Cong et al. 2015, Hirano, Gootenberg et al. 2016, Jiang, Taylor et al. 2016, Yamada, Watanabe et al. 2017, Sun, Yang et al. 2019, Schuler, Hu et al, 2022). We next compared characteristics of seven representive IscB and Cas9 proteins. FnCas9, from Francisella novicida, is the largest Cas9 identified by crystal structure and biochemistry experiments and carrys a large REC lobe with 775 amino acid (aa) (Hirano, Gootenberg et al. 2016, Acharya, Mishra et al. 2019). Compared to the ancestor OgeuIscB, the size of FnCas9 shows a 1.55-fold increase. However, the REC lobe domain of FnCas9 is 20.4-fold as large as that of IscB. Meanwhile, the variations in other-domains-size range from a 1.0-fold to a 3.08-fold. The required protospacer adjacent motif (PAM) likely becomes proportionally shorter for Cas9 with large size (FIG. IB and FIG. 8C). Cas9 sequences annotated in same genus or species showed wide range diversification in protein sizes (FIG. 1C). We reasoned that this diversity was driven by both horizontal and vertical transfer of type II CRISPR-Cas systems (Fonfara, Le Rhun et al. 2014). However, domainorganization of Cas9s with different sizes has maintained a remarkably high level of conservatism throughout long evolutionary history (Makarova, Wolf et aL 2020, Altae-Tran, Kannan et al, 2023). The REC lobe was reported as a mediator for regulating HNH domain during Cas9-activation (Chen, Dagdas et al, 2017, Shams, Higgins et al, 2021, Pacesa, Loeff et al. 2022). We proposed that REC lobe may also serve as a mediator for matching and balancing the protein-size expansion during the evolutionary process of Cas9s, thereby maintaining the functional conformation of the protein.
[0339] To evaluate characteristics of various Cas9 sequences, we first filtered 93 of 161,859 seqences with the size larger than 1629aa (FnCas9), then got 45 unique- sequences after removal of 48 repeat sequences with alignment (Table 1, Example 8). These 45 proteins carry an average size of 1712aa, in which CP041030.1 annotated from Francisella sp. LA112445 strain carries the predicted maximum REC lobe size reaching 841aa. The maximum REC lobe surpasses that of CjCas9, SpCas9 and FnCas9 by 491aa, 217aa and 65aa, respectively (FIG. 9). Next, we systemically analyzed 45,696 Cas9 sequences annotated from five kinds of organism (genus Francisella, species Streptococcus pyogenes, Streptococcus thermophilus, Streptococcus aureus, Neisseria meningitidis and Campylobacter jejuni) with removal of repeats (FIG. 10-15). Except for 5. pyogenes, we found clues of REC lobe expansions in Cas9 sequences (FIG. ID). Likewise, the predicted protein UHCS01000002.1 isolated from S. aureus carries 318aa more than SaCas9 and contains a predicted REC lobe with larger 268aa than that of SaCas9. REC lobe expansions appear at inside (for 5. thermophilus, S. aureus, N. meningitidis and C. jejuni) or terminus (for Francisella) of REC domains. Most notably, UHCS01000002.1 shows an 88.3% pairwise sequence identity with SpCas9 (FIG. 16), LR822033.1 and VBTKO 1000005.1 sequences isolated from .S’. thermophilus have 99.4% and 56.9% identity with SpCas9, respectively (FIG. 17). Domain exchange and recombination may happen between related species, and continuing transfers seem to also occur. These results also agree with that FnCas9 and SpCas9 were younger than StCas9 or SaCas9 or NmeCas9 (Fonfara, Le Rhun et al, 2014),
[0340] Together, bioinformatics analysis above shows that REC expansions appear in Cas9s from most microbial genus or species. However, these expansion- like Cas9 sequences contains many residue substitutions with each other outside of REC domain (FIG. ID). Unfortunetly, we did not found a pair of mutual comparisions to access Cas9 expansion. Therefore, we tried to address this question using the strategy of artificial building life to gain biological insights (Elowitz and Lim 2010, Alonso-Lerma, Jabalera et al. 2023).
[0341] Example 3: Design of SpCas9 variants with REC lobe expansion
[0342] As one of youngest Cas9s, it is unknown whether and how REC expansion could happen to SpCas9. To our knowledge, SpCas9 is currently the most extensively studied in terms of crystal structure and mechanism of function, which enable it to be the most widely used in academic and industrial fields. We sought to generate novel SpCas9 variants by enlarging the REC lobe size. Further alignment analysis of three Cas9s, SaCas9, StlCas9 and SpCas9, indicated that StlCas9 has lower identify with SpCas9 than SaCas9 (FIG. 18). Compared to SaCas9 (PDB 5axw), StlCas9 (PDB 6m0v) carries a unique wing region (Q167 to N182) composed of two 0 strands inside REC domain composed of RECI and REC2 domain linked by a large flexible linker with 50aa length (R196 to S245) (FIG. IE) (Zhang, Zhang et al, 2020). It was reported that chimeras generated by swapping PI domians of St3Cas9 and SpCas9 retained DNA cleavage activity (Nishimasu, Ran et al. 2014). Superposition of the bioinformatics analysis above, we therefore chose inserting the StlCas9 into SpCas9 to generarte REC expansion variants. Flexible linker regions in Cas9 appear to play a role in the inactive-to-active conformational transition of multiple domains, we selected two short- flexible-linkers from SaCas9 (T205 to D223, K420 to T436) to link insertion with receptor in variants (FIG. IF)
[0343] Next, we designed and constructed four variants (FIG. 1G), in which diferent REC elements were inserted into N- or C-terminus of REC domain in SpCas9. Human cells always have advantage than microbes in expressing large protein (Nevers, Glover et aL 2023). In order to avoid that Cas9 variants may fail in function in mammalian cells, we chose human embryonic kidney (HEK) 293T cells to evaluate variants activity with EGFP plasmid interfere assay (FIG. 1H). Unfortunately, we didn’t observe EGFP disruption activities from all four variants (FIG. II). Bridge helix (BH) domains ranges from 32aa to 37aa in IscBs and Cas9s, which indicates the length of BH strongly conserved. Therefore, we removed the BH-Stl in Left-BH-REC12 to generate Left-REC12 variant. The new variant still didn’t induce reduction of GFP+ cell population, but lead a 20% decrease in mean fluorescence intensity (MFI) of GFP+ cells (FIG. 1 J). Excetingly, we observed a significant decrease in the GFP+ cells in presence of Left-REC12 variant when the substrate plasmid concentration reduced to one-fifth of the previous level (FIG. IK). Then, we tried to use a human cell line bearing an integrated construct that constitutively expresses an deGFPprotein, HEK293-deGFP (Andreatta, Nahreini et al. 2001). to eveluate Left-REC12 activitiy on chromosome site. We observed 45.5 % to 70% relative GFP disruption activity to SpCas9 in presence of various doses of Left-REC12 variant (FIG. IL). We did not know what causes of great differences in EGFP disruption in presence of different copies of DNA substrate. We named Left-REC12 giant SpCas9, GS-Cas9 (FIG. IM). The GS-Cas9 protein has a length of 1780 aa and a 207kDa molecular weight , which were larger 1.3 times than that of SpCas9. The designed REC domain of GS-Cas9 consists of 1036 amino acids making it 1.66-fold as large as thesize of that in SpCas9. This represents not only the largest Cas9 protein size but also the largest REC domain size experimentally characterized to date.
[0344] Compared to SpCas9, GS-Cas9 exhibits lower activity in mammalian cells. We hypothesized that the reduced activity may be attributed to disturbance in the conformation of the RuvC domain following REC expansion, potentially blocking cleavage activity on the nontarget strand (NTS) (FIG. 19). The REC3 domain of the SpCas9 protein plays a crucial regulatory role in the conformational rearrangement of the HNH domain (Chen, Dagdas et aL 2017). Therefore, we hypothesize that the insertion of a foreign REC domain at the C-terminus of REC of SpCas9 may form a novel recombinant structure with disrupting right regulation of the HNH domain, consequently resulting in blocking its cleavage activity on the target strand (TS).
[0345] Example 4: REC expansion enhanced SpCas9 targeting specificity
[0346] It was predicted that the acquisition of REC domain could potentially enhance Cas9 specificity (Makarova, Wolf et al, 2020, Altae-Tran, Kannan et al, 2021). To assess this question, we investigate the performance of GS-Cas9 targeting to two endogenous genome loci in HEK293T cells (FIG. 2A). Efficiencies of GS-Cas9 at these two sites were comparable with SpCas9 (mean modified levels 84% vs 55.7% at EMX1 site, 88.8% vs 70% at P205 site for SpCas9 and GS-Cas9, respectively) (FIG. 2B). Alleles showed substitutions, insertions and deletions modified patterns (FIG. 2C and 2D). We did not observe different indel patterns between SpCas9 and GS-Cas9, which showed a prominent 1-bp insertion pattern (FIG. 2E and 2F). We next tested the top 3 known genomic off- target (OT) sites for EMX1 site as identified by genome-wide, unbiased identification of DSBs enabled by sequencing (GUIDE-seq) (Kleinstiver, Pattanayak et al, 2016). At all three off-target-sites, we observed lower editing levels mediated by GS-Cas9 (averaging 0.94%, 0.83% and 1.24% at the top 3 off-target sites, respectively) than that of SpCas9 (averaging 7.95%, 5.37% and 7.12%, respectively). Additionally, the ratio of on-target to off-target editing of GS-Cas9 was higher averaging -37.5- fold at across these 3 off-target sites than that of SpCas9 (FIG. 2G). To further evaluate the tolerance of GS-Cas9 for mismatched target sites, we chose four mutated guide sequences for VEGFA site introduced double-base mismatches at different PAM-distal positions reported in previous research (Slaymaker, Gao et al, 2016). Compared with SpCas9, GS-Cas9 induced undetectable modified reads with all mismatched guides, SpCas9 induced 0.36%-4.21% of modified reads (FIG. 2H). These data demonstrated that REC lobe expansion improved targeting specificity and retained comparable activity with SpCas9. Engineered high fidelity SpCas9 variants always suffer loss of cleavage activity relative to wild type (Kim, Kim et aL 2020, Kim, Kim et al. 2023), GS-Cas9 shows almost similar activity with SpCas9-HFl and HypaSpCas9.
[0347] Example 5: REC expansion can regulate the enzyme fused to N-terminus of SpCas9.
[0348] REC domains play various roles in regulating conformational rearrangement of Cas9s (Nishimasu, Ran et al, 2014, Nishimasu, Cong et aL 2015, Hirano, Gootenberg et al. 2016, Jiang, Taylor et al. 2016, Chen, Dagdas et al. 2017, Sun, Yang et al. 2019, Zhang, Zhang et al. 2020). We found that REC-domain-sizes often increase with increasing other domain-sizes of Cas9s (FIG. 1C and FIG. 8). For instance, the size of SpCas9-RuvC domain is 1.45-fold of that of SaCas9 (307 aa vs 212aa). We next chose ABE8e editor to investigate whether REC expansion could regulate the N-terminal catalytic domain. ABE8e contains a mutant adenine deaminase TadA8e fused to the N-terminus of nSpCas9 and exposed outside the spatial region of the Cas9 protein (Lapinaite, Knott Gavin et al, 2020, Richter, Zhao et al, 2020) (FIG. 3). Adenine deaminase-mediated off-target editing could lead various immune responses (i.e., oncogenic) (Song, Shiromoto et al, 2022). The activity of TadA is positively correlated with Cas9-independent off-target DNA editing and RNA editing activities of ABE (Richter, Zhao et al. 2020, Chen, Zhang et al. 2023, Neugebauer, Hsu et al. 2023). The exposure of TadA8e increases its freedom, thereby enhancing its chances of interacting with other DNA or RNA nucleotides, resulting in a higher occurrence of Cas9-independent off-target editing events (Liu, Zhou et al. 2020, Villiger, Schmidheini et al. 2021, Zeng, Yuan et al. 2023). We hypothesized that REC expansion could induce extra interaction with TadA8e to reduce its swinging (FIG. 4A). We next established a time saving, efficient, and cost-effective method, cell plasmid RNA editing and DNA editing reporter assay (cpREADER), to evaluate ABE variants editing activity (FIG. 4B, FIGS. 20-22, Example 8). Two variants ABE8e / left-BH-REC12 (ABE8e / LBR12) and ABE8e / right-REC12 (ABE8e / RR12) were constructed. We observed that only ABE8e / LBR12 induced comparable RNA editing and DNA editing activities with ABE8e (FIGS. 23A and 23B). But MFI reduction of DNA editing mediated by ABE8e / LBR12 is less 1.7-fold than that of ABE8e. ABE8e / left-BH-REC12 showed efficient editing activities, which indicates that SpCas9 / left-BH-REC12 above retains binding ability for the on-target. We therefore propose that stiff-BH with extra length in SpCas9 / left-BH-REC12 interrupts functional conformation of RuvC domain (FIG. 18) (Jiang, Taylor et al. 2016, Chen, Dagdas et al. 2017). To investigate whether the spatial position of TadA8e domain can be re-located, we tried to truncate the stiff-double-BH length in ABE8e / LBR12 to change the distance between TadA8e and GS-nCas9(D10A) (FIG. 4A). Compared to 2XBH, truncated BH induced not only decrease in RNA editing but also increase in DNA editing activity. Of note, EGFP activation of single-BH version mediated by RNA editing reduced 4.46-fold and 3.64-fold from double-BH version and ABE8e, respectively. We observed similar trends in transfection with or without loading sgRNA containing 20nt-guide. We did not observe differences in MFI mediated by RNA editing for these variants (FIG. 4C, 4D and 4E). However, MFI reduction induced by DNA editing from single-BH version showed almost a 2-fold increase relative to the double-BH version and ABE8e (FIG. 4F and 4G). Like GS-Cas9, ABE8e variant with single-BH configuration exhibited the best performance. This aligns with the natural evolutionary selection of Cas9s, in which near-fixed length BH enables adaptation to wide range of Cas9 sizes. Additionally, our results indicate the BH could be used as a distance adjuster for domain position in Cas9-derived architecture’s. We next investigated if length of guide sequence effect on ABE8e variants with various REC expansions. Double-BH version showed similar EGFP activation levels with ABE8e in presence of various gRNAs (FIG. 24). Compared to loading 20 nt-gRNA, ABE variants with truncated BH showed a 1.5-fold increase in percentage of GFP+ cells with loading 35nt guide sequence relative to ABE8e. All the three ABE architectures induced similar GFP+ levels in the presence of 40nt guide sequence (FIG. 4H). For MFI, truncated versions showed 1.66-fold and 1.49-fold of AB E8e with 35nt and 40nt, respectively (FIG. 25).
[0349] Then, we evaluated whether ABE variants could induce efficient edits at genome locus in HEK293T cells. We observed similar A-to-G editing rates at EMX1 site for these variants targeted with 20nt-gRNA (cumulative editing rates of 36.98%, 38.44% and 32.79% for ABE8e, Single-BH and 1.5-BH). Importantly, we observed decreases in indel levels for both REC expansion variants (averaging indels 7.09%, 3.63%, and 3.58% for ABE8e, Single-BH and 1.5- BH) (FIG. 41 and FIG. 26). We did not observe edits outside of protospacer, which suggest that enlarging REC lobe cannot enlarge R-loop length in presence of longer guide sequences. REC expansion variants induced modest lower A-to-G efficiencies in presence of 35nt-gRNA (cumulative editing rates of 41.10%, 28.78% and 23.50% for ABE8e, Single-BH and 1.5-BH, respectively), while indel levels was much lower 4.25-fold than that of ABE8e (2.72%, 0.95%, and 0.64% for ABE8e, Single-BH and 1.5-BH, respectively) (FIG. 4J, FIG. 27). Enlarged REC domain may provide protection for the ssDNA strand with reducing the ssDNA exposed to solvent to decrease double strand breaks. These results indicate that base editor variants with the expansion of the recognition lobe become more sensitive to gRNAs with extended lengths. We speculate that the excessive length of the gRNA might hinder its processing into 20nt within the cell, thereby reducing the efficiency of DNA editing (Ran, Hsu et aL 2013). We named single-BH version giant SpCas9 based ABE8e (GS-ABE8e). We further assessed whether the GS-ABE8e enables reducing off-target RNA editing activity on endogenous transcripts within the cell. After transfection of HEK293T cells by nSpCas9, ABE8e and GS-ABE8e, RNA was extracted from cells, after complementary DNA (cDNA)synthesis, IP90 transcript previously used to measure off-target RNA editing due to its abundance or sequence similarity to the native TadA tRNA substrate (Rees, Wilson et al, 201 , Neugebauer, Hsu et al, 2023) was amplified by RT-PCR and analyzed for A-to-I editing by high-throughput sequence. We found that GS- ABE8e drastically reduced A-to-I levels at all A positions inside the amplicon. Especially, A- to-I efficiencies at TAG motif reduced more than 2-fold compared to ABE8e (FIG. 4K). TadA9 (with N108Q and L145T variants in Tad8e) (Chen, Zhang et al. 2023) was compatible with GS-nCas9. GS-ABE9 retained similar MFI disruption activities with ABE9 (FIG. 4L). This suggests that previous strategies for engineering TadA (Rees and Liu 2018, Grilnewald, Zhou et al, 2019, Griinewald, Zhou et al. 2019, Rees, Wilson et al. 2019, Zhou, Sun et al. 2019, Doman, Raguram et al, 2020, Chen, Zhang et al, 2023, Huang, Lin et al, 2023) should be available for extensively improving accuracy and precision of GS-ABE8e. Despite being the most active and widely used editor in the field of precise gene singlebase editing (Richter, Zhao et al. 2020, Arbab, Matuszek et al. 2023, Huang, Heins et al. 2023), ABE8e still faces the dual challenges of editing accuracy and precision in its applications (Wang and Doudna 2023). Results above indicated that GS-Cas9 variant exhibited higher specificity than SpCas9, therefore we performed all subsequent experiments with GS-ABE8e to systemically comparing ABE8e to determine the characteristics of GS-ABE8e.
[0350] Example 6: Characterization of GS-ABE8e
[0351] We investigated whether GS-ABE8e was compatible with gRNA-scaffolds from related Cas9 systems using EGFP disruption assay (FIG. 28A and 28B). GS-ABE8e still exhibited the highest GFP disruption activity in the presence of SpCas9 gRNA-scaffold, which was similar to ABE8e. We observed similar activities for ABE8e and GS-ABE8e in presence of gRNAs with Sa- or Stl-gRNA-scaffold. Interestingly, GS-ABE8e performed higher activity than ABE8e in presence of gRNA-FnCas9-scaffold (FIG. 28C). Additionally, GS-ABE8e did not show activity with a gRNA carrying the SaCas9-gRNA-scaffold and the guide sequence targeting to a EGFP-DNA region bearing the promiscuous 5’-CGGAGT-PAM for SaCas9 and SpCas9 (FIG. 28D and 28E). We next evaluated the tolerance of GS-ABE8e for truncated guide sequence. Compared with ABE8e, GS-ABE8e exhibited lower relative activity with 12nt, lOnt and 8nt guides (FIG. 5A). This result indicates that GS-ABE8e is more sensitive to truncated guide sequences. To further determine whether GS- ABE8e is sensitive to mismatched target sites, we systematically mutated the EGFP guide sequence to introduce double-base mismatches at various positions (FIG. 5B). Compared with ABE8e, GS-ABE8e induced lower relative activity with mismatches located inside of the 1- to 10-base pair seed sequence.
[0352] Base editors could generate unexpected ultra-low edits outside of protospacer sequence (that is, out-of-protospacer) (Lei, Meng et al. 2021). EGFP reporter enable much lower 10-fold detection limit for base editing than NGS-based strategies (Ranzau, Rallapalli et al. 2023). To test whether REC expansion base editor reduce edits out-of-protospacer on non-target strand, we chose EGFP reporter to compared activities of GS-ABE8e and ABE8e on adenine base located at 4 position out-of -protospacer (that is, A (-4)). Notably, we observed that GS-ABE8e reduced 13.3-fold A (-4) edits thanABE8e (FIG. 5C and FIG. 29). In an R-loop assay (Doman. Raguram et aL 2020), GS-ABE8e mediated EGFP activation lower 6.5-fold than ABE8e (FIG. 5D), which contributes higher resolution of outcomes on low DNA-substrate concentration (FIG. 30). Anti-Cas9 protein AcrIIA4 can inhibit ABE activity by occupying both PAM- interacting and non-target DNA strand cleavage catalytic pockets of SpCas9 (Rauch, Silvis et al. 2017, Yang and Patel 2017, Liang, Sui et al. 2020, Zhang, Bamidele et al, 2022). In our EGFP disruption assay, we observed no MFI reduction induced by GS-ABE8e in presence of AcrIIA4 protein (MFI reduction 24% for ABE8e vs 0% for GS-ABE8e, respectively) (FIG. 31). These data reveal that REC lobe expansion enhances ABE8e performance in reducing bystander products. Importantly, GS-ABE8e retains similar DNA editing activity with ABE8e.
[0353] To extensively investigate the performance of GS-ABE8e, 11 endogenous human genome loci previous reported (Grunewald, Zhou et al, 2020, Richter, Zhao et al, 2020, Chen, Zhang et al. 2023, Gao, Guan et al, 2024) were tested. We compared ABE8e with GS-ABE8e side by side using 11 additional gRNA sites. The activity (defined as the editing level at the position with the highest A-to-G rate in each site) of GS-ABE8e ranged from 40.4% to 81.0%, which was similar with ABE8e (45.6% to 74.7%) at 10 of 11 sites (FIG. 6A-6I and FIG. 32). Compared with ABE8e, indel levels induced by GS-ABE8e at 8 of 11 target sites decreased 1.7-fold to 3.36-fold (0.46% to 2.49% for GS-ABE8e vs 1.07% to 8.63% for ABE8e, respectively) (FIG. 6J). Across all eleven sites, GS-ABE8e induces a 2-fold decrease of indels than ABE8e (FIG. 6K). We examined the base editing window of GS-ABE8e and ABE8e. Consistent with ABE8e, GS-ABE8e can efficiently edit A2-A8 positions, counting the PAM as positions 21-23 (FIG. 6L). REC expansion enhances targeting specificity of SpCas9, therefore we analyzed off-target activity of ABE8e in HEK293T cells at previously reported Cas9- dependent off-target sites (Grunewald, Zhou et al. 2020, Richter, Zhao et al. 2020). We assessed the top 2 or 5 known ABE off-target sites for EMX1 site and site P206 (VISTA enhancer). We observed a decrease at one of top two off-targets sites for site P206 when comparing GS-ABE8e to ABE8e (0.95% vs 2.70%). At top 5 EMX1 off-targets, we also observed a decrease of GS- ABE8e-mediated A-to-G editing (averaging 44.63% vs 31.1%, 1.28% vs 0.60%, 4.65% vs 2.58, 2.19% vs 1.37%, 0.27% vs 0.01% at the top five sites, respectively) (FIG. 6M and 60). GS- ABE8e also reduced the rate of indels at all off-targets sites (averaging 2.17% vs 1.04%, 0.1 % vs 0.01%, 0.11% vs 0.06%, 0.21% vs 0.12%, 0.01% vs 0%, at the top 5 sites, respectively) (FIG. 6N and 6P). The ratio of on- target to off-target editing was much higher than ABE8e at these sites (up to 38440-fold) (FIG. 6Q and 6R). Moreover, in an orthogonal R-loop assay for Cas-independent off-target editing, GS-ABE8e generated much lower rates of off-target effects (averaging 0.29% and up to 0.56%), but ABE8e generated much higher Cas9-independent off- target editing (averaging 3.1% and up to 6.8%) (FIG. 7). These data demonstrate that GS- ABE8e is a highly efficient ABE reducing unexpected Cas9-dependent indels and off-target edits. It has been challenging to achieve a balance between the target activity and bystander activity of ABE in previous studies. Though Cas9-independent RNA off-target editing and DNA off-target editing of Cas9-derived ABEs have been reduced through direction evolution of TadA, which failed in reduction of unexpected indels events (Richter, Zhao et al. 2020, Chen, Zhang et al. 2023). On the other side, ABE8e constructs also were compatible with previously described high-fidelity SpCas9 variants bearing different residue-mutations and succeeded in minimizing Cas-dependent off-target editing, which showed substantial loss in on-target editing (Talas, Simon et al. 2021, Alves, Ha et al. 2023, Sretenovic, Green et al, 2023). Our data demonstrate advantages of GS-nCas9 scaffold in improving overall performance of adenine base editor. We reasoned that GS-Cas9 has the potential to replace SpCas9 to balance high on-target editing with low bystander editing for therapeutic applications of base editors containing catalytic domain tethered to the N-terminus.
[0354] Example 7: GS-Cas9 scaffold is compatible with various TadA variants.
[0355] ABE9 carrying a TadA8e(N108Q / L145T) mutant could generate narrower editing window than ABE8e (Chen, Zhang et al. 2023). After evaluation at 9 endogenous sites, we observed that both ABE9 and GS-ABE9 showed lower activity than GS-ABE8e, but GS-ABE9 significantly reduced unexpected indel levels at tested sites and narrowed the editing window to 3 nucleotides (A4 to A6) (FIGS. 33A-33C). Additionally, GS-nCas9(D10A) also showed good capability with TadA8e(N46L) variant previous reported (Chen, Zhu et al, 2023) (FIG. 34). This data also strongly suggests that previous strategies reported to engineer TadA8e to improve editing precision could be used to improve GS-ABE8e performance in future research. The evaluation of base editing also showed no difference in PAM preference between GS- ABE8e and ABE8e (Table 2). We suggested that GS-Cas9 could be engineered to broaden PAM scopes using previous approaches (Chatterjee, Jakimo et al. 2020, Walton, Christie et al. 2020).
[0356] In our study, we systematically analyzed these 161,859 candidate sequences predicted as IscB and Cas9, and we found clues about recognition (REC) lobe expansion in Cas9. We proposed that the REC lobe, besides its recognizing function, may also have served as a mediator for multiple domains during the evolutionary process of Cas9s. Reaching this knowledge, we succeeded in expanding the REC domain size of SpCas9, the most widely used in academic and industry, with the creation of giant SpCas9 (GS-Cas9). Our findings uncovered the secret hidden in the natural evolution of Cas9, which contributes the novel operability to SpCas9. Most notably, we validated that REC expansion and BH domain play key roles in regulating or repositioning catalytic domain tethered to N-terminus of GS-nCas9 (D10A). Additionally, we also demonstrated that GS-Cas9 was compatible with various TadA derivatives, which means that enlarging REC domain could be expanded to almost all base editors based on SpCas9 scaffold. As a nuclease for dsDNA cleavage, SpCas9 may already be a mature enzyme, but SpCas9 has not yet evolved to maturity as a scaffold of further base editors to gain additional catalytic domains. It seems that we have now identified the first cause of its developmental deficiency. Our new insights on Cas9-expansion might inspire further foundational research to deepen our understanding of Cas9 topological and stimulate more innovative solutions to overcome the critical challenges associated with Cas9-based gene editors.
[0357] Example 8: Supplement Information
[0358] Analysis of Cas9 candidate sequences
[0359] Totally 161,859 sequences predicted by Han et al (Altae-Tran, Kannan et al. 2021) were used as database for analysis in this research. We can manually filter sequences with specific strings for further analysis. Specifically, we chose conserved residues to confirm alignment results. RuvC-I conserved D, RuvC-II conserved E, HNH Conserved H, and RuvC-III conserved D and H.
[0360] To group candidate sequences, proteins predicted were filtered with Francisella, Streptococcus pyogenes, Streptococcus thermophilus, Streptococcus aureus, Neisseria meningitidis, and Campylobacter jejuni, respectively. 105 sequences annotated from Francisella, 1926 sequences annotated from Streptococcus pyogenes, 214 sequences annotated from Streptococcus thermophilus, 7 sequences annotated from Streptococcus aureus, 1905 sequences annotated from Neisseria meningitidis, and 41539 sequences annotated from Campylobacter jejuni were obtained, respectively. Of note, only 1 of 161,859 sequences was annotated from Francisella novicida, so we selected genus Francisella to filter the database to generate F_Cas9 group. Crystal structure of FnCas9 (protein data base ID 5b2o) was identified(Hirano, Gootenberg et al. 2016). FnCas9 contains 1629 amino acids (aa) and is the largest Cas9 protein experimentally characterized so far.
[0361] The database was filtered with length>1629aa parameter to generate in total of 93 sequences. 48 of 93 sequences carry larger than 1700aa protein size and have not been experimentally characterized so far, which account for less than 0.03% of the total database (FIG. 8). Further, 47 sequences with larger than 1629aa were obtained after removal of repeated sequences, 23 of the 47-sequences were larger than 1700aa. After analysis of blast with NCBI protein database, WWFRO 1000017.1 and OLPQ01001336.1 sequences were clearly identified not Cas9 proteins and removed. Finally, 45 sequences with >1629aa and 21 sequences with >1700aa were obtained and used for extensive analysis (Table 1). We aligned the above 45 sequences (Table 1) with reference FnCas9 using Clustal Omega 1.2.2 algorithm and annotated domain-sizes in each sequence (FIG. 9). For each group of Cas9s from Francisella, Streptococcus pyogenes, Streptococcus thermophilus, Streptococcus aureus and Neisseria meningitidis, we manually aligned and did blast to identify and remove repeated sequences to generate final 8 sequences for Francisella, 17 sequences for Streptococcus pyogenes, 17 sequences for Streptococcus thermophilus, 4 sequences for Streptococcus aureus and 12 sequences for Neisseria meningitidis used for alignment analysis (FIG. 10 to 14 Of note, CABEX0010000001.1 and CABFEJ010000006.1 were removed in Streptococcus pyogenes group.
[0362] Additionally, sequences carrying larger size than the size of the reference identified by crystal structure and experiments made sense in this research.
[0363] Construction of cell plasmid RNA editing and DNA editing reporter assay (cpREADER)
[0364] We constructed a convenient method to detect ABE variants activities on RNA and DNA substrate. It was named of cell plasmid RNA editing and DNA editing reporter assay (cpREADER).
[0365] Adenosine deaminase could restore translation of UAG stop codon by acting on mRNA (Kaseniit, Katz et al. 2022). We created a fluorescence reporter EGFPstop by introducing premature TAG stop codon into the EGFP gene via G-to-A conversion on the sense strand of DNA (FIG. 20). The UAG stop codon prevents the translation of the EGFPstop. In presence of ABE variants, deaminase could edit the adenosine (A) in the UAG to an inosine (I), which enable the translation of the downstream EGFP CDS to generate green fluorescence signal (FIG. 19). We compared two EGFPstop variants, W58stop and Q81stop previous reported (Masaharu and Shun’ichi 2022, Zeng, Yuan et al, 2023) (FIG. 20). Results showed that W58stop could generate a much higher mean of fluorescence intensity (MFI) than Q81stop after A-to-I conversion in UAG codon (FIG. 20). EGFP activation showed a correlation with ABE plasmid dosage and ABE activity (FIGS. 21A-21D).
[0366] In DNA editing, ABE variant guided by a single guide RNA (sgRNA) can edit the adenine on the sense strand of EGFP gene and convert the A-to-I. In the presence of MMR in mammalian cells, the repair of the bottom strand results in the generation of the mutated EGFP gene nonsense strand, then mutated mRNA could be transcribed from the nonsense strand, and fluorescence is disrupted in the cells (Ito, Shiraishi et al. 2018, Averill and Jung 2023) (FIG.
[0367] 21E).
[0368] Table 1 List of Cas9 sequences with size larger than 1629aa.
[0369] Table 2. Summary of tested sites used for evaluation of base editing.
[0370] Table 3. Genomic sites.
[0371]
[0372]
[0373]
[0374] Table 4A Summary of sgRNA Plasmids Table 4B GFP Reporter Plasmids
[0375] Table 4C Gene Editor Plasmids
[0376] Table 5 Primers for amplicon PCR.
[0377]
[0378]
[0379] Table 6. Supplementary Sequences.
[0380]
[0381]
[0382] Table 7. Supplementary Sequences for guide RNA.
[0383] Example 9: Materials and methods for Examples 10-16
[0384] Analysis of sequences
[0385] Alignment analysis and annotation analysis were performed with Geneious Prime 11.0.18. Filtration of sequences was performed using Microsoft Excel. NCBI blastp (proteinprotein BLAST) was used to double check candidate-sequences.
[0386] If no other instructions are given, in the Examples 10-16, FnCas9 specifically refers to PDB:5B2O, SpCas9 specifically refers to PDB:4OO8, StlCas9 specifically refers to PDB:6M0V, SaCas9 specifically refers to PDB:5AXW, NmeCas9 specifically refers to PDB:6JDQ, CjCas9 specifically refers to PDB: 6JOO, OgeuIscB specifically refers to PDB:7UTN.
[0387] Construction of Plasmids
[0388] We used general cloning methods including Gibson assembly with GeneArt Gibson Assembly HiFi Master Mix (Thermo Fisher, A46628) and Quick ligation kit (New England Biolabs, M2200S) with a variety of type IIS restriction enzymes. One Shot™ MAX Efficiency™ DH5a-TlR Competent Cells (Thermo Fisher, 12297016) were used for DNA cloning. The sequences of cloned constructs were confirmed by sanger sequencing following extraction with QIAprep kit (QIAGEN, 27106). Plasmids including vectors purchased from Addgene used in this study can be found in Table 4A-4C. All human genome sites are available in Table 3. Oligonucleotides used in Examples 10-16 can be found in Table 5. Gene editors and gRNAs used in Examples 10-16 can be found in Table 6 and Table 7. EGFP expression plasmids containing amino acid substitutions were generated by standard PCR with Q5 Site- Directed Mutagenesis Kit (New England Biolabs, E0554S). Human-codon-optimized fragments used for REC expansions, oligonucleotides and sgRNA expression plasmids were synthesized by IDT (Integrated DNA Technologies, USA). Unless otherwise indicated, all sgRNAs were designed to target sites containing a 5’ guanine nucleotide.
[0389] Cell Culture and Transfection
[0390] HEK293T cells and HEK293-deGFP cells35(a gift from Prof. Eben Alsberg, University of Illinois at Chicago) were cultured in DMEM (Sigma, D6429) supplemented with 10% FBS and 1% Gibco™ Penicillin-Streptomycin (10,000 U / mL) at 37 °C and 5% CO2. Cells were seeded one day prior to transfection in 24-well plates. For screening Cas9 variants, cells were plated at a density of approximately 70,000 cells per well, and cells were transfected with 200ng of nuclease editor expression plasmid DNA, 200 ng of sgRNA plasmid, and 10 ng of reporter plasmid (pCMV-GFP plasmid). For the pbREADER assay, cells were plated at a density of approximately 70,000 cells per well, and cells were transfected with 40 fmol of base editor expression plasmid DNA, 200 ng of sgRNA plasmid, and 200 ng or 10 ng of reporter plasmid (200 ng of EGFPstop plasmid, 10 ng of pCMV-GFP plasmid). Unless otherwise noted, cells were plated at a density of approximately 50,000 cells per well for in vitro cell genome editing, cells were transfected with 70 fmol of nuclease or 80 fmol base editor expression plasmid DNA and 200 ng of sgRNA plasmid per well. Transfections were performed with Lipofectamine 2000 (Invitrogen, 11668027) and Lipofectamine 3000 (Thermo Fisher, L3000008) according to the manufacturer’s recommended protocol in the cpREADER assay and genome editing, respectively. Of note, the ratio of lipofectamine to DNA was set at 2.5:1. NEBioCalculator version 1.15.5 was used to convert dsDNA mass to moles of dsDNA.
[0391] The EGFPstop mRNA activation assay was performed according to the following protocol. mRNA carried a stop codon UAG in place of UGG at the 58-position, and mRNA was prepared as described in our previous study64. HEK293T cells were plated at a density of approximately 70,000 cells per well in 24-well plates. Cells were transfected with 200 ng of ABE8e plasmid or ABE9 plasmid using Lipofectamine 3000. After 6 hours, cells were transfected with 400 ng of mRNA using Lipofectamine 2000. Then 48 hours, cells were analyzed using a BD Accuri™ C6 Flow Cytometer.
[0392] Western blotting assay
[0393] HEK293T cells were lysed 3 d after transfection using RIPA buffer complemented with proteinase and phosphatase inhibitors (Pierce, Protease Phosphatase Inhibitor Tablets, Thermo Fisher Scientific: A32959). The total protein concentrations of cell lysate supernatants were quantified using a BCA protein assay kit (Thermo Fisher Scientific). In total, 20 pg of total protein per well was loaded to electrophoresis using a 10-well 4-12% Bis-Tris Gel (Invitrogen, NW04125BOX) and transferred using an XCell II™ Blot Module and PVDF (0.45-pm pore size)). Bolt™ Transfer Buffer (Invitrogen, BT0006) containing 5% methanol was selected as transferring buffer. 25 min and 70 min at 30V were used for transferring small protein (<50kDa) and large protein (> 150kDa), respectively. The gel was stained using Coomassie Blue to confirm transferring efficiency. The membranes were blocked with 5% Non-Fat Dry Milk (AmericanBio, AB 10109) for 1 h at room temperature and then divided and processed with different primary antibodies including anti-beta-actin (1 : 10,000; Abeam, ab49900) and the anti-Flag (1 :5000 dilution; Abeam, ab205606) separately overnight at 4 °C. Then, the membranes were incubated with Goat Anti-Rabbit IgG H&L (HRP) (1 : 10,000 dilution; Abeam, ab205718) for 1 h and visualized using the Bio-Rad imaging system (ChemiDoc™ MP Imaging System). Blot was incubated in chemiluminescent substrate (Thermo Scientific, A38554) for 5min prior to imaging.
[0394] Genomic DNA isolation
[0395] Cells were harvested 3 days post transfection. Genomic DNA was extracted from transfected cells using DNeasy Blood & Tissue Kit (Qiagen, 69504) following the manufacturer’s protocol. Extracted DNA was normalized to a final concentration of 20 ng per pl with ddH2O.
[0396] Flow Cytometry Detection
[0397] Adherent cells were treated with 100 pl of 0.25% Trypsin-EDTA (Gibco, 25200056) incubated for 5 min at 37 °C to completely detach cells. 400 pl of DMEM was used to stop trypsin digestion. Samples were applied to the BD Accuri™ C6 Flow Cytometer (BD Biosciences) directly, and GFP fluorescence was measured. Data analysis was performed with BD Accuri C6 Software and Microsoft Excel. Prism (GraphPad) was used to generate column graphs and for calculations. EGFP disruption or EGFPstop activation experiments in pbREADER assays were performed as previously described64. Briefly, transfected cells were analyzed 48h after transfection for loss or restoration of EGFP fluorescence. The background was determined by gating a negative control transfection.
[0398] RNA isolation from mammalian cells
[0399] HEK293T cells were transfected with nSpCas9-EGFP, ABE8e-EGFP or GS-ABE8e- EGFP plasmid. 48 hours post transfection, approximately 300,000 cells were harvested for mRNA isolation. RNA isolation was performed with the RNeasy Plus Mini Kit (QIAGEN, 74134) according to the manufacturer’s instructions. In short, RNA isolation began with removal of the culture medium and washing the cells with lx PBS (Thermo Fisher Scientific). 350 pl of RLT Plus Buffer was added into each well; cells were homogenized by pipetting and transferred into a DNA eliminator column, and the subsequent binding and washing steps for RNA isolation using the RNeasy columns were performed as recommended by the manufacturer. Upon elution of RNA from the RNeasy column with 45 pl of RNase (ribonuclease) free water (QIAGEN), 2 pl of RNase inhibitor (New England Biolabs, M0314S) was added to prevent RNA degradation, and RNA was stored at -80°C. cDNA Synthesis cDNA synthesis was performed using ProtoScript II First Strand cDNA Synthesis Kit (New England Biolabs, E6560) with 1 pg of RNA in a total reaction volume of 20 pl. Reactions were incubated for 2 hr at 42 °C. 2 pl cDNA was input into 50 pl NEBNext High-Fidelity 2X PCR Master Mix (New England Biolabs, M0541S) containing specific primers. The purification of target products was performed with GeneJET PCR Purification Kit (Thermo Fisher, K0702).
[0400] Preparation of genomic DNA amplicons for deep sequencing
[0401] Targeted regions flanking the on-target or off-targeted sites were amplified with specific primers and Q5 Hot Start High-Fidelity 2X Master Mix (NEB, M0494S) under the following thermal cycling conditions: one cycle, 98°C, 1 min; one cycle, 98°C, 30 s; 35 cycles, 98°C, 10 s, 65°C, 30 s, 72°C, 15s; one cycle, 72°C, 2 min; 4°C hold. 100 ng of isolated genomic DNA was input into each 50 pl of PCR. PCR products were analyzed on an agarose gel electrophoresis system to verify both size and purity. PCR products were purified with GeneJET PCR Purification Kit (Thermo Fisher, K0702) followed by normalizing to a final concentration of 20 ng per pl with ddH2O.
[0402] Analysis of HTS data for targeted amplicon sequencing
[0403] Targeted amplicon sequencing was carried out by Genwiz (Azenta, South Plainfield, NJ, US) using the Amplicon-EZ protocol. Batch analysis with the CRISPResso2 pipeline was used for targeted amplicon sequencing. For DNA analysis, a 30-bp window was used to quantify indels around the DNA nick site. Otherwise, the default parameters were used for analysis. The output file “NUCLEOTIDE_PERCENTAGE_TABLE.txt” was imported into Microsoft Excel for quantification of editing frequencies and “Indel_histogram.txt” for quantification of indel frequencies. Indel percentage was re-checked with BEanalyzer if requested. For analysis of RNA amplicon editing, no sgRNAflag was used. Instead, the output file “NUCLEOTIDE_PERCENTAGE_TABLE.txt” was imported into Microsoft Excel for analysis of A-to-G editing rates associated with each sample (inosine in RNA is read as a guanosine by polymerases).
[0404] Prism (GraphPad) was used to generate dot plots and bar plots of these data. For instances in the text where means have been calculated across multiple genomic or transcriptomic loci, the SDs reported represent the SD of the mean for all biological replicates.
[0405] Amplicon sequencing using Sanger sequencing
[0406] Sanger sequencing results of PCR amplicons were analyzed using EditR version 1.0. 10 (https : / / moriaritylab . shiny apps .io / editr_v 10 / ) .
[0407] Orthogonal R-loop assay
[0408] Orthogonal R-loop assays were performed to measure Cas9-independent off-target editing as described previously, with minor modifications. Under standard conditions, 200 ng of base editor plasmid, 300 ng of dSaCas9 plasmid, 100 ng of SpCas9 sgRNA plasmid and 100 ng of SaCas9 sgRNA plasmid were co-transfected into HEK293T cells using 1.5 pl of Lipofectamine 3000. Cells were cultured for 4 days after treatment followed by genomic DNA isolation.
[0409] Statistical Analysis and Graphical Illustrations.
[0410] Curve plotting and statistical analysis were performed using Prism 8 (GraphPad, La Jolla, CA). Data are shown as means ± standard error of the mean for groups of two or more replicates or as individual values with the mean indicated. Graphical illustrations were created using BioRender (https: / / biorender.com / ) and Office Power point.
[0411] Data availability
[0412] All unmodified reads for sequencing-based data are available from the NCBI Sequence Read Archive, under accession number PRJNA1121945 and PRJNA1121941. Key plasmids from this work are available from Addgene.
[0413] Supplementary Notes
[0414] Analysis of Cas9 candidate sequences
[0415] In total, 161,859 sequences predicted by Zhang and colleagues1were used as the database for analysis in this study. We filtered sequences with specific strings for further analysis. Specifically, we chose conserved residues to confirm alignment results, which included RuvC-I conserved D, RuvC-II conserved E, HNH Conserved H, and RuvC-III conserved D and H.
[0416] To group candidate sequences, predicted proteins were filtered with Francisella, Streptococcus pyogenes, Streptococcus thermophilus, Staphylococcus aureus, Neisseria meningitidis, and Campylobacter jejuni, respectively. 105 sequences isolated from Francisella, 1926 sequences isolated from S. pyogenes, 214 sequences isolated from S. thermophilus, 7 sequences isolated from S. aureus, 1905 sequences isolated from N. meningitidis, and 41539 sequences isolated from C. jejuni were obtained. Of note, only 1 of 161,859 sequences was isolated from F. novicida, so we selected genus Francisella to filter the database to generate the F_Cas9 group.
[0417] The database was filtered with length >1629aa parameter to generate in total of 93 sequences. 48 of 93 sequences carry protein larger than 1700aa protein size and have not been experimentally characterized to date, which accounts for less than 0.03% of the total database (FIG. 8). Further, 47 sequences with larger than 1629aa were obtained after removal of repeated sequences, 23 of the 47-sequences were larger than 1700aa. After analysis of blast with NCBI protein database, WWFR01000017.1 and OLPQ01001336.1 sequences were clearly identified as non-Cas9 proteins and removed. Finally, 45 sequences with > 1629aa and 21 sequences with >1700aa were obtained and used for extensive analysis (Table 1). We aligned the above 45 sequences with reference FnCas9 using the Clustal Omega 1.2.2 algorithm and annotated REC domain in each sequence (FIG. 9). For each group of Cas9 from Francisella, S. pyogenes, S. thermophilus, S. aureus and N. meningitidis, we aligned and did blast to identify and remove repeated sequences to generate final 8 sequences for Francisella, 17 sequences for 5. pyogenes, 17 sequences for .S’. thermophilus, 4 sequences for .S'. and 12 sequences for N. meningitidis used for alignment analysis (FIGS. 10-14). Of note, CABEXOO 10000001.1 and CABFEJ010000006.1 were removed in .S', pyogenes group.
[0418] On the other hand, 66,765 predicted sequences isolated from ‘metagenome’ were not used in this study because these couldn’t be really mapped to species.
[0419] Construction of plasmid-based RNA editing and DNA editing reporter assay (pbREADER)
[0420] We constructed a convenient method to detect activities of ABE variants on RNA and DNA substrate. It was named for plasmid-based RNA editing and DNA editing reporter assay (pbREADER).
[0421] Adenosine deaminase could restore translation of UAG stop codon by acting on mRNA2. We created a fluorescence reporter EGFPstop by introducing premature TAG stop codon into the EGFP gene via G-to-A conversion on the sense strand of DNA. The UAG stop codon prevents the translation of the EGFPstop mRNA (FIGS. 20 and 21) . In the presence of ABE variants, deaminase could edit the adenosine (A) in the UAG to an inosine (I), which enables the translation of the downstream EGFP CDS to generate green fluorescence signal. We compared two EGFPstop variants, W58stop and Q81stop previous reported3,4. Results showed that W 58stop could generate a much higher mean of fluorescence intensity (MFI) than Q81 stop after A-to-I conversion in UAG codon. EGFP activation showed a correlation with ABE plasmid dosage and ABE activity (FIGS. 43A-43B; and FIGS. 22C, 22D, 22F, and 22G).
[0422] In DNA editing, ABE variants guided by a single guide RNA (sgRNA) could edit the adenine on the sense strand of EGFP gene and convert the A-to-I. In the presence of MMR in mammalian cells, the repair of the bottom strand results in the generation of the mutated EGFP gene nonsense strand, then mutated mRNA could be transcribed from the nonsense strand, and fluorescence is disrupted in the cells5-6(FIGS. 22E, 22H, and 221). Activities of ABE8e and ABE9 also were evaluated by our pbREADER, which agreed with the previous research7.
[0423] Example 10: REC domains of Cas9 show high flexibility in size As a subset of RNA-guided endonucleases, the diversity of CRISPR-Cas9 orthologs provides rich material for studying its evolution2. However, CRISPR-Cas9 has not yet evolved to be an ideal scaffold for gaining additional catalytic domains3.The evolution hypothesis suggests that Cas9 originated from a predicted ancestor and underwent a complex evolution from small to large sizes in millions of years, involving the expansion of various domains through random insertions4(FIG. 35A). Among these domains, acquisition of the recognition (REC) domain is predicted to play a critical role in enhancing the specificity of Cas91’4, 5. Although the topological malleability of Cas9 was realized by random insertion of non-Cas9 domain using Mu transposon , the biological influence of natural domain expansion on Cas9 remains unclear. Despite many efforts to minimize Cas9 off-target cleavage through introducing amino acid substitutions over the past decade , there is currently no direct experimental evidence supporting the hypothesis that domain-expansions contribute to improving Cas9 performance. Meanwhile, compact editors have been paid excessive attention for overcoming the size limitations associated with viral-based delivery in gene editing131 14(FIG. 8A). Encouragingly, lipid nanoparticle (LNP) delivery has emerged as a safer and more efficient solution for the delivery of large cargos which drives us to explore various routes for engineering gene editors. Though domain expansion can increase size of payloads, exploring domain expansion can not only aid in refining the knowledge of Cas9 evolution and deep understanding of Cas9 structure, but also open a new pathway to artificially increase diversity of Cas9 and develop innovative solutions for improving accuracy and precision of Cas9-based gene editors.
[0424] Here, we tried learning knowledge from the natural evolution of Cas9 to improve the performance of gene editors. By combining bioinformatics analysis and protein engineering methods, we were inspired by nature evolution and created a natural-evolution-like nuclease, giant SpCas9 (GS-Cas9), carrying the largest REC domain experimentally identified to date. An unreported role of domain expansion was uncovered in this study, enlarging the REC domain regulated the catalytic activities of the domain tethered to the N-terminus of the Cas9 scaffold. ABE8e adenine base editors showed high compatibility with the enlarged Cas9- scaffold, which resulted in improving the precision of base editing. As an alternative strategy, GS-Cas9 scaffold has huge potential to be expanded to most SpCas9-based base editors to finetune their performance. The present disclosure demonstrated developing improved Cas9-based gene editors harnessing nature-evolution-like concept that is different from the strategies reported before. Zhang and colleagues predicted a total of 161,859 IsrB_IscB_Cas9 sequences from a prokaryotic database constructed by combining various databases1, among which there are 138,334 Cas9 sequences (FIG. 8B). However, only a minuscule fraction of these sequences carrying less than 1700aa has been experimentally characterized or structurally resolved in the past decade19, 20, 21, 22, 23, 24, 25, 26. To investigate which domain exhibits the representative expansion in natural evolution, we first compared characteristics of seven proteins (one IscB and six Cas9s) structurally identified to date. Compared to an ancestor OgeuIscB, FnCas9 (size 1629aa) shows a 3.28-fold increase in overall size and a 20.4-fold increase in REC domain size. Meanwhile, other domains increase less than 3.08-fold (FIG. 35B and FIG. 8C). To further evaluate the distribution of REC domain sizes in large Cas9s, we filtered sequences larger than FnCas9 to get a total of 45 unique sequences (Example 9 and Table 1) after the removal of 48 repeat sequences. These proteins were an average size of 1712aa, where CP041030.1 (size 1695aa) isolated from the Francisella sp. LAI 12445 strain carried a predicted 841aa REC indicating 22.1-fold larger than that of IscB. This REC surpasses that of CjCas9, SpCas9 and FnCas9 by 491aa, 217aa and 65aa, respectively (FIG. 9). The analysis indicates that REC shows higher flexibility in size and tolerance for continuous insertions than other domains. On the other hand, the protospacer adjacent motif (PAM) for large Cas9 likely becomes shorter than that of compact Cas9 (FIG. 35B and FIG. 8D). Additionally, Cas9 proteins isolated from the same genus or species show a wide range of protein size (FIG. 35C). We proposed that the REC lobe could also serve as a good mediator for matching and balancing protein-size expansion during the evolutionary process of Cas9s, thereby maintaining the functional conformation of multi-domain9'27’28.
[0425] To further explore possible insertion sites for REC expansion, we then systemically analyzed 45,696 Cas9 sequences isolated from five kinds of organisms (genus Francisella, species Streptococcus pyogenes, Streptococcus thermophilus, Staphylococcus aureus, Neisseria meningitidis and Campylobacter jejuni, in which at least one Cas9 has been structurally identified) by removing repeats. Except for 5. pyogenes Cas9 sequences, we found inspirations of REC lobe expansion. Likewise, the predicted protein UHCSO 1000002. 1 isolated from 5. aureus is 318aa longer than SaCas9 and contains a predicted REC that is 268aa longer than that of SaCas9. Insertions mainly appear in either the middle (for .S', thermophilus, S. aureus, N. meningitidis and C. jejuni) or boundaries (for Francisella, S. thermophilus and S. aureus) of REC domain (FIG. 35D, and FIGS. 10-15). For large Cas9 in Francisella, the expansion seems easier to occur at the C-terminus of REC. These positions might be used as sites for REC insertion. Most notably, UHCSO 1000002.1 shows an 88.3% pairwise sequence identity with SpCas9 (FIG. 16), while LR822033.1 and VBTKO 1000005.1 isolated from .S'. thermophilus have 99.4% and 56.9% sequences identity with SpCas9, respectively (FIG. 17). Since FnCas9 and SpCas9 were younger than StCas9, SaCas9, and NmeCas929, we suggested that SpCas9 might evolved from SaCas9 or StlCas9 through domain expansion. Though we didn’t capture expansion sequences derived from SpCas9, our analysis indicates that SpCas9 may have high compatibility with REC domain of SaCas9 or StlCas9.
[0426] Example 11: Building giant SpCas9 by enlarging REC domain
[0427] Based on the bioinformatic analysis above, we assumed that REC domain should have the ability to tolerate large-size insertions. SpCas9 has been extensively studied in terms of its crystal structure and mechanism of function, which enables it to be the most widely used in both academic and industrial fields13. Based on these inspirations above, we tried to explore whether we can create giant SpCas9 variants by enlarging the REC domain and gain biological insights30. Multiple- sequence alignments of .S’. thermophilus and S. aureus did not show conserved motifs near the boundaries of REC domains (FIGS. 12, 13, 18), which indicated that these sites might serve as recombination hot-spots compatible with foreign insertions.
[0428] Because domain combinations are often found in only one sequential order in the evolution of the majority of multi-domain proteins31132, we first tried inserting the REC lobe of StlCas9 at the termini of the REC domain in SpCas9 to generate expansion variants driven by these understandings. Given flexible linker regions in Cas9 appearing to play a role in the inactive- to-active conformational transition of multiple domains19, 20133, we selected two shortflexible-linkers from SaCas9 (residues T205 to D223 or K420 to T436) to fink insertions with receptor without using any non-Cas9 sequences (FIGS. 35E and 35F). We designed and constructed four variants in which different REC elements optimized for human using codon usages from IDT (FIG. 35G). To avoid Cas9 variants failing to function in mammalian cells, we chose human embryonic kidney (HEK) 293T cells to evaluate the activity of our variants with an EGFP plasmid interference assay modified from previous research34(FIG. 35H). Unfortunately, we did not observe EGFP disruption activity from any of the four variants (FIG. 351). We speculated that enlarged Cas9-variants might also not need extended BH domains. We also observed that the excessive length of the BH domain reduced the cleavage activity of SpCas9-2BH (FIG. 41A). Therefore, we removed the Stl-BH sequence in the variant Left- BH-REC12 to generate the Left-REC12 variant. The new variant still did not induce reduction of GFP+cell populations, but did lead to a 20% decrease in mean fluorescence intensity (MFI) of GFP+cells (FIG. 35J). Excitingly, we observed a significant decrease in GFP+cells in presence of the Left-REC12 variant when the substrate plasmid concentration was reduced to a lower level (FIG. 35K). Then, we tried to use a human cell line bearing an integrated construct that constitutively expresses a deGFP protein, HEK293-deGFP35, to evaluate Left- REC12 activity on chromosome sites. We observed 45.5% to 70% GFP disruption activity relative to SpCas9 in the presence of various doses of the Left-REC12 variant (FIG. 35L). We speculated that domain expansion might impact Cas9 kinetics then caused the great differences of EGFP disruption in presence of different mole concentrations of the DNA substrate36, 37. Then, we specially named Left-REC12 as giant SpCas9 (GS-Cas9). The GS-Cas9 protein has a length of 1780 aa and a 207kDa molecular weight, which is 1.3 times larger than SpCas9 and might form an unbalanced bilobed structure (FIG. 36A, FIG. 41B). The REC domain of GS- Cas9 consists of 1036 amino acids making it 1.66-fold larger than that of SpCas9. This represents not only the largest Cas9 protein size, but also the largest REC domain size experimentally characterized to date. Though others have built synthetic Cas9 scaffold using engineered Mu transposon system6’38, all these efforts introduced Mu-recognized sequences or non-Cas9 domains into variants. By contrast, GS-Cas9 is a nature-like Cas9 without being introduced any non-Cas9 domains. We also demonstrated an insertion site unidentified by transposon6’38, which suggested different compatibilities of SpCas9 with non- or nature-Cas9 domain. Importantly, SpCas9 and GS-Cas9 can allow for comparative study to deepen the understanding of influences induced by domain expansion on Cas9 properties.
[0429] We next evaluated the expression of GS-Cas9 using western blotting. The lower disruption efficiency of GS-Cas9 may be attributed in part to lower expression relative to that of SpCas9. However, SpCas9-2BH, GS-Cas9 showed similar expression levels (FIG. 41C). Deletions or mutations of REC domain in various Cas9 nucleases generally lead to reduced protein expression in human cells191 26, which suggests that topological changes may play a critical role in Cas9 expression or stability.
[0430] Example 12: Enlarging REC lobe enhanced editing precision on human genome
[0431] To test whether the enlarged RNA-guided nuclease could enhance editing precision, we investigated the editing performance of GS-Cas9 on seven endogenous genome loci in HEK293T cells (FIG. 36B). For wild type SpCas9, modified levels ranged from 35.8%-88.8% at the seven genomic loci tested, whereas for GS-Cas9 modified levels ranged from 2.6%-84% (FIG. 36C). Totally, GS-Cas9 have 55% on-target activity of SpCas9 across seven sites tested (FIG. 36D). Alleles showed patterns of substitutions, insertions, and deletions (FIG. 36E, FIG. 42); however, we did not observe different indel patterns between SpCas9 and GS-Cas9 (FIG. 36F). Next, we tested the top 3 known genomic off-target (OT) sites for EMX1 editing as identified by GUIDE-seq (genome-wide, unbiased identification of double-strand-breaks enabled by sequencing)7. At all three off-target-sites, we observed lower editing levels mediated by GS-Cas9 (averaging 0.94%, 0.83% and 1.24% at these 3 off-target sites, respectively) than that of SpCas9 (averaging 7.95%, 5.37% and 7.12%, respectively). Additionally, the ratio of on-target to off-target editing increased from averaging 12.7 for SpCas9 to averaging 50.2 for GS-Cas9 across these 3 off-target sites (FIG. 36G). To further evaluate the tolerance of GS-Cas9 for mismatched target sites, we chose four mutated guide sequences for the VEG FA site introducing double-base mismatches at different PAM-distal positions reported in previous research8. Compared with SpCas9, GS-Cas9 induced undetectable modified reads with all mismatched guides, while SpCas9 induced 0.36%-4.21% of modified reads (FIG, 36H) Two sites contain either a cytosine-rich homopolymeric sequence or a sequence with multiple TG repeats, where GS-Cas9 mediated significantly lower editing compared to SpCas9 (14-fold lower at the HBG2 site and 4-fold lower at the VEGFA site, respectively).
[0432] This data demonstrated that GS-Cas9 improved targeting specificity and retained detectable activity. FnCas9 possesses higher specificity than SpCas939, which may be attributed in part to its larger REC lobe. Excessive expression of gene editors could lead higher off-target edits40, domain expansion may improve editing accuracy by partially altering protein expression of Cas9. Engineered high-fidelity SpCas9 variants always suffer loss of editing activity compared to wild type12141142For instance, HypaSpCas9 demonstrates 60% activity relative to SpCas9, while GS-Cas9 shows similar activity levels to HypaSpCas9 in HEK293T cells.
[0433] Example 13: Enlarging REC domain reduced RNA editing of adenine deaminase TadA8e in base editor.
[0434] In contrast to SpCas9 and SaCas9, the RuvC domain interacts with the REC domain in FnCas9, a naturally occurring large Cas921(FIG. 37A). We also observed that REC-domain- sizes often increase with the increase of other domain-sizes of Cas9s (FIG. 37C and FIG. 8). For instance, the SpCas9-RuvC domain (307 residues) is 1.45-fold larger than that of SaCas9 (212 residues). These indications reveal that REC expansion might induce extra interaction between domains than the parent. Then we chose the ABE8e editor as a fusion model of Cas9 for gain-of-function to investigate whether REC expansion could impact the N-terminal catalytic domain activities. ABE8e fusion contains a adenine deaminase mutant TadA8e (166aa) with higher activity fused to the N-terminus of the nickase SpCas9 (D10A), in which TadA8e is exposed to the solvent with resulting in high mobility and no specific interaction with the SpCas9 scaffold43, 44(FIG. 37A). The exposure increases the freedom of TadA8e, thereby enhancing its chances of interacting with other DNA or RNA nucleotides, resulting in a higher occurrence of Cas9-independent off-target editing events4-‘’-46-47. Cas9-independent off-target DNA editing and RNA editing activities of ABE are positively correlated with the activity of the tethered adenine deaminase44, 48,49. We proposed a hypothetical model in which REC expansion might lead to a rearrangement of the original REC domain and shorten the distance between REC and N-terminal TadA by enlarging the coverage of the engineered REC lobe, thereby generating influences on TadA (FIG. 37B).
[0435] We first established a time saving and cost-effective method, plasmid-based RNA editing and DNA editing reporter assay (pbREADER), to evaluate editing activities of ABE8e variants (FIG. 38A, FIGS. 20, 21, 43A-43B, and 22C-22I; Example 9). Four variants ABE8e- 2BH, GS-ABE8e, GS-ABE8e-1.5BH and GS-ABE8e-2BH were constructed (FIG. 38B). We observed that GS-ABE8e-2BH induced comparable RNA editing and DNA editing activities with ABE8e, but MFI reduction of DNA editing mediated by GS-ABE8e-2BH was 2.3-fold less compared to ABE8e (FIGS. 44A and 44B). Meanwhile, we did not observe both RNA editing and DNA editing activities from the ABE8e / RR12 variant generated by inserting the REC domain at the C-terminus of Cas9 scaffold in ABE8e. Given the SpCas9 / Right-REC12 data, we suggest that the insertion at the C-terminus of the REC domain may lead to protein misfolding. GS-ABE8e-2BH showed efficient editing activities, which indicated that GS- Cas9-2BH retains on-target binding. As a long a-helix, BH also acts as rigid spacer of distance between REC and RuvC-I of Cas9so. Compared to GS-ABE8e-2BH, truncated BH induced not only a decrease in RNA editing but also increased DNA editing activity (FIGS. 38C-38F). Of note, EGFP activation of GS-ABE8e mediated by RNA editing reduced 4.46-fold and 3.64- fold compared to the GS-ABE8e-2BH and ABE8e, respectively. We observed similar trends in transfection with or without loading non-target sgRNA containing 20nt-guides. We did not observe differences in MFI mediated by RNA editing for these variants. However, MFI reduction induced by DNA editing from GS-ABE8e showed an almost 2-fold increase relative to GS-ABE8e-2BH (FIG. 38G). We also observed similar differences between ABE8e and ABE8e-2BH (FIG. 44C). Like GS-Cas9, ABE8e-expansion variants with the single-BH configuration exhibited the best performance. This aligns with the natural evolutionary selection of Cas9s, in which near-fixed length BH enables adaptation to wide range of Cas9 sizes. Unlike GS-Cas9, western blotting indicated that GS-ABE8e, GS-ABE8e-1.5BH and GS- ABE8e-2BH showed similar expression levels with ABE8e (FIG. 44D). We further assessed whether GS-ABE8e reduced off-target RNA editing activity on endogenous transcripts within the cell. After transfection of HEK293T cells by nickase SpCas9 (D10A), ABE8e, and GS- ABE8e, RNA was extracted from cells. After complementary DNA (cDNA)synthesis, the IP90 transcript previously used to measure off-target RNA editing due to its abundance or sequence similarity to the native TadA tRNA substrate49, 51was amplified by RT-PCR and analyzed for A-to-I editing by high-throughput sequencing. We found that GS-ABE8e drastically reduced A-to-I levels at all A positions inside the amplicon (FIG. 38H). Especially, A-to-I efficiencies at TAG motif reduced more than 2-fold compared to ABE8e. These results reveal that enlarging the REC domain succeeds in regulating adenine deaminase activities in ABE8e.
[0436] We then evaluated whether GS-ABE8e variants still could retain on-target editing activity on a genome locus in HEK293T cells. We observed similar A-to-G editing rates at the EMX1 site for these variants loaded with 20nt-gRNA (cumulative editing rates of 36.98%, 38.44% and 32.79% for ABE8e, GS-ABE8e, and GS-ABE8e-1.5BH, respectively) (FIG. 381 and FIG. 45). Importantly, we observed decreases in indel levels for both REC expansion variants (averaging indels of 7.09%, 3.63%, and 3.58% for ABE8e, GS-ABE8e, and GS- ABE8e-1.5BH, respectively). Enlarged REC domains may provide protection for the nontarget strand by reducing the ssDNA exposed to solvent, which may decrease double strand breaks and indels. REC expansion variants induced slightly lower A-to-G efficiencies in the presence of a 35nt-gRNA (cumulative editing rates of 41.10%, 28.78% and 23.50% for ABE8e, GS-ABE8e, and GS-ABE8e-1.5BH, respectively), while indel levels were around 4.25-fold lower compared to ABE8e (2.72%, 0.95%, and 0.64% for ABE8e, GS-ABE8e, and GS- ABE8e-1.5BH, respectively) (FIG. 38J and FIG. 46). We did not observe edits outside of the 20-nt protospacer, which suggests that enlarging the REC lobe does not enlarge the editing window in presence of longer guide sequences.
[0437] Here we showed that enlarging the REC domain enabled Cas9 scaffold to reduce the activities of N-end fused catalytic domain in ABE8e on RNA transcripts, which indicated an unreported role of REC domain expansion. Given the non-specific DNA mutations mediated by free HNH or RuvC nucleases52, 53, we suggested that REC acquisition or expansion might play a ‘peacemaker’ role during natural evolution to influence catalytic domains to work within a specific space and reduce non-specific cleavage (FIG. 37B). GS-Cas9 variant exhibited advantages in editing precision over SpCas9, therefore we performed all subsequent experiments with GS-ABE8e to systemically compare it with ABE8e and determine the characteristics of GS-ABE8e. Example 14: Enlarging REC domain contributed to reducing ABE8e unexpected editing in plasmid-based reporter.
[0438] The orthogonal R-loop assay is always used for evaluating Cas9-independent off-target editing induced by base editor54. Because the GS-Cas9 scaffold was a chimera of multiple Cas9s, we then investigated whether GS-ABE8e was compatible with gRNA- scaffolds from related Cas9 systems using the EGFP disruption assay (FIGS. 28A and 28B). GS-ABE8e still exhibited the highest GFP disruption activity in the presence of a SpCas9 gRNA-scaffold, which was like ABE8e. We observed similar activities for ABE8e and GS-ABE8e in the presence of gRNAs with Sa- or Stl-gRNA-scaffolds. Interestingly, GS-ABE8e showed greater activity than ABE8e in the presence of a gRNA-FnCas9-scaffold (FIG. 28C). Additionally, GS-ABE8e did not show activity with a gRNA carrying the SaCas9-gRNA- scaffold and a guide sequence targeting an EGFP-DNA region bearing the promiscuous 5’-CGGAGT-PAM for SaCas9 and SpCas9 (FIGS. 28D and 28E). These results indicate that GS-Cas9 scaffold retains its original orthogonality. sgRNA may suffer degradation from nuclease in cells55, we next evaluated the tolerance of GS-ABE8e for truncated guide sequences. Compared with ABE8e, GS-ABE8e exhibited lower relative activity with 12nt, lOnt and 8nt guides (FIG. 39A), which indicates that GS- ABE8e is more sensitive to truncated guide sequences. To further determine whether GS- ABE8e is still sensitive to mismatched target sites, we systematically mutated the EGFP guide sequence to introduce double-base mismatches at various positions (FIG. 39B). Compared with ABE8e, GS-ABE8e showed lower relative activity with mismatches located inside of the 1- to 10-base pair seed sequence. This revealed that enlarging REC domain also enhanced targeting specificity of ABE8e base editor.
[0439] Base editors could generate unexpected ultra-low edits outside of the protospacer sequence (that is, out-of-proto spacer) , which is lower than the detection limit of next generation sequencing (NGS)56. The EGFP reporter enables a much lower 10-fold detection limit for base editing than NGS-based strategies57. To evaluate whether GS-ABE8e has the potential to reduce edits out-of-protospacer on non-target strands, we compared the EGFP reporter activities of GS-ABE8e and ABE8e on an adenine base located at position 4 out-of- protospacer (that is, A (-4)). Notably, we observed that A (-4) out-of-proto spacer edits for GS- ABE8e reduced 13.3-fold compared to ABE8e (FIG. 39C and FIG. 29). In an orthogonal R- loop assay using reporter plasmid, GS-ABE8e mediated EGFP activation was 6.5-fold lower than ABE8e (FIG. 39D), which contributes to higher resolution outcomes on low EGFP-DNA- substrate concentrations (FIG. 30). Moreover, in an orthogonal R-loop assay for Cas9- independent off-target DNA editing on genome (FIG. 39E), GS-ABE8e expectedly generated much lower rates of off-target effects (averaging 0.29% and up to 0.56%), but ABE8e generated much higher Cas9-independent off-target editing (averaging 3.1% and up to 6.8%) (FIG. 39F). Reasonably, this improvement might benefit from extra interaction between the enlarged REC domain and TadA domain. Meanwhile, GS-ABE8e retains similar editing activity with ABE8e on locus site 1 (FIG. 39G).
[0440] Anti-Cas9 protein AcrIIA4 can inhibit ABE activity by occupying both PAM- interacting and non-target DNA strand cleavage catalytic pockets of SpCas958. In our EGFP disruption assay, we observed no MFI reduction induced by GS-ABE8e in the presence of AcrIIA4 (MFI reduction 24% for ABE8e vs 0% for GS-ABE8e, respectively) (FIG. 31). These data reveal that GS-ABE8e shows more sensitive to anti-Cas9 protein.
[0441] Example 15: GS-ABE8e showed lower indels and Cas9-dependent off-target editing than ABE8e on human genome.
[0442] It has been challenging to achieve a balance between the target activity and off-target activity of ABE in previous studies. Though Cas9-independent RNA off-target editing and DNA off-target editing of ABE8e has been reduced through directed evolution of TadA8e, TadA8e mutants were difficult to reduce unexpected indel events441 48. On the other side, ABE8e constructs were compatible with previously described high-fidelity SpCas9 variants (e.g., SpCas9-HFl, evoSpCas9, SuperFi-Cas9) bearing different residue-mutations and succeeded in minimizing Cas-dependent off-target editing, which unfortunately showed substantial loss in on-target editing59-6°.61-62. T0investigate the performance of GS-ABE8e more thoroughly, 11 endogenous human genome loci previously reported44148163, 64were tested. We compared ABE8e with GS-ABE8e side by side using 11 additional gRNA sites. The activity (defined as the editing level at the position with the highest A-to-G rate in each site) of GS-ABE8e ranged from 40.4% to 81.0%, which was similar with ABE8e (45.6% to 74.7%) at 10 of 11 sites (FIG. 40A and 40B). Compared with ABE8e, indel levels induced by GS-ABE8e at 11 target sites decreased 1.7-fold to 3.36-fold (0.46% to 2.49% for GS-ABE8e vs 1.07% to 8.63% for ABE8e, respectively) (FIG. 40C). Across all 12 tested sites (EMX1 site in FIG. 4i and 11 sites in FIG.6c), GS-ABE8e induced a 2-fold decrease of indels compared to ABE8e (FIG.40D). We examined the base editing window of GS-ABE8e and ABE8e. Consistent with ABE8e, GS-ABE8e can efficiently edit A2-A8 positions, counting the PAM as positions 21- 23 (FIG. 40E). REC expansion enhanced targeting specificity of SpCas9, therefore we analyzed off-target activity of GS-ABE8e in HEK293T cells at previously reported Cas9- dependent off- target sites44,63. We assessed the top 5 or 2 known ABE off-target loci for EMX1 and ABE site 2 (VISTA enhancer) editing. We observed a decrease in editing at one of the top two off-target sites for ABE site2 when comparing GS-ABE8e to ABE8e (0.95% vs 2.70%) (FIG. 40F). At the top 5 EMX1 off-target sites, we observed remarkable decreases in GS- ABE8e-mediated A-to-G editing (averaging 44.63% vs 31.1%, 1.28% vs 0.60%, 4.65% vs 2.58%, 2.19% vs 1.37%, 0.1% vs 0.00% at the top five sites for ABE8e and GS-ABE8e, respectively) (FIG. 40G). The ratio of on-target to off-target editing was much higher than ABE8e at these sites (FIGS. 40H and 401). GS-ABE8e also produced lower rates of indels at all off-targets sites (averaging 2.17% vs 1.04%, 0.13% vs 0.01%, 0.11% vs 0.06%, 0.21% vs 0.12%, 0.007% vs 0.005%, at the top 5 sites ABE8e and GS-ABE8e, respectively) (FIGS. 40J and 40K). GS-ABE8e also mediated lower A-to-G edits at 19 out of 20 positions on non-target strands across 12 sites (only positions with editing >0.1% were counted for this conclusion) (FIG. 47). These observations reveal that GS-ABE8e is a highly efficient adenine base editor with reducing unexpected indels and off-target edits.
[0443] TadA8e is reported with the much higher deoxyadenosine deaminase activity and leads higher off-target editing (e.g., Cas9-dependent, Cas9-independent, and RNA transcriptome) than that of variants with lower activity44. GS-Cas9 allows base editors to reduce unexpected edits even when carrying TadA8e. Importantly, our study demonstrated that the same goal can be achieved by adjusting the topological malleability of the Cas9 scaffold, unlike previous research which aimed to reduce unwanted edits through continuous mutagenesis of TadA8e51. GS-Cas9 has the potential to balance high on-target editing with low off-target editing and off- targets for therapeutic applications of base editors containing catalytic domain tethered to the N-terminus in future research. Here, our data demonstrated advantages of the GS-Cas9 scaffold in improving overall performance of adenine base editor.
[0444] Example 16: GS-Cas9 scaffold is compatible with various Ta A8e variants
[0445] Different substitutions introduced in deaminase can critically affect its compatibilities with Cas homologs44. We then extensively investigated whether various TadA8e-derived variants could be compatible with the GS-Cas9 scaffold. We first tested the possible compatibility of GS-Cas9 with other TadA9 variant (with N 108Q and L145T variants in Tad8e)48, GS-ABE9 retained slightly lower MFI disruption activities than ABE9 (FIG. 48). After evaluation at 9 endogenous sites, we observed that both ABE9 and GS-ABE9 showed moderately lower activity than GS-ABE8e, but GS-ABE9 significantly reduced unexpected indel levels at tested sites and narrowed the editing window from 7 nucleotides (A2 to A7) to 3 nucleotides (A4 to A6) (FIG. 49). This data reveals that specific residue can impact the compatibility of TadA with GS-Cas9 scaffold. Additionally, GS-Cas9 also showed good capabilities with TadA8e (N46L) and TadA8e (V106W) variants previously reportedS1, 65, which showed higher editing efficiencies than GS-ABE9 editor (FIG. 50). Excitedly, we observed that GS-ABE8e(V106W) reduced 3-fold EGFP activation mediated by RNA editing than ABE8e(V106W) (FIG. 51). These data further indicate that previous strategies reported for engineering TadA8e to improve editing precision can be used to improve GS-ABE8e performance in future research. The evaluation of base editing also showed no difference in PAM preference between GS-ABE8e and ABE8e (Table 2).
[0446] We tested REC lobe expansion in Cas9 after systemically analyzing thousands natural sequences and generated artificially enlarged SpCas9. It is demonstrated that engineering of Cas9 domain expansion can improve gene editing performance. Structural growth trajectory of II-C Cas9s was predicted in a recent study66, which indicates a vast potential for engineering the topological malleability of RNA-guided endonucleases. The ratio of REC lobe size to NUC lobe size may play crucial roles in regulating function of Cas9 or Cas9-based gene editors (FIG. 52). This approach is versatile and complementary to previous engineering strategies and holds great potential for advancing safer gene editing for clinical drug development and optimization.
[0447] Incorporation by Reference
[0448] All publications, patents, and patent applications mentioned herein are hereby incorporated by reference in their entirety as if each individual publication, patent or patent application was specifically and individually indicated to be incorporated by reference. In case of conflict, the present application, including any definitions herein, will control.
[0449] Also incorporated by reference in their entirety are any polynucleotide and polypeptide sequences which reference an accession number correlating to an entry in a public database, such as those maintained by The Institute for Genomic Research (TIGR) on the world wide web at tigr.org and / or the National Center for Biotechnology Information (NCBI) on the World Wide Web at ncbi.nlm.nih.gov.
[0450] Equivalents Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments of the invention described herein. Such equivalents are intended to be encompassed by the following claims.
Claims
What is claimed is:
1. A Cas9 variant comprising a Streptococcus pyogenes Cas9 (SpCas9) that comprises an insertion of a Streptococcus thermophilus Cas9 (StlCas9) recognition (REC) domain.
2. The Cas9 variant of claim 2, wherein the StlCas9 REC domain is inserted at the N- terminus of the SpCas9 REC domain.
3. The Cas9 variant of claim 1 or 2, wherein the StlCas9 REC domain is inserted between the SpCas9 Bridge helix (BH) domain and the SpCas9 REC domain.
4. The Cas9 variant of any one of claims 1-3, wherein the StlCas9 REC domain is inserted between residues 93D and 94D of the SpCas9.
5. The Cas9 variant of any one of claims 1-4, wherein the SpCas9 is a wild-type SpCas9.
6. The Cas9 variant of any one of claims 1-5, wherein the SpCas9 comprises an amino acid sequence of SEQ ID NO: 1.
7. The Cas9 variant of any one of claims 1-4, wherein the SpCas9 is a SpCas9 nickase (nSpCas9).
8. The Cas9 variant of claim 7, wherein the SpCas9 nickase is nSpCas9 (D10A).
9. The Cas9 variant of any one of claims 1-8, wherein the StlCas9 REC domain comprises an RECI domain and an REC2 domain.
10. The Cas9 variant of any one of claims 1-9, wherein the StlCas9 REC domain comprises residues 74-466 of StlCas9.
11. The Cas9 variant of any one of claims 1-10, wherein the StlCas9 REC domain comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%,85%, 90%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the sequence of SEQ ID NO: 19.
12. The Cas9 variant of any one of claims 1-11, wherein the StlCas9 REC domain comprises the amino acid of SEQ ID NO: 1 .
13. The Cas9 variant of any one of claims 1-12, wherein the Cas9 variant does not comprise a BH domain from StlCas9.
14. The Cas9 variant of any one of claims 1-13, wherein the StlCas9 REC domain is connected to the N-terminus of the SpCas9 REC domain via a linker.
15. The Cas9 variant of claim 14, wherein the linker comprises a linker from Staphylococcus aureus (SaCas9).
16. The Cas9 variant of claim 14 or 15, wherein the linker from SaCas9 comprises residues T205 to D223 of SaCas9.
17. The Cas9 variant of claim 15 or 16, wherein the linker from SaCas9 comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the sequence of SEQ ID NO: 20.
18. The Cas9 variant of any one of claims 15-17, wherein the linker from SaCas9 comprises the amino acid sequence of SEQ ID NO: 20.
19. The Cas9 variant of any one of claims 1-18, wherein the SpCas9 comprises an insertion of the amino acid sequence of SEQ ID NO: 8.
20. The Cas9 variant of any one of claims 1-19, comprising the amino acid sequence of SEQ ID NO: 21.
21. A nucleic acid encoding the Cas9 variant of any one of claims 1-20.
22. A nucleic acid of claim 21, wherein the nucleic acid is a DNA or RNA.
23. A vector comprising a nucleic acid of claim 21 or 22.
24. A composition or kit, comprising:(a) a Cas9 variant of any one of claims 1-20, a nucleic acid of claim 21 or 22, or a vector of claim 23; and(b) a Cas9 guide RNA (gRNA) .
25. The composition or kit of claim 24, wherein the Cas9 gRNA is a spCas9 gRNA.
26. A composition or kit, comprising:(a) a Cas9 variant of any one of claims 1-20, a nucleic acid of claim 21 or 22, or a vector of claim 23;(b) a Cas9 crRNA; and(c) a Cas9 tracrRNA.
27. The composition or kit of claim 26, wherein the Cas9 crRNA is a spCas9 crRNA, and the Cas9 tracrRNA is a spCas9 tracrRNA.
28. A method of cleaving or modifying a target DNA, comprising: contacting the target DNA or a cell comprising the target DNA with:(a) a Cas9 variant of any one of claims 1-20, a nucleic acid of claim 21 or 22, or a vector of claim 23; and(b) a Cas9 guide RNA (gRNA) ; wherein the Cas9 variant and the Cas9 gRNA form a complex that cleaves or modifies the target DNA.
29. The method of claim 28, wherein the Cas9 gRNA is a spCas9 gRNA.
30. A method of cleaving or modifying a target DNA, comprising: contacting the target DNA or a cell comprising the target DNA with:(a) a Cas9 variant of any one of claims 1-20, a nucleic acid of claim 21 or 20, or a vector of claim X;(b) a Cas9 crRNA; and(c) a Cas9 tracrRNA; wherein the Cas9 variant, the Cas9 crRNA, and the Cas9 tracrRNA form a complex that cleaves or modifies the target DNA.
31. The method of claim 30, wherein the Cas9 crRNA is a spCas9 crRNA, and the Cas9 tracrRNA is a spCas9 tracrRNA.
32. A method of cleaving or modifying a target DNA, comprising contacting the target DNA or a cell comprising the target DNA with a composition or kit of any one of claims 24- 27.
33. A Cas9-guided base editor comprising a base editor linked to a Cas9 variant, wherein the Cas9 variant comprises a SpCas9 nickase (nSpCas9) comprising an insertion of a StlCas9 REC domain.
34. The Cas9-guided base editor of claim 33, wherein the StlCas9 REC domain is inserted at the N-terminus of the nSpCas9 REC domain.
35. The Cas9-guided base editor of claim 33 or 34, wherein the StlCas9 REC domain is inserted between the BH domain and the REC domain of the nSpCas9.
36. The Cas9-guided base editor of any one of claims 33-35, wherein the StlCas9 REC domain is inserted between residues 93D and 94D of the nSpCas9.
37. The Cas9-guided base editor of any one of claims 33-36, wherein the nSpCas9 is nSpCas9 (D10A).
38. The Cas9-guided base editor of any one of claims 33-37, wherein nSpCas9 comprises an amino acid sequence of SEQ ID NO: 22.
39. The Cas9-guided base editor of any one of claims 33-38, wherein the StlCas9 REC domain comprises an RECI domain and an REC2 domain.
40. The Cas9-guided base editor of any one of claims 33-39, wherein the StlCas9 REC domain comprises residues 74-466 of StlCas9.
41. The Cas9-guided base editor of any one of claims 33-40, wherein the StlCas9 REC domain comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the sequence of SEQ ID NO: 19.
42. The Cas9-guided base editor of any one of claims 33-41, wherein the StlCas9 REC domain comprises the amino acid of SEQ ID NO: 19.
43. The Cas9-guided base editor of any one of claims 33-42, wherein the StlCas9 REC domain is connected to the N-terminus of the SpCas9 REC domain via a linker.
44. The Cas9-guided base editor of claim 43, wherein the linker comprises a linker from Staphylococcus aureus (SaCas9).
45. The Cas9-guided base editor of claim 43 or 44, wherein the linker from SaCas9 comprises residues T205 to D223 of SaCas9.
46. The Cas9-guided base editor of claim 44 or 45, wherein the linker from SaCas9 comprises an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the sequence of SEQ ID NO: 20.
47. The Cas9-guided base editor of any one of claims 44-46, wherein the linker from SaCas9 comprises the amino acid sequence of SEQ ID NO: 20.
48. The Cas9-guided base editor of any one of claims 33-47, wherein the Cas9 variant comprising half (0.5x), single (lx), or two (2x) StlCas9 BH domains immediately upstream of the StlCas9 REC domain.
49. The Cas9-guided base editor of claim 48, wherein the Cas9 variant comprising a single BH domain immediately upstream of the StlCas9 REC domain.
50. The Cas9-guided base editor of any one of claims 33-49, wherein the nSpCas9 comprises an insertion of an amino acid sequence selected from SEQ ID NOs: 3, 7 and 8.
51. The Cas9-guided base editor of claim 50, wherein the nSpCas9 comprises an insertion of the amino acid sequence of SEQ ID NOs: 8.
52. The Cas9-guided base editor of any one of claims 33-51, wherein the cas9 variant comprises an amino acid sequence selected from SEQ ID NOs: 23-25.
53. The Cas9-guided base editor of claim 52, wherein the cas9 variant comprises the amino acid sequence selected from SEQ ID NOs: 24.
54. The Cas9-guided base editor of any one of claims 33-53, wherein the base editor is linked to N-terminus of the Cas9 variant.
55. The Cas9-guided base editor of claim 54, wherein the base editor is an adenine deaminase.
56. The Cas9-guided base editor of claim 55, wherein the adenine deaminase is selected from TadA8e, TadA9, TadA8e(N46L), and TadA8e(V106W).
57. A nucleic acid encoding the Cas9-guided base editor of any one of claims 33-56.
58. The nucleic acid of claim 57, wherein the nucleic acid is a DNA or RNA.
59. A vector comprising a nucleic acid of claim 57 or 58.
60. A composition or kit, comprising:(a) a Cas9-guided base editor of any one of claims 33-56, a nucleic acid of claim 57 or58, or a vector of claim 59; and(b) a Cas9 guide RNA (gRN A) .
61. The composition or kit of claim 60, wherein the Cas9 gRNA comprises a gRNA- scaffold selected from SpCas9 gRNA-scaffold, SaCas9 gRNA-scaffold, StlCas9 gRNA- scaffold, and FnCas9 gRNA-scaffold.
62. The composition or kit of claim 60 or 61, wherein the Cas9 gRNA comprises a guide sequence of 8 nucleotides (nt), lOnt, 12nt, 20nt, 35nt, or 40 nt in length.
63. A composition or kit, comprising:(a) a Cas9-guided base editor of any one of claims 33-56, a nucleic acid of claim 57 or 58, or a vector of claim 59;(b) a Cas9 crRNA; and(c) a Cas9 tracrRNA.
64. The composition or kit of claim 63, wherein the Cas9 crRNA is selected from SpCas9 crRNA, SaCas9 crRNA, StlCas9 crRNA, and FnCas9 crRNA, and wherein the Cas9 tracrRNA is selected from the same species as the Cas9 crRNA.
65. The composition or kit of claim 63 or 64, wherein the Cas9 crRNA comprises a guide sequence of 8 nt, lOnt, 12nt, 20nt, 35nt, or 40 nt in length.
66. A method of modifying a target DNA or RNA comprising: contacting the target DNA or RNA or a cell comprising the target DNA or RNA with:(a) a Cas9-guided base editor of any one of claims 33-56, a nucleic acid of claim 57 or 58, or a vector of claim 59; and(b) a Cas9 guide RNA (gRNA) ; wherein the Cas9-guided base editor and the Cas9 gRNA form a complex that modifies the target DNA or RNA.
67. The method of claim 34, wherein the Cas9 gRNA comprises a gRNA-scaffold is selected from SpCas9 gRNA-scaffold, SaCas9 gRNA-scaffold, StlCas9 gRNA-scaffold, and FnCas9-gRNA-scaffold.
68. The method of claim 66 or 67, wherein the Cas9 gRNA comprises a guide sequence of 8 nt, lOnt, 12nt, 20nt, 35nt, or 40 nt in length.
69. A method of modifying a target DNA or RNA comprising: contacting the target DNA or a cell comprising the target DNA or RNA with:(a) a Cas9-guided base editor of any one of claims 33-56, a nucleic acid of claim 57 or58, or a vector of claim 59; and(b) a Cas9 crRNA; and(c) a Cas9 tracrRNA; wherein the Cas9-guided base editor, the Cas9 crRNA, and the Cas9 tracrRNA form a complex that modifies the target DNA or RNA.
70. The method of claim 69, wherein the Cas9 crRNA is selected from SpCas9 crRNA, SaCas9 crRNA, StlCas9 crRNA, and FnCas9 crRNA, and wherein the Cas9 tracrRNA is selected from the same species as the Cas9 crRNA.
71. The method of claim 69 or 70, wherein the Cas9 crRNA comprises a guide sequence of 8 nt, lOnt, 12nt, 20nt, 35nt, or 40 nt in length.
72. A method of modifying a target DNA or RNA comprising contacting the target DNA or a cell comprising the target DNA or RNA with a composition or kit of any one of claims 60-65.
73. The method of any one of claims 66-72, wherein modifying the target DNA or RNA comprises inducing a base change of the target DNA or RNA.
74. The method of claim 73, wherein the base change is a A-to-G or A-to-I change.
75. The method of any one of claims 28-32 and 66-74, wherein the cell is a mammalian cell.
76. A Cas9 variant comprising an enlarged SpCas9 that comprises an insertion of one or more domains selected from a non-catalytic BH domain and a REC domain.
77. The Cas9 variant of claim 76, wherein the enlarged SpCas9 comprises an insertion of one or more domains selected from a non-catalytic StlCas9 BH domain and a StlCas9 REC domain.
78. The Cas9 variant of claim 76 or 77, wherein the enlarged SpCas9 comprises an insertion of a StlCas9 REC domain.
79. A method of producing a Cas9 variant of any one of claims 1-20 and 76-78.
80. A method of producing a Cas9-guided base editor of any one of claims 33-56.
81. A method of regulating activity and / or specificity of a catalytic domain linked to aCas9 scaffold comprising expanding the Cas9 scaffold by inserting one or more domains selected from non-catalytic BH domain and REC domain.
82. The method of claim 81, wherein the catalytic domain is an N-terminal catalytic domain.
83. The method of claim 81 or 82, wherein the catalytic domain is a base editor.
84. The method of any one of claims 81-83, wherein the catalytic domain has reduced byproducts.
85. The method of any one of claims 81-84, wherein the Cas9 scaffold is a SpCas9.
86. The method of any one of claims 81-85, comprising expanding the Cas9 scaffold by inserting one or more domains selected from a non-catalytic BH domain and a REC domain from StlCas9.
87. A method of repositioning an N-terminal catalytic domain linked to Cas9, comprising adjusting the length of a BH domain.
88. The method of claim 87, wherein the BH domain is a half (0.5x), single (lx), or two (2x) BH domains.
89. A method of reducing bystander editing of a deaminase-based base editor, comprising modifying a Cas9 scaffold.
90. The method of claim 89, wherein modifying the Cas9 scaffold comprises inserting one or more domains selected from a non-catalytic BH domain and a REC domain into the Cas9 scaffold.
91. The method of claim 89 or 80, wherein modifying the Cas9 scaffold comprises inserting one or more domains selected from a non-catalytic StlCas9 BH domain and a StlCas9 REC domain into the Cas9 scaffold.
92. The method of any one of claims 89-91, wherein the Cas9 scaffold is a nSpCas9.
93. A method of improving Cas9 binding specificity, comprising enlarging Cas9 by inserting one or more domains selected from a non-catalytic BH domain and a REC domain.
94. The method of claim 93, comprising inserting a StlCas9 REC domain to the Cas9.
95. A Cas9 variant, comprising nSpCas9 (D10A) that comprises an insertion of BH and REC domains of StlCas9.
96. The Cas9 variant of claim 95, wherein the nSpCas9 (D10A) comprises an insertion of the amino acid sequence of SEQ ID NO: 3, 7, or 8.
97. The Cas9 variant of claim 95 or 96, wherein the nSpCas9 (D10A) comprises an insertion of the amino acid sequence of SEQ ID NO: 3, 7, or 8 between residues 93D and 94D of the nSpCas9 (D10A).
98. The Cas9 variant of any one of claims 95-97, comprising the amino acid sequence selected from SEQ ID NOs: 23-25.
99. The Cas9 variant of any one of claims 95-98, wherein the Cas9 variant is used as a binding scaffold to generate a base editor.
Citation Information
Patent Citations
Cas9 crystals and methods of use thereof
US20160319262A1
Crispr enzyme mutations reducing off-target effects
WO2016205613A1
Nme2Cas9 INLAID DOMAIN FUSION PROTEINS
WO2023081070A1
Engineered chimeric ISCB polypeptides and uses thereof
WO2023230483A2