Activity-enhanced engineered spcas9 PAM variant enzymes for efficient and specific genome editing

WO2026024892A3PCT designated stage Publication Date: 2026-04-23THE GENERAL HOSPITAL CORP
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
THE GENERAL HOSPITAL CORP
Filing Date
2025-07-23
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Existing CRISPR-Cas enzymes, such as SpCas9, are constrained by their requirement for specific protospacer adjacent motifs (PAMs), limiting their ability to target and edit regions of the genome harboring non-canonical PAMs, and relaxation of PAMs leads to increased off-target potential and reduced efficiency.

Method used

Development of activity-enhanced SpCas9 PAM variant enzymes capable of targeting non-canonical NGA, NGC, or NGT PAMs while minimizing canonical NGG PAMs, achieved through directed evolution and structure-guided mutagenesis to improve on-target editing efficiency and specificity.

Benefits of technology

The altered PAM variants offer reduced off-target editing and enhanced editing efficiency in cell culture and primary human cells, enabling access to orthogonal regions of the genome and boosting the practicality of base editing on non-NGG PAMs.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

Provided herein are variant SpCas9 enzymes with altered PAM requirements capable of targeting non-canonical NGA, NGC, or NGT PAMs while minimizing the canonical NGG PAM, compositions comprising the enzymes, and methods of use thereof to alter DNA, e.g., genomic sequences, in a cell, in a subject, or in vitro.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Activity-enhanced engineered SpCas9 PAM variant enzymes for efficient and specific genome editing

[0002] CLAIM OF PRIORITY

[0003] This application claims the benefit of U.S. Provisional Patent Application Serial No. 63 / 674,450, filed on July 23, 2024. The entire contents of the foregoing are hereby incorporated by reference.

[0004] FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT

[0005] This invention was made with Government support under Grant Nos. CA281401, HL142494, and HG012010 awarded by the National Institutes of Health. The Government has certain rights in the invention.

[0006] TECHNICAL FIELD

[0007] Provided herein are variant SpCas9 enzymes with altered PAM requirements capable of targeting non-canonical NGA, NGC, or NGT PAMs while minimizing the canonical NGG PAM, compositions comprising the enzymes, and methods of use thereof to alter DNA, e.g., genomic sequences, in a cell, in a subject, or in vitro.

[0008] BACKGROUND

[0009] The adaptation of CRISPR-Cas enzymes as genome editing technologies has resulted in many editing approaches whose efficiencies are highly dependent upon the precise positioning of the Cas protein. For instance, accurate positioning of a DNA double-strand break (DSB) can substantially impact on-target efficiencies when using nucleases for homology directed repair (HDR)1, or for allele-specific editing2, and additionally for next-generation nickase-based technologies like base editing3, prime editing4, and other approaches. These genome editing methods require a diversity of Cas enzymes that together offer sufficient reprogrammability to accelerate biological research and the pursuit of genome editing therapies.

[0010] CRISPR-Cas enzymes are constrained to targeting and editing only genomic sequences that encode compatible protospacer adjacent motifs (PAMs)5. The PAM is a short nucleotide sequence that is directly readout by Cas enzymes through protein:DNA contacts to initiate target site binding and eventual stable R-loop formation6. The PAM requirement stems from the natural role of Cas enzymes in the CRISPR bacterial immune system, where they must discriminate between invading DNA and ‘self DNA at the endogenous CRISPR locus (where these sequences do and don’t encode PAMs, respectively)7. Thus, Cas enzymes including Cas9, Cast 2a, and others have evolved to proofread DNA substrates for the presence of PAMs to minimize inadvertent cleavage of their endogenous ‘self bacterial genome. Whereas the PAM requirement is an important mechanism to limit self-targeting in natural CRISPR systems, for genome editing applications this restriction can limit the utility of Cas enzymes by rending most of the genome inaccessible. The commonly used Streptococcus pyogenes Cas9 (SpCas9) largely specifies an NGG PAM on the 3’ of the spacer (where N is any base) with minor preferences for NAG and NGA PAMs8. As a genome editing tool, SpCas9 is therefore constrained to regions of the genome harboring conventional PAM, rendering many potentially impact edits impossible to implement.

[0011] Approaches to overcome the caveat of PAM recognition have included characterizing Cas orthologs that naturally recognize alternate PAMs9, developing chimeric Cas enzymes that graft PAM-interacting domains between Cas enzymes10, engineering of CRISPR enzymes to alter the PAM requirement to non-canonical PAMs (from one PAM to another)11, or relaxation of the PAM to permit broader targeting with a single enzyme (typically preserving the ability to target the canonical PAM)12, and each of these approaches have their benefits and caveats. For instance, exploration of Cas orthologs necessitates thorough biochemical characterization and optimization of many enzymatic properties and many orthologs are naturally less efficacious editors in mammalian cells, making this approach less scalable and consequently less feasible. The most widely pursued engineering strategy to expand targeting range has been to develop modified Cas9 enzymes capable of accommodating of additional PAMs, leading to Cas9 variants with broader PAM tolerances12 l 4. However, enzymes with relaxed PAMs can exhibit increased off- target potential since a larger fraction of the genome is now accessible, along with reduced overall efficiency due to kinetic repercussions of prolonged genome searching15. Instead, a potentially safer approach to expand genome access would involve developing an inventory of Cas9 enzymes with altered PAM requirements that are selective against other PAMs, including minimizing activity on NGG. Enzymes with altered PAMs have been more challenging to engineer and are typically less efficient compared to SpCas9 on target sites harboring NGG PAMs, possibly because the simplest engineering trajectory in most directed evolution schemes is to relax rather than to alter the PAM requirement. Thus, there is a need to engineer efficacious and specific altered PAM enzymes that permit access to non-NGG regions of the genome in a selective manner.

[0012] SUMMARY

[0013] Described herein are compositions and methods that can be used to improve the safety and precision of genome editing by developing activity-enhanced SpCas9 PAM variant enzymes that overcome the typical caveats of PAM-relaxed technologies. We performed directed evolution experiments in bacteria to select for altered PAM enzymes capable of targeting non-canonical NGA, NGC, or NGT PAMs while minimizing the canonical NGG PAM. Candidate enzymes for altered PAM targeting were then subjected to extensive structure-guided mutagenesis to improve their on-target editing efficiencies without sacrificing their genome-wide specificities. Via GUIDE-seq2, we demonstrate that these altered PAM variant enzymes offer specificity advantages by reducing off-target editing compared to relaxed PAM technologies like SpG or SpRY. The altered PAM enzymes are highly effective as nucleases and both C-to-T and A-to-G base editors in cell culture and primary human cells, enabling access to orthogonal regions of the genome compared to wild-type SpCas9 and greatly boosting the practicality of base editing on non-NGG PAMs. Together, these altered PAM variants permit efficient and specific genome editing in a broad scope of impactful applications.

[0014] Thus, provided herein are Streptococcus pyogenes Cas9 (SpCas9) proteins (and fusion proteins comprising the proteins or DNA binding portions thereof, e.g., in base editors) comprising mutations as listed in Table A. In some embodiments, the proteins are isolated.

[0015] In some embodiments, the proteins comprise a sequence that is at least 80% identical to the amino acid sequence of SEQ ID NO: 1.

[0016] In some embodiments, the proteins further comprise one or more additional mutations that decrease nuclease activity selected from the group consisting of mutations at DIO, E762, D839, H983, or D986; and at H840 or N863. In some embodiments, the proteins comprise the mutations: (i) D10A or DION, and / or (ii) H840A, H840N, or H840 Y. In some embodiments, the proteins further comprise one or more additional mutations that alter or increase specificity selected from the group consisting of mutations listed in Table C or D.

[0017] Additionally provided herein are fusion proteins comprising the proteins described herein (or DNA binding domains thereof), fused to a heterologous functional domain, with an optional intervening linker, wherein the linker does not interfere with activity of the fusion protein.

[0018] In some embodiments, the heterologous functional domain is a transcriptional activation domain. In some embodiments, the transcriptional activation domain is from VP16, VP64, rTA, NF-KB p65, or the composite VPR (VP64-p65-rTA).

[0019] In some embodiments, the heterologous functional domain is a transcriptional silencer or transcriptional repression domain. In some embodiments, the transcriptional repression domain is a Krueppel-associated box (KRAB) domain, ERF repressor domain (ERD), or mSin3 A interaction domain (SID), or the transcriptional silencer is Heterochromatin Protein 1 (HP1).

[0020] In some embodiments, the heterologous functional domain is an enzyme that modifies the methylation state of DNA. In some embodiments, the enzyme that modifies the methylation state of DNA is a DNA methyltransferase (DNMT) or a TET protein, optionally wherein the TET protein is TET1.

[0021] In some embodiments, the heterologous functional domain is an enzyme that modifies a histone subunit.

[0022] In some embodiments, the enzyme that modifies a histone subunit is a histone acetyltransferase (HAT), histone deacetylase (HD AC), histone methyltransferase (HMT), or histone demethylase.

[0023] In some embodiments, the heterologous functional domain is a base editing domain, e.g., the fusion protein is a base editor. In some embodiments, the base editor comprises: (i) a cytidine deaminase domain, preferably selected from the group consisting of the apolipoprotein B mRNA-editing enzyme, catalytic polypeptide-like (APOBEC) family of deaminases, optionally APOBEC1, APOBEC2, AP0BEC3A, APOBEC3B, APOBEC3C, AP0BEC3D / E, APOBEC3F, APOBEC3G, AP0BEC3H, or APOBEC4; activation-induced cytidine deaminase (AID), optionally activation induced cytidine deaminase (AICDA); cytosine deaminase 1 (CDA1) or CDA2; cytosine deaminase acting on tRNA (CD AT), or an engineered variant thereof, optionally DddA-like cytidine deaminase or and engineered TadA-domain, or (ii) an adenosine deaminase domain, preferably selected from the group consisting of adenosine deaminase 1 (ADA1), ADA2; adenosine deaminase acting on RNA 1 (AD ARI), ADAR2, ADAR3; adenosine deaminase acting on tRNA 1 (ADAT1), ADAT2, ADAT3; and naturally occurring or engineered tRNA-specific adenosine deaminase (TadA), optionally wherein the base editor comprises ABEs 0.1, 0.2, 1.1, 1.2, 2.1, 2.2, 2.3, 2.4, 2.5, 2.6, 2.7, 2.8, 2.9, 2.10, 2.11, 2.12, 3.1, 3.2, 3.3, 3.4, 3.5, 3.6, 3.7, 3.8, 4.1, 4.2, 4.3, 5.1, 5.2, 5.3, 5.4, 5.5, 5.6, 5.7, 5.8, 5.9, 5.10, 5.11, 5.12, 5.13, 5.14, 6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 7.1, 7.2, 7.3, 7.4, 7.5, 7.6, 7.7, 7.8, 7.9, 7.10, 8, 8a, 8b, 8c, 8d, 8e, 8.8, 8.13, 8.17, 8.20, 9, 9e, ABERA1.0-5.2, 8r, ABEmax, 10, ABExl, ABEx2, ABEx3, ABEx4, or hpABE5.20.

[0024] In some embodiments, the heterologous functional domain is a biological tether. In some embodiments, the biological tether is MS2, Csy4 or lambda N protein.

[0025] In some embodiments, the heterologous functional domain is Fokl.

[0026] Also provided herein are nucleic acids encoding the proteins described herein, vectors comprising the nucleic acids, optionally wherein the nucleic acid is operably linked to one or more regulatory domains for expressing the proteins described herein, with mutations at one, two, three, four, five, or all six of the following positions: DI 135, SI 136, G1218, E1219, R1335, and / or T1337, wherein the mutations are listed in Table A, and optionally a nucleic acid encoding a guide RNA that complexes with the cas9 protein. In some embodiments, the nucleic acids are isolated. In some embodiments, the nucleic acids are in an expression vector.

[0027] Additionally provided herein are host cells, preferably mammalian host cells, comprising the nucleic acids described herein, and optionally expressing a protein as described herein.

[0028] Further provided herein are methods of altering the genome of a cell. The methods comprise expressing in the cell, or contacting the cell with, a protein or fusion protein as described herein, and a guide RNA having a region complementary to a selected portion of the genome of the cell.

[0029] In some embodiments, the protein or fusion protein comprises one or more of a nuclear localization sequence, cell penetrating peptide sequence, and / or affinity tag. In some embodiments, the cell is a stem cell, e.g., an embryonic stem cell, mesenchymal stem cell, or induced pluripotent stem cell; is in a living animal; or is in an embryo.

[0030] Also provided herein are methods of altering a double stranded DNA (dsDNA) molecule. The methods comprise contacting the dsDNA molecule with a protein or fusion protein as described herein, and a guide RNA that complexes with the cas9 protein, the guide RNA having a region complementary to a selected portion of the dsDNA molecule.

[0031] In some embodiments, the dsDNA molecule is in vitro.

[0032] In some embodiments, the fusion protein and RNA are in a ribonucleoprotein complex.

[0033] In some embodiments, the guide RNA comprises a spacer sequence at least 80%, 90%, 95%, or 99% identical to a gRNA sequence listed in Table E, preferably wherein the base editor comprises a corresponding variant listed in Table E.

[0034] Also provided herein are methods of treating or reducing risk of a disease in a subject, the method comprising administering to the subject, or to a cell from the subject, a guide RNA comprising a spacer sequence at least 80%, 90%, 95%, or 99% identical to a gRNA sequence listed in Table E, and a base editor, preferably a base editor comprising a variant listed in Table A, B, or E (more preferably a base editor listed in the corresponding row of Table E).

[0035] Also provided herein are methods for installing AELAH3447R variant for treating or reducing risk of Alzheimer’s disease in a subject, comprising administering to the subject, or to a cell from the subject, a guide RNA comprising a spacer sequence at least 80%, 90%, 95%, or 99% identical to a gRNA sequence listed in the following table, and a base editor, preferably a base editor comprising a variant listed in the following table:

[0036] Also provided herein are methods for installation of BAG3 C151R genetic variant for treating or reducing risk of heart failure in a subject, comprising administering to the subject, or to a cell from the subject, a guide RNA comprising a spacer sequence at least 80%, 90%, 95%, or 99% identical to a gRNA sequence listed in the following table, and a base editor, preferably a base editor comprising a variant listed in the following table:

[0037] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Methods and materials are described herein for use in the present invention; other, suitable methods and materials known in the art can also be used. The materials, methods, and examples are illustrative only and not intended to be limiting. All publications, patent applications, patents, sequences, database entries, and other references mentioned herein are incorporated by reference in their entirety. In case of conflict, the present specification, including definitions, will control.

[0038] Other features and advantages of the invention will be apparent from the following detailed description and figures, and from the claims. DESCRIPTION OF DRAWINGS

[0039] FIGs. 1A-H: Selection of SpCas9 PAM variants and validation in human cells, a) Schematic of the SpCas9 tertiary complex with gRNA and target DNA showing the amino acid residues that are integral to forming contacts with the bases of the non-target strand DNA comprising the PAM (PDB: 5UN3). TS: target strand, NTS: non-target strand, PAM: protospacer adjacent motif, b) Selection scheme for creating SpCas9 PAM variants capable of targeting non-canonical PAMs by cleavage of a toxin plasmid c) Third position PAM preference plots for PAM variants as characterized via HT-PAMDA. X-axis values are displayed as the mean HT-PAMDA logio(&) value across the four possible NGNN PAMs with the specified third position (left to right: A, C, G, or T). Y-axis values are the corresponding mean HT-PAMDA logio(A) values for each variant averaged across the 12 PAMs possible with any other base in the third NGNN position. SpCas9 variants that are the focus of this study are designated with colored markers, d) Editing of targeted major PAMs in HEK293T cells for groups of variants identified to have third position PAM preferences of NGA, NGC, and NGT. Heatmaps show the mean editing for n = 3 biological replicates, e) Mean editing of 16 NGNN PAMs in HEK293T cells for candidate third position PAM variants; n = 3 biological replicates, f, g, h) Ratios of mean editing in HEK293T cells across 4 target PAMs to the editing observed on 12 other non-target PAMs to demonstrate the variants’ preferences of f) A, g) C, or h) T in the third position of the PAM.

[0040] FIGs. 2A-H: Activity-enhancement mutations in SpCas9 PAM variants improve editing efficiency while preserving PAM specificity, a) Schematic of adding activity-enhancing amino acid substitutions to SpCas9. These amino acid positions are distinct from the initial substitutions installed that alter the PAM preference. Activity-enhancing mutations were chosen based on the potential to form additional interactions with the DNA backbone of the target and non-target strands b) Effects of activity-enhancing mutations Lil HR and A1322R in combination on varying PAM variant scaffolds. On NGA PAMs, VQR benefits substantially from the additional mutations. On NGC PAMs, VRER and MQKSER also exhibit enhanced levels of editing. Efficiency gains were more pronounced on target sites which previously showed low editing with the base variant; SD and mean shown, c) Activity enhancing mutations added to SpGA improve average editing efficiency across 4 NGAPAM sites, with limited average editing gains observed across 12 other NGN PAM sites (4 NGC, 4 NGG, 4 NGT). d) Editing activity on 3 total NGC PAMs and 3 non-NGC PAMs by an enhanced SpGC variant compared to the base SpGC and SpG, with rescue of the observed bias against the NGCC PAM target by the base SpGC. e) HT-PAMDA data of activity-enhanced variants comparing the average rate of NGNN PAMs for a particular base in the third position (X-axis) against the average rate across all other third position NGNN PAMs. Symbols are colored by the starting scaffold PAM variant, with each symbol representing an activity-enhanced variant. Variants later chosen to be best-in-class are denoted with a diamond. Clockwise from top-left: preference for NGA PAMs; preference for NGC PAMs; preference for NGT PAMs; preference for NGG PAMs. f) Mean editing in HEK293T cells across 16 NGNN PAMs for SpGA, SpGC, SpGT, SpG, and five chosen activity -enhanced variants that exhibited strong preferences for NGA, NGC, or NGT PAMs in the HT- PAMDA dataset, g) Comparison of the mean modified reads in HEK293T cells on major (4 gRNA) and minor (12 gRNA) PAMs for each group of activity -enhanced variants, h) Best-in-class third position PAM variants (eSpGA, eSpGC, and eSpGT) and WT SpCas9 editing in HEK293T cells across 16 NGNN PAMs compared to the relaxed PAM SpG.

[0041] FIGs. 3A-I: Genome-wide specificity analysis of altered PAM variants and application towards allele-specific editing, a-c) On-target editing at each of 4 gRNA target sites in HEK293T cells resulting from GUIDE-seq2 transfections containing the dsODN tag for groups of altered PAM variants and SpG for a) NGA, b) NGC, and c) NGT PAMs. Percent modified reads assessed by targeted sequencing; mean, SD, and datapoints shown for n = 3 technical replicates, d, e, f) Fraction of GUIDE-seq2 reads attributed to either the on-target site or off-targets for each tested nuclease for target sites containing d) NGA, e) NGC, or f) NGT PAMs. g) Schematic of allele-specific editing by designing the gRNA to place the pathogenic mutation within the third position of the PAM. Relaxed PAM variants can edit both alleles, whereas selective PAM variants preferentially edit the desired allele based on its PAM specificity, h) Allele-specific editing in K562 cells of gRNA targeted to heterozygous locations in the K562 genome. eSpGT nearly exclusively targets the NGT-containing allele; WT SpCas9, the NGG-containing allele; and eSpGC, the NGC-containing allele. Use of SpCas9 variants with orthogonal PAM preferences allows for the selective depletion of the desired allele, i) Editing at other K562 heterozygous locations by a number of SpCas9 altered PAM variants, WT SpCas9, and SpG demonstrates a varied level of exclusivity for the target PAM. Fraction of total mapped reads that are associated with each unedited allele or are edited are shown for each nuclease at each target SNP.

[0042] Figure 4A-H: SpCas9 PAM variant base editors to correct pathogenic mutations in human cells, a, b) Base editing efficiencies using SpCas9 PAM variants within a) ABE8e for A-to-G editing or b) BE4max constructs for C-to-T editing on targets with NGA, NGC, or NGT PAMs; n = 2 or 3 biological replicates, c) Base editing of the target base in an endogenous HEK293T E7V cell line model to generate the E7A allele. PAM variant SpCas9 ABE8e constructs were targeted to the CAC PAM (NRCH) or TGC (eSpGC, SpG) PAM based on their PAM preferences and editing efficiencies were assessed by targeted sequencing, d) Fraction of GUIDE-seq2 reads mapping to on-target or off-target sites for nuclease versions of the SpCas9 PAM variants used to edit the E7V cell line. GUIDE-seq2 performed within the HEK293T endogenous E7V cell line, e) Correction of the SERPINA1 E366K mutation using a lentivirally-integrated reporter cell line containing the target base and genomic context. Editing efficiencies of ABE8.20m-SpCas9 PAM variants shown using a A7 - AGC gRNA. f) Allele frequencies resulting from editing the SERPINA1 model site by the ABE8.20m-SpCas9 PAM variants. Pure edits of only the A7 and edits including bystander A5 (A5 and A7) are predicted to produce functional protein. Mean allele frequencies across n = 3 biological replicates shown, g) Conversion of the pathogenic C445X mutation in CYBB to C445R using an A9-AGT PAM gRNA and PAM variant base editors. Base conversion efficiencies resulting from editing in a patient-derived B cell line h) Fraction of GUIDE-seq2 reads mapped to on-target or off-target locations for eSpGT, eSpGT. l(HFl) and SpG. GUIDE-seq2 performed in WT HEK293T cells using the CYBB C445X gRNA, thus the gRNA was targeted to the WT allele in this experiment.

[0043] FIGs. 5A-L: Precise, efficient installation of protective genetic variants using altered PAM base editors, a) Adenine base editing of RELN H3447R necessitates the use of an NGA PAM variant. Editing efficiency on each adenine within the edit window by ABE8.20m or ABE8e variants shows ABE8e-mediated bystander editing in addition to on-target editing by all ABEs. Fraction of total edits determined by dividing the A-to-G editing value for a given base (target A4, bystander A10 and A12, or translationally silent bystander A2) by the sum of the editing percentages across all adenines within the spacer. Mean fractions are shown for n = 3 biological replicates, b) Fraction of on- or off-target reads by GUIDE-seq2 using the RELN H3447R A4-CGA gRNA with nuclease versions of the PAM variants used to install the protective edit, c) Installation of the BAG3 C151R edit can be accomplished with a variety of PAM variant base editors using NGT, NGA, or NGG PAMs. Editing efficiency resulting from the use of these guides (X-axis, from top, A3 - TGG, A5 - AGT, A7 - TGG) with either ABE8e (A3 gRNA) or ABE8.20m (A3, A5, and A7 gRNA) was assessed by targeted sequencing using the same amplicon, d) GUIDE-seq2 on- and off-target fractions of nucleases used to edit the BAG3 C151R locus, e) Comparison of the base editing efficiencies on SLC30A8 R325W by a WT SpCas9 or eSpGT CBE(BE4max) using two gRNA that vary by a single base, f) GUIDE-seq2 fractions comparing the off-target profiles of WT SpCas9 (C8 gRNA) and eSpGT (C7 gRNA) at the SLC30A8 R325W site, g) Comparison of PAM variant SpCas9-TadCBEd editing efficiency with WT SpCas9-TadCBEd on the CFB R32Q site using two gRNA that vary by a single base, h) Predicted numbers of off-target sites in the genome by Cas-OFFinder. PAM type selected for eSpGA was VQR SpCas9: 5’-NGA-3’; and for SpG it was XCas9 3.7 (TLIKDIV) SpCas9 5’-NG-3’. Cas-OFFinder set to 5 mismatches, 0 DNA bulge, 0 RNA bulge for all samples, i, j, k, 1) Installation of four protective variants to disrupt protein function in LPA, IL33, HSD17B13, or CIDEB, respectively, using gRNA that require a PAM variant SpCas9- TadCBEd or -ABE8e for C-to-T or A-to-G editing. Comparisons performed between the altered PAM SpCas9 variant and a relaxed PAM variant control, SpG.

[0044] FIGs. 6A-C. Optimization of adenine base editors to install the RELN H3447R protective genetic variant, a) Schematic of the RELN locus with potential base editor target sites to install an A-to-G edit to generate the H3447R allele. Putative target sites with 20 nt protospacers and 4 nt PAMs are highlighted, with the position of the target adenine (A#) shown, b) A-to-G base editing efficiencies using various TadA8e-derived base editors (ABE8e), different SpCas9 PAM variant enzymes, and gRNAs to install the AEZAH3447R protective genetic variant. The target adenine for each graph is indicated with numbering based on position in the target site protospacer (with a 20 protospacer, counting from the PAM distal end of the target site). Different gRNA configurations were explored including either 20 or 21 nt spacer sequences, and either the conventional (“conv.”) SpCas9 gRNA scaffold or a flip and extend (“F&E”) scaffold (with a U / T5C substitution and extended crRNA / tracrRNA duplex). Data points represent editing efficiencies from experiments in HEK 293T cells; mean, standard deviation, and individual datapoints shown for n = 2 or 3 technical replicates, c) A-to-G base editing efficiencies using TadA8e-, TadA8.20m-, or TadA8.8m-derived adenine base editors (ABEs), different SpCas9 PAM variant enzymes, and gRNAs to install the RELN H3447R protective genetic variant. The target adenine for each graph is indicated with numbering based on position in the target site protospacer (with a 20 protospacer, counting from the PAM distal end of the target site; gRNA ID). Editing efficiencies from experiments in HEK 293T cells; mean, standard deviation, and individual datapoints shown for n = 2 or 3 technical replicates.

[0045] FIGs. 7A-B. Optimization of adenine base editors to install the BAG3 C151R protective genetic variant, a) Schematic of the BAG3 locus with potential base editor target sites to install an A-to-G edit to generate the C151R allele. Putative target sites with 20 nt protospacers and 4 nt PAMs are highlighted, with the position of the target adenine (A#) shown. Two potential silent bystander mutations may cooccur with the on-target edit, located at BAG3 amino acid positions C151 and D148. b) A-to-G base editing efficiencies using various TadA8e- or TadA8.8-based base editors and gRNAs to install the BAG3 C151R protective genetic variant. The target adenine is highlighted with an asterisk. Editing efficiencies from experiments in HEK 293T cells; mean, standard deviation, and individual datapoints shown for n = 2 or 3 technical replicates.

[0046] FIGs. 8A-B: Structural basis for non-canonical PAM recognition by SpCas9- VQR and SpCas9-VRER. a) Mechanism by which SpCas9-VQR (PDB: 5B2R) uses R1335Q substitution to recognize the second position adenine base of the PAM. b) Mechanism by which SpCas9-VRER (PDB: 5B2T) uses the R1335E and T1337R substitutions to interact with an NGCG PAM.

[0047] FIGs. 9A-B: HT-PAMD A profiling of altered PAM SpCas9 variants, a) HT- PAMDANGNN profiles of SpCas9 PAM variants arising from the 6AA mutagenesis and bacterial selection, with each heatmap sorted by mean activity on NGNN PAMs with a A, C, G, or T in the third position, from most active to least active, b) Fourth position PAM preference plots for PAM variants as characterized via HT-PAMDA. X- axis values are displayed as the mean HT-PAMDA loglO(k) value across the four possible NGNN PAMs with the specified fourth position (left to right: A, C, G, or T). Y-axis values are the corresponding mean HT-PAMDA loglO(k) values for each variant averaged across the 12 PAMs possible with any other base in the fourth NGNN position. SpCas9 variants that were previously known to have high levels of activity on a particular 4th position PAM are shown in colored markers. MQKSER has been observed to efficiently edit NGNG; VRQR, NGAG, xCas9 NGNC.

[0048] FIGs. 10A-E: Comparison of existing SpCas9 variants on NGNN sites in human cells, a) Editing of SpCas9 variants known to edit non-canonical PAMs across 32 gRNAs in HEK293T cells covering NGNN PAMs. b) Average efficiencies of each SpCas9 variant based on the identity of the third base in the target PAM. Datapoints shown are each the mean editing efficiency of 3 biological replicates of editing at a given target site with the appropriate PAM. Dark lines represent the mean for each nuclease, c) Mean editing at each gRNA target across all NGNN PAM sites tested. Of the variants tested, SpG shows the highest average across all NGNN PAMs, while NRRH the lowest among PAM relaxed variants, d) Editing efficiency of SpCas9- VRER on target sites containing the preferred PAM (NGCG) and other non-target PAMs. e) Editing by SpCas9-VRQR on 8 NGAPAMs and 2 of each NGC, NGG, and NGT PAMs.

[0049] FIGs. 11A-B: Adding putative activity-enhancing mutations to SpCas9-VQR and WT SpCas9. a) Editing efficiency of SpCas9-VQR variants with additional single mutations meant to increase activity. Target sites with NGA PAMs were selected due to the known NGA PAM preference of SpCas9-VQR b) Editing efficiency of WT SpCas9 variants with additional single mutations meant to increase activity.

[0050] FIGs. 12A-B: Screening of potential editing improvement of additional activity-enhancing substitutions, a, b) Putative activity-enhancing mutations were added to a) WT SpCas9 or b) SpCas9-VQR to determine the consequence to editing activity on 4 target PAM sites.

[0051] FIGs. 13A-D: Effects of combining activity-enhancing mutations on PAM variants, a, b) Assessment of editing levels generated by the addition of a) single activity-enhancing mutations to SpCas9-VRER (PAM preference: NGCG) across 16 NGNN PAMs or b) the same substitutions combined with Lil HR + A1322R. c, d) Assessment of editing levels generated by the addition of c) single activity-enhancing mutations to SpCas9-VRQR (PAM preference: NGA, NGNG) across 16 NGNN PAMS or b) the same substitutions combined with LI 111R + A1322R. n = 2 biological replicates

[0052] FIGs. 14A-B: Addition of individual and combinations of activity-enhancing substitutions to MQKSER. a, b) Assessment of editing levels generated by the addition of a) single activity-enhancing mutations to SpCas9-MQKSER (PAM preference: NGC, NGNG) across 16 NGNN PAMs or b) the same substitutions combined with L1111R + A1322R.

[0053] FIGs. 15A-D: Evaluation of various activity-enhancing mutations on the editing efficiencies of SpGA and SpGC. a) Heatmap summarizing the mean (n = 3 biological replicates) editing efficiency in HEK293T cells of SpGA and SpGA combined with various activity-enhancing mutations on NGNN PAM targets, b, c, d) Human cell editing efficiencies at 6 target sites for SpGC (N / A indicates zero additional substitutions) and many derivatives including activity -enhancing substitutions compared to SpG. b) Heatmap shows the mean across 3 biological replicates for each target site, c) Summarizes the average editing on 3 on-target (NGC) PAMs compared to 3 non-target (NGW in this data set) PAMs for each variant, d) Summarizes the average editing of these SpGC derivatives based on the 4th position in the PAM for on-target NGC PAMs to measure the improvement on NGCC PAMs (which SpGC had been shown to have weak activity on previously). NGCW is mean editing across the NGCT-1 and NGCA-1 sites, while NGCC is the editing on the NGCC- 1 site.

[0054] FIGs. 16A-B: Addition of individual and combinations of activity-enhancing substitutions to VRAVQL (SpGT). a, b) Assessment of editing levels generated by the addition of a) single activity-enhancing mutations to SpCas9-VRAVQL (PAM preference: NGT) across 16 NGNN PAMs or b) the same substitutions combined with L1111R + A1322R.

[0055] FIGs. 17A-B: Fourth position NGNN PAM preferences and overall PAM selectivity of activity-enhanced SpCas9 PAM variants, a) HT-PAMDA data of activity-enhanced variants comparing the average rate of NGNN PAMs for a particular base in the fourth position (X-axis) against the average rate across all other fourth position NGNN PAMs. Symbols are colored by the starting scaffold PAM variant, with each symbol representing an activity-enhanced variant. Variants later chosen to be best-in-class are denoted with a diamond. From left to right, the graphs show variants stratified by: preference for NGNA PAMs; preference for NGNC PAMs; preference for NGNG PAMs; preference for NGNT PAMs. b) Comparison of the PAM selectivity of SpGH variants (YSREQM, LWKFEG, VRAVQL) and five activity-enhanced derivatives, along with an SpG benchmark, in HEK293T cells (using the data originally displayed in Figure 2.2f) based on their relationship between mean modified reads on specified major PAMs (4 gRNA, NGAPAMs, NGC PAMs, NGT PAMs) versus a custom PAM selectivity score. PAM selectivity score was calculated by subtracting the mean editing across all minor PAMs (12 gRNA, NGB PAMs, NGD PAMs, NGV PAMs) from 100; variants closer to 100 are deemed more PAM-selective based on this metric.

[0056] FIG. 18: GUIDE-seq2 tag integration. Measurement of integration efficiency of the GUIDE-seq2 double-stranded oligodeoxynucleotide (dsODN) tag at 16 endogenous target sites. At each site, the SpGH (depending on the target PAM), eSpGH, and SpG enzymes were compared in nuclease-based experiments in HEK293T cells. CRISPResso2 HDR mode was used to assess dsODN integration efficiency after targeted amplicon sequencing; mean, standard deviation, and individual datapoints shown for n = 3 technical replicates.

[0057] FIG. 19: Base editing efficiencies of SpCas9-ABE8.20m PAM variants. Base editing (A-to-G) efficiencies using SpCas9 PAM variants within ABE8.20m architectures on their target PAMs, clockwise starting from top left: NGA, NGC, NGT, NGG PAMs. n = 3 biological replicates

[0058] FIGs. 20A-F: Comparison of TadCBEd to BE4max for cytosine base editing using PAM variants. All plots describe the C-to-T editing efficiency in HEK293T cells of various base editor constructs. SpG was included in all samples as a positive control for targeting NGN PAM targets, a) Two NGA target sites edited by TadCBEd- or BE4max-SpGA. Site dependent variable efficiency based on the deaminase domain can be observed, b) SpGC -based TadCBEd or BE4max editors mediate efficient editing in 2 NGC target sites, again with TadCBEd showing increased activity relative to BE4max on one of the sites, as in panel a. c) Editing of WT SpCas9-based TadCBEd or BE4max on 2 NGG target sites, d) Editing of SpGT-based CBEs demonstrates the efficient editing on 2 NGT target sites and similar edit windows between BE4max and TadCBEd on these sites. Note: no SpGT-BE4max samples included in this plot, e) Comparison of MQKSER or MQKFER-based CBEs on a single NGCA target site, f) TadCBEd-based CBEs work demonstrably more efficiently than BE4max on this NGAG target site targeted by SpGA variants.

[0059] FIG. 21A-D: Editing efficiencies of PAM variant base editors across both on- and off-target PAMs. a) A-to-G editing efficiency of the most-edited A in the spacer in HEK293T cells for ABE8e constructs made with: eSpGA, eSpGC, WT SpCas9, SpGT and enhanced versions, and PAM relaxed SpG and SpRY tested across 16 NGNN PAMs. n = 2 for eSpGC panel, b) C-to-T editing efficiency in HEK293T cells for BE4max constructs made with: SpGA, eSpGA, SpGC, eSpGC, WT SpCas9, SpGT, eSpGT, and PAM relaxed SpG tested across 16 NGNN PAMs. c, d) Preference for a particular 3rd position PAM (NGA, NGC, NGG, or NGT) over all other 3rd positions (NGB, NGD, NGH, or NGV) was shown via a ratio for a selection of PAM variant c) ABE8e or d) BE4max tested. Ratio calculated by taking the mean editing of the most- edited base for 4 target sites with the specified PAM (e.g. NGA) and dividing by the mean editing of the most-edited base across the 12 other targets with non -target PAMs (e.g. NGB). Note: y-axis scales are kept consistent between panel c and d.

[0060] FIG. 22: Editing efficiencies of each base in the edit window varies by compatibility between SpCas9 variant and target PAM. Plots shown describe the editing in HEK293T cells of PAM variant ABE8e constructs tested on a panel of 16 NGNN guides. Target sites with PAMs not represented, NGCT and NGGA, only had a single A base in the edit window for the guides used in the experiment. Y-axis value calculated by dividing the A-to-G editing at each adenine base in the spacer with editing > 1% by the most-edited base at the target site. By comparing the relative editing efficiencies of each base, apparent disparities in the edit windows can be observed depending on the SpCas9 variant and target PAM, such as in NGAC-1. Variants capable of efficiently editing NGA PAMs, SpRY, eSpGA, and SpG retain higher relative editing activity on A10, which is near the edge of what would be the expected edit window for ABE8e, than PAM variants which are known to disfavor NGA PAMs, such as eSpGT or WT SpCas9.

[0061] FIGs. 23A-B: Titration of ABE8e-SpCas9 PAM variants, a) Editing efficiencies (most-edited base reported as each datapoint) of a titration of ABE8e- SpCas9 PAM variant plasmids in HEK293T cells. Total plasmid (ng) was normalized throughout each sample with the addition of a pCMV-null plasmid, and gRNA was scaled equivalently to the ABE8e titration. 4 gRNA, 1 representing each NGN PAM, were targeted by all eSpGH PAM variants, n = 2 biological replicates, b) Base editing efficiencies of the highest (100 ng) to lowest (4 ng) plasmid doses analyzed to compare the relative average editing of each PAM variant on each target site. Fold change was calculated by dividing the editing (average of 2 biological replicates) in the low dose by the editing in the high dose for each target site by each ABE8e- SpCas9 variant. After cutting the amount of transfected plasmid by 25x, the SpCas9 variants show a greater degree of retaining of editing efficiency on their target PAMs than non-target PAMs (WT SpCas9 on the NGG PAM, for example). eSpGC not following the same trend could potentially be explained by the use of NGD sites all with 4th position G, and eSpGC was shown to have increased affinity for 4th position G, which could raise the affinity for those non-target PAMs relative to what would be expected.

[0062] FIGs. 24A-E: GUIDE-seq2 tag integration for disease-relevant target sites, a- e) Measurement of integration efficiency of the GUIDE-seq2 double-stranded oligodeoxynucleotide (dsODN) tag at target sites corresponding to either a-b) a therapeutic edit to correct a pathogenic mutation or c-e) to install a protective SNR At each site, the relevant SpCas9 PAM variants (depending on the target PAM) were compared in nuclease-based experiments in HEK293T cells. Notably, HBB E7V samples (a) were conducted in the HEK293T endogenous E7V cell line model to ensure the presence of the correct target site, and CYBB C445X (b) was conducted in HEK293T cells with a WT allele at that sequence, meaning there is not a perfectly matched on-target spacer, likely contributing to the very low tag integration efficiency. CRISPResso2 HDR mode was used to assess dsODN integration efficiency after targeted amplicon sequencing; mean, standard deviation, and individual datapoints shown for n = 3 technical replicates.

[0063] FIGs. 25A-B: Additional protective variants created by base editing, a) Editing efficiency of PAM variant SpCas9-TadCBEd to install the GPR75 protective SNP (rsl48952285) at target base C9 using a TGC PAM. b) Editing efficiency of eSpGC- and SpG-BE4max to generate the C-to-T SNP in MSTN protective against muscle degradation. DETAILED DESCRIPTION

[0064] CRISPR-Cas enzymes have revolutionized gene editing capabilities. In addition to nuclease-based gene editing, next generation tools have leveraged Cas enzymes to create more precise edits, but the scope of these technologies is limited by their PAM requirements. Current options of WT SpCas9 or broadened PAM variants of SpCas9 do not properly satisfy the needs of a wide range of applications that otherwise would be well suited for genome editing interventions. WT SpCas9, with its NGG PAM, is too restrictive in the range of targetable sequences. Conversely, relaxed PAM variants are too permissive in their target acquisition, leading to increased potential for off-targets.

[0065] To overcome this challenge, here, we developed a novel suite of SpCas9 variant enzymes that enable efficient and specific targeting of NGN PAMs, drastically increasing the genomic coordinates compatible with precise Cas9 positioning. We engineered these variants using structure-guided mutagenesis and a bacterial selection to identify variants capable of targeting non-canonical PAMs. Variants had their PAM preferences profiled in high-throughput, and the most promising variants were further characterized in human cells for their editing activity and PAM preference. We used further rational mutagenesis to engineer variants with enhanced editing activity based on supplementing the binding affinity of the Cas9 to its target DNA. We demonstrated that these variants exhibit fewer genome-wide off-targets than an existing variant with relaxed PAM preference. Further, we applied the PAM selective properties of these SpCas9 variants in numerous pre-clinical applications to show their utility, including correction of pathogenic mutations in primary patient cells by mRNA delivery. The suite of variants function as ABEs and CBEs on a broad range of target sites including target sites of potential therapeutic interest previously inaccessible without the use of an error-prone variant of broad PAM tolerance. We observed superior on-target base editing when using PAM variant enzymes compared to WT SpCas9 or the PAM relaxed enzymes SpG or SpRY. Unbiased off-target analysis via GUIDE-seq2 revealed that the PAM variant enzymes were capable of highly specific editing. We demonstrated the wide-ranging potential impacts of expanded genome access by the precise, efficient installation of protective variants to prevent Alzheimer’s disease, heart disease, diabetes, and macular degeneration. Further, we explored the use of PAM variant base editors to improve the precision of generating protective variants by disrupting splice sites in genes linked to asthma and fatty liver disease.

[0066] We demonstrate that a small toolbox of variants with altered PAM selectivity can offer a best of both worlds, enabling increased targeting diversity without compromising as much on off-targeting potential. This study builds upon previous Cas9 engineering work and offers lessons for future strategies, in particular with regard to the ways in which our novel variants may recognize their target PAM. Using existing structural information of SpCas9 and PAM variants, we can formulate general hypotheses regarding the role of mutated side chains in altered PAM recognition. It has been observed in structures of SpCas9 PAM variants EQR, VQR, and VRER16,28that mutations in the PAM interacting domain serve to stabilize the Cas9 protein on its PAM in three primary ways. Residues can form base-specific interactions, the new side chains can sterically accommodate the distorted PAM, and new Cas9:DNA interactions can supplement binding affinity lost to sub-optimal interactions between the PAM and Cas9. Our altered PAM variants and activity- enhanced variants highlight these principles, and some consistent trends emerge from our PAM alteration strategy. We observed the substitution of R at 1337 for several variants, and this residue is positioned in SpCas9 such that it can closely interact with the 4thposition of the PAM. Variants such as MQKSER, VRQR, and VRER utilize T1337R and accordingly have a strong sequence specific interaction with a G in the 4thposition of the PAM. Our variants with altered 3rdposition PAMs (SpGA = YSREQM, SpGC = LWKFEG, SpGT = VRAVQL) exhibit relatively little variability based on the base present in the 4thposition of the PAM. This is in agreement with their T1337 substitutions being residues (M, G, L) unlikely to form strong basespecific bonds.

[0067] The most active SpGC variants with a strong preference for 3rdposition cytosine all encode R1335E. This is consistent with structures of VRER reporting that the positive charge of E directly forms an interaction with the nucleobase carbonyl groups16,28. This interaction is further enabled by the distortion of the PAM DNA to move the 3rdbase closer to the 1335 side chain. This distortion likely plays a major role in allowing shorter side chains like E or Q (compared to R) to directly recognize the non-G base. R1335Q, common to previously reported VQR, and now VRQR and YSREQM, contains negatively charged side chains that coordinate with the amino group on adenine and may repel other negatively charged regions present on other bases. VRAVQL uses a Q side chain in the R1335 position, with its positive charge, to stabilize the thymine base of its preferred NGT PAM, which offers a negatively charged group outside the ring.

[0068] Our activity-enhancing variants also add to the overall understanding of Cas9 target capture and how it may impact editing efficiency. PAM recognition initiates SpCas9-mediated DNA unwinding, and SpCas9 is stabilized on the genome by PAM binding prior to the gRNA:target base pairing sequence check39. If the interactions between the PAM and SpCas9 are not sufficiently favorable to stabilize the complex during this process, the protein dissociates and moves on with its surveillance of the genome. By mutating several SpCas9 residues that previously participated in basespecific interactions with the PAM, we have likely created an overall interaction that is lower affinity than wild type, and thus more likely to destabilize, as evidenced by the comparatively low editing efficiencies resulting from our initial PAM variants. Adding more non-specific interactions between the protein and DNA can increase the favorable energetic interactions to lead to more instances of Cas9 proceeding to the nuclease activating conformation, but in both major and minor PAM contexts. Importantly, all activity-enhanced variants still achieve selective PAM recognition, especially compared to relaxed SpCas9 variants. The evidence thus far suggests that there is an optimal level of binding affinity between Cas9 and its target for efficient cleavage in human cells. Beyond this level, additional contacts seem to have little effect or even detrimental effects. Careful calibration of the binding affinity of the enzyme to attain a binding affinity at or near the optimal level for activity, but without adding unnecessary contacts that boost previously inconsequential editing at off- targets has emerged as a key consideration in Cas9 activity enhancement.

[0069] Extension of the work conducted here could potentially yield SpCas9 variants capable of targeting additional useful PAMS, such as those with even more specific requirements or beyond the set of PAMs with a second position Guanine. A similar strategy could be carried out on Cas9 orthologs other than SpCas9 to further diversify the available options for specific genome targeting. The activity enhancing work could prove beneficial to Cas9 proteins with desirable properties but limited editing activity in their current state, such as many smaller Cas9 orthologs or newly discovered CRISPR systems. Safety is of critical importance in potential therapeutic editing applications, and the PAM preference of Cas9 plays a major role in governing the level of off- targets. With BEs, precise PAM positioning allows for potentially increased on-target activity, reduction of bystander edits, and a more desirable genome wide off-target profile. We foresee the use of these variants to bring these benefits to numerous targets of clinical need that are currently only served using a broadly targetable variant.

[0070] Engineered Cas9 Variants with Altered PAM Specificities

[0071] The SpCas9 variants engineered in this study safety and precision of genome editing by developing activity-enhanced SpCas9 PAM variant enzymes that overcome the typical caveats of PAM-relaxed technologies. We performed directed evolution experiments in bacteria to select for altered PAM enzymes capable of targeting non- canonical NGA, NGC, or NGT PAMs while minimizing the canonical NGG PAM. The altered PAM specificity SpCas9 variants described herein can efficiently target endogenous gene sites in both bacterial and mammalian, e.g., human, cells.

[0072] All of the SpCas9 variants described herein can be rapidly incorporated into existing and widely used vectors, e.g., by simple site-directed mutagenesis, and because they require only a small number of mutations contained within the PAM- interacting domain, the variants should also work with other previously described improvements to the SpCas9 platform (e.g., truncated sgRNAs (Tsai et al., Nat Biotechnol 33, 187-197 (2015); Fu et al., Nat Biotechnol 32, 279-284 (2014)), nickase mutations (Mali et al., Nat Biotechnol 31, 833-838 (2013); Ran et al., Cell 154, 1380- 1389 (2013)), dimeric FokI-dCas9 fusions (Guilinger et al., Nat Biotechnol 32, 577- 582 (2014); Tsai et al., Nat Biotechnol 32, 569-576 (2014)); and high-fidelity variants (Kleinstiver et al. Nature 2016).

[0073] SpCas9 Variants with Altered PAM Specificity

[0074] Provided herein are SpCas9 variants. The SpCas9 wild type sequence is as follows:

[0075] 10 20 30 40 50 60

[0076] MDKKYS I GLD I GTNSVGWAV ITDEYKVPSK KFKVLGNTDR HS IKKNLI GA LLFDSGETAE

[0077] 70 80 90 100 110 120

[0078] ATRLKRTARR RYTRRKNRI C YLQE I FSNEM AKVDDS FFHR LEES FLVEED KKHERHPI FG

[0079] 130 140 150 160 170 180

[0080] NIVDEVAYHE KYPTIYHLRK KLVDSTDKAD LRLIYLALAH MIKFRGHFLI EGDLNPDNSD 190 200 210 220 230 240

[0081] VDKLFIQLVQ TYNQLFEENP INASGVDAKA I LSARLSKSR RLENLIAQLP GEKKNGLFGN

[0082] 250 260 270 280 290 300

[0083] LIALSLGLTP NFKSNFDLAE DAKLQLSKDT YDDDLDNLLA QI GDQYADLF LAAKNLSDAI

[0084] 310 320 330 340 350 360

[0085] LLSDI LRVNT E ITKAPLSAS MIKRYDEHHQ DLTLLKALVR QQLPEKYKE I FFDQSKNGYA

[0086] 370 380 390 400 410 420

[0087] GYIDGGASQE E FYKFIKPI L EKMDGTEELL VKLNREDLLR KQRTFDNGS I PHQIHLGELH

[0088] 430 440 450 460 470 480

[0089] AI LRRQEDFY PFLKDNREKI EKI LTFRI PY YVGPLARGNS RFAWMTRKSE ETITPWNFEE

[0090] 490 500 510 520 530 540

[0091] VVDKGASAQS FIERMTNFDK NLPNEKVLPK HSLLYEYFTV YNELTKVKYV TEGMRKPAFL

[0092] 550 560 570 580 590 600

[0093] SGEQKKAIVD LLFKTNRKVT VKQLKEDYFK KIECFDSVE I SGVEDRFNAS LGTYHDLLKI

[0094] 610 620 630 640 650 660

[0095] IKDKDFLDNE ENEDI LEDIV LTLTLFEDRE MIEERLKTYA HLFDDKVMKQ LKRRRYTGWG

[0096] 670 680 690 700 710 720

[0097] RLSRKLINGI RDKQSGKTI L DFLKSDGFAN RNFMQLIHDD SLTFKEDIQK AQVSGQGDSL

[0098] 730 740 750 760 770 780

[0099] HE HI AN LAGS PAIKKGI LQT VKVVDELVKV MGRHKPENIV IEMARENQTT QKGQKNSRER

[0100] 790 800 810 820 830 840

[0101] MKRIEEGIKE LGSQI LKEHP VENTQLQNEK LYLYYLQNGR DMYVDQELDI NRLSDYDVDH

[0102] 850 860 870 880 890 900

[0103] IVPQS FLKDD S IDNKVLTRS DKNRGKSDNV PSEEVVKKMK NYWRQLLNAK LITQRKFDNL

[0104] 910 920 930 940 950 960

[0105] TKAERGGLSE LDKAGFIKRQ LVETRQITKH VAQI LDSRMN TKYDENDKLI REVKVITLKS

[0106] 970 980 990 1000 1010 1020

[0107] KLVSDFRKDF QFYKVRE INN YHHAHDAYLN AVVGTALIKK YPKLESE FVY GDYKVYDVRK

[0108] 1030 1040 1050 1060 1070 1080

[0109] MIAKSEQE I G KATAKYFFYS NIMNFFKTE I TLANGE IRKR PLIETNGETG E IVWDKGRDF

[0110] 1090 1100 1110 1120 1130 1140

[0111] ATVRKVLSMP QVNIVKKTEV QTGGFSKES I LPKRNSDKLI ARKKDWDPKK YGGFDS PTVA

[0112] 1150 1160 1170 1180 1190 1200

[0113] YSVLVVAKVE KGKSKKLKSV KELLGITIME RSS FEKNPID FLEAKGYKEV KKDLI IKLPK

[0114] 1210 1220 1230 1240 1250 1260

[0115] YSLFELENGR KRMLASAGEL QKGNELALPS KYVNFLYLAS HYEKLKGS PE DNEQKQLFVE

[0116] 1270 1280 1290 1300 1310 1320

[0117] QHKHYLDE I I EQI SE FSKRV I LADANLDKV LSAYNKHRDK PIREQAENI I HLFTLTNLGA

[0118] 1330 1340 1350 1360

[0119] PAAFKYFDTT IDRKRYTSTK EVLDATLIHQ S ITGLYETRI DLSQLGGD ( SEQ ID NO : 1 ) The SpCas9 variants described herein can include combinations of mutations as shown herein. In some embodiments, the SpCas9 variants are at least 80%, e.g., at least 85%, 90%, or 95% identical to the amino acid sequence of SEQ ID NO: 1, e.g., have differences at up to 5%, 10%, 15%, or 20% of the residues of SEQ ID NO: 1 replaced, e.g., with conservative mutations. In preferred embodiments, the variant retains desired activity of the parent, e.g., the nuclease activity (except where the parent is a nickase or a dead Cas9), and / or the ability to interact with a guide RNA and target DNA).

[0120] In some embodiments, the SpCas9 variants include a set of six mutations at D1135, S1136, G1218, E1219, R1335, and T1337, or four mutations at DI 135,

[0121] G1218, R1335, and T1337 and optionally one or more additional residues, as shown in Table A.

[0122] Table A. SpCas9 PAM variants with activity enhancing mutations _ _

[0123] YSREQM + A1322R + L1111R *eSpGA

[0124] YSREQM + A61R + A1285K

[0125] YSREQM + A1285K + A1322R

[0126] YSREQM + A1285K + G366R

[0127] YSREQM + S55R + G366R

[0128] YSREQM + D1332K + A1285K

[0129] YSREQM + G366R + L1111R

[0130] YSREQM + N394K + L1111R

[0131] YSREQM + D1332K + A1322R

[0132] YSREQM + D1332K + G366R

[0133] YSREQM + D1332K + L1111R

[0134] YSREQM + A61R + A1322R

[0135] YSREQM + N394K + A1322R

[0136] YSREQM + A1285K + L1111R

[0137] YSREQM + G366R + A1322R

[0138] YSREQM + D1332K + N394K

[0139] YSREQM + A61R + D1332K

[0140] YSREQM + S55R + D1332K

[0141] YSREQM + N394K + G366R

[0142] YSREQM + A1285K + N394K

[0143] YSREQM + A61R + L1111R

[0144] YSREQM + S55R + L1111R

[0145] YSREQM + S55R + A1322R

[0146] YSREQM + A1322R YSREQM + A61R + N394K

[0147] YSREQM + D1332K

[0148] YSREQM + S55R + A1285K

[0149] YSREQM + A61R + G366R

[0150] YSREQM + S55R + N394K

[0151] YSREQM + A61R + S55R

[0152] LWKFEG + A1285K + L1111R *eSpGC

[0153] LWKFEG + A1285K + N394K

[0154] LWKFEG + A1285K + G366R

[0155] LWKFEG + G366R + L1111R

[0156] LWKFEG + N394K + G366R

[0157] LWKFEG + L1111R + A1322R

[0158] LWKFEG + D1332K + L1111R

[0159] LWKFEG + N394K + L1111R

[0160] LWKFEG + G366R + A1322R

[0161] LWKFEG + A1285K

[0162] LWKFEG + D1332K

[0163] LWKFEG + N394K + A1322R

[0164] LWKFEG + A61R + N394K

[0165] LWKFEG + D1332K + A1322R

[0166] LWKFEG + S55R + L1111R

[0167] LWKFEG + A1285K + A1322R

[0168] LWKFEG + A61R + L1111R

[0169] LWKFEG + D1332K + A1285K

[0170] LWKFEG + S55R + N394K

[0171] LWKFEG + D1332K + G366R

[0172] LWKFEG + A61R + D1332K

[0173] LWKFEG + A61R + G366R

[0174] LWKFEG + A61R + A1322R

[0175] LWKFEG + S55R + A61R

[0176] LWKFEG + S55R + A1322R

[0177] LWKFEG + S55R + G366R

[0178] LWKFEG + S55R + A1285K

[0179] LWKFEG + D1332K + N394K

[0180] LWKFEG + S55R + D1332K

[0181] LWKFEG + A61R + A1285K

[0182] SpGT variants with enhanced activity (SpGT = VRAVQL)

[0183] VRAVQL + A1285K + G366R *eSpGT

[0184] VRAVQL + A61R + L1111R + A1322R + N497A + R661A

[0185] + Q695A + Q926A eSpGT.2-H Fl

[0186] VRAVQL + D1332K + A1285K

[0187] VRAVQL + N394K + A1322R VRAVQL + D1332K + A1322R

[0188] VRAVQL + D1332K + N394K

[0189] VRAVQL + A1285K + N394K

[0190] VRAVQL + N394K + G366R

[0191] VRAVQL + G366R + A1322R

[0192] VRAVQL + D1332K + G366R

[0193] VRAVQL + S55R + A1322R

[0194] VRAVQL + A1285K + A1322R

[0195] VRAVQL + A61R + A1322R

[0196] VRAVQL + A61R + G366R

[0197] VRAVQL + N394K + L1111R

[0198] VRAVQL + A1285K + L1111R

[0199] VRAVQL + A61R + D1332K

[0200] VRAVQL + D1332K + L1111R

[0201] VRAVQL + S55R + G366R

[0202] VRAVQL + S55R + A1285K

[0203] VRAVQL + G366R + L1111R

[0204] VRAVQL + S55R + D1332K

[0205] VRAVQL + S55R + N394K

[0206] VRAVQL + A61R + L1111R

[0207] VRAVQL + A61R + A1285K

[0208] VRAVQL + A61R + S55R

[0209] VRAVQL + S55R + L1111R

[0210] VRAVQL + A61R + N394K

[0211] Table B. Additional Enzymes _

[0212] SpGT variants with enhanced activity (SpGT = VRAVQL)

[0213] VRAVQL+A1285K Prev. reported

[0214] VRAVQL+D1332K Prev. reported

[0215] VRAVQL+N394K Prev. reported

[0216] *eSpGT.2 - Prev.

[0217] VRAVQL + A61R + L1111R + A1322R reported

[0218] VRAVQL + A1322R + L1111R Prev. reported

[0219] VRAVQ L + A61 R P rev. re po rted

[0220] VRQR+D1332K Prev. reported

[0221] VRQR+N394K Prev. reported

[0222] VRQR+L1111R+A1322R+G366R Prev. reported

[0223] VRQR+G366R Prev. reported

[0224] VRQR+A61R Prev. reported

[0225] VRQR+L1111R+A1322R+N394K Prev. reported

[0226] VRQR+L1111R+A1322R+A1285K Prev. reported

[0227] VRQR+S55R *eVRQR -Prev. reported VRQR+L1111R+A1322R+D1332K Prev. reported

[0228] VRQR+L1111R+A1322R+A61R Prev. reported

[0229] MQKSER variants with enhanced activity

[0230] MQKSER+A1322R Prev. reported

[0231] MQKSER+L1111R+A1322R+G366R Prev. reported

[0232] MQKSER+L1111R+A1322R+A1285K Prev. reported

[0233] MQKSER+L1111R+A1322R+N394K Prev. reported

[0234] MQKSER+L1111R+A1322R+D1332K Prev. reported

[0235] MQKSER+L1111R Prev. reported

[0236] *eMQKSER -Prev.

[0237] MQKSER+L1111R+A1322R+A61R reported

[0238] MQKSER+L1111R+A1322R+S55R Prev. reported

[0239] MQKSER+N394K Prev. reported

[0240] MQKSER+D1332K Prev. reported

[0241] MQKSER+S55R Prev. reported

[0242] VRER+L1111R+A1322R+G366R Prev. reported

[0243] VRER+L1111R+A1322R+D1332K Prev. reported

[0244] VRER+N394K Prev. reported

[0245] VRER+D1332K Prev. reported

[0246] VRER+A1285K *eVRER -Prev. reported

[0247] VRER+A61R Prev. reported

[0248] VRER+S55R Prev. reported

[0249] The parental variants are as follows: _

[0250] In some embodiments, the SpCas9 variants also include mutations at one of the following amino acid positions, which reduce or destroy the nuclease activity of the Cas9: DIO, E762, D839, H983, or D986 and H840 or N863, e.g., D10A / D10N and

[0251] H840A / H840N / H840Y, to render the nuclease portion of the protein catalytically inactive; substitutions at these positions could be alanine (as they are in Nishimasu al., Cell 156, 935-949 (2014)), or other residues, e.g., glutamine, asparagine, tyrosine, serine, or aspartate, e.g., E762Q, H983N, H983Y, D986N, N863D, N863S, or N863H (see WO 2014 / 152432). In some embodiments, the variant includes mutations at

[0252] D10A or H840A (which creates a single strand nickase), or mutations at D10A and H840A (which abrogates nuclease activity; this mutant is known as dead Cas9 or dCas9).

[0253] In some embodiments, the SpCas9 variants also include mutations at one or more amino acid positions that increase the specificity of the protein (i.e., reduce off- target effects). In some embodiments, the SpCas9 variants include one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or all thirteen mutations at the following residues: N497, K526, R661, R691, N692, M694, Q695, H698, K810, K848, Q926, K1003, and / or R0160. In some embodiments, the mutations are: N692A, Q695A, Q926A, H698A, N497A, K526A, R661A, R691A, M694A, K810A, K848A, K1003A, R0160A, Y450A / Q695A, L169A / Q695A, Q695A / Q926A, Q695A / D1135E, Q926A / D1135E, Y450A / D1135E, L169A / Y450A / Q695A, L169A / Q695A / Q926A, Y450A / Q695A / Q926A, R661A / Q695A / Q926A, N497A / Q695A / Q926A, Y450A / Q695A / D1135E, Y450A / Q926A / D1135E, Q695A / Q926A / D1135E, L169A / Y450A / Q695A / Q926A, L 169A / R661 A / Q695 A / Q926 A, Y450A / R661 A / Q695 A / Q926 A, N497A / Q695A / Q926A / D1135E, R661A / Q695A / Q926A / D1135E, and Y450A / Q695A / Q926A / D1135E; N692A / M694A / Q695A / H698A, N692A / M694A / Q695A / H698A / Q926A; N692A / M694A / Q695A / Q926A; N692A / M694A / H698A / Q926A; N692A / Q695A / H698A / Q926A; M694A / Q695A / H698A / Q926A; N692A / Q695A / H698A; N692A / M694A / Q695A; N692A / H698A / Q926A; N692A / M694A / Q926A; N692A / M694A / H698A; M694A / Q695A / H698A; M694A / Q695A / Q926A; Q695A / H698A / Q926A;

[0254] G582 A / V583 A / E584 A / D585 A / N588 A / Q926 A; G582A / V583A / E584A / D585A / N588A; T657A / G658A / W659A / R661A / Q926A; T657A / G658 A / W659A / R661 A; F491 A / M495 A / T496A / N497A / Q926 A; F491A / M495A / T496A / N497A; K918A / V922A / R925A / Q926A; or 918A / V922A / R925A; K855A; K810A / K1003A / R1060A; or K848A / K1003A / R1060A. See, e.g., US9512446B1; Kleinstiver et al., Nature. 2016 Jan 28;529(7587):490-5; SI ay maker et al., Science. 2016 Jan l;351(6268):84-8; Chen et al., Nature. 2017 Oct 19;550(7676):407-410; Tsai and Joung, Nature Reviews Genetics 17:300-312 (2016); Vakulskas et al., Nature Medicine 24: 1216-1224 (2018); Casini et al., Nat Biotechnol. 2018 Mar;36(3):265-271. In some embodiments, the variants do not include mutations at K526 or R691. In some embodiments, the SpCas9 variants include mutations at one, two, three, four, five, six or all seven of the following positions: L169A, Y450, N497, R661, Q695, Q926, and / or DI 135E, e.g., in some embodiments, the variant SpCas9 proteins comprise mutations at one, two, three, or all four of the following: N497, R661, Q695, and Q926, e.g., one, two, three, or all four of the following mutations: N497A, R661A, Q695A, and Q926A. In some embodiments, the variant SpCas9 proteins comprise mutations at Q695 and / or Q926, and optionally one, two, three, four or all five of L169, Y450, N497, R661 and DI 135E, e.g., including but not limited to Y450A / Q695A, L169A / Q695A, Q695A / Q926A, Q695A / D1135E, Q926A / D1135E, Y450A / D1135E, L169A / Y450A / Q695A, L169A / Q695A / Q926A, Y450A / Q695A / Q926A, R661A / Q695A / Q926A, N497A / Q695A / Q926A, Y450A / Q695A / D1135E, Y450A / Q926A / D1135E, Q695A / Q926A / D1135E, L 169A / Y450A / Q695 A / Q926 A,

[0255] L 169A / R661 A / Q695 A / Q926 A, Y450A / R661 A / Q695 A / Q926 A, N497A / Q695A / Q926A / D1135E, R661A / Q695A / Q926A / D1135E, and Y450A / Q695A / Q926A / D1135E. See, e.g., KI einstiver et al., Nature 529:490-495 (2016); WO 2017 / 040348; US 9,512,446).

[0256] In some embodiments, the SpCas9 variants also include mutations at one, two, three, four, five, six, seven, or more of the following positions: F491, M495, T496, N497, G582, V583, E584, D585, N588, T657, G658, W659, R661, N692, M694, Q695, H698, K918, V922, and / or R925, and optionally at Q926, preferably comprising a sequence that is at least 80% identical to the amino acid sequence of SEQ ID NO: 1 with mutations at one, two, three, four, five, six, seven, or more of the following positions: F491, M495, T496, N497, G582, V583, E584, D585, N588, T657, G658, W659, R661, N692, M694, Q695, H698, K918, V922, and / or R925, and optionally at Q926, and optionally one or more of a nuclear localization sequence, cell penetrating peptide sequence, and / or affinity tag.

[0257] In some embodiments, the proteins comprise mutations at one, two, three, or all four of the following: N692, M694, Q695, and H698; G582, V583, E584, D585, and N588; T657, G658, W659, and R661; F491, M495, T496, and N497; or K918, V922, R925, and Q926.

[0258] In some embodiments, the proteins comprise one, two, three, four, or all of the following mutations: N692A, M694A, Q695A, and H698A; G582A, V583A, E584A, D585A, and N588A; T657A, G658A, W659A, and R661 A; F491 A, M495A, T496A, and N497A; or K918A, V922A, R925A, and Q926A.

[0259] In some embodiments, the proteins comprise mutations: N692A / M694A / Q695A / H698A. In some embodiments, the proteins comprise mutations at F539, M763, and K890, e.g., F539S, M763I, K890N (Lee et al., Nature Communications volume 9, Article number: 3048 (2018)), and optionally at E1007, e.g., E1007L or E1007P (Kim et al., Nature Chemical Biology volume 19, pages972- 980 (2023)).

[0260] In some embodiments, the proteins comprise mutations: N692A / M694A / Q695A / H698A / Q926A; N692A / M694A / Q695A / Q926A; N692A / M694A / H698A / Q926A; N692A / Q695A / H698A / Q926A; M694A / Q695A / H698A / Q926A; N692A / Q695A / H698A; N692A / M694A / Q695A; N692A / H698A / Q926A; N692A / M694A / Q926A; N692A / M694A / H698A; M694A / Q695A / H698A; M694A / Q695A / Q926A; Q695A / H698A / Q926A;

[0261] G582 A / V583 A / E584 A / D585 A / N588 A / Q926 A;

[0262] G582A / V583A / E584A / D585A / N588A; T657A / G658A / W659A / R661A / Q926A; T657A / G658A / W659A / R661A; F491A / M495A / T496A / N497A / Q926A;

[0263] F491A / M495A / T496A / N497A; K918A / V922A / R925A / Q926A; or 918A / V922A / R925A. See, e.g., Chen et al., “Enhanced proofreading governs CRISPR-Cas9 targeting accuracy,” bioRxiv, doi.org / 10.1101 / 160036 (August 12, 2017) and Nature. 2017 Oct 19;550(7676):407-410; Nishimasu et al., Nature volume 550, pages407-410 (2017).

[0264] In some embodiments, the variant proteins include mutations at one or more of R780, K810, R832, K848, K855, K968, R976, H982, K1003, K1014, K1047, and / or R1060, e.g., R780A, K810A, R832A, K848A, K855A, K968A, R976A, H982A, K1003A, K1014A, K1047A, and / or R1060A, e.g., K855A; K810A / K1003A / R1060A; (also referred to as eSpCas9 1.0); or K848A / K1003A / R1060A (also referred to as eSpCas9 1.1) (see Slaymaker et al., Science. 2016 Jan l;351(6268):84-8).

[0265] The variant proteins can also include one or more mutations that increase activity, reduce off-target effects, and / or alter protospacer adjacent motif (PAM) or target adjacent motif (TAM) specificity (Tables C and D).

[0266] Table C: List of Exemplary High Fidelity and / or PAM-relaxed RGN Orthologs

[0267] * predicted based on UniRule annotation on the UniProt database.

[0268] Table D. List of Exemplary SpCas9 Activity-Altering Mutations The variant proteins described herein can be used in place of the SpCas9 proteins described in the foregoing references with guide RNAs that target sequences that have PAM sequences as described herein.

[0269] In addition, the variants described herein can be used in fusion proteins in place of the wild-type Cas9 or other Cas9 mutations (such as the dCas9 or Cas9 nickase described above) as known in the art, e.g., a fusion protein with a heterologous functional domain as described in WO 2014 / 124284. For example, the variants, preferably comprising one or more nucl ease-reducing or killing mutation, can be fused on the N or C terminus of the Cas9 to a transcriptional activation domain or other heterologous functional domains (e.g., transcriptional repressors (e.g., KRAB, ERD, SID, and others, e.g., amino acids 473-530 of the ets2 repressor factor (ERF) repressor domain (ERD), amino acids 1-97 of the KRAB domain of K0X1, or amino acids 1-36 of the Mad mSIN3 interaction domain (SID); see Beerli et al., PNAS USA 95: 14628-14633 (1998)) or silencers such as Heterochromatin Protein 1 (HP1, also known as swi6), e.g., HPla or HPIP; proteins or peptides that could recruit long non-coding RNAs (IncRNAs) fused to a fixed RNA binding sequence such as those bound by the MS2 coat protein, endoribonuclease Csy4, or the lambda N protein; enzymes that modify the methylation state of DNA (e.g., DNA methyltransferase (DNMT) or TET proteins); enzymes that modify histone subunits (e.g., histone acetyltransferases (HAT), histone deacetylases (HDAC), histone methyltransferases (e.g., for methylation of lysine or arginine residues) or histone demethylases (e.g., for demethylation of lysine or arginine residues)).

[0270] The Cas9 editing enzyme can be part of a base editor (e.g., a cytidine base editor or adenine base editor), e.g., a fusion protein comprising a base editing domain such as a deaminase (e.g., cytosine or adenosine deaminase) and a Cas9 DNA binding domain. Base Editors are known in the art and include cytosine base editors (CBE) and adenine base editors (ABE) that allow for the targeted deamination of cytosines and adenines, respectively, that are exposed on ssDNA by RNA-guided CRISPR-Cas proteins. Cytosine base editors (CBEs), such as BE3 or BE4max, catalyze the conversion of target C»G base pairs to T»A, while adenine base editors (ABEs), such as ABE7.10, ABEmax, or ABE8, convert target A»T base pairs to G»C. Base editing with canonical base editors requires the presence of a PAM located approximately

[0271] 15±2 base pairs from the target nucleotide(s). See, e.g., Komor, A.C. et al., Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and product purity, Sci Adv 3 (2017); Rees, H.A. et al., Improving the DNA specificity and applicability of base editing through protein engineering and protein delivery, Nat. Commun. 8, 15790 (2017); US2018 / 0073012, US2017 / 0121693, WO2017 / 070633, US2015 / 0166980, U.S. Patent No. 9,840,699; and U.S. Patent No. 10,077,453. Split ABEs or CBEs can also be used. Thus, in some embodiments, the heterologous functional domain is a base editor, e.g., a deaminase that modifies cytosine DNA bases, e.g., a cytidine deaminase from the apolipoprotein B mRNA-editing enzyme, catalytic polypeptide-like (APOBEC) family of deaminases, including APOBEC 1 , APOBEC2, APOBEC3 A, APOBEC3B, APOBEC3C, APOBEC3D / E, APOBEC3F, APOBEC3G, APOBEC3H, and APOBEC4 (see, e.g., Yang et al., J Genet Genomics. 2017 Sep 20;44(9):423-437); activation-induced cytidine deaminase (AID), e.g., activation induced cytidine deaminase (AICDA); cytosine deaminase 1 (CDA1) and CDA2; cytosine deaminase acting on tRNA (CD AT); and DddA-like cytidine deaminases (Huang et al., Cell. 2023 Jul 20;186(15):3182-3195.el4. The following table provides exemplary sequences; other sequences can also be used.

[0272] * from Saccharomyces cerevisiae S288C

[0273] In some embodiments, the heterologous functional domain is a deaminase that modifies adenosine DNA bases, e.g., the deaminase is an adenosine deaminase 1 (ADA1), ADA2; adenosine deaminase acting on RNA 1 (AD ARI), ADAR2, ADAR3 (see, e.g., Savva et al., Genome Biol. 2012 Dec 28;13(12):252); adenosine deaminase acting on tRNA 1 (ADAT1), ADAT2, ADAT3 (see Keegan et al., RNA. 2017 Sep;23(9): 1317-1328 and Schaub and Keller, Biochimie. 2002 Aug;84(8):791-803); and naturally occurring or engineered tRNA-specific adenosine deaminase (TadA) (see, e.g., Gaudelli et al., Nature. 2017 Nov 23;551(7681):464-471) (NP_417054.2 (Escherichia coli str. K-12 substr. MG1655); See, e.g., Wolf et al., EMBO J. 2002 Jul 15;21(14):3841-51). The following table provides exemplary sequences; other sequences can also be used. For example, the TadA domain has also been engineered to purposefully generate C-to-T edits in addition to, or instead of, the conventional A- to-G edits observed with TadA domains; these versions can also be used; see, e.g., Chen et al., Nature Biotechnology volume 41, pages663-672 (2023); Lam et al., Nature Biotechnology volume 41, pages686-697 (2023); and Neugebauer et al., Nature Biotechnology volume 41, pages673-685 (2023); examples include TadA8e, TadA8.20m, or TadA8.8m, e.g., as described in Gaudelli et al., Nature Biotechnology volume 38, pages892-900 (2020); and the engineered Tad-CBEs described in Wu et al., Nat Biotechnol (2025), doi.org / 10.1038 / s41587-025-02678-w.

[0274] A number of variants of base editors have been described, including ABEs 0.1, 0.2, 1.1, 1.2, 2.1, 2.2, 2.3, 2.4, 2.5, 2.6, 2.7, 2.8, 2.9, 2.10, 2.11, 2.12, 3.1, 3.2, 3.3, 3.4, 3.5, 3.6, 3.7, 3.8, 4.1, 4.2, 4.3, 5.1, 5.2, 5.3, 5.4, 5.5, 5.6, 5.7, 5.8, 5.9, 5.10, 5.11, 5.12, 5.13, 5.14, 6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 7.1, 7.2, 7.3, 7.4, 7.5, 7.6, 7.7, 7.8, 7.9, 7.10, 8, 8a, 8b, 8c, 8d, 8e, 8.8, 8.13, 8.17, 8.20, 9, 9e, ABERA1.0-5.2, 8r, ABEmax, 10, ABExl, ABEx2, ABEx3, ABEx4, and hpABE5.20, e.g., as described in Gaudelli et al., Nature. 2017 Nov 23; 551(7681): 464-471; Koblan et al., Nat Biotechnol. 2018 Oct;36(9):843-846); Tu et al., Mol Ther. 2022 Sep 7;30(9):2933-2941, Richter et al., Nat Biotechnol. 2020 Jul;38(7):883-891; Chen et al., Nat Chem Biol. 2023 Jan; 19(1): 101-110; Xiao et al., Nature Biotechnology volume 42, pagesl442-1453 (2024); Shang et al., bioRxiv 2024.11.23.624961; doi.org / 10.1101 / 2024.11.23.624961; Perrotta et al., bioRxiv 2024.05.17.594556; doi.org / 10.1101 / 2024.05.17.594556; Liao et al., bioRxiv 2025.05.14.653640; doi.org / 10.1101 / 2025.05.14.653640.

[0275] In some embodiments, the heterologous functional domain is an enzyme, domain, or peptide that inhibits or enhances endogenous DNA repair or base excision repair (BER) pathways, e.g., thymine DNA glycosylase (TDG; GenBank Acc Nos. NM_003211.4 (nucleic acid) and NP_003202.3 (protein)) or uracil DNA glycosylase (UDG, also known as uracil N-glycosylase, or UNG; GenBank Acc Nos. NM_003362.3 (nucleic acid) and NP_003353.1 (protein)) or uracil DNA glycosylase inhibitor (UGI) that inhibits UNG mediated excision of uracil to initiate BER (see, e.g., Mol et al., Cell 82, 701-708 (1995); Komor et al., Nature. 2016 May 19;533(7603)); or DNA end-binding proteins such as Gam, which is a protein from the bacteriophage Mu that binds free DNA ends, inhibiting DNA repair enzymes and leading to more precise editing (less unintended base edits). See, e.g., Komor et al., Sci Adv. 2017 Aug 30;3(8):eaao4774. See, e.g., Komor et al., Nature. 2016 May 19;533(7603):420-4; Nishida et al., Science. 2016 Sep 16;353(6305). pii: aaf8729; Rees et al., Nat Commun. 2017 Jun 6;8: 15790; or Kim et al., Nat Biotechnol. 2017 Apr;35(4):371-376) as are known in the art can also be used.

[0276] A number of sequences for domains that catalyze hydroxylation of methylated cytosines in DNA. Exemplary proteins include the Ten-Eleven-Translocation (TET)l-3 family, enzymes that converts 5-methylcytosine (5-mC) to 5- hydroxymethylcytosine (5-hmC) in DNA.

[0277] Sequences for human TET1-3 are known in the art and are shown in the following table:

[0278] * Variant (1) represents the longer transcript and encodes the longer isoform (a). Variant (2) differs in the 5' UTR and in the 3' UTR and coding sequence compared to variant 1. The resulting isoform (b) is shorter and has a distinct C-terminus compared to isoform a.

[0279] In some embodiments, all or part of the full-length sequence of the catalytic domain can be included, e.g., a catalytic module comprising the cysteine-rich extension and the 2OGFeDO domain encoded by 7 highly conserved exons, e.g., the Tetl catalytic domain comprising amino acids 1580-2052, Tet2 comprising amino acids 1290-1905 and Tet3 comprising amino acids 966-1678. See, e.g., Fig. 1 of Iyer et al., Cell Cycle. 2009 Jun 1;8(11): 1698-710. Epub 2009 Jun 27, for an alignment illustrating the key catalytic residues in all three Tet proteins, and the supplementary materials thereof for full length sequences (see, e.g., seq 2c); in some embodiments, the sequence includes amino acids 1418-2136 of Tetl or the corresponding region in Tet2 / 3.

[0280] Other catalytic modules can be from the proteins identified in Iyer et al., 2009.

[0281] In some embodiments, the heterologous functional domain is a biological tether, and comprises all or part of (e.g., DNA binding domain from) the MS2 coat protein, endoribonuclease Csy4, or the lambda N protein. These proteins can be used to recruit RNA molecules containing a specific stem-loop structure to a locale specified by the dCas9 gRNA targeting sequences. For example, a dCas9 variant fused to MS2 coat protein, endoribonuclease Csy4, or lambda N can be used to recruit a long non-coding RNA (IncRNA) such as XIST or HOTAIR; see, e.g., Keryer- Bibens et al., Biol. Cell 100:125-138 (2008), that is linked to the Csy4, MS2 or lambda N binding sequence. Alternatively, the Csy4, MS2 or lambda N protein binding sequence can be linked to another protein, e.g., as described in Keryer-Bibens et al., supra, and the protein can be targeted to the dCas9 variant binding site using the methods and compositions described herein. In some embodiments, the Csy4 is catalytically inactive. In some embodiments, the Cas9 variant, preferably a dCas9 variant, is fused to FokI as described in WO 2014 / 204578.

[0282] In some embodiments, the fusion proteins include a linker between the dCas9 variant and the heterologous functional domains. Linkers that can be used in these fusion proteins (or between fusion proteins in a concatenated structure) can include any sequence that does not interfere with the function of the fusion proteins. In preferred embodiments, the linkers are short, e.g., 2-20 amino acids, and are typically flexible (i.e., comprising amino acids with a high degree of freedom such as glycine, alanine, and serine). In some embodiments, the linker comprises one or more units consisting of GGGS (SEQ ID NO:2) or GGGGS (SEQ ID NO:3), e g., two, three, four, or more repeats of the GGGS (SEQ ID NO:2) or GGGGS (SEQ ID NO:3) unit. Other linker sequences can also be used.

[0283] Also provided herein are nucleic acids encoding the SpCas9 variants, vectors comprising the nucleic acids, optionally operably linked to one or more regulatory domains for expressing the variant proteins, and host cells, e.g., mammalian host cells, comprising the nucleic acids, and optionally expressing the variant proteins.

[0284] The variants described herein can be used for altering the genome of a cell; the methods generally include expressing the variant proteins in the cells, along with a guide RNA having a region complementary to a selected portion of the genome of the cell. Methods for selectively altering the genome of a cell are known in the art, see, e.g., US8,697,359; US2010 / 0076057; US2011 / 0189776; US2011 / 0223638; US2013 / 0130248; WO / 2008 / 108989; WO / 2010 / 054108; WO / 2012 / 164565;

[0285] WO / 2013 / 098244; WO / 2013 / 176772; US20150050699; US20150045546; US20150031134; US20150024500; US20140377868; US20140357530; US20140349400; US20140335620; US20140335063; US20140315985; US20140310830; US20140310828; US20140309487; US20140304853; US20140298547; US20140295556; US20140294773; US20140287938; US20140273234; US20140273232; US20140273231; US20140273230; US20140271987; US20140256046; US20140248702; US20140242702; US20140242700; US20140242699; US20140242664; US20140234972; US20140227787; US20140212869; US20140201857; US20140199767; US20140189896; US20140186958; US20140186919; US20140186843; US20140179770; US20140179006; US20140170753; Makarova et al., "Evolution and classification of the CRISPR-Cas systems" 9(6) Nature Reviews Microbiology 467-477 (1-23) (Jun. 2011); Wiedenheft et al., "RNA-guided genetic silencing systems in bacteria and archaea" 482 Nature 331-338 (Feb. 16, 2012); Gasiunas et al., "Cas9-crRNA ribonucleoprotein complex mediates specific DNA cleavage for adaptive immunity in bacteria" 109(39) Proceedings of the National Academy of Sciences USA E2579-E2586 (Sep. 4, 2012); Jinek et al., "A Programmable Dual- RNA-Guided DNA Endonuclease in Adaptive Bacterial Immunity" 337 Science 816- 821 (Aug. 17, 2012); Carroll, "A CRISPR Approach to Gene Targeting" 20(9) Molecular Therapy 1658-1660 (Sep. 2012); U.S. Appl. No. 61 / 652,086, filed May 25, 2012; Al-Attar et al., Clustered Regularly Interspaced Short Palindromic Repeats (CRISPRs): The Hallmark of an Ingenious Antiviral Defense Mechanism in Prokaryotes, Biol Chem. (2011) vol. 392, Issue 4, pp. 277-289; Hale et al., Essential Features and Rational Design of CRISPR RNAs That Function With the Cas RAMP Module Complex to Cleave RNAs, Molecular Cell, (2012) vol. 45, Issue 3, 292-302.

[0286] Delivery and Expression Systems

[0287] To use the Cas9 variants described herein, it may be desirable to express them from a nucleic acid that encodes them. This can be performed in a variety of ways. For example, the nucleic acid encoding the Cas9 variant can be cloned into an intermediate vector for transformation into prokaryotic or eukaryotic cells for replication and / or expression. Intermediate vectors are typically prokaryote vectors, e.g., plasmids, or shuttle vectors, or insect vectors, for storage or manipulation of the nucleic acid encoding the Cas9 variant for production of the Cas9 variant. The nucleic acid encoding the Cas9 variant can also be cloned into an expression vector, for administration to a plant cell, animal cell, preferably a mammalian cell or a human cell, fungal cell, bacterial cell, or protozoan cell. To obtain expression, a sequence encoding a Cas9 variant is typically subcloned into an expression vector that contains a promoter to direct transcription. Suitable bacterial and eukaryotic promoters are well known in the art and described, e.g., in Sambrook et al., Molecular Cloning, A Laboratory Manual (3d ed. 2001); Kriegler, Gene Transfer and Expression: A Laboratory Manual (1990); and Current Protocols in Molecular Biology (Ausubel et al., eds., 2010). Bacterial expression systems for expressing the engineered protein are available in, e.g., E. coh. Bacillus sp., and Salmonella (Palva et al., 1983, Gene 22:229-235). Kits for such expression systems are commercially available. Eukaryotic expression systems for mammalian cells, yeast, and insect cells are well known in the art and are also commercially available.

[0288] The promoter used to direct expression of a nucleic acid depends on the particular application. For example, a strong constitutive promoter is typically used for expression and purification of fusion proteins. In contrast, when the Cas9 variant is to be administered in vivo for gene regulation, either a constitutive or an inducible promoter can be used, depending on the particular use of the Cas9 variant. In addition, a preferred promoter for administration of the Cas9 variant can be a weak promoter, such as HSV TK or a promoter having similar activity. The promoter can also include elements that are responsive to transactivation, e.g., hypoxia response elements, Gal4 response elements, lac repressor response element, and small molecule control systems such as tetracycline-regulated systems and the RU-486 system (see, e.g., Gossen & Bujard, 1992, Proc. Natl. Acad. Sci. USA, 89:5547; Oligino et al., 1998, Gene Then, 5:491-496; Wang et al., 1997, Gene Then, 4:432-441; Neering et al., 1996, Blood, 88: 1147-55; and Rendahl et al., 1998, Nat. Biotechnol., 16:757-761).

[0289] In addition to the promoter, the expression vector typically contains a transcription unit or expression cassette that contains all the additional elements required for the expression of the nucleic acid in host cells, either prokaryotic or eukaryotic. A typical expression cassette thus contains a promoter operably linked, e.g., to the nucleic acid sequence encoding the Cas9 variant, and any signals required, e.g., for efficient polyadenylation of the transcript, transcriptional termination, ribosome binding sites, or translation termination. Additional elements of the cassette may include, e.g., enhancers, and heterologous spliced intronic signals. The particular expression vector used to transport the genetic information into the cell is selected with regard to the intended use of the Cas9 variant, e.g., expression in plants, animals, bacteria, fungus, protozoa, etc. Standard bacterial expression vectors include plasmids such as pBR322 based plasmids, pSKF, pET23D, and commercially available tag-fusion expression systems such as GST and LacZ.

[0290] Expression vectors containing regulatory elements from eukaryotic viruses are often used in eukaryotic expression vectors, e.g., SV40 vectors, papilloma virus vectors, and vectors derived from Epstein-Barr virus. Other exemplary eukaryotic vectors include pMSG, pAV009 / A+, pMTO10 / A+, pMAMneo-5, baculovirus pDSVE, and any other vector allowing expression of proteins under the direction of the SV40 early promoter, SV40 late promoter, metallothionein promoter, murine mammary tumor virus promoter, Rous sarcoma virus promoter, polyhedrin promoter, or other promoters shown effective for expression in eukaryotic cells.

[0291] The vectors for expressing the Cas9 variants can include RNA Pol III promoters to drive expression of the guide RNAs, e.g., the Hl, U6 or 7SK promoters. These human promoters allow for expression of Cas9 variants in mammalian cells following plasmid transfection.

[0292] Some expression systems have markers for selection of stably transfected cell lines such as thymidine kinase, hygromycin B phosphotransferase, and dihydrofolate reductase. High yield expression systems are also suitable, such as using a baculovirus vector in insect cells, with the gRNA encoding sequence under the direction of the polyhedrin promoter or other strong baculovirus promoters.

[0293] The elements that are typically included in expression vectors also include a replicon that functions in E. coh. a gene encoding antibiotic resistance to permit selection of bacteria that harbor recombinant plasmids, and unique restriction sites in nonessential regions of the plasmid to allow insertion of recombinant sequences.

[0294] Standard transfection methods are used to produce bacterial, mammalian, yeast or insect cell lines that express large quantities of protein, which are then purified using standard techniques (see, e.g., Colley et al., 1989, J. Biol. Chem., 264: 17619-22; Guide to Protein Purification, in Methods in Enzymology, vol. 182 (Deutscher, ed., 1990)). Transformation of eukaryotic and prokaryotic cells are performed according to standard techniques (see, e.g., Morrison, 1977, J. Bacteriol. 132:349-351; Clark-Curtiss & Curtiss, Methods in Enzymology 101 :347-362 (Wu et al., eds, 1983).

[0295] Any of the known procedures for introducing foreign nucleotide sequences into host cells may be used. These include the use of calcium phosphate transfection, polybrene, protoplast fusion, electroporation, nucleofection, liposomes, microinjection, naked DNA, plasmid vectors, viral vectors, both episomal and integrative, and any of the other well-known methods for introducing cloned genomic DNA, cDNA, synthetic DNA or other foreign genetic material into a host cell (see, e.g., Sambrook et al., supra). It is only necessary that the particular genetic engineering procedure used be capable of successfully introducing at least one gene into the host cell capable of expressing the Cas9 variant.

[0296] Alternatively, the methods can include delivering the Cas9 variant protein and guide RNA together, e.g., as a complex. For example, the Cas9 variant and gRNA can be can be overexpressed in a host cell and purified, then complexed with the guide RNA (e.g., in a test tube) to form a ribonucleoprotein (RNP), and delivered to cells. In some embodiments, the variant Cas9 can be expressed in and purified from bacteria through the use of bacterial Cas9 expression plasmids. For example, His- tagged variant Cas9 proteins can be expressed in bacterial cells and then purified using nickel affinity chromatography. The use of RNPs circumvents the necessity of delivering plasmid DNAs encoding the nuclease or the guide, or encoding the nuclease as an mRNA. RNP delivery may also improve specificity, presumably because the half-life of the RNP is shorter and there’s no persistent expression of the nuclease and guide (as you’d get from a plasmid). The RNPs can be delivered to the cells in vivo or in vitro, e.g., using lipid-mediated transfection or electroporation. See, e.g., Liang et al. "Rapid and highly efficient mammalian cell engineering via Cas9 protein transfection." Journal of biotechnology 208 (2015): 44-53; Zuris, John A., et al. "Cationic lipid-mediated delivery of proteins enables efficient protein-based genome editing in vitro and in vivo." Nature biotechnology 33.1 (2015): 73-80; Kim et al. "Highly efficient RNA-guided genome editing in human cells via delivery of purified Cas9 ribonucleoproteins." Genome research 24.6 (2014): 1012-1019.

[0297] The present invention includes the vectors and cells comprising the vectors. Editing disease-relevant mutations

[0298] Also provided herein are methods and compositions for editing disease relevant mutations in living cells, e.g., in cells in vitro or in vivo, using a variant describe herein and an appropriate gRNA. Specific examines include methods of installation of AEZ VH3447R for protection against autosomal dominant Alzheimer’s disease (e.g., in subjects with the Apoe4 genotype; installation of protective variant BAG3 C151R for heart failure; installation of protective variant SLC30A8 R325W for diabetes; installation of protective genetic variant R166H in HDAC7 for prevention of multiple sclerosis; installation of protective genetic variant in 1923 V in IFIH1 for prevention of diabetes; installation of protective genetic variant rs756654226 in ST6GALNAC5 for prevention of dementia; or installation of protective genetic variant rs756654226 in ST6GALNAC5 for prevention of dementia, and others, e.g., as listed in Table E. Any suitable base editor, including the enzymes described herein (e.g., in Tables A or B), and exemplary guide RNAs described herein (or variants thereof), including those shown in Table E, can be used. gRNA variants are at least 80%, 90%, 95%, or 99% identical to the listed gRNA sequence (e.g., having 1, or up to 2 or 3 mismatches from the listed gRNA sequence), optionally with +1, -1, -2, -3, or more nucleotides in length from the 5’ end from those shown in Table E, as long as they retain the ability to direct the base editor to the appropriate target sequence. Some of the listed gRNA spacer sequences include an additional mismatched 5’ G for more efficient transcription from a U6 promoter; where there are two sequences listed, the first one includes the mismatched G. Note that these sequences are shown with ’T’ instead of ‘U’ as encoded in DNA plasmids, but gRNA will have ‘U’s in these spots. In some embodiments, the gRNA are RNA / DNA hybrids.

[0299] TABLE E. Protective and Therapeutic Editing

[0300] As used in this context, to “treat” means to ameliorate at least one symptom of the disease. A treatment comprising administration of a genome editing system (i.e., comprising an SpCas9 variant and / or a gRNA as described herein, e.g., in Table E or elsewhere herein) can result in a reduction in a reduction in frequency or severity of symptoms; a reduction in the rate of progression of symptoms; and / or a return or approach to normal health. Thus, the methods can include administering a therapeutically effective amount of a genome editing system to a subject in need thereof. For example, the activity-enhanced PAM selective BEs described herein can be used for installing RELN H3447R for protection against autosomal dominant Alzheimer’s disease63(Fig. 5a), e.g., in subjects with the Apoe4 genotype, BAG3 C151R for heart failure64(Fig. 5b), and SLC30A8 R325W for diabetes65(Fig. 5c). Alzheimer’s disease inflicts an estimated 24 million people worldwide, stands to potentially increase as much as 4 fold by 205066, and currently lacks effective treatment strategies. Prevention of the most serious effects of the disease could potentially be possible via a recently described genetic variant (H3447R; the ColBos allele) in RELN, found to confer extreme resilience to disease progression in a patient with autosomal dominant Alzheimer’s disease (AD AD)63. Creating this genetic variant requires the use of a PAM variant base editor as there are no compatible NGG PAMs in the proximity of this edit. We tested ABE8e-eSpGA in HEK293T cells and achieved -60% A-to-G conversion of the target base (Fig. 5a).

[0301] The present methods can include delivery of nucleic acids, which can include naked mRNA or DNA, as well as expression constructs comprising sequences encoding an SpCas9 variant as described herein and / or gRNA.

[0302] Expression constructs comprising sequences encoding SpCas9 variant and / or gRNA can include viral vectors, including recombinant retroviruses, adenovirus, adeno-associated virus, lentivirus, and herpes simplex virus- 1, or recombinant bacterial or eukaryotic plasmids. Suitable expression constructs can include: a coding region; a promoter sequence, e.g., a promoter sequence that restricts expression to a selected cell type as described herein; an optional enhancer sequence; untranslated regulatory sequences, e.g., a 5' untranslated region (UTR), a 3' UTR; a polyadenylation site; and / or an insulator sequence. Such sequences are known in the art, and the skilled artisan would be able to select suitable sequences. See, e.g., Current Protocols in Molecular Biology, Ausubel, F.M. et al. (eds.) Greene Publishing Associates, (1989), Sections 9.10-9.14; Vancura (ed.), Transcriptional Regulation: Methods and Protocols (Methods in Molecular Biology (Book 809)) Humana Press; 2012 edition (2011) and other standard laboratory manuals. In some embodiments, the expression construct is capable of directing expression of the SpCas9 variant and / or gRNA nucleic acid preferentially in a selected tissue.

[0303] The constructs can include, e.g., a viral delivery vector, e.g., preferably an adeno-associated virus (AAV) vector that comprises sequences encoding an SpCas9 variant and / or gRNA. Adeno-associated virus is a naturally occurring defective virus that requires another virus, such as an adenovirus or a herpes virus, as a helper virus for efficient replication and a productive life cycle. (For a review see Muzyczka, N., Curr Top Microbiol Immunol, 1992. 158: p. 97-129. AAV vectors efficiently transduce various cell types and can produce long-term expression of transgenes in vivo. AAV vectors have been extensively used for gene augmentation or replacement and have shown therapeutic efficacy in a range of animal models as well as in the clinic; see, e.g., Mingozzi and High, Nat Rev Genet, 2011. 12(5): p. 341-55; Deyle and Russell, Curr Opin Mol Ther, 2009. 11(4): p. 442-7; Asokan et al., Mol Ther, 2012. 20(4): p. 699-708). AAV vectors containing as little as 300 base pairs of AAV can be packaged and can produce recombinant protein expression.

[0304] In some embodiments, the AAV vector can include (or include a sequence encoding) an AAV capsid polypeptide described in WO 2015 / 054653; and a sequence encoding an SpCas9 variant and / or gRNA as described herein. In some embodiments, the AAV capsid polypeptide is an Anc80 polypeptide, e.g., Anc80L27; Anc80L59; Anc80L60; Anc80L62; Anc80L65; Anc80L33; Anc80L36; or Anc80L44. Alternatively, AAV.CPP.21 or AAV.CPP.16 can be used, as described in Yao et al., Nat Biomed Eng. 2022 Nov;6(l l): 1257-1271. AAV vector natural serotypes with known CNS tropism include AAV1, 2, 5, 6, 8, 9, rh8 and rhlO ((Wang and Xiao, Int J Mol Sci. 2025 Feb 28;26(5):2213). Further modifications of capsid structure, including chimeric capsids and incorporation of peptides into the capsid can increase neuronal tropism (e.g. AAV2G9, AAV-D1, AAVPHP.B, AAV-DB-3; Matuszek et al., Mol Ther. 2025 May 7;33(5): 1988-2014). In some embodiments, the AAV incorporates inverted terminal repeats (ITRs), e.g., derived from the AAV2 or AAV9 serotype. It should be noted, however, that numerous modified versions of the AAV2 or AAV9 ITRs are used in the field. Modifications of these sequences are known in the art, or will be evident to skilled artisans, and are thus included in the scope of this disclosure.

[0305] AAV vectors containing as little as 300 base pairs of AAV can be packaged and can produce recombinant protein expression. Protocols for producing recombinant retroviruses and for infecting cells in vitro or in vivo with such viruses are known in the art, e.g., can be found in Ausubel, et al., eds., Current Protocols in Molecular Biology, Greene Publishing Associates, (1989), Sections 9.10-9.14, and other standard laboratory manuals. The use of AAV vectors to deliver constructs for expression in the brain has been described, e.g., in Iwata et al., Sci Rep. 2013;3: 1472; Hester et al., Curr Gene Ther. 2009;9(5):428-33; Doll et al., Gene Therapy 1996; 3(5):437-447; Foley et al., J Control Release. 2014;196:71-8; Liu et al., Metab Brain Dis. 2021 Jan;36(l):45-52; Ling et al., Nat Rev Drug Discov. 2023 Oct;22(10):789- 806; and Huang et al., Science. 2024 May 16; 384(6701): 1220-1227 (preprinted at Huang et al., bioRxiv. 2023 Dec 22:2023.12.20.572615). Thus, in some embodiments, the SpCas9 variant and / or gRNA encoding nucleic acid is present in a vector for gene therapy, such as an AAV vector. In some instances, the AAV vector is selected from the group consisting of AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAVrh8, AAVrhlO, AAV11, and AAV12. AAV1, 2, 5, 6, 8, 9, rh8, and rhlO have been shown to have strong affinity for the nervous system. In some embodiments, AAV2, AAV9, or AAVrhlO are used. AAV9 vectors are highly effective for direct in-brain injections and are being evaluated for treating neurological disorders28.

[0306] A vector as described herein can be a pseudotyped or engineered vector. Pseudotyping provides a mechanism for modulating a vector’ s target cell population. For instance, pseudotyped AAV vectors can be utilized in various methods described herein. Pseudotyped vectors are those that contain the genome of one vector, e.g., the genome of one AAV serotype, in the capsid of a second vector, e.g., a second AAV serotype. Methods of pseudotyping are well known in the art. For instance, a vector may be pseudotyped with envelope glycoproteins derived from Rhabdovirus vesicular stomatitis virus (VSV) serotypes (Indiana and Chandipura strains), rabies virus (e.g., various Evelyn-Rokitnicki-Abelseth ERA strains and challenge virus standard (CVS)), Lyssavirus Mokola virus, a rabies-related virus, vesicular stomatitis virus (VSV), Mokola virus (MV), lymphocytic choriomeningitis virus (LCMV), rabies virus glycoprotein (RV-G), glycoprotein B type (FuG-B), a variant of FuG-B (FuG- B2) or Moloney murine leukemia virus (MuLV). A virus may be pseudotyped for transduction of one or more neurons or groups of cells. In addition, the capsid can be engineered, e.g., altered to include one or more peptides that increase expression in the CNS, see, e.g., Yao et al., Nat Biomed Eng. 2022 Oct 10; Chatteijee et al., Gene Ther. 2022 Jun;29(6):390-397; Meng et al., Mol Ther Methods Clin Dev. 2021 Feb 27;21:28-41; Zhang et al., Biomaterials. 2022 Feb;281 : 121340; Gray, Cell Gene Ther. Insights 5, 1361-1368 (2019); Nonnenmacher et al., Mol. Ther. Methods Clin. Dev. 20, 366-378 (2021). Engineered vectors with capsids that have been altered to change their tropism can also be used. In some embodiments, the vector is enclosed in a AAV-BI-hTFRl capsid (Huang et al., Science. 2024 May 16; 384(6701): 1220-1227, preprinted at Huang et al., bioRxiv. 2023 Dec 22:2023.12.20.572615), or other capsids with affinity for the human transferrin receptor (TFRC), AAV-derived capsids or nanoparticles with affinity for components of the human blood-brain barrier, or otherwise have the capacity for crossing the human blood brain barrier, e.g., AAV.CPP.16 (Yao et al., Nat Biomed Eng. 2022 Nov;6(l l): 1257-1271) or variants of AAV9 (Wang et al. Mol. Ther. -Methods Clin. Dev. 9, 234-246 (2018)); using PB5-3 (Zhang et al., Biomaterials. 2022 Feb:281 : 121340). See also Liu et al., Metab Brain Dis. 2021 Jan;36(l):45-52.

[0307] Without limitation, illustrative examples of pseudotyped or engineered vectors include recombinant AAV2 / 1, AAV2 / 2, AAV2 / 5, AAV2 / 6, AAV2 / 7, AAV2 / 8, AAV9, AAVrhlO, AAV11, AAV12, and AAV-BI-hTFRl serotype or engineered vectors. It is known in the art that such vectors may be engineered to include a transgene encoding a protein or other transcript (e.g., the Cas9 and / or gRNA). For example, the present vectors can include a pseudotyped AAV9 or AAVrhlO viral vector including a nucleic acid as disclosed herein. See Viral Vectors for Gene Therapy: Methods and Protocols, ed. Machida, Humana Press, 2003.

[0308] In some instances, a particular AAV serotype vector may be selected based upon the intended use, e.g., based upon the intended route of administration.

[0309] Various methods for application of AAV vector constructs in gene therapy are known in the art, including methods of modification, purification, and preparation for administration to human subjects (see, e.g., Viral Vectors for Gene Therapy: Methods and Protocols, ed. Machida, Humana Press, 2003). For example, AAV based gene therapy targeted to cells of the CNS has been described (see, e.g., U.S. patents 6,180,613 and 6,503,888). High titer AAV preparations can be produced using techniques known in the art, e.g., as described in U.S. Pat. No. 5,658,776

[0310] Thus provided herein are AAV vectors encoding CRISPR / Cas9 genome editing systems, and the use of such vectors to treat or reduce the risk of a number of diseases, e.g., as listed in Table E. Exemplary AAV vector genomes can include: inverted terminal repeats (ITRs), a gRNA sequence and promoter sequences to drive its expression, and an SpCas9 variant coding sequence and another promoter to drive its expression. Cas9 expression is driven by a promoter known in the art. Expression of the gRNA in the AAV vector is also driven by a promoter known in the art. In some embodiments, a polymerase III promoter, such as a human U6 promoter, Hl promoter, 7sk promoter, or tRNA promoter, is used. Alternatively, the AAV can include ITRs, an SpCas9 variant coding sequence, and a gRNA coding sequence, with a single promoter (e.g., a PolII promoter) to drive expression of both, with a 2 A sequence between the SpCas9 variant and gRNA coding sequences. In this case, the Pol II promoter transcribes the entire sequence into a single mRNA. A single vector can be used to deliver an SpCas9 variant and gRNA; alternatively, a plurality of vectors are used, e.g., wherein one vector is used to deliver the SpCas9 variant, and another vector or vectors is used to deliver a gRNA.

[0311] The vector can also include one or more sequences that promote expression of an SpCas9 variant and / or gRNA, e.g., one or more promoter sequences; enhancer sequences, e.g., 5’ untranslated region (UTR) or a 3’ UTR; a polyadenylation site; and / or insulator sequences, operably linked to the SpCas9 variant and / or gRNA. In some embodiments, the promoter is a brain tissue specific promoter, e.g., a neuronspecific or glia-specific promoter. In certain embodiments, the promoter is a promoter of a gene selected to from: human choline acetyltransferase (ChAT) promoter (Santoscoy et al., Molecular Therapy Methods & Clinical Development, 29:532 - 540; 2023); neuronal nuclei (NeuN), ionized calcium-binding adapter molecule 1 (Iba-1), synapsin I (SYN), calcium / calmodulin-dependent protein kinase II, tubulin alpha I, neuron-specific enolase and platelet-derived growth factor beta chain. In some embodiments, the promoter is a pan-cell type promoter, e.g., EF-lalpha, cytomegalovirus (CMV), CMV immediate enhancer / chicken P-actin hybrid (CAG), chicken P-actin (CBA), beta glucuronidase (GUSB), ubiquitin C (UBC), or Rous sarcoma virus (RSV) promoter. The woodchuck hepatitis virus posttranscriptional response element (WPRE) can also be used.

[0312] In some embodiments, the AAV also has one or more additional mutations that increase delivery to the target tissue, e.g., the CNS, or that reduce off-tissue targeting, e.g., mutations that decrease liver delivery when CNS, heart, or muscle delivery is intended (e.g., as described in Pulicherla et al. (2011) Mol Ther 19:1070- 1078); or the addition of other targeting peptides, e.g., as described in Chen et al. (2008) Nat Med 15: 1215-1218 or Xu et al., (2005) Virology 341 :203-214 or US9102949; US 9585971; and US20170166926. See also Gray and Samulski (2011) “Vector design and considerations for CNS applications,” in Gene Vector Design and Application to Treat Nervous System Disorders ed. Glorioso J., editor. Washington, DC: Society for Neuroscience) 1-9, available at sfn.org / ~ / media / SfN / Documents / Short% 20Courses / 201 l%20Short%20Course%20I / 201 l_SCl_Gray.ashx. Alternatively, non-viral carriers can be used, e.g., encapsulated or associated with in a nanoparticle, e.g., a liposome, exosome, an extracellular vesicle, a polymer, a nanoparticle, a peptide, or a dendrimer. Preferably, the non-viral carrier is a lipid nanoparticle (LNP) (see, e.g., Farsani et al., Heliyon. 2024 Jan l l;10(2):e24606; Tuma et al., Biochemistry. 2023 Sep 20;62(24):3533-3547; Wei et al., Nature Communications 11 :3232 (2020)). In some embodiments, mRNA encoding the SpCas9 variant can be delivered with the gRNA, or ribonucleoprotein complexes comprising an SpCas9 variant protein complexed with the gRNA, can be delivered.

[0313] Also provided herein are compositions comprising the SpCas9 variants and / or gRNA described herein, in a carrier, e.g., a physiologically or pharmaceutically acceptable carrier. As used herein the language “pharmaceutically acceptable carrier” includes saline, solvents, dispersion media, coatings, antibacterial and antifungal agents, isotonic and absorption delaying agents, and the like, compatible with pharmaceutical administration. Supplementary active compounds can also be incorporated into the compositions.

[0314] Pharmaceutical compositions are typically formulated to be compatible with its intended route of administration. Examples of routes of administration include parenteral, e.g., intravenous, intradermal, subcutaneous, oral (e.g., inhalation), transdermal (topical), transmucosal, and rectal administration.

[0315] Methods of formulating suitable pharmaceutical compositions are known in the art, see, e.g., Remington: The Science and Practice of Pharmacy, 21st ed., 2005; and the books in the series Drugs and the Pharmaceutical Sciences: a Series of Textbooks and Monographs (Dekker, NY). For example, solutions or suspensions used for parenteral, intradermal, or subcutaneous application can include the following components: a sterile diluent such as water for injection, saline solution, fixed oils, polyethylene glycols, glycerine, propylene glycol or other synthetic solvents; antibacterial agents such as benzyl alcohol or methyl parabens; antioxidants such as ascorbic acid or sodium bisulfite; chelating agents such as ethylenediaminetetraacetic acid; buffers such as acetates, citrates or phosphates and agents for the adjustment of tonicity such as sodium chloride or dextrose. pH can be adjusted with acids or bases, such as hydrochloric acid or sodium hydroxide. The parenteral preparation can be enclosed in ampoules, disposable syringes or multiple dose vials made of glass or plastic. Pharmaceutical compositions suitable for injectable use can include sterile aqueous solutions (where water soluble) or dispersions and sterile powders for the extemporaneous preparation of sterile injectable solutions or dispersion. For intravenous administration, suitable carriers include physiological saline, bacteriostatic water, Cremophor EL™ (BASF, Parsippany, NJ) or phosphate buffered saline (PBS). In all cases, the composition must be sterile and should be fluid to the extent that easy syringability exists. It should be stable under the conditions of manufacture and storage and must be preserved against the contaminating action of microorganisms such as bacteria and fungi. The carrier can be a solvent or dispersion medium containing, for example, water, ethanol, polyol (for example, glycerol, propylene glycol, and liquid polyetheylene glycol, and the like), and suitable mixtures thereof. The proper fluidity can be maintained, for example, by the use of a coating such as lecithin, by the maintenance of the required particle size in the case of dispersion and by the use of surfactants. Prevention of the action of microorganisms can be achieved by various antibacterial and antifungal agents, for example, parabens, chlorobutanol, phenol, ascorbic acid, thimerosal, and the like. In many cases, it will be preferable to include isotonic agents, for example, sugars, polyalcohols such as mannitol, sorbitol, sodium chloride in the composition. Prolonged absorption of the injectable compositions can be brought about by including in the composition an agent that delays absorption, for example, aluminum monostearate and gelatin.

[0316] Sterile injectable solutions can be prepared by incorporating the active compound in the required amount in an appropriate solvent with one or a combination of ingredients enumerated above, as required, followed by filtered sterilization. Generally, dispersions are prepared by incorporating the active compound into a sterile vehicle, which contains a basic dispersion medium and the required other ingredients from those enumerated above. In the case of sterile powders for the preparation of sterile injectable solutions, the preferred methods of preparation are vacuum drying and freeze-drying, which yield a powder of the active ingredient plus any additional desired ingredient from a previously sterile-filtered solution thereof.

[0317] Oral compositions generally include an inert diluent or an edible carrier. For the purpose of oral therapeutic administration, the active compound can be incorporated with excipients and used in the form of tablets, troches, or capsules, e.g., gelatin capsules. Oral compositions can also be prepared using a fluid carrier for use as a mouthwash. Pharmaceutically compatible binding agents, and / or adjuvant materials can be included as part of the composition. The tablets, pills, capsules, troches and the like can contain any of the following ingredients, or compounds of a similar nature: a binder such as microcrystalline cellulose, gum tragacanth or gelatin; an excipient such as starch or lactose, a disintegrating agent such as alginic acid, Primogel, or com starch; a lubricant such as magnesium stearate or Sterotes; a glidant such as colloidal silicon dioxide; a sweetening agent such as sucrose or saccharin; or a flavoring agent such as peppermint, methyl salicylate, or orange flavoring.

[0318] For administration by inhalation, the compounds can be delivered in the form of an aerosol spray from a pressured container or dispenser that contains a suitable propellant, e.g., a gas such as carbon dioxide, or a nebulizer. Such methods include those described in U.S. Patent No. 6,468,798.

[0319] Systemic administration of a therapeutic compound as described herein can also be by transmucosal or transdermal means. For transmucosal or transdermal administration, penetrants appropriate to the barrier to be permeated are used in the formulation. Such penetrants are generally known in the art, and include, for example, for transmucosal administration, detergents, bile salts, and fusidic acid derivatives. Transmucosal administration can be accomplished through the use of nasal sprays or suppositories. For transdermal administration, the active compounds are formulated into ointments, salves, gels, or creams as generally known in the art.

[0320] Therapeutic compounds that are or include nucleic acids can be administered by any method suitable for administration of nucleic acid agents, such as a DNA vaccine. These methods include gene guns, bio injectors, and skin patches as well as needle-free methods such as the micro-particle DNA vaccine technology disclosed in U.S. Patent No. 6,194,389, and the mammalian transdermal needle-free vaccination with powder-form vaccine as disclosed in U.S. Patent No. 6,168,587. Additionally, intranasal delivery is possible, as described in, inter alia, Hamajima et al., Clin. Immunol. Immunopathol., 88(2), 205-10 (1998). Liposomes (e.g., as described in U.S. Patent No. 6,472,375) and microencapsulation can also be used. Biodegradable targetable microparticle delivery systems can also be used (e.g., as described in U.S. Patent No. 6,471,996).

[0321] In one embodiment, the therapeutic compounds are prepared with carriers that will protect the therapeutic compounds against rapid elimination from the body, such as a controlled release formulation, including implants and microencapsulated delivery systems. Biodegradable, biocompatible polymers can be used, such as ethylene vinyl acetate, polyanhydrides, polyglycolic acid, collagen, polyorthoesters, and polylactic acid. Such formulations can be prepared using standard techniques, or obtained commercially, e.g., from Alza Corporation and Nova Pharmaceuticals, Inc. Liposomal suspensions (including liposomes targeted to selected cells with monoclonal antibodies to cellular antigens) can also be used as pharmaceutically acceptable carriers. These can be prepared according to methods known to those skilled in the art, for example, as described in U.S. Patent No. 4,522,811.

[0322] The pharmaceutical compositions can be included in a container, pack, or dispenser together with instructions for administration.

[0323] Exemplary sequences

[0324] In some embodiments, the sequence of a protein or nucleic acid used in a composition or method described herein is at least 80%, 85%, 90%, 95%, 97%, 98%, or 99% identical to an exemplary sequence set forth herein. To determine the percent identity of two amino acid sequences, or of two nucleic acid sequences, the sequences are aligned for optimal comparison purposes (e.g., gaps can be introduced in one or both of a first and a second amino acid or nucleic acid sequence for optimal alignment and non-homologous sequences can be disregarded for comparison purposes). In a preferred embodiment, the length of a reference sequence aligned for comparison purposes is at least 80% of the length of the reference sequence, and in some embodiments is at least 90% or 100%. The amino acid residues or nucleotides at corresponding amino acid positions or nucleotide positions are then compared. When a position in the first sequence is occupied by the same amino acid residue or nucleotide as the corresponding position in the second sequence, then the molecules are identical at that position (as used herein amino acid or nucleic acid “identity” is equivalent to amino acid or nucleic acid “homology”). The percent identity between the two sequences is a function of the number of identical positions shared by the sequences, taking into account the number of gaps, and the length of each gap, which need to be introduced for optimal alignment of the two sequences.

[0325] The comparison of sequences and determination of percent identity between two sequences can be accomplished using a mathematical algorithm. For example, the percent identity between two amino acid sequences can be determined using the Needleman and Wunsch ((1970) J. Mol. Biol. 48:444-453 ) algorithm which has been incorporated into the GAP program in the GCG software package (available on the world wide web at gcg.com), using the default parameters, e.g., a Blossum 62 scoring matrix with a gap penalty of 12, a gap extend penalty of 4, and a frameshift gap penalty of 5.

[0326] In some embodiments, the sequence of a protein or nucleic acid used in a composition or method described herein has up to 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions or deletions as compared to a sequence set forth herein. In some embodiments, the substitutions are conservative substitutions. Conservative substitutions typically include substitutions within the following groups: glycine, alanine; valine, isoleucine, leucine; aspartic acid, glutamic acid, asparagine, glutamine; serine, threonine; lysine, arginine; and phenylalanine, tyrosine.

[0327] EXAMPLES

[0328] The invention is further described in the following examples, which do not limit the scope of the invention described in the claims.

[0329] Methods

[0330] Plasmids, oligonucleotides strains, and cloning

[0331] Plasmids used in this study are deposited with Addgene and available in Table F. Plasmid sequences were validated using Sanger sequencing and whole plasmid sequencing (Primordium Labs). SpCas9 nuclease human expression plasmid was cloned by inserting SpCas9 open reading frame from pX330 (Addgene 42230) into Notl and Agel sites of JDS246 (Addgene 43861). All human cell experiments use an SpCas9 or derivative variant harboring a C-terminal BP(SV40)NLS-3xFLAG- P2A-EGFP sequence. Standard molecular cloning and isothermal assembly techniques were used to modify all plasmids, including point mutations, addition of fused domains, and addition of P2A-EGFP. CBE constructs were generated by subcloning BE4max (Addgene 112099) into Notl and Agel locations in pCAG-CFP (Addgene 11179). ABEs were created by subcloning or mutating ABEmax or ABE8e (Addgene 112101; 185910). SpCas9 gRNA plasmids were created as previously described for human cell expression with U6 promoter sequences by annealing and ligating duplexed oligonucleotides containing the spacer sequences into a BsmBI- digested BPK1520 (Addgene 6577711). For T7 promoter-driven in vitro transcription of gRNA, plasmids were created by annealing and ligating oligonucleotide duplexes for each spacer into Bsal-digested MSP3485.

[0332] To generate the target plasmid libraries for HT-PAMDA, oligonucleotides encoding the spacer sequence alongside an 8 nt randomized PAM directly 3’ of the spacer were used, similarly to as previously described17. 2 libraries with unique spacer sequences were constructed using 2 separate oligonucleotides. To create a double stranded molecule, Klenow(exo) (NEB) was used, and the resulting DNA was digested with EcoRI before being ligated into the backbone for the library. The backbone was prepared by digestion of pl 1-lacY-wtxl (Addgene 69056) with both EcoRI and Sphl. The resulting plasmids were transformed into electrocompetent XL 1 -Blue E. coli cells. Bacterial cells were then recovered in 9 mL SOB with catabolite expression for 60 minutes at 37°C. Cultures were subsequently grown for 16 hours in 150 mL of Luria-Bertani (LB) medium with 100 pg / mL carbenicillin. Based on the number of transformants, library complexity was estimated at >105 unique PAMs. To prepare the plasmid libraries for use in the in-vitro cleavage reactions, they were linearized with Pvul (NEB).

[0333] Bacterial-based positive selection experiments

[0334] Plasmids containing the Cas9 and gRNA for bacterial expression were prepared as previously described by Kleinstiver et al11,54. To generate variants of SpCas9, plasmids were mutagenized by cloning oligonucleotides encoding randomized codons at positions DI 135, SI 136, G1218, E1219, R1335, and T1337 into the plasmid. Competent E. coli BW25141(XDE3) contained the positive selection plasmid, which included the target site. These cells were transformed with the Cas9 / gRNA plasmids, recovered for 60 minutes in SOB media, and plated on LB plates which were either non-selective (contained chloramphenicol) or selective (contained chloramphenicol + 10 mM arabinose. To perform the selection to identify SpCas9 variants capable of cleaving non-canonical PAMs, the SpCas9 6AA randomized position libraries were electroporated into E.coli BW25141(XDE3) cells that already contained the target site with a new PAM of interest in the positive selection plasmid. Surviving colonies had their SpCas9 plasmid isolated by miniprep. The isolated SpCas9 variants were then re-tested in the selection (individually retransformed into the E.coli cells) with the same PAM to remove false positives. In these follow-ups, generally -300 clones were transformed. Colonies which arose after the second selection had their SpCas9 plasmids sequenced to identify the 6AA residues.

[0335] Profiling the PAM requirements of SpCas9 enzymes

[0336] For HT-PAMDA experiments, transfections were performed in HEK293T cells approximately 16-20 hours after seeding 150,000 cells per well of a 24 well plate. About 600-700 ng of Cas9 variant plasmid was combined with 1.5 ul of TransIT-X2 in a total volume of 50 ul Opti-MEM. Reagent mixtures were gently added to the plate after light mixing and a 15 minute room temperature incubation.

[0337] The HT-PAMDA protocol has been extensively described elsewhere17, and these samples were performed in essentially the same manner. Briefly, lysates were generated containing the Cas9 variant by removal of media and resuspension in 100 pL HT-PAMDA lysis buffer (lx SIGMAFAST protease inhibitor cocktail, EDTA- free, 20 mM Hepes pH 7.5, 100 mM KC1, 5 mM MgCl2, 5% glycerol, 1 mM DTT, and 0.1% Triton X-100) 48 hours post-transfection. An estimate of the Cas9 variant quantity contained in each lysate was determined based on EGFP values measured in a 384 well format (10 pL sample) on a DTX 880 Multimode Plate Reader (Beckman Coulter) with an excitation wavelength of 485 nm and an emission wavelength of 535 nm. EGFP fluorescence was normalized to approximately 150 nM fluorescein (Sigma) based on a fluorescein titration performed in parallel. SpCas9 gRNA were generated in vitro using a Hindlll linearized gRNA T7-transcri ption plasmid template in a 16 hour reaction of the T7 RiboMAX Express Large Scale RNA Production Kit (Promega) at 37°C.

[0338] For the in vitro cleavage reactions, substrates consisting of plasmid libraries with eight nucleotide randomized PAMs on the 3’ end of the Cas target site were linearized. To complex the Cas9 (contained in the normalized lysate) with the gRNA, the following reaction for 3-10 min at 37°C was assembled: 4.375 pL of lysate, 3.5 pL of 2.5 pM IVT gRNA. For cleavage, the randomized PAM library (1.75 pL of 25 nM), was combined with the Cas9 RNP mixture at 37°C. Cleavage buffer final concentration was 10 mM HEPES pH 7.5, 150 mM NaCl, 5 mM MgCh in a total reaction volume of 17.5 pL. Reactions was stopped at timepoints of 1, 8, and 32 min by removing 5 pL aliquots into 5 pL of stop buffer (50 mM EDTA and 2 mg / mL Proteinase K), incubated for 10 min at room temperature and then 98°C for 5 min. All variants were assayed using two distinct PAM libraries with unique spacer sequences.

[0339] To prepare the reactions for sequencing, 3 ng of each reaction was used as the template in a PCR reaction that used primers encoding unique i5 and i7 barcodes. PCR products were then pooled based on the timepoint and purified twice with paramagnetic beads (AMPure XP, Beckman Coulter). Cleaned up pools of PCRs (diluted to 0.125 ng / pL) were then used as template in another PCR with barcoded primers. The combined library was then sequenced using a NextSeq sequencer (Illumina). Analysis of the reads to generate rate constants was conducted as previously described for HT-PAMDA.

[0340] Human cell culture and transfections

[0341] Human HEK293T cells (ATCC) were grown in a culture of Dulbecco’s Modified Eagle Medium (DMEM) supplemented with 10% heat-inactivated FBS and 1% pencillin / streptomycin. Mycoplasma testing was conducted on supernatant media monthly using MycoAlert PLUS (Lonza). Approximately 20-24 hours after seeding 20,000 HEK293T cells per well in a 96 well plate, plasmids were transfected into cells using TransIT-X2 (Minis). For nuclease experiments, 29 ng of nuclease plasmid and 12.5 ng of gRNA expression plasmid were combined with 0.3 pL of TransIT-X2 in a total volume of 15 pL Opti-MEM (Thermo Fisher). For base editing experiments, 70 ng of base editor plasmid and 30 ng of gRNA expression plasmid were combined with 0.72 pL of TransIT-X2 in a total volume of 15 pL Opti-MEM. All reactions were incubated at room temperature for 15 minutes before gentle addition to the HEK293T cells. Experiments were halted and genomic DNA was collected between 48 and 72 hours post-transfection. Genomic DNA was collected by removing the media from the wells and resuspending cells in 100 pL of a quick lysis buffer (20 mM HEPES pH 7.5, 100 mM KC1, 5 mM MgCh, 5% glycerol, 25 mM DTT, 0.1% Triton X-100, and 60 ng / pL Proteinase K (NEB). Lysate was heated at 65°C for 6 minutes, 98°C for 2 minutes, and was stored at -20°C. All transfections were performed in three independent biological replicates, unless otherwise noted. Assessment of nuclease, base editor, and prime editor activities in human cells

[0342] After transfection and harvesting of genomic DNA, the resulting gene editing was measured by targeted amplicon sequencing as previously described12. A 2-step PCR-based protocol created a library of barcoded, Illumina-competent molecules. Data analysis was performed by using the CRISPResso2 package with parameters defined in the section below.

[0343] NGS Data Analysis

[0344] CRISPResso296data analysis pipeline with the following parameters was used in the course of analyzing next-generation sequencing reads:

[0345] Nuclease experiments:

[0346] -min reads to use region 100

[0347] Adenine base editor experiments: -min_reads_to_use_region 100 -w 20 -cleavage_offset -10 - base editor output -conversion nuc from A -conversion nuc to G

[0348] Cytosine base editor experiments:

[0349] -min_reads_to_use_region 100 -w 20 -cleavage_offset -10 - base editor output

[0350] Base editing of patient-derived B cell lines

[0351] B cell lines (BCLs) from a patient (the patient was consented via NIH protocol 05-1-0213) with Chronic Granulomatous Disease with the C445X mutation were established by transformation with Epstein Barr virus as previously described58,97. BCLs were cultured in RPMI + 10% FBS. SpCas9 gRNA directing the base editor to C445X were synthesized (Synthego). Base editors were encoded in mRNAs created by IVT incorporating 100% substitution of the UTP content with pseudoUTP (Cellscript, LLC). mRNAs were post-translationally capped to >95% and poly(A) tailed to >200 A’s. Finally, mRNAs had double-strand RNA removed (Cellscript, LLC). BCLs were edited with ABE mRNA and gRNA by electroporation (MaxCyte Biosystem, Program BCL#3). BCLs were washed and resuspended (~2xl07cells / mL) using electroporation buffer. Samples were prepped by combining -0.25-0.5 xlO6BCLs, mRNA (-0.04 pg / pL final concentration), gRNA (-0.192 pg / pL final concentration), and ScriptGuard RNase inhibitor (1.6 U / pL final concentration; CELLSCRIPT, LLC) in a final volume of 25 pL. After electroporation, cells were moved to a 12-well plate and maintained at 0.5-1.0xl06 / mL for two days. Genomic DNA was then harvested for further analysis using the DNeasy kit (Qiagen).

[0352] GUIDE-seq2 to detect genome-wide off-targets

[0353] HEK 293T cells were prepared for transfection as described above, and the transfection mixture consisted of 29 ng of Cas9 plasmid, 12.5 ng gRNA plasmid, 1 pmol of GUIDE-seq double-strand oligodeoxynucleotide tag (dsODN, oSQT685 / 686, described in Tsai et al.98) and 0.3 pL of TransIT-X2. 72 hours after transfection, a DNAdvance Kit (Beckman Coulter) was used to harvest genomic DNA. Concentration of the genomic DNA was quantified by Qubit (ThermoFisher). As a quality control step to confirm the presence of the tag in the genome and to determine the on-target editing efficiency for the samples to be used in GUTDE-seq2, the on- target dsODN tag integration was measured by NGS sequencing as described above. To analyze the resulting reads, CRISPResso2 was run in non -pooled mode with the target spacer corresponding to each sample, the reference amplicon, and the amplicon edited with the dsODN tag sequence integrated in both the forward or reverse orientations was supplied as an alternate allele for HDR. CRISPResso2 custom parameters: -w 25 -g GUIDE -plot window size 50. To determine tag integration efficiency, the total number of reads containing the tag (sum of forward and reverse reads) was divided by the total number of reads mapped to that reference amplicon.

[0354] The GUIDE-seq2 assay (Lazzarotto & Li et al., in preparation) was performed as previously described but with small modifications. First, Tn5 transposases were complexed with barcoded i5 adapters. The adapter barcodes were 8 bases in length and had 10 nucleotide unique molecular indexes (UMIs). A reaction consisting of 36 pL hyperactive Tn5 (1.85 mg / mL), 15 pL annealed i5 adapter oligo, and 52 pL 2x Tn5 dialysis buffer (100 mM HEPES-KOH pH 7.2, 200 mM NaCl, 0.2 mM EDTA, 2 mM DTT, 0.2% Triton X-100, and 20% glycerol) was incubated for 60 minutes at 24°C. Genomic DNA was tagmented using the barcoded Tn5 complexes in the following reaction: -250 ng genomic DNA, 8 pL Tn5 complex, and 8 pL of 5x TAPS-DMF buffer (50 mM TAPS-NaOH, 25 mM MgCh, 50% dimethylformamide), with a total volume of 40 pL. After 7 minutes at 55°C, 5 pL of Proteinase K solution (50%, diluted in water) was added to stop the tagmentation. Reactions were then incubated at 55°C for 15 minutes before purification using SPRI-guanidine magnetic beads. Tagmentation reactions were analyzed for quality control using TapeStation High Sensitivity D5000 tapes (Agilent).

[0355] Tagmented DNA was then used in a single PCR step to prepare the library. Samples were divided into sense and antisense groups, with the primers corresponding to the possible direction of dsODN integration and the Tn5 adapter sequences. PCRs were purified with SPRI beads before pooling the sense and antisense samples into their own libraries. These two libraries were then further size selected by Pippin Prep (Sage Science) to 250-500 bp. Pure, size selected libraries were then pooled in equal amounts into a 2 nM final library. Library was sequenced with a NextSeq 1000 / 2000 P3 kit (Illumina) with cycle settings of 146, 8, 18, 146. Downsampling of reads was performed for the GUTDE-seq2 samples discussed in Fig- 3 to ensure equal number of reads per sample; downsampling was not performed for the samples discussed in Figs. 4 and 5. The publicly available GUIDE-seq analysis software was used to run the computational pipeline to analyze the data, with minor modifications to account for the GUTDE-seq2 workflow99.

[0356] K562 Nucleofection for allele-specific editing

[0357] K562 cells were nucleofected using the SF Cell Line 4D-NucleofectorTM X Kit (Lonza). K562 cells were sub-cultured at approximately 350,000 cells / mL two days prior to nucleofection. Plasmid DNA was prepared in a 96-well V-bottom plate in a total volume of 2 pL. K562 cells were collected, pelleted at 200 g for 5 minutes, and resuspended in complete nucleofection buffer. Buffer volume was based on the cell count of the K562 harvested, and the resuspended cells were to be concentrated such that 20 pL contained the proper number of cells per nucleofection. 20 pL of cells were then added to the prepared plasmid DNA, briefly mixed by pipetting, and added to the 16-well nucleofection cuvette and placed in the 4D nucleofector and the FF-120 program was executed. After a 10 minute incubation at room temperature, 80 pL of K562 media was added. A volume containing the desired number of cells was then placed in an uncoated, flat-bottom 96-well plate and K562 medium was added for a total volume of 100 pL.

[0358] After 72 hours, cells were transferred to V-bottom 96-well plates and centrifuged at 840 g for 5 minutes to pellet and resuspended in 50 pL of lysis buffer (comprised of the same buffer described above). Cell suspensions were then transferred to 96-well PCR plates, heated for 6 minutes at 65°C , heated for 2 minutes at 98°C, and genomic DNA was stored at -20°C until further use.

[0359] After Illumina sequencing using the same procedures as described above, gene editing efficiencies of the Cas9 nucleases on each allele was quantified. Each sample was analyzed in CRISPResso2 using both alleles in turn as reference sequences. CRISPResso2 output then provided the percentage of remaining, unmodified alleles for each allele in response to Cas9 editing. By comparing to a naive control sample that did not undergo editing, an estimate of the distribution of each allele could be created by subtracting the proportion of each allele that remained unmodified, and a measurement of the percentage of reads attributed to each allele was plotted.

[0360] Example 1.

[0361] To engineer SpCas9 variants capable of targeting alternate PAMs, we pursued a two-stage approach of first utilizing directed evolution in bacteria followed by activity enhancement via rational engineering. For directed evolution, we utilized insight from SpCas9 protein structures to focus a mutagenesis approach on the PAM- interacting domain. In wild-type (WT) SpCas9, base specific protein:DNA contacts are formed between the guanine bases of the NGG PAM and amino acid side chains R1333 and R1335 in the major groove6. Prior efforts to engineer the altered SpCas9 PAM variant enzymes VQR, EQR, and VRER, as well as the PAM relaxed enzyme SpG, coupled with detailed structural analyses, suggested that the 3rdand 4thposition PAM preference can be modified by altering six amino acid (AA) residues DI 135, SI 136, G1218, E1219, R1335, and T133711’16(Fig. la and Figs. 8a, b). The 2ndguanine of the PAM is recognized by R1333 and other nearby AAs that are mechanistically and functionally less well defined, complicating engineering approaches until key PAM modifying interactions are better elucidated.

[0362] Bacterial selection experiments to identify SpCas9 PAM variant enzymes

[0363] To develop new SpCas9 PAM variant enzymes with altered preferences at the 3rdand / or 4thpositions of the PAM, we performed saturation mutagenesis of these 6 AA positions to generate an SpCas9(6AA) library that kept R1333 fixed (reasoning that the resulting PAM variant enzymes harboring R1333 should retain at least moderate editing efficiency against sites with NGN PAMs) (Fig. lb). The SpCas9(6AA) plasmid library also encodes a guide RNA (gRNA) expression cassette. Bacterial selections were performed by transforming the SpCas9(6AA) library into E.coli strains harboring different selection plasmid encoding the ccdB toxin and gRNA target sites with various PAMs. Following the selection, the sequences of hundreds of SpCas9(6AA) variant enzymes were obtained from surviving bacterial colonies to determine what AA substitutions enabled targeting of non-NGG PAMs. Importantly, the bacterial selection requires only that a variant enzyme acquires the ability to target a new PAM to survive, leading to either PAM altered or PAM relaxed enzymes as solutions. Henceforth, SpCas9 PAM variant enzymes are named based on their AA identities at the SpCas9(6AA) positions (e.g. SpCas9-VRAVQL harbors AA substitutions DI 135V, S1136R, G1218A, S1219V, R1335Q, and T1337L).

[0364] Next, to investigate the activity and selectivity of the resultant hundreds of SpCas9(6AA) variant enzymes, we determined their complete PAM preferences via the high-throughput PAM determination assay (HT-PAMDA)17. HT-PAMDA profiles the relative rate of cleavage by SpCas9 variants on a substrate library encoding all 256 NNNN PAMs. Ranking all enzymes based on maximum activity on NGA, NGC, NGG, and NGT PAMs revealed that our selection produced enzymes that were effective at targeting each PAM class (Fig. 9a) Sorting this dataset also permitted identification of more selective enzyme variants that are active on certain PAMs and discriminate against other classes (e.g. active on intended NGC PAMs but selective against unintended NGD PAMs, where D is A, G, or T; Fig. 1c). Our results supported that selective enzymes could be obtained for NGC and NGT PAMs (with similar selectivity compared to WT SpCas9 on NGG PAMs), but that it was more difficult to identify enzymes selective for NGA PAMs (Fig. 1c). The inability to selectively specify adenine in the 3rdposition of the PAM may be partially contributed to the structural difference in nucleobases and how that structure influences the possible ways for a side chain to recognize the identity of the 4thPAM base. Each base, other than adenine, contains a strongly negative oxygen atom outside the aromatic ring structure, and amino acid side chains can interact strongly with this charge to impart specificity, similar to how R1335 interacts with guanine in WT SpCas9. Sorting this dataset for enzyme variants with strong preferences for 4thPAM position bases revealed potential for variants that favor or disfavor either C or G and limited preferences for A or T, likely via position 1337 AA residues interacting with the 4thPAM base (Fig. 9b).

[0365] Following our bacterial selections and biochemical characterization in vitro in human cell lysates, we sought to validate the editing efficiencies of several candidate altered PAM variant enzymes in human cells. The on-target editing efficiencies of engineered SpCas9 variants most selective or active forNGA, NGC, or NGT PAMs (by HT-PAMDA data sorting) were assessed in HEK293T cells across 4 target sites (Fig. Id). We observed that each enzyme could target the intended PAM suggested by its HT-PAMDA characterization, permitting editing of target sites with non-NGG PAMs. These results demonstrate the capability of our initial engineering strategy to discover diverse SpCas9 PAM variant enzymes that function in human cells.

[0366] For a PAM variant enzyme to be maximally selective for a PAM, the enzyme should theoretically accommodate the new PAM while minimizing interactions with others. To examine how PAM selective these initial candidate enzymes were, we performed experiments in HEK293T cells testing SpCas9 enzyme variants with gRNAs targeted to a panel of sites encompassing all 3rdand 4thposition variants of NGNN PAMs. We utilized SpG nuclease as a control, which is a previously engineered SpCas9 enzyme with a relaxed NGNN PAM preference12, as it resulted in the highest levels of editing across NGNN sites when compared to other PAM relaxed enzymes (SpRY, SpCas9-NG, -NRRH, -NRCH, or -NRTH12 14Figs. lOa-c). The PAM variant enzymes exhibited a range of activities, including some being more selective for specific NGA (e g. YSREQM or YSREQQ), NGC (MQKSER or LWKFEG), and NGT PAMs (VRAVQL, IRAVQL, or LRS VQL), and others exhibiting a more relaxed tolerance of PAMs akin to SpG (e.g. CWSHQR, LCRQQR, GWSMQR; Fig. le). These results demonstrate that both PAM selective and PAM relaxed enzymes are discoverable through the bacterial selection experiments, and that the human cell activities of the PAM variant enzymes generally agree with their expected NGNN preferences from HT-PAMDA.

[0367] Based on our initial observations in human cells and via HT-PAMDA, we identified PAM selective SpCas9 enzymes for subsequent development. YSREQM is a selective NGA variant with very limited editing of non-NGA PAMs, relatively uniform editing of NGA PAMs, and it was notably more selective in the HT-PAMDA assay than YSREQQ (Fig. 1c). YSREQQ’s advantage in the human cell experiment is likely attributable to its relatively high editing of the NGAG site, and it showed markedly higher activity on minor non-NGA PAMs in HT-PAMDA. LWKFEG was chosen as our preferred NGC variant due to its very high selectivity for NGCN PAMs observed in HT-PAMDA and moderate editing on NGC PAMs, though it was low efficiency on NGCC. VSREER (previously reported as SpCas9-VRERn) exhibited an extraordinary preference for NGCG PAMs and could be preferable when targeting sites with NGCG PAMs, though this is a narrow usable range (Fig. lOd). We chose VRAVQL for NGT PAMs due to its strong activity on target NGT PAMs in human cells and superior selectivity in HT-PAMDA compared to IRAVQL. We therefore arrived at a cohort of variants with a strong preference for their major PAM. The extent of the chosen PAM variants’ preference for their intended PAM class can be expressed by the average editing in HEK293T cells for the intended PAM (4 gRNA) over the average editing across minor PAMs (12 gRNA) (Fig. If-h) Taken together, these variants can cover the entire NGN PAM space, similar to SpG. Additional PAM variants that may be of use are VRQR18(for high editing on NGAG targets but with more NGNN preference) (Fig. lOe), VRER (for NGCG targets specifically), and MQKSER (for NGNG targets but with some NGCN activity as well Fig. le). For simplicity, we renamed these candidate enzymes based on their PAM preferences: YSREQM, renamed to SpGA, LWKFEG, renamed to SpGC, and VRAVQL, renamed to SpGT.

[0368] Structure-guided engineering to enhance on-target editing efficiencies

[0369] Although the candidate PAM selective enzymes were better able to discriminate between PAMs, their on-target editing efficiencies were suboptimal compared to engineered enzymes like SpG that can efficiently target sites harboring these non-canonical NGA, NGT, or NGC PAMs. Our engineering approach to identify enzymes capable of recognizing new nucleic acid substrates may attenuate on-target DNA binding or enzyme stability, similar to as previously described for other nucleases19. Previous evidence has suggested that on-target editing can be enhanced by the addition of non-specific DNA backbone contacts to attenuated SpCas9 PAM variant enzymes11,12,14’16’18, to improve enzyme stability or substrate interaction20, to improve the catalytic potency of SpCas9 and fidelity-enhanced derivatives21,22, or to enhance the efficiencies of naturally less active Cas9 orthologs23’24. These studies collectively suggest that on-target editing can be improved through rational mutation of SpCas9.

[0370] We therefore sought to explore rational mutations to supplement SpCas9 PAM variant enzymes and improve activity (Fig. 2a). Since SpCas9 exhibits robust activity in a variety of applications when targeting sites with NGG PAMs, there has been less general interest to explore approaches to improve on-target editing activities of SpCas9. However, now that altered PAM variants have arisen with a greater diversity of compatible PAMs, but comparably lower activity, we sought to further explore this space. We hypothesized that in the process of removing base-specific contacts to the PAM and distorting the native PAM interacting domain, the new PAM variants have a lower binding affinity to the target. This may result in lower observed nuclease editing because initiating and stabilizing the nascent R-loop is known to be an important contribution of stable PAM binding, along with the PAM-proximal base pairing ‘seed sequence’ by the gRNA25.

[0371] Our previous efforts to develop altered PAM variants led to the development of SpCas9-VQRn, which can recognize sites encoding NGA or NGNG PAMs (though the relative efficiencies on each of these PAMs varies). We later found that addition of G1218R to SpCas9-VRQR greatly enhanced editing, resulting in the enzyme SpCas9-VRQR18, suggesting that PAM variant enzymes can be modified to improve on-target activity by adding positive side AA chains for putative baseagnostic contacts and engaging with the 4thposition of the PAM, further strengthening the bond between Cas9 and non-NGG PAMs18. Here, SpCas9-VRQR was demonstrated to result in efficient editing of NGAN PAMs, NGNG PAMs, and has limited activity on NGG PAMs (Fig. lOe)

[0372] Later efforts attempted to elucidate domain functions in SpCas9 through deep mutational scanning, which revealed that modest enhancements in activity could be made with the addition of AA substitutions throughout SpCas920. Further, an SpCas9 variant with a relaxed PAM tolerance, SpCas9-NG (NGN PAM preference), utilized AA substitutions LI 111R and A1322R to form base-agnostic contacts with the nontarget strand DNA backbone in order to enhance on-target editing activities14.

[0373] Using available SpCas9 structures, we identified various AA residues whose substitution may form enhanced non-specific DNA contacts throughout the various stages of the SpCas9 reaction mechanism6 14 16 26 28. Given our previously demonstrated improvement of SpCas9-VQR with a G1218R substitution to create SpCas9-VRQR, we selected SpCas9-VQR as an example enzyme to assess putative activity-enhancing mutations along with previously identified activity enhancing mutations as controls11,14’18’20’28. Assessment of these modified VQR enzymes in human cells revealed that engineered enzymes modestly improved editing efficiencies for two target sites that started with lower editing, with lower improvements for two other target sites with higher baseline efficiencies (Fig. Ila). Since WT SpCas9 is naturally a robust genome editor, we tested similar mutations in WT and observed more minor alterations in on-target activity for three high activity sites, and some improvement on one weaker site (Fig. 11b). Testing of additional AA substitutions in SpCas9-VQR and WT identified more engineered variants with improved on-target editing, again following the trend of greater improvement when utilizing weaker gRNAs (Figs. 12a, b).

[0374] Next, we compared the impact of single AA substitutions alone or in the context of two additional activity-enhancing mutations previously shown to be necessary to improve the activity and expand the PAM tolerance of SpCas9-NG (Li l HR and A1322R)14. We first determined the impact of L1111R / A1322R substitutions on the on-target activity of altered PAM variant SpCas9-VQR, -VRER, and -MQKSER on four gRNA with variable baseline editing efficiencies. The editing improvements suggested that these mutations do indeed enhance the editing capability of these PAM variants and can even expand their viable PAMs, as seen in the VRER editing on targets other than NGCG (Fig. 2b). We then assessed single AA substitutions (arising from our previous mutagenesis sampling) without or with the LI 111R / A1322R substitutions using SpCas9-VRER, -VRQR, and -MQKSER across a range of target sites bearing putative active and inactive PAMs (Figs. 13, 14). The single substitutions largely improved editing on the PAMs where activity was expected, whereas the more non-specific LI 111R / A1322R substitutions generally improved efficiency across non-canonical PAMs as well (consistent with their important role in SpCas9-NG). Taken together, these data suggest that editing on low efficiency target sites can be improved via various rational AA substitutions to create engineered activity enhanced enzymes, that improving editing on already efficacious target sites is challenging, and that certain AA substitutions may enhance efficiency but at the expense of PAM selectivity. To explore the compatibility of enhancing mutations with our newly engineered PAM-selective enzymes SpGA, SpGC, and SpGT, we generated enhanced derivatives and assessed their activities in human cells. We began with SpGA because it was prone to weaker on-target editing in our initial experiments. The addition of single activity enhancing mutations substantially improved on-target editing, resulting in several more active SpGA enzymes without substantial loss in PAM specificity (Fig. 2c and Fig. 15a). The SpGA derivative encoding both LI 111R and A1322R mutations resulted in the highest levels of on-target editing for sites with NGA PAMs, but also came at the expense of elevated editing at sites bearing non-canonical PAMs (Fig. 2c).

[0375] Overall, the SpGA + LI 111R / A1322R variant was still highly selective for editing on NGA PAMs. We explored a similar approach to enhance the activity of SpGC with a particular focus on improving editing against sites with a cytosine in the 4thposition of the PAM. We tested 30 derivative SpGC variants containing some of the most consequential activity-enhancing substitutions evaluated thus far and observed improved editing efficiency on three NGC target sites while remaining biased against the non-NGC targets (Fig. 15b-c). Strikingly, activity-enhanced SpGC enzymes substantially improved editing of the target with an NGCC PAM, with one of the most potent combinations (SpGC + A1285K + LI 111R) improving the mean editing on an NGCC PAM target from 7.8% to 47.1% without substantially compromising PAM selectivity (Fig. 2d and Figs. 15b-d). Finally, although SpGT exhibited quite robust editing, we generated activity enhanced variants and assessed their on-target editing efficiencies. Addition of single activity-enhancing mutations to SpGT resulted in only marginal improvement of editing on NGT sites (Fig. 16a). More variants, this time in combination with the LI 111R / A1322R substitutions, demonstrated improved editing on NGT target sites, but also predictably raised the editing on minor PAMs substantially (Fig. 16b). These data suggest that meaningful improvement of SpGT on NGT PAMs - while limiting editing on targets with NGV PAMs - may require a deeper mutational screening.

[0376] Based on this extensive validation of this activity-enhancing strategy, we sought to establish the optimal combination of substitutions to add to these scaffolds (SpGA, SpGC, and SpGT) that enhance editing activity while retaining their PAM selectivity. To ensure deep coverage of variant possibilities, we tested up to double and triple combinations of the activity-enhancing mutations in our altered PAM variants and performed HT-PAMDA again to ascertain the complete PAM preference and relative activity of every variant on every possible NNNN PAM. Sorting of the PAM profiles by HT-PAMDA revealed many potential optimized PAM variants that have improved editing efficiencies without unwanted broadening of PAM specificity (Fig. 2e). Examination of 4thbase PAM position preference revealed little sequence preference by many of these variants, although there was some bias for SpGA- or SpGC-derived enzymes to prefer NGNG PAMs and SpGT-derived enzymes to favor NGNH (Fig. 17a).

[0377] We identified five enhanced SpGA, SpGC, and SpGT enzyme variants that maximize editing efficiency and retain PAM selectivity. For evaluation of their editing activity in human cells, we tested these variants using 16 gRNAs covering NGNN PAMs and compared to SpG (Fig. 2f). Notably, all enzyme variants that encoded activity-enhancing mutations led to higher genome editing efficiencies versus the original SpGA, SpGC, and SpGT enzymes. Activity-enhanced nucleases exhibited high levels of on-target editing with gRNAs within their predicted PAM class. They also retained their preference for the target PAM class with a modest reduction in PAM selectivity (Fig. 2g and Fig. 17b), consistent with the understanding that PAM recognition is a critical initial step in target site recognition and past results engineering PAM specificity. From among these validated options, we selected the following as our preferred variants due to their balance of high editing efficiency and PAM selectivity: SpGA + LI 111R + A1322R; SpGC + A1285K + LI 111R; SpGT + A1285K + G366R. These enhanced variants are noted throughout as eSpGA, eSpGC, and eSpGT, respectively. Additional promising variants of interest that we noted for more limited use cases include: eVRQR, VRQR + S55R; eVRER, VRER + A1285K; eMQKSER, MQKSER + A61R + LI 111R + A1322R; eSpGT.2, SpGT + A61R + LI 111R + A1322R. To demonstrate that our collection of three engineered variants and wild type SpCas9 could replace the use of a relaxed variant like SpG, they were directly compared in HEK293T cells on a range of NGNN gRNA targets. On nearly all PAMs, the activity-enhanced, selective variants resulted in editing efficiencies equivalent or better than SpG, suggesting that they might offer equivalently active but more specific alternatives to PAM relaxed enzymes (Fig. 2h). Specificity comparisons of PAM altered and PAM relaxed enzymes

[0378] One caveat of engineered Cas enzymes with expanded PAM tolerances, like SpG or SpRY, is that they can scan more of the genome and encounter additional putative off-target sites29. To investigate whether our engineered SpGA, SpGC, and SpGT enzymes (collectively SpGH enzymes) and their enhanced derivatives could minimize off-target editing by virtue of being restricted to a smaller fraction of the genome, we compared their specificities to SpG using a cell-based, unbiased, genome-wide off-target assay (GUIDE-seq2; Lazzarotto & Li et al, in preparation). Each nuclease was assessed using four different gRNAs, and the SpGH and particularly the eSpGH enzymes were generally capable of similar levels of on-target editing compared to SpG (Fig. 3a-c and Fig. 18). The SpGH and eSpGH enzymes also reduced the number of off-target sites detected by GLTDE-seq2 with a higher fraction of GUIDE-seq2 reads attributable to the on-target site when compared to SpG (Fig. 3d-f). These results support the safety advantage of Cas9 variants that have a major PAM preference over a relaxed PAM variant like SpG.

[0379] Beyond minimizing genome-wide off-targets, improved specificity enzymes resulting from stringent PAM selectivity could enable allele-specific editing approaches30,31. To achieve allele-specific editing, Cas nucleases must be capable of distinguishing between wild-type and mutant alleles that typically differ by only a single nucleotide (Fig. 3g). Allele-specific editing strategies typically attempt to position the nucleotide difference in the PAM-proximal region of the target site or within the PAM itself, where the latter may be a more generalizable approach given that PAM selectivity should reject mismatched alleles without regard to spacer sequence. Therefore, we leveraged the phased haplotype resolved genome of K562 cells to explore allele-specific editing with our altered PAM variants by intentionally targeting heterozygous SNPs positioned in a compatible PAM32. Here, the alleles containing PAMs corresponding to the preference of the PAM variant (i.e. NGT for eSpGT) will be edited, while the alleles harboring a minor PAM will be spared from targeting. We achieved successful allele-specific editing in multiple samples utilizing different altered PAM variants and with minimal editing of the non-targeted allele (Fig. 3h, i) While some of these results indicate the feasibility of tightly controlled allele-specific editing, many of the targets experienced some degree of leaky editing of the alleles containing the minor PAM (Fig. 3i). This is not necessarily unexpected given prior data that demonstrated that our altered PAM variants - while shown to have a very enriched preference for their major PAM - are not entirely incapable of targeting on minor PAMs and frequently resulted in low overall editing efficiency. If allele-specific editing is desired, these variants offer a viable option, but the efficacy may exhibit some site-specific variation.

[0380] Base editing with activity-enhanced PAM selective enzymes

[0381] Base editors (BEs), comprised of nickase Cas9 enzymes fused to engineered deaminases, are a genome editing technology that frequently requires engineered Cas9 PAM variant enzymes to appropriately position the deaminase edit window33,34. Although PAM relaxed enzymes have been shown to support flexible and efficacious on-target base editing, there remains concerns about off-target base editing resulting from expanded targeting ranges. To investigate the compatibility of our altered PAM variant enzymes as base editors in human cells, we constructed and tested both adenine and cytosine base editors (ABEs and CBEs, respectively). With the eSpGH enzymes we observed comparable on-target editing to SpG-BEs, though there were some target-specific differences depending on the ABE (ABE8e35or ABE8.20m36; Fig. 4a and Fig. 19a, respectively) or CBE construct used (BE4max37or TadCBEd38; Fig. 4b and Fig. 20, respectively). No major differences in edit window were observed between eSpGH- and SpG-BEs on target PAMs. Unexpectedly, the PAM preference of all tested SpCas9 base editors, including WT SpCas9, appeared far more relaxed than the established PAM preference as nucleases in these experiments (Figs. 21a, b, c, d).

[0382] We hypothesized that the observed base editing activity on minor PAMs was the result of brief, unstable R-loops that may form during non-ideal PAM interactions by the SpCas9 protein39. In the nuclease context, these would largely be unproductive interactions, as the R-loop may collapse prior to the conformational shift required for catalytic activity25, supported by our prior data clearly demonstrating the PAM preferences for these SpCas9 variant nucleases. However, a base editor can enact its activity on any ssDNA substrate, so an unstable R-loop may serve as a brief substrate for an adenine base positioned close to the deaminase40. This hypothesis is supported by an analysis of the editing efficiency that occurs at each base in the editing window when the PAM variants are targeted to non-preferred PAMs. We observed a slight narrowing of the edit window generally corresponding to the PAM preference of the ABE8e construct (Fig. 22). Stable, perfectly matched spacer sequences typically result in efficient editing of approximately 3 - 10 for ABE8e, with the most ideal positions being 4 - 8. This data suggests that the editing efficiency for bases outside of the peak edit window positions drops more when the SpCas9 PAM variant is used on a target site harboring a minor PAM. This data suggests that the base editor may not form a stable R-loop with the same frequency as it is for the major PAM targets, leading to less efficient editing on bases present outside the peak editing window, but the deaminase is still presented with a single strand DNA substrate in the ideal spacer positions.

[0383] To further explore this observation that SpCas9 enzymes were apparently more tolerant of minor PAMs as base editors, we performed a titration of PAM variant ABE8e constructs to determine the effects on editing in human cells. In our titration, we aimed to test if these prior results could be partially explained by an accumulation of A-to-G edits made over time by a high dose of ABE8e constructs. We hypothesized that at a lower effective concentration of ABE8e in the wells, the observed base editing would drop on unfavorable PAMs, but be retained on favorable PAMs. Our titration surveyed a 25x dilution of the amount of transfected ABE8e plasmid on 4 gRNA of NGA, NGC, NGG, and NGT PAMs. In general, the result suggest that at lower amounts of ABE8e delivered, the A-to-G editing still occurs on major and minor PAMs, albeit at a lower level (Fig. 23a). However, by comparing the relative reduction in editing efficiency between the high and low doses of ABE8e, we can observe that editing on target sites with minor PAMs drops more drastically than the editing efficiency of the targets with preferred PAMs, including in WT SpCas9 (Fig. 23b). The results suggest that, while we do still observe base editing in the low dose condition on minor PAMs, the target PAMs are preferentially edited generally in line with the SpCas9 variant’s PAM preference. To our knowledge, this observation of base editing outside of the PAM preferences of the SpCas9 enzyme has not yet been reported. Further exploration may be required to determine the extensibility of these results to different experimental conditions beyond transient plasmid transfection of cells in culture. Base editing of pathogenic SNPs

[0384] We investigated the potential of our activity enhanced and PAM-altered BEs to correct genetic mutations with therapeutic relevance and that require precise PAM positioning. Sickle cell disease (SCD) is caused by a A-to-T substitution in the HBB gene resulting in an E7V AA substitution41. Various genome editing strategies have been explored to treat SCD via the use of nucleases or BEs to up-regulate expression of the gamma globin genes (HBG1 / 2)42, prime editing to revert the mutation back to WT43, or via ABEs to generate a benign Makassar HBB-E7E allele44,45. In fact, the first CRISPR therapy to gain approval is the ex vivo activation of fetal hemoglobin in HSCs for the treatment of sickle cell anemia46,47. Ex vivo strategies, while transformative, are taxing on the patient and difficult to scale. Installation of the E7A allele by base editing requires SpCas9 PAM variant enzymes due to a lack of target sites that encode an NGG PAM at an appropriate distance from the edit. Previous efforts to generate the Makassar H B-E7E allele utilized a relaxed PAM variant enzyme, ABE8e-SpCas9-NRCH, which when paired with the A7 gRNA (CAC PAM) could effectively generate the intended edit but also resulted in many off-target edits due to expanded PAM compatibility of -NRCH44. To investigate whether our PAM altered enzymes could achieve similarly potent E7A editing to the PAM relaxed SpCas9-NRCH, we first generated a homozygous endogenous HEK293T cell line harboring the HBB-E N mutation. Using the HBB-E1N cell line, we compared ABE8e-SpCas9-NRCH with gRNA A7 (CAC PAM) to ABE8e-eSpGC and ABE8e- SpG using gRNA A9 (TGC PAM). We observed similarly efficient on-target editing to generate the E7A edit and comparable levels of translationally silent bystander edits (Fig. 4c). Comparison of the genome-wide specificities of these nucleases revealed that eSpGC with gRNA A9, due to its more restricted PAM preference, led to reduced off-target editing and resulted in a higher proportion of GUIDE-seq2 reads at the on-target site compared to SpCas9-NRCH with gRNA A7 or SpG with A9 (Fig. 4d and Fig. 24a).

[0385] Diseases of the liver are well-suited for treatment via genome editing given simplified delivery of nucleic acids by lipid nanoparticles48. A second genetic mutation that is inaccessible for base editing when using WT SpCas9 is the E366K (formerly E342K) AA substitution in the SERPINA1 gene that causes alpha-1 antitrypsin deficiency (AATD)49. AATD leads to non-functional AAT protein produced in the liver, which canonically plays a key role in regulating the activity of neutrophil elastase in the lung50. Without functional AAT, neutrophil elastase breaks down lung tissue, progressively damaging the alveoli and potentially leading to emphysema. Patients currently may rely on protein replacement therapy to raise serum AAT level, which requires a weekly IV infusion, or advanced patients may even need a lung transplant51. Mutations in the SERPINA1 gene can result in polymerization of AAT in liver. This harms the patient through AAT buildup in the liver and reduced secretion to the lung, where its activity is needed. The presence of these polymerized AAT proteins also provides further motivation to pursue a gene editing therapy to stop the accumulation of AAT, rather than a traditional gene therapy using a transgene52. AAT is primarily produced by the liver and transported to the lung, so rescue of the causal mutation in the liver may provide sufficient AAT activity to cure the disease. The most common mutation in SERPINA1 that causes AATD is the E366K AA substitution50(formerly known as E342K), which is caused by a C-to-T mutation and can thus be directly corrected using an ABE (by editing the opposite strand). There are no NGG PAMs appropriately positioned to facilitate a corrective A-to-G edit using ABEs (Fig. 4e). There is, however, an optimally placed NGC PAM that permits the use of an A7 gRNA, as previously described53.

[0386] Werder et al.53, previously assessed the potential of an SpCas9 PAM variant enzyme to correct the common PI*Z AATD mutation. The SpCas9 PAM variant enzyme they utilized was derived from our MQKSER enzyme54, but additionally encoded E1219F, A1322R, and D1332A AA substitutions derived from SpCas9-NG14(where the enzyme collectively encodes DI 135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R substitutions; MQKFRAER55). To minimize variables in our comparison of SpCas9 PAM variant enzymes, we elected to test our MQKSER enzyme and MQKFER only encoding the isolated SpCas9(6AA) substitutions, both in the context of a canonical ABE8.20m architecture, rather than the previously described MQKFRAER enzyme in the context of modified ABE8.20 architectures with variant TadA domains (ngcABEvar5 with TadA8.20 encoding 176, H123Y, D147Y, Q154; ngcABEvar9 with TadA8.20m encoding V82T, H123Y, D147T, Q154S), as previously described53.

[0387] We generated a poly-clonal HEK293T cell line harboring several lentivirally integrated copies of the PI*Z E366K allele, and then compared the on-target efficacy using gRNA A7 and ABE8.20m constructs encoding the SpCas9 PAM variants enzymes eSpGC, eMQKSER, and an MQKFER construct53. Amongst the 3 ABEs, we observed comparably high levels of on-target correction of E366K approaching 60% (Fig. 4e), and similar bystander editing and composition of edited alleles (Fig. 4e,f). While there are few pure alleles of only editing the target A7 base, alleles with editing of A7 and A5 produce a protein variant that has been reported to retain most of its activity56.

[0388] Next, we sought to validate our PAM selective enzymes in primary patient- derived cells. X-linked chronic granulomatous disease (X-CGD) is a rare inborn error of immunity caused by mutations in the CYBB gene that render patients susceptible to infections and other complications, and has thus been a target for various genome editing approaches57 59. We derived an X-CGD B-lymphoblastoid cell line (B-LCLs) from a patient harboring a CYBB C445X mutation. Assessment of ABE8e versions of SpG, eSpGT.2 (SpGT + A61R + LI 111R + A1322R), and a high-fidelity variant of eSpGT.2 (eSpGT.2-HFl, SpGT + A61R + L1111R + A1322R + HF1 mutations (N497A + R661 A + Q695A + Q926A)) delivered by mRNA with gRNA A9 (AGT PAM) in the C445X B-LCLs all yielded high levels (<60%) of A-to-G editing (eliminating the nonsense codon and resulting in a missense C445R; Fig. 4g). Although each ABE also resulted in bystander L444P editing, the potentially pathogenic bystander was dramatically reduced when using the eSpGT.2-HFl enzyme. Comparison of the off-target editing profiles of the 3 nucleases with CYBB C445X gRNA A9 via GUTDE-seq2 revealed superior on-target specificity with eSpGT.2 over SpG, and even further minimized off-target editing when using eSpGT ,2-HFl (Fig. 4h and Fig. 24b)

[0389] Installation of protective genetic variants using activity-enhanced PAM selective BEs

[0390] The pre-symptomatic prophylactic installation of natural genetic variants associated with protection against disease progression in afflicted or high-risk individuals is an emerging therapeutic strategy60,61. Such an approach has already borne out in clinical trials via base editing for familial hypercholesteremia62. Like other base editing approaches, the ability to efficiently and precisely install protective genetic variants is dependent on engineered PAM variant enzymes that can carefully position the BE deaminase domain over the target base. Thus, we aimed to install several previously identified natural genetic variants shown to be protective against serious disease.

[0391] To explore the use of our activity-enhanced PAM selective BEs, we first focused on installing RELN H3447R for protection against autosomal dominant Alzheimer’s disease63(Fig. 5a), BAG3 C151R for heart failure64(Fig. 5b), and SLC30A8 R325W for diabetes65(Fig. 5c). Alzheimer’s disease inflicts an estimated 24 million people worldwide, stands to potentially increase as much as 4 fold by 205066, and currently lacks effective treatment strategies. Prevention of the most serious effects of the disease could potentially be possible via a recently described genetic variant (H3447R; the ColBos allele) in RELN, found to confer extreme resilience to disease progression in a patient with autosomal dominant Alzheimer’s disease (AD AD)63. Creating this genetic variant requires the use of a PAM variant base editor as there are no compatible NGG PAMs in the proximity of this edit. We tested ABE8e-eSpGA in HEK293T cells and achieved -60% A-to-G conversion of the target base (Fig. 5a). Additionally, by using an ABE comprised of TadA8.20m rather than TadA8e, we attained similar on-target editing while nearly completely ablating missense bystander edits causing RELN N3449C and N3450D mutations, which result from the expanded edit window of TadA8e. This result highlights a benefit to precision PAM positioning. Off-target analysis via GUTDE-seq2 utilizing the same nucleases and gRNAs for this target site revealed that both eSpGA and SpG were highly specific for the on-target with no detectable off-targets, and that SpRY - with its extremely relaxed PAM - had increased off-targets (Fig. 5b and Fig. 24c).

[0392] We next sought to demonstrate the benefit of being able to test multiple editing strategies against target sites bearing different PAMs, towards achieving maximally efficient base editing even when there is a conveniently located NGG PAM for use with WT Cas9. The naturally occurring BAG3 C151R genetic variant (rs2234962) was identified to confer protection against heart failure64. Heart disease is a leading cause of death and dilated cardiomyopathy is one of the main causes of nonischemic heart failure67,68. A previous study explored base editing to prevent ischemic heart failure through the editing of the PCSK9 gene60, but fewer efforts have thus far been reported on non-ischemic heart failure. The missense BA G3 C151R variant results in altered BAG3 binding to maintenance proteins, leading to dose-dependent increases in proteotoxic stress in cardiomyocytes64. As such, we aimed to devise a base editing strategy to achieve the most efficient editing. This BAG3 variant is relatively common (-20% in European and South- Asian populations) and has not yet been described to have any negative health impacts, making it a promising candidate to evaluate as a gene editing intervention in non-ischemic heart failure64. Efficient editing of the target base with WT SpCas9 would necessitate the use of ABE8e’s expanded edit window via an A3 gRNA. However, utilization of enzymes compatible with NGT or NGA PAMs (and gRNAs A5 or A7, respectively; Fig. 5c) would permit the use of TadA domains with narrower edit windows like TadA8.20m and avoid the hyperactivity of ABE8e to further increase the safety profile of the editor. We sought to install this edit in HEK293T cells by assessing several ABEs compatible with the available PAMs. With WT SpCas9 ABEs and gRNA A3, we observed -50% base editing for ABE8e and -30% with ABE8.20m (Fig. 5c). Use of ABE8.20m-eSpGT and the A5 gRNA to target the NGT PAM in -50% editing of the target base.

[0393] Notably, the A7 NGA PAM target site encodes a 4-nucleotide NGAG PAM, which is very efficiently edited by SpCas9-eVRQR. Use of ABE8.20m-SpCas9- eVRQR resulted in the highest editing of this target at nearly 80% target base conversion (Fig. 5c). Off-target analysis via GUIDE-seq2 off-target analysis indicated that WT SpCas9, eSpGT, and eVRQR all had undetectable or very few off-targets (Fig. 5d and Fig. 24c).

[0394] More optimal placement of the base editing window may increase editing efficiency, even if shifted only by a single base pair. The SLC30A8 R325W genetic variant (rs 13266634) has been shown to increase B cell survival and may be protective against diabetes for individuals possessing the T allele65,69. To install the SLC30A8 R325W edit, there are two target sites that are shifted by a single nucleotide and that encode either NGG or NGT PAMs (Fig. 5e). Use of BE4max-eSpGT and the A7 gRNA to target the NGT PAM resulted in -72% on-target editing, whereas BE4max-WT and the A8 gRNA to target the NGG PAM led to -60% on-target editing (a -12% increase in edit efficiency; Fig. 5e). These results demonstrate that PAM variant enzymes can outperform WT SpCas9 when the edit window is more optimal. In addition to the increase in on-target editing, we detected substantially reduced off-target edits with eSpGT compared to WT SpCas9 when assessed by GUIDE-seq2 (Fig. 5f and Fig. 24d). The ability to precisely position BE target sites with PAM variant enzymes has obvious benefits for targeting sequences that lack NGG PAMs, but PAM variants should also be considered in cases where WT SpCas9 is capable of installing the edit at or near the same efficiency. For example, a gRNA targeting a site encoding an NGG PAM could be prone to a more promiscuous off-target profile; utilization of an alternate nearby gRNA compatible with a PAM variant may minimize off-target editing. We sought to install the CFB R32Q variant (rs641153), which has been linked to protection against age-related macular degeneration70. Complement factor B is an important part of the innate immune system, and the resulting protein of the R32Q variant has been shown to have reduced angiogenic activity71. Complete knockout of CFB has been shown to be detrimental, making the installation of this protective variant an attractive option over other interventions that may simply reduce or knockout function72. Most CFB protein is produced in the liver, and there is some evidence that local eye production is also involved in driving disease; however, both the liver and eye are tissues amenable to gene therapy and gene editing73,74. The R32Q variant can be installed by a C-to-T edit on the non-coding DNA strand (Fig. 5g). By using TadCBEd-eSpGA and the C6 gRNA to target the NGA PAM, or TadCBEd-WT and the C7 gRNA to target NGG PAM, we achieved similarly robust on-target editing, whereas TadCBEd-SpG and the C6 gRNA resulted in somewhat lower on- target editing (Fig. 5g). Nomination of potential off-target sites in silico using Cas- OFFinder75on these closely related guides suggests that the NGG target site has substantially higher off-target potential compared to the NGA target site (Fig. 5h).

[0395] Possible protective gene editing interventions are particularly promising for widespread diseases which cannot be easily controlled by lifestyle or existing treatments. Cardiovascular disease (CVD) is a leading cause of death, and the LPA gene is a well-validated risk factor towards the development of CVD76,77. Genetics are the primary factor determining the amount of Lp(a) in an individual, and current interventions have not been successful at lowering Lp(a). Approximately 1 in 5 people in the United States possess plasma Lp(a) levels that put them at an elevated risk for CVD78. Other strategies to prevent CVD driven by other risk factors, such as LDL cholesterol lowering therapies like PCSK9 inhibitors (or base editing to reduce PCSK9 levels) have shown promise79. Given that most of Lp(a) production takes place in the liver, gene editing has been proposed as a potential one-shot therapeutic to lower Lp(a)-related risk of CVD. Prior efforts to explore gene editing of Lp(a) have been reported, but nuclease-based methods led to detection of large genome rearrangements78. Base editing carries a lower inherent risk of double-strand breaks80, which reduces the risks of detrimental off-target effects, particularly chromosomal rearrangements81. To reduce Lp(a) levels and avoid intentional introduction of a double-strand break, we sought to install a protective base edit into the LPA gene. This protective variant (rs41267813) was identified among UK Biobank participants to be associated with lower Lp(a) concentrations82. Notably, this protective effect was observed even among individuals that also possessed known Lp(a)-elevating alleles elsewhere in the gene, with a 13x lowering in Lp(a) concentration among those possessing both a risk (rsl0455872) and protective allele compared to the absence of the protective allele82. Using eSpGA-TadCBEd, we achieved -40% editing of the target base using the only NGN PAM available in the surrounding genomic context (Fig. 5i). Interestingly, this BE also resulted in substantial editing of the CIO bystander, which produces an early termination codon. While in some cases the introduction of a stop codon would be detrimental, Lp(a) reduction is the expected outcome of this and other proposed therapies83. Further study would be needed to determine the functional consequences of the resulting alleles, but the availability of PAM variant BEs capable of targeting this locus enables the potential installation of this base edit.

[0396] Beyond coding variants found to be naturally protective against disease, there are also genetic variants that result in loss of function or reduced gene expression via splice site modifications. Although several of such variants have inspired current pharmacologic interventions, some may also be suitable candidates for exploring gene editing interventions. For example, an IL33 (rsl46597587-C) variant was recently identified in a study of Icelandic populations to result in protection against severe asthma84. Asthma is a common affliction that can generally be managed by standard- of-care inhaled corticosteroids and / or long acting beta agonist therapies, but 1-2% of the 300 million asthma patients worldwide still suffer from uncontrolled or persistent symptoms85. Additional therapeutic modalities such as antibodies are emerging, signaling a need for alternative treatment methods. Delivery via viral or non-viral vectors would be facilitated by the fact that most IL33 expression occurs in the lung stromal cells, making the target cells accessible to the vector upon administration. Recent developments in lipid nanoparticles have provided hope that lung delivery will soon be feasible86. The IL33 variant is a G>C substitution that disrupts a splice acceptor (Fig. 5j), leading to exon skipping. Although C-to-G base editors are a recently emerging technology87, we sought to disrupt this splice acceptor via the neighboring adenine base using ABEs. Positioning an ABE to disrupt the IL33 splice acceptor requires the use of a PAM variant, where the A8 gRNA target site would utilize an NGA PAM. We compared ABE8e versions of SpG, eSpGA, and eVRQR in HEK293T cells, and achieved nearly 50% editing at the target base without bystander editing (Fig. 5j). Notably, ABE8e-eSpGA was more active than SpG or eVRQR, where the latter two enzymes achieved -30% editing. Future studies are necessary to determine whether this edit alters IL33 splicing in a similar manner to the natural protective variant.

[0397] Lifestyle, diet, and environmental factors have led to a rise in metabolic disorders such as obesity, type 2 diabetes, and fatty liver disease (FLD)88. FLD is a leading cause for chronic liver disease and hepatocellular carcinoma in Western countries88. A broad range of environmental and genetic factors contribute to risk of either alcoholic or non-alcoholic fatty liver disease (NAFLD)88. In a large study surveying risk factors for FLD, an HSD17B13 allele was discovered to be protective against disease89. This variant (rs72613567:TA) was identified in individuals with lower ALT levels and was shown to reduce risk of alcoholic liver disease by 42% in heterozygotes and 53% homozygotes. Nonalcoholic liver disease risk dropped by 17% for heterozygotes and 30% for homozygotes. This allele has a variable frequency of - 22% worldwide, with higher prevalence in Asian populations and lower prevalence in African populations88. A subsequent study reported that liver damage risk was decreased in obese children carrying this allele, even among individuals that also carried major fatty liver disease risk alleles for genes PNPLA3, TM6SF2, or MBOAT790. HSD17B13 is part of a family of proteins involved in steroid metabolism, with different members of the family having differential expression in particular tissues. HSD17B13 is expressed primarily in the liver, where it has been found to associate with lipid droplets89. NAFLD patients often exhibit high expression of PISD17B139. While relatively little is known about the exact mechanism of this protective genetic variant, it is hypothesized to result in reduced HSD17B13 activity due to the disruption of the splice site between exon 6 and 7 due to the insertion of an A base89. Although technologies like prime editors, DNA-polymerase editors, or click editors91can mediate the insertion of a single base, the same splice site can also be disrupted using an ABE to convert the GT splice donor to GC. This single base edit would potentially result in a similar splicing defect to the natural variant and should knockdown or knockout HSD17B13 function in a one-and-done treatment. To modify this HSD17B13 splice donor, we utilized an A7 gRNA targeting an NGA PAM target site to utilize the only NGN PAM nearby (Fig. 5k). With ABE8e-eSpGA, we achieved >60% editing of the splice donor compared to -40% editing with ABE8e- SpG (Fig. 5k). We also observed bystander editing that is likely to be inconsequential as the bystander base is located in an intron and the goal of this edit is to generate a loss-of-function allele.

[0398] A separate genetic variant in the CIDEB gene has been similarly described to confer protection against liver disease92. Variants in CIDEB were found to associate with lower ALT levels and 53% reduced risk of NAFLD, and, strikingly, variants in the CIDEB gene were 2-3x more potent on a per allele basis than the HSD17B13:TA allele92. CIDEB is a structural protein in lipid droplets expressed highly and specifically in the liver and is thought to be involved in the fusion of small and large lipid droplets93. Interestingly, the protective effect appeared to be the most pronounced among carriers with obesity and is significant in the presence of a risk PNPLA3 allele92. Evidence also suggests that the effects of CIDEB variants are additive with the effects of HSD17B1392, which could make multiplex editing an intriguing future possibility to maximize genetic protection. In an induced steatosis cell model, silencing of CIDEB reduced lipid droplet size92. Unlike the HSD17B13. allele, several predicted CIDEB loss of function alleles have been studied and experimentally validated92. One validated allele was a splice site disruption as a result of a G>A variant (rsl46737422)92, which can be purposefully installed using CBEs targeting the opposite strand. The only NGN PAM available in the correct proximity is an NGC PAM target site that places the target base in position C6 of the protospacer (Fig. 51). Transfection of HEK293T cells using TadCBEd- eMQKSER resulted in approximately 15% editing of the target base, with two nearby bystander edits modified at similar efficiencies (Fig. 51). Notably, individuals homozygous for this variant have not been identified, so near-saturation levels of editing may not be needed to confer the protective effect or may be deleterious92. Finally, other protective variants can be installed that are both proposed to reduce function. Base editing of GPR75 could reduce function by creating the rsl48952285 SNP, which was discovered to be protective against the development of obesity94. Editing with our NGC PAM variants resulted in approximately 40% editing, though there are significant bystanders of unknown functional consequence (Fig. 25a). Another splice site disruption that recapitulates a protective allele can be installed within the MSTN gene, as previously evaluated in Walton et al. 202012using relaxed PAM variants. This protective variant, MSTN IVS1, may be useful in the prevention of muscular dystrophy effects by reducing muscle decay95. Using eSpGC- BE4max, we can also efficiently install this edit, though with likely fewer off-targets due to the use of a more PAM selective variant (Fig. 25b).

[0399] To further explore the ability to install the AELAH3447R protective genetic variant, we cloned additional gRNAs to file the locus using gRNAs A2 through A9 (numbered based on the position of the target adenine, counting from the 5’ or PAM distal end of a 20 nt protospacer; Fig. 6a). Alternate gRNA architectures were explored, including either 20 nt or 21 nt spacer sequences, and comparisons between the conventional SpCas9 sgRNA scaffold (conventional scaffold = GTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAAC TTGAAAAAGTGGCACCGAGTCGGTGC (SEQ ID NO:62)) and an alternate gRNA scaffold named ‘flip and extend’ (F&E) that includes an extended crRNA / tracrRNA duplex and a a U / T5C substitution to disrupt a poly-U sequence that acts as a transcriptional terminator (see Chen et al., Cell, 2013; 155(7): 1479-91; Dang et al., Genome Biology, 2015; 6:280) (sequence = GTTTCAGAGCTATGCTGGAAACAGCATAGCAAGTTGAAATAAGGCTAGTC CGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGC (SEQ ID NO:63)). We performed experiments in HEK 293T cells to assess the on-target editing when using various ABEs (comprised of different SpCas9 PAM variant enzymes fused to the TadA8e adenosine deaminase domain) paired with gRNA A2-A9. Analysis of editing was performed 3 days post-transfection, revealing a large range in editing depending on the combination of ABEs, gRNA spacer identity, gRNA spacer length, and / or gRNA scaffold utilized (Fig. 6b). Notably, certain conditions resulted in undetectable editing, while several resulted in >60% installation of the AELAH3447R variant. Particular combinations that were most efficacious include gRNA A4 with ABE8e- SpCas9(S55R)-VRQR (eVRQR; bearing S55R / D1135V / G1218R / R1335Q / T1337R substitutions in SpCas9), A3 with ABE8e-SpRY or ABE8e-SpRY-VRQR (bearing A61R / L1111R / N1317R / A1322R / D1135V / G1218R / R1333P / R1335Q / T1337R substitutions in SpCas9), A5 with ABE8e-SpRY-WT (bearing A61R / L1111RN1317R / A1322R / R1333P substitutions in SpCas9; also called SpRY- DSGERT instead of SpRY-WT), A6 with ABE8e-SpRY, A8 with ABE8e-SpRY-WT or ABE8e-SpRY, A7 with ABE8e-SpCas9-MRKCRS (bearing

[0400] DI 135M / S1136R / G1218K / E1219C / R1335R / T1337S substitutions in SpCas9; Silverstein et al., Nature volume 643, pages539-550 (2025)), ABE8e-SpG, or ABE8- eVRQR, or gRNA A9 with ABE8e-SpRY-VRQR or ABE8e-SpRY (Fig. 6b). Testing of ABEs comprised of different TadA domains (TadA8e, TadA8.20m, or TadA8.8m, e.g., as described in Gaudelli et al., Nature Biotechnology volume 38, pages892-900 (2020)) with some of the most efficacious enzymes with each of gRNAs A3 through A8 revealed efficient editing with most combinations (Fig. 6c), though some loss in efficiency was observed near the edge of the optimal ‘edit window’ (using gRNAs A3 A8). Together, these results demonstrate that the AELAH3447R protective genetic variant can be efficiently installed in human cells by screening for optimized ABE and gRNA combinations.

[0401] We also investigated additional ABE and gRNA combinations for installing the BAG3 C151R protective genetic variant. We designed and cloned gRNAs encoding spacers that positioned the target adenine in positions A3 through A8 (Fig. 7a). Experiments in HEK 293T cells were performed. Using each of the 6 different gRNAs, and also with 6 different ABEs comprised of either wild-type (WT) SpCas9, SpG, or SpRY, and either the TadA8.8m or TadA8e deaminase domains (Fig. 7b). On-target editing and bystander editing of nearby adenines, both of which lead to silent substitutions, were analyzed (Fig. 7a). Depending on the combination of gRNA and enzyme utilized, we observed a wide range of on-target editing efficiencies between ~0% and up to -70% (Fig. 7b). Notably, the ABE8.8-SpG or AB8.8-SpRY base editors combined with gRNA A5 resulted in high on-target editing and minimal bystander edits. Additional efficacious combinations included gRNA A3 with ABE8e-SpCas9, A5 with ABE8e-SpG or ABE8e-SpRY, A6 with ABE8.8-SpRY, ABE8e-SpCas9, ABE8.8m-SpCas9, or ABE8e-SpRY, A7 with ABE8e-SpG or ABE8.8m-SpG, or A8 with ABE8e-SpRY or ABE8e-SpG (Fig. 7b). These results reveal that screening for specific combinations of efficacious ABEs and gRNAs was necessary to be able to efficiently install the BAG3 C151R protective genetic variant.

[0402] Table F - list of plasmids #, SEQ ID NO:

[0403] References

[0404] 1. Schubert, M. S. et al. Optimized design parameters for CRISPR Cas9 and Cast 2a homology-directed repair. Sci. Rep. 11, 19482 (2021).

[0405] 2. Towards personalised allele-specific CRISPR gene editing to treat autosomal dominant disorders | Scientific Reports, nature.com / articles / s41598-017-16279-4.

[0406] 3. Rees, H. A. & Liu, D. R. Base editing: precision chemistry on the genome and transcriptome of living cells. Nat. Rev. Genet. 19, 770-788 (2018).

[0407] 4. Anzalone, A. V. et al. Search-and-replace genome editing without doublestrand breaks or donor DNA. Nature 576, 149-157 (2019).

[0408] 5. Jinek, M. et al. A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity. Science 337, 816-821 (2012).

[0409] 6. Anders, C., Niewoehner, O., Duerst, A. & Jinek, M. Structural basis of PAM- dependent target DNA recognition by the Cas9 endonuclease. Nature 513, 569-573 (2014).

[0410] 7. Marraffini, L. A. & Sontheimer, E. J. Self versus non-self discrimination during CRISPR RNA-directed immunity. Nature 463, 568-571 (2010).

[0411] 8. Jiang, W., Bikard, D., Cox, D., Zhang, F. & Marraffini, L. A. RNA-guided editing of bacterial genomes using CRISPR-Cas systems. Nat. Biotechnol. 31, 233- 239 (2013).

[0412] 9. Esvelt, K. M. et al. Orthogonal Cas9 proteins for RNA-guided gene regulation and editing. Nat. Methods 10, 1116-1121 (2013).

[0413] 10. Edraki, A. et al. A Compact, High-Accuracy Cas9 with a Dinucleotide PAM for In Vivo Genome Editing. Mol. Cell 73, 714-726. e4 (2019).

[0414] 11. Kleinstiver, B. P. et al. Engineered CRISPR-Cas9 nucleases with altered PAM specificities. Nature 523, 481-485 (2015).

[0415] 12. Walton, R. T., Christie, K. A., Whittaker, M. N. & Kleinstiver, B. P. Unconstrained genome targeting with near-PAMless engineered CRISPR-Cas9 variants. Science 368, 290-296 (2020).

[0416] 13. Miller, S. M. et al. Continuous evolution of SpCas9 variants compatible with non-G PAMs. Nat. Biotechnol. 38, 471-481 (2020).

[0417] 14. Nishimasu, H. et al. Engineered CRISPR-Cas9 nuclease with expanded targeting space. Science 361, 1259-1262 (2018). 15. Bravo, J. P. K. et al. Structural basis for mismatch surveillance by CRISPR- Cas9. Nature 603, 343-347 (2022).

[0418] 16. Hirano, S., Nishimasu, H., Ishitani, R. & Nureki, O. Structural Basis for the Altered PAM Specificities of Engineered CRISPR-Cas9. Mol. Cell 61, 886-894 (2016).

[0419] 17. Walton, R. T., Hsu, J. Y., Joung, J. K. & Kleinstiver, B. P. Scalable characterization of the PAM requirements of CRISPR-Cas enzymes using HT- PAMDA. Nat. Protoc. 1-37 (2021) doi : 10.1038 / s41596-020-00465-2.

[0420] 18. Kleinstiver, B. P. et al. High-fidelity CRISPR-Cas9 nucleases with no detectable genome-wide off-target effects. Nature 529, 490-495 (2016).

[0421] 19. Optimization of Protein Thermostability and Exploitation of Recognition Behavior to Engineer Altered Protein-DNA Recognition - ScienceDirect. sciencedirect.com / science / article / pii / S09692126203012837via%3Dihub.

[0422] 20. Deep mutational scanning of S. pyogenes Cas9 reveals important functional domains | Scientific Reports, nature. com / articles / s41598-017-17081-y.

[0423] 21. Vos, P. D. et al. Computationally designed hyperactive Cas9 enzymes. Nat. Commun. 13, 3023 (2022).

[0424] 22. Vos, P. D. et al. Mutational rescue of the activity of high-fidelity Cas9 enzymes. Cell Rep. Methods 4, 100756 (2024).

[0425] 23. High-throughput continuous evolution of compact Cas9 variants targeting single-nucleotide-pyrimidine PAMs | Nature Biotechnology. nature.com / articles / s41587-022-01410-2.

[0426] 24. Eggers, A. R. et al. Rapid DNA unwinding accelerates genome editing by engineered CRISPR-Cas9. Cell 187, 3249-3261.el4 (2024).

[0427] 25. Pacesa, M. et al. R-loop formation and conformational activation mechanisms of Cas9. Nature 609, 191-196 (2022).

[0428] 26. Nishimasu, H. et al. Crystal Structure of Cas9 in Complex with Guide RNA and Target DNA. Cell 156, 935-949 (2014).

[0429] 27. Huai, C. et al. Structural insights into DNA cleavage activation of CRISPR- Cas9 system. Nat. Commun. 8, 1375 (2017).

[0430] 28. Anders, C., Bargsten, K. & Jinek, M. Structural Plasticity of PAM Recognition by Engineered Variants of the RNA-Guided Endonuclease Cas9. Mol. Cell 61, 895-902 (2016). 29. Unraveling the mechanisms of PAMless DNA interrogation by SpRY-Cas9 | Nature Communications, nature.com / articles / s41467-024-47830-3.

[0431] 30. Christie, K. A. et al. Mutation-Independent Allele-Specific Editing by CRISPR-Cas9, a Novel Approach to Treat Autosomal Dominant Disease. Mol. Ther.

[0432] 28, 1846-1857 (2020).

[0433] 31. Gy orgy, B. et al. Allele-specific gene editing prevents deafness in a model of dominant progressive hearing loss. Nat. Med. 25, 1123-1130 (2019).

[0434] 32. Zhou, B. et al. Comprehensive, integrated, and phased whole-genome analysis of the primary ENCODE cell line K562. Genome Res. 29, 472-484 (2019).

[0435] 33. Komor, A. C., Kim, Y. B., Packer, M. S., Zuris, J. A. & Liu, D. R. Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage. Nature 533, 420-424 (2016).

[0436] 34. Gaudelli, N. M. et al. Programmable base editing of A»T to G»C in genomic DNA without DNA cleavage. Nature 551, 464-471 (2017).

[0437] 35. Richter, M. F. et al. Phage-assisted evolution of an adenine base editor with improved Cas domain compatibility and activity. Nat. Biotechnol. 38, 883-891 (2020).

[0438] 36. Gaudelli, N. M. et al. Directed evolution of adenine base editors with increased activity and therapeutic application. Nat. Biotechnol. 38, 892-900 (2020).

[0439] 37. Improving cytidine and adenine base editors by expression optimization and ancestral reconstruction | Nature Biotechnology, nature.com / articles / nbt.4172.

[0440] 38. Neugebauer, M. E. et al. Evolution of an adenine base editor into a small, efficient cytosine base editor with low off-target activity. Nat. Biotechnol. 41, 673- 685 (2023).

[0441] 39. Cofsky, J. C., Soczek, K. M., Knott, G. J., Nogales, E. & Doudna, J. A. CRISPR-Cas9 Bends and Twists DNA to Read Its Sequence. biorxiv.org / lookup / doi / 10.1101 / 2021.09.06.459219 (2021) doi:10.1101 / 2021.09.06.459219.

[0442] 40. Lapinaite, A. et al. DNA capture by a CRISPR-Cas9-guided adenine base editor. Science 369, 566-571 (2020).

[0443] 41. Sickle Cell Disease | New England Journal of Medicine. nejm.org / doi / full / 10.1056 / NE JMra 1510865. 42. Frangoul, H. et al. CRISPR-Cas9 Gene Editing for Sickle Cell Disease and P- Thalassemia. N. Engl. J. Med. 0, null (2020).

[0444] 43. Ex vivo prime editing of patient haematopoietic stem cells rescues sickle-cell disease phenotypes after engraftment in mice | Nature Biomedical Engineering. nature.com / articles / s41551-023-01026-0.

[0445] 44. Newby, G. A. et al. Base editing of haematopoietic stem cells rescues sickle cell disease in mice. Nature 595, 295-302 (2021).

[0446] 45. Quentin Blackwell, R., Oemijati, S., Pribadi, W., Weng, M.-I. & Liu, C.-S. Hemoglobin G Makassar: P6 Glu^Ala. Biochim. Biophys. Acta BB A - Protein Struct. 214, 396-401 (1970).

[0447] 46. Frangoul Haydar et al. Exagamglogene Autotemcel for Severe Sickle Cell Disease. N. Engl. J. Med. 390, 1649-1662 (2024).

[0448] 47. Commissioner, O. of the. FDA Approves First Gene Therapies to Treat Patients with Sickle Cell Disease. FDA fda.gov / news-events / press- announcements / fda-approves-first-gene-therapies-treat-patients-sickle-cell-disease (2023).

[0449] 48. Bottger, R. et al. Lipid-based nanoparticle technologies for liver targeting. Adv. Drug Deliv. Rev. 154-155, 79-101 (2020).

[0450] 49. Liu, P. et al. Improved prime editors enable pathogenic allele correction and cancer modelling in adult mice. Nat. Commun. 12, 2121 (2021).

[0451] 50. Molecular basis of alpha- 1 -antitrypsin deficiency - ScienceDirect. sci encedirect.com / science / arti cl e / pii / 0002934388901544?via%3Dihub.

[0452] 51. Alpha- 1 antitrypsin deficiency research and emerging treatment strategies: what’s down the road? - Franck F. Rahaghi, 2021. journals. sagepub. com / doi / full / 10.1177 / 20406223211014025?rfr_dat=cr_pub++0pub med&url_ver=Z39.88-2003&rfr_id=ori%3 Ari d%3Acrossref.org.

[0453] 52. Lomas, D. A., Evans, D. L., Finch, J. T. & Carrell, R. W. The mechanism of Z alpha 1-antitrypsin accumulation in the liver. Nature 357, 605-607 (1992).

[0454] 53. Werder, R. B. et al. Adenine base editing reduces misfolded protein accumulation and toxicity in alpha-1 antitrypsin deficient patient iPSC-hepatocytes. Mol. Ther. 29, 3219-3229 (2021).

[0455] 54. Joung, K. J. & KLEINSTIVER, B. Engineered crispr-cas9 nucleases with altered pam specificity. (2019). 55. Gaudelli, N. et al. Compositions and methods for treating alpha- 1 antitrypsin deficiency. (2023).

[0456] 56. Packer, M. S. et al. Evaluation of cytosine base editing and adenine base editing as a potential treatment for alpha-1 antitrypsin deficiency. Mol. Ther. 30, 1396-1406 (2022).

[0457] 57. Kohn, D. B. et al. Lentiviral gene therapy for X-linked chronic granulomatous disease. Nat. Med. 26, 200-206 (2020).

[0458] 58. De Ravin, S. S. et al. CRISPR-Cas9 gene repair of hematopoietic stem cells from patients with X-linked chronic granulomatous disease. Sci. Transl. Med. 9, eaah3480 (2017).

[0459] 59. De Ravin, S. S. et al. Enhanced homology-directed repair for highly efficient gene editing in hematopoietic stem / progenitor cells. Blood 137, 2598-2608 (2021).

[0460] 60. Lee, R. G. et al. Efficacy and Safety of an Investigational Single-Course CRISPR Base-Editing Therapy Targeting PCSK9 in Nonhuman Primate and Mouse Models. Circulation 147, 242-253 (2023).

[0461] 61. Protective alleles and modifier variants in human health and disease - PubMed. pubmed. ncbi.nlm.nih.gov / 26503796 / .

[0462] 62. Naddaf, M. First trial of ‘base editing’ in humans lowers cholesterol — but raises safety concerns. Nature 623, 671-672 (2023).

[0463] 63. Resilience to autosomal dominant Alzheimer’s disease in a Reelin-COLBOS heterozygous man | Nature Medicine, nature. com / articles / s41591-023-02318-3.

[0464] 64. Functional analysis of a common BAG3 allele associated with protection from heart failure | Nature Cardiovascular Research, nature.com / articles / s44161-023- 00288-w.

[0465] 65. Multiple genetic variants at the SLC30A8 locus affect local super-enhancer activity and influence pancreatic P-cell survival and function | bioRxiv. biorxiv.org / content / 10.1101 / 2023.07.13.548906v2.

[0466] 66. Mayeux, R. & Stern, Y. Epidemiology of Alzheimer Disease. Cold Spring Harb. Perspect. Med. 2, a006239 (2012).

[0467] 67. Virani, S. S. et al. Heart Disease and Stroke Statistics-2021 Update: A Report From the American Heart Association. Circulation 143, e254-e743 (2021). 68. Clinical and Mechanistic Insights Into the Genetics of Cardiomyopathy - ScienceDirect. sciencedirect.com / science / article / pii / S0735109716368231?via%3Dihub.

[0468] 69. Yang, Z. et al. Effects of a genetic variant rsl3266634 in the zinc transporter 8 gene (SLC30A8) on insulin and lipid levels before and after a high-fat mixed macronutrient tolerance test in U.S. adults. J. Trace Elem. Med. Biol. 77, 127142 (2023).

[0469] 70. Identification of candidate protective variants for common diseases and evaluation of their protective potential | BMC Genomics | Full Text. bmcgenomics. biomedcentral, com / arti cles / 10.1186 / s 12864-017-3964-3.

[0470] 71. Pilotti, C., Greenwood, J. & Moss, S. E. Functional Evaluation of AMD- Associated Risk Variants of Complement Factor B. Invest. Ophthalmol. Vis. Sci. 61, 19 (2020).

[0471] 72. Deficiency in Complement Factor B | New England Journal of Medicine. nej m . org / doi / 10.1056 / NEJMc 1306326?url_ver=Z39.88-

[0472] 2003&rfr_id=ori:rid:crossref.org&rfr_dat=cr_pub%20%200pubmed.

[0473] 73. Schnabolk, G. et al. Local Production of the Alternative Pathway Component Factor B Is Sufficient to Promote Laser-Induced Choroidal Neovascularization. Invest. Ophthalmol. Vis. Sci. 56, 1850-1863 (2015).

[0474] 74. Therapeutic in vivo delivery of gene editing agents - ScienceDirect. sciencedirect.com / science / article / pii / S00928674220039567via%3Dihub.

[0475] 75. Bae, S., Park, J. & Kim, J.-S. Cas-OFFinder: a fast and versatile algorithm that searches for potential off-target sites of Cas9 RNA-guided endonucleases. Bioinformatics 30, 1473-1475 (2014).

[0476] 76. Clarke, R. et al. Genetic variants associated with Lp(a) lipoprotein level and coronary disease. N. Engl. J. Med. 361, 2518-2528 (2009).

[0477] 77. Kamstrup, P. R., Tybjserg-Hansen, A. & Nordestgaard, B. G. Elevated lipoprotein(a) and risk of aortic valve stenosis in the general population. J. Am. Coll. Cardiol. 63, 470-477 (2014).

[0478] 78. LPA disruption with AAV-CRISPR potently lowers plasma apo(a) in transgenic mouse model: A proof-of-concept study: Molecular Therapy Methods & Clinical Development. cell.com / molecular-therapy-family / methods / fulltext / S2329- 0501(22)00151-6. 79. Chan, D. C. & Watts, G. F. The Promise of PCSK9 and Lipoprotein(a) as Targets for Gene Silencing Therapies. Clin. Ther. 45, 1034-1046 (2023).

[0479] 80. Genotoxic effects of base and prime editing in human hematopoietic stem cells | Nature Biotechnology, nature. com / articles / s41587-023-01915-4.

[0480] 81. Cullot, G. et al. CRISPR-Cas9 genome editing induces megabase-scale chromosomal truncations. Nat. Commun. 10, 1136 (2019).

[0481] 82. Said, M. A. et al. Genome-Wide Association Study and Identification of a Protective Missense Variant on Lipoprotein(a) Concentration. Arterioscler. Thromb. Vase. Biol. 41, 1792-1800 (2021).

[0482] 83. Tsimikas, S. et al. Lipoprotein(a) Reduction in Persons with Cardiovascular Disease. N. Engl. J. Med. 382, 244-255 (2020).

[0483] 84. Smith, D. et al. A rare IL33 loss-of-function mutation reduces blood eosinophil counts and protects from asthma. PLoS Genet. 13, el006659 (2017).

[0484] 85. Ijaz, H. M., Chowdhury, W., Lodhi, M. U., Gulzar, Q. & Rahim, M. A Case of Persistent Asthma Resistant to Available Treatment Options: Management Dilemma. Cureus 11, e4194.

[0485] 86. High-throughput barcoding of nanoparticles identifies cationic, degradable lipid-like materials for mRNA delivery to the lungs in female preclinical models | Nature Communications, nature.com / articles / s41467-024-45422-9.

[0486] 87. Kurt, I. C. et al. CRISPR C-to-G base editors for inducing targeted DNA transversions in human cells. Nat. Biotechnol. 39, 41-46 (2021).

[0487] 88. Motomura, T. et al. Is HSD17B13 Genetic Variant a Protector for Liver Dysfunction? Future Perspective as a Potential Therapeutic Target. J. Pers. Med. 11, 619 (2021).

[0488] 89. Abul-Husn, N. S. et al. A Protein-Truncating HSD17B13 Variant and Protection from Chronic Liver Disease. N. Engl. J. Med. 378, 1096-1106 (2018).

[0489] 90. Di Sessa, A. et al. The rs72613567: TA Variant in the Hydroxysteroid 17-beta Dehydrogenase 13 Gene Reduces Liver Damage in Obese Children. J. Pediatr.

[0490] Gastroenterol. Nutr. 70, 371-374 (2020).

[0491] 91. da Silva, J. F. et al. Click editing enables programmable genome writing using DNA polymerases and HUH endonucleases. BioRxiv Prepr. Serv. Biol.

[0492] 2023.09.12.557440 (2023) doi: 10.1101 / 2023.09.12.557440. 92. Germline Mutations in CIDEB and Protection against Liver Disease | New England Journal of Medicine. nejm.org / doi / full / 10.1056 / NEJMoa2117872.

[0493] 93. Ye, J. et al. Cideb, an ER- and lipid droplet-associated protein, mediates VLDL lipidation and maturation by interacting with apolipoprotein B. Cell Metab. 9, 177-190 (2009).

[0494] 94. Sequencing of 640,000 exomes identifies GPR75 variants associated with protection from obesity | Science, science.org / doi / 10.1126 / science.abf8683.

[0495] 95. Schuelke, M. et al. Myostatin mutation associated with gross muscle hypertrophy in a child. N. Engl. J. Med. 350, 2682-2688 (2004).

[0496] 96. Clement, K. et al. CRISPResso2 provides accurate and rapid genome editing sequence analysis. Nat. Biotechnol. 37, 224-226 (2019).

[0497] 97. Amoli, M. M., Carthy, D., Platt, H. & Ollier, W. E. R. EBV Immortalization of human B lymphocytes separated from small volumes of cryo-preserved whole blood. Int. J. Epidemiol. 37 Suppl 1, i41-45 (2008).

[0498] 98. Tsai, S. Q. et al. GUIDE-seq enables genome-wide profiling of off-target cleavage by CRISPR-Cas nucleases. Nat. Biotechnol. 33, 187-197 (2015).

[0499] 99. Open-source guideseq software for analysis of GUIDE-seq data | Nature Bi otechnology . nature . com / arti cl es / nbt .3534.

[0500] OTHER EMBODIMENTS

[0501] It is to be understood that while the invention has been described in conjunction with the detailed description thereof, the foregoing description is intended to illustrate and not limit the scope of the invention, which is defined by the scope of the appended claims. Other aspects, advantages, and modifications are within the scope of the following claims.

Claims

WHAT IS CLAIMED IS:

1. A Streptococcus pyogenes Cas9 (SpCas9) protein, comprising mutations as listed in Table A, preferably eSpGT, eSpGC, eSpGA, or eSpGT.2-HFl.

2. The protein of claim 1, comprising a sequence that is at least 80% identical to the amino acid sequence of SEQ ID NO: 1.

3. The protein of claim 1, further comprising one or more mutations that decrease nuclease activity selected from the group consisting of mutations at DIO, E762, D839, H983, or D986; and at H840 or N863.

4. The protein of claim 3, wherein the mutations are:(i) D10A or DION, and / or(ii) H840A, H840N, or H840Y5. The protein of claim 1, further comprising one or more additional mutations that alter or increase specificity selected from the group consisting of mutations listed in Table C or D.

6. A fusion protein comprising the protein of claim 1, fused to a heterologous functional domain, with an optional intervening linker, wherein the linker does not interfere with activity of the fusion protein.

7. The fusion protein of claim 6, wherein the heterologous functional domain is a transcriptional activation domain.

8. The fusion protein of claim 7, wherein the transcriptional activation domain is from VP16, VP64, rTA, NF-KB p65, or composite VPR (VP64-p65-rTA).

9. The fusion protein of claim 6, wherein the heterologous functional domain is a transcriptional silencer or transcriptional repression domain.

10. The fusion protein of claim 9, wherein the transcriptional repression domain is a Krueppel-associated box (KRAB) domain, ERF repressor domain (ERD), or mSin3 A interaction domain (SID), or the transcriptional silencer is Heterochromatin Protein 1 (HP1).

11. The fusion protein of claim 6, wherein the heterologous functional domain is an enzyme that modifies the methylation state of DNA.

12. The fusion protein of claim 11, wherein the enzyme that modifies the methylation state of DNA is a DNA methyltransferase (DNMT) or a TET protein, optionally wherein the TET protein is TET1.

13. The fusion protein of claim 6, wherein the heterologous functional domain is an enzyme that modifies a histone subunit.

14. The fusion protein of claim 13, wherein the enzyme that modifies a histone subunit is a histone acetyltransferase (HAT), histone deacetylase (HDAC), histone methyltransferase (HMT), or histone demethylase.

15. The fusion protein of claim 6, wherein the heterologous functional domain is a base editing domain.

16. The fusion protein of claim 15, wherein the base editing domain comprises:(i) a cytidine deaminase domain, or(ii) an adenosine deaminase domain.

17. The fusion protein of claim 6, wherein the heterologous functional domain is a biological tether.

18. The fusion protein of claim 17, wherein the biological tether is MS2, Csy4 or lambda N protein.

19. The fusion protein of claim 6, wherein the heterologous functional domain is Fokl.

20. A nucleic acid encoding the protein of claims 1-19.

21. A vector comprising the nucleic acid of claim 20, optionally wherein the nucleic acid is operably linked to one or more regulatory domains for expressing the Streptococcus pyogenes Cas9 (SpCas9) protein, with mutations at one, two, three, four, five, or all six of the following positions: D1135, S 1136, G1218, E1219, R1335, and / or T1337, wherein the mutations are listed in Table A, and optionally a nucleic acid encoding a guide RNA that complexes with the cas9 protein.

22. A host cell, preferably a mammalian host cell, comprising the nucleic acid of claim 20.

23. A method of altering the genome of a cell, the method comprising expressing in the cell, or contacting the cell with, the protein or fusion protein of any of claims 1-19, and a guide RNA having a region complementary to a selected portion of the genome of the cell.

24. The method of claim 23, wherein the protein or fusion protein comprises one or more of a nuclear localization sequence, cell penetrating peptide sequence, and / or affinity tag.

25. The method of claim 24, wherein the cell is a stem cell.

26. The method of claim 25, wherein the cell is an embryonic stem cell, mesenchymal stem cell, or induced pluripotent stem cell; is in a living animal; or is in an embryo.

27. A method of altering a double stranded DNA (dsDNA) molecule, the method comprising contacting the dsDNA molecule with the protein or fusion protein of claims 1 to 19, and a guide RNA that complexes with the cas9 protein, the guide RNA having a region complementary to a selected portion of the dsDNA molecule.

28. The method of claim 27, wherein the dsDNA molecule is in vitro.

29. The method of claim 27, wherein the fusion protein and RNA are in a ribonucleoprotein complex.

30. The method of any of claims 23-29, wherein the guide RNA comprises a spacer sequence at least 80%, 90%, 95%, or 99% identical to a gRNA sequence listed in Table E, preferably wherein the base editor comprises a corresponding variant listed in Table E.

31. A method of treating or reducing risk of a disease in a subject, the method comprising administering to the subject, or to a cell from the subject, a guide RNA comprising a spacer sequence at least 80%, 90%, 95%, or 99% identical to agRNA sequence listed in the following table, and a base editor, preferably a base editor comprising a variant listed in the following table:

32. The method of claim 31, for installing AEZ VH3447R variant for treating or reducing risk of Alzheimer’s disease in a subject, the method comprising administering to the subject, or to a cell from the subject, a guide RNA comprising a spacer sequence at least 80%, 90%, 95%, or 99% identical to a gRNA sequence listed in the following table, and a base editor, preferably a base editor comprising a variant listed in the following table:

33. The method of claim 31, for treating or reducing risk of heart failure in a subject, the method comprising administering to the subject, or to a cell from the subject, a guide RNA comprising a spacer sequence at least 80%, 90%, 95%, or 99% identical to a gRNA sequence listed in the following table, and a base editor, preferably a base editor comprising a variant listed in the following table:

Citation Information

Patent Citations

  • Crispr / CAS systems for genomic modification and gene modulation

    US20140273226A1

  • Crispr-cas enzymes with enhanced on-target activity

    US20210261932A1

  • Unconstrained Genome Targeting with near-PAMless Engineered CRISPR-Cas9 Variants

    US20210284978A1

  • Methods of editing single nucleotide polymorphism using programmable base editor systems

    US20210380955A1

  • Genome editing approaches to treat Spinal Muscular Atrophy

    US20240066102A1