Generation of custom crispr-cas9 PAM variant enzymes
Patent Information
- Application Number
- PCT/IB2025/050835
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-24
- Filing Date
- 2025-01-24
- Publication Date
- 2025-09-04
AI Technical Summary
Current CRISPR-Cas9 enzymes are limited by their requirement for specific protospacer adjacent motifs (PAMs), which restricts the number of genomic sites that can be edited and increases off-target editing, necessitating time-consuming protein engineering to produce generalist enzymes that are not optimally suited for specific targets.
Integration of high-throughput protein engineering with machine learning (ML) to predict and develop bespoke CRISPR-Cas9 PAM variant enzymes with customizable PAM specificities, using a PAM machine learning algorithm (PAMmla) to identify enzymes with tailored PAM requirements for precise genome editing.
The PAMmla-predicted enzymes demonstrate improved activity and specificity across various genomic sites, reducing off-target effects and enabling efficient, targeted genome editing, as shown in human cells and therapeutic applications.
Abstract
Description
[0001] Generation of Custom CRISPR-Cas9 PAM Variant Enzymes
[0002] CLAIM OF PRIORITY
[0003] This application claims the benefit of U.S. Provisional Patent Application Serial No. 63 / 624,377, filed on January 24, 2024. The entire contents of the foregoing are hereby incorporated by reference.
[0004] FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT
[0005] This invention was made with Government support under Grant Nos. HL142494 and CA281401 awarded by the National Institutes of Health. The Government has certain rights in the invention.
[0006] TECHNICAL FIELD
[0007] Provided herein are CRISPR-Cas9 PAM variant enzymes, methods of using the same, and machine-learning based methods of identifying variant enzymes with desired PAM specificity or activity.
[0008] BACKGROUND
[0009] CRISPR-Cas enzymes have been extensively engineered to modify their functions for genome editing applications4. The RNA-programmable nature of target site recognition by CRISPR-Cas enzymes has largely driven their broad adoption for various editing approaches. However, other key properties of these enzymes remain “hard-coded” in the protein sequence. For instance, DNA-targeting CRISPR-Cas nucleases initially recognize genomic targets through the readout of protospacer adjacent motifs (PAMs) that are typically short in sequence (~2-4 nucleotides)5. Cas9 enzymes form direct base-specific or non-specific protein:DNA contacts with the PAM though their PAM interacting (PI) domain, which enables the enzyme to undergo conformational changes to begin pairing the reprogrammable guide RNA (gRNA) spacer with the target site protospacer6,7(Fig. la). The commonly used Cas enzyme Streptococcus pyogenes Cas9 (SpCas9) requires a 3’ NGG PAM8,9. The PAM requirement provides target site selectivity by limiting enzyme access to only the fraction of the genome harboring the requisite PAM, leading to a reduction in undesired off-target editing. However, the necessity for Cas enzymes to recognize a PAM also limits the number of genomic sites which can be edited.
[0010] SUMMARY
[0011] Traditional protein engineering methods typically aim to produce single optimized proteins with broadly useful functionalities, as has been done previously to alter the properties of CRISPR-Cas enzymes. The process of engineering and characterizing Cas enzymes is time-consuming and cumbersome, necessitating the development of generalist enzymes that function across a range of targets. To enable more scalable reprogramming of efficient and safe genome editors that are more uniquely suited to specific genomic targets and that overcome the caveats / tradeoffs of generalist enzymes, high-throughput protein engineering was combined with machine learning (ML) to derive bespoke editors with superior properties. Given the importance of protospacer-adjacent motif (PAM) recognition to target site proof-reading, large datasets related to the PAM preferences of hundreds of novel engineered SpCas9 enzymes were created, facilitating development of a ML framework to predict CRISPR-Cas enzyme variants that recognize custom PAMs. Via structure / function- informed saturation mutagenesis followed by bacterial selections to identify active SpCas9 enzymes, the PAM specificities of over 700 enzyme variants were profiled using HT-PAMDA, and this data was used to train a neural network to relate PAM specificity to amino acid sequence. The resulting PAM ML algorithm (PAMmla) was then utilized to predict the PAM requirements for 64 million enzyme variants, leading to novel SpCas9 variants with unique PAM preferences. PAMmla-predicted enzymes performed as well as or better than evolution-based enzymes and highly optimized SpCas9 variants (e.g., SpG and SpRY) as nucleases and base editors across various sites in human cells and show consistently fewer genome-wide off targets.
[0012] Additionally, an in silico directed evolution method was developed that utilizes PAMmla predictions to guide evolution of custom Cas enzymes based on user-defined selection criteria. Using these methods, a custom editor was developed to target the dominant negative RHO P23H mutation causing certain types of retinitis pigmentosa. The present PAMmla methods integrate ML with protein engineering to derive a rich catalog of bespoke SpCas9-based enzymes, achieving a greater plasticity of the PAM interacting domain than previously explored, and provide a framework to quickly identify safe and effective SpCas9 variant enzymes, moving away from generalist technologies towards bespoke editors for a wide range of genome editing applications. Provided herein are isolated Streptococcus pyogenes Cas9 (SpCas9) variant proteins, comprising mutations at two, three, four, five, or all six of the following positions: D1135, S1136, G1218, E1219, R1335, and / or T1337. In some embodiments, the variant protein is predicted by the PAMmla machine learning algorithm to be active (with a rate constant, k, > 10'3) on at least one PAM. In some embodiments, the mutations are listed in any of Tables A, B, C, D, E, or E
[0013] In some embodiments, the isolated protein comprises a sequence that is at least 80% identical to the amino acid sequence of SEQ ID NO: 1.
[0014] In some embodiments, the isolated protein further comprises one or more mutations that decrease nuclease activity selected from the group consisting of mutations at DIO, E762, D839, H983, or D986; and at H840 or N863. In some embodiments, the mutations are: (i) D10A or DION, and / or (ii) H840A, H840N, or H840Y, optionally a combination of D10A or DION, and H840A, H840N, or H840Y.
[0015] In some embodiments, the isolated protein further comprises one or more mutations that increase specificity selected from the group consisting of mutations at N497, R661, N692, M694, Q695, H698, K810, K848, Q926, K1003, and R0160, and optionally at K526 and / or R691.In some embodiments, the isolated protein further comprises mutations N692A, Q695A, Q926A, H698A, N497A, R661A, M694A, K810A, K848A, K1003A, R0160A, Y450A / Q695A, L169A / Q695A, Q695A / Q926A, Q695A / D1135E, Q926A / D1135E, Y450A / D1135E, L169A / Y450A / Q695A, L169A / Q695A / Q926A, Y450A / Q695A / Q926A, R661A / Q695A / Q926A, N497A / Q695A / Q926A, Y450A / Q695A / D1135E, Y450A / Q926A / D1135E, Q695 A / Q926A / D 1135E, L 169A / Y450A / Q695 A / Q926 A, L169A / R661A / Q695A / Q926A,Y450A / R661A / Q695A / Q926A, N497A / Q695A / Q926A / D1135E, R661A / Q695A / Q926A / D1135E, and Y450A / Q695 A / Q926A / D 1135E; N692A / M694 A / Q695 A / H698 A, N692A / M694A / Q695A / H698A / Q926A; N692A / M694A / Q695A / Q926A;
[0016] N692A / M694A / H698A / Q926A; N692A / Q695A / H698A / Q926A; M694A / Q695A / H698A / Q926A; N692A / Q695A / H698A; N692A / M694A / Q695A; N692A / H698A / Q926A; N692A / M694A / Q926A; N692A / M694A / H698A; M694A / Q695A / H698A; M694A / Q695A / Q926A; Q695A / H698A / Q926A;
[0017] G582 A / V583 A / E584 A / D585 A / N588 A / Q926 A; G582A / V583A / E584A / D585A / N588A; T657A / G658AAV659A / R661A / Q926A;
[0018] T657A / G658A / W659A / R661A; F491A / M495A / T496A / N497A / Q926A;
[0019] F491A / M495A / T496A / N497A; K918A / V922A / R925A / Q926A; or 918A / V922A / R925A; K855A; K810A / K1003A / R1060A; or K848A / K1003A / R1060A, and optionally K526A and / or R691A.
[0020] Also provided herein are fusion proteins comprising a variant Cas9 protein as described herein fused to a heterologous functional domain, with an optional intervening linker, wherein the linker does not interfere with activity of the fusion protein.
[0021] In some embodiments, the heterologous functional domain is a transcriptional activation domain. In some embodiments, the transcriptional activation domain is from VP16, VP64, rTA, NF-KB p65, or the composite VPR (VP64-p65-rTA).
[0022] In some embodiments, the heterologous functional domain is a transcriptional silencer or transcriptional repression domain. In some embodiments, the transcriptional repression domain is a Krueppel-associated box (KRAB) domain, ERF repressor domain (ERD), or mSin3 A interaction domain (SID). In some embodiments, the transcriptional silencer is Heterochromatin Protein 1 (HP1).
[0023] In some embodiments, the heterologous functional domain is an enzyme that modifies the methylation state of DNA. In some embodiments, the enzyme that modifies the methylation state of DNA is a DNA methyltransferase (DNMT) or a TET protein. In some embodiments, the TET protein is TET1.
[0024] In some embodiments, the heterologous functional domain is an enzyme that modifies a histone subunit. In some embodiments, the enzyme that modifies a histone subunit is a histone acetyltransferase (HAT), histone deacetylase (HD AC), histone methyltransferase (HMT), or histone demethylase.
[0025] In some embodiments, the heterologous functional domain is a base editor.
[0026] In some embodiments, the base editor is (i) a cytidine deaminase domain, preferably selected from the group consisting of the apolipoprotein B mRNA-editing enzyme, catalytic polypeptide-like (APOBEC) family of deaminases, optionally APOBEC1, APOBEC2, AP0BEC3A, APOBEC3B, APOBEC3C, AP0BEC3D / E, APOBEC3F, APOBEC3G, AP0BEC3H, or APOBEC4; activation-induced cytidine deaminase (AID), optionally activation induced cytidine deaminase (AICDA); cytosine deaminase 1 (CDA1) or CDA2; cytosine deaminase acting on tRNA (CD AT), or an engineered variant thereof optionally DddA-like cytidine deaminase or engineered TadA-based cytosine base editors, or (ii) an adenosine deaminase, preferably selected from the group consisting of adenosine deaminase 1 (ADA1), ADA2; adenosine deaminase acting on RNA 1 (AD ARI), ADAR2, ADAR3; adenosine deaminase acting on tRNA 1 (ADAT1), ADAT2, ADAT3; and naturally occurring or engineered tRNA-specific adenosine deaminase (TadA).
[0027] In some embodiments, the heterologous functional domain is a biological tether. In some embodiments, the biological tether is MS2, Csy4 or lambda N protein.
[0028] In some embodiments, the heterologous functional domain is Fokl.
[0029] Also provided herein are isolated nucleic acids encoding the variant proteins described herein, and vectors comprising the isolated nucleic acids, optionally wherein the isolated nucleic acid is operably linked to one or more regulatory domains for expressing the isolated Streptococcus pyogenes Cas9 (SpCas9) protein, with mutations at one, two, three, four, five, or all six of the following positions: D1135, S 1136, G1218, E1219, R1335, and / or T1337, wherein the mutations are listed in any of Table A-F, and optionally a nucleic acid encoding a guide RNA that complexes with the cas9 protein.
[0030] Also provided herein are host cells, preferably mammalian host cells, comprising the nucleic acids.
[0031] Additionally, provided herein are methods of altering the genome of a cell comprising expressing in the cell, or contacting the cell with, an isolated variant protein or fusion protein as described herein, and a guide RNA having a region complementary to a selected portion of the genome of the cell. For example, the variants and guide RNAs described herein can be used for editing disease-relevant mutations in ApoE, HBB, and Rho.
[0032] In some embodiments, the isolated protein or fusion protein comprises one or more of a nuclear localization sequence, cell penetrating peptide sequence, and / or affinity tag. In some embodiments, the cell is a stem cell, e.g., an embryonic stem cell, mesenchymal stem cell, or induced pluripotent stem cell; is in a living animal; or is in an embryo.
[0033] Also provided herein are methods of altering a double stranded DNA (dsDNA) molecule, the method comprising contacting the dsDNA molecule with a variant protein or fusion protein as described herein, and a guide RNA that complexes with the cas9 protein, the guide RNA having a region complementary to a selected portion of the dsDNA molecule. In some embodiments, the dsDNA molecule is in vitro. In some embodiments, the fusion protein and RNA are in a ribonucleoprotein complex. Further, provided herein are methods for generating enzymes of desired functionality, the methods comprising: identifying active SpCas9 enzymes through structure-and- function-informed saturation mutagenesis and subsequent bacterial selection; profiling the active SpCas9 enzymes according to all possible protospacer-adjacent motif (PAM) requirements to amino acid sequences using a high-throughput PAM- determination assay (HT-PAMDA) to generate a first training data set for training a PAM machine-learning (ML) model to relate PAM requirements to amino acid sequence; generating a second training data set comprising random non-selected enzymes from a SpCas9(6AA) library; training the PAM ML model based on the first and the second training data sets to relate enzyme function to amino acid sequences; and using the PAM ML model to perform screening of amino acid combinations based on an amino acid sequence as input to obtain a PAM variant enzyme with predicted respective PAM requirements.
[0034] In some embodiments, the methods comprise storing the generated first and second training data set at an enzyme data store; and providing an interface for querying data from the enzyme data store.
[0035] In some embodiments, the methods comprise using the PAM ML model to predict PAM requirements of enzymes obtained through the interface of the enzyme data store; sorting the enzymes in the first and second training data sets according to activity and / or selectivity of the enzymes according to PAM requirements of the enzymes; and using the sorted enzymes, classifying the enzymes based on either activity or selectivity for each PAM combination for enzymes within the first and second training data sets.
[0036] In some embodiments, using the HT-PAMDA assay used to generate the first training data set comprises determining quantitative rate constants against all possible PAMs for each SpCas9 enzyme variant.
[0037] In some embodiments, training the PAM ML model comprises: assigning a label to each training enzyme of the first and second training data sets based on a most active PAM of the respective training enzyme; and performing a random selection of enzymes to define a training sample having balance across PAM classes.
[0038] In some embodiments, an enzyme is label as active when no PAM has a maximum rate constant k on a most active PAM, where k > 10'4. Additionally, provided herein are methods for computationally generating SpCas9 variant enzymes with customizable properties. The methods comprise computationally mutating a starting sequence to generate a library of sequences having random amino acid substitutions with a defined hamming distance from the starting sequence; using a protospacer-adjacent motif (PAM) machine-learning (ML) model trained to relate PAM requirements to amino acid sequence, generating PAM predictions for each member of the library; determining a score, using a fitness function, for each sequence in the library according to properties defined for evolution; determining a set of sequences from the library having a respective score matching a threshold score criterion for selection as a subsequent starting sequence for a subsequent evaluation iteration; and iteratively perform evaluation comprising performing the computational mutation and the PAM predictions generation until the determined scores of the sequences using the customizable fitness function satisfy a fitness criterion.
[0039] In some embodiments, the PAM ML model is trained based on a first data set and a second data set, wherein: the first data set is generated by: identifying active SpCas9 enzymes from a SpCas9(6AA) library through structure-and-function-informed saturation mutagenesis and subsequent bacterial selection; profiling the active SpCas9 enzymes according to all possible protospacer-adjacent motif (PAM) requirements to amino acid sequences using a high-throughput PAM-determination assay (HT- PAMDA) to generate the first data set, and the second data set is generated by: randomly determining non-selected enzymes from the SpCas9(6AA) library for the first data set to be included in the second data set. In some embodiments, the methods comprise receiving the fitness function through user input.
[0040] Further, provided herein are computer-implemented methods for automatic characterization of generated SpCas9 variant enzymes. The methods comprise: obtaining SpCas9 enzymes from a data store comprising SpCas9(6AA) enzymes; cloning the SpCas9 enzymes into a mammalian expression plasmid; sequencing the cloned SpCas9 enzymes by a whole-ORF sequencing; and subjecting the sequenced SpCas9 enzymes to a high-throughput PAM-determination assay (HT-PAMDA) for comprehensive PAM characterization of the sequenced SpCas9 enzymes.
[0041] In some embodiments, the enzymes are obtained from the data store through bacterial positive selection against 16 different substrates encoding PAMs or by randomly identifying unselected enzymes from the data store. In some embodiments, sequencing the cloned SpCas9 enzymes comprises: performing a whole-ORF sequencing process to perform tiled-amplicon sequencing across an entire coding sequence of an expression plasmid of an enzyme, and obtaining a complete sequence validation performed over up to 96 enzyme variants in parallel, to determine relationship between an amino acid sequence and a specific plasmid based on positions.
[0042] In some instances, one example method for generating enzymes of desired functionality may include operations such as: identifying active SpCas9 enzymes through structure-and-function-informed saturation mutagenesis and subsequent bacterial selection; profiling the active SpCas9 enzymes according to all possible protospacer-adjacent motif (PAM) requirements to amino acid sequences using a high-throughput PAM-determination assay (HT-PAMDA) to generate a first training data set for training a PAM machine-learning (ML) model to relate PAM requirements to amino acid sequence; generating a second training data set comprising random non-selected enzymes from a SpCas9(6AA) library; training the PAM ML model based on the first and the second training data sets to relate enzyme function to amino acid sequences; and using the PAM ML model to perform screening of amino acid combinations based on an amino acid sequence as input to obtain a PAM variant enzyme with predicted respective PAM requirements.
[0043] In some instances, an example method for computationally generating SpCas9 variant enzymes with customizable properties may include operations such as: computationally mutating a starting sequence to generate a library of sequences having random amino acid substitutions with a defined hamming distance from the starting sequence; using a protospacer-adjacent motif (PAM) machine-learning (ML) model trained to relate PAM requirements to amino acid sequence, generating PAM predictions for each member of the library; determining a score, using a fitness function, for each sequence in the library according to properties defined for evolution; determining a set of sequences from the library having a respective score matching a threshold score criterion for selection as a subsequent starting sequence for a subsequent evaluation iteration; and iteratively perform evaluation comprising performing the computational mutation and the PAM predictions generation until the determined scores of the sequences using the customizable fitness function satisfy a fitness criterion. In some instances, an example method for automatic characterization of generated SpCas9 variant enzymes may include operations such as: obtaining SpCas9 enzymes from a data store comprising SpCas9(6AA) enzymes; cloning the SpCas9 enzymes into a mammalian expression plasmid; sequencing the cloned SpCas9 enzymes by a whole-ORF sequencing; and subjecting the sequenced SpCas9 enzymes to a high- throughput PAM-determination assay (HT-PAMDA) for comprehensive PAM characterization of the sequenced SpCas9 enzymes.
[0044] The present disclosure also provides a computer-readable storage medium coupled to one or more processors and having instructions stored thereon which, when executed by the one or more processors, cause the one or more processors to perform operations in accordance with implementations of methods provided herein, for example, methods related to generating, training, and using a machine learning model in the context of generation of enzyme variants and characterization of the enzymes in accordance with implementations of the present disclosure.
[0045] The present disclosure further provides a system for implementing the methods provided herein. The system includes one or more processors, and a computer- readable storage medium coupled to the one or more processors having instructions stored thereon which, when executed by the one or more processors, cause the one or more processors to perform operations in accordance with implementations of the methods provided herein.
[0046] Additionally provided herein are methods of editing an Apolipoprotein E epsilon-4 (AP0E-&4 (Argl30, Argl76)) allele in a cell, e.g., in a subject, e.g., to reduce risk of developing Alzheimer’s disease. The methods comprise introducing two C to T point mutations to revert the allele to APOE-B2 (Cysl30, Cysl76) by contacting the cell with a first gRNA positioning a target cytosine at protospacer position 5 for Arg 130, and a second gRNA positioning a target cytosine at protospacer position 5 for Argl76, wherein the Argl30 gRNA utilizes an NGTA PAM and the Argl76 gRNA utilizes an NGGC PAM, and a cytosine base editor comprising an SpCas9 variant selected from the group consisting of MRKCRS, MRKQKC, MRKSKC, MRKMRS, MRKCRN, MRKCRQ, QRKCKK, MRKCRK, MRACRQ, MRKMRK, and MRKMRQ, preferably wherein the first and second guide RNAs comprise APOE spacer sequences listed in Table 1.
[0047] Also provided herein are methods of editing an HBB E7V mutation in a cell, e.g., a cell from or in a subject who has sickle cell disease. The methods comprise contacting the cell with an adenine base editor comprising an SpCas9 variant selected from the group consisting of LWQYQH, MWKYQS, LWKYQS, and MWKYQA, and a guide RNA, preferably an HBB guide RNA comprising a spacer sequence listed in Table 1. Further, provided herein are methods of allele-specific deletion of a RHO P23H allele in a cell, e.g., a cell in or from a subject who has retinitis pigmentosa (optionally a retinal cell). The methods comprise contacting the cell with an SpCas9 variant selected from the group consisting of MRRWMR, KRHWMR, MRRMYR, KRRYQR, MRAFMR, KRAYQR, KRRWQR, MRHWQR, IRSMQR, KRAWQR, LRKMYR, ARG I MR, MKRCMV, MWNVML, and MKKCMN, and a guide RNA, preferably a RHO-P23H guide RNA comprising a spacer sequence listed in Table 1. Also provided herein are methods of editing a CYBB T362I mutation in a cell, e.g., a cell from or in a subject who has X-linked Chronic Granulomatous Disease (X-CGD). The methods comprise contacting the cell with an adenine base editor comprising an SpCas9 variant KWRQLC, and a guide RNA, preferably a CYBB guide RNA comprising a spacer sequence listed in Table 1.
[0048] In some embodiments, the cell is in a living subject.
[0049] In some embodiments, the methods comprise contacting the cell with a nucleic acid encoding the cytosine base editor, adenine base editor, or SpCas9 variant.
[0050] In some embodiments, the nucleic acid comprises a viral vector, e.g., as described herein, e.g., an AAV.
[0051] In some embodiments, the methods comprise contacting the cell with a ribonucleoprotein (RNP) complex comprising the guide RNAs and the cytosine base editor, adenine base editor, or SpCas9 variant.
[0052] Also provided herein are compositions comprising the nucleic acids and / or RNPs as described herein.
[0053] It is appreciated that methods in accordance with the present disclosure can include any combination of the aspects and features described herein. That is, methods in accordance with the present disclosure are not limited to the combinations of aspects and features specifically described herein, but also include any combination of the aspects and features provided.
[0054] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Methods and materials are described herein for use in the present invention; other, suitable methods and materials known in the art can also be used. The materials, methods, and examples are illustrative only and not intended to be limiting. All publications, patent applications, patents, sequences, database entries, and other references mentioned herein are incorporated by reference in their entirety. In case of conflict, the present specification, including definitions, will control. Other features and advantages of the invention will be apparent from the following detailed description and figures, and from the claims.
[0055] DESCRIPTION OF DRAWINGS
[0056] FIGs. 1A-F. High-throughput characterization of engineered SpCas9 PAM variant enzymes, (a) Schematic of target site recognition by an SpCas9-gRNA complex, (b) Representation of the theoretical balance between targeting range and genome-wide specificity of engineered SpCas9 enzymes, where an ability to recognize a larger fraction of the genome via a relaxed PAM requirement may lead to increased off-target potential, (c) Schematic of the workflow to engineer SpCas9 PAM variant enzymes via directed evolution. SpCas9 enzymes were obtained from a saturation mutagenesis library (harboring 6 amino acids with NNS codons; SpCas9(6AA)) either via bacterial positive selection (against 16 different substrates encoding NGNN PAMs) or by randomly picking unselected library members. SpCas9 enzymes were cloned into a mammalian expression plasmid, sequenced by a whole- ORF sequencing workflow, and subjected to the HT-PAMDA assay41for comprehensive PAM characterization, (d) Heatmap representation of the PAM requirements of 655 PAM variant enzymes obtained through the 16 selection experiments on NGNN PAMs, determined using HT-PAMDA (where the rate constant (k) on a PAM is a measure of targeting efficiency). PAM profiles were hierarchically clustered, with the 8 largest clusters highlighted and analyzed using sequence logos to display the amino acid composition of enzymes from each cluster (right panel). More complete PAM profiles for representative enzymes from each cluster are shown (left panel). HT-PAMDA datasets are the mean of n = 2 biological replicates using different target sites, (e) Fraction of PAM variant enzymes maximally active against the specific NGNN PAM that they were selected / designed against (rank = 1st) or where the PAM selected against was within the top 4 most active PAMs (rank = 2nd- 4th), as determined by HT-PAMDA. Enzymes obtained from bacterial selections (left) and enzymes rationally designed based on most enriched amino acids from selections (right), (f) SpCas9 enzymes categorized by general PAM preference based on HT- PAMDA data (clustered as in panel d). Enzymes were labeled as inactive when no PAM had a k > 10'4.
[0057] FIGs. 2A-I. Development of a machine learning model to predict PAM preference from amino acid sequence, (a) Schematic a machine learning model that uses SpCas9 variant enzyme HT-PAMDA training data to predict the PAM requirements for novel enzymes bearing any amino acid at each of the 6 positions in the SpCas9 PAM-interacting (PI) domain, (b) Correlation between our PAM machine learning algorithm (PAMmla) model predictions and experimentally determined rate constants (by HT-PAMDA) on a test set comprising 20% of the HT-PAMDA dataset held out from training, (c) Model performance (predicting PAM ks using PAMmla compared to HT-PAMDA determined ks amongst different test sets by enzyme similarity to most similar sequence during training, (d) Receiver operating characteristic curve for binary classification of test set enzymes as active or inactive; enzymes are defined as inactive if the maximum HT-PAMDA rate constant on any PAM is < 10'4 3. (e) Classification results on the test set when the threshold for identifying inactive enzymes is maximum PAMmla predicted k < 10'4 3. (f-i) Comparison of experimentally determined PAM profiles (via HT-PAMDA; top panels in blue) to PAMmla predicted PAM profiles (bottom panels in red) with correlation between experimental and predicted ks (right panels), for previously published enzymes, including SpG11(panel f), VRER10(panel g), VRQR10,88(panel h), and xCas912(panel i). HT-PAMDA datasets are the mean of n = 2 biological replicates using different target sites.
[0058] FIGs. 3A-G. Characterization of the PAM requirements of PAMmla-predicted enzymes, (a) Workflow to generate and characterize PAMmla-predicted enzymes, including prediction of the PAM profiles for all 64 million enzymes in the SpCas9(6AA) library, and then sorting the predictions by predicted absolute activity and selectivity for each of the 16 NGNN PAMs and each of the 4 three-nucleotide PAMs (NGAN, NGCN, NGGN, NGTN). Up to the top 10 predicted enzymes for each PAM class (sorted for activity and / or selectivity) for a total of 281 enzymes were cloned into a human cell expression plasmid and PAM profiles were experimentally determined by HT-PAMDA. (b) Experimentally determined HT-PAMDA PAM profiles for 253 PAMmla predicted enzymes. Enzymes with no k > 10'4are not shown. HT-PAMDA profiles were clustered hierarchically and amino acid enrichment for the 10 largest clusters are shown in sequence logos (right). HT-PAMDA datasets are the mean of n = 2 biological replicates using different target sites, (c) Correlation between PAMmla predicted ks and experimentally determine ks (by HT-PAMDA) for 281 PAMmla predicted enzymes profiled by HT-PAMDA. (d) Distribution of enzyme amino acid hamming distances from the training set amongst the PAMmla predicted enzymes profiled by HT-PAMDA. (e) Categorization of PAMmla-predicted enzymes as shown in panel b; inactive enzymes had no k > 10'4as determined by HT-PAMDA. (f) Distribution of SpCas9 enzymes based on their maximal HT-PAMDA determined ks, with enzymes from 3 categories: randomly chosen from the SpCas9(6AA) library, evolved through a bacterial selection (the k plotted for each enzyme is its maximum rate constant on its most active PAM, and top PAMmla predicted enzymes to maximize activity on a PAM (the k for each enzyme on the PAM maximized by PAMmla). (g) Fraction of PAM variant enzymes maximally active against the specific NGNN PAM that they were selected / predicted against (rank = 1st) or where the PAM selected against was within the top 4 most active PAMs (rank = 2nd-4th), as determined by HT-PAMDA. The three categories of enzymes analyzed were: PAMmla enzymes by maximizing activity on the 16 NGNN PAMs, PAMmla enzymes by sorting for selectivity for each of the 16 NGNN PAMs, and enzymes from bacterial selections on each of the 16 NGNN PAMs.
[0059] FIGs. 4A-D. Genome editing in human cells with PAMmla-predicted enzymes as nucleases or base editors, (a) PAMmla predicted ks for NGNN PAMs for enzymes targeting one of seven distinct PAM categories; NGAT, NGCM, NGCN, NGTN, NGTC, NGDC, and NGTG. Hamming distances to the most similar enzyme in the training set are indicated in parentheses for each enzyme, (b) Nuclease-mediated genome editing efficiencies for each of the enzymes in panel a at endogenous target sites in HEK 293 T harboring the PAMs that they are predicted by PAMmla to target. Editing efficiencies are shown for enzymes from the training set (hamming distance = 0, shown with blue dots), enzymes predicted by PAMmla (shown in pink), as well as SpG (gray) and wild type SpCas9 (white). Editing efficiency was assessed by targeted amplicon sequencing and analyzed using CRISPResso284; data points are the mean of n = 3 biological replicates; between 3 and 10 distinct genomic target sites were selected for characterization, where the black line represents median editing across all target sites for that enzyme; the results for editing at individual loci are shown in FIG. 25. For each of the seven PAM classes, the base editing efficiencies of one enzyme was tested in the context of ABE8e54and TadCBEd55construct architectures (panel c and panel d, respectively). Base editing efficiencies were measured for each enzyme at 3 distinct endogenous target sites in HEK 293T cells harboring relevant PAMs. Data are plotted from n = 3 biological replicates; editing efficiencies were assessed by targeted amplicon sequencing; all edits at bases where any enzyme was observed to edit >5% efficiency are shown. Box plots show 25th, 50th, and 75thpercentiles; whiskers represent the range of the data. The results for A-to-G and C-to-T base editing at individual loci are shown in FIGs. 28A-H.
[0060] FIGs. 5A-G. Genome-wide analysis of SpCas9 enzyme specificity, (a)
[0061] Modification of endogenous on-target sites in HEK 293T cells from GUIDE-seq2 transfections containing the dsODN tag. Percent modification assessed by targeted sequencing; mean, SD, and individual data points shown for n = 3 technical replicates, (b) Normalized fraction of off-target sites detected by GUIDE-seq2 for PAMmla predicted enzymes, SpG, or SpRY, with normalized to the number of off- target sites for SpRY. (c, d) Fraction of GUIDE-seq2 reads attributed to on- and off- target sites for PAMmla predicted enzymes, SpG, and SpRY when using gRNAs targeted to sites with few or many off-target sites (panels c and d, respectively), (e) Schematic of the genomic region flanking the CYBB T326I mutation with the A8 sgRNA target site encoding an NGAT PAM shown, with intended edit position and bystander edit labeled in blue or red numbering, respectively, (f). Base editing efficiencies to correct the CYBB T326I mutation in a patient-derived B cell line. Base editing of the target and bystander base assessed by targeted sequencing; mean, SD, and individual data points shown for n = 3 independent biological replicates; all bases edited at >1% efficiency are shown, (g) Fraction of GUIDE-seq2 reads attributed to on- and off-target sites for the PAMmla predicted enzyme KWRQLC and the SpG variant in GUIDE-seq2 experiments using the CYBB T326I A8 sgRNA, but targeted to the WT allele of HEK 293T cells (see FIGs. 31d,e).
[0062] FIGs. 6A-I. In silico directed evolution of an allele-specific editor for the blindness-causing RHO P23H allele, (a) Schematic of the wild-type (WT) RHO and mutant P23H alleles, where a sgRNA target site positions the PAM over the P23H mutation (NGGG and NGTG PAMs on WT and mutant alleles, respectively), (b) Predicted PAM profiles of enzymes resulting from PAMmla-enabled ISDE, using WT SpCas9 a starting enzyme and seeking to maximize activity on NGTG while minimizing on NGGG (k < 10'3 7). Only the evolutionary trajectory leading to MRRWMR, is shown (see FIG. 39a for more complete evolution), (c) Predicted PAM profiles of ISDE enzymes starting from WT SpCas9, maximizing activity on NGTG, and minimizing on NGGG (k < 1 O'4). Only the evolutionary trajectory leading to KRHWMR, is shown (see FIG. 39b for more complete evolution), (d) Fitness function used to perform the in silico directed evolution experiment in panel b; Fitness of the top 10 enzymes from four rounds of evolution are shown and the evolutionary trajectory leading to MRRWMR is shown with connected lines, (e) Fitness function used to perform the in silico directed evolution experiment in panel c; Fitness of the top 10 enzymes from three rounds of evolution are shown and the evolutionary trajectory leading to KRHWMR is shown with connected lines. Only evolution rounds producing a new top variant are shown, (f) Modification of the WT RHO and P23H alleles in a heterozygous HEK 293 T cell line (harboring a 2: 1 P23HP23 allele ratio), comparing PAMmla evolved enzymes to WT SpCas9 and SpG (see also FIG. 34b). Editing assessed by targeted sequencing and CRISPResso2; mean and s.d. shown for n = 3 biological replicates; for reads containing indels that span the P23H mutation, edited counts were distributed using the ratio of WT to mutant as observed for the identifiable edited reads. Fractions of unidentifiable reads shown in FIG. 34d. (g) Fraction of GUIDE-seq2 reads attributed to on- and off-target sites for PAMmla predicted enzymes, SpG, and SpRY when paired with the RHO P23H sgRNA in homozygous P23H HEK 293T cells (see also FIGs. 35d,e). (h, i Mutations shared between MRRWMR and KRHWMR modelled on the structure of VRER (PDB: 5FW3)52interacting with NGTG or NGGG PAMs (panels h and i, respectively). Protein surface is colored by lipophilicity potential. Hydrogen bonds are represented by dashed lines and Van der Waals interactions are represented by green squiggles, (j, k) Force plots depicting SHapely Additive exPlanations51(SHAP) values for MRRWMR activity on NGTG or NGGG PAMs (panels j and k, respectively). (1) In vivo modification of the RHO P23H or WT alleles in heterozygous humanized P0-P2 mouse pups (harboring lx WT and lx P23H alleles). Plasmids encoding Cas9-P2A-mTagBFP2 and the human RHO sgRNA were subretinally injected and electroporation was performed, retinas were extracted, cells were sorted for BFP+ after 15-24 days, and editing was assessed by targeted sequencing and CRISPResso2. Mean and s.d. shown for n= rl, 10, and 4 mice injected with KRHWMR, MRRWMR, or SpG respectively; unidentifiable reads containing indels that span the P23H mutation were discarded; fractions of unidentifiable reads shown in FIG. 35g). FIG. 7 is a block diagram for an example method for generating enzymes of desired functionality in accordance with implementations of the present disclosure.
[0063] FIG. 8 is a block diagram for an example method for computationally generating SpCas9 variant enzymes with customizable properties in accordance with implementations of the present disclosure.
[0064] FIG. 9 is a block diagram for automatic characterization of generated SpCas9 variant enzymes in accordance with implementations of the present disclosure.
[0065] FIG. 10 is a schematic diagram of an example computing system 1000.
[0066] FIG. 11. Heatmaps representing the experimentally determined PAM profiles of SpCas9 enzyme variants from the SpCas9(6AA) library, where the enzymes are derived from Table C (which includes enzymes derived from bacterial selection assays, where the HT-PAMDA rate constant, k, was > 10'3for at least one PAM). The PAM profiles were determined by the HT-PAMDA assay; the legend represents the logw rate constants (k which are the mean of n = 2 replicate HT-PAMDA experiments performed using two distinct spacer sequences.
[0067] FIG. 12. Heatmaps representing the experimentally determined PAM profiles of ‘consensus’ SpCas9 enzyme variants from the SpCas9(6AA) library following bacterial selections, where the enzymes are derived from Table D (for enzymes where the HT-PAMDA rate constant, k, was > 10'3for at least one PAM). The PAM profiles were determined by the HT-PAMDA assay; the legend represents the logw rate constants (k which are the mean of n = 2 replicate HT-PAMDA experiments performed using two distinct spacer sequences.
[0068] FIG. 13. Heatmaps representing the experimentally determined PAM profiles of SpCas9 enzyme variants predicted using the PAMmla ML model, where the enzymes are derived from Table E (for enzymes where the HT-PAMDA rate constant, k. was > 10'3for at least one PAM). The PAM profiles were determined by the HT-PAMDA assay; the legend represents the logw rate constants (k which are the mean of n = 2 replicate HT-PAMDA experiments performed using two distinct spacer sequences. FIG. 14. Heatmap representations of the PAM profiles of various SpCas9 enzyme variants, predicted using the PAMmla ML model. The PAMmla-predicted PAM profiles for the top ten enzymes for each enzyme class from Table A are shown.
[0069] FIGs. 15A-D. Targeting range and characterization of previous engineered SpCas9 PAM variant enzymes, (a) Quantification of pathogenic and likely pathogenic single nucleotide variants (SNVs; from ClinVar89) that are theoretically revertible using ABE or CBE based on their proximity to an NGG PAM. SNVs were considered editable if a GG dinucleotide PAM was available at the appropriate distance upstream on the correct DNA strand, positioning the SNV anywhere between positions 5-9 of the spacer sequence (counting from the PAM-distal end of the spacer; typically called the ‘edit window’ of base editors), (b-d) Heatmap representations of the PAM profiles of SpCas9 enzymes determined using the HT-PAMDA assay41, for wild-type SpCas9 (panel b), for enzymes with altered PAM requirements (e.g., SpCas9-VRQR10,88, SpCas9-VRER10, and xCas912; panel c), and for enzymes with relaxed PAM requirements (e.g. SpCas9-NG14, SpG and SpRY11, and SpCas9- NR.RH / NR.CH / NRTH13); panel d). The logio rate constants Qi) are the mean of n = 2 replicate HT-PAMDA experiments performed using two distinct spacer sequences. Because the HT-PAMDA assay measures the relative depletion of substrates encoding various PAMs, it may underestimate rate constants for enzymes with highly relaxed PAM requirements such as SpRY and Cas9-NRRH41.
[0070] FIGs. 16A-D. Structure-informed saturation mutagenesis and bacterial positive selections for SpCas9 PAM variant enzymes, (a) Structural representation of the PAM-interacting (PI) domain of SpCas9 showing amino acid residues interacting with a canonical NGG PAM (from PDB ID: 4UN36). (b) Schematic of the bacterial positive selection assay. A plasmid encoding the SpCas9(6AA) library (with randomized NNS codons at SpCas9 positions DI 135, S 1136, G1218, E1219, R1335, and T1337), a gRNA expression cassette, and chloramphenicol resistance gene is transfected into an E. coli strain harboring a selection plasmid encoding an inducible toxic gene and the Cas9 target site (with protospacer adjacent to a non-canonical 4 nt PAM of interest). Selections were performed similar to previously described10,47’70, where the ccdB gene (encoding a DNA gyrase toxin) on the selection plasmid is induced by plating on arabinose-containing media. Bacterial colonies survive the selection when they harbor a plasmid that expresses an SpCas9 enzyme variant capable of cleaving the selection plasmid (by recognizing a non-canonical PAM), (c) Summary of the SpCas9 enzymes that survived the bacterial positive selections using selection plasmids encoding each of the 16 NGNN PAMs. The heatmaps depict the percent of SpCas9 enzymes from each of the 16x selections that contain each possible amino acid substitution at each of the six SpCas9(6AA) library positions. Each heatmap is labeled based on the PAM utilized in that set of bacterial selections; the number of enzymes selected from each set of selections is indicated, (d) Summary of the composition of amino acid residues at each of the six positions of the SpCas9(6AA) library.
[0071] FIG. 17. Workflow for a whole-ORF sequencing approach. Schematic depicting the workflow used to perform tiled-amplicon sequencing across the entire coding sequences of SpCas9-P2A-EGFP expression plasmids. This inexpensive sequencing method can be performed to obtain complete sequence validation of up to 96 SpCas9 enzyme variants in parallel, enabling linkage of amino acid sequence to the specific plasmid based on well position.
[0072] FIGs. 18A-C. PAM preferences of SpCas9 enzymes with relaxed NGN PAM requirements. (a,b) Heatmap representations of the PAM profiles of SpCas9 enzymes determined using the HT-PAMDA assay, for the previously described SpG enzyme (panel a), and for newly enzymes with NGN PAM requirements (panel b). SpCas9 variant enzymes are named based on their amino acids at each of the six positions in the SpCas9(6AA) library. The logio rate constants (k) are the mean of n = 2 replicate HT-PAMDA experiments using two distinct spacer sequences, (c) Nuclease-mediated genome editing at six endogenous target sites in HEK 293T cells harboring different PAMs. Editing efficiency was assessed by targeted amplicon sequencing and modified reads were analyzed using CRISPResso284; mean, standard deviation, and individual datapoints shown for n = 3 independent biological replicates.
[0073] FIGs. 19A-B. Necessity of R1335 for NGG PAM recognition. (a,b) Heatmap representations of the PAM profiles of SpCas9 enzyme variants determined using the HT-PAMDA assay, for SpCas9 enzymes obtained through bacterial selections that recognize NGG PAMs but do not contain R1335 (panel a), and for SpCas9 enzymes obtained through bacterial selections that contain R1335 but do not recognize or are not selective for NGG PAMs (panel b). SpCas9 variant enzymes are named based on their amino acids at each of the six positions in the SpCas9(6AA) library. The logio rate constants (k) are the mean of n = 2 replicate HT-PAMDA experiments using two distinct spacer sequences.
[0074] FIGs. 20A-C. Rational design of consensus SpCas9 enzymes resulting from the 16x bacterial selections using NGNN PAMs. (a) Comparison of most preferred PAM for each bacterial selected enzyme (maximum rate constant determined by HT- PAMDA) compared to the PAM on which each enzyme was selected for during the bacterial -based selection experiments, (b) Heatmap representations of the PAM profiles of SpCas9 ‘consensus enzymes’ determined using the HT-PAMDA assay. Consensus enzymes were rationally designed based on bacterial selection data from each NGNN PAM by combining the most enriched amino acid substitutions at each SpCas9(6AA) library position (see also Fig. 16c). If two amino acids were enriched approximately equally, multiple SpCas9 enzymes were cloned and tested. The logio rate constants (k) are the mean of n = 2 replicate HT-PAMDA experiments using two distinct spacer sequences, (c) Comparison of most preferred PAM for each enzyme (maximum rate constant determined by HT-PAMDA) compared to the PAM that the consensus enzymes were rationally designed to target. For panels a and c, outlined squares indicate the diagonal where there would be agreement between the most efficiently targeted PAM for the enzyme from HT-PAMDA data compared to the selection PAM (panel a) or the PAM for the consensus enzyme (panel c). For panels a-c, SpCas9 variant enzymes are named based on their amino acids at each of the six positions in the SpCas9(6AA) library.
[0075] FIGs. 21A-I. Machine learning models to predict PAM profile from amino acid sequence, a, Comparison of machine learning model architectures (linear regression, random forest, and neural network) and amino acid encodings (one-hot, one-hot plus all pairwise amino acid combinations, and Georgiev106). The R2value is shown between the experimentally determined k (via HT-PAMDA) and the predicted k (via each ML model) for an internal 5-fold cross-validation on the training set. Each validation set is sub-divided according to the minimum hamming distance (HD) of each variant to the nearest neighbor in the corresponding training set; thus, validation sets become more challenging as HD increases, b, Performance of the optimal PAM machine learning algorithm (PAMmla; comprised of a neural network with one hot encoding) on two additional 80% / 20% random train-test splits, c, Proportion of test set SpCas9 enzymes that have a predominant preference for A, C, G, T in the 3rdposition of the PAM, or are inactive (based on HT-PAMDA data), d, Comparison of test set ks broken down by nucleotide preference of each test variant at the 3rdposition of the PAM (comparing ks experimentally determined by HT-PAMDA versus predicted by PAMmla). Nucleotide preference is defined as the 3rdposition nucleotide of each enzyme variant’s most preferred PAM by HT-PAMDA. e, Proportion of test set SpCas9 enzymes that have a predominant preference for A, C, G, T in the 4thposition of the PAM, or are inactive (based on HT-PAMDA data), f, Comparison of test set ks broken down by preference of each test set variant at the 4thposition of the PAM (comparing ks experimentally determined by HT-PAMDA versus predicted by PAMmla). Nucleotide preference is defined as 4thposition nucleotide of each enzyme variant’s most preferred PAM by HT-PAMDA. g, Effect of random over-sampling by most active PAM. The PAMmla model was trained with and without randomly over- sampling the training set to balance the number of enzyme variants with different PAM preferences. R2values for the two models were compared on subsets of variants within the test set with different preferences at the 3rdand 4thpositions of the PAM. Over-sampling improved performance particularly for under-represented PAM classes (see panels c and e). h, Pearson’s correlations between HT-PAMDA replicates performed with distinct spacer sequences for a set of inactive versus active enzymes within the test set. True labels for active versus inactive enzymes were determined using a cutoff value for maximum k on any PAM of 10'4 3. Enzymes separated into active and inactive classes based on these criteria showed correlation between replicates only for active enzymes, indicating HT-PAMDA data for enzymes with maximum ks below this cutoff are likely due to non-reproducible noise in the HT- PAMDA assay, i, Correlation between ks experimentally determined by HT-PAMDA versus predicted by PAMmla for inactive variants (maximum HT-PAMDA k < IO'4,3) within the test set; PAMmla is not predictive for background noise in the HT-PAMDA determined PAM profiles of inactive enzymes. For all panels that utilize HT-PAMDA data, the logio rate constants (k) are the mean of n = 2 replicate HT-PAMDA experiments using two distinct spacer sequences. For all scatterplots, each datapoint represents the rate constant activity of one enzyme variant against on one of 64 possible NNNN PAMs.
[0076] FIG. 22. PAMmla predictions for 64 million Cas9 variants. PAMmla predicted PAM profiles for all possible 64 million Cas9 enzymes altered at amino positions 1135, 1136, 1218, 1219, 1335, and 1337 were generate and downsampled lOOx using scSampler81to preserve diversity. Enzymes were visualized using UMAP82based on predicted similarity in PAM preference (correlation of 64 PAMmla predicted rate constants). Points are colored according to the amino acid present at each of the 6 varied positions.
[0077] FIG. 23. Most efficiently targeted PAMs for top PAMmla predicted enzymes. Plot illustrating the proportion of PAMmla predicted enzymes where the most active PAM (ks experimentally assessed via the HT-PAMDA assay) matches the PAMmla predicted maximum k for that enzyme. Outlined squares in the plot indicate the diagonal where there would be agreement between the most efficiently targeted PAM for the enzyme from HT-PAMDA data compared to the PAMmla prediction.
[0078] FIGs. 24A-B. Comparison of experimental and predicted PAM profiles.
[0079] Comparison of PAMmla predicted rate constants (a) to HT-PAMDA determined rate constants (b) for variants chosen for human cell validation in figure 4.
[0080] FIGs. 25A-H: Genome editing in human cells with PAMmla predicted enzymes.
[0081] (a-g) Nuclease-mediated genome editing at endogenous target sites in HEK 293 T cells harboring different PAMs for wild-type (WT) SpCas9, SpG, and PAMmla predicted enzymes capable of targeting various classes of PAMs (NGAT, NGCM, NGCN, NGTN, NGTC, NGDC, and NGTG in panels a-g, respectively). SpCas9 variant enzymes are named based on their amino acids at each of the six positions in the SpCas9(6AA) library. Enzymes resulting from bacterial selections are colored in shades of blue (hamming distance (HD) of 0) and PAMmla predicted enzymes are shown in shades of red (HD for each enzyme indicated in parentheses). The site encoding an NGGC PAM was used as a negative control for editing with the PAMmla predicted enzymes. Editing efficiency was assessed by targeted amplicon sequencing and modified reads were analyzed using CRISPResso2; mean, standard deviation, and individual datapoints shown for n = 3 independent biological replicates, (h) Summary of editing efficiencies against the NGGC PAM control site in HEK 293T cells when performing experiments with PAMmla predicted enzymes amongst the classes tested in panels a-d and g (some of which are mostly predicted to exhibit low editing on NGGG PAMs).
[0082] FIG. 26. PAMmla feature importance for enzymes targeting different PAM classes. SHapely Additive exPlanations (SHAP)96analysis to investigate the impact of amino acid substitutions (i.e. PAMmla features) on model output for each of the 16 NGNN PAMs. SHAP values are shown for 200 enzymes sampled from the training set. Top 10 features with highest mean absolute SHAP values (greatest absolute impact on model output) are plotted for each PAM.
[0083] FIGs. 27A-I. Homology models of PAMmla predicted PAM-altering mutations, a, An E1219Y substitution may facilitate interaction with the amino group of bases in the 3rdposition of the PAM. b, R1335Q permits major groove readout of both bases of a C-G pair in the 3rdposition of the PAM. c, E1219C, R1335M, and T1337V substitutions form a hydrophobic pocket to promote van der Waals interactions with the methyl group of thymine in the 3rdposition of the PAM. Representation of the protein surface is colored by lipophilicity potential, d, T1337R results in direct major groove readout of guanine in the 4thposition of the PAM. e, T1337K facilitates major groove readout of oxygen group of bases in the 4thposition the PAM. f, R1335L and T1337C substitutions form a hydrophobic pocket to promote recognition of thymine in the 4thposition of the PAM. Protein surface is colored by lipophilicity potential, g, D1135L disrupts coordination with R1114, enabling improved flexibility of the R1114 side chain to contact the NTS backbone. WT SpCas9 is overlaid in grey, h, Substitution of G1218 to a positive residue establishes additional non-specific contacts with the NTS backbone, i, S1136W and D1135L result in a shift of the NTS and TS backbone towards the PAM-interacting domain, enabling novel base specific interactions in nearby regions. WT SpCas9 is overlaid in grey. For panels a-i, amino acid and PAM DNAbase substitutions were modeled on the structure of SpG (PDB: 8U3Y)17using Coot92, except for substitutions T1337R, T1337K, and T1337C which were modeled using SpCas9-VRER (PDB: 5FW3)52. Homology models were visualized using ChimeraX93.
[0084] FIGs. 28A-H. A-to-G base editing with PAMmla predicted enzymes, (a-g) A-to-G base editing efficiencies using ABE8e54enzymes for genome editing at endogenous target sites in HEK 293T cells harboring different PAMs for SpG, SpRY, and PAMmla predicted enzymes capable of targeting various classes of PAMs (NGAT, NGCM, NGCN, NGTN, NGTC, NGDC, and NGTG in panels a-g, respectively). SpCas9 variant enzymes are named based on their amino acids at each of the six positions in the SpCas9(6AA) library. Editing efficiency was assessed by targeted amplicon sequencing and modified reads were analyzed using CRISPResso2; mean, standard deviation, and individual datapoints shown for n = 3 independent biological replicates. Datapoints for all adenine bases in the ABE8e edit window (positions 4-9 of the spacer, counting from the PAM distal end) or any other adenine in the spacer sequence edited >5% are shown, (h) Summary of ABE8e base editing efficiencies across all genomic target sites from panels a-g at all adenines within the edit window or those for which any enzyme edited >5%, for the PAMmla predicted enzymes, SpG, and SpRY. The horizontal black line shows mean editing efficiency for each enzyme class.
[0085] FIGs. 29A-H. C-to-T base editing with PAMmla predicted enzymes, (a-g) C-to-T base editing efficiencies using TadCBEd55enzymes for genome editing at endogenous target sites in HEK 293T cells harboring different PAMs for SpG, SpRY, and PAMmla predicted enzymes capable of targeting various classes of PAMs (NGAT, NGCM, NGCN, NGTN, NGTC, NGDC, and NGTG in panels a-g, respectively). SpCas9 variant enzymes are named based on their amino acids at each of the six positions in the SpCas9(6AA) library. Editing efficiency was assessed by targeted amplicon sequencing and modified reads were analyzed using CRISPResso2; mean, standard deviation, and individual datapoints shown for n = 3 independent biological replicates. Datapoints for all cytosine bases in the TadCBEd edit window (positions 4- 9 of the spacer, counting from the PAM distal end) or any other cytosine in the spacer sequence edited >5% are shown, (h) Summary of TadCBEd base editing efficiencies across all genomic target sites from panels a-g at all cytosines within the edit window or those edited by any enzyme >5%, for the PAMmla predicted enzymes, SpG, and SpRY. The horizontal black line shows mean editing efficiency for each enzyme class.
[0086] FIGs. 30A-C. Base editing of sickle cell HBB E7V mutation with ML-derived PAM variants, (a) A to G editing of HBB E7V target adenine using PAMmla-derived editors (with the A9 gRNA encoding a spacer adjacent to TGCA PAM) versus SpCas9-NRCH13(using either A9 gRNA encoding a spacer adjacent to the TGCA PAM or previously reported A7 gRNA with spacer adjacent to the CACC PAM) (b) A to G editing on full edit window of the gRNA positioning target base at A7 (CACC PAM) (c) A to G editing on full edit window of the gRNA positioning target base at A9 (TGCA PAM). SpCas9 variant enzymes are named based on their amino acids at each of the six positions in the SpCas9(6AA) library. Editing efficiency was assessed by targeted amplicon sequencing and modified reads were analyzed using CRISPResso2; mean, standard deviation, and individual datapoints shown for n = 3 independent biological replicates.
[0087] FIGs. 31A-E. Genome-wide off-target analysis of PAMmla predicted enzymes, a, Quantification of GUIDE-seq2 double-stranded oligodeoxynucleotide (dsODN) tag integration at the on-target site, in nuclease-based experiments with SpG, SpRY, and PAMmla predicted enzymes targeting endogenous target sites in HEK 293T cells. SpCas9 variant enzymes are named based on their amino acids at each of the six positions in the SpCas9(6AA) library. dsODN integration efficiency was assessed by targeted amplicon sequencing and modified reads were analyzed using CRISPResso2; mean, standard deviation, and individual datapoints shown for n = 3 technical replicates, b, Venn diagram representations of the GUIDE-seq2 detected off-target sites that are shared between or unique to PAMmla generated, SpG, and SpRY nucleases, c, Nucleotide composition of PAMs adjacent to off-target spacers detected in GUIDE-seq2 experiments, not including the on-target reads. The y-axis represents the fraction of total off-target GUIDE-seq2 reads containing each nucleotide at each position of the PAM. d, Quantification of GUIDE-seq double-stranded oligodeoxynucleotide (dsODN) tag integration at the on-target site, in nuclease-based experiments with KWRQLC and SpG when using the CYBB T362I sgRNAbut targeting the wild-type genome of HEK 293 T cells. SpCas9 variant enzymes are named based on their amino acids at each of the six positions in the SpCas9(6AA) library. dsODN integration efficiency was assessed by targeted amplicon sequencing and modified reads were analyzed using CRISPResso2; mean, standard deviation, and individual datapoints are shown for n = 3 technical replicates, e, GUIDE-seq2 genome-wide specificity outputs for KWRQLC and SpG nucleases using the CYBB T362I targeted sgRNA; note that HEK 293T cells harbor the wild-type copy of the CYBB gene and are therefore an imperfect match to the sgRNA. Mismatched positions in the spacers of the off-target sites are highlighted in color; GUIDE-seq read counts from consolidated unique molecular events for each variant are shown to the right of the sequence plots.
[0088] FIGs. 32A-D. Design and validation of in silico directed evolution, a, Schematic of in silico directed evolution (ISDE) pipeline to rapidly identify bespoke SpCas9 enzymes with user-specifiable PAM profiles, b-d, Effect of ISDE parameter values on the identification of optimized PAMmla predicted enzymes, including varying the number of starting mutations per round (m) (panel b), random variants generated per round (panel c) and number of additional evolution rounds performed once a plateau is reached before decreasing m (panel d). Proof-of-concept PAMmla-ISDE runs were performed to identify enzymes with maximal activity against NGAT, NGCC, or NGTA PAMs. Aside from the parameter being tested, ISDE was run with default parameters of 1,000 random starting sequences, m = 4 starting mutations per enzyme, 5 = 1,000 sampled enzymes per round, n = 10 top variants to keep per round, and p = 1 additional round of evolution after a plateau is reached. The number of true top 10 predicted enzymes, determined by exhaustive sorting of PAMmla predictions, recovered by ISDE are shown. Top bar graphs represent the number of replicates in which the most optimal enzyme was recovered.
[0089] FIG. 33. MRRWMR evolved from SpG. Results of a PAMmla-enabled in silico directed evolution experiment starting with SpG (LWKQQR amino acids within the SpCas9(6AA) library), aiming to keep activity on NGTG PAMs above 10'2while minimizing activity on NGGG PAMs. Top 4 variants are kept per round of evolution. Two mutations are allowed per variant during mutagenesis. Predicted ks are shown for the top 10 enzymes from each selection round. The evolutionary trajectory of the enzyme with the highest fitness, MRRWMR, is highlighted in red.
[0090] FIGs. 34A-E. Characterization of PAMmla-ISDE generated enzymes in human cells, a, Nuclease-mediated genome editing at endogenous target sites in HEK 293T cells harboring different PAMs for wild-type (WT) SpCas9, SpG, and MRRWMR. b, Nuclease-mediated genome editing of the wild-type RHO or mutant RHO P23H alleles in a heterozygous RHO P23H HEK 293 T cell line using wild-type SpCas9, SpG, and various PAMmla generated enzymes. For reads containing indels that span the P23H mutation (and therefore could not be identified as WT or mutant), counts were distributed between WT and mutant alleles with the same ratio as WT:mutant ratio observed for the identifiable edited reads, c, Nuclease-mediated genome editing of the RHO target site in wild-type HEK 293T cells using wild-type SpCas9, SpG, and various PAMmla generated enzymes, d, Unidentifiable sequencing reads that were either P23H or WT due to deletions spanning the mutation for data shown in heterozygous P23H HEK 293T cells from data in Fig. 6f; edited reads were distributed based on the balance in identifiable reads, e, Ratio of editing efficiencies observed on mutant (P23H) versus WT RHO alleles, for each editor tested in Fig. 6f. Editing efficiencies in panels a-c,e were assessed by targeted amplicon sequencing and modified reads were analyzed using CRISPResso2; mean, standard deviation, and individual datapoints shown for n = 3 independent biological replicates.
[0091] FIGs. 35A-H. Specificity assessment of PAMmla-derived enzymes MRRWMR and KRHWMR. a, Quantification of GUIDE-seq2 double-stranded oligodeoxynucleotide (dsODN) tag integration at on-target sites in nuclease-based experiments with MRRWMR, SpG, and SpRY and sgRNAs targeting two different endogenous sites in HEK 293T cells. SpCas9 variant enzymes are named based on their amino acids at each of the six positions in the SpCas9(6AA) library. dsODN integration efficiency was assessed by targeted amplicon sequencing and modified reads were analyzed using CRISPResso2; mean, standard deviation, and individual datapoints are shown for n = 3 technical replicates, b, Venn diagram representations of the GUIDE-seq2 detected off-target sites that are shared between or unique to MRRWMR, SpG, and SpRY nucleases using the two sgRNAs targeted to sites with NGTG PAMs (similar to the RHO P23H on-target site), c, Fraction of GUIDE-seq2 reads attributed to on- and off-target sites for MRRWMR, SpG, and SpRY from experiments using the NGTG-2 or NGTG-3 sgRNAs. d, Quantification of GUIDE- seq2 dsODN tag integration at the on-target site for experiments in the homozygous RHO P23H cell line, when using the RHO P23H sgRNA and SpCas9-MRRWMR and -KRHWMR, SpG, and SpRY expression plasmids. dsODN integration efficiency was assessed by targeted amplicon sequencing and modified reads were analyzed using CRISPResso2; mean, standard deviation, and individual datapoints are shown for n = 3 technical replicates, e, GUIDE-seq2 genome-wide specificity outputs for SpCas9- MRRWMR and -KRHWMR, SpG, and SpRY nucleases using the RHO P23H targeted sgRNA in homozygous RHO P23H HEK 293T cells. Mismatched positions in the spacers of the off-target sites are highlighted in color; GUIDE-seq read counts from consolidated unique molecular events for each variant are shown to the right of the sequence plots, f, Venn diagram representation of the GUIDE-seq2 detected off- target sites that are shared between or unique to SpCas9-MRRWMR and -KRHWMR, SpG, and SpRY nucleases using the RHO P23H sgRNA. g, Unidentifiable sequencing reads unattributable to either WT or P23H alleles due to deletions spanning the base harboring the mutation, for data from heterozygous RHO P23H mice shown in Fig. 61. h, Ratio of in vivo editing efficiencies observed on mutant (P23H) versus WT RHO alleles, for each SpCas9 nuclease tested in Fig. 61.
[0092] FIGs. 36A-E. Analysis of factors contributing to MRRWMR and KRHWMR PAM preferences, a, Structural prediction of an alternative conformation of the S1136R mutation leading to additional hydrogen bonding with T at position 3 of the PAM. b-e, SHapely Additive exPlanations51(SHAP) values for PAMmla predictions for MRRWMR (panels b,c) and KRHWMR (panels d,e) interacting with NGTG (panels b,d) or NGGG PAMs (panels c,e) PAMs. Feature values are shown in gray (1 : mutation is present, 0: mutation is absent). Red represents features with positive impact on predicted rate constant and blue represent features with negative impact on predicted rate constant.
[0093] FIGs. 37A-C. in silico directed evolution of a PAM variant for multiplex base editing of the APOE-e4 Alzheimer’s risk allele, (a) Results of in silico directed evolution starting with the wild type SpCas9 sequence (DSGERT), aiming to maximize activity on NGTAPAMs while keeping activity on NGGC high (above 10"2 5). Predicted rate constants are shown for the top 10 variants surviving each round of selections, (b) Fitness function used to perform directed evolution in a. Fitness of each of the top predicted variants at each round of evolution is shown, (c) Multiplex cytosine base editing in a homozygous HEK 293 T APOE-e4 cell line. Transfections were performed with various PAM-variant cytosine base editors (TadCBEd55, BE491, evoFERNY69), and two simultaneous gRNAs targeting the Argl30 and Argl76 mutations. Target C for each edit is indicated with *. SpCas9 variant enzymes are named based on their amino acids at each of the six positions in the SpCas9(6AA) library. Editing efficiency was assessed by targeted amplicon sequencing and modified reads were analyzed using CRISPResso2; mean, standard deviation, and individual datapoints shown for n = 3 independent biological replicates.
[0094] FIGs. 38A-B. NGTA- and NGGC -targeting PAMmla variants tested as CBEs. (a,b) Cytosine base editing at an endogenous HEK 293 T sites bearing an NGGC PAM (a) or an NGTA PAM (b). Different cytosine base editors (CBEs) were selected for testing including BE4max91, evoFERNY69, and TadCBEd55. SpCas9 variant enzymes are named based on their amino acids at each of the six positions in the SpCas9(6AA) library. Editing efficiency was assessed by targeted amplicon sequencing and modified reads were analyzed using CRISPResso2; mean, standard deviation, and individual datapoints shown for n = 3 independent biological replicates.
[0095] FIGs. 39A-B. in silico directed evolution of PAM-selective SpCas9 enzymes, a, In silico directed evolution (ISDE) trajectory starting from wild-type (WT) SpCas9 and evolving towards maximizing efficiency on sites with NGTG PAMs, while minimizing activity on sites with NGGG PAMs (setting the PAMmla predicted k for NGGG to < 1 O'3 7). The top ten enzymes were kept per round of ISDE and the predicted ks are shown; four mutations were permitted per enzyme during mutagenesis, b, ISDE trajectory starting from WT SpCas9 and evolving towards maximizing efficiency on NGTG PAMs and minimizing activity on NGGG PAMs (k < 1 O'4). The top ten enzymes were kept per round of ISDE and the predicted ks are shown; four mutations were permitted per enzyme during mutagenesis. SpCas9 variant enzymes are named based on their amino acids at each of the six positions in the SpCas9(6AA) library; the evolutionary trajectories of the enzymes with the highest fitness are shown in red. DETAILED DESCRIPTION
[0096] For many CRIPSR-based applications (e.g., allele-specific editing, base editing, modifying regulatory elements, etc.), precise positioning of the Cas enzyme is paramount. Base editing is particularly sensitive to the placement of Cas9 on DNA with single base pair precision. Base editors have an activity window (edit window) which is defined as the range of single stranded DNA bases which are accessible for deamination when Cas9 forms a R-loop at the target site. The base editor edit window varies depending on the Cas enzyme, deaminase, and base editor construct, but for the majority of SpCas9-based base editors, most efficient editing typically occurs at approximately positions 5-8 in the protospacer (counting from the PAM-distal end). Thus, enzyme variants recognizing PAMs other than the canonical NGG are required to access much of the genome. For example, to highlight the challenge of PAM availability, of 47,159 pathogenic SNVs listed on ClinVar which are theoretically revertible using C-to-T or A-to-G base editors (CBEs and ABEs, respectively), only approximately one quarter have appropriately positioned NGG PAMs to position the optimal edit of the base editor over the target base (Fig. 15a). We found that approximately one quarter of annotated SNPs which could be theoretically reverted with either a cytosine or adenine base editor had an appropriately positioned NGG PAM to place the optimal edit window of the base editor over the target base. The total number of base-editable SNPs is likely even lower in practice due to the additional constraint of positioning the editor in such a way as to avoid harmful bystander edits or for the avoidance of off-target edits. Thus, there is a need to develop efficient and selective CRISPR-Cas enzymes for a variety of genome editing applications.
[0097] Previous studies have shown that the PAM requirement of SpCas9 can be modified through amino acid substitutions in the PI domain of SpCas9, leading to enzyme variants capable of recognizing different PAMs10-14. However, the plasticity of the PI domain has not been explored in a systematic manner, and the full extent of accessible PAMs via PI domain modification remains unknown. Engineered SpCas9 PAM variant enzymes can be broadly classified into two categories, including those which shift the PAM preference away from NGG (Figs. 15b, c), and those that expand editing to new PAMs while retaining some activity against NGG (Fig. 15d) (often referred to as “altered” or “relaxed” enzyme variants, respectively). Relatively few SpCas9-based enzyme variants such as VRER and VRQR have been developed which alter the specificity away from NGG to target NGCG or NGAG / NGNG PAMs, respectively (Fig. 15c). A much more common engineering trajectory is the minimization of the PAM requirement towards more generalist enyzmes. For instance, most engineered SpCas9 PAM variants harbor a relaxed PAM requirement, including SpCas9-NG (NGN PAM)14, xCas9 (reported originally as NGN12, later shown to target NGG and NGNC11,15), SpG and SpRY (NGN and NRN>NYN PAMs, respectively11, the SpCas9-NRRH, -NRCH, and -NRTH enzymes13(Fig. 15d). Both selective and relaxed enzyme variants have advantages and disadvantages. Relaxed variants allow target sites with a diversity of PAMs to be targeted with a single enzyme, making them broadly useful for research applications or uses involving multiplexed editing, necessitating compatibility with a wide range of PAMs.
[0098] However, enzymes with expanded PAM requirements typically exhibit increased off- target editing due to a larger fraction of the genome being accessible compared to wild-type (WT) SpCas911 16 17. In the context of base editors, protracted genome searching by PAM-relaxed editors can also lead to increased guide-independent off- target deamination18. Therefore PAM-relaxed editors may be less optimal in certain cases for therapeutic genome editing (Fig. lb). Additionally, since these enzymes are not specialized for a specific PAM of interest, they tend have sub-optimal activity on many PAMs that are within their theoretical targeting range (likely due to genome search kinetics or enzymatic attenuation resulting from engineering)17. Alternatively, P AM-selective variants can enable efficient on-target editing while reducing the potential for genome-wide off-targets. However, the paucity of available SpCas9 PAM variant enzymes does not permit comprehensive genome targeting (Fig. lb).
[0099] Thus, the development of a collection of selective SpCas9 PAM variant enzymes that each targeting a subset of PAMs, but that collectively enable broad genome targeting across a range of genomic sites, would be an optimal solution for therapeutic applications in order to maintain high editing specificity (Fig. lb).
[0100] Protein engineering and directed evolution are powerful approaches to obtain enzymes with modified properties1,2. The relationship between protein function and sequence is complex and selection strategies are often not straightforward, making certain engineering efforts challenging and time consuming. Thus, protein engineering approaches are often targeted to, or converge on, the generation of single generalist enzymes that are broadly useful across many applications. In contrast, the development of a comprehensive catalog of highly specialized enzymes could overcome caveats of generalist enzymes. However, the feasibility, cost, and scale of protein engineering approaches to create such collections is largely prohibitive. Machine learning (ML) has emerged as a powerful tool to augment protein engineering by learning the relationship between protein sequence and function3, reducing experimental time and costs, and increasing the number and diversity of protein sequences that can be interrogated. Thus, high-throughput functional assays combined with ML could reduce protein engineering barriers to move away from generalist enzymes towards the generation of large numbers of highly customized designer proteins.
[0101] Previously, PAM variant enzymes have been developed by rational design and by directed evolution10-14 19-25. Although both methods can produce useful enzymes, they are currently limited in their throughput and typically yield few enzymes per engineering campaign. Furthermore, rational protein design is limited by our knowledge of protein structure, and it is challenging to predict the outcome of many simultaneous mutations. While directed evolution is independent of our understanding of structure or mechanism, it often utilizes large protein libraries where mutations affecting the property of interest are sparse, rarely observed in combination, or limited in mutational scope. Furthermore, engineering typically results in new protein functions via stepwise mutations with the requirement that each intermediate mutant maintains its function. One potential reason that few altered specificity PAM variant enzymes have been developed may be that it is difficult to obtain enzymes that completely switch specificity through directed evolution due to the requirement of maintaining activity throughout the evolution trajectory, preventing access to a broader mutational space. Indeed, all reported relaxed PAM variant enzymes obtained through directed evolution maintain activity on the original PAM, supporting that relaxation may be the simplest engineering trajectory for an enzyme variant to access new a PAM space.
[0102] With sufficient training data, the relationship between protein sequence and function can be approximated using ML26-29, allowing for more rapid generation of novel sequences likely to have desired properties. This approach allows for computational “screening” of larger numbers of amino acid combinations than is possible even by high throughput experimental methods, increasing the probability of acquiring optimal variants across a deeper mutational space. ML has been applied in the context of engineering for various proteins including antibody design30-34, zinc finger DNA binding35-37, and AAV vector capsid tropism38-40among others, but has not been extensively explored for CRISPR-Cas enzyme engineering. A theoretical ML model capable of predicting the PAM preference of a Cas enzyme based on its amino acid sequence would enable systematic exploration of an exponentially large sequence space, accessible only by simultaneously mutating residues involved in PAM interaction. Training such a model on the enzyme variant’s activity across all possible PAMs (rather than a single PAM of interest) should enable the prediction of novel CRISPR-Cas enzymes with bespoke PAM requirements, which have advantages of precise and selective genome targeting.
[0103] Here we undertook an extensive and scalable protein engineering campaign to deeply profile the contribution of six amino acid positions in the PAM interacting (PI) domain of SpCas9 to the PAM requirement. By generating rich biochemical datasets involving kinetic information for hundreds of SpCas9 enzymes against all possible PAMs, this information enabled the training of a PAM machine learning algorithm (PAMmla) to relate amino acid sequence to enzyme function. PAMmla can predict PAM variant enzymes with tunable specificity (both relaxed and selective variants) for a wide range of PAMs, which we test in proof-of-concept experiments against therapeutic targets. As shown herein, the integration of ML with protein engineering methods allowed the development a collection of bespoke PAM-selective enzymes, for optimization of highly safe and effective CRISPR genome editing technologies.
[0104] Design, cloning, and testing of the SpCas9 PI domain libraries.
[0105] To construct a library of enzymes as input for directed evolution to select for SpCas9 enzymes capable of recognizing alternate PAMs, we identified 6 amino acid positions (DI 135, SI 135, G1218, E1219, R1335, and T1337) within the PI domain of SpCas9 for saturation mutagenesis (Fig 16a). Within the context of a wild-type SpCas9:DNA: single-guide RNA (sgRNA) tertiary complex, R1335 forms a base specific contact with the guanine nucleotide in the third position of the PAM6.
[0106] However, altering this amino acid alone is insufficient to change PAM specificity, and mutation of this residue without other nearby substitutions largely abrogates nuclease activity101 10. For the previously described SpCas9 PAM variant enzymes SPCas9- VRER (with DI 135V, G1218R, R1335E, and T1337R substitutions)101, SpCas9- VRQR (D1135V, G1218R, R1335Q, and T1337R)101’88, and SpG (D1135L, S1136W, G1218K, E1219Q, R1335Q, and T1337R)11, additional mutations were necessary at residue positions DI 135, SI 136, G1218, E1219, and T1337. These amino acid side chains do not directly contact the PAM DNA bases, but substitutions were necessary to alter PAM recognition11 10 101 102. Mutations at these positions are thought to be involved in altering the conformation of the PAM DNA backbone, allowing access to DNA bases by amino acids with side chains shorter than that of the wild-type arginine residue, and / or by creating new stabilizing interactions with the DNA backbone to compensate for weaker base-specific PAM interactions101 102. The T1337R mutation has been previously shown to enable a novel base-specific contact at the fourth position of the PAM, an interaction that is not typically observed wild-type SpCas9 and an NGG PAM10,101,102. We therefore included this position in our SpCas9(6AA) library to explore the potential of expanding the PAM from the canonical 3 nts to 4 nts, potentially imparting additional specificity.
[0107] We cloned the SpCas9(6AA) saturation mutagenesis library by first constructing an SpCas9 coding sequence harboring three IIS restriction enzyme cassettes located at the D 1135 / S 1136, G1218 / E 1219, and R1335 / T 1337 positions of SpCas9. The entry vector was then utilized in three sequential cloning steps involving restriction digests and subsequent ligations of oligonucleotide libraries harboring NNS codons. This process could be expedited using newer molecular approaches, including SpRYgests46for digesting the plasmid backbone followed by isothermal assembly71.
[0108] The R1333 residue, which forms a base specific contact with the guanine nucleotide in the second position of the PAM6, was not included in our SpCas9(6AA) library despite structural evidence of its importance for PAM interaction. Since substitutions of this residue likely also require additional compensatory mutations of nearby amino acids to change the preference for G at the second position of the PAM (similar to DI 135, SI 135, G1218, E1219, and T1337 being necessary to support R1335 substitutions that modify the positions of the PAM), we envisioned that a separate library would be necessary to efficiently access PAMs with modified second position nt preferences. Performing bacterial positive selections with our SpCas9(6AA) library on PAM substrates with A, C, or T at the second position of the PAM failed to produce viable enzymes with PAMs altered at the second nucleotide, supporting this hypothesis (FIGs. 11-14). We also explored the possibility of mutating R1333 residue in addition to the 6 residues included in SpCas9(6AA) library to create a SpCas9(7AA) library with the hope of altering the second position of the PAM. However, bacterial selections with the SpCas9(7AA) library on NAN, NCN, and NTN PAMs again failed to yield enzymes with specificity altered at the second position, indicating that additional mutations of other amino acids near the second position of the PAM are likely necessary to support recognition of non-NGNN PAMs. Thus, in the current study, we focused our efforts on altering the third and fourth positions of the PAM.
[0109] Engineered Cas9 Variants with Altered PAM Specificities
[0110] The SpCas9 variants engineered in this study greatly increase the range of target sites accessible by SpCas9. Machine-learning aided design and identification of variants that can target selected PAM containing sites, including variants that can improve activity against previously targetable sites, improve the prospects for accurate and high-resolution genome-editing. The altered PAM specificity SpCas9 variants described herein can efficiently target endogenous gene sites in both bacterial and mammalian, e.g., human, cells.
[0111] All of the SpCas9 variants described herein can be rapidly incorporated into existing and widely used vectors, e.g., by simple site-directed mutagenesis, and because they require only a small number of mutations contained within the PAM-interacting domain, the variants should also work with other previously described improvements to the SpCas9 platform (e.g., truncated sgRNAs (Tsai et al., Nat Biotechnol 33, 187- 197 (2015); Fu et al., Nat Biotechnol 32, 279-284 (2014)), nickase mutations (Mali et al., Nat Biotechnol 31, 833-838 (2013); Ran et al., Cell 154, 1380-1389 (2013)), dimeric FokI-dCas9 fusions (Guilinger et al., Nat Biotechnol 32, 577-582 (2014); Tsai et al., Nat Biotechnol 32, 569-576 (2014)); and high-fidelity variants (Kleinstiver et al. Nature 2016).
[0112] SpCas9 Variants with Altered PAM Specificity
[0113] Provided herein are SpCas9 variants. The SpCas9 wild type sequence is as follows:
[0114] 10 20 30 40 50 60
[0115] MDKKYSIGLD IGTNSVGWAV ITDEYKVPSK KFKVLGNTDR HSIKKNLIGA LLFDSGETAE
[0116] 70 80 90 100 110 120
[0117] ATRLKRTARR RYTRRKNRIC YLQEI FSNEM AKVDDSFFHR LEESFLVEED KKHERHPI FG
[0118] 130 140 150 160 170 180
[0119] NIVDEVAYHE KYPTIYHLRK KLVDSTDKAD LRLIYLALAH MIKFRGHFLI EGDLNPDNSD
[0120] 190 200 210 220 230 240
[0121] VDKLFIQLVQ TYNQLFEENP INASGVDAKA ILSARLSKSR RLENLIAQLP GEKKNGLFGN
[0122] 250 260 270 280 290 300 LIALSLGLTP NFKSNFDLAE DAKLQLSKDT YDDDLDNLLA QIGDQYADLF LAAKNLSDAI
[0123] 310 320 330 340 350 360
[0124] LLSDILRVNT EITKAPLSAS MIKRYDEHHQ DLTLLKALVR QQLPEKYKEI FFDQSKNGYA
[0125] 370 380 390 400 410 420
[0126] GYIDGGASQE EFYKFIKPIL EKMDGTEELL VKLNREDLLR KQRTFDNGSI PHQIHLGELH
[0127] 430 440 450 460 470 480
[0128] AILRRQEDFY PFLKDNREKI EKILTFRI PY YVGPLARGNS RFAWMTRKSE ETITPWNFEE
[0129] 490 500 510 520 530 540
[0130] WDKGASAQS FIERMTNFDK NLPNEKVLPK HSLLYEYFTV YNELTKVKYV TEGMRKPAFL
[0131] 550 560 570 580 590 600
[0132] SGEQKKAIVD LLFKTNRKVT VKQLKEDYFK KIECFDSVEI SGVEDRFNAS LGTYHDLLKI
[0133] 610 620 630 640 650 660
[0134] IKDKDFLDNE ENEDILEDIV LTLTLFEDRE MIEERLKTYA HLFDDKVMKQ LKRRRYTGWG
[0135] 670 680 690 700 710 720
[0136] RLSRKLINGI RDKQSGKTIL DFLKSDGFAN RNFMQLIHDD SLTFKEDIQK AQVSGQGDSL
[0137] 730 740 750 760 770 780
[0138] HEHIANLAGS PAIKKGILQT VKWDELVKV MGRHKPENIV IEMARENQTT QKGQKNSRER
[0139] 790 800 810 820 830 840
[0140] MKRIEEGIKE LGSQILKEHP VENTQLQNEK LYLYYLQNGR DMYVDQELDI NRLSDYDVDH
[0141] 850 860 870 880 890 900
[0142] IVPQSFLKDD SIDNKVLTRS DKNRGKSDNV PSEEWKKMK NYWRQLLNAK LITQRKFDNL
[0143] 910 920 930 940 950 960
[0144] TKAERGGLSE LDKAGFIKRQ LVETRQITKH VAQILDSRMN TKYDENDKLI REVKVITLKS
[0145] 970 980 990 1000 1010 1020
[0146] KLVSDFRKDF QFYKVREINN YHHAHDAYLN AWGTALIKK YPKLESEFVY GDYKVYDVRK
[0147] 1030 1040 1050 1060 1070 1080
[0148] MIAKSEQEIG KATAKYFFYS NIMNFFKTEI TLANGEIRKR PLIETNGETG EIVWDKGRDF
[0149] 1090 1100 1110 1120 1130 1140
[0150] ATVRKVLSMP QVNIVKKTEV QTGGFSKESI LPKRNSDKLI ARKKDWDPKK YGGFDSPTVA
[0151] 1150 1160 1170 1180 1190 1200
[0152] YSVLWAKVE KGKSKKLKSV KELLGITIME RSSFEKNPID FLEAKGYKEV KKDLI IKLPK
[0153] 1210 1220 1230 1240 1250 1260
[0154] YSLFELENGR KRMLASAGEL QKGNELALPS KYVNFLYLAS HYEKLKGSPE DNEQKQLFVE
[0155] 1270 1280 1290 1300 1310 1320
[0156] QHKHYLDEI I EQI SEFSKRV ILADANLDKV LSAYNKHRDK PIREQAENI I HLFTLTNLGA
[0157] 1330 1340 1350 1360
[0158] PAAFKYFDTT IDRKRYTSTK EVLDATLIHQ SITGLYETRI DLSQLGGD ( SEQ ID NO : 1 )
[0159] The SpCas9 variants described herein can include mutations at one, two, three, four, five, or all six of the following positions: DI 135, SI 136, G1218, E1219, R1335, and / or T1337, e.g., DI 135X / S1136X / G1218X / E1219X / R1335X / T1337X, where X is any amino acid (or at positions analogous thereto). In some embodiments, the SpCas9 variants are at least 80%, e.g., at least 85%, 90%, or 95% identical to the amino acid sequence of SEQ ID NO: 1, e.g., have differences at up to 5%, 10%, 15%, or 20% of the residues of SEQ ID NO: 1 replaced, e.g., with conservative mutations. Preferably the SpCas9 variants are at least 95% identical to the amino acid sequence of SEQ ID NO: 1. In preferred embodiments, the variant retains desired activity of the parent, e.g., the nuclease activity (except where the parent is a nickase or a dead Cas9), and / or the ability to interact with a guide RNA and target DNA).
[0160] To determine the percent identity of two nucleic acid sequences, the sequences are aligned for optimal comparison purposes (e.g., gaps can be introduced in one or both of a first and a second amino acid or nucleic acid sequence for optimal alignment and non-homologous sequences can be disregarded for comparison purposes). The length of a reference sequence aligned for comparison purposes is at least 80% of the length of the reference sequence, and in some embodiments is at least 90% or 100%. The nucleotides at corresponding amino acid positions or nucleotide positions are then compared. When a position in the first sequence is occupied by the same nucleotide as the corresponding position in the second sequence, then the molecules are identical at that position (as used herein nucleic acid “identity” is equivalent to nucleic acid “homology”). The percent identity between the two sequences is a function of the number of identical positions shared by the sequences, taking into account the number of gaps, and the length of each gap, which need to be introduced for optimal alignment of the two sequences. Percent identity between two polypeptides or nucleic acid sequences is determined in various ways that are within the skill in the art, for instance, using publicly available computer software such as Smith Waterman Alignment (Smith, T. F. and M. S. Waterman (1981) J Mol Biol 147: 195-7);
[0161] “BestFit” (Smith and Waterman, Advances in Applied Mathematics, 482-489 (1981)) as incorporated into GeneMatcher Plus™, Schwarz and Dayhof (1979) Atlas of Protein Sequence and Structure, Dayhof, M.O., Ed, pp 353-358; BLAST program (Basic Local Alignment Search Tool; (Altschul, S. F., W. Gish, et al. (1990) J Mol Biol 215: 403-10), BLAST-2, BLAST-P, BLAST-N, BLAST-X, WU-BL AST-2, ALIGN, ALIGN-2, CLUSTAL, or Megalign (DNASTAR) software. In addition, those skilled in the art can determine appropriate parameters for measuring alignment, including any algorithms needed to achieve maximal alignment over the length of the sequences being compared. In general, for proteins or nucleic acids, the length of comparison can be any length, up to and including full length (e.g., 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 100%). For purposes of the present compositions and methods, at least 80% of the full length of the sequence is aligned using the BLAST algorithm and the default parameters.
[0162] For purposes of the present invention, the comparison of sequences and determination of percent identity between two sequences can be accomplished using a Blossum 62 scoring matrix with a gap penalty of 12, a gap extend penalty of 4, and a frameshift gap penalty of 5.
[0163] Conservative substitutions typically include substitutions within the following groups: glycine, alanine; valine, isoleucine, leucine; aspartic acid, glutamic acid, asparagine, glutamine; serine, threonine; lysine, arginine; and phenylalanine, tyrosine.
[0164] In some embodiments, the SpCas9 variants include a set of mutations at DI 135, S1136, G1218, E1219, R1335, and T1337 shown in Table A. We utilized our PAMmla model to predict lists of SpCas9 variant enzymes that could exhibit maximal activity for certain PAMs (or combinations of PAMs), or that exhibited maximal selectivity for specific PAMs (minimizing activity against all other PAMs). Averaged PAMmla predicted rates from three separate training instances (on different randomized train test splits) were used to generate lists of top variants. PAMmla was used to generate predicted rate constants for each of the 64 possible three nucleotide PAM sequences for all possible amino acid combinations at positions at the six library positions. Predictions were then sorted by various criteria to maximize different enzyme properties. For enzyme lists described as “maximum targeting”, the average rate constant on the PAM of interest was used as the sorting criteria, and the 500 enzymes with the highest values are listed. For enzyme lists described as “sorted for selectivity”, the sorting criteria used was the rate constant on the PAM of interest divided by the sum of the rate constants on all 64 PAMs (this metric can be interpreted as the fraction of total activity that occurs on the PAM of interest). Again, the 500 enzymes with the highest values for this metric are listed. For enzymes from the list “maximizing NGTG and minimizing NGGG” list, enzymes were sorted to minimize rate constants on NGGG with the additional requirement that rate constants on NGTG be greater than 10'2. For enzyme lists maximizing targeting of two PAMs simultaneously, the specific sorting criteria are listed in the individual table headers. TABLE A - Variants with indicated PAM activity / specificity
[0165]
[0166]
[0167]
[0168]
[0169] In some embodiments, the SpCas9 variants do not comprise a set of mutations shown in WO 2016 / 141224 with wild type residues at SI 136 and E1219, e.g., do not include variants shown in Tables 7-9 of WO 2016 / 141224, e.g., do not include mutations at DI 135, G1218, R1335, and / or T1337, with wild type residues at SI 136 and E1219. In some embodiments, the SpCas9 variants do not comprise one of the following sets of mutations: DI 135V / R1335Q / T1337R (VQR variant);
[0170] DI 135V / G1218R / R1335Q / T1337R (VRQR variant); DI 135E / R1335Q / T1337R (EQR variant); or DI 135V / G1218R / R1335E / T1337R (VRER variant), with wild type residues at SI 136 and E1219. In some embodiments, the SpCas9 variants do not comprise one of the following sets of mutations VRQ, NRQ, YRQ, VRQL, VRQM, VRQR, VRQE, VRQQ, NRQL, NRQM, NRQR, NRQE, NRQQ, YRQL, YRQM, YRQR, YRQE, YRQQ, VRVQE, NRVQE, YRVQE, VVQE, NVQE, YVQE, VQR, VQR+L1 111H, VRQR+L1111H, VQR+N1317K, VRQR+N1317K, VQR+G1104K, VRQR+G1 104K, VQR+S1109T, or VRQR+S1109T, with wild type residues at SI 136 and E1219, or VQR+E1219K, VQR+E1219V, NQR+S1136N, NRQR+S1136N.
[0171] In some embodiments, the SpCas9 variants do not comprise a set of mutations shown in WO 2019 / 040650, e.g., shown in Tables 1, 2, or 3 of WO 2019 / 040650, e.g., do not include any of the following sets of mutations at
[0172] DI 135X / S1136X / G1218X / E1219X / R1335X / T1337X: SpCas9-LWKIQK, LWKIQK, IRAVQL, SWRVW, SWKVLK, TAHFKV, MSGVKC, LRSVRS, SKTLRP, MWVHLN, TWSMRG, KRRCKV, VRAVQL, VSSVRS, VRSVRS, SRMHCK, GRKIQK, GWKLLR, GWKOQK, VAKLLR, VAKIQK, VAKILR, GRKILR, VRKLLR, LRSVQL, IRAVQL, VRKIQK, VRMHCK, LRKIQK, LRSVQK, or VRKIQK variant (e.g, for NGTN PAMs); WMQAYG, MQKSER, LWRSEY, SQSWRS, LKAWRS, LWGWQH, MCSFER, LWMREQ, LWRVVA, HSSWVR, MWSEPT, GSNYQS, FMQWVN, YCSWVG, MCAWCG, LWLETR, FMQWVR, SSKWPA, LSRWQR, ICCCER, VRKSER, or ICKSER (e.g., for NGCN PAMs); or LRLSAR, AWTEVTR, KWMMCG, VRGAKE, MRARKE, AWNFQV, LWTTLN, SRMHCK, CWCQCV, AEEQQR, GWEKVR, NRAVNG, SRQMRG, LRSYLH, VRGNNR, VQDAQR, GWRQSK, AWLCLS, KWARVV, MWAARP, SRMHCK, VKMAKG, QRKTRE, LCRQQR, CWSHQR, SRTHTQ, LWEVIR, VSSVRS, VRSVRS, LRSVRS, IRAVRS, SRSVRS, LWKIQK, VRMHCK, or SRMHCK (e.g., for NGAN PAMs). In some embodiments, the spCas9 variants do not include DI 135L / S1136R / G1218S / E1219V / R1335X / T1337X, e.g., LRSVQL or LRSVRS. In some embodiments, the SpCas9 variants do not comprise VSREER or VSREQR. In some embodiments, the variants do not comprise only E1219V, R1335A, R1335A / L1111R / G1218R / A1322R / T1337R, R1335V / L1111R / D1135V / G1218R / E1219F / A1322R / T1337R .
[0173] In some embodiments, the variants comprise a consensus sequence of residues as shown in Fig. ID or Table B. Table B, Consensus sequences for top variants
[0174]
[0175]
[0176]
[0177]
[0178]
[0179]
[0180]
[0181]
[0182]
[0183]
[0184]
[0185]
[0186]
[0187]
[0188]
[0189]
[0190]
[0191] In some embodiments, the variants include one of the sets of mutations in Tables C-F.
[0192] Table C provides variants obtained from bacterial selections that have useful activity on one or more PAMs. HT-PAMDA was performed validating PAM preferences of all of these enzymes.
[0193] Table D provides consensus variants (cloned with most commonly observed amino acid substitutions from bacterial selections at each library position) which show useful activity on one or more PAMs. HT-PAMDA was performed validating PAM preferences of all of these enzymes. Table E provides variants obtained from sorting PAMmla predictions by various criteria (as in table A) that were subsequently validated by HT-PAMDA. Some of these variants may differ from what appears in table A since sorting was performed for the 3 training instances of PAMmla separately instead of taking average values.
[0194] HT-PAMDA was performed validating PAM preferences of all of these enzymes.
[0195] Table F provides variants that are a subset of all of the previous tables for which we also validated the editing efficiencies in human cells as nucleases and / or base editors. Human cell editing efficiencies of these enzymes appear in Fig. 4 and Figs. 18c, 25, 28, 29 and 34.
[0196] In some embodiments, the SpCas9 variants also include mutations at one of the following amino acid positions, which reduce or destroy the nuclease activity of the Cas9: DIO, E762, D839, H983, or D986 and H840 or N863, e.g., D10A / D10N and H840A / H840N / H840Y, to render the nuclease portion of the protein catalytically inactive; substitutions at these positions could be alanine (as they are in Nishimasu al., Cell 156, 935-949 (2014)), or other residues, e.g., glutamine, asparagine, tyrosine, serine, or aspartate, e.g., E762Q, H983N, H983Y, D986N, N863D, N863S, or N863H (see WO 2014 / 152432). In some embodiments, the variant includes mutations at D10A or H840A (which creates a single strand nickase), or mutations at D10A and H840A (which abrogates nuclease activity; this mutant is known as dead Cas9 or dCas9).
[0197] In some embodiments, the SpCas9 variants also include mutations at one or more amino acid positions that increase the specificity of the protein (i.e., reduce off-target effects). In some embodiments, the SpCas9 variants include one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or all thirteen mutations at the following residues: N497, K526, R661, R691, N692, M694, Q695, H698, K810, K848, Q926, K1003, and / or R0160. In some embodiments, the mutations are: N692A, Q695A, Q926A, H698A, N497A, K526A, R661A, R691A, M694A, K810A, K848A, K1003A, R0160A, Y450A / Q695A, L169A / Q695A, Q695A / Q926A, Q695A / D1135E, Q926A / D1135E, Y450A / D1135E, L169A / Y450A / Q695A, L169A / Q695A / Q926A, Y450A / Q695A / Q926A, R661A / Q695A / Q926A, N497A / Q695A / Q926A, Y450A / Q695A / D1135E, Y450A / Q926A / D1135E, Q695A / Q926A / D1135E, L169A / Y450A / Q695A / Q926A,
[0198] L 169A / R661 A / Q695 A / Q926A, Y450A / R661 A / Q695 A / Q926 A, N497A / Q695A / Q926A / D1135E, R661A / Q695A / Q926A / D1135E, and Y450A / Q695A / Q926A / D1135E; N692A / M694A / Q695A / H698A, N692A / M694A / Q695A / H698A / Q926A; N692A / M694A / Q695A / Q926A; N692A / M694A / H698A / Q926A; N692A / Q695A / H698A / Q926A;
[0199] M694A / Q695A / H698A / Q926A; N692A / Q695A / H698A; N692A / M694A / Q695A; N692A / H698A / Q926A; N692A / M694A / Q926A; N692A / M694A / H698A;
[0200] M694A / Q695A / H698A; M694A / Q695A / Q926A; Q695A / H698A / Q926A;
[0201] G582 A / V583 A / E584 A / D585 A / N588 A / Q926 A;
[0202] G582A / V583A / E584A / D585A / N588A; T657A / G658A / W659A / R661A / Q926A; T657A / G658A / W659A / R661A; F491A / M495A / T496A / N497A / Q926A;
[0203] F491A / M495A / T496A / N497A; K918A / V922A / R925A / Q926A; or 918A / V922A / R925A; K855A; K810A / K1003A / R1060A; or K848A / K1003A / R1060A. See, e.g., US9512446B1; Kleinstiver et al., Nature. 2016 Jan 28;529(7587):490-5; Slaymaker et al., Science. 2016 Jan l;351(6268):84-8; Chen et al., Nature. 2017 Oct 19;550(7676):407-410; Tsai and Joung, Nature Reviews Genetics 17:300-312 (2016); Vakulskas et al., Nature Medicine 24:1216-1224 (2018); Casini et al., Nat Biotechnol. 2018 Mar;36(3):265-271. In some embodiments, the variants do not include mutations at K526 or R691.
[0204] In some embodiments, the SpCas9 variants include mutations at one, two, three, four, five, six or all seven of the following positions: L169A, Y450, N497, R661, Q695, Q926, and / or DI 135E, e.g., in some embodiments, the variant SpCas9 proteins comprise mutations at one, two, three, or all four of the following: N497, R661, Q695, and Q926, e.g., one, two, three, or all four of the following mutations: N497A, R661A, Q695A, and Q926A. In some embodiments, the variant SpCas9 proteins comprise mutations at Q695 and / or Q926, and optionally one, two, three, four or all five of L169, Y450, N497, R661 and DI 135E, e.g., including but not limited to Y450A / Q695A, L169A / Q695A, Q695A / Q926A, Q695A / D1135E, Q926A / D1135E, Y450A / D1135E, L169A / Y450A / Q695A, L169A / Q695A / Q926A, Y450A / Q695A / Q926A, R661A / Q695A / Q926A, N497A / Q695A / Q926A, Y450A / Q695A / D1135E, Y450A / Q926A / D1135E, Q695A / Q926A / D1135E, L169A / Y450A / Q695A / Q926A, L169A / R661A / Q695A / Q926A,Y450A / R661A / Q695A / Q926A, N497A / Q695A / Q926A / D1135E, R661A / Q695A / Q926A / D1135E, and Y450A / Q695A / Q926A / D1135E. See, e.g., KI einstiver et al., Nature 529:490-495 (2016); WO 2017 / 040348; US 9,512,446).
[0205] In some embodiments, the SpCas9 variants also include mutations at one, two, three, four, five, six, seven, or more of the following positions: F491, M495, T496, N497, G582, V583, E584, D585, N588, T657, G658, W659, R661, N692, M694, Q695, H698, K918, V922, and / or R925, and optionally at Q926, preferably comprising a sequence that is at least 80% identical to the amino acid sequence of SEQ ID NO: 1 with mutations at one, two, three, four, five, six, seven, or more of the following positions: F491, M495, T496, N497, G582, V583, E584, D585, N588, T657, G658, W659, R661, N692, M694, Q695, H698, K918, V922, and / or R925, and optionally at Q926, and optionally one or more of a nuclear localization sequence, cell penetrating peptide sequence, and / or affinity tag.
[0206] In some embodiments, the proteins comprise mutations at one, two, three, or all four of the following: N692, M694, Q695, and H698; G582, V583, E584, D585, and N588; T657, G658, W659, and R661; F491, M495, T496, and N497; or K918, V922, R925, and Q926.
[0207] In some embodiments, the proteins comprise one, two, three, four, or all of the following mutations: N692A, M694A, Q695A, and H698A; G582A, V583A, E584A, D585A, and N588A; T657A, G658A, W659A, and R661A; F491A, M495A, T496A, and N497A; or K918A, V922A, R925A, and Q926A.
[0208] In some embodiments, the proteins comprise mutations: N692A / M694A / Q695A / H698A. In some embodiments, the proteins comprise mutations at F539, M763, and K890, e.g., F539S, M763I, K890N (Lee et al., Nature Communications volume 9, Article number: 3048 (2018)), and optionally at El 007, e.g., E1007L or E1007P (Kim et al., Nature Chemical Biology volume 19, pages972- 980 (2023)).
[0209] In some embodiments, the proteins comprise mutations: N692A / M694A / Q695A / H698A / Q926A; N692A / M694A / Q695A / Q926A; N692A / M694A / H698A / Q926A; N692A / Q695A / H698A / Q926A;
[0210] M694A / Q695A / H698A / Q926A; N692A / Q695A / H698A; N692A / M694A / Q695A; N692A / H698A / Q926A; N692A / M694A / Q926A; N692A / M694A / H698A; M694A / Q695A / H698A; M694A / Q695A / Q926A; Q695A / H698A / Q926A;
[0211] G582 A / V583 A / E584 A / D585 A / N588 A / Q926 A;
[0212] G582A / V583A / E584A / D585A / N588A; T657A / G658A / W659A / R661A / Q926A;
[0213] T657A / G658A / W659A / R661A; F491A / M495A / T496A / N497A / Q926A;
[0214] F491A / M495A / T496A / N497A; K918A / V922A / R925A / Q926A; or 918A / V922A / R925A. See, e.g., Chen et al., “Enhanced proofreading governs CRISPR-Cas9 targeting accuracy,” bioRxiv, doi.org / 10.1101 / 160036 (August 12, 2017) and Nature. 2017 Oct 19;550(7676):407-410; Nishimasu et al., Nature volume 550, pages407-410 (2017).
[0215] In some embodiments, the variant proteins include mutations at one or more of R780, K810, R832, K848, K855, K968, R976, H982, K1003, K1014, K1047, and / or R1060, e.g., R780A, K810A, R832A, K848A, K855A, K968A, R976A, H982A, K1003A, K1014A, K1047A, and / or R1060A, e.g., K855A; K810A / K1003A / R1060A; (also referred to as eSpCas9 1.0); or K848A / K1003A / R1060A (also referred to as eSpCas9 1.1) (see Slaymaker et al., Science. 2016 Jan 1;351(6268):84-8).
[0216] The variant proteins can also include one or more mutations that increase activity, reduce off-target effects, and / or alter protospacer adjacent motif (PAM) or target adjacent motif (TAM) specificity (Tables G and H).
[0217] Table G: List of Exemplary High Fidelity and / or PAM-relaxed RGN Orthologs
[0218] * predicted based on UniRule annotation on the UniProt database.
[0219] Table H. List of Exemplary SpCas9 Activity-Altering Mutations
[0220]
[0221] Also provided herein are isolated nucleic acids encoding the SpCas9 variants, vectors comprising the isolated nucleic acids, optionally operably linked to one or more regulatory domains for expressing the variant proteins, and host cells, e.g., mammalian host cells, comprising the nucleic acids, and optionally expressing the variant proteins.
[0222] The variants described herein can be used for altering the genome of a cell; the methods generally include expressing the variant proteins in the cells, along with a guide RNA having a region complementary to a selected portion of the genome of the cell. Methods for selectively altering the genome of a cell are known in the art, see, e.g., US8,697,359; US2010 / 0076057; US2011 / 0189776; US2011 / 0223638; US2013 / 0130248; WO / 2008 / 108989; WO / 2010 / 054108; WO / 2012 / 164565;
[0223] WO / 2013 / 098244; WO / 2013 / 176772; US20150050699; US20150045546; US20150031134; US20150024500; US20140377868; US20140357530; US20140349400; US20140335620; US20140335063; US20140315985; US20140310830; US20140310828; US20140309487; US20140304853; US20140298547; US20140295556; US20140294773; US20140287938; US20140273234; US20140273232; US20140273231; US20140273230; US20140271987; US20140256046; US20140248702; US20140242702; US20140242700; US20140242699; US20140242664; US20140234972; US20140227787; US20140212869; US20140201857; US20140199767; US20140189896; US20140186958; US20140186919; US20140186843; US20140179770; US20140179006; US20140170753; Makarova et al., "Evolution and classification of the CRISPR-Cas systems" 9(6) Nature Reviews Microbiology 467-477 (1-23) (Jun. 2011); Wiedenheft et al., "RNA-guided genetic silencing systems in bacteria and archaea" 482 Nature 331-338 (Feb. 16, 2012); Gasiunas et al., "Cas9-crRNA ribonucleoprotein complex mediates specific DNA cleavage for adaptive immunity in bacteria" 109(39) Proceedings of the National Academy of Sciences USA E2579-E2586 (Sep. 4, 2012); Jinek et al., "A Programmable Dual- RNA-Guided DNA Endonuclease in Adaptive Bacterial Immunity" 337 Science 816- 821 (Aug. 17, 2012); Carroll, "A CRISPR Approach to Gene Targeting" 20(9)
[0224] Molecular Therapy 1658-1660 (Sep. 2012); U.S. Appl. No. 61 / 652,086, filed May 25, 2012; Al- Attar et al., Clustered Regularly Interspaced Short Palindromic Repeats (CRISPRs): The Hallmark of an Ingenious Antiviral Defense Mechanism in Prokaryotes, Biol Chem. (2011) vol. 392, Issue 4, pp. 277-289; Hale et al., Essential Features and Rational Design of CRISPR RNAs That Function With the Cas RAMP Module Complex to Cleave RNAs, Molecular Cell, (2012) vol. 45, Issue 3, 292-302. The variant proteins described herein can be used in place of the SpCas9 proteins described in the foregoing references with guide RNAs that target sequences that have PAM sequences according to Tables 1, 2, or 3.
[0225] In addition, the variants described herein can be used in fusion proteins in place of the wild-type Cas9 or other Cas9 mutations (such as the dCas9 or Cas9 nickase described above) as known in the art, e.g., a fusion protein with a heterologous functional domains as described in WO 2014 / 124284. For example, the variants, preferably comprising one or more nuclease-reducing or killing mutation, can be fused on the N or C terminus of the Cas9 to a transcriptional activation domain or other heterologous functional domains (e.g., transcriptional repressors (e.g., KRAB, ERD, SID, and others, e.g., amino acids 473-530 of the ets2 repressor factor (ERF) repressor domain (ERD), amino acids 1-97 of the KRAB domain of K0X1, or amino acids 1-36 of the Mad mSIN3 interaction domain (SID); see Beerli et al., PNAS USA 95: 14628-14633 (1998)) or silencers such as Heterochromatin Protein 1 (HP1, also known as swi6), e.g., HPla or HPIP; proteins or peptides that could recruit long non-coding RNAs (IncRNAs) fused to a fixed RNA binding sequence such as those bound by the MS2 coat protein, endoribonuclease Csy4, or the lambda N protein; enzymes that modify the methylation state of DNA (e.g., DNA methyltransferase (DNMT) or TET proteins); enzymes that modify histone subunits (e.g., histone acetyltransferases (HAT), histone deacetylases (HD AC), histone methyltransferases (e.g., for methylation of lysine or arginine residues) or histone demethylases (e.g., for demethylation of lysine or arginine residues)).
[0226] In some embodiments, the heterologous functional domain is a base editor, e.g., a deaminase that modifies cytosine DNA bases, e.g., a cytidine deaminase from the apolipoprotein B mRNA-editing enzyme, catalytic polypeptide-like (APOBEC) family of deaminases, including APOBEC 1, APOBEC2, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D / E, APOBEC3F, APOBEC3G, APOBEC3H, and APOBEC4 (see, e.g., Yang et al., J Genet Genomics. 2017 Sep 20;44(9):423-437); activation-induced cytidine deaminase (AID), e.g., activation induced cytidine deaminase (AICDA); cytosine deaminase 1 (CDA1) and CDA2; cytosine deaminase acting on tRNA (CD AT); and DddA-like cytidine deaminases (Huang et al., Cell.
[0227] 2023 Jul 20;186(15):3182-3195.el4. The following table provides exemplary sequences; other sequences can also be used, including AID / CDA1, the cytosine deaminase domains from engineered CBEs (e.g. BE3, BE4max, evoAPOBECl- BE4max, AncBE4max, FERNY-BE4max, evoFERNY-BE4max, CDAl-BE4max, evoCDAl-BE4max), the cytosine deaminase domains from engineered TadA-based CBEs (e.g., TadCBEs or TadDEs (e.g, TadCBEd), CBE-Ts or CABE-Ts, Td-CBEs or Td-CGBEs, etc.), Sdd CBEs, or CBE6 enzymes, for example.
[0228] * from Saccharomyces cerevisiae S288C In some embodiments, the base editor is an adenine deaminase that modifies adenosine DNA bases, e.g., the deaminase is an adenosine deaminase 1 (ADA1), ADA2; adenosine deaminase acting on RNA 1 (AD ARI), ADAR2, ADAR3 (see, e.g., Savva et al., Genome Biol. 2012 Dec 28;13(12):252); adenosine deaminase acting on tRNA 1 (ADAT1), ADAT2, ADAT3 (see Keegan et al., RNA. 2017 Sep;23(9): 1317-1328 and Schaub and Keller, Biochimie. 2002 Aug;84(8):791-803); and naturally occurring or engineered tRNA-specific adenosine deaminase (TadA) (see, e.g., Gaudelli et al., Nature. 2017 Nov 23;551(7681):464-471) (NP_417054.2 (Escherichia coli str. K-12 substr. MG1655); See, e.g., Wolf et al., EMBO J. 2002 Jul 15;21(14):3841-51). The following table provides exemplary sequences; other sequences can also be used. For example, The TadA domain has also been engineered to purposefully generate C-to-T edits in addition to, or instead of, the conventional A- to-G edits observed with TadA domains; these versions can also be used; see, e.g., Chen et al., Nature Biotechnology volume 41, pages663-672 (2023); Lam et al., Nature Biotechnology volume 41, pages686-697 (2023); Neugebauer et al., Nature Biotechnology volume 41, pages673-685 (2023). In some embodiments, the deaminase comprises an engineered adenosine deaminase TadA monomer or dimer comprises a homodimeric or heterodimeric TadA domain from ABEmax, ABE7.10, or ABE8e; monomer or dimer TadA from ABE 0.1, 0.2, 1.1, 1.2, 2.1, 2.2, 2.3, 2.4,
[0229] 2.5, 2.6, 2.7, 2.8, 2.9, 2.10, 2.11, 2.12, 3.1, 3.2, 3.3, 3.4, 3.5, 3.6, 3.7, 3.8, 4.1, 4.2, 4.3, 5.1, 5.2, 5.3, 5.4, 5.5, 5.6, 5.7, 5.8, 5.9, 5.10, 5.11, 5.12, 5.13, 5.14, 6.1, 6.2, 6.3, 6.4,
[0230] 6.5, 6.6, 7.1, 7.2, 7.3, 7.4, 7.5, 7.6, 7.7, 7.8, 7.9, 7.10, ABEmax, ABE8.8, ABE8.13, ABE8.17, ABE8.20, ABE8e, ABE9, ABE9e, or K20A / R21A, V82G, or V106W variants thereof; E.coli TadA monomer, or homo- or heterodimers thereof fused to the N or C terminus, optionally comprising one or more mutations in either or both monomers, optionally TadA from mini AB Emax -V82G, miniABEmax-K20A / R21A, miniABEmax-V106W, or another variant.
[0231] In some embodiments, the heterologous functional domain is an enzyme, domain, or peptide that inhibits or enhances endogenous DNA repair or base excision repair (BER) pathways, e.g., thymine DNA glycosylase (TDG; GenBank Acc Nos.
[0232] NM_003211.4 (nucleic acid) and NP_003202.3 (protein)) or uracil DNA glycosylase (UDG, also known as uracil N-glycosylase, or UNG; GenBank Acc Nos.
[0233] NM_003362.3 (nucleic acid) and NP_003353.1 (protein)) or uracil DNA glycosylase inhibitor (UGI) that inhibits UNG mediated excision of uracil to initiate BER (see, e.g., Mol et al., Cell 82, 701-708 (1995); Komor et al., Nature. 2016 May 19;533(7603)); or DNA end-binding proteins such as Gam, which is a protein from the bacteriophage Mu that binds free DNA ends, inhibiting DNA repair enzymes and leading to more precise editing (less unintended base edits). See, e.g., Komor et al., Sci Adv. 2017 Aug 30;3(8):eaao4774.
[0234] See, e.g., Komor et al., Nature. 2016 May 19;533(7603):420-4; Nishida et al., Science. 2016 Sep 16;353(6305). pii: aaf8729; Rees et al., Nat Commun. 2017 Jun 6;8: 15790; or Kim et al., Nat Biotechnol. 2017 Apr;35(4):371-376) as are known in the art can also be used.
[0235] A number of sequences for domains that catalyze hydroxylation of methylated cytosines in DNA. Exemplary proteins include the Ten-Eleven-Translocation (TET)l-3 family, enzymes that converts 5 -methylcytosine (5-mC) to 5- hydroxymethylcytosine (5-hmC) in DNA.
[0236] Sequences for human TET1-3 are known in the art and are shown in the following table:
[0237] * Variant (1) represents the longer transcript and encodes the longer isoform (a). Variant (2) differs in the 5' UTR and in the 3' UTR and coding sequence compared to variant 1. The resulting isoform (b) is shorter and has a distinct C-terminus compared to isoform a. In some embodiments, all or part of the full-length sequence of the catalytic domain can be included, e.g., a catalytic module comprising the cysteine-rich extension and the 2OGFeDO domain encoded by 7 highly conserved exons, e.g., the Tetl catalytic domain comprising amino acids 1580-2052, Tet2 comprising amino acids 1290-1905 and Tet3 comprising amino acids 966-1678. See, e.g., Fig. 1 of Iyer et al., Cell Cycle. 2009 Jun 1 ;8(11): 1698-710. Epub 2009 Jun 27, for an alignment illustrating the key catalytic residues in all three Tet proteins, and the supplementary materials thereof for full length sequences (see, e.g., seq 2c); in some embodiments, the sequence includes amino acids 1418-2136 of Tetl or the corresponding region in Tet2 / 3.
[0238] Other catalytic modules can be from the proteins identified in Iyer et al., 2009.
[0239] In some embodiments, the heterologous functional domain is a biological tether, and comprises all or part of (e.g., DNA binding domain from) the MS2 coat protein, endoribonuclease Csy4, or the lambda N protein. These proteins can be used to recruit RNA molecules containing a specific stem-loop structure to a locale specified by the dCas9 gRNA targeting sequences. For example, a dCas9 variant fused to MS2 coat protein, endoribonuclease Csy4, or lambda N can be used to recruit a long noncoding RNA (IncRNA) such as XIST or HOTAIR; see, e.g., Keryer-Bibens et al., Biol. Cell 100: 125-138 (2008), that is linked to the Csy4, MS2 or lambda N binding sequence. Alternatively, the Csy4, MS2 or lambda N protein binding sequence can be linked to another protein, e.g., as described in Keryer-Bibens et al., supra, and the protein can be targeted to the dCas9 variant binding site using the methods and compositions described herein. In some embodiments, the Csy4 is catalytically inactive. In some embodiments, the Cas9 variant, preferably a dCas9 variant, is fused to FokI as described in WO 2014 / 204578.
[0240] In some embodiments, the fusion proteins include a linker between the dCas9 variant and the heterologous functional domains. Linkers that can be used in these fusion proteins (or between fusion proteins in a concatenated structure) can include any sequence that does not interfere with the function of the fusion proteins. In preferred embodiments, the linkers are short, e.g., 2-20 amino acids, and are typically flexible (i.e., comprising amino acids with a high degree of freedom such as glycine, alanine, and serine). In some embodiments, the linker comprises one or more units consisting of GGGS (SEQ ID NO:2) or GGGGS (SEQ ID NO:3), e.g., two, three, four, or more repeats of the GGGS (SEQ ID NO:2) or GGGGS (SEQ ID NO:3) unit. Other linker sequences can also be used. Delivery and Expression Systems
[0241] To use the Cas9 variants described herein, it may be desirable to express them from a nucleic acid that encodes them. This can be performed in a variety of ways. For example, the nucleic acid encoding the Cas9 variant can be cloned into an intermediate vector for transformation into prokaryotic or eukaryotic cells for replication and / or expression. Intermediate vectors are typically prokaryote vectors, e.g., plasmids, or shuttle vectors, or insect vectors, for storage or manipulation of the nucleic acid encoding the Cas9 variant for production of the Cas9 variant. The nucleic acid encoding the Cas9 variant can also be cloned into an expression vector, for administration to a plant cell, animal cell, preferably a mammalian cell or a human cell, fungal cell, bacterial cell, or protozoan cell.
[0242] To obtain expression, a sequence encoding a Cas9 variant is typically subcloned into an expression vector that contains a promoter to direct transcription. Suitable bacterial and eukaryotic promoters are well known in the art and described, e.g., in Sambrook et al., Molecular Cloning, A Laboratory Manual (3d ed. 2001); Kriegler, Gene Transfer and Expression: A Laboratory Manual (1990); and Current Protocols in Molecular Biology (Ausubel et al., eds., 2010). Bacterial expression systems for expressing the engineered protein are available in, e.g., E. coh. Bacillus sp., and Salmonella (Palva et al., 1983, Gene 22:229-235). Kits for such expression systems are commercially available. Eukaryotic expression systems for mammalian cells, yeast, and insect cells are well known in the art and are also commercially available. The promoter used to direct expression of a nucleic acid depends on the particular application. For example, a strong constitutive promoter is typically used for expression and purification of fusion proteins. In contrast, when the Cas9 variant is to be administered in vivo for gene regulation, either a constitutive or an inducible promoter can be used, depending on the particular use of the Cas9 variant. In addition, a preferred promoter for administration of the Cas9 variant can be a weak promoter, such as HSV TK or a promoter having similar activity. The promoter can also include elements that are responsive to transactivation, e.g., hypoxia response elements, Gal4 response elements, lac repressor response element, and small molecule control systems such as tetracycline-regulated systems and the RU-486 system (see, e.g., Gossen & Bujard, 1992, Proc. Natl. Acad. Sci. USA, 89:5547; Oligino et al., 1998, Gene Ther., 5:491-496; Wang et al., 1997, Gene Then, 4:432-441; Neering et al., 1996, Blood, 88: 1147-55; and Rendahl et al., 1998, Nat. Biotechnol., 16:757-761). In addition to the promoter, the expression vector typically contains a transcription unit or expression cassette that contains all the additional elements required for the expression of the nucleic acid in host cells, either prokaryotic or eukaryotic. A typical expression cassette thus contains a promoter operably linked, e.g., to the nucleic acid sequence encoding the Cas9 variant, and any signals required, e.g., for efficient polyadenylation of the transcript, transcriptional termination, ribosome binding sites, or translation termination. Additional elements of the cassette may include, e.g., enhancers, and heterologous spliced intronic signals.
[0243] The particular expression vector used to transport the genetic information into the cell can be selected with regard to the intended use of the Cas9 variant, e.g., expression in plants, animals, bacteria, fungus, protozoa, etc. Standard bacterial expression vectors include plasmids such as pBR322 based plasmids, pSKF, pET23D, and commercially available tag-fusion expression systems such as GST and LacZ.
[0244] Expression vectors containing regulatory elements from eukaryotic viruses are often used in eukaryotic expression vectors, e.g., SV40 vectors, papilloma virus vectors, and vectors derived from Epstein-Barr virus. Other exemplary eukaryotic vectors include pMSG, pAV009 / A+, pMTO10 / A+, pMAMneo-5, baculovirus pDSVE, and any other vector allowing expression of proteins, e.g., under the direction of the SV40 early promoter, SV40 late promoter, metallothionein promoter, murine mammary tumor virus promoter, Rous sarcoma virus promoter, polyhedrin promoter, or other promoters shown effective for expression in eukaryotic cells.
[0245] A preferred approach for in vivo introduction of nucleic acid into a cell is by use of a viral vector containing nucleic acid, e.g., a cDNA. Infection of cells with a viral vector has the advantage that a large proportion of the targeted cells can receive the nucleic acid. Additionally, molecules encoded within the viral vector, e.g., by a cDNA contained in the viral vector, are expressed efficiently in cells that have taken up viral vector nucleic acid. Viral vectors for use in the present methods and compositions include recombinant retroviruses, adenovirus, adeno-associated virus, alphavirus, and lentivirus, comprising the targeting peptides described herein and optionally a transgene for expression in a target tissue.
[0246] A preferred viral vector system useful for delivery of nucleic acids in the present methods is the adeno-associated virus (AAV). AAV is a tiny non-enveloped virus having a 25 nm capsid. No disease is known or has been shown to be associated with the wild-type virus. AAV has a single-stranded DNA (ssDNA) genome. AAV has been shown to exhibit long-term episomal transgene expression. Space for exogenous DNA in AAV is generally limited to an amount of nucleic acid that can physically fit inside the particle. For example, AAV types 1-5 can package up to 6 kb DNA, and in some reports AAV5 has been shown to package up to 8.9 kb DNA. An AAV vector such as that described in Tratschin et al., Mol. Cell. Biol. 5:3251-3260 (1985) can be used to introduce DNA into cells. A variety of nucleic acids have been introduced into different cell types using AAV vectors (see for example Hermonat et al., Proc. Natl. Acad. Sci. USA 81 :6466-6470 (1984); Tratschin et al., Mol. Cell. Biol. 4:2072- 2081 (1985); Wondisford et al., Mol. Endocrinol. 2:32-39 (1988); Tratschin et al., J. Virol. 51 :611-619 (1984); and Flotte et al., J. Biol. Chem. 268:3781-3790 (1993). There are numerous alternative AAV variants (over 100 have been cloned), and AAV variants have been identified based on desirable characteristics. The present disclosure contemplates uses of peptides that can be incorporated into an AAV capsid — thus providing capsid modified AAVs, e.g., AAVPR — for selectively transfecting endothelium, pericytes and SMC after delivery to a subject. Such AAV’s can also be used for delivery of a nucleic acid comprising an HDAC9-derived promoter as described herein. In some embodiments, an AAV suitable for use with a nucleic acid or a targeting peptide of the disclosure is AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AV6.2, AAV7, AAV8, rh.8, AAV9, rh.10, rh.39, rh.43 or CSp3; for CNS use, in some embodiments the AAV is AAV1, AAV2, AAV4, AAV5, AAV6, AAV8, or AAV9.
[0247] In cases where an AAV vector is used, the gene therapy construct can also include components such as inverted terminal repeats (ITRs) and Rep. cl.
[0248] The vectors for expressing the Cas9 variants can include RNA Pol III promoters to drive expression of the guide RNAs, e.g., the Hl, U6 or 7SK promoters. These human promoters allow for expression of Cas9 variants in mammalian cells following plasmid transfection.
[0249] Some expression systems have markers for selection of stably transfected cell lines such as thymidine kinase, hygromycin B phosphotransferase, and dihydrofolate reductase. High yield expression systems are also suitable, such as using a baculovirus vector in insect cells, with the gRNA encoding sequence under the direction of the polyhedrin promoter or other strong baculovirus promoters.
[0250] The elements that are typically included in expression vectors also include a replicon that functions in E. coh. a gene encoding antibiotic resistance to permit selection of bacteria that harbor recombinant plasmids, and unique restriction sites in nonessential regions of the plasmid to allow insertion of recombinant sequences.
[0251] Standard transfection methods are used to produce bacterial, mammalian, yeast or insect cell lines that express large quantities of protein, which are then purified using standard techniques (see, e.g., Colley et al., 1989, J. Biol. Chem., 264: 17619-22; Guide to Protein Purification, in Methods in Enzymology, vol. 182 (Deutscher, ed., 1990)). Transformation of eukaryotic and prokaryotic cells are performed according to standard techniques (see, e.g., Morrison, 1977, J. Bacteriol. 132:349-351; Clark- Curtiss & Curtiss, Methods in Enzymology 101 :347-362 (Wu et al., eds, 1983). Any of the known procedures for introducing foreign nucleotide sequences into host cells may be used. These include the use of calcium phosphate transfection, polybrene, protoplast fusion, electroporation, nucleofection, liposomes, microinjection, naked DNA, plasmid vectors, viral vectors, both episomal and integrative, and any of the other well-known methods for introducing cloned genomic DNA, cDNA, synthetic DNA or other foreign genetic material into a host cell (see, e.g., Sambrook et al., supra). It is only necessary that the particular genetic engineering procedure used be capable of successfully introducing at least one gene into the host cell capable of expressing the Cas9 variant.
[0252] Alternatively, the methods can include delivering the Cas9 variant protein and guide RNA together, e.g., as a complex. For example, the Cas9 variant and gRNA can be can be overexpressed in a host cell and purified, then complexed with the guide RNA (e.g., in a test tube) to form a ribonucleoprotein (RNP), and delivered to cells. In some embodiments, the variant Cas9 can be expressed in and purified from bacteria through the use of bacterial Cas9 expression plasmids. For example, His-tagged variant Cas9 proteins can be expressed in bacterial cells and then purified using nickel affinity chromatography. The use of RNPs circumvents the necessity of delivering plasmid DNAs encoding the nuclease or the guide, or encoding the nuclease as an mRNA. RNP delivery may also improve specificity, presumably because the half-life of the RNP is shorter and there’s no persistent expression of the nuclease and guide (as you’d get from a plasmid). The RNPs can be delivered to the cells in vivo or in vitro, e.g., using lipid-mediated transfection or electroporation. See, e.g., Liang et al. "Rapid and highly efficient mammalian cell engineering via Cas9 protein transfection." Journal of biotechnology 208 (2015): 44-53; Zuris, John A., et al. "Cationic lipid-mediated delivery of proteins enables efficient protein-based genome editing in vitro and in vivo." Nature biotechnology 33.1 (2015): 73-80; Kim et al. "Highly efficient RNA-guided genome editing in human cells via delivery of purified Cas9 ribonucleoproteins." Genome research 24.6 (2014): 1012-1019.
[0253] The present invention includes the vectors and cells comprising the vectors.
[0254] Editing disease-relevant mutations
[0255] Also provided herein are methods and compositions for editing disease relevant mutations in living cells, e.g., in cells in vitro or in vivo, using a variant described herein (e.g., as part of a base editor) and an appropriate gRNA. Specific example include methods of editing Apolipoprotein E, e.g., to provide reduced risk of developing Alzheimer’s disease (to simultaneously introduce the two C to T point mutations to revert the APOE-sA (Argl30, Argl76) Alzheimer’s risk allele to the APOE-&2 (Cysl30, Cysl76) Alzheimer’s protective allele); editing the therapeutically relevant HBB E7V mutation causing sickle cell disease; allele-specific deletion of the RHO P23H allele; and editing the CYBB T362I mutation in X-linked Chronic Granulomatous Disease (X-CGD). Enzymes described herein that can be used for these specific applications include the following:
[0256] APOE: cytosine base editors (e.g., BE4max, evoFERNY, and TadCBEd) comprising MRKCRS, MRKQKC, MRKSKC, MR. KM RS, MRKCRN, MRKCRQ, QRKCKK, MRKCRK, MRACRQ, MRKMRK, or MRKMRQ variant;
[0257] HBB: adenine base editor (e.g., ABE8e) comprising LWQYQH, MWKYQS, LWKYQS, or MWKYQA variant;
[0258] RHO: SpCas9 nuclease comprising MRRWMR, KRHWMR, MRRMYR, KRRYQR, MRAFMR, KRAYQR, KRRWQR, MRHWQR, IRSMQR, KRAWQR, LRKMYR, ARGIMR, MKRCMV, MWNVML, or MKKCMN variant;
[0259] CYBB: adenine base editor (e.g., ABE8e) comprising KWRQLC variant Exemplary guide RNA spacer sequences, which can preferably be used with the enzymes listed above, are provided in Table 1. For example, for editing HBB, gRNA HBB-E7V-A7 and HBB-E7V-A9 can be used. For editing ApoE, the APOE- R112 / 130 gRNA (use with NGTA enzymes) and APOE-R158 / 176 gRNA (use with NGGC enzymes) can be used. For these gRNAs, it would be preferable to use enzymes that have been prioritized to target both NGTA and NGGC with low or no activity on NGGG PAMs. For editing rho, the RHO-P23H gRNA can be used. For editing CYBB, the CYBB-T362I gRNA can be used. Machine-Learning Methods for Predicting PAM Variants
[0260] FIG. 7 is a block diagram for an example method for generating enzymes of desired functionality in accordance with implementations of the present disclosure. In some instances, the generation of enzymes can be performed as described in relation to Fig. 1c above.
[0261] In some instances, at 705, active SpCas9 enzymes are identified through structure- and-function-informed saturation mutagenesis and subsequent bacterial selection. At 710, the active SpCas9 enzymes are profiled according to all possible protospacer- adjacent motif (PAM) requirements to amino acid sequences using a high-throughput PAM-determination assay (HT-PAMDA) to generate a first training data set for training a PAM machine-learning (ML) model to relate PAM requirements to amino acid sequence. In some instances, including only active variants as part of the training data may be associated in imbalance in the data and thus, a second data set can be defined to learn combinations of amino acids that lead to both active and / or inactive enzymes. For example, the bacterial selection the 16 different NGNN PAMs can provide only enzymes that survived the bacterial selections and as such, the first training data set can be biased towards NGGN and NGTN enzymes, and variants preferring a G at the 4thposition (as described in relation to Fig. If). At 715, a second training data set comprising random non-selected enzymes from a SpCas9(6AA) library. At 720, the PAM ML model is trained based on the first and the second training data sets to relate enzyme function to amino acid sequences. In some instances, labels can be assigned to each training example, based on its most active PAM and / or randomly assigned to a selected training example to balance across PAM classes. At 725, the PAM ML model is used to perform screening of amino acid combinations based on an amino acid sequence as input to obtain a PAM variant enzyme with predicted respective PAM requirements.
[0262] In some instances, the PAM ML model can be trained according to any appropriate learning techniques that use the training set of data and can output related PAM properties to a given enzyme based on learned relationships between PAM requirements to amino acid sequences from the training data. The training of the model can be executed in a computer-controlled environment and can be performed in iterations over various subsets of the training data. In some instances, one or more computing devices can be used for processing the training data and learning the relationship between PAM properties and enzyme variants. . In some cases, the trained model an be validated based on a test data set, for example, a portion of the training data can be used for the validation as the test data set. Computing devices can be run implemented logic to trigger evaluation of the training examples and provide a trained model that meets a defined accuracy criterion, for example, a threshold level of likelihood of accurate prediction during the validation stage. In some instances, the generated training data including the first and the second training data sets can be stored at a data store and provided for accessing, for example, based on user queries. In some instances, generated enzymes based on the use of the PAM ML model can be stored in the data store and also made available for accessing. In some instances, the database can provide an interface that can receive queries related to enzymes or desired PAM preferences of an enzyme, and can provide results based on invoking data from the data store. In some cases, for a requested PAM preference, more than one enzymes may be provided. In some cases, if an enzyme is provided with a request, PAM preferences of the enzyme can be output and optionally, other enzymes having similar PAM preferences can be output. FIG. 8 is a block diagram for an example method for computationally generating SpCas9 variant enzymes with customizable properties in accordance with implementations of the present disclosure. In some instances, the generation of enzymes can be performed as described in relation to Fig. 6a above.
[0263] At 805, a starting sequence is computationally mutated to generate a library of sequences having random amino acid substitutions with a defined hamming distance from the starting sequence. At 810, a protospacer-adjacent motif (PAM) machinelearning (ML) model trained to relate PAM requirements to amino acid sequence is used to generate PAM predictions for each member of the library. In some instances, the PAM ML model can be substantially similar to the trained PAM ML model at 720 of FIG. 7. At 815, a score, using a fitness function (e.g., received through user input), is determined for each sequence in the library according to properties defined for evolution. At 820, a set of sequences from the library are determined, where the sequences in the set has a respective score matching a threshold score criterion for selection as a subsequent starting sequence for a subsequent evaluation iteration. At 825, evaluation comprising performing the computational mutation and the PAM predictions generation is iteratively initiated to be performed until the determined scores of the sequences using the customizable fitness function satisfy a fitness criterion. FIG. 9 is a block diagram for automatic characterization of generated SpCas9 variant enzymes in accordance with implementations of the present disclosure. In some instances the characterization of the enzymes can be used to generate training data for a PAM ML
[0264] At 905, SpCas9 enzymes are obtained from a data store comprising SpCas9(6AA) enzymes. For example, the data store can include enzymes generated from the methods as described in relation to FIG. 7 and / or 8.
[0265] At 910, the SpCas9 enzymes are cloned into a mammalian expression plasmid. At 915, the cloned SpCas9 enzymes are sequenced by a whole-ORF sequencing. At 920, the sequenced SpCas9 enzymes are subjected to a high-throughput PAM- determination assay (HT-PAMDA) for comprehensive PAM characterization of the sequenced SpCas9 enzymes.
[0266] Referring now to FIG. 10, a schematic diagram of an example computing system 1000 is provided. The system 1000 can be used for the operations described in association with the implementations described herein. For example, the system 1000 may be included in any or all of the server components discussed herein. The system 1000 includes a processor 1010, a memory 1020, a storage device 1030, and an input / output device 1040. The components 1010, 1020, 1030, 1040 are interconnected using a system bus 1050. The processor 1010 is capable of processing instructions for execution within the system 1000. In some implementations, the processor 1010 is a single-threaded processor. In some implementations, the processor 1010 is a multi -threaded processor. The processor 1010 is capable of processing instructions stored in the memory 1020 or on the storage device 1030 to display graphical information for a user interface on the input / output device 1040. The memory 1020 stores information within the system 1000. In some implementations, the memory 1020 is a computer-readable medium. In some implementations, the memory 1020 is a volatile memory unit. In some implementations, the memory 1020 is a non-volatile memory unit. The storage device 1030 is capable of providing mass storage for the system 1000. In some implementations, the storage device 1030 is a computer-readable medium. In some implementations, the storage device 1030 may be a floppy disk device, a hard disk device, an optical disk device, or a tape device. The input / output device 1040 provides input / output operations for the system 1000. In some implementations, the input / output device 1040 includes a keyboard and / or pointing device. In some implementations, the input / output device 1040 includes a display unit for displaying graphical user interfaces.
[0267] The features described can be implemented in digital electronic circuitry, or in computer hardware, firmware, software, or in combinations of them. The apparatus can be implemented in a computer program product tangibly embodied in an information carrier (e.g., in a machine-readable storage device, for execution by a programmable processor), and method steps can be performed by a programmable processor executing a program of instructions to perform functions of the described implementations by operating on input data and generating output. The described features can be implemented advantageously in one or more computer programs that are executable on a programmable system including at least one programmable processor coupled to receive data and instructions from, and to transmit data and instructions to, a data storage system, at least one input device, and at least one output device. A computer program is a set of instructions that can be used, directly or indirectly, in a computer to perform a certain activity or bring about a certain result. A computer program can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0268] Suitable processors for the execution of a program of instructions include, by way of example, both general and special purpose microprocessors, and the sole processor or one of multiple processors of any kind of computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. Elements of a computer can include a processor for executing instructions and one or more memories for storing instructions and data. Generally, a computer can also include, or be operatively coupled to communicate with, one or more mass storage devices for storing data files; such devices include magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and optical disks. Storage devices suitable for tangibly embodying computer program instructions and data include all forms of non-volatile memory, including by way of example semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks and removable disks; magnetooptical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, ASICs (application-specific integrated circuits).
[0269] To provide for interaction with a user, the features can be implemented on a computer having a display device such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor for displaying information to the user and a keyboard and a pointing device such as a mouse or a trackball by which the user can provide input to the computer.
[0270] The features can be implemented in a computer system that includes a back-end component, such as a data server, or that includes a middleware component, such as an application server or an Internet server, or that includes a front-end component, such as a client computer having a graphical user interface or an Internet browser, or any combination of them. The components of the system can be connected by any form or medium of digital data communication such as a communication network. Examples of communication networks include, for example, a LAN, a WAN, and the computers and networks forming the Internet.
[0271] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a network, such as the described one. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0272] In addition, the logic flows depicted in the figures do not require the particular order shown, or sequential order, to achieve desirable results. In addition, other steps may be provided, or steps may be eliminated, from the described flows, and other components may be added to, or removed from, the described systems. Accordingly, other implementations are within the scope of the following claims.
[0273] A number of implementations of the present disclosure have been described. Nevertheless, it will be understood that various modifications may be made without departing from the spirit and scope of the present disclosure. Accordingly, other implementations are within the scope of the following claims.
[0274] EXAMPLES
[0275] The invention is further described in the following examples, which do not limit the scope of the invention described in the claims. Methods
[0276] The following materials and methods were used in the Examples below.
[0277] Plasmids, oligonucleotides strains, and cloning
[0278] Target site sequences for gRNAs and epegRNAs are provided in Table 1.
[0279] A saturation mutagenesis plasmid library was generated by randomizing six amino acid positions in the SpCas9 coding sequence (DI 135, SI 136, G1218, E1219, R1335, and T1337). A parental SpCas9 bacterial expression plasmid pACYC-T7-SpCas9-T7- EGFPgRNAl (BPK848; Addgene plasmid ID 181745) was used to generate the library by cloning type IIS restriction enzyme cassettes into the sequence near amino acid positions DI 135 / S1136 (BspMI enzyme cassette), G1218 / E1219 (SapI enzyme cassette), and R1335 / T1337 (Bsal enzyme cassette) to create the library entry plasmid pACYC-T7-SpCas9(BspMESapI / BsaI_cassettes)-T7-EGFPgRNAl (BPK1807). The entry plasmid BPK1807 was then subjected to sequential cloning steps (restriction digests followed by ligations with library oligos encoding NNS codons (where ‘N’ is any nucleotide and ‘S’ is C or G) that were pre-annealed with adapter oligos to reform restriction site overhangs), creating the final saturation mutagenesis library plasmid pACYC-T7-SpCas9(6AA_NNS)-T7-EGFPgRNAl (MMW94; NNS at codons D1135, SI 136, G1218, E1219, R1335, and T1337).
[0280] Target plasmids for bacterial -based positive selection assays containing an arabinose- inducible ccdB toxin gene were generated by cloning duplexed oligonucleotides into Xbal and Sphl-digested pl l-lacY-wtxl (Addgene ID 69056) as previously described10. The derivative ccdB-expressing plasmids contain an EGFP-derived protospacer sequence (GGGCACGGGCAGCTTGCCGG, SEQ ID NO:4) adjacent to each of the 16 possible NGNN PAMs varying at positions 3 and 4.
[0281] Bacterial strains for positive selection assays were generated by separately transforming chemically competent BW25141(XDE3)70. E. coli with each of the 16 pl 1-lacY-wtxl plasmid derivatives harboring each of the NGNN PAMs. To make electrocomp etent cells harboring each toxic plasmid, single colonies from each transformation were grown overnight in 5 mL LB supplemented with 100 pg / mL carbenicillin and 10 mM dextrose. Overnight cultures were diluted in 500 mL LB + carbenicillin + dextrose and grown to an OD600 of 0.5-0.8. After chilling on ice for 30 minutes, cultures were pelleted at 4 °C at 4000g for 15 minutes and resuspended in 500 mL ice cold FEO. Pelleting and resuspension were repeated three more times using 250 mL ice cold H2O, 10 mL ice cold 10% glycerol, and 1 mL 10% glycerol. Aliquots were frozen in liquid nitrogen and stored at -80 °C.
[0282] SpCas9 gRNA expression plasmids for human cell experiments were cloned by digesting pUC19-U6-BsmBI_cassette-SpCas9gRNA (BPK1520; Addgene ID 65777)10with BsmBI at 55 °C overnight and performing ligations with annealed oligos encoding the gRNA spacer (Table 1).
[0283] SpCas9 nuclease expression plasmids for human cell experiments were generated by digesting pCMV-T7-SpCas9-P2A-EGFP (RTW3027; Addgene plasmid ID 139987)11with Pmll and Xhol, and inserting the modified SpCas9 PAM interacting domain via four PCR products with mutations contained in the primer overlaps followed by isothermal assembly71.
[0284] A-to-G base editor (ABE) expression plasmids utilizing the ABE8e architecture54were generated by digesting pCMV-T7-ABE8e-nSpCas9-P2A-EGFP (KAC978; Addgene plasmid ID 185910)72with Xcml and Pmll, and inserting the modified SpCas9 PAM interacting domain via PCR and isothermal assembly. C-to-T base editor (CBE) expression plasmids utilizing the TadCBEd architecture55were generated by digesting pCMV-T7-TadA-CDd-nSpCas9-P2A-EGFP (BKS327; Addgene plasmid ID PENDING) with Xcml and Pmll, and inserting the modified SpCas9 PAM interacting domain via PCR and isothermal assembly. For APOE editing experiments, BE4max68and evoFERNY69CBE architectures were also tested. For BE4max and evoFERNY, mutated PAM interacting domains were cloned into backbone plasmids pCMV-T7-BE4max-nSpCas9-P2A-EGFP (BKS248; Addgene plasmid ID PENDING) and pCMV-T7-evoFERNY-nSpCas9-P2A-EGFP (AHK366; Addgene plasmid ID PENDING) respectively using a similar strategy as used for TadCBEd and ABE8e.
[0285] SpCas9-based prime editor epegRNA73expression plasmids were generated via cloning into pUC19-U6-[BsmBI]-tevopreQl-term (LM1138) and second strand nicking gRNAs were generated via cloning into BPK1520 (Table 1).
[0286] A modified prime editor plasmid74harboring a co-translationally expressed EGFP protein (pCMV-T7-PEmax-P2A-EGFP; LM1589) was generated via SpRYgest46followed by isothermal assembly71.
[0287] Table 1. gRNA targets
[0288] #, SEQ ID NO:
[0289] Animal care and models
[0290] Our animal study followed the tenets of the Association for Research in Vision and Ophthalmology Statement for the Use of Animals in Ophthalmic and Vision Research and the guidelines of the Massachusetts Eye and Ear for Animal Care and Use (under IACUC protocol number 2021N000059). The humanized WT-hRHO-GFP and P23H- hRHO-RFP mice were gifted from Dr. Theodore G. Wensel94,95. The two strains were crossbred to generate heterozygous hRHO-WT-GFP / hRHO-P23H-RFP mouse line, in which the humanized WT-hRHO allele is fused to GFP and the humanized P23H- hRHO mutant allele is fused to the fluorescent protein TagRFPt. Mice are maintained on a 12 h: 12 h light / dark cycle.
[0291] Bacterial-based positive selection experiments
[0292] To perform positive selections, 100 pL of electrocompetent BW25141(ZDE3)75E. coll harboring pl 1-lacY-wtxl plasmid derivatives harboring each of the NGNN PAMs were each electrotransformed with 100 ng of the SpCas9(6AA_NNS) plasmid library (MMW94), which also expresses an gRNA targeting the protospacer sequence GGGCACGGGCAGCTTGCCGG (SEQ ID NO: 4) and chloramphenicol resistance marker. Following a 60-minute recovery in 3 mL Super Optimal broth with Catabolite repression (SOC) media, transformations were spread on LB plates containing either 25 pg / mL chloramphenicol and 10 mM dextrose (non-selective) or 25 pg / mL chloramphenicol + 10 mM arabinose (selective). Transformation efficiency was estimated based on colony count from non-selective plates.
[0293] Colonies from selective plates were picked and used as template for colony PCR to be transferred to mammalian expression plasmids for the HT-PAMDA assay. Colony PCR was performed on the PAM interacting domain sequence of -24-48 colonies that each of the 16 bacterial selections using primers oRAS122 and oBK591. PCR products were purified by paramagnetic beads (generated as previously described19,76and cloned into a gel-purified PvuII- and Xhol-digested mammalian SpCas9 expression vector pCAG-hSpCas9-P2A-EGFP (MSP2582) by isothermal assembly. We also cloned randomly chosen variants from the saturation mutagenesis SpCas9 (6AA_NNS) plasmid library (MMW94) without being subject to the bacterial selection strategy. The library was cloned en masse from MMW94 into PvuII- and Xhol-digested MS2582 by isothermal assembly. The cloning reactions for PAM variant enzyme plasmids derived from bacterial selections or that were randomly chosen were transformed into electrocompetent XL 1 -Blue E. coll and plated on LB + carbenicillin. Single colonies were mini prepped (Qiagen) for arrayed sequencing via multiplex PCR (described below).
[0294] Arrayed sequencing of SpCas9 variants via multiplex PCR
[0295] Two pools of staggered amplicons covering the entire expression construct (from CAG promoter, SpCas9 coding sequence, P2A linker, and EGFP ORFs) of MSP2582 were designed using Primal Scheme77(primalscheme.com / ) (Fig. 17). Primers were inspected manually to ensure that none overlapped with the sites of saturation mutagenesis, SpCas9 residues DI 135, SI 136, G1218, E1219, R1335, T1337. 33nt flaps overlapping with Illumina P5 and P7 adapter sequences were then added to each of the forward and reverse primers respectively forming two primer pools: oPooll (oRAS 127-156) and oPool2 (oRAS 157-186). Two separate multiplex PCR reactions using oPooll and oPool2 were performed for each plasmid to be sequenced using ~10 ng plasmid DNA as template. Multiplex PCR was performed as described by Quick et al11, with Pooll and Pool2 PCR products then pooled together to achieve a single PCR pool per plasmid. PCR pools were bead purified and distinct i5 and i7 barcodes were added to each PCR pool in a second round of PCR as previously described44. All barcoded PCR products were then combined into a single pool, bead purified, quantified by Qubit (Thermo Fisher), diluted to 0.3 ng / pL and sequenced a MiSeq sequencer using a 300-cycle v2 kit (Illumina).
[0296] Fastq files were analyzed using custom python script, “call variants.py” (script to be made available upon request) with gap_open_penalty set to 7. Briefly, reads were trimmed using TrimGalore with default settings and aligned to the reference vector map using Bowtie2. Pileup files containing aligned reads were analyzed using “identify variants. py” (script to be made available upon request) with parameters set to: MAX INSERTION FREQ = 0.2, MAX DELETION FREQ = 0.2, call variant when identity below = 0.95, include alleles above freq = 0.01, and MIXED VARIANT FR ACTION = 0.05. Samples were discarded if they contained insertions or deletions, point mutations other than the intended sites of saturation mutagenesis, or mixed reads indicating more than one plasmid per colony. Unique plasmid sequences containing only mutations in the intended 6 positions were used for further analysis by HT-PAMDA.
[0297] Profiling the PAM requirements of SpCas9 enzymes
[0298] SpCas9 gRNAs in vitro transcribed from roughly 1 pg of Hindlll linearized gRNA T7-transcription plasmid templates (RTW443 and RTW448; Addgene plasmid IDs 160136 and 160137, respectively41using the T7 RiboMAX Express Large Scale RNA Production Kit (Promega). The DNA template was degraded by the addition of 1 pL RQ1 DNase at 37 °C for 15 minutes. gRNAs were purified using beads and refolded by heating to 90 °C for 5 minutes and then cooling to room temperature for 15 minutes.
[0299] The high-throughput PAM determination assay (HT-PAMDA) was performed as previously described50with the exception of omitting the Exonucleasel digestion step following the first PCR. Alternative barcoding primers were used for different sequencing runs with either a four or five nucleotide unique barcode. Briefly, SpCas9 containing lysates were generated by transfecting 1.5xl05HEK 293T cells with approximately 700 ng Cas9 expression plasmid (containing -P2A-EGFP tag) and 1.5 pL TransIT-X2 transfection reagent (Minis). Cells were lysed ~48 hours posttransfection and EGFP signal was measured and normalized to a standard containing 150 nM Fluorescein (Sigma). 4.375 pL of normalized cell lysates were separately complexed with 8.75 pmol of in vitro transcribed gRNAs encoding two distinct spacers, and in vitro cleavage reactions were preformed using with the pre-formed RNPs and two distinct corresponding libraries encoding randomized PAMs (RTW554 and RTW555; Addgene plasmid IDs 160132 and 160133, respectively)50. Cleavage reactions were terminated at time points of 1, 4, and 32 minutes. Approximately 3 ng of digested PAM library for each SpCas9 variant and reaction timepoint was PCR amplified using Q5 polymerase (New England Biolabs; NEB) and barcoded with primers containing sample-specific 4 or 5 nucleotide barcodes. PCR products were pooled for each time point, purified twice using paramagnetic beads, and amplified with primers containing adapters and the Illumina i5 and i7 indexes. Libraries were quantified via qPCR using the Universal KAPA Illumina Library qPCR Quantification Kit (KAPA Biosystems) and sequenced on a NextSeq sequencer using either 150-cycle or 75-cycle NextSeq 500 / 550 High Output v2.5 kits (Illumina) at the Dana-Farber Cancer Institute Molecular Biology Core Facility. Sequencing reads were analyzed as previously described11using the HT-PAMDA data analysis pipeline available at github.com / kleinstiverlab / HT-PAMDA.
[0300] Additional analysis of the HT-PAMDA data was performed by clustering the PAM requirements of the characterized variants using the scipy.cluster.hierarchy.linkage() function with ‘optimal ordering’ parameter set to True, ‘method’ set to ‘average’ and ‘metric’ set to ‘correlation’. Flat clusters were generated from the hierarchical clustering using the scipy.cluster.hierarchy.fcluster() function with maximum cluster number set to 12. Sequence logos for the amino acid composition of HT-PAMDA data clusters were generated using Logomaker78.
[0301] Data preprocessing and PAM machine learning algorithm
[0302] The HT-PAMDA-calculated rate constants for each of the 64 three-nucleotide PAMs were compiled for all SpCas9 PAM variant enzymes obtained from bacterial selections (634 enzymes)and also those chosen randomly from the SpCas9(6AA_NNS) library (135 enzymes). Rate constants slower than 10'5were set to 10'5(approximately the detection limit of HT-PAMDA). Logio rate constants were normalized to center the mean at zero and standard deviation to one. Amino acid identity at the 6 randomized positions was encoded as either 1) a one-hot encoding or 2) a “Georgiev” numerical descriptor51(Georgiev encodings were obtained from code by Ofer & Linial79).
[0303] A random 20% of SpCas9 variant enzymes in the dataset were set aside to be used as a test set. We also tested the effect of balancing different classes of PAM variants prior to training; enzymes were divided into classes based on their most preferred four-nucleotide PAM and then all classes were randomly over-sampled to match the size of the largest class. Hyperparameters, amino acid encodings, and model architecture for the final model were chosen based on maximizing R2 score in an internal the 5-fold cross validation in the training set. Briefly, the training set was randomly sub-divided into 5 subsets prior to over-sampling and each 1 / 5 of the data was excluded as a validation set while the remaining 4 / 5 was subject to over-sampling and then training.
[0304] The PAM machine learning algorithm (PAMmla) neural network model was constructed using the Keras TensorFlow API80. The model consists of three DenseQ hidden layers of dimension 512, 256, and 128 with ReLU activation and dropout of 0.2 between each hidden layer. Training was performed for 100 epochs using MSE loss, a batch size of 32, and Adam optimizer with learning rate of 10'4and decay of 0.001. We also explored other potential models, including by assessing: (1) a linear model fit to the data using skleam.linear_model.LinearRegression(), or (2) a random forest model fit to the data using sklearn. ensemble. RandomForestRegressor() with max_depth=10, max_features=None, boostrap=True, and max_samples=0.8.
[0305] SHapely Additive exPlanations96(SHAP) values for the final PAMmla model were obtained using DeepSHAP97. A DeepExplainer object was fit on 200 variants sampled from the training dataset. SHAP values were visualized using the summary plot and force plot functions.
[0306] Visualization of PAMmla predictions
[0307] PAM predictions were generated using PAMmla for all 64 million possible combinations of amino acids at each of the 6 randomized library positions, DI 135, SI 136, G1218, E1219, R1335, and T1337. Predictions were filtered to remove variants with maximum rate constant below 10'4and 10'3: 10,255,193 and 1,890,023 sequences remained after filtering by the two cutoffs respectively. Filtered variants were then downsampled with diversity preservation using scSampler81with parameters: fraction = 0.01, random split = 256. UMAP82was then fitted to the data with parameters metric- correlation', n_neighbors=10.
[0308] Human cell culture and transfections
[0309] Human HEK 293T cells (ATCC) were cultured in Dulbecco’s Modified Eagle Medium (DMEM) supplemented with 10% heat-inactivated FBS (HI-FBS) and 1% penicillin / streptomycin. The supernatant media from cell cultures was analyzed monthly for the presence of mycoplasma using MycoAlert PLUS (Lonza).
[0310] HEK 293T cells were seeded at a density of -20,000 cells per well in 96-well plates -20 hours prior to transfections. For nuclease experiments, transfections were performed using 29 ng of nuclease expression plasmid, 12.5ng of gRNA expression plasmid and 0.3 pL of TransIT-X2 (Minis) in a total volume of 15 pL Opti-MEM (Thermo Fisher) according to manufacturer instructions. The transfection mixtures were incubated for -15 minutes at room temperature and distributed across the seeded HEK 293T cells. Base editor experiments with ABE8e-SpCas9 or TadCBEd-SpCas9 expression plasmids were performed using 70 ng of base editor plasmid, 30 ng of gRNA plasmid, and 0.72 pL of TransIT-X2. Genomic DNA was collected from all transfections after ~72 hours by removing media and resuspending in 100 pL of quick lysis buffer (20 mM Hepes pH 7.5, 100 mM KC1, 5 mM MgC12, 5% glycerol, 25 mM DTT, 0.1% Triton X-100, and 60 ng / pL Proteinase K (NEB)), heating the lysate for 6 minutes at 65 °C, heating at 98 °C for 2 minutes, as previously described11.
[0311] Generation of cell lines
[0312] We generated a various HEK 293T cell lines bearing therapeutically relevant mutations. First, we generated a cell line harboring the RHO P23H mutation via prime editing83. Transfections were performed as described above using HEK 293T cells with 70 ng of prime editor expression plasmid pCMV-T7-PEmax-P2A-EGFP, 38 ng of pegRNA expression plasmid, 12.5 ng of nicking gRNA expression plasmid, and 0.79 pL of TransIT-X2 (Minis) in a total volume of 20 pL Opti-MEM. Cells were grown for approximately 72 hours prior to extracting gDNA to assess editing efficiency in bulk transfected cells (by next-generation sequencing (NGS) as described below). To create the cell line, the top-performing PE (LM1589), pegRNA (AHK209), and ngRNA (AHK205) combination was re-transfected into low passage HEK 293T cells. Transfected cells were grown for approximately 72 hours prior to dilution plating into 96-well plates. Wells containing single colonies were identified and grown until confluent, and then transferred into 48-well plates with some cell mass reserved to extract genomic DNA (gDNA) for genotyping via PCR and NGS to verify introduction of RHO P23H.
[0313] We also generated a HEK 293T cell line bearing the s4 Alzheimers risk allele via prime editing, similar to as described above. Since HEK 293T cells naturally encode the s3 allele (C130, R176), we needed only to introduce the C130R mutation to achieve the s4 allele (R130, R176). Transfections were performed as described above using HEK 293T cells with 70 ng of prime editor expression plasmid pCMV-T7- PEmax-P2A-EGFP (LM1589), 38 ng of pegRNA expression plasmid (AHK5), 12.5 ng of nicking gRNA expression plasmid (AHK11), and 0.79 pL of TransIT-X2 (Minis) in a total volume of 20 pL Opti-MEM. Cells were grown for approximately 72 hours prior to extracting gDNA to assess editing efficiency in bulk transfected cells (by next-generation sequencing (NGS) as described below). To create the cell line, the top-performing PE (LM1589), pegRNA (AHK5), and ngRNA (AHK11) combination was re-transfected into low passage HEK 293T cells. Transfected cells were grown for approximately 72 hours prior to dilution plating into 96-well plates. Wells containing single colonies were identified and grown until confluent, and then transferred into 48-well plates with some cell mass reserved to extract genomic DNA (gDNA) for genotyping via PCR and NGS to verify introduction of APOE C130R. As described elsewhere (Hille LT et al., submitted), we also created a HEK 293T cell line encoding the 7 / 7>7> E7V mutation causative of sickle cell disease. Briefly, the E7V cell line was generated similar to as described above via prime editing using 70 ng pCMV-PEmax-P2A-hMLHldn (Addgene plasmid ID 174828), 38 ng of an evopreQi epegRNA expression plasmid (LLH439), 12.5 ng of nicking gRNA expression plasmid (LLH50), and 0.79 pL of TransIT-X2 (Minis) in a total volume of 20 pL Opti-MEM. Single cell clones were genotyped via NGS.
[0314] Assessment of nuclease, base editor, and prime editor activities in human cells
[0315] Genome editing efficiencies of nucleases and base editors were determined by targeted amplicon sequencing as previously described11. Briefly, a 2-step PCR-based protocol was utilized to construct Illumina-competent NGS libraries. On-target genome editing activities were analyzed using CRISPResso284in pooled mode with the following custom input parameters for nucleases: — min reads to use region 100 -w 3; for base editors: — min reads to use region 100 — quantification window size 10 — quantification window center -10 — base editor output — min_frequency_alleles_around_cut_to_plot 0.05; and for prime editors: — min reads to use region 100 -w 10.
[0316] Base editing of patient-derived B cell lines
[0317] Epstein Barr virus transformed B cell lines (BCLs) were established from a patient with Chronic Granulomatous Disease harboring the CYBB T362I mutation as previously described98" (the patient was consented via NIH protocol 05-1-0213). BCLs were maintained in RPMI + 10% fetal bovine serum. SpCas9 sgRNAs targeting the CYBB T362I mutation (spacer sequence in Table 1) were synthesized (Synthego). mRNAs encoding the ABE8e-KWRQLC and ABE8e-SpG were produced by in vitro transcription incorporating 100% substitution of the UTP content with pseudoUTP (CELLSCRIPT™). mRNAs were post-translationally capped to >95% and poly(A) tailed to >200 A’s and subsequently purified for removal of double-stranded RNA content (CELL SCRIPT™). Base editing experiments were performed in BCLs by electroporation (EP) (MaxCyte ATx, Program BCL#3) to deliver the ABE mRNA and synthetic sgRNA (Synthego). BCLs were washed with EP buffer (MaxCyte) and resuspended at ~2xl07cells / mL EP buffer. Approximately 0.25-0.5 xlO6BCLs per sample were combined with base editor mRNA (-0.04 pg / pL final concentration), sgRNA (-0.192 pg / pL final concentration), and ScriptGuard RNase inhibitor (1.6 U / pL final concentration; CELLSCRIPT™) in a volume of 25 pL. After EP, cells were transferred to 12-well tissue culture plate and cultured at 0.5-1.0xl06 / mL for a further two days before harvesting genomic DNA (DNeasy kit; Qiagen) for analysis of editing by targeted sequencing.
[0318] Specificity assessment using GUIDE-seq2
[0319] Approximately 20,000 HEK 293T cells were seeded per well in 96-well plates - 20 hours prior to transfection, performed using 29 ng of nuclease expression plasmid, 12.5 ng of gRNA expression plasmid, 1 pmol of the GUIDE-seq double-stranded oligodeoxynucleotide tag (dsODN; oSQT685 / 686)85, and 0.3 pL of TransIT-X2 (Minis). Genomic DNA was extracted -72 hours post transfection using the DNAdvance Kit (Beckman Coulter) according to manufacturer's instructions, and then quantified by Qubit (Thermo Fisher). On-target dsODN integration was assessed by PCR amplification, library preparation, and next-generation sequencing as described above, with data analysis via CRISPREsso284run in non-pooled mode by supplying the target site spacer, the reference amplicon, and both the forward and reverse dsODN-containing amplicons as ‘HDR’ alleles with custom parameters: -w 25 -g GUIDE — plot window size 50. The fraction of alleles bearing an integrated dsODN was calculated as the number of reads mapped to the forward dsODN amplicon plus the number of reads mapped to the reverse dsODN amplicon divided by the sum of the total reads mapped to all three amplicons.
[0320] GUIDE-seq2 reactions were performed essentially as described elsewhere (Lazzarottto et al., manuscript in preparation) with minor modifications. Briefly, the Tn5 transposase was prepared by combining 36 pL hyperactive Tn5 (1.85 mg / mL, purified as previously described86), 15 pL annealed i5 adapter oligos encoding 8 nucleotide (nt) barcodes and 10-nt unique molecular indexes (UMIs), with 52 pL 2x Tn5 dialysis buffer (100 mM HEPES-KOH pH 7.2, 200 mM NaCl, 0.2 mM EDTA, 2 mM DTT, 0.2% Triton X-100, and 20% glycerol) for 60 minutes at 24 °C. Tagmentation reactions were performed in 40 pL reactions for 7 minutes at 55 °C, containing approximately 250 ng of genomic DNA, 8 pL of the assembled Tn5 / i5 - transposome, and 8 pL of freshly prepared 5x TAPS-DMF buffer (50 mM TAPS- NaOH, 25 mM MgCh, and 50% dimethylformamide (DMF)). Tagmentation reactions were halted using 5 pL of a 50% proteinase K (NEB) solution (mixed with H2O) with incubation at 55 °C for 15 minutes, purified using SPRI-guanidine magnetic beads, and analyzed via TapeStation with High Sensitivity D5000 tapes (Agilent). Separate PCR reactions were performed using dsODN sense- and antisense-specific primers using Platinum Taq (Thermo Fisher), with a thermocycler program of 95 °C for 5 minutes, followed by 15 cycles of temperature cycling (95 °C for 30 s, 70 °C (-1 °C per cycle) for 120 s, and 72 °C for 30 s), 20 constant cycles (95 °C for 30 s, 55 °C for 60 s, and 72 °C for 30 s), an a final extension at 72 °C for 5 minutes. PCR products were purified using SPRI beads and analyzed via QIAxcel (Qiagen) prior to sample pooling to form single sense- and antisense- libraries. Libraries were purified using the Pippin Prep (Sage Science) DNA size selection system to achieve a size range of 250-500 base pairs. Sense- and antisense- libraries were quantified using Qubit (Thermo Fisher) and pooled in equal amounts to achieve a final concentration of 2 nM. The library was sequenced using NextSeql000 / 2000 P3 kit (Illumina) with cycle settings of 146, 8, 18, 146. Demultiplexed sequencing reads were down sampled to ensure equal numbers of reads for samples being compared using the same gRNA. Data analysis was performed using an updated version of the open-source GUTDE- seq2 analysis software87(github.com / tsailabSJ / guideseq / tree / V2) with max mismatches parameter set to 6.
[0321] Homology modeling of protein structures
[0322] Amino acid substitutions in SpCas9-derived PAMmla predicted enzymes were homology modeled in Coot (v0.9.8.93)92using the mutate and rotamer selection functions. Most amino acid and PAM DNA base substitutions were modeled using the structure for SpG (PDB: 8U3Y)17, except T1337R, T1337K, and T1337C substitutions or the MRRWMR enzyme variant, which were modeled using the structure for VRER (PDB: 5FW3)6. Homology models were visualized using ChimeraX (vl.8)93.
[0323] In silico directed evolution
[0324] Custom variants to target RHO P23H and APOE-s4 were evolved using a custom python script: “evolve vars.py” (made available upon request). For / / 7 / O, evolution was initiated from the starting amino acids DI 135, SI 136, G1218, E1219, R1335, and T1337 (wild-type SpCas9) with the following custom parameters; starting mutations per variant: 4; variants per round of evolution: 1000; n best variants to keep after each round: 10; decay mutation rate after n rounds plateau: 3; PAM to maximize: NGTG; and additional PAM cutoffs of NGGG < -3.7. For APOE , evolution was initiated from the starting amino acids DI 135, SI 136, G1218, E1219, R1335, and T1337 (wild-type SpCas9) with the following custom parameters; starting mutations per variant: 4; variants per round of evolution: 1000; n best variants to keep after each round: 10; decay mutation rate after n rounds plateau: 3; PAM to maximize: NGTA; and additional PAM cutoffs of NGGC > -2.5.
[0325] Sub-retinal injections, in vivo electroporation, and retinal cell collection
[0326] A DNA solution containing 1.6-2 pg of pCBh-Cas9-P2A-mTagBFP2 plasmids that express Cas9 PAM variant enzymes (WT SpCas9, RAS3575; SpG, RAS3583; SpCas9-MRRWMR, RAS3579; SpCas9-KRHWMR, RAS3594), and 0.8-1 pg sgRNA plasmids (AHK383) were injected into the subretinal space of neonatal pups (P0-P2) using established methods62. Briefly, pups were anesthetized by hypothermia. The fused upper and lower eyelids were separated. A small incision was made at the limbus using a 30-gauge needle, and 0.5 pl of plasmid DNA mix was injected into the sub-retinal space of right eye through the limbal incision using a Hamilton syringe with a 33-gauge blunt-ended needle. Left eyes were used as a negative control. The injected DNA plasmid was electroporated into retinal cells using a 7 mm diameter tweezer-type electrode (Model 520, BTX-Harvard Apparatus, Holliston, MA), and the electroporation parameters were set at five 90 V square pulses, 50 ms duration with 950 ms intervals (ECM830, BTX).
[0327] Two-three weeks post injection, the mice were euthanized. The retinas were dissected out through a corneal incision, placed into a drop of BGJB culture medium (ThermoFisher, # 12591038) in a petri dish, and examined for BFP expression under a fluorescent microscope. For dissociation, retinas were transferred into a tube containing 400 pl of solution (Img / ml pronase and 2mM EGTA in BGJB medium) and incubated at 37 °C for 30 minutes. Retinal samples were broken into single cells by pipetting up and down 20 times. Another 400 pl solution containing 100 u / mL DNase I, 0.5% BSA, 2mM EGTA in BGJB medium was added and incubated at room temperature for 10 minutes. Cell suspensions were filtered through a cell strainer (Falcon, #352235) and sorted for BFP positive cells using MA900 Multi-Application Cell Sorter. Transfection efficiency was -0.1-3%, measured by the percentage of BFP+ cells compared to all retinal cells. Only samples with more than 1,000 BFP+ sorted cells were used for analysis. Collected BFP+ cells were spun down at 15,000 x g for 15 minutes. Genomic DNA was extracted from cell pellet using QuickExtract (Biosearch Technologies, #SS000773), incubated at 60 °C overnight, and heat inactivated at 98 °C for 3 minutes. The humanized Rho P23H region was amplified using primers oRAS1384 / oRAS1385, and then sequenced and analyzed as described above for human cell samples.
[0328] Natural sequence models analysis
[0329] Five iterations of jackhmmer with bitscore 0.9 querying the UniRefl 00 database with the wildtype spCas9 sequence was performed to build a multiple sequence alignment. Any sequence that had more than 30% gaps in the alignment was excluded, and any columns in the alignment with more than 30% gaps were also excluded. A theta reweighting of 0.8 was used such that anything more than 80% similarity is down weighted relative to other sequences. This alignment was used to train an EVCouplings98model, and the delta hamiltonians from this model were used as the sequence scores to approximate protein fitness. Finally, a hybrid LLM and variational autoencoder (VAE) called TranceptEVE100was used to score sequences. To score the full Cas9 sequence, the sliding window method of length 1024 was used. The Large Tranception model was used for the LLM component and four separately trained EVE models" were averaged for the VAE component. The EVE models were trained with the same alignment used for EVCouplings with the same sequence and column coverage cutoffs and theta reweighting parameter.
[0330] Example 1. Scalable characterization of the PAM preferences of hundreds of SpCas9 variant enzymes
[0331] To generate large datasets related to the PAM requirements of SpCas9 enzymes to be used as training data for a ML model, we obtained novel SpCas9 variants with altered PAMs via bacterial selections prior to high throughput PAM characterization using the HT-PAMDA assay41. We first employed a structure and function-informed saturation mutagenesis approach to generate an SpCas9 library comprised of millions of enzyme variants with amino acid substitutions that should impact PAM specificity. Based on SpCas9 structures42and previous engineering efforts43-45, we selected 6 amino acid residues within the SpCas9 PAM interacting (PI) domain to randomize to all possible amino acids (Fig. 1c and Fig. 16a and section titled “Design, cloning, and testing of the SpCas9 PI domain libraries” above. We selected residues R1335 (which forms a base specific contact with the G nucleotide at the 3rdposition of the PAM in WT SpCas9) as well as DI 135, SI 135, G1218, E1219, and T1337 for saturation mutagenesis (leading to an SpCas9(6AA) library of plasmids), all of which have been shown to modulate specificity for the 3rd and 4th PAM bases in the context of the SpCas9-based enzymes SpCas9-VRER, SpCas9-VRQR and SpG41, 43’44. (see a|sosection above titled “Design, cloning, and testing of the SpCas9 PI domain libraries”). To identify SpCas9 PAM variant enzymes with activity on non-canonical PAMs, we performed an extensive set of bacterial -based positive selection assays (similar to those previously described10,43’46’47) to select for variants capable of cleaving target sites bearing each of the 16 possible NGNN PAMs with a fixed G nucleotide at the second position (Figs. 16b, c and Methods). Across the 16 selections using the SpCas9(6AA) library, we obtained 634 unique SpCas9 enzymes encoding different amino acids at the six PI domain positions (D1135, S 1136, G1218, E1219, R1335, and T1337). We observed context-dependent enrichment of amino acids depending on the both the identity of the 3rdor 4thnucleotide of the PAM used for the selection (Fig. 16c). These results suggest that multiple different enzymes can access the same PAM space, making the experimental identification of single optimal or consensus enzymes that exhibit maximal activity on specific PAMs challenging.
[0332] Next, to thoroughly characterize the PAM requirements of each of these 634 unique SpCas9 enzymes that survived in the bacterial selection, we transferred the PI domain sequence to a human expression plasmid that co-translationally expresses EGFP41. To economically sequence hundreds of new plasmids, we used an arrayed plasmid sequencing method for rapid and inexpensive whole ORF identification, similar to previously described48, with the addition of a multiplex PCR step49in order to enable linking of mutated residues across the entire length of the SpCas9 PI domain (Fig. 17). Using the human cell expression plasmids for each of the new SpCas9 PAM variant enzymes, we performed the high-throughput PAM-determination assay (HT- PAMDA)50to fully profile the PAM specificities of each of the 655 SpCas9 enzymes obtained from bacterial selections (Figs. 1c, d). The HT-PAMDA assay comprehensively measures rate constants (k) of nuclease cleavage for an enzyme across a library of all possible PAMs, providing rich kinetic data to quantify the global PAM profile of an enzyme.
[0333] HT-PAMDA experiments revealed that our selections identified enzymes from the SpCas9(6AA) library capable of recognizing each of the 16 NGNN PAMs (FIG. 11). However, few enzymes exhibited their highest activity on the specific PAM that they were selected on, highlighting a limitation of the bacterial-based selection approach (Fig. le). In contrast to the lack of consensus observed when grouping the enzymes based on the PAM that they were selected against which revealed little convergence on optimal amino acids (Fig. 16c), clustering the enzymes by their HT-PAMDA- determined PAM profiles revealed 8 main classes of PAM variants that were each associated with specific amino acid substitutions (Fig. Id). Single enzymes typical of each of the largest PAM preference clusters were selected for visualization (Fig. Id), however these PAM preferences are not exhaustive and do not capture the variation within or outside of these main clusters. Our results highlight how scalable enzyme selections and characterization methods can be deployed to identify enzymes with novel properties, while also revealing challenges for enzyme engineering schemes based solely on experimental methods (i.e. that enzymes recovered from selections rarely exhibit maximal activity against their query sequence, etc.).
[0334] The two most common classes of SpCas9 variant enzymes represented in our dataset from the SpCas9(6AA) library selections were those with NGG PAM (associated with an Arg at residue R1335, as in WT SpCas9) or relaxed PAM profiles for G at the fourth position similar ot SpG or Cas9-NG (associated with an Arg residue at T 1337) (Figs. Id, f). Examples of NGG and NGN variants recovered during selections include KWWERM and MWGAQR (named after amino acid identities at positions 1135, 1136, 1218, 1219, 1335 and 1337, a naming convention used henceforth) which differ from WT and SpG at 4 and 3 amino acid positions, respectively (Fig. Id). The large number of enzymes capable of targeting the NGG and NGN PAM space highlights the plasticity and tolerance of the PI domain to mutation at many positions while specifying similar PAMs. We confirmed the ability of several novel relaxed NGN enzyme variants to edit target sites bearing NGN PAMs in human cells, confirming that many of the variants are as active as SpG on a variety of PAMs, with subtle differences in preferences compared to SpG (Figs. 18a, b). Notably, although the T1337R substitution was observed with most NGN enzyme variants, this substitution was not necessary to specify an NGN PAM as evidenced by the activities of the QWRYVN and MWKSTS enzymes (Figs. 18a, b). Other substitutions that were slightly enriched amongst relaxed NGN variants included DI 135L and SI 136W, mutations found in SpG11(Fig. Id).
[0335] Residue R1335 in SpCas9 mediates a base-specific interaction with the G at the 3rdposition of the PAM in WT SpCas96. We found that although most of the enzyme variants in our dataset with an NGG PAM profile indeed encoded R1335 (Fig. Id), this residue was neither necessary nor sufficient to specify an NGG PAM (Figs. 19a, b). We observed several variant enzymes with NGG PAM preferences that contained R1335I, R133K, R1335N, or R1335V, although they exhibited notably weaker efficiency against NGG compared to enzymes with R1335 (Fig. 19a). We also observed variants containing R1335 that largely specified a PAM other than NGG (e.g. LKAWRS specifying NGT>NGV), or were not active on any PAM (Fig. 19b). These observations further highlight the previously uncharacterized plasticity of the SpCas9 PI domain.
[0336] In addition to NGG and NGN variants, our bacterial -based selections also identified a variety of variants with preferences for non-canonical nucleotides in the 3rdand 4thpositions of the PAM. For instance, enzymes that require A at the 3rdposition of the PAM (e.g., KWAVIG and LRDEQR), that require a C in the 3rdposition (e.g. KWTYER), and that require either pyrimidine base at the 3rdposition of the PAM (e.g. GWNWPS) (Fig. Id). Notably, in many cases we observed a nucleotide preference at the 4thnucleotide of the PAM, extending the PAM from the canonical 3 nt to 4 nt (specifying 3 PAM bases instead of 2). For instance, we observed enzymes with preference for T in the 4thposition of the PAM (e.g. MCRQCT and KWAVIG), for G in the 4thposition (e.g. MWGAQR, LRDEQR and KWTWYR), and a C in the 4thposition (e.g. MRVVKH) (Fig. Id). A preference for an extended PAM may hold advantages for minimizing off-targets, as has been demonstrated for other previously described enzymes10.
[0337] While our bacterial selection experiments led to the evolution of Cas9 variants with a wide range of PAM requirements, in general, an enzyme’s most efficiently targeted PAM did always not correlate with the PAM on which that variant was selected (Fig. le and Fig. 20a). We therefore sought to utilize the information gained from the selected variant sequences to rationally design more optimal PAM-selective enzymes that collectively targeted a broader range of NGNN PAMs. One potential simple model of PAM interaction at the protein:DNA level could be that each amino acid contributes independently to PAM recognition. In such a model, the most enriched amino acid at each position might be optimal for recognition of that PAM. To test this model, we cloned “consensus” enzymes for each of the 16 NGNN PAMs by combining the most enriched amino acids at each of the six positions recovered from each selection. If two amino acids were enriched equally, we cloned both enzymes. When assessing the global PAM requirements of these consensus enzymes using HT- PAMDA (Fig. 20b), we found that consensus enzymes often did not efficiently target the PAM for which they were designed, with only 4 / 31 consensus enzymes agreeing with their PAM class from which they were selected (Fig. le and Fig. 20c). The rationally designed consensus enzymes behaved more poorly on average than the selected enzymes (Fig. le and Fig. 20d), indicating that mutations at these six positions do not contribute to PAM preference independently, but rather interact with one another in an epistatic manner. These results suggest that selections alone are insufficient to systematically obtain enzymes with desired PAMs.
[0338] To approximate the fraction of the 64 million enzymes encoded with our SpCas9(6AA) library that retained editing activity on any PAM, we performed HT- PAMDA on 135 randomly selected library members (without performing bacterial selections). Unexpectedly, we found that -18% of the enzymes in the SpCas9(6AA) library permitted editing on one or more PAMs, many of them with non-canonical PAM preferences (Fig. If). This observation suggested that amongst the 64 million possible combinations of residues in our 6-position library, more than 10 million may be functional PAM variants, the vast majority of which remain uncharacterized.
[0339] Example 2. Machine learning to predict PAM preference from amino acid sequence
[0340] Given that we had experimentally assessed only a small number of enzymes from the SpCas9(6AA) library (-0.001%), and that our data supports a substantial fraction of the enzymes to be active on at least one PAM (>15%; Fig. If), we envisioned that many additional enzyme variants with useful PAM requirements remained to be characterized. This, together with the observation that enriched amino acid substitutions were often neither necessary nor sufficient to rationalize the PAM preference of most variants, led us to seek a method to more systematically investigate the relationship between amino acid sequence and PAM specificity. We therefore sought to explore the entire fitness landscape of our SpCas9(6AA) library via ML to predict the PAM requirements of all 64 million enzymes.
[0341] To train ML models, we utilized the rich kinetic data derived from our PAM profiling experiments. The HT-PAMDA assay outputs ideal data to train a ML model since quantitative rate constants are provided against all possible PAMs for each SpCas9 enzyme variant. We utilized two datasets to train our model, including HT-PAMDA from functional enzymes derived from our bacterial selections on the 16 different NGNN PAMs, and random (non- selected) enzymes from the SpCas9(6AA) library. We reasoned that including inactive variants in our training set, as opposed to only variants that are active on at least one PAM and thus survived the bacterial selections, would be necessary learn combinations of amino acids lead to both active and inactive enzymes. To account for the imbalanced nature of our training data (largely biased towards NGGN and NGTN enzymes, and variants preferring a G at the 4thposition; Fig. If), we assigned a label to each training example based on its most active PAM and randomly over-sampled to balance across PAM classes.
[0342] Using this data, we trained a model that to predict the k on each of the 64 PAMs (positions 2-4) when provided with a 6AA sequence as input (Fig. 2a). Linear regression, random forest, and neural network models were tested in combination with different feature encodings as inputs to the model (one-hot encoding of each amino acid substitution; one-hot encoding of all single plus pairwise amino acid combinations; and Georgiev51, a physiochemical descriptor that can improve performance of some ML models for protein engineering52) (Fig. 21a). Each combination of model architecture and amino acid encoding was compared using an internal 5-fold cross-validation (Fig. 21a). In all cases, the neural network combined with one hot encoding consistently outperformed other models, thus we termed this model the protospacer adjacent motif machine learning algorithm (PAMmla) and used it for subsequent analyses.
[0343] We evaluated PAMmla on a test set comprising 20% of HT-PAMDA data that was held out from training, revealing accurate rate constant predictions for unseen SpCas9 variant enzymes (Pearson’s r=0.91) (Fig. 2b). We evaluated PAMmla using two additional train test splits (Fig. 21b) which showed similar correlations (r = 0.92 and r = 0.92) demonstrating that these results are consistent regardless of the allocation of enzymes into training and testing sets. The test set was further sub-divided into progressively more challenging subsets containing variants with increasing numbers of mutations relative to the most similar variant in the training set, demonstrating generalizability to variants that are dissimilar to those on which the model was trained (Fig. 2c). Across different PAM classes, the model performed relatively consistently indicating that the model was able to generate accurate predictions even for PAM classes which had relatively few training examples (Figs. 21c-f). PAMmla models trained with over-sampling to balance PAM classes within the training set improved performance on under-represented PAM classes in the test set (Fig. 21g).
[0344] The model also accurately predicted combinations of amino acids that led to inactive enzymes (defined as inability to cleave any substrates with rate constant faster than 1 O'4 3) with AUROC = 0.99 (Fig. 2d). The model was able to accurately classify enzymes as active vs inactive (using a threshold maximum predicted rate constant of 1 O'4 3) for 149 / 154 enzymes (97%) in the test set (Fig. 2e). We selected a rate constant of k = 10'4 3to define a non-targetable PAM since the HT-PAMDA values for bona fide inactive enzymes show noise between rate constants of 10'5- 10'4 3. Our hypothesis that these slow constants are due to random noise and not experimentally reproducible weak enzyme activity is supported by the fact that the enzymes we categorized as “inactive” by this criteria have no correlation between HT-PAMDA experimental replicates using different spacer target sites (Pearson’s r = -0.1-0.1) whereas active variants generally have robust Pearson’s correlations of 0.7-0.9 between replicates (Fig. 21h). We assessed whether our ML model was overfit to experimental noise by measuring correlation between predictions and true rate constants for the inactive enzymes within the test set. The model did not have predictive power (Pearson’s r = 0.12) for the random noise in the PAM profiles of inactive enzymes (Fig. 21i). Thus, we concluded that for enzymes with low measured activity where the model produces strong correlations (for example the hamming distance = 4 subset) these predictions are likely meaningful and not an artifact of over-fitting.
[0345] As further validation, we tested PAMmla’s ability to recapitulate HT-PAMDA- determined PAM profiles of preexisting rationally designed or evolved SpCas9 PAM variants from the literature (not included in the training or test sets). We were able to generate accurate PAM profiles for the variants SpG, VRER, and VRQR (Pearson’s r = 0.99, 0.71, and 0.93, respectively) each of which contain two mutations relative to the closest member of the training set (Figs. 2f-h). We also predicted the PAM profile of xCas9, an enzyme which contains only a single PI domain mutation (E1219V) relative to WT SpCas9 with additional mutations outside the PI domain. PAMmla produced a predicted PAM profile that was remarkably similar to the HT-PAMDA determined profile for xCas9 (Fig. 2i), supporting previous evidence that the single E1219V mutation was the major contributor to xCas9’s altered PAM11.
[0346] Example 3. ML-assisted discovery of sequence-divergent PAM variant enzymes Next we sought to determine whether PAMmla could aid in the discovery of novel PAM variants. We utilized PAMmla to predict the PAM profiles for each of the 64 million possible enzymes in the SpCas9(6AA) library (Fig. 22). The resulting enzymes were sorted by their predicted PAM profiles to classify them by activity alone (highest activity) or selectivity (kpAM / sum(kaii PAMS)) for each of the 16 NGNN PAMs (Fig. 3a). Amongst the top 10 predicted variants, we chose between one and five variants for each criterion and PAM to experimentally validate PAMmla predictions using HT-PAMDA (for a total of 281 enzymes; Fig. 3b). Amongst this group of enzymes, the PAMmla predictions for each PAM correlated well with experimentally obtained rate constants (Pearson’s r = 0.90) (Fig. 3c). Most of the top predicted variants encoded 2 or 3 mutations (and some up to 4) relative to the most similar previously characterized variant (Fig. 3d), indicating that PAMmla can accurately predict the PAM requirements of enzymes highly dissimilar from the training set and enabling exploration of a vastly understudied sequence space. Clustering the experimentally determined PAM profiles from the top PAMmla predicted variants revealed 10 main clusters (Figs. 3b, e). Clusters comprised new enzyme variants with novel PAM requirements compared to those previously seen (e.g., NGTC, NGCT, NGC(A / C), and relaxed with NGG anti -preference), along with sequence-diversified examples of enzymes with PAM profiles similar to those obtained from bacterial selections (e.g. NGG, NGN, NGC). Both of these enzyme classes provide new insight into combinations of amino acids within the SpCas9(6AA) library that contribute to useful PAM profiles. For instance, the enzymes bearing the mutations E1219Y / R1335Q were largely associated with an NGC(A / C) PAMs preference, and enzymes encoding
[0347] D 1135L / S 1136W / E 1219W / R 1335E mutations prefer NGCT and NGTT P AMs (Fig. 3b). The discovery of enzymes with these unique PAM profiles is unlikely with bacterial -based selection methods, highlighting the utility of PAMmla towards investigating the plasticity of SpCas9 PAM interaction.
[0348] In addition to being dramatically less labor-intensive and costly to perform, the nomination of optimal SpCas9 variant enzymes using PAMmla that are capable of targeting a PAM of interest also showed a higher success rate of producing a PAM variant with desired properties than bacterial selections. For instance, enzymes from the bacterial selections were over-enriched for relaxed PAM variant enzymes and those that recognize NGG PAMs (likely due to the breadth of amino acid sequences compatible with these PAMs). These two classes of enzymes are easily avoided when using PAMmla, as enzymes obtained by sorting PAMmla predictions to maximize activity resulted in variants which, on average, had higher activity on the PAM of interest than variants obtained from bacterial selections (Fig. 3f). When we instead obtained enzymes by sorting PAMmla predictions for selectivity on a PAM of interest, the resulting variants were more specific for the target PAM than those obtained from bacterial selections (Fig. 3g). In contrast to bacterially selected or rationally designed variants, the majority of PAMmla predicted enzymes exhibited maximum activity on the PAM for which they were designed (Fig. 3g, Figs. 20a and 23). Thus, PAMmla can be used to obtain enzymes which are tailored to the PAM of interest rather than generalist PAM relaxed enzymes, which tend to be the most common bacterial selection-based engineering trajectory.
[0349] We investigated potential functional roles of the amino acid substitutions in PAMmla predicted enzymes. To investigate the potential functions of PAMmla predicted alterations to SpCas9, we modeled a subset of amino acid substitutions on the structures of SpCas9-VRER and SpG (Figs. 27A-I) and analyzed the relative importance of different mutations on PAMmla predictions for each PAM using SHapely Additive exPlanations (SHAP) analysis (Fig. 26). Mutations can be broadly divided into three categories, including those that alter nucleotide bias at the 3rdposition of the PAM (Figs. 27a-c), that alter nucleotide bias at the 4thposition of the PAM (Figs. 27d-f), and that result in general modification of PAM recognition by base-independent alterations (Figs. 27g-i).
[0350] In the first category, an E1219Y substitution can interact with the amino group of bases at the 3rdposition of the PAM, facilitating cytosine recognition (Fig. 27a). Additionally, R1335Q or R1335E can enable major groove readout of a C-G base pair (Fig. 27b). R1335M and E1219C / I / W (or other substitutions to hydrophobic amino acids) enable 3rdposition T readout by forming a hydrophobic pocket compatible with the thymine methyl group (Fig. 27c). SHAP values also suggest that SI 136R may be important for readout of NGT PAMs (Fig. 26), but further structural investigations are required.
[0351] In the second category, T1337R can read out G at the 4thposition of the PAM, as previously observed in SpCas9-VRER, SpG, and SpRY (Fig. 27d). The T1337K substitution may facilitate interaction with the -OH group of 4thposition nucleotides enabling readout of G or T (Fig. 27e). T1337C and R1337L or other hydrophobic amino acid substitutions allow for 4thposition T readout through hydrophobic interactions (Fig. 27f). T1337S may result in 4thposition C recognition through amino group readout, however more substantial repositioning of the PAM DNA would be required. There are many additional mutations indicated to be important based on SHAP values (Fig. 26) for which structural studies to investigate the possibility of interplay between multiple mutations to reposition either the PAM DNA backbone or PI domain are needed.
[0352] In the third category, DI 135 mutations to L, K, M, or A resulted in disruption of R1114 coordination, allowing R1114 to form novel nonspecific contacts with the nontarget strand (NTS) backbone (Fig. 27g). G1218 substitution to positive K, R, or H side chains may form additional nonspecific contacts with the NTS backbone (Fig. 27h), similar to the G1218K mutation in SpG and SpRY13. The DI 135L and SI 136W substitutions result in a shift of the NTS and TS towards the PI domain, potentially permitting the formation of novel contacts (Fig. 27i) as observed for SpG and SpRY13. This category of mutations typically resulted in a positive impact on model output across most PAMs, supporting their role in base-agnostic PAM interaction (Fig. 26).
[0353] Using SHAP51analysis to determine the relative importance of different mutations on PAMmla predictions (Fig. 26) and structural modeling of mutations on SpCas9- VRER52and SpG6(Figs. 27A-I), we observed that mutations in PAMmla enzymes can be divided into 3 categories that either: (1) alter 3rdPAM position preference (Figs. 27a-c); (2) alter 4th PAM position preference (Figs. 27d-f); or (3) modify PAM recognition via base-independent interactions (Figs. 27g-I). Generally, more selective PAMmla enzymes harbored mutations that form base-specific contacts with the DNA bases of the PAM, whereas PAMmla enzymes with broader PAM specificity shared non-specific activity-potentiating mutations with previously described PAM-relaxed enzymes that form nonspecific contacts with the DNA backbone3,6. Many PAMmla enzymes shared combinations of mutations from both categories, potentially achieving a balance between specifying novel PAMs and retaining high activity.
[0354] Example 4. Assessment of PAMmla predicted nucleases and base editors in human cells
[0355] We tested the editing efficiencies of PAMmla generated enzymes in HEK 293T cells, prioritizing those with preferences for PAMs which have not been previously described in the literature: NGAT, NGCM, NGCN, NGTN, NGTC, NGDC, and NGTG. In addition to the ML nominated enzymes, we also assessed several enzymes from the training set that were identified from our bacterial selection experiments. For both enzyme classes, we predicted their PAM requirements using PAMmla (Fig. 4a), which correlated well with their HT-PAMDA determined profiles (Fig. 24A-B). We tested the nuclease forms of these enzymes to generate insertion or deletion mutations (indels) across 32 endogenous target sites in HEK 293 T cells compared to SpG (an engineered SpCas9 nuclease enzyme variant capable of targeting sites with NGNN PAMs11). In general, ML nominated enzymes showed improved average editing efficiencies compared to SpG (Fig. 4b). For each site tested, we identified at least one (and often several) ML nominated enzymes that exhibited similar or higher editing efficiencies to SpG, while maintaining a more selective PAM preference (Figs. 25a- g). In some cases, ML nominated enzymes were substantially more efficient at creating indels compared to the most similar variants from the training set. For example, the PAMmla predicted enzyme LWKYQS, showed a ~1.5-fold (65% vs 45% average editing) increase in efficiency on NGCM PAM sites compared to its nearest neighbor from the training set, LWKYSS, which contains only one amino acid difference (Fig. 4b, NGCM variants). In other cases, ML nominated enzymes exhibited similar average editing compared to enzymes from the training set (e.g. NGTN variants). These observations varied by site (Figs. 25a-g), suggesting potential utility of testing multiple enzymes for each target of interest. ML-nominated enzymes generally exhibited low editing efficiencies against a negative control site encoding an NGGG PAM (Fig. 25h), consistent with PAMmla predictions and HT-PAMDA- determined rate constants for this PAM. Together, these observations demonstrate that PAMmla-predicted enzymes have improved human cell -based editing efficiencies compared to SpG or those derived from bacterial selections, and that PAMmla enzymes have altered rather than relaxed PAM specificities like SpG.
[0356] Base editing is one exemplary genome editing approach where precise target site positioning (and thus PAM availability) is critical, since the narrow ‘edit window’ of base editors must be positioned over the base of interest53. Thus, the advantageous properties of PAMmla predicted enzymes could be beneficial for base editing efficiency and precision. For each of the seven PAM categories tested above, we selected a single PAMmla-derived enzyme to assess A-to-G and C-to-T editing efficiencies (in the contexts of ABE8e54and TadCBEd55architectures, respectively). The PAMmla enzymes were selected for analysis as BEs based on nuclease editing efficiencies in human cells as well as selectivity for their PAM of interest based on HT-PAMDA profiles. For each base editor, we measured editing in HEK 293T cells at 3 genomic loci harboring the preferred PAM of the corresponding Cas9 variant (Figs. 4c, d, and Figs. 28A-H, 29A-H). ML-nominated enzymes functioned efficiently for both types of base editing in the ABE8e and TadCBEd enzyme contexts. We also compared base editing efficiencies to the commonly used relaxed PAM variant enzymes SpG and SpRY. The base editing efficiencies of PAMmla-generated BEs were superior to SpRY in both ABE8e and TadCBEd contexts, while PAMmla enzymes were more active compared to SpG as ABE8e constructs and similarly efficient to SpG as TadCBEd constructs depending on the site (Figs. 28h, 29h). We next tested the ability of ML-derived enzymes to edit at the therapeutically relevant HBB WIN mutation causing sickle cell disease. Base editing of the pathogenic E7V allele to the benign E7A (Makassar) allele has been previously shown to rescue sickle cell disease in mice56. We selected four PAMmla-derived enzymes which are predicted to target NGCA PAMs, LWQYQH, MWKYQS, LWKYQS, and MWKYQA, in the context of ABE8e to edit the E7V allele to E7A, using a guide that positions the target adenine at position 9 of the protospacer. We compared this editing strategy to a highly effective previously reported strategy56utilizing a guide positioning the target adenine at position 7 of the protospacer coupled with PAM-relaxed Cas9-NRCH-ABE8e recognizing a CACC PAM (Fig. 30A-C). We found that all PAMmla variants tested with the A9 strategy resulted in similar editing to the previously reported A7 strategy using Cas9-NRCH. Since SpCas9 NRCH should theoretically target both CACC PAMs (A7 strategy) and TGCA PAMs (A9 strategy), we additionally tested SpCas9-NRCH using the A9 strategy. In this case, the PAMmla-derived enzymes resulted in ~2x editing compared to SpCas9-NRCH when compared using the same A9 gRNA (Figs. 30A-C). Together, these observations demonstrate that ML-derived PAM specific editors are capable of efficient editing as both nucleases and base editors, and can often increase editing efficiency compared to published relaxed PAM enzymes. Additionally, our observation that the relative efficiencies of each editor varied depending on the target site highlights the utility of having a catalog of enzymes, rather than a single relaxed enzyme, to maximize editing efficiency for any site of interest.
[0357] Example 5. Specificity analysis of PAMmla-predicted enzymes
[0358] SpCas9 enzymes with relaxed PAM requirements are more prone to off-target editing, since they search a larger fraction of the genome and thus encounter a larger number of potential off-target sites11 17 18. In contrast, PAMmla-predicted enzymes with altered PAM requirements should encounter fewer potential off-targets due to their more limited PAM tolerances. To test this hypothesis, we performed a modified cellbased unbiased genome-wide off-target assay to identify nuclease-cleaved off-target sites (GUIDE-seq2; Lazzarotto et al., manuscript in preparation). We performed GUIDE-seq2 to compare the off-target profiles of ML-predicted altered PAM variant enzymes to the SpG and SpRY variants that exhibit varying degrees of PAM relaxation. Our GUIDE-seq2 experiments revealed comparable on-target editing between each of the three enzyme classes (Fig. 5a and Fig. 31a), but the total number of off-target sites detected for each gRNA were reduced for PAMmla-derived enzymes with altered PAM requirements relative to SpG and SpRY (ranging from 26%-93% reduction compared to SpG and 49%-96% reduction compared to SpRY; Fig 5b). The PAMmla enzymes resulted in a higher proportion of on-target to off- target GUIDE-seq reads relative to SpG and SpRY for all sites, including those with modest and high numbers of off-target sites (Figs. 5c and 5d, respectively). Off-target sites detected for each PAMmla-predicted enzyme were largely also detected in the SpG or SpRY sample with the same gRNA, though there were some off-targets specific to the altered PAM variant enzymes (Fig. 5e). The PAMs observed at each off-target site for each PAMmla-derived enzyme reflected their cognate HT-PAMDA- determined PAMs, while the aggregate PAMs observed for SpG and SpRY were more relaxed (Fig. 31c). These results demonstrate that PAMmla enzymes with altered PAMs can reduce genome-wide off-targets compared to PAM-relaxed enzymes.
[0359] Consistent with its more constrained PAM requirement, SpG had fewer off-targets than SpRY for all gRNAs except for the NGTA-2 site, which shows an unusually repetitive set of off-target sites with the majority containing a conserved NGTG PAM (Fig. 31c). The increase in off-targets with SpG compared to SpRY at this site is likely because SpG generally targets NGN PAMs more efficiently than SpRY11. Both SpG and SpRY have highest affinity for target sites with G in the fourth position of the PAM, while MKKCMN has relatively even affinity across all nucleotides at the fourth position. These observations support how MKKCMN exhibits similar on-target editing compared to SpG and SpRY but a reduction in editing across the large number of off-targets encoding NGTG PAMs. To verify the on- and off-target benefits of PAMmla versus generalist enzymes in human patient-derived cells, we sought to correct a CYBB mutation in cells from an individual with X-linked Chronic Granulomatous Disease (X-CGD). X-CGD patients are highly susceptible to invasive infections and hyperinflammation that results in significant morbidity and early mortality. We previously demonstrated that CYBB mutations may be correctable through ex vivo editing of hematopoietic stem cells57,58. To correct the CYBB T362I mutation via base editing, we identified a target site with an NGAT PAM that locates the causative G-to-A mutation at position 8 within the protospacer (Fig. 5e). Using patient-derived Epstein Barr virus-transformed B cells harboring the CYBB T362I mutation, we electroporated mRNA encoding either ABE8e-SpG or the PAMmla-derived ABE8e-KWRQLC enzyme along with the A8 sgRNA. Both ABEs resulted in highly efficient levels of T362I correction (up to >90%) with minimal bystander editing (Fig. 5f). GUIDE-seq2 assessment of off- targets using the CYBB T362I-targeted sgRNA in HEK 293T cells revealed only a single low-likelihood off-target site detected when using the PAMmla predicted KWRQLC enzyme compared to 5 off-target sites with SpG (Figs. 31d,e), resulting in a superior on-to-off-target ratio for the PAMmla enzyme compared to the PAM relaxed SpG (Fig. 5g).
[0360] Example 6. in silico ML-assisted directed evolution
[0361] Activity and selectivity (the metrics used to sort 64 million PAMmla predictions above) are just two examples of properties that can be optimized using PAMmla predictions. One strength of PAMmla is that it can be used to predict enzymes capable of more nuanced context-specific PAM requirements for enzymes that would be challenging, if not impossible, to select for using traditional bacterial selections. In some cases, we envisioned that it may be desirable, for example, to maximize activity on a PAM of interest while minimizing activity on a second PAM (e.g., an alternate PAM present either at an off-target site or another allele for allele-specific editing of heterozygous mutations57). However, sorting all possible PAMmla predictions for the 64 million SpCas9(6AA) enzymes is a time-consuming operation that is challenging to perform for large numbers of different sorting metrics. Additionally, extension of the PAMmla approach to libraries of enzymes with >6 amino acids would quickly expand the number of amino acid combinations making it impossible to store predictions for all library members. We therefore sought to develop an alternative method to rapidly identify PAM variants with customizable properties while minimizing the requisite computational resources and time. We developed an in silico directed evolution (ISDE) approach that utilizes PAMmla predictions to computationally engineer SpCas9 variant enzymes with customizable properties in a stepwise manner (Fig. 32A).
[0362] Activity and selectivity are two of several potential properties that can be optimized using PAMmla predictions. To prioritize enzymes with useful properties for experimental validation, we analyzed PAMmla predicted PAM profiles of all 64 million possible enzymes in our library for both activity on each of the 16 NGNN PAMs, and selectivity for each of the 16 NGNN PAMs. This led to a collection of enzymes that we envision will be useful for various genome editing applications. However, an advantage of prioritizing enzyme sequences using PAMmla is that users can also search for more nuanced context-specific PAM profiles. For example, it is possible to search PAMmla predictions for enzymes which efficiently target a combination of PAMs, or target one PAM while avoiding targeting a one or more other PAMs. There are many potential sorting metrics and what defines an optimal PAM variant is likely to often be specific to the desired edit of interest, making precomputation of all optimal enzymes challenging. To enable rapid generation of enzymes for any user-defined sorting criteria, we developed a PAMmla-guided in silico directed evolution (ISDE) method to permit users to rapidly explore the mutational fitness landscape (Fig. 32a).
[0363] Creating a rank-ordered list of all 64 million PAM profiles is computationally demanding and requires downloading or generating predictions for all 64 million SpCas9(6AA) enzymes. In contrast, evolving a variant using ISDE takes only seconds, allowing for interactive exploration of the SpCas9 mutational space and the PAM profiles of the resulting enzymes. In an ISDE campaign, a starting sequence is computationally mutated to generate a small sub-library (-1,000-100,000 sequences) bearing random amino acid substitutions with a defined hamming distance from the original sequence. PAM predictions are then generated for each member of the sublibrary using PAMmla. A customizable fitness function is used to score each variant according to the desired properties, and a selection step is performed where the most “fit” enzymes are isolated according to the chosen fitness metric. The resulting ISDE enzymes are then used as starting sequences for subsequent rounds of evolution, iterating the process until the fitness score plateaus. Although ISDE permits rapid and interactive enzyme prioritization, it does not ensure convergence on the most optimal predicted enzyme and carries the risk of landing in local maximums of the fitness landscape (similar to traditional experimental directed evolution). This risk can be mitigated by repeating the ISDE process from different starting sequences and varying ISDE parameters. To explore the effectiveness of ISDE at reaching the global fitness maximum, we compared enzymes obtained by exhaustive sorting to those obtained through ISDE for activity maximization on three PAMs (Figs. 32b-d). When using ISDE with default parameters (1,000 random starting sequences, m = 4 starting mutations per enzyme, 5 = 1,000 sampled enzymes per round, n = 10 top variants to keep per round, and p = 1 additional round of evolution after a plateau is reached), we recovered the true maximum enzyme for all 3 replicates for each of the 3 PAMs tested. Of the top 10 true maximum enzymes, we recovered between 5-10 depending on the PAM and replicate. We tested the effect of varying ISDE parameters and found that increasing the number of allowed mutations per enzyme (Fig. 32b) and sampling a greater number of mutants per round increased the probability of recovering top variants (Fig. 32c). Inclusion of additional rounds of evolution following fitness plateau increased the proportion of the top 10 variants recovered (Fig. 32d). In general, although a single ISDE run with default parameters should be sufficient to obtain the true maximum enzyme, additional rounds of evolution and mutations per variant can be utilized to increase the proportion of top enzymes recovered. Performing 3+ repetitions of ISDE and compiling results may be desirable to predict a larger number of maximally optimal enzymes. While it is currently possible to generate predictions for all 64 million possible enzymes from the PAMmla model, if the model is expanded to include additional amino acid positions, the possible sequence space becomes too large to generate all possible predictions. Therefore, methods such as ISDE will become necessary, where computation time does not depend on the size of the sequence space being explored.
[0364] ISDE enables prediction of optimal enzymes with greater computational efficiency, enabling rapid sorting by customizable fitness metrics (Figs. 32b-d and described above). First, a starting sequence was computationally mutated to generate a small sub-library (1000-100,000 sequences) bearing random amino acid substitutions with a defined hamming distance from the original sequence. PAM predictions were then generated for each member of the sub-library. A customizable fitness function was used to score each variant according to the desired properties to be evolved. The top few sequences with best scores were then selected as the starting variants for subsequent rounds of evolution. The process was continued until no further improvement in fitness is achieved.
[0365] Example 7. PAMmla-based engineering of an allele-specific editor for RHO P23H
[0366] We explored our in silico directed evolution approach to evolve a customized editor to target the human Rhodopsin (RHO) P23H mutation causative of retinitis pigmentosa58-60. As a dominant negative mutation, knock-out of the mutant allele while leaving the WT allele intact can rescue function in heterozygous genotypes61,62. Single nucleotide variants can be difficult to target in an allele-specific manner due to the propensity of SpCas9 to tolerate single base pair mismatches between the gRNA and target DNA57,63. However, if a disease-causing point mutation generates a novel PAM, SpCas9’s higher-specificity PAM proof-reading step can be exploited for allele-specific editing57,64,65. For RHO P23H, it is possible to design a gRNA such that the mutant allele encoding the point mutation harbors an NGTG PAM, while the WT allele encodes an NGGG PAM (Fig. 6b). In principle, neither WT SpCas9 nor SpG would be useful for this allele-specific editing strategy, given that SpCas9 should target only the WT allele and SpG should target both alleles equally.
[0367] To enable allele-specific deletion of the RHO P23H allele, we utilized PAMmla to evolve a custom enzyme to maximize the rate constant for editing an NGTG PAMs while requiring the rate constant on NGGG PAMs remain below a threshold value of IO'3 7. Using PAMmla-ISDE we utilized WT SpCas9 (DSGERT) as a substrate to evolve custom enzymes. Input parameters maximized activity on NGTG PAMs while minimizing on NGGG, mutating up to four amino acid positions per derivative enzyme, maximizing enzyme ks for NGTG, and requiring a k < 10-3.7 on NGGG. The PAMmla-ISDE fitness plateaued after four cycles of mutagenesis and prediction (Figs. 6b, d and Figs. 33 and 39A-B). The resulting enzyme predicted to have highest fitness according to our chosen metrics (maximizing NGTG while minimizing NGGG) was MRRWMR, which was predicted by PAMmla to have strong activity on NGTG and weak activity on NGGG ks of = 10'2 1and IO'3-7, respectively; Fig. 6b). We also performed a stricter evolution via IDSE to further minimize targeting of NGGG, requiring a k < 10'4on NGGG, which resulted in several enzymes including KRHWMR after 4 rounds of evolution (Figs. 6c, e and Fig. 39b). HT-PAMDA confirmed the PAM preference of MRRWMR (Figs. 6e, f). MRRWMR was also the most optimal enzyme when initiating our in silico directed evolution campaign from the SpG amino acid sequence (LWKQQR) by minimizing rate constant on NGGG while requiring that the rate constant on NGTG be greater than 10'2. Allowing only two mutations per enzyme in each round of evolution, and selecting the top 4 enzymes per round, we obtained MRRWMR in 5 evolution cycles (Fig. 33). This finding demonstrates the ability of the evolution strategy to converge to the same solution despite differing starting sequences and evolution parameters. In human cell experiments, the PAMmla-derived MRRWMR enzyme achieved comparable levels of on-target editing to SpG across three unrelated genomic sites harboring NGTG PAMs. Importantly, MRRWMR exhibited undetectable editing against a control NGGG site where SpG led to efficient editing (Fig. 6g), supporting that this enzyme should be useful for distinguishing the two RHO alleles. We then performed an analysis of MRRWMR genome-wide specificity via GUTDE-seq2. In a comparison to SpG and SpRY, MRRWMR preserved on-target editing in GUIDE- seq2 while substantially reducing off-target editing compared to the relaxed PAM enzymes (Fig. 6h). These results demonstrate that MRRWMR can achieve potent on- target editing while also minimizing off-target editing.
[0368] To determine whether the PAMmla-predicted MRRWMR enzyme was capable of allele-specific editing of RHO P23H, we generated a heterozygous HEK 293 T cell line bearing this mutation (HEK 293Ts are triploid for this locus, leading to a cell line with two copies of P23H and one copy of WT). We tested a total of four ISDE- derived enzymes from either campaign in the RHO P23H cell line and observed a preference for editing the mutant over the WT allele compared to SpG for all ISDE enzymes (Fig. 34b). The MRRWMR enzyme led to the most efficient mutant allele disruption and KRHWMR resulted in superior allele-specific discrimination with nearly undetectable editing of the WT allele in heterozygous and WT cells (Fig. 6f and Figs. 34c, d). The ISDE-derived MRRWMR and KRHWMR enzymes exhibited ~2.5- and ~40-fold preferences for mutant P23H over WT RHO alleles respectively (Fig. 34e).Use of other recently described Cas9 ortholog nucleases PrCas9, CoCas9, and GeCas9 harboring PAM preferences that should target the P23H site failed to elicit detectable editing at the RHO locus.
[0369] Beyond the use of PAMmla-designed enzymes, another potential option for achieving allele-selective editing of RHO P23H is to exploring other naturally occurring Cas9 orthologs with divergent PAM sequences. We compared MRRWMR and KRHWMR to previously reported PrCas9, CoCas9, and GeCas9 orthologs which can recognize either N5T or N4GT nucleotide PAMs and could theoretically discriminate WT and P23H RHO alleles. In experiments using a homozygous mutant RHO P23H HEK 293T cell line (harboring three P23H alleles), we observed high levels of P23H editing with SpG and the PAMmla-predicted MRRWMR and KRHWMR enzymes. However, we were unable to detect nuclease-mediated genome editing with any of the three orthologs when combined with their corresponding PAM-selective RHO P23H sgRNAs (each bearing 3 different lengths of sgRNA spacers). We did, however, observe genome editing with CoCas9 paired with two sgRNAs targeting other previously reported positive control genomic sites unrelated to RHO P23H26(max 23%, mean 14%). No positive control genomic target sites were previously reported for PrCas9 or GeCas9, since experiments were performed in vitro or to target plasmids in human cells25,26, and our experiments failed to detect editing with either PrCas9 or GeCas9. As has been shown previously for other Cas9 orthologs, challenges for GeCas9 and PrCas9 might include weaker DNA binding27, a reduced ability to target chromatinized genomic DNA28, or a requirement for higher magnesium concentrations than what is available in human cells compared to bacteria29.
[0370] Next, we performed off-target analyses via GUIDE-seq2 to investigate specificity improvements with the allele-selective PAMmla enzymes. On two unrelated genomic loci we observed enhanced specificity with MRRWMR over SpG and SpRY (Figs. 35a-c). When using the RHO P23H sgRNAs in a homozygous P23H cell line, both MRRWMR and KRHWMR also improved on-target specificity by minimizing off- target reads and the number of detected off-target sites compared to SpG and SpRY (Fig. 6g and Figs. 35d-f).
[0371] To examine potential structural determinants of the PAM-selective preferences of MRRWMR and KRHWMR, we modeled their mutations and investigated their contributions to PAMmla predicted rates on NGTG and NGGG PAMs using SHAP (Figs. 6h-k, Figs. 36a-e). For both enzymes, PAM selectivity for a 3rdPAM position T over G largely resulted from the E1219W and R1335M side chains forming a hydrophobic pocket, enabling discrimination via favorable hydrophobic interactions with the T3 methyl group of an NGTG PAM and unfavorable interaction with the polar G3 side chain of NGGG (Figs. 6h,i). SHAP analysis supports that these mutations positively impact recognition of NGTG and negatively impact activity on NGGG (Figs. 6j,k and Figs. 36b-e). Additional common mutations between the two PAMmla enzymes SI 136R and T1337R contribute to PAM selectivity and / or potentiation of activity (Figs. 36a-e).
[0372] We assessed the in vivo editing efficiency and allele-specificity of MRRWMR and KRHWMR in mouse retinas. Subretinal plasmid injections in humanized heterozygous WT-hRHO-GFP / P23H-hRHO-RFP mice at P0-P2 were performed prior to in vivo electroporation. Analysis of on-target editing in transfected cells with MRRWMR revealed mean on -target efficiency of 37% on the P23H allele (up to 59%) and 7.6% on the WT allele, leading to 4.8-fold selectivity for editing the mutant allele (Fig. 61 and Figs. 35g, h). Consistent with results in cells, KRHWMR resulted in lower in vivo editing (-20%) but exhibited increased specificity for the P23H over wild type allele (9.5-fold) (Fig. 61 and Fig. 35h). Together, these findings demonstrate how PAMmla-ISDE can predict novel enzymes for therapeutically relevant edits with no intervening engineering or evolution. PAMmla-nominated enzymes enable cellbased and in vivo allele-selective editing not possible with previously available SpCas9 enzyme variants or Cas9 orthologs.
[0373] Example 8. PAMmla-based prediction and validation of a multiplex base editor for the APOE-E4 Alzheimer’s risk allele
[0374] Another metric that we envisioned would be useful to optimize is simultaneous activity on two or more PAMs for multiplex editing. As a proof of concept, we aimed to develop an editor to simultaneously introduce the two C to T point mutations which revert the A POE-sA (Argl30, Argl76) Alzheimer’s risk allele to the APOE-P2 (Cysl30, Cysl76) Alzheimer’s protective allele66,67. We designed two gRNAs positioning each target cytosine at position 5 in the protospacer, with the Argl30 guide utilizing an NGTA PAM and the Argl76 guide utilizing an NGGC PAM. We performed in silico directed evolution starting from the sequence of wild type SpCas9 and maximizing activity on NGTA with a simultaneous constraint of maintaining activity on NGGC with rate constant of at least 10'2 5(Fig. 37a, b). We tested one of the top 10 resulting variants, MRKCRS, as well as MRKQKC and MRKSKC from other selection criteria using PAMmla (see figure 4) at HEK293T loci harboring NGTA and NGGC PAMs, in the context of different cytosine base editors including BE4max68, evoFERNY69, and TadCBEd55(Fig. 38A-B). We observed efficient editing for all variants, with similar editing as SpG on the NGTA site and increased editing compared to SpG on the NGGC site, suggesting that these variants might be effective at editing the APOE- s4 allele. We next tested the ability of the PAMmla derived variants to simultaneously a HEK 293 T cell line bearing the APOE- s4 allele at the endogenous locus introduced by prime editing. A single PAM variant CBE was transfected together with both gRNAs to induce multiplex editing (Fig. 37c). Editing at the APOE locus was less efficient than previously tested HEK loci, however, all variants induced a greater rate of multiplexed base editing than SpG at the desired cytosines, with MRKCRS editing with the greatest efficiency.
[0375] References
[0376] 1. Packer, M. S. & Liu, D. R. Methods for the directed evolution of proteins. Nature Reviews Genetics 2015 16:7 16, 379-394 (2015).
[0377] 2. Romero, P. A. & Arnold, F. H. Exploring protein fitness landscapes by directed evolution. Nature Reviews Molecular Cell Biology 2009 10: 12 10, 866-876 (2009).
[0378] 3. Yang, K. K., Wu, Z. & Arnold, F. H. Machine-learning-guided directed evolution for protein engineering. Nature Methods Preprint at doi.org / 10.1038 / s41592-019-0496-6 (2019).
[0379] 4. Liu, G., Lin, Q., Jin, S. & Gao, C. The CRISPR-Cas toolbox and gene editing technologies. Mol Cell 82, 333-347 (2022).
[0380] 5. Mojica, F. J. M., Diez-Villasenor, C., Garcia-Martinez, J. & Almendros, C. Short motif sequences determine the targets of the prokaryotic CRISPR defence system. Microbiology (N Y) 155, 733-740 (2009).
[0381] 6. Anders, C., Niewoehner, O., Duerst, A. & Jinek, M. Structural basis of PAM- dependent target DNA recognition by the Cas9 endonuclease. Nature (2014) doi: 10.1038 / naturel3579.
[0382] 7. Sternberg, S. H., Redding, S., Jinek, M., Greene, E. C. & Doudna, J. A. DNA interrogation by the CRISPR RNA-guided endonuclease Cas9. Nature (2014) doi: 10.1038 / naturel3011.
[0383] 8. Jinek, M. et al. A programmable dual -RNA-guided DNA endonuclease in adaptive bacterial immunity. Science (1979) (2012) doi: 10.1126 / science.1225829.
[0384] 9. Deltcheva, E. et al. CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III. Nature 2011 471 :7340 471, 602-607 (2011). 10. Kleinstiver, B. P. et al. Engineered CRISPR-Cas9 nucleases with altered PAM specificities. Nature (2015) doi: 10.1038 / naturel4592.
[0385] 11. Walton, R. T., Christie, K. A., Whittaker, M. N. & Kleinstiver, B. P. Unconstrained genome targeting with near-PAMless engineered CRISPR-Cas9 variants. Science (1979) (2020) doi: 10.1126 / science.aba8853.
[0386] 12. Hu, J. H. et al. Evolved Cas9 variants with broad PAM compatibility and high DNA specificity. Nature (2018) doi: 10.1038 / nature26155.
[0387] 13. Miller, S. M. et al. Continuous evolution of SpCas9 variants compatible with non-G PAMs. Nat Biotechnol (2020) doi: 10.1038 / s41587-020-0412-8.
[0388] 14. Nishimasu, H. et al. Engineered CRISPR-Cas9 nuclease with expanded targeting space. Science (1979) (2018) doi: 10.1126 / science.aas9129.
[0389] 15. Zhong, Z. et al. Improving Plant Genome Editing with High-Fidelity xCas9 and Non-canonical PAM-Targeting Cas9-NG. Mol Plant 12, 1027-1036 (2019).
[0390] 16. Zhang, W. et al. In-depth assessment of the PAM compatibility and editing activities of Cas9 variants. Nucleic Acids Res 49, 8785-8795 (2021).
[0391] 17. Hibshman, G. N. et al. Unraveling the mechanisms of PAMless DNA interrogation by SpRY Cas9. bioRxiv 2023.06.22.546082 (2023) doi:10.1101 / 2023.06.22.546082.
[0392] 18. Wu, Y. et al. Genome-wide analyses of PAM-relaxed Cas9 genome editors reveal substantial off-target effects by ABE8e in rice. Plant Biotechnol J 20, 1670- 1682 (2022).
[0393] 19. Kleinstiver, B. P. et al. Engineered CRISPR-Casl2a variants with increased activities and improved targeting ranges for gene, epigenetic and base editing. Nature Biotechnology 2019 37:3 37, 276-282 (2019).
[0394] 20. Chatterjee, P. et al. A Cas9 with PAM recognition for adenine dinucleotides. Nature Communications 2020 11 : 1 11, 1-6 (2020).
[0395] 21. Zhao, L. et al. PAM-flexible genome editing with an engineered chimeric Cas9. Nature Communications 2023 14: 1 14, 1-8 (2023).
[0396] 22. Chatterjee, P. et al. An engineered ScCas9 with broad PAM range and high specificity and activity. Nature Biotechnology 2020 38: 10 38, 1154-1158 (2020).
[0397] 23. Kleinstiver, B. P. et al. Broadening the targeting range of Staphylococcus aureus CRISPR-Cas9 by modifying PAM recognition. Nature Biotechnology 2015 33: 12 33, 1293-1298 (2015). 24. Goldberg, G. W. et al. Engineered dual selection for directed evolution of SpCas9 PAM specificity. Nature Communications 2021 12: 1 12, 1-16 (2021).
[0398] 25. Huang, T. P. et al. High-throughput continuous evolution of compact Cas9 variants targeting single-nucleotide-pyrimidine PAMs. Nature Biotechnology 2022 41 : 1 41, 96-107 (2022).
[0399] 26. Wu, Z., Johnston, K. E., Arnold, F. H. & Yang, K. K. Protein sequence design with deep generative models. Current Opinion in Chemical Biology Preprint at doi.org / 10.1016 / j.cbpa.2021.04.004 (2021).
[0400] 27. Frappier, V. & Keating, A. E. Data-driven computational protein design. Current Opinion in Structural Biology Preprint at doi.org / 10.1016 / j.sbi.2021.03.009 (2021).
[0401] 28. Wittmann, B. J., Johnston, K. E., Wu, Z. & Arnold, F. H. Advances in machine learning for directed evolution. Current Opinion in Structural Biology Preprint at doi.org / 10.1016 / j.sbi.2021.01.008 (2021).
[0402] 29. Hie, B. L. & Yang, K. K. Adaptive machine learning for protein engineering. Current Opinion in Structural Biology Preprint at doi.org / 10.1016 / j.sbi.2021.11.002 (2022).
[0403] 30. Makowski, E. K., Chen, H.-T. & Tessier, P. M. Simplifying complex antibody engineering using machine learning. Cell Syst 14, 667-675 (2023).
[0404] 31. Hie, B. L. et al. Efficient evolution of human antibodies from general protein language models. Nature Biotechnology 2023 1-9 (2023) doi: 10.1038 / S41587-023- 01763-2.
[0405] 32. Hie, B. L. et al. Efficient evolution of human antibodies from general protein language models. Nature Biotechnology 2023 1-9 (2023) doi: 10.1038 / S41587-023- 01763-2.
[0406] 33. Saka, K. et al. Antibody design using LSTM based deep generative model from phage display library for affinity maturation. Scientific Reports 2021 11 :1 11, 1- 13 (2021).
[0407] 34. Mason, D. M. et al. Optimization of therapeutic antibodies by predicting antigen specificity from antibody sequence via deep learning. Nature Biomedical Engineering 2021 5:6 5, 600-612 (2021).
[0408] 35. Gupta, A. et al. An improved predictive recognition model for Cys2-His2 zinc finger proteins. Nucleic Acids Res 42, 4800 (2014). 36. Aizenshtein-Gazit, S. & Orenstein, Y. DeepZF: improved DNA-binding prediction of C2H2-zinc-finger proteins by deep transfer learning. Bioinformatics 38, II62-II67 (2022).
[0409] 37. Ichikawa, D. M. et al. A universal deep-learning model for zinc finger design enables transcription factor reprogramming. Nature Biotechnology 2023 41 :8 41, 1117-1129 (2023).
[0410] 38. Bryant, D. H. et al. Deep diversification of an AAV capsid protein by machine learning. Nature Biotechnology 2021 39:6 39, 691-696 (2021).
[0411] 39. Ogden, P. J., Kelsic, E. D., Sinai, S. & Church, G. M. Comprehensive AAV capsid fitness landscape reveals a viral gene and enables machine-guided design. Science (1979) 366, 1139-1143 (2019).
[0412] 40. Eid, F.-E. et al. Systematic multi-trait AAV capsid engineering for efficient gene delivery. bioRxiv 2022.12.22.521680 (2022) doi: 10.1101 / 2022.12.22.521680.
[0413] 41. Walton, R. T., Hsu, J. Y., Joung, J. K. & Kleinstiver, B. P. Scalable characterization of the PAM requirements of CRISPR-Cas enzymes using HT- PAMDA. Nat Protoc (2021) doi: 10.1038 / s41596-020-00465-2.
[0414] 42. Anders, C., Niewoehner, O., Duerst, A. & Jinek, M. Structural basis of PAM- dependent target DNA recognition by the Cas9 endonuclease. Nature (2014) doi: 10.1038 / naturel3579.
[0415] 43. Kleinstiver, B. P. et al. Engineered CRISPR-Cas9 nucleases with altered PAM specificities. Nature (2015) doi: 10.1038 / naturel4592.
[0416] 44. Walton, R. T., Christie, K. A., Whittaker, M. N. & Kleinstiver, B. P. Unconstrained genome targeting with near-PAMless engineered CRISPR-Cas9 variants. Science 368(6488):290-296. (2020) doi: 10.1126 / science.aba8853.
[0417] 45. Nishimasu, H. et al. Engineered CRISPR-Cas9 nuclease with expanded targeting space. Science 361(6408): 1259-1262. (2018) doi: 10.1126 / science.aas9129.
[0418] 46. Christie, K. A. et al. Precise DNA cleavage using CRISPR-SpRYgests. Nature Biotechnology 2022 41 :3 41, 409-416 (2022).
[0419] 47. Chen, Z. & Zhao, H. A highly sensitive selection method for directed evolution of homing endonucleases. Nucleic Acids Res 33, el54-el54 (2005).
[0420] 48. Wittmann, B. J., Johnston, K. E., Almhjell, P. J. & Arnold, F. H. evSeq: Cost- Effective Amplicon Sequencing of Every Variant in a Protein Library. bioRxiv 2021.11.18.469179 (2021) doi: 10.1101 / 2021.11.18.469179. 49. Quick, J. et al. Multiplex PCR method for MinlON and Illumina sequencing of Zika and other virus genomes directly from clinical samples. Nat Protoc (2017) doi: 10.1038 / nprot.2017.066.
[0421] 50. Walton, R. T., Hsu, J. Y., Joung, J. K. & Kleinstiver, B. P. Scalable characterization of the PAM requirements of CRISPR-Cas enzymes using HT- PAMDA. Nat Protoc (2021) doi: 10.1038 / s41596-020-00465-2.
[0422] 51. Georgiev, A. G. Interpretable numerical descriptors of amino acid space. Journal of Computational Biology 16, 703-723 (2009).
[0423] 52. Wittmann, B. J., Yue, Y. & Arnold, F. H. Informed training set design enables efficient machine learning-assisted directed protein evolution. Cell Syst (2021) doi: 10.1016 / j.cels.2021.07.008.
[0424] 53. Rees, H. A. & Liu, D. R. Base editing: precision chemistry on the genome and transcriptome of living cells. Nature Reviews Genetics 2018 19: 12 19, 770-788 (2018).
[0425] 54. Richter, M. F. et al. Phage-assisted evolution of an adenine base editor with improved Cas domain compatibility and activity. Nat Biotechnol 38, 883-891 (2020).
[0426] 55. Neugebauer, M. E. et al. Evolution of an adenine base editor into a small, efficient cytosine base editor with low off-target activity. Nature Biotechnology 2022 41 :5 41, 673-685 (2022).
[0427] 56. Newby, G. A. et al. Base editing of haematopoietic stem cells rescues sickle cell disease in mice. Nature 2021 595:7866 595, 295-302 (2021).
[0428] 57. Christie, K. A. et al. Towards personalised allele-specific CRISPR gene editing to treat autosomal dominant disorders. Sci Rep (2017) doi: 10.1038 / s41598- 017-16279-4.
[0429] 58. Sung, C. H. et al. Rhodopsin mutations in autosomal dominant retinitis pigmentosa. Proceedings of the National Academy of Sciences 88, 6481-6485 (1991).
[0430] 59. Dryja, T. P. et al. A point mutation of the rhodopsin gene in one form of retinitis pigmentosa. Nature 1990 343:6256 343, 364-366 (1990).
[0431] 60. Hartong, D. T., Berson, E. L. & Dryja, T. P. Retinitis pigmentosa. The Lancet 368, 1795-1809 (2006).
[0432] 61. LaVail, M. M. et al. Ribozyme rescue of photoreceptor cells in P23H transgenic rats: Long-term survival and late-stage therapy. Proceedings of the National Academy of Sciences 97, 11488-11493 (2000). 62. Li, P. et al. Allele-Specific CRISPR-Cas9 Genome Editing of the Single-Base P23H Mutation for Rhodopsin-Associated Dominant Retinitis Pigmentosa, home- liebertpub-com.ezp-prodl.hul.harvard.edu / crispr 1, 55-64 (2018).
[0433] 63. Hsu, P. D. et al. DNA targeting specificity of RNA-guided Cas9 nucleases. Nature Biotechnology 2013 31 :9 31, 827-832 (2013).
[0434] 64. Shin, J. W. et al. Permanent inactivation of Huntington’s disease mutation by personalized allele-specific CRISPR / Cas9. Hum Mol Genet 25, 4566-4576 (2016).
[0435] 65. Courtney, D. G. et al. CRISPR / Cas9 DNA cleavage at SNP-derived PAM enables both in vitro and in vivo KRT12 mutation-specific targeting. Gene Therapy 2016 23: 1 23, 108-112 (2015).
[0436] 66. Farrer, L. A. et al. Effects of Age, Sex, and Ethnicity on the Association Between Apolipoprotein E Genotype and Alzheimer Disease: A Meta-analysis. JAMA 278, 1349-1356 (1997).
[0437] 67. Li, Z., Shue, F., Zhao, N., Shinohara, M. & Bu, G. APOE2: protective mechanism and therapeutic implications for Alzheimer’s disease. Mol Neurodegener 15, 1-19 (2020).
[0438] 68. Koblan, L. W. et al. Improving cytidine and adenine base editors by expression optimization and ancestral reconstruction. Nature Biotechnology 2018 36:9 36, 843-846 (2018).
[0439] 69. Thuronyi, B. W. et al. Continuous evolution of base editors with expanded target compatibility and improved activity. Nat Biotechnol 37, 1070 (2019).
[0440] 70. Kleinstiver, B. P., Fernandes, A. D., Gloor, G. B. & Edgell, D. R. A unified genetic, computational and experimental framework identifies functionally relevant residues of the homing endonuclease I-Bmol. Nucleic Acids Res 38, 2411 (2010).
[0441] 71. Gibson, D. G. et al. Enzymatic assembly of DNA molecules up to several hundred kilobases. Nature Methods 2009 6:5 6, 343-345 (2009).
[0442] 72. Alves, C. R. R. et al. Optimization of base editors for the functional correction of SMN2 as a treatment for spinal muscular atrophy. Nature Biomedical Engineering 2023 1-14 (2023) doi: 10.1038 / S41551-023-01132-Z.
[0443] 73. Nelson, J. W. et al. Engineered pegRNAs improve prime editing efficiency. Nat Biotechnol 40, 402 (2022).
[0444] 74. Chen, P. J. et al. Enhanced prime editing systems by manipulating cellular determinants of editing outcomes. Cell 184, 5635-5652. e29 (2021). 75. Kleinstiver, B. P., Fernandes, A. D., Gloor, G. B. & Edgell, D. R. A unified genetic, computational and experimental framework identifies functionally relevant residues of the homing endonuclease I-Bmol. Nucleic Acids Res (2010) doi: 10.1093 / nar / gkpl223.
[0445] 76. Rohland, N. & Reich, D. Cost-effective, high-throughput DNA sequencing libraries for multiplexed target capture. Genome Res 22, 939 (2012).
[0446] 77. Quick, J. et al. Multiplex PCR method for MinlON and Illumina sequencing of Zika and other virus genomes directly from clinical samples. Nat Protoc (2017) doi: 10.1038 / nprot.2017.066.
[0447] 78. Tareen, A. & Kinney, J. B. Logomaker: Beautiful sequence logos in Python. Bioinformatics (2020) doi: 10.1093 / bioinformatics / btz921.
[0448] 79. Ofer, D. & Linial, M. ProFET: Feature engineering captures high-level protein functions. Bioinformatics (2015) doi: 10.1093 / bioinformatics / btv345.
[0449] 80. Chollet, F. Keras, GitHub. Preprint at github.com / fchollet / keras (2015).
[0450] 81. Song, D., Xi, N. M., Li, J. J. & Wang, L. scSampler: fast diversity-preserving subsampling of large-scale single-cell transcriptomic data. Bioinformatics 38, 3126- 3127 (2022).
[0451] 82. Mclnnes, L., Healy, J. & Melville, J. UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction. (2018).
[0452] 83. Anzalone, A. V. et al. Search-and-replace genome editing without doublestrand breaks or donor DNA. Nature 2019 576:7785 576, 149-157 (2019).
[0453] 84. Clement, K. et al. CRISPResso2 provides accurate and rapid genome editing sequence analysis. Nature Biotechnology 2019 37:3 37, 224-226 (2019).
[0454] 85. Tsai, S. Q. et al. GUIDE-seq enables genome-wide profiling of off-target cleavage by CRISPR-Cas nucleases. Nature Biotechnology 2014 33:2 33, 187-197 (2014).
[0455] 86. Picelli, S. et al. Tn5 transposase and tagmentation procedures for massively scaled sequencing projects. Genome Res 24, 2033-2040 (2014).
[0456] 87. Tsai, S. Q., Topkar, V. V., Joung, J. K. & Aryee, M. J. Open-source guideseq software for analysis of GUIDE-seq data. Nature Biotechnology 2016 34:5 34, 483- 483 (2016).
[0457] 88. Kleinstiver, B. P. et al. High-fidelity CRISPR-Cas9 variants with undetectable genome-wide off-targets. Nature 529, 490 (2016). 89. Landrum, M. J. et al. ClinVar: Public archive of relationships among sequence variation and human phenotype. Nucleic Acids Res (2014) doi: 10.1093 / nar / gktl 113.
[0458] 90. Georgiev, A. G. Interpretable numerical descriptors of amino acid space. Journal of Computational Biology 16, 703-723 (2009).
[0459] 91. Gaudelli, N. M. et al. Programmable base editing of T to G C in genomic DNA without DNA cleavage. Nature (2017) doi: 10.1038 / nature24644.
[0460] 92. Emsley, P., Lohkamp, B., Scott, W. G. & Cowtan, K. Features and development of Coot. Acta Crystallogr D Biol Crystallogr 66, 486 (2010).
[0461] 93. Goddard, T. D. et al. UCSF ChimeraX: Meeting modem challenges in visualization and analysis. Protein Science 27, 14-25 (2018).
[0462] 94. Robichaux, M. A. et al. Subcellular localization of mutant P23H rhodopsin in an RFP fusion knock-in mouse model of retinitis pigmentosa. DMM Disease Models and Mechanisms 15, (2022).
[0463] 95. Chan, F., Bradley, A., Wensel, T. G. & Wilson, J. H. Knock-in human rhodopsin-GFP fusions as mouse models for human disease and targets for gene therapy. Proc Natl Acad Sci U S A 101, 9109-9114 (2004).
[0464] 96. Lundberg, S. M., Allen, P. G. & Lee, S.-I. A Unified Approach to Interpreting Model Predictions. doi:10.5555 / 3295222.3295230.
[0465] 97. Chen, H., Lundberg, S. M. & Lee, S. I. Explaining a series of models by propagating Shapley values. Nature Communications 2022 13: 1 13, 1-15 (2022).
[0466] 98. Hopf, T. A. et al. The EVcouplings Python framework for coevolutionary sequence analysis. Bioinformatics 35, 1582-1584 (2019).
[0467] 99. Frazer, J. et al. Disease variant prediction with deep generative models of evolutionary data. Nature 2021 599:7883 599, 91-95 (2021).
[0468] 100. Notin, P. et al. TranceptEVE: Combining Family-specific and Family-agnostic Models of Protein Sequences for Improved Fitness Prediction. doi: 10.1101 / 2022.12.07.519495.
[0469] 101. Anders, C., Bargsten, K. & Jinek, M. Structural Plasticity of PAM Recognition by Engineered Variants of the RNA-Guided Endonuclease Cas9. Mol Cell (2016) doi: 10.1016 / j.molcel.2016.02.020.
[0470] 102. Hirano, S., Nishimasu, H., Ishitani, R. & Nureki, O. Structural Basis for the Altered PAM Specificities of Engineered CRISPR-Cas9. Mol Cell (2016) doi: 10.1016 / j.molcel.2016.02.018. OTHER EMBODIMENTS
[0471] It is to be understood that while the invention has been described in conjunction with the detailed description thereof, the foregoing description is intended to illustrate and not limit the scope of the invention, which is defined by the scope of the appended claims. Other aspects, advantages, and modifications are within the scope of the following claims.
Claims
WHAT IS CLAIMED IS:
1. An isolated Streptococcus pyogenes Cas9 (SpCas9) protein, comprising mutations at two, three, four, five, or all six of the following positions: D1135, S 1136, G1218, E1219, R1335, and / or T1337, wherein the mutations are listed in any of Tables F, E, C A, B, or D.
2. The isolated protein of claim 1, comprising a sequence that is at least 80% identical to the amino acid sequence of SEQ ID NO: 1.
3. The isolated protein of claim 1, further comprising one or more mutations that decrease nuclease activity selected from the group consisting of mutations at DIO, E762, D839, H983, or D986; and at H840 or N863.
4. The isolated protein of claim 3, wherein the mutations are:(i) D10A or DION, and / or(ii) H840A, H840N, or H840Y5. The isolated protein of claim 1, further comprising one or more mutations that increase specificity selected from the group consisting of mutations at N497, R661, N692, M694, Q695, H698, K810, K848, Q926, K1003, and R0160, and optionally at K526 and / or R691.
6. The isolated protein of claim 5, comprising mutations N692A, Q695A, Q926A, H698A, N497A, R661A, M694A, K810A, K848A, K1003A, R0160A, Y450A / Q695A, L169A / Q695A, Q695A / Q926A, Q695A / D1135E, Q926A / D1135E, Y450A / D1135E, L169A / Y450A / Q695A, L169A / Q695A / Q926A, Y450A / Q695A / Q926A, R661A / Q695A / Q926A, N497A / Q695A / Q926A, Y450A / Q695A / D1135E, Y450A / Q926A / D1135E, Q695 A / Q926A / D 1135E, L 169A / Y450A / Q695 A / Q926 A, L169A / R661A / Q695A / Q926A,Y450A / R661A / Q695A / Q926A, N497A / Q695A / Q926A / D1135E, R661A / Q695A / Q926A / D1135E, and Y450A / Q695 A / Q926A / D 1135E; N692A / M694 A / Q695 A / H698 A, N692A / M694A / Q695A / H698A / Q926A; N692A / M694A / Q695A / Q926A; N692A / M694A / H698A / Q926A; N692A / Q695A / H698A / Q926A;M694A / Q695A / H698A / Q926A; N692A / Q695A / H698A; N692A / M694A / Q695A;N692A / H698A / Q926A; N692A / M694A / Q926A; N692A / M694A / H698A;M694A / Q695A / H698A; M694A / Q695A / Q926A; Q695A / H698A / Q926A;G582 A / V583 A / E584 A / D 585 A / N588 A / Q926 A;G582A / V583A / E584A / D585A / N588A; T657A / G658A / W659A / R661A / Q926A;T657A / G658 A / W659A / R661 A; F491 A / M495 A / T496A / N497A / Q926 A;F491A / M495A / T496A / N497A; K918A / V922A / R925A / Q926A; or 918A / V922A / R925A; K855A; K810A / K1003A / R1060A; or K848A / K1003A / R1060A, and optionally K526A and / or R691 A.
7. A fusion protein comprising the isolated protein of claim 1, fused to a heterologous functional domain, with an optional intervening linker, wherein the linker does not interfere with activity of the fusion protein.
8. The fusion protein of claim 7, wherein the heterologous functional domain is a transcriptional activation domain.
9. The fusion protein of claim 8, wherein the transcriptional activation domain is from VP16, VP64, rTA, NF-KB p65, or the composite VPR (VP64-p65-rTA).
10. The fusion protein of claim 7, wherein the heterologous functional domain is a transcriptional silencer or transcriptional repression domain.
11. The fusion protein of claim 10, wherein the transcriptional repression domain is a Krueppel-associated box (KRAB) domain, ERF repressor domain (ERD), or mSin3 A interaction domain (SID).
12. The fusion protein of claim 10, wherein the transcriptional silencer is Heterochromatin Protein 1 (HP1).
13. The fusion protein of claim 10, wherein the heterologous functional domain is an enzyme that modifies the methylation state of DNA.
14. The fusion protein of claim 10, wherein the enzyme that modifies the methylation state of DNA is a DNA methyltransferase (DNMT) or a TET protein.
15. The fusion protein of claim 14, wherein the TET protein is TET1.
16. The fusion protein of claim 7, wherein the heterologous functional domain is an enzyme that modifies a histone subunit.
17. The fusion protein of claim 16, wherein the enzyme that modifies a histone subunit is a histone acetyltransferase (HAT), histone deacetylase (HD AC), histone methyltransferase (HMT), or histone demethylase.
18. The fusion protein of claim 7, wherein the heterologous functional domain is a base editor.
19. The fusion protein of claim 18, wherein the base editor is (i) a cytidine deaminase domain, preferably selected from the group consisting of the apolipoprotein B mRNA-editing enzyme, catalytic polypeptide-like (APOBEC) family of deaminases, optionally APOBEC 1, AP0BEC2, AP0BEC3A, AP0BEC3B, AP0BEC3C, AP0BEC3D / E, APOBEC3F, AP0BEC3G, AP0BEC3H, or AP0BEC4; activation-induced cytidine deaminase (AID), optionally activation induced cytidine deaminase (AICDA); cytosine deaminase 1 (CDA1) or CDA2; cytosine deaminase acting on tRNA (CD AT), or an engineered variant thereof optionally DddA-like cytidine deaminase or engineered TadA -based cytosine base editors, or (ii) an adenosine deaminase, preferably selected from the group consisting of adenosine deaminase 1 (ADA1), ADA2; adenosine deaminase acting on RNA 1 (AD ARI), ADAR2, ADAR3; adenosine deaminase acting on tRNA 1 (ADAT1), ADAT2, ADAT3; and naturally occurring or engineered tRNA-specific adenosine deaminase (TadA).
20. The fusion protein of claim 7, wherein the heterologous functional domain is a biological tether.
21. The fusion protein of claim 20, wherein the biological tether is MS2, Csy4 or lambda N protein.
22. The fusion protein of claim 7, wherein the heterologous functional domain is Fokl.
23. An isolated nucleic acid encoding the protein of claims 1-22.
24. A vector comprising the isolated nucleic acid of claim 23.
25. A vector comprising the isolated nucleic acid of claim 23, which is operably linked to one or more regulatory domains for expressing the isolated Streptococcus pyogenes Cas9 (SpCas9) protein, with mutations at one, two, three, four, five, or all six of the following positions: D1135, S 1136, G1218, E1219, R1335, and / or T1337, wherein the mutations are listed in Table A, and optionally a nucleic acid encoding a guide RNA that complexes with the cas9 protein.
26. A host cell, preferably a mammalian host cell, comprising the nucleic acid of claim 23.
27. A method of altering the genome of a cell, the method comprising expressing in the cell, or contacting the cell with, the isolated protein or fusion protein of claim 1, and a guide RNA having a region complementary to a selected portion of the genome of the cell.
28. The method of claim 27, wherein the isolated protein or fusion protein comprises one or more of a nuclear localization sequence, cell penetrating peptide sequence, and / or affinity tag.
29. The method of claim 28, wherein the cell is a stem cell.
30. The method of claim 29, wherein the cell is an embryonic stem cell, mesenchymal stem cell, or induced pluripotent stem cell; is in a living animal; or is in an embryo.
31. A method of altering a double stranded DNA (dsDNA) molecule, the method comprising contacting the dsDNA molecule with the isolated protein or fusion protein of claim 1, and a guide RNA that complexes with the cas9 protein, the guide RNA having a region complementary to a selected portion of the dsDNA molecule.
32. The method of claim 31, wherein the dsDNA molecule is in vitro.
33. The method of claim 31, wherein the fusion protein and RNA are in a ribonucleoprotein complex.
34. A method for generating enzymes of desired functionality, the method comprising: identifying active SpCas9 enzymes through structure-and-function-informed saturation mutagenesis and subsequent bacterial selection; profiling the active SpCas9 enzymes according to all possible protospacer- adjacent motif (PAM) requirements to amino acid sequences using a high- throughput PAM-determination assay (HT-PAMDA) to generate a first training data set for training a PAM machine-learning (ML) model to relate PAM requirements to amino acid sequence; generating a second training data set comprising random non-selected enzymes from a SpCas9(6AA) library; training the PAM ML model based on the first and the second training data sets to relate enzyme function to amino acid sequences; and using the PAM ML model to perform screening of amino acid combinations based on an amino acid sequence as input to obtain a PAM variant enzyme with predicted respective PAM requirements.
35. The method of claim 34, comprising: storing the generated first and second training data set at an enzyme data store; and providing an interface for querying data from the enzyme data store.
36. The method of claim 34, comprising: using the PAM ML model to predict PAM requirements of enzymes obtained through the interface of the enzyme data store; sorting the enzymes in the first and second training data sets according to activity and / or selectivity of the enzymes according to PAM requirements of the enzymes; and using the sorted enzymes, classifying the enzymes based on either activity or selectivity for each PAM combination for enzymes within the first and second training data sets.
37. The method of claim 34, wherein using the HT-PAMDA assay used to generate the first training data set comprises determining quantitative rate constants against all possible PAMs for each SpCas9 enzyme variant.
38. The method of claim 34, wherein training the PAM ML model comprises: assigning a label to each training enzyme of the first and second training data sets based on a most active PAM of the respective training enzyme; and performing a random selection of enzymes to define a training sample having balance across PAM classes.
39. The method of claim 38, wherein an enzyme is labeled as active when no PAM has a maximum rate constant k on a most active PAM, where k > 10'4.
40. A method for computationally generating SpCas9 variant enzymes with customizable properties, the method comprising: computationally mutating a starting sequence to generate a library of sequences having random amino acid substitutions with a defined hamming distance from the starting sequence; using a protospacer-adjacent motif (PAM) machine-learning (ML) model trained to relate PAM requirements to amino acid sequence, generating PAM predictions for each member of the library; determining a score, using a fitness function, for each sequence in the library according to properties defined for evolution; determining a set of sequences from the library having a respective score matching a threshold score criterion for selection as a subsequent starting sequence for a subsequent evaluation iteration; and iteratively perform evaluation comprising performing the computational mutation and the PAM predictions generation until the determined scores of the sequences using the customizable fitness function satisfy a fitness criterion.
41. The method of claim 40, wherein the PAM ML model is trained based on a first data set and a second data set, wherein: the first data set is generated by: identifying active SpCas9 enzymes from a SpCas9(6AA) library through structure-and-function-informed saturation mutagenesis and subsequent bacterial selection; profiling the active SpCas9 enzymes according to all possible protospacer- adjacent motif (PAM) requirements to amino acid sequences using a high- throughput PAM-determination assay (HT-PAMDA) to generate the first data set,and the second data set is generated by: randomly determining non-selected enzymes from the SpCas9(6AA) library for the first data set to be included in the second data set.
42. The method of claim 40, comprising receiving the fitness function through user input.
43. A computer-implemented method for automatic characterization of generated SpCas9 variant enzymes, the method comprising: obtaining SpCas9 enzymes from a data store comprising SpCas9(6AA) enzymes; cloning the SpCas9 enzymes into a mammalian expression plasmid; sequencing the cloned SpCas9 enzymes by a whole-ORF sequencing; and subjecting the sequenced SpCas9 enzymes to a high-throughput PAM- determination assay (HT-PAMDA) for comprehensive PAM characterization of the sequenced SpCas9 enzymes.
44. The method of claim 43, wherein the enzymes are obtained from the data store through bacterial positive selection against 16 different substrates encoding PAMs or by randomly identifying unselected enzymes from the data store.
45. The method of claim 43, wherein sequencing the cloned SpCas9 enzymes comprises: performing a whole-ORF sequencing process to perform tiled-amplicon sequencing across an entire coding sequence of an expression plasmid of an enzyme, and obtaining a complete sequence validation performed over up to 96 enzyme variants in parallel, to determine relationship between an amino acid sequence and a specific plasmid based on positions.
46. A method of editing an Apolipoprotein E epsilon-4 (APOE-sA (Argl30, Argl76)) allele in a cell, the method comprising introducing two C to T point mutations to revert the allele to APOE-e.2 (Cysl30, Cysl76) by contacting the cell with a first gRNA positioning a target cytosine at protospacer position 5 for Argl30, and a second gRNA positioning a target cytosine at protospacer position 5 for Argl76, wherein the Argl30 gRNA utilizes an NGTA PAM and the Argl76 gRNA utilizesan NGGC PAM, and a cytosine base editor comprising an SpCas9 variant selected from the group consisting of MRKCRS, MRKQKC, MRKSKC, MRKMRS, MRKCRN, MRKCRQ, QRKCKK, MRKCRK, MRACRQ, MRKMRK, and MRKMRQ, preferably wherein the first and second guide RNAs comprise APOE spacer sequences listed in Table 1.
47. A method of editing an HBB WIN mutation in a cell, the method comprising contacting the cell with an adenine base editor comprising an SpCas9 variant selected from the group consisting of LWQYQH, MWKYQS, LWKYQS, and MWKYQA, and a guide RNA, preferably an HBB guide RNA comprising a spacer sequence listed in Table 1.
48. A method of allele-specific deletion of RHO P23H allele in a cell, the method comprising contacting the cell with an SpCas9 variant selected from the group consisting of MRRWMR, KRHWMR, MRRMYR, KRRYQR, MRAFMR, KRAYQR, KRRWQR, MRHWQR, IRSMQR, KRAWQR, LRKMYR, ARGIMR, MKRCMV, MWNVML, and MKKCMN, and a guide RNA, preferably a RHO- P23H guide RNA comprising a spacer sequence listed in Table 1.
49. A method of editing a CYBB T362I mutation in a cell, the method comprising contacting the cell with an adenine base editor comprising an SpCas9 variant KWRQLC, and a guide RNA, preferably a CYBB guide RNA comprising a spacer sequence listed in Table 1.
50. The method of any of claims 46 to 49, wherein the cell is in a living subject.
51. The method of any of claims 46 to 49, comprising contacting the cell with a nucleic acid encoding the cytosine base editor, adenine base editor, or SpCas9 variant.
52. The method of claim 51, wherein the nucleic acid comprises a viral vector.
53. The method of any of claims 46 to 49, comprising contacting the cell with a ribonucleoprotein (RNP) complex comprising the guide RNAs and the cytosine base editor, adenine base editor, or SpCas9 variant.
Citation Information
Patent Citations
Target-specific CRISPR mutant
US11441135B2
Engineered CRISPR-Cas9 Nucleases with Altered PAM Specificity
US20220169998A1