Cell-free in vitro methods for characterization of base editor properties and determination of base editor editing outcomes

In vitro methods for characterizing CRISPR base editors using synthetic or genomic DNA targets with barcodes enable rapid determination of editing profiles, addressing the inefficiencies of in-cell experimentation by accurately predicting editing outcomes.

WO2026015202A1PCT designated stage Publication Date: 2026-01-15EMD MILLIPORE CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/029391
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-10-10
Filing Date
2025-05-14
Publication Date
2026-01-15

AI Technical Summary

Technical Problem

Current methods for characterizing CRISPR base editors are labor-intensive, resource-intensive, and time-consuming, requiring extensive in-cell experimentation to determine editing characteristics such as editing window, sequence context effects, and methylation sensitivity, which are complicated by chromatin context and DNA repair preferences of cell lines.

Method used

In vitro methods for characterizing CRISPR base editors using synthetic, plasmid-based, or genomic DNA targets, with barcoded sequences to deconvolute editing patterns, sequence context preferences, and methylation sensitivity, allowing for rapid determination of editing profiles without cell-based assays.

Benefits of technology

These methods faithfully recapitulate editing patterns observed in cells, providing a rapid, lower-cost approach to characterize editing windows, sequence context specificity, and methylation sensitivity of CRISPR base editors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025029391_15012026_PF_FP_ABST
    Figure US2025029391_15012026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention is directed toward methods for the in vitro characterization of base editor proteins and the in vitro determination of base editor outcomes at a target site.
Need to check novelty before this filing date? Find Prior Art

Description

Cell-Free in vitro Methods for Characterization of Base Editor Properties and Determination of Base Editor Editing OutcomesRelated Applications

[0001] The present application claims the benefit of priority of U.S. Provisional Patent Application No.: 63 / 705,784, filed October 10, 2024, and U.S. Provisional Patent Application No.: 63 / 669,831 , filed July 11 , 2024, the entire contents of each of which is incorporated herein by reference.Sequence Listing

[0002] This application contains a Sequence Listing that has been submitted in XML format via Patent Center and is hereby incorporated by reference in its entirety. The XML copy, created on April 28, 2025, is named P24-130-WO-PCT_SL, and is 587 kilobytes in size.Background

[0003] CRISPR base editing systems enable the installation of precise substitutions within a target DNA through the catalytic conversion of one nucleobase into another. These systems are comprised of a CRISPR Cas effector, such as Cas9 or Cas12, along with a fused enzymatic domain that performs the nucleobase conversion. For example, for cytosine base editors, the fused enzymatic domain is a deaminase that catalyzes the conversion of cytosine to uracil, leading to C-to-T substitutions; for adenine base editors, the fused enzymatic domain is a deaminase that catalyzes the conversion of adenine to inosine, leading to A-to-G substitutions. The Cas effector directs the base editor molecule to the desired target DNA by base pairing between the target and a complementary guide RNA bound to the Cas effector.

[0004] The editing characteristics of a base editor may vary considerably depending on factors such as the Cas effector or fused enzymatic domain present in the editor, as well as the presence of any additional domains that may modify the activity of the base editor. Editing experiments are time-consuming and expensive; therefore, it is crucial that the researcher be able to appropriately design anexperiment to maximize the likelihood of successfully obtaining the desired edit. To do so, the base editor must be sufficiently characterized such that features including editing window, sequence context effects, and the effect of target site methylation are known. These characteristics are commonly determined by editing a large number of genomic sites in cells, which is costly and time-consuming.

[0005] Further, the editing results for any particular target sequence are difficult to predict and may depend on various factors including the position of the target editable residues within the target sequence, the number and position of additional editable residues in the sequence and / or the specific CRISPR / Cas system or catalytic effector domain used. Additional factors may include the extent of methylation in the target sequence and the specific target site nucleotide sequence.

[0006] Thus, what is needed in the art are compositions and methods to determine the editing characteristics of any base editor protein, without the need for intensive in-cell experimentation, and the probable outcome of any particular base editing reaction prior to performing expensive and time consuming in vivo research and diagnostics.Summary of the Invention

[0007] The editing outcomes at a target can be influenced greatly by the design of the base editor fusion protein and by the deaminase selected to perform the catalytic step. For example, the portion of the target DNA capable of being edited is termed the “editing window.” A base editor with a wide editing window will perform the nucleobase conversion across a large proportion of the target DNA, allowing for flexibility in guide RNA selection but raising the likelihood of editing at unintended residues. In contrast, a base editor with a narrow editing window will perform the nucleobase conversion over a smaller proportion of the target DNA, enabling more precision in editing but requiring more stringent placement of the edited residue relative to the guide RNA. Different deaminase domains may have differing substrate specificities, with certain broader nucleotide contexts favored or disfavored, or differing requirements for the presence or absence of methylation.

[0008] Defining the editing window, sequence context dependence, and methylation sensitivity are crucial steps during the development of new base editor variants. The typical method to establish these parameters is to perform editing at a large number of genomic targets in cell lines and to deconvolute the effects of sequence context, methylation, and position from the rate of editing at these targets. However, these experiments are labor-, resource-, and time-intensive, requiring extensive hands-on time, expensive transfection reagents, several days in culture, and large amounts of sequencing data. In addition, the data is complicated by differences in the chromatin context at the sites chosen, as well as the DNA repair preferences of the cell line used. Therefore, there is a need for rapid, lower-cost methods with fewer complicating variables for the determination of editing window, sequence context, and DNA methylation preferences.

[0009] When a researcher uses a base editor, there are often several possible editing outcomes depending on the number of cytosines in the target sequence. The researcher may be able to predict approximately which residues are most likely to be edited using their knowledge of the editor’s characteristics, but there is currently no method to experimentally determine precise editing profiles at a given target without performing the editing experiment in a cell. For editing in organisms that require extensive time to achieve a result, or for which experiments are particularly costly, there would be a benefit to having a rapid, lower cost in vitro method.

[0010] Characterization of a base editor’s characteristics, such as editing window, sequence context preferences, and methylation sensitivity, currently requires editing in cells at hundreds of target loci, which is time- and resource-intensive. Further, there is currently no way to experimentally determine the editing profile at a target site without introducing the base editor into a cell.

[0011] The present invention discloses methods for the in vitro characterization of base editor proteins and the in vitro determination of base editor outcomes at a target site. We demonstrate that these assays faithfully recapitulate editing patterns observed in cells.Schematic descriptions of these methods can be found in Fig. 1 . Briefly, the disclosed methods utilize in vitro editing of a DNA target molecule, whether synthetic, plasmid-based, genomic, or amplified from a genome by PCR. The edited DNA is then sequenced to determine the editing pattern and efficiency of each target (Fig. 1A). For characterization experiments, the DNA targets comprise a pool of possible targets, with editable residues at different positions, in different sequence contexts, and / or with and without methylation (Fig. 1 B and 1C). Barcoding of the individual targets enables deconvolution of the sequencing data. For editing pattern determination (Fig. 1 D), each target is edited individually.

[0012] Thus, the present invention is directed toward a method to characterize the editing window of a base editor in vitro.

[0013] The present invention is further directed toward a method to characterize the sequence context specificity of a base editor in vitro.

[0014] The present invention is still further directed toward a method to characterize the methylation sensitivity of a base editor in vitro.

[0015] The present invention is still further directed toward a method to experimentally determine the editing profile at a given target site in vitro.

[0016] The present invention is still further directed toward a kit that facilitates one or more of the methods disclosed herein to experimentally determine the editing profile at a given target site in vitro.

[0017] Certain embodiments of the present invention are directed toward an in vitro method for determining the editing profile of a plurality of one or more double stranded DNA target sequences for a CRISPR base editor, comprising: providing: i) a plurality of target sequences, each target sequence having a unique barcode sequence, a number of editable residues distributed along the non-target strand of the target sequence, a protospacer adjacent motif (PAM) and, optionally, wherein one or more of the editable residues is adjoining a thymine residue, a guanine residue, an adenine residue, or a cytosine residue; ii) a base editor protein comprising a Cas protein linked to acatalytic effector domain, and; iii) a guide RNA (gRNA) suitable for base pairing with the target sequences, wherein said base editor protein and said gRNA form a ribonucleoprotein (RNP); incubating the set of target sequences with the RNP under conditions suitable for base editing of the target sequences to create editing products; sequencing the editing products and quantifying the rate of editing at the position of each editable residue, thereby determining the editing profile of the base editor, including the editing window and, optionally, sequence context effects.

[0018] The method may further comprise, wherein the editable residue is cytosine and the catalytic effector domain is a deaminase capable of editing cytosine.

[0019] The method may still further comprise, wherein the editable residue is adenine and the catalytic effector domain is a deaminase capable of editing adenine.

[0020] The method may yet still further comprise, wherein the editable residue is guanine and the catalytic effector domain is a guaninespecific glycosylase.

[0021] The method may yet still further comprise, wherein said target sequences are one or more annealed oligonucleotides.

[0022] The method may yet still further comprise, wherein said target sequences are one or more plasmids.

[0023] The method may yet still further comprise, wherein the unique barcode sequence is between 5 and 10 nucleotides; wherein the unique barcode sequence is 6 nucleotides or wherein the target sequences comprise 2 to 10 editable residues.

[0024] The method may yet still further comprise, wherein the Cas protein is selected from a group consisting of a Cas9 and a Cas12 protein.

[0025] The method may yet still further comprise, wherein the Cas protein is selected from a group consisting of a nickase Cas protein or a catalytically inactive Cas protein.

[0026] The method may yet still further comprise, wherein said gRNA targets a single DNA sequence.

[0027] The method may yet still further comprise wherein the gRNA is comprised of a pool of gRNAs and said pool of gRNAs target two or more DNA sequences.

[0028] The method may yet still further comprise, wherein the gRNA is a single guide RNA (sgRNA) or is a crRNA:tracrRNA complex.

[0029] The method may yet still further comprise, wherein the gRNA is transcribed in vitro or chemically synthesized.

[0030] The method may yet still further comprise, wherein the sequencing is next generation sequencing or Sanger-based sequencing methods. The method may yet further still comprise, wherein sequences required for next-generation sequencing are added to the target sequences prior to sequencing.

[0031] The method may yet still further comprise, wherein synthetic target DNAs are designed to contain sequences required for nextgeneration sequencing prior to execution of in vitro editing reactions.

[0032] The method may yet still further comprise, wherein the addition of sequences required for next-generation sequencing is accomplished by adding these sequences to oligonucleotide primers and performing polymerase chain reaction (PCR).

[0033] The method may yet still further comprise, wherein the addition of sequences required for next-generation sequencing is accomplished by ligation of adapters, followed by PCR using primers directed against the ligated sequences.

[0034] The method may yet still further comprise, wherein the added sequences allow for the multiplexing of multiple samples into a single sequencing run.

[0035] The method may yet still further comprise, wherein the added sequences allow for the multiplexing of multiple target sequences into a single sequencing run.

[0036] Certain embodiments of the present invention are directed toward an in vitro method for predicting editing outcomes of CRISPR base editing for a chromosomal target sequence, comprising: providing: i) a target oligonucleotide sequence, the target sequence having at least one editable residue on the non-target strand of thetarget sequence and a protospacer adjacent motif (PAM); ii) a base editor protein comprising a Cas protein, linked to a catalytic effector domain, and; iii) a guide RNA (gRNA) suitable for base pairing with the target sequence, wherein said base editor protein and said gRNA form a ribonucleoprotein (RNP); incubating the target sequence with the RNP under conditions suitable for base editing of the editable nucleotides to create editing products; sequencing the editing products and quantifying the rate of editing at each editable residue.

[0037] The method may further comprise, wherein the editable residue is cytosine and the catalytic effector domain is a deaminase capable of editing cytosine.

[0038] The method may still further comprise, wherein the editable residue is adenine and the catalytic effector domain is a deaminase capable of editing adenine.

[0039] The method may yet still further comprise, wherein the editable residue is guanine and the catalytic effector domain is a guaninespecific glycosylase.

[0040] The method may yet still further comprise, wherein the Cas protein is selected from a group consisting of a Cas9 and a Cas12 protein.

[0041] The method may yet still further comprise, wherein the Cas protein is selected from a group consisting of a nickase Cas protein or a catalytically inactive Cas protein.

[0042] The method may yet still further comprise, wherein the target sequence is selected from a group consisting of genomic DNA, a PCR amplicon, a plasmid and a double stranded synthetic oligonucleotide.

[0043] The method may yet still further comprise, wherein the guide RNA is a single guide RNA (sgRNA).

[0044] The method may yet still further comprise, wherein the guide RNA is a crRNA:tracrRNA complex.

[0045] The method may yet still further comprise, wherein the gRNA is synthetic.

[0046] The method may yet still further comprise, wherein the guideRNA is transcribed in vitro.

[0047] The method may yet still further comprise, wherein the sequencing is next generation sequencing (NGS) or Sanger-based sequencing methods and wherein sequences required for nextgeneration sequencing are added to the target sequences.

[0048] The method may yet still further comprise, wherein synthetic target sequences are designed to contain sequences required for nextgeneration sequencing prior to execution of in vitro editing reactions.

[0049] The method may yet still further comprise, wherein the addition of sequences required for next-generation sequencing is accomplished by adding these sequences to oligonucleotide primers and performing polymerase chain reaction (PCR).

[0050] The method may yet still further comprise, wherein the addition of sequences required for next-generation sequencing is accomplished by ligation of adapters, followed by PCR using primers directed against the ligated sequences.

[0051] The method may yet still further comprise, wherein the added sequences allow for the multiplexing of multiple samples into a single sequencing run.

[0052] The method may yet still further comprise, wherein the added sequences allow for the multiplexing of multiple target sequences into a single sequencing run.

[0053] Certain embodiments of the present invention are directed toward an in vitro method for characterizing methylation sensitivity of a CRISPR base editor, comprising: providing: i) a set of one or more differentially methylated target oligonucleotide sequences, the target oligonucleotide sequences having a unique barcode sequence, having a number of editable residues on the non-target strand of the target oligonucleotide sequence and having at least one methylated target or bystander residue, and a protospacer adjacent motif (PAM); ii) a base editor protein comprising a Cas protein linked to a catalytic effector domain, and; iii) a guide RNA (gRNA) suitable for base pairing with the target oligonucleotide sequence, wherein said base editor protein and said gRNA form a ribonucleoprotein (RNP); incubating the target oligonucleotide sequence with the RNP under conditions suitable forbase editing of the editable nucleotides to create editing products; sequencing the editing products and quantifying the rate of editing at each editable nucleotide, thereby determining the effect of methylation on editing.

[0054] The method may further comprise, wherein the editable residue is cytosine and the catalytic effector domain is a deaminase capable of editing cytosine.

[0055] The method may the editable residue is adenine and the catalytic effector domain is a deaminase capable of editing adenine.

[0056] The method may yet still further comprise, wherein the editable residue is guanine and the catalytic effector domain is a guaninespecific glycosylase.

[0057] The method may yet still further comprise, wherein the unique barcode sequence is between 5 and 10 nucleotides.

[0058] The method may yet still further comprise, wherein the unique barcode sequence is 6 nucleotides.

[0059] The method may yet still further comprise, wherein the target oligonucleotide sequences comprise 2 to 10 editable residues.

[0060] The method may yet still further comprise, wherein the Cas protein is selected from a group consisting of a Cas9 and a Cas12 protein.

[0061] The method may yet still further comprise, wherein the Cas protein is selected from a group consisting of a nickase Cas protein or a catalytically inactive Cas protein.

[0062] The method may yet still further comprise, wherein the gRNA targets a single DNA sequence.

[0063] The method may yet still further comprise, wherein the gRNA is comprised of a pool of gRNAs and said pool of gRNAs target two or more DNA sequences.

[0064] The method may yet still further comprise, wherein the gRNA is a single guide RNA (sgRNA).

[0065] The method may yet still further comprise, wherein the gRNA is a crRNA:tracrRNA complex.

[0066] The method may yet still further comprise, wherein the gRNA is transcribed in vitro.

[0067] The method may yet still further comprise, wherein the gRNA is synthetic.

[0068] The method may yet still further comprise, wherein the sequencing is next generation sequencing (NGS) or Sanger-based sequencing methods.

[0069] The method may yet still further comprise, wherein sequences required for next-generation sequencing are added to the target oligonucleotide sequences.

[0070] The method may yet still further comprise, wherein the target oligonucleotide sequences are designed to contain sequences required for next-generation sequencing prior to execution of in vitro editing reactions.

[0071] The method may yet still further comprise, wherein the addition of sequences required for next-generation sequencing is accomplished by adding these sequences to oligonucleotide primers and performing polymerase chain reaction (PCR).

[0072] The method may yet still further comprise, wherein the addition of sequences required for next-generation sequencing is accomplished by ligation of adapters, followed by PCR using primers directed against the ligated sequences.

[0073] The method may yet still further comprise, wherein the added sequences allow for the multiplexing of multiple samples into a single sequencing run.

[0074] The method may yet still further comprise, wherein the added sequences allow for the multiplexing of multiple target sequences into a single sequencing run.Brief Description of the Figures

[0075] Fig. 1 A - D show a schematic representation of disclosed methods. In Panel A an overview of general process of the present invention is shown. In Panels B-D, exemplary targets or target pools are shown. Protospacer-adjacent motifs (NGG) are underlined. Horizontal dashes indicate that a DNA molecule may extend beyondthe displayed sequence. Indicated in blue are dinucleotide sequences to be investigated (Panel B), CpG sites with differential methylation to be investigated (Panel C), and cytosine residues within the target sequence of a genomic site of interest (Panel D). Figure 1 B discloses SEQ ID NOS 52-65, respectively, in order of appearance. Figure 1 D discloses SEQ ID NO: 66.

[0076] Fig. 2 shows editing efficiency by dinucleotide sequence context for every position of the target sequence.

[0077] Figs. 3A - 3G show in vitro editing of genomic sites by the minimal OBE (SEQ ID NO: 1 ) faithfully recapitulates editing outcomes as seen in cells. (A) EMX1-11 ; (B) EMX1-15; (C) HBB; (D) HEKSite2; € RNF2; (F) AAVS1 ; (G) CEL.

[0078] Figs. 4A- 4G show in vitro editing of genomic sites by the SSB variant CBE (SEQ ID NO: 2) faithfully recapitulates editing outcomes as seen in cells. (A) EMX1-11 ; (B) EMX1-15; (C) HBB; (D) HEKSite2; € RNF2; (F) AAVS1 ; (G) CEL.

[0079] Fig. 5 shows SEQ ID NOs: 1 , 2 and 3. The underlined portions of SEQ ID NO: 3 are NLS sequences PAAKRVKLD (SEQ ID NO: 4 (c- MYC)) and PKKKRKV SEQ ID NO: 5 ((SV40 large T antigen)). The amino acids in italics denote linker sequences, SGGSGGSGGS (SEQ ID NO: 6) and GGGGS (SEQ ID NO: 7).

[0080] Fig 6A - 6B show methylation results of (A) Dem targets and (B) CpG targets. Figures disclose SEQ ID NOS 41 and 45, respectively, in order of appearance.Detailed Description of the Invention

[0081] Definitions

[0082] Unless defined otherwise, all technical and scientific terms used herein have the meaning commonly understood by a person skilled in the art to which this invention belongs. The following references provide one of skill with a general definition of many of the terms used in this invention: Singleton, et al., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994); The Cambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossary of Genetics, 5th Ed., R. Rieger, et al. (eds.), Springer Verlag (1991 ); and Hale &Marham, The Harper Collins Dictionary of Biology (1991 ). As used herein, the following terms have the meanings ascribed to them unless specified otherwise.

[0083] When introducing elements of the present disclosure or the preferred embodiments(s) thereof, the articles "a," "an," "the" and "said" are intended to mean that there are one or more of the elements. The terms "comprising," "including" and "having" are intended to be inclusive and mean that there may be additional elements other than the listed elements.

[0084] Furthermore, the transitional phrases “comprising,” “consisting essentially of” and “consisting of” have the meanings as given in MPEP 2111.03 (Manual of Patent Examining Procedure; United States Patent and Trademark Office, 9thEd., Revision Feb 2023 [R-07.2022]). Any claims using the transitional phrase “consisting essentially of” will be understood as reciting only essential elements ( / .e., the basic and novel characteristics) of the invention and any other elements recited in dependent claims are understood to be non-essential to the invention recited in the claim from which they depend. Likewise, any additional elements over those claimed that are described in a prior art reference(s) are excluded from the claims by use of the transitional phrase “consisting essentially of” as being non-essential to the claimed invention.

[0085] As used herein, the term "endogenous sequence" refers to a chromosomal sequence that is native to the cell.

[0086] The term "exogenous" or “exogenous sequence,” as used herein, refers to a sequence that is not native to the cell or a chromosomal sequence whose native location in the genome of the cell is in a different chromosomal location.

[0087] A "gene," as used herein, refers to a DNA region (including exons and introns) encoding a gene product, as well as all DNA regions which regulate the production of the gene product, whether or not such regulatory sequences are adjacent to coding and / or transcribed sequences. Accordingly, a gene includes, but is not necessarily limited to, promoter sequences, terminators, translationalregulatory sequences such as ribosome binding sites and internal ribosome entry sites, enhancers, silencers, insulators, boundary elements, replication origins, matrix attachment sites, and locus control regions.

[0088] The terms “complementary” or “complementarity” refer to the association of double-stranded nucleic acids by base pairing through specific hydrogen bonds. The base paring may be standard Watson- Crick base pairing (e.g., 5'-A G T C-3' pairs with the complementary sequence 3'-T C A G-5'). The base pairing also may be Hoogsteen or reversed Hoogsteen hydrogen bonding. Complementarity is typically measured with respect to a duplex region and thus, excludes overhangs, for example. Complementarity between two strands of the duplex region may be partial and expressed as a percentage (e.g., 70%), if only some of the base pairs are complementary. The bases that are not complementary are “mismatched.” Complementarity may also be complete ( / .e., 100%), if all the base pairs of the duplex region are complementary.

[0089] The term “homologous” refers to the extent two or more sequences are identical. Two sequences are considered to be homologous if they will hybridize to the same sequence under a defined set of conditions. Defined conditions include but are not limited to buffer formulation, temperature and sequence concentration.

[0090] The term "heterologous" refers to an entity that is not endogenous or native to the cell of interest. For example, a heterologous protein refers to a protein that is derived from or was originally derived from an exogenous source, such as an exogenously introduced nucleic acid sequence. In some instances, the heterologous protein is a protein not normally produced by the cell of interest.

[0091] The terms "nucleic acid" and "polynucleotide" refer to a deoxyribonucleotide or ribonucleotide polymer, in linear or circular conformation, and in either single- or double-stranded form. For the purposes of the present disclosure, these terms are not to be construed as limiting with respect to the length of a polymer. Theterms can encompass known analogs of natural nucleotides, as well as nucleotides that are modified in the base, sugar and / or phosphate moieties (e.g., phosphorothioate backbones). In general, an analog of a particular nucleotide has the same base-pairing specificity; i.e., an analog of A will base-pair with T.

[0092] The term “synthetic nucleic acid” refers to a nucleotide sequence synthesized in vitro (for example, in a lab and either manually or with a nucleic acid synthesizer device) and in which the sequence is not found in nature. The sequence may be, for example, DNA or RNA or a modification thereof as described below, may be any length and may be any sequence of nucleotides so long as the sequence is not naturally occurring.

[0093] The term "nucleotide" refers to deoxyribonucleotides or ribonucleotides. The nucleotides may be standard nucleotides (i.e., adenosine, guanosine, cytidine, thymidine, and uridine) or nucleotide analogs. A nucleotide analog refers to a nucleotide having a modified purine or pyrimidine base or a modified ribose moiety. A nucleotide analog may be a naturally occurring nucleotide (e.g., inosine) or a non- naturally occurring nucleotide. Non-limiting examples of modifications on the sugar or base moieties of a nucleotide include the addition (or removal) of acetyl groups, amino groups, carboxyl groups, carboxymethyl groups, hydroxyl groups, methyl groups, phosphoryl groups, and thiol groups, as well as the substitution of the carbon and nitrogen atoms of the bases with other atoms (e.g., 7-deaza purines). Nucleotide analogs also include dideoxy nucleotides, 2'-O-methyl nucleotides, locked nucleic acids (LNA), peptide nucleic acids (PNA), and morpholinos.

[0094] The terms "polypeptide" and "protein" are used interchangeably to refer to a polymer of amino acid residues.

[0095] The term “inactivated,” in the context of the present invention, may refer to the deletion of one or more amino acids and / or the substitution of one or more amino acids or the addition of one and / or more amino acids in a target protein such that one or more functions of the protein are eliminated or reduced to a level where activity is lessthan 95%, less than 90%, less than 75%, less than 50%, less than 40%, less than 30%, less than 20%, less than 15%, less than 10%, less than 5%, less than 4%, less than 3%, less than 2% or less than 1 % of the level of activity of the active protein. In the present invention, the Cas-like protein (e.g., Cas9) is catalytically inactivated as the result of having two amino acids substituted with Alanine residues as detailed elsewhere in this specification thereby inhibiting its ability to cleave double-stranded DNA. Reduced activity may be referred to “partially inactivated.” The term “catalytically modified” in the context of the present invention refers to a subject protein that is inactivated or partially inactivated and means that one or more of the catalytic activities of the subject protein have been reduced or eliminated. The nuclease activity of the Cas protein of the present invention may be partially or fully inactivated making the Cas protein a nickase (cleaving only one DNA strand; discussed in greater detail, below) or catalytically inactive (cleaving none of the DNA strands), respectively.

[0096] The term “directly linked” with regard to proteins and polypeptides in the context of the present invention means that two proteins are joined together ( / .e., fused, e.g., via peptide bonds) to form a continuous protein or polypeptide with no added amino acid residues incorporated between the two joined / fused proteins.

[0097] The term “indirectly linked” in the context of the present invention means that one or more amino acids are incorporated between two joined / fused proteins. Additionally, two or more proteins may be indirectly linked by chemical moieties other than amino acids, the methods of which are known to one of skill in the art.

[0098] Techniques for determining nucleic acid and amino acid sequence identity are known in the art. Typically, such techniques include determining the nucleotide sequence of the mRNA for a gene and / or determining the amino acid sequence encoded thereby, and comparing these sequences to a second nucleotide or amino acid sequence. Genomic sequences can also be determined and compared in this fashion. In general, “identity” refers to an exact nucleotide-to-nucleotide or amino acid-to-amino acid correspondenceof two polynucleotides or polypeptide sequences, respectively. Two or more sequences (polynucleotide or amino acid) can be compared by determining their “percent identity." The percent identity of two sequences, whether nucleic acid or amino acid sequences, is the number of exact matches between two aligned sequences divided by the length of the shorter sequences and multiplied by 100. An approximate alignment for nucleic acid sequences is provided by the local homology algorithm of Smith and Waterman, Advances in Applied Mathematics 2:482-489 (1981). This algorithm can be applied to amino acid sequences by using the scoring matrix developed by Dayhoff, Atlas of Protein Sequences and Structure, M. O. Dayhoff ed., 5 suppl. 3:353-358, National Biomedical Research Foundation, Washington, D.C., USA, and normalized by Gribskov, Nucl. Acids Res. 14(6):6745-6763 (1986). An exemplary implementation of this algorithm to determine percent identity of a sequence is provided by the Genetics Computer Group (Madison, Wis.) in the "BestFit" utility application. Other suitable programs for calculating the percent identity or similarity between sequences are generally known in the art. For example, another alignment program is BLAST, used with default parameters. For example, BLASTN and BLASTP can be used using the following default parameters: genetic code=standard; filter=none; strand=both; cutoff=60; expect=10; Matrix=BLOSUM62;Descriptions=50 sequences; sort by=HIGH SCORE; Databases=non- redundant, GenBank+EMBL+DDBJ+PDB+GenBank CDS translations+Swiss protein+Spupdate+PIR. Details of these programs can be found on the GenBank website.

[0099] CRISPR / Cas Proteins and Systems

[0100] Understanding the present invention will be aided by understanding CRISPR / Cas protein systems in general and in the context of the present invention. Although much of the basics of CRISPR / Cas9 technology is understood by one of skill in the art, as detailed below, the present invention provides for the ability to better utilize the technology by providing methods for the in vitrocharacterization of base editor proteins and the in vitro determination of base editor outcomes at a target site.

[0101] The CRISPR / Cas9-based System

[0102] In its basic form, the CRISPR / Cas9 system introduces a double strand break close to the binding site of a guide RNA (gRNA). A Cas9 protein and a gRNA form a ribonucleoprotein (RNP) complex. This complex itself can be transfected into a recipient cell (for example, via the use of extracellular vesicles) or a plasmid or viral vector encoding the complex can be used. The gRNA targets the complex to the desired location in the genome through formation of an R-loop - a three-stranded nucleic acid structure composed of a DNA:RNA hybrid and a displaced single-stranded DNA - where the Cas9 protein (an endonuclease) causes a double-strand DNA break where modifications to the sequence can be made. Such modifications can take the form of non-homologous end joining (NHEJ) making random small insertions or deletions (indels) or homology-directed repair (HDR). NHEJ is useful for creating knockout mutations and HDR is useful for making desired modifications to the target sequence, also termed “precision editing.” While HDR enables the introduction of these precision edits, its use is hindered by low rates of installation of the desired edit and high rates of indels.

[0103] Base Editors

[0104] As discussed above, traditional CRISPR / Cas9 technology has challenges regarding efficiency and a high rate of undesired indels. Base editing was developed (Komor, 2016) to increase the efficiency of precision editing using the CRISPR / Cas9 technology while minimizing the rate of indels. Base editing allows for the direct, irreversible conversion of one target DNA base pair into another without requiring a dsDNA backbone break or a donor template. Rather, a target nucleobase is converted into another by a catalytic effector domain tethered to the RNP complex. For cytosine base editing, the tethered catalytic domain is a cytosine deaminase, which chemically converts a cytosine nucleobase (C) into a uracil (II). Uracil has the base-pairing properties of thymine (T). Thus, upon DNA replication or repair, thetargeted C:G pair is converted to a T:A pair. For adenine base editing, the tethered catalytic domain is an adenine deaminase, which chemically converts adenine (A) into inosine (I). Inosine has the basepairing properties of guanine (G). Thus, upon DNA replication or repair, the targeted A:T pair is converted to a G:C pair. For guanine base editing, the tethered catalytic domain is a guanine-specific glycosylase, which excises the guanine nucleobase from the DNA. Translesion synthesis at the abasic site leads to the installation of C or T, resulting in the installation of a G:C to C:G or T :A substitution.

[0105] RNA-Guided Endonucleases

[0106] While not limited by theory, the methods of the present invention may be used to better predict the outcome of CRISPR / Cas editing by 1 ) characterizing the editing window of a base editor, 2) characterizing the sequence context specificity of a base editor, 3) determining the editing profile at a given target site and / or 4) characterizing the methylation sensitivity of a base editor. In this regard, it may be helpful to more fully understand the basic technology for CRISPR / Cas editing. An overview is provided here.

[0107] RNA-guided endonucleases, such as Cas9, may comprise at least one nuclear localization signal (NLS), at least one nuclease domain, and at least one domain that interacts with a guide RNA (gRNA) to direct the endonuclease / deaminase complex of the present invention to a specific nucleobase for modification. Also known are nucleic acids encoding the RNA-guided endonucleases, as well as methods of using the RNA-guided endonucleases to modify chromosomal sequences of eukaryotic cells or embryos. The RNA- guided endonuclease interacts with specific gRNAs, each of which directs the endonuclease to a specific targeted site, at which site the catalytic effector domain modifies the target nucleobase, resulting in the installation of a substitution. Since the specificity is provided by the gRNA, the RNA-based endonuclease is, essentially, universal and can be used with different gRNAs to target different genomic sequences.

[0108] Many forms of guide RNAs (gRNAs) are known in the field. In general, a gRNA is an RNA molecule that can direct an RNA bindingprotein or endonuclease to a specific nucleic acid sequence / target site by means of base pairing. A gRNA may comprise a single RNA molecule, such as a CRISPRRNA (crRNA), or two RNA molecules, such as a crRNA and a tracrRNA (or, trRNA). In some embodiments, it may additionally comprise an accessory RNA or DNA molecule. In some embodiments, a crRNA and a tracrRNA may be covalently linked together to form a chimeric guide RNA or a single guide RNA (sgRNA). It is also well established in the field that a gRNA can be introduced alone or in combination with its cognate gRNA binding protein or endonuclease to a target DNA in different forms. It can be expressed from a DNA vector introduced into a target cell. It can be synthesized by in vitro transcription or by chemical reaction before being introduced to a target DNA. A person of skill in the art should know that many chemical modifications are possible on a gRNA by chemical synthesis. For example, certain modifications, such as 2’-0 methyl group modification and phosphorothioate linkage modification, may be introduced into a gRNA to alter its stability or performance.Nonstandard nucleotides, such as DNA and LNA, also may be introduced into a gRNA during chemical synthesis to alter its specificity or performance. A guide RNA that may have a different form or may contain a different chemical modification may be used in conjunction with the present disclosure without deviating from the spirit of the present disclosure.

[0109] RNA-guided endonucleases may comprise at least one nuclear localization signal (NLS), which permits entry of the endonuclease into the nuclei of eukaryotic cells and embryos such as, for example, nonhuman one cell embryos. RNA-guided endonucleases also comprise at least one nuclease domain and at least one domain that interacts with a gRNA. An RNA-guided endonuclease is directed to a specific nucleic acid sequence (or target site) by a gRNA. The gRNA interacts with the RNA-guided endonuclease as well as the target site such that, once directed to the target site, the associated catalytic effector domain can convert the target nucleobase to another nucleobase, thereby installing a substitution. Since the gRNA provides the specificity forthe targeted cleavage the RNA-guided endonuclease is universal (providing that its ability to cleave DNA is disabled) and can be used with different gRNAs to modify different target residues. RNA-guided endonucleases can be proteins, can be encoded by isolated nucleic acids (i.e., RNA or DNA), can be encoded by vectors comprising nucleic acids encoding the RNA-guided endonucleases, and can be protein-RNA complexes comprising the RNA-guided endonuclease plus a gRNA.

[0110] The RNA-guided endonuclease can be derived from a clustered regularly interspersed short palindromic repeats (CRISPR) / CRISPR- associated (Cas) system. The CRISPR / Cas system can be a type I, a type II, or a type III system, as known to one of ordinary skill in the art. Non-limiting examples of suitable CRISPR / Cas proteins include Cas3, Cas4, Cas5, Cas5e (or CasD), Cas6, Cas6e, Cas6f, Cas7, Cas8a1 , Cas8a2, Cas8b, Cas8c, Cas9, Casio, Cas10d, Cas12, CasF, CasG, CasH, Csy1 , Csy2, Csy3, Cse1 (or CasA), Cse2 (or CasB), Cse3 (or CasE), Cse4 (or CasC), Csc1 , Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1 , Cmr3, Cmr4, Cmr5, Cmr6, Csb1 , Csb2, Csb3,Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csz1 , Csx15, Csf1 , Csf2, Csf3, Csf4, and Cu1966. One of ordinary skill in the art will be able to modify any RNA-guided nuclease to inactivate its catalytic activity for use with a base editor system.

[0111] The RNA-guided endonuclease may be derived from a type II CRISPR / Cas system. More specifically, the RNA-guided endonuclease may be derived from a Cas9 protein. The Cas9 protein can be from, for example, Streptococcus pyogenes, Streptococcus thermophilus, Streptococcus sp., Nocardiopsis dassonvillei, Streptomyces pristinaespiralis, Streptomyces viridochromogenes, Streptomyces viridochromogenes, Streptosporangium roseum, Streptosporangium roseum, Alicyclobacillus acidocaldarius, Bacillus pseudomycoides, Bacillus selenitireducens, Exiguobacterium sibiricum, Lactobacillus delbrueckii, Lactobacillus salivarius, Microscilla marina, Burkholderiales bacterium, Polaromonas naphthalenivorans, Polaromonas sp., Crocosphaera watsonii, Cyanothece sp., Microcystisaeruginosa, Synechococcus sp., Acetohalobium arabaticum, Ammonifex degensii, Caldicelulosiruptor becscii, Candidatus Desulforudis, Clostridium botulinum, Clostridium difficile, Finegoldia magna, Natranaerobius thermophilus, Pelotomaculum thermopropionicum, Acidithiobacillus caldus, Acidithiobacillus ferrooxidans, Allochromatium vinosum, Marinobacter sp., Nitrosococcus halophilus, Nitrosococcus watsoni, Pseudoalteromonas haloplanktis, Ktedonobacter racemifer, Methanohalobium evestigatum, Anabaena variabilis, Nodularia spumigena, Nostoc sp., Arthrospira maxima, Arthrospira platensis, Arthrospira sp., Lyngbya sp., Microcoleus chthonoplastes, Oscillatoria sp., Petrotoga mobilis, Thermosipho africanus, or Acaryoch I oris marina.

[0112] In general, CRISPR / Cas proteins comprise at least one RNA recognition domain and / or RNA binding domain. RNA recognition and / or RNA binding domains interact with guide RNAs. CRISPR / Cas proteins also usually comprise nuclease domains ( / .e., DNase or RNase domains; but for the purposes of base editors one or more of these are disabled), DNA binding domains, helicase domains, RNase domains, protein-protein interaction domains, dimerization domains, as well as other domains.

[0113] The CRISPR / Cas-like protein can be a wild type CRISPR / Cas protein, a modified CRISPR / Cas protein, or a fragment of a wild type or modified CRISPR / Cas protein. The CRISPR / Cas-like protein can be modified to increase nucleic acid binding affinity and / or specificity, alter an enzymatic activity (e.g., inactivate its ability to cleave DNA), and / or change another property of the protein. For example, nuclease ( / .e., DNase, RNase) domains of the CRISPR / Cas-like protein can be modified, deleted, or inactivated. Alternatively, the CRISPR / Cas-like protein can be truncated to remove domains that are not essential for the function of the fusion protein. The CRISPR / Cas-like protein can also be truncated or modified to optimize the activity of the effector domain of the fusion protein.

[0114] In some embodiments, the CRISPR / Cas-like protein can be derived from a wild type Cas9 protein or fragment thereof. In otherembodiments, the CRISPR / Cas-like protein can be derived from a modified Cas9 protein. For example, the amino acid sequence of the Cas9 protein can be modified to alter one or more properties (e.g., nuclease activity, affinity, stability, etc.) of the protein. Alternatively, domains of the Cas9 protein not involved in RNA-guided cleavage can be eliminated from the protein such that the modified Cas9 protein is smaller than the wild type Cas9 protein.

[0115] In general, a Cas9 protein comprises at least two nuclease ( / .e., DNase) domains. For example, a Cas9 protein can comprise a RuvC- like nuclease domain and an HNH-like nuclease domain. The RuvC and HNH domains work together to cut single strands to make a double-stranded break in DNA (Jinek, et al., Science, 2012, 337: 816- 821). In some embodiments, the Cas9-derived protein can be modified to contain only one functional nuclease domain (either a RuvC-like or a HNH-like nuclease domain). For example, the Cas9-derived protein can be modified such that one of the nuclease domains is deleted or mutated such that it is no longer functional ( / .e., the nuclease activity is absent). In some embodiments in which one of the nuclease domains is inactive, the Cas9-derived protein is able to introduce a nick into a double-stranded nucleic acid (such protein is termed a "nickase"), but not cleave both strands of the double-stranded DNA. For example, an aspartate to alanine (D10A) conversion in a RuvC-like domain converts the Cas9-derived protein into a nickase. Likewise, a histidine to alanine (H840A or H839A) conversion in a HNH domain converts the Cas9-derived protein into a nickase. Each nuclease domain can be modified using well-known methods, such as site-directed mutagenesis, PCR-mediated mutagenesis, and total gene synthesis, as well as other methods known in the art. Modifying both of these domains results in inactivation of nuclease and nickase activity.

[0116] For in vivo eukaryotic use, the RNA-guided endonuclease may comprise at least one NLS. For in vitro use the RNA-guided nuclease may comprise a NLS to mimic components of and be better predictive of in vivo use. In general, an NLS comprises a stretch of basic amino acids. Nuclear localization signals are known in the art (see, e.g.,Lange, et al., J. Biol. Chem., 2007, 282:5101-5105). For example, in one embodiment, the NLS can be a monopartite sequence, such as PAAKRVKLD; (SEQ ID NO: 4) or PKKKRKV; (SEQ ID NO: 5). In another embodiment, the NLS can be a bipartite sequence. In still another embodiment, the NLS can be KRPAATKKAGQAKKKK; (SEQ ID NO: 8). The NLS can be located at the N-terminus, the C-terminus, or in an internal location of the RNA-guided endonuclease. One of ordinary skill in the art will be able to determine the best location for a particular use in view of the teachings of this specification.

[0117] Further, an RNA-guided endonuclease may additionally comprise at least one cell-penetrating domain. For example, the cellpenetrating domain can be a cell-penetrating peptide sequence derived from the HIV-1 TAT protein. As an example, the TAT cell-penetrating sequence can be GRKKRRQRRRPPQPKKKRKV; (SEQ ID NO: 9). For another example, the cell-penetrating domain can be TLM (PLSSIFSRIGDPPKKKRKV; SEQ ID NO: 10), a cell-penetrating peptide sequence derived from the human hepatitis B virus. In still another example, the cell-penetrating domain can be MPG (GALFLGWLGAAGSTMGAPKKKRKV; SEQ ID NO:11 or GALFLGFLGAAGSTMGAWSQPKKKRKV; SEQ ID NO: 12). In an additional example, the cell-penetrating domain can be Pep-1 (KETWWETWWTEWSQPKKKRKV; SEQ ID NO: 13), VP22, a cell penetrating peptide from Herpes simplex virus, or a polyarginine peptide sequence. The cell-penetrating domain can be located at the N-terminus, the C-terminus, or in an internal location of the protein. One of ordinary skill in the art will be able to determine the best location for a particular use in view of the teachings of this specification.

[0118] In certain embodiments, the RNA-guided endonuclease may be part of a protein-RNA complex comprising a gRNA. The gRNA interacts with the RNA-guided endonuclease to direct the endonuclease to a specific target site, wherein, for example, the 5' end of the guide RNA base pairs with a specific protospacer sequence.

[0119] Target Site

[0120] An RNA-guided endonuclease in conjunction with a gRNA is directed to a target site in the nucleotide sequence, wherein the RNA- guided endonuclease (enzymatically inactivated) binds the nucleotide sequence. The target site has no sequence limitation except that the sequence is immediately adjacent to a consensus sequence known as a protospacer adjacent motif (PAM). Examples of PAMs include, but are not limited to, NGG, NGGNG, and NNAGAAW (wherein N is defined as any nucleotide and W is defined as either A or T). As detailed above, the spacer region of the gRNA is complementary to the protospacer of the target sequence. Typically, the spacer region of the gRNA is about 19 to 21 nucleotides in length. Thus, in certain aspects, the sequence of the target site in the chromosomal sequence is 5'-Ni9- 21-NGG-3'.

[0121] The target site can be in the coding region of a gene, in an intron of a gene, in a control region of a gene, in a non-coding region between genes, etc. The gene can be a protein coding gene or an RNA coding gene. The gene can be any gene of interest.

[0122] With regard to the present invention, the target sequence may be a synthetic sequence designed, for example, to mimic (wholly or partly) or copy (wholly or partly) a natural sequence. The sequence may also be a synthetic sequence designed for a specific purpose (e.g., for an in vitro testing purpose or an assay development purpose) but without mimicking or without significantly mimicking a naturally occurring sequence. The target sequence may be a naturally occurring sequence or contain at least 99%, 98%, 95%, 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10% or less of a naturally occurring sequence. The sequence may be integrated into a carrier sequence such as a plasmid or vector. The plasmid or vector may be naturally occurring or synthetic, as is needed for the desired purpose.

[0123] The sequence can be any length; however, lengths of less than 10,000 bases, 5,000 bases 2,000 bases, 1 ,000 bases, 900 bases, 800, bases, 700, bases, 600 bases, 500, bases, 400, bases, 300 bases, 200 bases or 100 bases are suitable for use as target sequences for the present invention. Any one length of nucleic acid may contain one ormore target sequences as needed or desired for a particular use. In this regard, a single nucleotide sequence can contain multiple target sequences and, if desired, multiple assays (two or more) of the present invention may be performed in parallel or in series on the same nucleotide sequence on which mutable target sequences are found. The target sequence can be chemically modified to contain a modified base, a modified sugar moiety, a modified phosphodiester linkage, a fluorophore, a quencher, or a combination of a fluorophore and a quencher. The modifications may include, but are not limited to, 5- methylcytosine, 5-hydroxymethylcytosine, methyladenine (N6mdA), 8- oxoguanine (8-oxo-7,8-dihydroguanine), ribose, 2’-O-methyl group- containing ribose, locked nucleic acid (LNA), and phosphorothioate linkage. Many fluorophores and quenchers are suitable for attaching to a DNA oligo. Modifications may be made on a target nucleotide or on a non-target nucleotide.

[0124] Kits

[0125] Still another aspect of the present disclosure provides kits for carrying out the methods described above.

[0126] The kits provided herein generally include instructions for carrying out the processes detailed above. Instructions included in the kits may be affixed to packaging material, may be included as a package insert or as a downloadable file. While the instructions are typically written or printed materials, they are not limited to such. Any medium capable of storing such instructions and communicating them to an end user is contemplated by this disclosure. Such media include, but are not limited to, electronic storage media (e.g., magnetic discs, tapes, cartridges, chips), optical media (e.g., CD ROM), and the like. As used herein, the term “instructions” can include the address of an internet site that provides the instructions.

[0127] As various changes could be made in the above-described processes and kits without departing from the scope of the invention, it is intended that all matter contained in the above description and in the examples given below, shall be interpreted as illustrative and not in a limiting sense.

[0128] Next Gen Sequencing

[0129] Next generation sequencing (NGS), also called second generation, massively parallel, deep sequencing or high throughput DNA sequencing, is a DNA sequencing technology which has revolutionized genomic, medical and biological research. Using NGS, an entire human genome can be sequenced within a single day. In contrast, the previous Sanger sequencing technology, used to decipher the human genome, required over a decade to deliver the final draft.

[0130] NGS platforms perform sequencing of millions of small fragments of DNA in parallel. Bioinformatics analyses are used to piece together these fragments by mapping the individual reads. Behjati and Tarpey, Arch Dis Child Educ Pract Ed. 2013 Dec;98(6):236- 8. doi: 10.1136 / archdischild-2013-304340. Epub 2013 Aug 28. Although often associated with genomic research and diagnostics, NGS can also be used for less demanding work and automate routine procedures where speed and accuracy are needed.

[0131] NGS is a term referring to an array of related techniques. For example, as reviewed by Shendure and Ji (Shendure, J., Ji, H. Nextgeneration DNA sequencing. Nat Biotechnol 26, 1135-1145 (2008)) the concept of one type of NGS, cyclic-array sequencing, can be summarized as the sequencing of a dense array of DNA features by iterative cycles of enzymatic manipulation and imaging-based data collection.

[0132] Further, as reviewed by Slatko, et al. (Overview of Next Generation Sequencing Technologies, Curr. Protoc. Mol. Biol., 2018 April; 122(1 ), e59: which is representative of what was known by one of skill in the art at the time of the present invention), NGS can be roughly grouped into two major categories, sequencing by hybridization and sequencing by synthesis.

[0133] Sequencing by hybridization utilizes repeated hybridization and washing away of the unwanted non-hybridized DNA and determining if the hybridizing labeled fragments matched the sequence of the DNA probes on the filter. Sequencing by hybridization is largely used withtechnologies that depend upon using specific probes to interrogate sequences, such as in diagnostic applications. Slatko, et al.

[0134] Sequencing by synthesis (SBS) methods are varied but are based on Sanger sequencing methods (Sanger sequencing methods are well known to one of skill in the art, See, e.g., McGinn S, Gut IG. DNA sequencing - spanning the generations. N Biotechnol. 2013 May 25;30(4):366-72.), without the dideoxy terminators, in combination with repeated cycles of synthesis, imaging and methods to incorporate additional nucleotides in the growing chain. Current SBS methods differ from the approach of the original Sanger sequencing in that they rely on shorter reads, about 300 - 500 bases. Overlapping reads are “assembled” to create a consensus sequence. Slatko, et al.

[0135] The preferred method for the present method is NGS, however, the present invention is not limited by the method used to sequence the resultant edits from CRISPR-based editing.Exemplification

[0136] Example 1 : Characterization of CBE dinucleotide sequence preference and editing window

[0137] Double-stranded DNA targets were constructed by cloning annealed oligos into a plasmid backbone. Each set of annealed oligos contained a unique 6nt barcode, a target sequence containing three cytosine residues on the non-targeted strand distributed across the target sequence, and a 5’-NGG-3’ protospacer adjacent motif (PAM). Each set of annealed oligos was designed such that the residue preceding the cytosine residues was either thymine (TC dinucleotide) or a degenerate residue comprising adenine, cytosine or guanine (VC dinucleotide where V = not T for DNA and not U for RNA). Plasmids were pooled such that the NC dinucleotide frequency was equal for each value of N (where N = any one base) and the frequency of cytosine at each position was equal. Finally, plasmid pools were diluted to a concentration of 20 fmol / 10 pL and placed on ice until the time of experiment.

[0138] Cytosine base editor (CBE) proteins were constructed by fusing the C-terminal domain of human APOBEC3B to the amino-terminus ofan SpCas9 nickase protein (SEQ ID NO: 1). Proteins were expressed and purified from E. coli BL21AI by autoinduction and nickel column chromatography and were stored at -80°C before use in a buffer containing 10% glycerol, 300 mM KCI, 20 mM HEPES (pH 7.5), and 1 mM DTT.

[0139] Synthetic crRNAs were designed to target the pool of dsDNAs and ordered from IDT (Integrated DNA Technologies, Coralville, IA). crRNAs targeting dsDNAs harboring VC dinucleotides were synthesized with degenerate bases at the relevant positions. crRNAs were pooled such that the NC dinucleotide frequency was equal for each value of N and the frequency of cytosine at each position was equal. The crRNA pool was complexed with an equimolar amount of trRNA.

[0140] Ribonucleoproteins (RNPs) were assembled by incubating an equimolar amount of crRNA:tracrRNA complex with 19.8 pmol of CBE protein in a 33 pL final volume (6 pmol / 10 pL) at room temperature for 15 minutes. RNPs were placed on ice until the time of the experiment.

[0141] Reactions were set up on ice in triplicate. 10 pL of chilled plasmid pool dilution and 10 pL of chilled CBE RNP were aliquoted per well of a 96-well PCR plate. The plate was incubated in a thermocycler at 37°C for 10 min and 80°C for 15 min.

[0142] The target region was amplified by PCR using JumpStart Taq ReadyMix (MilliporeSigma, Burlington, MA) and the following cycling conditions: 94°C / 2m; 25 cycles of 94 °C / 30s, 58°C / 30s, 72 °C / 45s; 72 °C / 2m. Primers are listed in Table 1. PCR products underwent a second round of amplification using Illumina (San Diego, CA) index primers and JumpStart Taq ReadyMix and the following conditions: 95°C / 3m; 9 cycles of 95C / 30s, 55 °C / 30s, 72 °C / 30s; 72 °C / 5m. Indexed PCR products were quantified by PicoGreen (ThermoFisher, Waltham, MA), purified by Select-a-Size DNA Clean & Concentrator MagBeads (Zymo, Irvine, CA) using 1.2x beads by volume, and pooled according to DNA content. Pools were diluted to 4 nM. Sequencing was performed on an Illumina MiSeq instrument using a 300-cycle kit to obtain single-end reads. FASTQ files for each sample were analyzedusing a custom analysis script and the CRIS.py software package (Connelly and Pruett-Miller, Sci. Rep., 2019, doi.org / 10.1038 / s41598- 019-40896-w).Table 1. Primer sequences for Example 1

[0143] Results are presented in Fig. 2. Plotted is the rate of editing of cytosine for each NC dinucleotide, for each outcome. Results are organized by position in the target, numbered from the PAM. For DC dinucleotides (where D is adenine, guanine, or thymine), the only possible outcomes are DC (unedited) and DT (edited). For CC dinucleotides, possible outcomes include CC (unedited), TC (edited only at the first cytosine), CT (edited only at the second cytosine), and TT (edited at both cytosines). This data reveals that the CBE protein exhibits a marked dinucleotide sequence preference for TC dinucleotides, with considerable editing at CC dinucleotides, and minimal editing at AC and GC dinucleotides. The human APOBEC3B enzyme has been well characterized to exhibit a preference for TC substrates (Ito et al. J. Mol. Biol., 2017, doi.org / 10.1016 / j.jmb.2017.04.021 ; Burns, et a!., Nature, 2013, doi:10.1038 / nature11881); therefore, these in vitro cytosine base editing results recapitulate established knowledge. In addition, this data reveals that the tested CBE protein efficiently deaminates cytosine residues beginning at position 6 upstream of the PAM through position 18 upstream of the PAM.

[0144] Example 2: In vitro elucidation of editing pattern for a genomic target

[0145] Target DNA was amplified from genomic DNA isolated from HEK293 cells by PCR using JumpStart Taq ReadyMix (MilliporeSigma) and the following cycling conditions: 94 °C / 2m; 30 cycles of 94°C / 30s, TA / 30s, 72 °C / 45s; 72 °C / 2m using primers and annealing temperatures (TA) found in Table 2. PCR products were purified by Select-a-Size DNA Clean & Concentrator MagBeads (Zymo) using 1.2x beads by volume and quantified by Qubit (ThermoFisher). Purified PCR products were diluted with buffer containing 20 mM Tris, 100 mM NaCI, and 5 mM MgCl2 to a final concentration of 100 fmol / 10 pL and chilled on ice.Table 2. Primer sequences for Example 2

[0146] A minimal cytosine base editor (CBE) protein without UGI was constructed by fusing the C-terminal domain of human APOBEC3B to the amino terminus of an SpCas9 nickase protein (SEQ ID NO: 1). The minimal CBE was further modified to contain a T4 phage SSB at the amino terminus of the CBE (SEQ ID NO: 2). Proteins were expressed and purified from E. coli as in Example 1. Synthetic single guide RNAs (sgRNAs) targeting seven human genomic sites were purchased from MilliporeSigma or IDT. Spacer sequences of the sgRNAs are given in Table 3Table 3. sgRNA spacer sequences for Example 2

[0147] Ribonucleoprotein (RNP) complexes for the in vitro assay were prepared by mixing CBE protein with an equimolar amount of sgRNA in a buffer containing 20 mM Tris, 100 mM NaCI, and 5 mM MgCI2 for a final concentration of 6 pmol / 10 pL. RNP mixtures were incubated for 15 min at room temperature and placed on ice.

[0148] In vitro assays were performed by mixing 10 pL of the prepared RNP with 10 pL of the diluted purified PCR product in wells of a 96-well PGR plate on ice. Plates were incubated at 37°C for 45 minutes and 80°C for 15 min. Edited PCR product underwent a second round of amplification using Illumina index primers and Jumpstart Taq ReadyMix and the following conditions: 95°C / 3m; 9 cycles of 95°C / 30s, 55°C / 30s, 72 °C / 30s; 72 °C / 5m. Indexed PCR products were quantified by PicoGreen (ThermoFisher), pooled according to DNA content, and purified by Select-a-Size DNA Clean & Concentrator MagBeads (Zymo) using 1 ,2x beads by volume. Sequencing and analysis were performed as in Example 1.

[0149] A recombinant UGI containing a Bacillus phage UGI, a c-MYC nuclear localization signal (NLS; SEQ ID NO: 4: the first underlined portion of SEQ ID NO: 3), and a SV40 Large T antigen NLS (SEQ ID NO: 5, the second underlined portion of SEQ ID NO: 3) was purified from E. coli. Dextran sulfate sodium salt with average molecular weight greater than 500 kDa (Product number: D8906) was purchased from MilliporeSigma. A dextran sulfate solution was prepared by dissolving the chemical in water at 50 pg / pL and sterilized by filtration through a 0.22 urn filter. The stock solution was diluted with water to prepare working solutions of 1 pg / pL.

[0150] RNP complexes for transfection into cells were prepared by incubating for 15 min at room temperature: 40 pmol of CBE protein, 120 pmol of sgRNA, and 15 pg of UGI in buffer (20 mM HEPES, 100 mM KCI, 0.5 mM DTT, 0.1 mM EDTA, pH 7.5) for a final volume of 10 pL. RNPs were kept on ice until transfection. HEK293 cells were obtained from ATCC and grown at 37°C and 5% CO2 in DMEM supplemented with 10% PBS, 2 mM L-glutamine, 1 mM sodium pyruvate, and 0.1 mM non-essential amino acids. Cells were seeded at 1 .67 x 104cells / cm2of tissue culture surface area two days before transfection. At the time of transfection, cells were trypsinized to obtain a single-cell suspension, washed twice with Hank’s Balanced Salt Solution and resuspended in Nucleofector Solution (Lonza, Basel, Switzerland) at 2.5 x 105cells per 100 pL. Dextran sulfate solution was added to the cell suspension to a final concentration of 0.5 pg per 100 pL and mixed well by swirling. Nucleofection was performed by mixing 100 pL of prepared cell suspension with 10 pL of complexed CBE RNP by pipetting up and down six times before transferring to a cuvette for electroporation using program Q-001 on a Nucleofector 2b machine (Lonza). Nucleofected cells were immediately transferred to 6-well plates containing 2 mL pre-warmed media per well and grown for 3 days before harvest.

[0151] Genomic DNA was harvested by try psi n izi ng transfected cells and resuspending in 75 pL QuickExtract reagent (Lucigen, Middleton, Wl). Suspensions were incubated at 60 °C for 15m and 95 °C for 15m.Genomic regions targeted by the CBE were amplified by PCR, and libraries were prepared, sequenced, and analyzed as in Example 1.

[0152] Results are presented in Figs. 3 and 4. For both figures, data is plotted with the in vitro rate of editing on the x-axis and the in-cell rate of editing on the y-axis. For in-cell editing, any C-to-D (where D is A, G, or T) substitution is included, as base excision repair activity within the cell at the site of deamination by the CBE protein can result in the installation of C-to-R substitutions. Each data point represents a cytosine residue within the target sequence and is calculated as the average of two (in-cell rates) or three (in vitro rates) technical replicates. Each panel presents data from individual target sites: A) EMX1-11 ; B) EMX1-15; C) HBB; D) HEKSite2; E) RNF2; F) AAVS1 ; G) CEL.

[0153] The data show that there is good correlation between the editing rates in vitro and in the cell. Across the seven sites and both CBE variants, all cytosine residues that were unedited in cells had correspondingly low editing rates in vitro. Of the 56 cytosine residues with detectable levels of editing, the in vitro assay correctly predicted the rank of editing within the target for 45 residues (80%). 10 residues (18%) were mis-predicted by one rank; only a single residue (2%) was mis-predicted by more than one rank (see Fig. 2F).

[0154] Example 3: Determining Methylation Sensitivity

[0155] Double-stranded DNA targets were constructed by annealing synthetic oligos and diluting to a concentration of 100 fmol / 10 pL with a buffer containing 20 mM Tris, 100 mM NaCI, and 5 mM MgCh. Oligos were comprised of 1 ) a target sequence containing editable cytosine residues; 2) a protospacer-adjacent motif (PAM); and 3) PCR primer binding sites at the 5’ and 3’ ends of the molecules. Three variations on dsDNA targets were synthesized: first, targets were synthesized with C5 methylation on the editable cytosine residues of the nontargeted strand (NTS); second, targets were synthesized with C5 methylation within the gRNA-binding sequence on the targeted strand (TS); third, targets were synthesized without methylation on either strand. As acontrol for Cas9 binding, methylated targets also contained at least one unmethylated cytosine in the editing window. Exemplary target sequences are given in Table 4. Oligos were annealed such that: 1 ) only the NTS was methylated; 2) only the TS was methylated; 3) both strands were methylated; and 4) neither strand was methylated.Table 4. Target sequences. C5 methylated residues are indicated as [5MedC]

[0156] Cytosine base editor proteins were constructed and purified as in Examples 1 and 2. Synthetic single guide RNAs (sgRNAs) were purchased from MilliporeSigma or other commercial source. Spacer sequences of the sgRNAs were identical to the unmethylated targets (NTS) given in Table 4. Ribonucleoprotein (RNP) complexes for the in vitro assay were prepared by mixing CBE protein with an equimolar amount of sgRNA in a buffer containing 20 mM Tris, 100 mM NaCI, and 5 mM MgCh for a final concentration of 6 pmol / 10 pL. RNP mixtures were incubated for 15 min at room temperature and placed on ice.

[0157] In vitro assays were performed in triplicate by mixing 10 pL of the prepared RNP with 10 pL of the diluted annealed oligonucleotide in wells of a 96-well PCR plate on ice. Plates were incubated in a thermocycler at 37°C for 45 minutes and 80°C for 15 min. The oligonucleotides were amplified by PCR using JumpStart Taq ReadyMix (MilliporeSigma), or equivalent, and the following cycling conditions: 94°C / 2m; 25 cycles of 94 °C / 30s, 58 °C / 30s, 72 °C / 45s; 72 °C / 2m. Primers are listed in Table 5. PCR products underwent a second round of amplification using Illumina index primers and JumpStart Taq ReadyMix and the following conditions: 95 °C / 3m; 9 cycles of 95°C / 30s, 55 °C / 30s, 72 °C / 30s; 72 °C / 5m. Indexed PCR products were quantified by PicoGreen (ThermoFisher), or equivalent, purified by Select-a-Size DNA Clean & Concentrator Mag Beads (Zymo), or equivalent, using 1.2x beads by volume, and pooled according to DNA content. Pools were diluted to 4 nM. Sequencing was performed on an Illumina MiSeq instrument, or equivalent, to obtain single-end reads. FASTQ files for each sample were analyzed using, for example, a custom analysis script and the CRIS.py software package (Connelly & Pruett-Miller, Sci. Rep., 2019, doi.org / 10.1038 / S41598-019-40896-w).Table 5. PCR primer sequences

[0158] We anticipated that editing at the control, unmethylated cytosine residue would be equivalent for the methylated and unmethylated targets, indicating that methylation of the NTS does not affect Cas9 binding to the target dsDNA. Further, we predicted that editing rates would be unaffected by methylation of the TS. Finally, we predicted that editing of methylated cytosine residues would be lower than editing of the same residues on the unmethylated target, indicating that the base editor is sensitive to the methylation state of DNA.

[0159] Data is shown in Fig. 6. There are two types of methylation - Dem (Fig. 6A) and CpG (Fig 6B). Each target had residues that were always unmethylated; these are represented by solid-color bars in the plots and serve as a control. Residues that may be methylated are represented by the hatch pattern. There are four conditions for each methylation: fully unmethylated control, methylation only on the TS (guide-binding) strand, methylation only on the NTS (edited) strand, and methylation on both strands. The data show that methylation of the NTS (edited) strand has a strong impact on editing of the methylated cytosine residues, but not nearby unmethylated residues. Methylation of the TS (guide-binding) strand has minimal impact on editing.

Claims

Claims1 . An in vitro method for determining the editing profile of a plurality of one or more double stranded DNA target sequences for a CRISPR base editor, comprising: a. providing: i) a plurality of target sequences, each target sequence having a unique barcode sequence, a number of editable residues distributed along the non-target strand of the target sequence, a protospacer adjacent motif (PAM) and, optionally, wherein one or more of the editable residues is adjoining a thymine residue, a guanine residue, an adenine residue, or a cytosine residue; ii) a base editor protein comprising a Cas protein linked to a catalytic effector domain, and; iii) a guide RNA (gRNA) suitable for base pairing with the target sequences, wherein said base editor protein and said gRNA form a ribonucleoprotein (RNP); b. incubating the set of target sequences with the RNP under conditions suitable for base editing of the target sequences to create editing products; c. sequencing the editing products and quantifying the rate of editing at the position of each editable residue, thereby determining the editing profile of the base editor, including the editing window and, optionally, sequence context effects.

2. The method of Claim 1 , wherein the editable residue is cytosine and the catalytic effector domain is a deaminase capable of editing cytosine.

3. The method of Claim 1 , wherein the editable residue is adenine and the catalytic effector domain is a deaminase capable of editing adenine.

4. The method of Claim 1 , wherein the editable residue is guanine and the catalytic effector domain is a guanine-specific glycosylase.

5. The method of Claim 1 , wherein said target sequences are one or more annealed oligonucleotides.

6. The method of Claim 1 , wherein said target sequences are one or more plasmids.

7. The method of Claims 1-6, wherein said unique barcode sequence is between 5 and 10 nucleotides.

8. The method of Claims 1-6, wherein said unique barcode sequence is 6 nucleotides.

9. The method of Claims 1-8, wherein said target sequences comprise 2 to 10 editable residues.

10. The method of Claims 1-9, wherein said Cas protein is selected from a group consisting of a Cas9 and a Cas12 protein.11 . The method of Claims 1-10, wherein the Cas protein is selected from a group consisting of a nickase Cas protein or a catalytically inactive Cas protein.

12. The method of Claims 1 - 11 , wherein said gRNA targets a single DNA sequence.

13. The method of Claims 1 - 11 , wherein said gRNA is comprised of a pool of gRNAs and said pool of gRNAs target two or more DNA sequences.

14. The method of Claims 1 - 13, wherein said gRNA is a single guide RNA (sgRNA).

15. The method of Claims 1 - 13, wherein said gRNA is a crRNA:tracrRNA complex.

16. The method of Claims 1 - 15, wherein the gRNA is synthetic.

17. The method of Claims 1 - 15 wherein the gRNA is transcribed in vitro.

18. The method of Claims 1-17, wherein said sequencing is next generation sequencing or Sanger-based sequencing methods.

19. The method of Claims 1-18, wherein the target sequence is designed to contain the sequences required for next-generation sequencing.

20. The method of claims 1-18, wherein sequences required for nextgeneration sequencing are added to the target sequences.21 .The method of claim 20, wherein the addition of sequences required for next-generation sequencing is accomplished by adding these sequences to oligonucleotide primers and performing polymerase chain reaction (PCR).

22. The method of claim 20, wherein the addition of sequences required for next-generation sequencing is accomplished by ligation of adapters, followed by PCR using primers directed against the ligated sequences.

23. The method of claims 19 - 20, wherein the sequences required for nextgeneration sequencing allow for the multiplexing of multiple samples into a single sequencing run.

24. The method of claims s 19 - 20, wherein the sequences required for nextgeneration sequencing allow for the multiplexing of multiple target sequences into a single sequencing run.

25. An in vitro method for predicting editing outcomes of CRISPR base editing for a chromosomal target sequence, comprising: a. providing: i) a target sequence, the target sequence having at least one editable residue on the non-target strand of the target sequence and a protospacer adjacent motif (PAM); ii) a base editor protein comprising a Cas protein, linked to a catalytic effector domain, and; iii) a guide RNA (gRNA) suitable for base pairing with the target sequence, wherein said base editor protein and said gRNA form a ribonucleoprotein (RNP); b. incubating the target sequence with the RNP under conditions suitable for base editing of the editable nucleotides to create editing products; c. sequencing the editing products and quantifying the rate of editing at each editable residue.

26. The method of Claim 25, wherein the editable residue is cytosine and the catalytic effector domain is a deaminase capable of editing cytosine.

27. The method of Claim 25, wherein the editable residue is adenine and the catalytic effector domain is a deaminase capable of editing adenine.

28. The method of Claim 25, wherein the editable residue is guanine and the catalytic effector domain is a guanine-specific glycosylase.

29. The method of Claims 25 - 28, wherein said Cas protein is selected from a group consisting of a Cas9 and a Cas12 protein.

30. The method of Claims 25 - 29, wherein the Cas protein is selected from a group consisting of a nickase Cas protein or a catalytically inactive Cas protein.31 .The method of Claims 25 - 30, wherein said target sequence is selected from a group consisting of genomic DNA, a PCR amplicon, a plasmid and a double stranded synthetic oligonucleotide.

32. The method of Claims 25 - 31 , wherein the guide RNA is a single guide RNA (sgRNA).

33. The method of Claims 25 - 31 , wherein the guide RNA is a crRNA:tracrRNA complex.

34. The method of Claims 25 - 33, wherein the gRNA is synthetic.

35. The method of Claims 25 - 33, wherein the guide RNA is transcribed in vitro.

36. The method of Claims 25 - 35, wherein said sequencing is next generation sequencing (NGS) or Sanger-based sequencing methods.

37. The method of Claims 25 - 36, wherein the target sequence is designed to contain the sequences required for next-generation sequencing.

38. The method of claims 25 - 36, wherein sequences required for nextgeneration sequencing are added to the target sequences.

39. The method of claim 38, wherein the addition of sequences required for next-generation sequencing is accomplished by adding these sequences to oligonucleotide primers and performing polymerase chain reaction (PCR).

40. The method of claim 38, wherein the addition of sequences required for next-generation sequencing is accomplished by ligation of adapters, followed by PCR using primers directed against the ligated sequences.41 .The method of claims 37 - 40, wherein the sequences required for nextgeneration sequencing allow for the multiplexing of multiple samples into a single sequencing run.

42. The method of claims 37 - 40, wherein the sequences required for nextgeneration sequencing allow for the multiplexing of multiple target sequences into a single sequencing run.

43. An in vitro method for characterizing methylation sensitivity of a CRISPR base editor, comprising: a. providing: i) a set of one or more differentially methylated target oligonucleotide sequences, the target oligonucleotide sequences having a unique barcode sequence, having a number of editable residues on the non-target strand of the target oligonucleotide sequence and having at least one methylated target or bystander residue, and a protospacer adjacent motif (PAM); ii) a base editor protein comprising a Cas protein linked to a catalytic effector domain, and; iii) a guide RNA (gRNA) suitable for base pairing with the target oligonucleotide sequence, wherein said base editor protein and said gRNA form a ribonucleoprotein (RNP); b. incubating the target oligonucleotide sequence with the RNP under conditions suitable for base editing of the editable nucleotides to create editing products; c. sequencing the editing products and quantifying the rate of editing at each editable nucleotide, thereby determining the effect of methylation on editing.

44. The method of Claim 43, wherein the editable residue is cytosine and the catalytic effector domain is a deaminase capable of editing cytosine.

45. The method of Claim 43, wherein the editable residue is adenine and the catalytic effector domain is a deaminase capable of editing adenine.

46. The method of Claim 43, wherein the editable residue is guanine and the catalytic effector domain is a guanine-specific glycosylase.

47. The method of Claims 43 - 46, wherein said unique barcode sequence is between 5 and 10 nucleotides.

48. The method of Claims 43 - 47, wherein said unique barcode sequence is 6 nucleotides.

49. The method of Claims 43 - 48, wherein said target oligonucleotide sequences comprise 2 to 10 editable residues.

50. The method of Claims 43 - 49, wherein said Cas protein is selected from a group consisting of a Cas9 and a Cas12 protein.51 .The method of Claims 43 - 50, wherein the Cas protein is selected from a group consisting of a nickase Cas protein or a catalytically inactive Cas protein.

52. The method of Claims 43 - 51 , wherein said gRNA targets a single DNA sequence.

53. The method of Claims 43 - 51 , wherein said gRNA is comprised of a pool of gRNAs and said pool of gRNAs target two or more DNA sequences.

54. The method of Claims 43 - 53, wherein said gRNA is a single guide RNA (sgRNA).

55. The method of Claims 43 - 53, wherein said gRNA is a crRNA:tracrRNA complex.

56. The method of Claims 43 - 55, wherein the gRNA is transcribed in vitro.

57. The method of Claims 43 - 55, wherein the gRNA is synthetic.

58. The method of Claims 43 - 57, wherein said sequencing is next generation sequencing (NGS) or Sanger-based sequencing methods.

59. The method of Claims 43 - 58, wherein the target sequence is designed to contain the sequences required for next-generation sequencing.

60. The method of Claims 43 - 58, wherein sequences required for nextgeneration sequencing are added to the target oligonucleotide sequences.61 .The method of Claim 60, wherein the addition of sequences required for next-generation sequencing is accomplished by adding these sequences to oligonucleotide primers and performing polymerase chain reaction (PCR).

62. The method of Claim 60, wherein the addition of sequences required for next-generation sequencing is accomplished by ligation of adapters, followed by PCR using primers directed against the ligated sequences.

63. The method of Claims 59 - 62, wherein the added sequences allow for the multiplexing of multiple samples into a single sequencing run.

64. The method of Claims 59 - 62, wherein the added sequences allow for the multiplexing of multiple target sequences into a single sequencing run.