Method and compositions for quantification of base editor protein activity

An in vitro method using fluorescent-labeled dsDNA oligonucleotides allows for rapid and cost-effective assessment of base editor protein activity and detection of contaminating nucleases, addressing the inefficiencies of existing cell-based assays.

WO2025250226A1PCT designated stage Publication Date: 2025-12-04EMD MILLIPORE CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/019516
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-20
Filing Date
2025-03-12
Publication Date
2025-12-04

AI Technical Summary

Technical Problem

Current methods for evaluating base editor protein activity are time-consuming and expensive, relying on transfection into living cells and next-generation sequencing, and lack a rapid, low-cost assay to measure activity and detect contaminating nucleases.

Method used

An in vitro method using a dsDNA oligonucleotide substrate labeled with a fluorescent moiety and quencher, where base editing results in enzymatic conversion and cleavage, allowing for real-time fluorescence measurement of editing activity and detection of contaminating nucleases.

Benefits of technology

Enables rapid, cost-effective evaluation of base editor protein activity and identification of contaminating nucleases, reducing the time and cost associated with traditional methods while eliminating dependence on cellular DNA repair machinery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025019516_04122025_PF_FP_ABST
    Figure US2025019516_04122025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention discloses compositions and methods for an in vitro procedure for the quantification of the activity of base editor proteins.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD AND COMPOSITIONS FOR QUANTIFICATION OF BASE EDITOR PROTEIN ACTIVITYRelated Applications

[0001] The present application claims the benefit of priority of U.S. Provisional Application No. 63 / 662,005, filed June 20, 2024, and U.S. Provisional Application No. 63 / 652,822, filed May 29, 2024, the entire contents of each of which is incorporated herein by reference.Sequence Listing

[0002] The present application contains a Sequence Listing that has been submitted in XML format via Patent Center and is hereby incorporated by reference in its entirety. The XML copy, created on January 29, 2025, is named P24-103-WO-PCT_SL, and is 25.3 kilobytes in size.Background

[0003] Base editing systems enable the installation of precision substitutions within a target DNA without producing double stranded DNA breaks. These systems are comprised of a CRISPR effector, such as, for example, Cas9 or Cas12a, along with a fused enzymatic domain to chemically modify a nucleotide, thereby resulting in the installation of a nucleotide substitution. For some base editing systems, including cytosine and adenine base editing, this enzymatic domain is a deaminase that converts cytosine to uracil or adenine to inosine, leading to C-to-T or A-to-G substitutions. Guanine base editing systems utilize a glycosylase domain that excises a guanine nucleobase, leading to G-to-Y substitutions. The Cas effector directs the base editor molecule to the desired target DNA by base pairing between the target and a complementary guide RNA bound by the Cas effector. The Cas effector typically is engineered with a partially or fully inactive DNA cleavage function to prevent the generation of double-stranded breaks (DSBs) in the target DNA, thereby minimizing the generation of potentially deleterious insertions and deletions.

[0004] While plasmid delivery of base editing systems is prevalent, this mode of delivery carries an elevated risk of off-target effects due to extended, high- level expression of the base editor and guide RNA. This risk can be mitigated by using purified base editor protein and chemically synthesized guide RNA, delivered as a ribonucleoprotein (RNP). However, the efficacy of these delivered RNPs relies on both the purity and the activity of the purified protein. Contaminating nucleases may degrade the guide RNA or cause undesired cleavage of the cellular genomic DNA that may be mutagenic. Preparations with lower activity will yield reduced editing rates, which may make it more difficult for the researcher to obtain the desired variant. Further, there are many variations on the basic architecture of base editor systems which utilize alternative enzymatic domains or accessory domains to modulate the activity of the protein. Different base editor protein variants may have different editing activity, which may be important in variant selection or base editor variant development.

[0005] Editing activity of purified base editor protein is typically evaluated by transfecting the base editor protein-guide RNA complex as an RNP into a cell line and evaluating editing rates at the target site using next-generation sequencing. These experiments yield quality data but are time-consuming and relatively expensive. Therefore, there is a need for a rapid, low-cost assay with which to measure base editor protein activity, and an assay that provides insight into the level of contamination with nucleases.Summary of the Invention

[0006] The present invention discloses a method for the quantification of base editor protein activity in vitro. Currently, the primary method with which to evaluate base editor protein activity is delivery into living cells, followed by next generation sequencing, which is expensive and time-consuming. We hypothesized that in vitro editing of a dsDNA oligonucleotide substrate could be detected by fluorescent methods using DNA endonucleases, thereby enabling rapid, low-cost evaluation of protein activity.

[0007] This method is depicted in Figure 1. It is comprised of a dsDNA oligonucleotide containing a protospacer adjacent motif (PAM) and sequence containing an editable deoxynucleotide residue. For cytosine base editors, this residue is cytosine; for adenine base editors, this residue is adenine; for guanine base editors, this residue is guanine. The dsDNA oligonucleotide is labeled at the PAM-distal end with a fluorescent moiety on one strand and a quencher moiety on the other strand. Base editing by a base editor complexed with a guide RNA targeting the dsDNA oligonucleotide substrate results in the enzymatic conversion of the editable residue. For cytosine, the residue is converted to uracil; for adenine, the residue is converted to inosine; for guanine, the nucleobase is excised, resulting in an abasic site. The dsDNA oligonucleotide is cleaved at the site of the edit using endonuclease enzymes specific to the chemical modification. For cytosine base editing, the uracil nucleobase is removed and the phosphodiester bond at the resulting abasic site is cleaved; for adenine base editing, the phosphodiester bond is cleaved at the site of deoxyinosine; for guanine base editing, the phosphodiester bond at the abasic site is cleaved. This breakage of the dsDNA oligonucleotide substrate enables diffusion of one of the labeled DNA ends and release of the proximal association between the fluorophore and the quencher, resulting in increased fluorescent signal, which can be measured by a fluorescence detector.

[0008] To test this method, we incubated cytosine base editor RNPs with a dsDNA oligonucleotide substrate containing editable cytosine residues. The fluorescent signal increased with increasing incubation length, which could be monitored in real time. In addition, increasing ratios of RNP to dsDNA oligonucleotide substrate yielded increased fluorescence. These results demonstrate that in vitro editing of the dsDNA oligonucleotide substrate results in measurable editing using fluorescent signal in a dose-dependent manner.

[0009] This method also provides a mechanism by which to evaluate for contaminating DNA nuclease activity within preparations of base editor protein. These contaminating DNA nucleases could have deleterious effectson a cell if they were transfected alongside a base editor protein, especially in a clinical context. Further, any contaminating nuclease will increase the apparent activity of the preparation of base editor protein, making it appear more active than it is. To accomplish this, we prepared cytosine base editor RNPs using a guide RNA targeting the dsDNA oligonucleotide substrate and a second RNP using a nontargeting guide RNA. We found that very pure preparations of base editor protein yielded very little fluorescent signal using the nontargeting guide RNP; in contrast, preparations with substantial contaminating extraneous protein yielded much higher signal using the nontargeting guide RNP. These results demonstrate that the methods of the present invention can identify and control for the presence of contaminating DNA nuclease within base editor protein preparations.

[0010] Plate readers and qPCR machines are known to suffer from well position effects that affect the absolute level of fluorescence measured. We control for this by including two independent within-well normalization methods. The first uses an independent fluorophore in the assay buffer, which is not altered by the base editing reaction and does not exhibit spectral overlap with the fluorophore utilized on the dsDNA oligonucleotide substrate. Normalization of editing fluorescent signal to this second fluorescent signal reduces noise in the data and allows for sample-to-sample comparison. The second normalization method utilizes elevated temperature to fully melt the dsDNA oligonucleotide substrate, providing a measurement of the maximum fluorescence of the reaction. Enzymatic activity can then be calculated as a percentage of the maximum fluorescence at a defined time point, which can be used to compare multiple samples within an experiment or across multiple experiments.

[0011] In summary, the combination of these strategies results in a method that measures base editor protein activity in vitro without the need for transfection into cells or next-generation sequencing. This results in methods that are more cost effective and reduce the time required to obtain activity data from several days, with prior art methods, to several hours. The methods of the present invention also eliminate the dependence on cellularDNA repair machinery for editing outcomes. Further, the methods of the present invention can be used to evaluate purified base editor proteins for contaminating DNA nucleases.

[0012] In one aspect, the present invention is a method for measuring the activity of a base editor in vitro, the method comprising: providing: i) a double stranded DNA (dsDNA) oligonucleotide comprising a target nucleotide or nucleotides for editing, a protospacer adjacent motif (PAM), a fluorescent moiety, and a quencher moiety, wherein the proximity between the fluorophore and quencher prevents or reduces the emission of fluorescence, ii) a first base editor ribonucleoprotein (RNP) comprising a site-specific endonuclease fused to a deaminase and / or glycosylase domain and a guide RNA (gRNA) complementary to the target dsDNA, and iii) a second base editor RNP comprising a site-specific endonuclease fused to a deaminase and / or glycosylase domain and a guide RNA (gRNA) not complementary to the target dsDNA; contacting the dsDNA with the first base editor RNP to edit the dsDNA target nucleotide or nucleotides; wherein, upon base editing, the enzymatic breakage of the dsDNA at the site of one or more editable nucleotide residues results in a physical separation of fluorophore and quencher, resulting in a fluorescent signal increased over background fluorescence; separately, contacting the dsDNA with the second base editor RNP; and wherein any fluorescent signal elicited by the second base editor RNP above background fluorescence is not due to activity of the base editor but due to nonspecific nuclease activity.

[0013] In another aspect of the present invention, the fluorescent signal is quantified by an instrument capable of measuring fluorescence.

[0014] In still another aspect of the present invention the nonspecific normalized fluorescence achieved by contacting the dsDNA with the second RNP is subtracted from the normalized fluorescence achieved by contacting the dsDNA with the first RNP targeted to the dsDNA to calculate the activity specific to base editor activity.

[0015] In still another aspect of the present invention, the fluorescent moiety and the quencher moiety are on the same end of the dsDNA, on opposite strands.

[0016] In still another aspect of the present invention, both the fluorescent moiety and the quencher moiety are on different ends of the same strand of the dsDNA.

[0017] In still another aspect of the present invention, the site-specific endonuclease is a Cas domain.

[0018] In still another aspect of the present invention, the site-specific endonuclease is a Cas9 domain or a Cas12 domain.

[0019] In still another aspect of the present invention, the Cas domain has been altered to abrogate its endonuclease activity on one DNA strand.

[0020] In still another aspect of the present invention, the Cas domain has been altered to abrogate its endonuclease activity on both strands of DNA.

[0021] In still another aspect of the present invention, the guide RNA is a synthetic single guide RNA (sgRNA).

[0022] In still another aspect of the present invention, the guide RNA is an in vitro transcribed sgRNA.

[0023] In still another aspect of the present invention, the guide RNA is a crRNA complexed with a tracrRNA.

[0024] In still another aspect of the present invention, the activity is normalized using a free fluorescent molecule with a different spectrum from the target DNA-bound fluorophore that is not affected by the base editor RNP.

[0025] In still another aspect of the present invention, the activity is normalized by raising the temperature of the mixture, thereby fully separating the fluorophore and quencher.

[0026] In still another aspect of the present invention, the instrument capable of measuring fluorescence is a plate reader or a qPCR machine.

[0027] In still another aspect of the present invention, the base editor is a cytosine base editor (CBE).

[0028] In still another aspect of the present invention, the CBE is a fusion protein comprised of a site-specific endonuclease and a deaminase domain capable of catalyzing the conversion of cytosine to uracil.

[0029] In still another aspect of the present invention, the nucleotide residue capable of being edited by the CBE is cytosine.

[0030] In still another aspect of the present invention, the enzymatic activities that result in DNA strand breakage at the site of a successful edit are comprised of uracil DNA glycosylase and AP lyase activities.

[0031] In still another aspect of the present invention, the enzymatic activities that result in DNA strand breakage at the site of a successful edit are contained in separate reagents.

[0032] In still another aspect of the present invention, the enzymatic activities that result in DNA strand breakage at the site of a successful edit are contained in a single reagent.

[0033] In still another aspect of the present invention, the BE is an adenine base editor (ABE).

[0034] In still another aspect of the present invention, the ABE is a fusion protein comprised of a site-specific endonuclease and a deaminase domain capable of catalyzing the conversion of adenine to inosine.

[0035] In still another aspect of the present invention, the nucleotide residue capable of being edited by the ABE is adenine.

[0036] In still another aspect of the present invention, the enzymatic activity that results in DNA strand breakage at the site of a successful edit is performed by endonuclease V.

[0037] In still another aspect of the present invention, the base editor is a glycosylase-based guanine base editor (gGBE).

[0038] In still another aspect of the present invention, the gGBE is a fusion protein comprised of a site-specific endonuclease and a glycosylase domain capable of catalyzing the excision of guanine from DNA.

[0039] In still another aspect of the present invention, the nucleotide residue capable of being edited by the gGBE is guanine.

[0040] In still another aspect of the present invention, the enzymatic activity that results in DNA strand breakage at the site of a successful edit is comprised of an AP lyase activity.Brief Description of the Figures

[0041] Fig. 1 shows a schematic representation of one embodiment of the disclosed method.

[0042] Fig. 2A - C shows the results of an in vivo activity assay for many preparations of protein and correlation with in vivo editing rates. (A) Percent maximum fluorescence achieved for Protein A and Protein B. (B) BE-specific percent maximum fluorescence for Protein A and Protein B. (C) In vitro activity vs. in vivo editing.

[0043] Fig. 3A shows APOBEC3B-SpCas9 nickase amino acid sequence [SEQ ID NO: 1], also referred to herein as Protein A. Fig. 3B shows T4 SSB- APOBEC3B-SpCas9 nickase [SEQ ID NO: 2], also referred to herein as Protein B.Detailed Description of the Invention

[0044] Definitions

[0045] Unless defined otherwise, all technical and scientific terms used herein have the meaning commonly understood by a person skilled in the art to which this invention belongs. The following references provide one of skill with a general definition of many of the terms used in this invention: Singleton, et al., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994); The Cambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossary of Genetics, 5th Ed., R. Rieger, et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991). As used herein, the following terms have the meanings ascribed to them unless specified otherwise.

[0046] When introducing elements of the present disclosure or the preferred embodiments(s) thereof, the articles "a," "an," "the" and "said" are intended to mean that there are one or more of the elements. The terms "comprising,""including" and "having" are intended to be inclusive and mean that there may be additional elements other than the listed elements.

[0047] Furthermore, the transitional phrases “comprising,” “consisting essentially of” and “consisting of” have the meanings as given in MPEP 2111.03 (Manual of Patent Examining Procedure; United States Patent and Trademark Office, 9thEd., Revision Feb 2023 [R-07.2022]). Any claims using the transitional phrase “consisting essentially of” will be understood as reciting only essential elements ( / .e., the basic and novel characteristics) of the invention and any other elements recited in dependent claims are understood to be non-essential to the invention recited in the claim from which they depend. Likewise, any additional elements over those claimed that are described in a prior art reference(s) are excluded from the claims by use of the transitional phrase “consisting essentially of.”

[0048] As used herein, the term "endogenous sequence" refers to a chromosomal sequence that is native to the cell.

[0049] The term "exogenous" or “exogenous sequence,” as used herein, refers to a sequence that is not native to the cell or a chromosomal sequence whose native location in the genome of the cell is in a different chromosomal location.

[0050] A "gene," as used herein, refers to a DNA region (including exons and introns) encoding a gene product, as well as all DNA regions which regulate the production of the gene product, whether or not such regulatory sequences are adjacent to coding and / or transcribed sequences. Accordingly, a gene includes, but is not necessarily limited to, promoter sequences, terminators, translational regulatory sequences such as ribosome binding sites and internal ribosome entry sites, enhancers, silencers, insulators, boundary elements, replication origins, matrix attachment sites, and locus control regions.

[0051] The terms “complementary” or “complementarity” refer to the association of double-stranded nucleic acids by base pairing through specific hydrogen bonds. The base paring may be standard Watson-Crick base pairing (e.g., 5'-A G T C-3' pairs with the complementary sequence 3'-T C AG-5'). The base pairing also may be Hoogsteen or reversed Hoogsteen hydrogen bonding. Complementarity is typically measured with respect to a duplex region and thus, excludes overhangs, for example. Complementarity between two strands of the duplex region may be partial and expressed as a percentage (e.g., 70%), if only some of the base pairs are complementary. The bases that are not complementary are “mismatched.” Complementarity may also be complete ( / .e., 100%), if all the base pairs of the duplex region are complementary.

[0052] The term “homologous” refers to the extent two or more sequences are identical. Two sequences are considered to be homologous if they will hybridize to the same sequence under a defined set of conditions. Defined conditions include but are not limited to buffer formulation, temperature and sequence concentration.

[0053] The term "heterologous" refers to an entity that is not endogenous or native to the cell of interest. For example, a heterologous protein refers to a protein that is derived from or was originally derived from an exogenous source, such as an exogenously introduced nucleic acid sequence. In some instances, the heterologous protein is a protein not normally produced by the cell of interest.

[0054] The terms "nucleic acid" and "polynucleotide" refer to a deoxyribonucleotide or ribonucleotide polymer, in linear or circular conformation, and in either single- or double-stranded form. For the purposes of the present disclosure, these terms are not to be construed as limiting with respect to the length of a polymer. The terms can encompass known analogs of natural nucleotides, as well as nucleotides that are modified in the base, sugar and / or phosphate moieties (e.g., phosphorothioate backbones). In general, an analog of a particular nucleotide has the same base-pairing specificity; i.e., an analog of A will base-pair with T.

[0055] The term “synthetic nucleic acid” refers to a nucleotide sequence synthesized in vitro (for example, in a lab and either manually or with a nucleic acid synthesizer device) and in which the sequence is not found in nature. The sequence may be, for example, DNA or RNA or a modification thereof asdescribed below, may be any length and may be any sequence of nucleotides so long as the sequence is not naturally occurring.

[0056] The term "nucleotide" refers to deoxyribonucleotides or ribonucleotides. The nucleotides may be standard nucleotides ( / .e., adenosine, guanosine, cytidine, thymidine, and uridine) or nucleotide analogs. A nucleotide analog refers to a nucleotide having a modified purine or pyrimidine base or a modified ribose moiety. A nucleotide analog may be a naturally occurring nucleotide (e.g., inosine) or a non-naturally occurring nucleotide. Non-limiting examples of modifications on the sugar or base moieties of a nucleotide include the addition (or removal) of acetyl groups, amino groups, carboxyl groups, carboxymethyl groups, hydroxyl groups, methyl groups, phosphoryl groups, and thiol groups, as well as the substitution of the carbon and nitrogen atoms of the bases with other atoms (e.g., 7-deaza purines). Nucleotide analogs also include dideoxy nucleotides, 2'-O-methyl nucleotides, locked nucleic acids (LNA), peptide nucleic acids (PNA), and morpholinos.

[0057] The terms "polypeptide" and "protein" are used interchangeably to refer to a polymer of amino acid residues.

[0058] The term “inactivated,” in the context of the present invention, may refer to the deletion of one or more amino acids and / or the substitution of one or more amino acids or the addition of one and / or more amino acids in a target protein such that one or more functions of the protein are eliminated or reduced to a level where activity is less than 95%, less than 90%, less than 75%, less than 50%, less than 40%, less than 30%, less than 20%, less than 15%, less than 10%, less than 5%, less than 4%, less than 3%, less than 2% or less than 1 % of the level of activity of the active protein. In the present invention, the Cas-like protein (e.g., Cas9) is catalytically inactivated as the result of having two amino acids substituted with Alanine residues as detailed elsewhere in this specification thereby inhibiting its ability to cleave doublestranded DNA. Reduced activity may be referred to “partially inactivated.” The term “catalytically modified” in the context of the present invention refers to a subject protein that is inactivated or partially inactivated and means thatone or more of the catalytic activities of the subject protein have been reduced or eliminated. The nuclease activity of the Cas protein of the present invention may be partially or fully inactivated making the Cas protein a nickase (cleaving only one DNA strand; discussed in greater detail, below) or catalytically inactive (cleaving none of the DNA strands), respectively.

[0059] The term “directly linked” with regard to proteins and polypeptides in the context of the present invention means that two proteins are joined together ( / .e., fused, e.g., via peptide bonds) to form a continuous protein or polypeptide with no added amino acid residues incorporated between the two joined / fused proteins.

[0060] The term “indirectly linked” in the context of the present invention means that one or more amino acids are incorporated between two joined / fused proteins. Additionally, two or more proteins may be indirectly linked by chemical moieties other than amino acids, the methods of which are known to one of skill in the art.

[0061] T echniques for determining nucleic acid and amino acid sequence identity are known in the art. Typically, such techniques include determining the nucleotide sequence of the mRNA for a gene and / or determining the amino acid sequence encoded thereby, and comparing these sequences to a second nucleotide or amino acid sequence. Genomic sequences can also be determined and compared in this fashion. In general, “identity” refers to an exact nucleotide-to-nucleotide or amino acid-to-amino acid correspondence of two polynucleotides or polypeptide sequences, respectively. Two or more sequences (polynucleotide or amino acid) can be compared by determining their “percent identity." The percent identity of two sequences, whether nucleic acid or amino acid sequences, is the number of exact matches between two aligned sequences divided by the length of the shorter sequences and multiplied by 100. An approximate alignment for nucleic acid sequences is provided by the local homology algorithm of Smith and Waterman, Advances in Applied Mathematics 2:482-489 (1981 ). This algorithm can be applied to amino acid sequences by using the scoring matrix developed by Dayhoff, Atlas of Protein Sequences and Structure, M. O.Dayhoff ed., 5 suppl. 3:353-358, National Biomedical Research Foundation, Washington, D.C., USA, and normalized by Gribskov, Nucl. Acids Res. 14(6):6745-6763 (1986). An exemplary implementation of this algorithm to determine percent identity of a sequence is provided by the Genetics Computer Group (Madison, Wis.) in the "BestFit" utility application. Other suitable programs for calculating the percent identity or similarity between sequences are generally known in the art. For example, another alignment program is BLAST, used with default parameters. For example, BLASTN and BLASTP can be used using the following default parameters: genetic code=standard; filter=none; strand=both; cutoff=60; expect=10; Matrix=BLOSUM62; Descriptions=50 sequences; sort by=HIGH SCORE; Databases=non-redundant, GenBank+EMBL+DDBJ+PDB+GenBank CDS translations+Swiss protein+Spupdate+PIR. Details of these programs can be found on the GenBank website.

[0062] CRISPR / Cas Proteins and Systems

[0063] Understanding the present invention will be aided by understanding CRISPR / Cas protein systems in general and in the context of the present invention.

[0064] The CRISPR / Cas9-based System

[0065] In its basic form, the CRISPR / Cas9 system introduces a double strand break close to the binding site of a guide RNA (gRNA). A Cas9 protein and a gRNA form a ribonucleoprotein (RNP) complex. This complex itself can be transfected into a recipient cell (for example, via the use of extracellular vesicles) or a plasmid or viral vector encoding the complex can be used. The gRNA targets the complex to the desired location in the genome through formation of an R-loop - a three-stranded nucleic acid structure composed of a DNA:RNA hybrid and a displaced single-stranded DNA - where the Cas9 protein (an endonuclease) causes a double-strand DNA break where modifications to the sequence can be made. Such modifications can take the form of non-homologous end joining (NHEJ) making random small insertions or deletions (indels) or homology-directed repair (HDR). NHEJ is useful for creating knockout mutations and HDR is useful for making desiredmodifications to the target sequence, also termed “precision editing.” While HDR enables the introduction of these precision edits, its use is hindered by low rates of installation of the desired edit and high rates of indels.

[0066] Base Editors

[0067] As discussed above, traditional CRISPR / Cas9 technology has challenges regarding efficiency and a high rate of undesired indels. Base editing was developed (Komor, 2016) to increase the efficiency of precision editing using the CRISPR / Cas9 technology while minimizing the rate of indels. Base editing allows for the direct, irreversible conversion of one target DNA base pair into another without requiring a dsDNA backbone break or a donor template. Rather, a target nucleobase is converted into another by a catalytic effector domain tethered to the RNP complex. For cytosine base editing, the tethered catalytic domain is a cytosine deaminase, which chemically converts a cytosine nucleobase (C) into a uracil (U). Uracil has the base-pairing properties of thymine (T). Thus, upon DNA replication or repair, the targeted C:G pair is converted to a T:A pair. For adenine base editing, the tethered catalytic domain is an adenine deaminase, which chemically converts adenine (A) into inosine (I). Inosine has the base-pairing properties of guanine (G). Thus, upon DNA replication or repair, the targeted A:T pair is converted to a G:C pair. For guanine base editing, the tethered catalytic domain is a guanine-specific glycosylase, which excises the guanine nucleobase from the DNA. Translesion synthesis at the abasic site leads to the installation of C or T, resulting in the installation of a G:C to C:G or T:A substitution.

[0068] (I) RNA-Guided Endonucleases

[0069] RNA-guided endonucleases, such as Cas9, may comprise at least one nuclear localization signal (NLS), at least one nuclease domain, and at least one domain that interacts with a guide RNA (gRNA) to direct the endonuclease / deaminase complex of the present invention to a specific nucleobase for modification. Also known are nucleic acids encoding the RNA- guided endonucleases, as well as methods of using the RNA-guided endonucleases to modify chromosomal sequences of eukaryotic cells or embryos. The RNA-guided endonuclease interacts with specific gRNAs, eachof which directs the endonuclease to a specific targeted site, at which site the catalytic effector domain modifies the target nucleobase, resulting in the installation of a substitution. Since the specificity is provided by the gRNA, the RNA-based endonuclease is, essentially, universal and can be used with different gRNAs to target different genomic sequences.

[0070] Many forms of guide RNAs (gRNAs) are known in the field. In general, a gRNA is an RNA molecule that can direct an RNA binding protein or endonuclease to a specific nucleic acid sequence / target site by means of base pairing. A gRNA may comprise a single RNA molecule, such as a crRNA, or two RNA molecules, such as a crRNA and a tracrRNA (or, trRNA). In some embodiments, it may additionally comprise an accessory RNA or DNA molecule. In some embodiments, a crRNA and a tracrRNA may be covalently linked together to form a chimeric guide RNA or a single guide RNA (sgRNA). It is also well established in the field that a gRNA can be introduced alone or in combination with its cognate gRNA binding protein or endonuclease to a target DNA in different forms. It can be expressed from a DNA vector introduced into a target cell. It can be synthesized by in vitro transcription or by chemical reaction before being introduced to a target DNA. A person of skill in the art should know that many chemical modifications are possible on a gRNA by chemical synthesis. For example, certain modifications, such as 2’-0 methyl group modification and phosphorothioate linkage modification, may be introduced into a gRNA to alter its stability or performance. Nonstandard nucleotides, such as DNA and LNA, also may be introduced into a gRNA during chemical synthesis to alter its specificity or performance. A guide RNA that may have a different form or may contain a different chemical modification may be used in conjunction with the present disclosure without deviating from the spirit of the present disclosure.

[0071] The present disclosure provides for fusion proteins, wherein a fusion protein comprises a CRISPR / Cas-like protein or fragment thereof wherein the Cas-like protein is fully or partially catalytically inactivated, retaining its ability to bind DNA but without being able to generate double-strand breaks in the target DNA. Each fusion protein is guided to a specific target DNA sequenceby a specific gRNA, wherein the associated (e.g., tethered or otherwise linked) catalytic effector domain converts the target nucleobase to another nucleobase.

[0072] RNA-guided endonucleases may comprise at least one nuclear localization signal (NLS), which permits entry of the endonuclease into the nuclei of eukaryotic cells and embryos such as, for example, non-human one cell embryos. RNA-guided endonucleases also comprise at least one nuclease domain and at least one domain that interacts with a gRNA. An RNA-guided endonuclease is directed to a specific nucleic acid sequence (or target site) by a gRNA. The gRNA interacts with the RNA-guided endonuclease as well as the target site such that, once directed to the target site, the associated catalytic effector domain can convert the target nucleobase to another nucleobase, thereby installing a substitution. Since the gRNA provides the specificity for the targeted cleavage the RNA-guided endonuclease is universal (providing that its ability to cleave DNA is disabled) and can be used with different gRNAs to modify different target residues. RNA-guided endonucleases can be proteins, can be encoded by isolated nucleic acids ( / .e., RNA or DNA), can be encoded by vectors comprising nucleic acids encoding the RNA-guided endonucleases, and can be protein- RNA complexes comprising the RNA-guided endonuclease plus a gRNA.

[0073] The RNA-guided endonuclease can be derived from a clustered regularly interspersed short palindromic repeats (CRISPR)ZCRISPR- associated (Cas) system. The CRISPR / Cas system can be a type I, a type II, or a type III system, as known to one of ordinary skill in the art. Non-limiting examples of suitable CRISPR / Cas proteins include Cas3, Cas4, Cas5, Cas5e (or CasD), Cas6, Cas6e, Cas6f, Cas7, Cas8a1 , Cas8a2, Cas8b, Cas8c, Cas9, Cas10, Cas10d, Cas12, CasF, CasG, CasH, Csy1 , Csy2, Csy3, Cse1 (or CasA), Cse2 (or CasB), Cse3 (or CasE), Cse4 (or CasC), Csc1 , Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1 , Cmr3, Cmr4, Cmr5, Cmr6, Csb1 , Csb2, Csb3,Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csz1 , Csx15, Csf1 , Csf2, Csf3, Csf4, and Cu1966. One of ordinary skill in the artwill be able to modify any RNA-guided nuclease to inactivate its catalytic activity for use with a base editor system.

[0074] In one embodiment, the RNA-guided endonuclease is derived from a type II CRISPR / Cas system. In specific embodiments, the RNA-guided endonuclease is derived from a Cas9 protein. The Cas9 protein can be from, for example, Streptococcus pyogenes, Streptococcus thermophilus, Streptococcus sp., Nocardiopsis dassonvillei, Streptomyces pristinaespiralis, Streptomyces viridochromogenes, Streptomyces viridochromogenes, Streptosporangium roseum, Streptosporangium roseum, Alicyclobacillus acidocaldarius, Bacillus pseudomycoides, Bacillus selenitireducens, Exiguobacterium sibiricum, Lactobacillus delbrueckii, Lactobacillus salivarius, Microscilla marina, Burkholderiales bacterium, Polaromonas naphthalenivorans, Polaromonas sp., Crocosphaera watsonii, Cyanothece sp., Microcystis aeruginosa, Synechococcus sp., Acetohalobium arabaticum, Ammonifex degensii, Caldicelulosiruptor becscii, Candidatus Desulforudis, Clostridium botulinum, Clostridium difficile, Finegoldia magna, Natranaerobius thermophilus, Pelotomaculum thermopropionicum, Acidithiobacillus caldus, Acidithiobacillus ferrooxidans, Allochromatium vinosum, Marinobacter sp., Nitrosococcus halophilus, Nitrosococcus watsoni, Pseudoalteromonas haloplanktis, Ktedonobacter racemifer, Methanohalobium evestigatum, Anabaena variabilis, Nodularia spumigena, Nostoc sp., Arthrospira maxima, Arthrospira platensis, Arthrospira sp., Lyngbya sp., Microcoleus chthonoplastes, Oscillatoria sp., Petrotoga mobilis, Thermosipho africanus, or Acaryochloris marina.

[0075] In general, CRISPR / Cas proteins comprise at least one RNA recognition domain and / or RNA binding domain. RNA recognition and / or RNA binding domains interact with guide RNAs. CRISPR / Cas proteins also usually comprise nuclease domains i.e., DNase or RNase domains; but for the purposes of base editors one or more of these are disabled), DNA binding domains, helicase domains, RNase domains, protein-protein interaction domains, dimerization domains, as well as other domains.

[0076] The CRISPR / Cas-like protein can be a wild type CRISPR / Cas protein, a modified CRISPR / Cas protein, or a fragment of a wild type or modified CRISPR / Cas protein. The CRISPR / Cas-like protein can be modified to increase nucleic acid binding affinity and / or specificity, alter an enzymatic activity (e.g., inactivate its ability to cleave DNA), and / or change another property of the protein. For example, nuclease ( / .e., DNase, RNase) domains of the CRISPR / Cas-like protein can be modified, deleted, or inactivated. Alternatively, the CRISPR / Cas-like protein can be truncated to remove domains that are not essential for the function of the fusion protein. The CRISPR / Cas-like protein can also be truncated or modified to optimize the activity of the effector domain of the fusion protein.

[0077] In some embodiments, the CRISPR / Cas-like protein can be derived from a wild type Cas9 protein or fragment thereof. In other embodiments, the CRISPR / Cas-like protein can be derived from a modified Cas9 protein. For example, the amino acid sequence of the Cas9 protein can be modified to alter one or more properties (e.g., nuclease activity, affinity, stability, etc.) of the protein. Alternatively, domains of the Cas9 protein not involved in RNA- guided cleavage can be eliminated from the protein such that the modified Cas9 protein is smaller than the wild type Cas9 protein.

[0078] In general, a Cas9 protein comprises at least two nuclease ( / .e., DNase) domains. For example, a Cas9 protein can comprise a RuvC-like nuclease domain and an HNH-like nuclease domain. The RuvC and HNH domains work together to cut single strands to make a double-stranded break in DNA (Jinek, et al., Science, 337: 816-821). In some embodiments, the Cas9-derived protein can be modified to contain only one functional nuclease domain (either a RuvC-like or an HNH-like nuclease domain). For example, the Cas9-derived protein can be modified such that one of the nuclease domains is deleted or mutated such that it is no longer functional ( / .e., the nuclease activity is absent). In some embodiments in which one of the nuclease domains is inactive, the Cas9-derived protein is able to introduce a nick into a double-stranded nucleic acid (such protein is termed a "nickase"), but not cleave both strands of the double-stranded DNA. For example, anaspartate to alanine (D10A) conversion in a RuvC-like domain converts the Cas9-derived protein into a nickase. Likewise, a histidine to alanine (H840A or H839A) conversion in a HNH domain converts the Cas9-derived protein into a nickase. Each nuclease domain can be modified using well-known methods, such as site-directed mutagenesis, PCR-mediated mutagenesis, and total gene synthesis, as well as other methods known in the art. Modifying both of these domains results in inactivation of nuclease and nickase activity.

[0079] The RNA-guided endonuclease may comprise at least one NLS. In general, an NLS comprises a stretch of basic amino acids. Nuclear localization signals are known in the art (see, e.g., Lange, et al., J. Biol. Chem., 2007, 282:5101 -5105). For example, in one embodiment, the NLS can be a monopartite sequence, such as PKKKRKV (SEQ ID NO: 10) or PKKKRRV (SEQ ID NO: 11 ). In another embodiment, the NLS can be a bipartite sequence. In still another embodiment, the NLS can be KRPAATKKAGQAKKKK (SEQ ID NO: 12). The NLS can be located at the N- terminus, the C-terminus, or in an internal location of the RNA-guided endonuclease.

[0080] In some embodiments, the RNA-guided endonuclease can further comprise at least one cell-penetrating domain. In one embodiment, the cellpenetrating domain can be a cell-penetrating peptide sequence derived from the HIV-1 TAT protein. As an example, the TAT cell-penetrating sequence can be GRKKRRQRRRPPQPKKKRKV (SEQ ID NO: 13). In another embodiment, the cell-penetrating domain can be TLM (PLSSIFSRIGDPPKKKRKV; SEQ ID NO: 14), a cell-penetrating peptide sequence derived from the human hepatitis B virus. In still another embodiment, the cell-penetrating domain can be MPG (GALFLGWLGAAGSTMGAPKKKRKV; SEQ ID NO: 15 or GALFLGFLGAAGSTMGAWSQPKKKRKV; SEQ ID NO: 16). In an additional embodiment, the cell-penetrating domain can be Pep-1 (KETWWETWWTEWSQPKKKRKV; SEQ ID NO: 17), VP22, a cell penetrating peptide from Herpes simplex virus, or a polyarginine peptidesequence. The cell-penetrating domain can be located at the N-terminus, the C-terminus, or in an internal location of the protein.

[0081] In still other embodiments, the RNA-guided endonuclease can also comprise at least one marker domain. Non-limiting examples of marker domains include fluorescent proteins, purification tags, and epitope tags. In some embodiments, the marker domain can be a fluorescent protein. Non limiting examples of suitable fluorescent proteins include green fluorescent proteins (e.g., GFP, GFP-2, tagGFP, turboGFP, EGFP, Emerald, Azami Green, Monomeric Azami Green, CopGFP, AceGFP, ZsGreenl ), yellow fluorescent proteins (e.g., YFP, EYFP, Citrine, Venus, YPet, PhiYFP, ZsYellowl ,), blue fluorescent proteins (e.g., EBFP, EBFP2, Azurite, mKalamal , GFPuv, Sapphire, T-sapphire,), cyan fluorescent proteins (e.g., ECFP, Cerulean, CyPet, AmCyanl , Midoriishi-Cyan), red fluorescent proteins (mKate, mKate2, mPlum, DsRed monomer, mCherry, mRFP1 , DsRed- Express, DsRed2, DsRed-Monomer, HcRed-Tandem, HcRedl , AsRed2, eqFP611 , mRasberry, mStrawberry, Jred), and orange fluorescent proteins (mOrange, mKO, Kusabira-Orange, Monomeric Kusabira-Orange, mTangerine, tdTomato) or any other suitable fluorescent protein. In other embodiments, the marker domain can be a purification tag and / or an epitope tag. Exemplary tags include, but are not limited to, glutathione-S-transferase (GST), chitin binding protein (CBP), maltose binding protein, thioredoxin (TRX), poly(NANP), tandem affinity purification (TAP) tag, myc, AcV5, AU1 , AU5, E, ECS, E2, FLAG, HA, nus, Softag 1 , Softag 3, Strep, SBP, Glu-Glu, HSV, KT3, S, S1 , T7, V5, VSV-G, 6xHis (SEQ ID NO: 18), biotin carboxyl carrier protein (BCCP), and calmodulin.

[0082] In certain embodiments, the RNA-guided endonuclease may be part of a protein-RNA complex comprising a gRNA. The gRNA interacts with the RNA-guided endonuclease to direct the endonuclease to a specific target site, wherein, for example, the 5' end of the guide RNA base pairs with a specific protospacer sequence.

[0083] (II) Fusion Proteins

[0084] Another aspect of the present disclosure provides a fusion protein comprising a CRISPR / Cas-like protein, or fragment thereof, and a catalytic effector domain. The CRISPR / Cas-like protein is directed to a target site by a gRNA, at which site the effector modifies or affects the targeted nucleic acid sequence. In the present invention the “effector domain” can be a linked deaminase or glycosylase. The fusion protein can further comprise at least one additional domain chosen from a nuclear localization signal, a cellpenetrating domain, or a marker domain.

[0085] (a) CRISPR / Cas-Like Protein

[0086] The fusion protein comprises a CRISPR / Cas-like protein or a fragment thereof. CRISPR / Cas-like proteins are detailed above in section (I). The CRISPR / Cas-like protein can be located at the N-terminus, the C-terminus, or in an internal location of the fusion protein.

[0087] In some embodiments, the CRISPR / Cas-like protein of the fusion protein can be derived from a Cas9 protein. Cas9-derived proteins can be wild type, modified, or a fragment thereof. In some embodiments of the present invention, the Cas9-derived protein is modified to inactivate one or both functional nuclease domains (either a RuvC-like or an HNH-like nuclease domain). In some embodiments, both of the RuvC-like nuclease domain and the HNH-like nuclease domain can be modified or eliminated such that the Cas9-derived protein is unable to nick or cleave double stranded nucleic acid. In still other embodiments, all nuclease domains of the Cas9-derived protein can be modified or eliminated such that the Cas9-derived protein lacks all nuclease activity. Also, either the RuvC-like domain or the HNH-like domain can be inactivated independently of each other.

[0088] In any of the above-described embodiments, any or all of the nuclease domains can be inactivated by one or more deletion mutations, insertion mutations, and / or substitution mutations using well-known methods, such as site-directed mutagenesis, PCR-mediated mutagenesis, and total gene synthesis, as well as other methods known in the art.

[0089] (b) Target Site

[0090] An RNA-guided endonuclease in conjunction with a gRNA is directed to a target site in a dsDNA, wherein the RNA-guided endonuclease (enzymatically inactivated) binds the chromosomal sequence. The target site has no sequence limitation except that the sequence is immediately adjacent to a consensus sequence. This consensus sequence is also known as a protospacer adjacent motif (PAM). Examples of PAM include, but are not limited to, NGG (SEQ ID NO: 19), NGGNG (SEQ ID NO: 20), and NNAGAAW (SEQ ID NO: 21) (wherein N is defined as any nucleotide and W is defined as either A or T). As detailed above, the spacer region of the gRNA is complementary to the protospacer of the target sequence. Typically, the spacer region of the gRNA is about 19 to 21 nucleotides in length. Thus, in certain aspects, the sequence of the target site in the chromosomal sequence is 5'-Ni9-2i-NGG-3'.

[0091] The target site can be in the coding region of a gene, in an intron of a gene, in a control region of a gene, in a non-coding region between genes, etc. The gene can be a protein coding gene or an RNA coding gene. The gene can be any gene of interest. The target site can also be located on a plasmid or fragment of synthetic or amplified DNA.

[0092] (III) Kits

[0093] Still another aspect of the present disclosure provides kits for carrying out the methods described above.

[0094] The kits provided herein generally include instructions for carrying out the processes detailed above. Instructions included in the kits may be affixed to packaging material, may be included as a package insert or as a downloadable file. While the instructions are typically written or printed materials, they are not limited to such. Any medium capable of storing such instructions and communicating them to an end user is contemplated by this disclosure. Such media include, but are not limited to, electronic storage media (e.g., magnetic discs, tapes, cartridges, chips), optical media (e.g., CD ROM), and the like. As used herein, the term “instructions” can include the address of an internet site that provides the instructions.

[0095] As various changes could be made in the above-described processes and kits without departing from the scope of the invention, it is intended that all matter contained in the above description and in the examples given below, shall be interpreted as illustrative and not in a limiting sense.

[0096] Assay Components

[0097] (I) Assay Substrate

[0098] The present invention utilizes a double-stranded DNA (dsDNA) substrate. This substrate must contain the minimum sequences for base editing: a target sequence and a protospacer-adjacent motif (PAM). The target sequence can be any sequence, provided that it is complementary to the guide RNA used to target it and contains at least one editable residue. For cytosine base editing, the editable residue is cytosine; for adenine base editing, the editable residue is adenine; for guanine base editing, the editable residue is guanine. The dsDNA may be 20-30 nt in length, or 30-40 nt in length, or longer. The target sequence and PAM should be placed such that cleavage at the editing site results in diffusion of one or more cleavage products from the substrate.

[0099] The present invention utilizes fluorescent labels (fluorophores) on the dsDNA substrate. These labels may be attached covalently or noncovalently. A fluorophore is a molecule that, when exposed to light of a specific wavelength or range of wavelengths, emits light. This emitted light can be measured by instruments including plate readers and qPCR instruments. Examples of fluorescent moieties suitable for the present invention include, but are not limited to, Cy3, Cy5, 6-FAM (fluorescein), HEX, JOE, Texas Red, and VIC.

[0100] The present invention also utilizes quencher domains. A quencher domain, when in close spatial proximity to a fluorescent moiety, reduces or eliminates the amount of fluorescence emitted from the fluorescent moiety, i.e., reduces the intensity of the fluorescent emission. Examples of quencher domains suitable for the present invention include, but are not limited to: Deep Dark Quencher I & Deep Dark Quencher II (Eurgentec, Seraing, Belfium), Dabcyl, Eclipse (Epoch Biosciences, Bothell, WA), Iowa Black FQ & IowaBlack RQ (Integrated DNA Technologies, Coralville, Iowa), Black Hole Quencher 1 , Black Hole Quencher 2 & Black Hole Quencher 3 (Biosearch Technologies, Petaluma, CA), QSY-7 & QSY- 21 (Invitrogen, Waltham, MA) and Blackberry Quencher 650 (Berry and Associates, Petaluma, CA). See, e.g., Marras (Marras, S.A.E. Interactive Fluorophore and Quencher Pairs for Labeling Fluorescent Nucleic Acid Hybridization Probes. Mol Biotechnol 38, 247-255 (2008) the contents of which are representative of what one of skill in the art knew at the time of the present invention. The selected quencher must absorb light of the wavelength emitted by the fluorophore used.

[0101] The fluorophore and quencher should be placed such that, in the absence of base editor binding or activity, the light emitted by the fluorophore is meaningfully reduced by the quencher. The measured fluorescence may be reduced by at least 50%, at least 75%, or at least 90%. The fluorophore and quencher should also be placed such that cleavage at the site of base editing, or the activity of a nonspecific nuclease, results in separation of the fluorophore and quencher. The fluorophore and quencher may both be placed on the edited strand, on opposite sides of the editable residue. The fluorophore may be placed at the 5’ end of the edited strand with the quencher at the 3’ end of the unedited strand. The quencher may be placed at the 5’ end of the edited strand with the fluorophore at the 3’ end of the unedited strand.

[0102] The dsDNA substrate may be further modified to include alternative nucelobases or other modifications. These modifications should not include modifications that would interfere with cleavage of the backbone, as these modifications would prevent the intended cleavage at the editing site necessary to produce fluorescence and / or prevent the identification of nonspecific nuclease activity.

[0103] (II) Guide RNAs

[0104] As detailed above, a guide RNA is required to direct a base editor to its target sequence. The present invention utilizes two separate guide RNAs. A first guide RNA is complementary to the target sequence in the dsDNA substrate and directs the base editor to the substrate. A second guide RNAdoes not exhibit meaningful complementarity to the target sequence in the dsDNA substrate and is incapable of directing the base editor to the substrate. This second guide should have less than 50% identity to the first, on-target guide. Preferably, it should have at least two mismatches in the seed region of the guide.

[0105] (III) Assay setup

[0106] In the present invention, ribonucleoprotein complexes (RNPs) are assembled separately for the two guides. These separate RNPs are incubated with the dsDNA substrate in separate wells. Fluorescence resulting from the first, on-target RNP is the result of the combined activity of the base editor and any nonspecific nucleases present. Fluorescence resulting from the second, nontargeting RNP is due solely to the activity of said nonspecific nucleases. The difference in fluorescence between the first, on-target RNP and the second, nontargeting RNP yields the fluorescence due to base editing alone.

[0107] (IV) Data normalization

[0108] The present invention provides two independent methods for data normalization. The first method utilizes a free fluorophore in the assay mixture. Fluorescence of the free fluorophore is unchanged by the activity of the base editor, so it may be used to normalize fluorescence of the cleaved dsDNA substrate. This reduces well-to-well variation due to differences in the fluorescence detector. A second normalization method requires fully melting the dsDNA substrate using elevated temperature. When the fluorophore and quencher are on opposite strands and the dsDNA substrate is fully melted, the distance between fluorophore and quencher becomes too great for quenching to occur, and maximum fluorescence is observed. The measured fluorescence due to base editing can be expressed as a percentage of the maximum fluorescence due to melting.Exemplification

[0109] Example 1 - In vitro base editor assay predicts the activity of base editor RNPs in cells

[0110] Cytosine base editor (CBE) proteins were constructed by fusing the C- terminal domain of human APOBEC3B to the amino-terminus of an SpCas9 nickase protein (SEQ ID NO 1 , also referred to herein as Protein A). A variant of CBE was constructed by fusing the single-stranded DNA binding protein (SSB) domain from T4 bacteriophage to the amino terminus of the CBE described above (SEQ ID NO 2, also referred to herein as Protein B). All proteins were expressed and purified from E. coli BL21 Al by autoinduction, nickel column chromatography, and cation exchange chromatography and were stored at -80 °C in a buffer containing 10% glycerol, 300 mM KCI, 20 mM HEPES (pH 7.5), and 1 mM DTT, before use. Synthetic single guide RNAs (sgRNAs) were purchased from Integrated DNA Technologies (IDT; Coralville, IA). The spacer sequences of the sgRNA are listed in Table 1 .Table 1. sgRNA spacer sequences and PCR primer sequences

[0111] DNA oligonucleotides for the in vitro assay were purchased from MilliporeSigma (Non-targeted strand, NTS) and IDT (Targeted strand, TS). The NTS oligo was tagged at the 5’ end with a Texas Red fluorophore moiety. The TS oligo was tagged at the 3’ end with an Iowa Black RQ quencher moiety. In vitro assay oligo sequences are listed in Table 1 . Oligos wereannealed in reactions containing 5 pM NTS oligo, 5.5 pM TS oligo, 10 mM Tris, pH 8, 50 mM NaCI, and 1 mM EDTA in water. Reactions were annealed by incubation at 90°C for 5m, 75°C for 10m, and cooled to 20°C at a rate of 0.1 °C / s. Annealed oligos were diluted to 2 pM.

[0112] Ribonucleoprotein (RNP) complexes for use in the in vitro assay were prepared by incubating for 15 min at room temperature: 200 pmol of CBE protein and 200 pmol sgRNA in buffer (20 mM HEPES, 100 mM KCI, 0.5 mM DTT, 0.1 mM EDTA, pH 7.5) for a final volume of 20 pL. RNPs were prepared using the HEKSite2 sgRNA to target the annealed oligo substrate and, separately, using the HBB sgRNA as a nontargeting control. RNPs were kept on ice until use for transfection or in the in vitro assay.

[0113] In vitro assays were set up on ice and performed in 96-well white PCR plates. Reactions contained 6 pmol RNP, 2 pmol annealed oligo substrate, 1 Unit of USER enzyme (NEB), and 200 fmol fluorescein in 20 mM Tris, pH 8, 100 mM NaCI, and 5 mM MgCI2 in 20 pL reactions. Samples were tested in triplicate wells. Control wells included a buffer-only well (no RNP, fluorescein, or annealed oligo substrate) and duplicate No-RNP wells. Data was collected on a BioRad CFX Opus 96 instrument (BioRad; Hercules, California, USA), reading all fluorescence channels, at 37°C. A total of 100 data points was collected for each sample, with collection after every 5s of incubation. A melt curve was then performed with data collected every 1 °C, up to 95°C.

[0114] Data was analyzed by subtracting the fluorescence of the buffer control well from each sample at each time point or temperature. The fluorescence values were plotted with respect to time for the 37°C incubation and a time point selected for analysis at which the increase in fluorescence was approximately linear. The fluorescence values were plotted with respect to temperature for the melt curve and a temperature selected for analysis at which the fluorescence reached a plateau, indicating full melting of the oligo substrate. The fluorescence of each well at the selected time was divided by the fluorescence of that well at the selected temperature to calculate the % of maximum fluorescence released. Technical replicates were averaged and error calculated as standard deviation. The No-RNP control wells serve as anindication of background fluorescence due to the annealed oligo substrate. CBE-specific fluorescence was calculated by subtracting the % of maximum fluorescence using the non-targeting HBB sgRNA from the % of maximum fluorescence using the on-target HEKSite2 sgRNA and error calculated through arithmetic propagation of error.

[0115] RNP complexes for nucleofection were prepared by incubating for 15 min at room temperature: 40 pmol of CBE protein and 120 pmol EMX1-15 sgRNA in buffer (20 mM HEPES, 100 mM KCI, 0.5 mM DTT, 0.1 mM EDTA, pH 7.5) for a final volume of 10 pL. RNPs were kept on ice until use for transfection or in the in vitro assay.

[0116] HEK293 cells were obtained from ATCC and grown at 37 °C and 5% CO2 in DMEM supplemented with 10% FBS, 2 mM L-glutamine, 1 mM sodium pyruvate, and 0.1 mM non-essential amino acids. Cells were seeded at 1 .67 x 104cells / cm2 of tissue culture surface area two days before transfection. At the time of transfection, cells were trypsinized to obtain a single-cell suspension, washed twice with Hank’s Balanced Salt Solution and resuspended in Nucleofector Solution V (Lonza, Cambridge, MA) at 2.45 x 105cells per 100 pL. Nucleofection was performed by mixing 100 pL of prepared cell suspension with 10 pL of complexed CBE RNP by pipetting up and down six times before transferring to a cuvette for electroporation using program Q-001 on a Nucleofector 2b machine. Nucleofected cells were immediately transferred to 6-well plates containing 2 mL pre-warmed media per well and grown for 3 days before harvest. Each experimental condition was tested in two technical replicates.

[0117] Genomic DNA was harvested from transfected cells by obtaining a single cell suspension using Accutase™ (Stemcell Technologies, Cambridge, MA), pelleting the cells, and resuspending the pellet in 70 pL QuickExtract reagent (Lucigen; Middleton, Wisconsin, USA). Suspensions were incubated at 60 °C for 15m and 95 °C for 15m. The EMX1 -15 genomic site targeted by the CBE was amplified by PCR using JumpStart Taq ReadyMix (MilliporeSigma, Burlington, MA) and the following cycling conditions: 94 °C / 2m; 25 cycles of 94 °C / 30s, 62 C / 30s, 72 C / 45s; 72 °C / 2m. Primers arelisted in Table 1. PCR products underwent a second round of amplification using Illumina index primers and JumpStart Taq ReadyMix and the following conditions: 95 °C / 3m; 9 cycles of 95 °C / 30s, 55 ‘C / 30s, 72 °C / 30s; 72 ”C / 5m. Indexed PCR products were quantified by PicoGreen (ThermoFisher, Waltham, MA), pooled according to DNA content, and purified by Select-a- Size DNA Clean & Concentrator MagBeads (Zymo; Irvine, California, USA) using 1 ,2x beads by volume. Pools were diluted to 4 nM. Sequencing was performed on an Illumina MiSeq instrument using a 300-cycle kit to obtain single-end reads. FASTQ files for each sample were analyzed using the CRIS.py software package (Connelly & Pruett-Miller, Sci. Rep., 2019).

[0118] Results are presented in Figure 2. In Panel A, the % of maximum fluorescence achieved for each of nine to ten independent preparations of two CBE variants is plotted, for both the on-target HEKSite2 sgRNA (black bars) and the nontargeting HBB sgRNA control (gray bars). Values are average ± standard deviation for three replicates. All samples of a CBE variant were run at the same time, on a single plate; the two protein variants were run on successive days. The data for the on-target HEKSite2 sgRNA demonstrates the range of activities of different preps of the same protein, and the range of activities of different CBE variants. The data for the non-targeting HBB sgRNA demonstrates that most preps did not have substantial contaminating nonspecific DNase activity, with a few exceptions (for example, Prep 2 of Protein A).

[0119] In Panel B, the CBE-specific % of maximum fluorescence achieved for each of nine to ten independent preparations of two CBE variants is plotted, using the data from Panel A. This data again demonstrates the range of activities across preps of a single CBE variant, as well as between variants. It further demonstrates the impact of controlling for nonspecific DNase activity in the assay design. Compare, for example, the relative heights of the black bars in Panel A for Preps 2 and 3 of Protein A with the same preps in Panel B.

[0120] In Panel C, the rate of editing in cells for each protein prep is plotted against the in vitro activity assay results for the preps presented in Panels A and B. RNPs for nucleofection were prepared alongside RNPs for the in vitroassay from the same aliquots of protein and were delivered to cells the same day as the in vitro assay was performed. Preps of Protein A are indicated as gray circles (o) while preps of Protein B are indicated as black circles (•). The data demonstrates that there is strong correlation between the results for the in vitro activity assay and in-cell editing, even across multiple protein variants.

Claims

Claims\Ne Claim:1 . A method for measuring the activity of a base editor in vitro, the method comprising: a. providing: i) a double stranded DNA (dsDNA) oligonucleotide comprising a target nucleotide or nucleotides for editing, a protospacer adjacent motif (PAM), a fluorescent moiety, and a quencher moiety, wherein the proximity between the fluorophore and quencher prevents or reduces the emission of fluorescence, ii) a first base editor ribonucleoprotein (RNP) comprising a site-specific endonuclease fused to a deaminase and / or glycosylase domain and a guide RNA (gRNA) complementary to the target dsDNA, and iii) a second base editor RNP comprising a site-specific endonuclease fused to a deaminase and / or glycosylase domain and a guide RNA (gRNA) not complementary to the target dsDNA; b. contacting the dsDNA with the first base editor RNP to edit the dsDNA target nucleotide or nucleotides; c. wherein, upon base editing, the enzymatic breakage of the dsDNA at the site of one or more editable nucleotide residues results in a physical separation of fluorophore and quencher, resulting in a fluorescent signal increased over background fluorescence; d. separately, contacting the dsDNA with the second base editor RNP; and e. wherein any fluorescent signal elicited by the second base editor RNP above background fluorescence is not due to activity of the base editor but due to nonspecific nuclease activity.

2. The method of Claim 1 , wherein the fluorescent signal is quantified by an instrument capable of measuring fluorescence.

3. The method of Claim 2, wherein the nonspecific normalized fluorescence achieved by contacting the dsDNA with the second RNP is subtracted from the normalized fluorescence achieved by contacting the dsDNA with the first RNP targeted to the dsDNA to calculate the activity specific to base editor activity.

4. The method of Claim 1 , wherein the fluorescent moiety and the quencher moiety are on the same end of the dsDNA, on opposite strands.

5. The method of Claim 1 , wherein both the fluorescent moiety and the quencher moiety are on different ends of the same strand of the dsDNA.

6. The method of any one of Claims 1 - 5, wherein the site-specific endonuclease is a Cas domain.

7. The method of Claim 6, wherein the site-specific endonuclease is a Cas9 domain.

8. The method of Claim 6, wherein the site-specific endonuclease is a Cas12 domain.

9. The method of any one of Claims 6 - 8, wherein the Cas domain has been altered to abrogate its endonuclease activity on one DNA strand.

10. The method of any one of Claims 6 - 8, wherein the Cas domain has been altered to abrogate its endonuclease activity on both strands of DNA.11 . The method of any one of Claims 1 - 10, wherein the guide RNA is a synthetic single guide RNA (sgRNA).

12. The method of any one of Claims 1 - 10, wherein the guide RNA is an in vitro transcribed sgRNA.

13. The method of any one of Claims 1 - 10, wherein the guide RNA is a crRNA complexed with a tracrRNA.

14. The method of any one of Claims 1 - 13 wherein the activity is normalized using a free fluorescent molecule with a different spectrum from the target DNA-bound fluorophore that is not affected by the base editor RNP.

15. The method of any one of Claims 1 - 13, wherein the activity is normalized by raising the temperature of the mixture, thereby fully separating the fluorophore and quencher.

16. The method of any one of Claims 2 - 15, wherein the instrument capable of measuring fluorescence is a plate reader.

17. The method of any one of Claims 2 - 15, wherein the instrument capable of measuring fluorescence is a qPCR machine.

18. The method of any one of Claims 1 - 17, wherein the base editor is a cytosine base editor (CBE).

19. The method of Claim 18, wherein the CBE is a fusion protein comprised of a site-specific endonuclease and a deaminase domain capable of catalyzing the conversion of cytosine to uracil.

20. The method of any one of Claims 18 - 19, wherein the nucleotide residue capable of being edited by the CBE is cytosine.21 . The method of any one of Claims 18 - 20, wherein the enzymatic activities that result in DNA strand breakage at the site of a successful edit are comprised of uracil DNA glycosylase and AP lyase activities.

22. The method of any one of Claims 18 - 21 , wherein the enzymatic activities that result in DNA strand breakage at the site of a successful edit are contained in separate reagents.

23. The method of any one of Claims 18 - 21 , wherein the enzymatic activities that result in DNA strand breakage at the site of a successful edit are contained in a single reagent.

24. The method of any one of Claims 1 - 17, wherein the BE is an adenine base editor (ABE).

25. The method of Claim 24, wherein the ABE is a fusion protein comprised of a site-specific endonuclease and a deaminase domain capable of catalyzing the conversion of adenine to inosine.

26. The method of any one of Claims 24 - 25, wherein the nucleotide residue capable of being edited by the ABE is adenine.

27. The method of claims 24 - 26, wherein the enzymatic activity that results in DNA strand breakage at the site of a successful edit is performed by endonuclease V.

28. The method of any one of Claims 1 - 18, wherein the base editor is a glycosylase-based guanine base editor (gGBE).

29. The method of Claim 28, wherein the gGBE is a fusion protein comprised of a site-specific endonuclease and a glycosylase domain capable of catalyzing the excision of guanine from DNA.

30. The method of any one of Claims 28 - 29, wherein the nucleotide residue capable of being edited by the gGBE is guanine.The method of any one of Claims 28 - 30, wherein the enzymatic activity that results in DNA strand breakage at the site of a successful edit is comprised of an AP lyase activity.

Citation Information

Patent Citations

  • Method for identifying DNA base editing by means of cytosine deaminase

    EP3530737A1