Crispr repair in vitro screen (CRIS) off-target assessment
The method employs DNA binding protein-targeted enzymatic digestion and ANN analysis to accurately identify CRISPR-Cas off-target sites, addressing the limitations of current methods by enhancing sensitivity and efficiency in off-target detection.
Patent Information
- Application Number
- PCT/US2025/044171
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-29
- Filing Date
- 2025-08-29
- Publication Date
- 2026-03-05
AI Technical Summary
Current methods for identifying and reducing off-target effects in CRISPR-Cas gene editing are inadequate, particularly due to variations in DNA accessibility and the complexity of predicting off-target sites, leading to unwanted insertions and deletions.
A method involving DNA binding protein-targeted enzymatic digestion, followed by nucleic acid library generation and analysis using an artificial neural network (ANN) to identify on- and off-target cleavage sites, utilizing techniques like CUT&RUN and CUT&Tag, and computational workflows to enhance accuracy and sensitivity.
Enables robust identification of off-target sites from as few as 5,000 cells with high sensitivity and adaptability to diverse experimental methods, providing significant and detectable results for CRISPR-based repair analysis.
Smart Images

Figure US2025044171_05032026_PF_FP_ABST
Abstract
Description
[0001] 00B206.1663
[0002] CRISPR REPAIR IN VITRO SCREEN (CRIS) OFF-TARGET ASSESSMENT
[0003] PRIORITY
[0004] This application claims priority to U.S. Provisional Application No. 63 / 688,801, filed August 29, 2024, which is hereby incorporated by reference in its entirety.
[0005] FIELD OF THE INVENTION
[0006] The present disclosure is directed to materials and methods for CRISPR-based repair off-target analysis and computer analytics relating to the same.
[0007] REFERENCE TO A SEQUENCE LISTING
[0008] The instant application contains a Sequence Listing which has been submitted electronically and is hereby incorporated by reference in its entirety. Said sequence listing copy, created on August 28, 2025, is named P38843_SL.xml and is 24,576 bytes in size.
[0009] BACKGROUND
[0010] Gene editing was revolutionized by the clustered regularly interspaced short palindromic repeats (CRISPR)-Cas system through its more precise form of genetic manipulation. As the technology has progressed, however, an understanding has emerged that off-target effects are occurring when using the system, resulting in unexpected, unwanted, and even adverse alterations to the CRISPR-edited genome. Researchers have used various in silico prediction methods to identify potential off-target sites to improve the selection of sgRNA sequences, as the propensity to edit an off-target site is generally considered to be sgRNA- dependent. Such off-target prediction software can be classified into one of two groups based on their data output format. One group produces data describing the level of sgRNA alignment to the putative off-target sites in the genome. This first group of software includes CasOT, Cas-OFFinder, FlashFry, and Crisflash. The second group of in silico tools available today harness more complicated scoring models to assess computational nomination of the off-target sites. Available software includes MIT score, CCTop (Consensus Constrained TOPology prediction), CROP-IT (CRISPR / Cas9 off-target prediction and identification tool), CFD (cutting frequency determination), DeepCRISPR, and the Elevation software packages. Guo et al., “Off-target effects in CRISPR / Cas9 gene editing,” Front. Bioeng. Biotechol. Vol. 11, 2023 (doi.org / 10.3389 / fbioe.2023.1143157). 00B206.1663
[0011] Complicating the identification of off-target sites is the fact that the relationship between sequence similarity and CRISPR-induced cleavage frequency is understood to vary with the level of DNA accessibility. For example, the accessibility needed for efficient CRISPR-mediated chromatin cleavage is less than the accessibility needed for endogenous gene expression (Chung et al., “Computational Analysis Concerning the Impact of DNA Accessibility on CRISPR-Cas9 Cleavage Efficiency,” Mol. Ther. 28(1): 19-28, 2020). Yet, DNA accessibility below a certain threshold has been seen to completely abrogate the effect of the gRNA:target similarity on CRISPR-induced cleavage frequency. Despite such improvements in understanding, there remains a continued need to identify and reduce the likelihood of unwanted insertions and / or deletions (indels) due to off-target CRISPR-Cas binding.
[0012] Therefore, improved, cost-effective, rapid, reproducible materials and methods for CRISPR-based repair off-target analysis and computer analytics relating to the same are needed.
[0013] SUMMARY OF THE INVENTION
[0014] Identifying new methods as well as compositions for use in such methods, for cost-effectively and quickly mapping nuclease (e.g., Cas nuclease) on-target and / or off-target binding sites remains needed in the industry, especially in the context of CRISPR-based repair off-target analysis. Accordingly, the present disclosure is directed to materials and methods for CRISPR-based repair off-target analysis and computer analytics relating to the same. In certain embodiments, the present disclosure is directed to computer-implemented methods comprising: one or more processors; and a memory storing instructions, which when executed by the one or more processors, cause the one or more processor to identify on- and off-target cleavage sites of a Cas nuclease by: (a) aligning nucleic acid sequences obtained from a nucleic acid library to a genomic reference sequence, where the nucleic acid library is generated by DNA binding protein-targeted enzymatic digestion of nucleic acids from cells contacted by a Cas nuclease, an sgRNA, and a target DNA binding protein; (b) identifying peaks based on read coverage of the aligned nucleic acid sequences to the genomic reference sequence; (c) instantiating at least one one-dimensional data structure comprising values correlated to locations of the aligned obtained nucleic acid sequence to the genomic reference sequence correlated to the location of the identified peaks; (d) inputting the at least one, one-dimensional data structure to an artificial neural network (ANN) that has been trained using pre-identified peaks from known on- and off-target Cas nuclease sites to identify Cas nuclease off-target sites; 00B206.1663 and (e) generating, based on output from the ANN, a list of discovered Cas nuclease on and off-target sites.
[0015] In certain embodiments, the cells were subjected to fixation prior to generation of the nucleic acid library. In certain embodiments, the fixation comprises a 0.0001% - 0.1% fixative. In certain embodiments, fixation comprises contacting the cells with formaldehyde, paraformaldehyde, glutaraldehyde, disuccinimidyl glutarate (DSG), or a combination thereof. In certain embodiments, the fixation comprises contacting the cells with the fixative for between about 1 minute to about 5 minutes.
[0016] In certain embodiments of the methods of the present disclosure, the Cas nuclease is a Cas9 nuclease. In certain embodiments, the Cas nuclease is a Casl2a nuclease. In certain embodiments of the methods of the present disclosure, the DNA binding protein is a DNA repair factor. In certain embodiments, the DNA repair factor is selected from the group consisting of ATM, NBS, MRE11, Ku70, Ku80, and FANCD2. In certain embodiments, the DNA repair factor is ATM. In certain embodiments, the DNA repair factor is MRE11. In certain embodiments of the methods of the present disclosure, the cells are incubated in the presence of one or more DNA repair inhibitors prior to generation of the nucleic acid library. In certain embodiments, the DNA repair inhibitor is a DNA-dependent protein kinase (DNA- PK) inhibitor, an ATM inhibitor, or a combination of a DNA-PK inhibitor and an ATM inhibitor. In certain embodiments, the DNA repair inhibitor is selected from the group consisting of KU060648, KU55933, KU60019, AZD1390, AZ31, NU7026, NU7441, AZD764, and AZD7648. In certain embodiments, the DNA repair inhibitor is KU060648. In certain embodiments, the DNA repair inhibitor is AZD7648. In certain embodiments, the DNA repair inhibitor is AZD7648 and KU060648. In certain embodiments, the DNA repair inhibitor is the combination of a DNA-PK inhibitor and ATM inhibitor and the amount of DNA repair inhibitor used is at least 50% less of both the DNA-PK inhibitor and the ATM inhibitor as compared to a method using the DNA-PK inhibitor and ATM inhibitor alone.
[0017] In certain embodiments of the methods of the present disclosure, the cells are mammalian adherent cells. In certain embodiments, the cells are mammalian nonadherent cells. In certain embodiments, the cells are selected from the group consisting of T cells, K562 cells, induced pluripotent stem cells, and cells derived from induced pluripotent stem cells. In certain embodiments, the T cells are human CD4+ T cells or human CD8+ T cells.
[0018] In certain embodiments, the pluripotent cells are induced pluripotent stem cells. In certain embodiments of the methods of the present disclosure, the off-target site identified occurs with a frequency of at least 0.001% of the sequence reads as verified by a targeted 00B206.1663 approach. In certain embodiments, the off-target site identified occurs with a frequency of between 0.01% and 100% of the sequence reads as verified by a targeted approach. In certain embodiments, the off-target site identified occurs with a frequency of between 0.01% and 50% of the sequence reads as verified by a targeted approach. In certain embodiments, the off-target site identified occurs with a frequency of between 0.01% and 10% of the sequence reads as verified by a targeted approach. In certain embodiments, the off-target site identified occurs with a frequency of between 0.01% and 1% of the sequence reads as verified by a targeted approach. In certain embodiments, the off-target site identified occurs with a frequency of between 0.1% and 100% of the sequence reads as verified by a targeted approach. In certain embodiments, the off-target site identified occurs with a frequency of between 0.1% and 50% of the sequence reads as verified by a targeted approach. In certain embodiments, the off-target site identified occurs with a frequency of between 0.1% and 10% of the sequence reads as verified by a targeted approach. In certain embodiments, the off-target site identified occurs with a frequency of between 0.1% and 1% of the sequence reads as verified by a targeted approach. In certain embodiments, the off-target site identified occurs with a frequency of between 1% and 100% of the sequence reads as verified by a targeted approach. In certain embodiments, the off-target site identified occurs with a frequency of between 1% and 50% of the sequence reads as verified by a targeted approach. In certain embodiments, the off-target site identified occurs with a frequency of between 1% and 10% of the sequence reads as verified by a targeted approach.
[0019] In certain embodiments, the methods for identifying on and off-target cleavage sites of a Cas nuclease disclosed herein are used to assay samples comprising cells at leastl cell, at least 10 cells, at least 100 cells, at least 500 cells, at least 1000 cells, at least 5000 cells, at least 10,000 cells, at least 20,000 cells, at least 30,000 cells, at least 40,000 cells, at least 50,000 cells, at least 60,000 cells, at least 70,000 cells, at least 80,000 cells, at least 90,000 cells, at least 100,000 cells, at least 200,000 cells, at least 300,000 cells, at least 400,000 cells, at least 500,000 cells, at least 600,000 cells, at least 700,000 cells, at least 800,000 cells, at least 900,000 cells, at least IxlO6cells, at least IxlO7cells, or at least IxlO8cells per assay.
[0020] In certain embodiments of the methods of the present disclosure, the ANN is trained on the characteristics of the peaks obtained from known on- and off-target nucleotide- directed nuclease cut-sites validated by a targeted approach. In certain embodiments, the targeted approach is rhAmpSeq. In certain embodiments of the methods of the present disclosure, the peak calling comprises applying Threshold-based Method (TM) peak calling algorithm. In certain embodiments, the peak calling comprises applying a threshold to the read 00B206.1663 coverage, by the aligned nucleic acid sequences, of the genomic reference sequence. In certain embodiments, the threshold for peak width is at least about 200 bp and the threshold for peak depth is at least 4 reads. In certain embodiments, the methods described herein further comprise passing a filter over the called peaks, wherein the filter is configured to emphasize features in the called peaks associated with a breakpoint. In certain embodiments, the filter is configured for edge detection. In certain embodiments, the methods described herein comprise estimating a maximum likelihood that the filtered peaks correspond to sites of off-target sequences, wherein said estimating is based on the output from the ANN and a co-localized sgRNA alignment score.
[0021] BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee pursuant to 37 C.F.R. § 1.84.
[0023] FIGS. 1A-1C depicts differences between conventional ChlP-seq, CUT&RUN, and CUT&Tag strategies. FIG. 1 A depicts the steps needed to obtain nucleic acid sequences using ChlP-seq. This approach is the principle behind the Discover-seq method of Wienert et al. (2019) with Discover-seq+ having the addition of a specific DNA repair inhibitor as described by Zou et al. (2023). FIG. IB depicts the steps of CUT&RUN as described by Skene et al., (2017). FIG. 1C illustrates the steps of CUT&Tag as described by Kaya-Okur et al., (2019). These latter two methods have not previously been applied to CRISPR-Cas off-target discovery.
[0024] FIGS. 2A-2C show exemplary read profiles of: (2A) nucleic acid fragments obtained by DISCOVER-Seq using Cas9 or Casl2; (2B) nucleic acid fragments obtained by the CUT&RUN-based methods disclosed herein using Cas9; and (2C) nucleic acid fragments obtained by the CUT&RUN-based methods disclosed herein using Casl2.
[0025] FIG. 3 shows an exemplary method for identifying Cas9-derived off-target cutsites.
[0026] FIG. 4 shows an example of an artificial neural network, e.g., a convolutional neural network, architecture.
[0027] FIGS. 5A to 5D show examples of results of pre-processing / feature extraction performed on one dimensional vectors to be input to the artificial neural network disclosed herein. 00B206.1663
[0028] FIG. 6 depicts an overview of an exemplary approach according to the present disclosure comprising a subset of exemplary protocol parameters including: (1) contacting cells with a CRISPR-Cas gene editing system; (2) plating the cells in the presence of a DNA repair inhibitor; (3) collecting the cells after incubation in the presence of a DNA binding protein inhibitor, e.g., a DNA repair inhibitor; (4) performing a DNA binding protein-targeted enzymatic assay, e.g., CUT & RUN or CUT & Tag; (5) performing nucleic acid library preparation and sequencing; (6) performing machine learning-based computational approach to analyze the sequence data to determine on-target and off-target sites.
[0029] FIG. 7 depicts a barplot showing the on-target RPM for various combinations of CUT&Tag or CUT&RUN experiments with different fixation conditions.
[0030] FIG. 8 depicts a barplot showing the off-target RPM for various combinations of CUT&Tag or CUT&RUN experiments with different fixation conditions.
[0031] FIG. 9 depicts a coverage plot showing genomic locus centered on the on-target cut site (vertical black dotted line) + / - 250 bp (Chr6:43770576-43771075, GRCh38) with the VEGFA gene body below (grey horizontal bar represents VEGFA Exon 1 non-coding domain, and the black horizontal bar adjacent represents the VEGFA Exon 1 coding domain). FIG. 9 demonstrates that the methods of the present disclosure can provide significant and detectable results when employing Cas9 nuclease.
[0032] FIG. 10 depicts a coverage plot of Casl2 (As variant) CUT&RUN experiment showing genomic locus Chr6:43, 769, 752-43, 769, 937 (GRCh38) centered on an on-target cut site.
[0033] FIG. 11 depicts the RPMs per sample observed for the on-target site VEGFA Site 2 sgRNA, calculated as previously described, in the presence of different repair factors.
[0034] FIG. 12 depicts the RPMs per sample observed for the off-target site VEGFA Site 2 sgRNA, calculated as previously described, in the presence of different repair factors.
[0035] FIG. 13 depicts the RPMs for the on-target site with different combinations of inhibitors and concentrations as outlined in Example F, below.
[0036] FIG. 14 depicts the RPMs for the off-target site with different combinations of inhibitors and concentrations, as outlined in Example F, below.
[0037] FIG. 15 depicts a coverage plot of a CUT&Tag experiment in HEK293 cells showing genomic locus centered on the on-target cut site (vertical black dotted line) + / - 250 bp (Chr6:43770576-43771075, GRCh38) with the VEGFA gene body below (grey horizontal bar represents VEGFA Exon 1 non-coding domain, and the black horizontal bar adj acent represents the VEGFA Exon 1 coding domain). 00B206.1663
[0038] FIG. 16 depicts a coverage plot of a CUT&RUN experiment showing genomic locus centered on the on-target cut site (vertical black dotted line) + / - 250 bp (Chr6:43770576- 43771075, GRCh38) with the VEGFA gene body below (grey horizontal bar represents VEGFA Exon 1 non-coding domain, and the black horizontal bar adjacent represents the VEGFA Exon 1 coding domain).
[0039] FIG. 17 depicts a barplot showing the on-target RPM values for human CD8+ T cells with and without DNA-PK inhibitors, vehicle control, and nucleofection negative control.
[0040] FIG. 18 depicts a barplot showing the on-target RPM values for iPSCs with and without DNA-PK inhibitors, vehicle control, and nucleofection negative control.
[0041] FIG. 19 depicts on-target site RPM values which are plotted for titrated cell counts from 2,000,000 to 5,000 cells in a CUT&Tag experiment. Horizontal dotted line shows background signal level as mean of two control samples evaluated at the same on-target sites. RPM values decrease with cell number, but not in a linearly proportional manner.
[0042] FIG. 20 depicts off-target site RPM values, calculated as previously described, which are plotted for titrated cell counts from 2,000,000 to 5,000 cells in a CUT&Tag experiment. Horizontal dotted line shows background signal level as mean of two control samples evaluated at the same off-target sites. RPM values decrease with cell number, but not in a linearly proportional manner.
[0043] FIG. 21 illustrates that by using the ratio of RPM at the on-target site to cell number, it is possible to provide an indication of efficiency, in that at lower cell numbers, each cell is able to provide a disproportionately higher contribution to signal strength.
[0044] FIG. 22 illustrates that by using the ratio of RPM at off-target sites to cell number provides an indication of efficiency, in that at lower cell numbers, each cell is able to provide a disproportionately higher contribution to signal strength.
[0045] FIG. 23 depicts RPMs for selected off-target VEGFA Site 2 sites, representing a range of indel % values. As cell number decreases, off-target sites lose signal, but this signal loss is not linearly proportional to the loss in cell number. Additionally, the loss is consistently experienced between off-target sites with both high and low CRISPR off-target activity as measured by indel percentage.
[0046] FIG. 24A-24B depicts amplicon-based validation of the off-target site centered on Chrl :9689829 (GRCh38) for the VEGFA Site 2 sgRNA identified by an exemplary approach of the present disclosure and not identified previously using GUIDE-Seq, DISCOVER-Seq, or DISCOVER-Seq+ . The exemplary model of the present disclosure 00B206.1663 correctly identified the peak (FIG. 24A) with a high confidence (0.9995204 model output) given the peak outline and the indel % over background is 2.20% in K562 cells (FIG. 24B), validating the model’s finding.
[0047] FIG. 25 depicts an overview of an exemplary computational approach of the present disclosure and corresponding data flow after NGS reads have been aligned to the reference genome. Segments of the genome covered by NGS reads, or “peaks”, are identified and converted into vectors. The result of filtering and scaling these vectors can be plotted in 2- dimensional space (shown) where the x-axis represents genomic coordinates and the y-axis represents a histogram of NGS reads aligning to those coordinates, or “coverage”. These filtered and scaled vectors are passed through the ANN and result in a posterior float value representing the probability that the peak represents a cut site. This posterior float value can take the value between and inclusive of 0 and 1.
[0048] FIG. 26 compares peaks between CHiP-Seq (top track), CUT&RUN (middle track), and CUT&Tag (bottom track). The genomic loci are centered on the on-target cut site (vertical black dotted line) + / - 250 bp (Chr6:43770576-43771075, GRCh38) with the VEGFA gene body below (grey horizontal bar represents VEGFA Exon 1 non-coding domain, and the black horizontal bar adjacent represents the VEGFA Exon 1 coding domain). Due to the difference between the peak profiles surrounding the cut site, CUT&Tag requires a radically different and more flexible approach to identifying cut sites than CHiP-Seq or CUT&RUN. To address this need an ANN was developed that could recognize CUT&Tag cut site coverage profiles.
[0049] FIG. 27 illustrates examples of both positive and negative peaks. Positive peaks contain a characteristic valley between two peaks (left), while negative peaks can either be of low coverage (middle), or high coverage but incorrect profile (right). Peaks for training the ANN were aggregated across 66 samples using the VEGFA Site 2 sgRNA, resulting in 1020 positive peaks. 508915 negative peaks were aggregated from samples nucleofected with vehicle control. With a train-test split of 0.8:0.2, downsampling of the negative peaks to create class balance, and synthetic data approaches (primarily inversion of the peak signal), the final training set consisted of 9792 peaks, of which 4896 had positive labels and 4896 had negative labels. The corresponding test set consisted of 2448 peaks, of which 1224 had positive labels and 1224 had negative labels. These were also split in such a way as to avoid data leakage between the train and test set due to synthetic data: original peaks were split before synthetic data was generated. Data was then trained using Flux.jl, Adam training optimization (0.001 00B206.1663 learning rate), batch size of 64, and sending the data through training twice. After that, the model was fine-tuned using manual curation.
[0050] FIG. 28 depicts an exemplary AUROC (Area Under the Receiver Operating Characteristic curve), which is a measure of model performance describing the trade-off between the true positive rate and the false positive rate. The solid black line tracing through the top left quadrant represents the performance of the model, the lighter solid black line tracing the 45 degree angle from (0.0, 0.0) to (1.0, 1.0) represents the performance of a hypothetical baseline model that has the same classifying performance as random chance, and the dotted line demarcates a false positive rate of 1.0. The performance for the exemplary fine-tuned model described herein is 0.891985, with the ROC curve presented below. Other useful metrics of the model are accuracy (0.890931), precision (0.9741), recall (0.6918), and the Fl score (0.8086).
[0051] FIG. 29 depicts distribution of posterior estimates from an exemplary model. The distribution of exemplary model outputs for true positives are shown on the right, and the distribution of model outputs for true negatives are shown on the right. The model is confident in the vast majority of cases. Note, low scores for true positives may not be unfounded, as the ground truth dataset is not established and some “true” positives have been found to be false positives.
[0052] DETAILED DESCRIPTION
[0053] The present disclosure is directed to materials and methods for CRISPR-based repair off-target analysis and computer analytics relating to the same. In certain embodiments, such materials and methods can robustly identify off-targets arising from the CRISPR-Cas system inside cells from as few as 5,000 cells, or even fewer. In certain embodiments, the methods described herein can employ a computational workflow that is adaptable to different off-target profiles generated by diverse experimental methods and nucleases and is capable of predicting off-targets with high sensitivity.
[0054] In certain embodiments, the present disclosure is directed to materials and methods for CRISPR-based repair off-target analysis that take advantage of the binding of DNA binding proteins, e.g., DNA repair proteins, as a marker of CRISPR-based repair at one or more genomic loci. For example, but not by way of limitation, the recruitment of such DNA binding proteins to genomic loci can be induced by CRISPR-based repair, where the genomic loci can comprise an on-target or an off-target site. 00B206.1663
[0055] In certain embodiments, the locus at which a DNA binding protein is bound is determined using an enzymatic assay. For example, but not by way of limitation, the locus at which a DNA binding protein is bound to DNA can be determined by making use of targeted enzymatic digestion of DNA at or near the binding site of the DNA binding protein. In certain embodiments, such targeted enzymatic digestion is achieved by employing an antibody capable of binding the DNA binding protein to target such enzymatic digestion. For example, in certain embodiments, such antibody-targeted enzymatic digestion employs an antibody capable of binding the DNA binding protein to recruit an enzyme (e.g., a nuclease or a transposase) or a fusion protein comprising such an enzyme to the locus bound by the DNA binding protein. In certain embodiments, such antibody-targeted enzymatic digestion is performed as described herein for Cleavage Under Targets and Tagmentation (CUT&Tag) and in Kaya-Okur et al., “CUT&Tag for efficient epigenomic profiling of small samples and single cells,” Nature Comm. 10: 1930 (2019). In certain embodiments, such antibody-targeted enzymatic digestion is performed as described herein for Cleavage Under Targets and Release Using Nuclease (CUT&RUN)) and in Skene, et al, “An Efficient Targeted Nuclease Strategy for High- Resolution Mapping of DNA Binding Sites,” eLife. 2017 Jan 16;6:e21856.
[0056] In certain embodiments, the present disclosure is directed to computer analytics and associated computational workflows adaptable to analyze data generated by diverse experimental methods, e.g., the enzymatic assays described herein, that can identify on- and off-targets with high sensitivity based on such data. In certain embodiments, such computer analytics and associated computational workflows can identify on- and off-target sites based on sequencing data, e.g., next generation sequencing data. In certain embodiment, such sequencing data comprises an enrichment and / or pileup of sequencing reads obtained using the methods described herein. For example, but not by way of limitation, such enrichment and / or pileup of sequencing reads can exhibit a distinctive pattern that can be differentiated from background patterns (of reads or read pileups) not likely associated with / comprising cut sites. In certain embodiments, the computer analytics and computational workflows disclosed and used herein to assess on- and off-target sites can comprise the application of an artificial neural network (ANN), such as a convolutional neural network (CNN). In certain embodiments, such ANNs are trained on sequence data generated according to the methods disclosed herein (e.g., the newly modified CUT&RUN and / or CUT&Tag based methods disclosed herein), and / or generated or augmented by synthetic data techniques. 00B206.1663
[0057] Acronyms and Definitions
[0058] The following acronyms and definitions as used herein are defined below unless otherwise defined in the paragraph context wherein they occur.
[0059] ANN artificial neural network; an ANN is a system of interconnected computational units or nodes that transforms inputs into outputs through weighted connections, non-linear activations, and iterative learning processes, enabling the modeling of complex relationships in data.
[0060] AmpSeq amplicon sequencing using NGS
[0061] ATM inhibitor ataxia-telangiectasia mutated (ATM) inhibitor is a protein that inhibits the ATM protein involved in DNA damage repair
[0062] AUC Area under the curve
[0063] BAM file (*.bam) A type of file that represents NGS reads, typically from a FASTQ file, aligned to a reference sequence, typically in the form of a FASTA file. A BAM file is the compressed binary version of a SAM file that is used to represent aligned sequences up to 128 Mb.
[0064] BLENDER BLunt END findER - a customized open-source bioinformatics pipeline for identifying off-target sequences genome-wide from DISCOVER-Seq (ChlP-Seq) generated data. Optimized to find blunt-ends assumed and characteristically generated via DISCOVER-Seq.
[0065] BOWTIE Also known as B0WTIE1, BOWTIE is a memory-efficient short-read aligner. It aligns short DNA sequences (reads) to the human genome at a rate of over 25 million 35-bp reads per hour. Bowtie indexes the genome with a Burrows- Wheel er index to keep its memory footprint small: typically about 2.2 GB for the human genome (2.9 GB for paired-end)
[0066] B0WTIE2 Free open-source software that provides an ultrafast and memory-efficient tool for aligning sequencing reads to long reference sequences, such as genome reference sequences (e.g., hg38 human genome). For reads longer than about 50 bp Bowtie 2 is generally faster, more sensitive, and uses less memory than Bowtie 1 00B206.1663
[0067] BWA Burrow-Wheeler Aligner
[0068] BWA-MEM BWA maximal exact match
[0069] Cas CRISPR-associated protein
[0070] Cas9 CRISPR-associated protein 9
[0071] Cas 12 CRISPR-associated protein 12
[0072] ChIP chromatin immunoprecipitation
[0073] CNN convolutional neural network
[0074] CRIS CRISPR repair in vivo screen
[0075] CRISPR clustered regularly interspaced palindromic repeats crRNA CRISPR RNA
[0076] CUT&RUN Cleavage Under Targets and Release Using Nuclease; CUT&RUN sequencing is a method of analyzing protein interactions with DNA. It is an adaptation and improvement of chromatin endogenous cleavage
[0077] CUT&TAG Cleavage Under Targets and Tagmentation (CUT&Tag)
[0078] DDR DNA damage responses
[0079] DISCOVER-Seq Discovery of In Situ Cas off-targets and Verification by Sequencing; it is a type of sequencing involving inhibition of DNA-dependent protein kinase catalytic subunit which accumulates the repair protein MRE11 of the MRN complex at the CRISPR-Cas targeted sites enabling high-sensitivity mapping of off-target (OT) sites to positions of MRE11 binding using chromatin immunoprecipitation (ChIP) followed by sequencing (Zou et al., “Improving the sensitivity of in vivo CRISPR off-target detection with DISCOVER-SEQ+,” Nat. Meth. 20: 706-713 (2023)).
[0080] DISCOVER-Seq+ Inhibition of DNA-dependent protein kinase catalytic subunit accumulates the repair protein MRE11 at CRISPR-Cas-targeted sites, enabling high-sensitivity mapping of off-target sites to positions of MRE11 binding using chromatin immunoprecipitation followed by sequencing. This technique, termed DISCOVER-Seq+, discovered up to fivefold more CRISPR off-target sites in immortalized cell lines, primary 00B206.1663 human cells, and mice compared with previous methods. Zou et al. (2023).
[0081] DMSO dimethyl sulfoxide
[0082] DNA deoxyribonucleic acid
[0083] DNA PKcs DNA-dependent protein kinase catalytic subunit
[0084] DNA PK DNA-dependent protein kinase
[0085] DSB double-stranded DNA breaks
[0086] ENCODE DAC ENCODE data analysis center which contains genome regions that have anomalous, unstructured, or high signal in NGS experiments. Amemiya et al., “The ENCODE Blacklist: Identification of Problematic Regions of the Genome,” Set. Rep. 9: 9354 (2019).
[0087] FASTA text-based format for representing either nucleotide sequences or amino acid (protein) sequences
[0088] FASTQ FASTQ files are text files containing sequence data with a quality (Phred) score for each base, represented as an ASCII character. The quality score is an integer (Q) which is typically in the range of 2-40 but higher and lower values can sometimes be used.
[0089] FPR false positive rate gRNA guide RNA is a general term for all guide RNAs used with CRISPR. gRNA herein may refer to crRNA, tracrRNA, and / or sgRNA, which is a duplex made of a tracrRNA and a crRNA that hybridize together
[0090] HDR homology-directed repair
[0091] HI FBS heat-inactivated fetal bovine serum
[0092] HiFi High Fidelity
[0093] HS nuclease hypersensitive sites
[0094] ICE analysis inference of CRISPR edits
[0095] INDELS insertions / deletions
[0096] Indel frequency Indels are likely to represent between 16% and 25% of all sequence polymorphisms in humans. In fact, in most known genomes, including humans, indel frequency tends to be markedly lower than that of single nucleotide polymorphisms 00B206.1663
[0097] (SNP), except near highly repetitive regions, including homopolymers and microsatellites.
[0098] MACS is a peak calling algorithm for ChlP-seq data. Peak calling is the process used to predict the regions of the genome that, e.g., transcription factors, bind.
[0099] MACS2 Model-based analysis of ChlP-seq described in, e.g., Zhang, Y. et al., “Model-based analysis of ChlP-seq (MACS),” Genome Biol. 9, R137 (2008).
[0100] MAST Motif alignment and search tool mgRNA multi-target guide RNA
[0101] MMEJ microhomology -mediated end-joining
[0102] MNase micrococcal nuclease
[0103] MNase-seq (micrococcal nuclease sequencing) is used to map nucleosome positions in eukaryotic genomes to study the relationship between chromatin structure and DNA-dependent processes. Current protocols require at least two days to isolate nucleosome-protected DNA fragments. See Valouev et al. “Determinants of nucleosome organization in primary human cells,” Nature 474(7352): 516-20 (2011).
[0104] NBS1 Nijmegen breakage syndrome 1 gene, now NBN gene, encodes the nibrin 1 protein that is part of the MRN complex, which comprises MRE1 l-Rad50-NBSl), which senses doublestranded DNA breaks.
[0105] MRE11 a component of the MRX (or MRN) complex involved in double-stranded break repair and a DDR protein. It is designated as MRE11A to distinguish it from the pseudogene MRE11B, also known as MRE11P1.
[0106] MRX / MRN Mrel l-Rad50-Xrs2 / NBSl is a highly conserved complex comprising the components of Mrel l, Rad50, and Xrs2 (in yeast) or NBS1 (in humans) involved in DNA repair. MRX / MRN limits transcription and mediates chromatin anchoring to the nuclear pore complex (NPC). n.d. not determined
[0107] NC non-coding 00B206.1663
[0108] NGS next-generation sequencing
[0109] NHEJ non-homologous end joining
[0110] NPC Nuclear pore complex
[0111] NTC non-targeting control
[0112] OT Off-targets are one or more mismatches between the guide
[0113] RNA and the genome. “On-targets” is the perfect guide RNA- match to the reference genome. pA protein A pAG protein A and protein G pG protein G
[0114] PAM Protospacer Adjacent Motif or Partitioning Around Medoids
[0115] (PAM) algorithm.
[0116] PSSM Position specific scoring that can be used for CUT&RUN sequences rhAmpSeq™ type of amplicon sequencing using RNAse H2 cleavage. rhAmpSeq™ CRISPR panels are highly multiplexed primer sets for targeted amplicon sequencing, designed to assess CRISPR edits for confirming on- and off-target CRISPR gene editing experiments
[0117] RNA ribonucleic acid
[0118] RNP ribonucleoproteins
[0119] RPM reads per million mapped reads are typically used as a method to normalize NGS read count by sequencing depth. On-target RPM is calculated by summing the number of reads overlapping with the on-target cut site. Off-target RPM is calculated by the same approach as on-target RPM with stacked reads but instead of one on-target site the mean of the top 30 off-target sites (e.g., for VEGFA Site 2, the top 30 sites are sourced from Wienert et al., 2019, listed in Table 2) is calculated.
[0120] RT room temperature
[0121] SAMtools A software suite developed to manipulate SAM, BAM, and
[0122] CRAM alignment files (Li et al., “The sequence alignment / map format and SAMtools,” Bioinformatics 25: 2078-2079, 2009). 00B206.1663
[0123] Allows for operations on said alignment files, such operations including, but not limited to, conversion between formats, genomic interval set operations, and summary statistics.
[0124] Alignment files, such as BAM files, can be obtained by mapping raw NGS sequencing reads to reference sequences using aligners (e.g., using a BWA algorithm, such as BWA- MEM - see e.g., Li et al., “Fast and accurate long-read alignment with Burrows transform,” Bioinformatics 26: 589-95, 2010). sgRNA single guide RNA - a single molecule comprising crRNA and tracrRNA fused as a duplex tracrRNA trans-activating CRISPR RNA
[0125] TPR true positive rate
[0126] TSS transcription start sites
[0127] Uli-CUT&RUN ultra-low input CUT&RUN
[0128] VEGFA vascular endothelial growth factor A
[0129] WT wild type
[0130] The following definitions are provided and are to be used unless the context in which it is used indicates differently.
[0131] In the claims articles such as “a,” “an,” and “the” may mean one or more than one unless indicated to the contrary or otherwise evident from the context. Claims or descriptions that include “or” between one or more members of a group are considered satisfied if one, more than one, or all of the group members are present in, employed in, or otherwise relevant to a given product or process unless indicated to the contrary or otherwise evident from the context. The invention includes embodiments in which exactly one member of the group is present in, employed in, or otherwise relevant to a given product or process. The methods and compositions described herein include embodiments in which more than one, or all of the group members are present in, employed in, or otherwise relevant to a given product or process. Furthermore, it is to be understood that the methods and compositions encompass all variations, combinations, and permutations in which one or more limitations, elements, clauses, descriptive terms, etc., from one or more of the listed claims, are introduced into another claim. For example, any claim that is dependent on another claim can be modified to include one or more limitations found in any other claim that is dependent on the same base claim. Furthermore, where the claims recite a composition, it is to be understood that methods 00B206.1663 of using the composition for any of the purposes disclosed herein are included, and methods of making the composition according to any of the methods of making disclosed herein or other methods known in the art are included, unless otherwise indicated or unless it would be evident to one of ordinary skill in the art that a contradiction or inconsistency would arise.
[0132] Where elements are presented as lists, e.g., in Markush group format, it is to be understood that each member if a member comprises a subgroup of elements is also disclosed. Elements of such lists are also contemplated to include any combination or permutation of the list, wherein any one or more element(s) can be removed from the group. It should be understood that, in general, where a method, composition, or device, is / are referred to as comprising particular elements, features, etc., certain embodiments may “consist”, or “consist essentially” of, such elements, features, etc. For purposes of simplicity, those embodiments have not been specifically set forth explicitly herein. It is also noted that the claim term “comprising” is intended to be open and permits the inclusion of additional elements or steps, whereas “consisting” is closed and omits all other steps or components. The term “consisting essentially of’ limits the method, device, or composition by allowing in other non-essential steps, devices, and compositions that are not necessary to achieve the described methods, devices, and compositions as discussed herein.
[0133] The description may use the terms “aspect(s)” or “embodiment s),” which may each refer to one or more of the same or different embodiments. Furthermore, the terms “comprising,” “including,” “having,” and the like, as used concerning embodiments, are synonymous, and are generally intended as “open” terms (e.g., the term “including” should be interpreted as “including but not limited to,” the term “having” should be interpreted as “having at least,” the term “includes” should be interpreted as “includes but is not limited to,” etc.).
[0134] The terms “amplification” and “amplify” are meant to include and refer to an increase in the number of copies of a nucleic acid molecule, such as nucleic acids identified and / or obtained by the methods described herein. The resulting amplification products are called “amplicons.” Amplification of a nucleic acid molecule (such as a DNA or RNA molecule) refers to the use of a technique that increases the number of copies of a nucleic acid molecule (including fragments).
[0135] The term “binding” refers to an association between two substances or molecules, such as the hybridization of one nucleic acid molecule to another or itself, the association of an antibody with a peptide, or the association of a protein with another protein (for example the binding of a transcription factor to a cofactor) or nucleic acid molecule (for example the binding of a transcription factor to a nucleic acid, such as chromatin DNA). 00B206.1663
[0136] The term “stable binding” refers to binding that lasts for a long time, typically, hours or longer. For example, histones bind to DNA to form nucleosomes to maintain chromatin structure, which can last from several hours to the entire lifespan of a cell, depending on the specific context and cellular conditions.
[0137] The term “transient binding” refers to the temporary association of a protein with DNA typically ranging from milliseconds to seconds or minutes. For example, a Cas9 protein, when guided by a single-guide RNA (sgRNA), binds to its target DNA sequence to introduce a double-strand break (DSB). The binding duration can range from a few seconds to several minutes, depending on factors such as the efficiency of target recognition and the presence of mismatches. Additionally, transcription factors are believed to have a very short residence time of only a few seconds on their specific DNA or chromatin-binding sites. However, transcription factors and their binding sites are typically sufficiently abundant to provide sufficient signal for previous applications of CUT&RUN or CUT&Tag type methods.
[0138] The term “weak binding” refers to interactions between a protein and DNA that are characterized by low affinity and rapid dissociation. For example, some transcription factors exhibit weak binding to DNA; NF-KB can exhibit weak binding to DNA enabling it to quickly respond to external stimuli such as cytokines and stress signals. This transient binding is crucial for its role in rapid gene activation.
[0139] The term “strong binding” refers to interactions between a protein and DNA that are characterized by high affinity and slow dissociation. The stability of histone-DNA interactions is maintained through strong electrostatic interactions between the positively charged histone proteins and the negatively charged DNA backbone. This strong binding ensures nucleosomes remain intact over extended periods, providing structural integrity to chromatin and regulating access to genetic material.
[0140] The term “abundant binding sites” refers to the presence of numerous binding sites for a particular protein within the genome, typically at least thousands of sites across the human genome, as an example rate of sites. These sites can be widely distributed and frequently occupied by a particular protein. For example, CTCF has abundant binding sites throughout the genome; its widespread presence is essential for its role in maintaining genome architecture (typically tens of thousands of binding sites). GATA1 is also considered “sparse” (but can have several thousand binding sites so “sparse” is context-dependent, e.g., GATA1 binds to sparse binding sites in the genome, primarily at regulatory regions of genes involved in hematopoiesis (typically up to several thousand sites for GATA1)). 00B206.1663
[0141] The term “sparse binding sites” refers to the presence of relatively few binding sites for a particular protein within the genome, typically hundreds across the human genome as an example rate of sites. These sites may further be less frequently occupied and may be highly specific to certain genomic regions. For example, Cas9 nuclease binds to sparse binding sites in the genome. These sites are primarily the intended target sequences for genome editing, with a limited number of off-target sites. The specificity and sparsity of Cas9 binding sites are crucial for their role in precise genome editing, ensuring accurate modification of target genes while minimizing off-target effects. For highly specific sgRNAs, the number of sites generally ranges in the tens whereas for less specific sgRNAs, the number of binding sites can range in the hundreds.
[0142] As used herein, the term “control” refers to a reference standard. A control can be a positive control to assess the level of a positive response or a negative control that should produce no response or provide for example the background level for the experiment. A difference between a test sample and a control can be an increase or conversely a decrease. The difference can be a qualitative difference or a quantitative difference, for example, a statistically significant difference.
[0143] The terms “align” and “aligning” and “alignment” refer to lining up two nucleic acid sequences. For example, in one embodiment aligning the nucleic acid sequences of the undigested genomic DNA obtained from the methods described herein with a reference sequence suitable for the genomic DNA. For example, using a mouse genome for undigested mouse genomic DNA. Sequence alignment may be performed using any software-based alignment tool / program (aligner) configured to perform pairwise alignment between bases of the nucleic acid sequence of the undigested genomic DNA and the reference genome.
[0144] As used herein, the term “reference sequence” or “reference genome” is a nucleic acid sequence or genome used to line up with another sequence. The reference sequence is generally a “genomic reference sequence” that matches the cells (cell line) being analyzed by the disclosed methods. In the examples described herein, the reference sequence is the human genomic reference sequence of hg38 (e.g., NCBI Genome assembly GRCh38, GCF 000001405.40). The NCBI has 1,136 Homo sapiens genomes currently available. Reference sequences from other species can also be considered for analysis. In general, the reference sequence may be selected and / or designed to match a genome, or part thereof, of the cell line to be manipulated via CRISPR-Cas.
[0145] Time periods for culturing a cell line for analysis using the methods disclosed herein include from about 2 hours to about 48 hours in the presence of the DNA nuclease 00B206.1663 inhibitor. In other aspects, the ranges can be from 3 hours to 24 hours, 3 hours to 18 hours, and 3 hours to 12 hours as well as any 30-minute interval between 2 hours and 48 hours.
[0146] The term “DNA nuclease” as used herein is meant to include a tethered nuclease enzyme capable of cutting double-stranded DNA. For example, the method of CUT&RUN maps a chromatin protein by successive binding of a specific antibody and then tethering a protein A-micrococcal nuclease (pA-MNase) fusion protein in the permeabilized cells of interest without cross-linking. A protein G can be substituted for protein A and tethered to an MNase, i.e., pG-MNase. The tethered MNAse can also be a combination of both proteins A and G to make the tethered MNase, i.e., pAG-MNase. The tethered nuclease, MNase, is activated by the addition of calcium to the cells while they are in the permeabilization buffer. When the calcium is added, DNA fragments are released into the supernatant for DNA extraction, nucleic acid library preparation, and paired-end sequencing. Generally, previously validated DNA nucleases for use in the method are preferred in order to provide reproducible results.
[0147] In the “CUT&Tag” -based strategies described herein, a transposon, e.g., a transposon consisting of hyperactive Tn5 transposase, can be tethered (linked) to an antigen recognized by an antibody, e.g., protein A, protein G, or a protein A and G fusion. In certain embodiments, pAG-Tn5, which is a tethered nuclease produced, for example, by EpiCyper as CUT ANA™ is used. Various types of Tn5 transposases exist and can be used in the methods described herein including, but not limited to, hyperactive forms (EZ-Tn5™ by LGC Biosearch Technologies). A transposome is a transposase-transposon complex. A conventional way for transposon mutagenesis usually places the transposase on the plasmid. In some such systems, termed “transposomes”, the transposase can form a functional complex with a transposon recognition site that is capable of catalyzing a transposition reaction. In CUT&TAG, “tagmentation” combines “fragmentation” and “tag” as the hyperactive variant of the Tn5 transposase which can simultaneously cut the double-stranded DNA (dsDNA) and add adapter sequences (“tag”) - the transposase is modified to integrate sequencing adapters instead of transposons. The Tn5 will only fragment the dsDNA (cleave the dsDNA) and “tag” the DNA upon localization of an antibody to pAG (or other tethering protein) and the correct Mg2+concentration added to the media for cleavage activation of the transposase. Thus, the cells of interest must be under suitable conditions to activate the tethered nuclease to cleave the doublestranded DNA (dsDNA).
[0148] A DNA nuclease can be tethered (or linked) to a specific binding protein, such as an antibody, protein A, or protein G. Thus, in some embodiments, the methods can include 00B206.1663 contacting an uncrosslinked permeabilized cell with a specific binding protein that specifically recognizes the chromatin-associated factor of interest, wherein the specific binding agent is linked to at least one artificial transposome, under conditions that permit integration of a transposon into chromatin DNA. In embodiments using the modified CUT&RUN disclosed herein, the tethered DNA nuclease, is an endo-deoxyribonuclease, such as Micrococcal nuclease (MNase).
[0149] The term “activatable DNA nuclease” is meant to include a nuclease that can be switched from an inactive state to an active state. As discussed herein, the activatable nuclease is a tethered activatable nuclease. The activation switch can be initiated by the addition of an effector or by changing the conditions. In certain embodiments, the effector is a small molecule or atom, such as a Ca2+or Mg2+ion. The DNA nuclease can be any protein capable of inducing cleavage sites into DNA, either single or preferably double-stranded cleavage sites, provided this activity can be activated such as by an ion or other means. The tethered DNA nucleases used in the disclosed methods are capable of breaking the DNA in a largely sequenceindependent manner, generally at nucleosomal linker regions and nuclease hypersensitive sites. Many DNA nucleases however cleave the DNA in a sequence-specific manner, i.e. cleavage occurs mainly at recognition sequences of a few nucleotides. By inactive state, it is meant that the activity of the tethered nuclease is too low to be monitored, or is less than 10% of its maximal rate when active, preferably, less than 4% or less than 1%. The transition from an inactive state to an active state may be triggered by the addition of a chemical compound or by switching the temperature. For a micrococcal nuclease (MN), activation depends on Ca2+ions. An MNase introduces DNA double-stranded breaks in chromatin at nucleosomal linker regions and nuclease hypersensitive (HS) sites. An example of a particularly useful MNase is the sequence encoding the mature chain of Nuclease A (amino acids 83 to 231 of Genbank P00644. See e g., WO 2019 / 060907.
[0150] The phrase “DNA binding protein” is meant to include the class of proteins which can physically interact with DNA, such as DNA repair factors, histones, transcription factors, polymerases, endonucleases, and demethylases. A “DNA repair factor” is a factor that localizes to double-strand break (DSB) sites and is involved in the detection, signaling, and / or repair of DSBs. Such DNA repair factors either directly bind to the broken DNA ends or are recruited to the site of damage to facilitate repair. DNA repair factors include, for example, ATM (NCBI Gene ID No. 471), NBS1 (NCBI Gene ID No. 4683), MRE11 (NCBI Gene ID No. 4361), Ku70 (NCBI Gene ID No. 2547), Ku80 (NCBI Gene ID No. 7520), and FANCD2 (NCBI Gene ID No. 2177). 00B206.1663
[0151] The MRN complex (MRE11-RAD50-NBS1) is one of the primary sensors of DSBs, and its components (RAD50 (NCBI Gene ID No., 10111) and / or NBS1) can function instead of MRE11 as a DNA repair factor, as discussed herein. Any DNA repair factor that localizes and binds to double-strand break sites can be used in lieu of MRE11, such as but not limited to, RAD50, NBS1, Ku70, and Ku80 (which are part of the NHEJ pathway). There are two major pathways for DSB repair, homologous recombination (HR) and non-homologous end joining (NHEJ). Additional HR factors include BRCA1 / BARD1, CtIP, EXO1, DNA2, BLM, WRN, RPA, RAD51, RAD52, BRCA2, Rad55-Rad57, HELF, RADX, RAD54, and Srs2 (de Braganga et al., “Recent Insights Into Eukaryotic Double-Strand DNA Break Repair Unveiled by Single-Molecule Methods,” Trends in Genetics 39(12): 924-940, 2023). NHEJ factors can be used include but are not limited to 53BP1, Ku70-Ku80, DNA-PKcs, XRCC4, XLF, DNA ligase IV (LiglV), APLF, and PAXX (de Braganga et al., 2023).
[0152] The phrase “DNA repair inhibitor” is meant to include any inhibitor of a DNA repair factor. For example, but not by way of limitation, a DNA repair inhibitor can be a DNA- dependent protein kinase (DNA-PK) inhibitor or an ATM inhibitor. Exemplary DNA repair inhibitors include, for example, KU60648, KU55933, KU60019, AZD1390, AZ31, NU7026, NU7441, AZD764, and AZD7648. While combinations of inhibitors can be used, they should generally, but not exclusively, be used at lower doses in order to address their potential impact on cell viability. Alternatively, in certain embodiments, only a single DNA repair inhibitor is used in the methods disclosed herein. In certain embodiments, half a dose of one DNA repair inhibitor and half a dose of a second DNA repair inhibitor can be used in the methods disclosed herein to avoid impacting cell viability. In certain embodiments, e.g., when using more than one DNA repair inhibitor, 25% of one inhibitor and 25% of another inhibitor (relative to what is usually used when only using that particular inhibitor) can also be used in the methods disclosed herein. When using a combination of DNA repair inhibitors, the present disclosure contemplates using about 5% to 50% of the typical amount of the DNA repair inhibitor when the inhibitor is used by itself in the method disclosed herein. An exemplary combination of DNA repair inhibitors can be KU60648 and AZD7648.
[0153] The phrase “Cas nuclease OT sites” is meant to include the Cas nuclease off- target sites produced by a specific Cas nuclease such as Cas9 or Cas 12 nuclease in acting upon untargeted genomic sites creating cleavages that may lead to adverse outcomes. The off-target sites are gRNA-dependent. The Cas nuclease can be complexed with the gRNA before cell permeabilization or the Cas nuclease can be expressed by the cell into which the gRNA enters 00B206.1663 by cell permeabilization. Cells expressing a Cas nuclease can express the Cas nuclease after having been recombinantly transformed with a nucleic acid expressing the nuclease.
[0154] The term “Cas nuclease” as used herein is meant to include a species of Cas9 nuclease (a Class 2, type II Cas nuclease) and a species of Cas 12 nuclease (a class 2, type V Cas nuclease) (Hillary et al., “A Review on the Mechanism and Applications of CRISPR / Cas9 / Casl2 / Casl3 / Casl4 Proteins Utilized for Genome Engineering,” Mol. BiotechnoL 65(3): 311-325, 2023). Casl2 is different from Cas9 in that it can process the precursor crRNA by itself, without the need for tracrRNA or RNase III. This allows researchers to use Casl2 for multiplex genome editing. Casl2a, also known as Cpfl, is a smaller effector that doesn't need a tracrRNA and has RNase activity. It can create a doublestrand break (DSB) at specific sites within the genome, with a staggered 5' overhang. Cas 12 has several branches including Cas 12a, -b, -c, -d, -e, -h, and -i. Cas 12a orthologs from New Francis (Fn), Lachnospiraceaeebacteria (Lb), and Acidaminococus sp. (As) have been identified with effective DNase activity. The use of Casl2a enzyme orthologs is also contemplated with the methods disclosed. The term “cell” can be any cell of interest. Cells can be “eukaryotic cells.” Eukaryotic cells can be mammalian cells or plant cells. Cells of interest can be adherent and nonadherent cells when the cells are cultured. The mammalian cell can be a primate cell such as a human cell. Other mammalian cells may include porcine, equine, rodent, bovine, or ovine. Mammalian cells can include hematopoietic cells, pluripotent stem cells, and cells derived from pluripotent stem cells. Exemplary hematopoietic cells can include T cells, as well as subsets of T cells expressing certain markers.
[0155] By the phrase “nuclei derived from cultured cells” is meant that the nuclei from cells of interest (e.g., mammalian cells) cultured with the introduced gRNA and Cas nuclease complex have been isolated from the rest of the organelles and cell material. The isolated nuclei derived from cultured cells can be cryopreserved in suitable media for later isolation of the undigested genomic sites protected by the one or more DNA repair factors.
[0156] By “permeabilizing” and “permeabilization” are meant as a means to introduce CRISPR components into a target cell or cells, e.g., gRNA and as needed a Cas nuclease. Methods of permeabilizing cells to introduce the CRISPR components include electroporation, nucleofection (e.g., Nucleofector®), lipid-based transfection reagents, and microinjection. Both electroporation and nucleofection use an electrical pulse to create pores in the plasma membrane. Microinjection uses a needle to force a hole into a cell’s membrane and therefore is less efficient and has lower throughput. One form of nucleofection system that can be used is the TransIT-X2 Dynamic Delivery system which provides a means of delivering siRNA, 00B206.1663 miRNA, plasmid DNA, and or CRISPR / Cas9 components into a cell (Minis Bio, a part of Millipore Sigma, Cat. # MIR 6003).
[0157] The terms “fixing” and “fixation” are meant to include a method of subjecting a cell (e.g., a mammalian cell) of interest as used in the methods described herein to preserve the cells and terminate any ongoing biochemical reactions. Exemplary fixation methods include formaldehyde fixation of the cells. Examples of cell fixation that can be used include any one of the following reagents in various combinations:
[0158] A. 16% formaldehyde (Thermo Fisher Scientific, Thermo #28908) (zero-length cross-linking),
[0159] B. 16% paraformaldehyde (Electron Microscopy Sciences, EMS #15710) (zerolength cross-linking),
[0160] C. 8% glutaraldehyde (EMS #16000) (~5 angstrom spacer arm), and
[0161] D. 20 mM DSG (Thermo #A35392) (~7.7 angstrom spacer arm).
[0162] Cells can be fixed by, for example, administering either (A) or (B) fixatives at 0.0001-0.1% for 1 min to 20 min. to the cells. Alternatively, cells can be fixed by administering either the (C) fixative at 0.0001-0.1% for 30 sec to 10 min. followed by the addition of 0.0001- 0.1% of fixatives (A) or (B) for 30 sec to 10 min. In another aspect, cells can be fixed by exposing the cells to 0.1 mM to 2 mM of fixative (D) for 30 seconds to 10 min. followed by the addition of 0.0001-0.1% for 30 sec. to 10 min. of fixatives (A) or (B) to fixative (D) or to fixative (C). In another embodiment, the cells can be fixed with 0.0001 to 0.1% of fixative (C) for 30 sec. to 10 min. followed by exposure to 0.0001-0.1% of fixative (A) or (B) for 30 sec. to 10 min. Fixation used for the presently disclosed methods is lower / lighter than that required in ChIP Seq-based methods (~1% of the above fixatives). Lighter fixation reduces associated data artifacts or damage to samples, allowing for reduced sample input required.
[0163] The term “processing” as used herein refers to processing, by a computing device, of input to generate output data. The computing device comprises one or more processors and memory storing instructions that, when executed by the one or more processors, configure the one or more processors to perform the processing according to any of the computational methods disclosed herein. For example, the computing device may comprise one or more central processing units (CPUs) and / or one or more graphical processing units (GPUs).
[0164] The term “filter” as used herein refers to a type of processing, by a computing device, of data, using an algorithm that subsets original data to a smaller set of data wholly 00B206.1663 contained within the original data using a collection of parameters and / or criteria. For example, filtering may be performed by applying a peak calling algorithm to a BAM file.
[0165] A “feature” as used herein refers to an individual, measurable property or characteristic extracted / extractable from raw data, specifically designed to be used as input for a predictive model in machine learning
[0166] The term “feature extraction” as used herein refers to a type of processing that includes the transformation of data by an algorithm that emphasizes or identifies one or more features in the data. Feature-extracted data results in improved training and test results for a machine learning algorithm relative to the raw or non-feature-extracted data. Feature extraction may or may not include a type of filtering (e.g., peak calling).
[0167] CRISPR-Based Repair of On-Target and Off-Tar et Sites
[0168] Non-viral CRISPR-Cas based gene editing generally offers a safer genotoxicity profile in comparison to lentiviral based approaches. For example, while lentivectors can offer high infection rates in a wide range of cell-types, their inability to control the sites of integration and the copies of vector insertion can, in certain instances, restrict their application in clinical settings. CRISPR-Cas based systems, in contrast, are capable of performing targeted gene edits. The occurrence of off-target cleavages when using CRISPR-Cas based systems are generally considered to be dependent on sgRNA sequence and an in-silico strategy for sgRNA design, e.g., one incorporating certain materials and methods disclosed herein, can be used to identify and significantly minimize such off-target activity.
[0169] The ChlP-seq technique is considered the “gold standard” method for nonhistone proteins (e.g., transcription factors, repair factors, etc.), due to its high specificity and sensitivity, particularly for weakly or transiently binding proteins associated with low- abundance binding sites, such as MRE11 and other repair factor proteins. Traditionally, ChlP- seq-based approaches have been utilized to assess the occupancy of DNA-binding proteins with high sensitivity. Current ChlP-seq protocols, however, require a large amount of input material (e.g., about 1 million to 10 million cells - see, e.g., Wienert, Beeke, et al. "CRISPR off-target detection with DISCOVER-seq." Nature Protocols 15.5 (2020): 1775-1799.) and high sequencing depth to achieve a sufficient signal-to-noise ratio. The ChlP-seq protocols also require an appropriate control sample to reduce / eliminate false positive readouts. The choice of control sample for the ChlP-seq experiment could vary depending upon the biological context (e.g., cell type) of an experiment. However, the requirement of at least one million 00B206.1663 cells prevented the assay’s utilization for systems having fewer cells due to low signal and high background.
[0170] Newer methods of detecting protein-DNA interactions such as CUT&RUN and CUT&Tag offer several advantages over ChlP-seq. These techniques include reduced background noise and lower sample input requirements than ChlP-seq. For example, CUT&RUN can be used for profiling very abundantly / strongly / stably bound proteins with as little sample input as low as 1 pg DNA or other nucleic acid input, as low as 5,000 cells. CUT & Tag provides efficient epigenomic profiling of small samples and even single cells (Kaya- Okur et al., “CUT&Tag for efficient epigenomic profiling of small samples and single cells,” Nature Comm. 10: 1930 (2019)). However, CUT&RUN and CUT&Tag are designed and optimized for a limited number of applications. CUT&RUN and CUT&Tag are intended for use in identifying binding sites on DNA of stably, abundantly, strongly- and / or ubiquitously- bound proteins. CUT&RUN is recommended for profiling histones, histone modifications, transcription factors, and cofactors, all characterized by stable DNA binding at hundreds to thousands of binding sites across a given genome. For example, CUT&RUN has been successfully applied to profile chromatin and some transcription factors, where these transcription factors typically have thousands of putative binding sites. CUT&Tag has even more limited usage - it is strictly recommended for histone proteins, because histone proteins are not only abundant and ubiquitous, but also stably bound to DNA, allowing for effective profiling with CUT&Tag. This stable binding to DNA is critical for the stringent wash steps (e.g., 300 mM NaCl) necessary to limit Tn5's affinity for accessible DNA that is used with CUT&Tag.
[0171] Application of CUT&RUN and CUT&Tag on dynamic and low-abundance DNA binding proteins and / or for identifying low-abundance and / or sparse target binding sites has been limited given the limitations inherent to both techniques. In the case of CRISPR-based gene editing, off-target sites are expected to be few (i.e., in order of lOs / lOOs sites for a given cell genome) and the binding of DNA-bound double-strand break repair proteins (e.g., MRE11) to sites of double-stranded DNA breaks is transient.
[0172] CUT&RUN and CUT&Tag both rely on the successful recruitment of a fusion enzyme to precise, antibody-bound DNA-protein interaction sites. For example, CUT&RUN relies on the recruitment of pAG-MNase (Protein AG-Micrococcal Nuclease, see Meers et al., “Improved CUT&RUN chromatin profiling tools,” eLIFE, 8:e46314, 2019) and CUT&Tag relies on pAG-Tn5 (E coli transposase mutant (Tn5)). Both CUT&RUN and CUT&Tag have been developed and applied for profiling histone modifications and TF binding in diverse cell 00B206.1663 types. It was believed that both CUT&RUN and CUT&Tag would function sub-optimally in the case of weakly or sparsely bound DNA binding proteins, owing to the low availability of antibody-bound sites from CRISPR-based gene editing (even more so for inconsistent binding to off-target sites) and the less stable / more transient binding of DNA repair proteins relative to the transcription factors / histones / etc. for which CUT&RUN and CUT&Tag were developed. Additionally, upon successful localization to antibody-bound sites, CUT&RUN requires targeted enzymatic digestion of the DNA at the antibody-bound sites by pAG-MNase and CUT&Tag requires transposition by pAG-Tn5. These digestion and transposition steps are expected to work well for well-structured and stably bound proteins like histones, but again they were expected not to work for weakly or sparsely bound DNA-binding proteins. Moreover, transiently interacting proteins and obscured chromatin states are expected to result in incomplete digestion or transposition of the target DNA by the pAG-MNase / pAG-Tn5 and therefore, importantly, the failure to capture essential off-target sites. Lastly, the implementation of these protocols on a low cell number (e.g., less than 1 million, less than 100,000, even as low as 5,000-10,000 or lower, as may be the case required for certain cells of interest) was expected to yield significantly lower amounts of DNA for downstream sequencing library preparation from the extracted DNA produced by the modified methods disclosed herein for all of these reasons. It was also expected that both CUT&RUN and CUT&Tag methods would ultimately result in high technical noise (e.g., sequence reads not corresponding to Cas cut sites) and low biological signal (reads corresponding to Cas cut sites) that would render it unsuitable for purposes of assessing CRISPR on and off-target editing sites.
[0173] In view of the foregoing, there was little confidence prior to obtaining the results of the experiments described herein that strategies similar to either CUT&RUN (Cleavage Under Targets and Release Using Nuclease; see, e.g., FIG. IB) or CUT&Tag (Cleavage Under Targets and Tagmentation; see, e.g., FIG. 1C) would provide the sensitivity needed for accurate or reliable CRISPR off-target assessment. Instead, the modified methods described and illustrated herein were determined to produce a higher than expected signal, even when using fewer cells, and having less background noise.
[0174] Despite the expected outcomes described above, exemplary DNA binding protein-targeted enzymatic cleavage assays, e.g., CUT&RUN- and CUT&Tag-based protocols, as disclosed herein, were tested to see if CRISPR off-target sites could be identified and, surprisingly, both methods worked using a disclosed inhibitor even with cell numbers of about 5,000 to 1 million and are expected, based on the data presented herein, to work with any 00B206.1663 number in-between as well as even lower cell numbers as disclosed below. Surprisingly, the modified methods yielded reproducible identification of both CRISPR on- and off-target sites. As outlined herein, successful identification of on- and off-target sites was achieved using the modified CUT&RUN and CUT&Tag-based methods. As such, the modified CUT&RUN and CUT&Tag-based methods disclosed herein provide unexpected, yet robust capture of CRISPR on and off-target editing sites.
[0175] In certain embodiments, the method described herein comprise modifications to conventional CUT&RUN and / or CUT&Tag protocols. For example, but not by way of limitation, the method described herein contemplate modulation of the binding of DNA binding proteins to facilitate the methods of the present disclosure. In certain embodiments, inhibitors of such DNA binding proteins can be employed that extend the duration of the binding of the DNA binding protein. In certain embodiments, e.g., when the DNA binding protein is a DNA repair protein, the inhibitor can be a DNA repair inhibitor.
[0176] Another notable difference between the methods described herein and the methods of CHIP-Seq and CHIP-Seq+ is that CHIP-Seq and CHIP-Seq+ use cells fixed with 1% formaldehyde. In contrast, the methods described can use either fresh (non-fixed) cells, nuclei obtained from non-fixed cells, or cells that have been lightly fixed with less than about 0.1% formaldehyde.
[0177] Another distinction between the methods described herein using modified CUT&RUN and CUT&Tag-based approaches and conventional CUT&RUN and / or CUT&Tag protocols is that the methods described herein can have the off and on-target measurements done at 6 hours or longer with reproducible results, which was unexpected from existing CUT&RUN or CUT&Tag protocols / applications.
[0178] In certain embodiments, the methods for identifying on and off-target cleavage sites of a Cas nuclease disclosed herein is used to assay samples comprising at leastl cell, at least 10 cells, at least 100 cells, at least 500 cells, at least 1000 cells, at least 5000 cells, at least 10,000 cells, at least 20,000 cells, at least 30,000 cells, at least 40,000 cells, at least 50,000 cells, at least 60,000 cells, at least 70,000 cells, at least 80,000 cells, at least 90,000 cells, at least 100,000 cells, at least 200,000 cells, at least 300,000 cells, at least 400,000 cells, at least 500,000 cells, at least 600,000 cells, at least 700,000 cells, at least 800,000 cells, at least 900,000 cells, at least IxlO6cells, at least IxlO7cells, or at least IxlO8cells per assay.
[0179] In certain embodiments, the methods for identifying off-target cleavage sites of a Cas nuclease disclosed herein comprise identifying an off-target cleavage site occurring with a frequency of at least 0.001% of the sequence reads as verified by a targeted approach. In 00B206.1663 certain embodiments, the off-target site identified by the methods disclosed herein occurs with a frequency of between 0.01% and 100% of the sequence reads as verified by a targeted approach. In certain embodiments, the off-target site identified occurs with a frequency of between 0.01% and 50% of the sequence reads as verified by a targeted approach. In certain embodiments, the off-target site identified occurs with a frequency of between 0.01% and 10% of the sequence reads as verified by a targeted approach. In certain embodiments, the off-target site identified occurs with a frequency of between 0.01% and 1% of the sequence reads as verified by a targeted approach. In certain embodiments, the off-target site identified occurs with a frequency of between 0.1% and 100% of the sequence reads as verified by a targeted approach. In certain embodiments, the off-target site identified occurs with a frequency of between 0.1% and 50% of the sequence reads as verified by a targeted approach. In certain embodiments, the off-target site identified occurs with a frequency of between 0.1% and 10% of the sequence reads as verified by a targeted approach. In certain embodiments, the off-target site identified occurs with a frequency of between 0.1% and 1% of the sequence reads as verified by a targeted approach. In certain embodiments, the off-target site identified occurs with a frequency of between 1% and 100% of the sequence reads as verified by a targeted approach. In certain embodiments, the off-target site identified occurs with a frequency of between 1% and 50% of the sequence reads as verified by a targeted approach. In certain embodiments, the off-target site identified occurs with a frequency of between 1% and 10% of the sequence reads as verified by a targeted approach.
[0180] Computational Methods to Detect On-Target and Off-Target Sites
[0181] In certain embodiments, the present disclosure is directed to computational analysis for use in identifying on-target and / or off-target editing sites from sequences produced by CRISPR Cas-based gene editing. As noted above, such computational analysis strategies can be used in connection with the modified CUT&RUN and / or CUT&Tag-based methods disclosed herein. For example, the methods disclosed herein, including, but not limited to, the modified CUT&RUN and / or CUT&Tag-based methods disclosed herein, result in undigested sequences enriched at DNA binding protein binding sites, e.g., DNA-bound double-strand break repair protein (such as MRE11) binding sites. Sequencing (including, but not limited to next-generation sequencing (NGS)) of such cleaved sequences results in sequencing reads that are similarly enriched at said binding sites. This enrichment or pileup of sequencing reads has a distinctive pattern, based on the nature of the experimental protocol employed, that can be differentiated from background patterns (of reads, even read pile-ups) not likely associated 00B206.1663 with / comprising cut sites. In some examples, the computational methods disclosed and used herein to assess on- and off-target sites include the application of an artificial neural network (ANN), such as a convolutional neural network (CNN) trained on sequence data generated according to the methods disclosed herein (e.g., the newly modified CUT&RUN and / or CUT&Tag based methods disclosed herein), and / or generated or augmented by synthetic data techniques. The sequence data may have known on-target and / or off-target breakpoints to create a discriminative filter despite a relatively low number of training data points. In certain embodiments, training on such limited data is possible primarily due to a signal feature extraction disclosed herein, the parsimonious nature of the CNN model, and the application of synthetic data techniques, as disclosed herein and as will be discussed in more detail below.
[0182] The computational analysis strategies described herein not only improve the identification of on-target and / or off-target sites by incorporating training on specific types of data (e.g., data associated with the exemplary DNA binding protein-targeted enzymatic assays as described herein), but also can discern between true positive target sites, true negative target sites, false positive target sites, and false negative target sites more accurately than conventional strategies. Accordingly, the methods described herein represent practical applications of the recited computational approaches that are specifically adapted to the DNA- binding protein-targeted enzymatic assays, and their resultant one dimensional vector data structures, disclosed herein.
[0183] Existing computational approaches for identifying CRISPR off-target sites have been designed for data acquired using primarily the ChlP-Seq method (DISCOVER-Seq), most prominently, BLENDER (BLunt END findER). As such, existing computational approaches are designed / optimized to identify peaks characteristic of such protocols. For example, as shown in FIG. 2A, ChlP-Seq identified breakpoints generated by Cas9 or Casl2 are relatively symmetric and have a clear demarcation between forward and reverse reads. In contrast, the modified CUT&RUN peaks obtained using Cas9, as shown in the example of FIG. 2B, are less symmetric in the orientation of the CRISPR guide and less disciplined in forward and reverse reads (e.g., feature a clearer demarcation between forward and reverse reads). In other words, the modified CUT&RUN peaks obtained using Cas9 as an example of the methods of the present disclosure, are less balanced in terms of sequence pileup on either side of a given cut site, and this imbalance correlates with the guide RNA orientation. Similarly, the exemplary modified CUT&RUN peaks obtained using Casl2 in connection with certain methods disclosed herein, are shown in FIG. 2C and also result in distinct peak characteristics to both the ChlP-SEQ identified breakpoints and the modified CUT&RUN peaks obtained using Cas9. 00B206.1663
[0184] For example, the modified CUT&RUN peaks obtained using exemplary Casl2-based strategies of the present disclosure are generally flatter, with a small peak in the middle near the cut site, and they appear more symmetric than peaks obtained using Cas9. This supports the notion that the methods of the present disclosure, including, but not limited to the exemplary modified CUT&RUN peaks obtained using different Cas enzymes, can be reasonably expected to each produce a unique peak characteristic that may be distinguishable from each other and from background. Unique peak characteristics may also be reasonably expected between different strategies described herein, e.g., the various modified CUT&RUN-produced peaks and modified CUT&Tag-produced peaks. Thus, a computational approach that allows for the flexibility in determining / characterizing new peaks characteristic of all the strategies described herein, e.g., modified CUT&RUN- and / or CUT&Tag-based methods performed with various Cas enzymes is desirable. Such a computational approach can be combined with the DNA binding protein-targeted enzymatic assays, e.g., the modified CUT&RUN and CUT&Tag methods, described herein. The computational analysis including an ANN model, e.g., a CNN model, as disclosed herein even allows for such flexibility in training models for different Cas enzymes on limited data.
[0185] Therefore, the DNA binding protein-targeted enzymatic assays, e.g., both the modified CUT&RUN or CUT&Tag strategies disclosed herein, can be used to robustly detect CRISPR on- and / or off-target editing sites in a cost-effective and time-efficient manner and optionally combined further with the described computational approaches. The computational methods disclosed herein are configured to receive input of sequences of undigested nucleic acids resulting from the tethered nuclease (e.g., MNase)-based methods disclosed herein (e.g., the CUT&RUN and / or CUT&Tag based methods disclosed herein) and output information indicating and / or associated with predicted / detected on- and / or off-target Cas nuclease cut sites.
[0186] FIG. 3 shows an exemplary flowchart of the various computational methods disclosed herein. In step 300, for example, a library of undigested nucleic acid fragments obtained from the tethered nuclease-based methods disclosed herein is sequenced (e.g., by NGS) to generate a set of sequences (e.g., reads) corresponding to the obtained nucleic acid fragments. In step 310, the sequences (e.g., in the form of FASTQ files, etc.) are aligned to a reference sequence (e.g., reference genome, hg38). Alignment is performed using a software alignment tool (e.g., program / application) configured to align one or more sequences of the set of sequences to the reference sequence. Exemplary software alignment tools include, but are 00B206.1663 not limited to, BLAST (blast.ncbi.nlm.nih.gov / Blast.cgi); MAFFT (mafft.cbrc.jp / alignment / server), and Clustal Omega (www.ebi.ac.uk / jdispatcher / msa). In some examples, aligned sequence data may be split into those aligned with different portions of the reference sequence, such as different chromosomes or other desired portions, and the split aligned sequences may be processed as follows in parallel.
[0187] In step 320, peaks are called in the aligned sequence data to identify peaks in the signal associated with read pile-ups that indicate potential cut-sites. Various peak-calling algorithms / programs / filters may be applied, including, but not limited to, Model-based Analysis for ChlP-Seq (MACS), MACS version 2 (MACS2), Genome wide Event finding, and Motif discovery (GEM), MUltiScale enrichment Calling for ChlP-Seq (MUSIC), Bayesian Change point (BCP), Zero-Inflated Negative Binomial Algorithm (ZINBA), Thresholdbased Method (TM) or other suitable type of peak calling algorithms / programs / filters. In an example, coverage thresholding is performed, which has the benefit of simplicity and the ability to bias towards peak inclusion. Reads satisfying a coverage threshold (e.g., covering at least a threshold number of bases, such as at least 200 bp, at least 400 bp, at least 600 bp, at least 800 bp, at least 1000 bp, etc.) are considered peaks. The threshold amount of coverage may allow for gaps of up to a maximum bp length, (e.g., coverage is considered sufficiently continuous if gaps are less than 50 bp, less than 40 bp, less than 30 bp, less than 20 bp, less than 10 bp, less than 5 bp. etc.). The threshold (total coverage threshold and / or allowable gap threshold) may be chosen to be more or less inclusive of peaks and may be chosen based on the particular protocol, e.g., particular MNase-based protocol, used to obtain the nucleic acid fragments resulting in the sequencing data. For example, for the exemplary modified CUT&RUN-based method-generated sequences described herein, the thresholding may be, e.g., >200 bp, with 5 bp allowable gaps. For the exemplary modified CUT&Tag-based method generated sequences described herein, the thresholding may be, for example, >1000 bp, with 50 bp allowable gaps. In general, the thresholding may be applied to favor the inclusion of peaks (less stringent, e.g., lower required coverage, such as > 500 bp, >200 bp, with a higher allowable gap, e.g., <10 bp, <50 bp, etc.), in favor of identifying characteristic peaks with a trained ANN, for example, a CNN or other suitable type of ANN. A 200 bp coverage threshold is an example of less stringency, whereas a 5 bp allowable gap would be more stringent.
[0188] In certain embodiments, the called peaks (e.g., portions of aligned sequences corresponding to the called peaks) are converted to one-dimensional vectors representing the number of reads aligned to bases / locations (e.g., each base / location) of the reference sequence. 00B206.1663
[0189] A “signal”, as used herein with respect to these called peaks, refers to a one-dimensional vector data structure, which can grow dynamically, storing values indicating a number of reads aligned to given bases / locations of the reference sequence stored in a non-transitory medium (e.g., computer memory) by a processor.
[0190] In certain embodiments, the one dimensional vector data structures can be filtered and / or pre-processed to emphasize / extract features associated with cut peaks. For example, the one dimensional vector data structures can be expanded to accentuate intervals that contain relatively few reads. The one dimensional vector data structures correspond to the number of reads present at their alignment locations in the reference genome, and the one dimensional vector data structures can be expanded by expanding the locations (reference genome coordinates) by employing an expansion factor (e.g., 2x, 3x, 4x, etc.). In some instances, the one dimensional vector data structures are not expanded. For example, expansion may not, in certain instances, provide as much benefit in accentuating intervals for exemplary modified CUT&Tag-based method-derived one dimensional vectors relative to exemplary modified CUT&RUN-based method-derived one dimensional vectors.
[0191] In certain embodiments, features can be extracted from the one dimensional vector data structures (whether expanded or not). For example, but not by way of limitation, a kernel / filter configured for edge detection can be convolved with the one dimensional vector data structures (e.g., a kernel of [1, -2, 1]). In certain embodiments, Z-scoring and / or normalization can be applied to the resulting convolved one dimensional vector data structures, resulting in a Z-score normalized convolved one dimensional vector data structures as filtered / pre-processed one dimensional vector data structures.
[0192] With reference to the exemplary flow chart in FIG. 3, in step 340, the filtered / pre-processed one dimensional vector data structures can be input to a ANN, for example, a convolutional neural network (CNN) as shown in FIG. 4. The AAN, e.g., a CNN, can be trained on data generated using a DNA binding protein-targeted enzymatic assay, e.g., tethered-MNase-based methods, disclosed herein. In certain embodiments, the ANN, e.g., a CNN, can be trained on experimental data and / or synthetic data reflecting sequences generated using a DNA binding protein-targeted enzymatic assay, e.g., the tethered-MNase-based methods, disclosed herein using a Cas enzyme of interest. As such, different trained CNN models can be developed to specifically identify peak characteristics of a given Cas enzyme and / or particular a DNA binding protein-targeted enzymatic assay, e.g., tethered-MNase-based method such as the modified CUT&RUN and / or CUT&Tag based methods disclosed herein. 00B206.1663
[0193] To this end, FIG. 4 shows that the ANN, e.g., a CNN, generally includes a series of alternating convolutional and maxpool layers, where n indicates the number of alternating convolutional and maxpool layer pairs in the series (e.g., n = 2, 3, 4, 5, 6, etc.). Each convolutional layer can have channel numbers within the constraints of the previous and subsequent layers (e.g., the convolutional layers can be different or the same in these values). Similarly, the maxpool layers can be different or the same in size and / or stride. The ANN, e.g., a CNN, can then, in certain embodiments, include a flattening layer and a series of m dense layers (e.g., m = 1, 2, 3, 4, 5, etc.). In certain embodiments, the dense layers may be different from each other, m only indicates a number of the dense layers. Finally, in certain embodiments, the output of the dense layers can be passed through a softmax layer, so that the ANN output, e.g., a CNN output, identifies locations of predicted cut-sites and associated probabilities.
[0194] In certain embodiments, e.g., as illustrated in step 350 of FIG. 3, the ANN, e.g., a CNN, computes outputs corresponding to off-target site hits that satisfy a threshold. In certain embodiments, the threshold can take on any real number in the range of 0 to 1 and can be selected / tuned higher or lower depending on the desired trade-off between false positives and false negatives. In an example, the threshold may be selected as a function of the underlying distribution (e.g. the median output of the ANN model, e.g., a CNN model). In certain embodiments, the threshold is applied to a CNN softmax layer, as that layer provides values that are correlated with the underlying probability. Also, or alternatively, the threshold may be applied to further refined output according to the exemplary models described herein (e.g., see the Examples section, below).
[0195] It is intended that the computational analysis systems and methods described herein can be performed by software (stored in memory and / or executed on hardware), hardware, or a combination thereof. Hardware modules may include, for example, a general- purpose processor, a graphical processor unit (GPU), a field programmable gates array (FPGA), and / or an application specific integrated circuit (ASIC). Software modules (executed on hardware) can be expressed in a variety of programming languages (e.g., computer code), including Unix utilities, C, C++, Java™, JavaScript, Ruby, SQL, SAS®, the R programming language / software environment, Visual Basic™, and other object-oriented, procedural, or other programming language and development tools. Examples of computer code include, but are not limited to, micro-code or micro-instructions, machine instructions, such as produced by a compiler, code used to produce a web service, and files containing higher-level instructions that are executed by a computer using an interpreter. Additional examples of computer code include, but are not limited to, control signals, encrypted code, and compressed code. Each of 00B206.1663 the systems and methods described herein can include and / or employ one or more processors as described above.
[0196] Some embodiments described herein relate to devices with a non-transitory computer-readable medium (also can be referred to as a non-transitory processor-readable medium or memory) having instructions or computer code thereon for performing various computer-implemented operations. The computer-readable medium (or processor readable medium) is non-transitory in the sense that it does not include transitory propagating signals per se (e.g., a propagating electromagnetic wave carrying information on a transmission medium such as space or a cable). The media and computer code (also can be referred to as code) can be those designed and constructed for the specific purpose or purposes. Examples of non-transitory computer-readable media include, but are not limited to: magnetic storage media such as hard disks, floppy disks, and magnetic tape; optical storage media such as Compact Disc / Digital Video Discs (CD / DVDs), Compact Disc-Read Only Memories (CD-ROMs), and holographic devices; magneto-optical storage media such as optical disks; carrier wave signal processing modules; and hardware devices that are specially configured to store and execute program code, such as Application-Specific Integrated Circuits (ASICs), Programmable Logic Devices (PLDs), Read-Only Memory (ROM) and Random-Access Memory (RAM) devices. Other embodiments described herein relate to a computer program product, which can include, for example, the instructions and / or computer code discussed herein.
[0197] EXAMPLES
[0198] The following examples are provided to further describe the methods and compositions described herein. For the examples, the following materials and procedures were used unless otherwise indicated. Modifications to the examples would be evident to the skilled artisan depending on the cells used, the number of cells used, and so forth.
[0199] Materials and Methods Used in the Examples
[0200] Protocol for K562 Cells. K562s are maintained in RPMI-1640 medium supplemented with 10% fetal bovine serum (FBS), 2 mM L-Glutamine, and 1% Pen-Strep (RPMI complete media) at -IxlO6cells / mL at 37 °C with 5% CO2.
[0201] Nucleofection of the K562 cells using a Cas nuclease and gRNA (guide RNA) was performed as follows. K562 cells were collected and centrifuged at 300 x g for 5 minutes. The K562 cells were washed with IX PBS (phosphate-buffered saline) and counted for nucleofection. Depending on how many cells / guides / compounds are needed, Lonza recommends different formats as follows: 00B206.1663
[0202] 1. Cuvette Format (up to 10xl06cells per cuvette in 100 pL of nucleofection solution)
[0203] 2. 16-well Nucleocuvette Strip (less than IxlO6cells per well in 20 pL of nucleofection solution per well). Wells are used as they hold fewer cells than the cuvette (nucleocuvette).
[0204] 3. 4D-Nucleofector® 96-well Plate Format (less than IxlO6cells per well in 20 pL of nucleofection solution per well); allows for high-throughput screening.
[0205] A Cas (RNP or mRNA) and guide are added to the nucleofection cuvette. Nucleofection is performed using the appropriate program for K562 cells (e.g., Lonza program FF-120).
[0206] The K562 cells are immediately transferred to pre-warmed RPMI medium complete with inhibitor (1 pM AZD7848, for example), or without inhibitor for a control. In an example, approximately 0.5xl06-lxl06cells (e.g., in 20 pL of nucleofection solution, as in (2) or (3) above) are transferred to a volume of RPMI / inhibitor solution adding up to are added per 500 pL. The cells in RPMI / inhibitor are plated (into, e.g., 12-well, 6-well tissue culture plates - depending on how many cells are to be plated) and incubated at 37 °C with 5% CO2 for 2-48 hours.
[0207] Cells are harvested 2 to 48 hours after incubation in RPMI complete media with or without inhibitor for control. The K562 cells are collected and centrifuged at 300 x g for 5 minutes to form a pellet. After pellet formation, one of three steps can be used to treat the cells: A) fixation followed by nuclei extraction, B) nuclei extraction followed by fixation of the nuclei, or C) nuclei extraction only as follows:
[0208] A. Fixation followed by nuclei extraction: light fixation (with formaldehyde, paraformaldehyde, glutaraldehyde, disuccinimidyl glutarate (DSG), or some combination). By “light fixation” is meant using 0.1-0.0001% fixative for 1-20 minutes followed by 0.125 M glycine to quench fixation. After fixation, cells are again pelleted and resuspended in ice-cold nuclei extraction buffer (NEB) for 10 min on ice. Nuclei are centrifuged at 600 x g for 5 minutes at 4°C.
[0209] B. Nuclei extraction followed by fixation: cell pellets are resuspended in ice- cold nuclei extraction buffer (NEB) for 10 min on ice. Nuclei are centrifuged at 600 x g for 5 minutes at 4 °C. Pelleted nuclei are then lightly fixed (with formaldehyde, paraformaldehyde, glutaraldehyde, disuccinimidyl glutarate (DSG), or some combination). As defined in Paragraph A above, “light fixation” as used herein means using a 0.1- 00B206.1663
[0210] 0.0001% fixative for 1-20 minutes followed by 0.125 M glycine to quench fixation.
[0211] C. Nuclei extraction only: The K562 cell pellets are resuspended in ice-cold nuclei extraction buffer (NEB) for 10 min on ice. The nuclei are centrifuged at 600 x g for 5 minutes at 4 °C.
[0212] For all 3 steps, the NEB buffer used is 20 mM HEPES pH 7.9, 10 mM KC1, 0.1% Triton X-100, 20% glycerol, 0.5 mM spermidine, lx Roche cOmplete™, Mini, EDTA- free Protease Inhibitor.
[0213] Nuclei from A, B, and C above are either (1) cryopreserved at -80 or (2) go straight into CUT&RUN or CUT&Tag protocols below:
[0214] For CUT&RUN, materials referenced below are from CUTANA™ ChIC / CUT&RUN Kit Epicypher Cat# 14-1048, unless indicated otherwise. CUT&Run and CUT&Tag use a substrate upon which cells or cell nuclei can be immobilized. An example of such a substrate or surface is concanavalin A beads. However, it is noted that other means of cell or cell nuclei immobilization on a substrate (e.g., other magnetic beads) can be substituted for a concanavalin A bead. Not using beads as the substrate or any substrate would require centrifugation between the wash steps, which would likely lead to less robust signal / results. In these examples, concanavalin A beads are washed with Bead Activation Buffer and incubated with nuclei. The bead-bound nuclei are washed and incubated with the primary antibody (MRE11) in antibody buffer at 4 °C overnight on a nutator (VWR Nutating Mixer, Cat# 82007- 202; one-speed 24 rpm). The next day, wash the bead-bound nuclei+antibody with cell permeabilization buffer twice. Add pAG-MNase (to each reaction and incubate for 10 min. at room temperature (RT) (CUTANA™ ChIC CUT&RUN Kiy EpiCypher, Cat. #14-1048). The bead-bound nuclei+antibody+pAG-MNase are washed with the cell permeabilization buffer twice. MNase is activated by adding calcium chloride to a final concentration of 2 mM and incubating at 0 °C for 2 hours. The reaction is stopped by adding stop buffer and incubating at 37 °C for 10 minutes according to manufacturer instructions to release the cleaved DNA fragments. The cleaved DNA is then purified and from that, a DNA library is prepared.
[0215] For CUT&Tag, Concanavalin A beads are washed with bead activation buffer and incubated with the K562 nuclei. Bead-bound nuclei are washed and incubated with the primary antibody (MRE11) in antibody buffer at 4 °C overnight on a nutator (VWR Nutating Mixer, Cat# 82007-202; one-speed 24 rpm). Wash bead-bound nuclei+antibody with Washl 50 Buffer twice. Anti-rabbit secondary antibody (CUTANA™ Cat# 13-0047 and found in CUTANA™ CUT&Tag Kit Epicypher Cat# 14-1102; CUTANA™ pAG-Tn5 for CUT&Tag 00B206.1663
[0216] Cat# 15-1017 is sold separately; anti-rabbit secondary antibody to the MRE11 primary antibody) is added for signal amplification. The bead-bound nuclei+primary antibody is incubated with the secondary antibody for 30 min. at room temperature (RT). The bead-bound nuclei+antibody is then washed with Washl50 Buffer twice. pAG-Tn5 (CUT ANA™ CUT&Tag Kit Epicypher Cat# 14-1102, or sold separately as CUTANA™ pAG-Tn5 for CUT&Tag Cat# 15-1017) is added to each reaction and incubated for 1 hour atRT in Wash300. Wash the bead-bound nuclei+antibody+pAG-Tn5 with Wash300 Buffer twice. Activate the Tn5 by adding magnesium chloride to a final concentration of 10 mM and incubate at 37 °C for 1 hour (in Tagmentation Buffer). Add SDS release buffer and incubate at 58 °C for 1 hour to quench the tagmentation reaction. SDS Quench Buffer is added to neutralize the SDS present in the SDS release buffer. The next step would be the library preparation and indexing step.
[0217] The CUT&RUN buffers used are those of EpiCypher and are as follows and used the EpiCypher instructions unless indicated otherwise:
[0218] Bead Activation Buffer: 20 mM HEPES, pH 7.9, 10 mM KC1, 1 mM CaCh, 1 mM MnCh, and filter sterilize the buffer. Buffer can be stored at 4 °C for up to 6 months. Pre-Wash Buffer: 20 mM HEPES at pH 7.5, 150 mM NaCl, and filter sterilize the buffer. Store at 4 °C for up to 6 months.
[0219] Wash Buffer: Take the Pre-Wash Buffer and add 0.5 mM Spermidine* lx Roche cOmplete™, Mini, EDTA-free Protease Inhibitor (CPI, 1 tab / lOmL, Roche Cat. #11836170001) and filter sterilize the buffer. This buffer can be stored at 4 °C for up to 1 week.
[0220] Digitonin Buffer: Take the Wash Buffer (above) and add 0.01% digitonin. Prepare fresh each day and store at 4 °C as necessary.
[0221] Antibody Buffer: Take the Digitonin Buffer from above and add 2 mM EDTA. This should be prepared for use on the same day and stored at 4 °C as necessary.
[0222] Stop Buffer: 340 mM NaCl, 20 mM EDTA, 4 mM EGTA, 50 pg / mL RNase A, 50 pg / mL glycogen, and filter sterilize the buffer. The buffer can be stored at 4 °C for up to 6 months.
[0223] The CUT&Tag buffers used were prepared and used according to Epicypher instructions as follows:
[0224] Nuclear Extraction (NE) Buffer: (200 Wi reaction), 20 mM HEPES-KOH, pH 7.9; 10 mM KC1, 00B206.1663
[0225] 0.1% Triton X-100, 20% glycerol, 0.5 mM spermidine, lx Roche cOmplete™, Mini, EDTA-free Protease Inhibitor. After spermidine and CPI are added, the buffer can be stored at 4 °C for up to 1 week. NE buffer without spermidine and CPI is stable at 4 °C for up to 6 months.
[0226] Bead Activation Buffer'. (211 pL / reaction) 20 mM HEPES, pH 7.9; 10 mM KC1, 1 mM CaCh, 1 mM MnCh. The bead activation buffer is filter sterilized; it can be stored at 4 °C for up to 6 months.
[0227] Washl50 Buffer: For this buffer, which is also used to prepare Digitoninl50 Buffer, the buffer contains 20 mM HEPES, pH 7.5; 150 mM NaCl, 0.5 mM spermidine, lx Roche cOmplete™, Mini, EDTA-free Protease Inhibitor. The buffer is filter sterilized and can be stored at 4 °C for up to 1 week.
[0228] Digitoninl 50 Buffer: (450 pL / reaction) contains Washl50 Buffer + 0.01% Digitonin. Prepare fresh at the time of use and store at 4 °C.
[0229] Antibody 150 Buffer: (50 pL / reaction) contains Digitoninl 50 Buffer and 2 mM EDTA. Prepare fresh at the time of use and store at 4 °C.
[0230] Wash300 Buffer: This buffer is used to prepare Digitonin300 and Tagmentation Buffers. This buffer contains 20 mM HEPES, pH 7.5, 300 mM NaCl, 0.5 mM Spermidine, lx Roche cOmplete™, Mini, EDTA-free Protease Inhibitor (CPI-mini, 1 tab / 10 mL). Filter sterilize the buffer, which can be stored at 4 °C for up to 1 week.
[0231] Digitonin300 Buffer: (450 pL / reaction) contains Wash300 Buffer and 0.01% Digitonin. The Digitonin300 Buffer is to be prepared fresh for each day and stored at 4° C.
[0232] Tagmentation Buffer: (50 pL / reaction) contains Digitonin300 Buffer and 10 mM MgCh. The buffer can be stored at 4°C for up to 1 week.
[0233] TAPS Buffer: (50 pL / reaction) contains 10 mM TAPS, pH 8.5, and 0.2 mM EDTA. It can be stored at Room Temperature (RT) for up to 6 months. TAPS stands for N- [Tris(hydroxymethyl)methyl]-3-aminopropanesulfonic acid, [(2-Hydroxy-l, 1- bis(hydroxymethyl)ethyl)amino]-l-propanesulfonic acid.
[0234] SDS Release Buffer: (5 pL / reaction) contains 10 mM TAPS, pH 8.5, and 0.1% SDS (sodium dodecyl sulfate). It can be stored at RT for up to 6 months.
[0235] SDS Quench Buffer'. (15 pL / reaction) contains 0.67% Triton-X 100 in Molecular grade H2O and is stored at RT for up to 6 months. 00B206.1663
[0236] The following examples illustrate how the modified CUT&RUN and CUT&Tag systems can work and the data they can produce in conjunction with the various inhibitors used. The examples should not be viewed as limiting any of the claims.
[0237] EXAMPLE A - DNA Repair Factors
[0238] The initial DISCOVER-Seq technique (Wienert, et al. "Unbiased detection of CRISPR off-targets in vivo using DISCOVER-Seq." Science 364.6437 (2019): 286-289; “Wienert et al. (2019)”), reliant on ChlP-seq, identified MRE11 as the most robust repair factor involved in the repair of double-stranded breaks. Examination of other repair factors revealed that NBS1, RAD70, and Ku70 were found to be ineffective for off-target (OT) assessment using ChlP-seq due to the low coverage as demonstrated and discussed in Wienert et al., (2019) e.g., see Wienert’ s Figure 1.
[0239] However, in the modified method described herein, the data produced using MRE1 1 and other factors along with CUT&Tag were remarkably robust (Table 2 below) in contrast to data obtained when using ChlP-seq (Table 1). Notably both NBS1 and MRE11 produced robust results using the CUT&Tag method.
[0240] EXAMPLE B - Cas Titration of Cell Count Using Modified CUT&RUN
[0241] This experiment sought to establish the lower limit of detection for low- frequency off-target events (those <1%) which would allow researchers using the method to scale to a 96-well plate format and towards a semi-automated workflow. The system then allows for the screening of more compounds and more guide RNAs. 00B206.1663
[0242] In this experiment, 4xl06K562 cells were electroporated using the crRNA of
[0243] Table 3.
[0244] In a second tube, 8xl06K562s + a VEGFA gRNA + Cas9RNP were electroporated to allow entry into the K562 cells. NTC stands for “non-targeting control”. The Cas9-sgRNA complex was incubated with the cells in a 3:1 ratio (180 pmol sgRNA* and 60 pmol Cas9). The crRNA sequences of the sgRNA used in the experiment are displayed in Table 3. The reference genome is hg38 (human genome), which can be located using the University of California Santa Cruz genome browser gateway, for example. The sequences in Table 3 (below) appear in a 5’ to 3’ orientation. 00B206.1663
[0245] Cells were plated from these tubes at IxlO6cells per well of a 12-well plate in growth media (RPMI-1640, 10% HI FBS, and 2 mM L-glutamine). DMSO (negative control) or 1 pM KU60648 (inhibitor) was added to two wells, and the experiment was performed in duplicate. To the cells with the VEGFA guide, either DMSO was added or IpM KU60648, with the experiment being run in triplicate (3 wells for each). This methodology was used for both the titration experiments and 6- and 8-hour time course experiments discussed herein.
[0246] In this example, the DNA repair inhibitor used was KU60648 (Sigma SML 1257 - 5 mg). Cells were incubated under K562 cell culturing conditions and harvested 8 hours later for analysis.
[0247] At the time of cell harvest, the VEGFA and KU60648 group of cells were split into groups ranging from 10,000 cells to IxlO6cells. The 10,000 cell samples were analyzed in duplicate, whereas the others were not. RPM for off-target sites was calculated as a mean average of the top 8 off-target sites listed in Table 9. The data is as follows:
[0248] TABLE 4
[0249] EXAMPLE C - Cas9 Time Course
[0250] In this experiment, the method was analyzed to determine the advantageous time intervals for off-target assessment. The time course was analyzed over 6 hours to 16 hours.
[0251] The method using DISCOVER-Seq (Weinert et al., “Unbiased detection of CRISPR off-targets in vivo using DISCOVER-Seq,” Science 364: 286-289, 2019) was known to function best in the 8- to 24-hour timeframe and not as well at 4 hours or 24 hours. The Weinert et al. (2019) method also is sometimes referred to as ChlP-Seq and uses MRE11. DISCOVER-Seq was modified and became DISCOVER-Seq+ as published by Zou et al. (“Improving the sensitivity of in vivo CRISPR off-target detection with DISCOVER-Seq+,” 00B206.1663
[0252] Nature Meth. 20: 706-713, 2023). The Zou et al. (2023) method uses the DNA repair inhibitor KU60648. However, the DISCOVER-Seq+ system still requires about 12xl06cells for signal detection, similar to the number of cells required for signal detection when using DISCOVER- Seq. The number of cells needed for the analysis limits the use of the assay only to a system wherein adequate numbers of cells are available. The improvement of Zou et al. (2023) over Weinert et al. (2019) was that Zou’s method determined that 12 hours was the optimum time point for cell harvest after Cas9 delivery.
[0253] The KU60648 inhibitor is described as a dual inhibitor of DNA-PK and also PI3Ka, PI3KP, and PI3K5. Zou’s DISCOVER-Seq+ system had a substantially improved signal by using the KU60648 as the DNA repair inhibitor over the data obtainable using the older DISCOVER-Seq system (Zou et al., 2023).
[0254] When using Zou’s DISCOVER-Seq+ method, the off-target RPM assessments were made at 6, 12, and 24 hours after Cas9 delivery as described by Zou et al. (2023). Analysis was conducted using HEK293T cells with the VEGFA site 2 (5’ -gaccccctccaccccgcctc-3 ’ , SEQ ID NO: 13) and either in the presence of the DNA repair inhibitor KU60648 or in the absence of an inhibitor.
[0255] The modified method disclosed in this example used additional time points at 6, 8, and 14 hours. This experiment also utilized K562 cells instead of HEK293T cells. HEK293T cells are generally grown as an adherent monolayer, whereas K562 cells are suspended cells (non-adherent). However, it is expected that the modified method used in the experiment would yield similar results if the cells used were HEK293T cells despite being a cell line cultured as a monolayer in contrast to K562 cells which are non-adherent cells.
[0256] In this experiment, the cell time points analyzed were at 6, 8, 14, and 16 hours after Cas9 delivery to the K562 cells. K562 cells were maintained in growth media (i.e., RPMI- 1640, 10% HI FBS, 2 mM L-glutamine) at a concentration of IxlO6cells / mL until the time of electroporation. The K562 cells were electroporated using Lonza’s 4D-Nucleofector® Core Unit by centrifuging the K562 cells at 300x g for 5 minutes to pellet, washing in PBS, and resuspending the cells in the appropriate volume of Nucleofector™ Solution and Supplement per manufacturer’s instructions (e.g., 100 pL for a single nucleocuvette (82 pL Nucleofector™ Solution + 18 pL Supplement / reaction)). To the single nucleocuvette, 15 pL RNP complex was added. Cuvettes were then run using the specified settings for K562 cells on Lonza's 4D- Nucleofector® Core Unit following the manufacturer’s guidelines for cell permeabilization. Following nucleofection, cells were plated in growth media with or without a DNA repair inhibitor. 00B206.1663
[0257] For the CUT&RUN method described herein, on the morning of Day 1, 4xl06K562 cells + NTC or 8xl06K562 cells with VEGFA were electroporated and plated. Cas9- sgRNA complex was used again in a 3: 1 ratio (180 pmol sgRNA and 60 pmol Cas9). The DNA repair inhibitor KU60648 was used at 1 pM and the control was DMSO. Cells were harvested 6 hours or 8 hours after Cas9 delivery.
[0258] For the 14 and 16-hour time points, the cells were prepared and treated as follows. In the afternoon of Day 1, cells were electroporated as described above using the same 3: 1 ratio of the Cas9-sgRNA complex. The cells were plated as IxlO6K562 cells + NTC per well or 2xl06K562 cells + VEGFA per well. The DNA repair inhibitor, KU60648, was added at 1 pM. DMSO was used as the control. KU60648 was obtained from Sigma, Cat. No. SML 1257-5MG.
[0259] KU60648 is soluble in DMSO and is not water-soluble or PBS-soluble. Thus DMSO is used as the negative control. The amount of DMSO in pL is the sample as the pL volume of the inhibitor. For example, for 1 pM KU60648, a 1 :2000 dilution of a stock is prepared by adding 5 pL to 10 mL culture media. The same amount of DMSO control of 5 pL DMSO to 10 mL culture media is used.
[0260] Cells were harvested the morning of day 2, at either 14 hours or 16 hours. Cell nuclei were extracted. Isolation of cell nuclei was performed using a nuclear extraction buffer (20 mM HEPES pH 7.9, 10 mM KC1, 0.1% Triton X-100, 20% glycerol, 0.5 mM spermidine, lx Roche cOmplete™, Mini, EDTA-free Protease Inhibitor). The isolated nuclei were then cryopreserved at -80 °C. RPM for off-target sites was calculated as a mean average of the top 8 off-target sites listed in Table 9.
[0261] It was observed that the later time points of 14 hours and 16 hours when in the presence of the DNA repair inhibitor resulted in lower cell viability in the case of KU60648.
[0262] The data from the experiment demonstrated unexpectedly that the shorter incubation period of 6 hours is advantageous for the modified method using CUT&RUN and pA-MNase (protein A micrococcal nuclease) over the longer time points recommended and described for use for DISCOVER-Seq+. This shorter time period is not suggested by the data previously obtained by Zou et al. (2023). While cell viability plays a role in some of the observed data (more cells alive at the shorter incubation period time points in the presence of KU60648 when used), it is likely that DNA repair kinetics play a major role. The earlier time points are sufficient for substantial DNA editing to have occurred (on-target hypothetically occurring faster / earlier than off-target due to kinetics effects of mismatches). By the time the 00B206.1663 later time points are taken, MRE11 has had more time to dissociate from break sites on the nucleic acid sequence.
[0263] TABLE 5
[0264] EXAMPLE D - Cas 12a Nuclease, K562 Cells
[0265] Pilot experiments suggest that Casl2a nuclease-engineered K562 cells would provide different peak signals from those of Cas9-engineered cells (see, e.g., FIG. 2B vs FIG. 2C and FIG. 26). Similar experiments to those performed with Cas9-engineered cells in the context of Casl2a by designing guides with suitably high off-target activity, engineering K562 cells with Casl2a and these high off-target guides, performing the methods disclosed herein to obtain NGS reads that indicate putative off-target sites, and verifying these sites with a targeted rhAmpSeq assay. Without a Casl2a result-trained CNN, simple read pile-ups at the predicted sites will be used to identify high-confidence sites, which can then be used to train a model as described herein (naive and / or a model already trained for Cas9, as a model so trained may be closer to the ultimate Cas 12a trained model than a completely randomly initialized model) and refine the model’s predictions. This will allow for bootstrapping CNN training and continuously / iteratively verifying its accuracy using the rhAmpSeq assay.
[0266] EXAMPLE E - Cas9 Nuclease with Human T Cells
[0267] In this example, human CD8+cells (hCD8+) were obtained as follows. CD8+ cells were freshly collected according to Yau et al., “Measuring the effect of drug treatments on primary human CD8+ T cell activation and cytolytic potential,” STARProtoc. 2(2): 100549 (Jun. 18, 2021, available at doi: 10.1016 / j.xpro.2021.100549).
[0268] CUT&RUN-based procedure was performed as described herein and as follows. Cells were split to 5xl06after electroporation using Lonza’s 4D-Nucleofector® Core Unit according to the manufacturer’s instructions. The Cas9-sgRNA complex is the same as described and used for Example C, except it now was used in a 5: 1 ratio (300 pmol sgRNA* and 60 pmol Cas9). 00B206.1663
[0269] Immediately after electroporation with the Cas9-sgRNA complex, the 5xl06cells were split across three conditions and plated in Prime XV (Fuji Film Irvine Scientific Inc. Cat. 911541L) supplemented with hIL-7 (25 ng / ml), hIL-15 (50 ng / ml) and the presence of either (a) 1 pM KU60648, (b) 3 pM AZD7648, or (c) DMSO (control) resulting in ~1.5xl06cells per condition. In addition, a no electroporation (“No EP”) control was included to access cell viability after incubation. All cells were incubated for 6 hours at 37 °C after plating.
[0270] After incubation, the cell nuclei were isolated using a nuclear extraction buffer (20 mM HEPES pH 7.9, 10 mM KC1, 0.1% Triton X-100, 20% glycerol, 0.5 mM spermidine, and lx Roche cOmplete™, Mini, EDTA-free Protease Inhibitor). Isolated nuclei were then cryopreserved at -80 °C.
[0271] The results demonstrated that the addition of AZD7648 as the DNA repair inhibitor relative to the control does not impact the T cell viability at the 6-hour collection time point as follows:
[0272] TABLE 6
[0273] . Both on-target and off-target sites performed as expected in the hCD8+T cells. The experiment was conducted using about 400,000 nuclei for each sample analyzed.
[0274] EXAMPLE F - Method Using Different DNA Repair Inhibitors
[0275] In this example, the method described in the specification analyzed a variety of DNA repair inhibitors alone and in various combinations. The impact of the DNA inhibitors and inhibitor combinations was assessed for on-target RPM across only compounds (1 site) and the mean off-target RPM across only compounds (assessment was based on the top 8 off- target sites listed in Table 9 and then taking the mean average). The on-target and off-target analysis was conducted at 6 hours after the addition of the DNA repair inhibitor. The time is started from the time the DNA repair inhibitor is added to cells. The DNA repair inhibitors are added at most 15 minutes from the time the cells are electroporated (permeabilized) with Cas9 nuclease and gRNA. 00B206.1663
[0276] Five DNA repair inhibitors were analyzed at various concentrations as follows: KU60648 (1 pM and 2 pM), AZD7648 (1 pM, 5 pM, and 10 pM), NU7026 (20 and 40 pM), NU7441 (2.5 and 10 pM and M3814 (20 and 40 pM). The EpiCypher CUTANA™ kit was used in all the experiments according to manufacturer instructions, from which DNA was extracted from the cells to prepare libraries and sequenced. “No EP” stands for no electroporation.
[0277] TABLE 7 00B206.1663
[0278] The results in the table above and depicted in FIG. 13 and FIG. 14 demonstrate that the AZD7648 DNA repair inhibitor functioned better for both on-target and off-target RPM at the 1 pM amount than the other DNA repair inhibitors. AZD7648 at 1 pM also worked better than the 5.0 pM amount. Combinations of two or three DNA repair inhibitors lessened the signal for both on-target and off-target RPMs as indicated in the table above. However, the DNA-PK inhibitors (e.g., AZD7648, NU7026, and NU7441) significantly boost the signal of the method as compared to results obtained by the method in the absence of a DNA repair inhibitor.
[0279] EXAMPLE G - Fixation optimization
[0280] In this example, signal strength at on- and off-target sites by RPM using a single guide RNA (VEGFA Site 2, sgRNA sequence with PAM: GACCCCCTCCACCCCGCCTCNGG, on-target cut site: chr6:43770822-43770841, with the cut site nominally centered on chr6:43770825, SEQ ID: 13) across different fixation conditions. Using on-target RPM as a proxy of signal strength, this experiment identified advantageous fixation conditions for the approaches described herein. Fixation can negatively impact CUT&RUN and CUT&Tag signals by crosslinking chromatin in a way that reduces antibody accessibility and enzymatic activity, often resulting in weaker signals. Surprisingly, the instant example indicates that when fixation is performed within a range (e.g., <1% at varying time durations) signal is enhanced, rather than weakened. This unexpected 00B206.1663 improvement is indicative that low-level fixation may stabilize DNA repair factor interactions with DNA and / or preserve the integrity of DNA-protein complexes, thereby promoting more effective target recognition and signal generation.
[0281] Experiments were conducted using Cas9 nuclease (as RNP or mRNA) unless otherwise stated. Cells were subjected to fixation using paraformaldehyde (PF A) and other fixatives at a range of concentrations between 0.0001% and 0.1% (unless otherwise stated) prior to CUT&RUN or CUT&Tag. Specifically, fixation conditions included: 0.05% PFA for 1 minute, 0.01% PFA for 5 minutes, 0.001% PFA for 5 minutes, and 0.001% PFA for 10 minutes. In some experiments, cells were pre-treated with 0.005% glutaraldehyde (G) followed by fixation with PFA. In some experiments, cells were pre-treated with either 0.2 mM or 1 mM disuccinimidyl glutarate (DSG) followed by fixation with PFA. These conditions allowed for the evaluation of various crosslinking protocols to optimize performance for downstream analyses. CUT&RUN and CUT&Tag experiments were performed using MRE11 according to the manufacturer’s instructions (Epicypher) and sequencing libraries were generated. RPM is calculated as reads per million reads sequenced and normalizes samples to their sequenced depth of coverage. On-target RPM is calculated by summing the number of reads overlapping with the on-target cut site, in this case for VEGFA Site 2, inclusive of a 1001 bp window centered on the cut site. Off-target RPM is calculated by the same approach as on-target RPM with stacked reads but instead of one on-target site the mean of the top 30 off-target sites (sourced from Wienert et al., 2019, listed in Table 2) is calculated for the VEGFA Site 2 sgRNA (unless otherwise noted). Resulting data is presented in Table 8 as well as FIG. 7 and FIG. 8.
[0282] Table 8: on- and off-target RPM values with a collection of fixation conditions and durations. 00B206.1663
[0283] Table 9: off-target sites for VEGFA Site 2 sgRNA and associated indel percentages (as described in Wienert et al., 2019). 00B206.1663
[0284] EXAMPLE H - Exemplary Computational Analysis
[0285] A computational analysis for identifying off-target (OT) sites from NGS data of undigested DNA from the methods disclosed herein was developed. The computational analysis comprises aligning the NGS data from experiments as disclosed herein to an hg38 human genome as a reference genome using an aligner (Minimap2). Similar alignment programs that align DNA or mRNA sequences against a reference genome / sequence can be substituted for Minimap2 and are expected to yield similar results (e.g., BWA or Bowtie2). The resulting BAM file was input into a custom computational pipeline, which operated as follows:
[0286] • Reads from the BAM file were used to compute sequencing coverage across the reference genome. Peaks were called on the aligned reads using simple thresholds for coverage. Reads covering >200 bp were considered peaks, with an allowable gap of 5 bp between two peaks. This coverage was chosen to bias towards including more peaks, lowering the false negative rate. The result of the coverage thresholding was a series of 1 -dimensional vectors (referred to herein as signals) representing coverage. Each value in the vector represents a read depth of the aligned sequences at a given location of the reference genome, with 0 indicating no coverage (no sequences aligned to that location on the reference genome). FIG. 3A shows an example signal.
[0287] • Following this conversion from BAM file to signal, read intervals were expanded by 2x, resulting in an accentuation of intervals between reads that contain relatively few reads in the produced 1 -dimensional signals. The BAM file indicates a number of reads that map to intervals in the genome (e.g. chr6: 10, 000, 000-10, 000, 150), and these read intervals were expanded by 2x (chr6:20, 000, 000-20, 000, 300). Assuming, as an example, a cut site at chr6: 10,000,150, whereas before expansion there were reads on the other side of the cut site (chr6:10,000,151-chr6: 10,000,300), now there is a 1 base pair gap after the expansion at chr6:20,000,301. In aggregate, this phenomenon can be seen for example at the starred point (equivalent to 00B206.1663 chr6: 10,000, 150 / 1 in the preceding text example) in Fig. 5A where there is no dip in signal at the middle of the plot that goes to 0. After the expansion of the intervals in Fig. 5B (note the x-axis change) there is one point that goes all the way to 0 (equivalent to chr6:20,000,301 in the preceding text example). FIG. 5B shows the result of expanding the read intervals.
[0288] • Feature extraction was performed by applying a kernel of [1, -2, 1] on all signals, resulting in a convolved signal (FIG. 5C). In general, a kernel emphasizing / extracting edge features may be applied.
[0289] • Z-score normalization was applied to this convolved signal (FIG. 5D). Z- score normalization was applied to prepare the convolved signal for downstream training and application of our CNN.
[0290] • The Z-score normalized convolved signal was fed into a CNN, structured as illustrated in FIG. 4 and Table 4 below, and trained as described in Example I, below. The CNN included a series of convolutional layers interspersed with max pool layers before being flattened and passed through a series of dense layers and finally a softmax layer. For each convolution layer, a rectifier (reLU) layer was used as an activation layer, but another non-linear activation function may be used, such as a tanh, sigmoid, etc. In an example, the CNN included a series of 5 convolutional layers alternating with 5 max pool layers and 3 dense layers, as specified in Table 10 below:
[0291] TABLE 10 00B206.1663
[0292] • The CNN output soft-maxed values indicated a probability as to whether a signal corresponding to a location in the reference genome represents a cut site. TABLE 11 shows a portion of the exemplary output of the CNN that was obtained. The CNN output included, as shown in TABLE 11, locations of putative cut sites and corresponding soft-maxed values indicating probabilities that the identified locations are cut sites.
[0293] TABLE 11. EXAMPLE OUTPUT OF CNN Table 11 further includes a column indicating whether the putative cut site is a false positive (FP) or true positive (TP). The CNN outputs are unbiased in vitro nominations for CRISPR-derived off-target sites and are biased towards minimizing false negatives. This 00B206.1663 bias can result in a large number of false positives, which will be filtered out in downstream quantification of editing efficiency using NGS-based amplicon sequencing (rhAmpSeq™ panels).
[0294] RhAmpSeq panels will be performed on the CNN-identified off-target sites, and the number of insertions / deletions (indels) at the identified off-target sites will be used to determine the relative frequency of CRISPR activity at each site. False positives will be filtered out based on associated indel rates. RhAmpSeq panels are typically performed in replicate on edited cells and unedited control cells. Indel rates can be calculated from both edited and control cell replicates and compared using a T-test, Mann-Whitney U test, or Fisher’s exact test. Assuming sufficient statistical power, false positives can be removed with a desired degree of confidence by failure to reject the null hypothesis according to any of the previous tests. If indel rates are statistically higher for edited cells than unedited control cells, the identified site is a true positive off-target site, but if the indel rates are not statistically higher for edited cells than unedited control cells, then the identified site is a false positive (from the unbiased CRIS technique).
[0295] EXAMPLE I
[0296] The CNN was trained using Flux.jl, a machine learning package written in Julia, using experimentally generated data such as that from the above examples or a previously studied guide RNA and Cas combination (e.g., VEGFA Site 2 and Cas9). K562 cells were used for the data generation but the approach is cell type agnostic. The experimentally generated data was converted into 1-dimensional signals as previously described (e.g., aligned to the reference genome, thresholded for areas of sufficient coverage, expanded to reveal breaks between reads, passing over an edge-detection kernel one or more times, normalization such as z-score normalization, etc.). The 1 -dimensional signals were split into positive examples of cut sites as shown by previous studies (e.g., previous studies of VEGFA Site 2, taking the strongest hits from those papers as consensus to avoid too many false positives) or negative examples of cut sites that still contained sufficient coverage to be called a signal. Example references used for identifying positive examples include Tsai, Shengdar Q., et al. "GUTDE-seq enables genomewide profiling of off-target cleavage by CRISPR-Cas nucleases." Nature Biotechnology 33.2 (2015): 187-197; Wienert, Beeke, et al. "Unbiased detection of CRISPR off-targets in vivo using DISCOVER-Seq," Science 364(6437): 286-289 (2019); and Zou et al., "Improving the sensitivity of in vivo CRISPR off-target detection with DISCOVER-Seq+," Nature Methods 20(5): 706-13 (2023). 00B206.1663
[0297] The ratio of positive examples to negative examples was 578:2382 with 2978 total samples, a value that was bolstered by synthetic data generated from the positive examples. In particular, 108 positive sample data points were bolstered / augmented by synthetic samples acquired by signal reversal (2x) and shifting the signal up or down by 200 to 400 base pairs (3x). This would be 648 positive data points, but 70 data points were not included due to quality control (e.g. too few reads after shifting), resulting in the 578 positive samples. This synthetic data generation entailed transformations such as signal reversal, shifting of the signal up and down by 200 to 400 base pairs, and down sampling of the signal with Gaussian noise.
[0298] After splitting the data into positive and negative signals, the data was further split into train and test datasets at an 80:20 ratio for both the positive signals and the negative signals. Care was taken to not allow synthetic data from one sample split into the training dataset to also fall into the test dataset, thereby avoiding possible data leakage.
[0299] Positive and negative datasets were one-hot encoded to create a binary classification problem as follows. Positive and negative datasets were encoded as in a 2Xn matrix, where a 1 in a first row of a first column indicates that the first data point is a true positive (e.g., that it represents a true off-target cut site) and a 1 in the second-row first column indicates that the first data point is false positive (e.g., that it does not represent a true off-target cut site). There were n columns representing each of the n data points (in this case 2978). Each column has a single 1, so in this binary setting, where data points can only be positive or negative, each column also has a single 0. The model discriminates between data points that are true positive or false positive and therefore solves a binary classification problem. This model contrasts with the other approaches to training the model, which would be to use a regression and provide probabilities as ground truth for the model to be trained against.
[0300] The training dataset was then used to train the CNN previously described by b ackpropagation using the derived gradients of the weights which were updated for a binary cross entropy loss function using the Adam optimizer. Due to the small size of the model (<100,000 parameters or weights, 94,926 in the specific CNN described in Example H) and the small number of data points, the process of training was relatively quick, on average < 5 minutes with a GPU (graphics processing unit) and small enough that the CNN can be trained using a CPU (central processing unit). The small size of the model also allows for rapid forward propagation and rapid classification of peaks, with 4800 peak signals taking on average ~20 seconds to be processed by the CNN model using a CPU. 00B206.1663
[0301] PROPHETIC EXAMPLE 1 - Estimated maximum likelihood and ranking
[0302] Due to the known sequence similarity between a CRISPR-Cas9 guide sequence and off-target cut sites, we can use sequence alignment and in silico prediction to bolster our confidence in a given peak determined to be a cut site by our approach (e.g., output cut-site identified by our CNN). This will entail taking peak predictions from the CNN model and creating a secondary model that assesses the likelihood of an off-target cut site hit given: the CNN model’s output (location of the identified off-target cut site hit, optionally the softmax output from the model) and best alignment of the known guide to the underlying genomic sequence of the peak associated with the location. This secondary model may be a linear regression or logistic regression model fit to the data for which there is a known ground truth. Correlating alignment of the known guide to sequences of / near identified off-target cut site hits from the CNN model disclosed herein will allow for building confidence intervals around the identified off-target cut site hits, which would in turn allow for ranking the priority of these off-target sites, as well as selecting desired threshold for false positives in follow-up targeted amplicon-based NGS sequencing panels.
[0303] While various embodiments have been described above, it should be understood that they have been presented by way of example only, and not limitation. Where methods and / or schematics described above indicate certain events and / or flow patterns occurring in certain order, the ordering of certain events and / or flow patterns may be modified. While the embodiments have been particularly shown and described, it will be understood that various changes in form and details may be made. Additionally, certain of the steps may be performed concurrently in a parallel process when possible, as well as performed sequentially as described above. Although various embodiments have been described as having particular features and / or combinations of com-ponents, other embodiments are possible having any combination or subcombination of any features and / or components from any of the embodiments described herein. Furthermore, although various embodiments are described as having a particular entity associated with a particular compute device, in other embodiments different entities can be associated with other and / or different compute devices.
[0304] EXAMPLE J
[0305] In this example, K562 cells were used. Cells were centrifuged at 300 x g for 5 minutes, washed with IX PBS prior to nucleofection with Cas9 (as RNP or mRNA) and gRNA using one of three formats: cuvette, 16-well Nucleocuvette Strip, or 4D-Nucleofector® 96-well plate, with the Lonza program FF-120 for K562 cells. Post-nucleofection, cells were transferred 00B206.1663 to RPMI medium containing an inhibitor (1 pM AZD7848 unless otherwise stated) and incubated at 37 °C with 5% CO2 for 2 to 48 hours. Cells were harvested and processed by (A) fixation followed by nuclei extraction, (B) nuclei extraction followed by fixation, or (C) nuclei extraction only. Cells or nuclei were cryopreserved at -80 °C for downstream analysis using EpiCypher’s CUT ANA™ kit according to manufacturer instructions.
[0306] As illustrated in FIG 9, the coverage plot shows a genomic locus centered on the on-target cut site (vertical black dotted line) + / - 250 bp (Chr6:43770576-43771075, GRCh38) with the VEGFA gene body below (horizontal black bar). This example demonstrates that the strategies disclosed herein, e.g., a modified CUT&Tag-based approach that employs Cas9 nuclease, resulted in significant and detectable signals.
[0307] EXAMPLE K
[0308] In this example, samples employing Casl2 nuclease and a Casl2 guide (e.g., VEGFA_TTTV_3) sequence is ctaggaatattgaagggggc, were successfully assayed. Samples starting with "6H PFA" were fixed for 1 minute with either 0.1% or 1% PFA. Casl2 samples were treated with DMSO vehicle, KU60648, or AZD7648. Following treatment, cells or nuclei were cryopreserved at -80 °C for downstream analysis using EpiCypher’s CUT ANA™ kit according to manufacturer instructions.
[0309] As illustrated in FIG. 10, a coverage plot of a Cast 2 (As variant)-based CUT&RUN approach is depicted showing genomic locus Chr6:43, 769, 752-43, 769, 937 (GRCh38) centered on on-target cut site, which evidenced significant and detectable signals.
[0310] EXAMPLE L
[0311] The following example is directed to the use of DNA repair factors in onnection with the DNA binding protein-targeted enzymatic assays of the present disclosure. In this example, following the introduction of Cas9 nuclease and guide RNA to cells and collection after 2 to 48 hours, cells were fixed using 0.005% PFA for 5 minutes. CUT&Tag was performed to analyze the DNA repair factors ATM, NBS, MRE11, Ku70, Ku80, and FANCD2. Concanavalin A beads were washed with a bead activation buffer and incubated with the nuclei. The bead-bound nuclei were washed and incubated with the primary antibody specific to one of the repair factors (ATM, NBS, MREl 1, KU70, KU80, or FANCD2) in an antibody buffer at 4 °C overnight on a nutator. Downstream experimental processing was performed using EpiCypher’s CUT ANA™ kit according to manufacturer instructions. The subsequent steps involved library preparation and indexing for sequencing and analysis 00B206.1663
[0312] As illustrated in FIG. 11 (on-target) and FIG. 12 (off-target), RPMs per sample for the on-target and off-target sites corresponding to VEGFA Site 2 sgRNA, calculated as previously described, with different repair factors, where the underlying data is provided in Table 12.
[0313] Table 12.
[0314] EXAMPLE M
[0315] In this example, HEK293 cells were used as previously described. Aligned reads were visualized surrounding the on-target cut site in the HEK293 cell line, using the VEGFA Site 2 guide and a IDT V3 Cas9 RNP for the Cas9 nuclease.
[0316] As illustrated in FIG. 15, which is a coverage plot showing the genomic locus centered on the on-target cut site (vertical black dotted line) + / - 250 bp (Chr6:43770576- 43771075, GRCh38) with the VEGFA gene body below (horizontal black bar), this strategy resulted in significant and detectable signals.
[0317] EXAMPLE N
[0318] In this example, freshly collected human CD8+ T cells (hCD8+) were obtained, and cells were electroporated using Lonza’s 4D-Nucleofector® Core Unit according to the manufacturer’s instructions. The Cas9-sgRNA complex was prepared at a 5: 1 ratio (300 pmol sgRNA and 60 pmol Cas9) and used for nucleofection. After electroporation, 5 * 106cells were split across three conditions (1.5 x 106cells per condition) and plated in Prime XV medium supplemented with 25 ng / ml hIL-7 and 50 ng / ml hIL-15 in the presence of either 1 pM 00B206.1663
[0319] KU060648, 3 pM AZD7648, or DMSO (control). Additionally, a no-electroporation control was also included. Cells were incubated for 6 hours at 37 °C after plating and cryopreserved at -80 °C for downstream analysis using EpiCypher’s CUT&RUN CUT ANA™ kit according to manufacturer instructions
[0320] As illustrated in FIG. 16, which is a coverage plot showing genomic locus centered on the on-target cut site (vertical black dotted line) + / - 250 bp (Chr6:43770576- 43771075, GRCh38) with the VEGFA gene body below (horizontal black bar), this strategy resulted in significant and detectable signals. In addition, FIG. 17 depicts a barplot showing the on-target RPM values, calculated as previously described, for human CD8+ T cells with and without DNA-PK inhibitors, vehicle control, and nucleofection negative control. This figure corresponds to the underlying data in Table 13.
[0321] Table 13
[0322] EXAMPLE O
[0323] In this example, human induced pluripotent stem cells (iPSCs) were used. Cells were electroporated using Lonza’s 4D-Nucleofector® Core Unit according to the manufacturer’s instructions. After electroporation, 5 / I O cells were split across three conditions (1.5 x 106cells per condition) and plated in TeSR™-E8™ medium (Stemcell Technologies) supplemented with 10 pM Y-27632 (ROCK inhibitor) in the presence of either 1 pM KU060648, 3 pM AZD7648, or DMSO (control). A no electroporation control as well as a non-targeting guide control (NTC) were included. Cells were incubated for 6 hours at 37 °C after plating and cryopreserved at -80 °C for downstream analysis using EpiCypher’s CUT ANA™ kit according to manufacturer instructions
[0324] As Illustrated in FIG. 18, which is barplot showing the on-target RPM values, calculated as previously described, for iPSCs with and without DNA-PK inhibitors, vehicle control, and nucleofection negative control, this strategy resulted in significant and detectable signals. This figure corresponds to the underlying data in Table 14
[0325] Table 14. 00B206.1663
[0326] EXAMPLE P
[0327] In this example, the sensitivity of the methods disclosed herein are evaluated with respect to two senses of the word: (1) the ability to detect off-target sites with low cell number (“low cell number titration”); and (2) the ability to identify off-target sites not identified by prior art (“identify novel off-target sites”). For the first sensitivity, RPM was employed for on- and off-target sites as previously described, calculated for experiments in which the number of cells was titrated down, demonstrating the maintenance of on- and off-target signal despite lower cell number (see, e.g., FIG. 19 through FIG. 23 and Tables 15-16). For the second measure of sensitivity, we provide examples of off-target sites not identified by prior art that our approach identified (FIG. 24).
[0328] For CUT&Tag-based titration experiments, 10e6 K562 cells were electroporated with Cas9-sgRNA complexes targeting VEGFA Site 2 (SEQ ID NO: 13). After electroporation, cells were plated with either DMSO (negative control) or 1 pM KU60648. Upon collection between 2 to 48 hours, cells were fixed using 0.005% PFA for 5 minutes. CUT&Tag was then performed to analyze the DNA repair factors ATM or MRE11, with cell groups ranging from 5,000 cells to 2e6 cells. For CUT&RUN titration experiments, 4e6 K562 cells were electroporated Cas9-sgRNA complexes targeting VEGFA Site 2 (SEQ ID NO: 13). Cells were plated from these tubes at le6 cells per well of a 12-well plate in growth media (RPMI-1640, 10% HI FBS, and 2 mM L-glutamine). DMSO (negative control) or 1 pM KU60648 (inhibitor) was added to two wells, and the experiment was performed in duplicate. To the cells with the VEGFA guide, either DMSO was added or IpM KU60648, with the experiment being run in triplicate (3 wells for each). This methodology was used for both the titration experiments and 6- and 8-hour time course experiments discussed herein. In this example, the DNA repair inhibitor used was KU60648 (Sigma SML 1257 - 5 mg). Cells were incubated under K562 cell culturing conditions and harvested 8 hours later for analysis. At the 00B206.1663 time of cell harvest, the VEGFA and KU60648 group of cells were split into groups ranging from 10,000 cells to le6 cells. The 10,000 cell samples were analyzed in duplicate, whereas the others were not.
[0329] Titrated sample replicates were aggregated by taking the mean of the RPM values for on- / off-target sites for a given cell number. The background signal RPM was determined by taking the mean of the given on- / off-target sites for two control samples (100,000 cells each) electroporated with Cas9 and on which CUT&Tag was performed with the ATM antibody.
[0330] For identifying novel off-target sites, experiments were conducted using Cas9 nuclease (as RNP or mRNA) unless otherwise stated. In this example, K562 or HEK293 cells were electroporated (EP) with Cas9-sgRNA complexes using Lonza’s 4D-Nucleofector® Core Unit according to the manufacturer’s instructions. No inhibitors were used in this experiment. After electroporation, cells were plated in appropriate growth media and incubated at 37 °C with 5% CO2 for 96 hours. Following the incubation period, genomic DNA was extracted from the cells, and specific locus were amplified using PCR to verify model-determined off-target sites. Indels were subsequently counted using custom in-house software and the difference between incubated and control samples was reported to determine CRISPR activity at off-target sites.
[0331] As illustrated in FIG. 19 (on-target) and FIG. 20 (off-target), target site RPM values, calculated as previously described, were plotted for titrated cell counts from 2,000,000 to 5,000 cells in a CUT&Tag experiment. Horizontal dotted line shows background signal level as mean of two control samples evaluated at same off-target sites. RPM values decrease with cell number, but not in a linearly proportional manner. As illustrated in FIG. 21 and 22, the ratio of RPM at the on-target (FIG. 21) or off-target (FIG. 22) site to cell number provides an indication of efficiency, in that at lower cell numbers, each cell is able to provide a disproportionately higher contribution to signal strength. The data underlying FIG. 19 though FIG. 22 is included in Table 15.
[0332] Table 15 00B206.1663
[0333] As illustrated in FIG. 23, RPMs for selected off-target VEGFA Site 2 sites, representing a range of indel % values. As cell number decreases, off-target sites lose signal, but this signal loss is not linearly proportional to the loss in cell number. Additionally, the loss is consistently experienced between off-target sites with both high and low CRISPR off-target activity as measured by indel percentage. Indel percentages (as sourced from Wienert et al., 2019) are listed in Table 16.
[0334] Table 16
[0335] FIG. 24 illustrates amplicon-based validation of the off-target site centered on
[0336] Chrl :9689829 (GRCh38) for the VEGFA Site 2 sgRNA identified by our approach and not identified in previous publications using GUIDE-Seq (Tsai et al, 2015) or DISCOVER-Seq+ (Zou et al., 2023). Our model correctly identified the peak with a high confidence (0.9995204 model output) given the peak outline and we show that the indel % over background is 2.20% in K562 cells, validating the model’s finding.
[0337] EXAMPLE Q
[0338] In this example, peaks identified using previous approaches, such as DISCOVER-Seq which used ChiP-Seq, are compared to explain why an ANN, e.g., a CNN, finds particular use in the methods described herein. As a preliminary matter, BLENDER, discussed above, is not a suitable approach because early in its process “the bam files are traversed to identify all loci with two or more reads ending (for reverse reads) or starting (for forward reads) at the same point or having a 1 bp overlap” before using the rubric “For all loci with a valid PAM, the DISCOVER score is then calculated by summing up the read ends in a 00B206.1663
[0339] 10 bp window around the cut site.” (See, Supplementary Materials, Wienert et al., 2019). As is apparent from CUT&Tag coverage plots centered on the cut site, 10 bp surrounding the cut site will not contain two or more reads ending or starting at the same point, nor will the 10 bp window contain any reads to sum up for calculating the DISCOVER score. Therefore, the BLENDER approach will not work with CUT&Tag-based datasets and therefore is insufficiently flexible for alternative nuclease-based approaches.
[0340] In this example, an exemplary trained model of the present disclosure contains a total of 2,344,646 trainable parameters. Inputs to the model first pass through the two 1- dimensional convolutional layers with kernel size (5,1) and ReLU nonlinear activation, increasing the channel depth from 1 to 8 and then 8 to 16, and each followed by batch normalization. Max pooling layers with window size (2,1) are interleaved to reduce spatial dimensionality and reduce overfitting. A third convolutional layer further increases the channel depth to 32, followed by another max pooling operation. The output is then flattened and passed through three fully connected layers, reducing the feature dimension from 2240 to 1000, then to 100, and finally to 2, with ReLU nonlinear activation functions for each layer. The final output layer applies a softmax activation to produce class probabilities.
[0341] To provide additional context to an exemplary approach of the methods described in the instant application, FIG. 25 provides a high-level overview of an exemplary computational approach and data flow after NGS reads have been aligned to the reference genome. Segments of the genome covered by NGS reads, or “peaks”, are identified and converted into one dimensional vectors. The result of filtering and scaling these vectors can be plotted in 2-dimensional space (shown) where the x-axis represents genomic coordinates and the y-axis represents a histogram of NGS reads aligning to those coordinates, or “coverage”. These filtered and scaled vectors are passed through the ANN, e.g., a CNN, and result in a posterior float value representing the probability that the peak represents a cut site. This posterior float value can take the value between and inclusive of 0 and 1.
[0342] As illustrated in FIG 26, peaks between CHiP-Seq (top track), CUT&RUN (middle track), and CUT&Tag (bottom track) can be compared. The genomic locus is centered on the on-target cut site (vertical black dotted line) + / - 250 bp (Chr6:43770576-43771075, GRCh38) with the VEGFA gene body below (horizontal black bar). It is clear that, due to the difference between the peak profiles surrounding the cut site, CUT&Tag requires a radically different and more flexible approach to identifying cut sites than CHiP-Seq or CUT&RUN. To address this need, an ANN that could recognize CUT&Tag cut site coverage profiles was developed. 00B206.1663
[0343] Examples of both positive and negative peaks are provided in FIG. 27. Positive peaks contain a characteristic valley between two peaks (left), while negative peaks can either be of low coverage (middle), or high coverage but incorrect profile (right). These peaks were aggregated across 66 samples using the VEGFA Site 2 sgRNA, resulting in 1020 positive peaks. 508915 negative peaks were aggregated from samples nucleofected with vehicle control. With a train-test split of 0.8:0.2, downsampling of the negative peaks to create class balance, and synthetic data approaches (primarily inversion of the peak signal), the final training set consisted of 9792 peaks, of which 4896 had positive labels and 4896 had negative labels. The corresponding test set consisted of 2448 peaks, of which 1224 had positive labels and 1224 had negative labels. These were also split in such a way as to avoid data leakage between the train and test set due to synthetic data: original peaks were split before synthetic data was generated. Data was then trained using Flux.jl, Adam training optimization (0.001 learning rate), batch size of 64, and sending the data through training twice. After that, the model was fine-tuned using manual curation.
[0344] FIG. 28 depicts an AUROC (Area Under the Receiver Operating Characteristic curve), which is a measure of model performance describing the trade-off between the true positive rate and the false positive rate. The performance for the fine-tuned model is 0.891985, with the ROC curve presented below. Other useful metrics are accuracy (0.890931), precision (0.9741), recall (0.6918), and the Fl score (0.8086).
[0345] Finally, FIG. 29 depicts a distribution of posterior estimates from model. The distribution of model outputs for true positives are shown on the right, and the distribution of model outputs for true negatives are shown on the right. The model is confident in the vast majority of cases. Note, however, that low scores for true positives may not be unfounded, as the ground truth dataset is not established, and some “true” positives have indeed been found to be false positives.
[0346] EMBODIMENTS
[0347] The following methods represent various embodiments that are described herein.
[0348] Embodiment 1. A method of identifying a Cas nuclease off-target (OT) site, comprising the steps of: a. permeabilizing cells in the presence of a gRNA and a Cas nuclease; 00B206.1663 b. culturing the cells with the Cas nuclease and gRNA for at least about 2 hours to about 48 hours in the presence of a tethered nuclease and one or more DNA repair inhibitors; c. isolating nuclei from said cultured cells and performing CUT&RUN or CUT&Tag and isolating nucleic acid sequences from the isolated nuclei; d. amplifying the isolated nucleic acids and preparing a nucleic acid library produced by CUT&RUN or CUT&Tag; e. aligning the nucleic acid sequences of the CUT&RUN or CUT&Tag library to a genomic reference sequence; and f. identifying, based on distinct peaks in the aligned nucleic acid sequences, Cas nuclease off target sites in the obtained nucleic acid sequences of the CUT&RUN or CUT&Tag library.
[0349] Embodiment 2. The method of embodiment 1, wherein the isolated nucleic acid sequences are from are less than 1x106 isolated nuclei.
[0350] Embodiment s. The method of any of embodiments 1-2, wherein the gRNA and the Cas nuclease are in a complex prior to permeabilizing the cells.
[0351] Embodiment 4. The method of embodiment 3, wherein the cells express the Cas nuclease.
[0352] Embodiment 5. The method of any of embodiments 1-4, wherein the cells are recombinantly transformed to express the Cas nuclease.
[0353] Embodiment 6. The method of any of embodiments 1-5, wherein the cell is a mammalian cell.
[0354] Embodiment 7. The method of any of embodiments 1-6, wherein the cell is a human cell.
[0355] Embodiment 8. The method of embodiment 7, wherein the genomic reference sequence is hg38.
[0356] Embodiment 9. The method of any of embodiments 1-8, wherein the identifying of step f) is by distributing the aligned nucleic acid sequences into desired categories for parallel processing.
[0357] Embodiment 10. The method of any of embodiments 1-9, wherein permeabilizing is by electroporation or nucleofection.
[0358] Embodiment 11. The method of any of embodiments 1-10, wherein the nuclei derived from said cultured cells have been previously isolated and cryopreserved prior to steps c) to f). 00B206.1663
[0359] Embodiment 12. The method of any of embodiments 1-11, wherein said cultured cells or nuclei derived from said cultured cells are subject to fixation prior to steps c) to f).
[0360] Embodiment 13. The method of embodiment 12, wherein cell fixation is by formaldehyde, paraformaldehyde, glutaraldehyde, disuccinimidyl glutarate (DSG), or a combination thereof.
[0361] Embodiment 14. The method of embodiment 13, wherein cell fixation is by formaldehyde.
[0362] Embodiment 15. The method of any of embodiments 1 to 14, wherein the Cas nuclease is a Cas9 nuclease or a Casl2 nuclease.
[0363] Embodiment 16. The method of embodiment 15, wherein the Casl2 nuclease is Casl2a.
[0364] Embodiment 17. The method of any of embodiments 1-30, wherein the tethered nuclease is a tethered micrococcal nuclease (MNase) or a tethered transposase, and wherein an antibody is used in combination with the tethered MNase.
[0365] Embodiment 18. The method of embodiment 17, wherein the micrococcal nuclease is a pAG-MNase, a pA-MNase, a tethered NSB1, or a pG-MNase and the tethered transposase is protein A-Tn5.
[0366] Embodiment 19. The method of any of embodiments 1-30, wherein the DNA repair inhibitor is a DNA-dependent protein kinase (DNA-PK) inhibitor, an ATM inhibitor, or a combination of a DNA-PK inhibitor and an ATM inhibitor.
[0367] Embodiment 20. The method of embodiment 19, wherein the ATM inhibitor is KU0060648 or KU55933.
[0368] Embodiment 21. The method of any of embodiments 19-20, wherein the DNA-PK inhibitor is NU7026, NU7441, or AZD7648.
[0369] Embodiment 22. The method of any of embodiments 19-21, wherein the DNA-PK inhibitor is AZD7648.
[0370] Embodiment 23. The method of any of embodiments 1-30, wherein the DNA repair inhibitor is AZD7648 or KU0060648.
[0371] Embodiment 24. The method of any of embodiments 19-23, wherein the DNA repair inhibitor is the combination of a DNA-PK inhibitor and ATM inhibitor and the amount of DNA repair inhibitor used is at least 50% less of both the DNA-PK inhibitor and the ATM inhibitor as compared to a method using the DNA-PK inhibitor and ATM inhibitor alone. 00B206.1663
[0372] Embodiment 25. The method of any of embodiments 2-24, wherein the number of cells is about 100,000 or less.
[0373] Embodiment 26. The method of any of embodiments 2-25, wherein the number of cells is about 50,000 or less.
[0374] Embodiment 27. The method of any of embodiments 1-30, wherein the culturing is about 3 hours to about 16 hours.
[0375] Embodiment 28. The method of embodiment 27, wherein the culturing is about 5 hours to about 12 hours.
[0376] Embodiment 29. The method of any of embodiments 9-30, wherein the desired categories are chromosomes of the reference sequences.
[0377] Embodiment 30. The method of any of embodiments 1-30, wherein the cells are mammalian adherent cells or mammalian nonadherent cells.
[0378] Embodiment 31. The method of any of embodiments 1-30, wherein the cells are selected from T cells, induced pluripotent stem cells, or cells derived from induced pluripotent stem cells.
[0379] Embodiment 32. The method of embodiment 31, wherein the T cells are human CD4+ T cells or human CD8+ T cells.
[0380] Embodiment 33. The method of any of embodiments 31-32, wherein the pluripotent cells are induced pluripotent stem cells (iPSCs).
[0381] Embodiment 34. The method of embodiment 33, wherein the iPSC is KOLF2.1J.
[0382] Embodiment 35. A method of identifying a compound that has DNA repair inhibitory activity comprising the steps of: a. permeabilizing a first set of cells in the presence of a gRNA and a Cas nuclease to introduce the gRNA and the Cas nuclease into the first set of cells, wherein of the first set of cells are fewer than about IxlO6; b. culturing the first set of cells with the introduced gRNA and Cas nuclease for at least 3 hours to 48 hours in the presence of a known DNA repair inhibitor compound and a tethered nuclease under appropriate conditions for the tethered nuclease; c. obtaining a first set of nucleic acid sequences from said first set of cells or nuclei derived from said first set of cells using CUT&RUN or CUT&Tag; d. aligning the first set of nucleic acid sequences from the first set of cells or nuclei derived from said first set of cells to a genomic reference sequence; 00B206.1663 e. identifying, based on characteristics of peaks detected in the aligned first set of nucleic acid sequences, first Cas nuclease cut sites in the first set of nucleic acid sequences; f. permeabilizing a second set of cells in the presence of the gRNA and the Cas nuclease; g. culturing the second set of cells with the introduced gRNA and Cas nuclease for at least 3 hours to 48 hours in the presence of a compound and the tethered nuclease under the appropriate conditions for the tethered nuclease; h. obtaining a second set of nucleic acid sequences from said second set of cultured cells or nuclei derived from said second set of cultured cells using CUT&RUN or CUT&Tag; i. aligning the second set of nucleic acid sequences from the second set of cells or nuclei derived from said second set of cultured cells to the genomic reference sequence; and j. identifying, based on characteristics of peaks detected in the aligned second set of nucleic acid sequences, second Cas nuclease cut sites in the obtained second set of nucleic acid sequences; k. comparing the first Cas nuclease cut sites identified for the first set of nucleic acid sequences to the second Cas nuclease cut sites identified for the second set of nucleic acid sequences to ascertain an overlap in the identified first and second Cas nuclease cut sites; and l. optionally scoring cell viability of the second set of cells exposed to the compound relative to the cell viability of the first set of cells exposed to the known DNA repair inhibitor.
[0383] Embodiment 36. The method of embodiment 35, wherein the known DNA repair inhibitor is AZD7648 or KU60648.
[0384] Embodiment 37. The method of any of embodiments 35-36, wherein the first set of cells and the second set of cells express the Cas nuclease.
[0385] Embodiment 38. The method of any of embodiments 35 to 37, wherein the Cas nuclease is a Cas9 nuclease or a Casl2 nuclease.
[0386] Embodiment 39. The method of embodiment 38, wherein the Casl2 nuclease is Casl2a.
[0387] Embodiment 40. The method of any of embodiments 35 to 39, wherein the tethered nuclease is a tethered MNase or a tethered transposase. 00B206.1663
[0388] Embodiment 41. The method of embodiment 40, wherein the tethered MNase is pAG-MNase and the tethered transposase is protein A-Tn5.
[0389] Embodiment 42. The method of any of embodiments 34 to 39, wherein the characteristics of the peaks detected in the aligned first set of nucleic acid sequences and the characteristics of the peaks detected in the aligned second set of nucleic acid sequences are determined by a convolutional neural network (CNN) to be characteristics of cut-sites of the Cas nuclease, wherein the CNN was trained on corresponding characteristics of peaks obtained from known on- and off-target Cas nuclease cut-sites.
[0390] Embodiment 43. A method of identifying Cas-nuclease off-target sites, comprising steps of: a. isolating nucleic acid sequences of an undigested CUT&RUN or CUT&Tag library generated from cells incubated with a gRNA in the presence of a Cas nuclease, a tethered nuclease, and one or more DNA repair inhibitors; b. aligning the isolated nucleic acid sequences to a genomic reference sequence; c. calling peaks based on read coverage of the aligned isolated nucleic acid sequences to the genomic reference sequence; d. passing a filter over the called peaks, wherein the filter is configured to emphasize features, in the called peaks, associated with a breakpoint; e. inputting the filtered peaks to a convolutional neural network (CNN) trained on trained on corresponding features of peaks obtained from known on- and off-target Cas nuclease cut-sites; and f. generating, based on output from the CNN, a list of discovered Cas nuclease off-target sites.
[0391] Embodiment 44. The method of embodiment 43, further comprising the steps of: g. distributing the aligned nucleic acid sequences into desired categories for parallel processing, wherein steps b) to e) are performed via parallel processing.
[0392] Embodiment 45. The method of any of embodiments 43 or 44, wherein the peak calling comprises applying one or more peak calling algorithms selected from: i. Model-based Analysis for ChlP-Seq (MACS), ii. MACS version 2 (MACS2), iii. Genome wide Event finding and Motif discovery (GEM), iv. MUltiScale enrichment Calling for ChlP-Seq (MUSIC), v. Bayesian Change point (BCP), 00B206.1663 vi. Zero-Inflated Negative Binomial Algorithm (ZINBA), or vii. Threshold-based Method (TM).
[0393] Embodiment 46. The method of any of embodiments 43-45, wherein the peak calling comprises applying a threshold to the read coverage, by the aligned nucleic acid sequences, of the genomic reference sequence.
[0394] Embodiment 47. The method of embodiment 46, wherein the threshold is at least about 200 bp.
[0395] Embodiment 48. The method of any of embodiments 43-46, wherein the filter is configured for edge detection.
[0396] Embodiment 49. The method of any of embodiments 43-46, further comprising estimating a maximum likelihood that the filtered peaks correspond to sites of off- target sequences, wherein said estimating is based on the output from the CNN and a colocalized gRNA alignment score.
[0397] Embodiment 50. A computer-implemented method comprising: one or more processors; and a memory storing instructions, which when executed by the one or more processors, cause the one or more processor to identify on- and / or off-target cleavage sites of a Cas nuclease by: a. aligning nucleic acid sequences obtained from a nucleic acid library to a genomic reference sequence, where the nucleic acid library is generated by DNA binding protein-targeted enzymatic digestion of nucleic acids from cells contacted by a Cas nuclease, an sgRNA, and a target DNA binding protein; b. identifying peaks based on read coverage of the aligned nucleic acid sequences to the genomic reference sequence; c. instantiating at least one one-dimensional data structure comprising values correlated to locations of the aligned obtained nucleic acid sequence to the genomic reference sequence correlated to the location of the identified peaks; d. inputting the at least one, one-dimensional data structure to an artificial neural network (ANN) that has been trained using pre-identified peaks from known on- and off-target Cas nuclease sites to identify Cas nuclease off-target sites; and e. generating, based on output from the ANN, a list of discovered Cas nuclease on and off-target sites. 00B206.1663
[0398] Embodiment 51. The method of embodiment 50, wherein the cells were subjected to fixation prior to generation of the nucleic acid library.
[0399] Embodiment 52. The method of embodiment 51, wherein the fixation comprises a 0.0001% - 0.1% fixative.
[0400] Embodiment 53. The method of embodiment 51 or 52, wherein fixation comprises contacting the cells with formaldehyde, paraformaldehyde, glutaraldehyde, disuccinimidyl glutarate (DSG), or a combination thereof.
[0401] Embodiment 54. The method of any one of embodiments 51 to 53, wherein the fixation comprises contacting the cells with the fixative for between about 1 minute to about 5 minutes.
[0402] Embodiment 55. The method of any one of embodiments 50-54, wherein the Cas nuclease is a Cas9 nuclease.
[0403] Embodiment 56. The method of any one of embodiments 50-55, wherein the Cas nuclease is a Casl2a nuclease.
[0404] Embodiment 57. The method of any one of embodiments 50-56, wherein the DNA binding protein is a DNA repair factor.
[0405] Embodiment 58. The method of embodiment 57, wherein the DNA repair factor is selected from the group consisting of ATM, NBS, MRE11, Ku70, Ku80, and FANCD2.
[0406] Embodiment 59. The method of embodiment 58, wherein the DNA repair factor is ATM.
[0407] Embodiment 60. The method of embodiment 58, wherein the DNA repair factor is MRE11.
[0408] Embodiment 61. The method of any one of embodiments 50-60, wherein the cells were incubated in the presence of one or more DNA repair inhibitors prior to generation of the nucleic acid library.
[0409] Embodiment 62. The method of embodiment 61, wherein the DNA repair inhibitor is a DNA-dependent protein kinase (DNA-PK) inhibitor, an ATM inhibitor, or a combination of a DNA-PK inhibitor and an ATM inhibitor.
[0410] Embodiment 63. The method of embodiment 61 or 62, wherein the DNA repair inhibitor is selected from the group consisting of KU060648, KU55933, KU60019, AZDI 390, AZ31, NU7026, NU7441, AZD764, and AZD7648.
[0411] Embodiment 64. The method of any one of embodiments 61-63, wherein the DNA repair inhibitor is KU060648. 00B206.1663
[0412] Embodiment 65. The method of any one of embodiments 61-63, wherein the DNA repair inhibitor is AZD7648.
[0413] Embodiment 66. The method of any one of embodiments 61-63, wherein the DNA repair inhibitor is AZD7648 and KU060648.
[0414] Embodiment 67. The method of embodiment 61 or 62, wherein the DNA repair inhibitor is the combination of a DNA-PK inhibitor and ATM inhibitor and the amount of DNA repair inhibitor used is at least 50% less of both the DNA-PK inhibitor and the ATM inhibitor as compared to a method using the DNA-PK inhibitor and ATM inhibitor alone.
[0415] Embodiment 68. The method of any one of embodiments 50-67, wherein the cells are mammalian adherent cells or mammalian nonadherent cells.
[0416] Embodiment 69. The method of any one of embodiments 50-68, wherein the cells are selected from the group consisting of T cells, K562 cells, induced pluripotent stem cells, and cells derived from induced pluripotent stem cells.
[0417] Embodiment 70. The method of embodiment 69, wherein the T cells are human CD4+ T cells or human CD8+ T cells.
[0418] Embodiment 71. The method of embodiment 69, wherein the pluripotent cells are induced pluripotent stem cells.
[0419] Embodiment 72. The method of any one of embodiments 50-71, wherein the off-target site identified occurs with a frequency as low as 0.001% of the sequence reads as verified by a targeted approach.
[0420] Embodiment 73. The method of any one of embodiments 50-72, wherein the ANN was trained on the characteristics of the peaks obtained from known on- and off- target nucleotide-directed nuclease cut-sites validated by a targeted approach.
[0421] Embodiment 74. The method of embodiment 74, wherein the targeted approach is rhAmpSeq.
[0422] Embodiment 75. The method of any one of embodiments 50-74, wherein the peak calling comprises applying Threshold-based Method (TM) peak calling algorithm.
[0423] Embodiment 76. The method of any one of the embodiments 50-75, wherein the peak calling comprises applying a threshold to the read coverage, by the aligned nucleic acid sequences, of the genomic reference sequence.
[0424] Embodiment 77. The method of embodiment 76, wherein the threshold for peak width is at least about 200 bp and the threshold for peak depth is at least 4 reads. 00B206.1663
[0425] Embodiment 78. The method of any one of embodiments 50-77, further comprising passing a filter over the called peaks, wherein the filter is configured to emphasize features in the called peaks associated with a breakpoint.
[0426] Embodiment 79. The method of embodiment 78, wherein the filter is configured for edge detection.
[0427] Embodiment 80. The method of embodiment 78 or 79, further comprising estimating a maximum likelihood that the filtered peaks correspond to sites of off-target sequences, wherein said estimating is based on the output from the ANN and a co-localized sgRNA alignment score.
Claims
00B206.1663CLAIMS1. A computer-implemented method comprising: one or more processors; and a memory storing instructions, which when executed by the one or more processors, cause the one or more processor to identify on- and off-target cleavage sites of a Cas nuclease by: a. aligning nucleic acid sequences obtained from a nucleic acid library to a genomic reference sequence, where the nucleic acid library is generated by DNA binding protein-targeted enzymatic digestion of nucleic acids from cells contacted by a Cas nuclease, an sgRNA, and a target DNA binding protein; b. identifying peaks based on read coverage of the aligned nucleic acid sequences to the genomic reference sequence; c. instantiating at least one one-dimensional data structure comprising values correlated to locations of the aligned obtained nucleic acid sequence to the genomic reference sequence correlated to the location of the identified peaks; d. inputting the at least one, one-dimensional data structure to an artificial neural network (ANN) that has been trained using pre-identified peaks from known on- and off-target Cas nuclease sites to identify Cas nuclease off-target sites; and e. generating, based on output from the ANN, a list of discovered Cas nuclease on and off-target sites.
2. The method of claim 1, wherein the cells were subjected to fixation prior to generation of the nucleic acid library.
3. The method of claim 2, wherein the fixation comprises a 0.0001% - 0.1% fixative.
4. The method of claim 2 or 3, wherein fixation comprises contacting the cells with formaldehyde, paraformaldehyde, glutaraldehyde, disuccinimidyl glutarate (DSG), or a combination thereof.
5. The method of any of claims 2 to 4, wherein the fixation comprises contacting the cells with the fixative for between about 1 minute to about 5 minutes.
6. The method of any one of the preceding claims, wherein the Cas nuclease is a Cas9 nuclease.00B206.16637. The method of any one of the preceding claims, wherein the Cas nuclease is a Casl2a nuclease.
8. The method of any one of the preceding claims, wherein the DNA binding protein is a DNA repair factor.
9. The method of claim 8, wherein the DNA repair factor is selected from the group consisting of ATM, NBS, MRE11, Ku70, Ku80, and FANCD2.
10. The method of claim 9, wherein the DNA repair factor is ATM.
11. The method of claim 9, wherein the DNA repair factor is MRE11.
12. The method of any of the preceding claims, wherein the cells were incubated in the presence of one or more DNA repair inhibitors prior to generation of the nucleic acid library13. The method of claim 12, wherein the DNA repair inhibitor is a DNA-dependent protein kinase (DNA-PK) inhibitor, an ATM inhibitor, or a combination of a DNA-PK inhibitor and an ATM inhibitor.
14. The method of claim 12 or 13, wherein the DNA repair inhibitor is selected from the group consisting of KU060648, KU55933, KU60019, AZD1390, AZ31, NU7026, NU7441, AZD764, and AZD7648.
15. The method of any one of claims 12-14, wherein the DNA repair inhibitor is KU060648.
16. The method of any one of claims 12-14, wherein the DNA repair inhibitor is AZD7648.
17. The method of any one of claims 12-14, wherein the DNA repair inhibitor is AZD7648 and KU060648.
18. The method of claim 12 or 13, wherein the DNA repair inhibitor is the combination of a DNA-PK inhibitor and ATM inhibitor and the amount of DNA repair inhibitor used is at least 50% less of both the DNA-PK inhibitor and the ATM inhibitor as compared to a method using the DNA-PK inhibitor and ATM inhibitor alone.
19. The method of any one of the preceding claims, wherein the cells are mammalian adherent cells or mammalian nonadherent cells.00B206.166320. The method of any one of the preceding claims, wherein the cells are selected from the group consisting of T cells, K562 cells, induced pluripotent stem cells, and cells derived from induced pluripotent stem cells.
21. The method of claim 20, wherein the T cells are human CD4+ T cells or human CD8+ T cells.
22. The method of claim 20, wherein the pluripotent cells are induced pluripotent stem cells.
23. The method of any one of the preceding claims wherein the off-target site identified occurs with a frequency of at least 0.001% of the sequence reads as verified by a targeted approach.
24. The method of claim 23, wherein the off-target site identified occurs with a frequency of between 0.01% and 100% of the sequence reads as verified by a targeted approach.
25. The method of claim 24, wherein the off-target site identified occurs with a frequency of between 0.01% and 50% of the sequence reads as verified by a targeted approach.
26. The method of claim 25, wherein the off-target site identified occurs with a frequency of between 0.01% and 10% of the sequence reads as verified by a targeted approach.
27. The method of claim 26, wherein the off-target site identified occurs with a frequency of between 0.01% and 1% of the sequence reads as verified by a targeted approach.
28. The method of claim 23, wherein the off-target site identified occurs with a frequency of between 0.1% and 100% of the sequence reads as verified by a targeted approach.
29. The method of claim 28, wherein the off-target site identified occurs with a frequency of between 0.1% and 50% of the sequence reads as verified by a targeted approach.
30. The method of claim 29, wherein the off-target site identified occurs with a frequency of between 0.1% and 10% of the sequence reads as verified by a targeted approach.
31. The method of claim 30, wherein the off-target site identified occurs with a frequency of between 0.1% and 1% of the sequence reads as verified by a targeted approach.
32. The method of claim 31, wherein the off-target site identified occurs with a frequency of between 1% and 100% of the sequence reads as verified by a targeted approach.00B206.166333. The method of claim 23, wherein the off-target site identified occurs with a frequency of between 1% and 50% of the sequence reads as verified by a targeted approach.
34. The method of claim 33, wherein the off-target site identified occurs with a frequency of between 1% and 10% of the sequence reads as verified by a targeted approach.
35. The method of any one of the preceding claims, wherein the nucleic acid library is generated from: at least 1 cell; at least 10 cells; at least 100 cells; at least 500 cells; at least 1000 cells; at least 5000 cells; at least 10,000 cells; at least 20,000 cells; at least 30,000 cells; at least 40,000 cells; at least 50,000 cells; at least 60,000 cells; at least 70,000 cells; at least 80,000 cells; at least 90,000 cells; at least 100,000 cells; at least 200,000 cells; at least 300,000 cells; at least 400,000 cells; at least 500,000 cells; at least 600,000 cells; at least 700,000 cells; at least 800,000 cells; at least 900,000 cells; at least IxlO6cells; at least IxlO7cells; or at least IxlO8cells.
36. The method of any one of the preceding claims, wherein the ANN was trained on the characteristics of the peaks obtained from known on- and off-target nucleotide-directed nuclease cut-sites validated by a targeted approach.
37. The method of claim 36, wherein the targeted approach is rhAmpSeq.
38. The method of any one of the preceding claims, wherein the peak calling comprises applying Threshold-based Method (TM) peak calling algorithm.
39. The method of any one of the preceding claims, wherein the peak calling comprises applying a threshold to the read coverage, by the aligned nucleic acid sequences, of the genomic reference sequence.
40. The method of claim 39, wherein the threshold for peak width is at least about 200 bp and the threshold for peak depth is at least 4 reads.
41. The method of any one of the preceding claims, further comprising passing a filter over the called peaks, wherein the filter is configured to emphasize features in the called peaks associated with a breakpoint.
42. The method of claim 41, wherein the filter is configured for edge detection.00B206.166343. The method of claim 41 or 42, further comprising estimating a maximum likelihood that the filtered peaks correspond to sites of off-target sequences, wherein said estimating is based on the output from the ANN and a co-localized sgRNA alignment score.
44. A computer-implemented method comprising: one or more processors; and a memory storing instructions, which when executed by the one or more processors, cause the one or more processor to identify a compound that has DNA repair inhibitory activity by: a. aligning a first set of nucleic acid sequences obtained from a first set of cells comprising a gRNA and a Cas nuclease with a DNA repair inhibitor compound to a genomic reference sequence; b. identifying, based on characteristics of peaks detected in the aligned first set of nucleic acid sequences, first Cas nuclease off-target sites in the first set of nucleic acid sequences; c. aligning a second set of nucleic acid sequences obtained from a second set of cells comprising a gRNA and a Cas nuclease with a DNA repair inhibitor compound to the genomic reference sequence; and h. identifying, based on characteristics of peaks detected in the aligned second set of nucleic acid sequences, second Cas nuclease off-target sites in the obtained second set of nucleic acid sequences; i. comparing the first Cas nuclease off-target sites identified for the first set of nucleic acid sequences to the second Cas nuclease off-target sites identified for the second set of nucleic acid sequences to ascertain whether the compound has DNA repair inhibitory activity; and j. optionally scoring cell viability of the second set of cells exposed to the compound relative to the cell viability of the first set of cells exposed to the known DNA repair inhibitor.
45. The method of claim 44, wherein the known DNA repair inhibitor is AZD7648 or KU60648.00B206.166346. The method of claim 44, wherein the first set of cells and the second set of cells express the Cas nuclease.
47. The method of any one of claims 44 to 46, wherein the Cas nuclease is a Cas9 nuclease or a Cas 12a nuclease.
48. The method of any one of claims 44 to 47, wherein the characteristics of the peaks detected in the aligned first set of nucleic acid sequences and the characteristics of the peaks detected in the aligned second set of nucleic acid sequences are determined by an artificial neural network (ANN) to be characteristics of cut-sites of the Cas nuclease, wherein the ANN was trained on nucleic acid sequences of on-target Cas nuclease cut-sites obtained.
49. The method of any one of claims 44 to 48, wherein the compound is identified as having DNA repair inhibitory activity if the second Cas nuclease off-target sites identified for the second set of nucleic acid sequences is different in peak intensity from the first Cas nuclease off-target sites identified for the first set of nucleic acid sequences.
50. The method of any one of claims 44 to 48, wherein the compound is identified as having DNA repair inhibitory activity if the second Cas nuclease off-target sites identified for the second set of nucleic acid sequences is different in peak composition from the first Cas nuclease off-target sites identified for the first set of nucleic acid sequences.
Citation Information
Patent Citations
High efficiency targeted in SITU genome-wide profiling
WO2019060907A1
US202463688801P