Methods of using double-stranded cytosine deaminases to measure chromatin accessibility through base editing

Double-stranded DNA cytosine deaminases convert accessible cytosines to uracils for high-resolution chromatin accessibility profiling, overcoming limitations of existing methods by enabling nucleotide-level single-cell analysis and improving TF binding mapping.

WO2026020174A1PCT designated stage Publication Date: 2026-01-22THE BRIGHAM & WOMEN S HOSPITAL INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/038517
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-19
Filing Date
2025-07-21
Publication Date
2026-01-22

AI Technical Summary

Technical Problem

Current methods for measuring chromatin accessibility, such as DNase-seq and ATAC-seq, are limited by resolution and scalability, preventing analysis of co-occurrence of adjacent transcription factor binding or nucleosome positioning events, especially in single-cell applications, while alternative methods like Fiber-seq and NOMe-seq are expensive or incompatible with PCR amplification.

Method used

Employing double-stranded DNA cytosine deaminases (Ddd) to convert accessible cytosines to uracils, followed by PCR-based conversion to thymines, enabling high-resolution, single-molecule chromatin accessibility profiling compatible with next-generation sequencing.

Benefits of technology

Enhances chromatin accessibility measurement resolution by five-fold, allowing nucleotide-level analysis of single-cell chromatin accessibility and TF binding, compatible with existing sequencing pipelines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000025_0001
    Figure IMGF000025_0001
  • Figure IMGF000028_0001
    Figure IMGF000028_0001
  • Figure IMGF000028_0002
    Figure IMGF000028_0002
Patent Text Reader

Abstract

Described herein are compositions, kits, and methods for chromatin accessibility measurement using genome editing by double-stranded DNA cytosine deaminases (Ddds).
Need to check novelty before this filing date? Find Prior Art

Description

[0001]Attorney Docket No.29618-0483WO1 / BWH 2023-476 METHODS OF USING DOUBLE-STRANDED CYTOSINE DEAMINASES TO MEASURE CHROMATIN ACCESSIBILITY THROUGH BASE EDITING CLAIM OF PRIORITY This application claims the benefit of U.S. Provisional Patent Application Serial No.63 / 673,268, filed on July 19, 2024. The entire contents of the foregoing are hereby incorporated by reference. FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT This invention was made with Government support under Grant Nos. HL164409, HG012010, and GM143249 awarded by the National Institutes of Health. The Government has certain rights in the invention. TECHNICAL FIELD Described herein are compositions, kits, and methods for chromatin accessibility measurement using genome editing by double-stranded DNA cytosine deaminases (Ddds). BACKGROUND Genomic DNA comprises multiple overlapping codes that contain information specifying cellular function. Although the “genetic code” that governs how DNA encodes protein sequence through triplet codons was cracked more than 40 years ago, much remains to be learned about the code governing the epigenome, the functionalization of the genome that dictates how genes are regulated. Human tissue specialization as well as many disease states are controlled by the epigenetic state of cells, which is determined largely by genomic binding of ~2,000 transcription factors (TFs), the interaction between DNA and histones (chromatin accessibility), and DNA methylation. Fundamental gaps remain in our comprehension of how cell type- specific TF binding shapes epigenetic state. SUMMARY Provided herein are kits and methods for analyzing chromatin. For example, provided are methods comprising: providing a sample comprising one or more intact Attorney Docket No.29618-0483WO1 / BWH 2023-476 cell nuclei comprising chromatin; and contacting the intact nuclei concurrently with (i) one or more Ddd to convert accessible cytosines in the chromatin to uracils (C^U) and (ii) Tn5 transposase loaded with oligonucleotides comprising dsDNA sequencing adapters, wherein the oligonucleotides optionally further comprise tags, to fragment and tag the chromatin; and detecting conversion of C^U in the chromatin. Also provided are methods comprising providing a sample comprising one or more intact cell nuclei comprising chromatin; contacting the intact nuclei with one or more double-stranded DNA cytosine deaminase (Ddd) to convert accessible cytosines in the chromatin to uracils (C^U); then optionally contacting the intact nuclei with a Ddd inhibitor (Dddi) to inhibit the Ddd or purifying the DNA immediately after Ddd treatment to remove the Ddd, and detecting conversion of C^U in the chromatin. In some embodiments, detecting conversion of C^U in the chromatin comprises: amplifying the chromatin with a uracil-tolerant polymerase to convert the uracils to thymines (U^T), and performing sequencing to detect mutations of C^T, preferably wherein the sequencing is next generation sequencing. In some embodiments, detecting conversion of C^U in the chromatin comprises: contacting the nuclei with restriction enzyme DpnII, and comparing activity of the restriction enzyme DpnII in the chromatin to reference activity of the restriction enzyme DpnII in reference chromatin not treated with the Ddd. In some embodiments, detecting conversion of C^U in the chromatin comprises: contacting the nuclei with Tn5 transposase loaded with oligonucleotides comprising dsDNA sequencing adapters, and optionally further comprising tags, to fragment and tag the chromatin; amplifying the chromatin fragments with a uracil- tolerant polymerase to convert uracils to thymines (U^T); and sequencing the fragments, preferably wherein sequencing comprises using next generation sequencing. In some embodiments, the Ddd comprises DddSs from Simiaoa sunii. In some embodiments, the Ddd inhibitor comprises double-stranded DNA deaminase immunity protein (DddI) (from Simiaoa sunii or from Burkholderia cenocepacia (strain H111)). In some embodiments, the uracil-tolerant polymerase is PHUSION high fidelity DNA polymerase; KAPA HiFi Hotstart Uracil+ DNA B family Polymerase; Attorney Docket No.29618-0483WO1 / BWH 2023-476 KOD Multi Epi high fidelity DNA Polymerase; or Q5U Hot Start High-Fidelity DNA Polymerase. In some embodiments, the method further comprises comparing sequences obtained from sequencing the fragments to a reference genome for the cell; and identifying positions in the sequences that show conversion of a cytosine in the reference genome to a thymine in the fragment sequences; and designating the position in the genome as accessible. Additionally, provided herein are kits for use in a method as described herein. In some embodiments, the kits comprise reagents for isolating nuclei from a population of cells; double-stranded DNA cytosine deaminase (Ddd); a uracil-tolerant polymerase; and one or more reaction buffers; and optionally a Ddd inhibitor (Dddi). In some embodiments, the kits further comprise a Tn5 transposase and oligonucleotides comprising transposon tags and sequencing adaptors, optionally wherein the Tn5 transposase and oligonucleotides are present in pre-formed complexes. In some embodiments, the uracil-tolerant polymerase is PHUSION high fidelity DNA polymerase; KAPA HiFi Hotstart Uracil+ DNA B family Polymerase; KOD Multi Epi high fidelity DNA Polymerase; or Q5U Hot Start High-Fidelity DNA Polymerase. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Methods and materials are described herein for use in the present invention; other, suitable methods and materials known in the art can also be used. The materials, methods, and examples are illustrative only and not intended to be limiting. All publications, patent applications, patents, sequences, database entries, and other references mentioned herein are incorporated by reference in their entirety. In case of conflict, the present specification, including definitions, will control. Other features and advantages of the invention will be apparent from the following detailed description and figures, and from the claims. DESCRIPTION OF DRAWINGS FIGs.1A-M: DddA provides chromatin accessibility-dependent cytosine deamination. (A). Agilent Tapestation gel showing results of titration of DddA (top) Attorney Docket No.29618-0483WO1 / BWH 2023-476 and DddSs (bottom) activity using a double-stranded DNA (dsDNA) deamination assay. Appearance of a 90-nt band denotes dsDNA deamination, which impedes DpnII restriction digest. DddSs and DddA concentration and deamination activity are reported. DddSs and DddA deaminated cytosine dose-dependently and were inhibited by DddA-I in vitro. (B) DddA showed preferential editing in accessible chromatin in intact nuclei. (C) DddA edited at TC motifs proportionally to accessibility genome- wide in HCT116. (D) Enrichment of edited cytosines in accessible vs. inaccessible chromatin for 6 Ddd enzymes, separated by trinucleotide. Data is derived from ACCESS-WGS data, and each point represents the accessible vs. inaccessible foldchange in observed C2T editing for a particular cytosine-centered trinucleotide. (E) Dinucleotide-resolved edit fractions in DddSs-treated HCT116 cells binned by ATAC-seq coverage. Data derive from ACCESS-WGS data. (F) Motif logo showing the relative importance of each nucleotide surrounding the target cytosine in influencing DddSs editing efficiency in E. coli genomic DNA. (G) Schematic of ACCESS-ATAC, in which nuclei are concurrently treated with DddSs, which converts cytosine residues to uracils in proportion to chromatin accessibility, and Tn5. (H) Enrichment of transcription start sites (TSS) in K562 ATAC-seq and ACCESS-ATAC reads. Each row is a TSS, rows are ordered by the magnitude of read enrichment, and a summary plot above shows read enrichment in the 2 kb surrounding all TSS. (I) Scatter plot showing genome-wide correlation between ATAC-seq and ACCESS- ATAC coverage. Each dot represents a 10-kb bin, and the color refers to density of the bins. (J) Representative K562 ATAC-seq and ACCESS-ATAC signal tracks. Tn5 insertion counts from K562 ATAC-seq and ACCESS-ATAC in a representative 150- kb genomic region. (K) Metaplots showing K562 ATAC-seq Tn5 insertion signal, and ACCESS-ATAC Tn5 insertion and DddSs editing signal surrounding K562 ChIP-seq binding sites for CTCF, MAX and ATF3. Bottom plot shows merged non-normalized signal tracks to highlight the denser DddSs editing signal. Arrows denote footprint width. (L) Schematic of exemplary three-step ACCESS-ATAC pipeline. (M) Schematic of exemplary concurrent ACCESS-ATAC pipeline FIGs.2A-I: ACCESS-ATAC-seq improves transcription factor binding site prediction. (A) Schematic depicting deep learning-based modeling of DddSs editing bias using high-coverage E. coli genomic DNA editing data. (B) Scatter plot showing predicted and observed editing counts for held-out regions (30kb) of the E. Attorney Docket No.29618-0483WO1 / BWH 2023-476 coli genome. Each dot represents a single nucleotide, and colors refer to density of the nucleotides. (C) Representative tracks showing observed and predicted DddSs genomic DNA editing counts in a 30-kb stretch (top) and a 240-bp subregion (bottom) of the E. coli genome. (D) Single-nucleotide resolution signal for predicted, observed, and bias-corrected DddSs edit counts around K562 ELF4 ChIP-seq binding sites from K562 ACCESS-ATAC data. (E) Schematic depicting deep learning-based prediction of transcription factor binding sites using DNA sequence and open chromatin profiles. (F) Evaluation of auPRC comparing TFBS predictions with ChIP-seq binding sites for 1076 K562 motifs using ACCESS-ATAC Tn5 and DddSs signal, DddSs only, or Tn5 only, ATAC-seq Tn5 signal or DNA sequence only. Line shows median auPRC, boxes show quartiles, and lines show outliers. (G) Precision-recall curves of TFBS prediction performance for K562 ATF3 and ANDP ChIP-seq binding sites using the indicated tracks. auPRC for each track is reported in parentheses. (H) Representative tracks showing de novo footprint detection results using ACCESS-ATAC data in a 300-nt genomic region. Top: observed DddSs edit counts. Middle: predicted DddSs bias using the deep-learning model trained on the E. coli genome. Bottom: estimated p-value of TF footprints. (I) Jaccard index of detected footprints using ACCESS- ATAC DddSs signal or Tn5 signal or ATAC-seq Tn5 signal. For F and I, center line shows median, boxes show quartiles, and lines show outliers. FIGs.3A-D: Imputing allele-resolved transcription factor occupancy from ACCESS-ATAC data. (A) Schematic of Occupancy Pattern Inference by Editing (OccuPIE), a model to impute allelic TF occupancy from ACCESS-ATAC data. (B) Evaluation of Spearman correlation between OccuPIE-imputed median allelic occupancy and ChIP-seq read abundance for 98 HepG2 and 106 K562 TFs (C) Edit fractions for reads surrounding K562 STAT5A (left) and HepG2 USF1 (right) ChIP- seq-bound motifs. Reads are separated into bound, unbound, accessible, and unbound, inaccessible states according to maximum OccuPIE-assigned probability. (D) Evaluation of mean OccuPIE-imputed bound probability (x-axis) and ChIP-seq read abundance (y-axis) at K562 STAT5A (left) and HepG2 USF1 (right) ChIP-seq bound loci. FIGs.4A-F: Transcription factors show enriched co-occupancy of adjacent motifs with helical periodicity. (A) Schematic showing the procedure used to assess allelic co-occupancy enrichment. (B) Evaluation of magnitude (x-axis, Attorney Docket No.29618-0483WO1 / BWH 2023-476 median observed – expected co-occupancy) and statistical significance (y-axis, - log10(p)) of allelic co-occupancy for 2,373 TF pairs in K562. (C) Heat map showing median allelic co-occupancy enrichment for all 2,373 TF pairs in K562. In B-C, TF pairs that include a group of bHLH-LZ TFs with enriched co-occupancy, a set of repressor TFs with minimally enriched co-occupancy, and STAT5A-STAT5A are highlighted. (D) Analysis of K562 ACCESS-ATAC editing in 67,951 alleles at 460 K562 STAT5A ChIP-seq-bound loci with fixed 21-nt distance between motif centers. Top: motif logo showing positional nucleotide representation at the 460 loci. The STAT5A primary motifs are highlighted. Middle: Positional DddSs edit fractions for alleles at the 460 loci. STAT5A motif positions are highlighted, and a smoothed trendline is included. Bottom: Allelic editing patterns for all 67,951 reads at the 460 loci. Unedited and edited positions are shown. Alleles are separated by C2T-dominant and G2A-dominant editing patterns and sorted by observed – expected allelic co- occupancy. (E) Normalized median position-resolved observed – expected co- occupancy for all loci for which the listed TF is within 100-nt of any of the 64 analyzed TFs. The motif for the listed TF is always centered at 0 on the x-axis. The dominant period (Tp) and Bonferroni-corrected p-value from fast Fourier transform (FFT) analysis are shown for motifs upstream and downstream of the listed TF, and smoothed trendlines are included. (F) Evaluation of periodicity strength and Bonferroni-corrected p-value derived from FFT analysis of median position-resolved observed – expected co-occupancy measurements for 64 K562 and HepG2 TFs. For each TF, all loci bound by any other TF within 100-nt are included, analysis is performed separately for motifs upstream and downstream, and values are reported for the dominant period. (Fig5F_co-occupancy volcano). FIGs.5A-F: ACCESS-ATAC improves resolution of single-cell chromatin accessibility profiling. (A) Schematic of single-cell ACCESS-ATAC. Sequential treatment with DddSs, DddSs inhibitor, and Tn5 are used prior to encapsulation of nuclei following a minimally altered 10X scATAC workflow. (B) Barplot showing the number of valid cells after filtering by quality for each indexed sample within the scACCESS-ATAC experiment. (C) UMAP embedding of integrated scACCESS- ATAC data as colored by clusters. (D) Heatmap showing the genome-wide Spearman correlation among the two scACCESS-ATAC clusters, and K562 and HepG2 bulk ATAC-seq datasets. (E) Metaplots showing Tn5 insertions and DddSs editing at Attorney Docket No.29618-0483WO1 / BWH 2023-476 CTCF binding sites in K562 for different number of cells. (F) UMAP plot showing the scACCESS-ATAC data as colored by sgControl and sgZBTB7A treatments. DETAILED DESCRIPTION One key bottleneck in elucidating how human genomes encode epigenetic state is that, in spite of large consortia efforts such as ENCODE, chromatin accessibility data is not available at sufficient resolution and scale to match the staggering diversity of human cellular states and genotypes. Measurement of the accessibility of chromatin to DNA-modifying enzymes1has proven enormously valuable in dissecting cell type-specific functionalization of the genome. However, all current methods of genome-wide chromatin accessibility assessment have deficiencies. In the most common chromatin accessibility measurement techniques, DNase- seq2and ATAC-seq3, enzymatic cleavage is used to map the relative accessibility of nucleotides across the genome. ATAC-Seq uses hyperactive Tn5 transposase loaded in vitro with adapters for high-throughput DNA sequencing, which can fragment and tag a genome with sequencing adapters simultaneously (a process sometimes called “tagmentation”, see, e.g., Adey et al. Genome Biol.2010; 11:R119). This cleavage- based approach leads to two inherent limitations. First, a single sequencing read can provide at most two measurements of chromatin accessibility, from the two ends of the fragment. Second, these ends must be separated by >50-nucleotides in order to reliably map the resulting fragment to a reference genome. These limitations inherently constrain the resolution of cleavage-based techniques. They prevent analysis of co-occurrence of adjacent TF binding or nucleosome positioning events at an allelic level. Additionally, and crucially, while ATAC-seq can be scaled down to single-cell level (scATAC-seq), its resolution is inherently bounded. scATAC-seq cannot in principle surpass ~50-nt resolution, and typical resolution is far lower, precluding assessment of TF binding and nucleosome positioning in single cells. Instead, pseudo-bulk analysis, in which dozens to hundreds of similar cells are combined, is used to infer regulatory state of populations. Several non-cleavage-based chromatin accessibility measurement approaches have been developed to circumvent these inherent limitations; however, these techniques also have drawbacks. In Fiber-seq4and SMAC-seq20, chromatin is treated with a DNA N6-adenine methyltransferase, and long-read (e.g., nanopore) sequencing Attorney Docket No.29618-0483WO1 / BWH 2023-476 is used to distinguish methylated and unmethylated adenines. While these methods allow for single-molecule stenciling of chromatin accessibility, they rely on long-read sequencing platforms and are currently incompatible with PCR amplification due to erasure of the methylation. As such, Fiber-seq and SMAC-seq are prohibitively expensive for high-resolution genome-wide data and are incompatible with single-cell application. In NOMe-seq5 / single-molecule footprinting (SMF)21, chromatin is treated with a GpC methyltransferase, and either long-read sequencing22or bisulfite conversion is used to provide single-molecule chromatin accessibility profiling. When used with nanopore sequencing, NOMe-seq has the same drawbacks as Fiber-seq, and when paired with bisulfite, the use of a GpC methylase limits resolution to ~1 / 8 of nucleotides, bisulfite-induced DNA degradation limits sensitivity, and the lack of a simple approach to enrich reads in accessible chromatin makes it challenging to achieve high coverage in active chromatins. Thus, there is currently no approach capable of high-resolution, single-cell and single-molecule genome-wide chromatin accessibility profiling. Described herein are methods to improve the resolution of chromatin accessibility measurement through harnessing the genome editing abilities of double- stranded DNA cytosine deaminase (Ddd) bacterial toxins. Ddds are a recently discovered enzyme class and are uniquely capable of converting cytosines to uracils (C^U) in double-stranded DNA. The first discovered double-stranded DNA cytosine deaminase, DddA, natively acts as an inter-bacterial toxin and has been engineered into a tool for mitochondrial genome editing6. DddA has strong preference for editing at TC motifs, but enzyme engineering and ortholog mining efforts have identified Ddd enzymes with relaxed sequence preferences7-10. Aside from uses in mitochondrial editing, DddA has been fused with TFs in bacteria to map protein-DNA interaction sites.23The present methods harness this ability to mark nucleotides across the genome with accessible chromatin, providing both qualitative and quantitative advantages in epigenomic measurement. An exemplary method presented herein is referred to as an Accessible Chromatin by Cytosine Editing Site Sequencing (ACCESS) assay, in which intact nuclei are treated with purified Ddd enzymes to catalyze chromatin accessibility- dependent C^U editing genome-wide, and then immediately purified to remove the Ddd or treated with Ddd inhibitors (Dddi) to inactivate the Ddd. Crucially, through Attorney Docket No.29618-0483WO1 / BWH 2023-476 PCR-based conversion of uracils to thymines (U^T), e.g., using a uracil-tolerant polymerase, Ddd-treated samples are compatible with all nextgen sequencing pipelines. The accessibility of a given nucleotide in ACCESS is thus marked in sequenced reads as a C^T conversion, enabling basepair-resolution single-molecule stenciling of nucleosome positioning and TF binding. A number of uracil-tolerant polymerases have been described, including PHUSION DNA polymerase (a DNA-binding domain fused to a Pyrococcus-like proofreading polymerase, available from ThermoFisher); Kapa HiFi Hotstart Uracil+ DNA Polymerase (a type B polymerase that has a mutation in the uracil-binding pocket, available from the Roche Sequencing Store); KOD Multi Epi DNA Polymerase (a type B polymerase that has a mutation in the uracil-binding pocket, available from Sekisui diagnostics); and Q5U Hot Start High-Fidelity DNA Polymerase (a type B polymerase that has a mutation in the uracil-binding pocket, commercially available from New England Biolabs and described in WO / 2003 / 089637). See also Wardle et al., Nucleic Acids Res.2008 Feb; 36(3): 705– 711, US8481685, and WO2017121836. Ddd enzymes that can be used in the present methods include those described in Vaisvila et al., Mol Cell.2024 Mar 7;84(5):854-866.e7; Mok et al., Nature.2020 Jul;583(7817):631-637; and Huang et al., Cell.2023 Jul 20;186(15):3182-3195.e14. Exemplary Ddd enzymes include gp317; xp12da; AvDa02; BadTF3; DaDa01; APOBEC3A; MGYPDa408; MGYPDa687; MGYPDa917; MGYPDa624; KcDa01; HgDa01; TuDa01; AbcDa01; XcDa01; DddA; StsDa01; KsDa01; PwDa01; CaDa01; SrDa01; NgDa01; NsDa01; SzDa01; SpDa01; HmDa01; HmDa02; HmDa03; AmDa01; SjDa01; BsDa01; PpDa01; SaDa01; CpDa01; EcDa01; EcDa02; NgDa02; PaDa01; AsDa01; EcDa03; ScDa01; BpDa01; ScDa02; StsDa02; OTT-1508; NpDa01; HgmDa01; AdDa01; LbsDa01; BmDa01; BsDa02; PlDa01; HmDa04; AmDa02; AcDa01; MGYPDa01; MGYPDa13; BbDa01; OlDa01; BpDa02; AmDa03; MsDa01; KsDa02; WcDa01; BbDa02; PrDa01; VRDa02; VRDa03; VRDa04; VRDa05; VRDa06; MsDa02; SqDa01; TeDa01; StsDa03; SaDa02; AbDa01; MGYPDa23; PpDa02; EcDa04; MGYPDa829; MGYPDa02; MGYPDa03; BcDa01; IfDa01; PcDa01; StsDa04; AmDa04; XinDa01; XjaDa01; RhDa01; MGYPDa04; MGYPDa05; BaDa01; LbDa01; CbDa01; HcDa01; MGYPDa06; CseDa01; AvDa01; LbDa02; MGYPDa07; FbDa01; IfDa02; RsDa01; NoDa01; PfDa01; ScDa03; Attorney Docket No.29618-0483WO1 / BWH 2023-476 PsDa01; PvDa01; CdDa01; FbDa02; AaDa01; AzDa01; BdDa01; MGYPDa08; AbDa02; WWTPDa01; WWTPDa02; WWTPDa03; WWTPDa04; WWTPDa05; WWTPDa06; WWTPDa07; MGYPDa09; AoDa01; PaDa02; MGYPDa10; MGYPDa11; MGYPDa12; PbDa02; MGYPDa14; CrDa01; MGYPDa15; MGYPDa16; MGYPDa17; BaDa02; VsDa01; PdDa01; MGYPDa18; MGYPDa19; HmDa06; MGYPDa25; MGYPDa26; MsddA; AshDa01; MGYPDa21; PpDa03; SbDa01; BlDa01; PpDa04; CsDa01; MGYPDa22; FlDa01; PeDa01; MGYPDa24; AaDa02; MmgDa01; PbDa01; BcDa02; LsfDa01; RaDa01; SmgDa01; MmgDa02; SaDa03; HgmDa02; SsdA; CgmDa01; FbiDa01; PvmDa01; SoCaDa01; SoCaDa02; SoCaDa03; SoCaDa04; SoCaDa05; SoCaDa06; SoCaDa07; SoCaDa08; and SoCaDa09 (see Vaisvila et al., Mol Cell.2024 Mar 7;84(5):854-866.e7 and WO2023097226); or Ddd1, Ddd7, Ddd8, and Ddd (Huang et al., Cell.2023 Jul 20;186(15):3182-3195.e14). See also WO2023169410. In some embodiments, the DddA-like or dsDDDs (see Huang et al.). The sample cane be exposed to the Ddd, e.g., for 5-90 minutes, e.g., 15-60 minutes or 15-30 minutes. In some embodiments, an optimized concentration of DddSs of 250 nM is used. For example, DddA from Burkholderia cenocepacia (strain H111) can be used; an exemplary sequence is shown herein. DddI that can be used in the present methods include the DddI from Burkholderia cenocepacia (strain H111), an exemplary sequence for which is provided herein. The DddI is added in at least a 1:1 molar ratio with the Ddd, preferably at least 2:1 or 10:1 DddI to Ddd, preferably for at least 2 minutes, up to 5, 10, 1520, or 30 minutes or more. The intact nuclei can be obtained using known methods from any eukaryotic (nucleated) cell, including plant, animal, insect, bacterial, and fungal cells. ACCESS can be combined with ATAC-seq (ACCESS-ATAC) in bulk and single-cell applications. ATAC-seq is inherently limited in resolving single-cell TF binding and nucleosome positioning profiles, but ACCESS-ATAC eliminates this barrier, increasing the resolution of single-cell chromatin accessibility measurement over five-fold and paving the way for nucleotide-resolution single-cell chromatin analysis. ATAC-seq is described, e.g., in US Patent Nos.10,150,995, 10,619,207, and 10,738,357; Buenrostro et al., Nature Methods 10:1213–1218 (2013); Buenrostro et Attorney Docket No.29618-0483WO1 / BWH 2023-476 al., Current Protocols in Molecular Biology 21.29.1-21.29.9 (2015); and Buenrostro et al., Nature 523:486–490 (2015). ACCESS is, uniquely, a non-cleavage-based technique that is compatible with standard genomic workflows. Leveraging this attribute, ACCESS can be combined with ATAC-seq as demonstrated herein. In this ACCESS-ATAC workflow, ATAC transposition is used to enrich sequencing in accessible chromatin while ACCESS provides high-resolution stenciling of chromatin accessibility within each sequenced molecule. ACCESS-ATAC provides a 5-10-fold quantitative improvement in resolution over ATAC-seq through allowing measurement of 10-20 accessible nucleotides per NGS read. ACCESS-ATAC offers qualitative improvements in the ability to map single-molecule TF binding and nucleosome footprints, which ATAC- seq is inherently incapable of doing. ACCESS-ATAC is compatible with the most widely used commercial single- cell ATAC-seq platform (10X Genomics Chromium Single Cell ATAC), enabling nucleotide-resolution single-cell chromatin accessibility measurement. Thus, additionally provided herein are single-cell ATAC-seq and multiome (ATAC-seq and RNA-seq) assays that incorporate ACCESS methods. Such methods can comprise contacting the isolated nuclei with an insertional enzyme complex comprising a transposase, e.g., Tn5 transposase, that is loaded with dsDNA sequencing adapters to fragment and tag the chromatin; amplifying the chromatin fragments with a uracil- tolerant polymerase to convert uracils to thymines (U^T); and optionally sequencing the fragments. The present methods can include sequencing the DNA. As used herein, “sequencing” includes any method of determining the sequence of a nucleic acid. Any method of sequencing can be used in the present methods, including chain terminator (Sanger) sequencing and dye terminator sequencing. In preferred embodiments, Next Generation Sequencing (NGS), a high-throughput sequencing technology that performs thousands or millions of sequencing reactions in parallel, is used. Although the different NGS platforms use varying assay chemistries, they all generate sequence data from a large number of sequencing reactions run simultaneously on a large number of templates. Typically, the sequence data is collected using a scanner, and then assembled and analyzed bioinformatically. Thus, the sequencing reactions are performed, read, assembled, and analyzed in parallel; Attorney Docket No.29618-0483WO1 / BWH 2023-476 see, e.g., US 20140162897, as well as Voelkerding et al., Clinical Chem., 55: 641- 658, 2009; and MacLean et al., Nature Rev. Microbiol., 7: 287-296 (2009). Some NGS methods require template amplification and some do not. Amplification- requiring methods include pyrosequencing (see, e.g., U.S. Pat. Nos.6,210,89 and 6,258,568; commercialized by Roche); the Solexa / Illumina platform (see, e.g., U.S. Pat. Nos.6,833,246, 7,115,400, and 6,969,488); Life Technologies' IonTorrent platform; and the Supported Oligonucleotide Ligation and Detection (SOLiD) platform (Applied Biosystems; see, e.g., U.S. Pat. Nos.5,912,148 and 6,130,073). Methods that do not require amplification, e.g., single-molecule sequencing methods, include nanopore sequencing (e.g., using Oxford Nanopore Technologies’ platform); NANOSTRING, which uses oligonucleotide barcodes tagged with antibodies or in situ hybridization (ISH) probes to quantify RNA or protein expression in two- or three-dimensional space; HELISCOPE (U.S. Pat. Nos.7,169,560; 7,282,337; 7,482,120; 7,501,245; 6,818,395; 6,911,345; and 7,501,245); real-time sequencing by synthesis (see, e.g., U.S. Pat. No.7,329,492); single molecule real time (SMRT) DNA sequencing methods using zero-mode waveguides (ZMWs); and other methods, including those described in U.S. Pat. Nos.7,170,050; 7,302,146; 7,313,308; and 7,476,503). See, e.g., US 20130274147; US20140038831; Metzker, Nat Rev Genet 11(1): 31-46 (2010); Margulies et al (Nature 2005437: 376-80); Ronaghi et al (Analytical Biochemistry 1996242: 84-9); Shendure et al (Science 2005309: 1728- 32); Imelfort et al (Brief Bioinform.200910:609-18); Fox et al (Methods Mol Biol. 2009; 553:79-108); Appleby et al (Methods Mol Biol.2009; 513:19-39); Morozova et al (Genomics.200892:255-64); and US20180355424. Alternatively, hybridization-based sequence methods or other high-throughput methods can also be used, e.g., microarray analysis. KITS Further, provided herein are kits comprising reagents for use in the present methods. Such kits can include, for example: reagents for isolating nuclei from a population of cells; double-stranded DNA cytosine deaminase (Ddd); a Ddd inhibitor (Dddi); a uracil-tolerant polymerase; and suitable reaction buffers. In methods wherein ACCESS is combined with ATAC-SEQ, the kit can also comprise an insertional enzyme complex (i.e., comprising Tn5 transposase and transposon tags and sequencing adaptors), and additional or different suitable reaction buffers. Attorney Docket No.29618-0483WO1 / BWH 2023-476 The kit can also comprise one or more of the following: a cell lysis buffer; a surfactant (such as Tween20); a protease inhibitor cocktail; an insertional enzyme complex comprising an affinity tag; and an insert element comprising a nucleic acid, wherein the nucleic acid comprises a predetermined sequence. Exemplary sequences In some embodiments, the sequence of a protein or nucleic acid used in a composition or method described herein is at least 80%, 85%, 90%, 95%, 97%, 98%, or 99% identical, preferably at least 95% identical, to a sequence set forth herein. To determine the percent identity of two amino acid sequences, or of two nucleic acid sequences, the sequences are aligned for optimal comparison purposes (e.g., gaps can be introduced in one or both of a first and a second amino acid or nucleic acid sequence for optimal alignment and non-homologous sequences can be disregarded for comparison purposes). In a preferred embodiment, the length of a reference sequence aligned for comparison purposes is at least 80% of the length of the reference sequence, and in some embodiments is at least 90% or 100%. The amino acid residues or nucleotides at corresponding amino acid positions or nucleotide positions are then compared. When a position in the first sequence is occupied by the same amino acid residue or nucleotide as the corresponding position in the second sequence, then the molecules are identical at that position (as used herein amino acid or nucleic acid “identity” is equivalent to amino acid or nucleic acid “homology”). The percent identity between the two sequences is a function of the number of identical positions shared by the sequences, taking into account the number of gaps, and the length of each gap, which need to be introduced for optimal alignment of the two sequences. The comparison of sequences and determination of percent identity between two sequences can be accomplished using a mathematical algorithm. For example, the percent identity between two amino acid sequences can be determined using the Needleman and Wunsch ((1970) J. Mol. Biol.48:444-453 ) algorithm which has been incorporated into the GAP program in the GCG software package (available on the world wide web at gcg.com), using the default parameters, e.g., a Blossum 62 scoring matrix with a gap penalty of 12, a gap extend penalty of 4, and a frameshift gap penalty of 5. Attorney Docket No.29618-0483WO1 / BWH 2023-476 In some embodiments, the sequence of a protein or nucleic acid used in a composition or method described herein has up to 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions or deletions as compared to a sequence set forth herein. In some embodiments, the substitutions are conservative substitutions. DddA (from Burkholderia cenocepacia (strain H111)) MYEAARVTDPIDHTSALAGFLVGAVLGIALIAAVAFATFTCGFGVALLAGMMAGIGA QALLSIGESIGKMFSSQSGNIITGSPDVYVNSLSAAYATLSGVACSKHNPIPLVAQG STNIFINGRPAARKDDKITCGATIGDGSHDTFFHGGTQTYLPVDDEVPPWLRTATDW AFTLAGLVGGLGGLLKASGGLSRAVLPCAAKFIGGYVLGEAFGRYVAGPAINKAIGG LFGNPIDVTTGRKILLAESETDYVIPSPLPVAIKRFYSSGIDYAGTLGRGWVLPWEI RLHARDGRLWYTDAQGRESGFPMLRAGQAAFSEADQRYLTRTPDGRYILHDLGERYY DFGQYDPESGRIAWVRRVEDQAGQWYQFERDSRGRVTEILTCGGLRAVLDYETVFGR LGTVTLVHEDERRLAVTYGYDENGQLASVTDANGAVVRQFAYTNGLMTSHMNALGFT SSYVWSKIEGEPRVVETHTSEGENWTFEYDVAGRQTRVRHADGRTAHWRFDAQSQIV EYTDLDGAFYRIKYDAVGMPVMLMLPGDRTVMFEYDDAGRIIAETDPLGRTTRTRYD GNSLRPVEVVGPDGGAWRVEYDQQGRVVSNQDSLGRENRYEYPKALTALPSAHFDAL GGRKTLEWNSLGKLVGYTDCSGKTTRTSFDAFGRICSRENALGQRITYDVRPTGEPR RVTYPDGSSETFEYDAAGTLVRYIGLGGRVQELLRNARGQLIEAVDPAGRRVQYRYD VEGRLRELQQDHARYTFTYSAGGRLLTETRPDGILRRFEYGEAGELLGLDIVGAPDP HATGNRSVRTIRFERDRMGVLKVQRTPTEVTRYQHDKGDRLVKVERVPTPSGIALGI VPDAVEFEYDKGGRLVAEHGSNGSVIYTLDELDNVVSLGLPHDQTLQMLRYGSGHVH QIRFGDQVVADFERDDLHREVSRTQGRLTQRSGYDPLGRKVWQSAGIDPEMLGRGSG QLWRNYGYDGAGDLIETSDSLRGSTRFSYDPAGRLISRANPLDRKFEEFAWDAAGNL LDDAQRKSRGYVEGNRLLMWQDLRFEYDPFGNLATKRRGANQTQRFTYDGQDRLITV HTQDVRGVVETRFAYDPLGRRIAKTDTAFDLRGMKLRAETKRFVWEGLRLVQEVRET GVSSYVYSPDAPYSPVARADTVMAEALAATVIDSAKRAARIFHFHTDPVGAPQEVTD EAGEVAWAGQYAAWGKVEATNRGVTAARTDQPLRFAGQYADDSTGLHYNTFRFYDPD VGRFINQDPIGLNGGANVYHYAPNPVGWVDPWGLAGSYALGPYQISAPQLPAYNGQT VGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNN PEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKG GC (SEQ ID NO:1) DddSs (from Simiaoa sunii) MSLPEYDGTTTHGVLVLDDGTQIGFTSGNGDPRYTNYRNNGHVEQKSALYMRENNIS NATVYHNNTNGTCGYCNTMTATFLPEGATLTVVPPENAVANNSRAIDYVKTYTGTSN DPKISPRYKGN (SEQ ID NO:2) MGYPDa829 (metagenomic, source unknown) DFHTYYVGTESVLVHNNGGNCMKTASALAQGGSKGGSTAINLPEYDGKTTHGVLVLD NGTQVQLVSGNGDPRYTNYRNNGHVEQKAAIYMRENNISNATVYHNNTNGTCGYCNT MTATFLPEGATLTVVPPKNAVANNSRAIAYVKTYTGTSNDPKMSSRYKGN (SEQ ID NO:3) Attorney Docket No.29618-0483WO1 / BWH 2023-476 Double-stranded DNA deaminase immunity protein (DddI) (from Burkholderia cenocepacia (strain H111)) MYADDFDGEIEIDEVDSLVEFLSRRPAFDANNFVLTFEESGFPQLNIFAKNDIAVVY YMDIGENFVSKGNSASGGTEKFYENKLGGEVDLSKDCVVSKEQMIEAAKQFFATKQR PEQLTWSEL (SEQ ID NO:4) Ddd sequences used for in vitro transcription and translation. T7 promoter stub : AGGGCTTAAGTATAAGGAGGAAAAAAT (SEQ ID NO:5) Ddd enzyme sequence T7 terminator stub : GGTTAACTAGCATAACCCCTCT (SEQ ID NO:6) CrDa01_T7 AGGGCTTAAGTATAAGGAGGAAAAAATATGGGTGCGGGAGCGGAAGCGATATACAAA GCTAGTAAGGACGGGTTCAAGGAGGCGAAGAATGCGTTGCGCACTGTAAAGGACGAG AAGAACGGATTTGTAGACGGAGTGGCCGATACCAACGCTCCTTGGAAGGCGATAGAT GACTACAGAAACCAGCGCCAGGGGTTAGAAGTATTGCCGGAGGATTATGAGTTTATC AAAGGTGATGGAAAAAACACAGTGGCATTTAATGATACCTGCGACAAACGTTACTTC GGTGTCAACTCGACATTGCGCACAGATGCTGAGAAAGGTCTTGCAAAGAAATATTTT AACAAGCTGGTTGAAGAGGGTTACTTCCCGAAGGGAAGCATCTACCCTAGAGGAAAA GCTCAATTTATCACACATGCAGAGGGTTACACGCTGATTAAGGCATACGAAGAAAAT GGCATCGATATTGGAAAATCAGTTACGATATACGTTGATCGTCCGACTTGTGGATTT TGCCAAAATAACCTGCCCAAGTTGCAGGCAAGCATGGGCATTGACGTGTTGACTGTC ATCAATAAAAATGGCGATGTCTTCCGGTGCGATTTGCGGAAATACACAAGTATGACG GTTTCTGACGTGAACGCAGAGAAACTGATAAAGATAAAATCTAAGTAGGGTTAACTA GCATAACCCCTCT (SEQ ID NO:7) CseDa01_T7 AGGGCTTAAGTATAAGGAGGAAAAAATATGTACGATGCGTTCGGTAACGTTCGTAAC CAAAAGGAAACCCATCACAATCGCATTTTATACACGGGTCAACAATACGACAAAGAA TCAAATCAATACTATCTGCGGGCTAGATACTACAACCCAACCCTTGGCCGTTTCACT CAAGAGGATGTTTACCGTGGGGCGGGGTTAAATTTGTACGATTATTGCAAAGGCAAC CCGGTTATATGGTACGACCCTTCAGGCTACGAAAATTGTAAGATAGATAAAGGGAAT GTATCCAATAAAGAGAACGGATTCGAGGATGAGATCATCCAGGCCAAGATCGCCCAA CAAACCGCCTTCGAATTGGCGGAAAAGTATGGGTATAACGGAGCGCCACCGAAGAGA AAAACAGTTGCCTCAGATGGGGAGATAAATACGCTGTCTGGGTGGAAAAAATTAAAG GATAACACCGATTTCATCCGTGTGTCGCCCGAGGATATTATTGACAAATCTCAGGAG ATAGGACACAATCTTAGAAACGCTGGAGCAAACGATCAGGGCATTAAGGGTAAATAC AACGCTTCTCACGCTGAGAAACAATTGTCTTTAAAGACTGACAAACCTATTGGCATC TCTCAACCAATGTGCCAGGATTGCCAGAATTACTTTCGGTGTTTAGCTATTTCCGAA AGTAAGAACTTCGTTACTGCAGATACCAATATGGTACGTATTTTCAAGCCGGACGGT TCCGTGATCAATTACAAGTCCGACGACTCCTTTTCGATTATAAAAATAAAGAACAAG ATCAAATAAGGTTAACTAGCATAACCCCTCT (SEQ ID NO:8) Attorney Docket No.29618-0483WO1 / BWH 2023-476 DddSs_T7 AGGGCTTAAGTATAAGGAGGAAAAAATATGATTAGTTTACCAGAATATGATGGAACA ACTACACATGGAGTATTAGTTTTAGATGATGGAACACAAATAGGATTTACATCAGGT AATGGAGACCCACGGTATACTAATTATCGTAATAATGGTCATGTGGAACAGAAGTCA GCATTATATATGAGAGAAAATAATATATCCAATGCAACAGTATATCACAATAATACA AATGGTACTTGTGGTTATTGTAACACTATGACAGCAACTTTCTTGCCAGAAGGAGCA ACTTTAACTGTAGTTCCTCCTGAGAACGCCGTTGCAAATAATAGTAGGGCTATAGAT TATGTTAAAACGTATACGGGAACAAGTAATGACCCGAAGATAAGTCCAAGATATAAA GGAAACTGAGGTTAACTAGCATAACCCCTCT (SEQ ID NO:9) FlDa01_T7 AGGGCTTAAGTATAAGGAGGAAAAAATATGGTTTATGATAAACCGGAGCCTATCAAA AACCTGATCACATGGGTCTACGAAGGTGGCTCTTTCGTGCCGTCCGCGAAGATTATT GGCGAACATAAATTCTCTATAATAAATGACTACATAGGCCGTCCAATCCAAGTCTAT AATGAGGTAGGGGATGTCGTGTGGGAGACCGACTACGATATCTATGGTGGGCTTCGT AACCTGAAAGGAGACAAGTCTTTTATACCTTTCCGGCAATTGGGCCAATACGAGGAT GTAGAAACTGGCTTATACTATAACCGCCATCGTTATTACAATCCCGAAAGCGGTGGG TATATTAGCCAGGATCCTATAGGACTGCTTGGTGGTAGTGCGTCTTACAAGTACGTC CATGATTGTAACAACTGTGTAGATATTTTCGGGTTGAATCCTGTAATCTTTTCGGAA GAATTGAGTAAGATCGCACAGGAAGCGCATAATGTCCTGTTAGAACCCGGAAAGTCT CCTCGTGGATTTAACAACTCTACGGTCAGTGTAGCAAAGGTCGATGTCAACGGAGTT AGCACTTTATACGCATCGGGAAATGGAGCGTCGCTTTCTCCCGCCCAGCGTACTAAG TTGGTAGAATTAGGGGTCCCCGAAGAGAACATATTCAGTGGCAAACGGTTCAAAGAA ATCATTGATGGGGACACAGGAACCCTTACCAAGCTGTCTAACCACGCTGAAAGAGTG ATCGAACGCAATATACCGAAGGACGCCAGCATAAAAGAGTGGGGGATTAGTTGGGCA AGTAAGCAGAAAAATGAAATGTGCAACAACTGCAAGACCCACTTTGGATGCAAATAG GGTTAACTAGCATAACCCCTCT (SEQ ID NO:10) MGYPDa01_T7 AGGGCTTAAGTATAAGGAGGAAAAAATATGTATTTCGATGGGGAGACCGGCCTGCAT TACAATCGTTTCCGTTATTACGACCCGGTGGTGGGCCGCTTTGTTCACCAAGACCCC ATTGGATTTGGCGGCGGGAATAATTTCTACATATACGGCCCGAACAGTGGCAGTTGG TATGACCCTTTTGGCCTGGCAAAGAGACCTCCCCATAAATTAAAAGCTAAGACTAAG GACGAGAACGGGAATGTGAAAAGTGAACGCGACTACGTCTCTGGCGGAATGACTGAA GAGGATAAAAAACTTGGGTACCCACTTTGCTCCCTGGTTACGCATACGGAGAGAAAG GCCCTGAAGGAGGATGACTATTCACCGACGGACACGATTGAAATGCACGGGGAGTAT GCTCCCTGTTCGCACTGCAAAGGTGCGATGAATACAGCGGTCGATTCTGGTAGAGTC GCACGGATCTTATATTACTGGAAGGGAAAGATCTGGGAGGCAGGCGCTGCTGCTCGC AAACTTAGAAATAAGAAAAAGCGCAGCGCAGGAATGTGTTCCAATGATTAGGGTTAA CTAGCATAACCCCTCT (SEQ ID NO:11) MGYPDa029_T7 AGGGCTTAAGTATAAGGAGGAAAAAATatgGATTTCCATACGTATTATGTTGGTACA GAAAGTGTATTGGTCCACAACAACGGGGGAAACTGTATGAAAACGGCTTCCGCGCTG GCCCAAGGTGGTTCAAAGGGGGGATCTACCGCAATCAATCTGCCGGAATACGATGGA Attorney Docket No.29618-0483WO1 / BWH 2023-476 AAGACGACACACGGGGTCTTGGTCTTGGACAATGGGACTCAAGTTCAACTTGTGTCA GGCAATGGAGATCCTCGGTACACGAATTATCGCAACAATGGGCATGTCGAACAGAAA GCAGCGATCTATATGCGGGAAAATAACATATCAAACGCGACCGTTTACCATAATAAC ACGAATGGCACTTGCGGCTACTGTAACACCATGACCGCTACTTTCTTACCAGAGGGT GCGACTCTGACAGTAGTTCCCCCAAAGAACGCCGTTGCGAATAACAGCCGGGCTATT GCTTACGTCAAAACATATACCGGGACTAGCAATGACCCTAAAATGAGTTCGAGATAC AAAGGCAACtagGGTTAACTAGCATAACCCCTCT (SEQ ID NO:12) T7 promoter extension fw primer : GCGAATTAATACGACTCACTATAGGGCTTAAGTATAAGGAGGAAAA (SEQ ID NO:13) T7 terminator extension rv primer : CCAAATAAACCCCTCCGTTTAGAGAGGGGTTATGCTAGTTAACC (SEQ ID NO:14) DddSs and DddSs-I insert sequences used to clone pETDuet_DddSs-DddSs-I 6x His tag – CATCACCATCATCACCAC (SEQ ID NO:15) DddSs – ATTAGTTTACCAGAATATGATGGAACAACTACACATGGAGTATTAGTTTTAGATGAT GGAACACAAATAGGATTTACATCAGGTAATGGAGACCCACGGTATACTAATTATCGT AATAATGGTCATGTGGAACAGAAGTCAGCATTATATATGAGAGAAAATAATATATCC AATGCAACAGTATATCACAATAATACAAATGGTACTTGTGGTTATTGTAACACTATG ACAGCAACTTTCTTGCCAGAAGGAGCAACTTTAACTGTAGTTCCTCCTGAGAACGCC GTTGCAAATAATAGTAGGGCTATAGATTATGTTAAAACGTATACGGGAACAAGTAAT GACCCGAAGATAAGTCCAAGATATAAAGGAAAC (SEQ ID NO:16) DddSs-I : ATGTTAGTAGAACATTTTATGGGACAAAAGGAATGTGATAGTCTTGAAGAATTGAGA GAAGTTTTAAGCGAGAGGACAGAAAAAGGTGTTAATGAGTTTATTATATCAACGCAT GAACAATTTCCGTATATGATAATGTCTGTGAAGGAAAAATATGCATGTTTAAGCTAT TTTCGAGAAGAAGATGACCCAGGGTATTCTTCAGTTAATGCTAATCCTGTTTTGGAT GCAGATGGTATTAGTATATTTTACACAAATACTGATAGTGAAGAAATAGAAGTTGCA AATTATTCAATTGTAAAAATTGAAGATGCAGTGTCTGCAGTTGAGGAATTTTTTGAA ACACTACAATTACCAAAATGTATTGAATGGGAAGAATTA (SEQ ID NO:17) Construct: ATGGGCAGCAGCCATCACCATCATCACCACAGCCAGGATCCGATTAGTTTACCAGAA TATGATGGAACAACTACACATGGAGTATTAGTTTTAGATGATGGAACACAAATAGGA TTTACATCAGGTAATGGAGACCCACGGTATACTAATTATCGTAATAATGGTCATGTG GAACAGAAGTCAGCATTATATATGAGAGAAAATAATATATCCAATGCAACAGTATAT CACAATAATACAAATGGTACTTGTGGTTATTGTAACACTATGACAGCAACTTTCTTG CCAGAAGGAGCAACTTTAACTGTAGTTCCTCCTGAGAACGCCGTTGCAAATAATAGT AGGGCTATAGATTATGTTAAAACGTATACGGGAACAAGTAATGACCCGAAGATAAGT CCAAGATATAAAGGAAACTGAGCGGCCGCATAATGCTTAAGTCGAACAGAAAGTAAT CGTATTGTACACGGCCGCATAATCGAAATTAATACGACTCACTATAGGGGAATTGTG AGCGGATAACAATTCCCCATCTTAGTATATTAGTTAAGTATAAGAAGGAGATATACA Attorney Docket No.29618-0483WO1 / BWH 2023-476 TATGTTAGTAGAACATTTTATGGGACAAAAGGAATGTGATAGTCTTGAAGAATTGAG AGAAGTTTTAAGCGAGAGGACAGAAAAAGGTGTTAATGAGTTTATTATATCAACGCA TGAACAATTTCCGTATATGATAATGTCTGTGAAGGAAAAATATGCATGTTTAAGCTA TTTTCGAGAAGAAGATGACCCAGGGTATTCTTCAGTTAATGCTAATCCTGTTTTGGA TGCAGATGGTATTAGTATATTTTACACAAATACTGATAGTGAAGAAATAGAAGTTGC AAATTATTCAATTGTAAAAATTGAAGATGCAGTGTCTGCAGTTGAGGAATTTTTTGA AACACTACAATTACCAAAATGTATTGAATGGGAAGAATTATAG (SEQ ID NO:18) EXAMPLES The invention is further described in the following examples, which do not limit the scope of the invention described in the claims. Materials and Methods The following materials and methods were used in the Examples below. Expression of Ddd enzymes through in vitro transcription and translation Amino acid sequences of six Ddd enzymes were obtained from published work. Nucleotide sequence was codon optimized for E. coli expression using IDT Codon Optimization tool, and open reading frames containing T7 promoter and terminator stubs were ordered from Twist Bioscience. T7 promoter and terminator extension primers were used to PCR amplify each Ddd enzyme, and PCR purified products were used as input for in vitro transcription and translation using the NEB PURExpress® In Vitro Protein Synthesis kit. Activity of bulk protein obtained from in vitro transcription and translation was titrated using the in vitro dsDNA deamination assay, and concentrations of each Ddd enzyme that yielded ~50% in vitro dsDNA deamination were used in ACCESS-WGS. Bacterial expression and purification of Ddd enzymes A plasmid allowing for dual bacterial expression of DddA and DddA-I (pETDuet_DddA-DddA-I) was a kind gift from Joseph Mougous. We subcloned pETDuet_DddSs-DddSs-I by ordering E. coli codon optimized sequences for this enzyme (adding a His tag) and inhibitor based on Genbank sequences. The sequence of these inserts is provided herein. The enzymes and inhibitors were purified based on a published protocol56. The plasmids were transformed into BL21 Codon Plus cells for recombinant protein expression. The cells were cultured in LB / Ampicillin medium to an optimal cell density at OD600 = 0.6 and induced at 18°C overnight. The cells were harvested, resuspended in lysis buffer (50 mM Tris pH 8.0, 500 mM NaCl, 10 mM imidazole, 1 mM TCEP, 1mM PMSF, 1 tablet of protease inhibitor) and lysed by Attorney Docket No.29618-0483WO1 / BWH 2023-476 passing through a french press three times. After centrifugation at 27,000g for 40 min to remove debris, the lysate was passed through 4 mL of Ni-NTA resin (MCLAB Catalog #NINTA-100) three times. The resin was washed with 12 column volumes (CV) of wash buffer 1 (50 mM Tris pH 8.0, 500 mM NaCl, 10 mM imidazole, 1 mM TCEP) and 5 CV of wash buffer 2 (50 mM Tris pH 8.0, 500 mM NaCl, 20 mM imidazole, 1 mM TCEP). The deaminase-inhibitor complex was eluted using elution buffer (50 mM Tris pH 8.0, 500 mM NaCl, 300 mM imidazole, 1 mM TCEP) and dialyzed in 8 M urea buffer (50 mM Tris pH 7.5, 500 mM NaCl, 10 mM imidazole, 1 mM TCEP, 8 M urea) overnight. The dialyzed sample was passed through Ni-NTA resin three times and washed with 12 CV of 8 M urea buffer. For Deaminase (DddA / DddSs) Isolation: The Ni-NTA resin was washed sequentially with 6 CV of 6 M, 4 M, 2 M, and 1 M urea buffer. After a final wash with 12 CV of 0 M urea buffer, the deaminase was eluted using elution buffer. For Deaminase Inhibitor (DddA-I / DddSs-I) Isolation: The flowthrough from the Ni-NTA purification post 8 M urea dialysis and the 8 M urea wash fractions were pooled. The pooled fractions underwent step-dialysis using 6 M, 4 M, 2 M, 1 M, and 0 M urea buffer. The deaminase and deaminase inhibitor samples were each dialyzed in 0 M urea buffer overnight and concentrated. The proteins were further purified on a Superdex 7510 / 300 GL column (Cytiva, Catalog #17517401) with a mobile phase of 20 mM Tris pH 7.5, 200 mM NaCl, 2 mM DTT, 5% glycerol, flash frozen, and stored at - 80°C until use. All purified proteins were characterized by SDS-PAGE and ESI-MS (Q Exactive, Thermo Scientific). Data analysis of the ESI-MS spectra was performed using UniDec software57. In vitro dsDNA deamination assay Deamination of double-stranded DNA (dsDNA) by Ddd enzymes was measured through a DpnII restriction digest assay. A brief summary follows, and a full protocol is provided below. Two complementary 90-nt oligos were ordered from IDT with a central GATC sequence surrounded by A:T basepairs. These oligos were annealed at 25 uM in IDT Nuclease Free Duplex Buffer through heating to 95 deg C followed by gradual cooling to 25 deg C. Deamination was performed in a 10 uL reaction with final concentrations of 20 mM Tris-HCl pH 7.4, 200 mM NaCl, 1 mM DTT, 1 uM Attorney Docket No.29618-0483WO1 / BWH 2023-476 dsDNA substrate. Deamination was detected by DpnII (NEB) restriction digest. Deamination efficiency was measured using D1000 Tapestation (Agilent). ACCESS followed by whole genome sequencing (ACCESS-WGS) A full protocol for ACCESS-WGS is provided herein. A brief summary follows. HCT116 cells were harvested by trypsinization, and 2.5*10^5 cells per reaction were lysed for 3 minutes on ice followed by dilution and centrifugation to recover intact nuclei. Nuclei were then treated for 30 min at 37 deg C in a 25 uL reaction with a concentration of Ddd enzyme titrated to have ~50% deamination activity in the dsDNA deamination assay. Treated genomic DNA was purified using the Zymo Genomic DNA clean and concentrator 25 kit. Tagmentation was performed by 5 minute incubation at 55 deg C of 100 ng of ACCESS-treated genomic DNA in a 50 uL reaction using 0.5 uL of Tagmentase (Tn5 transposase) - loaded (Diagenode). Tagmented genomic DNA was purified using the Zymo DNA clean and concentrator 5 kit. Nextgen sequencing library prep was then performed using 2x Q5U mastermix (NEB). ACCESS-ATAC Full protocols for ACCESS-ATAC concurrent and sequential treatment are provided herein. A brief summary follows. K562 or HepG2 cells were harvested, and 2.5*10^5 cells per reaction were lysed for 3 minutes on ice followed by dilution and centrifugation to recover intact nuclei. For concurrent ACCESS-ATAC, nuclei were treated for 30 min at 37 deg C in a 25 uL reaction with 250 nM DddSs and 1.25 uL of Tagmentase (Tn5 transposase) - loaded (Diagenode). For sequential ACCESS-ATAC, nuclei were treated for 15-30 min at 37 deg C in a 25 uL reaction with 250 nM DddSs, then DddSs activity was quenched by treatment for 2 min at 37 deg C with 5- 10-fold molar excess of DddSs-I. Then nuclei were treated for an additional 30 min at 37 deg C with 1.25 uL of Tagmentase (Tn5 transposase) - loaded (Diagenode). ACCESS-ATAC-treated DNA was purified using the Zymo DNA clean and concentrator 5 kit. Nextgen sequencing library prep was then performed using 2x Attorney Docket No.29618-0483WO1 / BWH 2023-476 Q5U mastermix (NEB). Nextgen sequencing was performed using Element AVITI at Quintara Biosciences or Ultima Genomics UG100 at Ultima as specified. Single-cell ACCESS-ATAC K562 lentiviral CRISPR-Cas9 knockout lines were made for 12 gRNAs, each cloned into lentiCRISPRv2-FE PuroR (Addgene 186746). gRNA sequences are listed in Table 1. Lentiviral production and Puromycin selection were performed as described. On the day of the experiment, 2.0*10^5 of each cell gRNA-treated cell population, as well as HepG2-wt and K562-wt cells, were lysed, and the equivalent gRNA-treated populations of K562 and HepG2 nuclei were combined. Nuclei were treated as in the sequential ACCESS-ATAC protocol with 15-minute DddSs treatment time up through the DddSs-I quenching step. Barcoded Tn5 treatment was then performed using the Scale Biosciences scATAC Pre-Indexing Kit, with 30-minute Tn5 treatment. All nuclei were combined and processed as specified in the 10X Chromium Next GEM Single Cell ATAC Kit v2 with the following changes from the standard protocol.50,500 nuclei were loaded into the GEM generation step. Barcoding enzyme mix was replaced with 2 uL Q5U Hotstart DNA polymerase (NEB) and 0.5 uL ET SSB (NEB) in step 2.1. Eight linear amplification cycles were used in step 2.5. In step 4.1, Single Index N Set A primer was replaced with Scale Bio is700 primer, and 2x NEBNext Q5U mastermix was used in place of Amp Mix. Nextgen sequencing was performed using Element AVITI at Quintara Biosciences. Table 1. CRISPR-Cas9 gRNA used in single-cell ACCESS-ATAC Gene 19-20 bp gRNA spacer SEQ ID Scale Bio Subpool target sequence without G NO: Tn5 well Barcode Attorney Docket No.29618-0483WO1 / BWH 2023-476 Processing of ACCESS-ATAC-seq data To process the edit-rich ACCESS-ATAC-seq datasets, we employed an iterative alignment approach using Bowtie224(v2.5.3) parameter. In the first iteration, reads were aligned with the "--very-sensitive" preset and default maximum mismatch penalty (“--mp”) of 6. Unmapped reads from each iteration were subsequently re- aligned in further iterations, reducing the mismatch penalty by 2 in each step. This process included a total of five iterations, and alignments from all iterations were merged into a single dataset. Deduplication of aligned reads was performed using czid-dedup (v0.1.2). Multi-mapped reads were then removed with a custom Python script, retaining only uniquely aligned reads from chromosomes 1-22 and X. ACCESS-ATAC-seq quality control metrics We established three quality control modules for ACCESS-ATAC-seq: transcription start site enrichment (TSS Enrichment), accessible chromatin log fold change (AC- LFC), and peak-trough log fold change (PT-LFC). For TSS Enrichment, we adapted the analysis script from the ENCODE ATAC-seq pipeline (github.com / ENCODE- DCC / atac-seq-pipeline / ). AC-LFC and PT-LFC, on the other hand, are based on the editing activity of DddSs. To calculate AC-LFC, we utilized an in-house Python script to bin the genome by accessibility scores and determine the fraction of edited bases in ACCESS-ATAC-seq reads mapped to each accessibility bin. Chromatin accessibility scores for HepG2 (ENCFF262URW), K562 (ENCFF600FDO) and HCT116 (ENCFF962GFP) cell lines were obtained from existing ATAC-seq experiments in the ENCODE Project. AC-LFC was defined as the log fold change between the edit fraction in low-accessibility regions (accessibility<0.1) and high-accessibility regions (accessibility>1). For PT-LFC, we employed another in-house Python script to filter reads mapped to CTCF motifs and their flanking regions. Composite edit fractions were calculated for each position relative to the motif center. Using the resulting composite edit fraction curve, we defined signature peak areas (-39 to -14 and 14 to 39 positions relative to the motif center) and trough areas (-164 to -122 and 122 to Attorney Docket No.29618-0483WO1 / BWH 2023-476 164 positions relative to the motif center). PT-LFC was then calculated as the log fold change between the edit fraction of the peak and trough regions. Modelling DddSs bias with deep learning We implemented a convolutional neural network (CNN) model with two convolutional and two fully connected layers to predict the observed DddSs editing counts based on DNA sequence. To train the model, we generated genome-wide single-nucleotide resolution signal from the E. coli data by counting the DddSs edit events. We split the E. coli genome as training (from 1 to 3641652) and validation (from 3641652 to 4641651) dataset to avoid overfitting. We binned the genome as 128 bp windows and encoded the DNA sequence with one-hot encoding. The model was trained with mean squared error (MSE) and Adam optimizer (learning_rate=0.003, weight_decay=0.0001) with 200 epochs. We reduced the learning rate when the validation error stopped decreasing for ten epochs and saved the model with the lowest validation error. We took the DNA sequence from hg38 as input and predicted the DddSs edit counts using the model described above, generating an expected ACCESS profile for human genome. Benchmarking dataset for transcription factor binding sites prediction We collected 1333 motifs derived from 409 ChIP-seq peaks in K562 from Factorbook28. For each of the motifs, we ran FIMO algorithm (--thresh 0.0001) to obtain its motif-based predicted binding sites (MPBSs) using the reference genome hg38. These MPBSs were used as our training data. To obtain the true labels for each motif, we intersected the MPBSs with the corresponding ChIP-seq peaks, and considered MPBSs supported by ChIP-seq peaks as true positives, and MPBS without ChIP-seq evidence as true negatives. Prediction of TF binding sites using ATAC-seq and ACCESS-ATAC-seq data To fairly compare the performance of predicting TF binding sites between ATAC-seq and ACCESS-ATAC-seq data, we sub-sampled the alignment BAM files using samtools (v1.15.1) to have the same number of reads (10M) between ACCESS- ATAC-seq (concurrent version) and ATAC-seq. Next, we used Genrich downloaded from github.com / jsh58 / Genrich (v0.6.1) to identify peaks (-j -d 150 -D -p 0.05) after Attorney Docket No.29618-0483WO1 / BWH 2023-476 sorting the BAM files by reads name. We removed the peaks that overlapped with hg38 blacklist regions. Next, we used bedtools (v2.31.1) to identify the shared peaks between ACCESS-ATAC-seq and ATAC-seq. For each motif, we only considered the binding sites within the shared peaks. We split the chromosomes into training (chr2, chr3, chr5, chr7, chr10, chr11, chr12, chr13, chr14, chr15, chr16, chr17, chr18, chr19, chr21, chr22, chrX), validation (chr6, chr8, chr20) and test (chr1, chr4, chr9). We only kept the motifs that have at least 5% true positives in training, validation and test chromosomes to make the evaluation robust. After filtering, we obtained a total of 1076 motifs. We trained five models for each motif using different input data (DNA- only, DNA+ATAC-only, DNA+ATAC, DNA+ACCESS and DNA+ACCESS+ATAC). To account for the bias in ATAC-seq, we downloaded pre- calculated Tn5 motif bias from zenodo.org / records / 7121027#.ZCbw4uzMI8N. For each model, we trained it using the binary cross entropy (BCE) loss by 100 epochs with a minibatch size of 48 using the Adam optimizer (learning_rate=0.003, weight_decay=0.0001). Models with the lowest validation error were saved and evaluated on the testing chromosomes. De novo TF footprint detection To detect TF-DNA interaction, we first smoothed the observed and predicted DddSs edit counts by averaging the single-nucleotide signal with 5bp window. Next, we computed a footprint score for each nucleotide by comparing its signal with the left and right flank regions as follows: ^^^^^^^^   +  ^^^^^^^^^^^  =   ^^ ^^^^^^^^^^^^^^A higher footprint nucleotide being bound by a TF. To account score using predicted DddSs edit counts, denoted as background footprint score. To quantify the significance, we fitted a normal distribution using the background footprint score and estimated the p-values which were further adjusted using Benjamini-Hochberg (BH) approach to account for multiple tests. We used a cutoff of 0.01 to select footprints. To compare the performance of ACCESS and ATAC, we also used Tn5 insertions to identify TF footprints as described above. For evaluation, we overlapped the detected footprints with ChIP-seq peaks from 409 TFs in K562 and computed the Jaccard Attorney Docket No.29618-0483WO1 / BWH 2023-476 Index. A higher value presents a better consistency between TF footprints and ChIP- seq peaks. Pre-processing Factorbook motifs We retrieved motif position weight matrices (PWMs) from Factorbook and used FIMO to predict motif-based predicted binding sites (MPBSs) for 717 and 1,333 motifs in the HepG2 and K562 cell lines, respectively. The motifs were filtered using a master list of DNA-binding transcription factors (TFs). For motifs with multiple associated PWMs, we selected the PWM with the lowest E-value, followed by the highest count of motif sites. For each DNA-binding TF, a motif site was designated as “bound” if it overlapped with any narrow peaks in the corresponding ChIP-seq experiment. We further refined the bound motif sites by requiring each site to have a FIMO p-value less than 0.0001, a ChIP-seq q-value (adjusted p-value) less than 0.05, and at least 50 ACCESS-ATAC-seq reads mapped to it. This filtering resulted in 119 TF motifs for the HepG2 cell line and 140 TF motifs for the K562 cell line. Occupancy Pattern Inference by Editing (OccuPIE) model overview OccuPIE is a convolutional neural network model designed to infer the occupancy state of TFs at single-allele resolution. OccuPIE employs a sequential architecture comprising three convolutional layers followed by two fully connected layers. It takes a 301 × 7 matrix as input to represent a 150-nt flanking region on either side of the motif site. The input matrix adopts a one-hot encoding format, with four channels indicating the presence of A, T, G, and C in the genomic sequence, two channels indicating C-to-T and G-to-A edits in the ACCESS-ATAC-seq read, and one channel representing the read coverage across the 301 positions. This design enables OccuPIE to leverage both the sequence and edit patterns to predict the probability of a TF- associated allele (read) being in one of three bounding states: bound, unbound- accessible, or unbound-inaccessible. Details on the model architecture and the input data structure are provided herein. TF-specific binding state assignment Binding states were assigned to reads mapped to ChIP-seq-bound motifs using predefined rules that compared edit fractions in the footprint and immediate Attorney Docket No.29618-0483WO1 / BWH 2023-476 neighboring peak regions against TF-specific thresholds. Bound reads were characterized by high editing activity in the peaks and low editing in the footprint, while unbound-accessible reads displayed high editing in both the peaks and the footprint. Unbound-inaccessible reads were defined by low editing in both regions. The footprint and peak regions were determined based on the composite edit patterns specific to each TF. To evaluate the assigned states, we calculated the Spearman correlation between the fraction of bound reads and the log-transformed normalized ChIP-seq scores for the motif sites. Normalized ChIP-seq scores were computed by scaling raw signals to a 0–1 range. Thresholds for the peaks and footprints were derived from the composite motif edit profile and optimized to maximize the correlation between bound read fractions and log-transformed normalized ChIP-seq scores. Bound motif sites on Chromosome 1 were designated for defining the ranges of the footprint and peaks, as well as for optimizing the edit fraction thresholds. TF motifs in HepG2 and K562 cell lines achieving a Spearman correlation (rho) greater than 0.1 with the assigned states were selected for modeling with OccuPIE. A detailed description of the range identification and state assignment processes is provided herein. OccuPIE model training and evaluation For each TF, bound motif sites from 18 chromosomes (chr3–chr11, chr13–chr21) were used for training, while sites from three chromosomes (chr2, chr12, chr22) were reserved for testing. Models were trained for up to 500 epochs with an adaptive learning rate and early stopping enabled. To assess the contribution of ACCESS- ATAC-seq data, a baseline model was trained using only the genomic sequence input channels. With the testing data, read states were predicted by both the full model and the baseline model, and their performances were compared based on weighted F1 scores and the area under the precision-recall curve (AUPRC). Additionally, model performance was similarly evaluated through correlation analysis with log-normalized ChIP-seq scores. Instead of using the fraction of bound reads to indicate the binding status of a TF, the mean bound probability of all reads mapped to the site was used to represent the bound probability at the motif level. Models achieving a Spearman correlation (rho) greater than 0.2 (p-value < 0.05) were considered to have good Attorney Docket No.29618-0483WO1 / BWH 2023-476 concordance with ChIP-seq evidence and were included in downstream single-allele analyses. Further details on training parameters can be found herein. Co-occupancy analysis and periodicity detection We systematically identified pairs of adjacent bound motif sites for any two transcription factors (including self-pairs) with motif centers located within 100nt of each other. Motif pairs with overlapping footprint ranges were excluded from the analysis, as OccuPIE cannot reliably distinguish spatial differences between overlapping sites. To ensure precise center disposition calculations, motif centers were adjusted by the mean position of the footprint range relative to the FIMO- predicted center. For a given adjacent motif pair between TFA and TFB with N shared reads, we used TF-specific OccuPIE models to predict the read-level bound probability for both TFs. For the i-th shared read, the bound probabilities of TFA and TFBcan be denoted as ^^^^^^^ | ^^^ and ^^^^^^^ | ^^^ respectively. The “observed” co- between the two TFs was calculated as: 1 ' ^^^^^ = ^^^ × ^^^^^^^ ^^^^^ However, co- if the motif- level bound probability of either TFs is high. This “expected” co-occupancy probability was calculated as: 1 ' ' ^*+^,^^" = ^ % ^^^^^^^ | 1^ ^^^ ^ × ^ % ^^^^^^^ | ^^^^ To account between !^"− ^^*+^,^^") to quantify the extent to which observed co-occupancy deviated from expected values. Motif pairs where either TF had an extreme motif-level bound probability below 0.01 or above 0.99 were excluded from further analysis. For each pair of TFs, observed – expected co-occupancy values were pooled, and a one-sample T-test was performed to assess whether the mean difference was significantly different from zero. The p-values from these tests were corrected for multiple Attorney Docket No.29618-0483WO1 / BWH 2023-476 comparisons using the Bonferroni method, with a significance threshold of alpha = 0.05. For each TF, we applied the fast Fourier transformation (FFT) algorithm to detect periodic patterns in the observed – expected co-occupancy values between the TF and all other TFs across center displacements ranging from -100nt to 100 nt. To enable comparisons of periodicity across TFs, observed – expected values were normalized by scaling the 5th percentile to 0 and the 95th percentile to 1. Center displacement values were binned to the nearest integer, and the median observed – expected value was calculated for each bin. This pre-processing step ensured that the co-occupancy values were evenly distributed, optimizing the data for FFT analysis. FFT transformed the median observed – expected values into a spectrum of component frequencies, each characterized by an amplitude. The significance of these frequencies was evaluated using a randomization test. Null distributions for each frequency were generated by shuffling the normalized observed – expected values and performing FFT across 2000 iterations. The p-value for each frequency, representing the likelihood of observing an amplitude equal to or greater than the actual amplitude under the null hypothesis, was calculated as the proportion of randomization iterations in which the randomized amplitude exceeded or matched the real amplitude. We restricted the period detection range to 8–13nt to exclude excessively high or low frequencies. The dominant period of a TF was defined as the period corresponding to the most significant frequency within the detection range. P-values across TFs were corrected for multiple testing using the Bonferroni method (alpha = 0.05). Periodicity strength was defined as the ratio of the dominant frequency amplitude to the 95th percentile of its null distribution. Single-cell ACCESS-ATAC computational analysis We used barcoded Tn5 to assign each sequenced read to the corresponding CRISPR- KO and WT experiment. For each experiment, we grouped the cells and used the iterative approach as described above to align the reads to reference genome hg38. We filtered the BAM files by only retaining the properly paired mapped reads. Next, we converted the BAM files to fragment files and removed the duplicated reads with same start and end positions. To control the data quality, we used the package ArchR58to filter low-quality cells for each experiment independently based on TSS Attorney Docket No.29618-0483WO1 / BWH 2023-476 enrichment (> 2 for WT; > 2.5 for sgATXL7L3 and sgJMJD1C; > 3 for sgZBTB7A and > 4 for rest of the KO experiments) and the number of unique fragments (> 1000 and < 100000), obtaining 26147 high-quality ACCESS-ATAC single cells. We used the functions addGroupCoverages to create sample-specific Tn5 insertion coverages and addReproduciblePeakSet (cutOff = 0.01) to identify peaks for each sample. All peaks were merged to create a union peak set (n = 72166), and a count matrix was constructed with the function addPeakMatrix. Next, we used the package Signac59to analyze this count matrix. Specifically, we used the functions RunTFIDF to normalize the count matrix and RunSVD to perform dimensionality reduction. For visualization, we used the function RunUMAP to generate a 2D embedding after excluding the first component which has been shown to be highly correlated with the sequencing depth of the cells. We clustered the cells using the functions FindNeighbors and FindClusters, identifying two populations (denoted as C1 and C2). We split the BAM file for each experiment based on the clusters and merged the files from all conditions to create a bulk BAM file for each cluster. To annotate these two clusters, we downloaded ATAC-seq data of HepG2 and K562 from ENCODE. We sub-sampled the BAM files to 100M reads and computed the genome-wide Spearman correlation based on Tn5 read coverage. Identifying motif features The composite edit fraction for bound motifs is smoothed using a 15nt rolling average to reduce noise and highlight meaningful patterns. Here we denote the relative position to motif center as ., the edit fraction and the smoothed edit fraction at position x as / 0^.^ and / 0^1^^^^^.^ respectively. We calculated two key edit fraction values: ^minimum footprint edit fraction: / 02_456 = min: / 0^.^; , . ∈ [−12, 12]^ minimum of the left peak and right peak maximum edit fractions : max: / 0^1^^^ ^.^;, . ∈ [−75, −14] / 0 _1^C = mi ^+_4AB n Dmax: / 0^1^^^^^.^;, . ∈ [14, 75]Next, _1^C^ Attorney Docket No.29618-0483WO1 / BWH 2023-476 The footprint positions ^^, the left peak positions ^L+and the right peak positions ^+are then identified as: ^^^: any position where J^.^ < 0.55 for . ∈ [−12,12]^ ^L+: any position where J^.^ > 0.67 for . ∈ [−100, min^^^^]^ ^ +: any position where J^.^ > 0.67 for . ∈ [max:^^;, 100]OccuPIE training data states assignment For a given ACCESS-ATAC-seq read mapped to a motif site, we calculate the mean edit fraction of the footprint positions as R^^^^+ ^C^, the mean edit fraction of both the left and right peak positions as R+^ST, as well as the mean edit fraction of all positions (within ±150nt from motif center) as R^L^^SL. Then the read’s state can be assigned given the corresponding thresholds (UℎJWX) for each value: ^Unbound-inaccessible, if R+^ST < UℎJWX+^ST and R^L^^SL < UℎJWX^L^^SL^ Bound, if read is not unbound-inaccessible and R^^^^+ ^C^ < UℎJWX^^^^+ ^C^^ Unbound-accessible, if read is neither unbound-inaccessible nor bound The peak and footprint thresholds are further determined by the corresponding weights (Y): UℎJWX+^ST = R+^ST ∙ Y+^ST + R^^^^+ ^C^ ∙ ^1 − Y+^ST^The fraction standard : UℎJWX^L^^SL = R[C^^[C" + 2\[C^^[C"Finally, the weights for and are determined for each different motif. For a Y^^^^+ ^C^values from -2 to 2 with a step of 0.1. In each iteration, we assign states to reads and calculate the sparman correlation between the resulting bound read fractions and the log transformed normalized ChIP-seq scores. The weight combination that achieve the highest correlation is used to assign states for the training data. OccuPIE input matrix data structure Each read mapped to a motif site, representing a single allele, is processed into a 301 × 7 multi-channel one-hot encoded input matrix. This matrix is centered on the motif Attorney Docket No.29618-0483WO1 / BWH 2023-476 site, extending 150 nucleotides upstream and downstream. The input matrix consists of the following 7 channels: ^ Base Channels (4): These channels encode the A, T, G, and C nucleotides in the reference sequence. ^ Edit Channels (2): These channels encode “C-to-T” and “G-to-A” edits observed in the ACCESS-ATAC-seq read. ^ Coverage Channel (1): This channel encodes the read coverage relative to the motif center. OccuPIE model architecture and training parameters The OccuPIE model is built on the tensorflow frameowork with a sequential architecture. The model is mainly comprised of three convolutional layers and two dense layers. The convolutional layers contain 16, 32 and 641d filters with a fixed filter size of 12. OccuPIE implementation Identifying motif features The composite edit fraction for bound motifs is smoothed using a 15nt rolling average to reduce noise and highlight meaningful patterns. Here we denote the relative position to motif center as ., the edit fraction and the smoothed edit fraction at position x as / 0^.^ and / 0^1^^^^^.^ respectively. We calculated two key edit fraction values: ^minimum footprint edit fraction: / 02_456 = min: / 0^.^; , . ∈ [−12, 12]^ minimum of the left peak and fractions : . / 0 ^C = ^^+_4AB_1 min Dmax: / 0^1^^^^^.^;, . ∈ [14, 75]Next, _1^C^The footprint positions ^^, the left peak positions ^L+and the right peak positions ^+are ^^^: any position where J^.^ < 0.55 for . ∈ [−12,12]^ ^L+: any position where J^.^ > 0.67 for . ∈ [−100, min^^^^]^ ^ +: any position where J^.^ > 0.67 for . ∈ [max:^^;, 100] Attorney Docket No.29618-0483WO1 / BWH 2023-476 OccuPIE training data states assignment For a given ACCESS-ATAC-seq read mapped to a motif site, we calculate the mean edit fraction of the footprint positions as R^^^^+ ^C^, the mean edit fraction of both the left and right peak positions as R+^ST, as well as the mean edit fraction of all positions (within ±150nt from motif center) as R^L^^SL. Then the read’s state can be assigned given the corresponding thresholds (UℎJWX) for each value: ^Unbound-inaccessible, if R+^ST < UℎJWX+^ST and R^L^^SL < UℎJWX^L^^SL^ Bound, if read is not unbound-inaccessible and R^^^^+ ^C^ < UℎJWX^^^^+ ^C^^ Unbound-accessible, if read is neither unbound-inaccessible nor bound The peak and footprint thresholds are further determined by the corresponding weights (Y): UℎJWX+^ST = R+^ST ∙ Y+^ST + R^^^^+ ^C^ ∙ ^1 − Y+^ST^The fraction standard deviation of unbound motifs (within ±150nt from motif center): UℎJWX^L^^SL = R[C^^[C" + 2\[C^^[C"Finally, the weights determined for each different motif. For a given motif, we iteratively try Y+^STand Y^^^^+ ^C^values from -2 to 2 with a step of 0.1. In each iteration, we assign states to reads and calculate the sparman correlation between the resulting bound read fractions and the log transformed normalized ChIP-seq scores. The weight combination that achieve the highest correlation is used to assign states for the training data. OccuPIE input matrix data structure Each read mapped to a motif site, representing a single allele, is processed into a 301 × 7 multi-channel one-hot encoded input matrix. This matrix is centered on the motif site, extending 150 nucleotides upstream and downstream. The input matrix consists of the following 7 channels: ^ Base Channels (4): These channels encode the A, T, G, and C nucleotides in the reference sequence. Attorney Docket No.29618-0483WO1 / BWH 2023-476 ^ Edit Channels (2): These channels encode “C-to-T” and “G-to-A” edits observed in the ACCESS-ATAC-seq read. ^ Coverage Channel (1): This channel encodes the read coverage relative to the motif center. OccuPIE model architecture and training parameters The OccuPIE model is built on the TensorFlow framework with a sequential architecture. The input layer accepts data with dimensions corresponding to the preprocessed one-hot encoded dataset. The model includes three convolutional layers, each using ReLU activation and L2 regularization (coefficient 0.01). Filters in the first layer span the input height with a width of 15, while subsequent layers use a height of 1. The number of filters starts at 16 and increases linearly with the layer index. Each convolutional layer is followed by max pooling (1 ^ 2) and dropout (rate 0.25) to reduce overfitting. The convolutional layers are followed by a flattening layer, which converts the output into a one-dimensional vector for the dense layers. Two dense layers progressively reduce the node count, starting with 32 nodes and halving with each layer. Both dense layers use ReLU activation, L2 regularization, and dropout (rate 0.25). The final output layer includes three nodes (one for each class), with softmax activation to output class probabilities. The model is compiled with the Adam optimizer (learning rate 0.0001), categorical crossentropy loss, and accuracy as the primary metric. The model was trained using a batch size of 4096 for a maximum of 500 epochs. Early stopping was employed to monitor validation loss and halt training after 12 epochs of no improvement, restoring the best weights. A learning rate scheduler reduced the learning rate by a factor of 0.5 when validation loss plateaued for three epochs, with a minimum learning rate of 1 ^ 10-6. Class weights were calculated to address class imbalance, ensuring proportional weighting based on the inverse frequency of each class in the training dataset. Input data was expanded by an additional dimension to ensure compatibility with the convolutional layers. Attorney Docket No.29618-0483WO1 / BWH 2023-476 Exemplary Detailed dsDNA deamination assay 1. Substrate preparation: a. Resuspend the following oligos at 100 uM in TE buffer. Ddd_substrate_fw_FAM / 56- FAM / AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAG ATCAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA (SEQ ID NO:31) Ddd_substrate_rv : TTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTGATCTTTTTTTTT TTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTT (SEQ ID NO:32) b. Mix the following (scale as necessary): 2.5 uL Ddd_substrate_fw_FAM 2.5 uL Ddd_substrate_rv 5 uL IDT Nuclease Free Duplex Buffer (cat# 11-05-01-12) c. Anneal oligos In PCR block, heat to 95 deg C for 5 min, then cool to 25 deg with cooling rate of 1 deg / min (set this manually). Then cool to 4 deg C for temporary storage. Move to -20 for longer term storage. Note that this will produce 25 uM dsDNA Ddd substrate. 2. Deamination assay a. Make 5x Deamination buffer (100 mM Tris-HCl pH 7.4, 1M NaCl). Store at room temperature. b. Mix the following to create a 10 uL deamination reaction (can be made as mastermix omitting Ddd enzyme and dH2O). We recommend testing Ddd enzymes in the range of 50 nM to 5 uM to determine concentration with ~50% activity. We have used DddA at 500 nM and DddSs at 250 nM based on titration results. Include a control reaction with no Ddd enzyme. 2 uL 5x Deamination Buffer 1 uL 10 mM DTT 0.4 uL 25 uM dsDNA substrate (see above) Attorney Docket No.29618-0483WO1 / BWH 2023-476 XX uL Ddd enzyme XX uL dH2O (to 10 uL final volumen) c. Incubate reaction in PCR machine at 37C for 1 hour d. Store on ice and move directly to step below. 3. Detection of deamination: a. Perform DpnII digest as follows: 1 uL 10X DpnII Buffer (NEB) 0.5 uL DpnII (NEB) 1.5 uL dH2O 3 uL Deamination reaction from above b. Incubate at 37C for 1hr (PCR machine is best) c. Run on D1000 Tapestation. Calculate fraction deaminated as follows: [Concentration of 90-nt product, nM] / [(Concentration of 45-nt product, nM / 2) + (Concentration of 90-nt product, nM) d. We have found that the concentration of Ddd enzyme producing ~50% deamination in this dsDNA deamination assay is optimal as final Ddd enzyme concentration for ACCESS. Exemplary Detailed ACCESS and ACCESS-ATAC Protocol: Stock solutions and buffers 10% octylphenoxypolyethoxyethanol (IGEPAL) in nuclease-free dH2O. 1% Digitonin in nuclease-free dH2O. 10% Tween-20 in nuclease-free dH2O RSB buffer: 10 mM Tris-HCl (pH 7.5, Invitrogen, cat. no.15567027), 10 mM NaCl (Invitrogen, cat. no. AM9759) and 3 mM MgCl2(Invitrogen, cat. no. AM9530G) in nuclease-free dH2O. Lysis buffer (a cell lysis buffer that maintains chromatin intact, from Zhang et al 2022 BMC Genomics; also known as ATAC lysis buffer): Make the lysis buffer by adding 0.1% IGEPAL (Sigma, cat. no. I3021, 1:100 from 10% IGEPAL stock solution), 0.01% digitonin (Invitrogen, cat. no. BN2006, 1:100 from 1% digitonin Attorney Docket No.29618-0483WO1 / BWH 2023-476 stock solution), and 0.1% Tween-20 (Bio-Rad, cat. no.1610781, 1:100 from 10% Tween-20 stock solution) to RSB. Detergent percentages reported are final concentrations. ATAC-seq buffer (20 mM Tris-HCl (pH 7.5, Invitrogen, cat. no.15567027), 10 mM MgCl2 (Invitrogen, cat. no. AM9530G), 20% Dimethyl Formamide in nuclease-free dH2O.) RSB + 0.1% Tween-20: Make RSB + 0.1% Tween-20 by adding 0.1% Tween- 20 (Bio-Rad, cat. no.1610781, 1:100 from 10% Tween-20 stock solution) to RSB. Make 1.2-2 mL per sample to be tested to have plenty. Detergent percentages reported are final concentrations. a. Prepare intact nuclei from cells for DddA testing. At least 105(100,000) cells were harvested. The cells were then trypsinized and spun down. The media was aspirated, and ice-cold PBS added. Pipetting up and down was used to mix the sample well. Live cell counting was performed, and the results used to pipet 105(100,000) cells into a centrifuge tube. Then, the cells were spun down at 500xg for 5 min at 4°C. The supernatant was carefully removed, then each cell pellet was resuspended in ice-cold lysis buffer. After the cell pellets were resuspended in the lysis buffer, the samples were incubated on ice for 3 min, and then lysis stopped by adding RSB + 0.1% Tween-20 to each tube. The nuclei were centrifuged at 500 r.c.f for 10 min at 4 °C, then the supernatant was removed without disturbing the pellet. b. Treat with DddA and UGI. A deamination master mix comprising: 50% 2x ATAC-seq buffer (20 mM Tris-HCl (pH 7.5, Invitrogen, cat. no. 15567027), 10 mM MgCl2 (Invitrogen, cat. no. AM9530G), 20% Dimethyl Formamide in nuclease-free dH2O.) 10% of 1% digitonin stock solution for 0.1% final concentration 10% of 10% Tween-20 stock solution for 1% final concentration 4% UGI (NEB M0281S / M0281L) Ddd enzyme (optimal final concentration for DddSs in our hands is 250 nM, for DddA is 500 nM) Attorney Docket No.29618-0483WO1 / BWH 2023-476 Remainder of volume is RSB buffer + 0.1% Tween: 10 mM Tris-HCl (pH 7.5, Invitrogen, cat. no.15567027), 10 mM NaCl (Invitrogen, cat. no. AM9759) and 3 mM MgCl2 (Invitrogen, cat. no. AM9530G) in nuclease-free dH2O. Add Tween-20 to 0.1% final concentration The mix was freshly prepared and mixed with the pelleted nuclei, then pipetted up and down to mix thoroughly but gently. DddA + H2O were added and pipetted up and down to mix. The best signal to noise ratio was achieved with between 0.5 uM to 2 uM DddA. The samples were then incubated at 37oC in a heated tube water bath. After 30 minutes, either: A 5x molar excess of purified DddI was added and the sample incubated for 10 min at 237oC. Then the ATAC-seq transposome complex was added and the sample incubated for 30 minutes at 37oC, then ATAC-seq library prep using Q5U polymerase was performed. -OR- 2 volumes DNA binding buffer was added and the DNA purified right away for whole genome or targeted locus editing assessment. The treated DNA was cleaned up using Zymo Genomic DNA clean and concentrator (D4010 or D4065) or Zymo DNA Clean and Concentrator-5 columns (Zymo, cat. no. D4004). Purified gDNA was stored at -20. Whole genome sequencing was performed using a Nextera kit or targeted locus sequencing, using Q5U polymerase to fix C^U edits as C^T. Exemplary ACCESS-WGS protocol 1. Prepare stock solutions and buffers. The following buffers can be made in advance and stored. RSB: 10 mM Tris-HCl (pH 7.5, Invitrogen, cat. no.15567027), 10 mM NaCl (Invitrogen, cat. no. AM9759) and 3 mM MgCl2(Invitrogen, cat. no. AM9530G) in nuclease-free dH2O. RSB can be made in bulk and stored at 4 °C long-term. 10% IGEPAL in nuclease-free dH2O (Sigma, cat. no. I3021). Mix well and store at 4 °C long-term. 1% Digitonin in nuclease-free dH2O (Invitrogen, cat. no. BN2006). Mix well and store at 4 °C long-term. Attorney Docket No.29618-0483WO1 / BWH 2023-476 10% Tween-20 in nuclease-free dH2O (Bio-Rad, cat. no.1610781). Mix well and store at 4 °C long-term. 2x ATAC-seq buffer: 20 mM Tris HCl (pH 7.5), 10 mM MgCl2, 20% Dimethyl Formamide. The following buffers should be made on the day of the experiment: ATAC lysis buffer: Add 0.1% IGEPAL (Sigma, cat. no. I3021, 1:100 from 10% IGEPAL stock solution), 0.01% digitonin (Invitrogen, cat. no. BN2006, 1:100 from 1% digitonin stock solution), and 0.1% Tween-20 (Bio-Rad, cat. no.1610781, 1:100 from 10% Tween-20 stock solution) to RSB. Make 150 uL per sample to be tested. Detergent percentages reported are final concentrations. RSB + 0.1% Tween-20: Add 0.1% Tween-20 (Bio-Rad, cat. no.1610781, 1:100 from 10% Tween-20 stock solution) to RSB. Make 1.2-2 mL per sample to be tested. Detergent percentages reported are final concentrations. 2. Prepare intact nuclei. ^ We have performed ACCESS-ATAC using 2.5*10^5 cells in 25 uL reaction volume or 5*10^5 cells in 50 uL reaction volume. It is likely that the protocol would be successful at a range of cell concentrations. ^ For non-adherent cells, spin and resuspend in 1 mL ice-cold PBS. Pipet up and down to mix well. Perform cell counting while keeping cells on ice. ^ For adherent cells, trypsinize, quench, spin, and resuspend in 1 mL ice- cold PBS. Pipet up and down to mix well. Perform cell counting while keeping cells on ice. ^ Using the live cell count, pipet desired cells into one centrifuge tube. Add PBS up to at least 1mL (if there is >100 uL based on what you transfer, that’s fine). ^ Spin down at 500xg for 5 min at 4 deg. ^ Remove supernatant very carefully. By pipetting, thoroughly resuspend each cell pellet in 125 uL ice-cold ATAC Lysis Buffer. ^ After resuspending cell pellets in the lysis buffer, incubate on ice for 3 min, and then stop lysis by adding 1.3 ml RSB + 0.1% Tween-20 to each tube. Attorney Docket No.29618-0483WO1 / BWH 2023-476 ^ Centrifuge nuclei at 500 r.c.f for 10 min at 4 °C. Remove supernatant and make sure no more than 20 uL are left over. Do not disturb pellet. ^ Prepare thermal block at 37 deg 3. Treat nuclei with Ddd enzyme. a. Prepare ACCESS treatment mix (can be made as mastermix for multiple samples). Shown for 25 uL reaction, can be scaled to 50 uL reaction volume: 1.5 uL 2x ATAC-seq buffer 1.5 uL 1% digitonin stock solution for 0.1% final concentration 1.5 uL 10% Tween-20 stock solution for 1% final concentration 1 uL UGI (NEB M0281S / M0281L) 0.5 uL Ddd enzyme titrated to provide ~50% deamination in dsDNA deamination assay b. Gently resuspend nuclei pellet with 19 uL ACCESS treatment mix. Use pipet to measure reaction volume. If <25 uL, add RSB + 0.1% Tween-20 to final volume of 25 uL. c. Incubate at 37 deg C for 30 min, flicking gently to mix every 10-15 minutes. d. After incubation, immediately clean up treated DNA using Zymo Genomic DNA clean and concentrator 25 kit (D4064 / D4065). All centrifugation steps should be performed between 10,000 - 16,000 x g. Pre-heat 50 uL of Zymo Elution Buffer per sample to 50-60 deg C. -In a 1.5 ml microcentrifuge tube, add 2 volumes of ChIP DNA Binding Buffer to each volume of DNA sample. In this case, add 50 uL DNA binding buffer. Mix briefly by vortexing. -Transfer mixture to a provided Zymo-Spin™ Column in a Collection Tube. -Centrifuge for 30 seconds. Discard the flow-through. -Add 400 μl DNA Wash Buffer to the column. Centrifuge for 30 seconds. Repeat this wash step. -Add 50 uL heated Zymo Elution Buffer (50-60 deg C) directly to the column matrix Attorney Docket No.29618-0483WO1 / BWH 2023-476 and incubate at room temperature for one minute. Transfer the column to a 1.5 ml microcentrifuge tube and centrifuge for 30 seconds to elute the DNA. 4. Tagment genomic DNA ^ Dilute 100 ng of ACCESS-treated genomic DNA to a total volume of 20 uL (5 ng / uL). ^ Add 20 μl of ACCESS-treated genomic DNA at 5 ng / μl (100 ng total) to a PCR tube. ^ Add 25 μl of 2x ATAC-seq Buffer. ^ Add 5 μl of mix of 0.5 uL Tagmentase (Tn5 transposase) - loaded (Diagenode C01070012-30 / C01070012-200) + 4.5 uL dH2O (0.5 uL loaded Tn5 + 4.5 uL dH2O) ^ Pipette up and down 10 times to mix. ^ Perform incubation in a PCR machine as follows: 55C for 5 minutes, hold at 10C ^ After incubation, clean up treated DNA using Zymo DNA clean and concentrator 5 kit (D4003 / D4004) using protocol below. All centrifugation steps should be performed between 10,000 - 16,000 x g. Pre-heat 50 uL of Zymo Elution Buffer per sample to 50-60 deg C. -In a 1.5 ml microcentrifuge tube, add 5 volumes of DNA Binding Buffer to each volume of DNA sample. In this case, add 125 uL DNA binding buffer. Mix briefly by vortexing. -Transfer mixture to a provided Zymo-Spin™ Column in a Collection Tube. -Centrifuge for 30 seconds. Discard the flow-through. -Add 200 μl DNA Wash Buffer to the column. Centrifuge for 30 seconds. Repeat this wash step. -Add 50 uL heated Zymo Elution Buffer (50-60 deg C) directly to the column matrix and incubate at room temperature for one minute. Transfer the column to a 1.5 ml microcentrifuge tube and centrifuge for 30 seconds to elute the DNA. -Eluted DNA can be stored on ice temporarily or at -20 for 24-72 hours prior to library preparation. 5. Prepare ACCESS-WGS libraries for NGS. Attorney Docket No.29618-0483WO1 / BWH 2023-476 a. Perform qPCR to determine appropriate cycle count for library prep. Per reaction volumes are: 0.375 uL 20 uM Nextera N7 primer (see below for sequences) 0.375 uL 20 uM Nextera N5 primer (see below for sequences) 7.5 uL 2x NEBNext Q5U Master Mix (NEB M0597L) 0.75 uL 20x EvaGreen Dye (Biotium #31000) 5.5 uL dH2O 0.5 uL purified ACCESS-treated DNA qPCR settings: 72°C for 5min 98°C for 30 seconds 25 cycles of: 98°C for 10 seconds, 63°C for 30 seconds, 72°C for 30 seconds. Hold at 10°C Determine appropriate cycles using qPCR. Typically, perform cycles equivalent to qPCR Ct or add 1-2 cycles. b. PCR Amplification of ACCESS-treated DNA. Per-reaction volumes are: 1.25 uL 20 uM Nextera N7 primer (see below for sequences) 1.25 uL 20 uM Nextera N5 primer (see below for sequences) 25 uL 2x NEBNext Q5U Master Mix (NEB M0597L) 22.5 uL purified ACCESS- treated DNA Thermal Cycler settings: 72°C for 5min 98°C for 30 seconds XX cycles (based on qPCR) of: 98°C for 10 seconds, 63°C for 30 seconds, 72°C for 30 seconds. Hold at 10°C c. SPRI clean up using a 1-sided procedure. Attorney Docket No.29618-0483WO1 / BWH 2023-476 i.To the PCR reaction, add 55ul of mixed, ROOM TEMP, SPRI beads to each sample (this is a 1.1X SPRI). ii.Vortex briefly and incubate for 5min at room temp iii.Apply magnet to collect beads iv.Once solution is clear, use pipette to remove supernatant v.While still on magnet, add 180ul of 80% ETOH to each sample without mixing vi.Incubate for 30 sec at room temperature vii.Remove supernatant with pipette viii.Repeat steps EtOH wash steps for a second ethanol wash ix.Allow tubes to sit at room temp so the residual ethanol can evaporate, beads will turn from shiny to matte when dry (2-5 min), then proceed x.With tubes off the magnet, add 20ul Elution Buffer (can be from Zymo or other company) xi.Cap tubes and mix by vortexing xii.Incubate samples for 5 min at room temp xiii.Apply magnet to samples xiv.Transfer supernatant to a fresh labeled tube. xv.Samples are now ready for Tapestation D1000 and NGS Exemplary ACCESS-ATAC protocols 1. Prepare stock solutions and buffers. The following buffers can be made in advance and stored. RSB: 10 mM Tris-HCl (pH 7.5, Invitrogen, cat. no.15567027), 10 mM NaCl (Invitrogen, cat. no. AM9759) and 3 mM MgCl2 (Invitrogen, cat. no. AM9530G) in nuclease-free dH2O. RSB can be made in bulk and stored at 4 °C long-term. 10% IGEPAL in nuclease-free dH2O (Sigma, cat. no. I3021). Mix well and store at 4 °C long-term. 1% Digitonin in nuclease-free dH2O (Invitrogen, cat. no. BN2006). Mix well and store at 4 °C long-term. 10% Tween-20 in nuclease-free dH2O (Bio-Rad, cat. no.1610781). Mix well and store at 4 °C long-term. Attorney Docket No.29618-0483WO1 / BWH 2023-476 2x ATAC-seq buffer: 20 mM Tris HCl (pH 7.5), 10 mM MgCl2, 20% Dimethyl Formamide. The following buffers should be made on the day of the experiment: ATAC lysis buffer: Add 0.1% IGEPAL (Sigma, cat. no. I3021, 1:100 from 10% IGEPAL stock solution), 0.01% digitonin (Invitrogen, cat. no. BN2006, 1:100 from 1% digitonin stock solution), and 0.1% Tween-20 (Bio-Rad, cat. no.1610781, 1:100 from 10% Tween-20 stock solution) to RSB. Make 150 uL per sample to be tested. Detergent percentages reported are final concentrations. RSB + 0.1% Tween-20: Add 0.1% Tween-20 (Bio-Rad, cat. no.1610781, 1:100 from 10% Tween-20 stock solution) to RSB. Make 1.2-2 mL per sample to be tested. Detergent percentages reported are final concentrations. 2. Prepare intact nuclei. ^ We have performed ACCESS-ATAC using 2.5*10^5 cells in 25 uL reaction volume or 5*10^5 cells in 50 uL reaction volume. It is likely that the protocol would be successful at a range of cell concentrations. ^ For non-adherent cells, spin and resuspend in 1 mL ice-cold PBS. Pipet up and down to mix well. Perform cell counting while keeping cells on ice. ^ For adherent cells, trypsinize, quench, spin, and resuspend in 1 mL ice-cold PBS. Pipet up and down to mix well. Perform cell counting while keeping cells on ice. ^ Using the live cell count, pipet desired cells into one centrifuge tube. Add PBS up to at least 1mL (if there is >100 uL based on what you transfer, that’s fine). ^ Spin down at 500xg for 5 min at 4 deg. ^ Remove supernatant very carefully. By pipetting, thoroughly resuspend each cell pellet in 125 uL ice-cold ATAC Lysis Buffer. ^ After resuspending cell pellets in the lysis buffer, incubate on ice for 3 min, and then stop lysis by adding 1.3 ml RSB + 0.1% Tween-20 to each tube. ^ Centrifuge nuclei at 500 r.c.f for 10 min at 4 °C. Remove supernatant and make sure no more than 20 uL are left over. Do not disturb pellet. ^ Prepare thermal block at 37 deg Attorney Docket No.29618-0483WO1 / BWH 2023-476 3. Treat nuclei with DddSs and loaded Tn5. Concurrent protocol (recommended; FIGs.1G and 1M): a. Prepare ACCESS-ATAC treatment mix (can be made as mastermix for multiple samples). Shown for 25 uL reaction, can be scaled to 50 uL reaction volume: 1.5 uL 2x ATAC-seq buffer 1.5 uL 1% digitonin stock solution for 0.1% final concentration 1.5 uL 10% Tween-20 stock solution for 1% final concentration 1 uL UGI (NEB M0281S / M0281L) 1.25 uL Tagmentase (Tn5 transposase) - loaded (Diagenode C01070012-30 / C01070012-200) 0.75 uL DddSs (8.33 uM stock for 250 nM final concentration) b. Gently resuspend nuclei pellet with 20.5 uL ACCESS-ATAC treatment mix. Use pipet to measure reaction volume. If <25 uL, add RSB + 0.1% Tween-20 to final volume of 25 uL. c. Incubate at 37 deg C for 30 min, flicking gently to mix every 10-15 minutes. d. After incubation, immediately clean up treated DNA using Zymo DNA clean and concentrator 5 kit (D4003 / D4004) using protocol below. All centrifugation steps should be performed between 10,000 - 16,000 x g. Pre-heat 50 uL of Zymo Elution Buffer per sample to 50-60 deg C. 1. In a 1.5 ml microcentrifuge tube, add 5 volumes of DNA Binding Buffer to each volume of DNA sample. In this case, add 125 uL DNA binding buffer. Mix briefly by vortexing. 2. Transfer mixture to a provided Zymo-Spin™ Column in a Collection Tube. 3. Centrifuge for 30 seconds. Discard the flow-through. 4. Add 200 μl DNA Wash Buffer to the column. Centrifuge for 30 seconds. Repeat this wash step. 5. Add 50 uL heated Zymo Elution Buffer (50-60 deg C) directly to the column matrix and incubate at room temperature for one minute. Transfer the column to a 1.5 ml microcentrifuge tube and centrifuge for 30 seconds to elute the DNA. Sequential protocol (less recommended): Attorney Docket No.29618-0483WO1 / BWH 2023-476 a. Prepare ACCESS treatment mix (can be made as mastermix for multiple samples). Shown for 25 uL reaction, can be scaled to 50 uL reaction volume: 12.5 uL 2x ATAC-seq buffer 2.5 uL 1% digitonin stock solution for 0.1% final concentration 1.5 uL 10% Tween-20 stock solution for 1% final concentration 1 uL UGI (NEB M0281S / M0281L) 0.5 uL DddSs (11.63 uM stock for 250 nM final concentration in 23.25 uL final volume) b. Gently resuspend nuclei pellet with 19 uL ACCESS treatment mix. Use pipet to measure reaction volume. If <23.25 uL, add RSB + 0.1% Tween-20 to final volume of 23.25 uL. c. Incubate at 37 deg C for 15-30 min (see manuscript for effect of treatment time), flicking gently to mix every 10-15 minutes. d. Add 5-10X molar excess of DddSs-I in 0.5 uL. Pipet up and down gently to mix. For example, add 0.5 uL of 59.4 uM DddSs-I for 1.25 uM final concentration in 23.75 uL final volume. e. Incubate for 2 minutes at 37 deg C. f. Add 1.25 uL Tagmentase (Tn5 transposase) - loaded (Diagenode C01070012- 30 / C01070012-200). g. Incubate at 37 deg C for 30 min, flicking gently to mix every 10-15 minutes. h. After incubation, immediately clean up treated DNA using Zymo DNA clean and concentrator 5 kit (D4003 / D4004) using protocol below. All centrifugation steps should be performed between 10,000 - 16,000 x g. Pre-heat 50 uL of Zymo Elution Buffer per sample to 50-60 deg C. -In a 1.5 ml microcentrifuge tube, add 5 volumes of DNA Binding Buffer to each volume of DNA sample. In this case, add 125 uL DNA binding buffer. Mix briefly by vortexing. -Transfer mixture to a provided Zymo-Spin™ Column in a Collection Tube. -Centrifuge for 30 seconds. Discard the flow-through. -Add 200 μl DNA Wash Buffer to the column. Centrifuge for 30 seconds. Repeat this wash step. -Add 50 uL heated Zymo Elution Buffer (50-60 deg C) directly to the column matrix Attorney Docket No.29618-0483WO1 / BWH 2023-476 and incubate at room temperature for one minute. Transfer the column to a 1.5 ml microcentrifuge tube and centrifuge for 30 seconds to elute the DNA. -Eluted DNA can be stored on ice temporarily or at -20 for 24-72 hours prior to library preparation. 4. Prepare ACCESS-ATAC libraries for NGS. a. Perform qPCR to determine appropriate cycle count for library prep. Per reaction volumes are: 0.375 uL 20 uM Nextera N7 primer (see below for sequences) 0.375 uL 20 uM Nextera N5 primer (see below for sequences) 7.5 uL 2x NEBNext Q5U Master Mix (NEB M0597L) 0.75 uL 20x EvaGreen Dye (Biotium #31000) 5.5 uL dH2O 0.5 uL purified ACCESS-ATAC-treated DNA qPCR settings: 72°C for 5min 98°C for 30 seconds 25 cycles of: 98°C for 10 seconds, 63°C for 30 seconds, 72°C for 30 seconds. Hold at 10°C Determine appropriate cycles using qPCR. Typically, perform cycles equivalent to qPCR Ct or add 1-2 cycles. b. PCR Amplification of ACCESS-ATAC-treated DNA. Per-reaction volumes are: 1.25 uL 20 uM Nextera N7 primer (see below for sequences) 1.25 uL 20 uM Nextera N5 primer (see below for sequences) 25 uL 2x NEBNext Q5U Master Mix (NEB M0597L) 22.5 uL purified ACCESS-ATAC-treated DNA Thermal Cycler settings: 72°C for 5min Attorney Docket No.29618-0483WO1 / BWH 2023-476 98°C for 30 seconds XX cycles (based on qPCR) of: 98°C for 10 seconds, 63°C for 30 seconds, 72°C for 30 seconds. Hold at 10°C c. SPRI clean up using a 1-sided procedure. i. To the PCR reaction, add 55ul of mixed, ROOM TEMP, SPRI beads to each sample (this is a 1.1X SPRI). ii. Vortex briefly and incubate for 5min at room temp iii. Apply magnet to collect beads iv. Once solution is clear, use pipette to remove supernatant v. While still on magnet, add 180ul of 80% ETOH to each sample without mixing vi. Incubate for 30 sec at room temperature vii. Remove supernatant with pipette viii. Repeat steps EtOH wash steps for a second ethanol wash ix. Allow tubes to sit at room temp so the residual ethanol can evaporate, beads will turn from shiny to matte when dry (2-5 min), then proceed x. With tubes off the magnet, add 20ul Elution Buffer (can be from Zymo or other company) xi. Cap tubes and mix by vortexing xii. Incubate samples for 5 min at room temp xiii. Apply magnet to samples xiv. Transfer supernatant to a fresh labeled tube. xv. Samples are now ready for Tapestation D1000 and NGS Sequences used in PCR amplification Nextera N7 rimers Se SEQ ID NO: p quences CAAGCAGAAGACGGCATACGAGAT 33 Attorney Docket No.29618-0483WO1 / BWH 2023-476 Example 1. Ddd enzymes induce chromatin accessibility-dependent cytosine deamination Current approaches to measure chromatin accessibility have major limitations. The most common, cleavage-based techniques, DNase-seq and ATAC-seq, have inherently limited resolution because their interpretation relies on aligning a >50-nt sequence between two cleavage events. This prevents allelic level analyses such as TF co-binding measurement and prevents high-resolution single-cell assessment of TF binding and nucleosome positioning. Non-cleavage-based approaches (Fiber-seq, NOMe-seq) cannot enrich for accessible chromatin and employ low-depth long-read sequencing, precluding high-resolution genome-wide analysis. We developed methods, including the exemplary Accessible Chromatin by Cytosine Editing Site Sequencing (ACCESS) assay, in which intact nuclei are treated with purified Ddd enzymes to catalyze chromatin accessibility-dependent C^U Attorney Docket No.29618-0483WO1 / BWH 2023-476 editing genome-wide. ACCESS enables basepair-resolution single-allele stenciling of nucleosome positioning and TF binding. We combined ACCESS with ATAC-seq (ACCESS-ATAC) to increase the resolution and applications of single-cell chromatin accessibility measurement. The first double-stranded DNA cytosine deaminase, DddA, was discovered several years ago6. This enzyme natively acts as an inter-bacterial toxin and has been engineered into a tool for mitochondrial genome editing6. Since then, DddA has been engineered for altered sequence preference and activity, and many Ddd orthologs have been identified7–10. No reports to our knowledge have examined whether treatment of chromatin with Ddd enzymes elicits enriched editing in accessible regions. To assess chromatin accessibility dependence of Ddd editing, we produced purified DddA. Ddd enzymes occur in operons with Ddd inhibitors (DddI), which prevent toxic DNA damage to host bacteria. We co-expressed DddA and DddA-I in bacteria and purified both proteins by adapting a published protocol6. We confirmed and titrated DddA activity and DddA-I inhibition of this activity in an in vitro dsDNA deamination assay in which cytosine deamination to uracil prevents subsequent cleavage by the restriction enzyme DpnII (FIG.1A). We next titrated DddA editing activity in intact human HCT116 colorectal carcinoma nuclei at five known highly accessible loci (active gene promoters including the accessible PTMA promoter) and three known inaccessible loci (universally inaccessible “safe sites”11), incubating nuclei for one hour with DddA. Loci were then PCR-amplified using Q5U polymerase, which converts uracils to thymines, and editing efficiency was measured as C^T conversion fraction in NGS. At all doses tested, all five accessible loci had substantially higher editing than all three inaccessible loci at known DddA target TC motifs (up to10-15-fold); the PTMA promoter had substantially higher editing than the inaccessible locus, primarily at known DddA target TC motifs (FIG.1B). Editing was dose-dependent and robust, ranging from an average of 30-40% at optimal doses. We next performed whole genome sequencing after DddA treatment of HCT116 nuclei (ACCESS-WGS), finding that TC^TT edit fraction steadily increased in concert with published HCT116 ATAC-seq chromatin accessibility (FIG.1C). Altogether, these data indicate that DddA provides chromatin accessibility-dependent cytosine deamination. Attorney Docket No.29618-0483WO1 / BWH 2023-476 DddA’s preferential editing at TC dinucleotides limits its chromatin accessibility measurement resolution. We evaluated six Ddd enzymes reported to have low motif specificity8,9, including DddSs / Ddd1, CseDa01, FlDa01, CrDa01, MGYPDa01, and a highly homologous, metagenomically identified deaminase MGYPDa829, through in vitro transcription / translation of each enzyme followed by ACCESS-WGS. All six showed chromatin accessibility-dependent editing at certain trinucleotides; however, two highly homologous enzymes, the deaminases from Simiaoa Sunii, which we refer to as DddSs (also called Ddd18, LbsDa0124, and SsDddA25) and MGYPDa829, showed the most consistently enriched editing across sequence contexts (FIG.1D). We characterized DddSs in further detail, showing it to have a relatively permissive motif with considerably broader editing than DddA genome-wide (FIG.1E-F). All further ACCESS experiments employed DddSs, produced through bacterial expression. One observation we made is that, while DddSs editing increases with chromatin accessibility, read alignment to the genome using the standard Bowtie212parameters decreases with chromatin accessibility. We reasoned that high fractions of DddSs editing might prevent proper alignment. We found that re-aligning unaligned reads with iteratively lower mismatch penalties recovered reads from accessible regions of the genome while not substantially increasing alignment in DddSs- untreated samples. In all further ACCESS experiments described below, we use this iterative Bowtie2 alignment approach. Example 2. High-resolution profiling accessible chromatin with ACCESS- ATAC It is difficult to obtain the sequencing coverage needed for high-resolution dissection of TF binding using ACCESS-WGS since DNA fragmentation does not preferentially enrich for reads in accessible chromatin. And while DddSs editing increases steadily with chromatin accessibility genome-wide, the sequencing coverage needed for high-resolution dissection of TF binding and chromatin state across the entire genome is prohibitively expensive. We thus sought to combine ACCESS with ATAC-seq, using ATAC to focus sequencing depth on accessible chromatin and ACCESS to improve resolution within each sequenced allele. We compared two approaches to combine ACCESS with ATAC-seq (ACCESS-ATAC) in K562 cells, a cell line in which TF binding sites have been characterized extensively by ChIP-seq6. Attorney Docket No.29618-0483WO1 / BWH 2023-476 We tested 30-minute concurrent treatment of intact chromatin with Tn5 and DddSs (concurrent ACCESS-ATAC, Fig.1G, 1M), as well as a three-step sequential protocol treating intact nuclei with DddSs for variable times (15-, 30-, and 60-minute, FIG.1L), quenching DddSs activity with DddSs-I, and then performing 30-minute Tn5 treatment (sequential ACCESS-ATAC). Because dsDNA adapters transposed by Tn5 (ATAC) would be editing targets of DddSs, the three-step protocol treated intact nuclei with DddSs for one hour, quench DddSs activity with purified DddSs-I for 10 minutes, and then performed standard ATAC-seq (30 minute Tn5 transposition followed by DNA purification and amplification, using Q5U as the polymerase to convert U^T while amplifying edited, accessible fragments) (FIG.1L). We also compared the resulting data with 30-minute Tn5 treatment alone as an important baseline against the standard ATAC-seq protocol. We performed ACCESS-ATAC in two biological replicates each in HepG2 and K562 cells, as these lines have the most extensive ENCODE ChIP-seq comparison data13, obtaining ~250M 150+150-nt paired-end NGS reads per sample, as longer reads maximize allelic resolution of ACCESS. All approaches yielded NGS libraries with similar fragment diversity as measured by qPCR, suggesting that DddSs treatment does not substantially alter Tn5 integration or library amplification when using Q5U polymerase. We found that reads were strongly enriched in accessible chromatin, showing that ATAC performance was not strongly affected by prior DddSs treatment. We developed a metric of ACCESS data quality based on DddSs editing fraction surrounding CTCF ChIP-seq binding sites since CTCF binding profoundly alters chromatin accessibility and binds to a subset of common loci across cell types.26As expected, we found that ACCESS- ATAC data exhibited high editing in peak regions surrounding CTCF binding sites, low editing at the footprint where CTCF binding occludes editing, and a wave-like histone coordination pattern radiating distally from binding sites. We thus determined the signal strength of ACCESS data by measuring the DddSs edit fraction (C to T and G to A) in the peak regions and the signal-to-noise as the log-fold change in editing between the peak and valley regions. We observed that 30-minute concurrent and sequential ACCESS-ATAC yielded equivalent ACCESS signal strength (CTCF peak edit fraction) and signal-to-noise (peak / valley LFC), while shortening the DddSs treatment time below 15 minutes decreased signal strength and signal-to-noise. Attorney Docket No.29618-0483WO1 / BWH 2023-476 Next, we compared read enrichment surrounding transcriptional start sites (TSS), a commonly used metric for measuring ATAC-seq data quality. We found that concurrent ACCESS-ATAC treatment yielded a slightly higher TSS enrichment than ATAC-seq (13.7x and 12.5x, respectively), while sequential ACCESS-ATAC samples showed substantially decreased TSS enrichment (Fig.1H). Decreasing DddSs treatment time during sequential ACCESS-ATAC improved TSS enrichment, although not nearly to the level seen in the concurrent protocol (2.3x for 60-minute, 3.9x for 30-minute, and 7.6 for 15-minute). We concluded that, while concurrent and sequential protocols provide equivalent quality ACCESS data, concurrent ACCESS- ATAC best preserved the accessible chromatin enrichment of ATAC-seq. We next performed a deeper comparison of K562 concurrent ACCESS-ATAC and ATAC-seq data. We computed the genome-wide correlation between ACCESS- ATAC and ATAC-seq based on read coverage and observed a high Spearman correlation (r = 0.75; Fig.1I). Moreover, the global architecture of accessible peaks is conserved (Fig.1J), suggesting that ACCESS-ATAC does not substantially alter the ATAC signal. At single-nucleotide resolution, DddSs editing events show denser signal than Tn5 insertion frequencies while showing clear visual evidence of TF footprints (Fig.1J). To examine TF footprints more systematically, we produced metaplots centered on ChIP-seq binding events for six TFs. DddSs and Tn5 induce footprints with equivalent foldchange in signal between the peak and trough (Fig. 1K), although there are ~5x more DddSs editing events than there are Tn5 integrations at a given read depth. We observed that DddSs editing footprints are consistently less than half as wide as Tn5 integration footprints (Fig.1K), which is presumably due to the considerably smaller size of DddSs (14 kDa) as compared to the Tn5 dimer (106 kDa). Example 3. Transcription factor binding prediction using ACCESS-ATAC We next assessed whether using the ACCESS signal could improve TF binding site prediction. Accounting for Tn5 motif bias has been shown to be essential for accurate ATAC-seq footprinting analysis7-8. To provide a similar correction for DddSs motif bias, we performed DddSs editing followed by whole genome sequencing of purified non-methylated E. coli genomic DNA. We obtained >1,000X average coverage across the E. coli genome and editing at ~30% of all cytosines. We found that editing bias was relatively minor and primarily confined to the nucleotides Attorney Docket No.29618-0483WO1 / BWH 2023-476 surrounding the target cytosine (Fig.1K). To enable DddSs motif bias correction in ACCESS-ATAC experiments, we trained a convolutional neural network (CNN) model to predict the observed edit counts based on one-hot encoded E. coli DNA sequences (Fig.2A). Evaluation on held-out chromatin regions indicated that our model accurately predicted the edit counts for cytosines across the genome (Spearman r = 0.94) and recapitulated DddSs editing profiles at single-nucleotide resolution (Fig. 2B-C). We next assessed if this model, trained on E. coli DNA, could be transferred to the human genome to account for DddSs bias. We applied the model to predicted DddSs editing counts using the hg38 reference genome and generated the predicted, observed and bias-corrected signal around TF binding site centers (±50bp) identified by ChIP-seq. We found that for many TFs, the observed edit counts around their sequence motifs were strongly influenced by DddSs enzyme bias and a clearer footprint pattern was observed after correcting for bias (Fig.2D). These results demonstrate that our CNN model effectively captures DddSs motif bias and enables single-nucleotide resolution bias correction. To determine whether ACCESS-ATAC improves TF binding site prediction, we trained deep learning models using a CNN model similar to the maxATAC approach27, which uses DNA sequences and chromatin accessibility signals to predict bound motif instances. To create benchmarking resources, we collected 409 ENCODE TF ChIP-seq datasets in K562 along with 1076 MEME motifs derived from FactorBook28. For each motif, we applied the FIMO algorithm29to detect motif- predicted binding sites (MPBSs) based on Position Weight Matrices (PWMs) and overlapped them with corresponding ChIP-seq peaks to obtain the most likely binding sites within peaks to use as true labels. We noticed that, for most TFs, the labels were substantially imbalanced. The fraction of true negatives (MPBS in ACCESS-ATAC peaks but not overlapping ChIP-seq peaks) significantly outnumbered the fraction of true positives, indicating that correctly classifying these labels may pose challenges for predictive models. Next, we trained a CNN-based classifier to predict motif- dependent binding using different input features: i) DNA sequences alone (DNA); ii) DNA plus Tn5 integrations normalized for motif bias from ATAC-seq (ATAC-seq) or ACCESS-ATAC (ACCESS-ATAC (Tn5)) data; iii) DNA plus ACCESS-ATAC DddSs edit events normalized for motif bias (ACCESS-ATAC (DddSs)); or iv) a combination of DNA, DddSs, and Tn5 signal from ACCESS-ATAC (ACCESS- Attorney Docket No.29618-0483WO1 / BWH 2023-476 ATAC) (Fig.2E). The results were evaluated on three held-out chromosomes based on the area under the precision-recall curve (auPRC). We found that including DddSs edit signals significantly improved the accuracy of TF binding site prediction (median = 0.398 for ACCESS-ATAC, 0.397 for ACCESS-ATAC (DddSs)), compared with Tn5 (median = 0.36 for ATAC-seq, 0.35 for ACCESS-ATAC (Tn5)) and DNA sequences (median = 0.34; Fig.2F-G). One constraint of this supervised learning approach is that it relies on the existence of corresponding ChIP-seq data from the same context for training models, limiting its applicability in cases where such data is not available. In contrast, TF footprinting algorithms, such as HINT-ATAC30and PRINT31, search for footprint- like Tn5 transposase cleavage patterns in ATAC-seq data and subsequently assign TFs to footprints through motif matching. Since we observed that DddSs edit events produce denser signal than Tn5 at single-nucleotide resolution (Fig.1G), we asked if this could also benefit TF footprinting analysis. We implemented a statistical approach, similar to PRINT, to identify genomic regions with significant depletion of DddSs editing counts compared to the background signal (Fig.2H). For a comprehensive comparison, we also applied our method to detect TF footprints using Tn5 insertion patterns derived from conventional ATAC-seq and ACCESS-ATAC data. To quantitatively evaluate the performance of these results, we overlapped the identified TF footprints with peaks from 409 K562 ChIP-seq datasets and computed the Jaccard index to assess their consistency. Notably, our results revealed that TF footprints identified using DddSs editing exhibited significantly higher concordance with the ChIP-seq peaks than those derived from Tn5-based footprinting (Fig.2I). These results underscore the potential of DddSs editing to enhance TF footprinting analysis by providing higher-resolution and more accurate representations of TF binding sites. Example 4. Allele-resolved TF occupancy and co-occupancy imputation from ACCESS-ATAC data ACCESS-ATAC provides a stencil of bound TFs at a single-allele level, similar to SMF32and Fiber-seq4. Because, in contrast to these other methods, ACCESS-ATAC is compatible with short-read sequencing and accessible chromatin enrichment, we posited that ACCESS-ATAC could permit high-coverage single- molecule profiling of TF occupancy across the accessible human genome. We used Attorney Docket No.29618-0483WO1 / BWH 2023-476 Ultima Genomics sequencing33to generate high-coverage ACCESS-ATAC datasets in HepG2 and K562 cells to leverage the wealth of TF binding data in these cell lines34,35. After alignment and deduplication, we retained 330 million HepG2 and 612 million K562 reads. We developed the Occupancy Pattern Inference by Editing (OccuPIE) model to impute TF occupancy status at individual alleles from ACCESS-ATAC data using similar principles to a prior approach based on SMF data32(Fig.3A). OccuPIE assumes that a candidate TF binding site can exist in three states: actively occupied by the TF (bound); unoccupied but with accessible chromatin; and unoccupied with inaccessible chromatin. Using 98 HepG2 and 106 K562 ChIP-seq datasets, we used the composite editing patterns surrounding ChIP-seq-bound motif instances to define a TF-specific peak segmentation with three nucleotide ranges, one for the central footprint and two for the flanking regions. We then created a training set by assigning ACCESS-ATAC reads that spanned a ChIP-seq-bound motif into one of the three states through defining thresholds of editing in the footprint and peak regions. Bound reads were defined as having high editing in flanking regions and low footprint editing, unbound, accessible reads as having high editing in flanks and footprint, and unbound, inaccessible reads as having low editing in flanks and footprint. We used these assignments to train a deep learning model to output the occupancy probability for every motif-overlapping allele in our ACCESS-ATAC datasets. To assess whether the output of OccuPIE accurately reflects TF binding propensity, we compared the median imputed allelic occupancy probability at a given motif instance with the abundance of ChIP-seq reads at that locus (ChIP-seq score). For 87 of the HepG2 and 85 of the K562 TFs, there was a significant association between OccuPIE-imputed allelic occupancy and ChIP-seq score across bound loci (Bonferroni-corrected p-value <0.01, Fig.3B). This association was strongest for TFs with wide, robust footprints such as STAT5A and USF1 (Fig.3B-D), indicating that TFs that strongly impede editing when bound allow for the most robust occupancy imputation. We found that TFs show a wide dynamic range in their allelic occupancy. For example, HepG2 USF1 loci ranged from 1-97% occupancy, with a median of 36% of bound alleles (Fig.3D). OccuPIE-imputed occupancy showed consistently improved association with ChIP-seq score as compared with the predefined training labels (median 19% increased Spearman correlation), validating its utility. Attorney Docket No.29618-0483WO1 / BWH 2023-476 We next used OccuPIE to assess co-occupancy patterns among 64 TFs in HepG2 and 64 TFs in K562 with robustly imputed allelic occupancy (Spearman r > 0.2 with ChIP-seq score). We examined ChIP-seq-bound motif pairs within 100 nt, removing motif pairs for which the central footprints for the two TFs have overlapping nucleotides, as OccuPIE cannot independently assess occupancy for such pairs. Altogether, we assessed co-occupancy at a total of 658,230 adjacent motif pairs in HepG2 and 1,105,054 in K562. To control for factors such as motif strength that might cause a TF to bind more or less frequently at a given motif instance, we calculated an “expected” co-occupancy for every adjacent motif pair under the baseline assumption that binding occurs independently (Mean P(boundTF1) * Mean P(boundTF2)). For each motif pair, we then compared the observed co-occupancy for all ACCESS-ATAC alleles overlapping both motifs with this baseline value to determine whether the two TFs co-occupy more often than expected (observed – expected co-occupancy, Fig.4A). Our analysis supports the conclusion that TFs tend to co-occupy adjacent binding sites more frequently than would be expected by chance. We observe significantly enriched co-occupancy for >85% of TF pairs in HepG2 and K562 (Bonferroni-corrected p < 0.01, Fig.4B). TFs of the bHLH-LZ class such as USF1, MITF, and MAX show enriched co-occupancy with the majority of other TFs and especially strong co-occupancy with each other (Fig.4C). While these factors are known to homo- and heterodimerize36, the inter-motif distances explored in this analysis are larger than can easily be explained by dimerization. TFs that function as repressors such as REST, MAFF, and NR2C2 show the lowest co-occupancy enrichment. Of all analyzed adjacently bound TF pairs in K562, homotypic STAT5A- STAT5A pairs showed the most robustly enriched co-occupancy (Fig.4C). This enrichment was primarily driven by over 500 loci with ChIP-seq-bound STAT5A motif pairs for which the motif centers are separated by 20-25 nt. While these loci exhibited a wide range of STAT5A occupancy levels (median allelic occupancy probabilities from 10-75%), they showed uniformly highly enriched allelic co- occupancy (Fig.4D). STAT5A binds as a homodimer to a palindromic TTCNNNGAA consensus, and Stat factors including STAT5A are known to tetramerize upon persistent activating stimuli, enabling binding to lower affinity Attorney Docket No.29618-0483WO1 / BWH 2023-476 sites37–39. K562 cells possess BCR / ABL translocation, a known driver of Stat tetramerization40. Our allele-resolved analysis shows that, at loci with properly positioned motif pairs, activated STAT5A in K562 cells interacts with DNA predominantly as a tetrameric unit in all-or-nothing fashion. We further assessed whether co-occupancy is dependent on the distance between the two bound motifs. To maximize power for this analysis, we assessed spatially resolved co-occupancy enrichment for each TF when paired with any adjacent TF. Intriguingly, we found that, in addition to a steadily decreasing propensity for co-occupancy at increasing distances, certain TFs show enriched co- occupancy at periodic intervals (Fig.4E). Fast Fourier transform (FFT) analysis revealed that co-occupancy enrichment shows significant periodicity for 39 HepG2 and 32 K562 TFs in range of the 10.5 bp periodicity of the DNA helical turn (Fig.4F). TFs with periodic co-occupancy span multiple families1, suggesting that this property is not dependent on specific DNA-binding or protein-protein interaction domains. Periodic co-occupancy was often evident throughout the entire + / - 100-nt evaluated range, suggesting that this trend persists across a large number of helical turns of the DNA. We were unable to assess periodicity between specific TF pairs due to data sparsity. Altogether, our analysis reveals that many TFs show periodic co- occupancy on individual DNA molecules. Example 5. Single-cell ACCESS-ATAC Because the resolution of single cell ATAC-seq (scATAC-seq) is inherently limited, we next asked whether we could combine ACCESS with scATAC-seq (scACCESS-ATAC) to generate high-resolution chromatin accessibility profiles at the single-cell level (Fig.5A). We performed a proof-of-principle experiment in which we performed arrayed CRISPR-KO of 11 chromatin-associated proteins as well as a non-targeting control gRNA in HepG2 and K562 cells. We also collected two additional control datasets by performing scATAC and scACCESS-ATAC on these two cell lines without gRNA treatment. We adapted the 10X Chromium GEM Single Cell ATAC v2 protocol, adding DddSs treatment and DddSs-I quenching prior to treatment with barcoded Tn5 (sequential ACCESS-ATAC) and replacing the polymerase mix in the two PCR steps (in-droplet and after purification) with Q5U. The cells were assigned to the CRISPR-KO treatment they received based on the combination of cellular and Tn5 barcode. Using super-loading, we recovered a total of Attorney Docket No.29618-0483WO1 / BWH 2023-476 30,457 cells with an average of 2,387 unique fragments and TSS enrichment score of 8.3 per cell after controlling the data quality. These results indicate that DddSs treatment is compatible with the 10x scATAC-seq protocol. We next assessed whether HepG2 and K562 cells could be distinguished using scACCESS-ATAC data. To do this, we aggregated the cells from all experiments and identified 94,366 peaks, based on which a cell-by-peak count matrix was constructed using the ATAC profiles. We reduced the dimensionality using Latent Sematic Indexing (LSI) and projected the cells into a 2-dimensional space using Uniform Manifold Approximation and Projection (UMAP). We observed that different CRISPR-KO treatments were separated from the WT and enriched in the different areas in the UMAP embedding space, indicating that the KO of these 11 proteins differently altered the chromatin accessibility and that scACCESS-ATAC was able to capture these differences (Fig.5B). We next clustered the cells and identified two major populations, designated as C1 and C2, which presumably correspond to the two component cell lines (Fig.5B). To annotate the clusters, we generated the pseudo- bulk ATAC profiles for C1 and C2 and compared them to the ATAC-seq data for K562 and HepG2 cells obtained from ENCODE project. We observed a strong correlation between C1 and K562 (Spearman r = 0.86) and between C2 and HepG2 (Spearman r = 0.85) (Fig.5C). We therefore annotated C1 as K562, and C2 as HepG2. Together, these results demonstrate that scACCESS-ATAC can effectively distinguish between different cell states and types within complex cellular mixtures. Single-cell ATAC-seq has been used to compare TF binding activity in pseudo-bulk populations. To evaluate whether the increased resolution of scACCESS- ATAC could improve this application, especially for rare populations, we performed a proof-of-concept analysis. We subsampled K562 cells and assessed CTCF binding by plotting Tn5 insertions or DddSs edit counts surrounding K562 CTCF ChIP-seq- bound motifs for different number of cells (Fig.5D). As expected, we observed that the clarity of the peak and footprint profile improved as increasing cell numbers were included. Moreover, using DddSs editing signal consistently produced a cleaner profile than using Tn5 insertion signal at a given cell number. To quantify this improvement, we calculated the Pearson correlation between the subsampled CTCF profile and the profile obtained by using all cells (Fig.5E). We found that DddSs editing data requires 2-3-fold fewer cells than Tn5 insertion data to achieve a similar Attorney Docket No.29618-0483WO1 / BWH 2023-476 quality profile (similar Pearson correlation) across a wide range of subsampled inputs, demonstrating the superior performance of scACCESS-ATAC for detecting TF binding activity even with limited cell numbers. We next investigated the impact of CRISPR-KO on TF binding activity. We created a pseudo-bulk profile for each cell type and experimental condition by aggregating the scACCESS-ATAC data. Next, we performed differential footprinting analysis by comparing the ACCESS profiles between sgControl and sgZBTB7A based on the ChIP-seq peaks of ZBTB7A in K562. We observed that sgZBTB7A treatment significantly diminished ZBTB7A binding (Fig.5F). These findings demonstrate the utility of the additional information recoverable from the single-cell level editing patterns provided by scACCESS-ATAC in interrogating TF binding differences across cellular contexts. Example 6: Expand the editing range of ACCESS. We have currently tested seven Ddd enzymes (DddA, DddSs / Ddd1, CseDa01, FlDa01, CrDa01, MGYPDa01, and MGYPDa829), finding DddSs to show the broadest activity across motif contexts. We will expand our tests to profile 25+ Ddd orthologs reported in recent work8,9, drawing from evolutionary branches found to harbor Ddd enzymes with low motif-specificity. We will also test DddA variants evolved to show broader motif specificity7. By using an efficient in vitro transcription / translation approach followed by ACCESS-WGS, this scale of testing is quite feasible. We will investigate either switching to or combining additional Ddd enzymes with DddSs to give the most robust editing across genomic cytosines. Example 7: Optimize experimental steps in scACCESS-ATAC. There are hundreds of publications using single-cell ATAC-seq, so we anticipate that scACCESS-ATAC will become widely used. Experimentally, we will expand the generalizability of scACCESS-ATAC in several ways. We will optimize scACCESS-ATAC protocols using human CD34+ progenitors and mouse liver tissue to evaluate its performance on liquid and solid primary tissue and to obtain datasets with well-characterized rare subpopulations. We will evaluate scACCESS-ATAC using barcoded Tn5 to multiplex up to 24 distinctly treated HepG2 and K562 samples into a single run at superloaded cell density19. We will additionally adapt scACCESS- ATAC to the 10X Multiome assay to optimize compatibility in this setting. In all Attorney Docket No.29618-0483WO1 / BWH 2023-476 paradigms, we will collect 3 biological replicates with high coverage to serve as resources for other groups to analyze scACCESS-ATAC data. References 1. Weintraub, H. & Groudine, M. Chromosomal subunits in active genes have an altered conformation. Science 193, 848–56 (1976). 2. Boyle, A. P. et al. High-resolution mapping and characterization of open chromatin across the genome. Cell 132, 311–22 (2008). 3. Buenrostro, J. D., Giresi, P. G., Zaba, L. C., Chang, H. Y. & Greenleaf, W. J. Transposition of native chromatin for fast and sensitive epigenomic profiling of open chromatin, DNA-binding proteins and nucleosome position. Nat Methods 10, 1213–8 (2013). 4. Stergachis, A. B., Debo, B. M., Haugen, E., Churchman, L. S. & Stamatoyannopoulos, J. A. Single-molecule regulatory architectures captured by chromatin fiber sequencing. Science 368, 1449–1454 (2020). 5. Kelly, T. K. et al. Genome-wide mapping of nucleosome positioning and DNA methylation within individual DNA molecules. Genome Res.22, 2497– 2506 (2012). 6. Mok, B. Y. et al. A bacterial cytidine deaminase toxin enables CRISPR-free mitochondrial base editing. Nature 583, 631–637 (2020). 7. Mok, B. Y. et al. CRISPR-free base editors with enhanced activity and expanded targeting scope in mitochondrial and nuclear DNA. Nat. Biotechnol.40, 1378–1387 (2022). 8. Huang, J. et al. Discovery of deaminase functions by structure-based protein clustering. Cell 186, 3182-3195.e14 (2023). 9. Vaisvila, R. et al. Discovery of novel DNA cytosine deaminase activities enables a nondestructive single-enzyme methylation sequencing method for base resolution high-coverage methylome mapping of cell-free and ultra-low input DNA.2023.06.29.547047 Preprint at doi.org / 10.1101 / 2023.06.29.547047 (2023). 10. Mok, Y. G. et al. Base editing in human cells with monomeric DddA- TALE fusion deaminases. Nat. Commun.13, 4038 (2022). 11. Morgens, D. W. et al. Genome-scale measurement of off-target activity using Cas9 toxicity in high-throughput screens. Nat. Commun.8, 15178 (2017). Attorney Docket No.29618-0483WO1 / BWH 2023-476 12. Langmead, B. & Salzberg, S. L. Fast gapped-read alignment with Bowtie 2. Nat. Methods 9, 357–359 (2012). 13. Dunham, I. et al. An integrated encyclopedia of DNA elements in the human genome. Nature 489, 57–74 (2012). 14. Li, Z. et al. Identification of transcription factor binding sites using ATAC-seq. Genome Biol.20, 45 (2019). 15. He, H. H. et al. Refined DNase-seq protocol and data analysis reveals intrinsic bias in transcription factor footprint identification. Nat. Methods 11, 73–78 (2013). 16. Cazares, T. A. et al. maxATAC: Genome-scale transcription-factor binding prediction from ATAC-seq with deep neural networks. PLOS Comput. Biol. 19, e1010863 (2023). 17. Pratt, H. E. et al. Factorbook: an updated catalog of transcription factor motifs and candidate regulatory motif sites. Nucleic Acids Res.50, D141–D149 (2022). 18. Granja, J. M. et al. ArchR is a scalable software package for integrative single-cell chromatin accessibility analysis. Nat. Genet.53, 403–411 (2021). 19. Zhang, H. et al. txci-ATAC-seq, a massive-scale single-cell technique to profile chromatin accessibility.2023.05.11.540245 Preprint at doi.org / 10.1101 / 2023.05.11.540245 (2023). 20. Shipony, Z. et al. Long-range single-molecule mapping of chromatin accessibility in eukaryotes. Nat. Methods 17, 319–327 (2020). 21. Kelly, T. K. et al. Genome-wide mapping of nucleosome positioning and DNA methylation within individual DNA molecules. Genome Res.22, 2497– 2506 (2012). 22. Battaglia, S. et al. Long-range phasing of dynamic, tissue-specific and allele-specific regulatory elements. Nat. Genet.54, 1504–1513 (2022). 23. Gallagher, L. A. et al. Genome-wide protein-DNA interaction site mapping in bacteria using a double-stranded DNA-specific cytosine deaminase. Nat. Microbiol.7, 844–855 (2022). 24. Vaisvila, R. et al. Discovery of cytosine deaminases enables base- resolution methylome mapping using a single enzyme. Mol. Cell S1097276524000947 (2024) doi:10.1016 / j.molcel.2024.01.027. Attorney Docket No.29618-0483WO1 / BWH 2023-476 25. Swanson, E. G. et al. Deaminase-assisted single-molecule and single- cell chromatin fiber sequencing.2024.11.06.622310 Preprint at doi.org / 10.1101 / 2024.11.06.622310 (2024). 26. Wang, H. et al. Widespread plasticity in CTCF occupancy linked to DNA methylation. Genome Res.22, 1680–1688 (2012). 27. Cazares, T. A. et al. maxATAC: Genome-scale transcription-factor binding prediction from ATAC-seq with deep neural networks. PLOS Comput. Biol. 19, e1010863 (2023). 28. Pratt, H. E. et al. Factorbook: an updated catalog of transcription factor motifs and candidate regulatory motif sites. Nucleic Acids Res.50, D141–D149 (2022). 29. Grant, C. E., Bailey, T. L. & Noble, W. S. FIMO: scanning for occurrences of a given motif. Bioinformatics 27, 1017–1018 (2011). 30. Li, Z. et al. Identification of transcription factor binding sites using ATAC-seq. Genome Biol.20, 45 (2019). 31. Hu, Y. et al. Single-cell multi-scale footprinting reveals the modular organization of DNA regulatory elements.2023.03.28.533945 Preprint at doi.org / 10.1101 / 2023.03.28.533945 (2023). 32. Sönmezer, C. et al. Molecular Co-occupancy Identifies Transcription Factor Binding Cooperativity In Vivo. Mol. Cell 81, 255-267.e6 (2021). 33. Almogy, G. et al. Cost-efficient whole genome-sequencing using novel mostly natural sequencing-by-synthesis chemistry and open fluidics platform. 2022.05.29.493900 Preprint at doi.org / 10.1101 / 2022.05.29.493900 (2022). 34. Partridge, E. C. et al. Occupancy maps of 208 chromatin-associated proteins in one human cell type. Nature 583, 720–728 (2020). 35. Moore, J. E. et al. Expanded encyclopaedias of DNA elements in the human and mouse genomes. Nature 583, 699–710 (2020). 36. de Martin, X., Sodaei, R. & Santpere, G. Mechanisms of Binding Specificity among bHLH Transcription Factors. Int. J. Mol. Sci.22, 9150 (2021). 37. Moriggl, R. et al. Stat5 tetramer formation is associated with leukemogenesis. Cancer Cell 7, 87–99 (2005). Attorney Docket No.29618-0483WO1 / BWH 2023-476 38. Soldaini, E. et al. DNA binding site selection of dimeric and tetrameric Stat5 proteins reveals a large repertoire of divergent tetrameric Stat5a binding sites. Mol. Cell. Biol.20, 389–401 (2000). 39. Xu, X., Sun, Y. L. & Hoey, T. Cooperative DNA binding and sequence-selective recognition conferred by the STAT amino-terminal domain. Science 273, 794–797 (1996). 40. Buettner, R., Mora, L. B. & Jove, R. Activated STAT signaling in human tumors provides novel molecular targets for therapeutic intervention. Clin. Cancer Res. Off. J. Am. Assoc. Cancer Res.8, 945–954 (2002). 41. He, R. et al. Human transcription factor combinations mapped by footprinting with deaminase.2024.06.14.599019 Preprint at doi.org / 10.1101 / 2024.06.14.599019 (2024). 42. Mirny, L. A. Nucleosome-mediated cooperativity between transcription factors. Proc Natl Acad Sci U A 107, 22534–9 (2010). 43. Hashimoto, T. et al. A synergistic DNA logic predicts genome-wide chromatin accessibility. Genome Res. (2016) doi:10.1101 / gr.199778.115. 44. Zhu, F. et al. The interaction landscape between transcription factors and the nucleosome. Nature 562, 76–81 (2018). 45. Cui, F. & Zhurkin, V. B. Rotational positioning of nucleosomes facilitates selective binding of p53 to response elements associated with cell cycle arrest. Nucleic Acids Res.42, 836–847 (2014). 46. Avsec, Ž. et al. Base-resolution models of transcription-factor binding reveal soft motif syntax. Nat. Genet.53, 354–366 (2021). 47. de Boer, C. G. et al. Deciphering eukaryotic gene-regulatory logic with 100 million random promoters. Nat. Biotechnol.38, 56–65 (2020). 48. Duttke, S. H. et al. Position-dependent function of human sequence- specific transcription factors. Nature 631, 891–898 (2024). 49. Szczesnik, T., Chu, L., Ho, J. W. K. & Sherwood, R. I. A High- Throughput Genome-Integrated Assay Reveals Spatial Dependencies Governing Tcf7l2 Binding. Cell Syst.11, 315-327.e5 (2020). 50. Luscombe, N. M., Austin, S. E., Berman, H. M. & Thornton, J. M. An overview of the structures of protein-DNA complexes. Genome Biol.1, reviews001.1-reviews001.37 (2000). Attorney Docket No.29618-0483WO1 / BWH 2023-476 51. Morgunova, E. & Taipale, J. Structural perspective of cooperative transcription factor binding. Curr. Opin. Struct. Biol.47, 1–8 (2017). 52. Richter, W. F., Nayak, S., Iwasa, J. & Taatjes, D. J. The Mediator complex as a master regulator of transcription by RNA polymerase II. Nat. Rev. Mol. Cell Biol.23, 732–749 (2022). 53. Kim et al., Probing Allostery Through DNA Science.2013 Feb 15;339(6121):816-9. 54. Brennan, K. J. et al. Chromatin accessibility in the Drosophila embryo is determined by transcription factor pioneering and enhancer activation. Dev. Cell 58, 1898-1916.e9 (2023). 55. Hammelman, J., Krismer, K., Banerjee, B., Gifford, D. K. & Sherwood, R. I. Identification of determinants of differential chromatin accessibility through a massively parallel genome-integrated reporter assay. Genome Res.30, 1468–1480 (2020). 56. de Moraes, M. H. et al. An interbacterial DNA deaminase toxin directly mutagenizes surviving target populations. eLife 10, e62967 (2021). 57. Marty, M. T. et al. Bayesian deconvolution of mass and ion mobility spectra: from binary interactions to polydisperse ensembles. Anal. Chem.87, 4370– 4376 (2015). 58. Granja, J. M. et al. ArchR is a scalable software package for integrative single-cell chromatin accessibility analysis. Nat. Genet.53, 403–411 (2021). 59. Stuart, T., Srivastava, A., Madad, S., Lareau, C. A. & Satija, R. Single- cell chromatin state analysis with Signac. Nat. Methods 18, 1333–1341 (2021). OTHER EMBODIMENTS It is to be understood that while the invention has been described in conjunction with the detailed description thereof, the foregoing description is intended to illustrate and not limit the scope of the invention, which is defined by the scope of the appended claims. Other aspects, advantages, and modifications are within the scope of the following claims.

Claims

Attorney Docket No.29618-0483WO1 / BWH 2023-476 WHAT IS CLAIMED IS:

1. A method of analyzing chromatin, the method comprising: providing a sample comprising one or more intact cell nuclei comprising chromatin; contacting the intact nuclei concurrently with (i) one or more Ddd to convert accessible cytosines in the chromatin to uracils (C^U) and (ii) Tn5 transposase loaded with oligonucleotides comprising dsDNA sequencing adapters, wherein the oligonucleotides optionally further comprise tags, to fragment and tag the chromatin; and detecting conversion of C^U in the chromatin.

2. The method of claim 1, wherein detecting conversion of C^U in the chromatin comprises: amplifying the chromatin with a uracil-tolerant polymerase to convert the uracils to thymines (U^T), and performing sequencing to detect mutations of C^T, preferably wherein the sequencing is next generation sequencing.

3. The method of claim 2, wherein the uracil-tolerant polymerase is PHUSION high fidelity DNA polymerase; Kapa HiFi Hotstart Uracil+ DNA B family Polymerase; KOD Multi Epi high fidelity DNA Polymerase; or Q5U Hot Start High-Fidelity DNA Polymerase.

4. The method of claim 1, wherein detecting conversion of C^U in the chromatin comprises: contacting the nuclei with restriction enzyme DpnII, and comparing activity of the restriction enzyme DpnII in the chromatin in the nuclei to reference activity of the restriction enzyme DpnII in reference chromatin not treated with the Ddd.

5. A method of analyzing chromatin, the method comprising: providing a sample comprising one or more intact cell nuclei comprising chromatin; contacting the intact nuclei with one or more double-stranded DNA cytosine deaminase (Ddd) to convert accessible cytosines in the chromatin to uracilsAttorney Docket No.29618-0483WO1 / BWH 2023-476 (C^U); optionally contacting the intact nuclei with a Ddd inhibitor (Dddi) to inhibit the Ddd or purifying the DNA immediately after Ddd treatment to remove the Ddd; and detecting conversion of C^U in the chromatin.

6. The method of claim 5, wherein detecting conversion of C^U in the chromatin comprises: contacting the nuclei with Tn5 transposase loaded with oligonucleotides comprising dsDNA sequencing adapters, and optionally further comprising tags, to fragment and tag the chromatin; amplifying the chromatin fragments with a uracil-tolerant polymerase to convert uracils to thymines (U^T); and sequencing the fragments, preferably wherein sequencing comprises using next generation sequencing.

7. The method of claim 6, wherein the uracil-tolerant polymerase is PHUSION high fidelity DNA polymerase; Kapa HiFi Hotstart Uracil+ DNA B family Polymerase; KOD Multi Epi high fidelity DNA Polymerase; or Q5U Hot Start High-Fidelity DNA Polymerase.

8. The method of claim 5, wherein detecting conversion of C^U in the chromatin comprises: contacting the nuclei with restriction enzyme DpnII, and comparing activity of the restriction enzyme DpnII in the chromatin in the nuclei to reference activity of the restriction enzyme DpnII in reference chromatin not treated with the Ddd.

9. The method of any of claims 1-8, wherein the Ddd comprises DddSs from Simiaoa sunii.

10. The method of any of claims 1-8, wherein the Ddd inhibitor comprises double- stranded DNA deaminase immunity protein (DddI) (from Simiaoa sunii or from Burkholderia cenocepacia (strain H111)).Attorney Docket No.29618-0483WO1 / BWH 2023-476 11. The method of any of claims 1-10, further comprising: comparing sequences obtained from sequencing the fragments to a reference genome for the cell; and identifying positions in the sequences that show conversion of a cytosine in the reference genome to a thymine in the fragment sequences; and designating the position in the genome as accessible.

12. A kit for use in a method of any of claims 1-11, the kit comprising: reagents for isolating nuclei from a population of cells; double-stranded DNA cytosine deaminase (Ddd); a uracil-tolerant polymerase; and one or more reaction buffers; and optionally a Ddd inhibitor (Dddi).

13. The kit of claim 12, further comprising: a Tn5 transposase and oligonucleotides comprising transposon tags and sequencing adaptors, optionally wherein the Tn5 transposase and oligonucleotides are present in pre-formed complexes.

14. The kit of claims 12 or 13, wherein the uracil-tolerant polymerase is PHUSION high fidelity DNA polymerase; Kapa HiFi Hotstart Uracil+ DNA B family Polymerase; KOD Multi Epi high fidelity DNA Polymerase; or Q5U Hot Start High-Fidelity DNA Polymerase.

Citation Information

Patent Citations

  • Method for constructing high-resolution single cell Hi-C library with a lot of information

    US10900031B2

  • Methods of determining genome-wide DNA binding protein binding sites by footprinting with double stranded DNA deaminase

    WO2024065721A1

  • Method for DNA methylation detection of single cell

    WO2024146540A1