Footprinting protein occupancy along individual chromatin fibers

The use of a double-stranded DNA deaminase for cytosine deamination in chromatin mapping methods allows for sensitive and amplified sequencing of protein-DNA interactions, overcoming limitations of existing technologies and enabling precise chromatin structure analysis in single cells.

WO2026050388A1PCT designated stage Publication Date: 2026-03-05UNIV OF WASHINGTON
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/043751
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-28
Filing Date
2025-08-27
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

Current long-read chromatin mapping technologies are incompatible with DNA amplification, limiting the study of rare cell types and precious samples due to the use of enzymatic methylation that impedes targeted profiling and comprehensive single-cell approaches.

Method used

A method using a double-stranded DNA deaminase, such as SsDddA, to deaminate cytosine bases in accessible portions of chromatin, followed by sequencing and amplification, preserving chromatin architectures for precise mapping of protein-DNA interactions.

Benefits of technology

Enables sensitive chromatin mapping compatible with DNA amplification, allowing for precise mapping of protein occupancy and chromatin structures in single cells, resolving rare somatic variants, and identifying distinct chromatin architectures between normal and tumor cells.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025043751_05032026_PF_FP_ABST
    Figure US2025043751_05032026_PF_FP_ABST
Patent Text Reader

Abstract

Methods and kits for probing chromatin accessibility, including Deaminase Assisted single-molecule chromatin Fiber sequencing (DAF-seq) methods and kits. The disclosed approaches leverage DNA deaminase toxin A (DddA) to efficiently stencil protein occupancy along DNA molecules via selective deamination of sterically accessible cytosine bases. These deamination events result in C-to-T transitions during DNA amplification, such that the resultant sequences preserve the chromatin architectures of the underlying individual chromatin fibers as per-molecule DNA alterations. The disclosed approaches can be used for identifying where proteins are bound to genomic DNA, mappable at a single-cell, single-molecule, and single-nucleotide level, and as diagnostic tools in clinical settings.
Need to check novelty before this filing date? Find Prior Art

Description

FOOTPRINTING PROTEIN OCCUPANCY ALONG INDIVIDUAL CHROMATIN FIBERS CROSS-REFERENCE TO RELATED APPLICATION

[0001] This international application claims the benefit of, and priority to, U.S. Provisional Application No.63 / 687,924, filed August 28, 2024; the disclosure of which is incorporated by reference herein in its entirety for all purposes. STATEMENT REGARDING SEQUENCE LISTING

[0002] The Sequence Listing XML associated with this application is provided in XML format and is hereby incorporated by reference into the specification. The name of the XML file containing the sequence listing is 3915-P1364WOUW-Sequence- Listing.xml. The XML file is 53,344 bytes; was created on August 26, 2025; and is being submitted electronically via Patent Center with the filing of the specification. STATEMENT OF GOVERNMENT LICENSE RIGHTS

[0003] This invention was made with government support under Grant Nos. 1UM1DA058220 and 5DP5OD029630, awarded by the National Institutes of Health (NIH). The government has certain rights in the invention. BACKGROUND

[0004] Understanding the functional syntax of the non-coding genome is a daunting challenge which stems from the intrinsic epigenetic heterogeneity between individual cells, tissue and context-specific gene regulatory programs, and the >3,000,000 genetic variants observed between individuals. These regulatory programs are dictated by the epigenome, which uses non-coding genomic elements as the means of establishing and maintaining distinct patterns of gene activity. The functional activity of each gene regulatory element is dictated by the co-binding of numerous chromatin proteins such as nucleosomes and transcription factors (TFs) to the same molecule of DNA (i.e., an individual chromatin fiber). Consequently, obtaining a complete view of human gene regulation requires precise mapping of DNA-protein occupancy in conjunction with the underlying genetic sequence.3915-P1364WO.UW -1-

[0005] Recent advances in long-read single-molecule chromatin mapping approaches have enabled an unprecedented exploration of the complexities of the human epigenome: enabling haplotype resolved chromatin mapping, the resolution of Mendelian disorders, and the discovery of novel regulatory elements within repetitive portions of the genome such as segmental duplications. However, long-read chromatin-mapping technologies all rely on enzymatic methylation to label accessible nucleotides, labels which are incompatible with DNA amplification, impeding the targeted profiling of select regions of interest or comprehensive single-cell approaches. As a result, long-read chromatin mapping approaches necessitate high cell numbers, costly whole-genome sequencing, and limit the study of rare cell types and precious samples.

[0006] Accordingly, there is a need for novel and sensitive chromatin mapping methods that are compatible with DNA amplification. The present disclosure addresses these and other long-felt and unmet needs in the art. SUMMARY

[0007] This summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This summary is not intended to identify key features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter.

[0008] In an aspect, the disclosure provides a method for mapping protein-DNA interactions to a DNA sequence of a chromosome, the method comprising: contacting the chromosome with a DNA deaminase for a deaminase reaction that comprises cytosine deamination of enzyme-accessible portions of the chromosome; and sequencing the DNA sequence of the chromosome, wherein the DNA sequence comprises a plurality of extended footprints that correspond to positions of the chromosome that did not undergo cytosine deamination due to protein-DNA interactions of nucleosome-like proteins of the chromosome.

[0009] In embodiments, the method is performed with a single cell, a circulating tumor cell, a plurality of cells, a tissue, a plurality of tissues, an organ, a plurality of organs, an organism, or a plurality of organisms.

[0010] In embodiments, the DNA deaminase is SsDddA or a double-stranded DNA deaminase (dsDNA-deaminase) derived from Simiaoa sunii (DddAtox).3915-P1364WO.UW -2-

[0011] In embodiments, the DNA deaminase is SsDddA and comprises 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO:42.

[0012] In embodiments, the DNA deaminase is SsDddA and comprises 100% identity with SEQ ID NO:42.

[0013] In embodiments, the DNA deaminase is SsDddA5 and comprises 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO:43.

[0014] In embodiments, the DNA deaminase is SsDddA5 and comprises 100% identity with SEQ ID NO:43.

[0015] In embodiments, the deaminase reaction is stopped with addition of the DddA immunity protein, DddI, thereto.

[0016] In embodiments, the method further comprises amplifying the chromosome, or a portion thereof, by preferentially priming amplification on a primary DNA template that comprises C to T nucleotide transitions or G to A nucleotide transitions that correspond to locations of the chromosome that underwent cytosine deamination by the DNA-deaminase.

[0017] In embodiments, the sequencing comprises translocating the DNA sequence through a nanopore, circular consensus sequencing, sequencing by synthesis, or sequencing by binding.

[0018] In embodiments, the method further comprises amplifying the chromosome or a genome that comprises the chromosome with a whole genome amplification (WGA) method.

[0019] In embodiments, the WGA method comprises primary template-directed amplification (PTA).

[0020] In embodiments, the method further comprises aligning multiple sequence reads based on a unique molecular identifier (UMI) that comprises a unique cytosine deamination pattern formed by cytosine deamination of a primary DNA template that comprises C to T nucleotide transitions or G to A nucleotide transitions that correspond to locations of the chromosome that underwent cytosine deamination by the DNA-deaminase.

[0021] In embodiments, the method further comprises generating a chromatin accessibility map based on evidence of cytosine deamination, wherein the chromatin accessibility map comprises information that indicates one or more regions of chromatin that are not bound by proteins and one or more regions of chromatin that are bound by proteins.3915-P1364WO.UW -3-

[0022] In embodiments, the chromosome constitutes a portion of a genome of a cell derived from a subject.

[0023] In embodiments, the subject has or is suspected of having a condition, disease, or disorder.

[0024] In embodiments, the condition, disease, or disorder is somatic or germline.

[0025] In embodiments, the chromosome is edited.

[0026] In embodiments, an edit of the chromosome is produced by a Cas method or a CRISPR / Cas9 method.

[0027] In embodiments, the Cas method or the CRISPR / Cas9 method comprises adenine base editing.

[0028] In embodiments, the method further comprises comparing a chromatin accessibility map of a test cell to a chromatin accessibility map of a normal cell; wherein a difference in a DNA sequence and no difference in a protein-DNA interaction indicates the difference in the DNA sequence is independent of the protein-DNA interaction; and wherein a difference in a DNA sequence and a difference in a protein-DNA interaction indicates the difference in the DNA sequence is not independent of the protein-DNA interaction.

[0029] In embodiments, a protein of a protein-DNA interaction is a histone.

[0030] In embodiments, the method further comprises generating a database comprising data that corresponds to a chromatin structure, an underlying genomic DNA sequence, a correlation to a condition, a correlation to a disease, a correlation to a disorder, or any combination thereof.

[0031] In embodiments, sequence data is suitable for use to identify PCR duplicates.

[0032] In embodiments, identified PCR duplicates can be used to generate a consensus sequence of a sequence read.

[0033] In embodiments, sequence data can be used to reconstruct a genomic sequence of a sequenced sample, identify minimal residual disease, or both.

[0034] In an aspect, the disclosure provides a kit, comprising: one or more reagents configured for performance of the method; and an instructional material.

[0035] In embodiments, the kit further comprises the DNA deaminase.3915-P1364WO.UW -4-

[0036] In an aspect, the disclosure provides a non-transitory computer readable storage medium having stored thereon instructions which, when executed by one or more processors, cause the one or more processors to perform all or part of the method.

[0037] In an aspect, the disclosure provides a computational device or system comprising the non-transitory computer readable storage medium. DESCRIPTION OF THE DRAWINGS

[0038] The foregoing aspects and many of the attendant advantages of this disclosure will become more readily appreciated as the same become better understood by reference to the following detailed description, when taken in conjunction with the accompanying drawings.

[0039] FIGs 1A-1G show results from deaminase-assisted single-molecule chromatin fiber sequencing (DAF-seq), according to aspects of the disclosure. FIG. 1A shows a schematic for DAF-seq using the non-specific cytidine deaminase SsDddA to selectively stencil single-molecule protein occupancy using deaminated cytidines. FIG.1B shows a demonstration of how protein occupancy impacts cytidine deamination and how this results in C->T or G->A transitions relative to the reference depending on whether the top or bottom strand is used as the template. FIG. 1C shows a diagram showing deamination of cytidine to uridine and subsequent conversion to thymidine upon PCR amplification. Mass spectrometry quantification of cytidine in untreated genomic DNA, as well as genomic DNA treated with wild-type SsDddA and the variant SsDddA 5. FIG.1D shows motif logos showing the sequence context of cytidine deamination genome-wide after treatment of nuclei with wild-type SsDddA and the variant SsDddA 5. FIG.1E shows a genomic locus showing GM12878 DNase-seq, pseudo-bulked scATAC-seq, Fiber-seq and targeted DAF-seq data at the NAPA promoter. Region targeted for amplification are noted, as well as deamination rate, and coverage of chromatin features derived from the DAF-seq data. DAF-seq used only 1% of a SMRT cell for this sequencing, and only 78 / 25,768 DAF-seq molecules from this locus are displayed. FIG.1F shows (left) Median per-base deamination rates within the NAPA and WASF1 promoters, and a well-positioned nucleosome and highly occupied CTCF element, after treatment of GM12878 cells with various DAF-seq reaction conditions. (right) Violin plots showing percent deamination at the highest (4 uM) and lowest (0.25 uM) SsDddA reaction concentrations. Bases are indicated as TpC and non-TpC dinucleotide context. FIG. 1G shows enrichment of the3915-P1364WO.UW -5-targeted region versus genome-wide sequencing after targeted DAF-seq of four separate loci in GM12878 cells.

[0040] FIGs 2A-2F show DAF-seq disentangles the regulatory logic of single- molecule TF co-occupancy, according to aspects of the disclosure. FIG. 2A shows (top) targeted DAF-seq chromatin actuation and ChIP-seq data from the NAPA promoter in GM12878 cells. (bottom) Zoom-in of the NAPA promoter showing single-molecule DAF- seq profiles with deaminated bases are indicated. Only a subset of ‘top-strand’ reads is shown, with reads clustered based on their occupancy pattern at 11 well-defined TF binding elements within the NAPA promoter. FIG.2B shows a bar graph showing single-molecule protein occupancy measurements of the 11 binding elements within the NAPA promoter. FIG. 2C shows a bar graph showing the number of chromatin fibers with occupancy at different combinations of elements 1 and 2. FIG.2D shows thermodynamic stability of TF occupancy and co-occupancy on the NAPA promoter. (top) Heatmaps showing computed ΔG values for individual protein-DNA interactions and pair-wise interactions. Only pair- wise interactions that passed significance testing (P < 0.01, binomial test) are shown. (bottom) Cartoon diagram representing pair-wise interactions on the NAPA promoter. Stippled shaded squares represent thermodynamically favorable interactions relative to the reference state where no footprints are occupied, and solid shaded squares represent thermodynamically unfavorable interactions. The thickness of the curved lines connecting the numbered interactions is proportional to the absolute value of the interactions. FIG.2E shows computed ΔG values (methods) for individual protein-DNA interactions (top) and pair-wise protein-protein interactions (bottom). FIG. 2F shows computed ΔG values for three-way protein-protein interactions including elements 1 and 2 (top) and three-way protein-protein interactions not including either element 1 or element 2 (bottom).

[0041] FIGs 3A-3F show simultaneous single-molecule genomic and chromatin profiles, according to aspects of the disclosure. FIG. 3A shows (left) a diagram showing the evaluation of the underlying genetic architecture at each reference genome position. Specifically, reads are divided into ‘top-strand’ and ‘bottom-strand’ based on their predominance of C->T versus G->A mutations relative to reference. Below, the base content at a position at all reads from the top and bottom strands are quantified and used to evaluate the germline genetic content at that position. (right) Hexbin plots showing the base content at all of the targeted DAF-seq regions used in this manuscript. FIG. 3B shows haplotype-phased single-molecule and aggregate targeted DAF-seq data of the UBA13915-P1364WO.UW -6-promoter in GM12878 cells which have allelically skewed XCI. The single C / T germline variant used for phasing is indicated. FIG. 3C shows a bar graph showing chromatin actuation at each of the four UBA1 TSSs by haplotype. FIGs 3D, 3E show a heatmap showing single-molecule pair-wise co-actuation (FIG. 3D) and codependency (FIG. 3E) of the four UBA1 TSSs along the Xa. FIG.3F shows (left) a zoom-in of UBA1 isoform 4 TSS showing single-molecule protein occupancy at four binding elements, and (right) a bar graph showing single-molecule protein occupancy at each of the four UBA1 isoform 4 TSS binding elements by haplotype. Human genomes are estimated to contain a single- nucleotide variant approximately every 1,200 bp, making prior methods, such as ATAC- seq, blind to haplotype-selective chromatin features outside of the 50-150 bp surrounding each variant. As such, conclusions gained from the results of FIGs 3B-3F could not be reached with any prior short-read methods, such as ATAC-seq, as this 4.4 kb region contains only a single SNP which is over 500 bp upstream of the isoform 4 TSS. This means that short-read accessibility data at any of the four TSSs would be unphased, preventing the identification of Xa-specific TSS usage (isoforms 1-3) and escape from X inactivation (isoform 4). As such, the biological conclusion that these elements are being independently actuated along the Xa could not have been derived using any short-read based method.

[0042] FIGs 4A-4F show SLC39A4 haplotypes modulate single-molecule promoter actuation patterns, according to aspects of the disclosure. FIG. 4A shows liver eQTL data from GTEx of three single nucleotide polymorphisms (SNPs) within the SLC39A4 promoter as well as their linkage disequilibrium using 1,000 genomes data. Below are DNase-seq data at this region in GM12878 cells and primary liver tissue. FIG. 4B shows (top) a diagram showing targeted DAF-seq on the SLC39A4 promoter in two cell types heterozygous for rs2280838 followed by the Leiden clustering of reads based on their deamination profiles and subsequent projection of these reads into UMAP space. The figure also shows (bottom-right) a pie chart showing the relative contribution of each cluster to the total number of reads. FIG. 4C shows (left) aggregate chromatin actuation and nucleosome occupancy of reads from the different clusters identified in FIG. 4B. The figure shows (right) a stacked bar chart showing the relative contribution of liver or GM12878 reads to the different clusters, as well as the relative contribution of rs2280838- C or rs2280838-T. FIG. 4D shows results from sub-clustering of the reads from cluster 1 in FIG. 4B. The figure shows (left) aggregate chromatin actuation and nucleosome3915-P1364WO.UW -7-occupancy of reads from the different subclusters. The figure shows (right) a stacked bar chart showing the relative contribution of liver or GM12878 reads to the different sub- clusters, as well as the relative contribution of rs2280838-C or rs2280838-T. The figure shows (bottom) promoter modules defined using above aggregate profiles. FIG.4E shows (left) a barplot showing the number of liver reads actuated at 0, 1, 2, 3, 4, or 5 of the SLC39A4 promoter modules defined in 4d, split by haplotype. The figure shows (right) a stacked barplot showing the proportion of fibers with chromatin actuation at only a set number of the 5 modules. FIG.4F shows a diagram showing the preference of rs2280838- T to actuate module C to switch from the closed to open chromatin state at the SLC39A4 promoter.

[0043] FIGs 5A-5C show resolving the functional impact of low VAF mosaic mutations, according to aspects of the disclosure. The experiment was designed to mimic a low frequency somatic variant at a known frequency, applicable to the field of somatic mosaicism, which is hampered by the inability to assess the functional consequences of non-coding somatic variants. The results demonstrate that DAF-seq can accurately quantify the frequency of low-frequency somatic variants while simultaneously identifying resulting chromatin changes. The results enable accurate quantification of somatic variants and their functional impact. This is a direct demonstration of a real use case as can be applicable to the Somatic Mosaicism across Human Tissues (SMaHT) network, for which targeted DAF- seq can be implemented to determine the functional impact of somatic variants identified using deep whole genome sequencing of post-mortem human tissues. FIG.5A shows (left) a schematic showing the generation of the COLO829 BLT50 cell mixture, which is a mixture of the lymphoblastoid cell line COLO829BL and the melanoma cell line COLO829T derived from the same individual. Fiber-seq was performed on each of these cell lines separately, as is shown in FIG.5B. The figure shows (right) whole genome PCR- free Illumina®sequencing and targeted DAF-seq of the COLO829 BLT50 mixture showing the VAF at the chr17:19447245_6 CC>TT variant. FIG.5B shows Fiber-seq data from the COLO829BL and COLO829T cells showing the per-molecule and aggregate chromatin actuation data surrounding the chr17:19447245_6 CC>TT variant, as well as the impact of this variant on a CTCF binding element. Note the complete loss of per-molecule CTCF occupancy and chromatin actuation on the reads containing the variant sequence. FIG.5C shows targeted DAF-seq of the same region as FIG.5B in the COLO829 BLT50 cell mixture showing the single-molecule and aggregate chromatin patterns on reads3915-P1364WO.UW -8-containing the reference sequence (top) and the chr17:19447245_6 CC>TT variant (bottom). Note the complete loss of chromatin actuation and nucleosome positioning along the variant reads.

[0044] FIGs 6A-6E show chromosome-scale genomic phasing in single cells, according to aspects of the disclosure. FIG. 6A shows a schematic for single-cell DAF- seq. Specifically, permeabilized cells are treated with SsDddA and then sorted into individual wells of a plate, with each well being subjected to a custom PTA. PTA reads from each well are then sequenced using PacBio®HiFi®sequencing and mapped to GRCh38. Individual reads are then identified as arising from either the ‘top’ or ‘bottom’ strand based on their pattern of either C>T or G>A mutations relative to the reference, respectively. Overlapping reads are then combined to generate consensus reads for each ‘haplotype-strand’, which are then haplotype-phased across the entire genome using parental short-read data. FIG.6B shows (top) a swarm plot showing the sequencing depth of each cell in terms of gigabase pairs (Gbp). Note that only a fraction of the library was sequenced for each cell. (bottom) Number of ultra-long consensus reads >100,000 bp in length from each cell. FIG.6C shows (top) a density plot showing the size distribution of sequencing reads (top) or ‘haplotype-strand’ consensus reads (bottom) from four of the cells subjected to scDAF-seq. FIG.6D shows (top) genomic coverage of chromosome 13 from the raw sequencing reads from cell 2, as well as the coverage from the consensus reads from this same cell. The figure shows (bottom) a genomic locus showing the consensus reads from cell 2, with reads indicated based on whether they are from the top or bottom strand, and split based on their phasing. FIG. 6E shows a swarm plot showing genomic coverage of consensus reads from each cell for all reads, as well as haplotype- phased reads.

[0045] FIGs 7A-7F show single-molecule chromatin epigenome of a single-cell, according to aspects of the disclosure. FIG. 7A shows (left) enrichments of scDAF-seq actuated elements (MSP > 150 bp) within Fiber-seq peaks from the same cell line (GM24385). Peaks are grouped into 10% Fiber-seq actuation bins ranging from 10-20% to 90-100%. (right) Enrichments of scDAF-seq actuated elements (MSP > 150 bp) within transcriptional start sites (TSS). FIG. 7B shows example genomic loci showing single- molecule and aggregate Fiber-seq data in bulk GM24385 cells (top), as well single- molecule deamination patterns along all ‘haplotype-strand’ consensus reads at these positions within cells 2 and 4. FIG. 7C shows scDAF-seq single-molecule deamination3915-P1364WO.UW -9-patterns along top and bottom-strand consensus reads at the same NAPA promoter position shown in FIG. 2A, as well as the position of TF binding elements within this promoter. FIG. 7D shows violin plots displaying Jaccard distance comparisons of single-molecule scDAF-seq actuation patterns between haplotypes within each cell and between the same haplotype of different cells at genomic loci that are covered across both cells (left). Elements are defined using paired Fiber-seq data and are further divided into promoter- proximal (center) and promoter-distal (right) peaks. Chromatin actuation patterns were significantly more similar between haplotypes of the same cell than between the same haplotype within different cells when comparing all peaks (two-sided t-test, P = 0.006) and promoter-distal peaks (P = 0.005) but not when comparing promoter-proximal peaks (P = 0.44). FIG.7E shows violin plots showing Jaccard distance comparisons between the same haplotype of different cells for Fiber-seq peaks grouped into 10% Fiber-seq actuation bins as in FIG. 7A. FIG.7F shows Jaccard distance comparisons between the same haplotype of different cells for TSS-overlapping Fiber-seq peaks grouped by binned log2 full-length transcript gene expression from the same cell line (GM24385). Peaks in gene expression bin 0 produced no detectable transcripts.

[0046] FIGs 8A-8E show chromosome-length single fiber co-actuation and protein co-occupancy, according to aspects of the disclosure. FIG. 8A shows (top left) a diagram of chromatin loop formation. The N-terminal domains of CTCF bound to DNA in opposite orientations halts progression of the cohesin complex during loop extrusion. The figure shows (center) a heatmap showing Micro-C chromatin interaction frequency in HFFc6 cells. The figure shows (bottom) an example chromatin loop region identified by CTCF ChIA-PET with aggregate Fiber-seq data in bulk GM24385 cells, as well as single- molecule deamination patterns at CTCF sites within loop anchors. CTCF occupancy status is displayed for each ‘haplotype-strand’ consensus read covering each CTCF site. FIG.8B shows Box plots showing the percentage of scDAF-seq chromatin fibers co-occupied by CTCF within loop anchor pairs from the top 500 ChIA-PET interactions, stratified by CTCF motif orientation. FIG. 8C shows box plots showing scDAF-seq co-occupancy of CTCF sites in the + / - orientation at loop anchors as in FIG.8B, binning loop anchor pairs by ChIA-PET strength. CTCF co-occupancy at shuffled regions is also shown. FIG. 8D shows a diagram showing how regulatory element codependency scores are calculated. FIG.8E shows a difference in average single-molecule codependency between regulatory elements along the same chromatin fiber (haplotype-strand) and opposite haplotypes within3915-P1364WO.UW -10-the same cell, binned by genomic distance. Stars indicate bins with significantly increased codependency scores along the same chromatin fiber (one-sided t-test, P < 0.05).

[0047] FIG. 9 shows percent actuation of targeted DAF-seq Oxford Nanopore®data, percent actuation of Fiber-seq PacBio®HiFi data, and bulk ATAC-seq read coverage from GM12878 cells, K562 cells, and colon tissue, at each of 10 target regions. DETAILED DESCRIPTION

[0048] Prior short-read chromatin mapping technologies, including ATAC-seq and DNase-seq, are incapable of haplotype phasing reads (and, as a result, signal), except at sites spanning heterozygous variants. However, human genomes are estimated to contain a single-nucleotide variant approximately every 1,200 bp, making these and other prior methods blind to haplotype-selective chromatin features outside of the 50-150 bp surrounding each variant.

[0049] To overcome limitations of prior long-read chromatin-mapping technologies, the present disclosure provides a novel and sensitive chromatin mapping method that is compatible with DNA amplification. Double-stranded cytosine deaminases (e.g., DddA) are leveraged to stencil protein occupancy at near nucleotide resolution. Deamination of cytosine bases to uracil creates distinct modifications to the DNA sequence that can be directly targeted and amplified, effectively making copies of single-fiber chromatin architectures and exponentially amplifying signal, prior to sequencing.

[0050] The disclosure provides Deaminase Assisted Fiber-seq (DAF-seq), a method for detecting protein occupancy along DNA molecules via selective deamination of unoccupied cytosine bases. The disclosed DAF-seq method works by treating either individual cells or a group of cells with a double-stranded DNA deaminase, followed by genomic DNA extraction. Extracted DNA is then subjected to whole-genome amplification, or targeted amplification and sequenced using a DNA sequencing instrument. These deamination events result in C-to-T transitions during DNA amplification, preserving the chromatin architectures of the underlying individual chromatin fibers.

[0051] The disclosed methods represent a significant advancement in the field; the disclosed analyses cannot be performed using any prior method, such as scATAC-seq. As an example of an analysis enabled by DAF-seq, the disclosure provides a visual overview of codependency calculations for both transcription factor (TF) occupancy and3915-P1364WO.UW -11-actuated regulatory elements (see, e.g., FIG.8D). The methods also convey codependency scores with probability, a metric that is interpretable to a person having ordinary skill in the art.

[0052] DAF-seq can be applied to cells carrying rare somatic variants and genomic edits and single-fiber chromatin architectures precisely mapped within targeted loci to unprecedented depth (>1,000,000X coverage). The present disclosure also demonstrates the ability of DAF-seq to resolve the functional impacts of rare somatic variants, identifying distinct chromatin architectures between normal and tumor cells within a mixed population. Furthermore, DAF-seq can be combined with CRISPR / Cas9 or base editing to interrogate the functional impact of genomic variants at scale. As a further example, DAF-seq can be performed in individual cells to enable the assembly of single- cell genomic and epigenomic patterns. The disclosed methods also open new avenues for understanding the diverse mechanisms in which non-coding genetic variants drive human disease. MAPPING PROTEIN-DNA INTERACTIONS TO A DNA SEQUENCE (DAF-SEQ)

[0053] In an aspect, the disclosure provides a method for mapping protein-DNA interactions to a DNA sequence of a DNA molecule, such as a chromosome. The method comprises contacting the DNA molecule with a non-biased double-stranded DNA deaminase for deamination of cytosines of sterically-accessible portions of the DNA molecule and sequencing the DNA sequence of the DNA molecule. As a result of DNA deamination in the presence of DNA-bound nucleosome-like proteins, the DNA sequence lacks nucleotide changes within extended footprints that correspond to positions of the chromosome that did not undergo cytosine deamination due to protein-DNA interactions of the nucleosome-like proteins of the DNA molecule. The method, coined “DAF-seq,” enables unprecedented insight into genomic epigenetic structure and organization and can be performed with limited biological material, such as a single cell.

[0054] While the DAF-seq method can be performed with a single cell, it is also scalable to accommodate larger sample sizes. As such, in various embodiments, the method can be performed with a single cell, a plurality of cells, a tissue, a plurality of tissues, an organ, a plurality of organs, an organism, or a plurality of organisms.

[0055] In embodiments, the DAF-seq method utilizes a DNA deaminase, which can be SsDddA or a DNA deaminase derived from Simiaoa sunii (DddAtox). In3915-P1364WO.UW -12-embodiments, the DNA deaminase is SsDddA5 or a DNA deaminase derived from Simiaoa sunii (DddAtox) that comprises a T76I mutation, a T109I mutation, or both, relative to a wild-type polypeptide sequence of the DNA-deaminase.

[0056] For example, in embodiments, the DNA deaminase is SsDddA and comprises 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO:42. In embodiments, the DNA deaminase is SsDddA and comprises 100% identity with SEQ ID NO:42.

[0057] In other example embodiments, the DNA deaminase is SsDddA5 and comprises 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO:43. In embodiments, the DNA deaminase is SsDddA5 and comprises 100% identity with SEQ ID NO:43.

[0058] In embodiments, the deaminase reaction is stopped with addition of the DddA inhibitor protein, Double-stranded DNA deaminase immunity protein (DddI), to the deaminase reaction. The DddI can include, but is not necessarily limited to, a DddI protein obtained or derived from Simiaoa sunii.

[0059] In embodiments, the DAF-seq method further comprises amplifying the chromosome, or a portion thereof, by preferentially priming amplification on a primary DNA template that comprises C to T nucleotide transitions or G to A nucleotide transitions that correspond to locations of the chromosome that underwent cytosine deamination by the DNA deaminase. This can be accomplished, in embodiments, with whole genome amplification (WGA) such as primary template-directed amplification (PTA).

[0060] PTA is a WGA technique having certain advantages for various methods of the disclosure. Compared with other WGA methods, PTA produces improved and reproducible genome sequencing coverage and variant detection from a single genome of a single cell. PTA can be used for measuring genetic diversity from single cells, including examining the acquisition of genetic changes during normal development and aging, measuring the consequences of specific perturbations such as genome editing, and characterizing the evolution of clonal populations during cancer formation.

[0061] With PTA, the incorporation of exonuclease-resistant terminators in the reaction results in smaller double-stranded amplification products that undergo limited subsequent amplification, resulting in a quasilinear process. PTA produces more amplification originating from the primary template, and as such, errors have limited propagation from daughter amplicons during subsequent amplification compared to other3915-P1364WO.UW -13-methods. In addition, PTA has improved and reproducible genome coverage breadth and uniformity, as well as diminished allelic skewing.

[0062] In embodiments, the method further comprises aligning multiple sequence reads based on a unique molecular identifier (UMI) that comprises a unique cytosine deamination pattern formed by cytosine deamination of a primary DNA template. For example, as DAF-seq chromatin stencils are maintained upon DNA amplification, these deamination patterns can, surprisingly, be used as a UMI to identify reads arising from the same modified DNA template, which is effective for identifying PCR duplicates and stitching PTA-produced reads together. This enables true single-molecule footprinting across chromosome-length chromatin fibers, a ~10,000-fold improvement beyond what was previously possible with prior PCR-free chromatin stenciling methods, and a ~1,000,000-fold improvement beyond what was previously possible with prior Tn5-based chromatin stenciling methods.

[0063] In embodiments, the method further comprises generating a chromatin accessibility map based on evidence of cytosine deamination. The chromatin accessibility map comprises information that indicates one or more regions of chromatin that are not bound by proteins and one or more regions of chromatin that are bound by proteins. In embodiments, a protein of a protein-DNA interaction is a histone. However, other proteins, including but not limited to transcription factors or other regulatory factors, can be mapped by methods of the disclosure.

[0064] In embodiments, the chromosome constitutes a portion of a genome of a cell derived from a subject that is analyzed by the DAF-seq method. In embodiments, the subject has or is suspected of having a condition, disease, or disorder. In embodiments, the condition, disease, or disorder is somatic or germline.

[0065] In embodiments, the DAF-seq method can be implemented in the analysis of an edited genomic sequence, for example, a genomic sequence that comprises an edit produced by a Cas method or a CRISPR / Cas9 gene editing or genomic editing technique. In embodiments, the Cas method or the CRISPR / Cas9 method comprises adenine base editing.

[0066] In embodiments, the method further comprises comparing a chromatin accessibility map of a test cell to a chromatin accessibility map of a normal cell, which can be useful for differentiating between genetic and epigenetic origins of phenotype. Accordingly, in embodiments, a difference in a DNA sequence and no difference in a3915-P1364WO.UW -14-protein-DNA interaction indicates the difference in the DNA sequence is independent of the protein-DNA interaction and a difference in a DNA sequence and a difference in a protein-DNA interaction indicates the difference in the DNA sequence is not independent of the protein-DNA interaction.

[0067] In embodiments, the method further comprises generating a database comprising data that corresponds to a chromatin structure, an underlying genomic DNA sequence, a correlation to a condition, a correlation to a disease, a correlation to a disorder, or any combination thereof.

[0068] In embodiments, sequence data is suitable for use to identify PCR duplicates. In embodiments, identified PCR duplicates can be used to generate a consensus sequence of a sequence read. In embodiments, sequence data can be used to reconstruct a genomic sequence of a sequenced sample.

[0069] In embodiments, the DNA sequence comprises a cellular DNA sequence within a cell, a cellular DNA sequence of a mammalian cell, a cellular DNA sequence of a cancer cell, a portion of a gene sequence, a portion of a non-coding genomic DNA sequence, a chromatin sequence, a chromatin fiber sequence, a chromosome sequence, a plurality of chromosome sequences, a nuclear genome sequence, a mitochondrial DNA (mtDNA) sequence, a mitochondrial genome sequence, a plurality of mitochondrial genome sequences, an isolated DNA sequence, a plasmid DNA sequence, a circulating cell free DNA sequence, an extracellular DNA sequence, or any combination thereof.

[0070] In embodiments, the method is performed with a single cell, a circulating tumor cell, a plurality of cells, a tissue, a plurality of tissues, an organ, a plurality of organs, an organism, or a plurality of organisms. Healthy and / or diseased biological samples can be tested, including for the construction and use of assays for diagnostic, prognostic, or other medical uses, in various embodiments.

[0071] In embodiments, the method further comprises generating a database comprising data that corresponds to a chromatin structure, an underlying genomic DNA sequence, a correlation to a condition, a correlation to a disease, a correlation to a disorder, or any combination thereof.

[0072] In embodiments, sequence data can be used to identify PCR duplicates. In embodiments, identified PCR duplicates can be used to generate a consensus sequence of a sequence read. In embodiments, sequence data can be used to reconstruct a genomic sequence of a sequenced sample, such as a single cell.3915-P1364WO.UW -15-DAF-SEQ KITS, COMPUTATIONAL DEVICES, AND SYSTEMS

[0073] In an aspect, the disclosure provides a kit, comprising one or more reagents configured for performance of a method and an instructional material. In embodiments, the kit comprises a DNA deaminase.

[0074] In an aspect, the disclosure provides a non-transitory computer readable storage medium having stored thereon instructions which, when executed by one or more processors, cause the one or more processors to perform all or part of a method. In another aspect, the disclosure provides a computational device or system comprising the non- transitory computer readable storage medium. TERMINOLOGY

[0075] Unless stated otherwise, experimental hypotheses or forward-looking models and statements are not intended to be binding on the applicant or exhaustive of the range of possible experimental hypotheses or forward-looking models and statements, but rather are intended to be illustrative, non-limiting examples for aiding those in the art in the understanding and practice of elements of the disclosure.

[0076] As used herein and unless otherwise indicated, the terms “a” and “an” are taken to mean “one”, “at least one” or “one or more”. Unless otherwise required by context, singular terms used herein shall include pluralities and plural terms shall include the singular.

[0077] As used herein, the term “biological system” refers to a living or non- living mixture of biological factors. Non-limiting examples of biological systems include cells, cell cultures, cell-free systems, tissues, tissue cultures, organs, organ cultures, and organisms.

[0078] Unless the context clearly requires otherwise, throughout the description and the claims, the words “comprise”, “comprising”, and the like are to be construed in an inclusive sense as opposed to an exclusive or exhaustive sense; that is to say, in the sense of “including, but not limited to”.

[0079] Unless the context clearly requires otherwise, the phrase “consisting essentially of” limits the scope of a claim to the specified materials or steps and those that do not materially affect the basic and novel characteristic(s) of the claim.3915-P1364WO.UW -16-

[0080] Unless the context clearly requires otherwise, the phrase “consisting of” excludes any element, step, or ingredient not specified.

[0081] If an element is described or claimed herein such that it “comprises” a feature, that description or claim also includes embodiments wherein the element “consists essentially of” and embodiments wherein the element “consists of” the feature, unless something else is specifically stated to the contrary.

[0082] Unless stated otherwise herein, 1-letter abbreviations for nucleic acids are consistent with the nomenclature used in the art (i.e., A, Adenine; T, Thymine; C, Cytosine; G, Guanine; U, Uracil). Unless stated otherwise herein, 1-letter abbreviations for amino acids are consistent with the nomenclature used in the art (i.e., Alanine, A; Arginine, R; Asparagine, N; Aspartic acid, D; Cysteine, C; Glutamic acid, E; Glutamine, Q; Glycine, G; Histidine H; Isoleucine, I; Leucine, L; Lysine, K; Methionine, M; Phenylalanine, F; Proline, P; Serine, S; Threonine, T; Tryptophan, W; Tyrosine, Y; Valine, V).

[0083] Unless stated otherwise herein, the terms “nucleic acid,” “amino acid,” “nucleotide,” and “peptide” are inclusive and open-ended, and do not exclude from their scope any chemically modified or post-translationally modified versions of these structures, and also do not exclude from their scope any nuclear modified versions, for example, due to the presence of one or more radioisotopes in one or more of these structures.

[0084] Unless otherwise stated or the context clearly requires otherwise, methods of the disclosure can be performed, in whole or in part, in any order of steps, including steps that are performed subsequently, in parallel, and in combination. In addition, methods can be performed, in whole or in part, by humans optionally assisted by one or more machines such as one or more computational devices or systems (e.g., computer(s)). In at least some instances, methods can be performed by one or more humans with little or no substantive assistance by one or more machines. In at least some other instances, methods can be performed by one or more humans with substantive assistance by one or more machines, and in at least some instances, one or more machines can perform methods autonomously or semi-autonomously.

[0085] As used herein, an “instructional material” includes a publication, a recording, a diagram, or any other medium of expression which can be used to communicate the usefulness of one or more elements of a kit of the disclosure for carrying out a method of the disclosure, including methods for DAF-seq as described herein.3915-P1364WO.UW -17-Optionally, or alternately, the instructional material can describe one or more methods of detecting rare genetic or genomic elements in a cell or a tissue of a mammal. The instructional material of the kit can, for example, be affixed to a container which contains an identified compound or enzyme, such as a DNA deaminase, such as SsDddA or a dsDNA deaminase (DddAtox) derived from Simiaoa sunii. The DNA deaminase can be or can comprise, for example, SsDddA5 or a dsDNA deaminase (DddAtox) derived from Simiaoa sunii that comprises a T76I mutation, a T109I mutation, or both. The instructional material can be shipped together with a container which contains the identified compound or enzyme. Alternatively, the instructional material can be shipped separately from the container with the intention that the instructional material and the compound or enzyme be used cooperatively.

[0086] As used herein, the term “non-transitory machine-readable storage medium”, and similar terms, refers to a computer readable medium or processor readable medium that can store information as data, as well as instructions for performance of logic operations by one or more processors for manipulation of the data. The non-transitory machine-readable storage medium can include any type of memory, such as volatile memory like random access memory (RAM), dynamic random-access memory (DRAM), static random-access memory (SRAM), or non-volatile memory like read-only memory (ROM), flash memory, magnetic or optical disks, or compact-disc read-only memory (CD- ROM), among other devices used to store data or programs on a temporary or permanent basis. The non-transitory machine-readable storage medium can be configured to store instructions. The instructions are executable by the one or more processors to cause the computing device to perform any of the functions or methods described herein. The non- transitory machine-readable storage medium can also be configured to store a computational model.

[0087] As used herein, the term “computational device” refers to a device, e.g., an electronic device, capable of performing logic operations. Example computational devices include computers, laptops, and tablets, as are known in the art. A computing device includes one or more processors, a non-transitory computer readable medium, a communication interface, a display, and a user interface. Components of the computing device are linked together by a system bus, network, or other connection mechanism. The one or more processors can be any type of processor(s), such as a microprocessor, a digital signal processor, a multicore processor, etc., coupled to the non-transitory computer3915-P1364WO.UW -18-readable medium. The communication interface can include hardware to enable communication within the computational device and / or between the computational device and one or more other devices. The hardware can include transmitters, receivers, and antennas, for example. The communication interface can be configured to facilitate communication with one or more other devices, in accordance with one or more wired or wireless communication protocols. For example, the communication interface can be configured to facilitate wireless data communication for the computational device according to one or more wireless communication standards, such as one or more Institute of Electrical and Electronics Engineers (IEEE) 801.11 standards, ZigBee standards, Bluetooth standards, etc. As another example, the communication interface can be configured to facilitate wired data communication with one or more other devices. The communication interface can also include analog-to-digital converters (ADCs) or digital- to-analog converters (DACs) that the computational device can use to control various components. The display can be any type of display component configured to display data. As one example, the display can include a touchscreen display. As another example, the display can include a flat-panel display, such as a liquid-crystal display (LCD) or a light- emitting diode (LED) display. A user interface can be included as part of the computational device and can include one or more pieces of hardware used to provide data and control signals to the computing device. For instance, the user interface can include a mouse or a pointing device, a keyboard or a keypad, a microphone, a touchpad, or a touchscreen, among other possible types of user input devices. Generally, the user interface can enable an operator to interact with a graphical user interface (GUI) provided by the computing device (e.g., displayed by the display)

[0088] As used herein, the term “system” refers to one or more computational devices, or one or more elements thereof, configured to perform one or more tasks, methods, or processes. A system of the disclosure can comprise a computing device. The computing device includes one or more processors, a non-transitory computer readable medium, a communication interface, a display, and a user interface. Components of the computing device are linked together by a system bus, network, or other connection mechanism.

[0089] Words using the singular or plural number also include the plural and singular number, respectively. Additionally, the words “herein,” “above,” and “below” and3915-P1364WO.UW -19-words of similar import, when used in this application, shall refer to this application as a whole and not to any particular portions of the application.

[0090] Unless otherwise indicated, all numbers expressing quantities of components, molecular weights, and so forth used in the specification and claims are to be understood as being modified in all instances by the term “about.” Accordingly, unless otherwise indicated to the contrary, the numerical parameters set forth in the specification and claims are approximations that can vary depending upon the desired properties sought to be obtained by the present disclosure. At the very least, and not as an attempt to limit the doctrine of equivalents to the scope of the claims, each numerical parameter should at least be construed in light of the number of reported significant digits and by applying ordinary rounding techniques.

[0091] Notwithstanding that the numerical ranges and parameters setting forth the broad scope of the disclosure are approximations, the numerical values set forth in the specific examples are reported as precisely as possible. All numerical values, however, inherently contain a range necessarily resulting from the standard deviation found in their respective testing measurements.

[0092] All headings are for the convenience of the reader and should not be used to limit the meaning of the text that follows the heading, unless so specified.

[0093] All of the references cited herein are incorporated by reference. Aspects of the disclosure can be modified, if necessary, to employ the systems, functions, and concepts of the cited references and disclosure to provide yet further embodiments of the disclosure. These and other changes can be made to the disclosure in light of the detailed description.

[0094] It will be appreciated that, although specific embodiments of the disclosure have been described herein for purposes of illustration, various modifications can be made without deviating from the spirit and scope of the disclosure. Accordingly, the disclosure is not limited except as stated by the claims. SEQUENCES Table 1. Polynucleotide and polypeptide sequences of the disclosure.3915-P1364WO.UW -20-Name Molecule Sequence SEQ ID T e NO3915-P1364WO.UW -21-Name Molecule Sequence SEQ ID T e NO3915-P1364WO.UW -22-Name Molecule Sequence SEQ ID T e NO3915-P1364WO.UW -23-Name Molecule Sequence SEQ ID T e NO3915-P1364WO.UW -24-Name Molecule Sequence SEQ ID T e NO3915-P1364WO.UW -25-Name Molecule Sequence SEQ ID T e NO3915-P1364WO.UW -26-Name Molecule Sequence SEQ ID T e NO3915-P1364WO.UW -27-Name Molecule Sequence SEQ ID T e NO3915-P1364WO.UW -28-Name Molecule Sequence SEQ ID T e NO3915-P1364WO.UW -29-Name Molecule Sequence SEQ ID T e NO3915-P1364WO.UW -30-Name Molecule Sequence SEQ ID T e NOEXAMPLES Example 1. Comprehensive single-cell diploid chromatin fiber architectures using DAF-seq (overview)

[0095] Gene regulation is orchestrated by the co-binding of proteins along chromosome-length chromatin fibers within single cells, yet the heterogeneity of this occupancy between haplotypes and cells remains poorly resolved in diploid organisms. In this and the other Examples, Deaminase-Assisted single-molecule chromatin Fiber sequencing (DAF-seq), which enables single-molecule footprinting at near-nucleotide resolution while synchronously profiling single-molecule chromatin states and DNA sequence, is described.

[0096] DAF-seq illuminates cooperative protein occupancy at individual regulatory elements and resolves the functional impact of somatic variants and rare chromatin epialleles. Single-cell DAF-seq (scDAF-seq) enables the accurate reconstruction of the diploid genome and chromatin epigenome from individual cells, generating chromosome-length protein co-occupancy maps across 99% of each cell’s mappable genome. ScDAF-seq uncovers extensive chromatin plasticity both within and between single diploid cells, with chromatin actuation diverging by 61% between haplotypes within a cell, and 63% between cells. Moreover, it is uncovered that regulatory elements are preferentially co-actuated along the same fiber in a distance-dependent manner that mirrors cohesin-mediated loops. Overall, DAF-seq enables the comprehensive characterization of protein occupancy across entire chromosomes with single-nucleotide, single-molecule, single-haplotype, and single-cell precision.3915-P1364WO.UW -31-

[0097] Human gene regulation occurs at the level of an individual chromatin fiber within a single diploid cell. However, these regulatory patterns can markedly diverge between individual chromatin fibers based on the stochastic nature of protein-DNA interactions, cell-to-cell variability in the abundance of chromatin proteins, the presence of underlying genetic variants (germline or somatic), or via the coordinated co-occupancy or co-actuation of chromatin features across loci or entire chromosomes. The chromatin architecture of each fiber is related to its functional output, yet it remains unknown how rigidly these architectures are maintained. Although numerous methods exist for resolving single-molecule protein occupancy or single-cell chromatin accessibility, prior approaches are limited in either their ability to achieve deep sequencing coverage, high-resolution protein occupancy patterns, or comprehensive high-resolution patterns across a single cell.

[0098] Prior long-read methyltransferase stenciling approaches are all bulk assays, owing to the erasure of methylation marks during DNA amplification. Consequently, these prior approaches limit the field’s understanding of the chromatin landscape within a single cell to individual 10-100 kb fibers, which cover only ~0.001% of that cell’s genome. Furthermore, Tn5-dependent single-cell chromatin assays, including those that utilize deaminase footprinting, result in sparse single-cell chromatin data, owing to inherent sampling and signal-to-noise limitations with cleavage-based chromatin mapping.

[0099] These drawbacks limit these prior approaches to single-cell readouts of scattered ~100 bp chromatin fibers that cover only ~0.01% of a cell’s genome. Given these fundamental technical limitations, the field sorely lacks a detailed understanding of the basic principles guiding how chromatin is organized and regulated within a single cell, including the extent to which a cell’s chromatin epigenome varies.

[0100] To overcome these and other limitations in the art, this disclosure provides a single-molecule chromatin fiber sequencing method that enables the mapping of single- molecule protein occupancy patterns across nearly the entire genome of an individual cell with haplotype resolution (FIG. 1A). Specifically, this method leverages a variant of the double-stranded cytidine deaminase toxin A (DddA) to stencil protein occupancy in the form of deaminated cytidines. Upon DNA amplification, these deaminated cytidines create distinct modifications to the DNA sequence that can be used to track sequencing reads that arose from the same DNA template, thereby enabling a person having ordinary skill in the art to resolve the genetic composition and protein occupancy of each chromatin fiber.3915-P1364WO.UW -32-Example 2. Simiaoa sunii DddA (SsDddA) efficiently and specifically modifies accessible cytidines

[0101] As described herein, cytidine deaminases with activity towards at least double-stranded DNA (dsDNA), such as the recently described enzyme DddA, provide an innovative and valuable solution for mapping chromatin architectures at C / G base pairs. DddA modifies cytidine to uridine, resulting in C->T mutations upon amplification (FIGs 1B-1C), which are readily observable by both short- and long-read DNA sequencing. However, for cytidine deaminases to be effective for chromatin stenciling, they can have minimal sequence biases, be highly catalytically active, and be highly specific for accessible DNA. Although the originally described DddA enzymes have substantial TC sequence biases, DddA variants with reduced sequence bias when tethered to a Cas9 enzyme can be implemented in example embodiments.

[0102] To test whether DddA variants are suitable for chromatin stenciling, the production of two recombinant DddA variants from the bacterial species Simiaoa sunii (SsDddA) was developed using a bacterial system, and cytidine deamination activity on purified double-stranded DNA (dsDNA) was quantified using mass spectrometry (FIG. 1C). This revealed that recombinant SsDddA (SEQ ID NO.42) is highly catalytically active, deaminating 99.8% of cytidines within dsDNA. To confirm that SsDddA has minimal sequence bias within a chromatin context, nuclei were treated with SsDddA and long-read sequencing of whole-genome amplified DNA from these SsDddA-treated nuclei was performed. This demonstrated that, unlike other cytidine deaminases, cytidine deamination with SsDddA has no appreciable sequence bias (FIG. 1D). Furthermore, the impact of 5-methylcytidine (5mC) on SsDddA activity was quantified using DNA templates treated with the CpG methyltransferase M.SssI, demonstrating that SsDddA can deaminate 5mCpG, albeit with reduced activity.

[0103] To determine effective SsDddA reaction conditions for chromatin stenciling, GM12878 nuclei were treated with a range of enzyme concentrations and treatment times, and each condition evaluated using long-read sequencing of targeted amplicons from two genomic loci, the NAPA and WASF1 promoters, that have extensive paired DNase-seq, ATAC-seq, and single-molecule chromatin fiber sequencing (Fiber-seq) data (FIG.1E). Using 1% of a PacBio Revio SMRT cell, these regions were sequenced to 25,672x coverage for NAPA and 46,264x for WASF1. Cytidine deamination was highly specific to orthogonally defined regulatory elements and internucleosomal linker regions3915-P1364WO.UW -33-in a manner that mirrored the paired DNase-seq, ATAC-seq, and Fiber-seq data (FIG.1E). It was observed that treating nuclei with 4 μM SsDddA for 10 minutes provided high quality data, with a median of 82% and 73% deamination rates within accessible portions of the NAPA and WASF1 promoters, respectively, and a median of 2.6% and 2.3% deamination rates within well positioned CTCF and nucleosome footprints within these promoters, respectively (FIG. 1F). Furthermore, this reaction condition demonstrated minimal sequence biases (FIG. 1F) – enabling the near nucleotide-precise mapping of protein occupancy events, especially in cytidine rich CpG islands. In addition, deamination rates within the NAPA promoter were significantly higher than m6A rates from Fiber-seq performed on the same sample (one-sided Fisher’s exact test, P = 1.7 x 10-11), indicating that this reaction condition results in saturated SsDddA activity. Furthermore, it was observed that each molecule’s deamination pattern could be readily used as a unique molecular identifier (UMI) to identify reads arising from PCR duplicates, due to practically countless combinations of C->T mutations that could exist along a 5 kb fiber. Targeted DAF-seq was further benchmarked across 10 loci in human GM12878 cells, K562 cells, and frozen primary human post-mortem descending colon tissue, which showed strong agreement between DAF-seq, Fiber-seq, and ATAC-seq bulked chromatin accessibility measures (FIG.9).

[0104] Overall, it was observed that targeted DAF-seq can result in a 230,000- fold enrichment relative to untargeted genome-wide chromatin stenciling approaches (FIG. 1G), with many sequenced reads constituting unique molecules at a sequencing depth of 100,000.

[0105] Together, these findings demonstrate that recombinant SsDddA can be readily purified using a bacterial system, has minimal sequence bias, is highly catalytically active, and is highly specific for accessible DNA—features that make it well suited for chromatin stenciling with single-molecule and single-nucleotide precision. Example 3. DAF-seq disentangles the regulatory logic of single-molecule transcription factor (TF) co-occupancy

[0106] The single-molecule and single-nucleotide precision of DAF-seq, combined with its high sequencing depth, offers the potential to precisely delineate how individual transcription factors (TFs) occupy and co-occupy a given regulatory element. Furthermore, unlike Tn5-based enrichment strategies, or approaches that rely on harsh bisulfite conversion, targeted DAF-seq enables the single-molecule assessment of protein3915-P1364WO.UW -34-occupancy patterns across large stretches of regulatory DNA. To explore this, 11 high- confidence TF binding elements within the NAPA promoter were identified and their overall per-molecule occupancy quantified (FIGs 2A-2B), demonstrating that the occupancy of each element ranged from 13% (element 6) to 96% (element 11) of fibers (FIG. 2B). It was observed that occupancy at element 2 rarely occurred unless element 1 was also occupied along the same molecule (FIG. 2C), suggestive of a potential cooperative interaction between the proteins that occupy elements 1 and 2.

[0107] To quantify this, a thermodynamic formalism that relates the frequencies of DAF-seq reads representing different TF-bound states to the free energies of the TF- DNA and TF-TF interactions was implemented (FIGs 2D-2F.) (see, for example, Sherman MS, Cohen BA. Thermodynamic state ensemble models of cis-regulation. PLoS Comput Biol.2012;8(3):e1002407. doi: 10.1371 / journal.pcbi.1002407. Epub 2012 Mar 29. PMID: 22479169; PMCID: PMC3315449.). This approach assumes that the NAPA promoter is close to equilibrium and that the reads representing each of the protein-bound states will follow the Boltzmann Distribution, thereby enabling one to calculate the change in free energy (ΔG) of each singly bound, doubly bound, or triply bound TF state relative to the unbound state.

[0108] This demonstrated that the relative binding affinity of the protein-protein interaction between the proteins occupying elements 1 and 2 is 180,000 times stronger than the relative binding affinity of the protein-DNA interaction at element 2, indicating that element 2 is primarily occupied via a cooperative interaction (FIG.2E). In contrast, when elements 1 and 2 are co-occupied, it appears unfavorable for proteins to synchronously occupy elements 4, 5, 7, 8, 9, or 10 (FIG.2F), indicating that occupancy at elements 1 and 2 establishes a protein complex that is not reliant upon occupancy at the other elements within the NAPA promoter.

[0109] To discern whether element 1 or 2 is driving this cooperative interaction, a network graph approach was employed to quantify the conditional codependency between these elements, revealing that element 1 was by far the most essential, supporting a clear directional effect whereby element 1 drives the codependent occupancy at element 2. Of note, whereas element 1 contains a predicted high-affinity element for USF1 & USF2, the CAAT box in element 2 contains a predicted low-affinity sequence element for NFY- A, consistent with distance-dependent cooperative binding between USF1 / 2 and NFY-A, whereby USF1 / 2 binds and recruits NFY-A through its USF Specific Region (USR)3915-P1364WO.UW -35-domain. Favorable cooperative interactions were not limited to elements 1 and 2 (FIG. 2D), and three-way interactions that excluded elements 1 and 2 overall appeared favorable (FIG.2F), suggesting that the NAPA promoter is occupied by two multiprotein complexes, one occupying elements 1 and 2, and a second occupying elements 4, 5, 7, 8, 9, and 10.

[0110] Overall, these findings demonstrate that DAF-seq can accurately quantify single-molecule protein occupancy and co-occupancy patterns, revealing cooperative TF interactions within near single-nucleotide precision. Example 4. Synchronous single-molecule genomic and chromatin profiles using DAF-seq

[0111] Next, the ability of DAF-seq to disentangle SsDddA-induced deaminations from germline genetic variants was evaluated. This evaluation leveraged the fact that SsDddA only modifies one strand of a C / G base pair, and that reads can be readily partitioned into ‘top-strand’ and ‘bottom-strand’ reads relative to the reference, based on their predominance of C->T versus G->A changes, respectively (FIG.1B). Specifically, at a C / G base pair, although the top strand may be variably deaminated and converted to a T, the bottom strand will always remain a G, enabling one to readily distinguish whether a position was originally a C / G or T / A base-pair (FIG. 3A). Consequently, DAF-seq reads mapping to a reference position that has a C / G base-pair on both haplotypes will contain 100% G on the bottom-strand, and between 0 and 100% C on the top strand, with the base content of the top strand reflecting the frequency by which that site is deaminated (FIG. 3A).

[0112] This approach enabled the accurate delineation of germline variants, as well as the haplotype phasing of DAF-seq reads based on these germline variants, including instances where the only heterozygous germline variant in a read is a C / T or G / A variant (i.e., only the bottom or top strand can be accurately phased in these cases, respectively). To validate this, targeted DAF-seq was applied to a 4.4 kb region containing only a single C / T heterozygous variant (rs56269549) in GM12878 cells, which enabled to successfully phase DAF-seq reads from the bottom strand. This targeted region is located on the X chromosome and spans four UBA1 transcriptional start sites (TSSs), three of which are selectively accessible on only one of the haplotypes (i.e., Xa) in GM21878 cells with allelically skewed X chromosome inactivation. This haplotype-resolved DAF-seq data demonstrated that, whereas the canonical UBA1 TSS is accessible on both the Xa and Xi, the three upstream TSSs are selectively accessible on only the Xa, confirming that DAF-3915-P1364WO.UW -36-seq can accurately haplotype phase reads, even when they contain only a single heterozygous C / T variant (FIGs 3B-3C).

[0113] Leveraging this haplotype-phased data, it was evaluated whether these UBA1 TSSs are being independently actuated along the Xa. It was observed that although these four UBA1 TSSs showed strong co-actuation along the Xa (FIG. 3D), the actuation of each of these TSSs was largely occurring independently of each other (FIG. 3E). Furthermore, it was observed that protein occupancy at the major binding elements within the UBA1 isoform 4 TSS that escapes XCI was also largely occurring in an independent manner and did not show appreciable differences between the Xa and Xi (two-sided paired t-test, P = 0.22) (FIG.3F).

[0114] Together, these findings demonstrate that DAF-seq can accurately resolve chromatin patterns in a haplotype-aware manner and that actuation of the upstream UBA1 TSSs is being largely driven in an independent manner along the Xa. Example 5. Resolving chromatin transition states at the SLC39A4 promoter using DAF-seq

[0115] It was next sought to determine whether DAF-seq could resolve distinct chromatin epialleles formed at a regulatory element. To test this, the SLC39A4 promoter was studied, which is known to have haplotype- and cell-selective activity, and for which rare promoter variants can cause acrodermatitis enteropathica.

[0116] Targeted DAF-seq was performed on primary post-mortem frozen liver tissue and a lymphoblastoid cell line from individuals heterozygous for the common rs2280838-T haplotype, which is associated with modestly increased SLC39A4 transcript levels selectively in liver tissue (FIG. 4A). These fibers were sequenced to a depth of ~1,200,000x targeted coverage in each sample, the deduplicated DAF-seq reads from both samples combined, and clustered by their single-molecule deamination patterns (FIG.4B). This exposed distinct single-molecule patterns in nucleosome positioning and chromatin actuation along the SLC39A4 promoter that differed by haplotype and cell type (FIG.4C). Specifically, clusters 4, 5, and 6, which include variably positioned nucleosome arrays, were selective to reads from lymphoblastoid cells. In contrast, clusters 2 and 3, which include nucleosome arrays with exquisitely well-positioned nucleosomes within the SLC39A4 promoter, were represented by reads from both liver and lymphoblastoid cells, with a preference for liver reads from the rs2280838-C haplotype. Notably, only cluster 1 showed actuation of the SLC39A4 promoter, and this cluster was comprised almost3915-P1364WO.UW -37-exclusively of chromatin fibers from liver tissue, with 72% of those reads originating from the rs2280838-T haplotype (FIG. 4C), indicating that the association of rs2280838 with SLC39A4 transcript levels in liver is likely mediated through differences between the two haplotypes in their propensity for forming actuated chromatin at the SLC39A4 promoter.

[0117] Further sub-clustering of the 17% of fibers corresponding to cluster 1 revealed distinct single-molecule patterns of focal chromatin actuation within the SLC39A4 promoter that differed between the rs2280838-C and rs2280838-T haplotypes in liver (FIG. 4D). Specifically, the position immediately above rs2280838 was 2.0-fold more likely (one-sided Fisher’s exact test, P = 2.3 x 10-7) to be focally actuated along the rs2280838-T haplotype as opposed to the rs2280838-C haplotype in liver tissue (FIG.4E). Notably, this position is predominantly occupied by an exquisitely well-positioned nucleosome along non-actuated liver fibers, indicating that rs2280838-T is likely increasing SLC39A4 transcript levels by modulating the propensity of an overlying nucleosome to occlude the SLC39A4 promoter (FIG.4F).

[0118] Overall, these findings demonstrate that DAF-seq can capture rare chromatin epialleles with single-molecule and single-haplotype precision. Example 6. Quantifying the functional impact of non-coding mosaic mutations using DAF-seq

[0119] It was further sought to determine if DAF-seq can accurately capture the chromatin state along rare genetic alleles, as is often encountered when evaluating the genetic and functional impact of somatic variants with a low variant allele fraction (VAF). Single-molecule chromatin assays are well suited for measuring the functional impacts of mosaic variants, as unlike Tn5-based enrichment approaches, single-molecule chromatin fiber sequencing can independently measure VAF and allelic chromatin imbalance.

[0120] To determine if DAF-seq can accurately capture the chromatin state along rare genetic alleles, a cell-based model of somatic variation was implemented, which includes a 49:1 mixture of the B lymphoblast cell line COLO829BL (BL), and a melanoma tumor line COLO829T (T) derived from the same individual (FIG. 5A). Fiber-seq data from these two unmixed cell lines identified a CC>TT somatic dinucleotide mutation on one haplotype of COLO829T (chr17:19,447,245-19,447,246) that ablates an overlying CTCF binding element, causing selective loss of CTCF occupancy and chromatin accessibility on the variant haplotype relative to reference (FIG.5B). Illumina®PCR-free whole-genome sequencing of the 49:1 BL:T mixture identified 6 of 437 reads with the3915-P1364WO.UW -38-CC>TT variant for a variant allele fraction (VAF) of 1.4%. Application of targeted DAF- seq to a 3.8 kb region spanning the variant in the same 49:1 BL:T mixture readily exposed the presence of the variant, and identified 1,701 of 115,991 bottom strand (G-to-A) reads as containing the CC>TT variant, for a VAF of 1.5% (FIG.5A) - demonstrating minimal amplification biases using targeted DAF-seq. Furthermore, comparison of the chromatin architectures between the DAF-seq reads containing the reference or CC>TT variant sequence readily exposed that the variant reads lost CTCF occupancy, chromatin accessibility and nucleosome phasing at the targeted element (FIG.5C).

[0121] Overall, these results demonstrate that DAF-seq can accurately quantify the genetic and chromatin architecture of individual sequencing reads with minimal biases and showcase targeted DAF-seq as a powerful tool for functionally characterizing low VAF somatic variants within human tissues. Example 7. Comprehensive reconstruction of single-cell diploid genomes using single-cell DAF-seq

[0122] Having demonstrated that DAF-seq enables the accurate reconstruction of the genomic and chromatin architecture of individual reads, it was next sought to apply this technology to single cells to comprehensively resolve the gene regulatory architecture of an individual cell. To accomplish this, GM24385 lymphoblastoid cells were used, as a highly accurate diploid genome (HG002) exists for this cell line, enabling the benchmarking of single-cell genomic accuracy. Specifically, permeabilized GM24385 cells were treated with SsDddA and individual cells sorted into sample wells using fluorescence-activated cell sorting (FACS). Whole-genome amplification (WGA) was then performed separately on each cell using primary template-directed amplification (PTA) and the amplification products from 12 cells sequenced with long-read sequencing (FIG.6A). PTA is an isothermal whole genome amplification (WGA) method that reproducibly captures >95% of the genomes of single cells in a uniform and accurate manner, resulting in significantly improved variant calling sensitivity and precision (see, for example, V. Gonzalez-Pena, et al., Accurate genomic variant detection in single cells with primary template-directed amplification, Proc. Natl. Acad. Sci. U.S.A. 118 (24) e2024176118, (2021), which is incorporated by reference herein.). Eight cells were sequenced to a median depth of 12 Gb, two to ~22 Gb, one to 91 GB, and one to 133 Gb (N50 for sequencing fragment length of 4.0 kb) (FIGs 6B-6C).3915-P1364WO.UW -39-

[0123] Importantly, autosomes within diploid cells contain four ‘haplotype- strand’ templates at each genomic position, corresponding to the top and bottom strands of each of the two haplotypes. SsDddA treatment creates a unique deamination pattern along each ‘haplotype-strand’ template, as the precise locations of deaminase-induced mutations are stochastic due to incomplete SsDddA deamination and heterogeneous strand-specific protein occupancy patterns. Furthermore, PTA’s increased preference for priming on the primary template produces numerous partially overlapping amplicons originating from the same ‘haplotype-strand’ template. Consequently, it was reasoned that the unique molecular identifiers (UMIs) created by template-specific deamination events could be used to group and collapse reads arising from the same ‘haplotype-strand’, even in situations where the underlying genomic sequence lacked haplotype-unique variants (FIG. 6A), analogous to generating consensus reads from multiple sequencing passes during circular consensus sequencing (CCS).

[0124] To accomplish this, reads from each cell were mapped to GRCh38 and overlapping reads originating from the same ‘haplotype-strand’ template were collapsed to generate individual ‘consensus reads’ (FIG. 6D). This resulted in a single-cell consensus read N50 that was as high as 34.5 kb for the deepest sequenced cell (FIG.6C), with 5,608 consensus reads from that cell >100 kb in length (FIG. 6B) and a haplotype switch error rate of 2.4-3.3%, which is comparable to current assembly methods with similar coverage.

[0125] Overall, this resulted in an even coverage of consensus reads and readily exposed anomalous autosomal genomic loci harboring >4 ‘haplotype-strand’ templates, indicative of possible duplications present in that cell relative to GRCh38. In total, each cell had at least one read spanning between 60% to 99% of the mappable, autosomal portions of GRCh38 (FIGs 6D-6E) a rate that mirrored the sequencing depth of each cell. Moreover, between 27% and 80% of the mappable, autosomal portions of GRCh38 within a cell were covered by consensus reads that could be assigned to either the paternal or maternal haplotype using parental short-read data (FIG.6E).

[0126] Together, these data demonstrate that scDAF-seq enables the accurate reconstruction of the chromosome-scale haplotype-phased diploid genome from a single cell, with thousands of consensus reads >100,000 bp in length. Example 8. Widespread plasticity in the chromatin epigenome of a single cell

[0127] To date, it has been unknown the degree of plasticity that is permissible within a cell’s chromatin epigenome, as prior single cell chromatin assays rely on Tn5-3915-P1364WO.UW -40-based accessibility measurements, which are sparse and unable to measure the inaccessible chromatin state. However, scDAF-seq enables the genetic evaluation of up to 99% of each cell’s mappable genome, with up to 80% of the genome being haplotype phased within a single cell (FIG. 6). The disclosed scDAF-seq methods can be well suited for comprehensively evaluating chromatin patterns between haplotypes within a single cell, as well as along the same haplotype between cells.

[0128] To evaluate this, the ability of scDAF-seq to qualitatively and quantitatively measure single-molecule chromatin accessibility patterns genome-wide was first benchmarked. Specifically, it was demonstrated that pseudo-bulked scDAF-seq measures of chromatin accessibility are comparable to that of Fiber-seq, scATAC-seq, and ATAC-seq (FIGs 7A-7B). In addition, it was demonstrated that scDAF-seq single- molecule chromatin actuation measurements monotonically mirror single-molecule Fiber- seq chromatin actuation measurements across both accessible regulatory elements (FIG. 7A), as well as different euchromatic and heterochromatin genomic loci. Finally, it was established that scDAF-seq enables single-molecule and single-cell measures of TF occupancy and co-occupancy with near single-nucleotide resolution (FIG.7C). Together, these findings demonstrate that when using haplotype-phased samples, scDAF-seq permits profiling gapped chromatin architectures along single fibers that are up to >200 Mb in length (FIG. 7B), with the length and completeness of each fiber simply limited by the length of the underlying chromosome and library sequencing depth, respectively.

[0129] Bulk Fiber-seq performed on GM24385 cells identified 144,870 actuated regulatory elements within mappable autosomal regions. However, it was observed that, on average, only 46% of these regulatory elements were actuated on at least one haplotype within a single cell, indicating that within most cells, only a minority of regulatory elements are present in the open configuration. To compare these patterns between cells, haplotype- phased GM24385 regulatory elements that had sequencing coverage across a pair of cells subjected to scDAF-seq were identified, and the actuation status at each element between the two cells in a haplotype-aware manner was compared (FIG. 7D). This revealed that, on average, the actuation status of a regulatory element between two individual GM24385 cells differs by ~63% (range 60% to 67%) (FIG.7D). However, individual elements varied quite substantially in their degree of plasticity, with GM24385 regulatory elements that are consistently actuated in bulk Fiber-seq data only differing in their actuation status by ~9% between cells (FIG. 7E). Notably, Jaccard distances were 27% lower for promoter-3915-P1364WO.UW -41-proximal elements than for promoter-distal elements (FIG. 7D), suggesting that a regulatory element’s chromatin plasticity may be related to its function. Consistent with this, it was observed that genes with greater steady state transcriptional outputs had markedly fewer single-cell differences in their promoter actuation status, with the highest expressing genes in GM24385 differing in their actuation status by only ~16% between cells (FIG.7F).

[0130] Finally, to determine whether this plasticity is mediated by cell-to-cell variation in the trans environment, the actuation status of elements between the two haplotypes within the same cell was compared. This revealed that on average, the actuation status of a regulatory element between both haplotypes within an individual GM24385 cell differs by ~61% (range 59-65%) (FIG. 7D), with only a minority of these elements representing sites with consistent haplotype-selective chromatin across GM24385 cells, as measured using bulk Fiber-seq.

[0131] Overall, these findings demonstrate pervasive plasticity within the chromatin actuation status of a single cell’s regulatory landscape, and that this plasticity is directly related to the function of the chromatin epigenome. Furthermore, cell-to-cell differences in the abundance of trans-acting factors plays only a minor role in governing the plasticity in the chromatin epigenome of a cell. Example 9. Codependent single-molecule chromatin actuation is largely limited to ~100kb domains

[0132] 3D genome folding plays an integral role in modulating gene regulatory patterns by bringing pairs of regulatory elements into close proximity, thereby enabling elements to modulate their respective activity along the same chromatin fiber. Consistent with this model, bulk Fiber-seq measurements have shown that individual regulatory elements are preferentially co-actuated along the same chromatin fiber. However, these measurements are limited by the sequencing lengths of Fiber-seq (i.e., 10-100 kb) and are confounded by cell-to-cell differences in the abundance of trans-acting factors, which can result in molecules appearing to have codependent actuation simply because they are present within cells that have divergent levels of trans-acting factors.

[0133] In this example, it was sought to leverage the chromosome-length single- cell and single-molecule measurements of TF occupancy and chromatin actuation to test whether regulatory elements are indeed preferentially actuated along the same chromatin fiber. To validate this approach, CCCTC-Binding factor (CTCF) loop anchors were firstly3915-P1364WO.UW -42-focused on, as these by definition form along the same chromatin fiber, and are mechanistically mediated by CTCF binding sites arranged within a + / - asymmetric orientation (FIG.8A).

[0134] Overall, it was observed that CTCF loop anchors defined using Chromatin Interaction Analysis by Paired-End Tag Sequencing (ChIA-PET) exhibit significantly higher single-molecule CTCF co-occupancy than expected (two-sided t-test, P < 2.2x10-16). However, by leveraging the near single-nucleotide resolution of scDAF-seq it was found that CTCF co-occupancy at loop anchors is largely limited to CTCF binding elements bound in the + / - asymmetric orientation (FIGs 8A-8B), consistent with prior reports. Loop anchors with more 3D genome contacts show higher single-molecule CTCF co-occupancy, with the strongest 5003D genome contacts exhibiting a median CTCF co- occupancy of 71% (FIG. 8C). Overall, this establishes that scDAF-seq can accurately measure single-molecule co-occupancy events related to 3D genome folding.

[0135] To determine whether regulatory elements are preferentially co-actuated along the same chromatin fiber, the expected co-actuation of any two elements was first calculated based on their respective actuation percentages across the 12 cells with scDAF- seq data. It was then quantified whether the observed single-molecule co-actuation at these elements differed from this expected co-actuation (i.e., the ‘codependency score’), with positive codependency scores indicating elements that are preferentially co-actuated along the same molecule (FIG. 8D). To determine whether these codependency scores simply reflected differences in the trans environment between the 12 cells, a ‘pseudo codependency score’ was calculated, which uses the co-actuation status of the two elements on the different haplotypes within the same cell, under the assumption that human cells do not utilize transvection. Controlling for these ‘pseudo codependency scores’, it was observed that genome-wide, regulatory elements indeed are preferentially co-actuated along the same chromatin fiber in a distance dependent manner (FIG. 8E). However, preferential single-molecule co-actuation in GM24385 cells is limited to regulatory elements that are ~100 kb apart or closer (FIG.8E), a distance that mirrors that of cohesin- mediated chromatin loops and is an order of magnitude smaller than that of topologically associating domains.

[0136] Overall, these findings demonstrate that regulatory elements can be preferentially co-actuated along the same chromatin fiber over distances consistent with chromatin loops.3915-P1364WO.UW -43-Example 10. DAF-seq as a transformative method for studying the structure and function of the non-coding genome

[0137] This disclosure provides DAF-seq, a novel and transformative method that is suitable for studying the structure and function of the non-coding genome. DAF-seq is leveraged to reveal the first-ever comprehensive map of the diploid genome and chromatin epigenome from a single cell. As DAF-seq chromatin stencils are maintained upon DNA amplification, these deamination patterns can be used as a Unique Molecular Identifier (UMI) to identify reads arising from the same modified DNA template. This is effective for identifying PCR duplicates and stitching PTA-produced reads together, thereby enabling true single-molecule footprinting across chromosome-length chromatin fibers, a ~10,000-fold improvement beyond what was previously possible with PCR-free chromatin stenciling methods, and a ~1,000,000-fold improvement beyond what was previously possible with Tn5-based chromatin stenciling methods.

[0138] In addition, the non-sequence specific nature of SsDddA with the reaction conditions enables DAF-seq to have near single-nucleotide resolution, a significant improvement over prior methods that rely on enzymes with significant sequence preferences. Furthermore, whereas prior short-read based deaminase stenciling methods are severely limited by biased read mapping issues, pairing high-resolution deaminase stenciling with long-read sequencing as per the present disclosure enables 99.9% of the reads to be mapped back to the genome, including possibly due to unmodified DNA within nucleosome footprints that allow for seeding read mapping.

[0139] As such, scDAF-seq enables the comprehensive mapping of the chromatin epigenome across nearly the entire genome of a single-cell in a haplotype-aware manner. This enables the study of single-fiber co-actuation and protein co-occupancy measurements of genomic elements located >200 Mb apart, a distance and resolution unobtainable with prior technologies. Importantly, although scDAF-seq uses more sequencing per cell than traditional single-cell chromatin assays, it is demonstrated that fundamental principles of gene regulation can be derived from just 12 cells, breaking the previously accepted paradigm that high quality single-cell chromatin analyses require data from millions of cells. Furthermore, it is contemplated herein that further reduction in the cost of long-read sequencing can permit scDAF-seq to capture the diversity of cells within a sample.

[0140] Using DAF-seq, the present disclosure elucidates widespread heterogeneity in the primary chromatin architecture of individual regulatory elements at3915-P1364WO.UW -44-the single-molecule and single-cell level. Specifically, critical chromatin transition states that illuminate TF and regulatory element cooperativity patterns (FIGs 2 and 3) are observed, patterns that, to-date, have largely necessitated time-consuming genome and / or epigenome editing methods to resolve. Furthermore, as DAF-seq does not require Tn5- based enrichment, it can accurately measure diverse chromatin states, including quantifying regulatory elements present within a closed state. These features enable a person having ordinary skill in the art to comprehensively measure the accessible chromatin epigenome of a single cell, for example, exposing that the accessible chromatin landscape of individual GM24385 cells can diverge by ~63% within a cell line. As scDAF- seq measures the number of haplotype-strands at each location along the genome, per the examples of this disclosure, one can confidently state that all twelve of the cells assayed were in the G1 phase of the cell cycle. These findings indicate that GM24385 cells are tolerant of widespread plasticity in their accessible chromatin epigenome. Furthermore, one can leverage the single-cell haplotype-resolved data to directly account for the impact of the cellular trans environment in this finding, showing that cell-to-cell differences in the cellular trans environment only modestly contribute to this heterogeneity. Together, these results enable the further study of the dynamics underlying this heterogeneity, including how a cell stores an epigenetic memory that a specific regulatory element should become accessible within a population of GM24385 cells.

[0141] DAF-seq significantly improves how the functional impact of somatic variants are studied across a range of disciplines. Specifically, sequencing reads arising from the same DNA template can be readily identified and leveraged to construct a highly accurate consensus sequence of the primary template, analogous to duplex sequencing methods. In addition, DAF-seq can accurately distinguish SsDddA-induced deamination events from germline and somatic variants, permitting DAF-seq to accurately identify and quantify low VAF somatic variants using an economical experimental design (i.e., genomic PCR followed by amplicon sequencing using only a small fraction of a sequencing run per target). Furthermore, as DAF-seq is compatible with long-read sequencing, it can be leveraged to interrogate somatic variants in complex genomic regions. However, unlike ddPCR or other amplicon-based approaches, DAF-seq enables the simultaneous quantification of somatic genetic variants and their functional impact on chromatin patterns with improved resolution and throughput relative to existing targeted methods or genome- wide methods. This would not have been possible with prior Tn5-based enrichment3915-P1364WO.UW -45-strategies, as the functional impact of a variant on chromatin accessibility can disrupt the ability to directly quantify that variant’s VAF using the same sequencing reads.

[0142] It is demonstrated herein that DAF-seq can be performed on frozen primary human tissue as well as cultured cells, and that DAF-seq chromatin stencils are compatible with all current short- and long-read sequencing platforms. These results position DAF-seq as a transformative experimental tool for resolving the functional impact of the millions of genetic variants associated with both common and rare disease risk.

[0143] In addition, it is demonstrated herein that the high sequencing accuracy of Single Molecule Real-Time (SMRT) DNA Sequencing (e.g., PacBio®HiFi®sequencing), coupled with the high fidelity and primary template preference provided by PTA, enables scDAF-seq to generate highly accurate single-cell consensus sequences from each of the four ‘haplotype-strand’ template within a single diploid cell. The process of generating consensus template sequences dramatically improves the read coverage within a single cell, readily exposing genomic locations with >4 ‘haplotype-strand’ templates that correspond to duplications in that cell relative to the reference.

[0144] It is further shown herein that scDAF-seq enables the generation of thousands of ultra-long consensus reads from a single cell, and the generation of these ultra- long consensus reads can be readily increased by simply sequencing the library from each cell to higher depths. These single-cell ultra-long consensus reads enable the evaluation of single-cell genomic and epigenomic variation within the most complex regions of the genome and, significantly, lay the groundwork for assembling complete telomere-to- telomere genomes and chromatin epigenomes from single cells. Example 11. Methods Bacterial strains and culture conditions

[0145] All bacterial strains used in this disclosure were grown in Lysogeny Broth (LB) at 37 °C or on LB medium solidified with agar (RPI, cat# L24030-100.0). Filter sterilized kanamycin (Gold Biotechnology®, cat# K-120) (100 mg / L for plasmid propagation, or 30 mg / L for protein expression), and IPTG (ThermoFisher®, cat# R0393) (0.5 mM) were added to culture when necessary. E. coli strains DH5α (NEB®, cat# C2987H) and BL21(DE3) (NEB®, cat# C2527H) were used for cloning and producing plasmids, and protein expression respectively. Cloning and purification of SsDddA3915-P1364WO.UW -46-

[0146] The genes for SsDddA WT, SsDddA5, and corresponding immunity protein (SsDddI), which is required for SsDddA purification, were codon optimized for E. coli expression and synthesized as gBlocks with corresponding restriction enzyme recognition sites flanking each end by IDT. The SsDddI was inserted between NdeI and XhoI, and WT SsDddA or SsDddA5 was inserted between NcoI and NotI with a N-terminal 6xHis tag of the vector pColADuet-1 (LifescienceMarket®, #PVT0105). The deaminases were cloned into the vector after the immunity protein was successfully cloned into the vector. The whole plasmid sequence was confirmed by PlasmidSaurus.

[0147] The purification of deaminases was performed as previously described with the following modifications: protein expression was induced at 16 ˚C overnight, cells were resuspended in Ni-NTA Buffer A (50 mM Tris-HCl, pH 7.8, 600 mL NaCl, 10% Glycerol, 10 mM 2-Mercaptoethanol, 0.1% Triton-100) with protease inhibitor cocktail (Thermo Scientific®, cat#PIA32955). The cell lysate was loaded on a HisTrap®HP His tag protein purification column (5 mL, Cytiva®# 17524801) and the deaminase-immunity protein complex was eluted during a gradient from 100% Ni-NTA Buffer A to 100% Ni- NTA Buffer B (50 mM Tris-HCl, pH 8, 600 mL NaCl, 500 mM imidazole, 10% Glycerol) using the NGC Quest®10 Plus Chromatography System®(Bio-Rad®, #7880003) and the fractions of corresponding A280 peaks were pooled. The major peak eluted during the gradient was pooled, verified using SDS-PAGE, and collected for the denaturing and renaturing steps to separate the immunity protein and the deaminase. The pooled protein complex samples were added to denaturing buffer (50 mM Tris-HCl pH 7.8, 20 mM imidazole, 500 mM NaCl, 6 M guanidine HCl, prepared from Guanidine-HCl (ThermoFisher®, cat# 24110), and 5 mM 2-Mercaptoethanol) at 1:10 (v:v) ratio and incubated overnight. The mixture was then loaded back to the 5 mL HisTrap column, with 50 mL of denaturing buffer. The deaminase was renatured during a gradient from 100% denaturing buffer to 100% renaturing buffer (50 mM Tris-HCl pH 7.8, 500 mM NaCl, 10 µM ZnCl2, and 10 mM 2-Mercaptoethanol) at 1 ml / min and washed with additional 50 mL renaturing buffer. The renatured deaminase was further purified using HiLoad®Superdex®200 pg preparative SEC column (120 mL, Cytiva®# 28989335) with storage buffer (50 mM Tris-HCl pH 7.8, 500 mM NaCl, 10 µM ZnCl2, 1 mM DTT, and 10% Glycerol). The fractions were evaluated by SDS-PAGE, and the purest fractions were pooled, aliquot into 20 µL stocks, flash frozen with liquid nitrogen, and stored at -80 ˚C. All Tris-based3915-P1364WO.UW -47-buffers were prepared from 1 M Tris-HCl (pH8, molecular biology grade ultrapure, ThermoScientific®, cat# J22638-K2). Mass Spectrometry validation of SsDddA activity

[0148] SsDddA activity was validated via quantification of deoxycytidine deamination by UHPLC-MS / MS. Samples for quantification were treated as previously described with modifications. In brief, 2 ng / mL of stable isotope-labeled 2-deoxycytidine triphosphate (MilliporeSigma®, cat# 646229) and 2-deoxyadenosine triphosphate (MilliporeSigma®, cat# 646237) were used as references and added to 50 ng of DNA from each sample before any mass spectrometry sample preparation. The DNA from each sample was mixed with 0.02 U phosphodiesterase I (Worthington®, cat# LS003926), 1 U Benzonase (MilliporeSigma®, cat# E1014), and 2 U Quick CIP (NEB®, cat# M0525S) in digestion buffer (10 mM Tris, 1 mM MgCl, pH 8 at RT) for 3 hours at 37˚C, with a total reaction volume of 50 µL. Single nucleotides were separated from the enzymes by collecting the flow-through of a Nanosep®centrifugal filter (MWCO 3 kDa, Pall®, cat# OD003C33). The UHPLC-MS / MS analysis of cytosine and adenosine was performed on an ACQUITY®Premier®UPLC System coupled with a XEVO-TQ-XS triple quadrupole mass spectrometer. UPLC was performed on a ZORBAX®Eclipse Plus®C18 column (2.1 × 50 mm I.D., 1.8 μm particle size) (Agilent®, cat# 959757-902) using solvent A consisting of 0.1% ammonium hydroxide in 100% acetonitrile (v / v) and solvent B consisting of 0.1 M ammonium acetate in water, with the following gradient at 0.3 mL / min: 0-1 min 100% A, 1-6 min 100-30% A and 0-70% B, 6-7 min 30-5% A and 70-95% B, 7-8 min 5-100% A and 95-0% B, 8-10 min 100% A. MS / MS analysis was operated in positive ionization mode with 3000 V capillary voltage as well as 350°C and 1000 L / hour nitrogen drying gas. A multiple reaction monitoring (MRM) mode was adopted with the following m / z transition: 227.9 -> 94.82, 227.9 -> 98.98, 227.9 -> 111.98, 227.9 -> 116.99 for dC (collision energy, 32, 18, 6, 12 eV respectively); 238.9 -> 100.8, and 238.9 -> 118.9 for isotope-labeled deoxycytidine (collision energy 32 and 6 eV respectively); 252.10 -> 136.09 for dA (collision energy, 14 eV), 267.1 -> 146.1 for isotope-labeled dA (collision energy, 14 eV) was monitored as well as control. MassLynX®was used to quantify the data. All reagents used for mass spectrometer analysis are molecular grade level or above. Bulk Whole-Genome Amplification (WGA) DAF-seq

[0149] 2 million K562, HG002, or GM12878 cells were permeabilized as previously described with the difference of permeabilized cells being resuspended in3915-P1364WO.UW -48-freshly prepared Buffer C (15 mM Tris, pH 8.0; 15 mM NaCl; 60 mM KCl; 1mM EDTA, pH 8.0; 0.5 mM EGTA, pH 8.0; 0.5 mM Spermidine, 10 nM ZnCl2). The permeabilizedcells were treated with 0.25 µM WT SsDddA or SsDddA5 at 25̊ C for 10 min. The reactionwas quenched with 5% SDS (ThermoFisher®, cat# AM9820) before gDNA extraction using HMW DNA Extraction kit (Promega®, cat# A2920). The gDNA was then subjected to whole genome amplification with REPLI-G®Mini®kit (Qiagen®, cat#150023) according to the manufacturer’s protocol before being prepared for sequencing. Nuclei isolation

[0150] GM12878, COLO829BL, and COLO829T cell lines were permeabilized using a digitonin containing isotonic buffer. Briefly, 800,000-1,000,000 cells per sample were added to a 1.5 mL tube (Eppendorf®, 022363204) and centrifuged at 400g for 5 min at 4°C. The supernatant was removed, and the cell pellet was resuspended in 100 µL of chilled isotonic Perm Buffer (20 mM Tris-HCl pH 7.4, 150 mM NaCl, 3 mM MgCl2, 0.05% digitonin) by pipette-mixing 10 times. Cells were incubated on ice for 5 min, after which they were diluted with 1 mL of isotonic Wash Buffer (20 mM Tris-HCl pH 7.4, 150 mM NaCl, 3 mM MgCl2, 10 nM ZnCl2) by pipette-mixing five times. Cells were centrifuged at 400g for 5 min at 4°C and the supernatant was removed. The cell pellet was resuspended in chilled isotonic Wash Buffer. Cells were counted using a Cellometer®Spectrum®Cell Counter (Nexcelom®) using ViaStain®acridine orange / propidium iodide solution (Nexcelom®, C52-0106-5). Tissue Preparation

[0151] Liver, heart, and colon tissue were separately homogenized using a Dounce homogenizer with ten strokes using pestle A, followed by ten strokes using pestle B.1.5 mL of cold homogenization buffer was added to the sample, which was then filtered through a 70-micron filter. Cells were counted using a Cellometer®Spectrum Cell Counter (Nexcelom®) using ViaStain®acridine orange / propidium iodide solution (Nexcelom®, C52-0106-5). 200,000 cells were added to a 1.5 mL tube (Eppendorf®, 022363204) and centrifuged at 500g for 5 min at 4°C. The supernatant was removed and the cell pellet was resuspended in 60 µL of chilled Buffer A w / ZnCl2(15 mM Tris, pH 8.0; 15 mM NaCl; 60 mM KCl; 1mM EDTA, pH 8.0; 0.5 mM EGTA, pH 8.0; 0.5 mM Spermidine, 10 nM ZnCl2) by pipette-mixing, followed by addition of 60 ul of 2X lysis buffer (0.025% IGEPAL, 15 mM Tris, pH 8.0; 15 mM NaCl; 60 mM KCl; 1mM EDTA, pH 8.0; 0.5 mM EGTA, pH 8.0; 0.5 mM Spermidine, 10 nM ZnCl2). Cells were incubated on ice for 10 minutes and3915-P1364WO.UW -49-nuclei were counted using a Cellometer®Spectrum Cell Counter (Nexcelom®) using ViaStain®acridine orange / propidium iodide solution (Nexcelom®, C52-0106-5). ~100,000 nuclei were used as input to the SsDddA reaction. DddA Reactions

[0152] All targeted DAF-seq reactions utilized SsDddA WT. Enzyme optimization experiments targeted NAPA and WASF1 and treated 100,000 GM12878 permeabilized cells at 25°C for 10 minutes and 20 minutes, at enzyme concentrations of 0.25 μM, 1 μM, and 4 μM. All subsequent SsDddA reactions were performed at 25°C for 10 minutes with a 4 μM enzyme concentration. Reactions were neutralized by the addition of 20% sodium dodecyl sulfate (SDS) to a final concentration of 5%. Genomic DNA was extracted using the Monarch®Genomic DNA Purification Kit (New England Biolabs®, T3010S) and quantified using a Qubit®1X dsDNA High-Sensitivity kit (Invitrogen®, Q33231). PCR Amplification of SsDddA-treated genomic DNA

[0153] Regions-of-interest were amplified from SsDddA-treated genomic DNA in 50 ul reactions consisting of 25 μL LongAmp®Hot Start®Taq 2X Master Mix (New England Biolabs®, M0533L), 2 μL each of 10 μM forward and reverse primers, and 30- 100 ng of gDNA. PCR conditions: 94°C for 60s, 35 cycles of 94°C for 30s, 30s annealing, and 65°C extension, followed by a final extension of 65°C for 10 min. Annealing temperatures and extension times varied by target. Amplicons were purified using the Monarch®PCR & DNA Cleanup Kit (New England Biolabs®, T1030L). In general, primers are designed by targeting AT rich regions and avoiding being within 500 bp of an actuated regulatory element. A pool of primers was used for amplification, which included primers with the random incorporation of A or G at C-complementary positions. In general, primers were screened to avoid common polymorphisms and repeat elements. Primer pairs were tested for amplification of a single band of the predicted size after gel electrophoresis, using gDNA from DAF-seq treated nuclei. After sequencing, primer pairs were evaluated for biases in terms of their amplification of ‘top-strand’ or ‘bottom-strand’ fibers, and only those pairs with comparable amplification of both strands were used for downstream applications. Library preparation and sequencing

[0154] Libraries were prepared as previously described. Multiplexed library preparation was performed using the SMRTbell®prep kit 3.0 (PacBio®, cat#102-141-700)3915-P1364WO.UW -50-and SMRTbell®barcoded adapter plate 3.0 (PacBio®, cat#102-009-200). Final sequencing libraries were sequenced on the PacBio®Revio®platform using v3.2 chemistry. DAF-seq alignment and preprocessing

[0155] PacBio®HiFi®reads were converted to FastQ format using samtools fastq (v1.17, parameters: -T) and aligned to hg38 using minimap2 (v2.22-r1101, parameters: --MD -Y -y -a -x map-pb). C-to-T and G-to-A changes were identified by comparing the sequencing read and the hg38 reference using pysam (v0.21.0). Secondary and supplementary alignments were filtered out. The designation of the original SsDddA modified DNA strand as “top” or “bottom” was identified by quantifying C-to-T and G-to- A changes. Reads with at least 90% of these changes being either C-to-T or G-to-A were classified as “CT” or “GA”, respectively, and retained for subsequent analyses. This cutoff allowed to accurately assign read strands in the presence of germline or somatic variants. C-to-T and G-to-A changes were converted to the IUPAC DNA ambiguity codes Y and R in CT and GA reads, respectively. Modified reads were realigned to hg38 using the same parameters. Identification of modification sensitive patches (MSPs) and the generation of nucleosome and accessibility pileups (e.g., chromatin accessibility maps) was done using the fibertools commands ddda-to-m6a (v0.6.4, default parameters) and add-nucleosomes (v0.6.4, parameters: -n 60, -c 70, --min-distance-added 15, -d 10). MSPs > 150bp in length were identified as actuated elements. DAF-seq Target Enrichment

[0156] The proportion of total primary read alignments that mapped within the target region was calculated for both targeted DAF-seq and Fiber-seq data from the same cell line (GM12878). DAF-seq target enrichment was calculated as the proportion of on- target reads in DAF-seq over Fiber-seq. Targeted DAF-seq multi-region benchmarking

[0157] Targeted DAF-seq was performed as described above. Purified PCR products were sequenced using the Oxford®Nanopore®platform. Reads were aligned to hg38 using minimap2 with the “map-ont” preset. Deamination events and read strands were identified as described above. For each library and strand, outlier reads with deamination proportions above or below 1.5 times the interquartile range were filtered out. MSPs were identified using fibertools add-nucleosomes (v0.6.4, parameters: -n 55, -c 65, --min- distance-added 5). PCR Duplicate Identification3915-P1364WO.UW -51-

[0158] DAF-seq PCR duplicate reads were identified by comparing the deamination status of every position susceptible to deamination (each C on top strand reads, each G on bottom strand reads). This combination of deamination statuses was treated as a unique identifier and all reads sharing this identifier were grouped. One read from each group was randomly selected to be marked as unique while the remaining reads in each group were marked as duplicates. Unique GM12878 and Liver tissue reads identified in this manner were randomly selected for SLC39A4 analyses (FIGs 4A-4F). For comparison, duplicate reads were separately identified using the PacBio®tool pbmarkdup which grouped reads with 98% sequence identity at the first and last 500 bp (v1.0.2, default parameters). CpG Methylation Analysis

[0159] NAPA and UBA1 regions were amplified for 30 PCR cycles as described above using 70 ng of HEK293 genomic DNA as template. 1 μg of amplicons from each target was treated with M.SssI CpG methyltransferase (New England Biolabs®, M0226S) per the vendor’s recommendations.600 ng of M.SssI treated and untreated amplicons were treated with SsDddA for 10 minutes at 25°C. DNA treated with SsDddA and M.SssI, and M.SssI only (5mC negative control) were re-amplified for 20 PCR cycles. These two samples along with M.SssI treated DNA (5mC positive control) were sequenced using PacBio®HiFi®as described above. Amplicons were purified after each PCR using the Monarch®PCR & DNA Cleanup Kit (New England Biolabs®, T1030L).

[0160] Sequencing reads from each target were aligned to reference FASTA files containing only the target region. Primary alignments from the M.SssI + SsDddA treatment beginning and ending within 100bp of the region boundaries were used in analysis. Genomic positions were grouped into CpG and non-CpG cytidine (or guanidine for bottom strand) categories, and the deamination occurrences at each position within each read were aggregated. Positions within 28bp of the target region boundaries overlapped priming sites and were omitted from analysis. Positions CpG methylation of the 5mC positive control was quantified using pb-CpG-tools (aligned_bam_to_cpg_scores, v2.3.1). Deamination Motif Analysis

[0161] Sequence bias in SsDddA activity was evaluated for GM12878 WGA realigned data using pysam (v0.21.0), considering only primary alignments. For each deaminated base, denoted in the read sequence by the ambiguity codes Y and R, the seven base (7mer) reference sequence centered on the deaminated base was tracked. Sequences3915-P1364WO.UW -52-originating from GA reads were reverse-complemented to orient the 7mer to the cytidine context. All 7mers were combined to generate a position weight matrix (PWM) which was used to generate a sequence logo (Logomaker v0.8). Transcription Factor Footprinting and Codependency

[0162] DAF-seq footprint density was calculated as the density of all single- molecule stretches of three consecutive non-deaminated cytidines within the NAPA promoter. Examples of this disclosure used FIMO to scan the hg38 reference sequence of an applicable region for transcription factor motifs contained within the JASPAR database and filtered the results for motifs with a q-value <= 0.05. The examples then used Fibertools footprint to identify single-fiber TF footprints for each motif. Motifs were filtered to those footprinted on >= 5% of fibers on both top (CT) and bottom (GA) strands. Footprinted motifs that overlapped either motif by 80% were merged using Bedops (v2.4.41), and the resulting elements that overlapped by 90% were merged to combine elements encompassed entirely within a larger element. Occupancy within each merged element was quantified as the proportion of fibers footprinted at a motif contained within the merged element. TF co- occupancy and codependency were calculated as described previously. Briefly, for each pair of footprinted elements, the expected co-occupancy was calculated as the product of their proportion of fibers bound, while the observed co-occupancy was calculated as the proportion of fibers with an accessible element (MSP > 150 bp) spanning both elements and bound at each element. The essentiality of each element was quantified by constructing codependency graphs, with nodes representing TF elements and edge weights representing codependency scores. Graphs were constructed that omitted individual elements and limited each analysis to fibers accessible but unbound at that respective element. A baseline codependency graph containing all elements was also constructed. The total codependency of each graph was then quantified by summing all edge weights and normalizing by the number of edges. Finally, essentiality was calculated as the ratio of total codependency of the baseline graph over element-excluded graphs, with a higher ratio indicating a larger reduction in codependency following the loss of TF binding at that element. Thermodynamic analysis of the NAPA promoter

[0163] A thermodynamic formalism that relates the frequencies of DAF-seq reads representing the different TF-bound states of the NAPA promoter to the free energies of the TF-DNA and TF-TF interactions that occur on the NAPA promoter was implemented. It was assumed that the NAPA promoter is close to equilibrium and that3915-P1364WO.UW -53-therefore the reads representing each of the protein-bound states will follow the Boltzmann Distribution. The unbound state, the state in which none of the footprints are occupied, was used as the reference state in all analyses. Thus, in these analyses each ΔG is a unitless value that relates the frequency of a protein-bound state relative to the state in which no proteins are bound to NAPA. Because footprint 11 (the CTCF site) was occupied more than 99% of the time, this footprint was excluded from the analyses because one cannot compute a reliable value for its interaction with DNA, nor would one be able to detect interactions between CTCF and the other footprints. The custom Python script written for this analysis is available on GitHub. Analysis of individual TF-DNA interactions

[0164] First, the change in free energy of each singly bound TF state relative to the unbound state was computed. The unbound state, the state in which none of the footprints are occupied, was used as the reference state for all analyses. The Boltzmann Factor relates the relative probabilities of a TF bound state and the unbound reference state to the change in free energy between the states as, ^^(^^^^^ ^^^^^^^^^^ ^^^^^^^^^^) ௱ீ^^^(^^^ ^ ^ ^ ^ ^ ^ = ^^ିோ்^ ^ ^ ^ ^ ^ ^^^^^^^^^)

[0165] In all analyses the RT term was ignored because it factors out of the analyses. The ratio of the probability of a singly bound TF state to the unbound state can be computed directly from DAF-seq data as the ratio of the read counts representing each state, ^^(^^^^^ ^^^^^^^^^^ ^^^^^^^^^^) ^^^^^^^^ ^^^^^^^^^^ ^^^^ ^^^^ ^^^^^^^^^^ ^^^^^^^^^^^^ ^ ^ ^ ^ ^ ^ ^ ^ ^ ^ = ^( ^ ^ ^ ^ ^ ^ ^ ^ ^ ^^^^^) ^^^^^^^^ ^^^^^^^^^^ ^^^^ ^^^^^^^^^^^^^^ ^^^^^^^^^^

[0166] The change in free energy associated with each singly bound TF state is then computed as, ^^^^^^^^ ^^^^^^^^^^ ^^^^ ^^^^ ^^^^^^^^^^ ^^^^^^^^^^^^^^^ = − ln ^ ^Analysis of two-way and three-way TF interactions3915-P1364WO.UW -54-

[0167] The ΔGij values were calculated for each pair wise interaction between bound TFs on the Napa promoter using ^^(^^^^^^ ^^^^^^^^^^ ^^^^^^^^^^) ^^^^^^^^ ^^^^^^^^^^ ^^^^ ^^^^ ^^^^^^^^^^ ^^^^^^^^^^^^ ^ ^ ^ ^ ^ = ^^( ^ ^ ^ ^ ^^^^^ ^^^^^^^^^^) ^^^^^^^^ ^^^^^^^^^^ ^^^^ ^^^^^^^^^^^^^^ ^^^^^^^^^^

[0168] The counts of the TFijbound state and the previously computed values of ΔGiand ΔGias, ^^^^^^^^ ^^^^^^^^^^ ^^^^ ^^^^ ^^^^^^^^^^ ^^^^^^^^^^^^^^^^ = − ln ^ ^^^^^^^^^^ ^^^^^^^^^^ ^^^^ ^^^^^^^^^^^^^^ ^^^^^^^^^^ ^ −^^^^^ − ^^^^^

[0169] were as, ^^^^^^^^ ^^^^^^^^^^ ^^^^ ^^^^ ^^^^^^^^^^ ^^^^^^^^^^^^^^^^^ = − ln^ ^^^^^^^^^^^ ^^^^^^^^^^ ^^^^ ^^^^^^^^^^^^^^ ^^^^^^^^^^ ^ −^^^^^ − ^^^^^ − ^^^^^

[0170] To test for the significance of the ΔGij values it was tested whether the observed read count of the TFijbound state was significantly higher or lower than the expected read count of the TFij bound state assuming that ΔGij=0. To perform this test, the Binomial Distribution was used, where ^^^^ = ^^^^^^^^ ^^^^^^^^^^ ^^^^ ^^^^^^^^^^^^^^ ^^^^^^^^^^ ∗ ^^ି൫௱ீ^ା௱ீೕ൯^^

[0171] Because n is large, the Normal Distribution was used as an approximation for the Binomial, where ^^^^^^^^^^^^^^^^ ^^^^^^^^^^^^ − ^^^^^^ = ^^3915-P1364WO.UW -55-

[0172] A Bonferroni-corrected P-value threshold for the z-scores corrected for the number of individual ΔGijvalues being tested (i.e., 45). Likewise, for testing the significance of ΔGijk values two-sided binomial tests were performed, assuming ΔGijk=0, where ^^^^ = ^^^^^^^^ ^^^^^^^^^^ ^^^^ ^^^^^^^^^^^^^^ ^^^^^^^^^^ ∗ ^^ି൫௱ீ^ା௱ீೕା௱ீೖା௱ீ^ೕା௱ீ^ೖା௱ீೕೖ൯^^ = + ^^^^^^^^^^^^^ + ^^^^^^^^^^^^^ + ^^^^^^^^^^^^^ + ^^^^^^^^^^^^^^

[0173] A Bonferroni-corrected P-value threshold corrected for the number of individual ΔGijk values being tested (i.e., 120) was used. Identification of heterozygous positions

[0174] Top and bottom strand base calls were counted at each genomic position within each target region for each primary read alignment. Positions within 25bp of the target region boundaries overlapped priming sites and were omitted from analysis. Top strand and bottom strand base call proportions were plotted in R using the ggplot2 hexbin package (parameters: bins = 30). UBA1 Transcription Start Site Co-actuation and Codependency

[0175] UBA1 bottom strand reads were assigned haplotypes according to base calls at chrX:47,194,331 (rs56269549). The four UBA1 transcriptional start sites (TSSs) were identified previously using full-length transcript data. MSPs > 150bp long were intersected with each TSS using Bedops (bedmap, parameters: --ec --fraction-ref 0.8 --echo --echo-map-id, v2.4.41). TSS co-actuation was calculated as the proportion of fibers with actuated elements overlapping both TSSs. Codependency was calculated as above using TSS co-actuation instead of TF co-occupancy. DAF-seq Clustering

[0176] Liver and GM12878 reads were assigned haplotypes according to base calls at chr8:144,416,180 (rs2280839). Reads from each sample with a Hamming distance of 3 or less were identified as duplicates and one read from each group was randomly selected as a unique read. The deamination status at each applicable genomic position was recorded and used as a feature for clustering DAF-seq reads. Clustering analysis was performed using Scanpy (v1.10.3). Briefly, 5000 unique reads from each haplotype of each tissue were randomly selected. A neighborhood graph was computed (pp.neighbors,3915-P1364WO.UW -56-parameters: n_pcs=0, n_neighbors=200) and visualized with Uniform Manifold Approximation and Projection (UMAP). Clustering was performed using the Leiden algorithm (tl.leiden, parameters: flavor=“igraph”). Clusters containing fewer than 1000 fibers were removed. Fibers from cluster 1 were re-clustered using the same parameters. Sub-clusters containing fewer than 200 fibers were removed. COLO829 cell mixture

[0177] Pure populations of COLO829BL cells (cat# CRL-1980, lot# 70022927) and COLO829 (referred to as COLO829T) cells (cat# CRL-1974, lot# 70024393) were obtained from ATCC and expanded. All cells were grown at 37°C with 5% CO2. COLO829BL lymphoblastic suspension cells were grown in RPMI-1640 (Fisher®cat# 11875093) with 10% fetal bovine serum (FBS, Fisher®cat# 10082147) shaking at 90 rpm. COLO829T adherent melanoma cells were grown in RPMI-164010% FBS. Both cell lines were expanded and cryopreserved in freezing media consisting of RPMI-164010% FBS with 10% DMSO (Sigma Aldrich®cat# D2650-100ML) at a cell density of 3 million viable cells per mL. Cryopreserved COLO828T and COLO829BL cells were thawed and mixed in a 1:49 ratio using viable cell counts based on Countess cell counter readings (Invitrogen®Thermo Fisher®) with trypan blue (Fisher®cat# 15250061). Cells were aliquoted into cryovials with constant swirling to maintain the homogeneity of the mix. Mixed cells were cryopreserved at a cell density of 2.55 million viable cells per mL. Single-cell Fluorescence Activated Cell Sorting (FACS)

[0178] DddA-treated cells were sorted on a BD FACSAriaII®using sequential gating on forward-scatter (FSC) and side-scatter (FSC) to detect doublets and debris. Individual cells were sorted into wells of a 96-well plate containing 3 ul of Cell Buffer (BioSkryb Genomics®, 100183). Cells were immediately flash frozen on dry ice. Single-cell PTA Whole Genome Amplification

[0179] After dispensation of individual nuclei into independent plate wells, the ResolveServices®(SM) team performed a custom protocol derived from the Services Custom PacBio Long Read Amplification (BioSkryb Genomics®101157) assay. Briefly, the ResolveDNA®workflow was customized to increase amplicon size. After 2.5 hours of isothermal amplification, samples were quantified using qubit HS DNA (Thermofisher®Q33231) and processed for library preparation and sequencing. Single-cell read collapsing3915-P1364WO.UW -57-

[0180] PacBio®HiFi®reads were converted to FastQ format using Samtools fastq (v1.17, parameters: -T) and aligned to hg38 using minimap2 (v2.22-r1101, parameters: --MD -Y -y -a -x map-pb). Soft-clipped bases were trimmed from the alignments to remove chimeric portions of each read. The original template strand of clipped reads was identified as “top” or “bottom” strand by quantifying the proportion of C-to-T and G-to-A changes relative to the hg38 reference using pysam (v0.21.0). Primary alignments with at least 90% of these changes were classified as “CT” or “GA” and retained for subsequent analyses.

[0181] The hg38 reference genome was partitioned into overlapping 150 bp bins with a 25 bp sliding window. For each bin, the sequence Hamming distance between reads that completely spanned that bin was computed, ignoring insertion and deletions. Similar reads were identified as those sharing a minimum of 11 bins (400 bp in total) and with >= 99% identical sequence within >= 80% of shared bins. Similar reads were sequentially grouped together such that each read shared similarity with at least one other read within the group. All reads within each group were required to meet the sequence similarity requirements with each other (i.e., reads in disagreement with >= 1 read within the group were grouped separately).

[0182] Consensus sequences were generated for each group as follows: for each reference position represented within the read group, the composition of base calls at that position was quantified. In cases with disagreement between reads, bases comprising >= 50% of all base calls were used, except in cases where the two most common bases are C & T or G & A and comprised >= 50% of all base calls, in which case C or G was used to overcome putative spurious post-amplification deamination. Deletions and insertions were ignored unless they were present in all reads at that position. When insertions were present in every read the shortest insertion was used. Consensus sequences were subjected to a second round of merging in which overlapping consensus sequences could be merged if they overlapped by a minimum of 7 bins (300 bp in total) and had >= 99% identical sequence within >= 80% of shared bins. In cases of base-level disagreements, the base call of the consensus sequence with the most raw reads in the applicable bin was chosen. Consensus reads were assigned a random identifier and converted to FASTA format. Consensus reads were realigned to hg38 as above and C-to-T and G-to-A changes were converted to the IUPAC DNA ambiguity codes Y and R in CT and GA reads, respectively. Consensus read phasing and coverage quantification3915-P1364WO.UW -58-

[0183] Consensus reads were assigned haplotypes with Whatshap haplotag (v2.3, parameters: --ignore-read-groups, --output-haplotag-list) using previously phased parental variants. The mappable hg38 genome was computed as regions not overlapping the ENCODE blacklist (accession ENCFF356LFX) or hg38 regions with unreliable coverage in HG002 Fiber-seq data. Regions of greater than 400bp contiguous bases that consisted of more than 1 consensus read per haplotype strand in at least one cell were identified using samtools depth (v1.17) and excluded from the hg38 mappable genome. Coverage calculations were performed using Samtools depth in combination with Bedtools (v2.31.0) and Bedops (v2.4.41). Read length statistics were calculated in python using custom scripts (see Zenodo entry). Single-cell chromatin actuation

[0184] Fibertools add-nucleosomes (v0.5.4, parameters: -n 60, -c 70, --min- distance-added 10) was used to calculate deamination autocorrelation and identify modification-sensitive patches (MSPs) within consensus reads. HG002 regulatory elements were identified previously from Fiber-seq data using the FIRE pipeline. scDAF- seq regulatory elements were classified as actuated if they were overlapped by an MSP of >150bp on either strand. These overlaps were required to span at least 50% of either the length of the MSP or the length of the FIRE peak. Read lengths, deamination rates, Jaccard distances, and percent actuation were computed in python using custom scripts (see Zenodo entry). FIRE peaks within 250 bp of an Ensembl canonical transcript TSS in the Gencode 45 release were classified as promoter-proximal. FIRE peaks beyond 250 bp of a canonical transcript TSS were classified as promoter-distal. Enrichments of deaminated cytidines and actuated regulatory elements by genomic repeat class was performed by intersecting these features with RepeatMasker (see, for example, Smit, AFA, Hubley, R. & Green, P “RepeatMasker”) annotations within the mappable autosomal portions of GRCh38. Regulatory element codependency was calculated as described previously. Briefly, for each pair of Fiber-seq FIRE peaks, the expected co-actuation was calculated as the product of fibers actuated at each peak, while the observed co-actuated was calculated as the proportion of fibers actuated at both elements. Codependency scores were binned by genomic distance between peaks using a log2 scale. Codependency scores of each bin were calculated as the mean of codependency scores within that distance bin for peak pairs covered by at least 8 fibers. Full-length transcript data processing3915-P1364WO.UW -59-

[0185] GM24385 ISO-seq reads were aligned to hg38 using pbmm2 align (PacBio®; parameters: --preset ISOSEQ) and the aligned reads were collapsed using Iso- Seq collapse (PacBio®; parameters: --do-not-collapse-extra-5exons). Transcript starts were intersected with GM24385 Fiber-seq promoter-proximal peaks using Bedops bedmap and transcript counts were quantified using custom scripts. Accessibility enrichment scores

[0186] Bulk ATAC-seq datasets were aligned to hg38 using Bowtie2, deduplicated for enrichments using GATK MarkDuplicates, and peaks were called separately for each dataset using macs2 (parameters: --nomodel --shift -100 --extsize 200). MSP and TSS enrichment scores were calculated for scDAF-seq and Fiber-seq libraries by mapping actuated elements (MSP > 150 bp) within 2 kb of GM24385 Fiber-seq FIRE peaks and TSSs using Bedops (v2.4.41). TSS enrichment scores were calculated for ATAC-seq libraries as above by mapping deduplicated sequencing reads to sample-specific TSS- overlapping peak sets. Single-cell ATAC-seq fragments were used as pseudobulked input to peak calling and TSS enrichment. All scATAC-seq datasets were pseudobulked using all sequenced fragments with the exception of GM12878 which was filtered to only include sequenced fragments from cell passing the Cell Ranger ATAC pipeline (10x Genomics®v.2.1.0). Genomic density at these positions was calculated using Bedtools genomecov (v2.31.0) and signal was aggregated across all regions using custom python scripts. TSS signal was oriented by strand. Single-cell CTCF co-occupancy

[0187] CTCF motifs within the GRCh38 reference were identified using FIMO. These motifs were then filtered to include only those that filly overlapped CTCF ChIP-seq peaks from the lymphoblastoid cell line GM12878 (ENCODE accessions ENCFF356LIU and ENCFF960ZGP), GM12878 CTCF ChIA-PET anchor regions (ENCODE accession ENCFF780PGS), and GM24385 Fiber-seq FIRE peaks using Bedtools (v2.31.0) and Bedops (v2.4.41). It was decided to limit CTCF footprinting to the core CTCF binding motif, which includes modules two and three. CTCF footprints were identified within actuated elements that completely overlapped modules two and three of a motif using Fibertools footprint. Thus, fibers with partial deamination within a footprint region are classified as ‘unbound’ at that site. CTCF co-occupancy analyses was limited to FIRE peaks actuated on at least 30% of fibers in Fiber-seq data and required a minimum of four haplotype-strand fibers covering both CTCF motifs. For each loop, the percentage of fibers3915-P1364WO.UW -60-co-occupied by CTCF was quantified within anchor regions (anchor regions often contain multiple CTCF motifs) for each motif orientation observed within the loop anchor pair. Data availability

[0188] DNA sequencing data have been deposited to the NCBI BioProject database under accession number PRJNA1203351.

[0189] Example 12. Supplementary Information PCR Design Considerations Amplicon Selection

[0190] When designing primers for DAF-seq, it is recommended to target 2-7 kilobase regions of DNA. When possible, it is recommended to avoid targeting regions with a high abundance of repeat elements and to avoid the design of primers that overlap repeat elements or common SNPs. In addition, one should avoid designing primers that directly overlap with accessible regulatory elements within the sample of interest. Primer Design

[0191] One can find success designing primers to have limited (≤2) guanine nucleotides, to use degenerate bases - specifically, an unspecified purine nucleotide (i.e., R, wherein R is A or G) - in place of guanine nucleotides, and to be around 20-25 nucleotides in length. In addition, it is suggested to aim to minimize the number of cytosine nucleotides when designing primers, where possible. It is also suggested to minimize or avoid guanine or cytosine nucleotides within the final 3 base pairs at the 3’ end of the primer. Off-Target Analysis

[0192] Primers can be analyzed to ensure that they avoid targeting undesired sites. Off-target alignments can be identified using a suitable tool, such as BLAST. The primer aligns with 100% identity to the desired site on chromosome 3 but also aligns with 100% identity to an off-target site on chromosome 22. In this example, care should be taken when designing a second primer to avoid amplification of an off-target amplicon. Hairpin and Dimer Analysis

[0193] Primers should be analyzed to ensure that they will not form unwanted hairpin structures or dimers. Possible hairpins or dimers can be assessed through the IDT Oligo Analyzer tool. For these structures, a predicted delta G of -9 kcal / mol or less suggests a higher risk of forming these troublesome structures. For example, a primer can meet other3915-P1364WO.UW -61-criteria for good design but can be predicted to form homodimers and so should be removed from selection. PCR Reaction Conditions

[0194] One can find success using any of various PCR master mixes including LongAmp Taq 2x Master Mix (New England Biolabs®, M0287), NEBNext®Q5U®Master Mix (New England Biolabs®, M0597), and repliQa HiFi ToughMix (Quantabio®, 95200). For sites that prove difficult to target, experimentation with different PCR enzymes can facilitate success. For example, specific amplification of a targeted region (Hg38; chr4:184646373-184651535) proved challenging when using NEBNext Q5U Master Mix, which may be attributed to the presence of repeating elements in the site. Using these PCR enzymes, the gel image shown displays a large amount of non-specific or off-target amplification. However, by testing a different PCR master mix (here, repliQa HiFi ToughMix), it became possible to significantly increase targeted amplification of the desired region. Annealing Temperature Optimization

[0195] For primer quality control, it is suggested that one optimize the annealing temperature by running a gradient PCR with untreated genomic DNA (gDNA) followed by gel electrophoresis. Given the use of degenerate bases in the primers, one can aim to select the lowest viable annealing temperature that produces a strong clear band at the appropriate size for the targeted amplicon on gel electrophoresis.

[0196] Gradient PCR was performed on two sets of primers (NAPA: hg38; chr19:47514488-47518985, UBA1: hg38; chrX:47190561-47194939). A final annealing temperature of 56°C was selected for both sets of primers based on these tests. Testing on gDNA from DddA-treated Nuclei

[0197] After an optimal annealing temperature has been identified, it is suggested that primers be tested using the identified reaction conditions on gDNA from DddA-treated nuclei. Primers can be tested on gDNA from DddA-treated nuclei at the annealing temperature determined by gradient PCR to verify presence of appropriately sized amplicons by gel electrophoresis. PCR Product Quality Control

[0198] An initial verification of target amplification and PCR product quality control can be established by gel electrophoresis and long-read sequencing (e.g., nanopore sequencing), such as PacBio®or Oxford Nanopore Technologies®(ONT®) sequencing3915-P1364WO.UW -62-(such as through commercial options like Plasmidsaurus). Both NEBNext®Q5U and repliQa®HiFi ToughMix effectively amplified the target region (hg38; chr22:37029577- 37033638) and yielded a single band of the expected length on gel electrophoresis. However, sequencing through the Plasmidsaurus Standard Premium PCR service revealed that a substantial portion of the NEBNext®Q5U library consisted of short, truncated amplicons. By contrast, the repliQa®HiFi ToughMix library consisted of primarily full- length amplicons and yielded substantially more even coverage across the target region. Specifically, 54.80% of reads from the NEBNext®Q5U library that mapped to the target region were full-length whereas 75.52% of reads from the repliQa HiFi ToughMix library that mapped to the target region were full-length, where full-length is defined as reads that align within a window of 30 base pairs on either side of the expected ends of the target region and have less than 30 base pairs of soft clipping.

[0199] In addition, it was observed that when using 35 PCR cycles, there were fragments that clearly arose from chimeric reads, as these individual reads contained regions with C->T mutations as well as G->A mutations. To avoid these chimeric reads, the number of PCR cycles was reduced from 35 to 30 which significantly reduced the number of chimeric reads at the target region (hg38; chr4:184646373-184651535). This resulted in a reduction in the number of chimeric reads from 15.0% of the full-length reads using 35 cycles to 1.3% of the full-length reads using 30 cycles.

[0200] Verifying targeting efficiency (total mapped reads and full-length reads) and deamination rate, as well as assessing for strand bias, chimera rate, jackpot effects, and mutation rate (that is, the rate of mismatches per read that are not C>T mutations on the top strand and that are not G>A mutations on the bottom strand), is recommended. Cell Input Considerations

[0201] The protocol for performing DAF-seq was developed with 100,000 permeabilized cells or nuclei as input into the DddA reaction. Higher inputs to the deaminase reaction were also tested. Specifically, 100,000, 250,000, 500,000, and 1,000,000 permeabilized cells were treated at 25°C for 10 minutes at an enzyme concentration of 4 μM. Reactions were terminated by adding 20% SDS to reach a final concentration of 5%. Genomic DNA was extracted using the Monarch Spin gDNA Extraction Kit (New England Biosciences®, T3010). DNA quantification was performed using a Qubit 1x dsDNA High Sensitivity kit (Invitrogen®, Q33231). For PCR amplification, hg38; chr3:179229955-179234791 was targeted and amplified in 50 uL3915-P1364WO.UW -63-reactions made up of 25 uL repliQa HiFi ToughMix (Quantabio®, 95200), 1.5 uL each of 10 μM forward and reverse primers (F: AATCACTATATTTCCATACTACTCA (SEQ ID NO:17), R: TACACCTAATRCAATAACARCCT (SEQ ID NO:18)), and 15-70 ng of DddA-treated gDNA. PCR conditions were as follows: 98°C for 30 seconds; 30 cycles of 98°C denaturation for 10 seconds, 53°C annealing for 5 seconds, and 68°C extension for 25 seconds; followed by a final extension at 68°C for 1 minute. PCR products were purified using the Monarch Spin PCR & DNA Cleanup Kit (New England Biosciences®, T1130). Amplicons were sequenced using the standard Premium PCR service from Plasmidsaurus. It was identified that the deamination rate decreased by only 20% with increasing input cell / nuclei number from 100,000 to 1,000,000. Specifically, whereas the deamination rate was 29% with 100,000 cells, this decreased to 26% with 250,000, 24% with 500,000 cells, and 23% with 1,000,000 cells. Overall, this demonstrates that analytical differences in the cell number input have a minimal effect on the chromatin stenciling of DAF-seq. NON-LIMITING EMBODIMENTS

[0202] While general features of the disclosure are described and shown and particular features of the disclosure are set forth in the claims, the following non-limiting embodiments relate to features, and combinations of features, that are explicitly envisioned as being part of the disclosure. The following non-limiting Embodiments contain elements that are modular and can be combined with each other in any number, order, or combination to form a new non-limiting Embodiment, which can itself be further combined with other non-limiting Embodiments.

[0203] Embodiment 1. A method for mapping protein-DNA interactions to a DNA sequence of a chromosome, the method comprising: contacting the chromosome with a DNA deaminase for a deaminase reaction that comprises cytosine deamination of enzyme-accessible portions of the chromosome; and sequencing the DNA sequence of the chromosome, wherein the DNA sequence comprises a plurality of extended footprints that correspond to positions of the chromosome that did not undergo cytosine deamination due to protein-DNA interactions of nucleosome-like proteins of the chromosome.

[0204] Embodiment 2. The method of Embodiment 1 or any other Embodiment, wherein the method is performed with a single cell, a circulating tumor cell, a plurality of cells, a tissue, a plurality of tissues, an organ, a plurality of organs, an organism, or a plurality of organisms.3915-P1364WO.UW -64-

[0205] Embodiment 3. The method of any one of Embodiments 1-2 or any other Embodiment, wherein the DNA deaminase is SsDddA or a dsDNA deaminase derived from Simiaoa sunii (DddAtox).

[0206] Embodiment 4. The method of any one of Embodiments 1-3 or any other Embodiment, wherein the DNA deaminase is SsDddA and comprises 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO:42.

[0207] Embodiment 5. The method of any one of Embodiments 1-4 or any other Embodiment, wherein the DNA deaminase is SsDddA and comprises 100% identity with SEQ ID NO:42.

[0208] Embodiment 6. The method of any one of Embodiments 1-5 or any other Embodiment, wherein the DNA deaminase is SsDddA5 and comprises 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO:43.

[0209] Embodiment 7. The method of any one of Embodiments 1-6 or any other Embodiment, wherein the DNA deaminase is SsDddA5 and comprises 100% identity with SEQ ID NO:43.

[0210] Embodiment 8. The method of any one of Embodiments 1-7 or any other Embodiment, wherein the deaminase reaction is stopped with addition of the DddA immunity protein, DddI, thereto.

[0211] Embodiment 9. The method of any one of Embodiments 1-8 or any other Embodiment, wherein the method further comprises amplifying the chromosome, or a portion thereof, by preferentially priming amplification on a primary DNA template that comprises C to T nucleotide transitions or G to A nucleotide transitions that correspond to locations of the chromosome that underwent cytosine deamination by the DNA deaminase.

[0212] Embodiment 10. The method of any one of Embodiments 1-9 or any other Embodiment, wherein the sequencing comprises translocating the DNA sequence through a nanopore, circular consensus sequencing, sequencing by synthesis, or sequencing by binding.

[0213] Embodiment 11. The method of any one of Embodiments 1-10 or any other Embodiment, further comprising amplifying the chromosome or a genome that comprises the chromosome with a whole genome amplification (WGA) method.

[0214] Embodiment 12. The method of any one of Embodiments 1-11 or any other Embodiment, wherein the WGA method comprises primary template-directed amplification (PTA).3915-P1364WO.UW -65-

[0215] Embodiment 13. The method of any one of Embodiments 1-12 or any other Embodiment, further comprising aligning multiple sequence reads based on a unique molecular identifier (UMI) that comprises a unique cytosine deamination pattern formed by cytosine deamination of a primary DNA template that comprises C to T nucleotide transitions or G to A nucleotide transitions that correspond to locations of the chromosome that underwent cytosine deamination by the DNA deaminase.

[0216] Embodiment 14. The method of any one of Embodiments 1-13 or any other Embodiment, further comprising generating a chromatin accessibility map based on evidence of cytosine deamination, wherein the chromatin accessibility map comprises information that indicates one or more regions of chromatin that are not bound by proteins and one or more regions of chromatin that are bound by proteins.

[0217] Embodiment 15. The method of any one of Embodiments 1-14 or any other Embodiment, wherein the chromosome constitutes a portion of a genome of a cell derived from a subject.

[0218] Embodiment 16. The method of any one of Embodiments 1-15 or any other Embodiment, wherein the subject has or is suspected of having a condition, disease, or disorder.

[0219] Embodiment 17. The method of any one of Embodiments 1-16 or any other Embodiment, wherein the condition, disease, or disorder is somatic or germline.

[0220] Embodiment 18. The method of any one of Embodiments 1-17 or any other Embodiment, wherein the chromosome is edited.

[0221] Embodiment 19. The method of any one of Embodiments 1-18 or any other Embodiment, wherein an edit of the chromosome is produced by a Cas method or a CRISPR / Cas9 method.

[0222] Embodiment 20. The method of any one of Embodiments 1-19 or any other Embodiment, wherein the Cas method or the CRISPR / Cas9 method comprises adenine base editing.

[0223] Embodiment 21. The method of any one of Embodiments 1-20 or any other Embodiment, further comprising comparing a chromatin accessibility map of a test cell to a chromatin accessibility map of a normal cell; wherein a difference in a DNA sequence and no difference in a protein-DNA interaction indicates the difference in the DNA sequence is independent of the protein-DNA interaction; and wherein a difference in3915-P1364WO.UW -66-a DNA sequence and a difference in a protein-DNA interaction indicates the difference in the DNA sequence is not independent of the protein-DNA interaction.

[0224] Embodiment 22. The method of any one of Embodiments 1-21 or any other Embodiment, wherein a protein of a protein-DNA interaction is a histone.

[0225] Embodiment 23. The method of any one of Embodiments 1-22 or any other Embodiment, further comprising generating a database comprising data that corresponds to a chromatin structure, an underlying genomic DNA sequence, a correlation to a condition, a correlation to a disease, a correlation to a disorder, or any combination thereof.

[0226] Embodiment 24. The method of any one of Embodiments 1-23 or any other Embodiment, wherein sequence data is suitable for use to identify PCR duplicates.

[0227] Embodiment 25. The method of any one of Embodiments 1-24 or any other Embodiment, wherein identified PCR duplicates can be used to generate a consensus sequence of a sequence read.

[0228] Embodiment 26. The method of any one of Embodiments 1-25 or any other Embodiment, wherein sequence data can be used to reconstruct a genomic sequence of a sequenced sample, identify minimal residual disease, or both.

[0229] Embodiment 27. A kit, comprising: one or more reagents configured for performance of the method of any one of Embodiments 1-26 or any other Embodiment; and an instructional material.

[0230] Embodiment 28. The kit of Embodiment 27 or any other Embodiment, comprising the DNA deaminase.

[0231] Embodiment 29. A non-transitory computer readable storage medium having stored thereon instructions which, when executed by one or more processors, cause the one or more processors to perform all or part of the method of any one of Embodiments 1-26 or any other Embodiment.

[0232] Embodiment 30. A computational device or system comprising the non-transitory computer readable storage medium of Embodiment 29 or any other Embodiment.

[0233] While illustrative embodiments have been illustrated and described, it will be appreciated that various changes can be made therein without departing from the spirit and scope of the disclosure.3915-P1364WO.UW -67-

Claims

CLAIMS The embodiments of the invention in which an exclusive property or privilege is claimed are defined as follows:

1. A method for mapping protein-DNA interactions to a DNA sequence of a chromosome, the method comprising: contacting the chromosome with a DNA deaminase for a deaminase reaction that comprises cytosine deamination of enzyme-accessible portions of the chromosome; and sequencing the DNA sequence of the chromosome, wherein the DNA sequence comprises a plurality of extended footprints that correspond to positions of the chromosome that did not undergo cytosine deamination due to protein-DNA interactions of nucleosome- like proteins of the chromosome.

2. The method of claim 1, wherein the method is performed with a single cell, a circulating tumor cell, a plurality of cells, a tissue, a plurality of tissues, an organ, a plurality of organs, an organism, or a plurality of organisms.

3. The method of claim 1, wherein the DNA deaminase is SsDddA or a dsDNA deaminase derived from Simiaoa sunii (DddAtox).

4. The method of claim 3, wherein the DNA deaminase is SsDddA and comprises 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO:

42.

5. The method of claim 4, wherein the DNA deaminase is SsDddA and comprises 100% identity with SEQ ID NO:

42.

6. The method of claim 3, wherein the DNA deaminase is SsDddA5 and comprises 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO:

43.

7. The method of claim 6, wherein the DNA deaminase is SsDddA5 and comprises 100% identity with SEQ ID NO:

43.

8. The method of claim 1, wherein the deaminase reaction is stopped with addition of the DddA immunity protein, DddI, thereto.3915-P1364WO.UW -68-9. The method of claim 1, wherein the method further comprises amplifying the chromosome, or a portion thereof, by preferentially priming amplification on a primary DNA template that comprises C to T nucleotide transitions or G to A nucleotide transitions that correspond to locations of the chromosome that underwent cytosine deamination by the DNA deaminase.

10. The method of claim 1, wherein the sequencing comprises translocating the DNA sequence through a nanopore, circular consensus sequencing, sequencing by synthesis, or sequencing by binding.

11. The method of claim 1, further comprising amplifying the chromosome or a genome that comprises the chromosome with a whole genome amplification (WGA) method.

12. The method of claim 9, wherein the WGA method comprises primary template-directed amplification (PTA).

13. The method of claim 1, further comprising aligning multiple sequence reads based on a unique molecular identifier (UMI) that comprises a unique cytosine deamination pattern formed by cytosine deamination of a primary DNA template that comprises C to T nucleotide transitions or G to A nucleotide transitions that correspond to locations of the chromosome that underwent cytosine deamination by the DNA deaminase.

14. The method of claim 1, further comprising generating a chromatin accessibility map based on evidence of cytosine deamination, wherein the chromatin accessibility map comprises information that indicates one or more regions of chromatin that are not bound by proteins and one or more regions of chromatin that are bound by proteins.

15. The method of claim 1, wherein the chromosome constitutes a portion of a genome of a cell derived from a subject.

16. The method of claim 1, wherein the subject has or is suspected of having a condition, disease, or disorder.3915-P1364WO.UW -69-17. The method of claim 1, wherein the condition, disease, or disorder is somatic or germline.

18. The method of claim 1, wherein the chromosome is edited.

19. The method of claim 18, wherein an edit of the chromosome is produced by a Cas method or a CRISPR / Cas9 method.

20. The method of claim 19, wherein the Cas method or the CRISPR / Cas9 method comprises adenine base editing.

21. The method of claim 1, further comprising comparing a chromatin accessibility map of a test cell to a chromatin accessibility map of a normal cell; wherein a difference in a DNA sequence and no difference in a protein-DNA interaction indicates the difference in the DNA sequence is independent of the protein- DNA interaction; and wherein a difference in a DNA sequence and a difference in a protein-DNA interaction indicates the difference in the DNA sequence is not independent of the protein- DNA interaction.

22. The method of claim 1, wherein a protein of a protein-DNA interaction is a histone.

23. The method of claim 1, further comprising generating a database comprising data that corresponds to a chromatin structure, an underlying genomic DNA sequence, a correlation to a condition, a correlation to a disease, a correlation to a disorder, or any combination thereof.

24. The method of claim 1, wherein sequence data is suitable for use to identify PCR duplicates.

25. The method of claim 24, wherein identified PCR duplicates can be used to generate a consensus sequence of a sequence read.

26. The method of claim 24, wherein sequence data can be used to reconstruct a genomic sequence of a sequenced sample, identify minimal residual disease, or both.3915-P1364WO.UW -70-27. A kit, comprising: one or more reagents configured for performance of the method of claim 1; and an instructional material.

28. The kit of claim 27, comprising the DNA deaminase.

29. A non-transitory computer readable storage medium having stored thereon instructions which, when executed by one or more processors, cause the one or more processors to perform all or part of the method of claim 1.

30. A computational device or system comprising the non-transitory computer readable storage medium of claim 29.3915-P1364WO.UW -71-