Joint analysis of DNA damage and RNA expression in single cells
Paired-Damage-seq addresses the lack of single-cell resolution analysis by jointly analyzing DNA damage and transcriptome, providing insights into cellular programs and disease outcomes.
Patent Information
- Application Number
- PCT/US2025/036769
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-08
- Filing Date
- 2025-07-08
- Publication Date
- 2026-01-15
AI Technical Summary
Current methods lack the capability to analyze DNA damage and gene expression jointly at single-cell resolution, hindering the understanding of their functional impacts and relationships with cellular programs, particularly in the context of aging and cancer research.
The development of Paired-Damage-seq, a single-cell multiomics sequencing technology that simultaneously analyzes DNA damage and transcriptome from the same cells, enabling the identification of DNA damage sites and linking them with functional outputs.
Enables the identification of DNA damage signals and gene expression information at cell type and state resolution, facilitating insights into pathological processes and potential interventions in aging and cancer research.
Smart Images

Figure IMGF000049_0001 
Figure IMGF000070_0001 
Figure IMGF000070_0002
Abstract
Description
[0001] JOINT ANALYSIS OF DNA DAMAGE AND RNA EXPRESSION IN SINGLE
[0002] CELLS
[0003] CROSS-REFERENCE TO RELATED APPLICATION
[0004] This application claims the benefit of U.S. Provisional Application No. 63 / 668,570, filed on July 8, 2024, the entire contents of which are incorporated by reference herein in their entirety.
[0005] GOVERNMENT SUPPORT
[0006] This invention was made with government support under grant number GM 154011 and R00HG011483 awarded by the National Institutes of Health. The government has certain rights in the invention.
[0007] BACKGROUND
[0008] In multicellular organisms, cells with virtually identical genome can perform diverse biological functions - which requires faithful preservation of genome templates and precise regulation of gene expression. To make this possible, multiple DNA repair machineries safeguard the inherently unstable genome material against constant challenges from various internal and environmental stresses. In addition, a complex collection of chemical and molecular assemblies known as epigenome, orchestrate the spatiotemporal control of genetic information flow for intended functions. However, errors can arise during these processes and accumulate over time, ultimately leading to deleterious consequences. For example, defects in DNA repair machinery and rises in mutational burden are frequently observed in cancers and age-related diseases. Disrupted epigenetic programs can also change cells’ behaviors, thereby resulting in altered phenotypic plasticity that can be pathogenic. However, it has not been clear to what extent the genome and epigenome erosions contribute to the disease outcome primarily due to the lack of methods capable of analyzing DNA damage with cellular programs jointly at single-cell resolution. Accordingly, there is a great need in the art for compositions and methods that can help understand the functional impacts of DNA damage and their relationships with gene expression.
[0009] SUMMARY OF THE INVENTION
[0010] The present invention is based, at least in part, on the discovery that the methods comprising a single-cell multiomics sequencing technology, Paired-Damage-seq, are effective in analyzing simultaneously DNA damage with transcriptome from the same cells. These methods enable the identification of DNA damage signals (including e.g., the preferred genomic distribution of DNA damage or DNA damage hotspots) at cell type and state resolution and directly linking DNA damage with functional outputs in single cells. The methods of the present disclosure have broad applications in aging and cancer research to understand pathological processes’ causes and develop interventions.
[0011] In certain aspects, provided herein are methods of determining DNA damage sites in the genome and obtaining gene expression information of a cell or a nucleus, the method comprising: a) obtaining one or more nuclei, b) labeling the DNA damage; c) performing in situ reverse transcription of RNA in the one or more nuclei to generate cDNA; d) barcoding the DNA damage to generate genomic DNA fragments comprising a barcode, and barcoding the cDNA to generate cDNA comprising a barcode, optionally wherein the barcode for DNA damage is identical to the barcode for cDNA; e) preparing a DNA library from the barcoded genomic DNA fragments, and preparing an RNA library from the barcoded cDNA; f) sequencing the DNA library and the RNA library (e.g., via next generation sequencing), thereby determining the DNA damage sites in the genome and obtaining the gene expression information of the cell or the nucleus.
[0012] In some embodiments, the methods further comprise, after obtaining one or more nuclei, fixing the one or more nuclei. In some embodiments, the one or more nuclei are fixed by contacting with formaldehyde or formalin. In other embodiments, the one or more nuclei are fixed by being exposed to UV light.
[0013] In some embodiments, the methods further comprise, after obtaining one or more nuclei, permeabilizing the one or more nuclei. In some embodiments, the one or more nuclei are permeabilized by contacting with a detergent. In some embodiments, the detergent comprises any one selected from Triton X-100, NP-40, digitonin, and a combination thereof.
[0014] In some embodiments, the methods further comprise, after obtaining one or more nuclei, disrupting or depleting nucleosomes from the one or more nuclei. In some embodiments, the nucleosomes are disrupted or depleted by contacting with detergent. In some embodiments, the detergent is a strong detergent such as Sodium Dodecyl Sulfate or Lithium Diiodosalicylate. In some embodiments, the methods comprise labeling DNA damage by tagmentation. In some embodiments, the targeted tagmentation comprises a transposase, optionally Tn5 or a variant (e.g., hyperactive variant) thereof. In preferred embodiments, tagmentation is a targeted tagmentation. In some embodiments, the targeted tagmentation comprises a transposase (e.g., Tn5) and the transposase is a fusion protein or a conjugate further comprising a protein that recognizes the DNA damage (e.g., protein A-Tn5 conjugate or fusion protein).
[0015] In some embodiments, the tagmentation comprises incorporating an oligonucleotide comprising a barcode (e.g., barcode 1 or BC1). In some embodiments, the oligonucleotide further comprises a first restriction site.
[0016] In some embodiments, the tagmentation comprises: a) incorporating a labeled nucleotide (e.g., biotin-labeled dUTP) to the DNA damage in the one or more nuclei (e.g., using a DNA polymerase); and b) labeling the DNA damage by incubating the one or more nuclei with (i) a transposase, which further comprises a protein that targets the transposase to the labeled nucleotide (e.g., an anti-biotin antibody fused or conjugated to a transposase), and (ii) an oligonucleotide, which optionally comprises a barcode and / or a first restriction site.
[0017] In some embodiments, the tagmentation comprises: a) incorporating a labeled nucleotide to the DNA damage in the one or more nuclei (e.g., using a DNA polymerase, e.g., biotin-labeled dUTP); b) contacting the one or more nuclei with a protein that binds to the labeled nucleotide at the DNA damage (e.g., anti-biotin antibody); and c) labeling the DNA damage by incubating the one or more nuclei with (i) a transposase (e.g., Tn5), which further comprises a protein (e.g., fused to or conjugated to Protein A) that binds to the protein in b), and (ii) an oligonucleotide, which optionally comprises a barcode and / or a first restriction site.
[0018] In some embodiments, the tagmentation comprises labeling the DNA damage in two or more aliquots, each with an aliquot- specific oligonucleotide, optionally comprising: a) dividing the one or more nuclei into (n) aliquots (e.g., (n) = 1, 2... 12 or more); and b) labeling the DNA damage in each aliquot with an aliquot- specific oligonucleotide, which optionally comprises a barcode (e.g., barcode 1 or BC1) and, optionally, a first restriction site, by incubating each aliquot with (i) a transposase, or a fusion protein or a conjugate thereof, and (ii) an aliquot- specific oligonucleotide. In some embodiments, the methods comprise the in situ reverse transcription to generate cDNA, and the in situ hybridization further comprises incorporating an oligonucleotide comprising a barcode (e.g., barcode 1' or BC1') and, optionally, a second restriction site.
[0019] In some embodiments, the methods comprise barcoding the DNA damage and cDNA, and the barcoding comprises a single-cell barcoding.
[0020] In some embodiments, the methods comprise barcoding the DNA damage and cDNA, and the barcoding comprises a ligation-based combinatorial barcoding. In some embodiments, the ligation-based combinatorial barcoding comprises: a) pooling all aliquots of nuclei; b) dividing the pooled nuclei into (n) aliquots (e.g., (n) = 1, 2...96... 192...288...384 or more); c) ligating the DNA damage or the cDNA with an aliquot- specific barcode (e.g., BC2); and / or d) optionally repeating steps a) to c) at least 1, 2, or 3 times with a different aliquot- specific barcode (e.g., BC2, BC3, BC4, etc.), optionally whereby the DNA damage and the cDNA from the same cell comprise the same combination of the aliquot- specific barcodes.
[0021] In some embodiments, the methods comprise preparing DNA library, which comprises: a) lysing the one or more nuclei; b) adding a polynucleotide tail to the genomic DNA fragments (e.g., encompassing the DNA damage sites, e.g., by a terminal deoxynucleotidyl transferase (TdT)); c) amplifying the genomic DNA fragments using primers, one of which comprises (i) a sequence complementary to the polynucleotide tail, or a portion thereof, and / or (ii) a third restriction site; d) digesting the amplified genomic DNA fragments with restriction enzymes that cut the second restriction site (e.g., eliminating contaminating cDNAs) and the third restriction site; and / or e) generating the genomic DNA fragments comprising a sequencing adaptor, optionally (i) by ligating the sequencing adaptor to the genomic DNA fragments, or (ii) by a tagmentation (e.g., incubating the genomic DNA fragments with a transposase (e.g., Tn5) and a sequencing adaptor). In some embodiments, the methods comprise preparing RNA library, which comprises: a) lysing the one or more nuclei; b) adding a polynucleotide tail to the cDNA (e.g., by a terminal deoxynucleotidyl transferase (TdT)); c) amplifying the cDNA using primers, one of which comprises (i) a sequence complementary to the polynucleotide tail, or a portion thereof, and / or (ii) a third restriction site; d) digesting the amplified cDNA with a restriction enzyme that cuts the first restriction site (e.g., eliminating contaminating genomic DNA fragments); and e) generating the cDNA comprising a sequencing adaptor, optionally (i) by ligating the sequencing adaptor to the cDNA, or (ii) by a tagmentation (e.g., incubating the cDNA with a transposase (e.g., Tn5) and a sequencing adaptor).
[0022] In some embodiments, the methods comprise preparing DNA library and RNA library, which comprises: a) lysing the one or more nuclei to release the nucleic acid comprising genomic DNA fragments (e.g., encompassing DNA damage) and cDNA (e.g., generated from mRNA); b) adding a polynucleotide tail to the nucleic acid (e.g., by a terminal deoxynucleotidyl transferase (TdT)); c) amplifying the nucleic acid using primers, one of which comprises (i) a sequence complementary to the polynucleotide tail, or a portion thereof, and / or (ii) a third restriction site; d) dividing the nucleic acid into a DNA library and an RNA library; e) for the DNA library: i) enriching for genomic DNA fragments by digesting the nucleic acid with restriction enzymes that cut the second restriction site (e.g., eliminating cDNAs) and the third restriction site; and ii) generating the genomic DNA fragments comprising a sequencing adaptor by ligating the sequencing adaptor to the genomic DNA fragments; f) for the RNA library: i) enriching for cDNA by digesting the nucleic acid with a restriction enzyme that cuts the first restriction site (e.g., eliminating genomic DNA fragments); and e) generating the cDNA comprising a sequencing adaptor by a tagmentation (e.g., incubating the cDNA with a transposase (e.g., Tn5) and a sequencing adaptor).
[0023] Numerous embodiments are further provided that can be applied to any aspect of the methods of the present disclosure and / or combined with any other embodiments described herein.
[0024] For example, in some embodiments, the methods comprise the DNA damage selected from an oxidative DNA damage, double- stranded DNA breaks, DNA intrastrand crosslinking, and ribonucleotide in DNA. In some embodiments, the DNA damage is oxidative DNA damage, and the oxidative DNA damage comprises 8-oxoguanine (8-oxoG), repair intermediates apurinic / apyrimidinic (AP) sites, and / or single-stranded DNA breaks.
[0025] In some embodiments, the DNA damage comprises an oxidized base, and the oxidized base comprises an oxidized purine and / or an oxidized pyrimidine.
[0026] In some embodiments, the DNA damage may be processed to generate a certain intermediate that allows labeling.
[0027] In some embodiments, the DNA damage comprises an oxidized base and labeling the DNA damage comprises:
[0028] 1) contacting the one or more nuclei with an enzyme that removes an oxidized base; and / or
[0029] 2) contacting the one or more nuclei with an apurinic / apyrimidinic (AP) endonuclease, thereby generating a single-strand DNA break (SSB) at the DNA damage site.
[0030] In some embodiments, the DNA damage comprises an oxidized base and labeling the DNA damage comprises: a) an enzyme that removes an oxidized base selected from a DNA repair protein formamidopyrimidine [fapy]-DNA glycosylase (Fpg) or a glycosylase, optionally wherein the glycosylase is selected from 8-Oxoguanine DNA glycosylase (OGGI), endonuclease III, MutY homolog DNA glycosylase (MYH), NTHL1, NEIL1, NEIL2, and NEIL3; and / or b) an AP endonuclease selected from endonuclease IV, APE1, APE2, APN1, and APN2.
[0031] In some embodiments, the DNA damage comprises double-strand breaks (DSBs), and labeling the DNA damage comprises: a) contacting the one or more nuclei with Klenow fragment DNA polymerase and T4 Polynucleotide kinase; b) incubating the one or more nuclei with a double-stranded DNA adaptor comprising a label (e.g., biotin, labeled nucleotide) and a DNA ligase (e.g., T4 DNA ligase) to label the DSB site with the adaptor; and c) performing a tagmentation, optionally a targeted tagmentation (e.g., using an anti-label antibody (e.g., anti-biotin antibody) and protein A-Tn5).
[0032] In some embodiments, the labeling DNA damage comprises incorporating a labeled nucleotide to the DNA damage in the one or more nuclei by contacting the one or more nuclei with a DNA polymerase, a Bst polymerase, or a terminal deoxynucleotidyltransferase.
[0033] In some embodiments, a) the labeled nucleotide is selected from a labeled dUTP, a labeled dCTP, a labeled dATP, a labeled dTTP, and a labeled dGTP, optionally a labeled dUTP; and / or b) the labeled nucleotide comprises a biotin or a fluorophore.
[0034] In some embodiments, a) the labeled nucleotide is a biotin-labeled dUTP; and / or b) the protein that binds the labeled nucleotide is an antibody or streptavidin.
[0035] In some embodiments, the polynucleotide tail comprises an array of a single type of nucleotide, optionally an array of dCTP.
[0036] In some embodiments, the polynucleotide tail is added to the genomic DNA fragments and / or cDNA: a) by a terminal deoxynucleotidyltransferase (TdT); b) by a DNA ligase; c) by a DNA polymerase; or d) by a DNA or an RNA oligonucleotide comprising a reactive chemical group (e.g., an azide group or an alkyne group, e.g., suitable for click chemistry), which attaches to the 3’ end of the genomic DNA fragments and / or cDNA.
[0037] In some embodiments, the genomic DNA fragments and / or cDNA are amplified using a polymerase chain reaction.
[0038] In some embodiments, the ligation-based combinatorial barcoding is selected from SPLiT-seq, Sci-RNA-seq3, Paired- seq / Tag, and SHARE-seq.
[0039] In some embodiments, the third restriction site is recognized by a type IIS endonuclease, optionally selected from FokI, Acul, AsuHPI, Bbvl, Bpml, BpuEI, BseMII, BseRI, BseXI, Bsgl, BslFI, BsmFI, BsPCNI, BstVlI, BtgZI, Ecil, Eco57I, FaqI, Gsul, HphI, Mmel, NmeAIII, Schl, TaqII, TspDTI, and TspGWI.
[0040] In some embodiments, the ligase used in the methods is a T3, T4, or T7 DNA ligase.
[0041] In some embodiments, the methods further comprise obtaining information of the genome- wide histone modifications and / or chromatin state (e.g., location of histone variants, chromatin-associated factors (e.g., transcription factors, chromatin remodelers, etc.) in a single nucleus.
[0042] In some embodiments, the obtaining information of the genome-wide histone modifications and / or chromatin state comprises performing ChlP-chip, ChlP-seq, and / or Paired-Tag. In some embodiments, the histone modification is selected from H3K4mel, H3K4me2, H3K4me3, H3K27ac, H3K27me2, H3K27me3, H3K9me3, H3K9me2, H3K9Ac, H3K14Ac, H3K18Ac, H3K23Ac, H3K27me3, H4K16ac, H4K20mel, H4K20me2, H4K20me3, H4K27me3, HlK26mel, H3S10Ph, and H2A.XS139.
[0043] In some embodiments, the one or more nuclei are of a fungus, a bacterium, or a mammalian cell (e.g., a mouse cell, a rat cell, a monkey cell, a human cell).
[0044] In some embodiments, the one or more nuclei are of a cell type selected from a neuronal cell (e.g., excitatory neurons (ExN), inhibitory neurons (InN), and non-neuron cell types, including oligodendrocyte precursor cells (OPC), oligodendrocytes (ODC), astrocytes (AST), microglia cells (MiG), endothelial cells (Endo), and vascular leptomeningeal cells (VLMC)), a bone cell, an endothelial cell, a fibroblast, a blood cell (e.g., erythrocyte, lymphocyte, platelet), an epithelial cell, a muscle cell, a stem cell, an epidermal cell, adipocyte, and hepatocyte.
[0045] In some embodiments, the one or more nuclei are of a cancer cell.
[0046] In some embodiments, the one or more nuclei are of a proliferating cell, a quiescent cell, or a senescent cell.
[0047] In some embodiments, the one or more nuclei are from a young subject, a middle- aged subject, or an old subject.
[0048] In some embodiments: a) the subject is a mouse, wherein the young mouse is of the age between 1-9 months (preferably 1-4 months), the middle-aged mouse is of the age between 10-17 months (preferably 10-15 months), and the old mouse is of the age 18 months and higher (preferably 18-24 months); and / or b) the subject is a human, wherein the young human is of the age between the birth to 39 years, the middle-aged human is of the age between 40-59 years, and the old human is of the age 60 years and older.
[0049] In some embodiments, the one or more nuclei are treated with an agent, or the one or more nuclei are obtained from one or more cells treated with an agent. In some embodiments, the agent alters the cellular physiology (e.g., altering the integrity of DNA, RNA, and / or protein; altering histone modification and / or chromatin structure; altering gene expression (e.g., transcription and / or translation); altering cell morphology; altering cellular proliferation; altering the identity and / or activity of the cell, etc.; or any combination of two or more of such alterations. In some embodiments, the agent is selected from a toxin, a DNA damage inducing agent (e.g., H2O2, reactive oxygen species (ROS)), a chemotherapeutic agent, a small molecule, an alkylating agent, a hormone, a cell signaling molecule, a cytokine, a chemokine, and any combination of two or more thereof.
[0050] In some embodiments, the agent induces oxidative DNA damage.
[0051] In certain aspects, further provided herein are methods of determining the effect of an agent on a cell, the method comprising contacting the cell with the agent, and determining DNA damage sites in the genome and obtaining gene expression information of the cell according to the methods of the present disclosure.
[0052] BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Fig. lA-Fig. IL present joint profiling of oxidative and single-stranded DNA damage (and repair intermediates) with transcriptome in single cells. Fig. 1A presents schematic of the Paired-Damage- seq workflow with “labeling-by-repair” strategy to label damage sites. Fig. IB presents uniform manifold approximation and projection (UMAP) embedding showing single cells based on Paired-Damage- seq transcriptomic profiles. Each dot represents an individual cell and is colored according to the treatment conditions. Ctrl, control; 0 hr, 2 hr, 6 hr, 24 hr and 48 hr indicate the duration posttreatment in fresh medium. Fig. 1C presents representative genome browser view of pseudo-bulk and single-cell DNA damage signals of Paired-Damage- seq in HeLa cells. Bulk damage CUT&Tag, CLAPS-seq, ATAC-seq, and non-targeting tagmentation control signals are also shown. Fig. ID-Fig. IE present violin plots showing the (Fig. ID) unique number of transcripts and (Fig. IE) unique number of DNA fragments per cell across datasets generated with different single-cell multiomic assays. For all box plots, hinges were drawn from the 25th to 75th percentiles, with the middle line denoting the median and whiskers denoting a maximum 2x the interquartile range. N = 64,794 (Paired-Damage-seq), 1,133 (Paired-seq RNA), 1,091 (Paired-seq DNA), 408 (Paired-Tag RNA), 468 (Paired-Tag DNA), 5,000 (Droplet Paired- Tag RNA), 4,455 (Droplet Paired-Tag DNA), 302 (CoTECH RNA), 3,031 (CoTECH DNA), 11,614 (10X Multiome) cells. Fig. IF presents scatter plots showing the fraction of DNA reads mapped to human and mouse reference genome for each cell barcode in the speciesmixing experiment. Fig. IG-Fig. 1H present scatter plots that show the correlation of read counts from two technical replicates of Paired-Damage-seq (Fig. 1G) RNA profiles and (Fig. 1H) DNA profiles. Fig. II presents line plot showing the correlation of DNA damage signals (at 10-kb non-overlapping bins) between Bulk Paired-Damage-seq DNA profile and aggregated single-cell Paired-Damage-seq DNA signals from different numbers of cells. Fig. 1J presents heatmaps showing the reads densities on DNA damage hotspots for untreated, pre-repair control and non-targeting tagmentation control profiles. Fig. IK presents scatter plot showing the correlation between detected DNA damage reads densities with the numbers of Nt.BbvCI cutting sites per lOk-bp non-overlapping bins for nickase-treated nuclei. Fig. IL shows schematics for 2nd adaptor tagging of DNA and RNA libraries. For DNA libraries, amplified products were digested with a type IIS restriction enzyme - FokI, and the cohesive end was then used to ligate the P5 adaptor. For RNA libraries, N5 adaptor was added by tagmentation (see also Extended Data Fig. 1 of Zhu et al. (2021) Nat Methods 18(3):283-292, which is incorporated herein by reference in its entirety).
[0054] Fig. 2A-Fig. 2F present distribution of oxidative DNA damage hotspots in HeLa cells. Fig. 2A presents Distribution of DNA damage hotspots on chromosome 1 in HeLa cells. Colored bars are peak counts in 250-kb non-overlapping bins. ATAC-seq and chromosome compartments (4D Nucleosome: 4DNFIBMVFFOF) are also shown; Zoom-in view of DNA damage signals, ATAC-seq, H3K9me3 CUT&Tag, CLAPS-seq signals are also shown. Fig. 2B presents averaged DNA damage levels (RPKM) of different chromatin states (ENCODE: ENCSR098REA) in HeLa cells with and without H2O2 treatments. EnhBiv, Bivalent enhancer; EnhAl, Active enhancer 1; EnhA2, Active enhancer 2; EnhGl, Genic enhancer 1; EnhG2, Genic enhancer 2; EnhWk, Weak enhancer; TssBiv, Bivalent / Poised enhancer; TssA, Active TSS; TssFlnk, Flanking TSS; TssFlnkD, Flanking TSS downstream; TssFlnkU, Flanking TSS upstream; Tx, Strong transcription; TxWk, Weak transcription; ReprPC, Repressed PolyComb; ReprPCWk, Weak Repressed PolyComb; Het, Heterochromatin; Quies, Quiescent / Low; ZNF_Rpts, ZNF genes & repeats. Fig. 2C-Fig. 2D present scatterplot showing the correlation of changes in Paired-Damage-seq DNA levels and change in (Fig. 2C) ATAC-seq signals and (Fig. 2D) H3K9me3 CUT&Tag signals (RPKM, in 250-kb non-overlapping bins) in 48-hr post H2O2 HeLa cells compared to control group. Pearson correlation coefficients are indicated. Fig. 2E shows frequency density plots showing the distribution of Spearman’s correlation coefficients between gene expression levels and DNA damage peak signals within 500-kb range across metacells. To estimate background noise level, the distribution of correlations between shuffled DNA damage peak signals and gene expression levels is shown in the grey area. Fig. 2F presents enrichments of different groups of DNA damage peaks in histone modification ChlP-seq peaks (H3K4mel, H3K4me3, H3K9me3, H3K27ac, H3K27me3 and H3K36me3) of HeLa cells. N = 102,705 (Union), 4,362 (Negative) and 11,814 (Positive) peaks. Odd ratios for peaks displaying positive and negative relationships with gene expression are shown in red and blue, respectively; odd ratios for union peaks are shown in green. Odd ratios were calculated using Fisher’s exact test. Error bars represent the 95% confidence interval of the odds ratio.
[0055] Fig. 3A-Fig. 3G present cell-type resolved DNA damage landscapes of the mouse brain. Fig. 3A presents uniform manifold approximation and projection (UMAP) embedding showing single cells based on Paired-Damage- seq transcriptomic profiles. Each dot represents an individual cell and is colored according to the RNA-based annotation. Fig. 3B presents representative genome browser view of gene expression and DNA damage signals at representative cell-type-specific marker genes loci. Public ATAC-seq signals are also shown. Fig. 3C presents violin plots showing the DNA damage levels detected in compartment A and compartment B (RPKM, in 250-kb non-overlapping bins), respectively, for different mouse brain cell types; for all box plots, hinges were drawn from the 25th to 75th percentiles, with the middle line denoting the median, whiskers denoting a maximum 2x the interquartile range and outliers indicated with dots; n = 5,678 (Compartment A) and 5,233 (Compartment B). Fig. 3D presents barplots showing the relative enrichment of DNA damage peaks of different cell types in different genomic regions. Fig. 3E-Fig. 3F present scatter plots showing the relationships between DNA damage levels and (Fig. 3E) H3K27ac levels (RPKM, in 250-kb non-overlapping bins) in compartment A, and (Fig. 3F) H3K9me3 levels (RPKM, in 250-kb non-overlapping bins) in compartment B for excitatory neurons. Pearson correlation coefficients are also shown. Fig. 3G presents heatmap showing the DNA damage signals over public H3K27ac peaks in mouse excitatory neurons.
[0056] Fig. 4A-Fig. 4H present cell-type specific DNA damage hotspots associated with diverse cellular functions. Fig. 4A presents heat map showing cell-type specific DNA damage peaks. Top enriched GREAT GO terms for each group are shown; P value, one-sided Fisher’s exact test. Fig. 4B presents top enriched de novo motifs for DNA damage peaks of each cell type. P-values: one-sided binomial test. Fig. 4C presents heatmap showing the enrichment of known motifs for cell-type specific peaks. P-value: one-sided binomial test. Fig. 4D presents heat map showing the associations of DNA damage peaks in the human orthologues with the GWAS traits or diseases for different mouse brain cell types. -: FDR < 0.1, *: FDR < 0.05, **: FDR < 0.01. Displayed are all the 8 subclasses and traits or diseases with at least one significant association by linkage disequilibrium score regression analysis (FDR < 0.1). Fig. 4E presents genome browser view of lnpp5d locus in microglia cells. DNA damage peaks overlapped with ATAC-seq peaks decreased in aged mice are highlighted in light blue. Signals from inhibitory neurons are also shown as a control. Cis-regulatory elements (CREs) connected with Inpp5d gene by snATAC-seq co-accessiblity analysis are shown in violet. The AD-associated SNPs are also indicated. Fig. 4F-Fig. 4H present scatter plots showing the correlations of DNA damage levels (RPKM, in 100-kb non-overlapping bins) with change of H3K9me3 levels (RPKM, in 100-kb non-overlapping bins) between 18- month and 3-month for (Fig. 4F) excitatory neurons, (Fig. 4G) oligodendrocytes, and (Fig. 4H) astrocytes. Only the genomic bins with the top 1% highest damage levels are shown. Genomic bins with AH3K9me3 < -0.2 are shown in red, > 0.2 are shown in blue. Pearson correlation coefficients are indicated.
[0057] Fig. 5A-Fig. 5J present overview and validation of Paired-Damage-seq. Fig. 5A presents schematics of biotin labeling of DNA damaged sites with “Labeling-by-repair” strategy. Fig. 5B presents normalized Paired-Damage-seq DNA signal (RPKM) and DNase- seq on DHS (DNase I hypersensitive sites). Fig. 5C presents scatter plots showing the fraction of RNA reads mapped to human and mouse reference genome for each cell barcode in the species-mixing experiment. Barcodes with less than 75% reads from the same species were identified as mixed cells. Fig. 5D presents scatter plot showing the Pearson’s correlation coefficient of Paired-Damage-seq RNA dataset with in-house generated nucleus RNA-seq from HeLa cells. Fig. 5E presents scatter plots showing the Pearson’s correlation coefficients of pair-wise correlations between bulk and aggregated single-cell Paired-Damage-seq DNA dataset, AP-seq (AP-sites) and CLAPS-seq (8-Oxoguanine) datasets from HeLa cells. Fig. 5F presents dotplots showing the relative enrichments of model DNA sequences treated with different buffer conditions. Nickases Nt.AlwI and Nt.BstNBI were used to generate SSBs as positive control; technical replicates n = 3 for all conditions. Fig. 5G presents line plots showing the normalized DNA signal enrichments of ATAC-seq, non-targeting tagmentation control and Paired-Damage-seq DNA signal on DHSs (DNase I hypersensitive sites) of HeLa cells. Fig. 5H presents barplots showing the relative DNA damage levels (normalized by spike-in mouse 3T3 cells) in HeLa cells labeled with different enzyme combinations; technical replicates n = 3 for all combinations. Fig. 51 presents scatter plot showing the correlation between detected DNA damage reads densities and the numbers of Nt.BbvCI cutting sites per lOk-bp non-overlapping bins for control nuclei. Fig. 5J presents scatter plots showing the Spearman’s correlation coefficients of pair- wise correlations between bulk and aggregated single-cell Paired-Damage-seq DNA dataset, AP-seq (AP-sites) and CLAPS-seq (8-Oxoguanine) datasets from HeLa cells. ATAC-seq and non-targeting tagmentation control are also shown for comparisons. Fig. 6A-Fig. 6L present distribution of oxidative DNA damage hotspots. Fig. 6A presents experiment design of H2O2 treatment on HeLa cells. Fig. 6B presents uniform manifold approximation and projection (UMAP) embedding showing single cells based on Paired-Damage- seq DNA profiles. Each dot represents an individual nucleus profiled by Paired-Damage- seq and is colored according to the treatment conditions. Fig. 6C presents violin plots showing the averaged detected signal levels in compartment A and compartment B (RPKM, in 250-Kb non-overlaping bins) for DNA damage, non-targeting tagmentation control and ATAC-seq signals of HeLa cells; for all box plots, hinges were drawn from the 25th to 75th percentiles, with the middle line denoting the median, whiskers denoting a maximum 2x the interquartile range and outliers indicated with dots; n = 5,495 (Compartment A) and 4,593 (Compartment B). Fig. 6D presents line plots showing the DNA damage signals around genic regions of genes with different expression levels. Fig. 6E presents barplots showing the numbers of DNA damage peaks in control and H2O2 treated HeLa cells. Fig. 6F presents barplots showing the relative enrichments of DNA damage peaks of control and H2O2 treated HeLa cells in different genomic regions. Fig. 6G presents Upset plots showing the intersection size of damage peaks in control and H2O2 treated HeLa cells. Fig. 6H presents heatmaps showing the reads densities on DNA damage hotspots from cells with treatment of varying concentrations of H2O2. Fig. 61 presents weighted gene co-expression network analysis (hdWGCNA) dendrograms for the co-expression networks constructed. Fig. 6J presents module eigengene as a function of pseudotime for the representative coexpression module Ml (with decreased expression levels) and M3 (with increased expression levels). Fig. 6K presents the enriched GO Terms for co-expression module Ml and M3. Fig. 6L presents Upset plots showing the intersection size of damage peaks in untreated and H2O2 treated HeLa cells. The non-targeting tagmentation control are also shown for comparison.
[0058] Fig. 7A-Fig. 7E present accumulation of DNA damage induced by oxidative stress. Fig. 7A presents relative enrichment of DNA damage peaks in control and 48-hr post H2O2 treatment HeLa cells in different short tandem repeat (STR) subfamilies. Fig. 7B presents barplots showing the relative enrichment of conserved and induced DNA damage peaks in different genomic regions. Fig. 7C presents line plots showing the DNA damage levels (RPKM) on conserved and induced peaks in HeLa cells of different treatments. Fig. 7D presents line plots showing the DNA damage levels (RPKM) on simple repeats, Z-form DNA and putative G-quadruplex sequences in control and 48-hr post H2O2 treatment HeLa cells. Fig. 7E presents line plots showing DNA damage signals, ATAC-seq signals and RNA-seq signals around endogenous retroviruses (ERV) long terminal repeats (LTR) regions in compartment A and compartment B of control and 48-hr post H2O2 treatment HeLa cells.
[0059] Fig. 8A-Fig. 8D present relationships between DNA damage levels and epigenome signature changes. Fig. 8A-Fig. 8B present scatter plots showing the relationships between changes in Paired-Damage-seq DNA levels and change in ATAC-seq signals (RPKM, in 250-kb non-overlapping bins) in (Fig. 8A) 0-hr post H2O2 treatment, and (Fig. 8B) 6-hr post H2O2 treatment HeLa cells compared to control group. Fig. 8C-Fig. 8D present scatter plots showing the correlation of changes in Paired-Damage-seq DNA levels and change in H3K9me3 CUT&Tag signal (RPKM, in 250-kb non-overlapping bins) in (Fig. 8C) 0-hr post H2O2 treatment, and (Fig. 8D) 6-hr post H2O2 treatment HeLa cells compared to control group. Pearson correlation coefficients are also shown.
[0060] Fig. 9A-Fig. 9D present clustering of mouse cerebral cortex cells based on Paired- Damage-seq RNA profile. Fig. 9A presents dot plots showing the expression of marker genes for each mouse brain cell type measured from Paired-Damage-seq RNA profiles. The size of the dots represents the fraction of cells positively detect the transcripts and the color of the dots represents the average levels. Fig. 9B presents UMAP co-embedding of single nuclei transcriptomic profile from Paired-Damage-seq and reference snRNA-seq datasets on mouse motor cortex regions. Fig. 9C presents Heatmap showing the overlap coefficients between cell type annotations based on Paired-Damage-seq RNA profiles and the previously published snRNA-seq dataset. Fig. 9D presents line plots showing the normalized DNA signal enrichments of ATAC-seq, non-targeting tagmentation control and Paired-Damage-seq DNA signals on ATAC-seq peak regions of mouse brain.
[0061] Fig. lOA-Fig. IOC present distribution of DNA damage signals on coding genes, LINE1 and ERV elements. Fig. 10A presents line plots showing the DNA damage levels around genic regions of genes with different expression levels in each brain cell type, respectively. Fig. 10B presents line plots showing the DNA damage signals around long interspersed nuclear elements- 1 (LINE1) and endogenous retroviruses (ERV) elements in compartment A and compartment B, respectively, for different brain cell types. Fig. IOC presents violin plots showing the average detected signal levels in compartments A and B (RPKM, in 250-kb non-overlapping bins) for non-targeting tagmentation control and ATAC- seq of mouse brain; for all box plots, hinges were drawn from the 25th to 75th percentiles, with the middle line denoting the median, whiskers denoting a maximum 2x the interquartile range and outliers indicated with dots; n = 4089 (compartment A), n = 5095 (compartment B). Fig. IIA-Fig. 11D present distribution of DNA damage signal on enhancers in the mouse brain. Fig. IIA-Fig. 11B present scatter plots showing the relationships between DNA damage levels and (Fig. 11A) H3K27ac levels (RPKM, in 250-kb non-overlapping bins) in compartment A, and (Fig. 11B) H3K9me3 levels (RPKM, in 250-kb non-overlapping bins) in compartment B for inhibitory neurons, oligodendrocytes and microglia cells. Pearson correlation coefficients are also shown. Fig. 11C presents heatmap showing the DNA damage signals over public H3K27ac peaks in different mouse brain cell types. Fig. 11D presents Venn plots showing the overlaps between public ATAC-seq peaks (with and without H3K27ac peaks) and DNA damage peaks in different mouse brain cell types; P value, two- sided Fisher’s exact test.
[0062] Fig. 12A-Fig. 12D present relationships between DNA damage and epigenome erosion. Fig. 12A presents barplots showing the numbers of cell-type specific DNA damage peaks that could be mapped to hg38. The mapped peaks (reproducible, <lkb) were used for the GWAS trait enrichment analysis. Fig. 12B presents genome browser view of Tpcnl locus in microglia cells. DNA damage peaks overlapped with ATAC-seq peaks decreased in aged mice are highlighted in light blue. Signals from inhibitory neurons are also shown as a control. Fig. 12C-Fig. 12D present scatter plots showing the correlation of DNA damage levels (RPKM, in 100-kb non-overlapping bins) with change in H3K9me3 levels (RPKM, in 100-kb non-overlapping bins) between 18-month and 3-month for (Fig. 12C) inhibitory neuron and (Fig. 12D) oligodendrocyte precursor cells. Only the genomic bins with the top 1% highest damage levels were shown. Genomic bins with AH3K9me3 < -0.2 are shown in red, > 0.2 are shown in blue. Pearson correlation coefficients are indicated.
[0063] DETAILED DESCRIPTION OF THE INVENTION
[0064] Maintenance of genome integrity is paramount to all molecular programs in multicellular organisms. However, the genome is constantly challenged by various endogenous and environmental stresses throughout the lifespan. Understanding the functional consequences of DNA damage requires investigating their genomic distributions and influences on regulatory programs.
[0065] Emerging evidence suggests the crosstalk between DNA damage and epigenetic alterations. The DNA repair process exhibits a chromatin dynamic similar to DNA replication, which involves shared chaperons promoting the removal of pre-existing histones and deposition of new ones. The ‘relocalization of chromatin modifiers’ theory hypothesizes that chromatin factors move away from genes to damaged sites in response to DNA damage signaling and fail to return, resulting in progressive loss of youthful epigenetic patterns. In addition, DNA damage signaling could trigger chromatin decondensation, and reduced H3K9me2 levels facilitates homologous recombination repair. Furthermore, the DNA base excision repair pathway erases both oxidative damaged bases and the epigenetic cytosine methylation in mammalian cells. This process promotes extensive endogenous DNA damage in neuronal enhancers and safeguards cell identities, but could also lead to base mutations. Thus, it is not surprising that cells’ epigenetic landscapes reconfigure in response to formation and repair of DNA damage. However, whether DNA damage-associated epigenetic changes are transient or could persist long-term, and whether such alterations occur in a programmed manner remained unclear due to the lack of methods capable of analyzing DNA damage with cellular programs jointly at single-cell resolution.
[0066] Oxidized DNA base damage is a prevalent endogenous lesion; if accumulated, it can result in lesions such as apurinic and / or apyrimidinic sites, single-strand and double-strand DNA breaks (DSBs) and base mutations. Oxidative DNA damage is repaired by base excision repair in several steps: (1) DNA glycosylases cleave the N-glycosidic bond connecting damaged base and the deoxyribose, leaving an abasic site; (2) apyrimidinic endonucleases process the abasic site and generate a gap and (3) replacement synthesis by DNA polymerases incorporates new deoxynucleotide triphosphates (dNTPs) and DNA ligases seal the resulting nick. Sequencing technologies were developed to pinpoint the locations of oxidative DNA damage in bulk cell populations.
[0067] However, measuring the averaged signals obscures investigating the distribution preferences and functional consequences of DNA damage hotspots for different cell types. Recent advances in single-cell genomics facilitated the analysis of cell-type-specific molecular programs, including transcriptome and epigenome. Single-cell multiomics approaches, by jointly measuring multiple molecular types, empowered the dissection of the crosstalk across distinct molecular modalities. However, mapping DNA damage in single cells is challenging due to the inability to amplify damage signals. The inherent stochasticity of DNA damage formation further complicates the investigation of their relationships with other molecular modalities. Accordingly, the identification of DNA damage hotspots with pathological consequences is hindered by both the complex cell-type compositions within organs and the high background noise levels due to the stochasticity of damage formation.
[0068] To address these challenges, presented herein are methods comprising Paired- Damage-seq for joint analysis of DNA oxidative damage and repair intermediates with gene expression (transcriptome) in single cells. Further presented herein are exemplary applicatons of the methods to cancer cells (e.g., HeLa cells) and brain tissue (e.g., mouse cerebral cortex). Integrated analysis revealed the association between damage formation and epigenetic changes - selective genome vulnerability attributed to base sequence contents and local chromatin states. This selective genome vulnerability, in turn, can predict cell types and dysregulated molecular programs that contribute to disease risks.
[0069] Unique Features of the Invention
[0070] 1. A high-throughput single-cell method to map oxidative DNA damage simultaneously with gene expression from single cells. From 10,000 - 1,000,000 cells can be analyzed at a time.
[0071] 2. Enabled the analysis of DNA damage signals in different cell types and states from heterogeneous samples.
[0072] 3. Enabled the study of relationships between DNA damage and RNA expression directly from the same cells.
[0073] 4. Enabled the study of relationships between DNA damage and additional epigenome features by integrating with existing datasets or datasets generated with existing methods.
[0074] 5. It can be applied to tissue samples.
[0075] Accordingly, presented herein are methods comprising Paired-Damage- seq, which is a high-throughput method for single-cell joint analysis of DNA damage (e.g., oxidative DNA damage) and repair intermediates with gene expression for the analysis of DNA damage landscapes at single-cell resolution with cell identity-reserved. Analyzed herein are cultured HeLa cells and heterogeneous mouse brain tissues, which revealed selective genome and celltype vulnerability driven by base sequence contents and local epigenetic states. For the cultured HeLa cells, by introducing epigenome fluctuations with exogeneous oxidative stress, discovered herein is that the degree of epigenetic information loss is associated with the accumulation of inducible oxidative DNA damage. For the mouse brain, observations indicate that a subset of genome regions most susceptible to genotoxin are also hotspots for age-associated epigenome loss. In addition, such cell-type specific genome vulnerability could increase the risk of pathologic gene regulation programs by DNA damage accumulation and epigenome erosion over time.
[0076] It has been revealed by previously developed bulk sequencing methods that DNA damage, including 8-oxoG and AP sites, are not randomly distributed and open chromatin regions are more prone to be oxidized. Despite the different capturing approaches and genomes analyzed, the results presented herein agreed well with these consistent observations. However, analyzing DNA damage from bulk level can only reveal the most conserved hotspots across different cell types and states within the population, which are largely driven by local sequence content and the conserved / stable chromatin states. Indeed, for most dynamics processes, such as initiation and progression of age-related pathology, including cancer, the underlying chromatin programs are highly diverse even within the same tissues, necessitating the measurements at single-cell resolution. Based on single-cell wholegenome amplification, LCS-WGA can efficiently capture damage-associated singlenucleotide variants from individual neurons. The knowledge on DNA damage alone could not reveal cell identity nor identify the damaged loci associated with specific cellular state alternations. Paired-Damage- seq directly linked DNA damage with transcriptome, allowing the investigation of the influences of damage formation and repair on gene programs in different cell types from heterogeneous tissue samples.
[0077] Recent studies in single-cell dissection of human Alzheimer’s disease (AD) revealed the links between DNA damage and epigenome erosion during disease progression. These works used an estimation of damage burden by measuring chimeric gene fusions, which are formed by double- stranded breaks; however, DSBs per se is the most harmful form of DNA damage, and it is challenging to distinguish the consequences associated with the specific DSB-related molecular pathway from the decayed gene regulation programs. By consistently generate DSBs in mice can facilitate aging phenotype, which is reversable by epigenome reprogramming, supporting that the “relocalization of chromatin remodelers” could be the driving factor of epigenome erosion over time. However, this model expects the rates of epigenetic information loss to be generally uniform across the whole genome, but whether or not epigenetic erosion initiates in pathology-associated “pioneering” loci remained unclear. The data presented herein revealed cell-type specific DNA damage hotspots and suggested the relationships between DNA damage accumulation and epigenetic information loss in both in vitro cultured cells and primary tissue samples. With the Dictionary Learning integration103, Paired-Damage-seq datasets can be further bridged with additional modalities for the dissection of the relationships and cumulative impacts across these genomic layers.
[0078] The current Paired-Damage-seq protocol achieved high-throughput single-cell analysis with combinatorial barcoding; however, it could be readily adapted to droplet-based systems for fast and accessible DNA damage analyses from heterogeneous samples. The methods of the present disclosure can be used to study oxidative DNA damage as well as other DNA damage types by adopting different labeling approaches. In addition, the observed DNA damage hotspots are combined results of the DNA damage formation and repair. Similar to analyzing DNA methylation dynamics, measuring DNA repair together with DNA damage and transcriptome jointly further facilitates dissecting the dynamics of genome / epigenome erosion and the investigation of their functional impacts on cells’ molecular programs in diseases and during aging.
[0079] In some aspects, the methods and libraries described herein link DNA damage to gene expression changes and / or epigenetic changes. This methodology can provide information on how cells respond to genetic damage. The methods and libraries described herein can identify or predict disease or disorder risks, and / or lead to development of therapies to prevent or treat such disease or disorder.
[0080] Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments of the invention described herein. For example, various tagmentation techniques, barcoding techniques, sequencing techniques, or any other techniques that are known in the art can be used in the methods of the present disclosure.
[0081] Definitions
[0082] The articles “a” and “an” are used herein to refer to one or to more than one (z.e. to at least one) of the grammatical object of the article. By way of example, “an element” means one element or more than one element.
[0083] The “tag” or “adapter” (as used herein, also referred to as “adaptor”) may comprise a nucleic acid, a protein, and / or a small molecule. In some embodiments, a tag or an adapter may comprise a single-stranded or double- stranded oligonucleotide. In some embodiments, a tag or an adapter may be ligated to the end of a DNA or RNA molecule. In other embodiments, a tag or an adapter may be inserted within the DNA or RNA molecule. In certain embodiments, the tag or adapter may comprise a barcode and / or a restriction site that can be recognized and cleaved by an endonuclease.
[0084] The term “sequencing adapter” (also referred to as “sequencing adaptor”) is art- recognized. If typically refers to a short DNA sequence that is ligated to the ends of DNA fragments during library preparation for next-generation sequencing (NGS). They serve as binding sites for primers and the flow cell during sequencing, enabling the amplification and identification of DNA fragments. These adapters also often include index sequences (barcodes) for multiplexing and molecular identifiers for error corrections.
[0085] The term “TSO” (also referred to as “template- switching oligonucleotide”) is art- recognized. TSO allows tagging the 5 '-end of captured mRNAs using a reverse transcriptase enzyme with terminal transferase activity. Sequencing for low-input samples can use utilize template- switching reverse transcription. Template switching permits ligation-free incorporation of a 5 '-adapter during reverse transcription.
[0086] All numerical ranges provided herein are understood to be shorthand for all of the decimal and fractional values within the range. For example, a range of 1 to 50 is understood to include any number, combination of numbers, or sub-range from the group consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50, as well as all intervening decimal values between the aforementioned integers such as, for example, 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, and 1.9 and all intervening fractional values between the aforementioned integers such as, for example, 1 / 2, 1 / 3, 1 / 4, 1 / 5, 1 / 6, 1 / 8, and 1 / 9, and all multiples of the aforementioned values. With respect to sub-ranges, "nested sub-ranges" that extend from either end point of the range are specifically contemplated. For example, a nested sub-range of an exemplary range of 1 to 50 may comprise 1 to 10, 1 to 20, 1 to 30, and 1 to 40 in one direction, or 50 to 40, 50 to 30, 50 to 20, and 50 to 10 in the other direction.
[0087] DNA damage
[0088] The methods of present disclosure can be used to analyze and map various types of DNA damage in the genome, including but not limited to oxidative DNA damage, misincorporated deoxyuridines, DNA double-strand breaks (DSBs), DNA single-strand breaks, or any combination thereof. In particular, the methods described herein are useful in identifying DNA damage (which may be optionally processed to a form or an intermediate that can be detectably labeled), which can be incorporated into and / or used in combination with the methods of the present disclosure, e.g., Paired-Damage- seq method. See also Bai el al. (2025) Nature Methods 22:962-972, which is incorporated herein by reference in its entirety.
[0089] Oxidative DNA damage
[0090] Base oxidation is a frequent insult that can arise from endogenous metabolic processes as well as from exogenous sources such as ionizing radiation. At background levels, a human cell is estimated to undergo 100 to 500 such modifications per day, most commonly resulting in 8-oxo-7,8-dihydroguanine (8-oxoG) and related products, which are then processed into repair intermediates.
[0091] Oxidative damage is processed by cells in a two-step process through the base excision repair (BER) pathway. The damaged base is first recognized and excised by 8- oxoguanine DNA glycosylase 1 (OGGI), leaving an apurinic site (AP-site). Glycohydrolysis is highly efficient, with an 8-oxoG half-life of 11 min. AP-sites are removed through backbone incision by AP-lyase (APEX1), and end processing through flap-endonuclease 1 (FEN1), and the base is subsequently replaced with an undamaged nucleotide. Alternatively, in short-patch base excision repair, replacement is dependent on polymerase beta. Other sources of AP-sites include spontaneous depurination and excision of non-oxidative base modifications, such as uracil. Cells are reported to typically present with a steady state of ~ 15,000 to ~ 30,000 AP-sites per cell, which includes the associated beta-elimination product. Left unrepaired, 8-oxoG can compromise transcription, DNA replication, and telomere maintenance. Also, AP-sites can lead to genomic instability and compromise genomic processes. Moreover, damaged sites provide direct and indirect routes to C-to-A mutagenesis.
[0092] Oxidative DNA damage can be identified, processed, and / or mapped in the genome by the methods described herein. Alternatively, AP-seq can also be used to process oxidative DNA damage and incorporated into the methods of the present disclosure (see Poetsch et al. (2018) Genome biology, 19:215, which is incorporated herein by reference in its entirety). Deoxyuridine
[0093] Uracil in DNA can be generated by cytosine deamination or dUMP misincorporation; however, its distribution in the human genome is poorly understood. Deoxyuridines can be labeled selectively and enriched using pull-down technology (see Shu et al. (2018) Nat Chem Biol, 14(7):680-687, which is incorporated herein by reference in its entirety) for genome-wide uracil profiling.
[0094] Double-strand breaks (DSBs) & Single-strand breaks (SSBs)
[0095] DNA double-strand breaks can be caused by exogenous and endogenous physical and chemical agents, and appear during apoptosis, meiotic crossing-over, and gene rearrangements. Replication fork stalling and collapse also cause DSBs, and are considered the major endogenous source of breaks in cycling cells. Unresolved DSBs pose a serious threat to genomic stability, potentially leading to the formation of oncogenic mutations, including translocations, deletions, and amplifications. Despite extensive knowledge on mechanisms of DSBs sensing and repair, the genome-wide landscape of DSBs in different cells and conditions remains largely unknown. Exemplary methods for identifying DSBs in the genomes include chromatin immunoprecipitation (ChIP) coupled to microarray (ChlP- chip) and next-generation sequencing (ChlP-seq). Certain methods may also utilize phosphorylated histone variant H2A.X (yH2A.X) as a marker of DSBs. Alternatively, the recruitment of replication protein A (RPA) can be used to map DNA damage. Other approaches include those involving capturing or direct labeling of single-stranded DNA (ssDNA) followed by microarray analysis, assuming ssDNA is a good proxy for DSBs. Methods based on labeling breaks with the terminal deoxynucleotydil transferase (TdT) enzyme can be used to detect DSBs at defined locations in vitro and in purified genomic DNA. Experimental and computational approach to directly map DSBs genome-wide, based on direct in situ breaks labeling, enrichment on streptavidin, and next-generation sequencing (BLESS) is described by Crosetto et al. (2013) Nat Methods, 10(4):361-365. Additional methods, DSB-Seq and SSB-Seq (see Baranello et al. (2018) Protocol, 1672:155-166), can also be incorporated into the methods of the present disclosure. Both Crosetto and Baranello are incorporated herein by reference in their entirety.
[0096] Labeling DNA Damage
[0097] The methods of the present disclosure comprise labeling DNA damage. For example, as described above, DNA breaks can be labeled with labeled nucleotides using enzymes such as TdT. Alternatively, the sites of DNA may be labeled using an antibody that binds to the markers of DNA damage, e.g., DNA repair proteins or phosphorylated histone variant H2A.X (yH2A.X).
[0098] Exemplary methods for labeling DNA damage comprising oxidized bases are provided herein (e.g., see above and working Examples).
[0099] Exemplary methods for labeling DNA damage comprising DSBs may comprise:
[0100] 1) contacting the one or more nuclei with Klenow fragment DNA polymerase and T4 Polynucleotide kinase, followed by incubating with the chemically synthesized doublestranded DNA adaptor comprising a label (e.g., biotin), to label the DSB sites; and
[0101] 2) performing the targeted tagmentation with an anti-label antibody (e.g., an anti-biotin antibody) and protein A-Tn5, followed by reverse transcription and barcoding as described herein.
[0102] In preferrred embodiments, DNA damage is labeled by tagmentation or a variant thereof. Tagmentation involves an enzymatic reaction where a transposome (a complex of transposase enzyme and DNA adapter) fragments DNA and simultaneously appends sequencing adapters to the fragments. This process replaces traditional methods that involve separate steps for fragmentation, end-repair, and ligation.
[0103] For example, tagmentation comprises an exemplary transposase Tn5 (or a hyperactive variant thereof), which forms a complex with an oligonucleotide comprising a doublestranded 19-bp mosaic end sequence recognized by Tn5 transposase, as well as a single- stranded overhang on the transfer strand that comprises an adapter used for subsequent processing. This ssDNA overhang can be any length; however, shorter sequences improve efficiency of tagmentation. The Tn5 enzyme is loaded with a mix of adapters with forward or reverse adapter overhangs. For standard tagmentation workflows, this includes a 1:1 molar ratio of the two adapter species and a 1 : 1 molar ratio of the total adapter content to Tn5 monomer. The tagmentation reaction involves the binding of transposome complexes to the target DNA at high density, e.g., one insertion every ~500 bp. Each tagmentation event results in the cleavage of the DNA backbone on both strands staggered by 9 bp. The 3' end of the transfer strand is then covalently appended to the 5' end of the nick in the target DNA backbone at each of the cut sites. After the tagmentation, the Tn5 enzyme remains tightly bound to the target DNA and must be removed to enable end repair, where the bottom strand acts as a priming site to copy through the adapter on the transfer strand. End repair results in the copying of the 9-bp region between the two cuts in the target DNA backbone, where the pair of adjacent fragments produced by a single tagmentation event overlap by the 9-bp segment when aligned to a reference genome.
[0104] Tagmentation can be targeted (also referred to as targeted tagmentation), which targets the transposome complex to certain sites in the genome (e.g., to sites of DNA damage). Targeted tagmentation is commonly used in methods like CUT&Tag (Cleavage Under Targets and Tagmentation). Various tags used to label the target site and its binding partner can be used. For example, for a DNA damage site labeled with dUTP, an antibody against dUTP can be used in combination with a protein A-Tn5 fusion protein to target Tn5 to the DNA damage site comprising dUTP.
[0105] Barcoding
[0106] Nucleic acids (e.g., DNA, cDNA, RNA, genomic DNA comprising DNA damage, etc.) can be “barcoded”, which is a method of species identification using a certain short genetic sequence, analogous to a Universal Product Code (UPC) barcode, to distinguish one species from another. Various barcoding methods are known in the art (reviewed by e.g., Serrano et al. (2022) Nature Reviews Cancer 22:609-624).
[0107] In preferred embodiments, methods of the present disclosure comprise single-cell barcoding.
[0108] In some embodiments, single-cell barcoding is performed after physically separating individual cells. For example, a flow containing diluted single-cell suspensions and another flow containing barcoded beads are merged in the flow focus, which continuously generates a droplet of a single cell with a unique barcode (Macosko et al. (2015) Cell 161:1202-1214; Prakadan et al. (2017) Nat. Rev. Genet. 18:345-361, each of which is incorporated herein by reference in it entirety). Here, the droplets (nanoliter- scale aqueous compartments formed by precisely combining aqueous and oil flows in a microfluidic device) serve as tiny reaction chambers for PCR and reverse transcription. Although droplet generators limit the physical isolation of multiple molecules (DNA, RNA, or protein), capturing by unique tags (adapters) on the bead enables demultiplexing and simultaneous separation of each tag.
[0109] In other embodiments, single-cell barcoding is performed without physically isolating cells. For example, in a method referred to as ligation-based combinatorial barcoding, split- and-pooling techniques (Split&Pooling) generate single-cell barcodes in the cell mixture by combinatorial chemistry (examples include those disclosed in Cao et al. (2017) Science 357:661-667; Rosenberg et al. (2018) Science 360:176-18, each of which is incorporated herein by reference in it entirety). The barcode process starts with permeabilized but preserved intact cells being deposited into several wells in a mixture. Molecules are uniquely barcoded by adapter ligation in a cell mixture, and they are then pooled and split into several wells followed by another round of unique barcode ligation. Two to three rounds of this process enable enough complexity to cover unique barcodes for all single cells. Because this method works with the ligation of nucleic acid barcodes, it supports the ligation of unique barcodes for different molecules simultaneously.
[0110] Nucleic Acid Library
[0111] Various methods for preparing nucleic acid libraries (e.g., DNA, cDNA, RNA, genomic DNA comprising DNA damage, etc.) are well known in the art, some of which are reviewed at least by Hess et al. (2020) Biotechnology Advances, 41 : 107537, which is incorporated herein by reference in its entirety. Preparation of nucleic acid library allows efficient downstream processing, including next generation sequencing by converting nucleic acid samples into a library of uniformly sized, adapter-ligated DNA fragments.
[0112] A conventional library construction may comprise the following steps: fragmentation of DNA, if necessary; end repair of fragmented DNA; addition of adapter sequences; and optional library amplification.
[0113] Four different methods are commonly used to generate fragmented nucleic acid (e.g., gDNA): enzymatic digestion, sonication, nebulization, and hydrodynamic shearing. Endonucleolytic digestion is easy and fast, but it is often difficult to accurately control the fragment length distribution. The other three methods employ physical methods to introduce double- strand breaks into DNA, which are believed to occur randomly resulting in an unbiased representation of the DNA in the library. The resulting DNA fragment size distribution can be controlled by agarose gel electrophoresis or automated DNA analysis. However, the fragmentation step can be omitted if fragments of desired length are obtained in a method preceding the library preparation.
[0114] Once the nucleic acid fragments are generated, a sequencing adapter can be added to the fragments by either ligation-based method or a tagmentation-based method as described below.
[0115] Ligation-based library preparation
[0116] Following fragmentation, the nucleic acids are repaired to generate blunt-ended, 5'- phosphorylated DNA ends compatible with the sequencing platform- specific adapter ligation strategy. Often, the end repair is accomplished by exploiting the 5'— 3' polymerase and the 3'- 5' exonuclease activities of T4 DNA polymerase, while T4 Polynucleotide Kinase ensures the 5'-phoshorylation of the blunt-ended DNA fragments, preparing these fragments for subsequent adapter ligation. Alternatively, other methods known in the art as well as those described herein can be used.
[0117] Depending on the sequencing platform used, the blunt-ended nucleic acid fragments can either directly be used for adapter-ligation, or need the addition of a single A overhang at the 3' ends of the DNA fragments to facilitate subsequent ligation of platform- specific adapters with compatible single T overhangs. Often, this A-addition step is catalyzed by Klenow Fragment (minus 3' to 5' exonuclease) or other polymerases with terminal transferase activity.
[0118] T4 DNA ligase or the like then adds the double- stranded adapters to the end-repaired library fragments, followed by an optional reaction of cleanup and nucleic acid size selection to remove free library adapters and adapter dimers. The methods for size selection may include agarose gel isolation, the use of magnetic beads, or advanced column-based purification methods. Adapter-dimers that can occur during the ligation and will subsequently be co-amplified with the adapter-ligated library fragments should be removed from the libraries before sequencing, as they may reduce the capacity of the sequencing platform for real library fragments and may reduce sequencing quality. Some sequencing platforms require a narrow distribution of library fragments for optimal results, which in many cases can only be achieved by excising the respective fragment section from the gel. This can also serve to deplete adapter-dimers. Tagmentation-based library preparation
[0119] As described in the tagmentation section above, a transposase such as Tn5 or a variant thereof can add adapters to the nucleic acid fragments in a single step.
[0120] Nucleic acid fragment libraries, once obtained by either the ligation-based method or tagmentation-based method, can be qualified and quantified. Depending on the concentration and adapter design of the sequencing library, it can either be directly diluted and used for sequencing, or subjected to optional library amplification. In the library amplification step, high-fidelity DNA polymerases can be employed to either generate the entire adapter sequence needed for subsequent clonal amplification and binding of sequencing primers, with overlapping PCR primers, and / or to produce higher yields of the DNA libraries. Optimal library amplification requires DNA polymerase with high fidelity and minimal sequence bias.
[0121] To enable efficient use of the sequencing capacity, sequencing libraries generated from different samples can be pooled and sequenced in the same sequencing run. This is enabled by ligation DNA fragments to adaptors with characteristic barcodes, i.e., short stretches of nucleotide sequences that are distinct for each sample.
[0122] In some embodiments, during the nucleic acid library preparation, certain species of nucleic acid (e.g., cDNA or genomic DNA) can be enriched by digestion with one or more restriction enzyme(s), whose recognition sites were engineered into the desired nucleic acid sequences. Such a method is presented herein (see e.g., Examples 1 and 9, Fig. IL), as well as in e.g., Extended data Fig. 1 of Zhu et al. (2021) Nat Methods, 18(3):283-292, which is incorporated herein by reference in its entirety.
[0123] Sequencing
[0124] Any of a variety of sequencing reactions known in the art can be used to directly sequence the nucleic acids. Examples of sequencing reactions include those based on techniques developed by Maxam and Gilbert (1977) Proc. Natl. Acad. Sci. USA 74:560 or Sanger (1977) Proc. Natl. Acad Sci. USA 74:5463. It is also contemplated that any of a variety of automated sequencing procedures can be utilized (Naeve (1995) Biotechniques 19:448-53), including sequencing by mass spectrometry (see, e.g., PCT International Publication No. WO 94 / 16101; Cohen et al. (1996) Adv. Chromatogr. 36:127-162; and Griffin et al. (1993) Appl. Biochem. Biotechnol. 38:147-159). Notably, mass spectrometry (e.g., LC-MS, LC-MS / MS) may be used to sequence DNA (see Chowdhury and Guengerich (2013) Curr Protoc Nucleic Acid Chem 7:Unit-7.1611) or RNA (Zhang et al. (2019) Nucleic Acids Research 47:el25). In certain embodiments, sequencing of nucleic acids can be accomplished using methods including, but not limited to, sequencing by hybridization (SBH), sequencing by ligation (SBL), quantitative incremental fluorescent nucleotide addition sequencing (QIFNAS), pyrosequencing, fluorescent in situ sequencing (FISSEQ), FISSEQ beads (U.S. Pat. No. 7,425,431), wobble sequencing (PCT / US05 / 27695), multiplex sequencing (U.S. Ser. No. 12 / 027,039, filed Feb. 6, 2008; Porreca et al. (2007) Nat. Methods 4:931), polymerized colony (POLONY) sequencing (U.S. Pat. Nos. 6,432,360, 6,485,944 and 6,511,803, and PCT / US05 / 06425); nanogrid rolling circle sequencing (ROLONY) (U.S. Ser. No. 12 / 120,541, filed May 14, 2008), and the like. High- throughput sequencing methods, e.g., on cyclic array sequencing using platforms such as Roche 454, Illumina Solexa or MiSeq or HiSeq, AB-SOLiD, Helicos, Polonator platforms and the like, can also be utilized. High-throughput sequencing methods are described in U.S. Ser. No. 61 / 162,913, filed Mar. 24, 2009. A variety of light-based sequencing technologies are known in the art (Landegren et al. (1998) Genome Res. 8:769-76; Kwok (2000) Pharmocogenom. 1:95-100; and Shi (2001) Clin. Chem. 47:164-172) (see, for example, U.S. Pat. Publ. Nos. 2013 / 0274117, 2013 / 0137587, and 2011 / 0039304).
[0125] Next-generation sequencing (NGS), which is also called “massively-parallel sequencing”, enables the sequencing of many DNA strands at the same time, instead of one at a time as with traditional Sanger sequencing by capillary electrophoresis (CE). Because of the speed, throughput, and accuracy of NGS, NGS enables the interrogation of hundreds to thousands of nucleic acid molecules at one time in multiple samples, as well as discovery and analysis of different types of genomic features in a single sequencing run, from single nucleotide variants (SNVs), to copy number and structural variants, and even RNA fusions. NGS provides the ideal throughput per run, and studies can be performed quickly and cost- effectively. Additional advantages of NGS include lower sample input requirements, higher accuracy, and ability to detect variants at lower allele frequencies than with Sanger sequencing.
[0126] Analyzing the whole genome using next-generation sequencing (NGS) delivers a base-by-base view of all genomic alterations, including single nucleotide variants (SNV), insertions and deletions, copy number changes, and structural variations. Paired-end wholegenome sequencing involves sequencing both ends of a DNA fragment, which increases the likelihood of alignment to the reference and facilitates detection of genomic rearrangements, repetitive sequences, and gene fusions. In some embodiments, the Illumina “Phased Sequencing” platform, which employs a combination of long and short pair-ends, can be used. In other embodiments, the third- generation single-molecule sequencing technologies (e.g., ONT and PacBio) can produce much longer reads of DNA sequences.
[0127] In some embodiments, the “Deep Sequencing” or high-coverage version of Illumina NGS can be used. Deep Sequencing refers to sequencing a sample multiple times, sometimes hundreds or even thousands of times. The Deep Sequencing allows detection of rare clonal types, cells, or microbes comprising as little as 1% of the original sample. Illumina’s NovaSeq performs such whole-genome sequencing efficiently and cost-effectively, and its scalable output generates up to 6 Tb and 20 billion reads in dual flow cell mode with simple streamlined automated workflows.
[0128] Other Methods for Genome-Wide Studies
[0129] The methods of the present disclosure (e.g., Paired-Damage-seq) can be used in combination with any one or more methods that analyze the features (e.g., gene expression, DNA methylation, histone modification, chromatin accessibility and structure) of a cell, preferably a single cell.
[0130] In some embodiments, the methods of the present disclosure may be used in combination with any one or more selected from: ChlP-chip (ChIP on chip; use of microarrays to detect DNA fragments enriched by chromatin immunoprecipitation (ChIP)), ChlP-seq (Chromatin Immunoprecipitation sequencing), Paired-Tag (Zhu et al. (2021) Nat Methods, 18(3): 283-292, which is incorporated herein by reference in its entirety), G&T-seq, DR-seq, SIDR, TARGET-seq, scTrio-seq, scM&T-seq, scMT-seq, sci-CAR, SNARE-seq, Paired-seq, scNMT-seq, PEA / STA, PLAYR, CITE-seq, REAP-seq, RAID, and ECCITE-seq (reviewed by Lee et al. (2020) Experimental & Molecular Medicine, 52:1428-1442, which is incorporated herein by reference in its entirety).
[0131] Agents
[0132] One or more nuclei, or one or more cells from which the one or more nuclei are derived, may be treated with an agent prior to the methods of the present disclosure. In some embodiments, the agent alters the cell dynamics (e.g., chromatin remodeling, gene expression, inducing DNA damage, cell behavior, cell differentiation state, etc.). In some embodiments, the agent may be selected from a toxin, a DNA damage inducing agent (e.g., H2O2, reactive oxygen species (ROS)), a chemotherapeutic agent, a small molecule, an alkylating agent, a hormone, a cell signaling molecule, a cytokine, a chemokine, and any combination or two or more thereof.
[0133] A chemotherapeutic agent may be, but is not limited to, those selected from among the following groups of compounds: platinum compounds, cytotoxic antibiotics, antimetabolites, anti-mitotic agents, alkylating agents, arsenic compounds, DNA topoisomerase inhibitors, taxanes, nucleoside analogues, plant alkaloids, and toxins; and synthetic derivatives thereof. Exemplary compounds include, but are not limited to, alkylating agents: cisplatin, treosulfan, and trofosfamide; plant alkaloids: vinblastine, paclitaxel, docetaxol; DNA topoisomerase inhibitors: teniposide, crisnatol, and mitomycin; anti-folates: methotrexate, mycophenolic acid, and hydroxyurea; pyrimidine analogs: 5- fluorouracil, doxifluridine, and cytosine arabinoside; purine analogs: mercaptopurine and thioguanine; DNA antimetabolites: 2'-deoxy-5-fluorouridine, aphidicolin glycinate, and pyrazoloimidazole; and antimitotic agents: halichondrin, colchicine, sorafenib, doxorubicin, and rhizoxin. Compositions comprising one or more chemotherapeutic agents (e.g., FLAG, CHOP) may also be used. FLAG comprises fludarabine, cytosine arabinoside (Ara-C) and G- CSF. CHOP comprises cyclophosphamide, vincristine, doxorubicin, and prednisone. In another embodiments, PARP (e.g., PARP-1 and / or PARP-2) inhibitors are used and such inhibitors are well-known in the art (e.g., Olaparib, ABT-888, BSI-201, BGP-15 (N-Gene Research Laboratories, Inc.); INO-lOOl (Inotek Pharmaceuticals Inc.); PJ34 (Soriano et al., 2001; Pacher et al., 2002b); 3-aminobenzamide (Trevigen); 4-amino-l,8-naphthalimide; (Trevigen); 6(5H)-phenanthridinone (Trevigen); benzamide (U.S. Pat. Re. 36,397); and NU1025 (Bowman et al.). The mechanism of action is generally related to the ability of PARP inhibitors to bind PARP and decrease its activity. PARP catalyzes the conversion of .beta.-nicotinamide adenine dinucleotide (NAD+) into nicotinamide and poly-ADP-ribose (PAR). Both poly (ADP-ribose) and PARP have been linked to regulation of transcription, cell proliferation, genomic stability, and carcinogenesis (Bouchard V. J. et.al. Experimental Hematology, Volume 31, Number 6, June 2003, pp. 446-454(9); Herceg Z.; Wang Z.-Q. Mutation Research / Fundamental and Molecular Mechanisms of Mutagenesis, Volume 477, Number 1, 2 Jun. 2001, pp. 97-110(14)). Poly(ADP-ribose) polymerase 1 (PARP1) is a key molecule in the repair of DNA single-strand breaks (SSBs) (de Murcia J. et al. 1997. Proc Natl Acad Sci USA 94:7303-7307; Schreiber V, Dantzer F, Ame J C, de Murcia G (2006) Nat Rev Mol Cell Biol 7:517-528; Wang Z Q, et al. (1997) Genes Dev 11:2347-2358). Knockout of SSB repair by inhibition of PARP1 function induces DNA double-strand breaks (DSBs) that can trigger synthetic lethality in cancer cells with defective homology-directed DSB repair (Bryant H E, et al. (2005) Nature 434:913-917; Farmer H, et al. (2005) Nature 434:917-921). The foregoing examples of chemotherapeutic agents are illustrative, and are not intended to be limiting.
[0134] Cancer
[0135] As described herein, the methods of the present disclosure may involve one or more nuclei of a cancer cell. Cancer, tumor, or hyperproliferative disease refer to the presence of cells possessing characteristics typical of cancer-causing cells, such as uncontrolled proliferation, immortality, metastatic potential, rapid growth and proliferation rate, and certain characteristic morphological features. In some embodiments, cancer cells are highly differentiated. In other embodiments, cancer cells are poorly differentiated. Cancer cells are often in the form of a tumor, but such cells may exist alone within an animal, or may be a non-tumorigenic cancer cell, such as a leukemia cell. Cancers include, but are not limited to, B cell cancer, (e.g., multiple myeloma, Diffuse large B-cell lymphoma (DLBCL), Follicular lymphoma, Chronic lymphocytic leukemia (CEE), small lymphocytic lymphoma (SLL), Mantle cell lymphoma (MCL), Marginal zone lymphomas, Burkitt lymphoma, Waldenstrom's macroglobulinemia, Hairy cell leukemia, Primary central nervous system (CNS) lymphoma, Primary intraocular lymphoma, the heavy chain diseases, such as, for example, alpha chain disease, gamma chain disease, and mu chain disease, benign monoclonal gammopathy, and immunocytic amyloidosis), T cell cancer (e.g., T- lymphoblastic lymphoma / leukemia, non-Hodgkin lymphomas, Peripheral T-cell lymphomas, Cutaneous T-cell lymphomas (e.g., mycosis fungoides, Sezary syndrome), Adult T-cell leukemia / lymphoma, Angioimmunoblastic T-cell lymphoma, Extranodal natural killer / T-cell lymphoma, Enteropathy-associated intestinal T-cell lymphoma (EATL), Anaplastic large cell lymphoma (ALCL), Hodgkin lymphoma), melanomas, breast cancer, lung cancer, bronchus cancer, colorectal cancer, prostate cancer, pancreatic cancer, stomach cancer, ovarian cancer, urinary bladder cancer, brain or central nervous system cancer, peripheral nervous system cancer, esophageal cancer, cervical cancer, uterine or endometrial cancer, cancer of the oral cavity or pharynx, liver cancer, kidney cancer, testicular cancer, biliary tract cancer, small bowel or appendix cancer, salivary gland cancer, thyroid gland cancer, adrenal gland cancer, osteosarcoma, chondrosarcoma, cancer of hematologic tissues, and the like. Other nonlimiting examples of types of cancers applicable to the methods encompassed by the present invention include human sarcomas and carcinomas, e.g., fibrosarcoma, myxosarcoma, liposarcoma, chondrosarcoma, osteogenic sarcoma, chordoma, angiosarcoma, endotheliosarcoma, lymphangiosarcoma, lymphangioendotheliosarcoma, synovioma, mesothelioma, Ewing's tumor, leiomyosarcoma, rhabdomyosarcoma, colon carcinoma, colorectal cancer, pancreatic cancer, breast cancer, ovarian cancer, prostate cancer, squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland carcinoma, sebaceous gland carcinoma, papillary carcinoma, papillary adenocarcinomas, cystadenocarcinoma, medullary carcinoma, bronchogenic carcinoma, renal cell carcinoma, hepatoma, hepatocellular carcinoma (HCC), bile duct carcinoma, liver cancer, choriocarcinoma, seminoma, embryonal carcinoma, Wilms' tumor, cervical cancer, bone cancer, brain tumor, testicular cancer, lung carcinoma, small cell lung carcinoma (SCLC), bladder carcinoma, epithelial carcinoma, glioma, astrocytoma, medulloblastoma, craniopharyngioma, ependymoma, pinealoma, hemangioblastoma, acoustic neuroma, oligodendroglioma, meningioma, neuroblastoma, retinoblastoma; leukemias, e.g., acute lymphocytic leukemia and acute myelocytic leukemia (myeloblastic, promyelocytic, myelomonocytic, monocytic and erythroleukemia); chronic leukemia (chronic myelocytic (granulocytic) leukemia and chronic lymphocytic leukemia); and polycythemia vera, lymphoma (Hodgkin's disease and non-Hodgkin's disease), multiple myeloma, Waldenstrom's macroglobulinemia, and heavy chain disease. In some embodiments, cancers are epithlelial in nature and include but are not limited to, bladder cancer, breast cancer, cervical cancer, colon cancer, gynecologic cancers, renal cancer, laryngeal cancer, lung cancer, oral cancer, head and neck cancer, ovarian cancer, pancreatic cancer, prostate cancer, or skin cancer. In other embodiments, the cancer is breast cancer, prostate cancer, lung cancer, or colon cancer. In still other embodiments, the epithelial cancer is non- small-cell lung cancer, nonpapillary renal cell carcinoma, cervical carcinoma, ovarian carcinoma e.g., serous ovarian carcinoma), or breast carcinoma. The epithelial cancers may be characterized in various other ways including, but not limited to, serous, endometrioid, mucinous, clear cell, Brenner, or undifferentiated.
[0136] Exemplary Embodiments
[0137] In some aspects, provided herein are methods of determining one or more DNA damage sites in the genome and obtaining gene expression information of a cell or a nucleus, the methods comprising: contacting one or more nuclei with (i) optionally, means for removing an oxidized base and / or means for generating a single-strand DNA break (SSB) at the DNA damage sites; (ii) means for labeling the DNA damage sites; and (iii) means for reverse transcribing RNA in the one or more nuclei. In some embodiments, prior to the contacting steps, the methods comprise disrupting or depleting nucleosomes from the one or more nuclei. In some embodiments, the methods further comprise sequencing the DNA and RNA of the one or more nuclei, thereby determining the DNA damage site or sites in the genome and obtaining the gene expression information. In some embodiments, prior to the contacting steps, the one or more nuclei are treated with an agent or obtained from one or more cells treated with an agent that causes DNA damage. In some embodiments, the one or more nuclei is a single nucleus.
[0138] In some aspects, provided herein are methods of determining a DNA damage site in the genome and obtaining gene expression information of a cell or a nucleus, the methods comprising: obtaining one or more nuclei; disrupting or depleting nucleosomes from the one or more nuclei; contacting the one or more nuclei with an enzyme that removes an oxidized base; contacting the one or more nuclei with an apurinic / apyrimidinic (AP) endonuclease, thereby generating a single-strand DNA break (SSB) at the DNA damage site; labeling the DNA damage site, comprising: (i) incorporating a labeled nucleotide to the DNA break, (ii) contacting the one or more nuclei with a protein that binds the labeled nucleotide, and (iii) performing a tagmentation (e.g., CUT&Tag), which generates a genomic DNA fragment comprising a first tag, wherein the first tag comprises a first restriction site and a first barcode; and reverse transcribing RNA in the one or more nuclei using a primer comprising a second tag, thereby generating cDNA comprising the second tag, wherein the second tag comprises a second restriction site and a second barcode.
[0139] Further provided herein are following exemplary embodiments.
[0140] 1. A method of determining DNA damage sites in the genome and obtaining gene expression information of a cell or a nucleus, the method comprising: a) obtaining one or more nuclei; b) optionally fixing the one or more nuclei; c) disrupting or depleting nucleosomes from the one or more nuclei; d) contacting the one or more nuclei with an enzyme that removes an oxidized base; e) contacting the one or more nuclei with an apurinic / apyrimidinic (AP) endonuclease, thereby generating a single-strand DNA break (SSB) at the DNA damage site; f) labeling the DNA damage site, comprising: i) incorporating a labeled nucleotide to the DNA break; ii) contacting the one or more nuclei with a protein that binds the labeled nucleotide; and iii) performing a tagmentation (e.g., CUT&Tag), which generates a genomic DNA fragment comprising a first tag, wherein the first tag comprises a first restriction site and a first barcode; g) reverse transcribing RNA in the one or more nuclei using a primer comprising a second tag, thereby generating cDNA comprising the second tag, wherein the second tag comprises a second restriction site and a second barcode.
[0141] 2. The method of embodiment 1, further comprising: barcoding the genomic DNA fragment and / or cDNA with at least one additional tag comprising a barcode, thereby generating a genomic DNA fragment comprising the first tag and at least one additional tag, and / or cDNA comprising the second tag and at least one additional tag.
[0142] 3. The method of embodiment 1 or 2, further comprising: a) optionally lysing the one or more nuclei; b) adding a polynucleotide tail to the genomic DNA fragments and cDNA; and c) amplifying the genomic DNA fragments and / or cDNA using primers, wherein one of the primers comprises a sequence complementary to the polynucleotide tail, or a portion thereof, and a third restriction site.
[0143] 4. The method of any one of embodiments 1-3, further comprising: a) dividing the nucleic acid comprising the genomic DNA fragments and cDNA into a DNA library and an RNA library; b) for the DNA library: i) enriching for the genomic DNA fragments; and ii) contacting genomic DNA fragments with a sequencing adaptor and a ligase to generate the genomic DNA fragments comprising the sequencing adaptor; c) for the RNA library: i) enriching for the cDNA; and ii) performing a tagmentation (e.g., CUT&Tag), which generates the cDNA comprising a sequencing adaptor.
[0144] 5. The method of any one of embodiments 1-4, further comprising: sequencing the DNA molecules in the DNA library and RNA library, thereby determining the DNA damage sites in the genome and obtaining the gene expression information of the single nucleus. 6. The method of any one of embodiments 1-5, wherein the method is for determining DNA damage sites in the genome and obtaining gene expression information in a single nucleus or a single cell.
[0145] 7. The method of any one of embodiments 1-6, wherein the method comprises: a) obtaining one or more nuclei; b) optionally fixing the one or more nuclei; c) disrupting or depleting nucleosomes from the one or more nuclei; d) contacting the one or more nuclei with an enzyme that removes an oxidized base; e) contacting the one or more nuclei with an apurinic / apyrimidinic (AP) endonuclease, thereby generating a single-strand DNA break (SSB) at the DNA damage site; f) labeling the DNA damage site, comprising: i) incorporating a labeled nucleotide to the DNA break; ii) contacting the one or more nuclei with a protein that binds the labeled nucleotide; and iii) performing a tagmentation (e.g., CUT&Tag), which generates a genomic DNA fragment comprising a first tag, wherein the first tag comprises a first restriction site and a first barcode; g) reverse transcribing RNA in the one or more nuclei using a primer comprising a second tag, thereby generating cDNA comprising the second tag, wherein the second tag comprises a second restriction site and a second barcode; h) barcoding the genomic DNA fragment and / or cDNA with at least one additional tag comprising a barcode, thereby generating a genomic DNA fragment comprising the first tag and at least one additional tag, and / or cDNA comprising the second tag and at least one additional tag; i) lysing the one or more nuclei; j) adding a polynucleotide tail to the genomic DNA fragments and cDNA; k) amplifying the genomic DNA fragments and / or cDNA using primers, wherein one of the primers comprises a sequence complementary to the polynucleotide tail, or a portion thereof, and a third restriction site; l) dividing the nucleic acid comprising the genomic DNA fragments and cDNA into a DNA library and an RNA library; m) for the DNA library: i) enriching for the genomic DNA fragments; and iv) contacting genomic DNA fragments with a sequencing adaptor and a ligase to generate the genomic DNA fragments comprising the sequencing adaptor; n) for the RNA library: i) enriching for the cDNA; and iv) performing a tagmentation (e.g., CUT&Tag), which generates the cDNA comprising a sequencing adaptor; o) sequencing the DNA molecules in the DNA library and RNA library, thereby determining the DNA damage sites in the genome and obtaining the gene expression information of the single nucleus.
[0146] 8. The method of any one of embodiments 1-7, wherein the DNA damage is selected from an oxidative DNA damage, double- stranded DNA breaks, DNA intrastrand crosslinking, and ribonucleotide in DNA.
[0147] 9. The method of any one of embodiments 1-8, wherein the DNA damage is oxidative DNA damage, optionally wherein the oxidative DNA damage comprises comprises 8- oxoguanine (8-oxoG), repair intermediates apurinic / apyrimidinic (AP) sites, and / or singlestranded DNA breaks.
[0148] 10. The method of any one of embodiments 1-9, wherein the fixing the one or more nuclei comprises contacting the one or more nuclei with formaldehyde, or exposing the one or more nuclei to UV light.
[0149] 11. The method of any one of embodiments 1-10, wherein the membrane of the one or more nuclei is permeabilized.
[0150] 12. The method of any one of embodiments 1-11, wherein the oxidized base is an oxidized purine or oxidized pyrimidine.
[0151] 13. The method of any one of embodiments 1-12, wherein the enzyme that removes an oxidized base is selected from a DNA repair protein Fpg or a glycosylase, optionally wherein the glycosylase is selected from 8-Oxoguanine DNA glycosylase (OGGI), endonuclease III, MutY homolog DNA glycosylase (MYH), NTHL1, NEIL1, NEIL2, and NEIL3.
[0152] 14. The method of any one of embodiments 1-13, wherein the AP endonuclease is selected from endonuclease IV, APE1, APE2, APN1, and APN2.
[0153] 15. The method of any one of embodiments 1-14, wherein incorporating a labeled nucleotide to the DNA break comprises contacting the DNA break with a DNA polymerase or a terminal deoxynucleotidyltransferase, optionally Bst polymerase. 16. The method of any one of embodiments 1-15, wherein the labeled nucleotide is selected from a labeled dUTP, a labeled dCTP, a labeled dATP, a labeled dTTP, and a labeled dGTP, optionally a labeled dUTP.
[0154] 17. The method of any one of embodiments 1-16, wherein the label is a biotin or a fluorophore.
[0155] 18. The method of any one of embodiments 1-17, wherein the labeled nucleotide is a biotin-labeled dUTP.
[0156] 19. The method of any one of embodiments 1-18, wherein the protein that binds the labeled nucleotide is selected from an antibody and streptavidin.
[0157] 20. The method of any one of embodiments 1-19, further comprising adding at least one antibody that binds the protein that binds the labeled nucleotide (e.g., secondary antibody).
[0158] 21. The method of any one of embodiments 1-20, wherein the tagmentation comprises contacting a transposome comprising a tag or an adaptor with one or more nuclei.
[0159] 22. The method of embodiment 21, wherein the transposome comprises protein A-linked Tn5.
[0160] 23. The method of any one of embodiments 1-22, wherein the polynucleotide tail comprises an array of single nucleotides, optionally an array of dCTP.
[0161] 24. The method of any one of embodiments 1-23, wherein the polynucleotide tail is added to the genomic DNA fragments and cDNA using a) a terminal deoxynucleotidyltransferase (TdT); b) a DNA ligase and DNA or RNA oligonucleotide; c) a DNA polymerase and a random primer; or d) a DNA or RNA oligonucleotide with a reactive chemical group that attaches to the 3’ end of the DNA or RNA (e.g., an azide group or an alkyne group, e.g., suitable for click chemistry).
[0162] 25. The method of any one of embodiments 1-24, wherein the genomic DNA fragments and / or cDNA are amplified using a polymerase chain reaction.
[0163] 26. The method of any one of embodiments 1-25, wherein, in step (h), the barcoding the genomic DNA fragment and / or cDNA comprises one additional tag, thereby generating a genomic DNA fragment comprising the first tag and the third tag, and / or cDNA comprising the second tag and the third tag.
[0164] 27. The method of any one of embodiments 1-25, wherein, in step (h), the barcoding the genomic DNA fragment and / or cDNA comprises two additional tags, each comprising a barcode, thereby generating a genomic DNA fragment comprising a first tag, a third tag, and a fourth tag, and a cDNA comprising a second tag, a third tag, and a fourth tag.
[0165] 28. The method of any one of embodiments 1-27, wherein the barcoding the genomic DNA fragment and / or cDNA comprises a ligation-based barcoding, wherein the one or more nuclei are contacted with a ligase and at least one additional tag comprising a barcode.
[0166] 29. The method of embodiment 28, wherein the ligation-based barcoding comprises a combinatorial barcoding, optionally selected from SPLiT-seq, Sci-RNA-seq3, Paired- seq / Tag, and SHARE-seq.
[0167] 30. The method of any one of embodiments 1-29, wherein for the DNA library, the enriching for the genomic DNA fragments comprises: i) cleaving the cDNA by contacting the nucleic acid with an endonuclease that cleaves the second restriction site; ii) amplifying the genomic DNA fragments; and iii) contacting the genomic DNA fragments with an endonuclease that cleaves the third restriction site, optionally wherein the genomic DNA fragments are contacted with an endonuclease that cleaves the third restriction site and an endonuclease that cleaves the second restriction site.
[0168] 31. The method of any one of embodiments 1-30, wherein for the RNA library, the enriching for the cDNA comprises: i) cleaving the genomic DNA fragments with an endonuclease that cleaves the first restriction site; ii) amplifying the cDNA; and iii) optionally repeating i).
[0169] 32. The method of any one of embodiments 1-31, wherein the third restriction site is recognized by a type IIS endonuclease.
[0170] 33. The method of embodiment 32, wherein the type IIS endonuclease is selected from FokI, Acul, AsuHPI, Bbvl, Bpml, BpuEI, BseMII, BseRI, BseXI, Bsgl, BslFI, BsmFI, BsPCNI, BstVlI, BtgZI, Ecil, Eco57I, FaqI, Gsul, HphI, Mmel, NmeAIII, Schl, TaqII, TspDTI, and TspGWI.
[0171] 34. The method of any one of embodiments 1-33, wherein the ligase is a T3, T4, or T7 DNA ligase.
[0172] 35. The method of any one of embodiments 1-34, further comprising obtaining information of the genome-wide histone modifications and / or chromatin state (e.g., location of histone variants, chromatin-associated factors (e.g., transcription factors, chromatin remodelers, etc.) in a single nucleus.
[0173] 36. The method of embodiment 35, wherein the obtaining information of the genomewide histone modifications and / or chromatin state comprises performing ChlP-chip, ChlP- seq, and / or Paired-Tag.
[0174] 37. The method of embodiment 35 or 36, wherein the histone modification is selected from H3K4mel, H3K4me2, H3K4me3, H3K27ac, H3K27me2, H3K27me3, H3K9me3, H3K9me2, H3K9Ac, H3K14Ac, H3K18Ac, H3K23Ac, H3K27me3, H4K16ac, H4K20mel, H4K20me2, H4K20me3, H4K27me3, HlK26mel, H3S10Ph, and H2A.XS139.
[0175] 38. The method of any one of embodiments 1-37, wherein the one or more nuclei are of a fungus, a bacterium, or a mammalian cell (e.g., a mouse cell, a rat cell, a monkey cell, a human cell).
[0176] 39. The method of any one of embodiments 1-38, wherein the one or more nuclei are of a cell type selected from a neuronal cell (e.g., excitatory neurons (ExN), inhibitory neurons (InN), and non-neuron cell types, including oligodendrocyte precursor cells (OPC), oligodendrocytes (ODC), astrocytes (AST), microglia cells (MiG), endothelial cells (Endo), and vascular leptomeningeal cells (VLMC)), a bone cell, an endothelial cell, a fibroblast, a blood cell (e.g., erythrocyte, lymphocyte, platelet), an epithelial cell, a muscle cell, a stem cell, an epidermal cell, adipocyte, and hepatocyte.
[0177] 40. The method of any one of embodiments 1-39, wherein the one or more nuclei are of a cancer cell.
[0178] 41. The method of any one of embodiments 1-40, wherein the one or more nuclei are of a proliferating cell, a quiescent cell, or a senescent cell.
[0179] 42. The method of any one of embodiments 1-41, wherein the one or more nuclei are treated with an agent, or the one or more nuclei are obtained from one or more cells treated with an agent that changes cell dynamics.
[0180] 43. The method of embodiment 42, wherein the agent is selected from a toxin, a DNA damage inducing agent (e.g., H2O2, reactive oxygen species (ROS)), a chemotherapeutic agent, a small molecule, an alkylating agent, a hormone, a cell signaling molecule, a cytokine, a chemokine, and any combination of two or more thereof.
[0181] 44. The method of embodiment 42 or 43, wherein the agent induces oxidative DNA damage. EXAMPLES
[0182] Example 1: Materials and Methods
[0183] Cell culture and processing
[0184] HeLa S3 (human, ATCC CCL-2.2) cells were cultured according to standard procedures in DMEM (Gibco, 10569010) supplemented with 10% fetal bovine serum (Gibco, 16000044) and 1% penicillin- streptomycin (Gibco, 10378016) at 37 °C with 5% CO2. Cells were not authenticated or tested for mycoplasma. Cells were sown on plates 48 hours before treatment with H2O2 at the density of 3xl05cells per culture dish. Subconfluent cells were treated with H2O2, in a complete medium, at the final concentration of 200 pM (Thermo Scientific Chemicals, 202465000) for 2 hours at 37°C 5% CO2. After treatment, the medium was replaced with complete medium without H2O2 for additional incubation.
[0185] To prepare single-cell suspensions, HeLa S3 cells were collected by centrifugation at 0 hour, 2 hours, 6 hours, 24 hours and 48 hours after medium replacement, washed with PBS (Gibco, 10010-23) and counted using a BioRad TC20 cell counter. Cells without treatment were also collected as control. Next, cells were resuspended in ImL cold NIB-HEPES buffer (10 mM HEPES pH 7.3 (Thermo Scientific Chemicals, J16924.K2), 10 mM NaCl (Sigma, S7653), 3 mM MgCh (Sigma, 63069), lx Protease Inhibitor (Roche, 05056489001), 0.5 U / pL RNase OUT (Invitrogen, 10777-019) and 0.5 U pL-1SUPERase Inhibitor (Invitrogen, AM2694)) with 0.1% IGEPAL CA630 (Sigma, 18896), 0.1% Tween-20 (Sigma, P1379- 25ML), incubated on ice for 10 mins and centrifuged for 10 mins at 600 g, 4 °C. Cells were then washed with 1 mL cold NIB-HEPES buffer and centrifuged for 10 mins at 600 g, 4 °C before proceeding to Paired-Damage- seq experiments.
[0186] Processing of biospecimens
[0187] The frontal cortex collected from 2-month-old male C57BL / 6J mice was purchased from Jackson Laboratories. Single-nuclei suspensions were prepared from douncing of the frozen tissues, in Douncing Buffer with protease / RNase inhibitor cocktail (DBI: 0.25 M sucrose (Sigma, S7903), 25 mM KC1 (Sigma, P9333), 5 mM MgCh (Sigma, 63069), 10 mM Tris-HCl pH 7.5 (Invitrogen, 15567027), 1 mM DTT (Sigma, D9779), lx Protease Inhibitor (Roche, 05056489001), 0.5 U / pL RNase OUT (Initrogen, 10777-019) and 0.5 U / pL SUPERase Inhibitor (Invitrogen, AM3694)), supplemented with 0.1% Triton-XlOO (Sigma T9284) / The cell suspension were then filtered by 30-pm Cell-Trie (Sysmex) and spun down for 10 mins, at 1,000 g and 4 °C. After the cell pellets were washed with DBI and spun down again, cold NIB-HEPES buffer was added to resuspend the nuclei pellets, incubated on ice for 10 mins and centrifuged for 10 mins at 600 g. 4 °C. Next, cells were washed with 1 mL cold NIB-HEPES buffer and centrifuged for 10 mins at 600 g, 4 °C, counted by BioRad TC20 cell counter and proceeded to Paired-Damage-seq experiments immediately.
[0188] Annealing of barcodes plates and library adaptors
[0189] To prepare the barcode plates (Barcode-plate-R02 and Barcode-plate-R03, Tables 1A- 1F), 6 pL of each of the barcoded oligos (100 pM) were distributed into two 96-well plates (Eppendorf, 0030603303). Then, 44 pL of Linker-R02 (for Barcode-plate-R02) or Linker- R03 (for Barcode-plate-R03) (12.5 pM) was added to each well of the two plates. The plates were sealed and annealed with the following program: 95 °C for 5 mins, and slowly cool down to 20 °C with a ramp of -0.1 °C / s (stock solution plates). The stock solution plates were then divided into ten new 96-well plates, with each well of the working plates containing 5 pL of annealed barcoded adaptor ready for ligation reaction. To prepare P5 adaptor mix for second adaptor tagging of DNA libraries, P5-Fokl was mixed with P5c- NNDC-Fokl and P5H-FokI was mixed with P5Hc-NNDC-FokI, the final concentration 50 pM for both of them. The oligo mixtures were then annealed with the following program: 95 °C for 5 mins, then slowly cool down to 20 °C with a ramp of -0.1 °C / s. The annealed P5 complex and P5H complex were then mixed on ice at the ratio of 1:3, and stored at -20 °C. The sequences of DNA oligo used here are listed in Tables 1A-1F.
[0190] Assembly of transposomes
[0191] Barcoded Tn5 adaptor oligos were mixed with a pMENTs oligo with 50 pM final concentration. The oligo mixtures were then annealed with the following program: 95 °C for 5 mins, followed by slowly cool down to 20 °C with a ramp of -0.1 °C / s. Next, 1 pL of annealed barcoded Tn5 adaptor was mixed with 6 pL of unloaded protein A-Tn5 (0.5 mg / mL), briefly vortexed and quickly spun down. The mixtures were incubated at room temperature for 30 mins then at 4 °C for an additional 10 mins. To prepare assembled protein A-Tn5 complex for bulk CUT&Tag experiments, Adaptor-A and Adaptor-B oligos were mixed with pMENTs oligo at 50 pM final concentration, individually, followed by same annealing program as above. Next, 0.5 pL of annealed Adaptor-A and 0.5 pL of annealed Adaptor-B were mixed with 6 pL of unloaded protein A-Tn5 (0.5 mg / mL), individually, followed by same incubation as above. The protein A-Tn5 complex was then diluted into 0.05 mg / mL. Assembled Tn5 complex for ATAC-seq was prepared similarly as the protein A-Tn5 complex for bulk CUT&Tag, except using unloaded Tn5 (0.5 mg / mL) for assembling, and the assembled Tn5 complex was diluted into 0.05 mg / mL final concentration. To prepare Tn5-N5 complex, Adaptor-A oligo was mixed with pMENTs oligo at 50 pM final concentration, followed by the same annealing and assembling steps as above. The Tn5-N5 complex was diluted into 0.05 mg / mL. The assembled Tn5 or protein A-Tn5 complex can be stored at -20 °C for up to 6 months. The sequences of DNA oligo used here are listed in Tables 1A-1F. Protein A-Tn5 and CUT&Tag technology are also described in EP3688157B1 and US 11,885,814 B2, each of which is incorporated herein by reference in its entirety.
[0192] Paired-Damage-seq procedures.
[0193] Cell fixation and nucleosome depletion. Nuclei were suspended in 2.5 mL NIB-HEPES buffer, 123 pL 37% formaldehyde (Sigma, F8775) was added into nuclei suspension and incubated at room temperature for 10 mins with gentle rotation. After the incubation, fixed cells were spun-down at 4 °C, 600 g for 8 mins, and the cells were washed with cold NIB- Tris Buffer (10 mM Tris-HCl pH 7.5 (Invitrogen, 15567027), 10 mM NaCl (Sigma, S7653), 3 mM MgCh (Sigma, 63069), lx Protease Inhibitor (Roche, 05056489001), 0.5 U / pL RNase OUT (Invitrogen, 10777-019) and 0.5 U / pL SUPERase Inhibitor (Invitrogen, AM2694), 0.1% IGEPAL CA630 (Sigma, 18896), 0.1% Tween-20 (Sigma, P1379-25ML)), and spun- down again at 4 °C, 600 g for 10 mins. Nuclei were resuspended in 800 pL lx NEBuffer r2.1 (NEB, B6002S) supplemented with 0.3% SDS (Invitrogen, 15553-035) and incubated at 42 °C, 300 r.p.m. for 30 mins in the Thermomixer (Eppendorf). Nuclei were spun-down at 600 g, for 10 mins and washed with 200 pL NIB-Tris with buffer twice for SDS quenching, and 200 pL lx NEBuffer r2.1 with 0.1% Triton X-100 (Sigma T9284) for one time.
[0194] DNA damage enzymatic labeling. Nuclei were suspended in 84 pL enzymatic labeling buffer (71 pL H2O, 10 pL lOx NEBuffer 2 (NEB, B7002S), 1 pL SUPERase Inhibitor, 0.5 pL Rnase OUT and 1 pL 10% BSA (Sigma, A1595-50ML) and 0.1% Triton X-100). 8 pL Fpg (NEB, M0240L) and 8 pL Endo IV (NEB, M0304L) were added into nuclei suspension and incubated at 37 °C for 16 hrs. Nuclei were spun-down at 600 g, for 10 mins and suspended in 79 pL enzymatic labeling buffer (66 pL H2O, 10 pL lOx NEBuffer 3 (NEB, B7003S), 1 pL SUPERase Inhibitor, 0.5 pL Rnase OUT and 1 pL 10% BSA (Sigma, A1595-50ML) and 0.1% Triton X-100). Next, 5 pL Bst full-length polymerase (NEB, M0328S), 10 pL Taq DNA Ligase (NEB, M0208L), 2 pL of P-Nicotinamide adenine dinucleotide (NAD+, NEB, B9007S), 4 pL mixed dNTPs (5 pM for each dATP (NEB, N0440S), dCTP (NEB, N0441S), dGTP (NEB, N0442S), biotin-dUTP (Thermo Scientific, R0081)) were added into nuclei suspension and incubated at 37 °C for 6 hrs.
[0195] Antibody staining and targeted tagmentation. Anti-biotin antibody (Abeam, ab234284, 2 pg for each tube) and barcoded protein A-Tn5 (1 pL of 0.5 mg / mL for each tube) were assembled in 20 pL lx Med Buffer #1 (20 mM HEPES pH 7.5 (Invitrogen, 15630106), 300 mM NaCl, 0.5 mM Spermidine (Sigma, S2626-1G), lx Protease Inhibitor cocktail, 0.5 U / pL SUPERase IN, 0.5 U / pL RNase OUT, 0.01% IGEPAL CA630, 0.01% Digitonin (Millipore, 300410-250MG) and 2 mM EDTA (Thermo Scientific Chemicals, 15575020)), and the mixtures were rotated for 60 mins at room temperature. Nuclei were spun-down at 600 g, for 10 mins and washed with 200 pL Med Buffer #1 twice. Next, 3.6 million nuclei were aliquoted into 12 Maximum Recovery tubes (Axygen, MCT-150-L-C) and resuspended in 50 pL of Med Buffer #1. Complex of antibody -protein A-Tn5 was mixed with nuclei suspension and incubated at 4 °C with rotation overnight. The nuclei were then spun down at 600 g, 4 °C for 10 mins, and resuspended in 50 pL of Med Buffer #2 (20 mM HEPES pH 7.5, 300 mM NaCl, 0.5 mM Spermidine, lx Protease Inhibitor cocktail, 0.5 U pL-1SUPERase IN, 0.5 U pL-1RNase OUT, 0.01% IGEPAL CA630 and 0.01% Digitonin), and this was repeated twice. The tagmentation reaction was initiated by adding 2 pL of 250 mM MgCh and was carried out at 550 r.p.m. at 37 °C for 60 mins in a ThermoMixer (Eppendorf). The reaction was quenched by adding 16.5 pL of 40.5 mM EDTA. Nuclei were then spun down at 1,000 g, 4 °C for 10 mins and proceeded to reverse transcription immediately.
[0196] Reverse transcription. Nuclei pellets were resuspended in 20 pL of reverse transcription mix with the corresponding #1-#12 barcoded reverse transcription primers (lx RT Buffer, 1 M Betaine (Sigma, B0300-1VL), 2 pM TSO, 5 mM DTT (Sigma, D9779), 6.25 mM MgCh, 0.5 mM dNTPs (NEB, N0447S), 0.5 U / pL SUPERase IN, 0.5 U / pL RNase OUT, 2.5 pM barcoded T15 primer and 2.5 pM barcoded N6 primer, and 1 U / pL Maxima Reverse H Minus Reverse Transcriptase (Invitrogen, EP0751)). The reverse transcription was performed in a thermocycler with the following program: Step 1: 42 °C x 90 min; Step 2: 50 °C x 10 min; Step 3: 8 °C x 12 s, 12°C x 45 s, 25 °C x 45 s, 37 °C x 45 s, 45 °C x 45 mins, go to Step 2 for additional two times; Step 3: 50 °C x 10 mins and hold at 4 °C. After the reaction, the nuclei were transferred and pooled into 1.5 mL tubes prewashed with 5% BSA in PBS and cooled on ice for 2 mins, with 4.8 pL of 5% Triton X-100. Nuclei were then spun down at 1,000 g, 4 °C for 10 mins and proceeded to ligation-based combinatorial barcoding immediately.
[0197] Ligation-based combinatorial barcoding. Nuclei were resuspended in 1 mL of lx NEBuffer r3.1 (NEB, B6003S) and then transferred to ligation mix (2,262 pL of H2O, 500 pL of lOx T4 DNA Ligase Buffer (NEB, B0202S), 50 pL of 10 mg / mL BSA, 100 pL of lOx NEBuffer r3.1 and 100 pL of T4 DNA Ligase (NEB, M0202L)). Each 20 pL of the ligation reaction mix was distributed to two Barcode-plate-R02 using a multichannel pipette and incubate at 300 r.p.m. at 37 °C for 30 mins in a ThermoMixer. Then, 5 pL of R02-Blocking-Solution (264 pL of 100 pM Blocker-R02 oligo, 250 pL of lOx T4 DNA Ligase Buffer, 486 pL of H2O) was added to each well using a multichannel pipette and continued incubation for an additional 30 mins. The nuclei were then pooled and spun down at 1,000g, 4 °C for 10 mins. The second round of ligation was then carried out similarly to the first round, except that after 30 mins of the ligation reaction, Termination-Solution (264 pL of 100 pM Quencher-R03, 250 pL of 0.5 M EDTA and 236 pL of H2O) was added to quench the reaction without additional incubation. Typically, 100,000 to 300,000 nuclei could be recovered after ligation-based barcoding. Nuclei were then resuspended in PBS, counted and aliquoted to sub-libraries containing 5,000 to 10,000 nuclei. Sub-libraries were diluted to 35 pL with lx NIB-Tris buffer. Then, 5 pL of 4 M NaCl, 5 pL of 10% SDS (Invitrogen, 15553-035) and 5 pL of 10 mg / mL Protease K (NEB, P8107S) were added and nuclei were lysed at 850 r.p.m. at 55 °C for 2 hrs in a ThermoMixer. The lysed solution was cooled down to room temperature and then purified with lx SPRI beads (Beckman Coulter, B23319) and eluted in 12.5 pL of H2O. The purified DNA can be stored at -20 °C or -80 °C for up to 4 weeks. An exemplary ligation-based single-cell barcoding, SPLiT-seq, is described in US11427856B2, which is incorporated herein by reference in its entirety.
[0198] Preamplification of barcoded DNA / cDNA. First, 1.5 pL of lOx TdT buffer and 0.5 pL of 1 mM dCTP (NEB, N0447S) were added into 12.5 pL of purified DNA / cDNA mix and denatured at 95 °C for 5 mins, and then quickly chilled on ice for 5 mins. Then, 1 pL of TdT (NEB, M0315S) was added and incubated at 37 °C for 30 mins followed by heat deactivation at 75 °C for 20 mins. Anchor Mix (6 pL of 5x KAPA HiFi buffer, 0.6 pL of 10 mM dNTPs, 0.6 pL of 10 pM Anchor-Fokl-GSH-Oligo and 0.6 pL of KAPA HiFi HS (KAPA, KK2502)) was added and the linear amplification was performed in a thermocycler with the following program: Step 1: 98 °C x 3 mins; Step 2: 98 °C x 15 s, 47 °C x 60 s, 68 °C x 2 mins, 47 °C x 60 s, 68 °C x 2 mins, and repeat Step 2 an additional 15 times; Step 3: 72 °C x 10 mins and hold at 12 °C. Preamplification mix (4 pL of 5x KAPA HiFi buffer, 0.5 pL of 10 mM dNTPs, 2 pL of 10 pM PA-F and PA-R primers, 0.5 pL KAPA HiFi HS) was then added and preamplification was performed with the following program: Step 1: 98 °C x 3 mins; Step 2: 98 °C x 20 s, 65 °C x 20 s, 72 °C x 2.5 mins, and repeat Step 2 an additional 4 times; Step 3: 72 °C x 2 mins and hold at 12 °C. Amplified products were purified with SPRI double-size selection (10 pL + 32.5 pL, 0.2x + 0.65x) and were eluted in 35 pL of H2O.
[0199] Endonuclease digestion and second adaptor tagging. First, 17 pL of each of the purified amplified products was transferred into two tubes for DNA and RNA library construction, respectively. For the DNA part, 2.5 pL of lOx rCutsmart buffer (NEB, B6004S), 1 pL of Sbfl-HF (NEB, R3642), and 4.5 pL of H2O were added to a DNA-tube. The digestion reaction was incubated at 37 °C for 60 mins. 1.8x SPRI beads were used to purify the digestion product and eluted in 11 pL H2O. Second round-preamplification mix (4 pL of 5x KAPA HiFi buffer, 0.5 pL of 10 mM dNTPs, 2 pL of 10 pM PA-F and PA-R primers, 0.5 pL KAPA HiFi HS) was then added and preamplification was performed with the following program: Step 1: 98 °C x 3 mins; Step 2: 98 °C x 20 s, 65 °C x 20 s, 72 °C x 2.5 mins, and repeat Step 2 an additional 4 times; Step 3: 72 °C x 2 mins and hold at 12 °C. Amplified products were purified with 1.8x SPRI beads and were eluted in 20 pL of H2O. Then, 2.5 pL of lOx rCutsmart buffer, 1 pL of Sbfl-HF (NEB, R3642), 1 pL of FokI (NEB, R0109S) were added to the DNA-tube. The digestion reaction was incubated at 37 °C for 60 mins. 1.25x (31.3 pL for DNA and 25 pL for RNA) SPRI beads were used to purify the digestion product and eluted in 10 pL H2O. Next, 2 pL of lOx T4 DNA Ligase Buffer, 2 pL of P5 Adaptor Mix, 4 pL of H2O and 2 pL of T4 DNA Ligase were added and ligation reactions were carried out with the program: 4 °C for 10 mins, 10 °C for 15 mins, 16 °C for 15 mins, 25 °C for 45 mins. The ligation product was then purified with 1.25x SPRI beads and eluted in 30 pL of H2O. For the RNA part, 2 pL of lOx rCutsmart buffer and 1 pL of Notl-HF (NEB, R3189) were added to an RNA-tube. The digestion reaction was incubated at 37 °C for 60 min. 1.8x SPRI beads were used to purify the digestion product and eluted in 11 pL H2O. Second round-preamplification mix (4 pL of 5x KAPA HiFi buffer, 0.5 pL of 10 mM dNTPs, 2 pL of 10 pM PA-F and TSO primers, 0.5 pL KAPA HiFi HS) was then added and preamplification was performed with the following program: Step 1: 98 °C x 3 min; Step 2: 98 °C x 20 s, 65 °C x 20 s, 72 °C x 2.5 min, and repeat Step 2 an additional 6 times; Step 3: 72 °C x 2 mins and hold at 12 °C. Amplified products were purified with 1.8x SPRI beads and were eluted in 17 pL of H2O. Next, 2 pL of lOx rCutsmart buffer and 1 pL of Notl-HF (NEB, R3189) were added to an RNA-tube. The digestion reaction was incubated at 37 °C for 60 min. 1.25x SPRI beads were used to purify the digestion product and eluted in 10 pL H2O. 10 pL of 2x Tagmentation Buffer (66 mM Tris-Ac, pH 7.8 (ThermoFisher Scientific, BP- 152), 22 mM MgAc (Sigma, M2545), 133 mM KAc (Sigma, P5708) and 32% DMF (EMD Millipore, DX1730)), and 0.5 pL of 0.05 mg / mL Tn5-N5 were added and tagmentation reactions were carried out at 550 r.p.m. at 37 °C for 30 mins in a ThermoMixer followed by clean-up using QIAquick PCR purification kit and elution in 30 pL of 0.1 x Elution Buffer (QIAGEN).
[0200] Indexing PCR and sequencing, PCR mix (30 pL of purified P5-tagged product for DNA or N5-tagged product for RNA, 10 pL of 5x Q5 buffer, 1 pL of 10 mM dNTPs, 0.5 pL of 50 pM P5 primer for DNA or N5 primer for RNA, 2.5 pL of 10 pM P7 primer, 5 pL of H2O and 1 pL of NEB Q5 DNA Polymerase (NEB, M0491)) was prepared and ran the following program: Step 1: 72 °C x 5 min, 98 °C x 30 s; Step 2: 98 °C x 10 s, 63 °C x 30 s, 72 °C x 1 min, and repeat Step 2 an additional 8-10 times to reach 10 nM concentration; Step 3: 72 °C x 1 min and hold at 12 °C. The libraries were cleaned up using 0.85x (42.5 pL) SPRI beads. The final libraries were sequenced using NovaSeq 6000 platform (Illumina) with the following read lengths: PE 100 + 8 + 8 + 100 or PE 150 + 8 + 8 + 150 (Read 1 + Index 1 + Index 2 + Read 2).
[0201] Bulk CUT&Tag
[0202] Bulk CUT&Tag assay was performed similar to previously described (Kaya-Okur,
[0203] H.S. et al. (2019) Nat Commun 10, 1930, which is incorporated herein by reference in its entirety) to generate DNA damage or H3K9me3 profiles. For DNA damage CUT&Tag, nuclei were treated similar as the Paired-Damage- seq procedures before antibody staining. Next, nuclei were washed and counted, 1 million permeabilized nuclei were spun down at
[0204] I,000 g, 4 °C for 10 mins. Biotin (Abeam, ab234284) or H3K9me3 (Abeam, ab8898) antibody-protein A-Tn5 (1 pL of 0.5 mg / ml) were added and the mixtures were rotated at 4 °C overnight. Nuclei were spun down at 600 g, 4 °C for 10 mins, then resuspended in 50 pL of Med Buffer # 2 and this was repeated for two times. The tagmentation reaction was initiated by adding 2 pL of 250 mM MgCh (Sigma, M1028) and was carried out at 550 r.p.m., 37 °C for 60 mins in a ThermoMixer (Eppendorf). The reaction was quenched by adding 16.5 pL of 40.5 mM EDTA (Thermo Scientific, 15575020). Nuclei were then spun down at 1,000 g, 4 °C for 10 mins and proceeded to nuclei lysis and DNA purification by PCR cleanup kit (Zymo, D4033) for PCR amplification with N5 primer and N7 primer.
[0205] ATAC-seq
[0206] An optimized ATAC-seq protocol similar previously described (see Buenrostro, J.D., et al. (2013) Nat Methods 10:1213-1218, which is incorporated herein by reference in its entirety) was used to generate chromatin accessibility profile. Briefly, nuclei were isolated counted, 50k permeabilized nuclei were spun down at 1,000 g, 4 °C for 10 mins. Nuclei were resuspended in 45 pL of l.lx Tagmentation Buffer (36.7 mM Tris-Ac, pH 7.8 (ThermoFisher Scientific, BP- 152), 12.1 mM MgAc (Sigma, M2545), 73.3 mM KAc (Sigma, P5708) and 17.8% DMF (EMD Millipore, DX1730)). Tn5 was then added, and reactions were carried out at 37 °C, 550 r.p.m. for 30 mins in a ThermoMixer. Reactions were quenched by adding 25 pL of 40 mM EDTA and nuclei were then spun down at 1,000g, 4 °C for 10 mins and proceeded to nuclei lysis and DNA purification by PCR cleanup kit for PCR amplification with N5 primer and N7 primer.
[0207] Nucleus RNA-seq
[0208] Bulk ATAC-seq assay was performed similar to previously described to generate transcriptome profile. Briefly, nuclei were isolated similarly to Paired-Tag procedures (e.g., as described in WO2021262671 A3 and US2023227813A1, which is incorporated herein by reference in its entirety) and then counted, and 15k permeabilized nuclei were spun down at 1,000 g, 4 °C for 10 mins. Nuclei pellets were resuspended in 20 pL of reverse transcription mix (lx RT Buffer, 1 M Betaine (Sigma, B0300-1VL), 2 pM TSO, 5 mM DTT (Sigma, D9779), 6.25 mM MgCl2(Sigma, 63069), 0.5 mM dNTPs (NEB, N0447S), 0.5 U / pL SUPERase IN (Invitrogen, AM2694), 0.5 U / pL RNase OUT (Invitrogen, 10777-019), 2.5 pM barcoded T15 primer and 2.5 pM barcoded N6 primer, and 1 U / pL Maxima Reverse H Minus Reverse Transcriptase (Invitrogen, EP0751)). The reverse transcription was performed in a thermocycler with the following program: Step 1: 42 °C x 90 min; Step 2: 50 °C x 10 mins; Step 3: 8 °C x 12 s, 12°C x 45 s, 25 °C x 45 s, 37 °C x 45 s, 45 °C x 45 mins, go to Step 2 for additional two times; Step 3: 50 °C x 10 mins and hold at 4 °C. After the reaction, the nuclei were cooled on ice for 2 mins, with 0.4 pL of 5% Triton-XlOO (Sigma, T9284). Nuclei were then lysed and proceeded to cDNA amplification with TSO primer and R01_PAF, followed by 1.25x SPRI beads purification and elution in 11 pL of H2O. 10.5 pL of 2x Tagmentation Buffer (66 mM Tris-Ac, pH 7.8 (ThermoFisher Scientific, BP- 152), 22 mM MgAc (Sigma, M2545), 133 mM KAc (Sigma, P5708) and 32% DMF (EMD Millipore, DX1730)), 0.5 pL of 0.05 mg / mL Tn5-N5 was added and tagmentation reactions were carried out at 550 r.p.m., 37 °C for 30 mins in a ThermoMixer (Eppendorf) followed by clean-up using QIAquick PCR purification kit and elution in 30 pL of 0.1 x Elution Buffer (QIAGEN) for PCR amplification with N5 primer and P7 primer.
[0209] Data analysis procedures
[0210] Pre-processing of Paired-Damage- seq data. Cellular barcodes and the linker sequences are read by Read 2. The first bases for barcodes BC#1, BC#2 and BC#3 should locate within the 84th-87th, 47th-50th and 10th-13th bases of Read 2 and the exact positions were identified by matching the linker sequences adjacent to the cellular barcodes. A bowtie reference index was generated with all possible cellular barcode combinations (192 x 192 x 12) and barcode sequences were mapped to the cellular barcode reference using bowtie with the parameters: - v l -m l —norc. Nextera adaptor sequences were trimmed from 3' of DNA and RNA libraries, Poly-dT sequences were further trimmed from 3' of RNA libraries and low-quality reads (minimal base calling quality: Q = 30) were excluded from further analysis.
[0211] Evaluation of collision rate. Reads from species-mixing test were extracted based on cellular barcodes (BC#1 = 11 or 12 of HeLa cells Paired-Damage- seq dataset) and mapped to a reference genome using STAR with the combined reference genome (GRCh37 for human and GRCm38 for mouse). Duplicates were removed based on the mapped position, UMI and cellular barcode, and sub-library indices. For evaluation of the collision rate, nuclei with less than 75% UMIs mapped to one species were classified as mixed cells.
[0212] Reads mapping. Adaptor-trimmed reads were first mapped to a mouse GRCm38 (mmlO) or human GRCh38 (hg38) reference genome with STAR (v.2.6.0a) for RNA or bowtie2 for DNA. Mapped DNA reads were further filtered by MAPQ > 10. Duplicates were removed based on the mapped position, UMI and cellular barcode, and sub-library indices. BC#1 was used for the identification of sample origin identities. Low-coverage nuclei were removed from further analysis (<1,000 transcripts and <1,000 unique DNA reads for HeLa cells, <200 transcripts and <1,000 unique DNA reads for mouse brain cells).
[0213] Clustering of Paired-Damage-seq RNA profiles. RNA alignment files were converted to a matrix with cells as columns and genes as rows. The clustering of single cells based on RNA profiles was performed with the Seurat package. Briefly, cell-to-gene counts were normalized with SCTransform and variable genes were selected for dimension reduction by principal component analysis (PCA), visualized with UMAP and clustered with the Leiden algorithm. To validate the annotation results, reference-mapping (see Hao, Y. et al. (2021) Cell 184:3573-3587 e3529, which is incorporated herein by reference in its entirety) of Paired- Damage-seq RNA profile to published snRNA-seq dataset was performed. To compare the clustering results, The overlap coefficients ( ) were calculated according to the number of cells with the labels from the Paired-Tag dataset (A), from Yao et al. (B) and from reference mapping ( / ?) (z = row index, j = column index):
[0214] Peak calling. To identify DNA damage hotspots, peak calling was performed using MACS3 with the following parameters: ‘-q 0.01 —nomodel —shift 200 —extsize 500 -keep-dup all — nolambda —nomdel’. To identify induced peaks in HeLa cells, control HeLa cells dataset were used as control file for peak calling using MACS3 with the same parameters. Identified peaks were classified in control HeLa cells as conserved peaks. To identify cell-type specific peaks for all the mouse cortex cell types, the “tl.marker_regions” function was used from SnapATAC2 to aggregate the DNA damage signal across cells and utilizes z-scores to identify specifically enriched peaks, with p-value cutoff = 0.05.
[0215] DNA damage signal-gene expression correlation analysis. To minimize the influences from stochastic DNA damage signal, the SEACell software was used to group the single cells into 300 metacells based on transcriptome profiles. To analyze the associations between DNA damage levels and gene expression levels, the ranked correlation score between the raw count of each gene and the DNA damage read density of peaks within the 500kb span of the gene in constructed SEACells was calculated using the “add_cor_scores” function of SnapATAC2. To estimate the background noise levels, the identities of metacells were shuffled and recalculated the corresponding PCCs between DNA damage levels and gene expression levels.
[0216] Motif enrichment and GO analysis. Motif enrichment for each peak sets was performed using HOMER (v4.9) and the ‘findMotifsGenome.pl’ function. GO analysis for each peak set or peak-associated CREs was performed using Genomic Regions Enrichment of Annotations Tool (GREAT) with default parameters, and biological process terms were used for annotation.
[0217] GWAS enrichment. The comparison of mouse DNA damage hotspots with GWASs of human phenotypes were performed similar to previous study, briefly: liftOver software was used with parameters ‘-minMatch=0.5’ to convert the DNA damage peaks from mmlO to hg38 genomic coordinates; the human hg38 elements were lifted back to mm 10 and only kept the regions that could be mapped to original loci, peaks larger then 1-kb in the human genome were removed from further analysis. The following GWAS summary statistics for quantitative traits were used: Major Depressive Disorder, Post Traumatic Stress Disorder, Neuroticism, Lymphocyte count, Basophil count, Amyotrophic Lateral Sclerosis, Monocyte count, Anorexia Nervosa, Alzheimer’s Disease, Autism Spectrum Disorder, Obsessive Compulsive Disorder, Atopic Dermatitis, Breast cancer, Educational Attainment, Stroke. To compare with background control from union of DNA damage peaks for all cell types, the linkage disequilibrium score regression (World Wide Web at github.com / bulik / ldsc) was performed to estimate the enrichment coefficients of each annotation.
[0218] Data availability
[0219] Raw sequencing and processed data generated in this study are available from NCBI Gene Expression Omnibus (GEO) (World Wide Web at ncbi.nlm.nih.gov / geo / ) under the accession number GSE268567 (reviewer’s token: oxitwycarfivfkx). Other external datasets were downloaded from the GEO with the following accession numbers: AP-seq (GSE121005), CLAPS-seq (GSE181312), Paired-seq and Paired-Tag (GSE152020), Droplet Paired-Tag (GSE156683), snATAC-seq and Paired-Tag of aging mouse brain (GSE187332), CoTECH (GSE158435), ENCODE (World Wide Web at encodeproject.org / ) with the following accession numbers: HeLa DNase-seq (ENCFF977IGB), HeLa H3K4me3 ChlP-seq (ENCFF578NOK), HeLa H3K36me3 ChlP-seq (ENCFF248WXB), HeLa H3K4mel ChlP- seq (ENCFF360CQR), HeLa H3K9me3 ChlP-seq (ENCFF712ATO), HeLa H3K27ac ChlP- seq (ENCFF392EDT), HeLa H3K27me3 ChlP-seq (ENCFF512TQI) and HeLa 18-state model chromatin states (ENCSR098REA); 4DN Data Portal (World Wide Web at data.4dnucleome.org / )) with the following accession numbers: HeLa cell Hi-C (4DNFIBMVFF0F), mouse cortex Hi-C (4DNFIB16WAKX); Mouse Brain scRNA-seq (World Wide Web at portal.brain-map.org / atlases-and-data / maseq); and the lOx Genomics website (World Wide Web at 10xgenomics.com / ). Code availability
[0220] Custom scripts used for analyzing Paired-Damage-seq datasets are available from GitHub (World Wide Web at github.com / czhulab / Paired-Damage-seq).
[0221] Example 2: Paired-Damage-seq for joint analysis of oxidative DNA damage and repair intermediates with transcriptome in single cells
[0222] More than 70% of endogenous lesions in human cells are oxidative damage 8- oxoguanine (8-oxoG) and its repair intermediates apurinic / apyrimidinic (AP) sites and singlestrand DNA breaks (SSBs), which can arise during metabolism or base hydrolysis. To capture the genome fragments with these prevalent and closely related DNA breaks, in situ “labeling-by-repairing” strategy was designed to attach biotin to the damaged sites, based on the in vitro base excision repair (BER reaction for the detection of DNA damaged sites (Fig. 5A). For the sake of conciseness, hereafter, 8-oxoG and its repair intermediates AP site and SSB are collectively referred to as DNA damage. To minimize the labeling bias raised from chromatin epigenetic states, a nucleosome depletion step was first performed on formaldehyde crosslinked cells or nuclei (Fig. 5B). Next, the fixed nuclei were treated with DNA repair proteins Fpg and Endo IV to remove pre-existed 8-oxoG and AP sites to give SSBs. Nick translation with Bst polymerase was then performed to incorporate biotin- modified dUTPs at the pre-existed and newly generated SSBs, followed by repairs of nicks with Taq DNA ligase (Fig. 5A). The above steps were performed in situ and was compatible with different single-cell barcoding approaches. Next, cells were aliquoted into 12 tubes and incubated with the anti-biotin antibody, followed by targeted tagmentation with proteinA-Tn5 (PAT), attaching the tube-specific DNA barcodes (BC#1) to the biotin-labeled DNA damage sites. In situ reverse transcription (RT) was then performed with primers containing the same set of DNA barcodes (BC#1) as in the PAT. Cells were further labeled by two rounds of ligation-based combinatorial barcoding: DNA fragments and cDNA molecules from the same cells would contain the identical combinations of barcodes (BCs#l, #2, and #3) (Tables 1A- 1F). Cells were then lysed and barcoded DNA / cDNA molecules were purified, followed by library amplification steps similar to Paired-Tag. After sequencing, the cellular barcodes simultaneously attached to DNA and cDNA fragments were used to assign each read to individual cells (Fig. 1A and Example 1).
[0223] As a proof-of-concept, Paired-Damage-seq was performed on cultured HeLa cells and obtained 64,794 cells passed quality filtering (Fig. IB, Fig. 1C, and Example 1). At PCR duplication rates of 51.47% (DNA) and 24.29% (RNA), Paired-Damage- seq captured the median numbers of 5,320 DNA damage loci and 4,453 transcripts per cell, showing comparable sensitivity as prior single-cell multiomics sequencing methods for joint RNA and epigenome analyses (Fig. ID and Fig. IE). A species mixing experiment was performed by spike-in mouse cells (NIH / 3T3) into human HeLa cells, 94.7% of barcodes were unambiguously assigned to a single species (Fig. IF and Fig. 5C). Then the aggregated single-cell signals were compared with DNA damage and RNA datasets generated with prior methods for analysis of bulk cell population. For the RNA modality, Paired-Damage- seq showed good agreement in the detected gene expression levels compared to bulk RNA-seq profile of the same cell line (Pearson’s correlation coefficient [PCC]: 0.87) (Fig. 1G and Fig. 5D). For the DNA damage modality, the whole genome reads distributions patterns also correlated well with the public bulk dataset generated by CLAPS-seq (a chemical labelingbased approach for 8-oxoG sequencing) (PCC: 0.72) (Fig. 1H and Fig. 5E). To evaluate the sensitivity of Paired-Damage-seq in detecting DNA damage signals, the numbers of cells for DNA modality were down-sampled and reliable profiles (PCC > 0.9) could be obtained by aggregating signals from 1,000 cells (Fig. II). These results demonstrate that Paired- Damage-seq robustly measures oxidative damage and resulting breaks with gene expression from single cells.
[0224] In Paired-Damage-seq, different samples could be multiplexed in one experiment to minimize the risk of batch effects. HeLa cells were treated with H2O2 for 2 hrs to introduce oxidative stress, and continued incubation to 48 hr. Cells collected from different posttreatment time points were barcoded with BC#1 and multiplexed in a single Paired-Damage- seq experiment (Fig. 6A). Cells without treatment and collected at 0- and 2- hrs after treatments exhibited similar transcriptional states, which changed 6 hrs after the stress treatment (Fig. IB). In contrast to gene expression, dimension reductions of cells based solely on DNA damage signals displayed a salt-and-pepper pattern; albeit with a weak gradient distribution across different treatments, unsupervised clustering could not resolve these cell states (Fig. 6B). This was expected as the large scale and stochastic acquired DNA damage could overwhelm the subset of genome regions showing functional influences; thus, identifying cell states based on the noisy single-cell damage signals alone is challenging.
[0225] Example 3: Chromatin states for transcription affects genome susceptibility
[0226] The endogenous DNA damage levels resulted from the balance between the formation and repair of damage. In line with the observation that in cancer genome, single nucleotide variants (SNVs) are preferentially accumulated in heterochromatin, slightly higher DNA damage levels were detected in compartment B over compartment A (average RPKM: 0.37 vs. 0.34), agreeing with the DNA repair activities are less efficient in quiescent chromatin regions (Fig. 2A and Fig. 6C). Interestingly, the DNA damage levels in gene body regions of highly transcribed genes were found to be higher than the corresponding loci in less transcribed or repressed genes (Fig. 6D). This result demonstrate that the chromatin states such as open chromatin and R-loop conformation during transcription renders the DNA more susceptible to endogenous genotoxin. However, even for the highly expressed genes, the average genebody damage level was still lower compared to compartment B. Thus, DNA repair machinery could selectively protect genic regions from the active chromatin states- coupled genome susceptibility, regardless of the genes’ expression levels.
[0227] The average DNA damage levels for different chromatin states regions were calculated from the ENCODE reference annotation. Similar to the results above, genic enhancers (EnhGl / 2) and transcribed regions (Tx / TxWk) were shown to be hotspots of oxidative DNA damage (Fig. 2B). Interestingly, for distal enhancers, comparable damage levels as transcribing regions were observed; this may have resulted from enhancer RNA (eRNA) transcription or due to open chromatin states being more sensitive to damaging stress (Fig. 2B). On the other hand, CREs associated with H3K27me3, such as bivalent enhancers / promoters (EnhBiv, TssBiv) and polycomb-associated heterochromatin (ReprPC / ReprPCWk) showed the lowest damage levels (Fig. 2B). Thus, similar to low- activity genes, these non-coding regions responsible for cell and tissue identities were also considered as targets for selective protection, irrespective of their activity states. However, the constitutive heterochromatin was not included in the transcription-associated prioritized protection (Fig. 2B). The data collectively demonstrate that chromatin epigenetic states associated with active transcription render the genome more susceptible to DNA-damaging stress. To mitigate this risk, mechanisms for prioritized DNA repair of transcribing and regulatory regions could exist while acting independently of their activity states.
[0228] Example 4: Short-tandem repeats (STRs) are hotspots of stress-induced DNA oxidation
[0229] To analyze the oxidative stress-induced DNA damage hotspots, the single-cell damage signal according to their treatment conditions was aggregated and performed peak calling with MACS3. 30,477 peaks were detected from control cells without treatment, which are highly enriched in simple repeats (also known as short tandem repeats, STRs), long interspersed nuclear elements (LINEs) and CpG islands, while depleted from promoters and exons (Fig. 6E and Fig. 6F). By down-sampling the reads into the same depth for all treatment conditions, the increasing numbers of peaks (31,495 - 42,017) were detected from cells collected 0 to 48 hrs after treatment (Fig. 6E-Fig. 6G). Noteworthy, for all treatment conditions, >75% of damaged regions overlapped with at least one of the other conditions, further supporting that the hotspots for DNA damage formation are not randomly distributed, which is driven by conserved factors among different cell populations including base sequence composition (Fig. 6G). For example, chromatin regions associated with zinc-finger genes and repeats (ZNF_Rpts) showed high damage levels; as expected, the G-rich STR subfamilies (such as (AC)n / (GT)n and (TG)n / (CA)n) displayed the highest associations with DNA damage peaks (Fig. 2B and Fig. 7A).
[0230] To further identify hydrogen peroxide treatment-induced damage, differential peak enrichment analysis was performed and identified 1,105 inducible peaks in cells collected 48 hrs after treatment, as compared to control cells (Example 1). The inducible peaks displayed the highest enrichment in CpG islands, low complexity repeats, and STRs (Fig. 7B-Fig. 7C). For example, upon treatment, DNA damage levels increased in simple repeats, predicted Z- form DNA, and G-quadruplex-forming sequences (PQS) regions (Fig. 7D). Non-canonical chromatin structures, including G-quadruplexes (G4-DNA), are known to affect gene expression levels. ROS-mediated 8-oxoG and AP site generation in PQS-containing promoters facilitates the formation of G4-DNA for fast gene activation. In line with this, a positive association between gene expression levels and local DNA damage levels was observed (Fig. 6D). However, the influences of DNA damage formation on chromatin activity were bidirectional: for example, oxidative stress increased the DNA damage levels on endogenous retroviruses (ERVs) in both compartments A and B; while during this process, chromatin accessibilities reduced at compartment A-residing ERVs and remained mostly unchanged (if not slightly increased) for the repressed ERVs in compartment B (Fig. 7E). These results demonstrate that a subset of G-rich STRs could sense oxidative and alter gene activities epigenetic ally by direct formation of non-canonical DNA structure, independent of “canonical” chromatin modifiers.
[0231] Example 5: DNA damage accumulation is associated with epigenetic memory loss
[0232] The epigenetic consequences of oxidative DNA damage formation raised the question of whether the damage-epigenome crosstalk is limited in “sandbox” (for example, formation of G4-DNA in response to PQS guanine oxidation) or could pose threads to chromatin states across the whole genome. As both H3K9me3-associated heterochromatin and enhancers showed high DNA damage levels (Fig. 2B), bulk ATAC-seq and H3K9me3 CUT&Tag (Cleavage Under Targets and Tagmentation) were performed on HeLa cells collected 0-, 6-, and 48- hrs after oxidative stress treatment. The changes of DNA damage levels negatively correlated with both chromatin accessibility (PCC = -0.356) and H3K9me3 levels (PCC = - 0.326) (Fig. 2C, Fig. 2D). This antagonistic relationship was upheld with different window resolutions: PCCs = -0.100 and -0.174 at 10-kb, and -0.459 and -0.428 at 1-mb bins size, for chromatin accessibility and H3K9me3, respectively. Moreover, the epigenetic changes happened immediately after the treatment (2 hrs incubation with H2O2) but did not recover even after 1-2 cell divisions (cell-doubling time: 17.5-32.3 hours for HeLa cells81) (Fig. 8). Therefore, the association between DNA damage accumulation and epigenetic changes were not limited to specific STR families and could widely affect various regulatory regions.
[0233] To understand the influences of DNA damage on gene expression, the multiomics profiles were considered and correlative analysis between DNA damage and gene expression levels was performed. Considering the highly noisy nature of DNA damage, the single cells were first grouped into “metacells” using the SEACell software based on their transcriptional similarities and then calculated the PCCs between gene expression and DNA damage levels (in peaks within the 500-kb range of the gene) across “metacells”. Compared to the background, 11,814 positive and 4,362 negative linked gene-damage peak “pairs” were observed, suggesting oxidative DNA damage is more likely to disrupt repressive regulatory states (Fig. 2E). This could be due to the chromatin decondensation as required by loading of DNA repair machineries. As expected, the DNA damage peaks showing stronger influences on gene expression were enriched in CREs, evidenced by their high associations with histone marks H3K4mel, H3K4me3, and H3K27ac (Fig. 2F). H3K27me3 marked regions were depleted from transcription-influencing DNA damage peaks, echoing the above result that polycomb-associated chromatin regions are included for prioritized DNA repair, albeit with lower degrees of transcription-associated susceptivity. Collectively, the results presented herein revealed the links between the accumulation of oxidative DNA damage and the loss of epigenetic memory.
[0234] Example 6: Brain cell type-specific DNA damage landscape
[0235] The results that the chromatin state is a key determinant for DNA damage hotspots prompted the exploration as to whether the epigenetic states defining cell type identified in complex tissues also contribute to selective cell-type genome vulnerability. Paired-Damage- seq was performed on the mouse cerebral cortex and 73,813 cells were sequenced. Cell clustering based on RNA profiles revealed all major brain subclasses, annotated based on their marker gene expression, including excitatory neurons (ExN), inhibitory neurons (InN), and non-neuron cell types, including oligodendrocyte precursor cells (OPC), oligodendrocytes (ODC), astrocytes (AST), microglia cells (MiG), endothelial cells (Endo), and vascular leptomeningeal cells (VLMC). (Fig. 3A and Fig. 9A). To validate the cell clustering results, reference mapping was performed using the Azimuth software: cells from the Paired-Damage- seq matched well with those from the reference snRNA-seq dataset (Fig. 9B-Fig. 9C). Therefore, the Paired-Damage- seq RNA profiles could assist in the robust integration with various public single-cell RNA-seq atlases. Next, the DNA damage signals were aggregated for each cell group according to the transcriptome -based clustering results and obtained the mouse brain cell-type-specific maps for hotspots of oxidative DNA damage and its repair intermediates. For the sake of statistical power, the analysis at the subclass level was performed by merging the closely related cell sub-types. Interestingly, the DNA damage levels were found to be higher for marker gene loci that are highly expressed in certain nonneuron cell types, including Pdgfra in OPCs, Mobp in oligodendrocytes, Slcla2 for astrocytes and Csflr for microglia cells; echoing the observations from HeLa cells (Fig. 3B, Fig. 6B and Fig. 10A). However, this relationship disappeared in neuron cells and in non- brain- specific endothelial cells and VLMCs (vascular and leptomeningeal cells). Hence, the impacts of transcription-associated chromatin states on genome vulnerability showed intratissue and inter-cell-type heterogeneities.
[0236] For all mouse brain cell types, the average DNA damage levels were higher in compartment B (median DNA damage levels [RPKM]: 1.14 - 1.21) compared to Compartment A (median DNA damage levels [RPKM]: 0.99 - 1.04) (Fig. 3C). This is conservation with the cultured HeLa cells suggesting the nuclei compartmentalization by organizing certain genome regions closer to nuclei membrane, increasing their susceptibility to oxidative reactive species (ROS) which are generated in the mitochondria. Albeit with varying degree of enrichment scores, for all cell types, DNA damage hotspots preferred CpG- Island, LINE, low complexity repeats, LTR, simple repeats and satellites, similar to that in HeLa cells, thereby confirming the contribution from base compositions (Fig. 3D and Fig. 6F). For LINE1, the high DNA damage levels were observed at 5’UTRs of LINE1 for both neuron and non-neuron cells, which was consistent with that mouse LINE1 5’UTRs contain a variable number of GC-rich tandemly repeats (Fig. 10B). For ERVs, higher average damage levels were once again observed in compartment B compared to compartment A, similar to HeLa cells (Fig. 10B and Fig. 7F). Thus, a subset of DNA damage hotspots is conserved across cell-types or species, which resulted from the specific base sequence compositions.
[0237] Example 7: Neuronal cis-regulatory elements are DNA damage hotspots
[0238] Recent studies revealed that DNA repair hotspots are enriched in a subset of well- defined regulatory regions, including enhancers, to protect the identities and functions of postmitotic cells (such as neurons). To examine whether such pattern is established during the formation of DNA damage, the cell-type resolved DNA damage signals were compared with public snATAC-seq and H3K27ac Paired-Tag datasets. At 250k-bp resolution, DNA damage levels were positively correlated with H3K27ac levels for ExN, InN, ODC, and MiG (PCC: 0.45-0.51) (Fig. 3E and Fig. 11A). However, no strong relationship of DNA damage levels with the constitutive histone mark H3K9me3 (PCC: -0.15-0.06) was observed; although the higher average damage burdens were consistently observed in these regions (Fig. 3F and Fig. 11B). At higher resolution, DNA damage signals again tended to be enriched in H3K27ac peak regions: neurons and non-neuron brain cell types (including ODC and MiG cells) showed the highest enrichment; while Endo and VLMCs, which are not brain- specific cell types, displayed weakly enrichments (Fig. 3G and Fig. 11D). Interestingly, for the brainspecific cell types, a higher fraction active enhancers (H3K27ac+ / ATAC-seq+peaks) overlapped with DNA damage peaks compared to weak enhancers (H3K27ac7ATAC-seq+peaks). Again, for endothelial cells and VLMCs, an elevated association between DNA damage hotspots and active enhancers was not observed (Fig. 11D). These results demonstrate a brain- specific mechanism in coordinating enhancer activities via the formation and repair of DNA damage for certain cell types, potentially related to TET-mediated active DNA demethylation.
[0239] Example 8: The divergent impacts of selective cell-type genome vulnerability
[0240] The prevalence of DNA damage in H3K27ac-marked brain enhancers hinted for possible regulatory functions; differential peak analysis was performed to explore whether DNA damage hotspots’ influences are cell identity-dependent. Based on the transcriptome annotation, 3,701-36,903 cell-type- specific peaks were identified with the SnapATAC2 software: among them fewer peaks from the non-proliferating brain cell types (ExN, InN, ODC and AST) were detected compared to proliferating cells (OPC and MiG) and non-brain- specific cell types (Endo and VLMC) (Fig. 4A). GREAT analysis revealed cCRE-associated DNA damage peaks correlated with specific brain functions, for example: adult locomotory behavior (ExN), long-term synaptic potentiation (InN), and myelin assembly (ODC). The DNA damage hotspots in endothelial cells were not enriched in enhancers. However, they still displayed association with functional terms including regulation of VEGF signaling pathway and gamma-aminobutyric acid transport (Fig. 4A). Gamma-aminobutyric acid (GABA) is the primary inhibitory neurotransmitter in the central nervous system but is also produced in tissues outside the nervous system including pancreases and intestines, suggesting the potential endothelial cell-specific genome vulnerability across different tissues. The motif analysis was performed herein on the cell type-specific damage peaks and revealed potential regulators could be affected by DNA damage (Fig. 4B-Fig. 4C). For example, the YY 1 binding motif contains guanine-guanine dinucleotide, which was found enriched in multiple cell types. PAX6 plays an important role in neurogenesis during development, and its binding motif was enriched in ExN-specific DNA damage peaks. The observation indicated that the poly-dG sequence (predicted to be Zfp740 binding motif) was highly enriched for both non-brain- specific cell types (Endo and VLMC) and proliferating brain cell types (OPC and AST) (Fig. 4B). The depletion of this sequence motif intrinsically susceptible to oxidative DNA damage hinted a selected protection of post-mitotic cells in the brain. These results demonstrate that brain cell types display diverse patterns of selective genome vulnerability, shaped by both conserved and brain- specific mechanisms, posing varying functional influences for different cell identities.
[0241] To explore whether the selective genome vulnerability could be associated with disease risks, the human orthologues of the mouse DNA damage hotspots were identified with reciprocal homology searches and linkage disequilibrium score regression analysis was performed to predict the associated diseases and traits (Fig. 4D, Fig. 12A and Example 1). The statistical power of such analysis was limited by the intrinsic high background noise level of DNA damage profiles; however, diseases and traits associated with the dysregulation of specific cell types were still observed (Fig. 4D). For example, the highest enriched disease among them was Major Depressive Disorder (MDD) in excitatory neurons; it has been found DNA oxidation was elevated in patients with MDD compared to healthy individuals. Thus, the DNA damage hotspots may acquire oxidation faster and potentially results in the pathology-associated gene regulation programs. Although at a less statistically significant level, the analysis showed the association of damage hotspots with neurodegenerative disorders, including amyotrophic lateral sclerosis in ODC and AST, and Alzheimer’s disease (AD) in microglia. For example, INPP5D regulates inflammasome in microglia, and its expression level is associated with AD risk. Multiple MiG-specific DNA damage hotspots in the mouse Inpp5d loci were observed, which showed decreased chromatin accessibility in aged mouse brain (Fig. 4E). Another example is another AD and dementia- associated gene Tpcnl, which again showed elevated DNA damage levels compared to other cell types (i.e., inhibitory neuron) and decreased chromatin accessibility in aged mouse brain in enhancers within its intron regions86(Fig. 12B). Interestingly, an association of endothelial DNA damage hotspots with breast cancer was observed, suggesting an inter-tissue conserved genome susceptibility for this cell type (Fig. 4D). Therefore, the accumulation of DNA damage was non-random and exhibited a cell-type specific pattern. Determined by both base sequence contents and cell-type specific chromatin states, the selective cell-type genome vulnerability increased the risk of specific cell types to the corresponding disease-associated dysregulated gene expression programs.
[0242] Finally, to investigate the relationships between DNA damage and epigenetic alternations in regions outside the well-defined active CREs, the distribution patterns of DNA damage and the heterochromatin H3K9me3 histone mark were compared. The age-associated heterochromatin loss was previously observed in the mouse brain (Zhang et al. (2022) Cell Res 32:1008-1021). Here, the differences of H3K9me3 levels between 18- and 3-month-old mice were calculated for each brain cell type and compared with the DNA damage levels detected in 2-month-old mice. At 100-kb resolution, DNA damage levels and H3K9me3 loss displayed neither positive nor negative genome-wide associations. However, when focused on the top 1% damaged regions, negative associations were observed for excitatory neurons (PCC: -0.19), inhibitory neurons (PCC: -0.11), oligodendrocytes (PCC: -0.31), and astrocytes (PCC: -0.18), but not for the oligodendrocyte precursor cells that undergoes proliferation in the adult brain (PCC: -0.03) (Fig 4G-Fig. 4H and Fig. 12C-Fig. 12D). Thus, consistent formation and repair of DNA damage in hotspot regions could lead to progressive loss of epigenome information during aging, which could be alleviated by epigenome “replication” during cell division.
[0243] Example 9: Materials and Methods for Examples 10-16
[0244] Examples 10-16 further confirm the results provided in Examples 1-8 (see also Bai et al. (2025) Nature Methods 22:962-972, (and information available online including supplementary information and material (World Wide Web at doi.org / 10.1038 / s41592- 025-02632-3), source data, and extended data (World Wide Web at doi.org / 10.1038 / s41592- 025-02632-3)) the entire contents of each of which are incorporated herein by reference in their entirety). Cell culture and processing
[0245] HeLa S3 cells (ATCC CCL-2.2) were cultured in DMEM (Gibco, cat. no. 10569010) with 10% fetal bovine serum (Gibco, cat. no. 16000044) and 1% penicillin- streptomycin (Gibco, cat. no. 10378016) at 37 °C with 5% CO2. No authentication or mycoplasma testing was performed for the cells. Cells were seeded at 3 x 105per dish 48 h before 200 pM H2O2 (Thermo Scientific, cat. no. 202465000) treatment for 2 h. After treatment, the medium was replaced with fresh complete medium for recovery. Cells were collected at 0, 2, 6, 24 and 48 h posttreatment, washed with PBS (Gibco, cat. no. 10010-23), and counted using a BioRad TC20 cell counter. Cells were then resuspended in cold NIB-HEPES buffer (10 mM HEPES, 10 mM NaCl, 3 mM MgCh, protease and RNase inhibitors, 0.1% IGEPAL CA630, 0.1% Tween-20) and incubated on ice for 10 min. After centrifugation at 600g for 10 min at 4 °C, the pellet was washed with cold NIB-HEPES buffer, centrifuged again and proceeded to Paired-Damage-seq.
[0246] Processing of biospecimens
[0247] Frontal cortex from 2-month-old male C57BL / 6J mice (Jackson Laboratory) were used. Single-nuclei suspensions were prepared by douncing frozen tissues in Douncing Buffer (0.25 M sucrose, 25 mM KC1, 5 mM MgC12, 10 mM Tris-HCl (pH 7.5), 1 mM dithiothreitol (DTT), lx Protease Inhibitor (Roche), 0.5 U pl- 1 RNase OUT, 0.5 U pl-1 SUPERase Inhibitor (Invitrogen) and 0.1% Triton X100). The suspension was filtered through a 30-pm Cell-Trie filter (Sysmex) centrifuged at 1,000g for 10 min at 4 °C, washed with Douncing Buffer and centrifuged again. Nuclei were then resuspended in cold NIB- HEPES buffer, incubated on ice for 10 min and centrifuged at 600g for 10 min at 4 °C. After a final wash with cold NIB-HEPES buffer, nuclei were counted using a BioRad TC20 cell counter and processed for Paired-Damage-seq.
[0248] Annealing of barcodes plates and library adapters
[0249] This step is adopted from Paired-Tag. To prepare barcode plates (Barcode-plate-R02 and Barcode -plate-R03), 6 pl of each 100 pM barcoded oligo was dispensed into two 96-well plates (Eppendorf, cat. no. 0030603303). Each well received 44 pl of either Linker-R02 (for Barcode-plate-R02) or Linker-R03 (for Barcode-plate-R03) at a concentration of 12.5 pM. The plates were sealed and subjected to the following program: 95 °C for 5 min, followed by gradual cooling to 20 °C at -0.1 °C s-1 . The stock plates were aliquoted into ten 96-well working plates, with 5 pl of annealed barcoded adapter per well. For the P5 adapter mix, P5- Fokl was combined with P5c-NNDC-FokI and P5H-FokI with P5Hc-NNDC-FokI, each at a final concentration of 50 pM. The mixtures were annealed using the program: 95 °C for 5 min, cooling to 20 °C at -0.1 °C s-1 . The annealed P5 and P5H complexes were then mixed at a 1:3 ratio on ice and stored at -20 °C.
[0250] Assembly of transposomes
[0251] This step is adopted from Paired-Tag. For DNA libraries, barcoded Tn5 adapter oligos were combined with a pMENTs oligo to a final concentration of 50 pM and annealed using the program: 95 °C for 5 min, cooled to 20 °C at -0.1 °C s-1 . Subsequently, 1 pl of annealed adapter was mixed with 6 pl of unloaded protein A-Tn5 (0.5 mg ml-1 ), vortexed briefly, spun down and incubated at room temperature for 30 min followed by 10 min at 4 °C. To prepare Tn5-N5 complexes, Adapter-A oligo was mixed with pMENTs oligo at 50 pM, followed by the same annealing and assembly process. For bulk CUT&Tag, protein A-Tn5 coupled with adapter A and B oligos were separately prepared at 0.5 mg ml-1 and then mixed at 1:1 ratio. For ATAC-seq, Tn5 coupled with adapters A and B were assembled similarly at 0.5 mg ml-1 and diluted to 0.05 mg ml-1 . All assembled complexes were stored at -20 °C for up to 6 months. qPCR quantification of DNA damage on model sequences
[0252] Generation of model sequences. Lambda_500bp was amplified from lambda DNA (Lambda_500bp_F and Lambda_500bp_R) with NEB Q5 DNA Polymerase (NEB, cat. no. M0491) and dNTPs (NEB, cat. no. N0447S); Lambda_350bp was amplified from lambda DNA (Lambda_350bp_F and Lambda_350bp_R) with NEB Q5 DNA Polymerase (NEB, cat. no. M0491) and dNTPs (NEB, cat. no. N0447S) supplied with 1% biotin-dUTP (Thermo Scientific, cat. no. R0081).
[0253] Treatment of model sequences. Lambda_500bp was spiked with 1% biotin-labeled lambda_350bp. The model sequences mixture was treated as nuclei in ‘cell fixation and nucleosome depletion’ in Paired-Damage-seq procedures. Purified DNA was treated as ‘DNA damage enzymatic labeling’ in Paired-Damage-seq, and purified with 1.8x Ampure XP beads. To induce nicks at different positions on lambda_500bp, it was treated with 10U Nt.AIwI or Nt.BstNBI at 37 °C for 60 min. Lambda_500bp with or without nicks was spiked with 1% biotin-labeled lambda_350bp. The model sequences mixture was treated as nuclei in ‘cell fixation and nucleosome depletion’ in Paired-Damage- seq procedures. All the DNA was purified with 1.8x Ampure XP beads.
[0254] DNA enrichment for qPCR. Enrichments were performed with Dynabeads MyOne Streptavidin Cl beads (ThermoFisher Scientific, cat. no. 65001, 10 pl per sample). Beads were washed three times with 100 pl of lx BW-T buffer (5 mM pH 8.0 Tris-HCl, 1 M NaCl, 0.5 mM EDTA, 0.05% Tween-20) and resuspended beads in 50 pl of 2x BW buffer (10 mM pH 8.0 Tris-HCl, 2 M NaCl, 1 mM EDTA). Then 50 pl of beads were added to each sample and rotated at 10 rpm for 60 min. Samples were washed three times with 100 pl of lx BW-T buffer, once with lx STE buffer (10 mM pH 8.0 Tris-HCl, 50 mM NaCl, 1 mM EDTA) and once with ultrapure water. Beads were resuspended in 20 pl of ultrapure water and incubated at 95 °C for 5 min to elute DNA from beads. qPCR quantification of enrichment. Quantitative PCRs (qPCRs) were performed using lambda_500bp primers and lambda_350bp primers with Luna Universal qPCR Master Mix (NEB, cat. no. M3003S). Input and pull-down were analyzed.
[0255] Paired-Damage-seq procedures
[0256] Cell fixation and nucleosome depletion. This and some of following steps are adopted from Paired-Tag and SIMPLE-seq. Nuclei were suspended in 2.5 ml NIB-HEPES buffer, and 123 pl of 37% formaldehyde (Sigma, cat. no. F8775) was added. The suspension was gently rotated at room temperature for 10 min. Fixed nuclei were centrifuged at 600g, 4 °C for 8 min, washed with cold NIB-Tris buffer (10 mM Tris-HCl pH 7.5, 10 mM NaCl, 3 mM MgC12, lx Protease Inhibitor, 0.5 U pl-1 RNase OUT, 0.5 U pl-1 SUPERase Inhibitor, 0.1% IGEPAL CA630, 0.1% Tween-20) and centrifuged again under the same conditions for 10 min. The nuclei were resuspended in 800 pl of lx NEBuffer r2.1 (NEB, cat. no. B6002S) with 0.3% SDS and incubated at 42 °C, and spun at 300 rpm for 30 min using a ThermoMixer (Eppendorf). After incubation, the nuclei were centrifuged at 600g for 10 min and sequentially washed twice with 200 pl of NIB-Tris buffer to quench SDS and once with 200 pl of lx NEBuffer r2.1 containing 0.1% Triton X100.
[0257] DNA damage enzymatic labeling. Nuclei were suspended in 84 pl of enzymatic labeling buffer (71 pl of H2O, 10 pl of lOx NEBuffer 2 (NEB, cat. no. B7002S), 1 pl of SUPERase Inhibitor, 0.5 |iil of RNase OUT and 1 pl of 10% BSA (Sigma, cat. no. A1595-50ML) and 0.1% Triton X100). 8 l of Fpg (NEB, cat. no. M0240L) and 8 pl of Endo IV (NEB, cat. no. M0304L) were added into nuclei suspension and incubated at 37 °C for 16 h. Nuclei were spun down at 600g, for 10 min and suspended in 79 pl of enzymatic labeling buffer (66 pl of H2O, 10 pl of lOx NEBuffer 3 (NEB, cat. no. B7003S), 1 pl of SUPERase Inhibitor, 0.5 pl of RNase OUT and 1 pl 10% BSA (Sigma, cat. no. A1595-50ML) and 0.1% Triton X100). Next, 5 pl of Bst full-length polymerase (NEB, cat. no. M0328S), 10 pl of Taq DNA Ligase (NEB, cat. no. M0208L), 2 pl of P-Nicotinamide adenine dinucleotide (NAD+ , NEB, cat. no. B9007S), 4 pl of mixed dNTPs (5 pM for each dATP (NEB, cat. no. N0440S), dCTP (NEB, cat. no. N0441S), dGTP (NEB, cat. no. N0442S) and biotin-dUTP (Thermo Scientific, cat. no. R0081)) were added into nuclei suspension and incubated at 37 °C for 6 h.
[0258] Antibody staining and targeted tagmentation. Anti-biotin antibody (Abeam, cat. no. ab234284, 10 pl of 0.2 mg ml-1per tube) and barcoded protein A-Tn5 (1 pl of 0.5 mg ml-1per tube) were mixed in 20 pl of lx Med Buffer number 1 (20 mM HEPES pH 7.5, 300 mM NaCl, 0.5 mM Spermidine, lx Protease Inhibitor cocktail, 0.5 U pl-1SUPERase IN, 0.5 U pl-1RNase OUT, 0.01% IGEPAL CA630, 0.01% Digitonin and 2 mM EDTA).
[0259] The mixture was rotated at room temperature for 60 min. Nuclei were centrifuged at 600g, 4 °C for 10 min, washed twice with 200 pl of Med Buffer number 1 and aliquoted (3.6 million nuclei) into 12 Maximum Recovery tubes (Axygen, MCT-150-L-C). Each aliquot was resuspended in 50 pl of Med Buffer number 1, mixed with the antibody-protein A-Tn5 complex and incubated overnight at 4 °C with rotation. The nuclei were centrifuged at 600g, 4 °C for 10 min, washed twice with 50 pl of Med Buffer number 2 (similar to Med Buffer number 1 but without EDTA) and resuspended. Tagmentation was initiated by adding 2 pl of 250 mM MgC12 and incubating at 37 °C, at 550 rpm for 60 min in a ThermoMixer. The reaction was quenched with 16.5 pl of 40.5 mM EDTA. Nuclei were centrifuged at 1,000g at 4 °C for 10 min and proceeded immediately to reverse transcription.
[0260] Reverse transcription. Nuclei pellets were resuspended in 20 pl of reverse transcription mix containing barcoded reverse transcription primers (numbers 1-12), prepared with lx reverse transcription Buffer, 1 M Betaine, 2 pM TSO, 5 mM DTT, 6.25 mM MgC12, 0.5 mM dNTPs, 0.5 U pl-1SUPERase IN, 0.5 U pl-1RNase OUT, 2.5 pM barcoded T15 primer, 2.5 pM barcoded N6 primer and 1 U pl-1Maxima Reverse H Minus Reverse Transcriptase (Invitrogen, cat. no. EP0751). The reaction was carried out in a thermocycler using the following program: step 1, 42 °C x 90 min; step 2, 50 °C x 10 min; step 3, 8 °C x 12 s, 12 °C x 45 s, 25 °C x 45 s, 37 °C x 45 s, 45 °C x 45 min, go to step 2 two more times; step 3 50 °C x 10 min and hold at 4 °C. The nuclei were pooled into 1.5-ml tubes prewashed with 5% BSA in PBS, cooled on ice for 2 min and treated with 4.8 pl of 5% Triton X100. After centrifugation at 1,000g at 4 °C for 10 min, the nuclei were immediately processed for ligation-based combinatorial barcoding.
[0261] Ligation-based combinatorial barcoding. Nuclei were resuspended in 1 ml of lx NEBuffer r3.1 (NEB, cat. no. B6003S) and transferred to a ligation mix containing: 2,262 pl of H2O, 500 pl of lOx T4 DNA Ligase Buffer (NEB, cat. no. B0202S), 50 pl of 10 mg ml-1BSA, 100 pl lOx NEBuffer r3.1, 100 pl of T4 DNA Ligase (NEB, cat. no. M0202L). The ligation reaction mix (20 pl per well) was distributed across two Barcode-plate-R02 plates using a multichannel pipette and incubated at 300 rpm at 37 °C for 30 min in a ThermoMixer. Afterward, 5 pl of R02-Blocking-Solution (264 pl of 100 pM Blocker-R02 oligo, 250 pl of lOx T4 DNA Ligase Buffer, 486 pl of H2O) was added to each well using a multichannel pipette. The reaction was incubated for an additional 30 min. The nuclei were pooled and centrifuged at 1,000g at 4 °C for 10 min. A second round of ligation was performed similarly to the first, except that after 30 min of ligation Termination-Solution (264 pl of 100 pM Quencher-R03, 250 pl of 0.5 M EDTA, 236 pl of H2O) was added to quench the reaction without further incubation. Nuclei were resuspended in PBS, counted and aliquoted into sublibraries containing 5,000-10,000 nuclei each and diluted to 35 pl. Next, 5 pl of 4 M NaCl, 5 pl of 10% SDS and 5 pl of 10 mg ml-1 Proteinase K (NEB, cat. no. P8107S) were added. The mixture was incubated at 850 rpm at 55 °C for 2 h in a ThermoMixer. The lysate was cooled to room temperature, purified with lx SPRI beads (Beckman Coulter, cat. no. B23319) and eluted with 12.5 pl of H2O. The purified product can be stored at -20 or -80 °C for up to 4 weeks.
[0262] Preamplification of barcoded DNA / cDNA. Purified DNA / cDNA (12.5 pl) was mixed with 1.5 pl of lOx TdT buffer and 0.5 pl of 1 mM dCTP (NEB, cat. no. N0447S), denatured at 95 °C for 5 min, chilled on ice for 5 min and incubated with 1 pl of TdT (NEB, cat. no. M0315S) at 37 °C for 30 min. The reaction was then heat-inactivated at 75 °C for 20 min. Anchor Mix, consisting of 6 pl of 5x KAPA HiFi buffer, 0.6 pl of 10 mM dNTPs, 0.6 pl of 10 pM Anchor-Fokl-GSH-Oligo and 0.6 pl of KAPA HiFi HS (KAPA, KK2502), was added to the reaction. Linear amplification was performed at 98 °C for 3 min, followed by 16 cycles of 98 °C for 15 s, 47 °C for 60 s, 68 °C for 2 min, 47 °C for 60 s and 68 °C for 2 min. The final extension was at 72 °C for 10 min, with a hold at 12 °C. The reaction was supplemented with 4 pl of 5x KAPA HiFi buffer, 0.5 pl of 10 mM dNTPs, 2 pl of 10 pM PA-F and PA-R primers, and 0.5 pl of KAPA HiFi HS. Preamplification was carried out for five cycles at 98 °C for 20 s, 65 °C for 20 s and 72 °C for 2.5 min, with a final extension at 72 °C for 2 min. Amplified DNA was purified using SPRI beads with double-size selection (10 + 32.5 pl, 0.2x + 0.65x). The final product was eluted in 35 pl of H2O.
[0263] Endonuclease digestion and second adapter tagging. 17 pl of each purified amplified product was transferred into two tubes for DNA and RNA library construction. To the DNA tube, 2.5 pl of lOx rCutSmart buffer, 1 pl of Sbfl-HF and 4.5 pl of H2O were added, incubating at 37 °C for 60 min. The digestion product was purified using 1.8x SPRI beads and eluted in 11 pl of H2O. For the second preamplification, 4 pl of 5x KAPA HiFi buffer, 0.5 pl of 10 mM dNTPs, 2 pl of 10 pM PA-F and PA-R primers, and 0.5 pl of KAPA HiFi HS were added. The preamplification program included: 98 °C for 3 min; 98 °C for 20 s, 65 °C for 20 s, 72 °C for 2.5 min (five cycles); 72 °C for 2 min; hold at 12 °C. Amplified products were purified using 1.8x SPRI beads and eluted in 20 pl of H2O. Next, 2.5 pl of lOx rCutSmart buffer, 1 pl each of Sbfl-HF and FokI were added, and the reaction was incubated at 37 °C for 60 min. The product was purified with 1.25x SPRI beads and eluted in 10 pl of H2O. Ligation was performed by adding 2 pl of lOx T4 DNA Ligase Buffer, 2 pl of P5 Adapter Mix, 4 pl of H2O and 2 pl of T4 DNA Ligase, with the program: 4 °C for 10 min, 10 °C for 15 min, 16 °C for 15 min and 25 °C for 45 min. The ligation product was purified with 1.25x SPRI beads and eluted in 30 pl of H2O. To the RNA tube, 2 pl of lOx rCutSmart buffer and 1 pl of NotLHF were added, incubating at 37 °C for 60 min. The digestion product was purified using 1.8x SPRI beads and eluted in 11 pl of H2O. For the second preamplification, 4 pl of 5x KAPA HiFi buffer, 0.5 pl of 10 mM dNTPs, 2 pl of 10 pM PA-F and TSO primers, and 0.5 pl of KAPA HiFi HS were added. The preamplification program included: 98 °C for 3 min; 98 °C for 20 s, 65 °C for 20 s, 72 °C for 2.5 min (seven cycles); 72 °C for 2 min; hold at 12 °C. Amplified products were purified with 1.8x SPRI beads and eluted in 17 pl of H2O. Subsequently, 2 pl of lOx rCutSmart buffer and 1 pl of NotLHF were added, followed by incubation at 37 °C for 60 min. The product was purified with 1.25x SPRI beads and eluted in 10 pl of H2O. For tagmentation, 10 pl of 2x Tagmentation Buffer and 0.5 pl of 0.05 mg ml-1 Tn5-N5 were added. The reaction was carried out at 550 rpm at 37 °C for 30 min and cleaned using a QIAquick PCR purification kit, with elution in 30 pl of O.lx Elution Buffer.
[0264] Indexing PCR and sequencing. The PCR mix was prepared with 30 pl of purified P5-tagged DNA or N5-tagged RNA product, 10 pl of 5x Q5 buffer, 1 pl of 10 mM dNTPs, 0.5 pl of 50 pM P5 primer (DNA) or N5 primer (RNA), 2.5 pl of 10 pM P7 primer, 5 pl of H2O and 1 pl of NEB Q5 DNA Polymerase. The PCR program included: 72 °C for 5 min; 98 °C for 30 s, 98 °C for 10 s, 63 °C for 30 s, 72 °C for 1 min (repeated 8-10 times to reach 10 nM concentration); 72 °C for 1 min; hold at 12 °C. Libraries were cleaned using 0.85x (42.5 pl) SPRI beads and sequenced on the NovaSeq X platform (Illumina) with read lengths of either paired end 1OO + 8 + 8 + 1OO (read 1 + index 1 + index 2 + read 2).
[0265] Bulk CUT&Tag
[0266] Bulk CUT&Tag assay was performed similarly to previously described (see Kaya- Okur et al. (2019) Nat. Commun. 10:1930, which is incorporated herein by reference in its entirety) to generate DNA damage or H3K9me3 profiles. For DNA damage CUT&Tag, nuclei were treated similarly to the Paired-Damage-seq procedures before antibody staining. Next, nuclei were washed and counted, 1 million permeabilized nuclei were spun down at
[0267] I,000g at 4 °C for 10 min. Biotin (Abeam, cat. no. ab234284) or H3K9me3 (Abeam, cat. no. ab8898) antibody -protein A-Tn5 (1 pl of 0.5 mg ml-1) were added, and the mixtures were rotated at 4 °C overnight. Nuclei were spun down at 600g at 4 °C for 10 min, then resuspended in 50 pl of Med Buffer no. 2 and this was repeated twice. The tagmentation reaction was initiated by adding 2 pl of 250 mM MgC12 (Sigma, cat. no. M1028) and was carried out at 550 rpm at 37 °C for 60 min in a ThermoMixer (Eppendorf). The reaction was quenched by adding 16.5 pl of 40.5 mM EDTA (Thermo Scientific, cat. no. 15575020). Nuclei were then spun down at 1,000g at 4 °C for 10 min and proceeded to nuclei lysis and DNA purification by PCR cleanup kit (Zymo, cat. no. D4033) for PCR amplification with N5 primer and N7 primer.
[0268] ATAC-seq
[0269] An optimized ATAC-seq protocol similarly to previously described (see Buenrostro,
[0270] J.D., et al. (2013) Nat Methods 10:1213-1218, which is incorporated herein by reference in its entirety) was used to generate chromatin accessibility profile. Briefly, nuclei were isolated and counted, 50,000 permeabilized nuclei were spun down at 1,000g at 4 °C for 10 min. Nuclei were resuspended in 45 pl of l.lx Tagmentation Buffer (36.7 mM Tris-Ac, pH 7.8 (ThermoFisher Scientific, cat. no. BP- 152), 12.1 mM MgAc (Sigma, cat. no. M2545), 73.3 mM KAc (Sigma, cat. no. P5708) and 17.8% DMF (EMD Millipore, cat. no.
[0271] DX1730)). Tn5 was then added, and reactions were carried out at 37 °C at 550 rpm for 30 min in a ThermoMixer. Reactions were quenched by adding 25 pl of 40 mM EDTA and nuclei were then spun down at 1,000g at 4 °C for 10 min and proceeded to nuclei lysis and DNA purification by PCR cleanup kit for PCR amplification with N5 and N7 primers.
[0272] Nucleus RNA-seq
[0273] Bulk nuclei RNA-seq assay was performed similarly to previously described to generate transcriptome profile. Briefly, nuclei were isolated similarly to Paired-Tag procedures and then counted, and 15,000 permeabilized nuclei were spun down at 1,000g at 4 °C for 10 min. Nuclei pellets were resuspended in 20 pl of reverse transcription mix (lx reverse transcription Buffer, 1 M Betaine (Sigma, cat. no. B0300-1VL), 2 pM TSO, 5 mM DTT (Sigma, cat. no. D9779), 6.25 mM MgCh (Sigma, cat. no. 63069), 0.5 mM dNTPs (NEB, cat. no. N0447S), 0.5 U pl-1SUPERase IN (Invitrogen, cat. no. AM2694), 0.5 U pl-1RNase OUT (Invitrogen, cat. no. 10777-019), 2.5 pM barcoded T15 primer and 2.5 pM barcoded N6 primer and 1 U pl-1Maxima Reverse H Minus Reverse Transcriptase (Invitrogen, cat. no. EP0751)). The reverse transcription was performed in a thermocycler with the following program: step 1, 42 °C x 90 min; step 2, 50 °C x 10 min; step 3, 8 °C x 12 s, 12 °C x 45 s, 25 °C x 45 s, 37 °C x 45 s, 45 °C x 45 min, go to step 2 two more times; step 3, 50 °C x 10 min and hold at 4 °C. After the reaction, the nuclei were cooled on ice for 2 min, with 0.4 pl of 5% Triton X100 (Sigma, cat. no. T9284). Nuclei were then lysed and proceeded to cDNA amplification with TSO primer and R01 PAF, followed by 1.25x SPRI beads purification and elution in 11 pl of H2O. Then 10.5 pl of 2x Tagmentation Buffer (66 mM Tris-Ac, pH 7.8 (ThermoFisher Scientific, cat. no. BP- 152), 22 mM MgAc (Sigma, cat. no. M2545), 133 mM KAc (Sigma, cat. no. P5708) and 32% DMF (EMD Millipore, cat. no. DX1730)), 0.5 pl of 0.05 mg ml-1Tn5-N5 was added and tagmentation reactions were carried out at 550 rpm at 37 °C for 30 min in a ThermoMixer (Eppendorf) followed by cleanup using QIAquick PCR purification kit and elution in 30 pl of 0.1 x Elution Buffer (QIAGEN) for PCR amplification with N5 and P7 primers.
[0274] Cell treatment with hydrogen peroxide of high dosage HeLa S3 (human, ATCC CCL-2.2) cells were cultured according to standard procedures in DMEM (Gibco, cat. no. 10569010) supplemented with 10% fetal bovine serum (Gibco, cat. no. 16000044) and 1% penicillin- streptomycin (Gibco, cat. no. 10378016) at 37 °C with 5% CO2. Cells were not authenticated or tested for mycoplasma. Cells were sown on plates 48 h before treatment with H2O2 at the density of 3 x 105cells per culture dish. Subconfluent cells were treated with H2O2, in a complete medium, at the final concentration of 5 or 25 mM (Thermo Scientific Chemicals, cat. no. 202465000) for 2 h at 37 °C, 5% CO2. After treatment, the cells were collected.
[0275] Nontargeting tagmentation control
[0276] After fixation and nucleosome depletion as in Paired-Damage-seq, Bulk pA-Tn5-only CUT&Tag assay was performed similar to bulk CUT&Tag to generate un-targeting DNA profile. Nuclei were treated similarly to the Paired-Damage-seq procedures before antibody staining. Next, nuclei were washed and counted, 1 million permeabilized nuclei were spun down at 1,000g, 4 °C for 10 min, and resuspended in 50 pl of Med Buffer no. 2. Assembled protein A-Tn5 (1 pl of 0.5 mg ml-1) were added followed by overnight incubation. The tagmentation reaction was initiated by adding 2 pl of 250 mM MgCh (Sigma, cat. no. M1028) and was carried out at 550 rpm at 37 °C for 60 min in a ThermoMixer (Eppendorf). The reaction was quenched by adding 16.5 pl of 40.5 mM EDTA (Thermo Scientific, cat. no. 15575020). Nuclei were then spun down at 1,000g at 4 °C for 10 min, and proceeded to nuclei lysis and DNA purification by PCR cleanup kit (Zymo, cat. no. D4033) for PCR amplification with N5 and N7 primers.
[0277] In situ prerepair
[0278] After fixation and nucleosome depletion as in Paired-Damage-seq, nuclei were suspended in 84 pl of enzymatic labeling buffer (71 pl of H2O, 10 pl of lOx NEBuffer 2 (NEB, cat. no. B7002S), 1 pl of SUPERase Inhibitor, 0.5 pl of RNase OUT and 1 pl of 10% BSA (Sigma, cat. no. A1595-50ML) and 0.1% Triton X100). 8 pl of Fpg (NEB, cat. no. M0240L) and 8 pl of Endo IV (NEB, cat. no. M0304L) were added into nuclei suspension and incubated at 37 °C for 16 h. Nuclei were spun down at 600g for 10 min and suspended in 79 pl of enzymatic labeling buffer (66 pl of H2O, 10 pl of lOx NEBuffer 3 (NEB, cat. no. B7003S), 1 pl of SUPERase Inhibitor, 0.5 pl of RNase OUT and 1 pl of 10% BSA (Sigma, cat. no. A1595-50ML) and 0.1% Triton X100). Next, 5 pl of Bst full-length polymerase (NEB, cat. no. M0328S), 10 pl of Taq DNA Ligase (NEB, cat. no. M0208L), 2 pl of 0- nicotinamide adenine dinucleotide (NAD+ , NEB, cat. no. B9007S), 4 pl of mixed dNTPs (5 pM for each dATP (NEB, cat. no. N0440S), dCTP (NEB, cat. no. N0441S), dGTP (NEB, cat. no. N0442S), dTTP (NEB, cat. no. N0443S)) were added into nuclei suspension and incubated at 37 °C for 6 h.
[0279] SSB induction by Nickase NtBbvCI
[0280] After fixation and nucleosome depletion as in Paired-Damage-seq, nuclei were spun down at 600g for 10 min and then washed twice with lx NEB buffer 2.1 (50 mM NaCl, 10 mM Tris-HCl and 10 mM MgCh). After removal of supernatant, the nuclei were resuspended in 50 pl of lx NEB buffer 2.1 and treated with 50 U of Nt.BbvCI for 2 h at 37 °C. The nuclei were collected by centrifugation at 600g for 10 min.
[0281] Data analysis procedures
[0282] Preprocessing of Paired-Damage-seq data. Cellular barcodes and the linker sequences are in read 2. The first bases of BC1, BC2 and BC3 locate within the 84th-87th, 47th-50th and 10th-13th positions of read 2 and the exact positions were identified by matching the linker sequences. A reference index for all possible cellular barcode combinations (192 x 192 x 12) was generated and the barcode sequences were mapped to the cellular barcode reference using bowtie with the parameters: ‘-v 1 -m 1 -norc’. Nextera adapter sequences were removed from 3' of DNA and RNA libraries, Poly-dT sequences were further removed from 3' of RNA libraries, with TrimGalore . Low-quality reads with Q < 30 were excluded from downstream analysis.
[0283] Evaluation of collision rate. Reads from species-mixing test were extracted based on cellular barcodes (BC1 = 11 or 12 of HeLa cells Paired-Damage-seq dataset) and mapped to a reference genome using STAR with the combined reference genome (GRCh37 for human and GRCm38 for mouse). Duplicates were removed based on the mapped positions, UMI, cellular barcodes and sub library indices. For estimation of barcodes collision rate, nuclei with less than 75% UMIs mapped to one species were classified as mixed cells. The effective barcode collision rate is calculated according to the previously reported method under the assumption that the loading of cells into wells is Poisson.
[0284] Briefly: the numbers of barcodes contain at least one human cell, at least one mouse cell and contain both human and mouse cells are denoted as N1, N2 and N12. respectively. The barcode
[0285] Reads mapping. Adapter- trimmed reads were first mapped to a mouse GRCm38 (mm 10) or human GRCh38 (hg38) reference genome with STAR (v.2.6.0a) for RNA or bowtie2 for DNA. Mapped DNA reads were further filtered by MAPQ > 10. Duplicates were removed based on the mapped position, UMI, cellular barcode and sublibrary indices. Low- coverage nuclei (<1,000 transcripts and <1,000 unique DNA reads for HeLa, <200 transcripts and <1,000 unique DNA reads for mouse brain) were removed from downstream analysis. Bigwig files were generated with deepTools.
[0286] Clustering of Paired-Damage- seq RNA profiles. RNA bam files were converted to sparse matrix with cells as columns and genes as rows. The cell clustering based on transcriptome profiles was performed with Seurat. Briefly, the cell-to-gene counts matrix were normalized with SCTransform function and variable genes were used for principal component analysis, visualized with uniform manifold approximation and projection (UMAP). To validate the annotation results, reference mapping of Paired-Damage-seq RNA profile to published snRNA-seq dataset was performed. The overlap coefficients (O) was calculated according to the number of cells with the labels from the Paired-Tag dataset (A), from Yao et al. (Yao et al. (2021) Nature 598:103-110, which is incorporated herein by reference in its entirety) (B) and from reference mapping (R) (i, row index; j, column index):
[0287] Peak calling. Peak calling was performed using MACS3 (ref. 51) with the following parameters: ‘-q 0.01 -nomodel -shift 200 -extsize 500 -keep-dup all -nolambda -nomdel’. A cutoff of P < 0.05 (Benjamini-Hochberg procedure) was used to filter the results. To identify induced peaks in HeLa cells, control HeLa cells dataset was used as control file for peak calling using MACS3 with the same parameters. Peaks identified in control HeLa cells were classified as conserved peaks. To identify cell-type- specific peaks for all the mouse cortex cell types, the ‘tl.marker_regions’ function from SnapATAC2 was used to aggregate the DNA damage signal across cells and used z-scores to identify specifically enriched peaks, with a P value cutoff of 0.05. DNA damage signal-gene expression correlation analysis. To minimize the influences from stochastic DNA damage signal, the SEACell software was used to group the single cells into 300 metacells based on transcriptome profiles. To analyze the associations between DNA damage levels and gene expression levels, the following was calculated: the ranked correlation score between the raw count of each gene and the DNA damage read in peaks within the 500-kb span of the gene in constructed SEACells using the ‘add_cor_scores’ function of SnapATAC2. To estimate the background noise levels, the identities of metacells were shuffled and recalculated the corresponding PCCs between DNA damage levels and gene expression levels.
[0288] Motif enrichment and GO analysis. Motif enrichment for DNA damage hotspots of each cell type was performed using the ‘findMotifsGenome. pl’ function of HOMER (v.4.9). GO term enrichment analysis for DNA damage hotspots or peak-associated CREs was performed using GREAT with default parameters.
[0289] GWAS enrichment. The comparison of mouse DNA damage hotspots with genome-wide association studies (GWASs) of human phenotypes were performed similarly to a previous study, briefly: the DNA damage peaks were converted from mmlO to hg38 coordinates using the liftOver software with parameters ‘-minMatch = 0.5’; the human hg38 coordinates were then lifted back to mmlO and only kept the peaks that could be mapped to original loci. Next, peaks wider than 1 kb in the human genome were excluded from further analysis. The following GWAS summary statistics for quantitative traits were used: major depressive disorder, posttraumatic stress disorder, neuroticism, lymphocyte count, basophil count, amyotrophic lateral sclerosis, monocyte count, anorexia nervosa, Alzheimer’s disease, autism spectrum disorder, obsessive compulsive disorder, atopic dermatitis, breast cancer, educational attainment and stroke. Linkage disequilibrium score regression (World Wide Web at github.com / bulik / ldsc) was performed to estimate the enrichment coefficients of each trait compared to background control.
[0290] Data availability
[0291] Raw sequencing and processed data generated in this study are available from NCBI
[0292] Gene Expression Omnibus (GEO) (World Wide Web at ncbi.nlm.nih.gov / geo / ) under the accession number GSE268567. Other external datasets were downloaded from the GEO with the following accession numbers: AP-seq (GSE121005), CLAPS-seq (GSE181312), Paired- seq and Paired-Tag (GSE152020), Droplet Paired-Tag (GSE224560), snATAC-seq and Paired-Tag of aging mouse brain (GSE187332), CoTECH (GSE158435), ENCODE (World Wide Web at encodeproject.org / ) with the following accession numbers: HeLa DNase-seq (ENCFF977IGB), HeLa H3K4me3 chromatin immunoprecipitation with sequencing (ChlP- seq) (ENCFF578NOK), HeLa H3K36me3 ChlP-seq (ENCFF248WXB), HeLa H3K4mel ChlP-seq (ENCFF360CQR), HeLa H3K9me3 ChlP-seq (ENCFF712ATO), HeLa H3K27ac ChlP-seq (ENCFF392EDT), HeLa H3K27me3 ChlP-seq (ENCFF512TQI) and HeLa 18- state model chromatin states (ENCSR098REA); 4DN Data Portal (https : / / data.4dnucleome. org / )) with the following accession numbers: HeLa cell Hi-C (4DNFIBMVFF0F), mouse cortex Hi-C (4DNFIB16WAKX); Mouse Brain scRNA-seq (World Wide Web at portal.brain-map.org / atlases-and-data / maseq) and the lOx Genomics website (World Wide Web at 10xgenomics.com / ).
[0293] Code availability
[0294] Custom scripts used for analyzing Paired-Damage-seq datasets are available from GitHub (World Wide Web at github.com / czhulab / Paired-Damage-seq).
[0295] Example 10: Single-cell joint analysis of DNA damage and transcription
[0296] More than 70% of endogenous lesions in human cells are apyrimidinic sites and single-strand DNA breaks (SSBs), which can arise during metabolism or base hydrolysis, for example, the repair of oxidative DNA damage (such as 8-oxoguanine (8-oxoG)) by base excision repair and the removal of mis-incorporated ribonucleotides by ribonucleotide excision repair. An in situ ‘labeling-by-repair’ strategy was designed to attach biotin to the damaged sites, built on the in vitro base excision repair reaction for the detection of DNA damaged sites (Fig. 1A). For the sake of conciseness, hereafter, 8-oxoG, apyrimidinic site and SSBs is collectively referred to as DNA damage. By treating model DNA sequences with reaction conditions required by ‘labeling-by-repair’, it was validated that the sample handling processes did not introduce detectable false positive labeling (Fig. 5F). To minimize the labeling bias raised from chromatin epigenetic states, the nucleosome depletion step was performed on formaldehyde crosslinked cells (Fig. 5G). Next, cells were treated with DNA repair proteins Fpg and Endo IV to remove pre-existing 8-oxoG and cleave apyrimidinic sites, giving SSBs. Nick translation with Bst polymerase was then performed to incorporate biotin-modified deoxyuridine triphosphates (dUTPs) at the pre-existing and newly generated SSBs, followed by nick sealing with Taq DNA Ligase (Fig. 1A). Fpg enzyme, in addition to cleave 8-oxoG, can also convert 5'-dRP to phosphate group and Endo IV is able to remove the 3 '-phosphate group, providing 3'-OH and 5 '-phosphate ends required by Bst DNA polymerase and Taq DNA ligase (Fig. 5H). The labeled cells were then counted and pipette- aliquoted into 12 tubes; each tube contains 300,000-500,000 nuclei. The cells are then incubated with biotin antibody, followed by targeted tagmentation with protein A-Tn5, attaching the tube-specific DNA barcodes (BC1) to the DNA damage sites. In situ reverse transcription is performed with primers containing the same set of BC1 barcodes. The above steps were performed in intact nuclei and were compatible with different single-cell barcoding approaches. The cells were labeled by two rounds of ligation-based combinatorial barcoding to introduce BC2 and BC3 (see Rosenberg et al. (2018) Science 360:176-182, which is incorporated herein by reference in its entirety): DNA fragments and complementary DNA (cDNA) molecules from the same cells receive the identical combination of barcodes (BCs 1, 2 and 3). Cells were then lysed and barcoded DNA / cDNA molecules are purified, followed by library amplification steps similar to Paired-Tag. After sequencing, the cellular barcodes simultaneously attached to DNA and cDNA fragments were used to assign each read to individual cells.
[0297] Validation experiments were performed on cells with artificially generated DNA damage sites (Example 9, Figs. 1J-1K, and Figs. 5-6). The fixed HeEa nuclei were treated with nickase Nt.BbvCI to generate SSBs in the genome, and performed bulk Paired-Damage- seq experiment. The detected DNA damage levels showed a well positive correlation with the densities of nickase cutting sites in 10,000-base pair (bp) nonoverlapping bins (Spearman’s p = 0.67); while control cells without Nt.BbvCI treatment displayed a weakly negative correlation (Spearman’s p = -0.27) (Fig. IK and Fig. 51). Next, HeEa cells were treated with varying concentrations of H2O2 (0 pM, 200 pM, 5 mM and 25 mM) and as expected, higher oxidant concentrations increased the detected DNA damage levels (Fig. 6H). Thus, Paired- Damage-seq can detect the induced DNA damage sites. To further validate the detected signal without treatments are from endogenous DNA damage, prerepair of nuclei was performed to reduce the level of endogenous lesions. An obvious signal decrease on DNA damage peak regions was observed from prerepaired nuclei compared to control group (Fig. 1 J). These results collectively indicate that Paired-Damage- seq faithfully detects DNA damage signal from bulk cell populations.
[0298] Next, Paired-Damage- seq was performed on cultured HeEa cells and obtained 64,794 cells after filtering (minimal numbers of transcripts 1,000, and minimal numbers of DNA fragments 1,000; Figs. 1B-1C). At the moderate saturation level (PCR duplication rates of 51.47 and 24.29%, for DNA and RNA libraries, respectively), Paired-Damage- seq captured the median numbers of 5,320 DNA damage loci and 4,453 transcripts per cell. Compared to previous single-cell multiomics sequencing methods for joint RNA and epigenome analyses, Paired-Damage-seq displayed similar sensitivity in RNA detection, despite different DNA modalities being measured (Figs. ID- IE). The barcode collision rate is estimated at 11.4% from a species-mixing experiment (Fig. IF and Fig. 5C). The aggregated single-cell signals were compared with DNA damage and RNA datasets generated with previous methods for bulk cell population analysis. For the RNA modality, Paired-Damage-seq showed good agreement in the detected gene expression levels compared to the bulk RNA-seq profile of the same cell line (Pearson’s correlation coefficient (PCC) 0.87) (Fig. 1G and Fig. 5D). For the DNA damage modality, the whole-genome reads distribution patterns also correlated well with public bulk dataset generated by CLAPS-seq (a chemical labeling-based approach for 8- oxoG sequencing) (PCC 0.72) (Fig. 1H and Fig. 5 J) . On the other hand, the signals from assay for transposase-accessible chromatin using sequencing (ATAC-seq) and nontargeting tagmentation control poorly correlated with both Paired-Damage-seq DNA and the refence DNA damage profiles (Fig. IE). To evaluate the sensitivity of Paired-Damage-seq in detecting DNA damage signals, the numbers of cells for DNA modality were downsampled and damage profiles could be robustly obtained by aggregating signals from the minimal of 1,000 cells (PCC > 0.9) (Fig. Ik). These results indicate that Paired-Damage-seq robustly measures oxidative damage and SSBs with gene expression from single cells.
[0299] In Paired-Damage-seq, different samples could be multiplexed in one experiment to minimize the risk of batch effects. Treating cells with subtoxic concentration of H2O2 can induce senescence-like growth arrest without affecting the viability for most cells. HeLa cells were treated with this subtoxic 200 pM H2O2 for 2 hours to introduce oxidative stress, followed by medium replacement and continued incubation for 48 hours. Cells collected at different time points after treatment were barcoded with BC1 and multiplexed in a single experiment (Fig. IB and Figs. 6A-6B). Cells without H2O2 treatment and those collected at 0 and 2 hours after replacing the H2O2-containing medium exhibited similar transcriptional states. For cells collected at 6 hours and longer after replacing the medium, their transcriptional states were distinct from control cells; the affected genes were enriched in hippo signaling and cAMP-related pathways46 (Fig. IB and Figs. 2I-2K). In contrast to gene expression, dimension reductions of cells based solely on DNA damage signals displayed a salt-and-pepper pattern, albeit with a weak gradient distribution across different treatments: unsupervised clustering could not resolve these cell states (Fig. 6B). This is expected as the large-scale and stochastic-acquired DNA damage could overwhelm the subset of genome regions showing functional influences; thus, identifying cell states based on the noisy singlecell damage signals alone is challenging.
[0300] Example 11: Chromatin states affect genome susceptibility
[0301] The abundances of endogenous DNA damage result from the balance between the formation and repair of damage. In line with the observations that in cancer genome singlenucleotide variants are preferentially accumulated in heterochromatin, due to the condensed structure presents as a barrier for efficient DNA repair, a higher average DNA damage level in compartment B was detected compared to compartment A (average reads per kilobase per million mapped reads (RPKM) 0.37 versus 0.34), which is opposite to ATAC-seq signals (Fig. 2A and Fig. 6C). It was found that in cells without stress treatments, the gene body region DNA damage levels of highly transcribed genes were higher than those in less transcribed or repressed genes (average RPKM 0.357 (high), 0.326 (low) and 0.289 (repressed)) (Fig. 6D). This result indicates that the chromatin states such as open chromatin or R-loop conformation during transcription render the DNA more susceptible to endogenous genotoxins. However, even for the highly expressed genes, the average gene body damage level was still not higher than the average level of compartment B (average RPKM 0.357 (gene body) versus 0.361 (B)). Thus, DNA repair machinery could selectively protect genic regions from the accessible chromatin-coupled genome susceptibility, regardless of whether the gene is highly expressed or not.
[0302] The average DNA damage levels for different chromatin states regions were then calculated from the ENCODE. Similar to the results above, genic enhancers (average RPKM: genic enhancer 1, 0.391 and genic enhancer 2, 0.392) and transcribed regions (average RPKM, strong transcription 0.378 and weak transcription 0.367) were hotspots for oxidative DNA damage (Fig. 2B). For distal enhancers, comparable damage levels to transcribing regions were observed; this may result from enhancer RNA transcription or due to the genome regions of open chromatin configuration are more sensitive to damaging stresses (Fig. 2B). On the other hand, czs -regulatory elements (CREs) associated with H3K27me3, such as bivalent enhancers and / or promoters (average RPKM: bivalent enhancer, 0.271 and bivalent or poised enhancer, 0.244) and polycomb-associated heterochromatin (average RPKM: repressed PolyComb, 0.266 and weak repressed PolyComb, 0.284) showed the lowest damage levels (Fig. 2B). Thus, the noncoding elements responsible for cell and tissue identities are also considered as targets for selective protection, irrespective of their activity states. On the other hand, the constitutive heterochromatin was not included in this prioritized protection (Fig. 2B). The data presented herein demonstrate that chromatin epigenetic states associated with active transcription render the genome more susceptible to DNA-damaging stresses. To mitigate this risk, mechanisms for prioritized DNA repair of transcribing and regulatory regions could exist while acting independently of their activation states.
[0303] Example 12: Short tandem repeats are hotspots of inducible DNA oxidation
[0304] To analyze the oxidative stress-induced DNA damage hotspots, the single-cell damage signals according to their treatment conditions were aggregated and performed peak calling with MACS3. 30,477 peaks were detected from cells without treatment (fold enrichment >3 and P < 0.05), which were highly enriched in simple repeats (also known as short tandem repeats, STRs) (log2(ob served (obs) / expected (exp)) 2.107), long interspersed nuclear elements (LINEs) (log2(obs / exp) 0.136) and CpG islands (log2 (obs / exp) 0.334), while depleted from promoters (log2 (obs / exp) -0.979) and exons (log2 (obs / exp) -0.796) (Figs. 6E-6F). By downsampling the reads into the same depth for all conditions, the increasing numbers of peaks (31,495-42,017; fold enrichment >3 and P < 0.05) were detected from cells collected 0 to 48 h after treatment (Fig. 6E and Fig. 6L). Something to note is that, for all conditions, >75% of damaged regions overlapped with at least one of the other conditions and 5% of covered regions overlapped with nontargeting tagmentation control, supporting the hotspots for DNA damage formation are not randomly distributed but are driven by conserved factors among these cell populations, such as base sequence compositions (Fig. 6L). For example, chromatin regions associated with zinc-finger genes and repeats showed high damage levels; as expected, the G-rich short tandem repeat subfamilies (such as (AC)n / (GT)n and (TG)n / (CA)n) displayed highest associations with DNA damage peaks (Fig. 2B and Fig. 7A).
[0305] To identify regions that could further increase in damage levels on hydrogen peroxide treatment-induced stress responses, peak calling was performed on cells collected at different posttreatment time points against control cells and identified the union of 2,399 peaks. The inducible peaks displayed highest enrichment in CpG islands (log2(obs / exp) 2.339), low complexity repeats (log2(obs / exp) 3.193) and STRs (log2(obs / exp) 4.322) (Figs. 7B-7C). For example, on treatment, DNA damage levels increased in simple repeats, predicted Z-form DNA and G-quadruplex-forming sequences (PQS) regions (to 6.36-, 1.35- and 5.73-fold, respectively, after background subtraction) (Fig. 7D). Noncanonical chromatin structures, including G-quadruplexes (G4-DNA), are known to affect gene expression. ROS-mediated 8- oxoG and apyrimidinic site generation in PQS -containing promoters facilitates the formation of G4-DNA for fast gene activation. In line with this, a positive association was also observed between gene expression levels and local DNA damage levels (Spearman’s p = 0.36). Besides, oxidative stress treatment increased the damage levels on endogenous retroviruses in both compartments A and B (to 1.72-fold (A), 1.41-fold (B), after background subtraction); while during this process, chromatin accessibilities reduced at compartment A- residing endogenous retroviruses and remained mostly unchanged for the repressed elements in compartment B (average RPKM 0.322-0.313 (A) versus 0.212-0.216 (B)) (Fig. 7E). These results indicate a subset of repeat elements that could sense oxidative stress and alter gene activities epigenetically.
[0306] Example 13: DNA damage formation associates with epigenetic memory loss
[0307] The epigenetic consequences of oxidative DNA damage formation raised the question of whether the damage-epigenome crosstalk is limited in the ‘sandbox’ (for example, formation of G4-DNA in response to PQS guanine oxidation) or could pose threats to epigenetic memory across the whole genome. As both enhancers and H3K9me3- associated heterochromatin showed high DNA damage levels (Fig. 2B), bulk ATAC-seq and H3K9me3 CUT&Tag were performed on HeLa cells collected 0, 6 and 48 hours after stress treatment. The changes in DNA damage levels negatively correlated with the changes in both chromatin accessibility (PCC -0.356) and H3K9me3 levels (PCC -0.326) (Figs. 2C-2D). This antagonistic relationship was upheld with different window resolutions: PCCs -0.100 and -0.174 at 10-kb, and -0.459 and -0.428 at 1-Mb bin sizes, for chromatin accessibility and H3K9me3, respectively. Moreover, the epigenetic changes happened immediately after the treatment (2 hours incubation with H2O2) and lasted for at least 1-2 cell divisions (celldoubling time 17.5-32.3 hours for HeLa cells) (Fig. 8). Therefore, the association between DNA damage formation and epigenetic changes is not limited to specific short tandem repeat families but also in various regulatory regions.
[0308] To understand the influences of DNA damage on gene expression, the multiomics profiles were considered and performed correlative analysis between DNA damage and gene expression. Considering the highly noisy nature of DNA damage signals, the single cells were first grouped into ‘metacells’ based on their transcriptional similarities using the SEACells software, and then calculated the Spearman’s correlation coefficients between gene expression levels and DNA damage levels (in peaks within the 500-kb range of the gene) across ‘metacells’. Compared to the background, 11,814 positive and 4,362 negative linked gene-damage peak ‘pairs’ (false discovery rate (FDR) <0.05) were observed, suggesting oxidative DNA damage is more likely to disrupt repressive regulatory states (Fig. 2E). This could be due to chromatin decondensation required by loading of DNA repair machineries. As expected, DNA damage peaks showing stronger influences on gene expression are enriched in CREs, evidenced by their high associations with histone marks H3K4mel, H3K4me3 and H3K27ac (Fig. 2F). Collectively, the results presented herein revealed the links between oxidative DNA damage formation and epigenetic memory loss.
[0309] Example 14: Brain cell-type-specific DNA damage landscape
[0310] The results that the chromatin state is a key determinant of DNA damage hotspots prompted the exploration of whether the epigenetic states defining cell types in complex tissues also contribute to selective cell-type genome vulnerability. Paired-Damage- seq was performed on mouse cerebral cortex and sequenced 73,813 cells (minimal numbers of transcript 200, and minimal numbers of DNA fragments 1,000). Cell clustering and annotation based on RNA profiles revealed all main brain subclasses, including excitatory neurons (ExN), inhibitory neurons (InN) and nonneuronal cell types: oligodendrocyte (ODC) precursor cells (OPC), ODCs, astrocytes (AST), microglia cells (MiG), endothelial cells (Endo) and vascular leptomeningeal cells (VLMC). (Fig. 3A and Fig. 9A). To validate the cell clustering results, reference mapping was performed to the reference single-nucleus RNA sequencing (snRNA-seq) dataset using the Azimuth software (Figs. 9B-9C). Next, the DNA damage signals were aggregated for each cell group according to the transcriptome-based clustering results and obtained the mouse brain cell-type- specific maps for hotspots of oxidative DNA damage and SSBs. For the sake of statistical power, the analysis at the subclass level was performed by merging the closely related cell subtypes. It was found that the DNA damage levels were higher at gene that are highly expressed in certain nonneuron cell types (average RPKM highdow: 1.43- (OPC), 1.54- (ODC), 1.41- (AST) and 1.71-fold (MiG)), including Pdgfra in OPCs, Mobp in ODCs, Slcla2 in ASTs and Csflr in MiG, echoing the observations in HeLa cells; meanwhile the enrichment of genome- wide DNA damage signals was not observed on ATAC-seq peak regions (Fig. 3B, Fig. 5G, Fig. 9D, and Fig. 10A). However, this relationship disappeared in neuron cells and nonbrain- specific cell types including endothelial cells and VLMCs. Hence, the impacts of transcription-associated chromatin states on genome vulnerability show intra-tissue and inter-cell-type heterogeneities. For all mouse brain cell types, the average DNA damage levels were higher in compartment B (median DNA damage levels (RPKM) 1.14—1.21) compared to compartment A (0.99-1.04) (Fig. 3C and Fig. 10C). This pattern was conserved with the cultured HeLa cells, suggesting that in addition to inefficient DNA repair activity due to condensed heterochromatin structure, the nuclei compartmentalization by organizing certain genome regions closer to nuclei membrane may also increases their susceptibility to oxidative reactive species, which are generated in the mitochondria. Albeit with varying degrees of enrichment scores, for all cell types, DNA damage hotspots prefer CpG islands (log2(obs / exp) 0.042- 0.776), LINEs (log2(obs / exp) 0.210-0.622), low complexity repeats (log2(obs / exp) 0.129- 0.926), long terminal repeats (log2(obs / exp) 0.269-0.488), simple repeats (log2(obs / exp) 1.314-1.897) and satellites (log2(obs / exp) 1.521-2.142), similar to that in HeLa cells and confirming the contribution from base compositions (Fig. 3D and Fig. 6F). For LINE1, the high DNA damage levels were observed at 5' untranslated regions of LINE1 for both neuronal and nonneuronal cells, which is consistent with the fact that mouse LINE1 5' untranslated regions contain a variable number of GC-rich tandem repeats (Fig. 10B). Thus, a subset of base sequence compositions-associated DNA damage hotspots is conserved across cell types and species.
[0311] Example 15: Neuronal CREs are DNA damage hotspots
[0312] Recent studies revealed that DNA repair hotspots are enriched in a subset of regulatory regions including enhancers, to protect the identities and functions of postmitotic cells (such as neurons). To examine whether such pattern is already established during the formation of DNA damage, the cell-type-resolved DNA damage signals were compared with public snATAC-seq and H3K27ac Paired-Tag datasets. At 250-kb resolution, DNA damage levels are positively correlated with H3K27ac levels for ExN, InN, ODC and MiG (PCC 0.45-0.51) (Fig. 3E and Fig. 11 A). No strong relationship of DNA damage levels with the constitutive heterochromatin histone mark H3K9me3 (PCC -0.15-0.06) was observed; although the higher average damage burdens were constantly observed in these regions (Fig. 3F and Fig. 1 IB). At a finer resolution, DNA damage signals again tended to be enriched in H3K27ac peak regions: neurons and nonneuronal brain cell types (including ODC and MiG cells) showing the highest enrichment; while Endo and VLMCs, which are not brain- specific cell types display weak enrichments (Fig. 3G and Fig. 11C). A higher fraction of active enhancers (H3K27ac+ / ATAC-seq+ peaks) overlapped with DNA damage peaks compared to weak enhancers (H3K27ac- / ATAC-seq+ peaks) for the brain-specific cell types (ExN, 32.2 versus 15.4%; ODC, 15.2 versus 7.9%). Again, for endothelial cells and VLMCs, this relationship was not observed (Endo, 1.7 versus 2.1%; VLMC, 1.3 versus 1.3%) (Fig. 11D). These results indicate a brain- specific mechanism in coordinating enhancer activities via the formation and repair of DNA damage for certain cell types, possibly related to the active DNA demethylation pathway.
[0313] Example 16: The cell-type-specific genome vulnerability
[0314] The prevalence of DNA damage in H3K27ac-marked brain enhancers hinted at possible regulatory functions. Thus, differential peak analysis was performed to explore whether DNA damage hotspots’ influences are cell identity dependent. 3,701-36,903 cell- type-specific peaks were identified with the SnapATAC2 software (P < 0.05). Among them, detected were fewer marker peaks from the nonproliferating brain cell types (ExN, InN, ODC and AST) compared to proliferating cells (OPC and MiG) and nonbrain- specific cell types (Endo and VLMC) (Fig. 4A). Genomic Regions Enrichment of Annotations Tool (GREAT) analysis revealed candidate CREs67 -associated DNA damage peaks correlate with specific brain functions, for example: adult locomotory behavior (ExN), long-term synaptic potentiation (InN) and myelin assembly (ODC). Motif analysis was performed on the cell- type- specific damage peaks and revealed potential regulators could be affected by DNA damage (Figs. 4B-4C). For example, the YY1 binding motif contains guanine-guanine dinucleotide, which was found enriched in multiple cell types. PAX6 plays an important role in neurogenesis during development, and its binding motif was enriched in ExN-specific DNA damage peaks. It was also observed that the poly-dG sequence (predicted to be the Zfp740 binding motif) was highly enriched for both proliferating brain cell types (OPC and MiG) and nonbrain- specific cell types (Endo and VLMC) (Fig. 4B). These results indicate that brain cell types display diverse patterns of selective genome vulnerability, shaped by both conserved and brain- specific mechanisms, posing varying functional influences for different cell identities.
[0315] To explore whether the selective genome vulnerability could be associated with disease risks, the human orthologues of mouse DNA damage hotspots were identified with reciprocal homology searches and performed linkage disequilibrium score regression analysis to predict the associated traits (Fig. 4D, Fig. 12A, and Example 9). The statistical power is limited by intrinsic high background level of DNA damage profiles, however, traits associated with specific cell types were still observed (Fig. 4D). For example, the highest enriched trait among them was major depressive disorder in ExN (P = 0.0038); it has been found DNA oxidation was elevated in patients with major depressive disorder compared to healthy individuals. Thus, the DNA damage hotspots may acquire oxidation faster and potentially result in the pathology-associated gene programs. At the less statistically notable levels, the associations of damage hotspots with neurodegenerative disorders were observed, including amyotrophic lateral sclerosis in ODC (P = 0.086) and AST (P = 0.070), and Alzheimer’s disease in MiG (P = 0.055). For example, INPP5D regulates inflammasome in microglia, and its expression level is associated with Alzheimer’s disease risk. Multiple MiG- specific DNA damage hotspots were observed in mouse lnpp5d loci that showed decreased chromatin accessibility in aged mouse brain (Fig. 4E). Tpcnl is another Alzheimer’s disease- associated gene with elevated DNA damage levels compared to other cell types (that is, InN) and decreased chromatin accessibility in enhancers within its intron in aged mice (Fig. 12B). The association of endothelial DNA damage hotspots were observed with breast cancer, hinting an inter-tissue conserved genome susceptibility (Fig. 4D). Therefore, the accumulation of DNA damage was nonrandom and exhibited a cell-type-specific pattern; such selective genome vulnerability increases the risk of gene program dysregulation in different disease-associated cell types.
[0316] Finally, to investigate the relationships between DNA damage and epigenetic alterations in regions outside active CREs, the distribution patterns of DNA damage and the heterochromatin H3K9me3 histone mark were compared. The age-associated heterochromatin loss was previously discovered in the mouse brain: The differences in H3K9me3 levels between 18- and 3-month-old mice for each brain cell type were calculated and compared with the DNA damage levels detected in 2-month-old mice. At 100-kb resolution, DNA damage levels and H3K9me3 loss displayed neither positive nor negative genome- wide associations. However, when focusing on the top 1% most damaged regions, negative associations were observed for ExN (PCC -0.19), InN (PCC -0.11), ODC (PCC -0.31) and AST (PCC -0.18), but not for the OPC that undergo proliferation in the adult brain (PCC -0.03) (Figs. 4F-4H and Figs. 12C-12D). Thus, the persistent formation and repair of DNA damage could be a potential cause for the progressive loss of epigenome information in postmitotic cells. Pinpointing the damage hotspots could facilitate the identification of vulnerable targets before the pathological changes occur for early interventions.
[0317] Presented herein are methods comprising Paired-Damage- seq, which is a high- throughput method for single-cell joint analysis of oxidative and single-stranded DNA damage with gene expression to reveal cell-type-resolved DNA damage landscapes. Both HeLa cells and mouse brain were analyzed and revealed the selective genome vulnerability is driven by local epigenetic states. By introducing epigenome fluctuations with exogenous oxidative stress in the HeLa cells it was discovered that the degree of epigenetic information loss is associated with the accumulation of inducible oxidative DNA damage. It was also observed herein that in mouse brain, a subset of genome regions most susceptible to genotoxins were also hotspots for age-associated epigenome decay and such cell-type- specific genome vulnerability could increase the risk of pathologic gene programs by DNA damage accumulation and epigenome erosion over time.
[0318] Previously developed bulk sequencing methods revealed that DNA damage are not randomly distributed and open chromatin regions are prone to be oxidized. Despite the different genomes were analyzed, the results presented herein agreed well with these observations. However, analyzing DNA damage from bulk level only reveals the most conserved hotspots across different cell types and states within the population analyzed. Indeed, chromatin programs underlying most biological processes (such as the initiation and progression of age-related pathology) are highly dynamic, necessitating the measurements of DNA damage at single-cell resolution. Based on whole-genome amplification, linear copy and split-based whole-genome amplification can capture single-nucleotide variants from individual neurons. However, the knowledge on DNA damage alone could not reveal cell identity nor identify the loci associated with specific cellular state alterations. Paired- Damage- seq links DNA damage with transcriptome, enables investigation of the influences of DNA damage on gene programs in heterogeneous samples.
[0319] Recent single-cell studies of human Alzheimer’ s disease revealed the links between DNA damage and epigenome erosion during disease progression. These studies estimated damage burden by measuring chimeric gene fusions arose from DSBs. However, DSBs per se is the one of the most harmful types of DNA damage, it is challenging to distinguish the consequences associated with the specific DSB-related molecular pathway from the decayed regulatory programs. Persistently generating DSBs in mice can facilitate aging phenotypes and can be reversed by epigenome reprogramming. This supports the hypothesis that the ‘relocalization of chromatin modifiers’ drives epigenome erosion while expecting the rates of epigenetic memory loss to be uniformly across the genome. Whether or not epigenetic erosion initiates in pathology-associated ‘pioneering’ loci remains unclear; the data presented herein revealed the DNA damage hotspots exhibit cell-type- specific patterns. With Dictionary Learning, Paired-Damage-seq datasets can be integrated to additional modalities for dissection of relationships and cumulative impacts across different genomic layers.
[0320] The current Paired-Damage-seq is based on combinatorial barcoding and could be readily adapted to droplet-based systems for fast and accessible DNA damage analyses. While the focus was on the prevalent oxidative DNA damage, apyrimidinic sites and SSBs by in situ ‘BER’, this could be extended to other DNA damage types by adopting different labeling strategies described herein or known in the art. Moreover, the observed DNA damage hotspots were combined results from DNA damage formation and repair. Similar to coanalyzing DNA methylation formation and reversal, measuring DNA damage and repair jointly with transcriptome is expected to further facilitate dissecting the dynamics of genome and / or epigenome erosion and the investigation of their functional impacts on cells’ molecular programs in diseases and during aging.
[0321] Tables 1A-1F. Oligonucleotide sequences. The tables below include the primers, oligonucleotides, and DNA barcodes sequences used in this study. Any of these sequences can be used in any of the compositions and methods described herein, as appropriate.
[0322] Table 1A. Paired-Damage-seq Primer Sequences
[0323] Table IB. DNA barcode sequences (Barcode#01 or BC#01)
[0324] Table 1C. DNA barcode sequences (Barcode#02 or BC#02; Plate 1)
[0325] Table ID. DNA barcode sequences (Barcode#02 or BC#02; Plate 2)
[0326] Table IE. DNA barcode sequences (Barcode#03 or BC#03; Plate 1)
[0327] Table IF. DNA barcode sequences (Barcode#03 or BC#03; Plate 2)
[0328] References
[0329] 1. Schumacher, B., Pothof, J., Vijg, J. & Hoeijmakers, J.H.J. The central role of DNA damage in the ageing process. Nature 592, 695-703 (2021).
[0330] 2. Consortium, E.P. et al. Expanded encyclopaedias of DNA elements in the human and mouse genomes. Nature 583, 699-710 (2020).
[0331] 3. Sancar, A., Lindsey-Boltz, L.A., Unsal-Kacmaz, K. & Linn, S. Molecular mechanisms of mammalian DNA repair and the DNA damage checkpoints. Annu Rev Biochem 73, 39-85 (2004).
[0332] 4. Tiwari, V. & Wilson, D.M., 3rd DNA Damage and Associated DNA Repair Defects in Disease and Premature Aging. Am J Hum Genet 105, 237-257 (2019).
[0333] 5. Feinberg, A.P. Phenotypic plasticity and the epigenetics of human disease. Nature 447, 433- 440 (2007).
[0334] 6. Dabin, J., Fortuny, A. & Polo, S.E. Epigenome Maintenance in Response to DNA Damage. Mol Cell 62, 712-727 (2016).
[0335] 7. Oberdoerffer, P. et al. SIRT1 redistribution on chromatin promotes genomic stability but alters gene expression during aging. Cell 135, 907-918 (2008).
[0336] 8. Yang, J.H. et al. Loss of epigenetic information as a cause of mammalian aging. Cell 186, 305-326 e327 (2023). Lu, Y.R., Tian, X. & Sinclair, D.A. The Information Theory of Aging. Nat Aging 3, 1486- 1499 (2023). Qian, M.X. et al. Acetylation-mediated proteasomal degradation of core histones during DNA repair and spermatogenesis. Cell 153, 1012-1024 (2013). Kriaucionis, S. & Heintz, N. The nuclear DNA base 5-hydroxymethylcytosine is present in Purkinje neurons and the brain. Science 324, 929-930 (2009). Tahiliani, M. et al. Conversion of 5 -methylcytosine to 5-hydroxymethylcytosine in mammalian DNA by MLL partner TET1. Science 324, 930-935 (2009). He, Y.F. et al. Tet-mediated formation of 5 -carboxylcytosine and its excision by TDG in mammalian DNA. Science 333, 1303-1307 (2011). Ito, S. et al. Tet proteins can convert 5 -methylcytosine to 5 -formylcytosine and 5- carboxylcytosine. Science 333, 1300-1303 (2011). Pfaffeneder, T. et al. The discovery of 5 -formylcytosine in embryonic stem cell DNA. Angew Chem Int Ed Engl 50, 7008-7012 (2011). Maiti, A. & Drohat, A.C. Thymine DNA glycosylase can rapidly excise 5-formylcytosine and 5-carboxylcytosine: potential implications for active demethylation of CpG sites. J Biol Chem 286, 35334-35338 (2011). Wu, W. et al. Neuronal enhancers are hotspots for DNA single-strand break repair. Nature 593, 440-444 (2021). Reid, D.A. et al. Incorporation of a nucleoside analog maps genome repair sites in postmitotic human neurons. Science 372, 91-94 (2021). Luquette, L.J. et al. Single-cell genome sequencing of human neurons identifies somatic point mutation and indel enrichment in regulatory elements. Nat Genet 54, 1564-1571 (2022). Wang, D. et al. Active DNA demethylation promotes cell fate specification and the DNA damage response. Science 378, 983-989 (2022). Ding, Y., Fleming, A.M. & Burrows, C.J. Sequencing the Mouse Genome for the Oxidatively Modified Base 8-Oxo-7,8-dihydroguanine by OG-Seq. J Am Chem Soc 139, 2569-2572 (2017). Wu, J., McKeague, M. & Sturla, S.J. Nucleotide-Resolution Genome- Wide Mapping of Oxidative DNA Damage by Click-Code-Seq. J Am Chem Soc 140, 9783-9787 (2018). Poetsch, A.R., Boulton, S.J. & Luscombe, N.M. Genomic landscape of oxidative DNA damage and repair reveals regioselective protection from mutagenesis. Genome Biol 19, 215 (2018). Amente, S. et al. Genome-wide mapping of 8-oxo-7,8-dihydro-2'-deoxyguanosine reveals accumulation of oxidatively-generated damage at DNA replication origins within transcribed long genes of mammalian cells. Nucleic Acids Res 47, 221-236 (2019). Liu, Z.J., Martinez Cuesta, S., van Delft, P. & Balasubramanian, S. Sequencing abasic sites in DNA at single-nucleotide resolution. Nat Chem 11, 629-637 (2019). Fang, Y. & Zou, P. Genome-Wide Mapping of Oxidative DNA Damage via Engineering of 8- Oxoguanine DNA Glycosylase. Biochemistry 59, 85-89 (2020). Cao, B. et al. Nick-seq for single-nucleotide resolution genomic maps of DNA modifications and damage. Nucleic Acids Res 48, 6715-6725 (2020). Gorini, F. et al. The genomic landscape of 8-oxodG reveals enrichment at specific inherently fragile promoters. Nucleic Acids Res 48, 4309-4324 (2020). An, J. et al. Genome-wide analysis of 8-oxo-7,8-dihydro-2'-deoxyguanosine at single- nucleotide resolution unveils reduced occurrence of oxidative damage at G-quadruplex sites. Nucleic Acids Res 49, 12252-12267 (2021). Xiao, S., Fleming, A.M. & Burrows, C.J. Sequencing for oxidative DNA damage at single- nucleotide resolution with click-code-seq v2.0. Chem Commun (Camb) 59, 8997-9000 (2023). Tang, F. et al. mRNA-Seq whole-transcriptome analysis of a single cell. Nat Methods 6, 377- 382 (2009). Cusanovich, D.A. et al. Multiplex single cell profiling of chromatin accessibility by combinatorial cellular indexing. Science 348, 910-914 (2015). Jin, W. et al. Genome-wide detection of DNase I hypersensitive sites in single cells and FFPE tissue samples. Nature 528, 142-146 (2015). Buenrostro, J.D. et al. Single-cell chromatin accessibility reveals principles of regulatory variation. Nature 523, 486-490 (2015). Rotem, A. et al. Single-cell ChlP-seq reveals cell subpopulations defined by chromatin state. Nat Biotechnol 33, 1165-1172 (2015). Hainer, S.J., Boskovic, A., McCannell, K.N., Rando, O.J. & Fazzio, T.G. Profiling of Pluripotency Factors in Single Cells and Early Embryos. Cell 177, 1319-1329 el311 (2019). Harada, A. et al. A chromatin integration labelling method enables epigenomic profiling with lower input. Nat Cell Biol 21, 287-296 (2019). Kaya-Okur, H.S. et al. CUT&Tag for efficient epigenomic profiling of small samples and single cells. Nat Commun 10, 1930 (2019). Carter, B. et al. Mapping histone modifications in low cell number and single cells using antibody-guided chromatin tagmentation (ACT-seq). Nat Commun 10, 3747 (2019). Ku, W.L. et al. Single-cell chromatin immunocleavage sequencing (scChIC-seq) to profile histone modification. Nat Methods 16, 323-325 (2019). Wang, Q. et al. CoBATCH for High-Throughput Single-Cell Epigenomic Profiling. Mol Cell 76, 206-216 e207 (2019). Ai, S. et al. Profiling chromatin states using single-cell itChlP-seq. Nat Cell Biol 21, 1164- 1172 (2019). Guo, H. et al. Single-cell methylome landscapes of mouse embryonic stem cells and early embryos analyzed using reduced representation bisulfite sequencing. Genome Res 23, 2126- 2135 (2013). Mooijman, D., Dey, S.S., Boisset, J.C., Crosetto, N. & van Oudenaarden, A. Single-cell 5hmC sequencing reveals chromosome-wide cell-to-cell variability and enables lineage reconstruction. Nat Biotechnol 34, 852-856 (2016). Zhu, C. et al. Single-Cell 5-Formylcytosine Landscapes of Mammalian Early Embryos and ESCs at Single-Base Resolution. Cell Stem Cell 20, 720-731 e725 (2017). Wu, X., Inoue, A., Suzuki, T. & Zhang, Y. Simultaneous mapping of active DNA demethylation and sister chromatid exchange in single cells. Genes Dev 31, 511-523 (2017). Mulqueen, R.M. et al. Highly scalable generation of DNA methylation profiles in single cells. Nat Biotechnol 36, 428-431 (2018). Nagano, T. et al. Single-cell Hi-C reveals cell-to-cell variability in chromosome structure. Nature 502, 59-64 (2013). Mulqueen, R.M. et al. High-content single-cell combinatorial indexing. Nat Biotechnol 39, 1574-1580 (2021). Stoeckius, M. et al. Simultaneous epitope and transcriptome measurement in single cells. Nat Methods 14, 865-868 (2017). Peterson, V.M. et al. Multiplexed quantification of proteins and transcripts in single cells. Nat Biotechnol 35, 936-939 (2017). Cao, J. et al. Joint profiling of chromatin accessibility and gene expression in thousands of single cells. Science 361, 1380-1385 (2018). Zhu, C. et al. Joint profiling of histone modifications and transcriptome in single cells from mouse brain. Nat Methods 18, 283-292 (2021). Swanson, E. et al. Simultaneous trimodal single-cell measurement of transcripts, epitopes, and chromatin accessibility using TEA-seq. Elife 10 (2021). Mimitou, E.P. et al. Scalable, multimodal profiling of chromatin accessibility, gene expression and protein levels in single cells. Nat Biotechnol 39, 1246-1258 (2021). Luo, C. et al. Single nucleus multi-omics identifies human cortical cell regulatory genome diversity. Cell Genom 2 (2022). Wen, X. et al. Single-cell multiplex chromatin and RNA interactions in ageing human brain. Nature 628, 648-656 (2024). Wu, H. et al. Simultaneous single-cell three-dimensional genome and gene expression profiling uncovers dynamic enhancer connectivity underlying olfactory receptor choice. Nat Methods (2024). Zhou, T. et al. GAGE-seq concurrently profiles multiscale 3D genome organization and gene expression in single cells. Nat Genet (2024). Fabyanic, E.B. et al. Joint single-cell profiling resolves 5mC and 5hmC and reveals their distinct gene regulatory effects. Nat Biotechnol (2023). Bai, D. et al. Simultaneous single-cell analysis of 5mC and 5hmC with SIMPLE-seq. Nat Biotechnol (2024). Tubbs, A. & Nussenzweig, A. Endogenous DNA Damage as a Source of Genomic Instability in Cancer. Cell 168, 644-656 (2017). Riedl, J., Fleming, A.M. & Burrows, C.J. Sequencing of DNA Lesions Facilitated by Site- Specific Excision via Base Excision Repair DNA Glycosylases Yielding Ligatable Gaps. J Am Chem Soc 138, 491-494 (2016). Shu, X. et al. Genome-wide mapping reveals that deoxyuridine is enriched in the human centromeric DNA. Nat Chem Biol 14, 680-687 (2018). Rosenberg, A.B. et al. Single-cell profiling of the developing mouse brain and spinal cord with split-pool barcoding. Science 360, 176-182 (2018). Zhu, C. et al. An ultra high-throughput method for single-cell joint analysis of open chromatin and transcriptome. Nat Struct Mol Biol 26, 1063-1070 (2019). Xie, Y. et al. Droplet-based single-cell joint profiling of histone modifications and transcriptomes. Nat Struct Mol Biol 30, 1428-1433 (2023). Xiong, H., Luo, Y., Wang, Q., Yu, X. & He, A. Single-cell joint detection of chromatin occupancy and transcriptome enables higher-dimensional epigenomic reconstructions. Nat Methods 18, 652-660 (2021). Schuster-Bockler, B. & Lehner, B. Chromatin organization is a major influence on regional mutation rates in human cancer cells. Nature 488, 504-507 (2012). Fousteri, M. & Mullenders, L.H. Transcription-coupled nucleotide excision repair in mammalian cells: molecular mechanisms and biological effects. Cell Res 18, 73-84 (2008). Poetsch, A.R. The genomics of oxidative DNA damage, repair, and resulting mutagenesis. Comput Struct Biotechnol J 18, 207-219 (2020). Roadmap Epigenomics, C. et al. Integrative analysis of 111 reference human epigenomes. Nature 518, 317-330 (2015). Milano, L., Gautam, A. & Caldecott, K.W. DNA damage and transcription stress. Mol Cell 84, 70-79 (2024). Zhang, Y. et al. Model-based analysis of ChlP-Seq (MACS). Genome Biol 9, R137 (2008). Khristich, A.N. & Mirkin, S.M. On the wrong DNA track: Molecular mechanisms of repeat- mediated genome instability. J Biol Chem 295, 4134-4170 (2020). Fleming, A.M. & Burrows, C.J. Oxidative stress-mediated epigenetic regulation by G- quadruplexes. NAR Cancer 3, zcab038 (2021). Fleming, A.M., Ding, Y. & Burrows, C.J. Oxidative DNA damage is epigenetic by regulating gene transcription via base excision repair. Proc Natl Acad Sci U S A 114, 2604-2609 (2017). Fleming, A.M., Zhu, J., Ding, Y. & Burrows, C.J. 8-Oxo-7,8-dihydroguanine in the Context of a Gene Promoter G-Quadruplex Is an On-Off Switch for Transcription. ACS Chem Biol 12, 2417-2426 (2017). Fleming, A.M. & Burrows, C.J. Interplay of Guanine Oxidation and G-Quadruplex Folding in Gene Promoters. J Am Chem Soc 142, 1115-1136 (2020). Buenrostro, J.D., Giresi, P.G., Zaba, E.C., Chang, H.Y. & Greenleaf, W.J. Transposition of native chromatin for fast and sensitive epigenomic profiling of open chromatin, DNA-binding proteins and nucleosome position. Nat Methods 10, 1213-1218 (2013). Eiu, Y. et al. Multi-omic measurements of heterogeneity in HeEa cells across laboratories. Nat Biotechnol 37, 314-322 (2019). Persad, S. et al. SEACells infers transcriptional and epigenomic cellular states from singlecell genomics data. Nat Biotechnol 41, 1746-1757 (2023). Hao, Y. et al. Integrated analysis of multimodal single-cell data. Cell 184, 3573-3587 e3529 (2021). Yao, Z. et al. A transcriptomic and epigenomic cell atlas of the mouse primary motor cortex. Nature 598, 103-110 (2021). Li, P.W., Li, J., Timmerman, S.L., Krushel, L.A. & Martin, S.L. The dicistronic RNA from the mouse LINE-1 retrotransposon contains an internal ribosome entry site upstream of each ORF: implications for retrotransposition. Nucleic Acids Res 34, 853-864 (2006). Zhang, Y. et al. Single-cell epigenome analysis reveals age-associated decay of heterochromatin domains in excitatory neurons in the mouse brain. Cell Res 32, 1008-1021 (2022). Zeisel, A. et al. Molecular Architecture of the Mouse Nervous System. Cell 174, 999-1014 el022 (2018). Zhang, K., Zemke, N.R., Armand, E.J. & Ren, B. A fast, scalable and versatile tool for analysis of single-cell omics data. Nat Methods 21, 217-227 (2024). McLean, C.Y. et al. GREAT improves functional interpretation of cis-regulatory regions. Nat Biotechnol 28, 495-501 (2010). Li, Y.E. et al. An atlas of gene regulatory elements in adult mouse cerebrum. Nature 598, 129-136 (2021). Erdo, S.L. & Wolff, J.R. gamma-Aminobutyric acid outside the mammalian brain. J Neurochem 54, 363-372 (1990). Heins, N. et al. Glial cells generate neurons: the role of the transcription factor Pax6. Nat Neurosci 5, 308-315 (2002). Lee, B.T. et al. The UCSC Genome Browser database: 2022 update. Nucleic Acids Res 50, D1115-D1122 (2022). Bulik-Sullivan, B.K. et al. LD Score regression distinguishes confounding from polygenicity in genome-wide association studies. Nat Genet 47, 291-295 (2015). Szebeni, A. et al. Elevated DNA Oxidation and DNA Repair Enzyme Expression in Brain White Matter in Major Depressive Disorder. Int J Neuropsychopharmacol 20, 363-373 (2017). Chou, V. et al. INPP5D regulates inflammasome activation in human microglia. Nat Commun 14, 7552 (2023). Bellenguez, C. et al. New insights into the genetic etiology of Alzheimer's disease and related dementias. Nat Genet 54, 412-436 (2022). Mingard, C., Wu, J., McKeague, M. & Sturla, S.J. Next-generation DNA damage sequencing. Chem Soc Rev 49, 7354-7377 (2020). Amente, S. et al. Genome-wide mapping of genomic DNA damage: methods and implications. Cell Mol Life Sci 78, 6745-6762 (2021). Zhu, Q., Niu, Y., Gundry, M. & Zong, C. Single-cell damagenome profiling unveils vulnerable genes and functional pathways in human genome toward DNA damage. Sci Adv 7 (2021). Dileep, V. et al. Neuronal DNA double-strand breaks lead to genome structural variations and 3D genome disruption in neurodegeneration. Cell 186, 4404-4421 e4420 (2023). Xiong, X. et al. Epigenomic dissection of Alzheimer's disease pinpoints causal variants and reveals epigenome erosion. Cell 186, 4422-4437 e4421 (2023). Hao, Y. et al. Dictionary learning for integrative, multimodal and scalable single-cell analysis. Nat Biotechnol 42, 293-304 (2024). Langmead, B., Trapnell, C., Pop, M. & Salzberg, S.L. Ultrafast and memory-efficient alignment of short DNA sequences to the human genome. Genome Biol 10, R25 (2009). Dobin, A. et al. STAR: ultrafast universal RNA-seq aligner. Bioinformatics 29, 15-21 (2013). Langmead, B. & Salzberg, S.L. Fast gapped-read alignment with Bowtie 2. Nat Methods 9, 357-359 (2012). Hafemeister, C. & Satija, R. Normalization and variance stabilization of single-cell RNA-seq data using regularized negative binomial regression. Genome Biol 20, 296 (2019). Leland Mclnnes, J.H., Nathaniel Saul, Lukas GroBberger UMAP: Uniform Manifold Approximation and Projection. Journal of Open Source Software 3, 861 (2018). Heinz, S. et al. Simple combinations of lineage-determining transcription factors prime cis- regulatory elements required for macrophage and B cell identities. Mol Cell 38, 576-589 (2010). Howard, D.M. et al. Genome-wide meta-analysis of depression identifies 102 independent variants and highlights the importance of the prefrontal brain regions. Nat Neurosci 22, 343- 352 (2019). Duncan, L.E. et al. Largest GWAS of PTSD (N=20 070) yields genetic overlap with schizophrenia and sex differences in heritability. Mol Psychiatry 23, 666-673 (2018). Luciano, M. et al. Association analysis in over 329,000 individuals identifies 116 independent variants influencing neuroticism. Nat Genet 50, 6-11 (2018). Astle, W.J. et al. The Allelic Landscape of Human Blood Cell Trait Variation and Links to Common Complex Disease. Cell 167, 1415-1429 el419 (2016). van Rheenen, W. et al. Genome-wide association analyses identify new risk variants and the genetic architecture of amyotrophic lateral sclerosis. Nat Genet 48, 1043-1048 (2016). Watson, H.J. et al. Genome-wide association study identifies eight risk loci and implicates metabo-psychiatric origins for anorexia nervosa. Nat Genet 51, 1207-1214 (2019). Jansen, I.E. et al. Genome-wide meta-analysis identifies new loci and functional pathways influencing Alzheimer's disease risk. Nat Genet 51, 404-413 (2019). Grove, J. et al. Identification of common genetic risk variants for autism spectrum disorder. Nat Genet 51, 431-444 (2019). Fernandez de la Cruz, L. et al. Suicide in obsessive-compulsive disorder: a population-based study of 36 788 Swedish patients. Mol Psychiatry 22, 1626-1632 (2017). Paternoster, L. et al. Multi-ancestry genome-wide association study of 21,000 cases and 95,000 controls identifies new risk loci for atopic dermatitis. Nat Genet 47, 1449-1456 (2015). Michailidou, K. et al. Association analysis identifies 65 new breast cancer risk loci. Nature 551, 92-94 (2017). Okbay, A. et al. Genome-wide association study identifies 74 loci associated with educational attainment. Nature 533, 539-542 (2016). Malik, R. et al. Multiancestry genome-wide association study of 520,000 subjects identifies 32 loci associated with stroke and stroke subtypes. Nat Genet 50, 524-537 (2018).
[0337] Incorporation by Reference
[0338] All publications, patents, and patent applications mentioned herein are hereby incorporated by reference in their entirety as if each individual publication, patent or patent application was specifically and individually indicated to be incorporated by reference. In case of conflict, the present application, including any definitions herein, will control.
[0339] Also incorporated by reference in their entirety are any polynucleotide and polypeptide sequences which reference an accession number correlating to an entry in a public database, such as those maintained by The Institute for Genomic Research (TIGR) on the world wide web at tigr.org and / or the National Center for Biotechnology Information (NCBI) on the World Wide Web at ncbi.nlm.nih.gov.
[0340] Equivalents
[0341] Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments of the invention described herein. Such equivalents are intended to be encompassed by the following claims.
Claims
What is claimed is:
1. A method of determining DNA damage sites in the genome and obtaining gene expression information of a cell or a nucleus, the method comprising: a) obtaining one or more nuclei, b) labeling the DNA damage; c) performing in situ reverse transcription of RNA in the one or more nuclei to generate cDNA; d) barcoding the DNA damage to generate genomic DNA fragments comprising a barcode, and barcoding the cDNA to generate cDNA comprising a barcode, optionally wherein the barcode for DNA damage is identical to the barcode for cDNA; e) preparing a DNA library from the barcoded genomic DNA fragments, and preparing an RNA library from the barcoded cDNA; f) sequencing the DNA library and the RNA library (e.g., via next generation sequencing), thereby determining the DNA damage sites in the genome and obtaining the gene expression information of the cell or the nucleus.
2. The method of claim 1, further comprising, after obtaining one or more nuclei:(a) fixing the one or more nuclei (e.g., by contacting with formaldehyde or by being exposed to UV light);(b) permeabilizing the one or more nuclei (e.g., by contacting with a detergent, e.g., Triton X-100, NP-40, digitonin, etc.); and / or(c) disrupting or depleting nucleosomes from the one or more nuclei (e.g., by contacting with detergent, e.g., Sodium Dodecyl Sulfate, Lithium Diiodo salicylate).
3. The method of claim 1 or 2, wherein the DNA damage is labeled by tagmentation, optionally wherein the tagmentation comprises a transposase or a variant thereof (e.g., Tn5), optionally wherein the tagmentation is a targeted tagmentation that comprises a transposase, which is a fusion protein or a conjugate further comprising a protein that recognizes the DNA damage.
4. The method of claim 3, wherein the tagmentation comprises incorporating an oligonucleotide comprising a barcode (e.g., barcode 1 or BC1) and, optionally, a first restriction site, into the DNA damage.
5. The method of claim 3 or 4, wherein the tagmentation comprises: a) incorporating a labeled nucleotide to the DNA damage in the one or more nuclei (e.g., using a DNA polymerase, e.g., biotin-labeled dUTP); and b) labeling the DNA damage by incubating the one or more nuclei with (i) a transposase, which further comprises a protein that targets the transposase to the labeled nucleotide (e.g., an anti-biotin antibody fused or conjugated to a transposase), and (ii) an oligonucleotide, which optionally comprises a barcode and / or a first restriction site.
6. The method of claim 3 or 4, wherein the tagmentation comprises: a) incorporating a labeled nucleotide to the DNA damage in the one or more nuclei (e.g., using a DNA polymerase, e.g., biotin-labeled dUTP); b) contacting the one or more nuclei with a protein that binds to the labeled nucleotide at the DNA damage (e.g., anti-biotin antibody); and c) labeling the DNA damage by incubating the one or more nuclei with (i) a transposase (e.g., Tn5), which further comprises a protein (e.g., fused to or conjugated to Protein A) that binds to the protein in b), and (ii) an oligonucleotide, which optionally comprises a barcode and / or a first restriction site.
7. The method of any one of claims 3-6, wherein the tagmentation comprises labeling the DNA damage in two or more aliquots, each with an aliquot- specific oligonucleotide, optionally comprising: a) dividing the one or more nuclei into (n) aliquots (e.g., (n) = 1, 2... 12 or more); and b) labeling the DNA damage in each aliquot with an aliquot- specific oligonucleotide, which optionally comprises a barcode (e.g., barcode 1 or BC1) and, optionally, a first restriction site, by incubating each aliquot with (i) a transposase, or a fusion protein or a conjugate thereof, and (ii) an aliquot- specific oligonucleotide.
8. The method of any preceding claim, wherein the in situ reverse transcription to generate cDNA further comprises incorporating an oligonucleotide, which comprises a barcode (e.g., barcode T or BC1') and, optionally, a second restriction site.
9. The method of any preceding claim, wherein the barcoding the DNA damage and cDNA comprises a single-cell barcoding.
10. The method of any preceding claim, wherein the barcoding the DNA damage or the cDNA comprises a ligation-based combinatorial barcoding, which optionally comprises: a) pooling all aliquots of nuclei; b) dividing the pooled nuclei into (n) aliquots (e.g., (n) = 1, 2...96... 192...288...384 or more); c) ligating the DNA damage or the cDNA with an aliquot- specific barcode (e.g., BC2); and / or d) optionally repeating steps a) to c) at least 1, 2, or 3 times with a different aliquotspecific barcode (e.g., BC2, BC3, BC4, etc.), optionally whereby the DNA damage and the cDNA from the same cell comprise the same combination of the aliquot- specific barcodes.
11. The method of any preceding claim, wherein the preparing DNA library comprises: a) lysing the one or more nuclei; b) adding a polynucleotide tail to the genomic DNA fragments (e.g., encompassing the DNA damage sites, e.g., by a terminal deoxynucleotidyl transferase (TdT)); c) amplifying the genomic DNA fragments using primers, one of which comprises (i) a sequence complementary to the polynucleotide tail, or a portion thereof, and / or (ii) a third restriction site; d) digesting the amplified genomic DNA fragments with restriction enzymes that cut the second restriction site (e.g., eliminating contaminating cDNAs) and the third restriction site; and / or e) generating the genomic DNA fragments comprising a sequencing adaptor, optionally (i) by ligating the sequencing adaptor to the genomic DNA fragments, or (ii) by a tagmentation (e.g., incubating the genomic DNA fragments with a transposase (e.g., Tn5) and a sequencing adaptor).
12. The method of any preceding claim, wherein the preparing RNA library comprises: a) lysing the one or more nuclei; b) adding a polynucleotide tail to the cDNA (e.g., by a terminal deoxynucleotidyl transferase (TdT)); c) amplifying the cDNA using primers, one of which comprises (i) a sequence complementary to the polynucleotide tail, or a portion thereof, and / or (ii) a third restriction site; d) digesting the amplified cDNA with a restriction enzyme that cuts the first restriction site (e.g., eliminating contaminating genomic DNA fragments); and e) generating the cDNA comprising a sequencing adaptor, optionally (i) by ligating the sequencing adaptor to the cDNA, or (ii) by a tagmentation (e.g., incubating the cDNA with a transposase (e.g., Tn5) and a sequencing adaptor).
13. The method of any preceding claim, wherein the preparing DNA library and RNA library comprises: a) lysing the one or more nuclei to release the nucleic acid comprising genomic DNA fragments (e.g., encompassing DNA damage) and cDNA (e.g., generated from mRNA); b) adding a polynucleotide tail to the nucleic acid (e.g., by a terminal deoxynucleotidyl transferase (TdT)); c) amplifying the nucleic acid using primers, one of which comprises (i) a sequence complementary to the polynucleotide tail, or a portion thereof, and / or (ii) a third restriction site; d) dividing the nucleic acid into a DNA library and an RNA library; e) for the DNA library: i) enriching for genomic DNA fragments by digesting the nucleic acid with restriction enzymes that cut the second restriction site (e.g., eliminating cDNAs) and the third restriction site; and ii) generating the genomic DNA fragments comprising a sequencing adaptor by ligating the sequencing adaptor to the genomic DNA fragments; f) for the RNA library: i) enriching for cDNA by digesting the nucleic acid with a restriction enzyme that cuts the first restriction site (e.g., eliminating genomic DNA fragments); ande) generating the cDNA comprising a sequencing adaptor by a tagmentation (e.g., incubating the cDNA with a transposase (e.g., Tn5) and a sequencing adaptor).
14. The method of any preceding claim, wherein the DNA damage is selected from an oxidative DNA damage, double- stranded DNA breaks, DNA intrastrand crosslinking, and ribonucleotide in DNA, optionally wherein the DNA damage is oxidative DNA damage, and the oxidative DNA damage comprises 8-oxoguanine (8-oxoG), repair intermediates apurinic / apyrimidinic (AP) sites, and / or single-stranded DNA breaks.
15. The method of any preceding claim, wherein the DNA damage comprises an oxidized base, and the oxidized base comprises an oxidized purine and / or an oxidized pyrimidine.
16. The method of any preceding claim, wherein the DNA damage comprises an oxidized base and labeling the DNA damage comprises:1) contacting the one or more nuclei with an enzyme that removes an oxidized base; and / or2) contacting the one or more nuclei with an apurinic / apyrimidinic (AP) endonuclease, thereby generating a single-strand DNA break (SSB) at the DNA damage site.
17. The method of claim 16, wherein: a) the enzyme that removes an oxidized base selected from a DNA repair protein formamidopyrimidine [fapy]-DNA glycosylase (Fpg) or a glycosylase, optionally wherein the glycosylase is selected from 8-Oxoguanine DNA glycosylase (OGGI), endonuclease III, MutY homolog DNA glycosylase (MYH), NTHL1, NEIL1, NEIL2, and NEIL3; and / or b) the AP endonuclease selected from endonuclease IV, APE1, APE2, APN1, and APN2.
18. The method of any preceding claim, wherein the DNA damage comprises doublestrand breaks (DSBs), and labeling the DNA damage comprises: a) contacting the one or more nuclei with Klenow fragment DNA polymerase and T4 Polynucleotide kinase;b) incubating the one or more nuclei with a double- stranded DNA adaptor comprising a label (e.g., biotin, labeled nucleotide) and a DNA ligase (e.g., T4 DNA ligase) to label the DSB site with the adaptor; and c) performing a tagmentation, optionally a targeted tagmentation (e.g., using an antilabel antibody (e.g., anti-biotin antibody) and protein A-Tn5).
19. The method of any preceding claim, wherein the labeling DNA damage comprises incorporating a labeled nucleotide to the DNA damage in the one or more nuclei by contacting the one or more nuclei with a DNA polymerase, a Bst polymerase, or a terminal deoxynucleotidyltransferase.
20. The method of any preceding claim, wherein a) the labeled nucleotide is selected from a labeled dUTP, a labeled dCTP, a labeled dATP, a labeled dTTP, and a labeled dGTP, optionally a labeled dUTP; and / or b) the labeled nucleotide comprises a biotin or a fluorophore.
21. The method of any preceding claim, wherein a) the labeled nucleotide is a biotin-labeled dUTP; and / or b) the protein that binds the labeled nucleotide is an antibody or streptavidin.
22. The method of any preceding claim, wherein the polynucleotide tail comprises an array of a single type of nucleotide, optionally an array of dCTP.
23. The method of any one of claims 11-22, wherein the polynucleotide tail is added to the genomic DNA fragments and / or cDNA: a) by a terminal deoxynucleotidyltransferase (TdT); b) by a DNA ligase; c) by a DNA polymerase; or d) by a DNA or an RNA oligonucleotide comprising a reactive chemical group (e.g., an azide group or an alkyne group, e.g., suitable for click chemistry), which attaches to the 3’ end of the genomic DNA fragments and / or cDNA.
24. The method of any preceding claim, wherein the genomic DNA fragments and / or cDNA are amplified using a polymerase chain reaction.- I l l -25. The method of claim 10, wherein the ligation-based combinatorial barcoding is selected from SPLiT-seq, Sci-RNA-seq3, Paired-seq / Tag, and SHARE-seq.
26. The method of any one of claims 11-23, wherein the third restriction site is recognized by a type IIS endonuclease, optionally selected from FokI, Acul, AsuHPI, Bbvl, Bpml, BpuEI, BseMII, BseRI, BseXI, Bsgl, BslFI, BsmFI, BsPCNI, BstVlI, BtgZI, Ecil, Eco57I, FaqI, Gsul, HphI, Mmel, NmeAIII, Schl, TaqII, TspDTI, and TspGWI.
27. The method of any one of claims 10-13, wherein the ligase is a T3, T4, or T7 DNA ligase.
28. The method of any preceding claim, further comprising obtaining information of the genome- wide histone modifications and / or chromatin state (e.g., location of histone variants, chromatin-associated factors (e.g., transcription factors, chromatin remodelers, etc.) in a single nucleus.
29. The method of claim 26, wherein the obtaining information of the genome-wide histone modifications and / or chromatin state comprises performing ChlP-chip, ChlP-seq, and / or Paired-Tag.
30. The method of claim 26 or 27, wherein the histone modification is selected from H3K4mel, H3K4me2, H3K4me3, H3K27ac, H3K27me2, H3K27me3, H3K9me3, H3K9me2, H3K9Ac, H3K14Ac, H3K18Ac, H3K23Ac, H3K27me3, H4K16ac, H4K20mel, H4K20me2, H4K20me3, H4K27me3, HlK26mel, H3S10Ph, and H2A.XS139.
31. The method of any preceding claim, wherein the one or more nuclei are of a fungus, a bacterium, or a mammalian cell (e.g., a mouse cell, a rat cell, a monkey cell, a human cell).
32. The method of any preceding claim, wherein the one or more nuclei are of a cell type selected from a neuronal cell (e.g., excitatory neurons (ExN), inhibitory neurons (InN), and non-neuron cell types, including oligodendrocyte precursor cells (OPC), oligodendrocytes (ODC), astrocytes (AST), microglia cells (MiG), endothelial cells (Endo), and vascular leptomeningeal cells (VEMC)), a bone cell, an endothelial cell, a fibroblast, a blood cell (e.g.,erythrocyte, lymphocyte, platelet), an epithelial cell, a muscle cell, a stem cell, an epidermal cell, adipocyte, and hepatocyte.
33. The method of any preceding claim, wherein the one or more nuclei are of a cancer cell.
34. The method of any preceding claim, wherein the one or more nuclei are of a proliferating cell, a quiescent cell, or a senescent cell.
35. The method of any preceding claim, wherein the one or more nuclei are from a young subject, a middle-aged subject, or an old subject, optionally wherein: a) the subject is a mouse, wherein the young mouse is of the age between 1-9 months (preferably 1-4 months), the middle-aged mouse is of the age between 10-17 months (preferably 10-15 months), and the old mouse is of the age 18 months and higher (preferably 18-24 months); and / or b) the subject is a human, wherein the young human is of the age between the birth to 39 years, the middle-aged human is of the age between 40-59 years, and the old human is of the age 60 years and older.
36. The method of any preceding claim, wherein the one or more nuclei are treated with an agent, or the one or more nuclei are obtained from one or more cells treated with an agent, optionally wherein the agent alters the cellular physiology (e.g., altering the integrity of DNA, RNA, and / or protein; altering histone modification and / or chromatin structure; altering gene expression (e.g., transcription and / or translation); altering cell morphology; altering cellular proliferation; altering the identity and / or activity of the cell, etc.; or any combination of two or more of).
37. A method of determining the effect of an agent on a cell, the method comprising contacting the cell with the agent, and determining DNA damage sites in the genome and obtaining gene expression information of the cell according to the method of any preceding claim.
38. The method of claim 34 or 35, wherein the agent is selected from a toxin, a DNA damage inducing agent (e.g., H2O2, reactive oxygen species (ROS)), a chemotherapeuticagent, a small molecule, an alkylating agent, a hormone, a cell signaling molecule, a cytokine, a chemokine, and any combination of two or more thereof.
39. The method of any one of claims 34-36, wherein the agent induces oxidative DNA damage.