Methods and systems for profiling chromatin architecture
The method of DNA crosslinking, cleavage, ligation, and amplification provides comprehensive chromatin architecture analysis, addressing the limitations of existing techniques by revealing detailed chromatin structures and variations, enhancing disease understanding.
Patent Information
- Application Number
- PCT/US2025/022193
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-28
- Filing Date
- 2025-03-28
- Publication Date
- 2025-10-02
AI Technical Summary
Current methods for analyzing chromatin architecture in heterogeneous tissue samples are inadequate, failing to provide comprehensive insights into gene regulatory programs during development and disease pathogenesis.
A method involving DNA crosslinking, cleavage, ligation, transposition, and amplification of chromosomal fragment DNA sequences, followed by cell profiling through encapsulation, barcoding, and sequencing, to generate high-resolution chromatin structure data.
Enables detailed profiling of chromatin architecture, revealing multi-scale structures correlated with gene activity, identifying chromatin hubs, and detecting copy number variations and ecDNA in tumor samples, thereby enhancing understanding of disease mechanisms.
Smart Images

Figure US2025022193_02102025_PF_FP_ABST
Abstract
Description
METHODS AND SYSTEMS FOR PROFILING CHROMATIN ARCHITECTURERELATED APPLICATION DATA
[0001] This application claims the benefit of priority under 35 U.S.C. § 119(e) of U.S. Patent Application No. 63 / 571,382, filed on March 28, 2024, which is hereby incorporated by reference in its entirety and for all purposes.GOVERNMENT SUPPORT CLAUSE
[0002] This invention was made with government support under HG011585, CA258248, NS080939, and HG011922 awarded by National Institutes of Health. The government has certain rights in the invention.BACKGROUND
[0003] Comprehensive analysis of chromatin architecture is crucial for understanding the gene regulatory programs during development and in disease pathogenesis, yet current methods often inadequately address the unique challenges presented by analysis of heterogeneous tissue samples. Provided herein, inter alia, are methods and systems to address these and other problems in the art.BRIEF SUMMARY
[0004] In an aspect is provided a method of amplifying a chromosomal fragment DNA sequence, the method including: (a) contacting a plurality of cell nuclei with a DNA crosslinking agent wherein each of the plurality of cell nuclei includes chromosomal DNA thereby forming crosslinked chromosomal DNA within each of the plurality of cell nuclei;(b) contacting the crosslinked chromosomal DNA with a DNA endonuclease thereby forming cleaved chromosomal DNA within each of the plurality of cell nuclei, wherein each of the crosslinked chromosomal DNAs includes a plurality of double stranded cleaved sites; (c) contacting the cleaved chromosomal DNA with a ligase thereby forming ligated chromosomal DNA within each of the plurality of cell nuclei, wherein each of the ligated chromosomal DNAs includes at least one ligated site that forms a non-endogenous ligation DNA sequence, wherein the non-endogenous ligation DNA sequence is not present in the endogenous chromosomal DNA within each of the plurality of cell nuclei; (d) contacting the ligated chromosomal DNA with a transposase, a 5' hybridization DNA sequence and a 3’ hybridization DNA sequence thereby forming a plurality of transposed chromosomal DNA sequences within each of the plurality of cell nuclei, wherein each of the transposed chromosomal DNA sequences includes from 5' to 3' a first DNA hybridization sequence, achromosomal fragment DNA sequence, and a second DNA hybridization sequence; (e) separating each of the plurality of cell nuclei into a separate compartment, wherein each compartment contains one cell nucleus and an identifying library nucleic acid sequence, wherein the identifying library nucleic acid sequence includes from 5' to 3' a forward primer sequence, a barcode sequence, and a 5' complementary hybridization sequence; (f) lysing each cell nucleus in each compartment and allowing the first DNA hybridization sequence to hybridize with the 5' complementary hybridization sequence to form a first hybridized template DNA sequence; (g) contacting the first hybridized template DNA with a DNA polymerase under conditions conducive to amplification and producing a first amplified chromosomal DNA including from 5' to 3' the forward primer sequence or complement thereof, the barcode sequence or complement thereof, the 5' complementary hybridization sequence or complement thereof, the chromosomal fragment DNA sequence or complement thereof, and the second DNA hybridization sequence or complement thereof, thereby amplifying the chromosomal fragment DNA sequence.
[0005] In another aspect is provided a method of cell profiling the method including: (a) providing or obtaining a cross-linked chromatin complex of a cell or a cellular nucleus; (b) removing one or more histone proteins from the cross-linked chromatin complex; (c) digesting the cross-linked chromatin complex or portions thereof after (b), thereby generating a digested chromatin; (d) ligating the digested chromatin, thereby generating a ligated chromatin; (e) fragmenting the ligated chromatin, thereby generating a fragmented chromatin; (f) encapsulating the fragmented chromatin in a droplet; (g) barcoding the fragmented chromatin with a barcode, thereby generating a barcoded fragmented chromatin; and (h) sequencing the barcoded fragmented chromatin.BRIEF DESCRIPTION OF THE DRAWINGS
[0006] FIGS. 1A-1H show a representative overview and performance of Droplet Hi-C. FIG. 1A: Schematic of the Droplet Hi-C workflow. FIGS. 1B-1C Comparison of throughput, sample preparation time, and cost among different single-cell Hi-C methods. FIG. ID: UMAP visualization of Droplet Hi-C data from adult mouse cortex. FIG. IE: Genome-wide Spearman’ s correlation coefficients (SCC) between compartment scores from different cell types. FIG. IF: Cell type-specific Droplet Hi-C contacts map from chromosome 1 and compartment score at 100-kb resolution. The color bar represents the imputed contact number. FIG. 1G: Cell type-specific Droplet Hi-C contacts map and boundary probability of example region (chromosome 1: 55.2 - 59.5Mb) at 25-kb resolution. FIG. 1H: Cell type-specific Droplet Hi-C contacts map surrounding gene Satb2 (chrl : 55.5 - 57.2Mb) at 10-kb resolution, along with genome browser view showing transcriptome and histone modifications profiles in the same cell types from the public datasets.
[0007] FIGS. 2A-2J show multi-scale chromatin structures are correlated with cell typespecific gene activity in adult mouse cortex. FIG. 2A: Heatmap showing compartment score of differential compartments among all the cell types. FIG. 2B: Distribution of Pearson correlation coefficients between compartment score and histone modification signals at each 100-kb bin across differential compartments; n = 895 (H3K27ac and H3K27me3). All box plot hinges were drawn from the 25th to 75th percentiles, with the middle line denoting the median and whiskers denoting 2x the interquartile range. FIG. 2C: Comparison of single-cell insulation scores from example cell types surrounding gene Pdgfra. A bulk contacts map at 10-kb resolution is shown above. FIG. 2D: Schematics for correlation analysis between boundary probability and nearby gene expression level. FIG. 2E Comparison of Pearson correlation coefficients between gene expression level and boundary probability at variable domain boundaries. Genes are classified as constant (n = 521), housekeeping (n = 512) and variable (n = 448) genes. P values are from Wilcoxon signed-rank test. All box plot hinges were drawn from the 25th to 75th percentiles, with the middle line denoting the median and whiskers denoting 2x the interquartile range. FIG. 2F: Comparison of histone modification signal enrichment at loop anchors across different cell types; n =265,277 (all groups); All box plot hinges were drawn from the 25th to 75th percentiles, with the middle line denoting the median and whiskers denoting 2x the interquartile range. P values, Wilcoxon signed-rank test. FIG. 2G: Top enriched gene ontology terms for genes at loop anchors in selected cell types VIPGA and OGC. P values and fold enrichment were calculated by binomial test. Benjamini-Hochberg FDRs were then calculated to select overrepresented GO terms. FIG. 2H: Diagram of method for identifying chromatin hubs involving multi-way interactions from Droplet Hi-C data. FIG. 21: Box plots showing the overlap between chromatin hubs and super-enhancers in matched (n = 16) and unmatched (n = 240) mouse cortical cell types. Odd ratio is calculated by Fisher exact test. P values were calculated by one-sided Wilcoxon signed-rank test. FIG 2J: Box plots showing the overlap between chromatin hubs and cell type-specific marker genes in matched (n = 18) and unmatched (n = 306) cortical cell types, similar as FIG. 21 (n = 2 biologically independent experiments). All box plot (FIGS. 2B, 2E, 2F, 21, and 2J) hinges were drawn from the 25th to 75th percentiles, with the middle line denoting the median and whiskers denoting 2x the interquartile range.
[0008] FIGS. 3A-3I show Droplet Hi-C illuminates copy number, structural variations and ecDNA in tumor samples. FIG. 3A: Copy number inferred from pseudo-bulk Droplet Hi-C profiles in COLO320DM and 1217 COLO320HSR, with heatmap showing representative single-cell CNV in each sample. The ecMYC bins in COLO20DM are highlighted in pink. FIG. 3B: Example of a sample-specific SV on chromosome 6, predicted with EagleC.Cartoon illustration explaining the rearranged contacts pattern is shown aside. FIG. 3C: Comparison of genome-wide contact maps and adjusted normalized inter-chromosomal interaction frequency (adjnTIF) between COLO320DM and COLO320HSR. FIG. 4D: Circos plot showing trans-interaction profiles of genomic bin containing MYC in COLO320DM or COLO320HSR. FIGS. 3E-3G: Distribution of single-cell hub index (FIG. 3E), copy number (FIG. 3F), and trans-to-cis contacting bin ratio (FIG. 3G) of ecMYC in COLO320DM and COLO320HSR. n = 1,352 (Hub index, COLO320DM), 1,366 (Hub index, COLO320HSR), 1,426 (Inferred copy number and trans-to-cis contacting bin ratio, COLO320DM), 1,535 (Inferred copy number and trans to cis contacting bin ratio, COLO320HSR); P values are from Wilcoxon signed-rank test. FIG. 3H: Schematics of deep learning-based ecDNA caller. FIG. 31: Dot plot showing genome-wide ecDNA and HSR prediction results from deep learning-based ecDNA caller for COLO320DM and COLO320HSR cells.
[0009] FIGS. 4A-4G show Droplet Hi-C reveals ecDNA dynamics under drug treatment. FIG. 4A: Illustration of the erlotinib treatment procedure for GBM39 cells. FIG. 4B: UMAP embedding visualization and clustering analysis of Droplet Hi-C data from GBM39 before - and after-erlotinib treatment. FIGS. 4C-4D: Pie charts showing the cell proportion of each cluster in GBM39 (FIG. 4C) and GBM39-ER (FIG. 4D). FIG. 4E: Comparison of contact maps and percentages of ecDNA-positive cells among all clusters for ecEGFR, ecMYC and ecMDM2. Genomic coordinates for key oncogenes are displayed aside. FIG. 4F: UMAP visualization of the ecDNA positive cell distribution within each cluster. FIG. 4G: Pseudobulk contacts map at 10-kb resolution showing ecMYC local structure in both GBM39 and GBM39-ER samples. Heatmaps representing representative single-cell inferred CNV in the same genomic range are shown below.
[0010] FIGS. 5A-5O show Droplet Hi-C depicts ecDNA variations in patient samples. FIG. 5A: Schematics of GBM cellular state analysis for Droplet Hi-C data by co-embedding with reference lOx Multiome data. FIG. 5B: UMAP visualization of Droplet Hi-C data on GBM patient sample. FIG. 5C: Representative single-cell inferred CNV on chromosome 7, 10 and 19 in tumor-like and non-tumor populations identified from Droplet Hi-C. FIG. 5D:Line plot showing copy number among single cells in tumor-like and non-tumor populations. Lines with color highlighted the median profiles. The single-cell copy number profile examples are colored in gray (n = 20). FIG. 5E: Genome-wide contact maps from tumor-like and non-tumor populations. Color bar shows the raw contact number. FIG. 5F: The ecDNA prediction results in GBM patient sample. The percentage of cells predicted to contain ecEGFR in tumor- like population is shown in a pie chart. FIG. 5G: An example of tumor-like population-specific SV. Genome browser view showing the Droplet Hi-C read coverage at associated genes displaying below. Predicted breakpoints of SV are highlighted in pink. Gene expression level (RPKM) of associated genes in different populations from lOx Multiome is also shown. FIGS. 5H-5I: Two- dimensional representation of cellular states based on scRNA-seq data from lOx Multiome (FIG. 5H) and Droplet Hi-C data (FIG 51). Each quadrant corresponds to one cellular state. FIG. 5 J: The percentage of ecEGFR-positive cells across four different GBM cellular states, n = 563 (AC-like), 600 (MES-like), 247 (NPC-like), 346 (OPC-like). FIG. 5K: Inferred copy number at 100-kb resolution of regions on ecEGFR among different GBM cellular states. FIGS. 5L-5M: Violin plots showing EGFR ATAC gene score (FIG. 5L) and EGFR gene expression level (FIG. 5M) across four different GBM cellular states from 10X Multiome dataset. Only cells with GBM cellular states score pass cutoff (± 0.5) are shown, n = 244 (AC-like), 410 (MES-like), 125 (NPC-like), 263 (OPC-like). FIG. 5N: Heatmap showing Spearman’ s correlation coefficients (SCC) between GBM cellular states score and gene expression level for all the genes located on ecEGFR. FIG. 50: Two-dimensional representation of GBM cellular states, colored by ecDNA genes expression levels.
[0011] FIGS. 6A-6I show Joint profiling of single-cell Hi-C and transcriptome. FIG. 6A: Schematic of the Paired Hi-C molecule barcoding step with 10X Multiome kit. FIG. 6B: UMAP visualization of mouse cortex single-cell transcriptome from Paired Hi-C. FIG. 6C: Comparison of pseudo-bulk contact maps between Droplet Hi-C and Paired Hi-C at the region surrounding gene Erbb4, along with compartment score profiles from Paired Hi-C in representative cell types. Color bar shows the imputed contact number. A violin plot of Erbb4 expression level in representative cell types is also shown. FIG. 6D: Comparison of Spearman’ s correlation coefficients (SCC) between compartment score across cell types in Droplet Hi-C and Paired Hi-C. n =15 (match); 225 (unmatch). FIG. 6E: Heatmap showing marker genes expression levels and compartment scores of corresponding bins among all clusters. FIG. 6F: Single-cell inferred copy number heatmaps of regions harboring ecMYC inGBM39 and GBM39-ER. Single-cell UMI counts heatmaps of representative genes on ecDNA are also shown. FIG. 6G: SCC between gene expression and copy number for ecMYC genes in 5’, 3 ’variable regions and shared regions in GBM39-ER. Illustration of ecMYC variable regions is shown on the left, n = 9 (5’ gene); 10 (shared); 11 (3’ gene). FIG. 6H: Comparison of cell proportions in four GBM cellular states between GBM39 and GBM39-ER samples. FIG. 61: Comparison of expression levels of EGFR and MDM2 among four GBM cellular states, n = 4,307 (AC-like), 6,791 (MES-like), 4,945 (NPC-like), 3,123 (OPC-like).
[0012] FIGS. 7A-7H show a comparison between Droplet Hi-C and in situ Hi-C on cultured cells. FIG. 7A: Scatter plots showing proportion of human and mouse DNA reads in each cell in the Droplet Hi-C species mixing experiment. FIG. 7B: UMAP visualization of Droplet Hi-C data on three human cell lines mixing sample. FIG. 7C: Genome-wide Spearman correlation coefficients (SCC) of compartment score between pseudo-bulk Droplet Hi-C data and bulk in-situ Hi-C datasets. FIG. 7D: Comparison of contact frequency by distance between Droplet Hi-C and in situ Hi-C datasets of Hela S3, GM12878 and K562 cells. FIG. 7E: Comparison of contact maps between Droplet Hi-C and in situ Hi-C for Hela S3, GM12878 and K562 cells on chromosome 11 at 100-kb resolution. FIG. 7F Scatter plot showing genome-wide SCC of compartment score between Droplet Hi-C and in situ Hi-C for Hela S3, GM12878 and K562 cells. FIGS. 7G-7H Comparison of compartment score (FIG. 7G) and insulation score (FIG. 7H) profiles on chromosome 11: 23 - 28 Mb from Droplet Hi- C and in situ Hi-C datasets for Hela S3, GM12878 and K562 cells.
[0013] FIGS. 8A-8L show performance of Droplet Hi-C on adult mouse cortex. FIG. 8A: Comparison of library complexity between two datasets generated with lOx Genomics Chromium Single Cell ATAC kit v 1.1 and v2. FIG. 8B: Distribution of cis-short, cis-long or trans-contacts number per nucleus. FIG. 8C: Comparison of cis-long contact (>lkb) numbers distribution across different cell types. FIG. 8D: Pseudo-bulk and representative single-cell genome- wide contact maps from mouse cortex. Color bar showed the raw contact number. FIGS. 8E-8F: Comparison of contact frequency by distance (FIG. 8E) and contacts ratio (FIG. 8F) on mouse cortex data among Droplet Hi-C, sn-m3C-seq and Dip-C. The cis-long interactions in (FIG. 8F) represent contacts separated more than 1 kb. FIG. 8G: Comparison of multi-scale genome organization on chromosome 1 between Droplet Hi-C and sn-m3C- seq. FIGS. 8H-8L: The compartmentalization strength (FIG. 8H), compartment scores (FIGS.81-8 J) and insulation scores (FIGS. 8K-8L) among Droplet Hi-C, sn-m3C-seq and Dip-C are also shown.
[0014] FIGS. 9A-9E show validation of cell type identities in mouse cortex characterized by Droplet Hi-C. FIG. 9A: UMAP visualization of public mouse cortex single-nucleus RNA- seq data used to co-embed Droplet Hi-C. FIG. 9B: Distribution of single-cell prediction scores for Droplet Hi-C cells, embedded with Droplet Paired-Tag RNA data. FIG. 9C: Comparison of single-cell prediction scores across different cell types. FIG. 9D: Heatmap showing gene associating domain (GAD) scores for known marker genes among all the cell types from Droplet Hi-C data. FIG. 9E: The overlap scores between the joint clusters and the original annotations from Droplet Hi-C and sn-m3C-seq data in adult mouse cortex.
[0015] FIGS. 10A-10B show Droplet Hi-C illuminates CNV, SVs and ecDNA in cancer cells. FIG. 10A: Contact maps of regions containing ecMYC in COLO320DM and COLO320HSR. Inferred copy number from pseudo-bulk profiles as well as representative single cells are shown below. FIG. 10B: Single-cell hub index calculated with normal versus shuffle contact profiles in COLO320DM and COLO320HSR. P values are calculated from paired sample Wilcoxon signed-rank test.
[0016] FIGS. 11A-11D show a flowchart of representative ecDNA caller algorithms. FIG.11 A: Workflow of multivariate logistic regression model-based ecDNA caller used to predict ecDNA identity for each genomic region, and cells containing the candidate ecDNA. FIG.1 IB: ecDNA prediction results from the logistic regression model-based ecDNA caller in COLO320DM and COLO320HSR sample. FIG. 11C: Schematic of deep learning-based ecDNA caller used to predict HSR / ecDNA identity of genomic regions, and cells containing the candidate HSR / ecDNA. FIG. 1 ID: Confusion matrix of prediction results versus original label from deep learning-based ecDNA caller on COLO320DM and COLO320HSR validation datasets.
[0017] FIGS. 12A-12F show dynamics of ecDNA in GBM39 under erlotinib treatment. FIG. 12A: Scatter plot showing inferred copy number on chromosomes containing candidate ecDNA (chromosome 7, 8, 12) across clusters in GBM39 and GBM39-ER. ecDNA- associated regions are highlighted in pink. FIG. 12B: Genome-wide contact maps for GBM39 and GBM39-ER cells. FIG. 12C: Genome-wide ecDNA prediction results from deep learning-based ecDNA caller for GBM39 and GBM39-ER. FIG. 12D: Representative single-cell contact map and percentage of cells containing rare ecDNA (ecChrl8) found by deep learning-based ecDNAcaller. FIG. 12E: Genome-wide contact maps at cluster and representative single-cell levels. FIG 12F: Comparison of ecMYC interaction profiles for ecDNA positive cells within CO from different samples.
[0018] FIGS. 13A-13H show Cell state-specific chromatin structure and transcriptome differences in GBM patient samples. FIG. 13 A: Another example of tumor- like populationspecific SV on chromosome 1 from GBM patient sample. SV are highlighted with black rectangle. FIG. 13B: Comparison of A / B compartment profiles on chromosome 3 between tumor-like cells and non-tumor populations. FIG. 13C: Heatmap showing Spearman correlation coefficients (SCC) of Droplet Hi-C compartment score across different cellular states. FIG. 13D: Illustration of AML and MDS patient treatment procedure. FIG. 13E: Pseudo-bulk genome- wide contact maps for BMMC samples from AML and MDS patient before and after treatment. FIG. 13F: Proportion of cells containing known tumor- associated mutations in BMMC samples before and after treatment. Aggregated genome-wide contact maps from mutant-carrying cells are also shown. Color bar shows the raw contact number. FIGS. 13G-13H: An example of a before-treated sample-specific loop between MYC and its enhancer region (BENC) on ecMYC. Normalized aggregate peak analysis (APA) scores of before- and after-treated samples are shown aside.
[0019] FIG. 14 shows a representative workflow of Paired Hi-C. This detailed reperesentative workflow describes the end-to-end procedure of Paired Hi-C and time cost for each step.
[0020] FIGS. 15A-15F show joint chromatin organization and gene expression analysis using 1370 Paired Hi-C. FIG. 15A: Distribution of UMI, gene number, and contacts per nucleus in Paired Hi-C data from adult mouse cortex. FIG. 15B: UMAP co-embedding of single-nucleus transcriptome profiles from Paired Hi-C experiment and the reference BICCN and SCENIC+ datasets. FIG. 15C: The overlap scores of shared annotations between BICCN single-nucleus RNA-seq dataset and Paired Hi-C single-nucleus RNA-seq dataset. FIG. 15D: Dot plot showing expression level of marker genes in each cell type. FIG. 15E: Box plot showing Spearman correlation coefficients (SCC) of compartment score between sn-m3C-seq and Paired Hi-C. FIG. 15F: Overlap of co-embedding annotations between using Droplet Paired-Tag single-nucleus RNA-seq dataset as reference and Paired Hi-C single- nucleus RNA-seq dataset as reference.
[0021] FIGS. 16A-16C show Paired Hi-C illuminates alterations in the transcriptome associated with variations in chromatin structure. FIG. 16A: Comparison of expression levelfor ecMYC trans-contacting genes between GBM39 and GBM39-ER cells. P values, Wilcoxon signed-rank test. FIG. 16B: Summary of changes in A / B compartments in GBM39 and GBM39-ER sample. FIG. 16C: Expression of genes classified by the underlying A / B compartment transitions in GBM39 and GBM39-ER cells.DETAILED DESCRIPTIONDEFINITIONS
[0022] While various embodiments and aspects of the present invention are shown and described herein, it will be obvious to those skilled in the art that such embodiments and aspects are provided by way of example only. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in practicing the invention.
[0023] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described. All documents, or portions of documents, cited in the application including, without limitation, patents, patent applications, articles, books, manuals, and treatises are hereby expressly incorporated by reference in their entirety for any purpose.
[0024] The abbreviations used herein have their conventional meaning within the chemical and biological arts. The chemical structures and formulae set forth herein are constructed according to the standard rules of chemical valency known in the chemical arts.
[0025] Unless defined otherwise, technical and scientific terms used herein have the same meaning as commonly understood by a person of ordinary skill in the art. See, e.g., Singleton et al., DICTIONARY OF MICROBIOLOGY AND MOLECULAR BIOLOGY 2nd ed„ J. Wiley & Sons (New York, NY 1994); Sambrook et al., MOLECULAR CLONING, A LABORATORY MANUAL, Cold Springs Harbor Press (Cold Springs Harbor, NY 1989). Any methods, devices and materials similar or equivalent to those described herein can be used in the practice of this invention. The following definitions are provided to facilitate understanding of certain terms used frequently herein and are not meant to limit the scope of the present disclosure.
[0026] " Nucleic acid" refers to nucleotides (e.g., deoxyribonucleotides or ribonucleotides) and polymers thereof in either single-, double- or multiple-stranded form, or complements thereof; or nucleosides (e.g., deoxyribonucleosides or ribonucleosides). In embodiments, “nucleic acid” does not include nucleosides. The terms “polynucleotide,” “oligonucleotide,”“oligo” or the like refer, in the usual and customary sense, to a linear sequence of nucleotides. The term “nucleoside” refers, in the usual and customary sense, to a glycosylamine including a nucleobase and a five-carbon sugar (ribose or deoxyribose). Non limiting examples, of nucleosides include cytidine, uridine, adenosine, guanosine, thymidine and inosine. The term “nucleotide” refers, in the usual and customary sense, to a single unit of a polynucleotide, i.e., a monomer. Nucleotides can be ribonucleotides, deoxyribonucleotides, or modified versions thereof. Examples of polynucleotides contemplated herein include single and double stranded DNA, single and double stranded RNA, and hybrid molecules having mixtures of single and double stranded DNA and RNA. Examples of nucleic acid, e.g. polynucleotides contemplated herein include any types of RNA, e.g. mRNA, siRNA, miRNA, and guide RNA and any types of DNA, genomic DNA, plasmid DNA, and minicircle DNA, and any fragments thereof. The term “duplex” in the context of polynucleotides refers, in the usual and customary sense, to double strandedness. Nucleic acids can be linear or branched. For example, nucleic acids can be a linear chain of nucleotides or the nucleic acids can be branched, e.g., such that the nucleic acids comprise one or more arms or branches of nucleotides. Optionally, the branched nucleic acids are repetitively branched to form higher ordered structures such as dendrimers and the like.
[0027] Nucleic acids, including e.g., nucleic acids with a phosphothioate backbone, can include one or more reactive moieties. As used herein, the term reactive moiety includes any group capable of reacting with another molecule, e.g., a nucleic acid or polypeptide through covalent, non-covalent or other interactions. By way of example, the nucleic acid can include an amino acid reactive moiety that reacts with an amino acid on a protein or polypeptide through a covalent, non-covalent or other interaction.
[0028] The terms also encompass nucleic acids containing known nucleotide analogs or modified backbone residues or linkages, which are synthetic, naturally occurring, and non- naturally occurring, which have similar binding properties as the reference nucleic acid, and which are metabolized in a manner similar to the reference nucleotides. Examples of such analogs include, without limitation, phosphodiester derivatives including, e.g., phosphoramidate, phosphorodiamidate, phosphorothioate (also known as phosphothioate having double bonded sulfur replacing oxygen in the phosphate), phosphorodithioate, phosphonocarboxylic acids, phosphonocarboxylates, phosphonoacetic acid, phosphonoformic acid, methyl phosphonate, boron phosphonate, or O-methylphosphoroamidite linkages (see Eckstein, OLIGONUCLEOTIDES AND ANALOGUES: A PRACTICAL APPROACH,Oxford University Press) as well as modifications to the nucleotide bases such as in 5-methyl cytidine or pseudouridine.; and peptide nucleic acid backbones and linkages. Other analog nucleic acids include those with positive backbones; non- ionic backbones, modified sugars, and non-ribose backbones (e.g. phosphorodiamidate morpholino oligos or locked nucleic acids (LNA) as known in the art), including those described in U.S. Patent Nos. 5,235,033 and 5,034,506, and Chapters 6 and 7, ASC Symposium Series 580, CARBOHYDRATE MODIFICATIONS IN ANTISENSE RESEARCH, Sanghui & Cook, eds. Nucleic acids containing one or more carbocyclic sugars are also included within one definition of nucleic acids. Modifications of the ribose-phosphate backbone may be done for a variety of reasons, e.g., to increase the stability and half-life of such molecules in physiological environments or as probes on a biochip. Mixtures of naturally occurring nucleic acids and analogs can be made; alternatively, mixtures of different nucleic acid analogs, and mixtures of naturally occurring nucleic acids and analogs may be made. In embodiments, the intemucleotide linkages in DNA are phosphodiester, phosphodiester derivatives, or a combination of both.
[0029] Nucleic acids can include nonspecific sequences. As used herein, the term "nonspecific sequence" refers to a nucleic acid sequence that contains a series of residues that are not designed to be complementary to or are only partially complementary to any other nucleic acid sequence. By way of example, a nonspecific nucleic acid sequence is a sequence of nucleic acid residues that does not function as an inhibitory nucleic acid when contacted with a cell or organism.
[0030] A polynucleotide is typically composed of a specific sequence of four nucleotide bases: adenine (A); cytosine (C); guanine (G); and thymine (T) (uracil (U) for thymine (T) when the polynucleotide is RNA). Thus, the term “polynucleotide sequence” is the alphabetical representation of a polynucleotide molecule; alternatively, the term may be applied to the polynucleotide molecule itself. This alphabetical representation can be input into databases in a computer having a central processing unit and used for bioinformatics applications such as functional genomics and homology searching. Polynucleotides may optionally include one or more non-standard nucleotide(s), nucleotide analog(s) and / or modified nucleotides.
[0031] The term “complement,” as used herein, refers to a nucleotide (e.g., RNA or DNA) or a sequence of nucleotides capable of base pairing with a complementary nucleotide or sequence of nucleotides. As described herein and commonly known in the art the complementary (matching) nucleotide of adenosine is thymidine and the complementary(matching) nucleotide of guanosine is cytosine. Thus, a complement may include a sequence of nucleotides that base pair with corresponding complementary nucleotides of a second nucleic acid sequence. The nucleotides of a complement may partially or completely match the nucleotides of the second nucleic acid sequence. In embodiments, the complement partially matches the nucleotides of the second nucleic acid sequence. In embodiments, the complement completely matches the nucleotides of the second nucleic acid. Where the nucleotides of the complement completely match each nucleotide of the second nucleic acid sequence, the complement forms base pairs with each nucleotide of the second nucleic acid sequence. Where the nucleotides of the complement partially match the nucleotides of the second nucleic acid sequence only some of the nucleotides of the complement form base pairs with nucleotides of the second nucleic acid sequence. Examples of complementary sequences include coding and a non-coding sequences, wherein the non-coding sequence contains complementary nucleotides to the coding sequence and thus forms the complement of the coding sequence, A further example of complementary sequences are sense and antisense sequences, wherein the sense sequence contains complementary nucleotides to the antisense sequence and thus forms the complement of the antisense sequence.
[0032] As described herein the complementarity of sequences may be partial, in which only some of the nucleic acids match according to base pairing, or complete, where all the nucleic acids match according to base pairing. Thus, two sequences that are complementary to each other, may have a specified percentage of nucleotides that are the same (i.e., about 60% identity, preferably 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or higher identity over a specified region).
[0033] The term "amino acid" refers to naturally occurring and synthetic amino acids, as well as amino acid analogs and amino acid mimetics that function in a manner similar to the naturally occurring amino acids. Naturally occurring amino acids are those encoded by the genetic code, as well as those amino acids that are later modified, e.g., hydroxyproline, y- carboxyglutamate, and O-phosphoserine. Amino acid analogs refers to compounds that have the same basic chemical structure as a naturally occurring amino acid, i.e., an a carbon that is bound to a hydrogen, a carboxyl group, an amino group, and an R group, e.g., homoserine, norleucine, methionine sulfoxide, methionine methyl sulfonium. Such analogs have modified R groups (e.g., norleucine) or modified peptide backbones, but retain the same basic chemical structure as a naturally occurring amino acid. Amino acid mimetics refers to chemical compounds that have a structure that is different from the general chemical structure of anamino acid, but that functions in a manner similar to a naturally occurring amino acid. The terms “non-naturally occurring amino acid” and “unnatural amino acid” refer to amino acid analogs, synthetic amino acids, and amino acid mimetics which are not found in nature.
[0034] Amino acids may be referred to herein by either their commonly known three letter symbols or by the one-letter symbols recommended by the IUPAC-IUB Biochemical Nomenclature Commission. Nucleotides, likewise, may be referred to by their commonly accepted single-letter codes.
[0035] The terms "polypeptide," "peptide" and "protein" are used interchangeably herein to refer to a polymer of amino acid residues, wherein the polymer may In embodiments be conjugated to a moiety that does not consist of amino acids. The terms apply to amino acid polymers in which one or more amino acid residue is an artificial chemical mimetic of a corresponding naturally occurring amino acid, as well as to naturally occurring amino acid polymers and non-naturally occurring amino acid polymers. A "fusion protein" refers to a chimeric protein encoding two or more separate protein sequences that are recombinantly expressed as a single moiety.
[0036] An amino acid or nucleotide base "position" is denoted by a number that sequentially identifies each amino acid (or nucleotide base) in the reference sequence based on its position relative to the N-terminus (or 5’-end). Due to deletions, insertions, truncations, fusions, and the like that must be considered when determining an optimal alignment, in general the amino acid residue number in a test sequence determined by simply counting from the N-terminus will not necessarily be the same as the number of its corresponding position in the reference sequence. For example, in a case where a variant has a deletion relative to an aligned reference sequence, there will be no amino acid in the variant that corresponds to a position in the reference sequence at the site of deletion. Where there is an insertion in an aligned reference sequence, that insertion will not correspond to a numbered amino acid position in the reference sequence. In the case of truncations or fusions there can be stretches of amino acids in either the reference or aligned sequence that do not correspond to any amino acid in the corresponding sequence.
[0037] The terms "numbered with reference to" or "corresponding to," when used in the context of the numbering of a given amino acid or polynucleotide sequence, refers to the numbering of the residues of a specified reference sequence when the given amino acid or polynucleotide sequence is compared to the reference sequence. An amino acid residue in a protein "corresponds" to a given residue when it occupies the same essential structuralposition within the protein as the given residue. One skilled in the art will immediately recognize the identity and location of residues corresponding to a specific position in a protein (e.g., BS2) in other proteins with different numbering systems. For example, by performing a simple sequence alignment with a protein (e.g., BS2) the identity and location of residues corresponding to specific positions of the protein are identified in other protein sequences aligning to the protein. For example, a selected residue in a selected protein corresponds to glutamic acid at position 138 when the selected residue occupies the same essential spatial or other structural relationship as a glutamic acid at position 138. In some embodiments, where a selected protein is aligned for maximum homology with a protein, the position in the aligned selected protein aligning with glutamic acid 138 is the to correspond to glutamic acid 138. Instead of a primary sequence alignment, a three dimensional structural alignment can also be used, e.g., where the structure of the selected protein is aligned for maximum correspondence with the glutamic acid at position 138, and the overall structures compared. In this case, an amino acid that occupies the same essential position as glutamic acid 138 in the structural model is said to correspond to the glutamic acid 138 residue.
[0038] "Conservatively modified variants" applies to both amino acid and nucleic acid sequences. With respect to particular nucleic acid sequences, "conservatively modified variants" refers to those nucleic acids that encode identical or essentially identical amino acid sequences. Because of the degeneracy of the genetic code, a number of nucleic acid sequences will encode any given protein. For instance, the codons GCA, GCC, GCG and GCU all encode the amino acid alanine. Thus, at every position where an alanine is specified by a codon, the codon can be altered to any of the corresponding codons described without altering the encoded polypeptide. Such nucleic acid variations are "silent variations," which are one species of conservatively modified variations. Every nucleic acid sequence herein which encodes a polypeptide also describes every possible silent variation of the nucleic acid. One of skill will recognize that each codon in a nucleic acid (except AUG, which is ordinarily the only codon for methionine, and TGG, which is ordinarily the only codon for tryptophan) can be modified to yield a functionally identical molecule. Accordingly, each silent variation of a nucleic acid which encodes a polypeptide is implicit in each described sequence.
[0039] As to amino acid sequences, one of skill will recognize that individual substitutions, deletions or additions to a nucleic acid, peptide, polypeptide, or protein sequence which alters, adds or deletes a single amino acid or a small percentage of amino acids in the encodedsequence is a "conservatively modified vari nt" where the alteration results in the substitution of an amino acid with a chemically similar amino acid. Conservative substitution tables providing functionally similar amino acids are well known in the art. Such conservatively modified variants are in addition to and do not exclude polymorphic variants, interspecies homologs, and alleles of the disclosure.
[0040] The following eight groups each contain amino acids that are conservative substitutions for one another:1) Alanine (A), Glycine (G);2) Aspartic acid (D), Glutamic acid (E);3) Asparagine (N), Glutamine (Q);4) Arginine (R), Lysine (K);5) Isoleucine (I), Leucine (L), Methionine (M), Valine (V);6) Phenylalanine (F), Tyrosine (Y), Tryptophan (W);7) Serine (S), Threonine (T); and8) Cysteine (C), Methionine (M) (see, e.g., Creighton, Proteins (1984)).
[0041] The terms "identical" or percent "identity," in the context of two or more nucleic acids or polypeptide sequences, refer to two or more sequences or subsequences that are the same or have a specified percentage of amino acid residues or nucleotides that are the same (i.e., about 60% identity, preferably 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or higher identity over a specified region, when compared and aligned for maximum correspondence over a comparison window or designated region) as measured using a BLAST or BLAST 2.0 sequence comparison algorithms with default parameters described below, or by manual alignment and visual inspection (see, e.g., NCBI web site www.ncbi.nlm.nih.gov / BLAST / or the like). Such sequences are then said to be "substantially identical." This definition also refers to, or may be applied to, the compliment of a test sequence. The definition also includes sequences that have deletions and / or additions, as well as those that have substitutions. As described below, the preferred algorithms can account for gaps and the like. Preferably, identity exists over a region that is at least about 25 amino acids or nucleotides in length, or more preferably over a region that is 50-100 amino acids or nucleotides in length.
[0042] "Percentage of sequence identity" is determined by comparing two optimally aligned sequences over a comparison window, wherein the portion of the polynucleotide orpolypeptide sequence in the comparison window may comprise additions or deletions (i.e., gaps) as compared to the reference sequence (which does not comprise additions or deletions) for optimal alignment of the two sequences. The percentage is calculated by determining the number of positions at which the identical nucleic acid base or amino acid residue occurs in both sequences to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the window of comparison and multiplying the result by 100 to yield the percentage of sequence identity.
[0043] A "comparison window", as used herein, includes reference to a segment of any one of the number of contiguous positions selected from the group consisting of, e.g., a full length sequence or from 20 to 600, about 50 to about 200, or about 100 to about 150 amino acids or nucleotides in which a sequence may be compared to a reference sequence of the same number of contiguous positions after the two sequences are optimally aligned. Methods of alignment of sequences for comparison are well-known in the art. Optimal alignment of sequences for comparison can be conducted, e.g., by the local homology algorithm of Smith and Waterman (1970) Adv. Appl. Math. 2:482c, by the homology alignment algorithm of Needleman and Wunsch (1970) J. Mol. Biol. 48:443, by the search for similarity method of Pearson and Lipman (1988) Proc. Nat’l. Acad. Sci. USA 85:2444, by computerized implementations of these algorithms (GAP, BESTFIT, FASTA, and TFASTA in the Wisconsin Genetics Software Package, Genetics Computer Group, 575 Science Dr., Madison, WI), or by manual alignment and visual inspection (see, e.g., Ausubel et al., Current Protocols in Molecular Biology (1995 supplement)).
[0044] An example of an algorithm that is suitable for determining percent sequence identity and sequence similarity are the BLAST and BLAST 2.0 algorithms, which are described in Altschul et al. (1977) Nuc. Acids Res. 25:3389-3402, and Altschul et al. (1990) J. Mol. Biol. 215:403-410, respectively. Software for performing BLAST analyses is publicly available through the National Center for Biotechnology Information (www.ncbi.nlm.nih.gov / ). This algorithm involves first identifying high scoring sequence pairs (HSPs) by identifying short words of length W in the query sequence, which either match or satisfy some positive-valued threshold score T when aligned with a word of the same length in a database sequence. T is referred to as the neighborhood word score threshold (Altschul et al., supra). These initial neighborhood word hits act as seeds for initiating searches to find longer HSPs containing them. The word hits are extended in both directions along each sequence for as far as the cumulative alignment score can be increased.Cumulative scores are calculated using, for nucleotide sequences, the parameters M (reward score for a pair of matching residues; always > 0) and N (penalty score for mismatching residues; always < 0). For amino acid sequences, a scoring matrix is used to calculate the cumulative score. Extension of the word hits in each direction are halted when: the cumulative alignment score falls off by the quantity X from its maximum achieved value; the cumulative score goes to zero or below, due to the accumulation of one or more negativescoring residue alignments; or the end of either sequence is reached. The BLAST algorithm parameters W, T, and X determine the sensitivity and speed of the alignment. The BLASTN program (for nucleotide sequences) uses as defaults a word length (W) of 1 1, an expectation (E) or 10, M=5, N=-4 and a comparison of both strands. For amino acid sequences, the BLASTP program uses as defaults a word length of 3, and expectation (E) of 10, and the BLOSUM62 scoring matrix (see Henikoff and Henikoff (1989) Proc. Natl. Acad. Sci. USA 89:10915) alignments (B) of 50, expectation (E) of 10, M=5, N=-4, and a comparison of both strands.
[0045] The BLAST algorithm also performs a statistical analysis of the similarity between two sequences (see, e.g., Karlin and Altschul (1993) Proc. Natl. Acad. Sci. USA 90:5873- 5787). One measure of similarity provided by the BLAST algorithm is the smallest sum probability (P(N)), which provides an indication of the probability by which a match between two nucleotide or amino acid sequences would occur by chance. For example, a nucleic acid is considered similar to a reference sequence if the smallest sum probability in a comparison of the test nucleic acid to the reference nucleic acid is less than about 0.2, more preferably less than about 0.01, and most preferably less than about 0.001.
[0046] An indication that two nucleic acid sequences or polypeptides are substantially identical is that the polypeptide encoded by the first nucleic acid is immunologically cross reactive with the antibodies raised against the polypeptide encoded by the second nucleic acid, as described below. Thus, a polypeptide is typically substantially identical to a second polypeptide, for example, where the two peptides differ only by conservative substitutions. Another indication that two nucleic acid sequences are substantially identical is that the two molecules or their complements hybridize to each other under stringent conditions, as described below. Yet another indication that two nucleic acid sequences are substantially identical is that the same primers can be used to amplify the sequence.
[0047] The phrase "specifically (or selectively) binds" to an antibody or "specifically (or selectively) immunoreactive with," when referring to a protein or peptide, refers to a bindingreaction that is determinative of the presence of the protein, often in a heterogeneous population of proteins and other biologies. Thus, under designated immunoassay conditions, the specified antibodies bind to a particular protein at least two times the background and more typically more than 10 to 100 times background. Specific binding to an antibody under such conditions requires an antibody that is selected for its specificity for a particular protein. For example, polyclonal antibodies can be selected to obtain only a subset of antibodies that are specifically immunoreactive with the selected antigen and not with other proteins. This selection may be achieved by subtracting out antibodies that cross-react with other molecules. A variety of immunoassay formats may be used to select antibodies specifically immunoreactive with a particular protein. For example, solid-phase ELISA immunoassays are routinely used to select antibodies specifically immunoreactive with a protein (see, e.g., Harlow & Lane, Using Antibodies, A Laboratory Manual (1998) for a description of immunoassay formats and conditions that can be used to determine specific immunoreactivity) .
[0048] A "ligand" refers to an agent, e.g., a polypeptide or other molecule, capable of binding to a receptor or antibody, antibody variant, antibody region or fragment thereof.
[0049] Techniques for conjugating therapeutic agents to antibodies are well known (see, e.g., Amon et al., "Monoclonal Antibodies For Immunotargeting Of Drugs In Cancer Therapy", in Monoclonal Antibodies And Cancer Therapy, Reisfeld et al. (eds.), pp. 243-56 (Alan R. Liss, Inc. 1985); Hellstrom et al., “Antibodies For Drug Delivery”in Controlled Drug Delivery (2ndEd.), Robinson et al. (eds.), pp. 623-53 (Marcel Dekker, Inc. 1987); Thorpe, "Antibody Carriers Of Cytotoxic Agents In Cancer Therapy: A Review" in Monoclonal Antibodies ‘84: Biological And Clinical Applications, Pinchera et al. (eds.), pp. 475-506 (1985); and Thorpe et al., "The Preparation And Cytotoxic Properties Of Antibody- Toxin Conjugates", Immunol. Rev., 62:119-58 (1982)). As used herein, the term “antibodydrug conjugate’" or “ADC” refers to a therapeutic agent conjugated or otherwise covalently bound to to an antibody.
[0050] For specific proteins described herein, the named protein includes any of the protein’s naturally occurring forms, variants or homologs that maintain the protein transcription factor activity (e.g., within at least 50%, 80%, 90%, 95%, 96%, 97%, 98%, 99% or 100% activity compared to the native protein). In some embodiments, variants or homologs have at least 90%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity across the whole sequence or a portion of the sequence (e.g. a 50, 100, 150 or 200continuous amino acid portion) compared to a naturally occurring form. Tn other embodiments, the protein is the protein as identified by its NCBI sequence reference. In other embodiments, the protein is the protein as identified by its NCBI sequence reference, homolog or functional fragment thereof.
[0051] The term "gene" means the segment of DNA involved in producing a protein; it includes regions preceding and following the coding region (leader and trailer) as well as intervening sequences (introns) between individual coding segments (exons). The leader, the trailer as well as the introns include regulatory elements that are necessary during the transcription and the translation of a gene. Further, a "protein gene product" is a protein expressed from a particular gene.
[0052] The terms "plasmid", "vector" or "expression vector" refer to a nucleic acid molecule that encodes for genes and / or regulatory elements necessary for the expression of genes. Expression of a gene from a plasmid can occur in cis or in trans. If a gene is expressed in cis, the gene and the regulatory elements are encoded by the same plasmid. Expression in trans refers to the instance where the gene and the regulatory elements are encoded by separate plasmids.
[0053] The terms "transfection", "transduction", "transfecting" or "transducing" can be used interchangeably and are defined as a process of introducing a nucleic acid molecule or a protein to a cell. Nucleic acids are introduced to a cell using non-viral or viral-based methods. The nucleic acid molecules may be gene sequences encoding complete proteins or functional portions thereof. Non-viral methods of transfection include any appropriate transfection method that does not use viral DNA or viral particles as a delivery system to introduce the nucleic acid molecule into the cell. Exemplary non-viral transfection methods include calcium phosphate transfection, liposomal transfection, nucleofection, sonoporation, transfection through heat shock, magnetifection and electroporation. In some embodiments, the nucleic acid molecules are introduced into a cell using electroporation following standard procedures well known in the art. For viral-based methods of transfection any useful viral vector may be used in the methods described herein. Examples for viral vectors include, but are not limited to retroviral, adenoviral, lentiviral and adeno-associated viral vectors. In some embodiments, the nucleic acid molecules are introduced into a cell using a retroviral vector following standard procedures well known in the art. The terms "transfection" or "transduction" also refer to introducing proteins into a cell from the external environment. Typically, transduction or transfection of a protein relies on attachment of a peptide orprotein capable of crossing the cell membrane to the protein of interest. See, e.g., Ford et al. (2001) Gene Therapy 8:1-4 and Prochiantz (2007) Nat. Methods 4:119-20.
[0054] A "label" or a "detectable moiety" is a composition detectable by spectroscopic, photochemical, biochemical, immunochemical, chemical, or other physical means. For example, useful labels include 32P, fluorescent dyes, electron-dense reagents, enzymes (e.g., as commonly used in an ELISA), biotin, digoxigenin, or haptens and proteins or other entities which can be made detectable, e.g., by incorporating a radiolabel into a peptide or antibody specifically reactive with a target peptide. Any appropriate method known in the art for conjugating an antibody to the label may be employed, e.g., using methods described in Hermanson, Bioconjugate Techniques 1996, Academic Press, Inc., San Diego.
[0055] When the label or detectable moiety is a radioactive metal or paramagnetic ion, the agent may be reacted with another long-tailed reagent having a long tail with one or more chelating groups attached to the long tail for binding to these ions. The long tail may be a polymer such as a polylysine, polysaccharide, or other derivatized or derivatizable chain having pendant groups to which the metals or ions may be added for binding. Examples of chelating groups that may be used according to the disclosure include, but are not limited to, ethylenediaminetetraacetic acid (EDTA), diethylenetriaminepentaacetic acid (DTP A), DOTA, NOTA, NETA, TETA, porphyrins, polyamines, crown ethers, bis- thiosemicarbazones, polyoximes, and like groups. The chelate is normally linked to the PSMA antibody or functional antibody fragment by a group, which enables the formation of a bond to the molecule with minimal loss of immunoreactivity and minimal aggregation and / or internal cross-linking. The same chelates, when complexed with non-radioactive metals, such as manganese, iron and gadolinium are useful for MRI, when used along with the antibodies and carriers described herein. Macrocyclic chelates such as NOTA, DOTA, and TETA are of use with a variety of metals and radiometals including, but not limited to, radionuclides of gallium, yttrium and copper, respectively. Other ring-type chelates such as macrocyclic polyethers, which are of interest for stably binding nuclides, such as223Ra for RAIT may be used. In certain embodiments, chelating moieties may be used to attach a PET imaging agent, such as an A1-18F complex, to a targeting molecule for use in PET analysis.
[0056] "Contacting" is used in accordance with its plain ordinary meaning and refers to the process of allowing at least two distinct species (e.g. antibodies and antigens) to become sufficiently proximal to react, interact, or physically touch. It should be appreciated, however, that the resulting reaction product can be produced directly from a reaction between theadded reagents or from an intermediate from one or more of the added reagents which can be produced in the reaction mixture.
[0057] The term "contacting" may include allowing two species to react, interact, or physically touch, wherein the two species may be, for example, a pharmaceutical composition as provided herein and a cell. In embodiments contacting includes, for example, allowing a pharmaceutical composition as described herein to interact with a cell.
[0058] A "cell" as used herein, refers to a cell carrying out metabolic or other function sufficient to preserve or replicate its genomic DNA. A cell can be identified by well-known methods in the art including, for example, presence of an intact membrane, staining by a particular dye, ability to produce progeny or, in the case of a gamete, ability to combine with a second gamete to produce a viable offspring. Cells may include prokaryotic and eukaryotic cells. Prokaryotic cells include but are not limited to bacteria. Eukaryotic cells include, but are not limited to, yeast cells and cells derived from plants and animals, for example mammalian, insect (e.g., spodoptera) and human cells.
[0059] The term "recombinant" when used with reference, e.g., to a cell, nucleic acid, protein, or vector, indicates that the cell, nucleic acid, protein or vector, has been modified by the introduction of a heterologous nucleic acid or protein or the alteration of a native nucleic acid or protein, or that the cell is derived from a cell so modified. Thus, for example, recombinant cells express genes that are not found within the native (non-recombinant) form of the cell or express native genes that are otherwise abnormally expressed, under expressed or not expressed at all. Transgenic cells and plants are those that express a heterologous gene or coding sequence, typically as a result of recombinant methods.
[0060] The term "isolated", when applied to a nucleic acid or protein, denotes that the nucleic acid or protein is essentially free of other cellular components with which it is associated in the natural state. It can be, for example, in a homogeneous state and may be in either a dry or aqueous solution. Purity and homogeneity are typically determined using analytical chemistry techniques such as polyacrylamide gel electrophoresis or high performance liquid chromatography. A protein that is the predominant species present in a preparation is substantially purified.
[0061] The term "heterologous" when used with reference to portions of a nucleic acid indicates that the nucleic acid comprises two or more subsequences that are not found in the same relationship to each other in nature. For instance, the nucleic acid is typically recombinantly produced, having two or more sequences from unrelated genes arranged tomake a new functional nucleic acid, e.g., a promoter from one source and a coding region from another source. Similarly, a heterologous protein indicates that the protein comprises two or more subsequences that are not found in the same relationship to each other in nature (e.g., a fusion protein).
[0062] The term "exogenous" refers to a molecule or substance (e.g., a compound, nucleic acid or protein) that originates from outside a given cell or organism. For example, an "exogenous promoter" as referred to herein is a promoter that does not originate from the cell or organism it is expressed by. Conversely, the term "endogenous" or "endogenous promoter" refers to a molecule or substance that is native to, or originates within, a given cell or organism.
[0063] As defined herein, the term "inhibition", "inhibit", "inhibiting" and the like in reference to cell proliferation (e.g., cancer cell proliferation) means negatively affecting (e.g., decreasing proliferation) or killing the cell. In some embodiments, inhibition refers to reduction of a disease or symptoms of disease (e.g., cancer, cancer cell proliferation). Thus, inhibition includes, at least in part, partially or totally blocking stimulation, decreasing, preventing, or delaying activation, or inactivating, desensitizing, or down-regulating signal transduction or enzymatic activity or the amount of a protein (e.g. a cancer-associated protein). Similarly an "inhibitor" is a compound or protein that inhibits a receptor or another protein, e.g.,, by binding, partially or totally blocking, decreasing, preventing, delaying, inactivating, desensitizing, or down-regulating activity (e.g., a receptor activity or a protein activity).
[0064] The term "expression" includes any step involved in the production of the polypeptide including, but not limited to, transcription, post-transcriptional modification, translation, post-translational modification, and secretion. Expression can be detected using conventional techniques for detecting protein (e.g., ELISA, Western blotting, flow cytometry, immunofluorescence, immunohistochemistry, etc.).
[0065] “Biological sample” or “sample” refer to materials obtained from or derived from a subject or patient. A biological sample includes sections of tissues such as biopsy and autopsy samples, and frozen sections taken for histological purposes. Such samples include bodily fluids such as blood and blood fractions or products (e.g., serum, plasma, platelets, red blood cells, and the like), sputum, tissue, cultured cells (e.g., primary cultures, explants, and transformed cells) stool, urine, synovial fluid joint tissue, synovial tissue, synoviocytes, fibroblast-like synoviocytes, macrophage-like synoviocytes, immune cells, hematopoieticcells, fibroblasts, macrophages, T cells, etc. A biological sample is typically obtained from a eukaryotic organism, such as a mammal such as a primate e.g., chimpanzee or human; cow; dog; cat; a rodent, e.g., guinea pig, rat, mouse; rabbit; or a bird; reptile; or fish.
[0066] A “control” or “standard control” refers to a sample, measurement, or value that serves as a reference, usually a known reference, for comparison to a test sample, measurement, or value. For example, a test sample can be taken from a patient suspected of having a given disease (e.g. cancer) and compared to a known normal (non-diseased) individual (e.g. a standard control subject). A standard control can also represent an average measurement or value gathered from a population of similar individuals (e.g. standard control subjects) that do not have a given disease (i.e. standard control population), e.g., healthy individuals with a similar medical background, same age, weight, etc. A standard control value can also be obtained from the same individual, e.g. from an earlier-obtained sample from the patient prior to disease onset. For example, a control can be devised to compare therapeutic benefit based on pharmacological data (e.g., half-life) or therapeutic measures (e.g., comparison of side effects). Controls are also valuable for determining the significance of data. For example, if values for a given parameter are widely variant in controls, variation in test samples will not be considered as significant. One of skill will recognize that standard controls can be designed for assessment of any number of parameters (e.g. RNA levels, protein levels, specific cell types, specific bodily fluids, specific tissues, etc).
[0067] The term “chromosomal fragment DNA sequence” is used herein according to its plain ordinary meaning and refers to a nucleic acid sequence that includes a nucleic acid sequence that is the same or a complement of a chromosomal sequence.
[0068] The term “DNA crosslinking agent” is used herein according to its plain ordinary meaning and refers to a molecule that contains two or more reactive ends capable chemically attaching to specific functional groups of a nucleic acid sequence (e.g., DNA) and / or a protein. In embodiments, the DNA crosslinking agent is a formaldehyde.
[0069] The term “double stranded cleaved site” is used herein according to its plain ordinary meaning and refers to a break in a double stranded nucleic acid sequence (e.g., DNA) caused by a DNA endonuclease.
[0070] The term “DNA endonuclease” is used herein according to its plain ordinary meaning and refers to an enzyme that cleaves a phosphodiester bond within a nucleic acid sequence (e.g., DNA). In embodiments, the DNA endonuclease is a restriction endonuclease. The term “restriction endonuclease” refers to an endonuclease that cleaves DNA intofragments at or near specific recognition sites (e.g., restriction sites). In embodiments, the DNA endonuclease is a DpnII restriction endonuclease, an Mbol restriction endonuclease, an Nlalll restriction endonuclease, a Ddel restriction endonuclease, or a DpnII restriction endonuclease.
[0071] The term “ligated chromosomal DNA’' is used herein according to its plain ordinary meaning and refers to a chromosomal DNA sequence that is formed by a ligase joining two separate (e.g., cleaved) chromosomal DNA sequences. In embodiments, the ligated chromosomal DNA includes an endogenous ligation DNA sequence or a non-endogenous ligation DNA sequence.
[0072] The term “ligase” is used herein according to its plain ordinary meaning and refers to an enzyme that catalyzes the formation of a phosphodiester bond between two nucleotides of separate nucleic acid strands (e.g., DNA), thereby forming a single continuous strand. In embodiments, the ligase is a T4 DNA ligase.
[0073] The term “endogenous ligation DNA sequence” is used herein according to its plain ordinary meaning and refers to a naturally occurring sequence of ligated chromosomal DNA that can be found in a cell nucleus as disclosed herein. In embodiments, the endogenous ligation DNA sequence includes two DNA sequences that are endogenously proximal in location or three dimensional space in a cell nucleus. In embodiments, the endogenous ligation DNA sequence includes two DNA sequences that are on the same chromosome.
[0074] The term “non-endogenous ligation DNA sequence” is used herein according to its plain ordinary meaning and refers to a sequence of ligated chromosomal DNA that is not found in a cell nucleus that is the subject of the methods described herein. In embodiments, the non-endogenous ligation DNA sequence includes two DNA sequences present in a cell nucleus chromosomal sequence that is the subject of the methods described herein that are not endogenously proximal in location in a cell nucleus that is the subject of the methods described herein. In embodiments, the non-endogenous ligation DNA sequences includes two DNA sequences present in a cell nucleus chromosomal sequence that is subject of the methods described herein that are endogenously distant from each other in a cell nucleus that is subject to the methods described herein. In embodiments, the non-endogenous ligation DNA sequence includes a chromosomal DNA sequence and an extrachromosomal DNA sequence. In embodiments, the non-endogenous ligation DNA sequence includes an extrachromosomal DNA sequence or two extrachromosomal DNA sequences that areendogenously distant from each other in a cell nucleus that is the subject of the methods described herein.
[0075] The term “transposase” is used herein according to its plain ordinary meaning and refers to an enzyme that fragments (i.e., cuts or cleaves) a nucleic acid sequence (e.g., a ligated chromosomal DNA), thereby creating two ends and ligates a tag nucleic acid sequence (e.g., a 5' hybridization DNA sequence and / or a 3' hybridization DNA sequence) to each end of the fragmented nucleic acid sequence. In embodiments, the combined fragmentation of a nucleic acid sequence and subsequent tagging of the fragmented sequence is referred to herein as tagmentation. In embodiments, the transposase is Tn5 transposase.
[0076] The term “transposed chromosomal DNA sequence” is used herein according to its plain ordinary meaning and refers to a chromosomal DNA sequence that has undergone tagmentation. In embodiments, the transposed chromosomal DNA sequence includes from 5’ to 3' a first DNA hybridization sequence, a chromosomal fragment DNA sequence, and a second DNA hybridization sequence. In embodiments, a ligated chromosomal DNA is contacted with a transposase, a 5' hybridization DNA sequence and a 3' hybridization DNA sequence, thereby forming a transposed chromosomal DNA sequence.
[0077] The term “hybridization DNA sequence” is used herein according to its plain ordinary meaning and refers to a nucleic acid sequence that binds (e.g., hybridizes) or is capable of binding to a complementary DNA sequence. In embodiments, the hybridization sequence and its complement are sufficiently complementary to hybridize. In embodiments, the hybridization sequence and its complement are perfectly complementary to hybridize. The term “5' hybridization DNA sequence” refers to a hybridization DNA sequence on the 5' end of a nucleic acid. The term “3' hybridization DNA sequence” refers to a hybridization DNA sequence on the 3' end of a nucleic acid. The term “complementary hybridization sequence” refers to a nucleic acid sequence that is complementary to and hybridizes with a hybridization DNA sequence. The term “5' complementary hybridization DNA sequence” refers to a complementary hybridization DNA sequence on the 5' end of a nucleic acid. The term “3' complementary hybridization DNA sequence” refers to a hybridization DNA sequence on the 3' end of a nucleic acid.
[0078] The term “hybridized template DNA sequence” is used herein according to its plain ordinary meaning and refers to a nucleic acid sequence including a hybridization DNA sequence bound (e.g., hybridized) to its complementary hybridization sequence. In embodiments, the hybridization sequence and its complementary hybridization sequence aresufficiently complementary to hybridize. In embodiments, the hybridization sequence and its complementary hybridization sequence are perfectly complementary to hybridize. In embodiments, the hybridized template DNA sequence is formed to allow for amplification or extension of a nucleic acid sequence.
[0079] The term “library nucleic acid sequence’' is used herein according to its plain ordinary meaning and refers to a nucleic acid sequence that includes a primer sequence and a hybridization sequence. In embodiments, the primer sequence is an initiation point for DNA synthesis. In embodiments, the primer sequence is a forward primer sequence or reverse primer sequence. In embodiments, the library nucleic acid sequence further comprises a barcode sequence. In embodiments, the library nucleic acid sequence is an identifying library nucleic acid sequence. In embodiments, the library nucleic acid sequence includes from 5’ to 3’ a 3' complementary hybridization sequence and a reverse primer sequence.
[0080] The term “identifying library nucleic acid sequence” is used herein according to its plain ordinary meaning and refers to a library nucleic acid sequence used to identify chromosomal fragment DNA sequences from a specific cell nucleus and / or separate compartment. In embodiments, the identifying library nucleic acid sequence comprises from 5’ to 3' a forward primer sequence, a barcode sequence, and a 5’ complementary hybridization sequence.
[0081] The term “forward primer sequence” is used herein according to its plain ordinary meaning and refers to a sequence that hybridizes to a primer for extension and / or amplification.
[0082] The term “barcode sequence” is used herein according to its plain ordinary meaning and refers to a nucleic acid sequences that is added to a chromosomal fragment DNA sequence to identify which cell nucleus the chromosomal fragment DNA originated.
[0083] The term “amplified chromosomal DNA” is used herein according to its plain ordinary meaning and refers to a nucleic acid sequence including a chromosomal fragment DNA sequence that has been amplified. In embodiments, the amplified chromosomal DNA includes from 5' to 3' the forward primer sequence or complement thereof, the barcode sequence or complement thereof, the 5' complementary hybridization sequence or complement thereof, the chromosomal fragment DNA sequence or complement thereof, and the second DNA hybridization sequence or complement thereof. In embodiments, the amplified chromosomal DNA includes comprising from 5’ to 3’ the forward primer sequence or complement thereof, the barcode sequence or complement thereof, the 5’ complementaryhybridization sequence or complement thereof, the chromosomal fragment DNA sequence or complement thereof, the second DNA hybridization sequence or complement thereof, the 3' complementary hybridization sequence or complement thereof and the reverse primer sequence or complement thereof.
[0084] The term “copy number variation” is used herein according to its plain ordinary meaning and refers to a structural variation in a genome in which sections of the genome are repeated. In embodiments, the number of repeats in the genome varies between cell types or tissue types. In embodiments, the number of repeats in the genome varies between organisms. In embodiments, the copy number variation may include a duplication event or a deletion event.
[0085] The term “structural variation” is used herein according to its plain ordinary meaning and refers to the variation in structure of a chromosome between one or more cells. In embodiments, the structural variation is a deletion, a duplication, a copy number variant, an insertion, an inversion, or a translocation of a nucleic acid sequence within the chromosome.
[0086] The term “extrachromosomal DNA” or “ecDNA” is used herein according to its plain ordinary meaning and refers to a DNA sequence that is not found in a chromosome. In embodiments, the ecDNA is found inside the cell nucleus.
[0087] The terms “sequencing”, “sequence determination”, “determining a nucleotide sequence”, and the like are used herein according to their plain ordinary meaning and include determination of partial as well as full sequence information, including the identification, ordering, or locations of the nucleotides that comprise the polynucleotide being sequenced, and inclusive of the physical processes for generating such sequence information. That is, the term includes sequence comparisons, fingerprinting, and like levels of information about a target polynucleotide (e.g., a fragmented chromosomal DNA sequence), as well as the express identification and ordering of nucleotides in a target polynucleotide. The term also includes the determination of the identification, ordering, and locations of one, two, or three of the four types of nucleotides within a target polynucleotide. Sequencing methods, such as those outlined in U.S. Pat. No. 5,302,509 can be carried out using the nucleotides described herein. The sequencing methods are preferably carried out with the target polynucleotide arrayed on a solid substrate. Multiple target polynucleotides can be immobilized on the solid support through linker molecules, or can be attached to particles, e.g., microspheres, which can also be attached to a solid substrate. The solid substrate is in the form of a chip, a bead, awell, a capillary tube, a slide, a wafer, a filter, a fiber, a porous media, or a column. In embodiments, the sequencing is sequencing-by-synthesis (SBS).
[0088] The term “sequencing-by-synthesis” or “SBS” is used herein according to its plain ordinary meaning and refers to a next-generation sequencing (NGS) method in which the identity of the sequence of a nucleic acid (e.g., chromosomal fragment DNA) is determined via DNA synthesis. In embodiments, the DNA synthesis includes sequential incorporation of chemically modified nucleotides based on complementarity to the DNA template strand (e.g., fragmented chromosomal DNA sequence). In embodiments, the chemically modified nucleotides are chemically modified deoxynucleoside triphosphates (dNTPs). In embodiments, each chemically modified dNTP includes a fluorescent tag and a reversible terminator. In embodiments, the reversible terminator blocks incorporation of a subsequent nucleotide (e.g., chemically modified dNTP). In embodiments, a chemically modified nucleotides bind to the DNA template strand (e.g., fragmented chromosomal DNA sequence) through natural complementarity. In embodiments, during each SBS cycle, the fluorescence of the chemically modified dNTP is detected. In embodiments, the reversible terminator is cleaved from the chemically modified dNTP, thereby allowing incorporation of a subsequent nucleotide (e.g., chemically modified dNTP). Sequencing by synthesis methods are well known in the art. See, e.g., U.S. Patent Nos. 5,302,509; 7,033,764; 7,416,844; 8,241,573; and Seo et al., PNAS, 2005; 102(17):5926-31, each of which are incorporated herein in their entirety and for all purposes. Chemically modified dNTPs are well known in the art. See, e.g., U.S. Patent Nos. 9,593,373 and 8,399,188; and PCT Patent Pub. Nos. WO 2017 / 058953 and WO 2017 / 205336; each of which is incorporated herein by reference in its entirety and for all purposes.
[0089] The term “sequencing reaction mixture” is used herein according to its plain ordinary meaning and refers to an aqueous mixture that contains the reagents necessary to allow a dNTP or dNTP analogue to add a nucleotide to a DNA strand by a DNA polymerase. In embodiments, the sequencing reaction mixture includes a buffer. In embodiments, the buffer includes an acetate buffer, 3-(N-morpholino) propanesulfonic acid (MOPS) buffer, N- (2-Acetamido)-2-aminoethanesulfonic acid (ACES) buffer, phosphate-buffered saline (PBS) buffer, 4-(2-hydroxyethyl)-l -piperazineethanesulfonic acid (HEPES) buffer, N-(l,l- Dimethyl-2-hydroxyethyl)-3-amino-2-hydroxypropanesulfonic acid (AMPSO) buffer, borate buffer (e.g., borate buffered saline, sodium borate buffer, boric acid buffer), 2-Amino-2- methyl- 1,3 -propanediol (AMPD) buffer, N-cyclohexyl-2-hydroxyl-3-aminopropanesulfonicacid (CAPSO) buffer, 2-Amino-2-methyl-l -propanol (AMP) buffer, 4-(Cyclohexylamino)-l - butanesulfonic acid (CABS) buffer, glycine-NaOH buffer, N-Cyclohexyl-2- aminoethanesulfonic acid (CHES) buffer, tris(hydroxymethyl)aminomethane (Tris) buffer, or a N-cyclohexyl-3-aminopropanesulfonic acid (CAPS) buffer. In embodiments, the buffer is a borate buffer. In embodiments, the buffer is a CHES buffer. In embodiments, the sequencing reaction mixture includes nucleotides, wherein the nucleotides include a reversible terminating moiety and a label covalently linked to the nucleotide via a cleavable linker. In embodiments, the sequencing reaction mixture includes a buffer, DNA polymerase, detergent (e.g., Triton X), a chelator (e.g., EDTA), or salts (e.g., ammonium sulfate, magnesium chloride, sodium chloride, or potassium chloride).
[0090] The term “sequencing cycle” is used herein according to its plain ordinary meaning and refers to incorporating one or more nucleotides (e.g., nucleotide analogues) to the 3’ end of a polynucleotide with a polymerase, and detecting one or more labels that identify the one or more nucleotides incorporated. The sequencing may be accomplished by, for example, sequencing by synthesis, pyrosequencing, and the like. In embodiments, a sequencing cycle includes extending a complementary polynucleotide by incorporating a first nucleotide using a polymerase, wherein the polynucleotide is hybridized to a template nucleic acid, detecting the first nucleotide, and identifying the first nucleotide. In embodiments, to begin a sequencing cycle, one or more differently labeled nucleotides and a DNA polymerase can be introduced. Following nucleotide addition, signals produced (e.g., via excitation and emission of a detectable label) can be detected to determine the identity of the incorporated nucleotide (based on the labels on the nucleotides). Reagents can then be added to remove the 3’ reversible terminator and to remove labels from each incorporated base. Reagents, enzymes and other substances can be removed between steps by washing. Cycles may include repeating these steps, and the sequence of each cluster is read over the multiple repetitions.
[0091] The term “hybridize” is used herein according to its plain ordinary meaning and refers to the annealing of one single- stranded nucleic acid (such as a primer) to another nucleic acid based on the well-understood principle of sequence complementarity. In embodiments, the other nucleic acid is a single-stranded nucleic acid. The propensity for hybridization between nucleic acids depends on the temperature and ionic strength of their milieu, the length of the nucleic acids and the degree of complementarity. The effect of these parameters on hybridization is described in, for example, Sambrook J., Fritsch E. F., Maniatis T., Molecular cloning: a laboratory manual, Cold Spring Harbor Laboratory Press, New York(1989). As used herein, hybridization of a primer, or of a DNA extension product, respectively, is extendable by creation of a phosphodiester bond with an available nucleotide or nucleotide analogue capable of forming a phosphodiester bond, therewith. For example, hybridization can be performed at a temperature ranging from 15° C. to 95° C. In some embodiments, the hybridization is performed at a temperature of about 20° C., about 25° C., about 30° C., about 35° C., about 40° C., about 45° C., about 50° C., about 55° C., about 60° C., about 65° C., about 70° C., about 75° C., about 80° C., about 85° C., about 90° C., or about 95° C. In other embodiments, the stringency of the hybridization can be further altered by the addition or removal of components of the buffered solution. In some embodiments nucleic acids, or portions thereof, that are configured to hybridize are often about 80% or more, 81% or more, 82% or more, 83% or more, 84% or more, 85% or more, 86% or more, 87% or more, 88% or more, 89% or more, 90% or more, 91% or more, 92% or more, 93% or more, 94% or more, 95% or more, 96% or more, 97% or more, 98% or more, 99% or more or 100% complementary to each other over a contiguous portion of nucleic acid sequence. A specific hybridization discriminates over non-specific hybridization interactions (e.g., two nucleic acids that a not configured to specifically hybridize, e.g., two nucleic acids that are 80% or less, 70% or less, 60% or less or 50% or less complementary) by about 2-fold or more, often about 10-fold or more, and sometimes about 100-fold or more, 1000-fold or more, 10,000-fold or more, 100,000-fold or more, or 1,000,000-fold or more. Two nucleic acid strands that are hybridized to each other can form a duplex which comprises a doublestranded portion of nucleic acid.
[0092] The terms “dark cycle” and “limited-extension cycle” and “LE cycle” are used herein according to their plain ordinary meaning and refer to incorporating with a polymerase one or more nucleotides (e.g., native nucleotides) to the 3’ end of a polynucleotide under a set of conditions that are different from a sequencing cycle. In embodiments, during a dark cycle the identity of a nucleotide is not determined following incorporation of the nucleotide. In embodiments, the identity of one or more (but not all) nucleotides is optionally determined upon incorporation. In embodiments, during a dark cycle, a native nucleotide (e.g., dATP, dCTP, dTTP, or dGTP) is incorporated into a polynucleotide. Due to it being a native nucleotide having no reversible terminator moiety, the polymerase does not temporarily halt, and the incorporated nucleotide is not detected or identified, and polymerization continues. In embodiments, during a dark cycle a nucleotide analogue comprising a label (e.g., dATP*, dCTP*, dTTP*, or dGTP*, wherein ‘*’ indicates a labeled nucleotide) may be used and isincorporated into a polynucleotide. The identity of the incorporated nucleotide may he determined to ensure cluster synchronization. The native nucleotides may be any number of naturally occurring or modified nucleotides. In embodiments, the nucleotides include a reversible blocking group(i.e., a reversible terminator moiety). In embodiments, a dark cycle includes the incorporation of one or more nucleotides that are unidentified, and optionally one or more nucleotides that are identified.
[0093] The term “extension” or “elongation” is used herein according to their plain ordinary meanings and refer to synthesis by a polymerase of a new polynucleotide strand complementary to a template strand by adding free nucleotides (e.g., dNTPs) from a reaction mixture that are complementary to the template in the 5'-to-3' direction. Extension includes condensing the 5'-phosphate group of the dNTPs with the 3'-hydroxy group at the end of the nascent (elongating) DNA strand.
[0094] The term “sequencing read” is used herein according to its plain ordinary meaning and refers to an inferred sequence of base pairs (or base pair probabilities) corresponding to all or part of a single DNA fragment. Sequencing technologies vary in the length of reads produced. A sequencing read may include 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, or more nucleotide bases. Reads of length 20-40 base pairs (bp) are referred to as ultra- short. Typical sequencers produce read lengths in the range of 100-500 bp. Read length is a factor which can affect the results of biological studies. For example, longer read lengths improve the resolution of de novo genome assembly and detection of structural variants.
[0095] One of skill in the art will understand which standard controls are most appropriate in a given situation and be able to analyze data based on comparisons to standard control values. Standard controls are also valuable for determining the significance (e.g. statistical significance) of data. For example, if values for a given parameter are widely variant in standard controls, variation in test samples will not be considered as significant.
[0096] It is understood that the examples and embodiments described herein are for illustrative purposes only and that various modifications or changes in light thereof will be suggested to persons skilled in the art and are to be included within the spirit and purview of this application and scope of the appended claims. All publications, patents, and patent applications cited herein are hereby incorporated by reference in their entirety for all purposes.METHODS
[0097] The compositions and methods provided herein including embodiments thereof may be used, inter alia, to amplify a chromosomal fragment DNA sequence. The inventors developed a high-throughput, single-cell method for amplifying and analyzing chromosomal fragment DNA sequences. The methods provided herein including embodiments thereof may be used, inter alia, in chromatin conformation profiling of single cells from heterogenous tissue. The inventors found the methods provided herein can be used, inter alia, to detect copy number variations, structural variations, and extrachromosomal DNA. Furthermore, the compositions and methods provided herein including embodiments thereof can be used, inter alia, to jointly profile chromatin architecture and the transcriptome of single cells from heterogenous tissue. The methods provided herein can be used, inter alia, to jointly analyze chromatin structure and gene expression to understand the regulation of gene activation and / or repression. The methods provided herein including embodiments thereof, address critical gaps in chromatin analysis of heterogenous tissues and enhances understanding of gene regulation at a single cell level.
[0098] Thus, in an aspect is provided a method of amplifying a chromosomal fragment DNA sequence, the method including: (a) contacting a plurality of cell nuclei with a DNA crosslinking agent wherein each of the plurality of cell nuclei includes chromosomal DNA thereby forming crosslinked chromosomal DNA within each of the plurality of cell nuclei;(b) contacting the crosslinked chromosomal DNA with a DNA endonuclease thereby forming cleaved chromosomal DNA within each of the plurality of cell nuclei, wherein each of the crosslinked chromosomal DNAs includes a plurality of double stranded cleaved sites; (c) contacting the cleaved chromosomal DNA with a ligase thereby forming ligated chromosomal DNA within each of the plurality of cell nuclei, wherein each of the ligated chromosomal DNAs includes at least one ligated site that forms a non-endogenous ligation DNA sequence, wherein the non-endogenous ligation DNA sequence is not present in the endogenous chromosomal DNA within each of the plurality of cell nuclei; (d) contacting the ligated chromosomal DNA with a transposase, a 5' hybridization DNA sequence and a 3' hybridization DNA sequence thereby forming a plurality of transposed chromosomal DNA sequences within each of the plurality of cell nuclei, wherein each of the transposed chromosomal DNA sequences includes from 5' to 3’ a first DNA hybridization sequence, a chromosomal fragment DNA sequence, and a second DNA hybridization sequence; (e) separating each of the plurality of cell nuclei into a separate compartment, wherein eachcompartment contains one cell nucleus and an identifying library nucleic acid sequence, wherein the identifying library nucleic acid sequence includes from 5' to 3' a forward primer sequence, a barcode sequence, and a 5' complementary hybridization sequence; (f) lysing each cell nucleus in each compartment and allowing the first DNA hybridization sequence to hybridize with the 5' complementary hybridization sequence to form a first hybridized template DNA sequence; (g) contacting the first hybridized template DNA with a DNA polymerase under conditions conducive to amplification and producing a first amplified chromosomal DNA including from 5' to 3' the forward primer sequence or complement thereof, the barcode sequence or complement thereof, the 5’ complementary hybridization sequence or complement thereof, the chromosomal fragment DNA sequence or complement thereof, and the second DNA hybridization sequence or complement thereof, thereby amplifying the chromosomal fragment DNA sequence.
[0099] In embodiments, the first amplified chromosomal DNA includes from 5' to 3’ the forward primer sequence, the barcode sequence, the 5' complementary hybridization sequence, the complement of the chromosomal fragment DNA sequence, and the complement of the second DNA hybridization sequence. In embodiments, the first amplified chromosomal DNA includes from 5’ to 3' a complement of the forward primer sequence, a complement of the barcode sequence, a complement of the 5' complementary hybridization sequence, the chromosomal fragment DNA sequence, and the second DNA hybridization sequence.
[0100] In embodiments, the method further includes: (h) contacting the first amplified chromosomal DNA with a second library nucleic acid sequence including from 5' to 3' a 3’ complementary hybridization sequence and a reverse primer sequence and allowing the second DNA hybridization sequence to hybridize with the 3' complementary hybridization sequence thereby forming a second hybridized template DNA sequence; and (i) contacting the second hybridized template DNA sequence with a DNA polymerase under conditions conducive to amplification, thereby amplifying the first amplified chromosomal DNA to produce a second amplified chromosomal DNA including from 5’ to 3’ the forward primer sequence or complement thereof, the barcode sequence or complement thereof, the 5’ complementary hybridization sequence or complement thereof, the chromosomal fragment DNA sequence or complement thereof, the second DNA hybridization sequence or complement thereof, the 3' complementary hybridization sequence or complement thereof and the reverse primer sequence or complement thereof.
[0101] In embodiments, the second library acid sequence includes from 5' to 3' a 3' complementary hybridization sequence or a complement thereof and a reverse primer sequence or a complement thereof. In embodiments, the second library acid sequence includes from 5' to 3’ a 3' complementary hybridization sequence and a reverse primer sequence. In embodiments, the second library acid sequence includes from 5' to 3' a complement of a 3' complementary hybridization sequence and a complement of a reverse primer sequence.
[0102] In embodiments, each chromosomal fragment DNA sequence is from about 200 nucleotides to about 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 225 nucleotides to about 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 250 nucleotides to about 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 275 nucleotides to about 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 300 nucleotides to about 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 325 nucleotides to about 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 350 nucleotides to about 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 375 nucleotides to about 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 400 nucleotides to about 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 425 nucleotides to about 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 450 nucleotides to about 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 475 nucleotides to about 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 500 nucleotides to about 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 525 nucleotides to about 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 550 nucleotides to about 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 575 nucleotides to about 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 600 nucleotides to about 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is fromabout 625 nucleotides to about 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 650 nucleotides to about 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 675 nucleotides to about 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 700 nucleotides to about 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 725 nucleotides to about 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 750 nucleotides to about 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 775 nucleotides to about 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 800 nucleotides to about 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 825 nucleotides to about 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 850 nucleotides to about 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 875 nucleotides to about 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 900 nucleotides to about 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 925 nucleotides to about 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 950 nucleotides to about 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 975 nucleotides to about 1000 nucleotides in length.
[0103] In embodiments, each chromosomal fragment DNA sequence is from about 200 nucleotides to about 975 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 200 nucleotides to about 950 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 200 nucleotides to about 925 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 200 nucleotides to about 900 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 200 nucleotides to about 875 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 200 nucleotides to about 850 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 200 nucleotides to about 825 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 200 nucleotides toabout 800 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 200 nucleotides to about 775 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 200 nucleotides to about 750 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 200 nucleotides to about 725 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 200 nucleotides to about 700 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 200 nucleotides to about 675 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 200 nucleotides to about 650 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 200 nucleotides to about 625 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 200 nucleotides to about 600 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 200 nucleotides to about 575 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 200 nucleotides to about 550 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 200 nucleotides to about 525 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 200 nucleotides to about 500 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 200 nucleotides to about 475 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 200 nucleotides to about 450 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 200 nucleotides to about 725 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 200 nucleotides to about 400 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 200 nucleotides to about 975 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 200 nucleotides to about 375 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 200 nucleotides to about 350 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 200 nucleotides to about 325 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 200 nucleotides to about 300 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 200 nucleotides to about 275 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is fromabout 200 nucleotides to about 250 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from about 200 nucleotides to about 225 nucleotides in length.
[0104] In embodiments, each chromosomal fragment DNA sequence is from 200 nucleotides to 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 225 nucleotides to 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 250 nucleotides to 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 275 nucleotides to 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 300 nucleotides to 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 325 nucleotides to 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 350 nucleotides to 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 375 nucleotides to 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 400 nucleotides to 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 425 nucleotides to 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 450 nucleotides to 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 475 nucleotides to 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 500 nucleotides to 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 525 nucleotides to 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 550 nucleotides to 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 575 nucleotides to 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 600 nucleotides to 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 625 nucleotides to 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 650 nucleotides to 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 675 nucleotides to 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 700 nucleotides to 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 725 nucleotides to 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 750 nucleotides to 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 775 nucleotides to 1000 nucleotides in length.In embodiments, each chromosomal fragment DNA sequence is from 800 nucleotides to 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 825 nucleotides to 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 850 nucleotides to 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 875 nucleotides to 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 900 nucleotides to 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 925 nucleotides to 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 950 nucleotides to 1000 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 975 nucleotides to 1000 nucleotides in length.
[0105] In embodiments, each chromosomal fragment DNA sequence is from 200 nucleotides to 975 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 200 nucleotides to 950 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 200 nucleotides to 925 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 200 nucleotides to 900 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 200 nucleotides to 875 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 200 nucleotides to 850 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 200 nucleotides to 825 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 200 nucleotides to 800 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 200 nucleotides to 775 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 200 nucleotides to 750 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 200 nucleotides to 725 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 200 nucleotides to 700 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 200 nucleotides to 675 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 200 nucleotides to 650 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 200 nucleotides to 625 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 200 nucleotides to 600 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 200 nucleotides to 575 nucleotides in length. In embodiments, each chromosomal fragmentDNA sequence is from 200 nucleotides to 550 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 200 nucleotides to 525 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 200 nucleotides to 500 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 200 nucleotides to 475 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 200 nucleotides to 450 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 200 nucleotides to 725 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 200 nucleotides to 400 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 200 nucleotides to 975 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 200 nucleotides to 375 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 200 nucleotides to 350 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 200 nucleotides to 325 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 200 nucleotides to 300 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 200 nucleotides to 275 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 200 nucleotides to 250 nucleotides in length. In embodiments, each chromosomal fragment DNA sequence is from 200 nucleotides to 225 nucleotides in length.
[0106] In embodiments, the plurality of cell nuclei are imaged following step (a), step (b), step (c) and / or step (d) to determine if each cell nuclei is intact (e.g., not lysed). In embodiments, the plurality of cell nuclei are imaged following step (a), step (b), step (c) and step (d) to determine if each cell nuclei is intact (e.g., not lysed). In embodiments, the plurality of cell nuclei are imaged following step (a), step (b), step (c) or step (d) to determine if each cell nuclei is intact (e.g., not lysed). In embodiments, the plurality of cell nuclei are imaged following step (a) to determine if each cell nuclei is intact (e.g., not lysed). In embodiments, the plurality of cell nuclei are imaged following step (b) to determine if each cell nuclei is intact (e.g., not lysed). In embodiments, the plurality of cell nuclei are imaged following step (c) to determine if each cell nuclei is intact (e.g., not lysed). In embodiments, the plurality of cell nuclei are imaged following step (d) to determine if each cell nuclei is intact (e.g., not lysed).
[0107] In embodiments, each separate compartment is a droplet or a liquid volume within a well of a microwell plate. In embodiments, each separate compartment is a droplet. Inembodiments, each separate compartment is a liquid volume within a well of a microwell plate.
[0108] In embodiments, each separate compartment includes one cell nucleus and an identifying library nucleic acid sequence. In embodiments, each separate compartment includes only one cell nucleus and an identifying library nucleic acid sequence. In embodiments, each separate compartment includes no more than one cell nucleus and an identifying library nucleic acid sequence. In embodiments, each separate compartment includes a single cell nucleus and an identifying library nucleic acid sequence. In embodiments, each separate compartment includes only a single cell nucleus and an identifying library nucleic acid sequence. In embodiments, each separate compartment includes no more than a single cell nucleus and an identifying library nucleic acid sequence.
[0109] In embodiments, the identifying library nucleic acid sequence is bound to a solid support. In embodiments, the solid support is the surface of a well, the surface of a microwell, or the surface of a solid phase particle. In embodiments, the solid support is the surface of a well. In embodiments, the solid support is the surface of a microwell. In embodiments, the solid support is the surface of a solid phase particle. In embodiments, the solid phase particle is a gel bead or a polymer bead. In embodiments, the solid phase particle is a gel bead. In embodiments, the solid phase particle is a polymer bead.
[0110] In embodiments, the barcode sequence in each separate compartment is unique.
[0111] In embodiments, the DNA endonuclease is a DpnII restriction endonuclease, an Mbol restriction endonuclease, an Nlalll restriction endonuclease, a Ddel restriction endonuclease, or a DpnII restriction endonuclease. In embodiments, the DNA endonuclease is a DpnII restriction endonuclease. In embodiments, the DNA endonuclease is an Mbol restriction endonuclease. In embodiments, the DNA endonuclease is an Nlalll restriction endonuclease. In embodiments, the DNA endonuclease is a Ddel restriction endonuclease. In embodiments, the DNA endonuclease is a DpnII restriction endonuclease.
[0112] In embodiments, step (a) further includes contacting the crosslinked chromosomal DNA with from 1 to 5 additional DNA endonucleases. In embodiments, step (a) further includes contacting the crosslinked chromosomal DNA with from 2 to 5 additional DNA endonucleases. In embodiments, step (a) further includes contacting the crosslinked chromosomal DNA with from 3 to 5 additional DNA endonucleases. In embodiments, step (a) further includes contacting the crosslinked chromosomal DNA with from 4 to 5 additional DNA endonucleases.
[0113] In embodiments, step (a) further includes contacting the crosslinked chromosomal DNA with from 1 to 4 additional DNA endonucleases. In embodiments, step (a) further includes contacting the crosslinked chromosomal DNA with from 1 to 3 additional DNA endonucleases. In embodiments, step (a) further includes contacting the crosslinked chromosomal DNA with from 1 to 2 additional DNA endonucleases.
[0114] In embodiments, the additional DNA endonucleases are the same DNA endonuclease or different DNA endonucleases. In embodiments, the additional DNA endonucleases are the same DNA endonuclease. In embodiments, the additional DNA endonucleases are different DNA endonucleases.
[0115] In embodiments, the additional DNA endonucleases include a DpnII restriction endonuclease, an Mbol restriction endonuclease, an Nlalll restriction endonuclease, a Ddel restriction endonuclease, or a DpnII restriction endonuclease. In embodiments, the additional DNA endonuclease is a DpnII restriction endonuclease, an Mbol restriction endonuclease, an Nlalll restriction endonuclease, a Ddel restriction endonuclease, or a DpnII restriction endonuclease. In embodiments, the additional DNA endonuclease is a DpnII restriction endonuclease. In embodiments, the additional DNA endonuclease is an Mbol restriction endonuclease. In embodiments, the additional DNA endonuclease is an Nlalll restriction endonuclease. In embodiments, the additional DNA endonuclease is a Ddel restriction endonuclease, or a DpnII restriction endonuclease.
[0116] In embodiments, the ligase is a T4 DNA ligase.
[0117] In embodiments, the transposase is a Tn5 transposase.
[0118] In embodiments, the method further includes amplifying a plurality of different chromosomal fragment DNA sequences from each of the plurality of cell nuclei, thereby producing a plurality of amplified chromosomal fragment DNA sequences from each of the plurality of cell nuclei.
[0119] In embodiments, the method further includes sequencing the plurality of amplified chromosomal fragment DNA sequences. In embodiments, the sequencing is a next-generation sequencing method. In embodiments, the sequencing is a sequencing-by-synthesis (SBS) method.
[0120] In embodiments, the method further includes: (z) aligning the sequence of each of the amplified chromosomal fragment DNA sequences to a reference genome; and (zz) determining a relative abundance of each of the amplified chromosomal fragment DNAsequences including non -endogenous ligation DNA sequences within the amplified chromosomal fragment DNA sequences.
[0121] In embodiments, the method further includes aligning the sequence of each of the amplified chromosomal fragment DNA sequences to a reference genome. In embodiments, the method further includes determining a relative abundance of each of the amplified chromosomal fragment DNA sequences including non-endogenous ligation DNA sequences within the amplified chromosomal fragment DNA sequences. In embodiments, the method further includes determining a relative abundance of each of the amplified chromosomal fragment DNA sequences within the amplified chromosomal fragment DNA sequences.
[0122] In embodiments, the aligning and determining are computer implemented.
[0123] In embodiments, the method further includes computing a three-dimensional structure of the chromatin of the plurality of cell nuclei based on the aligning and the determining.
[0124] In embodiments, the method further includes detecting a copy number variation within the chromosomal DNA of each of the plurality of cell nuclei, a structural variation within the chromosomal DNA of each of the plurality of cell nuclei, and / or extrachromosomal DNA within the chromosomal DNA of each of the plurality of cell nuclei based on the aligning and determining. In embodiments, the method further includes detecting a copy number variation within the chromosomal DNA of each of the plurality of cell nuclei, a structural variation within the chromosomal DNA of each of the plurality of cell nuclei, and extrachromosomal DNA within the chromosomal DNA of each of the plurality of cell nuclei based on the aligning and determining.
[0125] In embodiments, the method further includes detecting a copy number variation within the chromosomal DNA of each of the plurality of cell nuclei, a structural variation within the chromosomal DNA of each of the plurality of cell nuclei, or extrachromosomal DNA within the chromosomal DNA of each of the plurality of cell nuclei based on the aligning and determining. In embodiments, the method further includes detecting a copy number variation within the chromosomal DNA of each of the plurality of cell nuclei based on the aligning and determining. In embodiments, the method further includes detecting a structural variation within the chromosomal DNA of each of the plurality of cell nuclei based on the aligning and determining. In embodiments, the method further includes detecting extrachromosomal DNA within the chromosomal DNA of each of the plurality of cell nuclei based on the aligning and determining.
[0126] In embodiments, the method further includes detecting extrachromosomal DNA in close proximity to chromosomal DNA within each of the plurality of cell nuclei. In embodiments, the method further includes detecting interchromosomal contacts and / or intrachromosomal contacts within each of the plurality of cell nuclei. In embodiments, the method further includes detecting interchromosomal contacts and intrachromosomal contacts within each of the plurality of cell nuclei. In embodiments, the method further includes detecting interchromosomal contacts and / or intrachromosomal contacts within each of the plurality of cell nuclei. In embodiments, the method further includes detecting interchromosomal contacts within each of the plurality of cell nuclei. In embodiments, the method further includes detecting intrachromosomal contacts within each of the plurality of cell nuclei. In embodiments, the method further includes detecting an evolutionary shift in a population of cells.
[0127] In embodiments, the method can differentiate between extrachromosomal DNA and a homogenously staining region (HSR).
[0128] In embodiments, the method is performed before, during, or after administration of a therapeutic agent. In embodiments, the method is performed before administration of a therapeutic agent. In embodiments, the method is performed during administration of a therapeutic agent. In embodiments, the method is performed after administration of a therapeutic agent. In embodiments, the method is performed before, during, or after administration of an anti-cancer agent. In embodiments, the method is performed before administration of an anti-cancer agent. In embodiments, the method is performed during administration of an anti-cancer agent. In embodiments, the method is performed after administration of an anti-cancer agent.
[0129] In embodiments, the method further includes detecting a copy number variation within the chromosomal DNA of each of the plurality of cell nuclei, a structural variation within the chromosomal DNA of each of the plurality of cell nuclei, or extrachromosomal DNA within the chromosomal DNA of each of the plurality of cell nuclei in response to an exogenous stimulus (e.g. administration of a therapeutic agent, an anti-cancer agent, etc.).
[0130] In embodiments, the method further includes detecting a copy number variation within the chromosomal DNA of each of the plurality of cell nuclei, a structural variation within the chromosomal DNA of each of the plurality of cell nuclei, or extrachromosomal DNA within the chromosomal DNA of each of the plurality of cell nuclei in response to administration of a therapeutic agent. In embodiments, the method further includes detectinga copy number variation within the chromosomal DNA of each of the plurality of cell nuclei in response to administration of a therapeutic agent. In embodiments, the method further includes detecting a structural variation within the chromosomal DNA of each of the plurality of cell nuclei in response to administration of a therapeutic agent. In embodiments, the method further includes detecting extrachromosomal DNA within the chromosomal DNA of each of the plurality of cell nuclei in response to administration of a therapeutic agent.
[0131] In embodiments, the method further includes detecting a copy number variation within the chromosomal DNA of each of the plurality of cell nuclei, a structural variation within the chromosomal DNA of each of the plurality of cell nuclei, or extrachromosomal DNA within the chromosomal DNA of each of the plurality of cell nuclei in response to administration of an anti-cancer agent. In embodiments, the method further includes detecting a copy number variation within the chromosomal DNA of each of the plurality of cell nuclei in response to administration of an anti-cancer agent. In embodiments, the method further includes detecting a structural variation within the chromosomal DNA of each of the plurality of cell nuclei in response to administration of an anti-cancer agent. In embodiments, the method further includes detecting extrachromosomal DNA within the chromosomal DNA of each of the plurality of cell nuclei in response to administration of an anti-cancer agent.
[0132] In embodiments, each of the compartments of step (e) further include a plurality of endogenous cellular mRNA, a plurality of messenger RNA (mRNA) capture sequences and a reverse transcriptase, wherein each mRNA capture sequence includes a second barcode sequence and a poly-adenine (poly- A) tail complementary sequence; wherein the method further includes allowing the plurality of poly-A tail complementary sequences to hybridize with a plurality of the endogenous mRNAs to form a plurality of hybridized template mRNA sequences; and allowing the reverse transcriptase to contact the plurality of hybridized template mRNA sequences under conditions conducive to reverse transcription, thereby forming a plurality of barcoded transcript sequences.
[0133] In embodiments, each of the separate compartments of step (e) further include
[0134] In embodiments, the method further includes detecting the barcoded transcript sequences. In embodiments, the method further comprising includes quantifying barcoded transcript sequences. In embodiments, the method further includes detecting the barcoded transcript sequences and quantifying barcoded transcript sequences. In embodiments, the method further includes analyzing the quantified barcoded transcript sequences. In embodiments, the analyzing of the quantified barcoded transcript sequences is computerimplemented. In embodiments, the method further includes computing a gene expression pattern of each cell nucleus. In embodiments, the method further includes computing a gene expression profile of each cell nucleus.
[0135] In embodiments, the method further includes analyzing the gene expression of the plurality of cell nuclei based on the quantifying of barcoded transcript sequences; and analyzing the three dimensional structure of the chromatin of the plurality of cell nuclei based on the aligning and determining, thereby generating regulation of gene activation and / or gene repression data.
[0136] In embodiments, the chromosomal DNA fragment sequence includes extrachromosomal DNA (ecDNA).
[0137] In embodiments, the plurality of cell nuclei include a plurality of cell nuclei from a heterogenous population of cells.
[0138] In embodiments, the heterogenous population of cells includes a cancer cell.
[0139] In embodiments, the heterogenous population of cells is a heterogenous population of human cells. In embodiments, the heterogenous population of cells is a heterogenous population of mammalian cells.
[0140] In embodiments, the plurality of cell nuclei are isolated from cancer cells or wherein the plurality of cell nuclei are within cancer cells prior to lysing. In embodiments, the plurality of cell nuclei are isolated from cancer cells. In embodiments, the plurality of cell nuclei are within cancer cells prior to lysing. In embodiments, the cancer cells are isolated from a patient tumor sample.
[0141] In another aspect is provided a method of cell profiling the method including: (a) providing or obtaining a cross-linked chromatin complex of a cell or a cellular nucleus; (b) removing one or more histone proteins from the cross-linked chromatin complex; (c) digesting the cross-linked chromatin complex or portions thereof after (b), thereby generating a digested chromatin; (d) ligating the digested chromatin, thereby generating a ligated chromatin; (e) fragmenting the ligated chromatin, thereby generating a fragmented chromatin; (f) encapsulating the fragmented chromatin in a droplet; (g) barcoding the fragmented chromatin with a barcode, thereby generating a barcoded fragmented chromatin; and (h) sequencing the barcoded fragmented chromatin.
[0142] In embodiments, the method further includes determining a spatial proximity between fibers of the cross-linked chromatin.
[0143] In embodiments, the spatial proximity is determined across the genome of the cell or the cellular nucleus. In embodiments, the spatial proximity is determined across the genome of the cell. In embodiments, the spatial proximity is determined across the cellular nucleus.
[0144] In embodiments, the barcode is specific to the cell or the cellular nucleus. In embodiments, the barcode is specific to the cell. In embodiments, the barcode is specific to the cellular nucleus.
[0145] In embodiments, the barcoding includes providing a barcode bound to a bead in the droplet and capturing the fragmented chromatin on the barcode.
[0146] In embodiments, the bead is degradable upon application of a stimulus.
[0147] In embodiments, the droplet includes a GEM® by 10X GENOMICS™.
[0148] In embodiments, the droplet is generated with the aid of a droplet microfluidic device.
[0149] In embodiments, (f) and (g) are performed using a platform sold under trademark CHROMIUM NEXT GEM CHIP H® by 10X GENOMICS™.
[0150] In embodiments, the method includes performing single-cell chromatin conformation profiling on a population of cells including the cell or the cellular nucleus. In embodiments, the method includes performing single-cell chromatin conformation profiling on a population of cells including the cell. In embodiments, the method includes performing single-cell chromatin conformation profiling on a population of cells including the cellular nucleus.
[0151] In embodiments, a chromatin architecture of the cell or the cellular nucleus is mapped on a single cell level in the population of cells including the cell or the cellular nucleus. In embodiments, a chromatin architecture of the cell or the cellular nucleus is mapped on a single cell level in the population of cells including the cell. In embodiments, a chromatin architecture of the cell or the cellular nucleus is mapped on a single cell level in the population of cells including the cellular nucleus.
[0152] In embodiments, the cell or the cellular nucleus is derived from a tissue of a subject, and wherein the method further includes deriving and analyzing a population of cells from the tissue.
[0153] In embodiments, the method further includes diagnosing the subject with a condition.
[0154] In embodiments, the method further includes prescribing a medicine, an action, or a combination thereof to the subject, based on results of the cell profiling. In embodiments, the method further includes prescribing a medicine to the subject, based on results of the cell profiling. In embodiments, the method further includes prescribing an action to the subject, based on results of the cell profiling.
[0155] In embodiments, the method further includes detecting a copy number variation (CNV), a structural variation (SV), an extrachromosomal DNA (ecDNA), or any combination thereof in the population of cells. In embodiments, the method further includes detecting a copy number variation (CNV) in the population of cells. In embodiments, the method further includes detecting a structural variation (SV) in the population of cells. In embodiments, the method further includes detecting an extrachromosomal DNA (ecDNA), in the population of cells.
[0156] In embodiments, the method further includes revealing clonal dynamics or an oncogenic event in the population of cells. In embodiments, the method further includes revealing clonal dynamics in the population of cells. In embodiments, the method further includes revealing an oncogenic event in the population of cells.
[0157] In embodiments, the method further includes analyzing the transcriptome of the cell or the cellular nucleus. In embodiments, the method further includes analyzing the transcriptome of the cell. In embodiments, the method further includes analyzing the transcriptome of the cellular nucleus.
[0158] In embodiments, the method further includes identifying an architecture of the cross-linked chromatin complex, obtaining a transcriptomic dataset from the cell or the cellular nucleus, and identifying a correlation between the architecture and the transcriptomic dataset. In embodiments, the method further includes identifying an architecture of the crosslinked chromatin complex. In embodiments, the method further includes obtaining a transcriptomic dataset from the cell or the cellular nucleus. In embodiments, the method further includes identifying a correlation between the architecture and the transcriptomic dataset.
[0159] In embodiments, the method further includes recording one or more data sets from the cell screening.
[0160] In embodiments, the method further includes training a computational model using the one or more datasets.
[0161] In embodiments, the computational model includes a Machine Learning (ML) model, an Artificial Intelligence (Al) model, Deep Learning, or any combination thereof. In embodiments, the computational model includes a Machine Learning (ML) model. In embodiments, the computational model includes an Artificial Intelligence (Al) model. In embodiments, the computational model includes Deep Learning.
[0162] In embodiments, the method further includes fixating the cell or the cellular nucleus prior to (a).
[0163] In embodiments, the method includes single cell or single nucleus sequencing. In embodiments, the method includes single cell sequencing. In embodiments, the method includes single nucleus sequencing.
[0164] In embodiments, the method further includes constructing a single cell or a single nucleus library. In embodiments, the method further includes constructing a single cell library. In embodiments, the method further includes constructing a single nucleus library.
[0165] In embodiments, the method further includes analysis of chromatin domains, analysis of chromatin loops, or both. In embodiments, the method further includes analysis of chromatin domains and analysis of chromatin loops. In embodiments, the method further includes analysis of chromatin domains. In embodiments, the method further includes analysis of chromatin loops.
[0166] In embodiments, the method further includes identifying structural variations (SV) in the cross-linked chromatin complex.
[0167] In embodiments, the method further includes single cell or single nucleus genotyping. In embodiments, the method further includes single cell genotyping. In embodiments, the method further includes single nucleus genotyping.
[0168] In embodiments, the method further includes analysis of gene expression levels and copy numbers of ecDNA in the cell or the cellular nucleus. In embodiments, the method further includes analysis of gene expression levels and copy numbers of ecDNA in the cell. In embodiments, the method further includes analysis of gene expression levels and copy numbers of ecDNA in the cellular nucleus. In embodiments, the method further includes analysis of gene expression levels in the cell or the cellular nucleus. In embodiments, the method further includes analysis of gene expression levels in the cell. In embodiments, the method further includes analysis of gene expression levels in the cellular nucleus. In embodiments, the method further includes analysis of copy numbers of ecDNA in the cell or the cellular nucleus. In embodiments, the method further includes analysis of copy numbersof ecDNA in the cell. In embodiments, the method further includes analysis of copy numbers of ecDNA in the cellular nucleus.EXAMPLESExample 1: Abstract
[0169] Here, we introduce Droplet Hi-C, which employs a commercial microfluidic device for high-throughput, single-cell chromatin conformation profiling in droplets. Using Droplet Hi-C, we mapped the chromatin architecture at single-cell resolution from the mouse cortex and analyzed gene regulatory programs in major cortical cell types. Additionally, we used this technique to detect copy number variation (CNV), structural variations (S Vs) and extrachromosomal DNA (ecDNA) in cancer cells, revealing clonal dynamics and other oncogenic events during treatment. We further refined this technique to allow for joint profiling of chromatin architecture and transcriptome in single cells, facilitating a more comprehensive exploration of the links between chromatin architecture and gene expression in both normal tissues and tumors. Thus, Droplet Hi-C not only addresses critical gaps in chromatin analysis of heterogeneous tissues but also emerges as a versatile tool enhancing our understanding of gene regulation in health and disease.Example 2: Introduction
[0170] In eukaryotic cells, chromosomes are organized within the nucleus in a non-random manner1-4. Each chromosome is folded into a complex and dynamic structure, forming chromatin loops5, topologically associating domains6 7, and A / B compartments8-10. Chromatin organization plays a crucial role in both embryonic development11, 12and the progression of various diseases13, 14. During development, the chromatin structure influences which genes are turned on or off by the distal regulatory elements15-18. Abnormal chromatin organization can lead to aberrant gene expression or silencing, contributing to cancer development and progression14. Indeed, components of the nuclear machinery that regulates chromatin organization are among the most frequently mutated genes in human cancers19-22. Thus, comprehensive analysis of the chromatin organization is crucial for understanding the gene regulatory programs involved in both normal development and disease pathogenesis.
[0171] Analyzing chromatin organization in primary tissues and tumor biopsies presents unique challenges due to the complexity of chromatin dynamics and the technical limitations of current methodologies. First, the biospecimens available for assays are often heterogeneous, containing a mixture of different cell types. Existing techniques such as insitu Hi-C for analysis of 3D chromatin organization typically require bulk tissues as input and could not resolve specific chromatin organization patterns relevant to particular cell types, especially in tumors where cancer cells coexist with stromal and immune cells. Second, the resulting data obtained from bulk chromatin organization assays are complex, and relating changes in chromatin organization to specific functional outcomes can be difficult.Addressing these challenges requires the development of more sensitive single-cell chromatin analysis, better methods for handling tissue samples, and advanced computational tools for data analysis.
[0172] Great strides have been made in recent years in single-cell Hi-C technologies1-4. In general, single-cell Hi-C methods can be categorized into low-throughput microwell-based approaches and high-throughput combinatorial indexing-based approaches. In the case of microwell-based single-cell Hi-C methods, cells / nuclei are individually dispensed into microwells, and library construction is carried out in each microwell in parallel, frequently with the help of automated liquid handlers. Notably, several single-cell Hi-C methods have been developed, such as those by Nagano et al. (2013)23, Nagano et al. (2017)24, Stevens et al. (2017)25, Flyamer et al. (2017)26, Tan et al. (2018)27and third-generation sequencing (TGS) based scNanoHi-C28. Recently, additional micro well-based single-cell Hi-C methods have emerged that combine Hi-C analysis with assays for one or more additional molecular modalities in a single cell. For example, Methyl-HiC29and single-nucleus methyl-3C sequencing (sn-m3C-seq)30capture chromatin interactions and DNA methylation patterns from the same cell. HiRES31 jointly performs Hi-C and RNA sequencing to explore the functional relationship between 3D genome organization and transcriptome dynamics. In general, microwell-based techniques are limited in scale, and difficult to adopt broadly due to lengthy procedures and high cost. On the other hand, combinatorial indexing-based singlecell Hi-C methods utilize combinatorial barcoding strategies to achieve high throughput and scalability. Previously, a single-cell combinatorial indexed Hi-C (sciHi-C)32method allowed the generation of single-cell Hi-C libraries from a few thousand cells, albeit with limited genomic coverage in each cell. Building upon the same strategy, GAGE-seq (genome architecture and gene expression by sequencing)33was developed to simultaneously profile chromatin interactions and geneexpression in single cells, to achieve high throughput, multimodality, and high coverage per cell. However, the lengthy and largely manual combinatorial indexing procedure still poses a challenge for its general adoption.
[0173] Here, we introduce a highly scalable, generally accessible, droplet-based single-cell Hi-C method, Droplet Hi-C, which combines an in situ chromosomal conformation capture (3C) assay with commercially available droplet microfluidics, to simultaneously capture the 3D genome structure from tens of thousands of individual cells in a single experiment. We demonstrate the utility of Droplet Hi-C data in resolving cell type-specific chromatin architecture in complex tissues such as the mouse brain. We further show that Droplet Hi-C can be used to identify aberrant chromatin structure in cancer cells. Most excitingly, we used Droplet Hi-C to identify extra-chromosomal DNA (ecDNA) and map their chromatin interactome in tumor cells at single-cell resolution. Finally, we extended this method to enable the simultaneous capture of transcriptome and chromatin architecture in single cells. Example 3: Results
[0174] Development of Droplet Hi-C
[0175] Droplet Hi-C builds upon the in situ Hi-C procedure5. It captures spatial proximity genome-wide between chromatin fibers in formaldehyde-crosslinked cells / nuclei through restriction digestion and ligation in situ (FIG. 1A). After SDS treatment to remove histone proteins, DNA fragmentation and capture is then carried out in a commercially available microfluidic platform (i.e., lOx Genomics single cell ATAC kit), where cell-specific DNA barcodes are added to the DNA fragments. This is followed by sequencing library construction and next-generation sequencing (NGS) (FIG. 1 A). The whole procedure lasts about 10 hours from fixed cells / nuclei to final sequencing libraries, and 8 samples can be processed in parallel, enabling profiling of 40,000 or more cells simultaneously (FIGS. 1A- 1C). The throughput of Droplet Hi-C surpasses plate -based single-cell Hi-C methods by an order of magnitude, offering the shortest experimental duration and hands-on time (FIG. IB). Given the widespread use of commercial microfluidic systems, ease of the experimental procedure, and relatively low cost (FIG. 1C), Droplet Hi-C has the potential to be rapidly adopted by the research community.
[0176] To test the performance of Droplet Hi-C, we first applied it to a mixture of human HeLa S3 cell line and mouse embryonic stem cell (mESC) line. Equal number of cells from two cell lines were combined after cross-linking and subjected to Droplet Hi-C. The results showed that 1,773 human and 3,489 mouse high-quality cells (with total contacts number > 1,000) were recovered after shallow sequencing, along with 284 potential doublets (FIG. 7A). We further carried out Droplet Hi-C with a mixture of three human cell lines (K562, GM12878 and HeLa S3). After filtering, we obtained 3,668 high-quality cells with a medianof 73,813 unique contacts in each cell (duplication rate 19.0%), including 1 ,780 HeLa cells (median 66,828 contacts), 1,209 GM12878 cells (median 76,965 contacts) and also 679 K562 cells (median 77,415 contacts). Using Higashi34for cell clustering, three distinct populations can be clearly distinguished (FIG. 7B). By performing genome wide correlation analysis with bulk in-situ Hi-C references, each population can be confidentially assigned to one cell type (FIGS 7C-7H). On the chromosome level, cell type-specific chromatin structures can be observed, such as the distinct Hi-C pattern of chromosomal rearrangements on chromosome 11 in HeLa S3 cells (FIG. 7E). When focusing on a region encompassing the K562-specific upregulated gene LGR435, we observed an active A compartment at the LGR4 genomic locus that is specifically shown in the K562 cells (FIG. 7G). The differences in compartment scores as well as insulation scores among different cell lines can also be faithfully reproduced as in the published datasets (FIGS. 7G-7H). In summary, these results demonstrate that Droplet Hi-C can accurately distinguish cell type-specific chromatin organization.
[0177] Droplet Hi-C reveals cell type-specific chromatin structure in mouse cortex
[0178] To demonstrate the feasibility of applying Droplet Hi-C to primary heterogeneous tissues, we performed this assay using the cortex dissected from 8-week-old mice. In total, we generated 6,235 high-quality single-cell profiles of chromatin architecture from two replicates, with a median of 175,021 unique contacts per cell (with 22,456 cis-long contacts (> 1 kb) and 9,954 trans contacts per cell, at 58% duplication rate) (FIGS. 8A-8C). On the single-cell level, despite the sparse contact signals, we can still discern an enrichment of intra-chromosomal interactions (FIG. 8D). To benchmark the performance of Droplet Hi-C, we compared our data with previous single-cell Hi-C methods, Dip-C36and sn-m3C-seq37, data from the mouse cortex. The decay curve of contact probability by genomic distance showed a similar trend among these assays, with sn-m3C-seq profiles showing a slightly lower fraction of extreme long-range interactions than Droplet Hi-C and Dip-C (FIG. 8E). Droplet Hi-C yields a higher fraction of cis-long (> 1 kb) interactions than Dip-C, and lower than sn-m3C-seq (FIG. 8F). Among these methods, similar cis long-trans ratios 155 in chromatin contacts were detected (FIG. 8F).
[0179] To verify the fidelity of Droplet Hi-C profiles, we examined chromatin structures on contact heatmaps spanning multiple resolutions, and found them to be nearly identical between Droplet Hi-C and sn-m3C-seq (FIG. 8G). At the chromatin compartment level, although Droplet Hi-C used Tn5 transposase for chromatin fragmentation in nucleus, bias in interactions within the active A or inactive B compartments is not observed (FIG. 8H). Thecompartment scores and insulation scores from Droplet Hi-C data are highly correlated with other two single-cell Hi-C methods (FIGS. 8I-8L). Overall, these results indicated that Droplet Hi-C can reliably and robustly detect chromatin features in complex tissue.
[0180] Chromatin interactions within the gene body have been shown to be positively associated with gene expression levels, and therefore were used to represent features of single-cell Hi-C data at the gene level38. To annotate cell types, we first summarized the chromatin contacts on the gene body region of each gene as gene associating domain (GAD) scores39, and then generated a cell-by-gene GAD score feature matrix. We then leveraged single-nuclei RNA-seq data from previous Droplet Paired-Tag40on mouse cortex to coembed and annotate our Droplet Hi-C data (FIG. 9A). We successfully resolved 20 cell types, including 5 non-neuronal types, 9 glutamatergic neuronal types and 6 GABAergic neuronal types (FIG. ID). The cell type prediction results show an overall high prediction score (FIGS. 9B-9C). The cell type identities were confirmed based on both high GAD scores on known cell type markers (FIG. 9D) and concordant annotations when compared to the published cortex sn-m3C-seq data (FIG. 9E). Among these cell types, neuronal cells exhibit greater similarity to each other than to non-neuronal cells (FIG. IE). The hierarchical chromatin organization features, such as compartments, domains, and chromatin loops, are distinctly discernible in the aggregated single-cell profiles across various cell types (FIGS. 1F-1H). For example, the glutamatergic neurons marker gene Satb2 displayed long-range interactions with the region nearby Gm32949 gene that are specific in glutamatergic neurons (ITL6GL, Cortex L6 excitatory neurons; PTGL, Cortex NP excitatory neurons), but not in GABAergic (VIPGA, MGE-derived neurogliaform cells Vip+; PVGA, MGE-derived neurogliaform cells Pvalb+) or non-neuronal cells (OGC, oligodendrocytes; OPC, oligodendrocytes precursor cells) (FIG. 1H). Accompanying with this glutamatergic neurons-specific contact, the Satb2 gene also displayed an elevated expression level as well as association with an active histone modification, H3K27ac, in ITL6GL and PTGL cells when analyzed in conjunction with our Droplet Paired-Tag data (FIG. 1H). By contrast, in other cell types, this gene remained inactive and was even silenced with the presence of repressive H3K27me3 signals.
[0181] To systematically probe the differences in 3D genome structure across different cell types and elucidate their relationships with chromatin state and gene expression, we first examined differences in A / B compartments at 100-kb resolution among cell types. These compartments were found to correspond with transcriptionally active and inactive chromatin states. In total, we identified 895 genomic regions displaying variable compartment scoresacross the brain cell types (FIG. 2A). When combined with single-cell histone modification profiles previously reported for the mouse cortical region, the compartment scores displayed a positive correlation with H3K27ac signals and a negative correlation with H3K27me3 signals across different cell types, suggesting that variable compartments are associated with dynamic chromatin states (FIG. 2B).
[0182] Chromatin domains orchestrate enhancer-promoter interactions and regulate cell type-specific gene activities18-36. We defined domain boundaries in each of the 20 cortical cell types, and subsequently summarized the boundaries across all the cells to obtain the boundary probability for each cell type. Notably, genes with highest expression variation among cell types were often found to reside within or near cell type-specific domain boundaries34. For example, for the Pdgfra gene, which plays a crucial role in controlling oligodendrocyte differentiation41, we found an OPC specific-chromatin domain boundary overlapping the gene body of Pdgfra (FIG. 2C). We identified and fetched genes within a 100-kb window centering on variable domain boundaries, and then checked corresponding single-nuclei RNA-seq datasets to classify such genes as constant genes, housekeeping genes and variable genes based on their expression levels across different cell types (FIG. 2D). The results showed that the genes with variable expression levels had higher positive correlation with boundary probability than the constant genes and housekeeping genes (FIG. 2E).
[0183] Finally, we examined chromatin loops within each cell type at 10-kb resolution. To check whether loops are associated with functional chromatin states, we compared both H3K27ac and H3K27me3 histone modification signals at loop anchors across all the cell types. Notably, the active chromatin mark H3K27ac showed a higher enrichment at loop anchors identified in the corresponding cell types than other cell types, while repressive chromatin mark H3K27me3 showed the opposite trend, and was more likely to be depleted at the corresponding cell type-specific loop anchors (FIG. 2F). To systematically assess the functions of genes located near cell type-specific loop anchors, we performed GREAT analysis42for VIPGA- and OGC- specific loop anchors. The VIPGA-specific loop anchors are near genes related to neurotransmitters and ion transport, consistent with the signal transduction functions in neurons, while OGC-specific loop anchors are near genes related to myelination and axon ensheathment (FIG. 2G). Taken together, the above analyses showed that Droplet Hi-C can accurately reveal cell type-specific chromatin organization in heterogeneous tissues.
[0184] Droplet Hi-C detects structural alterations and ecDNA in cancer cells
[0185] Cancer cells often exhibit large-scale genomic rearrangements, including copy number variation (CNV) as well as structural variations (SVs), which have been implicated in the initiation and progression of cancer43. In previous studies, Hi-C has been shown to detect these genomic variations in bulk tumor tissues with high accuracy and efficiency44’45. However, these bulk analyses could not resolve the heterogeneity and evolution of tumor populations in response to therapies. Droplet Hi-C overcomes this problem through detection of genomic variations in individual cancer cells. As a proof of principle, we performed Droplet Hi-C with two colorectal cancer cell lines, COLO320DM and COLO320HSR, derived from the same patient, both carrying similar copies of MYC amplicon. In COLO320DM, the MYC oncogene is located on ecDNA, while In COLO320HSR, it exists on chromosomes as tandem repeats, termed homogeneously staining regions (HSR). By inferring the copy number of each 1-Mb bin in single cells as well as aggregated pseudo-bulk profiles, we detected elevated copies of the MYC locus from both samples, as anticipated (FIG. 3A). We also observed fluctuations in the copy numbers across the entire genome (FIG. 3A). We used the inferred DNA copy number of each genomic region to correct bias in the contact matrices and then identified SVs. Our analysis revealed distinct SVs between these two cancer lines, and we showed one example of duplication event on chromosome 6 (FIG. 3B).
[0186] It has been difficult to distinguish whether the MYC exists on the ecDNA or as HSR from whole genome DNA sequencing alone. We ran AmpliconArchitect algorithm using whole-genome sequencing (WGS) data of COLO320HSR, and AmpliconArchitect classified the HSR as a circular amplicon. While the genome-wide interaction patterns have been proposed as potential indicators of ecDNA46, aggregate Hi-C profiles showed similar patterns of inter-chromosomal interactions between the amplified MYC locus and the rest of the genome in both COLO320DM and COLO320HSR (FIG. 3C). Although measurement like adjusted normalized inter-chromosomal interaction frequency (adjnTIF) has been proposed to determine ecDNA identity44, we found that adjnTIF was equally high at the 1-Mb bin containing MYC in both cell lines, indicating that this value is likely driven by abnormally high copy number rather than other specific property of ecDNA (FIG. 3C).
[0187] We therefore tried to discern distinctions between MYC ecDNA and HSR by analyzing their interactome obtained from the Droplet Hi-C dataset. In COLO320DM, relatively uniform interactions happened within ecDNA (FIG. 10A). In COLO320HSR, wefound two deletions, evidenced by significantly lower numbers of sequencing reads at those regions. Single-cell profiles confirmed that these two regions appeared uniformly across all the cells (FIG. 10A). Interestingly, we found that ecDNA interacts with other chromosomes more evenly, while HSR displayed a preference towards certain chromosomes (FIG. 3D). To quantitatively measure the evenness of inter-chromosomal contacts, we devised the hub index as the Gini coefficient of inter-chromosomal contacts for the 1 Mb MYC bin in each single cell. COLO320DM shows a significantly smaller hub index, indicating a more uniform transinteracting profile (FIG. 3E). We also quantified other parameters which can be obtained from single-cell Hi-C data, including inferred copy number (FIG. 3F) and the trans-to-cis contacting bin ratio (FIG. 3G). Both of these parameters exhibited significant differences between COLO320DM and COLO320HSR (FIGS. 3F-3G). COLO320DM cells had approximately 47% more copies than COLO320HSR cells at the MYC genomic bin, which aligns with previous research findings47(FIG. 3F). Additionally, COLO320DM displayed a higher trans-to-cis contacting ratio for the MYC bin compared to COLO320HSR (FIG. 3G). These results collectively suggest that MYC ecDNA exhibits distinct interaction patterns from MYC HSR at the single-cell resolution, making it a potential indicator for determining whether a genomic bin is likely to be ecDNA or not within an individual cell.
[0188] The measurement of trans-interaction evenness also paves the way to investigate whether multiple ecDNA tend to cluster into “hubs” in a single cell, since the agglomerated ecDNAs will show contacting preference against specific chromosomes compared to diffused ecDNAs. By shuffling the inter-chromosomal contacts across genome to generate a uniform distribution as reference for the 1Mb MYC bin in COLO320DM and COLO320HSR, we found that the hub index of inter-chromosomal contacts in COLO320HSR is significantly higher, demonstrating that this genomic bin showed stronger bias in contacts for HSR, consistent with the tandem repeat aggregation property of HSR (FIG. 10B). However, COLO320DM exhibits a slightly lower hub index compared to the random background, indicating that ecDNA does not have a strong nuclear localization bias in these cells (FIG. 10B), and that previously reported ecDNA hubs in COLO320DM are likely transient or infrequent features48.
[0189] The observed differences in contact profiles between ecDNA and HSR hint at potential ways to distinguish these aberrant chromosomal features from single-cell Hi-C data. To this end, we fit a multivariate logistic regression model using hub index, estimated copy number, trans-to-cis contacting bin ratio for ecDNA identification. This regression model is then used to predict ecDNA probability for each 1Mb bin that contains ecDNA among the entire cell population, as well as whether each 1Mb bin contains ecDNA in each single cell (FIG. 11A). Specifically, the MYC locus (chr8 :127Mb- 128Mb) in COLO320DM cells is treated as ground truth and in COLO320HSR cells as negative in the training dataset. On the validation dataset, the model reached a 0.99 specificity, a 0.97 precision, a 0.89 accuracy and a 0.71 recall. We first applied the regression model to 1,426 COLO320DM cells and 1,535 COLO320HSR cells, which identified two consecutive bins (chr8: 126Mb-128Mb) as ecDNA. In COLO320DM, 72.7% of cells were predicted to be ecDNA positive (FIG. 1 IB). In contrast, only 20.5% of cells were predicted to be ecDNA positive in COLO320HSR cells (FIG. 1 IB). We also analyzed the 3,423 cells obtained from the adult mouse cortex, and found that no bin contains ecDNA. Taken together, these results show that our logistic regression model based ecDNA caller algorithm can identify known ecDNA from COLO320DM, although the false positive rate remains a problem in the COLO320HSR samples.
[0190] To further improve the accuracy of our ecDNA caller, we developed a deep learning model based on convolutional neural networks (CNN) (FIG. 3H). The deep learning model utilizes the sliced contact matrix, comprising five consecutive 1Mb genomic bins as its input. At the heart of this process, a neural network is trained using these input variables to identify ecDNA within each central bin, providing predictions for the presence of ecDNA or HSR in 1Mb bins, both at the cell population and single-cell levels (FIG. 11C, Methods). Our deep learning model has demonstrated remarkable performance, achieving a recall rate of 0.80, while maintaining a 0.99 precision in ecDNA prediction, surpassing that of the logistic regression model. Furthermore, it achieves 0.99 specificity and 0.93 accuracy, as well as a faster processing speed (FIG. 1 ID). By applying the deep learning-based ecDNA caller to the COLO320DM and COLO320HSR validation dataset, we were able to identify MYC ecDNA in 96% of COLO320DM cells and HSR in 3.86% of COLO320DM cells (FIG. 31).Conversely, MYC HSR was identified in 0.20% of COLO320HSR cells, along with HSR in 99.54% of COLO320HSR cells (FIG. 31). Given that the deep learning-based ecDNA caller pipeline can directly distinguish between ecDNA and HSR, while also outperforming the regression model-based algorithm, we decided to employ the deep learning model-based ecDNA caller algorithm in all the subsequent studies.
[0191] Droplet Hi-C uncovers ecDNA heterogeneity and tumor cell evolution in response to drug treatment
[0192] Previous studies showed that glioblastoma (GBM) cells carrying extrachromosomal oncogenic variant of epidermal growth factor receptor (EGFRvIII) can develop drug resistance after treatment with tyrosine kinase inhibitors such as erlotinib (ER), which coincides with the loss of EGFR ecDNA49. To further characterize the dynamics of EGFR ecDNA abundance and heterogeneity during the acquisition of drug resistance, we employed Droplet Hi-C to profile both naive GBM39 cells and cells after more than 30-days of treatment of erlotinib (GBM39-ER) (FIG. 4A). With 9,204 high-quality nuclei profiled, we were able to resolve five different clusters of cells based on chromatin architecture in each cell (FIGS. 4B, 12 A). After erlotinib treatment, a notable evolutionary shift is observed in the GBM39 cells. Subpopulation cluster 2 (C2) nearly vanished, subpopulation cluster 0 (CO) expanded, and several new subpopulations including clusters 1, 3, and 4 (Cl, C3, C4) emerged (FIGS. 4C-4D).
[0193] At the sample level, elevated inter-chromosomal contacts are observed for EGFR locus and, surprisingly, MYC locus before treatment. After erlotinib treatment, the extensive inter-chromosomal contacts at EGFR locus diminished, but increased at the MYC locus and a new locus on chromosome 12 encompassing MDM2 (FIG. 12B). We utilized our ecDNA caller to identify candidate ecDNAs in the Droplet Hi-C data from both GBM39 and GBM39-ER cells. In GBM39, we could detect ecDNA containing the EGFR locus (ecEGFR) and ecDNA containing the MYC locus (ecMYC) (FIG. 12C). We also detected a low- frequency ecDNA on chromosome 18 (1.5%, ecChrl8) (FIG. 12D). In GBM39-ER cells, we discovered ecDNA harboring the MDM2 locus (ecMDM2) and also ecMYC (FIG. 12C).
[0194] At the cluster level, we showed the percentage of cells containing these three ecDNA candidates in every cluster, revealing dynamics of ecDNA-harboring cancer cell populations during development of drug resistance (FIG. 4E). Specifically, ecEGFR disappeared nearly completely after erlotinib treatment (Cl, C3, C4) as previously reported49, whereas ecMDM2 appeared only in erlotinib-resistant populations and is dominant in GBM39-ER subgroup (C3). And ecMYC showed low frequency in GBM39 population (C2) and increased in transition state clusters (CO, Cl, C3). We also plotted the contact maps at 25-kb resolution for these three ecDNAs in every cluster. We observed that the structure and size of the ecMYC amplicon were altered, unlike ecEGFR and ecMDM2, suggesting a more intricate transformation at the MYC ecDNA during the drug treatment (FIG. 4E).
[0195] Surprisingly, when we plotted the distribution of each ecDNA among all the identified clusters (FIG. 4F), we found coexistence of distinct ecDNA species in some tumorcells (FIG. 12E). In cluster CO, a small population of cells (4.7%) harbored both ecMYC and ecEGFR. These cells may potentially serve as the drug-resistant pool of cells during erlotinib treatment, given that the gain of ecMYC may increase cellular fitness. This observation raises the intriguing question of whether GBM39-ER was selected from a small population of cells by erlotinib treatment. Since the simultaneous increase in copy number can indicate whether two adjacent genomic regions are co-present in the same ecDNA, we focused on the ecMYC and characterized its structure in GBM39 cells before and after erlotinib treatment at 10-kb resolution. Although the copy number varies across different cells, the boundaries of ecMYC in GBM39 are almost identical (FIG. 4G). Strikingly after treatment, the ecMYC boundaries as well as internal structure displayed huge variability, reflecting the heterogeneity of MYC ecDNA at single-cell resolution. Boundary variations in ecDNA did not exhibit specificity with relation to clusters (FIG. 4G). We next selected the same number of cells from either GBM39 or GBM39-ER within the transition cluster (CO) and compared the aggregated contacts map at ecMYC regions to see whether an intermediate state for ecMYC structure can be found. Both profiles, although from the same cluster, still resembled their original samples and showed the same differences in ecMYC structure, instead of showing a similar intermediate structure (FIG. 12F). Therefore, we reasoned that for GBM39 cells, erlotinib treatment induced the formation of unique cell populations containing distinct MYC ecDNA, instead of selectively enriching a pre-existing cell population harboring ecMYC with growth advantage upon treatment.
[0196] Droplet Hi-C detects ecDNA in primary tumor samples
[0197] ecDNA is not only a biomarker for poor prognosis but also a crucial driver in GBM50. As Droplet Hi-C successfully revealed dynamic changes of ecDNA in GBM39 cells associated with drug treatment, we further sought to apply Droplet Hi-C to an IDH-wildtype GBM tumor sample (FIG. 5A). By employing only chromatin structure information for cell embedding, we segregated the cells into tumor-like and non-tumor cells (FIG. 5B), featuring expected CNV events in tumor-like cells such as loss of chromosome 10 and extensive amplification of chromosome 7 (FIG. 5C). We found that the genomic region with the most variable copy numbers on chromosome 7 carried the oncogene EGFR (FIGS. 5C-5D), and also displayed an ecDNA-like contact pattern in the contact map of the tumor-like population (FIG. 5E). Utilizing our ecDNA caller, the tumor-like population had significant enrichment of ecEGFR as expected (FIG. 5F).
[0198] In addition to CNV and ecDNA, Droplet Hi-C also identified tumor-like cellspecific chromatin organization and SVs in the patient sample (FIGS. 5G and 13A-13B). Based on the sequencing reads coverage and abnormally high interaction frequency, IKZF1 gene, encoding a tumor suppressor, was found to translocate to ecEGFR without the TSS sequence, leading to its transcriptional disruption in the tumor cell population (FIG. 5G). Among the genes located on the ecDNA, the VOPP1 gene exhibited the strongest interaction frequency with IKZF1, suggesting their fusion (FIG. 5G). Interestingly, this fusion event did not show any discernible effect on the transcriptional activity of VOPP1 , based on the aggregated expression profiles from single-cell RNA-seq data (FIG. 5G).
[0199] Previous single-cell RNA-seq profiling of gliomas defined four major cellular states for tumor cells, including neural progenitor-like (NPC-like) cells, oligodendrocyte progenitor-like (OPC-like) cells, astrocyte-like (AC-like) cells and mesenchymal-like (MES- like) cells51. To investigate whether ecDNA is associated with specific cellular states, we obtained single-cell gene expression together with chromatin accessibility at single-cell resolution using 10X multiome. We co-embedded the Droplet Hi-C data with single-cell RNA-seq, and used k-nearest neighbors to assign one of the four malignant cellular states to individual cells from the Droplet Hi-C dataset (FIGS. 5H-5I and 13C). We subsequently characterized the differential ecDNA features across these states. We found that the fractions of cells containing ecEGFR were similar among the four cell states (FIG. 5 J), but the cells in AC- and MES-like states contained higher copy numbers of ecEGFR (FIG. 5K). The increase of copy number was associated with higher EGFR expression level and chromatin accessibility in these states (FIGS. 5L-5M). Furthermore, expression levels of genes on the ecDNA showed strong correlation with the cellular states score (FIG. 5N). Consistent with previous reports51, the expression level of EGFR exhibited a high correlation with AC-like state scores (FIG. 5N). We also found that genes on the same ecDNA displayed varying correlation levels with different cell states, and that such correlation could not be fully explained by ecDNA copy number, indicating a complex regulatory mechanism governing ecDNA genes expression across cellular states. Examples besides EGFR included the preferential expression of known co-amplified gene SEC61 G in MES-like state, and the preferential expression of LANCL2 in AC- and MES-like states (FIG. 50). The abundance of ecDNA seems to be more closely related to the differentiation states (AC- and MES-like), although mechanism dictating varied expression levels of ecDNA genes besides copy number effects requires further investigation.
[0200] To further demonstrate the utility of Droplet Hi-C in clinical tumor samples, we adopted it to analyze the bone marrow mononuclear cells (BMMC) from an acute myeloid leukemia (AML) and myelodysplastic syndrome (MDS) patient before and after treatment with azacitidine and ventoclax. Azacitidine is a DNA methyltransferase inhibitor known to slow the growth of AML cells in part by promoting their differentiation into more mature cells (FIG. 13D). Venetoclax is an inhibitor of the anti- apop totic protein BCL2. The patient entered a remission from secondary AML after completing one cycle of therapy. The contact maps revealed that the before-treated BMMC initially harbored a strikingly large (~5 MB) ecDNA containing the MYC gene. After treatment, the ecMYC can no longer be detected by Droplet Hi-C (FIG. 13E). To classify tumor-like and non-tumor cells in BMMC, we leveraged known mutations detected by targeted bulk DNA sequencing to distinguish malignant cells. After drug treatment, the proportion of cells harboring mutations also diminished (FIG. 13F). Clinical karyotype analysis showed the persistence of trisomy 8 which was present while the patient had MDS prior to AML transformation. Thus, we speculated that remaining mutant 437 reads might originate from pre-leukemic MDS cells.
[0201] Expression of the oncogene MYC is up-regulated by numerous tissue-specific enhancers through long-range interactions in different cancers52. One of the evolutionarily conserved enhancer clusters (blood enhancer cluster, BENC), which is ~1.8 Mb downstream of MYC, has been shown to activate MYC expression in hematopoietic stem cells and AML52. The long- range interactions between MYC and BENC were observed in the before-treated BMMC sample. The detection of focal amplification of BENC is also consistent with a previous report52(FIG. 13G). Specifically, both the MYC gene and BENC were located on the ecMYC. The interaction between MYC and BENC is obvious in the before-treated sample and disappears after treatment, suggesting its role in tumor initiation and progression (FIG. 13H). Thus, in this patient’s tumor, ecDNA may promote MYC expression not only through increased DNA copy number but also cancer-specific enhancer-promoter interactions.
[0202] In summary, Droplet Hi-C facilitates the detection of ecDNA dynamics in tumors before and after drug treatment. Besides ecDNA, identification of canonical chromatin structures such as SVs and chromatin loops also helps to elucidate the ectopic regulatory programs responsible for tumor progression and drug resistance. Droplet Hi-C thereforeshows great potential for advancing understanding of mechanisms driving tumor clonal evolution and sensitivity to treatment.
[0203] Droplet-based joint profiling of chromatin architecture and transcriptome Single-cell joint profiling of chromatin organization and transcriptome can facilitate the investigation of the relationships between gene expression and chromatin architecture. We modified the above Droplet Hi-C protocol to be compatible with the lOx Genomics Chromium Single Cell Multiome kit, to achieve simultaneous profiling of RNA and Hi-C from single nuclei (FIGS. 6A and 14), termed Paired Hi-C. The major modification in Paired Hi-C is that a milder SDS treatment condition is used before restriction enzymes digestion to minimize RNA degradation. Meanwhile, we used a lower formaldehyde concentration for cross-linking to increase the complexity of the chromatin contact profiles. Third, we noticed that the length of cDNA became shorter post SDS treatment, thus we revised the cDNA library preparation protocol to recover more cDNA molecules. Like Droplet Hi-C, Paired Hi- C also offers benefits in terms of throughput and hands-on time over other microwell- or combinatorial indexing-based joint Hi-C / RNA-seq single-cell assays (FIG. 14).
[0204] To demonstrate the utility of Paired Hi-C, we applied it to a cortex specimen collected from an 8-week old mouse. Overall, we obtained 12,361 joint single-cell profiles of transcriptome and chromatin structure with 42,210 chromatin contacts per cell (duplication rate 66.2%), and a median of 3,914 UMI per cell and 1,746 genes per cell (duplication rate 74%, FIG. 15A). From Paired Hi-C RNA modality, we successfully identified 20 major cell types in the mouse cortex, including 9 excitatory neuron subtypes, 6 inhibitory neuron subtypes, and 5 non-neuronal subtypes (FIG. 6B). The annotation is transferred from Droplet Paired-Tag. The accuracy of cell type identification was confirmed by overlapping with other single-nuclei RNA-seq datasets and the unique expression patterns of marker genes (FIGS. 15B-15D). While single-cell 3D genome features were relatively limited in Paired Hi-C, the compartment score computed from Paired Hi-C data for each cell type (as defined by single- nuclei RNA-seq) exhibited a stronger correlation with the matched cell type's Droplet Hi-C and sn-m3C-seq data (FIGS. 6C-6D and 15E). At cell type-specific marker genes, we can also detect concordant changes of compartment and gene expression (FIG. 6E). As an example, Erbb4 genes were located within the A compartment in PVGA and VIPGA GABAergic neurons, while they were placed in the B compartment in OPC non-neuronal types and ITL5GL glutamatergic neurons. The expression level of Erbb4 was notably higher in PVGA and VIPGA than OPC and ITL5GL (FIG. 6C). We also utilized transcriptomeinformation from Paired Hi-C and co-embed it with Droplet Hi-C. The results agree well with the cell annotations based on Droplet Paired-Tag (FIG. 15F). As such, we can augment the number of chromatin contacts for each cell type by combining Paired Hi-C and Droplet Hi-C.
[0205] The heterogeneity in the ecMYC boundaries at single-cell level within GBM39-ER prompted the question of whether this variable ecM Y C boundary could result in varying ecDNA gene expression within individual cells. By applying Paired Hi-C in GBM39 and GBM39-ER cells, we analyzed ecDNA structure and the expression level of its related genes in single cells. We found that in general, the genes located on the ecMYC regions within GBM39 and GBM39-ER displayed elevated expression levels where their underlying copy number surges, indicating a higher probability of being part of ecMYC (FIG. 6F). For genes located 5’ ecMYC boundaries such as CASC8 and PCAT1, their expression exhibited a strong correlation with the copy number of underlying DNA segments (FIG. 6F). In the case of ecMYC constant genes, such as PVT1, expression was constantly high and also positively correlated with local contact numbers (FIG. 6F). Unexpectedly, the expression of the CCDC26 gene, located at the 3’ end of the ecMYC variable region, was high in GBM39-ER cells despite low copy number (FIG. 6F). When we conducted correlation analysis between gene expression levels and copy numbers for the genes at the 5’ and 3' ends of the variable ecMYC boundary separately in GBM-ER cells, the findings revealed that genes located within the 5' ecMYC variable region exhibited a more pronounced increase in expression with rising copy numbers. In contrast, genes within the 3' ecMYC variable region showed much lower correlations between gene expression levels and copy numbers (FIG. 6G). These results highlight the complex relationship between amplicon structure and gene expression.
[0206] Besides ecMYC structural heterogeneity and gene expression variation, we found that ecMYC displayed different trans-interaction patterns in GBM39 and GBM39-ER cells. These different trans-interaction patterns of ecDNA were also associated with distinct transcriptional responses. Specifically, we observed that GBM39-specific interacting genes showed higher expression levels in GBM39 than in GBM39-ER cells, while GBM39-ER- specific interacting genes had higher expression levels in GBM39-ER cells (FIG. 16A). This suggests a potential mobile regulatory role of ecDNA in global chromosomal transcription46.
[0207] In the GBM patient sample, we found that ecEGFR showed higher copy numbers in AC-like and MES-like cell populations as well as higher expression levels of EGFR, which may render the cellular states changes under erlotinib treatment. Using previously defined methods, we utilized Paired Hi-C single-cell RNA-seq data to classify GBM39 and GBM39-ER cells into four GBM cellular states, and found a significant reduction of differentiated-like population after erlotinib treatment (FIG. 6H). We previously demonstrated that GBM39-ER cells were characterized by the loss of ecEGFR and the emergence of ecMDM2, and we sought to determine whether the expressions of these genes enriched in particular states within the resistant cell population. As found in the GBM tumor sample, EGFR expression was higher in AC-like and MES-like cells (FIG. 61). Interestingly, MDM2 expression was enriched in the progenitor states (OPC- and NPC-like), which became more prevalent after treatment (FIG. 61). These results highlight the unique interplay between ecDNA dynamics, transcriptionally defined cellular states, and drug resistance.
[0208] Besides ecDNA, chromatin architecture also undergoes a dramatic shift during erlotinib treatment, along with changes in gene expression. In GBM39 cells following erlotinib treatment, we observed 1,066 A compartments transitioned into B compartments, whereas 2,796 B compartments shifted to A compartments. The majority of compartment identities remained stable throughout the treatment, including 10,930 A compartments and 11,876 B compartments (FIG. 16B). This compartment shift is also accompanied with gene expression changes. Specifically, genes situated within compartments that changed from A to B exhibited a significant decrease in expression, while those transitioning from B to A showed a significant increase in expression. In contrast, gene expression within stable compartments displayed only slight expression alterations (FIG. 16C).
[0209] In summary, by incorporating transcrip tomic data with chromatin structure information, Paired Hi-C enables direct interrogation of the effects of chromatin reorganization and structural alterations like ecDNA on gene expression, providing a valuable tool for characterizing gene regulatory programs in development and disease pathogenesis.Example 4: Discussion
[0210] In recent years, there has been a significant surge in the development and application of single-cell Hi-C methods. However, current single-cell Hi-C methods still face challenges when applying to heterogeneous tissues and tumor biopsies. Our Droplet Hi-C leverages the commercially available microfluidic platform to provide high-throughput and fast single-cell Hi-C assays with minimal hands-on time, and relatively low costs. Utilizing Droplet Hi-C technology on the adult mouse cortex, we identified cell type-specific hierarchical chromatin structures and established correlations with epigenetic modifications. Additionally, our study unveiled alterations in chromatin organization, CNV, SVs, and ecDNA in cancer cells and tumor patient samples, especially during drug treatment. Wefurther extend Droplet Hi-C to Paired Hi-C, which achieves joint profiling of chromatin architecture and transcriptome in single cells. This enables the association of gene expression phenotypes with 3D genomic structure. The complexity of the Hi-C modality in Paired Hi-C is notably lower when compared to Droplet Hi-C. Nonetheless, we have shown that it is possible to leverage transcriptome data obtained from Paired Hi-C experiments for coembedding with Droplet Hi-C data. This approach enabled us to resolve and annotate cell types and supply Hi-C contact information for each cell type to enhance the resolution of chromatin organization.
[0211] Our droplet-based single-cell Hi-C methods offer advantages in scalability, speed, cost, and adaptability. Additionally, while Droplet Hi-C has lower ratio of cz -long (> 1 kb) interactions than some in situ Hi-C methods (FIG. 8H), it performs similarly to or better than other single-cell Hi-C methods that don’t enrich ligated DNA fragments (FIG. 8H). While cis- short interactions (< 1 kb) provide limited information on chromatin contacts, they are valuable for analyzing CNVs and SVs in cancer cells. Biotin enrichment of cw-long interactions could improve chromatin architecture capture efficiency and reduce sequencing costs.
[0212] EcDNA was first identified in 1965, described as double minutes (DM), and is prevalent in cancer as they often harboring oncogenes53, 54. Unlike the kilobase-sized circular DNA (eccDNA) found in healthy somatic tissues55, ecDNA varies in size from dozens of kilobases to megabases, making it 100 to 1,000 times larger56. The ecDNA is characterized by high amplification and a circular structure57. Current single-cell ecDNA detection methods can be broadly categorized into two primary groups. The first group involves the use of computational tools and WGS data obtained from NGS58or TGS59, 60to identify circularization sites. However, distinguishing between ecDNA and HSR using this approach can be challenging, because they have the same breakpoint sequences. The second group is based on Circle-seq61-63, a strategy that utilizes exonuclease to degrade linear DNA to retain circular DNA for sequencing. This method allows for the differentiation of ecDNA from linear tandem repeats. However, Circle-seq-based strategies have been associated with a high false-positive rate due to incomplete degradation of linear DNA and multiple displacement amplification (MDA). Accurate identification of ecDNA using its multiple features at single-cell level remainschallenging. Here, we have developed an innovative ecDNA detection algorithm that utilizes both trans and cis interactions data from Droplet Hi-C datasets, enabling precise ecDNA identification while effectively distinguishing it from HSR. With our deep learning-based ecDNA caller algorithm, we can reliably detect ecDNA in both cancer cell lines and clinical tumor samples in each single cell, and calculate their proportion within the cell population. Furthermore, the increased occurrence of intra-ecDNA contacts provides valuable support for pinpointing ecDNA boundaries at the single-cell level. For instance, we observed variations in the ecMYC boundary within the GBM39-ER cell population. Additionally, we are capable of detecting SVs events within both ecDNA and HSR.
[0213] Pan-cancer studies have demonstrated the detection of ecDNA in numerous human cancer types and that ecDNA is associated with inferior clinical outcomes across all amplification classes64’68. Further, new studies have begun to highlight unique therapeutic vulnerabilities of tumors that harbor ecDNA69-70. Thus, given the independent prognostic value of ecDNA and its potential to be targeted, recent calls have been made to include ecDNA profiling in clinical specimens71. Droplet Hi-C and deep learning-based ecDNA caller algorithm offer a reliable toolkit for the identification of ecDNA in clinical samples as demonstrated in patient GBM and AML tumors. This toolkit may be particularly useful in heterogeneous patient samples that harbor ecDNA species whose structures and copy numbers may evolve subclonally through treatment. Importantly, our ecDNA caller algorithm possesses the unique capability to differentiate between amplifications occurring on ecDNA and HSRs. Such differentiation may carry critical therapeutic and clinical implications. In our study, we observed the disappearance of ecEGFR, emergence of ecMDM2 and structural dynamics of ecMYC in GBM39 cells following treatment with erlotinib, a tyrosine kinase inhibitor commonly used in the management of EGFR-positive tumors. Notably, the effectiveness of erlotinib treatment in GBM was poor, possibly attributed to the ecDNA evolution. Additionally, in an AML patient sample, we noted that ecMYC could no longer be detected after treatment with azacitidine and venetoclax. Collectively, our methods can be used in ecDNA detection in clinical samples and hold promise for advancing research related to ecDNA in tumor diagnosis and treatment.Example 5: Methods
[0214] Cell cultureHeLa S3 (human, ATCC CCL-2.2) cells were cultured according to standard procedures inhigh glucose DMEM (Sigma- Aldrich, D6429) supplemented with 10% fetal hovine serum (FBS; Omega Scientific, FB-02) and 1% Penicillin-Steptomycin-Glutamine (ThermoFisher Scientific, 10378016) at 37°C with 5% CO2.
[0215] mESCs were maintained in feeder- free and serum-free 2i medium at 37 °C with 5% CO2. To isolate nuclei, mESCs were dissociated with Accutase (Innovative Cell Technologies, AT 104), collected by centrifugation.
[0216] K562 cells were cultured at 37°C, 5% CO2 in RPMI-1640 (ATCC, 30-2001) supplemented with 1% Penicillin-Steptomycin-Glutamine and 10% FBS.
[0217] GM12878 cells were cultured at 37°C, 5% CO2 in RPMI-1640 supplemented with 1% Penicillin-Steptomycin-Glutamine and 15% FBS.
[0218] COLO320DM (ATCC, CCL-220) and COLO320HSR (ATCC, CCL-220.1) were maintained in RPMI-1640 supplemented with 10% FBS (ThermoFisher, SI 155OH) and 1% penicillin- streptomycin (ThermoFisher, 15140122). Media was refreshed every 3-4 days, and cells were passaged once per week using Accutase (VWR, AT 104) for dissociation.
[0219] As described previously, patient-derived xenograft (PDX) model GBM39 (Mayo Clinic Hospital) was cultured in DMEM / F12 (ThermoFisher, 11320033), 20ng / mL EGF (STEMCELL Technologies, 78006), 20ng / ml FGF (STEMCELL Technologies, 78134.1), 10% B27 Supplement without Vitamin A (ThermoFisher, 12587010), and 1% penicillinstreptomycin solution at 37°C with 5% CO2. GBM39-ERL was under continuous treatment with 5pM erlotinib (LC Laboratories, E-4007). GBM39 cell lines were cultured as neurospheres.
[0220] Adherent cells were washed once with lx PBS, trypsinized with 0.25% trypsin- EDTA (ThermoFisher Scientific, 25200056), spun down at 500x g for 5 min, resuspended in culture medium, and spun down again at 500 x g for 5 min. Suspension cells were spun down at 500x g for 5 min to be collected from the culture medium. The cell pellets were resuspended in 1 million cells / mL lx PBS (pH=7.4) for crosslinking.
[0221] Nuclei preparation from the mouse cortex
[0222] All animal work described in this manuscript has been approved and conducted under the oversight of the Institutional Animal Care and Use Committee at the University of California, San Diego. Male C57BL / 6J mice were purchased from the Jackson Laboratory (000664) at 7 weeks of age and were housed in the animal facility at University of California, San Diego, under a 12-h light / 12-h dark cycle in a temperature-controlled room with ad libitum access to water and food until euthanasia and tissue collection at 8 weeks of age. Theisocortex was dissected from 8-week-old male mice, snap-frozen in liquid nitrogen and stored at -80 °C before proceeding to nuclei extraction.
[0223] Single-nuclei suspensions were prepared from fresh tissues by dounce homogenization in douncing buffer (0.25 M sucrose (Sigma, S7903), 25 mM KC1 (Invitrogen, AM9640G), 5 mM MgCl (Invitrogen, AM9530G), 10 mM Tris-HCl (pH 7.5) (ThermoFisher Scientific, 15567027), 1 mM DTT (Sigma, D9779), lx protease inhibitor (Roche, 5056489001), 0.5 U / pL RNaseOUT (Invitrogen, 10777019), 0.5 U / pLSUPERaseln inhibitor (Invitrogen, AM2694) and 0.1% Triton-XlOO (Sigma, 93443)). The nuclei suspension was then filtered through a 30-pm Cell-Trie filter (Sysmex) and centrifuged for 10 min at 300 x g at 4°C. Cell pellets was washed once with douncing buffer without Triton- XlOO, centrifuged again and resuspended in 1 million cells / mL lx PBS (pH=7.4) for crosslinking.
[0224] Experimental protocol for Droplet Hi-C
[0225] A brief description of the Droplet Hi-C experimental procedure is provided below.
[0226] Fixation
[0227] Cells were crosslinked in 1% formaldehyde, which was diluted from 37% formaldehyde with lx PBS (pH=7.4), and incubated at room temperature for 10 min. After crosslinking, the reactions were quenched in 200 mM glycine and incubated at room temperature for 5 min. Quenched reactions were spun down at 1,000 x g for 5 min at 4°C, resuspended using 1% BSA in lx PBS (pH=7.4) to wash twice. One million cells were aliquoted into each tube. The cells were spun once again at 1,000 x g for 5 min, supernatant was removed, and the pellet was flash frozen in liquid nitrogen, and finally stored indefinitely at -80°C.
[0228] In situ Hi-C
[0229] Cell pellets were lysed with pre-cold 300 pL lysis buffer (10 mM Tris-HCl, pH 8.0 (ThermoFisher Scientific, 15568025), 10 mM NaCl (Sigma, S5150), 0.2% Igepal CA630 (Sigma, 18896), lx protease inhibitor (Roche, 5056489001)) on ice for 45 min, then centrifuged at 1,000 x g for 5 min at 4°C to collect nuclei, and washed once with 200 pL lysis buffer. The nuclei were resuspended with 50 pL 0.5% SDS and incubated at 62°C for 10 min on Thermomixer. Then, we added 145 pL of nuclease-free H2O and 25 pL 10% Triton X-100 (Sigma, 93443) to quench SDS and incubated samples at 37°C for 15 min on Thermomixer at 300 rpm. The samples were added 27 pL of lOx CutSmart Buffer and three restriction enzymes, including 50 U DpnII (NEB, R0543L), 62.5U MboII and 7.5U Nlalll (NEB,R0125L), and then incubated at 37°C for 90 min on Thermomixer at 550 rpm. The enzymes were deactivated at 65 °C for 20 min and then cooled down to room temperature. The nuclei were collected by centrifugation at 1000 x g for 5 min at 4°C, washed once with 200 pL ligation buffer (100 pL lOx T4 DNA ligase buffer (NEB, B0202S), 5 pL 20 mg / mL bovine serum albumin (BSA; NEB, B9000S), 865 pL H2O) and ligation reaction was performed with 200 pL ligation buffer and 20 pL T4 DNA ligase (NEB, M0202L) at 37°C for 40 min on Thermomixer at 300 rpm.
[0230] Single-cell Hi-C sequencing Library construction
[0231] The ligated nuclei pellets were resuspended in 1 mL 1% BSA in PBS (diluted from 10% BSA in PBS (Sigma, A1595) using lx PBS (pH=7.4)) each tube, added 1 pL l,000x 7- AAD (Invitrogen, A1310) and sorted by fluorescence- activated nuclei sorting with an SH800 cell sorter (Sony) for the isolation of single nuclei. Nuclei were collected in collection buffer (5% BSA in PBS) at 5 °C and immediately centrifuged for 5 min at 1,000 x g and 4 °C to collect nuclei. The nuclei were washed once with lx Nuclei Buffer (diluted from 20x Nuclei Buffer (lOx Genomics, PN-2000207)), resuspended in 10 pL lx Nuclei Buffer, and counted on an RWD ClOO-Pro fluorescence cell counter with DAPI staining. According to the targeted nuclei recovery number, 3,000-15,000 nuclei were aliquoted into PCR tubes and added Transposition Mix following user guide of Chromium Next GEM Single Cell ATAC Reagent Kits vl. 1 or Chromium Next GEM Single Cell ATAC Reagent Kits v2 (lOx Genomics). The incubation time for tagmentation is 60 min for both vl.l kit and v2 kit. We also extended index PCR elongation time from 20s to 1 min. The double-sided size selection was changed to 1.14x SPRIselect to only remove small fragments.
[0232] All sequencing experiments were performed with an Illumina NextSeq 2000 or NovaSeq 6000 sequencer.
[0233] Experimental protocol for Paired Hi-C
[0234] A brief description of the Paired Hi-C experimental procedure is provided below.
[0235] Fixation
[0236] Cells were crosslinked in 0.6% formaldehyde, which was diluted from 37% formaldehyde with lx PBS (pH=7.4), and incubated at room temperature for 10 min. After crosslinking, the reactions were quenched in 200 mM glycine and incubated at room temperature for 5 min. Quenched reactions were spun down at 1 ,000g for 5 min at 4°C, resuspended using 1 % BSA in l x PBS (pH=7.4) to wash twice, and then 1 million cells were aliquoted into each tube. The cells were spun once again at 1 ,000 x g for 5 min, supernatantwas removed, and the pellet was flash frozen in liquid nitrogen, and finally stored indefinitely at -80°C.
[0237] In situ Hi-C
[0238] Cell pellets were lysed with pre-cold 300 pL lysis buffer (10 mM Tris-HCl, pH 8.0 (ThermoFisher Scientific, 15568025), 10 mM NaCl (Sigma, S5150), 0.2% Igepal CA630 ((Sigma, 18896), lx protease inhibitor (Roche, 5056489001), 1 U / pL RNaseOUT (Invitrogen, 10777-019), 1 U / pL SUPERaseln inhibitor (Invitrogen, AM2694)) on ice for 45 min, then centrifuged at 1,000 x g for 5 min at 4°C to collect nuclei, and washed once with 200 pL lysis buffer. The nuclei were resuspended with 50 pL 0.5% SDS, which was diluted from 10% SDS using lx PBS (pH=7.4) with 1 U / pL RNaseOUT and 1 U / pL SUPERaseln inhibitor, and incubated at 37°C for 60 min on Thermomixer at 300rpm. Then, we added 197 pL quenching buffer (137.59 pL of nuclease-free H2O, 25 pL 10% Triton X-100, 27 pL lOx CutSmart, 1 U / pL RNaseOUT and 1 U / pL SUPERaseln inhibitor) and incubated samples at 37°C for 15 min on Thermomixer at 300 rpm. The samples were centrifuged at l,000x g for 5 min at room temperature, supernatant was removed, and the pellets were resuspended with 250 pL digestion buffer (lx CutSmart Buffer, 1 U / pL RNaseOUT, 1 U / pL SUPERaseln inhibitor, 50 U DpnII, 62.5U MboII and 5U Nlalll), and then incubated at 37°C for 90 min on Thermomixer at 550 rpm. The enzymes were deactivated at 65°C for 20 min and then cold to room temperature. The nuclei were collected by centrifugation at 1000 x g for 5 min at 4°C, washed once with 200 pL ligation buffer (827.5 pL H2O, 100 pL lOx T4 DNA ligase buffer, 5 pL 20 mg / mL BSA, 25 pL SUPERaseln inhibitor and 12.5 pL RNaseOUT), and ligation reaction was performed with 200 pL ligation buffer and 20 pL T4 DNA ligase at 37°C for 40 min on Thermomixer at 300 rpm.
[0239] Single cell sequencing library construction
[0240] The ligated nuclei pellets were resuspended in 1 mL 1% BSA in PBS (diluted from 10% BSA in PBS using lx PBS (pH=7.4)) with 1 U / pL RNaseOUT and 1 U / pL SUPERaseln inhibitor each tube, added 1 pL l,000x 7-AAD (Invitrogen, Al 310) and sorted by fluorescence-activated nuclei sorting with an SH800 cell sorter (Sony) for the isolation of single nuclei. Nuclei were collected in collection buffer (5% BSA in PBS with 5 U / pL Recombinant RNAsin (Promega, PAN2515)) at 5 °C and immediately centrifuged for 5 min at 1,000 xg and 4 °C to collect nuclei. The nuclei were washed once with lx Nuclei Buffer (diluted from 20x Nuclei Buffer (lOx Genomics, PN-2000207) with 1 mM DTT and 1 U / pL Recombinant RNAsin), resuspended in 10 pL lx Nuclei Buffer, and counted on an RWDClOO-Pro fluorescence cell counter with DAPI staining. According to the targeted nuclei recovery number, 3,000-15,000 nuclei were aliquoted into PCR tubes and added Transposition Mix following user guide of Chromium Next GEM Single Cell Multiome AT AC + Gene Expression (lOx Genomics). The tagmented nuclei were loaded onto a Chromium Next GEM Chip J for droplet generatiom with a Chromium X microfluidic system (lOx Genomics). The GEMs were collected, reverse transcription and cell barcoding were performed using a thermal cycler, and the reaction was quenched following the user guide of Chromium Next GEM Single Cell Multiome AT AC + Gene Expression (lOx Genomics). After GEM cleanup and pre-amplification, the SPRlselect products were eluted into 100 pL buffer EB (Qiagen, 19086) instead of 160 pL. For scHi-C library construction, we extended index PCR elongation time from 20s to 1 min, and the double sided size selection was changed to 1.55x SPRlselect to only remove small fragments. For scRNA seq library construction, we used 0.8x SPRlselect to purify cDNA twice instead of 0.6x SPRlselect to keep short cDNA fragments.
[0241] All sequencing experiments were performed with an Illumina NextSeq 2000 or NovaSeq 6000 sequencer.
[0242] Processing of Droplet Hi-C data
[0243] Custom scripts for demultiplexing, mapping and extracting single cell contacts from Droplet Hi-C data were publicly available at github.com / Xieeeee / Droplet-Hi-C. First, depending on indexing primers used, Droplet Hi-C fastq files were demultiplexed using Illumina bcl2fastq (v2.19.0.316) or 10X Genomics cellranger-atac mkfastq (v2.0.0). After demultiplexing, a custom script was used to extract cellular barcode sequence from Read2, and barcodes are aligned to the whitelist provided by lOx Genomics using Bowtie (vl.3.0)72. Aligned barcode sequences were appended to the beginning of the read name to record cellular identity of each read. Next, sequencing adapters were detected and trimmed with Trim-Garole (v0.6.10)73. Cleaned reads were then mapped to the human (hg38) or mouse (mmlO) reference genome using BWA-MEM (vO.7.17)74, with arguments ‘-SP5M’ specified. Finally, after mapping, valid contacts were parsed, sorted and deduplicated by Pairtools (v0.3.0)75, with barcodes information stored in a separated column in the pairs file. Contacts were balanced and stored in cool format using Cooler (v0.8.10) for visualization and downstream analysis76. High-quality cells were selected based on the number of total unique contacts per cell in each library.
[0244] Analysis of Droplet Hi-C data
[0245] We here described the general analysis strategy and workflows for our Droplet Hi-C data. Dataset-specific manipulations in the manuscript, if any, will be indicated.
[0246] Visualization of Hi-C contact matrices
[0247] Bulk, pseudo-bulk, or single cell contact matrices in cool format were visualized and plotted using Cooler along with Matplotlib (v3.5.1)77.
[0248] Embedding and clustering of single-cell Droplet Hi-C dataset on cultured cells
[0249] We used Higashi to infer low-dimensional cell embeddings for all our Droplet Hi-C datasets without imputations, including human cell lines mix (HeLa / GM12878 I K562, FIG.7 A), GBM39 / GBM39-ER (FIG. 4B), and GBM patient sample (FIG. 5B). For visualization, the L2 norm of cell embeddings were projected to two-dimensional space with uniform manifold approximation and projection (UMAP). To identify cells with similar identity, we performed Leiden clustering using igraph (vO.9.9) and leidenalg (vO.8.8)78, 79.
[0250] Imputation of single-cell chromatin contacts on adult mouse cortex Droplet Hi-C dataset
[0251] For the adult mouse cortex Droplet Hi-C data, we used scHiCluster (vl.3.4) to perform contact matrix imputation for individual cell at three different resolutions: 100 kb (for cells embedding and visualization), 25 kb (for domain boundaries calling) and 10 kb (for cells embedding and loops calling)80. In brief, contacts from individual cell were first masked for ENCODE blacklist regions81. scHiCluster then performed linear convolution and random walk with restart to impute the sparse single cell contact matrices for each chromosome.Considering storage efficiency, we output the imputed results for the whole chromosome at lOOkb, for 10.05Mb window at 25kb, and for 5.05Mb window at lOkb resolution for each cell.
[0252] Embedding and annotation of single-cell Droplet Hi-C dataset on adult mouse cortex
[0253] For the imputed adult mouse cortex datasets, we adopt a previously described method to co-embed Hi-C data with gene expression data (snRNA-seq) generated from the same tissue and to assign cell-type identities39. Briefly, single-cell gene associating domain (scGAD) scoreRij is calculated as the raw number of interactions at gene body region for gene i in cell j with the 10 kb imputed matrix. This method has been implemented in scHicluster as the ‘gene_score’ module. By this, the single-cell chromatin interaction profiles can then berepresented as a cell-by-gene scGAD score matrix, which serves as the input for integration with snRNA-seq data.
[0254] To perform integration, we began by normalizing and selecting the top 2,000 variable genes in reference snRNA-seq data. These genes were subjected to dimension reduction by principal-component analysis (PC A) using Scanpy (vl.7.2), retaining 30 principal components82. The derived PCA model was then applied to transform the scGAD score matrix.Subsequent integration was conducted with a tailored version of Seurat Integration using canonical correlation analysis (CCA)83. For the visualization of co-embedded data, we calculated the nearest neighbor graph (k = 25) and UMAP embedding with Scanpy.
[0255] To annotate cell identities within the Droplet Hi-C dataset, we identified the 15 nearest snRNA-seq neighbors for each cell via PyNNDescent (v0.5.6) using Euclidean distances84. These distances were scaled and converted into standardized scores so that the total score is added up to 1 , and neighbor with the minimal distance would have the highest score. The identity of each Droplet Hi-C cell was inferred by nominating the cell type label that garnered the highest standardized score among its top nearest neighbors.
[0256] To validate the cell type annotations, co-embedding of Droplet Hi-C and sn-m3C- seq datasets was performed using scHiCluster, which facilitates the low-dimensional representation of chromatin contacts information. For each cell, imputed matrices at 100 kb resolution were binarized and flattened. The binarized data from individual cells across both datasets were then concatenated into a larger matrix. PCA was applied to this matrix on a per- chromosome basis. The PCA results for each chromosome were concatenated for a second round of PCA to generate the final cell embeddings. Batch effects were corrected using harmonypy (v0.0.9)85. After batch correction, Leiden clustering was employed to identify coembedding clusters. The validity of original annotation was verified by calculating the overlap coefficients (O / ,7) between the clusters and the original annotations from Droplet Hi-C(A) and sn-m3C-seq dataset (B):, where i indicates query cell type from Droplet Hi-C, j indicates reference cell type from sn-m3C-seq, and k indicates the co-embedding cluster.
[0257] Analysis of A / B compartments
[0258] On sample levels, we used the ‘eigs_cis’ module from cooltools (vO.5.1 ) to perform eigen decomposition on balanced contact matrices and calculate the compartment score at 100 kb resolution86. The orientation of the resulting eigenvectors was adjusted to correlate positively with the CpG density of the corresponding 100 kb genomic bins, thereby determining the sign of the compartments (A / B). Similarity between samples is calculated as the Spearman correlation coefficient of compartment scores across all autosomes (FIG. 7C). To illustrate chromatin interactions within and between compartments, we used the ‘saddle’ module from cooltools to calculate the average observed versus expected contact frequency, categorized by compartment scores.
[0259] To generate pseudo-bulk profiles for distinct cell types or clusters within a sample, we adopted a uniform approach outlined before to minimize bias and facilitate direct comparisons. Initially, the raw contact matrices at the sample level were normalized for distance effects. Next, Pearson correlation coefficients were computed on the distance- normalized matrices, and PCA was then performed on the correlation matrix. The model fitted at the sample level was utilized to transform the raw contact matrices of various cell subtypes or clusters. To compare similarity between cell types or clusters, we calculated the Spearman correlation coefficient of compartments score using compartment intersected with variable gene regions (FIGS. IE and 13C). To identify differential compartments, we employed dchic (v2.1)87, selecting only those with an adjusted P value < 0.01 and a Manhattan distance greater than the 2.5th percentile of the standard normal distribution for further downstream analysis.
[0260] For joint analysis of compartment and cell type-specific histone modifications, we utilized the recently published adult mouse cortex Droplet Paired-Tag dataset40. We retained only cell types that contain >100 cells, and were shared between datasets to ensure comparability across the analyses. For all differential compartments, we calculated the Pearson correlation the same 100 kb genomic bins across various cell types.
[0261] Analysis of chromatin domains
[0262] On sample levels, we used the 'insulation' module from cooltools to compute insulation score for raw contact matrices at 25 kb resolution. To delineate cell type-specific variable domain boundaries in mouse cortex data, we employed TopDom (v0.0.2) on 25 kb resolution imputed contact matrices for each individual cell88. We defined the 'boundary probability' for a given bin as the proportion of cells within a particular cell type thatdesignate the bin as a domain boundary. The presence or absence of a domain boundary is summarized by n cell types into a n x 2 contingency tables. We then computed the chi-square statistics and P value for each bin, and the domain boundaries were classified as variable if they exhibited a false discovery rate (FDR) < 0.001 and a boundary probability difference > 0.05.
[0263] For correlation analysis of gene expression surrounding variable domain boundaries, we first identified genes within a lOOkb window centered on the variable domain boundaries. We then calculated Pearson correlation coefficients to quantify the relationship between boundary probabilities and expression levels (RPKM) of these genes from Droplet Paired-Tag data across shared cell types. The genes were further categorized into “housekeeping”, “constant”, and “variable”. In brief, the housekeeping gene list is taken from previous report89. The top 2,500 variable genes from Droplet Paired-Tag RNA modality are selected as variable group. Genes with similar expression level as the top 2,500 variable genes but are not within the “housekeeping” or “variable” group is treated as “constant”.
[0264] Analysis of chromatin loops
[0265] For loops calling on sample levels, we adopted the 'dots' module in cooltools, which implements the principle of the commonly used HiCCUPS loop calling strategy, to call loops on the 10 kb raw contact matrices90.
[0266] A modified version of SnapHiC workflow implemented in scHiCluster is used to perform loops calling on the 10-kb imputed contact matrices for mouse cortex cell types91. To compare histone modification enrichment at loop anchors, we calculate H3K27ac or H3K27me3 CPM across all cell types at all loop anchors identified. When histone modification profiles and loop 91 anchors are from the same cell types, the enrichment is classified as “matched”, otherwise 911 classified as “unmatched”.
[0267] Gene Ontology annotation of loop anchors was performed using rGREAT (v 1.26.0) in R92. Gene Ontology biological process was selected for annotations. The result is used to generate plots in Fig. 2g.
[0268] Infer copy number variation (CNV) with Droplet Hi-C data
[0269] Copy number is inferred from Hi-C data using the ‘calculate-cnv’ module in NeoLoopFinder (vO.4.3)93. For each single cell, the output residuals from the generalized additive model in ‘calculate-cnv’ was directly used to estimate copy number ratio. On bulk or pseudo-bulk level, an additional HMM segmentation is performed using the ‘segment-cnv’module on the copy number ratio to determine the boundaries of copy number ratio segments. We assume all samples used in this manuscript for CNV calculation are diploid, therefore the inferred copy number is equal to 2 x copy number ratio. Genome-wide CNV heatmap is plotted at 10 Mb resolution, where chromosome level CNV is plotted at 1 Mb resolution, and regional CNV 926 profile is plotted at 10 kb resolution in R using pheatmap (vl.0.12). The inferred copy number 927 is further used to assist identifying ecDNA candidates and the associated significant trans-928 interactions, and to correct the bias in contact matrices for structural variation identification.
[0270] Identify structural variation (SV) from Droplet Hi-C data
[0271] Structural variations (SV) were identified using the 'predicts V’ module in EagleC (vO.1.9)94. In short, CNV effects on contact matrices at 5, 10, and 50 kb resolutions were removed using ‘correct-cnv’ module in NeoLoopFinder. ‘predictSV’ then uses a deep learning model to predict SV at each resolution on the corrected matrices, and combines all results to obtain a uniform, high-resolution SV list at 5 kb resolution, ‘annotate-gene-fusion’ module is applied on the final SV list to annotate gene fusion events.
[0272] Metrics to define chromatin interactions
[0273] We compiled a suite of metrics to assess the cis- and trans-interaction patterns at specific genomic loci. These metrics include: (1) contact evenness or “hub index”, quantified as the Gini coefficient for trans -interactions across all chromosomes excluding chromosome Y given an 1 Mb genomic bin, and was derived using ineq (vO.2-13) in R; (2) trans-to-cis contacting bin ratio ( / ?,), which compares the quantity of interacting bins within the samechromosome (Nc) to those across dilferent chromosomes (NT) for a given bin i: -V.This ratio represents the rrans-interaction tendency while it is not confounded by copy number variation; (3) copy number-adjusted tran. -chromosomal interaction (adjnTJF), which is also used to measure a genomic locus’s interaction activity, was calculated as described before46.
[0274] Develop ecDNA caller for identifying ecDNA candidates
[0275] We derived two ecDNA callers to identify ecDNA candidates, one is based on the logistic regression model and the other one is based on the convolutional neural network (CNN). Both models predict genome-wide 1 Mb bins which contain ecDNA, as well as cells with presence of ecDNA population-wide. Detailed methodologies for both algorithms are delineated as follows.
[0276] Logistic regression model-based ecDNA caller
[0277] We trained a multivariate logistic regression model using the glm function in R with inferred copy number, hub index, and trans-to-cis contacting bin ratio as predictive variables. The training dataset comprised the well-defined MYC ecDNA in COLO320DM(chr8: 127Mb- 128Mb) and EGFR ecDNA in GBM39 (chr7:55Mb-56Mb) as positive data, with the same loci in COLO320HSR and GBM39-ER as negative control. We established a probability threshold for classifying ecDNA presence. Specifically, in COLO320DM / COLO320HSR data, this threshold was determined to be 0.5. In GBM39 / GBM39-ER data the threshold was determined to be 0.95.
[0278] Deep learning-based ecDNA caller
[0279] Data preprocessing
[0280] To train the deep learning-based ecDNA caller, we also selected the well-defined MYC ecDNA in COLO320DM (chr8 :127Mb- 128Mb) and EGFR ecDNA inGBM39 (chr7:55Mb-56Mb) as positive data, with the same loci in COLO320HSR and GBM39-ER as negative control. We used all autosomes and chromosome X to create Hi-C contact matrix at 1 Mb resolution (3,044x3,044) for each cell. We then randomly selected 90% of cells as the training data and 973 kept the remaining 10% of cells as the validation data. For each 1Mb bin, we retained its local 974 5 Mb neighborhood region (including the center bin itself), and used both cis- and trans-975 contacts (i.e., a 5x3,044 matrix) as its feature in the neural network model.
[0281] Model architecture
[0282] Our proposed model consists of two sequentially placed convolutional modules that extract features from the binarized Hi-C contact matrices. Each convolutional module consists of a multi-channel convolutional layer (8 channels in the first module, 16 channels in the second module), a batch normalization layer, a rectified linear unit (ReLU) activation function95, and a max-pooling96layer sequentially. Convolution kernels scan along the direction of rows in each layer. Large convolutional kernel sizes (5x45 in the first module and 1x45 in the second layer) are set to improve pattern capture because of sparsity. Maxpooling reduces the size of the matrix to half on its second dimension to keep the most important features and thus improves learning efficiency and propagation speed. Strides and paddings of each layer were designed to balance computational efficiency and information retention and must be compatible with the matrix shapes and max-pooling layers. To enhance robustness and prevent overfitting, a dropout layer97with probability of 0.5 is insertedbetween convolutional modules. Subsequently, a two-layer fully connected (dense) network integrates information from multiple sources and makes ternary predictions. The first fully connected layer with a 218 hidden size receives (A) the flattened output of the convolutional modules, (B) flattened 5x5 small contact matrix of the corresponding 5 Mb region (whose diagonal entries denotes intra-bin contacts) from binarized 5x3044 Hi-C contact matrices, (C) the hub index (used to measure inequality and heterogeneity of interactions with different chromosomes for certain bin and calculated by Python package “pygini”) of the center bin calculated from interaction counts aggregated per chromosome, and (D) LI normalized row means of binarized 5x3044 Hi-C contact matrices. The output passes through a batch normalization layer, a Gaussian error linear unit (GELU) activation function98, and a dropout layer with probability of 0.5. Finally, he second fully connected layer with hidden size 64 produces predictions with a subsequent softmax activation function, in which each prediction contains three probabilities of “None”, “ecDNA”, “HSR” sequentially.
[0283] Model training and validation
[0284] With the package PyTorch", the training process spanned 40 epochs using the minibatch training strategy (batch size=32). To ensure robust optimization, we applied the AdamW100'101optimization algorithm with 0.001 weight decay and 0.001 learning rate, and implemented the“hard” bootstrapped cross-entropy loss with parameter equals to O.9930, which is calculated from one-hot-encoded class labels and predictive softmax102probabilities. To efficiently reduce false positives, we compensated the difference in quantities of different labels and in addition forced the model to favor negative prediction (“None”) using biased class weights, which is incorporated into the loss calculation. Thus, the model receives a greater loss as a penalty when generating any positive prediction (“ecDNA” or “HSR”). We mainly utilized confusion matrix and subsequently derived binary and ternary precision, recall, specificity, and accuracy as our evaluation metrics, and binary results (“None” or “ecDNA” only) from logistic regression as the baseline performance.
[0285] Prediction of ecDNA
[0286] The model scans each Hi-C contact matrix from beginning to end without strides to generate SoftMax probabilities on each 1 Mb bin (except the first two bins of chromosome 1 and the last two bins of chromosome X as no padding to the Hi-C matrix is applied) on each single cell. Argmax (the arguments of the maxima) function is used to determine the finalpredicted class, which is either 0 (“None”), 1 (“ecDNA”), or 2 (“HSR”). Subsequently, results are aggregated to calculate the proportion of cells with ecDNA and / or HSR among the entire cell population.
[0287] Identify significant trans-interactions of ecDNA candidate loci
[0288] To identify genomic regions preferentially interact with ecDNA, we first quantify the number of interactions ( ) of 500 kb genomic interval i on different chromosomes with ecDNA candidate loci, treating each chromosome separately. The observed interaction frequency (P() is the roportion of Nt relative to all contacts on the same chromosome:The expected interaction frequency for interval i was calculated as the ratio of the inferred copy number (CM) to the total copy number on the same chromosome:, based on the null hypothesis that interaction frequency for genomic regions with ecDNA are only weighted by their underlying copy number. To identify significant interactions, observed- versus-expected P value was calculated based binomial distribution model. Multiple testing correction is done by Bonferroni adjustment. Regions with adjusted P value < 0.05 are selected as significant interacting regions.
[0289] Analysis of 10X Multiome datasets
[0290] 10X Multiome fastq files for the GBM patient sample were demultiplexed and pre- processed with cellranger- arc (v2.0.0). Clustering of single-nucleus transcriptomic or chromatin accessibility data were performed in R using Seurat (v4.1.0) or Signac (vl.6.0)103,104. For transcriptomic data, gene counts were normalized and scaled, and the top 2,500 variable genes were selected for dimension reduction by PCA. The first 30 principal components were used for UMAP visualization and Louvain clustering. Potential doublets were identified and removed using Scrublet105. For chromatin accessibility data, cell-by-peak matrices were normalized by the two-step term frequency-inverse document frequency (TF- IDF). The top 95% of genomic bins were selected for linear dimension reduction, again followed by UMAP visualization and Louvain clustering. Gene activity scores were computed by AT AC signal density in promoter and gene body regions.
[0291] Analysis of GBM cellular states
[0292] To classify single cells within GBM lOx Multiome or Paired Hi-C RNA datasets into one of four predefined malignant cellular states, we adopted a previously described two- dimensional visualization technique51. Briefly, we quantified the gene enrichment score(SCj(i)) for each cell (?) against one of the four gene sets (G / ) associated with the particular cellular state. This score was calculated as the relative averaged expression (Exp) of Gj in cell i compared to a group of genes (Gjcont) with a similar level of expression as control:nt are the number of genes in Gj and Gjcont- After obtaining scores for all four cellular states, the cells were stratified into OPC / NPC versus AC / MES categories using the d .ifferenti ■al . score lFor further refinement,OPC / NPC cells were assigned an identity valueandAC / MES cells were similarly categorized with+ -0. The distribution of cellular states was then plotted in the two-dimensional representation with D on the y-axis and C on the x-axis.
[0293] For cellular state identification in Droplet Hi-C data, co-embedding with the reference snRNA-seq was performed using the scGAD score. After determining the top 15 nearest snRNA-seq neighbors and their scaled distance-based similarity score (Dm) for each Droplet Hi-C cell, theHi-C gene enrichment score (HSCjtp) was computed as the sum of the neighbor- weighted enrichment scores:.
[0294] Single-cell genotyping by Droplet Hi-C dataFrequently observed mutations in AML were sequenced using Hematologic Malignancy Comprehensive Panel for the patient BMMC sample. To detect malignant cells carrying the detected mutations from Droplet Hi-C data, cellular barcode information was appended to the mapped bam files as an extra tag. To ensure comparability between pre- and post-treated1079 samples, bam files were subsampled to a uniform sequencing depth, we employed a1080 previously reported custom script to count reads for both the wild-type and mutant alleles, considering only barcodes that met predefined quality standards106. For each cell and mutational site, we then summarized the detected mutant and wild-type reads. Given the limited sequencing coverage, we defined potential malignant cells as cell carrying at least one mutant read.
[0295] Processing and Analysis of Paired Hi-C data
[0296] Preprocessing and analysis of Hi-C modality in Paired Hi-C are identical as described for Droplet Hi-C. For RNA modality, fastq were demultiplexed by cellranger-arc, but preprocessed using cellranger (v6.1.2). After obtaining the cell-by-gene matrix, clustering and visualization were performed as described for lOx Multiome RNA dataset. Since barcodes in the same gel bead for RNA and Hi-C modalities are different, we performed manual pairing to match cell barcodes based on the 10X Multiome barcodes whitelist provided in cellranger-arc (lOx Genomics).
[0297] Integration of the Paired Hi-C RNA dataset with reference datasets was performed by Seurat. First, normalization of gene counts was performed, and the top 2,000 shared variable genes across datasets were selected as integration features. Subsequent canonical correlation analysis allowed projecting all nuclei into a unified embedding space. Anchors (pairs of corresponding cells from distinct datasets) were then discerned through mutual nearest neighbors searching. Anchors with low confidence were excluded, and the shared neighbor between anchors and query cells are computed. Louvain clustering was applied to the shared neighbor graph to discern co-embedded clusters. We calculated overlap coefficients as described in previous section to compare clustering results from different datasets.
[0298] Analysis of gene expression levels and copy numbers of ecDNA
[0299] To calculate the correlation between gene expression level and inferred copy number of genes on ecDNA, we first refined the range of ecDNA at 10 kb resolution. This is based on the observation of the increased intra-ecDNA interaction frequency than interactions with regions on linear genome, irrespective of their genomic distance. Specifically, we enumerated the local interactions within each 10 kb segment of the 1 Mb ecDNA candidates. Subsequently, the contact numbers were smoothed across adjacent bins. By calculating the differential contact numbers between neighboring bins, we determined the changing points of interaction with an average difference cutoff across the entire 1 Mb region. The outermost local maxima and minima were designated as ecDNA boundaries.
[0300] With the ecDNA boundaries established, we categorized genes with gene bodies residing within ecDNA, as “ecDNA genes”. In the case of MYC ecDNA (ecMYC), we further classified genes on GBM39 ecMYC as “shared genes”, and genes on non-overlapping regions between GBM39 and GBM39-ER as “ecMYC variable genes”, considering the greater ecDNA size in GBM39-ER at pseudo-bulk level. The Spearman correlationcoefficient between gene expression levels (RPKM) and the average inferred copy number over all 10 kb segments at gene body across all cells was calculated, for a comprehensive representation of the interplay between gene expression and copy number.
[0301] Ethical approval
[0302] The GBM specimen collection was approved by the Institutional Review Board (IRB) at the University of Minnesota. The AML and MDS specimen collection was approved by the IRB at the University of California, San Diego. Each patient was consented by a dedicated research coordinator prior to collection. Samples were obtained with informed consent in accordance with the Declaration of Helsinki and appropriate Ethics Committee approval from each partner institution.
[0303] Data availability
[0304] Raw and processed sequencing data generated in this study have been submitted to GEO (accession number GSE253407, with reviewer token mpgtmwcgvpgfvwj). The processed data can also be accessed as supplementary files in GEO. Datasets for bulk in-situ Hi-C on cultured cells were downloaded from the 4DN data portal with the following accession number: HeLa (4DNFIIT9PKS9 and 4DNFI2XGP1BW), K562 (4DNFIIX5BNC9 and 4DNFI244AS29), and GM12878 (4DNFIIS73OJN and 4DNFI3O82QVV). Other external datasets were downloaded from NCBI GEO with the following accession numbers: Droplet Paired-Tag dataset on mouse frontal cortex (GSE152020), 10X Multiome dataset on mouse cortex (GSE210749), 10X Multiome dataset on COLO320DM / COLO320HSR (GSE160148). The BICCN whole mouse brain sn-m3C-seq datasets and BICCN MOp lOx snRNA-seq data were downloaded via the NeMO archive (assets.nemoarchive.org / dat- sig83t9; assets.nemoarchive.org / dat-chlnqb7). Source data are provided with this paper.
[0305] Code availability
[0306] Scripts and code are available at github.com / Xieeeee / Droplet-Hi-C. The ecDNA callers 1149 are available at github.com / HuMingLab / ecDNAcaller.
[0307] Additional Detailed Methods
[0308] Additional descriptions and details regarding the compositions and methods provided herein are found in Chang et al., Nat Biotechnol, 2024 (doi: 10.1038 / s41587-024- 02447-1), which is incorporated herein in its entirety and for all purposes.Example 6
[0309] This method, referred to as Droplet Hi-C, allows a user to rapidly profile the chromatin organization at single cell resolution in a complex tissue or heterogeneous cellpopulation. Our method combines an in situ chromosomal conformation capture (3C) assay with commercially available droplet microfluidics, to simultaneously capture the 3D genome structure from tens of thousands of individual cells in a single experiment. Using Droplet Hi- C, we can map the chromatin architecture at single-cell resolution from the heterogeneous tissues or mixed cell populations, such as mouse cortex, and use the results to delineate gene regulatory programs in the constituent cell types. Additionally, we can use this technique to detect copy number variation (CNV), structural variations (SVs) and extrachromosomal DNA (ecDNA) in cancer cells, to reveal clonal dynamics and other oncogenic events. We further refined this technique to allow for joint profiling of chromatin architecture and transcriptome in single cells, facilitating a more comprehensive exploration of the links between chromatin architecture and gene expression in both normal tissues and tumors. We also developed bioinformatics algorithms to process the data and determine aberrant chromatin architecture especially ecDNA in the cancer cells.
[0310] This is the first method to profile 3D genome architecture in nano droplets. It offers at least an order of magnitude greater in throughput over existing methods, and is rapid, easy to do, and highly scalable.
[0311] There are two categories of existing methods, including low-throughput microwellbased approaches (Dip-C, scNanoHi-C, sn-m3C-seq and HiRES) and high-throughput combinatorial indexing-based approaches (sciHi-C and GAGE-seq). It is still very difficult to apply them to primary tissues and tumor biopsies especially in average molecular biology labs. As a matter of fact, the high labor intensity and cost per cell make the low-throughput microwell-based approaches difficult to cover complex cell types in the biospecimens. The lengthy and largely manual combinatorial indexing procedure still poses a challenge for high- throughput combinatorial indexing-based approaches general adoption.
[0312] Further we developed a method to identify ecDNA from single cell Hi-C data, which can be used to characterize the clonality of cancer cells during drug treatment and cancer development, to identify cancer drivers.References
[0313] 1. Dixon, J.R., Gorkin, D.U. & Ren, B. Chromatin domains: the unit of chromosome organization. Molecular cell 62, 668-680 (2016).
[0314] 2. Rowley, M.J. & Corces, V.G. The three-dimensional genome: principles and roles of long-distance interactions. Current opinion in cell biology 40, 8-14 (2016).
[0315] 3. Dekker, J. & Heard, E. Structural and functional diversity of topologicallyassociating domains. FEBS letters 589, 2877-2884 (2015).
[0316] 4. Bonev, B. & Cavalli, G. Organization and function of the 3D genome. Nature Reviews Genetics 17, 661-678 (2016).
[0317] 5. Rao, S.S. et al. A 3D map of the human genome at kilobase resolution reveals principles of chromatin looping. Cell 159, 1665-1680 (2014).
[0318] 6. Dixon, J.R. et al. Topological domains in mammalian genomes identified by analysis of chromatin interactions. Nature 485, 376-380 (2012).
[0319] 7. Nora, E.P. et al. Spatial partitioning of the regulatory landscape of the X- inactivation centre. Nature 485, 381-385 (2012).
[0320] 8. Lieberman- Aiden, E. et al. Comprehensive mapping of long-range interactions reveals folding principles of the human genome, science 326, 289-293 (2009).
[0321] 9. Dekker, J., Marti-Renom, M.A. & Mirny, L.A. Exploring the three- dimensional organization of genomes: interpreting chromatin interaction data. Nature Reviews Genetics 14, 390-403 (2013).
[0322] 10. Gibcus, J.H. & Dekker, J. The hierarchy of the 3D genome. Molecular cell 49, 773-782 (2013).
[0323] 11. Du, Z. et al. Allelic reprogramming of 3D chromatin architecture during early mammalian development. Nature 547, 232-235 (2017).
[0324] 12. Ke, Y. et al. 3D chromatin structures of mature gametes and structural reprogramming during mammalian embryogenesis. Cell 170, 367-381. e320 (2017).
[0325] 13. Xu, J. et al. Subtype-specific 3D genome alteration in acute myeloid leukaemia. Nature611, 387-398 (2022).
[0326] 14. Xu, Z. et al. Structural variants drive context-dependent oncogene activation in cancer. Nature 612, 564-572 (2022).
[0327] 15. Dixon, J.R. et al. Chromatin architecture reorganization during stem cell differentiation. Nature 518, 331-336 (2015).
[0328] 16. Zheng, H. & Xie, W. The role of 3D genome organization in development and cell differentiation. Nature Reviews Molecular Cell Biology 20, 535-550 (2019).
[0329] 17. Ghosh, R.P. & Meyer, B.J. Spatial Organization of Chromatin: Emergence of Chromatin Structure During Development. Annual Review of Cell and Developmental Biology 37, 199-232 (2021).
[0330] 18. Gorkin, D.U., Leung, D. & Ren, B. The 3D genome in transcriptional regulation andpluripotency. Cell stem cell 14, 762-775 (2014).
[0331] 19. Morgan, M.A. & Shilatifard, A. Chromatin signatures of cancer. Genes & development 29, 238-249 (2015).
[0332] 20. Papaemmanuil, E. et al. Genomic classification and prognosis in acute myeloid leukemia. New England Journal of Medicine 374, 2209-2221 (2016).
[0333] 21. Kandoth, C. et al. Mutational landscape and significance across 12 major cancer types. Nature 502, 333-339 (2013).
[0334] 22. Blum, A., Wang, P. & Zenklusen, J.C. Snapshot: TCGA- Analyzed Tumors. Cell 173, 530 (2018).
[0335] 23. Nagano, T. et al. Single-cell Hi-C reveals cell-to-cell variability in chromosome structure. Nature 502, 59-64 (2013).
[0336] 24. Nagano, T. et al. Cell-cycle dynamics of chromosomal organization at singlecell resolution. Nature 547, 61-67 (2017).
[0337] 25. Stevens, T.J. et al. 3D structures of individual mammalian genomes studied by single-cell Hi-C. Nature 544, 59-64 (2017).
[0338] 26. Flyamer, I.M. et al. Single-nucleus Hi-C reveals unique chromatin reorganization at oocyte-to-zygote transition. Nature 544, 110-114 (2017).
[0339] 27. Tan, L„ Xing, D„ Chang, C.-H., Li, H. & Xie, X.S. Three-dimensional genome structures of single diploid human cells. Science 361 , 924-928 (2018).
[0340] 28. Li, W. et al. scNanoHi-C: a single-cell long-read concatemer sequencing method to reveal high-order chromatin structures within individual cells. Nature Methods 20, 1493-1505 (2023).
[0341] 29. Li, G. et al. Joint profiling of DNA methylation and chromatin architecture in single cells. Nature Methods 16, 991-993 (2019).
[0342] 30. Lee, D.-S. et al. Simultaneous profiling of 3D genome structure and DNA methylation in single human cells. Nature Methods 16, 999-1006 (2019).
[0343] 31. Liu, Z. et al. Linking genome structures to functions by simultaneous singlecell Hi-C and RNA-seq. Science 380, 1070-1076 (2023).
[0344] 32. Ramani, V. et al. Massively multiplex single-cell Hi-C. Nature Methods 14, 263-266 (2017).
[0345] 33. Zhou, T. et al. Concurrent profiling of multiscale 3D genome organization and gene expression in single mammalian cells. bioRxiv, 2023.2007.2020.549578 (2023).
[0346] 34. Zhang, R., Zhou, T. & Ma, J. Multiscale and integrative single-cell Hi-Canalysis with Higashi. Nature Biotechnology 40, 254-261 (2022).
[0347] 35. Salik, B. et al. Targeting RSPO3-LGR4 Signaling for Leukemia Stem Cell Eradication in Acute Myeloid Leukemia. Cancer Cell 38, 263-278. e266 (2020).
[0348] 36. Tan, L. et al. Changes in genome architecture and transcriptional dynamics progress independently of sensory experience during post-natal brain development. Cell 184, 741-758. e717 (2021).
[0349] 37. Liu, H. et al. Single-cell DNA methylome and 3D multi-omic atlas of the adult mouse brain. Nature 624, 366-377 (2023).
[0350] 38. Zhang, C. et al. tagHi-C reveals 3D chromatin architecture dynamics during mouse hematopoiesis. Cell Reports 32 (2020).
[0351] 39. Shen, S., Zheng, Y. & Kele§, S. scGAD: single-cell gene associating domain scores for exploratory analysis of scHi-C data. Bioinformatics 38, 3642-3644 (2022).
[0352] 40. Xie, Y. et al. Droplet-based single-cell joint profiling of histone modifications and transcriptomes. Nature Structural & Molecular Biology 30, 1428-1433 (2023).
[0353] 41. Zhu, Q. et al. Genetic evidence that Nkx2.2 and Pdgfra are major determinants of the timing of oligodendrocyte differentiation in the developing CNS. Development 141, 548-555 (2014).
[0354] 42. McLean, C.Y. et al. GREAT improves functional interpretation of cis- regulatory regions. Nature Biotechnology 28, 495-501 (2010).
[0355] 43. Aaltonen, L.A. et al. Pan-cancer analysis of whole genomes. Nature 578, 82- 93 (2020).
[0356] 44. Dixon, J.R. et al. Integrative detection and analysis of structural variation in cancer genomes. Nature Genetics 50, 1388-1398 (2018).
[0357] 45. Harewood, L. et al. Hi-C as a tool for precise detection and characterisation of chromosomal rearrangements and copy number variation in human tumours. Genome Biology 18, 125 (2017).
[0358] 46. Zhu, Y. et al. Oncogenic extrachromosomal DNA functions as mobile enhancers to globally amplify chromosomal transcription. Cancer cell 39, 694-707. e697 (2021).
[0359] 47. Liang, Z. et al. Chromatin- associated RNA Dictates the ecDNA Interactome in the Nucleus. bioRxiv, 2023.2007.2027.550855 (2023).
[0360] 48. Hung, K.L. et al. ecDNA hubs drive cooperative intermolecular oncogene expression. Nature 600, 731-736 (2021).
[0361] 49. Nathanson, D. A. et al. Targeted Therapy Resistance Mediated by Dynamic Regulation of Extrachromosomal Mutant EGFR DNA. Science 343, 72-76 (2014).
[0362] 50. Purshouse, K. et al. Oncogene expression from extrachromosomal DNA is driven by copy number amplification and does not require spatial clustering in glioblastoma stem cells. eLife 11, e80207 (2022).
[0363] 51. Neftel, C. et al. An integrative model of cellular states, plasticity, and genetics for glioblastoma. Cell 178, 835-849. e821 (2019).
[0364] 52. Lancho, O. & Herranz, D. The MYC enhancer-ome: long-range transcriptional regulation of MYC in cancer. Trends in cancer 4, 810-822 (2018).
[0365] 53. Cox, D„ Yuncken, C. & Spriggs, A. MINUTE CHROMATIN BODIES IN MALIGNANT TUMOURS OF CHILDHOOD. The Lancet 286, 55-58 (1965).
[0366] 54. Spriggs, A.I., Boddington, M.M. & Clarke, C.M. Chromosomes of human cancer cells. British medical journal 2, 1431-1435 (1962).
[0367] 55. Mpller, H.D. et al. Circular DNA elements of chromosomal origin are common in healthy human somatic tissue. Nature Communications 9, 1069 (2018).
[0368] 56. Paulsen, T., Kumar, P., Koseoglu, M.M. & Dutta, A. Discoveries of Extrachromosomal Circles of DNA in Normal and Tumor Cells. Trends in Genetics 34, 270- 278 (2018).
[0369] 57. Wu, S. et al. Circular ecDNA promotes accessible chromatin and high oncogene expression. Nature 575, 699-703 (2019).
[0370] 58. Turner, K.M. et al. Extrachromosomal oncogene amplification drives tumour evolution and genetic heterogeneity. Nature 543, 122-125 (2017).
[0371] 59. Fan, X. et al. SMOOTH-seq: single-cell genome sequencing of human cells on a third-generation sequencing platform. Genome Biology 22, 195 (2021).
[0372] 60. Chang, L. et al. Single-cell third-generation sequencing-based multi-omics uncovers gene expression changes governed by ecDNA and structural variants in cancer cells. Clinical and Translational Medicine 13, el 351 (2023).
[0373] 61. Chamorro Gonzalez, R. et al. Parallel sequencing of extrachromosomal circular DNAs and transcriptomes in single cancer cells. Nature Genetics 55, 880-890 (2023).
[0374] 62. Mpller, H.D., Parsons, L., .Iqrgensen. T.S., Botstein, D. & Regenberg, B.Extrachromosomal circular DNA is common in yeast. Proceedings of the National Academy of Sciences 1 12, E31 14-E3122 (2015).
[0375] 63. Koche, R.P. et al. Extrachromosomal circular DNA drives oncogenic genomeremodeling in neuroblastoma. Nature Genetics 52, 29-34 (2020).
[0376] 64. Kim, H. et al. Extrachromosomal DNA is associated with oncogene amplification and poor outcome across multiple cancers. Nature Genetics 52, 891-897 (2020).
[0377] 65. Wu, S., Bafna, V., Chang, H.Y. & Mischel, P.S. Extrachromosomal DNA: an emerging hallmark in human cancer. Annual Review of Pathology: Mechanisms of Disease 17, 367-386 (2022).
[0378] 66. Verhaak, R.G., Bafna, V. & Mischel, P.S. Extrachromosomal oncogene amplification in tumour pathogenesis and evolution. Nature Reviews Cancer 19, 283-288 (2019).
[0379] 67. Wu, S., Bafna, V. & Mischel, P.S. Extrachromosomal DNA (ecDNA) in cancer pathogenesis. Current Opinion in Genetics & Development 66, 78-82 (2021).
[0380] 68. Curtis, E.J., Rose, J.C., Mischel, P.S. & Chang, H.Y. ExtrachromosomalDNA: Biogenesis and Functions in Cancer. Annual Review of Cancer Biology 8, null (2024).
[0381] 69. Chowdhry, S. et al. Abstract 1520: Replication stress and the inability to repair damaged DNA, the potential “Achilles' heel” of ecDNA-i- tumor cells. Cancer Research 82, 1520-1520 (2022).
[0382] 70. Bei, Y. et al. Amplicon structure creates collateral therapeutic vulnerability in cancer. bioRxiv, 2022.2009.2008.506647 (2022).
[0383] 71. Noorani, I., Mischel, P.S. & Swanton, C. Leveraging extrachromosomal DNA to fine-tune trials of targeted therapy for glioblastoma: opportunities and challenges. Nature Reviews Clinical Oncology 19, 733-743 (2022).
[0384] 72. Langmead, B., Trapnell, C., Pop, M. & Salzberg, S.L. Ultrafast and memoryefficient alignment of short DNA sequences to the human genome. Genome Biology 10, R25 (2009).
[0385] 73. Martin, M. Cutadapt removes adapter sequences from high-throughput sequencing reads. EMBnet. journal 17, 10-12 (2011).
[0386] 74. Li, H. Aligning sequence reads, clone sequences and assembly contigs with BWA-MEM. arXiv preprint arXiv: 1303.3997 (2013).
[0387] 75. Open2C et al. Pairtools: from sequencing data to chromosome contacts. bioRxiv, 2023.2002.2013.528389 (2023).
[0388] 76. Abdennur, N. & Mirny, L.A. Cooler: scalable storage for Hi-C data and other genomically labeled arrays. Bioinformatics 36, 31 1-316 (2019).
[0389] 77. Hunter, J.D. Matplotlib: A 2D graphics environment. Computing in science & engineering 9, 90-95 (2007).
[0390] 78. Traag, V.A., Waltman, L. & van Eck, N.J. From Louvain to Leiden: guaranteeing well-connected communities. Scientific Reports 9, 5233 (2019).
[0391] 79. Csardi, G. & Nepusz, T. The igraph software package for complex network research. InterJournal, complex systems 1695, 1-9 (2006).
[0392] 80. Zhou, J. et al. Robust single-cell Hi-C clustering by convolution- and random- walk-based imputation. Proceedings of the National Academy of Sciences 116, 14011-14018 (2019).
[0393] 81. Amemiya, H.M., Kundaje, A. & Boyle, A.P. The ENCODE Blacklist: Identification of Problematic Regions of the Genome. Scientific Reports 9, 9354 (2019).
[0394] 82. Wolf, F.A., Angerer, P. & Theis, F.J. SCANPY: large-scale single-cell gene expression data analysis. Genome Biology 19, 15 (2018).
[0395] 83. Liu, H. et al. DNA methylation atlas of the mouse brain at single-cell resolution. Nature 598, 120-128 (2021).
[0396] 84. Dong, W., Moses, C. & Li, K. in Proceedings of the 20th international conference on World wide web 577-586 (2011).
[0397] 85. Korsunsky, I. et al. Fast, sensitive and accurate integration of single-cell data with Harmony. Nature Methods 16, 1289-1296 (2019).
[0398] 86. Open2C et al. Cooltools: enabling high-resolution Hi-C analysis in Python. BioRxiv, 2022.2010. 2031.514564 (2022).
[0399] 87. Chakraborty, A., Wang, J.G. & Ay, F. dcHiC detects differential compartments across multiple Hi-C datasets. Nature Communications 13, 6827 (2022).
[0400] 88. Shin, H. et al. TopDom: an efficient and deterministic method for identifying topological domains in genomes. Nucleic Acids Research 44, e70-e70 (2015).
[0401] 89. Hounkpe, B.W., Chenou, F., de Lima, F. & De Paula, Erich V. HRT Atlas vl.O database: redefining human and mouse housekeeping genes and candidate reference transcripts by mining massive RNA-seq datasets. Nucleic Acids Research 49, D947-D955 (2020).
[0402] 90. Durand, N.C. et al. Juicer provides a one-click system for analyzing loopresolution Hi-C experiments. Cell systems 3, 95-98 (2016).
[0403] 91 . Yu, M. et al. SnapHiC: a computational pipeline to identify chromatin loops from single-cell Hi-C data. Nature Methods 18, 1056-1059 (2021).
[0404] 92. Gu, Z. & Hiibschmann, D. rGREAT: an R / bioconductor package for functional enrichment on genomic regions. Bioinformatics 39 (2022).
[0405] 93. Wang, X. et al. Genome-wide detection of enhancer-hijacking events from chromatin interaction data in rearranged genomes. Nature methods 18, 661-668 (2021).
[0406] 94. Wang, X., Luan, Y. & Yue, F. EagleC: A deep-learning framework for detecting a full range of structural variations from bulk and single-cell contact maps. Science Advances 8, eabn9215 (2022).
[0407] 95. Nair, V. & Hinton, G.E. in Proceedings of the 27th international conference on machine learning (1CML-10) 807-814 (2010).
[0408] 96. Krizhevsky, A., Sutskever, I. & Hinton, G.E. Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems 25(2012).
[0409] 97. Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I. & Salakhutdinov, R. Dropout: a simple way to prevent neural networks from overfitting. The journal of machine learning research 15, 1929-1958 (2014).
[0410] 98. Hendrycks, D. & Gimpel, K. Gaussian error linear units (gelus). arXiv preprint arXiv: 1606.08415 (2016).
[0411] 99. Paszke, A. et al. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems 32 (2019).
[0412] 100. Kingma, D.P. & Ba, J. Adam: A method for stochastic optimization. arXiv preprintarXiv: 1412.6980 (2014).
[0413] 101. Loshchilov, I. & Hutter, F. Decoupled weight decay regularization. arXiv preprint arXiv: 1711.05101 (2017).
[0414] 102. Bridle, J.S. in Neurocomputing: Algorithms, architectures and applications 227-236 (Springer, 1990).
[0415] 103. Cao, Y. et al. Integrated analysis of multimodal single-cell data with structural similarity. Nucleic Acids Research 50, el21-el21 (2022).
[0416] 104. Stuart, T., Srivastava, A., Madad, S., Lareau, C.A. & Satija, R. Single-cell chromatin state analysis with Signac. Nature Methods 18, 1333-1341 (2021).
[0417] 105. Wolock, S.L., Lopez, R. & Klein, A.M. Scrublet: computational identification of cell doublets in single-cell transcriptomic data. Cell systems 8, 281-291. e289 (2019).
[0418] 106. van Galen, P. et al. Single-Cell RNA-Seq Reveals AML Hierarchies Relevant to Disease Progression and Immunity. Cell 176, 1265-1281.el224 (2019).Example 7: Additional Methods
[0419] Multi-way chromatin interactions analysis
[0420] Multi-way interactions are extracted with ‘pairtools parse2’ function in Pairtools, which is designed to rescue complex chromatin ligation events. After deduplication, only trans interactions and cis interactions with genomic distance >10 kb were retained. Read pairs overlapping with the ENCODE mmlO blacklist regions were removed. For each pairend read on autosomes supporting multi-way contacts, we collected the 10-kb genomic bins containing their anchors. Finally, we defined one pair-end read to support multi-way contacts, if it contacts >=3 unique 10-kb bins (can be both intra-chromosomal and inter-chromosomal).
[0421] To define contact hubs for multi-way interactions in each mouse brain cell type, we followed the below strategy: First, the reference genome (mmlO for mouse cortex) was segmented into consecutive 10-kb bins. Next, a binary indicator is assigned to each 10-kb bin in each single cell, depending on whether it is involved in a multi-way contact. Then, for each 10-kb bin, we calculated the frequency of single cells involving in multi-way contacts among all cells belonging to the same cell type. We excluded the top 1% bins with the highest frequency or bins with sequence mappability score <0.8 when calculating the mean and standard deviation for frequency distribution, then converted the frequency distribution into Z-score. Finally, a 10-kb bin is defined as in multi-way contact hub if its Z-score >1.96.Custom scripts to analyze multi-way interactions hub is available at: github.com / HuMingLab / Multiway.hub.
[0422] Enrichment of chromatin features at multi-way chromatin hubs
[0423] After identifying the multi-way chromatin hubs in each mouse cortex cell type, we performed enrichment analysis against candidate cis-regulatory elements (cCREs), superenhancers and cell-type-specifically expressed genes. Specifically, cCREs are calculated and converted into 10-kb genomic bins from snATAC-seq data92. Super-enhancers (SE) are called using HOMER 'findPeaks -style super' on the identified cCREs within each cell type with default parameters93, and converted into 10-kb genomic bins. For each cell type A involving multi-way contact hub, we calculated its overlap with cCREs or super-enhancers from another cell type (B), and created a 2x2 contingency table. We then performed the Fisher’s exact test to evaluate the significance of enrichment and reported the Log2 odds ratio as enrichment score. We enumerated all pairs of cell types, and compared the Log2 odds ratio in matched (AA) versus unmatched (AB, where fi / l l cell types. We repeated the same enrichment analysis for cell type-specific marker genes from published datasets40, except thatwe used the 10-kb bins overlap with marker genes transcription start site (TSS) for enrichment analysis.
[0424] ecDNA hub analysis
[0425] To generate a background distribution for ecDNA hub analysis, we shuffled chromosome identities for all trans interactions in each cell, and calculated the hub index using R package ineq based on the shuffled contact matrices. P value is calculated from Wilcoxon signed-rank test.
[0426] To generate a background distribution for ecDNA hub analysis, we shuffled chromosome identities for all trans interactions in each cell, and calculated the hub index using R package ineq based on the shuffled contact matrices. P value is calculated from Wilcoxon signed-rank test.P EMBODIMENTS
[0427] P Embodiment 1. A method of cell profiling the method comprising: (a) providing or obtaining a cross-linked chromatin complex of a cell or a cellular nucleus; (b) removing one or more histone proteins from the cross-linked chromatin complex; (c) digesting the cross-linked chromatin complex or portions thereof after (b), thereby generating a digested chromatin; (d) ligating the digested chromatin, thereby generating a ligated chromatin; (e) fragmenting the ligated chromatin, thereby generating a fragmented chromatin; (f) encapsulating the fragmented chromatin in a droplet; (g) barcoding the fragmented chromatin with a barcode, thereby generating a barcoded fragmented chromatin; and (h) sequencing the barcoded fragmented chromatin.
[0428] P Embodiment 2. The method of P embodiment 1 , further comprising determining a spatial proximity between fibers of the cross-linked chromatin.
[0429] P Embodiment 3. The method of P embodiment 2, wherein the spatial proximity is determined across the genome of the cell or the cellular nucleus.
[0430] P Embodiment 4. The method of any one of P embodiments 1-3, wherein the barcode is specific to the cell or the cellular nucleus.
[0431] P Embodiment 5. The method of any one of P embodiments 1-4, wherein barcoding comprises providing a barcode bound to a bead in the droplet and capturing the fragmented chromatin on the barcode.
[0432] P Embodiment 6. The method of P embodiment 5, wherein the bead is degradable upon application of a stimulus.
[0433] P Embodiment 7. The method of P embodiment 5, wherein the droplet comprises a GEM® by 10X GENOMICS™.
[0434] P Embodiment 8. The method of any one of P embodiments 1-7, wherein the droplet is generated with the aid of a droplet microfluidic device.
[0435] P Embodiment 9. The method of any one of P embodiments 1-8, wherein (f) and (g) are performed using a platform sold under trademark CHROMIUM NEXT GEM CHIP H® by 10X GENOMICS™.
[0436] P Embodiment 10. The method of any one of P embodiments 1-9, wherein the method comprises performing single-cell chromatin conformation profiling on a population of cells comprising the cell or the cellular nucleus.
[0437] P Embodiment 11. The method of P embodiment 10, wherein a chromatin architecture of the cell or the cellular nucleus is mapped on a single cell level in the population of cells comprising the cell or the cellular nucleus.
[0438] P Embodiment 12. The method of any one of P embodiments 1-11, wherein the cell or the cellular nucleus is derived from a tissue of a subject, and wherein the method further comprises deriving and analyzing a population of cells from the tissue.
[0439] P Embodiment 13. The method of P embodiment 12, further comprising diagnosing the subject with a condition.
[0440] P Embodiment 14. The method of P embodiment 12 or 13, further comprising prescribing a medicine, an action, or a combination thereof to the subject, based on results of the cell profiling.
[0441] P Embodiment 15. The method of any one of P embodiments 10-14, further comprising detecting a copy number variation (CNV), a structural variation (SV), an extrachromosomal DNA (ecDNA), or any combination thereof in the population of cells.
[0442] P Embodiment 16. The method of P embodiment 15, further comprising revealing clonal dynamics or an oncogenic event in the population of cells.
[0443] P Embodiment 17. The method of any one of P embodiments 1-16, further comprising analyzing the transcriptome of the cell or the cellular nucleus.
[0444] P Embodiment 18. The method of P embodiment 17, further comprising identifying an architecture of the cross-linked chromatin complex, obtaining a transcriptomic dataset from the cell or the cellular nucleus, and identifying a correlation between the architecture and the transcriptomic dataset.
[0445] P Embodiment 19. The method of any one of P embodiments 1-18, further comprising recording one or more data sets from the cell screening.
[0446] P Embodiment 20. The method of P embodiment 19, further comprising training a computational model using the one or more datasets.
[0447] P Embodiment 21. The method of P embodiment 20, wherein the computational model comprises a Machine Learning (ML) model, an Artificial Intelligence (Al) model, Deep Learning, or any combination thereof.
[0448] P Embodiment 22. The method of any one of P embodiments 1-21, further comprising fixating the cell or the cellular nucleus prior to (a).
[0449] P Embodiment 23. The method of any one of P embodiments 1-22, the method comprises single cell or single nucleus sequencing.
[0450] P Embodiment 24. The method of P embodiment 23, further comprising constructing a single cell or single nucleus library.
[0451] P Embodiment 25. The method of any one of P embodiments 1-24, further comprising analysis of chromatin domains, analysis of chromatin loops, or both.
[0452] P Embodiment 26. The method of any one of P embodiments 1-25, further comprising identifying structural variations (SV) in the cross-linked chromatin complex.
[0453] P Embodiment 27. The method any one of P embodiments 1 -26, further comprising single cell or single nucleus genotyping.
[0454] P Embodiment 28. The method of any one of P embodiments 1-27, further comprising analysis of gene expression levels and copy numbers of ecDNA in the cell or the cellular nucleus.EMBODIMENTS
[0455] Embodiment 1. A method of amplifying a chromosomal fragment DNA sequence, the method comprising: (a) contacting a plurality of cell nuclei with a DNA crosslinking agent wherein each of the plurality of cell nuclei comprises chromosomal DNA thereby forming crosslinked chromosomal DNA within each of the plurality of cell nuclei;(b) contacting the crosslinked chromosomal DNA with a DNA endonuclease thereby forming cleaved chromosomal DNA within each of the plurality of cell nuclei, wherein each of the crosslinked chromosomal DNAs comprises a plurality of double stranded cleaved sites; (c) contacting the cleaved chromosomal DNA with a ligase thereby forming ligated chromosomal DNA within each of the plurality of cell nuclei, wherein each of the ligated chromosomal DNAs comprises at least one ligated site that forms a non-endogenous ligationDNA sequence, wherein the non-endogenous ligation DNA sequence is not present in the endogenous chromosomal DNA within each of the plurality of cell nuclei; (d) contacting the ligated chromosomal DNA with a transposase, a 5' hybridization DNA sequence and a 3’ hybridization DNA sequence thereby forming a plurality of transposed chromosomal DNA sequences within each of the plurality of cell nuclei, wherein each of the transposed chromosomal DNA sequences comprises from 5' to 3' a first DNA hybridization sequence, a chromosomal fragment DNA sequence, and a second DNA hybridization sequence; (e) separating each of the plurality of cell nuclei into a separate compartment, wherein each compartment contains one cell nucleus and an identifying library nucleic acid sequence, wherein the identifying library nucleic acid sequence comprises from 5’ to 3' a forward primer sequence, a barcode sequence, and a 5' complementary hybridization sequence; (f) lysing each cell nucleus in each compartment and allowing the first DNA hybridization sequence to hybridize with the 5' complementary hybridization sequence to form a first hybridized template DNA sequence; (g) contacting the first hybridized template DNA with a DNA polymerase under conditions conducive to amplification and producing a first amplified chromosomal DNA comprising from 5' to 3’ the forward primer sequence or complement thereof, the barcode sequence or complement thereof, the 5’ complementary hybridization sequence or complement thereof, the chromosomal fragment DNA sequence or complement thereof, and the second DNA hybridization sequence or complement thereof, thereby amplifying the chromosomal fragment DNA sequence.
[0456] Embodiment 2. The method of embodiment 1, further comprising: (h) contacting the first amplified chromosomal DNA with a second library nucleic acid sequence comprising from 5' to 3' a 3' complementary hybridization sequence and a reverse primer sequence and allowing the second DNA hybridization sequence to hybridize with the 3' complementary hybridization sequence thereby forming a second hybridized template DNA sequence; and (i) contacting the second hybridized template DNA sequence with a DNA polymerase under conditions conducive to amplification, thereby amplifying the first amplified chromosomal DNA to produce a second amplified chromosomal DNA comprising from 5’ to 3’ the forward primer sequence or complement thereof, the barcode sequence or complement thereof, the 5’ complementary hybridization sequence or complement thereof, the chromosomal fragment DNA sequence or complement thereof, the second DNA hybridization sequence or complement thereof, the 3' complementary hybridization sequence or complement thereof and the reverse primer sequence or complement thereof.
[0457] Embodiment 3. The method of embodiment 1 or 2, wherein each chromosomal fragment DNA sequence is from about 200 nucleotides to about 1000 nucleotides in length.
[0458] Embodiment 4. The method of any one of embodiments 1-3, wherein each separate compartment is a droplet or a liquid volume within a well of a microwell plate.
[0459] Embodiment 5. The method of any one of embodiments 1-4, wherein the identifying library nucleic acid sequence is bound to a solid support.
[0460] Embodiment 6. The method of any one of embodiments 1-5, wherein the solid support is the surface of a well, the surface of a microwell, or the surface of a solid phase particle.
[0461] Embodiment 7. The method of embodiment 6, wherein the solid phase particle is a gel bead or a polymer bead.
[0462] Embodiment 8. The method of any one of embodiments 1-7, wherein the barcode sequence in each separate compartment is unique.
[0463] Embodiment 9. The method of any one of embodiments 1-8, wherein the DNA endonuclease is a DpnII restriction endonuclease, an Mbol restriction endonuclease, an Nlalll restriction endonuclease, a Ddel restriction endonuclease, or a DpnII restriction.
[0464] Embodiment 10. The method of any one of embodiments 1 -9, wherein the ligase is a T4 DNA ligase.
[0465] Embodiment 11. The method of any one of embodiments 1-10, the transposase is a Tn5 transposase.
[0466] Embodiment 12. The method of any one of embodiments 1-11, further comprising amplifying a plurality of different chromosomal fragment DNA sequences from each of the plurality of cell nuclei thereby producing a plurality of amplified chromosomal fragment DNA sequences from each of the plurality of cell nuclei.
[0467] Embodiment 13. The method of embodiment 12, further comprising sequencing the plurality of amplified chromosomal fragment DNA sequences.
[0468] Embodiment 14. The method of embodiment 13, wherein the sequencing is a next-generation sequencing method.
[0469] Embodiment 15. The method of embodiment 13 or 14, wherein the sequencing is a sequencing-by-synthesis (SBS) method.
[0470] Embodiment 16. The method of any one of embodiments 13-15, wherein the method further comprising: (z) aligning the sequence of each of the amplified chromosomal fragment DNA sequences to a reference genome; and (z'r) determining a relative abundance ofeach of the amplified chromosomal fragment DNA sequences including non-endogenous ligation DNA sequences within the amplified chromosomal fragment DNA sequences.
[0471] Embodiment 17. The method of embodiment 16, wherein the aligning and determining are computer implemented.
[0472] Embodiment 18. The method of embodiment 16 or 17, further comprising computing a three-dimensional structure of the chromatin of the plurality of cell nuclei based on the aligning and the determining.
[0473] Embodiment 19. The method of any one of embodiments 16-18, further comprising detecting a copy number variation within the chromosomal DNA of each of the plurality of cell nuclei, a structural variation within the chromosomal DNA of each of the plurality of cell nuclei, and / or extrachromosomal DNA within the chromosomal DNA of each of the plurality of cell nuclei based on the aligning and determining.
[0474] Embodiment 20. The method of any one of embodiments 1-19, wherein each of the compartments of step (e) further comprise a plurality of endogenous cellular mRNA, a plurality of messenger RNA (mRNA) capture sequences and a reverse transcriptase, wherein each mRNA capture sequence comprises a second barcode sequence and a poly-adenine (poly-A) tail complementary sequence; wherein the method further comprises allowing the plurality of poly-A tail complementary sequences to hybridize with a plurality of the endogenous mRNAs to form a plurality of hybridized template mRNA sequences; and allowing the reverse transcriptase to contact the plurality of hybridized template mRNA sequences under conditions conducive to reverse transcription, thereby forming a plurality of barcoded transcript sequences.
[0475] Embodiment 21. The method of embodiment 20, further comprising detecting the barcoded transcript sequences.
[0476] Embodiment 22. The method of embodiment 21, further comprising quantifying barcoded transcript sequences.
[0477] Embodiment 23. The method of any one of embodiments 1-22, wherein the chromosomal DNA fragment sequence comprises extrachromosomal DNA (ecDNA).
[0478] Embodiment 24. The method of any one of embodiments 1-23, wherein the plurality of cell nuclei comprise a plurality of cell nuclei from a heterogenous population of cells.
[0479] Embodiment 25. The method of any one of embodiments 1 -24, wherein the heterogenous population of cells comprises a cancer cell.
[0480] Embodiment 26. The method of any one of embodiments 1 -25, wherein the heterogenous population of cells is a heterogenous population of human cells.
[0481] Embodiment 27. The method of any one of embodiments 1-26, wherein the plurality of cell nuclei are isolated from cancer cells or wherein the plurality of cell nuclei are within cancer cells prior to lysing.
Claims
CLAIMSWHAT IS CLAIMED IS:
1. A method of amplifying a chromosomal fragment DNA sequence, the method comprising:(a) contacting a plurality of cell nuclei with a DNA crosslinking agent wherein each of the plurality of cell nuclei comprises chromosomal DNA thereby forming crosslinked chromosomal DNA within each of the plurality of cell nuclei;(b) contacting the crosslinked chromosomal DNA with a DNA endonuclease thereby forming cleaved chromosomal DNA within each of the plurality of cell nuclei, wherein each of the crosslinked chromosomal DNAs comprises a plurality of double stranded cleaved sites;(c) contacting the cleaved chromosomal DNA with a ligase thereby forming ligated chromosomal DNA within each of the plurality of cell nuclei, wherein each of the ligated chromosomal DNAs comprises at least one ligated site that forms a non- endogenous ligation DNA sequence, wherein the non-endogenous ligation DNA sequence is not present in the endogenous chromosomal DNA within each of the plurality of cell nuclei;(d) contacting the ligated chromosomal DNA with a transposase, a 5' hybridization DNA sequence and a 3' hybridization DNA sequence thereby forming a plurality of transposed chromosomal DNA sequences within each of the plurality of cell nuclei, wherein each of the transposed chromosomal DNA sequences comprises from 5' to 3' a first DNA hybridization sequence, a chromosomal fragment DNA sequence, and a second DNA hybridization sequence;(e) separating each of the plurality of cell nuclei into a separate compartment, wherein each compartment contains one cell nucleus and an identifying library nucleic acid sequence, wherein the identifying library nucleic acid sequence comprises from 5’ to 3' a forward primer sequence, a barcode sequence, and a 5’ complementary hybridization sequence;(f) lysing each cell nucleus in each compartment and allowing the first DNA hybridization sequence to hybridize with the 5' complementary hybridization sequence to form a first hybridized template DNA sequence;(g) contacting the first hybridized template DNA with a DNA polymerase under conditions conducive to amplification and producing a first amplified chromosomal DNA comprising from 5’ to 3' the forward primer sequence or complement thereof, thebarcode sequence or complement thereof, the 5' complementary hybridization sequence or complement thereof, the chromosomal fragment DNA sequence or complement thereof, and the second DNA hybridization sequence or complement thereof, thereby amplifying the chromosomal fragment DNA sequence.
2. The method of claim 1, further comprising:(h) contacting the first amplified chromosomal DNA with a second library nucleic acid sequence comprising from 5' to 3’ a 3' complementary hybridization sequence and a reverse primer sequence and allowing the second DNA hybridization sequence to hybridize with the 3' complementary hybridization sequence thereby forming a second hybridized template DNA sequence; and(i) contacting the second hybridized template DNA sequence with a DNA polymerase under conditions conducive to amplification, thereby amplifying the first amplified chromosomal DNA to produce a second amplified chromosomal DNA comprising from 5’ to 3’ the forward primer sequence or complement thereof, the barcode sequence or complement thereof, the 5' complementary hybridization sequence or complement thereof, the chromosomal fragment DNA sequence or complement thereof, the second DNA hybridization sequence or complement thereof, the 3' complementary hybridization sequence or complement thereof and the reverse primer sequence or complement thereof.
3. The method of claim 1, wherein each chromosomal fragment DNA sequence is from about 200 nucleotides to about 1000 nucleotides in length.
4. The method of claim 1, wherein each separate compartment is a droplet or a liquid volume within a well of a microwell plate.
5. The method of claim 1, wherein the identifying library nucleic acid sequence is bound to a solid support.
6. The method of claim 1, wherein the solid support is the surface of a well, the surface of a microwell, or the surface of a solid phase particle.
7. The method of claim 6, wherein the solid phase particle is a gel bead or a polymer bead.
8. The method of claim 1, wherein the barcode sequence in each separate compartment is unique.
9. The method of claim 1 , wherein the DNA endonuclease is a DpnII restriction endonuclease, an Mbof restriction endonuclease, an Nlalll restriction endonuclease, a Ddel restriction endonuclease, or a DpnII restriction endonuclease.
10. The method of claim 1 , wherein the ligase is a T4 DNA ligase.
11. The method of claim 1 , the transposase is a Tn5 transposase.
12. The method of claim 1, further comprising amplifying a plurality of different chromosomal fragment DNA sequences from each of the plurality of cell nuclei thereby producing a plurality of amplified chromosomal fragment DNA sequences from each of the plurality of cell nuclei.
13. The method of claim 12, further comprising sequencing the plurality of amplified chromosomal fragment DNA sequences.
14. The method of claim 13, wherein the sequencing is a next-generation sequencing method.
15. The method of claim 13, wherein the sequencing is a sequencing -bysynthesis (SBS) method.
16. The method of claim 13, wherein the method further comprising:(z) aligning the sequence of each of the amplified chromosomal fragment DNA sequences to a reference genome; and(z7) determining a relative abundance of each of the amplified chromosomal fragment DNA sequences including non-endogenous ligation DNA sequences within the amplified chromosomal fragment DNA sequences.
17. The method of claim 16, wherein the aligning and determining are computer implemented.
18. The method of claim 16, further comprising computing a three- dimensional structure of the chromatin of the plurality of cell nuclei based on the aligning and the determining.
19. The method of claim 16, further comprising detecting a copy number variation within the chromosomal DNA of each of the plurality of cell nuclei, a structural variation within the chromosomal DNA of each of the plurality of cell nuclei, and / or extrachromosomal DNA within the chromosomal DNA of each of the plurality of cell nuclei based on the aligning and determining.
20. The method of claim 1, wherein each of the compartments of step (e) further comprise a plurality of endogenous cellular mRNA, a plurality of messenger RNA (mRNA) capture sequences and a reverse transcriptase, wherein each mRNA capture sequence comprises a second barcode sequence and a poly-adenine (poly- A) tail complementary sequence;wherein the method further comprises allowing the plurality of poly- A tail complementary sequences to hybridize with a plurality of the endogenous mRNAs to form a plurality of hybridized template mRNA sequences; and allowing the reverse transcriptase to contact the plurality of hybridized template mRNA sequences under conditions conducive to reverse transcription, thereby forming a plurality of barcoded transcript sequences.
21. The method of claim 20, further comprising detecting the barcoded transcript sequences.
22. The method of claim 21 , further comprising quantifying barcoded transcript sequences.
23. The method of claim 1, wherein the chromosomal DNA fragment sequence comprises extrachromosomal DNA (ecDNA).
24. The method of claim 1, wherein the plurality of cell nuclei comprise a plurality of cell nuclei from a heterogenous population of cells.
25. The method of claim 1, wherein the heterogenous population of cells comprises a cancer cell.
26. The method of claim 1 , wherein the heterogenous population of cells is a heterogenous population of human cells.
27. The method of claim 1, wherein the plurality of cell nuclei are isolated from cancer cells or wherein the plurality of cell nuclei are within cancer cells prior to lysing.
Citation Information
Patent Citations
Single tube bead-based DNA co-barcoding for accurate and cost-effective sequencing, haplotyping, and assembly
US20210115595A1
Single cell whole genome libraries and combinatorial indexing methods of making thereof
US20230323426A1