Methods for identifying and using extrachromosomal DNA

By employing chromatin interaction analysis techniques, the method effectively identifies ecDNA and its interaction with target genes in cancer cells, addressing the limitations of existing methods and offering insights into cancer biology and therapy.

JP7672345B2Active Publication Date: 2025-05-07JACKSON LAB THE
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2021564693
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-04-30
Filing Date
2020-04-29
Publication Date
2025-05-07
Estimated Expiration
2040-04-29

AI Technical Summary

Technical Problem

Current methods are inadequate for identifying extrachromosomal circular DNA (ecDNA) and assessing its role in modulating target gene transcription in cancer cells.

Method used

A method involving chromatin interaction analysis, such as ChIA-PET or Hi-C, to detect physical interactions between nonlinear ecDNA and linear chromosomes, allowing for the identification of ecDNA and its interaction with target genes.

Benefits of technology

This approach enables the accurate identification of ecDNA and its modulation of target gene transcription, providing insights into cancer progression and potential therapeutic targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007672345000002
    Figure 0007672345000002
  • Figure 0007672345000003
    Figure 0007672345000003
  • Figure 0007672345000004
    Figure 0007672345000004
Patent Text Reader

Abstract

The present invention encompasses, in part, methods for identifying extrachromosomal circular DNA (ecDNA), as well as methods for identifying and assessing interactions between ecDNA and oncogene transcription.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] Related Applications This application claims the benefit under 35 U.S.C. § 119(e) of U.S. Provisional Application No. 62 / 840,735, filed April 30, 2019, the entire contents of which are incorporated herein by reference in their entirety.

[0002] Government involvement This invention was made with Government support under Grant Nos. NCI P30CA034196, U54 DK107967, UM1 HG009409, and R01 CA190121 awarded by the National Institutes of Health (NHI). The Government has certain rights in this invention.

[0003] FIELD OF THEINVENTION The present invention relates in some embodiments to methods for identifying extrachromosomal circular DNA (ecDNA) and using ecDNA to assess transcription of target genes. [Background technology]

[0004] Extrachromosomal DNA (ecDNA) is an extrachromosomal circular chromatin element [Cox, D. et al. (1965) Lancet 1, 55-58, Spriggs, AI et al. (1962) Br Med J 2, 1431-1435]. Initially described as “double microchromatin bodies” in cell karyotypes by microscopic imaging [Cox, D. et al. (1965) Lancet 1, 55-58], ecDNA was thought to be a form of gene amplification associated with in vitro drug resistance [Spriggs, AI et al. (1962) Br Med J 2, 1431-1435, deCarvalho, AC et al. (2018) Nat Genet 50, 708-717, Nathanson, DA et al. (2014) Science 343, 72-76]. More recently, ecDNA has been found to be common in primary cancers [Xu, K. et al. (2018) Acta Neuropathol, doi:10.1007 / s00401-018-1912-1] and to constitute a bona fide mechanism and adaptive reservoir for the amplification of oncogenes [Turner, KM et al. (2017) Nature 543, 122-125]. ecDNA can be derived through chromosomal genome shattering events and, as a result, can consist of hundreds of DNA segments [Ma, K. et al. (2012) Int J Mol Sci 13, 11974-11999]. Their accumulation in cancer cells provides a competitive advantage in response to selective pressures in the tumor microenvironment and in response to cytotoxic therapeutic agents [Spriggs, AI et al. (1962) Br Med J 2, 1431-1435, Kohl, NE et al. (1983) Cell 35, 359-367]. Rapid fluctuations in ecDNA levels resulting from disruption of the inheritance pattern [Spriggs, AI et al. (1962) Br Med J 2, 1431-1435] likely contribute to tumor development mechanisms.Glioblastoma (GBM) is an aggressive brain tumor in which ecDNA is frequently observed [Rausch, T. et al. (2012) Cell 148, 59-71, Xue, Y. et al. (2017) Nat Med 23, 929-937]. Previous analysis of a set of unique GBM-derived neurosphere cultures using standard genomic approaches detected multiple ecDNAs harboring oncogenes, including EGFR, MYC, and CDK4 [Spriggs, AI et al. (1962) Br Med J 2, 1431-1435]. Although the presence of ecDNAs and their structural information are beginning to be characterized by these standard approaches, the mechanisms by which ecDNAs are deployed to modulate cancer progression and contribute to cancer drug resistance remain unknown. Summary of the Invention [Means for solving the problem]

[0005] According to an aspect of the present invention, a method of identifying extrachromosomal DNA (ecDNA) in a cell is provided, the method comprising: (a) detecting a chromatin interaction between a nonlinear DNA molecule and at least one linear chromosome of a chromosome pair, the interaction comprising contact between the nonlinear DNA molecule and at least one linear chromosome of the chromosome pair; and (i) the presence of a significantly higher frequency of detected chromatin interactions in the cell, (ii) contact between the nonlinear DNA molecule and at least one linear chromosome of each of the chromosome pair in the cell, and (iii) an increase in the average copy number per cell of the nonlinear DNA molecule in a plurality of cells over time identifies the nonlinear DNA molecule as ecDNA. In some embodiments, the method also comprises determining the frequency of the detected chromatin interactions. In certain embodiments, the method also comprises determining the size of the nonlinear DNA. In some embodiments, the method also comprises determining the copy number of the nonlinear DNA molecule in the cell. In some embodiments, the method also comprises determining the average copy number per cell of the nonlinear DNA molecule in a plurality of cells at a first time point, and comparing the determined average to the average copy number per cell of the nonlinear DNA of a control. In certain embodiments, the average copy number per cell of the control nonlinear DNA molecule is the average copy number per cell of the nonlinear DNA molecule determined in a plurality of cells at different time points. In some embodiments, the method also includes determining the sequence of at least a portion of the nonlinear DNA. In certain embodiments, the method also includes identifying the presence of an oncogene sequence in the determined sequence. In some embodiments, the cell is a cancer cell. In some embodiments, the cell is a precancerous cell. In some embodiments, the cell is obtained from a plurality of cells including a cancer cell. In certain embodiments, the cell is obtained from a subject. In some embodiments, the subject is at least one of diagnosed with cancer, suspected of having cancer, and at risk of having cancer. In some embodiments, the cell is obtained from a cell culture.In certain embodiments, the cell is a vertebrate cell, hi some embodiments, the cell is a mammalian cell, optionally a human cell.

[0006] According to another aspect of the present invention, a method for identifying oncogenes modulated by ecDNA is provided, comprising: (a) detecting an interaction between ecDNA and one or more target genes located in a cell, the detection comprising directly measuring chromatin interaction between ecDNA and a control element of one or more target genes; (b) identifying one or more of the target genes involved in the interaction, the transcription of which is modulated by the interaction detected in step (a); and (c) determining whether one or more of the target genes identified in step (b) are oncogenes, wherein determining the target genes as oncogenes identifies the oncogenes as oncogenes modulated by ecDNA. In some embodiments, the identification in step (b) comprises measuring the level of transcription of the identified target genes and comparing the measured level with the transcription level of the target genes of a control. In certain embodiments, the modulation of transcription is an increase in transcription. In some embodiments, the means of detection in step (a) comprises ChIA-PET. In some embodiments, the means of detection in step (a) comprises Hi-C. In some embodiments, the control element comprises a promoter of the target gene. In certain embodiments, the target gene is located on a linear chromosome. In some embodiments, the target gene is located in an ecDNA. In certain embodiments, the target gene is located in a second ecDNA. In some embodiments, the cell is a cancer cell. In some embodiments, the cell is a precancerous cell. In certain embodiments, the cell is obtained from a plurality of cells including a cancer cell. In some embodiments, the cell is obtained from a subject. In some embodiments, the subject has been diagnosed with cancer, is suspected of having cancer, and / or is at risk of having cancer. In some embodiments, the cell is obtained from a cell culture. In certain embodiments, the cell is a vertebrate cell.In some embodiments, the cells are mammalian cells, optionally human cells.

[0007] According to another aspect of the present invention, a method for determining the oncogene status of a cancer is provided, comprising: (a) identifying oncogenes modulated by ecDNA in a cancer cell; and (b) determining one or more of the levels and effects of ecDNA modulation of oncogenes as a determination of the oncogene status of the cancer. In some embodiments, the identifying means in step (a) comprises: (i) detecting an interaction between ecDNA and one or more target genes located in the DNA of a cancer cell, wherein the detecting comprises directly measuring chromatin interactions between ecDNA and a control element of one or more target genes; (ii) identifying the target genes in the detected interactions, the transcription of the one or more target genes being modulated by the interaction detected in step (i); and (iii) determining whether one or more of the target genes identified in step (ii) are oncogenes, wherein determining the target genes as oncogenes identifies the oncogenes as oncogenes modulated by ecDNA in a cancer cell. In certain embodiments, the identifying in step (ii) comprises measuring the level of transcription of the identified target gene and comparing the measured level with the transcription level of a control target gene. In some embodiments, the modulation of transcription is an increase in transcription. In some embodiments, the means of detection in step (i) comprises a ChIA-PET method. In some embodiments, the means of detection in step (i) comprises a Hi-C method. In certain embodiments, the control element comprises a promoter of the target gene. In some embodiments, activation of the promoter increases transcription of the target gene. In some embodiments, the target gene is located on a linear chromosome. In certain embodiments, the target gene is located in an ecDNA. In some embodiments, the target gene is located in a second ecDNA. In some embodiments, the cell is a cancer cell. In some embodiments, the cell is a precancerous cell.In certain embodiments, the cell is obtained from a plurality of cells including cancer cells. In some embodiments, the cell is obtained from a subject. In some embodiments, the subject is at least one of diagnosed with cancer, suspected of having cancer, and at risk of having cancer. In some embodiments, the cell is obtained from a cell culture. In some embodiments, the cell is a vertebrate cell. In certain embodiments, the cell is a mammalian cell, optionally a human cell. In some embodiments, the means for detecting one or more of the levels and effects of ecDNA modulation of the oncogene comprises directly measuring the interchromosomal chromatin contact frequency between the modulating ecDNA and the modulated oncogene. In some embodiments, the means for detecting one or more of the levels and effects of ecDNA modulation of the oncogene comprises determining the level of transcription of the oncogene modulated by one or more ecDNAs, and the level of transcription determines the oncogene status of the cancer. In certain embodiments, the method also includes repeating steps (a) and (b) in cancer cells obtained from a second plurality of cells comprising cancer, and comparing one or more levels or effects detected in the cancer cells obtained from the first plurality of cells with the levels or effects detected in the cancer cells obtained from the second plurality of cells, respectively, and a difference in one or both of the levels and effects indicates a change in the oncogene status of the cancer. In some embodiments, the method also includes contacting the second plurality of cancer cells with a candidate therapeutic agent after determining the oncogene status of the cancer of the first plurality of cancer cells and before determining the oncogene status of the second plurality of cancer cells, and determining the effect of contacting with the candidate therapeutic agent on the oncogene status of the second plurality of cancer cells. In some embodiments, the first and second plurality of cells are obtained from a subject.In certain embodiments, the subject is one or more of diagnosed with cancer, suspected to have cancer, and at risk of having cancer. In some embodiments, the first and second plurality of cells are obtained from cell culture. In some embodiments, the cancer cells are mammalian cells, and optionally human cells. In certain embodiments, the first and second plurality of cells include cancer cells. In some embodiments, the method also includes a step of assisting in the selection of a cancer treatment based at least in part on the determined oncogene status of the cancer. In certain embodiments, the method also includes a step of identifying one or more additional oncogenes modulated by one or more additional ecDNAs in the cancer cells, and a step of determining one or more of the level and effect of ecDNA modulation of the identified one or more additional oncogenes as a determination of the oncogene status of the cancer.

[0008] According to another aspect of the present invention, there is provided a method for identifying extrachromosomal DNA (ecDNA) in a cell, comprising the steps of directly detecting physical interactions between a nonlinear DNA molecule and at least one linear chromosome using a chromatin interaction analysis method, the analysis method comprising: (a) contacting the cell with a fixative and performing chromatin proximity ligation on DNA isolated from the cell; (b) performing chromatin immunoprecipitation; (c) generating a library from the DNA immunoprecipitated in step (b); (d) sequencing the library generated in step (c) to generate sequencing data; and (e) analyzing the sequencing data generated in step (d) to detect ecDNA. In certain embodiments, steps (a)-(c) are performed using ChIA-PET or HI-C. In some embodiments, step (b) comprises the use of an anti-RNAPII antibody in the chromatin immunoprecipitation. In some embodiments, steps (c) and (d) comprise the use of paired-end tags and high-throughput sequencing, respectively. In certain embodiments, the method also includes determining the frequency of the detected physical interactions. In some embodiments, the method also includes determining the size of the nonlinear DNA molecule. In some embodiments, the method also includes determining the copy number of the nonlinear DNA molecule in the cell. In some embodiments, the method also includes determining the average copy number per cell of the nonlinear DNA molecule in a plurality of cells at a first time point, and comparing the determined average to the average copy number per cell of a control nonlinear DNA molecule. In certain embodiments, the average copy number per cell of the control nonlinear DNA molecule is the average copy number per cell of the nonlinear DNA molecule determined in a plurality of cells at different time points. In some embodiments, the method also includes determining the sequence of at least a portion of the nonlinear DNA molecule. In some embodiments, the method also includes identifying the presence of a tumor gene sequence in the determined sequence. In certain embodiments, the cell is a cancer cell.In some embodiments, the cell is a precancerous cell. In some embodiments, the cell is obtained from a plurality of cells including cancer cells. In certain embodiments, the cell is obtained from a subject. In some embodiments, the subject has been diagnosed with, is suspected of having, and / or is at risk of having cancer. In certain embodiments, the cell is obtained from a cell culture. In some embodiments, the cell is a vertebrate cell. In some embodiments, the cell is a mammalian cell, optionally a human cell. In certain embodiments, the method also includes subjecting a plurality of cells to steps (a)-(e).

[0009] According to yet another aspect of the present invention, a method for selecting a treatment for reducing cancer in a subject is provided, comprising identifying the presence of one or more specific ecDNA-modulated oncogenes in cancer cells obtained from the subject, and selecting one or more treatments based on the identified ecDNA-modulated oncogenes. In some embodiments, the identifying step includes any of the above-mentioned methods for identifying extrachromosomal DNA (ecDNA) in a cell, the method including (a) detecting chromatin interactions between a nonlinear DNA molecule and at least one linear chromosome of a chromosome pair, the interaction including contact between the nonlinear DNA molecule and at least one linear chromosome of the chromosome pair, and the presence of (i) a significantly higher frequency of detected chromatin interactions in the cell, (ii) contact between the nonlinear DNA molecule and at least one linear chromosome of each of the chromosome pair in the cell, and (iii) an increase in the average copy number per cell of the nonlinear DNA molecule over time in a plurality of cells identifies the nonlinear DNA molecule as ecDNA. In some embodiments, the identifying step includes any of the embodiments of the aforementioned methods of identifying extrachromosomal DNA (ecDNA) in a cell, the method including the step of directly detecting a physical interaction between a nonlinear DNA molecule and at least one linear chromosome using chromatin interaction analysis, the analysis method including the steps of (a) contacting the cell with a fixative and performing chromatin proximity ligation on DNA isolated from the cell; (b) performing chromatin immunoprecipitation; (c) generating a library from the DNA immunoprecipitated in step (b); (d) sequencing the library generated in step (c) to generate sequencing DNA; and (e) analyzing the sequencing data generated in step (d) to detect ecDNA. [Brief description of the drawings]

[0010] [Figure 1A]Figure 1A-B are diagrams and traces showing that ecDNA signatures can be distinguished by the distribution of trans-chromosomal interaction frequencies (nsTIFs) across 23 chromosomes. Figure 1A provides a circos plot of trans-interactions mediated by ecDNA regions across all 23 chromosomes in HF-3016 and HF-3177 ecDNA(+) cell lines. Strong connections between ecMYC, ecEGFR, and ecCDK4 regions are shown. Figure 1B shows the distribution of genome-wide normalized summed TIFs (nsTIFs) with 50 Kb bin sizes in ecDNA(+) HF-3016 and HF-3177 cell lines, and ecDNA(-) HF-3035 line. Elevated nsTIFs are observed at chromosomes 7, 8, and 12. The distribution of nsTIFs along the entire chromosomes 7, 8 and 12 is shown below, and the regions with elevated nsTIF levels match well with known ecEGFR, ecMYC and ecCDK4 regions. [Figure 1B-1] Figure 1A-B are diagrams and traces showing that ecDNA signatures can be distinguished by the distribution of trans-chromosomal interaction frequencies (nsTIFs) across 23 chromosomes. Figure 1A provides a circos plot of trans-interactions mediated by ecDNA regions across all 23 chromosomes in HF-3016 and HF-3177 ecDNA(+) cell lines. Strong connections between ecMYC, ecEGFR, and ecCDK4 regions are shown. Figure 1B shows the distribution of genome-wide normalized summed TIFs (nsTIFs) with 50 Kb bin sizes in ecDNA(+) HF-3016 and HF-3177 cell lines, and ecDNA(-) HF-3035 line. Elevated nsTIFs are observed at chromosomes 7, 8, and 12. The distribution of nsTIFs along the entire chromosomes 7, 8 and 12 is shown below, and the regions with elevated nsTIF levels match well with known ecEGFR, ecMYC and ecCDK4 regions. [Figure 1B-2]Figure 1A-B are diagrams and traces showing that ecDNA signatures can be distinguished by the distribution of trans-chromosomal interaction frequencies (nsTIFs) across 23 chromosomes. Figure 1A provides a circos plot of trans-interactions mediated by ecDNA regions across all 23 chromosomes in HF-3016 and HF-3177 ecDNA(+) cell lines. Strong connections between ecMYC, ecEGFR, and ecCDK4 regions are shown. Figure 1B shows the distribution of genome-wide normalized summed TIFs (nsTIFs) with 50 Kb bin sizes in ecDNA(+) HF-3016 and HF-3177 cell lines, and ecDNA(-) HF-3035 line. Elevated nsTIFs are observed at chromosomes 7, 8, and 12. The distribution of nsTIFs along the entire chromosomes 7, 8 and 12 is shown below, and the regions with elevated nsTIF levels match well with known ecEGFR, ecMYC and ecCDK4 regions. [Figure 2A]Figure 2A-D show profiles, graphs, and boxplots demonstrating evidence that ecDNA shows a strong enhancement of broad-span H3K27ac modifications. Figure 2A provides H3K27ac modification enrichment profiles within ±3 Kb of chromosomal non-coding anchors interacting with promoters of ec- oncogenes found in each of the ecDNA(+) cell lines across HF-2354, HF-2927, HF-3016, and HF-3177. The number of anchors found in each line is labeled, and each region is shown as a row ordered by signal intensity (top to bottom, from high to low intensity) detected in their corresponding cell line. Figure 2B shows the agreement between the distribution of chromatin interaction frequency and H3K27ac signal density across the ecEGFR region (chr7: 54,929,292-55,441,765) in HF-2927. Lower panel: H3K27ac signal density profile of the ecEGFR region in ecDNA(+) HF-2927 (top). For comparison, the H3K27ac signal density profile obtained from the same region in the ecDNA(-) HF-3035 strain is shown (bottom). Regions are highlighted to show differences in signal intensity and span. Figure 2C-D show box plots depicting the fold enrichment (Figure 2C) and span size distribution (Figure 2D) of H3K27ac peaks (group A, n = 17, 16, 96, and 70), their corresponding trans-interacting chromosomal anchors (group B, n = 166, 634, 913, and 745), as well as the remaining whole genome peaks (group C, n = 38,259, 40,413, 48,751, and 56,041) in ecDNA obtained from each of the four ecDNA(+) strains. In the ecDNA(-) HF-3035 strain, group A (n=182) refers to H3K27ac peaks found in the collective ecDNA-equivalent regions, and group C (n=53,529) represents the remaining genome-wide peaks detected. Y-axes are log2 and log10 scales, respectively, in Figure 2C and D. Center line is median, boxes are first and third quartiles, whiskers are 1.5× interquartile range (IQR), and points are outliers. *: P value < 0.005 (one-sided Wilcoxon rank sum test).With respect to sample size and exact P value for each paired comparison. In each of Figures 2C and 2D, from left to right on the x-axis, the first A, B, and C show data from HF-2354, the second A, B, and C show data from HF-2927, the third A, B, and C show results from HF-3016, the fourth A, B, and C show data from HF-3177, and the last A and C show data from HF-3035. [Figure 2B]Figure 2A-D show profiles, graphs, and boxplots demonstrating evidence that ecDNA shows a strong enhancement of broad-span H3K27ac modifications. Figure 2A provides H3K27ac modification enrichment profiles within ±3 Kb of chromosomal non-coding anchors interacting with promoters of ec- oncogenes found in each of the ecDNA(+) cell lines across HF-2354, HF-2927, HF-3016, and HF-3177. The number of anchors found in each line is labeled, and each region is shown as a row ordered by signal intensity (top to bottom, from high to low intensity) detected in their corresponding cell line. Figure 2B shows the agreement between the distribution of chromatin interaction frequency and H3K27ac signal density across the ecEGFR region (chr7: 54,929,292-55,441,765) in HF-2927. Lower panel: H3K27ac signal density profile of the ecEGFR region in ecDNA(+) HF-2927 (top). For comparison, the H3K27ac signal density profile obtained from the same region in the ecDNA(-) HF-3035 strain is shown (bottom). Regions are highlighted to show differences in signal intensity and span. Figure 2C-D show box plots depicting the fold enrichment (Figure 2C) and span size distribution (Figure 2D) of H3K27ac peaks (group A, n = 17, 16, 96, and 70), their corresponding trans-interacting chromosomal anchors (group B, n = 166, 634, 913, and 745), as well as the remaining whole genome peaks (group C, n = 38,259, 40,413, 48,751, and 56,041) in ecDNA obtained from each of the four ecDNA(+) strains. In the ecDNA(-) HF-3035 strain, group A (n=182) refers to H3K27ac peaks found in the collective ecDNA-equivalent regions, and group C (n=53,529) represents the remaining genome-wide peaks detected. Y-axes are log2 and log10 scales, respectively, in Figure 2C and D. Center line is median, boxes are first and third quartiles, whiskers are 1.5× interquartile range (IQR), and points are outliers. *: P value < 0.005 (one-sided Wilcoxon rank sum test).With respect to sample size and exact P value for each paired comparison. In each of Figures 2C and 2D, from left to right on the x-axis, the first A, B, and C show data from HF-2354, the second A, B, and C show data from HF-2927, the third A, B, and C show results from HF-3016, the fourth A, B, and C show data from HF-3177, and the last A and C show data from HF-3035. [Figure 2C]Figure 2A-D show profiles, graphs, and boxplots demonstrating evidence that ecDNA shows a strong enhancement of broad-span H3K27ac modifications. Figure 2A provides H3K27ac modification enrichment profiles within ±3 Kb of chromosomal non-coding anchors interacting with promoters of ec- oncogenes found in each of the ecDNA(+) cell lines across HF-2354, HF-2927, HF-3016, and HF-3177. The number of anchors found in each line is labeled, and each region is shown as a row ordered by signal intensity (top to bottom, from high to low intensity) detected in their corresponding cell line. Figure 2B shows the agreement between the distribution of chromatin interaction frequency and H3K27ac signal density across the ecEGFR region (chr7: 54,929,292-55,441,765) in HF-2927. Lower panel: H3K27ac signal density profile of the ecEGFR region in ecDNA(+) HF-2927 (top). For comparison, the H3K27ac signal density profile obtained from the same region in the ecDNA(-) HF-3035 strain is shown (bottom). Regions are highlighted to show differences in signal intensity and span. Figure 2C-D show box plots depicting the fold enrichment (Figure 2C) and span size distribution (Figure 2D) of H3K27ac peaks (group A, n = 17, 16, 96, and 70), their corresponding trans-interacting chromosomal anchors (group B, n = 166, 634, 913, and 745), as well as the remaining whole genome peaks (group C, n = 38,259, 40,413, 48,751, and 56,041) in ecDNA obtained from each of the four ecDNA(+) strains. In the ecDNA(-) HF-3035 strain, group A (n=182) refers to H3K27ac peaks found in the collective ecDNA-equivalent regions, and group C (n=53,529) represents the remaining genome-wide peaks detected. Y-axes are log2 and log10 scales, respectively, in Figure 2C and D. Center line is median, boxes are first and third quartiles, whiskers are 1.5× interquartile range (IQR), and points are outliers. *: P value < 0.005 (one-sided Wilcoxon rank sum test).With respect to sample size and exact P value for each paired comparison. In each of Figures 2C and 2D, from left to right on the x-axis, the first A, B, and C show data from HF-2354, the second A, B, and C show data from HF-2927, the third A, B, and C show results from HF-3016, the fourth A, B, and C show data from HF-3177, and the last A and C show data from HF-3035. [Figure 2D]Figure 2A-D show profiles, graphs, and boxplots demonstrating evidence that ecDNA shows a strong enhancement of broad-span H3K27ac modifications. Figure 2A provides H3K27ac modification enrichment profiles within ±3 Kb of chromosomal non-coding anchors interacting with promoters of ec- oncogenes found in each of the ecDNA(+) cell lines across HF-2354, HF-2927, HF-3016, and HF-3177. The number of anchors found in each line is labeled, and each region is shown as a row ordered by signal intensity (top to bottom, from high to low intensity) detected in their corresponding cell line. Figure 2B shows the agreement between the distribution of chromatin interaction frequency and H3K27ac signal density across the ecEGFR region (chr7: 54,929,292-55,441,765) in HF-2927. Lower panel: H3K27ac signal density profile of the ecEGFR region in ecDNA(+) HF-2927 (top). For comparison, the H3K27ac signal density profile obtained from the same region in the ecDNA(-) HF-3035 strain is shown (bottom). Regions are highlighted to show differences in signal intensity and span. Figure 2C-D show box plots depicting the fold enrichment (Figure 2C) and span size distribution (Figure 2D) of H3K27ac peaks (group A, n = 17, 16, 96, and 70), their corresponding peaks of trans-interacting chromosomal anchors (group B, n = 166, 634, 913, and 745), as well as the remaining whole genome peaks (group C, n = 38,259, 40,413, 48,751, and 56,041) in ecDNA obtained from each of the four ecDNA(+) strains. In the ecDNA(-) HF-3035 strain, group A (n=182) refers to H3K27ac peaks found in the collective ecDNA-equivalent regions, and group C (n=53,529) represents the remaining genome-wide peaks detected. Y-axes are log2 and log10 scales, respectively, in Figure 2C and D. Center line is median, boxes are first and third quartiles, whiskers are 1.5× interquartile range (IQR), and points are outliers. *: P value < 0.005 (one-sided Wilcoxon rank sum test).With respect to sample size and exact P value for each paired comparison. In each of Figures 2C and 2D, from left to right on the x-axis, the first A, B, and C show data from HF-2354, the second A, B, and C show data from HF-2927, the third A, B, and C show results from HF-3016, the fourth A, B, and C show data from HF-3177, and the last A and C show data from HF-3035. [Figure 3A] Figures 3A-C provide Venn diagrams, schematics, and tables showing ecDNA-mediated trans-interacting genes and their associated interaction networks. Figures 3A-B show Venn diagrams presenting the number of ecDNA-connected genes (Figure 3A) and oncogenes (Figure 3B) in each of the four ecDNA(+) cell lines, as well as their overlap. Figure 3C shows the process flow for defining interaction anchors, nodes, hubs, and communities (see Methods in the Examples section). Genomic regions connected by chromatin loops were defined as anchors. Non-overlapping anchors were merged as nodes with connectivity scores (number of nodes they connect), among which highly connected nodes (mean connectivity score of +3 standard deviations or more) were classified as hubs. Hubs and nodes with extensive connectivity were collectively defined as communities. The number of communities, hubs, and oncogenes associated with ecDNA are summarized for each of the ecDNA(+) lines. [Figure 3B]Figures 3A-C provide Venn diagrams, schematics, and tables showing ecDNA-mediated trans-interacting genes and their associated interaction networks. Figures 3A-B show Venn diagrams presenting the number of ecDNA-connected genes (Figure 3A) and oncogenes (Figure 3B) in each of the four ecDNA(+) cell lines, as well as their overlap. Figure 3C shows the process flow for defining interaction anchors, nodes, hubs, and communities (see Methods in the Examples section). Genomic regions connected by chromatin loops were defined as anchors. Non-overlapping anchors were merged as nodes with connectivity scores (number of nodes they connect), among which highly connected nodes (mean connectivity score of +3 standard deviations or more) were classified as hubs. Hubs and nodes with extensive connectivity were collectively defined as communities. The number of communities, hubs, and oncogenes associated with ecDNA are summarized for each of the ecDNA(+) lines. [Figure 3C] Figures 3A-C provide Venn diagrams, schematics, and tables showing ecDNA-mediated trans-interacting genes and their associated interaction networks. Figures 3A-B show Venn diagrams presenting the number of ecDNA-connected genes (Figure 3A) and oncogenes (Figure 3B) in each of the four ecDNA(+) cell lines, as well as their overlap. Figure 3C shows the process flow for defining interaction anchors, nodes, hubs, and communities (see Methods in the Examples section). Genomic regions connected by chromatin loops were defined as anchors. Non-overlapping anchors were merged as nodes with connectivity scores (number of nodes they connect), among which highly connected nodes (mean connectivity score of +3 standard deviations or more) were classified as hubs. Hubs and nodes with extensive connectivity were collectively defined as communities. The number of communities, hubs, and oncogenes associated with ecDNA are summarized for each of the ecDNA(+) lines. [Figure 4A]Figure 4A-B shows an overview of the ChIA-PET analysis in five GBM-derived neurosphere cell lines. Figure 4A shows the steady-state expression levels of genes amplified within ecDNA from ecDNA(+) GBM-derived cell lines. Figure 4B shows the number of RNAPII binding sites and long-range chromatin interactions detected in each of the five cell lines from the ChIA-PET data. Interactions are reported in three categories: significant cis interactions in the whole genome (PET count ≥ 3, P value < 0.05, and FDR < 0.05, see Methods in the Examples section), interactions between ecDNA regions (intra-ecDNA interactions), and interactions between ecDNA regions and 23 linear chromosomes (ecDNA trans-chromosomes). Only interactions with RNAPII binding sites detected in both anchors are reported. [Figure 4B] Figure 4A-B shows an overview of the ChIA-PET analysis in five GBM-derived neurosphere cell lines. Figure 4A shows the steady-state expression levels of genes amplified within ecDNA from ecDNA(+) GBM-derived cell lines. Figure 4B shows the number of RNAPII binding sites and long-range chromatin interactions detected in each of the five cell lines from the ChIA-PET data. Interactions are reported in three categories: significant cis interactions in the whole genome (PET count ≥ 3, P value < 0.05, and FDR < 0.05, see Methods in the Examples section), interactions between ecDNA regions (intra-ecDNA interactions), and interactions between ecDNA regions and 23 linear chromosomes (ecDNA trans-chromosomes). Only interactions with RNAPII binding sites detected in both anchors are reported. [Figure 5A]Figure 5A-C show schematics, distributions, and boxplots of results showing the discovery of ecDNA signatures by ChIA-PET analysis. Figure 5A is a schematic showing the chromatin interaction analysis used to detect amplified genomic regions and their associated chromatin contacts in ecDNA. RNAPII ChIA-PET assays were performed to capture all RNAPII-associated chromatin. ecDNA harbors actively expressed oncogenes, is not constrained by chromatin territory, and makes extensive contacts with other chromosomal regions, which can be used to discover ecDNA-specific signatures within active transcriptional hubs and characterize their co-regulated genes. Figure 5B shows the distribution of normalized sum of trans interaction frequencies (nsTIFs) across 23 chromosomes at 50 Kb bin sizes in HF-2927 and HF-2354. The distribution of normalized nsTIFs (shown in the expanded nsTIF plots) for chromosomes 7 and 8, respectively, in their corresponding cell lines reveals the location of ecDNA predicted to encompass EGFR and MYC, respectively. The left side of Figure 5B shows a genome-wide 2D chromatin contact heat map showing distinct line pairs in regions on chromosomes 7p11 and 8q24, indicating strong contacts throughout the genome. The right side of Figure 5B shows a circos plot of transchromosomal contact frequencies across all 23 chromosomes mediated by ecEGFR and ecMYC regions. Figure 5C provides a boxplot presentation of normalized nsTIFs between known ecDNA regions and chromosomal DNA regions with copy number gains of 3 or more in four ecDNA(+) cell lines. From left to right, n=31, 3,423, 11, 5, 15, 82, 25, 833. nsTIFs in ecDNA regions are statistically higher than nsTIFs in regions with copy number gains. P values ​​(one-sided Wilcoxon rank sum test) are 4E-22, 4.5E-4, 4.4E-10, and 7.1E-18 for HF-2354, HF-2927, HF-3016, and HF-3177, respectively. Center line is median, boxes are 1st and 3rd quartiles, whiskers are 1.5 × interquartile range (IQR), points are outliers. [Figure 5B] Figure 5A-C show schematics, distributions, and boxplots of results showing the discovery of ecDNA signatures by ChIA-PET analysis. Figure 5A is a schematic showing the chromatin interaction analysis used to detect amplified genomic regions and their associated chromatin contacts in ecDNA. RNAPII ChIA-PET assays were performed to capture all RNAPII-associated chromatin. ecDNA harbors actively expressed oncogenes, is not constrained by chromatin territory, and makes extensive contacts with other chromosomal regions, which can be used to discover ecDNA-specific signatures within active transcriptional hubs and characterize their co-regulated genes. Figure 5B shows the distribution of normalized sum of trans interaction frequencies (nsTIFs) across 23 chromosomes at 50 Kb bin sizes in HF-2927 and HF-2354. The distribution of normalized nsTIFs (shown in the expanded nsTIF plots) for chromosomes 7 and 8, respectively, in their corresponding cell lines reveals the location of ecDNA predicted to encompass EGFR and MYC, respectively. The left side of Figure 5B shows a genome-wide 2D chromatin contact heat map showing distinct line pairs in regions on chromosomes 7p11 and 8q24, indicating strong contacts throughout the genome. The right side of Figure 5B shows a circos plot of transchromosomal contact frequencies across all 23 chromosomes mediated by ecEGFR and ecMYC regions. Figure 5C provides a boxplot presentation of normalized nsTIFs between known ecDNA regions and chromosomal DNA regions with copy number gains of 3 or more in four ecDNA(+) cell lines. From left to right, n=31, 3,423, 11, 5, 15, 82, 25, 833. nsTIFs in ecDNA regions are statistically higher than nsTIFs in regions with copy number gains. P values ​​(one-sided Wilcoxon rank sum test) are 4E-22, 4.5E-4, 4.4E-10, and 7.1E-18 for HF-2354, HF-2927, HF-3016, and HF-3177, respectively. Center line is median, boxes are 1st and 3rd quartiles, whiskers are 1.5 × interquartile range (IQR), points are outliers. [Figure 5C] Figure 5A-C show schematics, distributions, and boxplots of results showing the discovery of ecDNA signatures by ChIA-PET analysis. Figure 5A is a schematic showing the chromatin interaction analysis used to detect amplified genomic regions and their associated chromatin contacts in ecDNA. RNAPII ChIA-PET assays were performed to capture all RNAPII-associated chromatin. ecDNA harbors actively expressed oncogenes, is not constrained by chromatin territory, and makes extensive contacts with other chromosomal regions, which can be used to discover ecDNA-specific signatures within active transcriptional hubs and characterize their co-regulated genes. Figure 5B shows the distribution of normalized sum of trans interaction frequencies (nsTIFs) across 23 chromosomes at 50 Kb bin sizes in HF-2927 and HF-2354. The distribution of normalized nsTIFs (shown in the expanded nsTIF plots) for chromosomes 7 and 8, respectively, in their corresponding cell lines reveals the location of ecDNA predicted to encompass EGFR and MYC, respectively. The left side of Figure 5B shows a genome-wide 2D chromatin contact heat map showing distinct line pairs in regions on chromosomes 7p11 and 8q24, indicating strong contacts throughout the genome. The right side of Figure 5B shows a circos plot of transchromosomal contact frequencies across all 23 chromosomes mediated by ecEGFR and ecMYC regions. Figure 5C provides a boxplot presentation of normalized nsTIFs between known ecDNA regions and chromosomal DNA regions with copy number gains of 3 or more in four ecDNA(+) cell lines. From left to right, n=31, 3,423, 11, 5, 15, 82, 25, 833. nsTIFs in ecDNA regions are statistically higher than nsTIFs in regions with copy number gains. P values ​​(one-sided Wilcoxon rank sum test) are 4E-22, 4.5E-4, 4.4E-10, and 7.1E-18 for HF-2354, HF-2927, HF-3016, and HF-3177, respectively. Center line is median, boxes are 1st and 3rd quartiles, whiskers are 1.5 × interquartile range (IQR), points are outliers. [Figure 6A] Figures 6A-C provide heat maps showing ChIA-PET assay detection of chromatin topology changes reflected by genomic structural variants. Figure 6A shows spatial chromatin topology as measured by chromatin contacts common in ChIA-PET data, which can be visualized by 2D contact heat maps. Heat maps of chromosome 2 from all five GBM patient-derived neurosphere cell lines are shown. Figure 6B provides 2D contact heat maps of genomic regions with deletions of PTEN (top) and CDKN2A and CDKN2B (bottom) in HF-2927 and HF-3035, respectively. TAD regions are delimited by blue lines. Deletions of gene loci are represented by loss of chromatin contacts. Figure 6C shows additional structural variants, a deletion in the DMD gene, a complex rearrangement in chromosome 3, and a double translocation t(3;6), visualized as aberrant contact patterns by 2D heat maps. [Figure 6B] Figures 6A-C provide heat maps showing ChIA-PET assay detection of chromatin topology changes reflected by genomic structural variants. Figure 6A shows spatial chromatin topology as measured by chromatin contacts common in ChIA-PET data, which can be visualized by 2D contact heat maps. Heat maps of chromosome 2 from all five GBM patient-derived neurosphere cell lines are shown. Figure 6B provides 2D contact heat maps of genomic regions with deletions of PTEN (top) and CDKN2A and CDKN2B (bottom) in HF-2927 and HF-3035, respectively. TAD regions are delimited by blue lines. Deletions of gene loci are represented by loss of chromatin contacts. Figure 6C shows additional structural variants, a deletion in the DMD gene, a complex rearrangement in chromosome 3, and a double translocation t(3;6), visualized as aberrant contact patterns by 2D heat maps. [Figure 6C]Figures 6A-C provide heat maps showing ChIA-PET assay detection of chromatin topology changes reflected by genomic structural variants. Figure 6A shows spatial chromatin topology as measured by chromatin contacts common in ChIA-PET data, which can be visualized by 2D contact heat maps. Heat maps of chromosome 2 from all five GBM patient-derived neurosphere cell lines are shown. Figure 6B provides 2D contact heat maps of genomic regions with deletions of PTEN (top) and CDKN2A and CDKN2B (bottom) in HF-2927 and HF-3035, respectively. TAD regions are delimited by blue lines. Deletions of gene loci are represented by loss of chromatin contacts. Figure 6C shows additional structural variants, a deletion in the DMD gene, a complex rearrangement in chromosome 3, and a double translocation t(3;6), visualized as aberrant contact patterns by 2D heat maps. [Figure 7A]Figures 7A-C show heat maps, circus plots, and schematics showing that ecDNA is bound by RNAPII and mediates a wide range of extrachromosomal, intrachromosomal, and transchromosomal interactions. Figure 7A provides a 2D contact heat map comparison between ecDNA(+)HF-2927 and HF-2354 and ecDNA(-)HF-3035 cell lines. Figure 7B shows the cis-interaction and RNAPII binding intensity profiles of the ecEGFR region (chr7:54,860,254-55,535,856) in HF-2927 compared to the non-ecDNA EGFR gene coding region in HF-3035, and of two segments of the ecMYC region (chr8:128,032,011-128,806,493 and chr8:129,573,241-130,968,628) in HF-2354 compared to the non-ecDNA MYC coding gene in HF-3035. Figure 7B provides circos plots of the defined ecDNA regions in ecDNA(+) cell lines, HF-2927 (left) and HF-3177 (right). From inner to outer circle: intra-ecDNA interaction loops between different regions in ecDNA, blue: distribution of intra-ecDNA interaction frequency, green: distribution of ecDNA-chromosomal trans-interaction frequency, orange (third outer ring): H3K27ac fold enrichment intensity, brown (second outer ring): RNAPII binding enrichment intensity. Signal tracks are at 1 Kb resolution. High concordance between H3K27ac signals and interaction frequency is highlighted in grey. Figure 7C shows genomic features (promoter, intergenic, and intragenic regions) associated with trans-interaction anchors derived from ecDNA and their chromosomal targets. [Figure 7B]Figures 7A-C show heatmaps, circus plots, and schematics showing that ecDNA is bound by RNAPII and mediates a wide range of extrachromosomal, intrachromosomal, and transchromosomal interactions. Figure 7A provides a 2D contact heatmap comparison between ecDNA(+)HF-2927 and HF-2354 and ecDNA(-)HF-3035 cell lines. Figure 7B shows the cis-interaction and RNAPII binding intensity profiles of the ecEGFR region (chr7:54,860,254-55,535,856) in HF-2927 compared to the non-ecDNA EGFR gene coding region in HF-3035, and of two segments of the ecMYC region (chr8:128,032,011-128,806,493 and chr8:129,573,241-130,968,628) in HF-2354 compared to the non-ecDNA MYC coding gene in HF-3035. Figure 7B provides circos plots of the defined ecDNA regions in ecDNA(+) cell lines, HF-2927 (left) and HF-3177 (right). From inner to outer circle: intra-ecDNA interaction loops between different regions in ecDNA, blue: distribution of intra-ecDNA interaction frequency, green: distribution of ecDNA-chromosomal trans-interaction frequency, orange (third outer ring): H3K27ac fold enrichment intensity, brown (second outer ring): RNAPII binding enrichment intensity. Signal tracks are at 1 Kb resolution. High concordance between H3K27ac signals and interaction frequency is highlighted in grey. Figure 7C shows genomic features (promoter, intergenic, and intragenic regions) associated with trans-interaction anchors derived from ecDNA and their chromosomal targets. [Figure 7C]Figures 7A-C show heatmaps, circus plots, and schematics showing that ecDNA is bound by RNAPII and mediates a wide range of extrachromosomal, intrachromosomal, and transchromosomal interactions. Figure 7A provides a 2D contact heatmap comparison between ecDNA(+)HF-2927 and HF-2354 and ecDNA(-)HF-3035 cell lines. Figure 7B shows the cis-interaction and RNAPII binding intensity profiles of the ecEGFR region (chr7:54,860,254-55,535,856) in HF-2927 compared to the non-ecDNA EGFR gene coding region in HF-3035, and of two segments of the ecMYC region (chr8:128,032,011-128,806,493 and chr8:129,573,241-130,968,628) in HF-2354 compared to the non-ecDNA MYC coding gene in HF-3035. Figure 7B provides circos plots of the defined ecDNA regions in ecDNA(+) cell lines, HF-2927 (left) and HF-3177 (right). From inner to outer circle: intra-ecDNA interaction loops between different regions in ecDNA, blue: distribution of intra-ecDNA interaction frequency, green: distribution of ecDNA-chromosomal trans-interaction frequency, orange (third outer ring): H3K27ac fold enrichment intensity, brown (second outer ring): RNAPII binding enrichment intensity. Signal tracks are at 1 Kb resolution. High concordance between H3K27ac signals and interaction frequency is highlighted in grey. Figure 7C shows genomic features (promoter, intergenic, and intragenic regions) associated with trans-interaction anchors derived from ecDNA and their chromosomal targets. [Figure 8A]Figure 8A-D provide circos plots, schematics, and micrographs showing that ecDNA regions exhibit strong cis- and trans-interactions with strong H3K27ac enrichment. Figure 8A shows circos plots of defined ecDNA regions in HF-2354 (left) and HF-3016 (right) ecDNA(+) cell lines. From inner circle to outer circle: intra-ecDNA interaction loops between different regions of ecDNA, blue: distribution of intra-ecDNA interaction frequency, green: distribution of ecDNA-chromosome trans-interaction frequency, third outer circle in orange: H3K27ac enrichment intensity fold, second outer circle in brown: RNAPII binding enrichment intensity. Signal tracks are at 1 Kb resolution. High concordance between high H3K27ac enrichment and interaction frequency is highlighted in grey. Figure 8B shows the percentage of chromosome interacting anchors connected by ecDNA (orange) versus anchors that do not have ecDNA connections (blue) that overlapped with H3K27ac peaks. The left of Figure 8B shows chromosome anchors in intragenic regions (G). The right of Figure 8B shows chromosome anchors in intergenic regions (I) (outside gene coding regions). Figure 8C shows H3K27ac peaks of cell-specific chromosomes interacting with ecMYC promoter. The genomic locations of these broad peaks are shown in 10 Kb windows with the start positions of chr3:42,090,000, chr5:148,938,000, chr1:224,353,000, chr1:33,906,500, chr19:18,408,000, chr1:207,060,000 in HF-3177, chr3:195,902,000, chr10:10,0120,000 in HF-3016, and chr7:63,920,000 in HF-2354. Figure 8D provides an image showing H3K27ac immunostaining of metaphase chromosomes from HF-2927 cells. Distinct DAPI-positive extrachromosomal spots overlap with H3K27ac staining spots. [Figure 8B]Figure 8A-D provide circos plots, schematics, and micrographs showing that ecDNA regions exhibit strong cis- and trans-interactions with strong H3K27ac enrichment. Figure 8A shows circos plots of defined ecDNA regions in HF-2354 (left) and HF-3016 (right) ecDNA(+) cell lines. From inner circle to outer circle: intra-ecDNA interaction loops between different regions of ecDNA, blue: distribution of intra-ecDNA interaction frequency, green: distribution of ecDNA-chromosome trans-interaction frequency, third outer circle in orange: H3K27ac enrichment intensity fold, second outer circle in brown: RNAPII binding enrichment intensity. Signal tracks are at 1 Kb resolution. High concordance between high H3K27ac enrichment and interaction frequency is highlighted in grey. Figure 8B shows the percentage of chromosome interacting anchors connected by ecDNA (orange) versus anchors that do not have ecDNA connections (blue) that overlapped with H3K27ac peaks. The left of Figure 8B shows chromosome anchors in intragenic regions (G). The right of Figure 8B shows chromosome anchors in intergenic regions (I) (outside gene coding regions). Figure 8C shows H3K27ac peaks of cell-specific chromosomes interacting with ecMYC promoter. The genomic locations of these broad peaks are shown in 10 Kb windows with the start positions of chr3:42,090,000, chr5:148,938,000, chr1:224,353,000, chr1:33,906,500, chr19:18,408,000, chr1:207,060,000 in HF-3177, chr3:195,902,000, chr10:10,0120,000 in HF-3016, and chr7:63,920,000 in HF-2354. Figure 8D provides an image showing H3K27ac immunostaining of metaphase chromosomes from HF-2927 cells. Distinct DAPI-positive extrachromosomal spots overlap with H3K27ac staining spots. [Figure 8C]Figure 8A-D provide circos plots, schematics, and micrographs showing that ecDNA regions exhibit strong cis- and trans-interactions with strong H3K27ac enrichment. Figure 8A shows circos plots of defined ecDNA regions in HF-2354 (left) and HF-3016 (right) ecDNA(+) cell lines. From inner circle to outer circle: intra-ecDNA interaction loops between different regions of ecDNA, blue: distribution of intra-ecDNA interaction frequency, green: distribution of ecDNA-chromosome trans-interaction frequency, third outer circle in orange: H3K27ac enrichment intensity fold, second outer circle in brown: RNAPII binding enrichment intensity. Signal tracks are at 1 Kb resolution. High concordance between high H3K27ac enrichment and interaction frequency is highlighted in grey. Figure 8B shows the percentage of chromosome interacting anchors connected by ecDNA (orange) versus anchors that do not have ecDNA connections (blue) that overlapped with H3K27ac peaks. The left of Figure 8B shows chromosome anchors in intragenic regions (G). The right of Figure 8B shows chromosome anchors in intergenic regions (I) (outside gene coding regions). Figure 8C shows H3K27ac peaks of cell-specific chromosomes interacting with ecMYC promoter. The genomic locations of these broad peaks are shown in 10 Kb windows with the start positions of chr3:42,090,000, chr5:148,938,000, chr1:224,353,000, chr1:33,906,500, chr19:18,408,000, chr1:207,060,000 in HF-3177, chr3:195,902,000, chr10:10,0120,000 in HF-3016, and chr7:63,920,000 in HF-2354. Figure 8D provides an image showing H3K27ac immunostaining of metaphase chromosomes from HF-2927 cells. Distinct DAPI-positive extrachromosomal spots overlap with H3K27ac staining spots. [Figure 8D]Figure 8A-D provide circos plots, schematics, and micrographs showing that ecDNA regions exhibit strong cis- and trans-interactions with strong H3K27ac enrichment. Figure 8A shows circos plots of defined ecDNA regions in HF-2354 (left) and HF-3016 (right) ecDNA(+) cell lines. From inner circle to outer circle: intra-ecDNA interaction loops between different regions of ecDNA, blue: distribution of intra-ecDNA interaction frequency, green: distribution of ecDNA-chromosome trans-interaction frequency, third outer circle in orange: H3K27ac enrichment intensity fold, second outer circle in brown: RNAPII binding enrichment intensity. Signal tracks are at 1 Kb resolution. High concordance between high H3K27ac enrichment and interaction frequency is highlighted in grey. Figure 8B shows the percentage of chromosome interacting anchors connected by ecDNA (orange) versus anchors that do not have ecDNA connections (blue) that overlapped with H3K27ac peaks. The left of Figure 8B shows chromosome anchors in intragenic regions (G). The right of Figure 8B shows chromosome anchors in intergenic regions (I) (outside gene coding regions). Figure 8C shows H3K27ac peaks of cell-specific chromosomes interacting with ecMYC promoter. The genomic locations of these broad peaks are shown in 10 Kb windows with the start positions of chr3:42,090,000, chr5:148,938,000, chr1:224,353,000, chr1:33,906,500, chr19:18,408,000, chr1:207,060,000 in HF-3177, chr3:195,902,000, chr10:10,0120,000 in HF-3016, and chr7:63,920,000 in HF-2354. Figure 8D provides an image showing H3K27ac immunostaining of metaphase chromosomes from HF-2927 cells. Distinct DAPI-positive extrachromosomal spots overlap with H3K27ac staining spots. [Figure 9A]Figure 9A-E provide box plots and schematics showing that ecDNA-mediated chromatin interactions target oncogenes for active transcription in spatially condensed nuclear networks. Figure 9A shows the distribution of RNA expression (FPKM) between chromosomal genes that trans-interact with ecDNA (n=1,887, 1,270, 1,483, and 1,157) and genes that do not have trans-chromosomal interactions (n=483, 194, 653, and 597) from each of the four ecDNA(+) cell lines. * indicates a P value less than 0.005 (one-tailed Wilcoxon rank sum test). With respect to sample size and exact P value for each paired comparison. In Figure 9A, from left to right on the X-axis, the first +,- indicates the results of HF-2354, the second +,- indicates the results of HF-2927, the third +,- indicates the results of HF-3016, and the fourth +,- indicates the results of HF-3177. Figure 9B shows the distribution of gene expression (FPKM) of chromosomal genes according to increasing degrees of ecDNA contact frequency (0 to 9). For each ecDNA (+) strain, the 95% confidence interval of the fitted value is shown shaded. The smoothed FPKM is represented as the fitted solid line. Figure 9B shows the results of HF-2354, HF-2927, HF-3016, and HF-3177, with a set of four boxes each for each trans-interaction frequency (1 to 10) from left to right. Figure 9C shows box plots showing expression levels of ecDNA-connected tumor genes (n=87, 56, 78, and 54) versus the whole transcriptome (n=21,186, 18,988, 19,206, and 19,180). P values ​​for paired comparisons by one-tailed Wilcoxon rank sum test are 1.2E-7, 1.4E-5, 9.0E-8, and 1.1E-5. In Figure 9A-C, center line is median, boxes are first and third quartiles, whiskers are 1.5× interquartile range (IQR), and points are outliers. In Figure 9C, from left to right on the X-axis, a tumor, wild type (WT), shows the results for HF-2354, a second tumor, WT, shows the results for HF-2927, a third tumor, WT, shows the results for HF-3016, and a fourth tumor, WT, shows the results for HF-3177.Figure 9D shows that oncogenes are clustered in spatial proximity through ecDNA-mediated chromatin interactions. Two examples of communities mediated by ecDNA-connected hubs in HF-3016 and HF-3177 are shown. Trans interactions between nodes are represented by tan lines whose thickness is represented by log10(iPET counts). Blue circles are gene (promoter) nodes, purple circles are gene (promoter) nodes annotated as oncogenes, and grey circles are intergenic nodes. Circle size is proportional to the connectivity score (number of edges across the whole genome network), except for the ecDNA circle, which was manually adjusted to a smaller size. Selected genes are labeled to show variability across communities. Figure 9E is a schematic of a model showing that ecDNA functions as a mobile enhancer and engages in extensive cis- and trans-chromosomal interactions to recruit oncogenes to active transcriptional hubs and promote global transcriptional amplification in cancer cells. [Figure 9B]Figure 9A-E provide box plots and schematics showing that ecDNA-mediated chromatin interactions target oncogenes for active transcription in spatially condensed nuclear networks. Figure 9A shows the distribution of RNA expression (FPKM) between chromosomal genes that trans-interact with ecDNA (n=1,887, 1,270, 1,483, and 1,157) and genes that do not have trans-chromosomal interactions (n=483, 194, 653, and 597) from each of the four ecDNA(+) cell lines. * indicates a P value less than 0.005 (one-tailed Wilcoxon rank sum test). With respect to sample size and exact P value for each paired comparison. In Figure 9A, from left to right on the X-axis, the first +,- indicates the results of HF-2354, the second +,- indicates the results of HF-2927, the third +,- indicates the results of HF-3016, and the fourth +,- indicates the results of HF-3177. Figure 9B shows the distribution of gene expression (FPKM) of chromosomal genes according to increasing degrees of ecDNA contact frequency (0 to 9). For each ecDNA (+) strain, the 95% confidence interval of the fitted value is shown shaded. The smoothed FPKM is represented as the fitted solid line. Figure 9B shows the results of HF-2354, HF-2927, HF-3016, and HF-3177, with a set of four boxes each for each trans-interaction frequency (1 to 10) from left to right. Figure 9C shows box plots showing expression levels of ecDNA-connected tumor genes (n=87, 56, 78, and 54) versus the whole transcriptome (n=21,186, 18,988, 19,206, and 19,180). P values ​​for paired comparisons by one-tailed Wilcoxon rank sum test are 1.2E-7, 1.4E-5, 9.0E-8, and 1.1E-5. In Figure 9A-C, center line is median, boxes are first and third quartiles, whiskers are 1.5× interquartile range (IQR), and points are outliers. In Figure 9C, from left to right on the X-axis, a tumor, wild type (WT), shows the results for HF-2354, a second tumor, WT, shows the results for HF-2927, a third tumor, WT, shows the results for HF-3016, and a fourth tumor, WT, shows the results for HF-3177.Figure 9D shows that oncogenes are clustered in spatial proximity through ecDNA-mediated chromatin interactions. Two examples of communities mediated by ecDNA-connected hubs in HF-3016 and HF-3177 are shown. Trans interactions between nodes are represented by tan lines whose thickness is represented by log10(iPET counts). Blue circles are gene (promoter) nodes, purple circles are gene (promoter) nodes annotated as oncogenes, and grey circles are intergenic nodes. Circle size is proportional to the connectivity score (number of edges across the whole genome network), except for the ecDNA circle, which was manually adjusted to a smaller size. Selected genes are labeled to show variability across communities. Figure 9E is a schematic of a model showing that ecDNA functions as a mobile enhancer and engages in extensive cis- and trans-chromosomal interactions to recruit oncogenes to active transcriptional hubs and promote global transcriptional amplification in cancer cells. [Figure 9C]Figure 9A-E provide box plots and schematics showing that ecDNA-mediated chromatin interactions target oncogenes for active transcription in spatially condensed nuclear networks. Figure 9A shows the distribution of RNA expression (FPKM) between chromosomal genes that trans-interact with ecDNA (n=1,887, 1,270, 1,483, and 1,157) and genes that do not have trans-chromosomal interactions (n=483, 194, 653, and 597) from each of the four ecDNA(+) cell lines. * indicates a P value less than 0.005 (one-tailed Wilcoxon rank sum test). With respect to sample size and exact P value for each paired comparison. In Figure 9A, from left to right on the X-axis, the first +,- indicates the results of HF-2354, the second +,- indicates the results of HF-2927, the third +,- indicates the results of HF-3016, and the fourth +,- indicates the results of HF-3177. Figure 9B shows the distribution of gene expression (FPKM) of chromosomal genes according to increasing degrees of ecDNA contact frequency (0 to 9). For each ecDNA (+) strain, the 95% confidence interval of the fitted value is shown shaded. The smoothed FPKM is represented as the fitted solid line. Figure 9B shows the results of HF-2354, HF-2927, HF-3016, and HF-3177, with a set of four boxes each for each trans-interaction frequency (1 to 10) from left to right. Figure 9C shows box plots showing expression levels of ecDNA-connected tumor genes (n=87, 56, 78, and 54) versus the whole transcriptome (n=21,186, 18,988, 19,206, and 19,180). P values ​​for paired comparisons by one-tailed Wilcoxon rank sum test are 1.2E-7, 1.4E-5, 9.0E-8, and 1.1E-5. In Figure 9A-C, center line is median, boxes are first and third quartiles, whiskers are 1.5× interquartile range (IQR), and points are outliers. In Figure 9C, from left to right on the X-axis, a tumor, wild type (WT), shows the results for HF-2354, a second tumor, WT, shows the results for HF-2927, a third tumor, WT, shows the results for HF-3016, and a fourth tumor, WT, shows the results for HF-3177.Figure 9D shows that oncogenes are clustered in spatial proximity through ecDNA-mediated chromatin interactions. Two examples of communities mediated by ecDNA-connected hubs in HF-3016 and HF-3177 are shown. Trans interactions between nodes are represented by tan lines whose thickness is represented by log10(iPET counts). Blue circles are gene (promoter) nodes, purple circles are gene (promoter) nodes annotated as oncogenes, and grey circles are intergenic nodes. Circle size is proportional to the connectivity score (number of edges across the whole genome network), except for the ecDNA circle, which was manually adjusted to a smaller size. Selected genes are labeled to show variability across communities. Figure 9E is a schematic of a model showing that ecDNA functions as a mobile enhancer and engages in extensive cis- and trans-chromosomal interactions to recruit oncogenes to active transcriptional hubs and promote global transcriptional amplification in cancer cells. [Figure 9D]Figure 9A-E provide box plots and schematics showing that ecDNA-mediated chromatin interactions target oncogenes for active transcription in spatially condensed nuclear networks. Figure 9A shows the distribution of RNA expression (FPKM) between chromosomal genes that trans-interact with ecDNA (n=1,887, 1,270, 1,483, and 1,157) and genes that do not have trans-chromosomal interactions (n=483, 194, 653, and 597) from each of the four ecDNA(+) cell lines. * indicates a P value less than 0.005 (one-tailed Wilcoxon rank sum test). With respect to sample size and exact P value for each paired comparison. In Figure 9A, from left to right on the X-axis, the first +,- indicates the results of HF-2354, the second +,- indicates the results of HF-2927, the third +,- indicates the results of HF-3016, and the fourth +,- indicates the results of HF-3177. Figure 9B shows the distribution of gene expression (FPKM) of chromosomal genes according to increasing degrees of ecDNA contact frequency (0 to 9). For each ecDNA (+) strain, the 95% confidence interval of the fitted value is shown shaded. The smoothed FPKM is represented as the fitted solid line. Figure 9B shows the results of HF-2354, HF-2927, HF-3016, and HF-3177, with a set of four boxes each for each trans-interaction frequency (1 to 10) from left to right. Figure 9C shows box plots showing expression levels of ecDNA-connected tumor genes (n=87, 56, 78, and 54) versus the whole transcriptome (n=21,186, 18,988, 19,206, and 19,180). P values ​​for paired comparisons by one-tailed Wilcoxon rank sum test are 1.2E-7, 1.4E-5, 9.0E-8, and 1.1E-5. In Figure 9A-C, center line is median, boxes are first and third quartiles, whiskers are 1.5× interquartile range (IQR), and points are outliers. In Figure 9C, from left to right on the X-axis, a tumor, wild type (WT), shows the results for HF-2354, a second tumor, WT, shows the results for HF-2927, a third tumor, WT, shows the results for HF-3016, and a fourth tumor, WT, shows the results for HF-3177.Figure 9D shows that oncogenes are clustered in spatial proximity through ecDNA-mediated chromatin interactions. Two examples of communities mediated by ecDNA-connected hubs in HF-3016 and HF-3177 are shown. Trans interactions between nodes are represented by tan lines whose thickness is represented by log10(iPET counts). Blue circles are gene (promoter) nodes, purple circles are gene (promoter) nodes annotated as oncogenes, and grey circles are intergenic nodes. Circle size is proportional to the connectivity score (number of edges across the whole genome network), except for the ecDNA circle, which was manually adjusted to a smaller size. Selected genes are labeled to show variability across communities. Figure 9E is a schematic of a model showing that ecDNA functions as a mobile enhancer and engages in extensive cis- and trans-chromosomal interactions to recruit oncogenes to active transcriptional hubs and promote global transcriptional amplification in cancer cells. [Figure 9E]Figure 9A-E provide box plots and schematics showing that ecDNA-mediated chromatin interactions target oncogenes for active transcription in spatially condensed nuclear networks. Figure 9A shows the distribution of RNA expression (FPKM) between chromosomal genes that trans-interact with ecDNA (n=1,887, 1,270, 1,483, and 1,157) and genes that do not have trans-chromosomal interactions (n=483, 194, 653, and 597) from each of the four ecDNA(+) cell lines. * indicates a P value less than 0.005 (one-tailed Wilcoxon rank sum test). With respect to sample size and exact P value for each paired comparison. In Figure 9A, from left to right on the X-axis, the first +,- indicates the results of HF-2354, the second +,- indicates the results of HF-2927, the third +,- indicates the results of HF-3016, and the fourth +,- indicates the results of HF-3177. Figure 9B shows the distribution of gene expression (FPKM) of chromosomal genes according to increasing degrees of ecDNA contact frequency (0 to 9). For each ecDNA (+) strain, the 95% confidence interval of the fitted value is shown shaded. The smoothed FPKM is represented as the fitted solid line. Figure 9B shows the results of HF-2354, HF-2927, HF-3016, and HF-3177, with a set of four boxes each for each trans-interaction frequency (1 to 10) from left to right. Figure 9C shows box plots showing expression levels of ecDNA-connected tumor genes (n=87, 56, 78, and 54) versus the whole transcriptome (n=21,186, 18,988, 19,206, and 19,180). P values ​​for paired comparisons by one-tailed Wilcoxon rank sum test are 1.2E-7, 1.4E-5, 9.0E-8, and 1.1E-5. In Figure 9A-C, center line is median, boxes are first and third quartiles, whiskers are 1.5× interquartile range (IQR), and points are outliers. In Figure 9C, from left to right on the X-axis, a tumor, wild type (WT), shows the results for HF-2354, a second tumor, WT, shows the results for HF-2927, a third tumor, WT, shows the results for HF-3016, and a fourth tumor, WT, shows the results for HF-3177.Figure 9D shows that oncogenes are clustered in spatial proximity through ecDNA-mediated chromatin interactions. Two examples of communities mediated by ecDNA-connected hubs in HF-3016 and HF-3177 are shown. Trans interactions between nodes are represented by tan lines whose thickness is represented by log10(iPET counts). Blue circles are gene (promoter) nodes, purple circles are gene (promoter) nodes annotated as oncogenes, and grey circles are intergenic nodes. Circle size is proportional to the connectivity score (number of edges across the whole genome network), except for the ecDNA circle, which was manually adjusted to a smaller size. Selected genes are labeled to show variability across communities. Figure 9E is a schematic of a model showing that ecDNA functions as a mobile enhancer and engages in extensive cis- and trans-chromosomal interactions to recruit oncogenes to active transcriptional hubs and promote global transcriptional amplification in cancer cells. [Figure 10A] Figure 10A-C provide diagrams, schematics, and box plots showing that ecDNA trans-connected genes are actively transcribed. Figure 10A shows high correlation between all expressed RNAs measured from five GBM-derived neurosphere cell lines, representing the consistency of the RNA-seq analysis. Figure 10B shows that based on their ecDNA connectivity status, genes are classified into three distinct categories: Group I: genes with promoters that connect to ecDNA, Group II: genes with promoters that connect to other promoters in trans but do not have ecDNA connections, and Group III: genes with no trans interactions. Figure 10C provides box plots of steady-state RNA expression (FPKM) of genes from Groups I, II, and III in HF-2354, HF-2927, HF-3017, and HF-3177 ecDNA(+) cell lines. * represents significant P-value based on one-tailed Wilcoxon rank sum test. Exact P values ​​for paired comparisons were determined and genes were classified into groups I, II, and III, respectively. [Figure 10B]Figure 10A-C provide diagrams, schematics, and box plots showing that ecDNA trans-connected genes are actively transcribed. Figure 10A shows high correlation between all expressed RNAs measured from five GBM-derived neurosphere cell lines, representing the consistency of the RNA-seq analysis. Figure 10B shows that based on their ecDNA connectivity status, genes are classified into three distinct categories: Group I: genes with promoters that connect to ecDNA, Group II: genes with promoters that connect to other promoters in trans but do not have ecDNA connections, and Group III: genes with no trans interactions. Figure 10C provides box plots of steady-state RNA expression (FPKM) of genes from Groups I, II, and III in HF-2354, HF-2927, HF-3017, and HF-3177 ecDNA(+) cell lines. * represents significant P-value based on one-tailed Wilcoxon rank sum test. Exact P values ​​for paired comparisons were determined and genes were classified into groups I, II, and III, respectively. [Figure 10C] Figure 10A-C provide diagrams, schematics, and box plots showing that ecDNA trans-connected genes are actively transcribed. Figure 10A shows high correlation between all expressed RNAs measured from five GBM-derived neurosphere cell lines, representing the consistency of the RNA-seq analysis. Figure 10B shows that based on their ecDNA connectivity status, genes are classified into three distinct categories: Group I: genes with promoters that connect to ecDNA, Group II: genes with promoters that connect to other promoters in trans but do not have ecDNA connections, and Group III: genes with no trans interactions. Figure 10C provides box plots of steady-state RNA expression (FPKM) of genes from Groups I, II, and III in HF-2354, HF-2927, HF-3017, and HF-3177 ecDNA(+) cell lines. * represents significant P-value based on one-tailed Wilcoxon rank sum test. Exact P values ​​for paired comparisons were determined and genes were classified into groups I, II, and III, respectively. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0011] The present invention relates, in part, to methods for identifying extrachromosomal circular DNA (ecDNA) and the role of ecDNA in diseases such as cancer. Certain methods of the present invention include characterizing ecDNA and its oncogenic changes in cancer genomes. It has now been determined that chromatin interaction assays, such as but not limited to ChIA-PET chromatin interaction assays, can be utilized to advance the identification of ecDNA and characterize genome-wide ecDNA-mediated chromatin contacts that functionally affect transcriptional programs in diseases such as but not limited to cancer. Studies have been performed using ecDNA in neurosphere cultures derived from glioblastoma patients, some of which are described herein. In these studies, ecDNA was identified by the presence of their extensive interchromosomal interactions. The focal points of ecDNA-chromatin contacts were indicated by widespread and high-level signals that primarily converge on chromosomal promoters, indicating a critical regulatory role in genome-wide activation of chromosomal gene transcription. In some embodiments of the present invention, the signals included H3K27ac signals. Deciphering the chromosomal targets of ecDNA revealed its association with actively expressed oncogenes that are spatially convergent within the ecDNA-chromatin connectivity network. Results of the studies performed showed that ecDNA, in addition to being a signature of oncogene amplification, functions as a mobile transcriptional amplification element that activates oncogene expression in cancer.

[0012] Identification of ecDNA To identify the chromatin organization of ecDNA and how it contributes to the regulation of gene transcription, chromatin interaction assays were utilized to test and investigate both general spatial chromatin organization and long-range chromatin interactions mediated by protein factors in the same cell line. Non-limiting examples of chromatin interaction assays that can be used in some embodiments of the present invention are ChIA-PET, ChIP, and Hi-C. It has now been demonstrated that known ecDNAs can be identified through their strong and aberrant intra- and inter-molecular genome-wide chromatin contacts. In addition, studies conducted to decipher the RNA polymerase II (RNAPII)-mediated ecDNA connectome and their chromosomal partners have identified a relationship between ecDNA and actively expressed autosomal oncogenes. This discovery indicates a mechanism for ecDNA to function as a mobile transcriptional enhancer to promote tumor progression.

[0013] In addition to providing a detailed characterization of the ecDNA-targeted chromatin interactome in cancer genomes, the use of the chromatin interaction assays disclosed herein provides an effective means to precisely map amplified genomic domains within ecDNA based on strong chromatin contacts within and between ecDNA, and between ecDNA and linear DNA. Previous methods used to characterize ecDNA utilized either whole genome sequencing and image-based or computational analysis of DNA copy number data. The embodiments of the invention disclosed herein differ from previous methods that rely on structural analysis or microscopic imaging approaches of regions with increased copy number, at least in that the methods provided herein can be used to directly measure the frequency of chromatin contacts between chromosomes through chromatin interaction assays, such as, but not limited to, ChIA-PET and Hi-C. The embodiments of the methods of the invention provide an unbiased approach that can be used to identify one or more ecDNA signatures, such as, but not limited to, ecDNA size, size comparison between different ecDNAs, ecDNA copy number, ecDNA sequence information, and ecDNA sequence content. In addition, embodiments of the methods of the invention can be used to determine and / or evaluate one or more characteristics, such as, but not limited to, the frequency and pattern of contacts between different regions of ecDNA molecules and chromosomal DNA molecules. The use of chromatin interaction assays in the methods of the invention provides insight into the physical structure and continuity of ecDNA molecules.

[0014] Certain aspects of the present invention include a method for identifying one or more ecDNAs in a cell or a plurality of cells. The method may include a means for detecting chromatin interactions between a nonlinear DNA molecule and at least one linear chromosome. In some embodiments of the present invention, the method includes a step of detecting chromatin interactions between a nonlinear DNA molecule and at least one chromosome of a chromosome pair. The term "detecting chromatin interactions" as used herein means detecting one or more of the following: the frequency of chromatin interactions, the ecDNA involved in the interaction, the target gene involved in the interaction, and other characteristics of the chromatin interactions. The characteristics of the chromatin interactions may include, but are not limited to, one or more of the following: the size of the nonlinear DNA, the copy number of the nonlinear DNA molecule in the cell. In some embodiments, the method of the present invention includes a step of comparing one or more characteristics of the chromatin interactions, for example, a step of determining an average copy number per cell of the nonlinear DNA molecule in a plurality of cells at a first determination time point, and a step of comparing the determined average with the average copy number per cell of the nonlinear DNA of a control. In some embodiments of the present invention, comparing one or more characteristics of chromatin interaction may include determining the average copy number per cell of the nonlinear DNA molecule in a plurality of cells at a first determination time point, and comparing the determined average with the average copy number per cell of the control nonlinear DNA. In some embodiments of the present invention, the average copy number per cell of the control nonlinear DNA molecule is the average copy number per cell of the nonlinear DNA molecule determined in a plurality of cells at a time point different from another determination time point. Other non-limiting examples of detecting chromatin interaction characteristics include determining the sequence of at least a portion of the nonlinear DNA, and identifying the presence of a tumor gene sequence in the determined sequence.

[0015] Certain characteristics are identified herein that can be used to identify ecDNA in a cell. The characteristics are identified by one or more determined features of chromatin interactions in the cell. For example, identifying a chromatin interaction as including contact between ecDNA and at least one linear chromosome includes identifying (i) a significantly higher frequency of chromatin interactions detected in the cell, (ii) contact between the nonlinear DNA molecule and at least one linear chromosome (e.g., but not limited to, at least one linear chromosome of each of a chromosome pair in the cell), and (iii) an increase in the average copy number per cell of the nonlinear DNA molecule over time in a plurality of cells. In some embodiments of the present invention, the presence of these three characteristics (characteristics i-iii) identifies the nonlinear DNA molecule as ecDNA.

[0016] Several characteristics of ecDNA are identified herein, and in some embodiments of the method of the present invention, the presence of these characteristics in a nonlinear DNA molecule in a cell confirms the identity of the nonlinear DNA molecule as ecDNA. Certain embodiments of the method of the present invention include determining one, two, or three of the following characteristics as a means to confirm the identity of the detected nonlinear DNA as ecDNA. One such characteristic of ecDNA identified herein is that chromatin interactions involving ecDNA and its target genes occur at a significantly higher level and frequency than other types of chromatin interactions in a cell. As used herein, the term "significantly higher frequency" of detected chromatin interactions means that the number of such chromatin interactions detected is statistically significantly higher than the number of chromatin interactions detected when the chromatin interactions do not involve ecDNA.

[0017] Another characteristic of ecDNA identified herein is the presence of contact between ecDNA and at least one linear chromosome in a cell. In some embodiments of the present invention, the characteristic of ecDNA includes the presence of contact between ecDNA and at least one linear chromosome of each chromosome pair in a cell. For example, but not intended to be limiting, a nonlinear DNA molecule identified as having contact with at least one chromosome in a cell is identified as ecDNA. In another non-limiting example, a nonlinear DNA molecule identified as having contact with at least one chromosome in each of the chromosomes in a cell (e.g., each of the 23 pairs of chromosomes in a human diploid cell) is identified as ecDNA. In the latter example, there will be ecDNA interaction with a gene target located in at least one linear chromosome in each of the 23 chromosome pairs in a human cell.

[0018] A third characteristic of ecDNA identified herein is that the average copy number per cell of ecDNA increases over time. Thus, as a non-limiting example, the average copy number of nonlinear DNA is determined in cell samples obtained from the cell population. At a later time point, a second cell sample is obtained from the cell population, and the average copy number of nonlinear DNA is determined and compared to the average number determined in the first sample. The increase in the average number of nonlinear DNA in the later sample supports the conclusion that the nonlinear DNA is ecDNA. In some embodiments of the invention, the increase in the average number may be at least a 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 125%, 150%, 200%, 250%, or 500% increase, including all percentages within the ranges indicated. In some embodiments of the invention, the increase in the average number may be at least a 500%, 1000%, 1500%, 2000%, or 5000% increase.

[0019] In some embodiments of the present invention, a cell or a plurality of cells are obtained from a sample, culture, or subject at two or more different time points. The length of time between obtaining two cell samples can be independently selected based on factors including, but not limited to, the convenience of the subject, the convenience of the medical practitioner, the status or stage of cancer, the rate of cancer development, the rate of tumor growth, etc. In some embodiments of the present invention, the time interval between obtaining two cell samples is at least 1 day, 2 days, 3 days, 4 days, 5 days, 6 days, 7 days, 8 days, 9 days, 10 days, 11 days, 12 days, 13 days, 14 days, 15 days, 16 days, 17 days, 18 days, 19 days, 20 days, 21 days, 22 days, 23 days, 24 days, 25 days, 26 days, 27 days, 28 days, 29 days, and 30 days. In some embodiments of the invention, the time interval between obtaining two cell samples is at least 1 week, 2 weeks, 3 weeks, 4 weeks, 5 weeks, 6 weeks, 7 weeks, 8 weeks, 9 weeks, 10 weeks, 11 weeks, 12 weeks, 13 weeks, 14 weeks, 15 weeks, 16 weeks, 17 weeks, 18 weeks, 19 weeks, 20 weeks, 21 weeks, 22 weeks, 23 weeks, 24 weeks, 25 weeks, 26 weeks, 27 weeks, 28 weeks, 29 weeks, 30 weeks, 31 weeks, 32 weeks, 33 weeks, 34 weeks, 35 weeks, 36 weeks, 37 weeks, 38 weeks, 39 weeks, 40 weeks, 41 weeks, 42 weeks, 43 weeks, 44 weeks, 45 weeks, 46 weeks, 47 weeks, 48 ​​weeks, 49 weeks, 50 weeks, 51 weeks, and 52 weeks.In some embodiments of the invention, the time interval between obtaining two cell samples is at least 1 month, 2 months, 3 months, 4 months, 5 months, 6 months, 7 months, 8 months, 9 months, 10 months, 11 months, 12 months, 13 months, 14 months, 15 months, 16 months, 17 months, 18 months, 19 months, 20 months, 21 months, 22 months, 23 months, 24 months, 25 months, 26 months, 27 months, 28 months, 29 months, 30 months, 31 months, 32 months, 33 months, 34 months, 35 months, 36 months, 37 months, 38 months, 39 months, 40 months, 41 months, 42 months, 43 months, 44 months, 45 months, 46 months, 47 months, 48 ​​months, 49 months, 50 months, 51 months, 52 months, 53 months, 54 months, 55 months, 56 months, 57 months, 58 months, 59 months, 60 months, 61 months, 62 months, 63 months, 64 months, 65 months, 66 months, 67 months, 68 months, 69 months, 70 months, 71 months, 72 months, 73 months, 74 months, 75 months, 76 months, 77 months, 78 months, 79 months, 80 months, 81 months, 82 months, 83 months, 84 months, 85 months, 86 months, 87 months, 88 months, 89 months, 90 months, 91 months, 92 months, 93 months, 94 months, 95 months, 96 months, 97 months, 98 months, 99 months, 100 months, 3 months, 24 months, 25 months, 26 months, 27 months, 28 months, 29 months, 30 months, 31 months, 32 months, 33 months, 34 months, 35 months, 36 months, 37 months, 38 months, 39 months, 40 months, 41 months, 42 months, 43 months, 44 months, 45 months, 46 months, 47 months, 48 ​​months, or more. In some embodiments of the present invention, the time interval between obtaining two cell samples is at least 1 year, 2 years, 3 years, 4 years, 5 years, 6 years, 7 years, 8 years, 9 years, 10 years, or more. It will be understood that more than two cell samples can be taken for use in certain embodiments of the method of the present invention, and the time interval between obtaining any two samples can be selected independently, but need not necessarily be the same as the time interval between obtaining other cell samples.

[0020] ecDNA and target tumor genes The studies described herein show a regulatory relationship between ecDNA and target genes, in a non-limiting example, oncogenes. As used herein, the term "target oncogene" refers to an oncogene whose activity is modulated by ecDNA. The ecDNA that modulates the transcription of the target oncogene may do so through a gene regulatory system (GRS). The GRS comprises a gene regulatory element in the ecDNA. The ecDNA gene regulatory element may be a gene activator or, in some cases, a gene silencer. It will be understood that the interaction between the ecDNA and its target gene may include binding of the ecDNA gene regulatory element sequence with a transcription factor, which also binds to the gene regulatory element of the target oncogene. The binding of the transcription factor to the gene regulatory element, which comprises a specific short region of DNA, stimulates the transcription of the target oncogene.

[0021] The transcription factor that binds to the ecDNA gene regulator and the target oncogene may comprise a complex of polypeptides, acting as a "connector" between the ecDNA gene regulator and the gene regulatory element of the target oncogene. The term "gene regulatory element" refers to a DNA sequence, such as, but not limited to, a promoter or enhancer sequence, that is responsible for and / or involved in the transcription of the target oncogene. As used herein, the interaction between the ecDNA gene regulator and its target oncogene includes contact with the transcription factor of the ecDNA gene regulator, which also contacts with the gene regulatory element of the target oncogene, e.g., the promoter sequence of the target oncogene. The ecDNA / oncogene interaction enhances the transcription of the target oncogene, which may result in or promote the development of cancer. In some embodiments, the cancer may be present in a cell that contains the ecDNA and its target oncogene. A cancer cell resulting from the ecDNA / oncogene interaction may be present in the subject. Various target oncogenes may have enhanced transcription due to ecDNA / oncogene interactions, and cells may contain one, two, or more different oncogenes whose transcription is increased by ecDNA / oncogene interactions. Two or more cells in a subject may contain the same target oncogene, or may contain different target oncogenes whose transcription is modulated by one or more ecDNAs. Cancer in a subject may be caused and / or maintained by the activity of one, two, or more different oncogenes, each of which is modulated by one or more ecDNAs.

[0022] In cases where two or more different oncogenes are activated in a subject's cancer, two or more different cancer therapeutics directed to different oncogenes and / or different ecDNA / oncogene interactions can be used to effectively treat the cancer in the subject. In some embodiments of the invention, a cancer therapeutic can be selected and / or administered to a subject based at least in part on the presence or absence of an interaction between the ecDNA and a particular oncogene and the modulation of that oncogene by the interaction.

[0023] In addition to the methods of the present invention that can be used to identify ecDNA, certain methods of the present invention can be used to evaluate the status of cancer in a cell by evaluating the presence or absence of interaction between ecDNA and its target oncogene. These methods are based in part on improved understanding of three-dimensional genome organization and its role in gene regulation. Certain embodiments of the present invention include methods for identifying the interaction between ecDNA and at least one target oncogene in a cell. In addition, certain embodiments of the methods of the present invention can be used to evaluate the frequency and effect of ecDNA / oncogene interaction. In some embodiments of the present invention, the interaction of ecDNA with its target oncogene enhances (also referred to herein as "increases") the transcription of the target oncogene.

[0024] Chromatin interaction assays Certain aspects of the present invention include the use of chromatin interaction assays.In certain embodiments of the method of the present invention, chromatin interaction assays are used to determine the structural characteristics of chromosomes, identify ecDNA in cells, and determine the interaction between ecDNA and target genes, such as, but not limited to, oncogenes.The chromatin interaction assays and analysis disclosed herein also allow the determination of the transcriptional regulation of target oncogenes by ecDNA.

[0025] A non-limiting example of a means for assaying chromatin interactions in cells is chromatin interaction analysis by paired-end tag sequencing (ChIA-PET), which combines ChIP with chromatin composition capture (3C) technology (see Fullwood, et al. 2009, Nature, Vol. 426 (7269):58-64, the contents of which are incorporated herein by reference in their entirety). The ChIA-PET method allows for the detection of interactions between distant DNA regions that interact with each other via a protein or protein complex of interest. In a non-limiting example of the ChIA-PET method, chromatin from cells is crosslinked, digested, and removed using an antibody against the protein of interest. Linker sequences are ligated to the ends of DNA, and the presence of the linker sequences promotes ligation with each other (Zhang et al., 2012, Methods Vol. 58, No. 3:289-299, the contents of which are incorporated herein by reference in their entirety). This results in hybrid DNA fragments from two different regions of the genome. The resulting library is sequenced, and the results identify the DNA regions that interact with each other and with the protein of interest. ChIA-PET has previously been used to map transcription factor interactions, and its use in the method embodiments of the present invention allows the identification of ecDNA, ecDNA target genes, and the modulation of target gene transcription by ecDNA / target gene interactions. Chromatin interaction analysis using ChIA-PET is now used to discover chromatin interactions across the genome. Previous visual and computational methods were generally not suitable for detecting weak or dynamic interactions, but this shortcoming is improved through the use of ChIA-PET methods.

[0026] Another non-limiting example of the means for identifying and evaluating chromatin interactions used in certain embodiments of the present invention is the Hi-C evaluation method. Hi-C-based methods can be used in embodiments of the present invention due in part to their ability to provide unbiased whole genome coverage that can measure the chromatin interaction strength between any two given genomic loci. In certain embodiments of the present invention, Hi-C data can be used to evaluate whole genome chromatin organization, such as topologically associating domains (TADs), which are linear continuous regions of genome that are related in three-dimensional space. To identify TADs from Hi-C data, various algorithms known in the art are routinely used (see, for example, Dixon et al, 2012 Nature 485 (7398):376-80, the contents of which are incorporated herein by reference in their entirety).

[0027] The following is a general overview of the elements in Hi-C analysis. The genome of a cell is cross-linked, thereby preserving the interactions between genomic loci. There are several fixation methods known in the art that are suitable for use in the Hi-C method to cross-link the genome of a cell. After cross-linking, the cross-linked genome is cut using a restriction enzyme, and the size of the resulting fragments determines the resolution of the interaction mapping by the Hi-C method. Non-limiting examples of restriction enzymes that can be used are those that cut every 4000 bp, such as EcoR1 or HindIII, resulting in about 1 million fragments in the human genome. For higher resolution interaction mapping, restriction enzymes that cut more frequently can also be used. After the digestion step, the pieces are randomly ligated under conditions that are suitable for ligation between cross-linked interacting fragments, but not between non-cross-linked fragments. The interacting loci can then be quantified by amplifying the ligated junctions using methods such as, but not limited to, polymerase chain reaction (PCR), see, for example, Naumova, et al., 2012 Methods. 58 (3):192-203 and Gavrilov, et al., 2013 PLOS One. 8 (3):e60403, the contents of each of which are incorporated herein by reference in their entirety. Certain Hi-C methods can include high-throughput sequencing to find the nucleotide sequence of the fragments, for example, the sequence of ecDNA.

[0028] Additional steps and procedures that may be included in the ChIA-PET and / or HI-C methods used in embodiments of the present invention are described herein, such as computational methods, analytical methods, data evaluation methods, etc. In addition to chromatin interaction analysis methods related to ChIA-PET and HI-C, other chromatin interaction assays and analysis methods are known and used in the art and are suitable for use in the method embodiments of the present invention. Non-limiting examples of additional chromatin interaction assays / analysis methods include 4C (Nat Genet. 2006 Nov;38(11):1348-54), Hi-C (Lieberman-Aiden, E. et al. 326, 289-293 (2009), Capture Hi-C (Nat Genet. 2015 Jun;47(6):598-606), PLAC-seq (Cell research. 2016;26(12):1345-8), and HiChIP (Nat Methods. 2016 Nov;13(11):919-922), the contents of each of which are incorporated herein by reference in their entirety.

[0029] Interaction Analysis - General Information and Non-Limiting Embodiments It will be understood that the following description of interaction analysis is not intended to be limiting, but is intended to illustrate the results and information obtained using embodiments of the method of the present invention. Other specific cell types, regions, tumor genes, etc. can be used in embodiments of the present invention. Additional information regarding the analysis of interactions between ecDNA and their target tumor genes is provided in the Examples section herein. Such information includes, but is not limited to, the presentation of the use of embodiments of the method of the present invention for the discovery of known and unknown ecDNA regions using ChIA-PET data, ChIP-seq library construction, data analysis, etc.

[0030] Certain embodiments of the method of the present invention included quantification of the degree of chromatin contact. In a non-limiting example, quantification is performed using a measure that represents the genome-wide trans-interaction frequency (nsTIF) normalized across all 23 chromosomes. ecDNA regions were identified as having highly elevated nsTIF levels in regions with ecDNA. In addition, high nsTIF regions linked with ecDNA segments show trans-contacts across the genome, indicating dynamic ecDNA connectivity from extrachromosomal genetic elements. The genome-wide contact pattern observed in this example is caused by the mobility of ecDNA. In this non-limiting example, a validation step was performed to verify that the high nsTIFs identified were specific to the extrachromosomal nature of ecDNA. In this example, the results confirmed that the elevated contact frequency of ecDNA across the genome was determined by its autonomous capabilities, rather than being the result of DNA dosing action alone.

[0031] In another non-limiting example, high frequency of cis interactions was detected within the genomic region of ecDNA. In HF-2927 ecDNA(+) cells, the observed cis interaction intensity within an approximately 530Kb ecDNA region was 2,879, a 240-fold increase compared to only 12 in the same region in ecDNA(-) HF-3035 cells. This high increase in contacts directly reflected both the size and genomic structure of this ecEGFR. Similarly, results showed strong cis interactions within and between two segments of a given ecMYC region in HF-2354. Extensive RNAPII-tethered chromatin contacts (defined as RNAPII binding detected at DNA regions connected by interactions, referred to as anchors) were detected both in cis between different regions within ecDNA (referred to as intraecDNA) and in trans with other genes or regulatory elements on linear chromosomes (referred to as trans-interactions). Extrachromosomal connectivity patterns were identified that clearly show pairs of loops with high frequency of interactions and focal points of strong contacts, which were predicted to collectively result from contacts between different ecDNA molecules and folding within individual ecDNAs. Among the trans interactions between ecDNAs and their chromosomal partners, anchors on ecDNA were identified as being primarily in intragenic or intergenic non-coding regions, and those trans-interacting chromosomal anchors were primarily localized to promoters. This juxtaposition of these interactions subserves the transcriptional function of these contacts.

[0032] In a non-limiting example of evaluating the interaction between ecDNA and transcriptional control regions, H3K27ac profiling is performed to show active enhancers and promoters.The control of oncogenes amplified in ecDNA is examined by evaluating their transchromosomal interaction regions.These ecDNA-connected non-coding chromosomal anchors show high overlap with H3K27ac peaks, which is significantly higher than that of the trans-interaction non-coding chromosomal anchors that do not have ecDNA contact, supporting the conclusion that the transcription of oncogenes on ecDNA is further enhanced by binding with enhancers on linear chromosomes through chromatin contact.

[0033] In a particular embodiment of the present invention, ecDNA interaction is evaluated by observing the coincidence of high frequency contact foci with H3K27ac peaks in ecDNA, and the results support the conclusion that these interaction anchors behave like active enhancers. In this non-limiting example of the interaction evaluation method, H3K27ac peaks in the 530Kb ecEGFR region co-align with regions of high interaction frequency in HF-2927, and compared with H3K27ac peaks in the chromosomal EGFR region of ecDNA(-) cells (HF-3035), they show a pattern of close clusters over a wider genome span, supporting the conclusion that enhancer signals were concentrated at chromatin contact sites of ecDNA. In this example, immunostaining of metaphase HF2927 cells using an antibody targeting H3K27ac showed overlapping signals between H3K27ac and DAPI signals indicative of ecDNA, thereby confirming the relationship between enhancer function and ecDNA.

[0034] In certain embodiments of the present invention, the method includes quantitatively evaluating the increase in H3K27ac signal associated with ecDNA-mediated transchromatin interactions, and it was found that H3K27ac peaks associated with ecDNA chromatin interaction anchors have significantly higher enrichment compared to those of genome-wide H3K27ac peaks that do not have ecDNA contacts. In this non-limiting example of the evaluation, the strong enhancement of H3K27ac signal was confirmed to be specific to ecDNA.

[0035] In another non-limiting example of the use of the method of the present invention, a method was used to evaluate the enhancer signature observed in ecDNA, reminiscent of "super enhancers". In this non-limiting example, a test was performed on the span size of the H3K27ac peaks detected in ecDNA and their trans-interacting chromosomal anchors. It was found that the H3K27ac peaks in ecDNA had significantly longer spans than the chromosomal H3K27 peaks that do not have ecDNA contacts. Sequence analysis showed enrichment of binding motifs of transcription factors important for controlling RNAPII global transcription and cell proliferation, including JUN, FOS, and ATF. Taken together, the convergence of both cis and trans RNAPII signals with strong enhancement of H3K27ac signals supports the conclusion that ecDNA molecules can connect to the RNA polymerase machinery extensively throughout the genome, supporting their function as genome-wide transcription amplifiers.

[0036] In a non-limiting example, the method of the present invention is used to determine whether the increase in enhancer signal associated with ecDNA trans-interaction leads to active transcription.In this example, RNA expression is tested, and the results show that ecDNA-interacting genes have significantly higher levels of expression compared to either genes that do not have other trans-chromosomal contacts, or genes that do not have ecDNA contacts but have trans-chromosomal interactions with other genes.Furthermore, in this example method, the expression levels of ecDNA-connected genes are identified as positively correlated with the frequency of their ecDNA contacts (measured by the number of independent trans-interactions), supporting the discovery that ecDNA connectivity is highly associated with transcriptional activity and highly enhanced H3K27ac signatures, suggesting that ecDNA may act as a global transcriptional amplification mechanism.

[0037] Besides the recruitment of individual oncogenes, ecDNA was identified as a focal point where multiple oncogenes are spatially proximate together via interactions with ecDNA. Many of these ecDNA-connected oncogenes reside within each of the chromatin networks, supporting the conclusion that oncogene coaggregation is a structure-based mechanism adopted by ecDNA to achieve coordinated transcriptional coactivation to promote tumorigenesis.

[0038] Cancer status assessment It has now been shown that ecDNA can enhance extrachromosomal and chromosomal gene transcription through chromatin interactions. An embodiment of the present invention includes a method for identifying ecDNA in a cell. Certain embodiments of the present invention provide a method for determining the effect of ecDNA on target genes, e.g., oncogenes, whose transcription is modulated by one or more ecDNAs. It has now been identified that ecDNA can enhance the expression of extrachromosomal and chromosomal gene transcription through chromatin interactions. This discovery, combined with the abundance and diversity of ecDNA, identifies ecDNA, ecDNA / target oncogene interactions, and target genes as targets for therapeutic intervention in diseases such as cancer. The method of the present invention is based, in part, on identifying interactions between genetic structure and epigenetic outcomes in tumor evolution. An embodiment of the method of the present invention provides a means for identifying ecDNA, identifying interactions between ecDNA and target genes, and identifying the effects of ecDNA interactions with target genes. The role of ecDNA activity in cancer and the unique genomic dynamics of this extrachromosomal structure provide new approaches for targeting ecDNA and their activated chromosomal and activated ecDNA target genes, for use, for example, in therapeutic applications.

[0039] Gene expression programs that establish and participate in the status or state of a cell include, but are not limited to, the activity of one or more transcription factors that bind to ecDNA gene regulators and gene regulatory elements of the target oncogene. A non-limiting example of a particular genomic element is an enhancer element that can bind to a transcription factor and loop over long distances to contact and regulate a particular gene. The interaction between ecDNA gene regulator sequences and gene regulatory elements, such as, but not limited to, the promoter of the target oncogene, and the interaction between ecDNA and gene regulatory elements on linear chromosomes are studied herein. Certain embodiments of the invention can be used to obtain information about the identity of extrachromosomal DNA (ecDNA) that participates in intracellular gene regulation processes through interactions with linear genomic elements, such as oncogene promoters. In addition, certain methods of the invention can be used to evaluate the regulatory interactions between two or more ecDNAs that participate in the gene expression program of a cell, including in some cases aberrant gene expression programs, such as those present in cancer cells.

[0040] Identifying candidate therapeutic agents and selecting cancer treatments In some aspects of the present invention, a method is provided for identifying the status of one or more oncogenes in cancer cells. The method of the present invention can be used to determine the level of transcription of one or more target oncogenes, where an increase in the level of one or more target oncogenes identifies the possibility of cancer. In some embodiments of the present invention, a plurality of cancer cells can be a source for obtaining cells for use in comparative studies and for testing candidate treatments. For example, and not intended to be limiting, the plurality of cancer cells can be cancer cells in culture or a subject and maintained in the same environment. In some embodiments of the present invention, one or more cancer cells from such a culture or subject are included in the method of the present invention to evaluate the status of the cells with respect to ecDNA / oncogene interactions. Different one or more cancer cells are contacted with a therapeutic agent or a candidate therapeutic agent, and the contacted cells are included in the method of the present invention to evaluate the status of the cells with respect to ecDNA / oncogene interactions. The ecDNA / oncogene interactions determined in non-contacted and contacted cancer cells can be determined and compared with each other or with a suitable control to obtain information regarding the effect of a therapeutic agent or a candidate therapeutic agent on ecDNA / oncogene interactions and the status of the cancer.

[0041] As used herein, the term "status" when used in reference to cancer cells refers to the presence or absence of one or more specific ecDNA / oncogene interactions. For example, and not intended to be limiting, the initial status of a cancer may include interactions between ecDNA and oncogenes A and B, and as the cancer progresses, its status may be determined to include interactions between ecDNA and oncogenes A, B, and C.

[0042] In some embodiments of the present invention, identifying tumor genes modulated by ecDNA provides information that can be used to help select a treatment for a subject with cancer. In some embodiments, a subject can be screened for a predisposition to cancer or to determine the stage of cancer present or suspected to be present in the subject. Method embodiments of the present invention can be used to screen for cancer or cancer status in a subject, and such methods can include one or more of identifying ecDNA / tumor gene interactions in cells and determining the modulating effect of ecDNA on target tumor genes. Such methods can be used to identify the status of cancer in a subject. In addition, the methods of the present invention can be used to evaluate the effect of a candidate agent on the modulating effect of ecDNA on its target tumor genes in cancer, and the results of the evaluation can be used to help select a treatment for cancer.

[0043] Embodiments of the methods of the present invention can be used to assess the status of cancer in one or more of a cell, a tissue, a subject, and a plurality of cells (or populations). As used herein, the term "cancer" is used to refer to a malignant neoplasm. Exemplary cancers include acoustic neuroma, adenocarcinoma, adrenal cancer, anal cancer, angiosarcoma, appendix cancer, biliary tract cancer (e.g., cholangiocarcinoma), bladder cancer, breast cancer (e.g., adenocarcinoma of the breast, papillary carcinoma of the breast, breast cancer, medullary carcinoma of the breast), brain cancer (e.g., meningioma, glioblastoma, glioma (e.g., astrocytoma, oligodendroglioma), medulloblastoma), cervical cancer (e.g., cervical adenocarcinoma), colorectal cancer (e.g., colon cancer, rectal ... colorectal adenocarcinoma), connective tissue cancer, epithelial carcinoma, ependymoma, endothelial sarcoma (e.g., Kaposi's sarcoma, multiple idiopathic hemorrhagic sarcoma), endometrial cancer (e.g., uterine carcinoma, uterine sarcoma), esophageal cancer (e.g., esophageal adenocarcinoma, Barrett's adenocarcinoma), Ewing's sarcoma, eye cancer (e.g., intraocular melanoma, retinoblastoma), generalized eosinophilia, gallbladder cancer, stomach cancer (e.g., gastric adenocarcinoma), gastrointestinal cancer, head and neck cancer (e.g., head and neck squamous cell carcinoma, oral cancer), pharyngeal cancer, hematopoietic cancer (e.g. leukemia, e.g. acute lymphocytic leukemia (ALL), lymphoma, e.g. Hodgkin's lymphoma (HL) and non-Hodgkin's lymphoma (NHL), multiple myeloma (MM), hemangioblastoma, kidney cancer (e.g. nephroblastoma, also known as Wilms' tumor, renal cell carcinoma), liver cancer (e.g. hepatocellular carcinoma (HCC), malignant hepatoma), lung cancer (e.g. bronchial Cancers of the present invention include, but are not limited to, myeloid leukemia, small cell lung cancer (SCLC), non-small cell lung cancer (NSCLC), adenocarcinoma of the lung, leiomyosarcoma (LMS), lipocytosis (e.g., systemic mastocytosis), malignant mesothelioma, muscle cancer, myeloproliferative disorders (MPD), neuroblastoma, neurofibroma, neuroendocrine carcinoma, osteosarcoma, ovarian cancer, papillary adenocarcinoma, pancreatic cancer, penile cancer, prostate cancer, rectal cancer, rhabdomyosarcoma, salivary gland cancer, skin cancer, melanoma, small bowel cancer, soft tissue sarcoma, sebaceous gland carcinoma, small intestine cancer, sweat gland carcinoma, synovium, testicular cancer, thyroid cancer, urethral cancer, vaginal cancer, and vulvar cancer.

[0044] Cancer may be primary or metastatic, and may be considered as early or late stage cancer, or the stage of cancer in a subject may be characterized by one or more known cancer staging classifications and conventions in the art. In some aspects of the invention, cancer is a first cancer in a subject, and in certain aspects of the invention, cancer may be a recurrence or reoccurrence of a previous cancer. In some cases, embodiments of the method of the invention may be used to assess the status of cancer in a subject who has not been treated with a cancer therapy. In certain embodiments, the method of the invention is used to assess the status of cancer in a subject who has been or is currently being treated with one or more cancer therapies. Non-limiting examples of cancer therapy include surgery, radiation therapy, chemotherapy, immunotherapy, dietary therapy, or other therapeutic approaches known in the art.

[0045] Certain embodiments of the present invention include methods for aiding in determining and / or selecting one or more treatment protocols for a subject. For example, and not intended to be limiting, some embodiments of the present invention can be used to aid in selecting a cancer treatment in a subject based at least in part on the status of ecDNA / oncogene interactions identified in cancer cells obtained from the subject. Using embodiments of the methods of the present invention to determine the status of cancer in a subject allows for the selection of one or more treatments based on the identified ecDNA / oncogene interactions. For example, and not intended to be limiting, the methods of the present invention can be used to detect the status of cancer in a subject through the identification of one or more ecDNA / target oncogene interactions in cancer cells obtained from the subject. The methods of the present invention can also be used to identify one or more specific oncogenes and other components in the detected ecDNA / interactions, and this information can be used to aid in selecting a cancer treatment in a subject. For example, if interactions between ecDNA and oncogene A and oncogene B are detected in cancer cells from a subject, this information can help select cancer treatments that result in one or more of: (i) reducing the interaction of ecDNA with oncogene A; and (ii) reducing the interaction of ecDNA with oncogene B. Based on the ecDNA / oncogene interaction information determined using embodiments of the methods of the present invention, the cancer in the subject can be classified by the specific oncogene / ecDNA interaction, and an appropriate treatment to reduce the specific oncogene / ecDNA interaction can be selected and administered to the subject.

[0046] In certain embodiments of the present invention, methods are provided that allow for determining the effectiveness of a cancer treatment administered to cancer cells or to a subject having, suspected of having, or at increased risk of having cancer. In a non-limiting example, the method embodiments of the present invention are used to determine an initial status of cancer in cancer cells obtained from a subject. The status of cancer is determined to include identified interactions between one or more ecDNAs and oncogene A, oncogene B, and oncogene C. A cancer treatment is selected for the subject based at least in part on the identified interactions of the ecDNAs with the three oncogenes. After administering the selected treatment to the subject, the method of the present invention is used to determine the subsequent status of cancer cells obtained from the subject after the treatment. The status of cancer determined in the cancer cells obtained previously, and the status of cancer determined in the cancer cells obtained after administration of the treatment, may indicate the effectiveness of the treatment against cancer in the subject. For example, the observation of an interaction between ecDNA and oncogene A, but no interaction between ecDNA and oncogene B or oncogene C, in cancer cells obtained from a treated subject, supports the efficacy of a cancer treatment in the subject and can confirm the efficacy of a treatment that reduces the enhancing interaction between ecDNA and oncogene B and oncogene C.

[0047] Non-limiting examples of oncogenes activated by ecDNA are epidermal growth factor receptor (EGFR), mouse double minute type 2 (MDM2), cyclin-dependent kinase 4 (CDK4), and cMYC. Drugs that can be used to treat cancers in which a particular oncogene is activated include, but are not limited to, EGFR inhibitors, MDM2 inhibitors, CDK4 inhibitors, and cMYC inhibitors. In a non-limiting example, a cancer identified using the method of the present invention as containing EGFR amplification due to the interaction of ecDNA with the EGFR oncogene can be treated by administering a tyrosine kinase inhibitor (TKI) drug to a subject with the cancer. Non-limiting examples of TKIs are gefitinib and erlotinib. In another non-limiting example, a cancer identified using the method of the present invention as containing increased CDK4 transcription caused by the interaction of ecDNA with the CDK4 oncogene can be treated by administering one or both of palbociclib and ribociclib. Based on the teachings presented herein, one of skill in the art would be able to select other art-known treatments based, at least in part, on the identification of one or more interactions between ecDNA and oncogenes that result in increased transcription of the oncogenes.

[0048] A non-limiting example of a cancer treatment may include administering to a subject diagnosed with, at increased risk of, or believed to have cancer an effective amount of an agent that disrupts and reduces the interaction between ecDNA and the target oncogene of the ecDNA. In an ecDNA gene regulatory system (GRS) in a cell, the ecDNA contains a sequence that may be referred to as a "gene regulator" sequence, non-limiting examples of which are gene actuator sequences and gene silencer sequences in the ecDNA. The GRS containing ecDNA also contains a "transcription factor", which is a complex that serves as a "contact point" between the ecDNA and its target oncogene. The transcription factor may constitute a complex of polypeptides that connects the gene regulator of the ecDNA to the target oncogene of the ecDNA through the transcription factor binding to the "gene regulatory element" of the target oncogene. For example, but not intended to be limiting, there is a promoter that controls the transcription of the target oncogene. In some embodiments, the method of the present invention includes identifying a candidate target present in the interaction between ecDNA and its target oncogene, and disruption of the identified candidate target disrupts the interaction between ecDNA and its target oncogene and reduces the enhancement of transcription of the target oncogene by ecDNA. As a non-limiting example, a polypeptide included in a transcription factor complex may be identified as a candidate target that reduces the enhancement of transcription of the target oncogene by ecDNA when contacted with a therapeutic agent that disrupts GRS.

[0049] In some embodiments of the present invention, the method includes contacting the cancer cells with a therapeutic agent that disrupts a candidate target in the GRS, reduces the interaction between the ecDNA and its target oncogene, and thereby reduces the enhancement of the transcription of the target oncogene by the ecDNA, and / or administering such a therapeutic agent to the subject. In some embodiments of the present invention, the candidate target includes a gene actuator in the ecDNA. As used herein, the terms "gene actuator" and "gene enhancer" may be used interchangeably to refer to a gene regulator. In some aspects of the present invention, the candidate target is one or more components of a transcription factor. In some embodiments of the present invention, the candidate target is a gene regulator element, such as, but not limited to, a promoter element of a target oncogene of the ecDNA. Thus, in certain embodiments of the present invention, one or more of the ecDNA activator, transcription factor, and gene regulator element may be identified as a candidate target to which one or more therapeutic agents are directed to treat cancer.

[0050] In some embodiments of the present invention, the therapeutic agent may be administered in combination with a second therapeutic agent. In some embodiments, the agent is administered in combination with the cancer therapeutic agent, for example, before, after, or during the administration or administration of the cancer therapeutic agent, or in combination with another cancer treatment, such as, but not limited to, one or more of radiation therapy, chemotherapy, and surgery. In some embodiments, the agent of the present invention is administered to a subject undergoing conventional chemotherapy and / or radiation therapy. In some embodiments, the cancer therapeutic agent is a chemotherapeutic agent. In some embodiments, the cancer therapeutic agent is an immunotherapeutic agent. In some embodiments, the cancer therapeutic agent is a radiation therapy agent.

[0051] cell It will be understood that the cell included in the method of the present invention can be one of a plurality of cells. As used herein, the term "plurality" of cells can refer to a cell population. The plurality of cells can all be of the same type and / or all have the same disease or condition. As a non-limiting example, a cell can be obtained from a liver cell population, and other cells obtained from this cell population are also liver cells. In some embodiments of the present invention, the plurality of cells can be a mixture of cell populations, meaning that not all cells are of the same type. In another non-limiting example, the cell can be a cancer cell obtained from a plurality of cancer cells. The cell used in the embodiment of the method of the present invention can be one or more of a single cell, an isolated cell, a cell that is one of a plurality of cells, a cell that is in a network of two or more interconnected cells, a cell that is one of two or more cells that are in physical contact with each other, and the like.

[0052] In some aspects of the invention, the cells may be obtained from a living animal, e.g., a mammal, or may be isolated cells. The isolated cells may be primary cells, e.g., cells recently isolated from an animal (e.g., cells that have not undergone population doublings and / or passages or have undergone only a few times after isolation), or cells of a cell line capable of long-term growth in culture (e.g., more than 3 months) or indefinite growth in culture (immortalized cells). In some embodiments of the invention, the cells are somatic cells. Somatic cells can be obtained from an individual, e.g., a human, and can be cultured according to standard cell culture protocols known to those skilled in the art. The cells can be obtained from surgical specimens, tissues, or cell biopsies, etc. The cells can be obtained from any organ or tissue of interest, including, but not limited to, skin, lung, cartilage, brain, breast, blood, blood vessels (e.g., arteries or veins), fat, pancreas, liver, muscle, digestive tract, heart, bladder, kidney, urethra, and prostate. In some embodiments of the invention, the cells are HF-3035 cells or HF-2354 cells.

[0053] In some embodiments, the cells used in conjunction with the present invention may be healthy normal cells that are not known to have a disease, disorder, or abnormal condition. In some embodiments, the host cells used in conjunction with the methods and compositions of the present invention are abnormal cells, such as cells obtained from subjects diagnosed with a disorder, disease, or condition, including, but not limited to, degenerated cells, neurological disease-bearing cells, cell models of a disease or condition, damaged cells, and the like. In some embodiments of the present invention, the cells may be control cells. In some aspects of the present invention, the host cells may be model cells of a disease or condition.

[0054] The cell that can be used in certain embodiments of the present invention is a human cell. Non-limiting examples of cells that can be used in the method of the present invention are one or more of eukaryotic cells, vertebrate cells, which can be mammalian cells in some embodiments of the present invention. Non-limiting examples of cells that can be used in the method of the present invention are vertebrate cells, invertebrate cells, and non-human primate cells. Further non-limiting examples of cells that can be used in the method of the present invention are one or more of rodent cells, canine cells, feline cells, avian cells, fish cells, cells obtained from wild animals, cells obtained from livestock animals, and other suitable cells of interest. In some embodiments, the cell is an embryonic stem cell or an embryonic stem cell-like cell. In some embodiments, the cell is a neuronal cell, a glial cell, or other types of central nervous system (CNS) or peripheral nervous system (PNS) cells. In some embodiments, the cell is an astrocyte. In some embodiments of the present invention, the cell is a natural cell, and in certain embodiments of the present invention, the cell is an engineered cell.

[0055] The cells useful in the embodiments of the method of the present invention can be maintained in cell culture after they are isolated. The cells can be genetically modified or not genetically modified in various embodiments of the present invention. The cells can be obtained from normal or diseased tissue. In some embodiments, the cells are obtained from a donor and their state or type is modified ex vivo using the method of the present invention. In certain embodiments of the present invention, the cells can be floating cells in culture, free cells obtained from a subject, cells obtained in a solid biopsy obtained from a subject, an organ, or solid culture, etc.

[0056] The population or plurality of isolated cells in any embodiment of the present invention may be composed primarily or essentially entirely of cells of a particular cell type or state. In some embodiments, the isolated cell population is composed of at least 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% of cells of a particular type or state (i.e., the population is at least 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% pure), for example, as determined by expression of one or more markers or any other suitable method.

[0057] Control Certain embodiments of the methods of the invention are used to assess one or more of the following: the effect of ecDNA on a target oncogene, the status of a cell with respect to ecDNA / oncogene interaction, the effect of a candidate therapeutic agent on the interaction between ecDNA and its target oncogene, etc. Such assessment of ecDNA / target oncogene characteristics in a cell, tissue, and / or subject can be performed by comparing the results obtained in a sample cell, tissue, or subject with the results obtained in a control cell, tissue, or subject, respectively. As a non-limiting example, some embodiments of the invention include determining the status of one or more ecDNA target oncogenes in a sample cancer cell and a control cancer cell, and comparing the results as a measure of the difference in the status of the sample cancer cell and the control cancer cell. In another non-limiting example, the status of ecDNA / target oncogene interaction is identified in a subject with cancer, followed by administering a candidate therapeutic agent intended to disrupt the identified ecDNA / target oncogene interaction to the subject, and comparing the status before and after administration of the candidate therapeutic agent. It should be understood that results obtained from subjects who have not yet been contacted with a candidate therapeutic agent may be referred to as "control results," and the non-contacted subjects may be referred to as "control subjects."

[0058] As used herein, a control can be as described above and can also be a predefined value that can take various forms. It can be a single cut-off value, such as a median or mean. It can be constructed based on a comparison group. Other examples of a comparison group can include cells or subjects with a particular cancer or ecDNA / target tumor gene status, and cells or subjects without a particular cancer or ecDNA / target tumor gene status. Another comparison group can be subjects from a group with a family history of cancer and subjects from a group without such a family history. For example, a predefined value can be set that divides the population to be tested into groups evenly (or unequally) based on the test results. A person skilled in the art can select a suitable control group and value to be used in the comparison method of the present invention.

[0059] The candidate therapeutic agent identification methods of the invention can be carried out in a cell or cells present in a subject or in a host cell in culture or in vitro. The candidate therapeutic agent identification methods of the invention carried out in a subject can include delivering a candidate agent intended to disrupt ecDNA / target oncogene interaction to the subject's cells, and evaluating the ecDNA / target oncogene interaction and oncogene status (before and / or after delivery of the candidate therapeutic agent). The results of contacting the host cells, tissues, and / or subjects with the candidate therapeutic agent can be measured and compared to a control value as a determination of the effectiveness of the candidate drug in disrupting ecDNA / target oncogene interaction.

[0060] composition The composition used in the method of the present invention may be, but is not necessarily, a pharmaceutical composition. The term "pharmaceutical composition" as used herein means a composition that is generally safe, non-toxic, and not biologically or otherwise undesirable, and includes at least one pharma- ceutical acceptable carrier that is useful for preparing a pharmaceutical composition. Pharmaceutical compositions can be used in certain embodiments of the method of the present invention, a non-limiting example of which is for administering a candidate therapeutic agent to a cell or subject to disrupt ecDNA / target oncogene interaction.

[0061] In certain aspects of the invention, the pharmaceutical composition comprises one or more therapeutic agents or candidate therapeutic agents, together with one or more other additional molecules, therapeutic agents, candidate agents, candidate treatments, and treatment regimens that are also administered to the cells and / or the subject. The pharmaceutical composition used in the method embodiments of the invention may comprise an effective amount of the candidate therapeutic agent to perform one or more of the following: reduce ecDNA / target oncogene interaction, alter the status of target oncogene transcription in cancer cells, etc. In some embodiments of the invention, the pharmaceutical composition of the invention may comprise a pharmaceutically acceptable carrier.

[0062] Pharmaceutically acceptable carriers include diluents, fillers, salts, buffers, stabilizers, solubilizers and other materials that are well known in the art.Exemplary pharma-ceutically acceptable carriers are listed in U.S. Patent No. 5,211,657, and others are known to those skilled in the art.In certain embodiments of the present invention, such preparations may contain salts, buffers, preservatives, compatible carriers, aqueous solutions, water, etc.

[0063] The delivery of therapeutic agent to cells or subjects can be achieved by various means described herein and other means known in the art. Such administration can be performed once or multiple times. When administered to a subject multiple times, one or more therapeutic agents can be administered by a single route or different routes. For example, and not intended to be limiting, the first (or first few) administrations can be directly administered to the tissue of the subject to be treated, and subsequent administrations can be systemic.

[0064] The amount of therapeutic agent delivered to a cell or subject may be, in certain embodiments of the invention, an amount that statistically significantly reduces the interaction of ecDNA with its target oncogene. A suitable amount can be readily determined by the practitioner using methods known in the art, e.g., in conjunction with clinical trials, and using the teachings provided herein, without the need for undue experimentation. EXAMPLES

[0065] Example 1 Within the nucleus, chromosomes undergo extensive folding into chromatin loops that occupy distinct chromatin territories [Zheng, S. et al. (2013) Genes Dev 27, 1462-1472]. Such highly organized 3D chromatin architectures provide the topological basis for many genomic functions, including transcription, by bringing distant regulatory elements and their targeted genes into spatial proximity [Cremer, T. & Cremer, M. (2010) Cold Spring Harb Perspect Biol 2, a003889]. Alterations in chromatin organization as a result of chromosomal rearrangements have been implicated in a number of human diseases, particularly cancer [Sexton, T. & Cavalli, G. (2015) Cell 160, 1049-1059]. To understand the chromatin organization of ecDNA and how it contributes to gene transcription regulation, we applied chromatin interaction assessment methods, such as ChIA-PET [Taberlay, PC et al. (2016) Genome Res 26, 719-731]. We designed a method to incorporate both general spatial chromatin organization [Zhang, Y. et al. (2013) Nature 504, 306-310] and long-range chromatin interactions mediated by protein factors into the same neurosphere cell line. Studies have shown that known ecDNAs are easily recognizable through their strong and unusual intra- and intermolecular genome-wide chromatin contacts. In addition, in deciphering the RNA polymerase II (RNAPII)-mediated ecDNA connectome and their chromosomal partners, a relationship between ecDNA and actively expressed autosomal oncogenes was identified. This relationship supports the finding that ecDNA functions as a mobile transcriptional enhancer that promotes tumor progression.

[0066] method Culture of neurosphere cells derived from GBM patient tumors Neurosphere cell lines were generated and cultured as described [deCarvalho, AC et al. (2018) Nat Genet 50, 708-717]. Brain tumor specimens were obtained with written informed consent from patients under a protocol approved by the Henry Ford Hospital Institutional Review Board. Briefly, tumor specimens were sectioned and cultured as neurospheres in DMEM / F12 medium (11330-032, Gibco) supplemented with N-2 supplement (17502-048, Gibco) and growth factors (EGF and FGF-basic). Cells from passages 15–26 were harvested for experiments.

[0067] ChIA-PET experiments and data analysis Ten million cells were double crosslinked with 1.5 mM EGS (21565, Thermo Fisher) for 45 min, followed by 1% formaldehyde (F8775, Sigma) for 20 min at room temperature (RT), then quenched with 0.125 M glycine (G8898, Sigma) for 10 min. Crosslinked cells were washed twice with 1× PBS, lysed in 100 μL of 0.55% SDS, incubated at room temperature, 62°C, and 37°C for 10 min each, followed by quenching the SDS by adding 25 μL of 25% Triton-X 100 at 37°C for 30 min, and chromatin was fragmented by adding 50 μL of AluI (R0137L, NEB), 50 μL of 10× CutSmart buffer, and 275 μL of HO at 37°C overnight. The pelleted digested nuclei were resuspended in 500 μL of dA-tailing solution containing 50 μL of 10× CutSmart buffer, 10 μL BSA (B9000S, NEB), 10 μL of 10 mM dATP (N0440S, NEB), 10 μL of Klenow (3′- 5′ exo-) (M0202L, NEB), and 420 μL of HO, incubated for 1 hour at room temperature, and then subjected to proximity ligation by adding 200 μL of 5× ligation buffer (B6058S, NEB), 6 μL of biotinylated cross-linker (200 ng / μL), 10 μL of T4 DNA ligase (M0202L, NEB), and 284 μL of HO and incubating overnight at 16° C. The ligated chromatin was then sheared by sonication and immunoprecipitated with anti-RNAPII antibody (920102, Biolegend).Tagmentation of immunoprecipitated DNA, biotin selection, library preparation, and sequencing were performed as described [Tang, Z. et al. (2015) Cell 163, 1611-1627].

[0068] ChIA-PET data were processed using ChIA-PET utilities, an extensible reimplementation of the ChIA-PET tool [Li, G. et al. (2010) Genome Biol 11, R22] (see code availability). After removing sequencing adapters, paired-end reads with cross-linkers were identified and tags adjacent to the linkers were extracted. Identified tags (16 bp or longer) were mapped to hg19 using BWA alignment [Li, H. & Durbin, R. (2009) Bioinformatics 25, 1754-1760] and mem [Li, H. (2013) arXiv:1303.3997 [q-bio.GN]] depending on tag length. Uniquely mapped non-redundant paired-end tags (PETs) were classified as interchromosomal (left and right tags originate from different chromosomes), intrachromosomal (left and right tags are genomic spans of >8Kb), and self-ligation PETs (left and right tags are genomic spans of ≤8Kb). Both interchromosomal and intrachromosomal PETs were extended by 500bp. PETs overlapping at both ends were then clustered as iPET-2,3... Interactions overlapping with chr M, chr Y were not tested in this study. Interactions where the anchors overlapped with the blacklist (see below for how they were defined) were filtered to remove possible noise caused by genomic sequence organization and tagmentation bias created by Tn5 digestion. For intrachromosomal interactions, statistical assessment of interaction significance was performed using ChiaSigScaled, an extensible reimplementation of ChiaSig [Paulsen, J. et al. (2014) Nucleic Acids Res 42, e143].Significant interactions defined as iPET ≥ 3 and FDR < 0.05 and interchromosomal interactions as iPET ≥ 2 were used in all downstream analyses, except for nsTIF analysis, which used all reported interchromosomal interactions (see section on finding known ecDNA regions from in situ ChIA-PET data below). RNAPII binding peaks were called with all uniquely mapped reads using MACS2 (options: --keep-dup all --nomodel --extsize 250) [Liu, T. (2014) Methods Mol Biol 1150, 81-95]. To define intra-ecDNA interactions, ecDNA regions were used to collect all interactions within the reported regions. For ecDNA-mediated trans-chromosomal interactions, only interactions originating from chromosomes other than the chromosome where ecDNA resides were included. RNAPII binding status at both anchors of the interaction was examined, and RNAPII-mediated interactions were defined as interactions with RNAPII binding at both anchors. Interactions were further classified based on anchors overlapping with GENCODE gene models (19 releases, all pseudogenes, all RNAs other than miRNAs excluded), prioritizing promoter (P) regions (defined as ±2.5kb of the TSS), followed by genic regions (G). Anchors not overlapping with any intragenic regions were classified as intergenic (I). ecDNA-interacting genes were annotated using oncogenes derived from the union list of NCG 6 [Repana, D. et al. (2019) Genome Biol 20, 1] and COSMIC v87 [Forbes, SA et al. (2015) Nucleic Acids Res 43, D805-811].

[0069] Blacklist Area To remove biases introduced by the ChIA-PET experimental procedure, such as over-tagging by Tn5 at certain genomic loci, eight human cell ChIA-PET libraries generated by different antibody enrichments (four with anti-CTCF and four with anti-RNAPII) were used to create a greylist and downsample to represent the same number of reads totaling 75,215,727 tags. Peaks were called from the combined dataset with an FDR of less than 0.05 using MACS2.1.0.20151222 [Liu, T. (2014) Methods Mol Biol 1150, 81-95], which resulted in 153,735 peaks. These peak regions on the autosomes and X chromosome were candidates for greylisting. Regions with short peaks (likely due to Tn5 tagmentation) were further filtered by the following criteria, while retaining the highest confidence as measured by q-value and pileup: region length less than 600 bp, top 1% pileup, bottom 10% q-value, and top 10% peak enrichment fold. These three quantities or vectors (reciprocal length, enrichment fold, and pileup) of the filtered peaks (1,119 regions with q-value less than 1E-165 and enrichment fold between 10 and 50) were taken and each of the vectors was scaled (using the R command "scale(center=F,scale=T)"). The three vectors were then normalized separately, so that the average of each vector was 1. The scoring function was defined as s = enrichment fold + pileup + 1 / length. The grey list consisted of 321 peak regions with score s above the average. In addition, four regions that were seen as artifacts by visual inspection were included.The final blacklist adopted was the concatenation of the ChIA-PET greylist and the publicly available blacklist from Kundaje lab (github.com / kundajelab / HiC-pipeline / blob / master / hic_flexibleWindow-pipeline / data / reference_genomes / hg19 / wgEncodeHg19ConsensusSignalArtifactRegions.bed.gz).

[0070] Discovery of known ecDNA regions from in situ ChIA-PET data The processed interaction data obtained by RNAPII ChIA-PET from each cell line were analyzed to generate a genome-wide interaction frequency (IF) matrix M NxN ={M ij The IFs were aggregated to 100,000 |i,j=1,2,..,N}. The hg19 genome was segmented into 60,739 non-overlapping bins at 50 Kb intervals from the beginning of the chromosome (the last bin of each chromosome may not represent the full 50 Kb). Bins overlapping with the blacklist (see below) were removed from the IF matrix. Known ecDNA regions exhibited a large number of interactions, both within the ecDNA region and broadly across all 23 chromosomes, resulting in very high IF sums, especially between different chromosomal regions amplified in ecDNA. The method used exploited these properties to test whether known ecDNA regions [deCarvalho, AC et al. (2018) Nat Genet 50, 708-717] could be revealed from the ChIA-PET interaction data, and for possible prediction of further genomic regions amplified in ecDNA.

[0071] The sum of transchromosomal IFs (TIFs) was calculated for all bins and normalized to allow comparison across different libraries. This vector of data was scaled (divided by its magnitude and multiplied by the length of the vector) so that the mean was equal to 1. The i-th bin of this normalized vector was then denoted as nsTIF i The bins with the highest nsTIFs were compared to known ecDNA regions in the ecDNA(+) data. To understand the distribution of nsTIFs in ChIA-PET data obtained from ecDNA(-) cells, the genome-wide distribution of nsTIFs was tested in HF-3035 cells (Figure 1B) as well as other pluripotent cell lines (data not shown) and found that they were all below 20. Therefore, a threshold was introduced as a first pass to determine ecDNA candidate regions. An additional nsTIF threshold was also introduced to prioritize bins as candidates for ecDNA. Based on knowledge, the gene regions amplified by ecDNA did not exceed 0.1% of the genome size, so a low threshold (t l ) in the top 0.1% of nsTIFs i Empirically, a high threshold (t h ) is t l If is less than 25, set it to 25 (i.e., 25 times higher than expected), otherwise, t h =max(nsTIF). i t l All bins with a size greater than i were put into a list of candidates. i t h Any region was not considered ecDNA if its nsTIF score was below 1. In this study, the highest nsTIF obtained from the ecDNA(-) data was less than 21, substantially lower than that in the ecDNA(+) data set (~38-82).

[0072] The list of candidates was grouped based on their genomic distance such that groups located near each other were grouped together, i.e., the minimum distance between two groups was 1 Mb. The goal of the grouping was to score all candidates according to their fixed connections (measured by IF) to candidates of different groups. The score was based on the number of groups to which each candidate was connected. The normalized TIF of a candidate was calculated based on the number of groups to which it was connected, and the normalized TIF was calculated based on the number of groups to which it was connected, and the normalized TIF of the candidate ... T "Connectivity" is defined as when the normalized interchromosomal elements of matrix M are higher than T ij Then the normalized TIF vector for bin i is T i (The same normalization as for nsTIF, i.e.

[0073]

number

[0074] ). Then, if their interaction frequency is relatively high, for example, there is at least one bin, e.g., the k-th bin, in another candidate set, and T ik t T (where t T is 100 and the top 0.1% T i A connection exists if i is the higher of the averages of i and . A connection score of +1 is then added to candidate i.

[0075] Next, the set of candidates with the highest nsTIFs was first examined to see if all of their connection scores were zero. If this condition was true, then the nsTIFs were hIf there were no more than five groups with a connectivity score of 0 or higher, all of the candidate regions in this single group were proposed to be ecDNA regions, otherwise no ecDNA regions were predicted (noisy data may present many groups with a connectivity score of 0 and above the high nsTIF threshold, which is likely to be a false positive). If this condition was false, i.e., there were multiple groups with a non-zero connectivity score, regions from these groups with a connectivity score above zero were predicted to be ecDNA. Additional studies were performed to assign the specificity of regions amplified in ecDNA from ChIA-PET data.

[0076] ChIP-seq library construction and data analysis Two million cells were crosslinked and lysed in the same manner as ChIA-PET. After lysis, the nuclear pellet was sonicated and immunoprecipitated with anti-H3K27ac antibody (39133, Active Motif). 4ng of DNA from both antibody immunoprecipitation and input was subjected to end repair, A-tailing, and adapter ligation using KAPA Hyper Prep Kit (KK8505, KAPA Biosystems). Adapter-ligated DNA fragments were PCR amplified with KAPA Library Amplification ReadyMix (KK2612, Kapa Biosystems) and sequenced on the Illumina platform with 75bp single-end sequencing. Raw reads were quality cleaned using Trim Galore version 0.4.3 (options: --stringency 3 -q 30 -e .20 --length 15) and mapped to the hg19 genome using BWA 0.7.12 (command: aln) [Li, H. & Durbin, R. (2009) Bioinformatics 25, 1754-1760]. Uniquely mapped and duplicated de-duplicated reads were used for peak calling (FDR less than 0.05) with MACS2.1.0.20151222 (options: --nomodel --extsize 250 -B --SPMR -g hs) [Liu, T. (2014) Methods Mol Biol 1150, 81-95]. Peaks defined with FDR less than 0.05 were used in all analyses. For the H3K27ac analyses in Figures 2C-D, we further applied the requirement of P <0.001.

[0077] immunostaining Unfixed metaphase cells were dropped onto slides and preincubated in KCM buffer (120 mM KCl, 20 mM NaCl, 10 mM Tris-HCl, pH 8.0, 0.5 mM EDTA, 0.1% (vol / vol) Triton X-100) for 10 min at room temperature. Slides were blocked in 1% (wt / vol) BSA / KCM buffer for 30 min at room temperature, followed by incubation with primary H3K27ac antibody (39133, Active Motif) in 2% BSA overnight at 4°C, two incubations in KCM buffer for 10 min at room temperature, and incubation with goat anti-rabbit Alexa Fluor 488 secondary antibody (A32731, Invitrogen) for 30 min at room temperature. After washing twice with KCM buffer, slides were crosslinked with 4% (vol / vol) formaldehyde / KCM for 15 min, coverslipped with 50 μL of Prolong Gold Antifade (Invitrogen), and sealed with clear nail polish. Slides were scanned under a Leica STED 3X / DLS confocal microscope.

[0078] Transcription factor motif analysis 414 known motifs were examined in 206 target sequences against a normalized background of 71,232 H3K27ac peak regions using Homer2 [Heinz, S. et al. (2010) Mol Cell 38, 576-589]. Search results are selected based on the following criteria: q-value < 0.001, enrichment > 1.5, % for target sequences with motif > 25%.

[0079] RNA-seq library construction and data analysis Total RNA was isolated in biological replicates using the AllPrep DNA / RNA Mini Kit (80204, QIAGEN). Strand-specific RNA libraries were generated from 300 ng of total RNA using the KAPA Stranded mRNA Sequencing Kit (KK8502, KAPA Biosystems) according to the manufacturer's instructions. Libraries were sequenced on the Illumina platform with 75 bp paired-end sequencing. Raw sequencing reads were trimmed using Trim Galore version 0.4.3 (options: --stringency 3 -q 20 -e .20 --length 15 --paired) and aligned to the hg19 genome using hisat 2.1.0 (options: --dta-cufflinks). Transcripts were assembled using Cufflinks 2.2.1 [Trapnell, C. et al. (2010) Nat biotech 28, 511-515] and final expression levels were quantified using Cuffdiff (options: --library-type fr-firststrand) [Trapnell, C. et al. (2013) Nat biotech 31, 46-53]. Correlations between sequencing data obtained from different samples were analyzed using the R / pheatmap package (version 1.0.2.).

[0080] Defining ecDNA-mediated chromatin interaction communities A collection was performed for all trans-chromosomal interactions supported by RNAPII binding at both anchors. All anchors bound to RNAPII from HF-2354, HF-2927, HF-3016, and HF-3177 cell lines were pooled, and anchors overlapping with blacklist regions were removed and merged into 29,721 non-overlapping interaction nodes (Figure 3C). For each node, a connectivity score was defined, i.e., the number of interaction partner nodes to which the node is connected. Hubs were defined as nodes with a connectivity score higher than 3 standard deviations above the mean. With this parameter, a hub has more than 10 connections to other nodes. In total, 69, 106, 82, and 99 hubs were defined from HF-2354, HF-2927, HF-3016, and HF-3177 cell lines, respectively. Using the hub-to-hub network, communities were generated using the cluster_edge_betweenness function in the igraph library of R [Csardi, G. & Nepusz, T. (2006) Computer Science]. The number of communities associated with ecDNA is 8, 8, 10, and 4 in HF-2354, HF-2927, HF-3016, and HF-3177 cell lines, respectively.

[0081] Data Availability Statement All data described in this study have been deposited in the NCBI Gene Expression Omnibus GSE124769 using the following link www.ncbi.nlm.nih.gov / geo / query / acc.cgi?acc=GSE124769 with the reviewer linker key “uxkhaeyctdixts.”

[0082] Code Availability Code for ecDNA detection from ChIA-PET data is available at www.dropbox.com / sh / 2crbfjt1kr2yyws / AACLiUg6Ch9y6FurbbUzmcbPa?dl=0. ChIA-PET tool (code available at github.com / cheehongsg / CPU) ChiaSigScaled (code available at github.com / cheehongsg / ChiaSigScaled) Results and Discussion ChIA-PET analysis was performed on five GBM patient-derived neurosphere cell lines whose ecDNA status had already been established from whole genome sequencing data and confirmed by fluorescence in situ hybridization (FISH) analysis. Four of the five neurosphere lines were ecDNA(+) (HF-2354, HF-2927, HF-3016, and HF-3177) and one line was ecDNA(-) (HF-3035) [deCarvalho, AC et al. (2018) Nat Genet 50, 708-717]. Based on the high levels of RNA expressed from genes amplified within ecDNA (Figure 4A), we determined that ecDNA was highly associated with RNAPII within active chromatin domains. We used RNAPII chromatin immunoprecipitation to elicit RNAPII-associated chromatin and ChIA-PET assays to characterize the ecDNA-chromatin interactome (Figure 5A). The obtained ChIA-PET data detected long-range chromatin interactions between both RNAPII binding sites, regulatory elements [Zhang, Y. et al. (2013) Nature 504, 306-310] (Figure 4B), as well as non-enriched chromatin contacts within spatial topologically chromatin associated domains (TADs) [Dixon, JR et al. (2012) Nature 485, 376-380] (Figure 6A). Chromosomal structural variants, such as deletions of PTEN and CDKN2A and CDKN2B at chromosomes 10q23 and 9p21, resulted in the elimination of chromatin contacts (Figure 6B). Other examples included the detection of a 600 Kb deletion at chrX:31.4-32 Mb common fragile site involving the DMD gene in HF-2927 [Ma, K. et al. (2012) Int J Mol Sci 13, 11974-11999], a 15 Mb extensive rearrangement at chr3:168-183 Mb, and a 3.5 and 11.5 Mb double transposition event between chr3 and chr6 in the HF-2354 genome (Figure 6C).

[0083] HF-2927 has ecDNA containing chr7p11 / EGFR, designated ecEGFR [deCarvalho, AC et al. (2018) Nat Genet 50, 708-717], while HF-2354 contains chr8q24 / MYC ecDNA, designated ecMYC. In two neurosphere lines, HF-3016 and HF-3177, derived from primary and recurrent GBM from the same patient, three genes were found to be amplified extrachromosomally [deCarvalho, AC et al. (2018) Nat Genet 50, 708-717], among which chr7p11 / EGFR and ch12q14.1 / CDK4 were shown to be co-amplified on ecDNA, while chr8q24 / MYC was also found on ecDNA (designated ecEGFR, ecCDK4, and ecMYC, respectively). All ecDNA loci exhibited extensive contacts across the genome (Figure 5B, Figure 1A), suggesting high trans connectivity of these ecDNA regions to regions across all chromosomes. To quantify the extent of transchromatin contacts, we developed a measure that faithfully describes the normalized genome-wide trans interaction frequency (nsTIF) across all 23 chromosomes and applied this measure to each of the five neurosphere cell lines (see Methods herein). ecDNA regions exhibited highly elevated nsTIF levels (maximum nsTIFs were approximately 38-82) in all four ecDNA(+) lines, but not in HF-3035 ecDNA(-) cells (nsTIFs less than 21) (Figure 5B, Figure 1B). nsTIF spike regions closely matched ecDNA regions. Furthermore, high nsTIF regions linked to ecDNA segments exhibited trans contacts across the genome (Figure 5B), suggesting dynamic ecDNA connectivity from extrachromosomal genetic elements. The characteristic genome-wide contact patterns may be explained by the mobility of ecDNA.In the two diamplicon lines, HF-3016 and HF-3177, results showed elevated nsTIF levels for ecMYC, ecEGFR, and ecCDK4 (Figure 1B), as well as cross-interactions between the three loci, suggesting that the dominant ecDNA in these lines harbors all three oncogenes or that they are in close intermolecular proximity. To validate that the high nsTIFs were specific to the extrachromosomal nature of the ecDNA and not its amplification status, nsTIF values ​​from all genomic regions with copy number gains ≥3 (range 3–6) were compared to nsTIFs from the ecDNA regions. The nsTIF values ​​obtained from chromosomal amplified segments were significantly lower than those from ecDNA regions (median nsTIFs of 1.5–4 vs. 24–43) as expected for chromosomal territories that are restricted within the chromosomal territory (one-sided Wilcoxon rank sum test, P value < 0.0005) ( Figure 5C ), confirming that the increased genome-wide ecDNA contact frequency is not explained by DNA dosage effects alone but is determined by its autonomous capacity.

[0084] In addition to the high trans interaction frequency, an extremely high frequency of cis interactions was detected within the genomic region of ecDNA. In HF-2927 ecDNA(+) cells, the cis interaction intensity observed within a ∼530 Kb ecDNA region was 2,879, a 240-fold increase compared to only 12 in the same region in ecDNA(−) HF-3035 cells (Figure 7A). This high increase in contacts directly reflected both the size and genomic structure of this ecEGFR. Similarly, strong cis interactions were observed within and between two segments of the defined ecMYC region in HF-2354 (Figure 7A). Extensive RNAPII-tethered chromatin contacts (defined as RNAPII binding detected at DNA regions connected by interactions, referred to as anchors) were detected both in cis between different regions within ecDNA (referred to as intraecDNA) (Figure 7B, Figure 8A) and in trans with other genes or regulatory elements on linear chromosomes (referred to as trans interactions) (Figure 7C). The intra-extrachromosomal connectivity patterns clearly show distinct pairs of loops with high frequency of interactions and a concentration of strong contacts (Figure 7B, Figure 8A), which may collectively result from contacts between different ecDNA molecules and folding within individual ecDNA. This hypothesis is supported by primary and matched recurrent neurosphere lines (HF-3016 vs. HF-3177), both of which contain ecMYC, ecEGFR, and ecCDK4, but 92% and 95% of the ecDNA intra-loops detected in HF-3016 and HF-3177 are found exclusively in the respective cells. The ecDNA intra-loops detected in HF-3016 are mainly located at the ecCDK4 locus (inner circle in Figure 8A), whereas in HF-3177 the ecMYC region shows stronger looping (inner circle in Figure 7B), which may indicate that ecDNA arises from a different, separate structure, although it may contain similar oncogene effectors.Among the trans interactions between ecDNA and their chromosomal partners, the anchors on ecDNA were predominantly (75-93%) located in intragenic or intergenic noncoding regions, and the trans interacting chromosomal anchors were predominantly (79-84%) localized to promoters (defined as TSS ± ​​2.5 Kb) (Figure 7C). Such juxtaposition of these interactions suggests a transcriptional function for these contacts.

[0085] To investigate how ecDNA interactions relate to transcriptional regulatory regions, H3K27ac profiling was performed to indicate active enhancers and promoters. The regulation of oncogenes amplified in ecDNA was first tested by assessing their trans-chromosomal interacting regions. These ecDNA-connected non-coding chromosomal anchors showed high overlap with H3K27ac peaks (61-80% intergenic), which was significantly higher than that by trans-interacting non-coding chromosomal anchors without ecDNA contacts (38-69% intergenic, Figure 8B, P value 0.019, one-sided Wilcoxon rank sum test). Specifically, 73% (144 out of 196) of chromosomal non-coding regions interacting with promoters of oncogenes present on ecDNA overlapped with H3K27ac peaks in the corresponding cell lines (Figure 2A), suggesting that transcription of oncogenes on ecDNA is further enhanced by binding to enhancers on linear chromosomes through chromatin contacts. Enhancer contacts can be variable and dynamic among different ecDNA(+) lines. In the case of ecMYC, the MYC promoter interacted with nine different H3K27ac enhancers in three ecMYC(+) cell lines (Figure 8C).

[0086] Co-occurrence was observed between high frequency contact foci and H3K27ac peaks in ecDNA (Figure 7B, Figure 8A), suggesting that these interacting anchors behave like active enhancers. H3K27ac peaks in the 530 Kb ecEGFR region co-aligned with regions of high interaction frequency in HF-2927 and showed a pattern of closer clusters over a broader genomic span compared to H3K27ac peaks in the chromosomal EGFR region in ecDNA(-) cells (HF-3035) (Figure 2B), indicating that enhancer signals are aggregated at chromatin contact sites of ecDNA. Immunostaining of metaphase HF2927 cells using an antibody targeting H3K27ac showed overlapping signals between H3K27ac and DAPI signals indicative of ecDNA, confirming the relationship between enhancer function and ecDNA (Figure 8D).

[0087] To quantitatively demonstrate the increase in H3K27ac signals associated with ecDNA-mediated trans-chromatin interactions across all four ecDNA(+) cell lines, a comparison of enrichment folds from all detected H3K27ac peaks (FDR<0.05, P<0.001) was performed between ecDNA regions (referred to as group A), their corresponding trans-interacting chromosomal partners (group B), and whole-genome H3K27ac peaks with no ecDNA contacts (group C) in each of the four ecDNA(+) cell lines. H3K27ac peaks associated with ecDNA chromatin interacting anchors had significantly higher enrichment (median, group A: 58-138 and group B: 43-91) than those of whole-genome H3K27ac peaks with no ecDNA contacts (median 10-12, P-value 5E-09-2.3E-164, one-sided Wilcoxon rank sum test) (Figure 2C). They were also higher than the fold enrichment (median: 9–11) from regions corresponding to ecMYC, ecEGFR, and ecCDK4 found in the ecDNA(−) HF-3035 cell line, confirming that the strong enhancement of H3K27ac signals was specific to ecDNA.

[0088] Enhancers with very high intensity and large domains of H3K27ac signals have been termed "super enhancers" [Whyte, WA et al. (2013) Cell 153, 307-319] and have been found to promote the transcription of oncogenes in cancer [Hnisz, D. et al. (2013) Cell 155, 934-947]. To assess whether the enhancer signature observed in ecDNA is reminiscent of a "super enhancer," we examined the span size of the H3K27ac peaks detected in ecDNA and their trans-interacting chromosomal anchors. H3K27ac peaks in ecDNA were found to have significantly longer spans than chromosomal H3K27 peaks without ecDNA contacts (median spans 2-3.5 Kb in group A and 1.5-2.1 Kb in group B versus 700-800 bp in group C, P values ​​9.6E-08-4.4E-153, one-sided Wilcoxon rank sum test) (Figure 2D). Sequence analysis of group A H3K27ac peaks in ecDNA compared to other H3K27ac peak regions showed enrichment for binding motifs of transcription factors important in the control of RNAPII general transcription and cell proliferation, including JUN, FOS, and ATF (q-values ​​<0.001, enrichment >1.5). Taken together, the convergence of RNAPII signals in both cis and trans with strong enhancement of H3K27ac signals suggests that ecDNA molecules can connect to the RNA polymerase machinery extensively throughout the genome, supporting their function as genome-wide transcription amplifiers.

[0089] We next determined through analysis of RNA expression from the same four lines whether the increase in enhancer signals associated with ecDNA trans interactions results in active transcription. Taken together, the results showed the detection of 1,887, 1,270, 1,483, and 1,157 chromosomal genes whose promoters were in contact with ecDNA in the HF-2354, HF-2927, HF-3016, and HF-3177 ecDNA(+) cell lines, respectively. EcDNA interacting genes showed significantly higher levels of expression (median FPKM 12-14) compared to either genes with no other trans-chromosomal contacts (median FPKM 0.7-4, P-values ​​4.3E-34-1.8E-55, one-sided Wilcoxon rank sum test) (Figure 9A) or genes with no ecDNA contacts but with trans-chromosomal interactions with other genes (median FPKM 8-9, P-values ​​6.1E-04-1.5E-08, one-sided Wilcoxon rank sum test) (Figure 10). Furthermore, the expression levels of genes connected to ecDNA positively correlate with the frequency of their ecDNA contacts (measured by the number of independent trans interactions) (Figure 9B). Taken together, ecDNA connectivity was highly associated with transcriptional activity, and a highly enhanced H3K27ac signature, suggesting that ecDNA may function as a global transcriptional amplification mechanism.

[0090] In total, 4,763 genes were in contact with ecDNA across the four ecDNA(+) cell lines, of which 877 (18%) were common to two or more cell lines (Figure 3A). The functions of these 877 genes were significantly enriched in translation initiation (FDR 2.06E-04), cell-cell adhesion (FDR 0.003), and transcription (FDR 0.02), biological processes involved in cell communication and amplification (DAVID online analysis [Huang da, W. et al. (2009) Nat Protoc 4, 44-57]). Of the 20 genes found in common across all four ecDNA(+) cell lines, more than half were genes functionally associated with tumorigenesis, control of apoptosis, cell growth or proliferation, including ERBB2, DNAJB4, MCL1, DDIT4 and BAD, JUND and its transcriptional cofactor FOS, and the non-coding RNA gene MALAT1. We tested for the presence of a set of 736 chromosomal tumor genes in the ecDNA connectivity network [Forbes, SA et al. (2015) Nucleic Acids Res 43, D805-811; Repana, D. et al. (2019) Genome Biol 20, 1] and found that 87, 56, 78, and 54 annotated tumor genes were within the ecDNA-mediated chromatin interactome across the four ecDNA(+) cell lines, respectively (Figure 3B). This represents a 2.1- to 2.5-fold enrichment over random predictions (P<0.05, one-sided Wilcoxon rank-sum test), supporting the hypothesis that ecDNA recruits additional oncogenes through chromatin interactions for coactivation in cancer cells. Consistent with an action of ecDNA on transcriptional activation, trans-interacting oncogenes showed a 6- to 10-fold increase in FPKM compared to the median transcription levels across all genes in each of the four ecDNA(+) lines (P-values ​​1.3E-14 to 2.1E-24, one-sided Wilcoxon rank-sum test) (Figure 9C).In summary, 216 of the 736 annotated oncogenes were determined to have exhibited transchromosomal interactions with ecDNA in at least one of the four ecDNA(+) cell lines (Figure 3B). Notably, ERBB2 and MALAT1 exhibited chromatin connections with ecDNA in all four neurosphere lines. ERBB2 is a canonical oncogene in breast cancer and shares structural similarity with EGFR, which is altered in 55%-60% of glioblastomas. ERBB2 and EGFR can form heterodimers to activate their downstream signaling pathways [Qian, X., et al. (1994) Proc Natl Acad Sci USA 91, 1500-1504]. ERBB2 is a rare genomic alteration in GBM and is expressed in nearly half of GBMs [Zhang, C. et al. (2016) J Natl Cancer Inst 108, doi:10.1093 / jnci / djv375; Liu, G. et al. (2004) Cancer Res 64, 4980-4986] but not in non-neoplastic brain cells (GEPIA online data [Tang, Z. et al. (2017) Nucleic Acids Res 45, W98-W102]). Expression of MALAT1 in GBM can result in WNT signaling [Vassallo, I. et al. (2016) Oncogene 35, 12-21], which activates endothelial transdifferentiation and increased migratory capacity [Hu, B. et al. (2016) Cell 167, 1281-1295].

[0091] Besides the recruitment of individual oncogenes, ecDNA also appears to be a focal point where numerous oncogenes are in spatial proximity together via interactions with ecDNA. Among 8, 11, 10, and 4 genome-wide interaction networks defined by extensive communication between hubs, in each of the four ecDNA(+) cell lines (Methods, Figure 3C), all but three in HF-2927 have interaction hubs originating from ecDNA. Many of these ecDNA-connected oncogenes are present within each of the chromatin networks. Specifically, individual communities in HF-3016 and HF-3177 can connect up to 10–12 additional oncogenes (Figure 9D), which is significantly higher than random expectation (P<0.05, one-sided Wilcoxon rank sum test). Such coaggregation of oncogenes was shown to be a structure-based mechanism that ecDNA employs to achieve coordinated transcriptional coactivation to promote tumorigenesis. Notably, even though ecDNA from HF-3016 and HF-3177 originated from primary and recurrent GBM from the same patient, the combinatorial coherence of the ecDNA network did not have many overlapping genes, indicating that the heterogeneity of cancer clones may be further expanded by the ecDNA-chromatin network.

[0092] Taken together, we conducted studies using chromatin interaction assays, ChIA-PET assays, to characterize the ecDNA transcriptional interactome and its regulation in cancer cells. Through multi-omics integrated analysis of genome-wide ecDNA connectivity focal points, trans-interacting chromosomal target genes, H3K27ac binding and RNA expression, we showed that ecDNA may act as a mobile enhancer element that can preferentially target oncogenes for transcriptional coactivation in cancer cells (Figure 9E). These findings outline a novel ecDNA mechanism that provides cancer cells with a competitive advantage to drive tumor progression and tumor evolution. Furthermore, the identification of oncogenes targeted by ecDNA may reveal candidates for targeted inhibitory treatment strategies, and oncogene clusters transcriptionally activated by ecDNA may be a mechanism to prioritize efficacy against a given tumor type.

[0093] In addition to providing a detailed characterization of ecDNA-targeted chromatin interactomes in cancer genomes, the use of chromatin interaction assays, such as but not limited to ChIA-PET assays, provides an effective means to precisely map amplified genomic domains within ecDNA based on their strong chromatin contacts with linear chromosomes. Existing methods used to characterize ecDNA are either through imaging-based analysis [Turner, KM et al. (2017) Nature 543, 122-125] or structural analysis of regions with copy number gain [Deshpande, V. et al. (2018) bioRxiv doi.org / 10.1101 / 457333]. Compared to these methods, directly measuring interchromosomal chromatin contact frequency through chromatin interaction assays, of which ChIA-PET and Hi-C are non-limiting examples, provides an unbiased approach to reveal ecDNA signatures of distinct size, copy number, or sequence content. Furthermore, the frequency and pattern of contacts between different regions of ecDNA molecules provide insight into their physical structure and continuity, as well as the ability of three-dimensional chromatin organization to aid in characterizing genome structural variation and assembly [Spielmann, M. et al. (2018) Nat Rev Genet 19, 453-467, Dixon, JR et al. (2018) Nat Genet 50, 1388-1398].

[0094] Taken together, the experimental results indicate that ecDNA can enhance the expression of extrachromosomal and chromosomal gene transcription through chromatin interactions. Combined with the prevalence and diversity of ecDNA, these findings provide an additional level of complexity of ecDNA action in cancer. Importantly, these results provide insight into the interplay between genetic structure and epigenetic outcomes in tumor evolution. Given the prevalence of ecDNA in cancer and the unique genomic dynamics of this extrachromosomal structure, this supports therapeutic targeting of ecDNA and their activated chromosomal genes.

[0095] Equivalent Although several embodiments of the invention have been described and illustrated herein, those skilled in the art will readily envision various other means and / or structures for performing the functions described herein and / or obtaining one or more of the results and / or advantages thereof, and each of such variations and / or modifications is considered to be within the scope of the present invention. More generally, those skilled in the art will readily appreciate that all parameters, dimensions, materials, and configurations described herein are meant to be exemplary, and that the actual parameters, dimensions, materials, and / or configurations will depend on the specific application for which the teachings of the present invention are used. Those skilled in the art will recognize, or can ascertain using no more than routine experimentation, numerous equivalents to the specific embodiments of the invention described herein. Thus, the foregoing embodiments have been presented for illustrative purposes only, and it will be understood that, within the scope of the appended claims and their equivalents, the invention may be practiced otherwise than as specifically described and claimed. The present invention is directed to each individual feature, system, article, material, and / or method described herein. In addition, any combination of two or more such features, systems, articles, materials, and / or methods is included within the scope of the present invention, unless such features, systems, articles, materials, and / or methods are mutually inconsistent.

[0096] When a range of values ​​is provided, it is understood that each intervening value is encompassed. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the invention.

[0097] All definitions defined and used herein are understood to take precedence over dictionary definitions, definitions in documents incorporated by reference, and / or ordinary meanings of the defined terms.

[0098] The indefinite articles "a" and "an" as used herein in the specification and claims are to be understood to mean "at least one" unless expressly indicated otherwise. The term "and / or" as used herein in the specification and claims is to be understood to mean "either or both" of the elements so connected, i.e., the elements are present conjunctively in some instances and disjointly in other instances. Other elements than the elements specifically identified in the "and / or" clause may be present as needed, whether or not associated with the specifically identified elements, unless expressly indicated otherwise.

[0099] All references, patents, and patent applications, and publications cited or referred to in this application are hereby incorporated by reference in their entirety. Aspects of the invention [Aspect 1] A method for identifying extrachromosomal DNA (ecDNA) in a cell, comprising: (a) detecting a chromatin interaction between a nonlinear DNA molecule and at least one linear chromosome of a chromosome pair, the interaction comprising a contact between the nonlinear DNA molecule and the at least one linear chromosome of the chromosome pair; (i) a significantly higher frequency of detected chromatin interactions in said cells; (ii) contact in the cell between the nonlinear DNA molecule and at least one linear chromosome of each of the chromosome pairs; and (iii) an increase in the average copy number per cell of the nonlinear DNA molecule over time in a plurality of cells; the presence of identifies the nonlinear DNA molecule as ecDNA. A method comprising: [Aspect 2] The method described in aspect 1, further comprising a step of determining the frequency of the detected chromatin interactions. [Aspect 3] The method described in aspect 1 or 2, further comprising a step of determining the size of the nonlinear DNA. [Aspect 4] The method according to any one of aspects 1 to 3, further comprising a step of determining the copy number of the nonlinear DNA molecule in the cell. [Aspect 5] A method according to any of aspects 1 to 4, further comprising the steps of determining an average copy number per cell of the nonlinear DNA molecule in a plurality of cells at a first time point, and comparing the determined average with an average copy number per cell of the nonlinear DNA in a control. [Aspect 6] The method described in aspect 5, wherein the control average copy number per cell of the nonlinear DNA molecule is the average copy number per cell of the nonlinear DNA molecule determined in the plurality of cells at different time points. [Aspect 7] The method according to any one of aspects 1 to 6, further comprising determining the sequence of at least a portion of the nonlinear DNA. [Aspect 8] The method described in aspect 7, further comprising a step of identifying the presence of a tumor gene sequence in the determined sequence. [Aspect 9] A method according to any one of aspects 1 to 8, wherein the cell is a cancer cell. [Aspect 10] The method described in aspect 9, wherein the cells are obtained from a plurality of cells including the cancer cells. [Aspect 11] A method according to any one of aspects 1 to 10, wherein the cells are obtained from a subject. [Aspect 12] The method of aspect 11, wherein the subject is at least one of: diagnosed with cancer, suspected of having cancer, and at risk of having cancer. [Aspect 13] The method according to any one of aspects 1 to 10, wherein the cells are obtained from a cell culture. [Aspect 14] A method according to any one of aspects 1 to 13, wherein the cell is a vertebrate cell. [Aspect 15] The method of any one of aspects 1 to 14, wherein the cell is a mammalian cell, and optionally a human cell. [Embodiment 16] A method for identifying a tumor gene modulated by ecDNA, comprising: (a) detecting an interaction between ecDNA and one or more target genes located in a cell, comprising directly measuring chromatin interactions between the ecDNA and regulatory elements of the one or more target genes; (b) identifying one or more of the target genes involved in the interaction, the transcription of which is modulated by the interaction detected in step (a); (c) determining whether one or more of the target genes identified in step (b) are oncogenes, wherein determining the target genes as oncogenes identifies the oncogenes as oncogenes that are modulated by ecDNA; A method comprising: [Aspect 17] The method of aspect 16, wherein the identification in step (b) comprises measuring the transcription level of the identified target gene and comparing the measured level with the transcription level of the target gene in a control. [Aspect 18] The method according to aspect 16, wherein the modulation of transcription is an increase in transcription. [Aspect 19] A method according to any one of aspects 16 to 18, wherein the detection means in step (a) includes a ChIA-PET method. [Aspect 20] A method according to any one of aspects 16 to 18, wherein the detection means in step (a) includes the Hi-C method. [Aspect 21] The method of any one of aspects 16 to 20, wherein the control element comprises a promoter of the target gene. [Aspect 22] A method according to any one of aspects 16 to 21, wherein the target gene is located on a linear chromosome. [Aspect 23] A method according to any one of aspects 16 to 21, wherein the target gene is located on a second ecDNA. [Aspect 24] A method according to any one of aspects 16 to 23, wherein the cell is a cancer cell. [Aspect 25] The method described in aspect 24, wherein the cells are obtained from a plurality of cells including the cancer cells. [Aspect 26] A method according to any one of aspects 16 to 25, wherein the cells are obtained from a subject. [Aspect 27] The method of aspect 26, wherein the subject is at least one of: diagnosed with cancer, suspected of having cancer, and at risk of having cancer. [Aspect 28] The method according to any one of aspects 16 to 25, wherein the cells are obtained from a cell culture. [Aspect 29] The method of any one of aspects 16 to 28, wherein the cell is a vertebrate cell. [Aspect 30] The method of any one of aspects 16 to 29, wherein the cell is a mammalian cell, and optionally a human cell. [Aspect 31] A method for determining the oncogene status of a cancer, comprising: (a) identifying oncogenes modulated by ecDNA in cancer cells; and (b) determining one or more of the level and effect of said ecDNA modulation of said oncogenes as a determination of the oncogene status of said cancer. A method comprising: [Aspect 32] The identifying means in step (a) (i) detecting an interaction between ecDNA and one or more target genes located in DNA of a cancer cell, said detecting comprising directly measuring chromatin interactions between said ecDNA and regulatory elements of said one or more target genes; (ii) identifying target genes in the detected interactions, the transcription of which is modulated by the interaction detected in step (i); (iii) determining whether one or more of the target genes identified in step (ii) are oncogenes, wherein determining the target genes as oncogenes identifies the oncogenes as oncogenes that are modulated by ecDNA in cancer cells; 32. The method of embodiment 31, comprising: [Aspect 33] The method of aspect 32, wherein the identification in step (ii) comprises measuring the transcription level of the identified target gene and comparing the measured level with the transcription level of the target gene in a control. [Aspect 34] The method of aspect 31, wherein the modulation of transcription is an increase in transcription. [Aspect 35] A method according to any one of aspects 32 to 34, wherein the detection means in step (i) includes a ChIA-PET method. [Aspect 36] A method according to any one of aspects 32 to 34, wherein the detection means in step (i) includes the Hi-C method. [Aspect 37] The method of any one of aspects 32 to 36, wherein the control element comprises a promoter of the target gene. [Aspect 38] A method according to any one of aspects 32 to 37, wherein the target gene is located on a linear chromosome. [Aspect 39] A method according to any one of aspects 32 to 37, wherein the target gene is located on a second ecDNA. [Aspect 40] A method according to any one of aspects 31 to 39, wherein the cell is a cancer cell. [Aspect 41] The method described in aspect 40, wherein the cells are obtained from a plurality of cells including the cancer cells. [Aspect 42] A method according to any one of aspects 31 to 41, wherein the cells are obtained from a subject. [Aspect 43] The method of aspect 42, wherein the subject is at least one of: diagnosed with cancer, suspected of having cancer, and at risk of having cancer. [Aspect 44] A method according to any one of aspects 31 to 41, wherein the cells are obtained from a cell culture. [Aspect 45] A method according to any one of aspects 31 to 44, wherein the cell is a vertebrate cell. [Aspect 46] The method of any one of aspects 31 to 45, wherein the cell is a mammalian cell, optionally a human cell. [Aspect 47] A method according to any of aspects 31 to 46, wherein the means for detecting one or more of the level and effect of the ecDNA modulation of the oncogene comprises directly measuring the interchromosomal chromatin contact frequency between the modulating ecDNA and the modulated oncogene. [Aspect 48] A method according to any of aspects 31 to 47, wherein the means for detecting one or more of the level and effect of ecDNA modulation of an oncogene comprises determining the level of transcription of the oncogene modulated by the one or more ecDNAs, the level of transcription determining the oncogene status of the cancer. [Aspect 49] The method further includes repeating steps (a) and (b) in cancer cells obtained from a second plurality of cells comprising cancer, and comparing one or more levels or effects detected in the cancer cells obtained from the first plurality of cells to the levels or effects detected in the cancer cells obtained from the second plurality of cells, respectively; The method of any of embodiments 41 to 48, wherein a difference in level and / or effect is indicative of a change in said oncogenic status of said cancer. [Aspect 50] After determining the oncogene status of the first plurality of cancer cells in the cancer, and before determining the oncogene status of the second plurality of cancer cells, contacting the second plurality of cancer cells with a candidate therapeutic agent; determining an effect of said contacting with said candidate therapeutic agent on said oncogenic status of said second plurality of cancer cells; 50. The method of embodiment 49, further comprising: [Aspect 51] A method according to any of aspects 41 to 50, wherein the first and second plurality of cells are obtained from a subject. [Aspect 52] The method of aspect 51, wherein the subject is one or more of: diagnosed with cancer, suspected of having cancer, and at risk of having cancer. [Aspect 53] A method according to any of aspects 41 to 50, wherein the first and second plurality of cells are obtained from a cell culture. [Aspect 54] A method according to any one of aspects 31 to 42, wherein the cancer cells are mammalian cells, and optionally human cells. [Aspect 55] A method according to any of aspects 41 to 54, wherein the first and second plurality of cells comprise cancer cells. [Aspect 56] Assisting in the selection of a treatment for the cancer based at least in part on the determined tumor genetic status of the cancer. 56. The method of any of embodiments 31 to 55, further comprising: [Aspect 57] A method according to any of aspects 31 to 56, further comprising the steps of identifying one or more additional oncogenes modulated by one or more additional ecDNAs in a cancer cell, and determining one or more of the level and effect of the ecDNA modulation of the identified one or more additional oncogenes as a determination of the oncogene status of the cancer. [Embodiment 58] A method for identifying extrachromosomal DNA (ecDNA) in a cell, comprising the step of directly detecting a physical interaction between a nonlinear DNA molecule and at least one linear chromosome using a chromatin interaction assay, the chromatin interaction assay comprising: a) contacting cells with a fixative and performing chromatin proximity ligation on DNA isolated from said cells; b) performing chromatin immunoprecipitation; c) generating a library from the DNA immunoprecipitated in step (b); d) sequencing the library generated in step (c) to generate sequencing data; e) analyzing the sequencing data generated in step (d) to detect ecDNA; A method comprising: [Aspect 59] The method described in aspect 58, wherein steps (a) to (c) are carried out using ChIA-PET or HI-C. [Aspect 60] A method described in aspect 58 or 59, wherein step (b) comprises the use of an anti-RNAPII antibody in the chromatin immunoprecipitation. [Aspect 61] A method according to any of aspects 58 to 60, wherein steps (c) and (d) comprise the use of paired-end tags and high-throughput sequencing, respectively. [Aspect 62] A method according to any of aspects 58 to 61, further comprising a step of determining the frequency of the detected physical interactions. [Aspect 63] A method according to any of aspects 58 to 62, further comprising a step of determining the size of the nonlinear DNA molecule. [Aspect 64] A method according to any one of aspects 58 to 63, further comprising a step of determining the copy number of the nonlinear DNA molecule in the cell. [Aspect 65] A method according to any of aspects 58 to 64, further comprising the steps of determining an average copy number per cell of the nonlinear DNA molecule in a plurality of cells at a first time point, and comparing the determined average with an average copy number per cell of the nonlinear DNA molecule in a control. [Aspect 66] The method described in aspect 65, wherein the average copy number per cell of the control nonlinear DNA molecule is the average copy number per cell of the nonlinear DNA molecule determined in the plurality of cells at different time points. [Aspect 67] A method according to any of aspects 58 to 66, further comprising determining the sequence of at least a portion of the nonlinear DNA molecule. [Aspect 68] The method described in aspect 67, further comprising a step of identifying the presence of a tumor gene sequence in the determined sequence. [Aspect 69] A method described in any one of aspects 58 to 68, wherein the cell is a cancer cell. [Aspect 70] The method described in aspect 69, wherein the cells are obtained from a plurality of cells including cancer cells. [Aspect 71] A method according to any one of aspects 58 to 70, wherein the cells are obtained from a subject. [Aspect 72] The method of aspect 71, wherein the subject is at least one of: diagnosed with cancer, suspected of having cancer, and at risk of having cancer. [Aspect 73] A method according to any one of aspects 58 to 72, wherein the cells are obtained from a cell culture. [Aspect 74] A method according to any of aspects 58 to 73, wherein the cell is a vertebrate cell. [Aspect 75] A method according to any one of aspects 58 to 74, wherein the cell is a mammalian cell, optionally a human cell. [Aspect 76] A method according to any of aspects 58 to 75, further comprising a step of performing steps (a) to (e) on a plurality of cells. [Aspect 77] A method for selecting a treatment for reducing cancer in a subject, comprising: identifying the presence of one or more specific ecDNA-modulated oncogenes in cancer cells obtained from the subject; and selecting one or more treatments based on the identified ecDNA-modulated oncogenes. A method comprising: [Aspect 78] The method described in aspect 77, wherein the identifying step includes a method described in any one of aspects 1 to 15 and 58 to 76.

Claims

1. 1. A method for identifying extrachromosomal DNA (ecDNA) in a cell, comprising: (a) detecting a chromatin interaction between a nonlinear DNA molecule and at least one linear chromosome of a chromosome pair, the interaction comprising a contact between the nonlinear DNA molecule and the at least one linear chromosome of the chromosome pair; (i) a significantly higher frequency of detected chromatin interactions in said cells; (ii) contact in the cell between the nonlinear DNA molecule and at least one linear chromosome of each of the chromosome pairs; and (iii) an increase in the average copy number per cell of the nonlinear DNA molecule over time in a plurality of cells; identifies the nonlinear DNA molecule as ecDNA. Including, The method further comprises determining a frequency of the detected chromatin interactions.

2. The method of claim 1 , further comprising determining a size of the nonlinear DNA.

3. 3. The method of claim 1 or 2, further comprising determining the copy number of the nonlinear DNA molecule in the cell.

4. 4. The method of claim 1, further comprising determining an average copy number per cell of the nonlinear DNA molecule in a plurality of cells at a first time point, and comparing the determined average with an average copy number per cell of the nonlinear DNA molecule in a control.

5. 5. The method of claim 4, wherein the control average copy number per cell of the nonlinear DNA molecule is the average copy number per cell of the nonlinear DNA molecule determined in the plurality of cells at different time points.

6. 6. The method of claim 1, further comprising determining the sequence of at least a portion of the nonlinear DNA.

7. The method of claim 6, further comprising identifying the presence of an oncogene sequence in the determined sequence.

8. The method of claim 1 , wherein the cell is a cancer cell.

9. The method of claim 8 , wherein the cells are obtained from a plurality of cells that includes the cancer cells.

10. 10. The method of any one of claims 1 to 9, wherein the cells are obtained from a subject.

11. 11. The method of claim 10, wherein the subject has been diagnosed with, is suspected of having, and is at risk of having cancer.

12. 10. The method of any one of claims 1 to 9, wherein the cells are obtained from a cell culture.

13. The method of any one of claims 1 to 12, wherein the cell is a vertebrate cell.

14. 14. The method of any one of claims 1 to 13, wherein the cell is a mammalian cell, optionally a human cell.

Citation Information

Patent Citations

  • Methods and composition for visualizing and interfering with chromosomal tethering of extrachromosomal molecules

    US20060051739A1