Methods to diagnose, monitor, classify and treat a cancer comprising a transcription factor gene fusion
By isolating and enriching cell-free DNA with histone modification-specific antibodies, the method effectively identifies cancers with transcription factor gene fusions, enhancing diagnostic accuracy and monitoring capabilities without invasive procedures.
Patent Information
- Application Number
- PCT/IB2025/057011
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-04-18
- Filing Date
- 2025-07-10
- Publication Date
- 2026-01-15
AI Technical Summary
Current methods for detecting and monitoring cancers driven by transcription factor gene fusions, such as tRCC, are inadequate due to the diversity of potential breakpoints and fusion partners, leading to false negatives and the need for invasive tissue biopsies, and lack sensitivity and specificity in distinguishing these cancers from other subtypes.
The method involves isolating cell-free DNA from a sample, enriching it with antibodies specific to histone modifications (H3K27ac and/or H3K4me3), and detecting nucleotide sequences associated with gene regulatory elements to reliably identify cancers like tRCC, prostate cancer, and synovial sarcoma using liquid biopsies.
This approach allows for accurate diagnosis, monitoring, and classification of cancers with transcription factor gene fusions, reducing the need for invasive biopsies and improving treatment selection and monitoring efficacy.
Smart Images

Figure IB2025057011_15012026_PF_FP_ABST
Abstract
Description
METHODS TO DIAGNOSE, MONITOR, CLASSIFY AND TREAT A CANCER COMPRISING A TRANSCRIPTION FACTOR GENE FUSIONCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims benefit of U.S. Provisional Application No. 63 / 669,465 filed July 10, 2024, and U.S. Provisional Application No. 63 / 791,036 filed April 18, 2025, which are incorporated herein by reference.SEQUENCE LISTING
[0002] The present specification makes reference to a Sequence Listing submitted electronically as an .xml file name “2025-07-08 DFCI_3540.WOlWO Sequence Listing” on 10 July 2025. The .xml file was generated on 8 July 2025 and is 39 MB in size. The entire contents of the sequence listing are herein incorporated by reference.FIELD
[0003] Provided herein are methods of detecting, diagnosing and / or monitoring a cancer comprising a transcription factor gene fusion (e.g., translocation renal cell carcinoma (tRCC)) in a subject by isolating and enriching cell-free DNA (cfDNA) and detecting the presence of one or more nucleotide sequences comprising gene regulatory elements specifically associated with the cancer comprising a transcription factor gene fusion (e.g., tRCC).BACKGROUND
[0004] Gene fusion events that result in the creation of an oncogenic transcription factor can lead to fusion-driven malignancies. For example, one of the most frequent chromosomal abnormalities associated with acute myeloid leukemia (AML) is a translocation which involves the AML1 gene on chromosome 21 and the ETO gene on chromosome 8. This translocation generates an AML1-ETO fusion transcription factor. Other cancers driven by gene fusions involving transcription factors include translocation renal cell carcinoma (tRCC) (e.g., a TFE3 gene fusion), Ewings Sarcoma (e.g., an EWS-FLI gene fusion), synovial sarcoma (e.g., an SSX2 fusion), prostate cancer (e.g., a TMPRSS2-ERG gene fusion), and childhood rhabdomyosarcoma (RMS) (e.g., &PAX3-FOXO1 gene fusion).
[0005] Cell-free DNA (cfDNA) in the blood can be used to detect cancer and monitor tumor burden. Current cfDNA- based methods rely heavily on the analysis of DNA alterations, such as point mutations or copy number variations (CNVs). However, DNA alterations are not commonly found in cancers caused by oncogenic transcription factor gene fusions, rendering current methods that detect point mutations and CNVs unsuitable for the effective detection and / or monitoring of transcription factor gene fusion-driven cancers. Moreover, transcription factor gene fusions are challenging to directly identify in cfDNA. In the case of several transcription factor gene fusion-driven cancers including tRCC, detecting the gene fusion (such as a TFE3 gene fusion in tRCC) is particularly difficult, owing to the diversity of potential breakpoints and fusion partners.
[0006] Correct identification of cancer types and subtypes is often essential for determining suitable treatment options and maximizing positive treatment outcomes. It is therefore important to distinguish between cancers caused by oncogenic transcription factor gene fusions from other subtypes of cancers affecting the same organ or tissue. This is not generally possible with histology alone.
[0007] For example, there are over 40 different subtypes of renal cell carcinoma (RCC). tRCC is an aggressive subtype of RCC driven by gene fusions involving an MiT / TFE family transcription factor, most commonly TFE3. Approximately 5% of RCC in adults and up to 50% in children is tRCC. Despite harboring a distinctive genetic landscape, tRCC shares histological similarities with other RCC subtypes, complicating accurate diagnosis and an estimation of its true incidence. tRCC is most commonly misdiagnosed as clear cell RCC (ccRCC).
[0008] More than 20 different MiT / TFE fusion partners have been reported to date. In addition, genomic breakpoint locations can vary substantially between patients. Both of these factors can lead to false negatives with current genetic detection assays. Moreover, while TFE3 gene fusions can be detected by break-apart fluorescence in situ hybridization (FISH) or nextgeneration sequencing (NGS), these tests require invasive tissue biopsies and therefore are not a first choice in clinical practice.
[0009] Similarly, there are over 100 subtypes of sarcoma, which are generally grouped into two main types: soft tissue sarcoma and bone sarcoma. Synovial sarcoma is a soft tissue sarcoma that is often characterized by a translocation between chromosomes X and 18, leading to the formation of the SS18:SSX transcription factor fusion genes. Currently, diagnosis is typically confirmed only following a solid tumor biopsy which is an invasive technique, followed by molecular or cytogenetic testing for SS18-SSX fusion. Commonly used methods for diagnosis can be time consuming and include FISH, RT-PCR and NGS (Cancer Res 2002,62: 135, Genes Chromosomes Cancer 2001;30:l).
[0010] Prostate cancer is a very common cancer in men. Prostate cancer is usually detected through blood tests that check for prostate-specific antigen (PSA) levels. However, this test has a low degree of sensitivity and specificity. For instance, PSA levels may be elevated as a result of other prostatic diseases. A biopsy is typically needed to accurately diagnose prostate cancer and determine its grade. Certain subtypes of prostate cancer are characterized by a TMPRSS2- ERG fusion gene. The TMPRSS2-ERG fusion gene has potential as a predictive marker for this subtype of prostate cancer and may indicate alternative forms of treatment.
[0011] There remains a need for diagnostic tests that can be used to identify cancer subtypes which are driven a transcription factor gene fusion.SUMMARY
[0012] Identified herein are nucleotide sequences comprising gene regulatory elements specifically associated with cancers comprising a transcription factor gene fusion. These nucleotide sequences can be used to differentiate cancers comprising a transcription factor gene fusion from other common subtypes of cancers affecting the same organ or tissue. The nucleotide sequences have been discovered using RNA sequencing (RNA-seq), chromatin immunoprecipitation (ChIP) sequencing (ChlP-seq), and methylated DNA immunoprecipitation sequencing (MeDIP-seq).
[0013] In particular, it has been discovered that isolated cell-free DNA (cfDNA) can be enriched for histone H3 protein acetylated at the lysine at residue 27 (H3K27ac)-associated DNA and / or histone H3 protein methylated at the lysine at residue 4 (H3K4me3)-associated DNA toreliably detect the presence of or level of one or more nucleotide sequences comprising gene regulatory elements specifically associated with a cancer comprising a transcription factor gene fusion, including tRCC, prostate cancer and synovial sarcoma, in a liquid biopsy sample such as serum, plasma or blood. Such enriched cfDNA can be used in a method to detect, diagnose, and monitor these cancers in a subject.
[0014] Accordingly, in one aspect, a method for diagnosing a cancer comprising a transcription factor gene fusion (e.g., tRCC) in a subject is provided that comprises: (a) isolating cell-free (cfDNA) from a sample obtained from the subject; (b) enriching the cfDNA with a means for specifically binding histone H3 protein acetylated at the lysine at residue 27 (H3K27ac) for H3K27ac-associated DNA and / or with a means for specifically binding histone H3 protein methylated at the lysine at residue 4 (H3K4me3) for H3K4me3 -associated DNA; and (c) detecting the presence of or level of one or more nucleotide sequences comprising gene regulatory elements specifically associated with the cancer comprising the transcription factor gene fusion in the cfDNA enriched for H3K27ac-associated DNA and / or H3K4me3 -associated DNA, wherein the presence or level of the one or more nucleotide sequences comprising gene regulatory elements specifically associated with the cancer comprising the transcription factor gene fusion in the enriched cfDNA is indicative of the subject suffering from the cancer comprising the transcription factor gene fusion.
[0015] In a further aspect, a method for monitoring tumor progression in a subject suffering from a cancer comprising a transcription factor gene fusion (e.g., tRCC) is provided that comprises: (a) measuring, from a first sample obtained from the subject at a first time point, the presence and / or level of one or more nucleotide sequences comprising gene regulatory elements specifically associated with the cancer comprising the transcription factor gene fusion; (b) measuring, from a second sample obtained from the subject at a second or subsequent time point, the presence and / or level of the one or more nucleotide sequences comprising gene regulatory elements specifically associated with the cancer comprising the transcription factor gene fusion; and (c) comparing the measurements from steps (a) and (b), thereby monitoring the tumor progression, wherein the measuring steps are carried out according to a method disclosed herein.
[0016] In a further aspect, a method for classifying a cancer subtype (e.g. RCC subtype such as tRCC) is provided that comprises: (a) detecting, from a sample derived from a subject suffering from a cancer (e.g., RCC), the presence and / or level of the one or more nucleotide sequences comprising gene regulatory elements specifically associated with a cancer comprising a transcription factor gene fusion (e.g., tRCC); and (b) identifying the type or subtype of the cancer as the cancer comprising the transcription factor gene fusion based on the presence of the one or more nucleotide sequences comprising gene regulatory elements specifically associated with the cancer comprising the transcription factor gene fusion, wherein the detecting step is carried out according to a method disclosed herein.
[0017] In a further aspect, a method of selecting a therapy for a subject suffering from a cancer comprising a transcription factor gene fusion (e.g., tRCC) is provided that comprises: (a) isolating cell-free DNA (cfDNA) from a sample obtained from the subject; (b) enriching the cfDNA with a means for specifically binding histone H3 protein acetylated at the lysine at residue 27 (H3K27ac) for H3K27ac-associated DNA and / or with a means for specifically binding histone H3 protein methylated at the lysine at residue 4 (H3K4me3) for H3K4me3- associated DNA; and (c) detecting the presence of one or more nucleotide sequences comprising gene regulatory elements specifically associated with the cancer comprising the transcription factor gene fusion in the cfDNA enriched for H3K27ac-associated DNA and / or H3K4me3- associated DNA, wherein the presence of the one or more nucleotide sequences comprising gene regulatory elements specifically associated the cancer comprising the transcription factor gene fusion in the enriched cfDNA indicates that the subject’s cancer is the cancer comprising the transcription factor gene fusion; and (d) selecting a therapy that is effective against the cancer comprising the transcription factor gene fusion.
[0018] In a further aspect, a method of treating a cancer comprising a transcription factor gene fusion (e.g., tRCC) in a subject is provided that comprises: (a) isolating cell-free DNA (cfDNA) from a sample obtained from the subject; (b) enriching the cfDNA with a means for specifically binding histone H3 protein acetylated at the lysine at residue 27 (H3K27ac) for H3K27ac-associated DNA and / or with a means for specifically binding histone H3 protein methylated at the lysine at residue 4 (H3K4me3) for H3K4me3 -associated DNA; and (c)detecting the presence of one or more nucleotide sequences comprising gene regulatory elements specifically associated with the cancer comprising the transcription factor gene fusion in the cfDNA enriched for H3K27ac-associated DNA and / or H3K4me3 -associated DNA, wherein the presence of the one or more nucleotide sequences comprising gene regulatory elements specifically associated with the cancer comprising the transcription factor gene fusion in the enriched cfDNA is indicative of the subject suffering from the cancer comprising the transcription factor gene fusion, and (d) administering to the subject a therapeutically effective amount of a therapeutic agent effective in the treatment of the cancer comprising the transcription factor gene fusion.
[0019] In a further aspect, a method of monitoring the effectiveness of a therapy in a subject with a cancer comprising a transcription factor gene fusion (e.g., tRCC) is provided that comprises: (a) providing a first sample obtained from the subject at a first time point; (b) isolating cell-free DNA (cfDNA) from the first sample; (c) enriching the cfDNA with a means for specifically binding histone H3 protein acetylated at the lysine at residue 27 (H3K27ac) for H3K27ac-associated DNA and / or with a means for specifically binding histone H3 protein methylated at the lysine at residue 4 (H3K4me3) for H3K4me3 -associated DNA; and (d) determining levels of one or more nucleotide sequences comprising gene regulatory elements specifically associated with the cancer comprising the transcription factor gene fusion in the cfDNA enriched for H3K27ac-associated DNA and / or H3K4me3 -associated DNA; (e) providing a second or subsequent sample obtained from the subject and repeating steps (b)-(d) with the second sample; and (f) comparing the levels of the one or more nucleotide sequences obtained in steps (d) and (e), wherein an increase in the level of the one or more nucleotide sequences determined in step (e) relative to step (d) indicates that the therapy is not effective against the subject’s cancer, and no change or a decrease in the level of the one or more nucleotide sequences determined in step (e) relative to step (d) indicates that the therapy is effective against the subject’s cancer.
[0020] In a further aspect, there is provided herein a computer program comprising computer program code configured to cause one or more physical computing devices to perform a method of determining whether a subject’s cancer comprises a transcription factor gene fusion, themethod comprising: (a) receiving sequencing read data obtained from cfDNA isolated from a sample obtained from the subject and enriched for H3K27ac-associated DNA and optionally sequencing read data obtained from cfDNA isolated from the sample and enriched for H3K4me3 -associated DNA; (b) aligning the sequencing read data obtained in (a) with whole human genome read data; (c) receiving a first predetermined list of transcription factor fusion binding sites, a second predetermined list of gene regulatory elements associated with H3K27ac, and optionally a third predetermined list of gene regulatory elements associated with H3K4me3; (d) for each site or element present in one of the lists received in (c) determining the number of aligned sequencing reads obtained in (b); (e) aggregating the number of sequencing reads for each element in each list; (f) receiving a reference value; (g) for each aggregate, normalizing the aggregate based on the reference value; (h) calculating a combined score based on the two or more aggregates; and (i) comparing the combined score to a predetermined score to determine whether the cancer comprises the transcription factor gene fusion.
[0021] In a further aspect, there is provided herein a computer program comprising a computer program code configured to cause one or more physical computing devices to perform a method of obtaining a list of transcription factor fusion binding sites, the method comprising: obtaining a first reference dataset and a second reference data set, wherein the first reference dataset comprises a plurality of reference sequencing reads of DNA fragments associated with a fusion transcription factor encoded by a transcription factor gene fusion and the second reference dataset comprises a plurality of reference sequencing reads of DNA fragments associated with a corresponding wildtype transcription factor; carrying out a peak analysis on the first reference dataset and the second reference dataset to identify a plurality of sites that overlap between the first reference dataset and the second reference data set; determining significant overlap sites; and selecting, from the significant overlap sites, the list of transcription factor fusion binding sites by selecting sites having a log2fold change in either direction greater than 1.
[0022] In a further aspect, there is provided herein a computer program comprising a computer program code configured to cause one or more physical computing devices to perform a method of obtaining a list of gene regulatory elements associated with H3K27ac, the method comprising: obtaining a first reference dataset and a second reference data set, wherein the first reference dataset comprises reference sequencing reads of DNA fragments associated withH3K27ac of a first cancer cell population comprising a transcription gene fusion and the second reference dataset comprising a plurality of reference sequencing reads of DNA fragments associated with H3K27ac in a second cancer cell population of the same tissue type as the cancer comprising the transcription factor gene fusion, wherein the second cancer cell population does not comprise the transcription factor gene fusion; carrying out a differential peak analysis on the first reference dataset and the second reference dataset to identify a plurality of differentially marked sites present only in the first cancer cell population comprising the transcription gene fusion; determining significant differentially marked sites of the differentially marked sites; and selecting, from the significant differentially marked sites, the list of gene regulatory elements associated with H3K27ac by selecting sites having a log2fold change in either direction greater than 1.
[0023] In a further aspect, there is provided herein a computer program comprising a computer program code configured to cause one or more physical computing devices to perform a method of obtaining a list of gene regulatory elements associated with H3K4me3, the method comprising: obtaining a first reference dataset and a second reference data set, wherein the first reference dataset comprises a plurality of reference sequencing reads of DNA fragments associated with H3K4me3 of a first cancer cell population comprising a transcription gene fusion and the second reference dataset comprising a plurality of reference sequencing reads of DNA fragments associated with H3K4me3 in a second cancer cell population of the same tissue type as the cancer comprising the transcription factor gene fusion, wherein the second cancer cell population does not comprise the transcription factor gene fusion; carrying out a differential peak analysis on the first reference dataset and the second reference dataset to identify a plurality of differentially marked sites present only in the first cancer cell population comprising the transcription gene fusion; determining significant differentially marked sites of the differentially marked sites; and selecting, from the significant differentially marked sites, the list of gene regulatory elements associated with H3K4me3 by selecting sites having a log2fold change in either direction greater than 2.
[0024] In a further aspect, there is provided herein a computer program comprising a computer program code configured to cause one or more physical computing devices to perform a method of analysis of sample results for a subject when the code is run on the one or morephysical computing devices, the method of analysis comprising: (a) receiving data indicative of the presence or level of one or more nucleotide sequences comprising gene regulatory elements specifically associated with a cancer comprising a transcription factor gene fusion (e.g. tRCC) in a sample collected from the subject; (b) obtaining a statistical model; (c) comparing the data and to the statistical model to generate a score; and (d) outputting the score from the statistical model.
[0025] The described methods can be advantageously applied to the detection of subtypes of renal cell carcinoma, prostate cancer, and synovial sarcoma, which comprise a transcription factor gene fusion, e.g., TFE3, TMPRSS2-ERG and SSX2 fusions, respectively. Therefore, in certain aspects, the cancer comprising the transcription factor gene fusion is tRCC, and the gene regulatory elements associated with tRCC are gene regulatory elements associated with a TFE3 gene fusion. In other aspects, the cancer comprising the transcription factor gene fusion is a prostate cancer, and the gene regulatory elements associated with prostate cancer are gene regulatory elements associated with a TMPRSS2-ERG gene fusion. In further aspect, the cancer comprising the transcription factor gene fusion is a synovial sarcoma, and the gene regulatory elements associated with the synovial sarcoma are gene regulatory elements associated with an SSX2 gene fusion.BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Aspects will be described, by way of example, with reference to the following drawings.
[0027] FIG. 1 shows principal component analysis (PCA) plots of the H3K27ac peaks in tRCC and ccRCC.
[0028] FIG. 2 shows volcano plots of differentially marked peaks for (A) H3K27ac and (B) MeDIP between tRCC and ccRCC.
[0029] FIG. 3 shows heatmaps of normalized (A) H3K27ac and (B) MeDIP tag densities at differential H3K27ac and MeDIP regions, respectively, between tRCC and ccRCC cell lines located ±2 kb from peak center.
[0030] FIG. 4 shows aggregated H3K27ac cfChlP-seq signals at (A) cell line informed tRCC-specific H3K27ac sites and (B) TFE3 binding sites and comparisons between tRCC, ccRCC and healthy plasma samples.
[0031] FIG. 5 illustrates a comparison of individual (H3K27ac or TFE3 binding sites alone) and combined classifier performances in (A) discriminating tRCC and healthy plasma and (B) discriminating tRCC and ccRCC.
[0032] FIG. 6 shows a plot monitoring three tRCC patients with multiple time points via plasma epigenomic profiling.
[0033] FIG. 7 shows volcano plots of differentially marked peaks between tRCC and ccRCC for (A) H3K4me3, (B) H3K27ac, and (C) MeDIP.
[0034] FIG. 8 shows heatmaps of normalized (A) H3K4me3 and (B) H3K27ac tag densities at differential H3K4me3 and H3K27ac regions, respectively, between tRCC and ccRCC cell lines.
[0035] FIG. 9 illustrates schematically the identification of TFE3 fusion-occupied TFBS through the intersection of (A) known TFE3 TFBS from the GTRD database and TFE3 fusion peaks obtained by TFE3 ChlP-seq of tRCC cell lines and (B) the aggregated H3K27ac- associated signals at TFE3 fusion-occupied TFBS across various RCC cell lines.
[0036] FIG. 10 shows aggregated cfChIP signal comparisons between tRCC, ccRCC, and healthy plasma (HP) samples for (A) H3K4me3 signals at cell line-informed H3K4me3- associated tRCC-up sites and (B) H3K27ac signals at cell line-informed H3K27ac-associated tRCC-up sites and TFE3 fusion-occupied TFBS.
[0037] FIG. 11 illustrates a comparison of individual cfChIP H3K4me3 and H3K27ac signals at cell line-informed sites and combined classifier performances in (A) discriminating tRCC and ccRCC, and (B) discriminating tRCC and healthy plasma.
[0038] FIG. 12 shows longitudinal tracking of the tRCC integrated epigenomic score (light grey filled circle) and ctDNA fraction (dark grey filled circle) in three tRCC patients (Patient I,FIG. 12A; Patient II, FIG. 12B; and Patient III, FIG. 12C). SD: stable disease; PR: partial response; PD: progressive disease; ; NED: no evidence of disease; Nx: nephrectomy.
[0039] FIG. 13 shows the percentage change in (A) tRCC integrated epigenomic score and (B) CNA-based ctDNA fraction between consecutive plasma draws that are grouped by radiographic response during that same interval.
[0040] FIG. 14 shows aggregated H3K27ac cfChlP-seq signals in prostate cancer samples with and without the TMPRSS2-ERG gene fusion.
[0041] FIG. 15 illustrates a comparison of classifier performances in discriminating prostate cancer with and without the TMPRSS2-ERG gene fusion.
[0042] FIG. 16 illustrates the fraction of genome altered (FGA) in cancers comprising a transcription factor gene fusion and cancers that do not comprise a transcription gene fusion.
[0043] FIG. 17 illustrates the fraction of genome altered (FGA) in RCCs lacking a TFE3 gene fusion compared to RCCs comprising a TFE3 gene fusion.
[0044] FIG. 18 illustrates the fraction of genome altered (FGA) in synovial sarcomas lacking an SSX2 gene fusion compared to synovial sarcomas comprising an SSX2 gene fusion.DETAILED DESCRIPTION
[0045] In order for the following description to be more readily understood, certain terms are first defined below. Additional definitions may be set forth throughout the specification.
[0046] Unless otherwise required by context, singular terms shall include pluralities and plural terms shall include the singular. Thus, as used in this specification and the appended claims, the singular forms “a”, “an” and “the” include plural referents unless the context clearly dictates otherwise. For example, “a nucleotide sequence” is understood to represent one or more nucleotide sequences. As such, the terms “a” (or “an”), “one or more,” and “at least one” can be used interchangeably herein.
[0047] Throughout this specification and aspects, the words “have” and “comprise”, or variations such as “has”, “having”, “comprises”, or “comprising”, will be understood to implythe inclusion of a stated integer or group of integers but not the exclusion of any other integer or group of integers.
[0048] Unless specifically stated or obvious from context, as used herein, the term “or” is understood to be inclusive and covers both “or” and “and”. Furthermore, “and / or” where used herein is to be taken as specific disclosure of each of the two specified features or components with or without the other. Thus, the term “and / or” as used in a phrase such as “A and / or B” herein is intended to include “A and B”, “A or B”, “A” (alone), and “B” (alone). Likewise, the term “and / or” as used in a phrase such as “A, B, and / or C” is intended to include “A and / or B and / or C” and to thus encompass each of the following aspects: A, B, and C; A, B, or C; A or C; A or B; B or C; A and C; A and B; B and C; A (alone); B (alone); and C (alone).
[0049] It is understood that wherever aspects are described herein with the language “comprising”, otherwise analogous aspects described in terms of “consisting of’ and / or “consisting essentially of’ are also provided. In other words, if an aspect is described as comprising A, B and C, aspects consisting essentially of A, B and C are also contemplated, as are aspects consisting of A, B and C.
[0050] The term “about” refers to an interval of accuracy that a person skilled in the art will understand to still ensure the technical effect of the feature in question. The term indicates a deviation from the indicated numerical value of ±10%, or ±5% of the indicated numerical value, or ±1% of the indicated numerical value.
[0051] The terms “cancer”, “malignancy”, “neoplasm”, “tumor”, and “carcinoma” are used interchangeably to refer to a disease, disorder, or condition in which cells exhibit or exhibited relatively abnormal, uncontrolled, and / or autonomous growth, so that they display or displayed an abnormally elevated proliferation rate and / or aberrant growth phenotype.
[0052] The term “cancer comprising a transcription factor gene fusion” and grammatical equivalents thereof may refer to any cancer that is associated with a gene fusion event (e.g., as the result of a chromosomal translocation) involving a transcription factor. The term includes the specific cancers listed in Table 1 herein. Throughout the specification, the term may refer to any and all such cancers. Alternatively, in specific aspects, the term may refer to a subtype ofrenal cell carcinoma, prostate cancer, or synovial sarcoma that are associated with a transcription factor gene fusion, e.g., a TFE3 gene fusion in the case of renal cell carcinoma (referred to as translocation tRCC), a TMPRSS2-ERG gene fusion in the case of prostate cancer, or an SSX2 gene fusion in the case of synovial sarcoma.
[0053] The phrase “nucleotide sequences comprising gene regulatory elements specifically associated with a cancer comprising a transcription factor gene fusion” and grammatical equivalents thereof refer to genomic loci that are overrepresented, and as such are detectable at a higher frequency, in cfDNA isolated by ChIP and / or MeDIP from cancer cells comprising the transcription factor gene fusion as compared to cancer cells that do not. More specifically, this includes nucleotide sequences that are associated with histone modifications including methylation and acetylation and / or DNA methylation and may drive gene expression, including promoters, enhancers, and CpG islands. In particular, such nucleotide sequences comprise nucleotide sequences associated with a transcription factor gene fusion (e.g. , a TFE3 gene fusion, a TMPRSS2-ERG gene fusion, or an SSX2 gene fusion) including gene regulatory elements bound by the neomorphic transcription factor encoded by the transcription factor gene fusion and optionally gene regulatory elements of any downstream genes that are controlled by the neomorphic transcription factor.
[0054] The terms “diagnosis” or “detection” of / for a cancer includes determining whether a subject has or likely has cancer. Diagnosis can include a determination relating to the risk, type, stage, malignancy, or other classification of the cancer. In some instances, a diagnosis can be or include a determination relating to prognosis and / or likely response to one or more general or particular therapeutic agents or regimens.
[0055] Unless otherwise defined herein, technical and scientific terms used herein have the same meaning as commonly used and / or understood by one of ordinary skill in the art to which this application belongs. In case of conflict, the present specification, including definitions, will control.
[0056] Generally, nomenclature used in connection with, and techniques of, cell and tissue culture, molecular biology, virology, immunology, microbiology, genetics, analytical chemistry,synthetic organic chemistry, medicinal and pharmaceutical chemistry, and protein and nucleic acid chemistry and hybridization described herein are those well-known and commonly used in the art. Enzymatic reactions and purification techniques are performed according to manufacturer’s specifications, as commonly accomplished in the art or as described herein.Transcription factor gene fusions in cancer
[0057] Transcription factors are considered to account for a large proportion of human fusion genes involved in cancer. Gene fusion events were originally identified in hematological malignancies. Advances in sequencing platforms, e.g., next-generation sequencing (NGS) have revealed a variety of cancers that comprise a transcription factor gene fusion. For example, transcription factor fusions have also been identified in a wide array of solid tumors, including sarcomas and carcinomas. Transcription factor gene fusions are considered attractive targets for therapy as they are expressed exclusively in tumor tissue.
[0058] Several examples of both solid and hematological cancers have been found to comprise a transcription factor gene fusion. Accordingly, a cancer comprising a transcription factor gene fusion may be a hematological cancer. Specifically, the hematological cancer may be a leukemia or a lymphoma. Alternatively, a cancer comprising a transcription factor gene fusion may be a solid cancer. Specifically, the solid cancer may be a sarcoma or a carcinoma. For example, the solid cancer may be selected from breast cancer, colorectal cancer, brain cancer, head and neck cancer, prostate cancer, kidney cancer, liver cancer, lung cancer, bowel cancer, bile duct cancer, skin cancer (e.g., melanoma), bladder cancer, cervical cancer, gastrointestinal cancer, thyroid cancer, and bone cancer.
[0059] Examples of cancers comprising a transcription factor gene fusion are provided in Table 1.Table 1. Exemplary transcription factor gene fusions in cancer
[0060] For example, the cancer may be a renal cell carcinoma (RCC), such as translocation renal cell carcinoma (tRCC). The tRCC may comprise a TFE3 gene fusion. Alternatively, the cancer may be a prostate cancer. The prostate cancer may comprise a TMPRSS2-ERG gene fusion. Or the cancer may be a synovial sarcoma. The synovial sarcoma may comprise an SSX2 gene fusion.Sample
[0061] A sample obtained from a subject can be or may include cells, tissue, or bodily fluids. Typically, the sample is a liquid such as blood, plasma, serum, or urine. Liquid biopsies canoften be obtained by non-invasive or minimally invasive means and therefore are easier to collect than solid tumor biopsies.Isolating cell-free DNA
[0062] When cancer cells die and undergo apoptosis, DNA and / or chromatin is released into the blood and other bodily fluids such as urine as cell-free DNA (cfDNA). Using cfDNA for cancer detection, diagnosis, and / or monitoring is particularly advantageous for solid tumors because it avoids the need for an invasive tissue biopsy. cfDNA includes all extracellular DNA that may be freely circulating and / or present in exosomes.
[0063] cfDNA associated with the presence of tumors may be present as circulating mononucleosomes (Sanchez et al., (2021) JCI Insight. 2021 Apr 8; 6(7): el44561 doi: 10.1172 / jci. insight.144561). Typically, the isolated chromatin associated DNA is mononucleosomic.
[0064] cfDNA can be isolated from a sample such as blood, plasma, serum, urine or other bodily fluids by any means known in the art. For example, cells and cellular debris may be removed by centrifugation.Enriching for chromatin-associated DNA
[0065] The described methods may involve the step of enriching the cfDNA for chromatin- associated DNA. Chromatin immunoprecipitation (ChIP) can be performed as part of this enrichment step. ChIP usually involves cross-linking of chromatin-bound proteins (e.g., histones), e.g., using a crosslinking agent such as formaldehyde, optionally followed by a fragmentation step (e.g., sonication or nuclease treatment) to obtain DNA fragments. Immunoprecipitation can be then carried out using a means for specifically binding chromatin, e.g, one or more antibodies that specifically bind to a chromatin-bound protein of interest. The DNA can then be released from the proteins and analyzed using various methods. X-ChIP methods utilize fixed chromatin fragmented by sonication, while the N-ChIP methods utilize native chromatin, which can be unfixed and nuclease digested.
[0066] Formaldehyde is a commonly used cross-linking agent. One advantage of using formaldehyde can be the ease of reversibility of the cross-links and its ability to form bonds that span a distance of approximately 2 angstroms. This means that formaldehyde can bind molecules in close association with each other. Other cross-linking agents that have been used include methylene blue, acridine orange, cisplatin, dimethylarsinic acid, potassium chromate, and ultraviolet (UV) light and lasers. The disclosed methods can use any known cross-linking agent known in the art.
[0067] The chromatin-associated DNA can be fragmented or unfragmented. If fragmented, the fragmentation will desirably result in fragments of 100-400 base pairs in size, e.g., 100-300 base pairs in size. Fragmentation may be carried out by sonication. An alternative fragmentation method to sonication can be nuclease digestion of the chromatin, e.g., in N-ChIP methods. Purification of chromatin can be achieved by cesium chloride (CsCl) gradient centrifugation. The present methods can use any method known in the art for fragmenting DNA. Fragmentation (e.g., sonication or nuclease digestion) of the chromatic-associated DNA may not be required.
[0068] Chromatin architecture, nucleosomal positioning, and access to DNA for gene transcription is largely controlled by histone proteins. A nucleosome comprises two identical subunits, each of which contains four histone proteins (H2A, H2B, H3 and H4). The disclosed methods comprise the immunoprecipitation of chromatin-associated DNA that is associated with modified histone proteins. Histone proteins may be modified in different ways which influence DNA interactions. Certain modifications correspond to the disruption of histone-DNA interaction, causing nucleosomes to unwind. In this “open” conformation, known as euchromatin, DNA is accessible for the binding of transcriptional machinery and subsequent gene activation. In contrast, other modifications can strengthen histone-DNA interactions which create a tightly packed chromatin structure called heterochromatin. In this compact form, transcriptional machinery cannot access DNA, resulting in gene silencing. Consequently, histone modification influences chromatin architecture and gene activation / inactivation. Histone proteins may be modified by methylation, acetylation and phosphorylation among other modifications.
[0069] Accordingly, a means for specifically binding chromatin for use in the disclosed methods can be by binding a modified histone, such an acetylated histone or a methylated histone e.g., histone H3. The means for specifically binding chromatin can be an antibody. For example, the methods disclosed herein may comprise contacting a sample with an antibody that specifically binds to an acetylated histone (e.g., acetylated histone H3). The methods disclosed herein may comprise contacting a sample with an antibody that specifically binds to histone H3 protein acetylated at the lysine at residue 27 (H3K27ac). Alternatively, or in addition, the methods disclosed herein may comprise contacting a sample with an antibody that specifically binds to a methylated histone (e.g., methylated histone H3). For instance, the methods disclosed herein may comprise contacting a sample with an antibody that specifically binds to a histone methylated at the lysine at residue 4 (H3K4me3). For example, the methods disclosed herein may comprise contacting a sample with an antibody that specifically binds to an acetylated histone (e.g., acetylated histone H3 such as H3K27ac) and an antibody that specifically binds to a methylated histone (e.g., methylated histone H3 such as H3K4me3). Such antibodies may be used sequentially and / or simultaneously to isolate chromatin-associated DNA.
[0070] Commercially available antibodies for use with the methods disclosed herein include, but are not limited to, the anti-H3K27ac antibody C15410196 and the anti-H3K4me3 antibody ThermoFisher #PA5-27029.
[0071] The chromatin-associated DNA can be isolated by immunoprecipitation. Immunosorbents commonly used to separate the antigen-antibody complex from a sample include salmon sperm DNA-protein A-Sepharose®, protein G, magnetic beads, and other engineered immunoprecipitation systems known to those of skill in the art.
[0072] Once isolated, the chromatin associated DNA may be further purified and / or amplified. The amplification may be by PCR or an isothermal amplification. Purification can be by any means known in the art, including organic extraction, silica-column-based techniques, ethanol precipitation, and anion-exchange chromatography.Enriching for methylated DNA
[0073] The methods described herein may involve the step of enriching the cfDNA for methylated DNA. Single-stranded, fragmented genomic DNA containing 5-methyl cytosine can be specifically captured from the rest of the genomic DNA using methylated DNA immunoprecipitation (MeDIP). DNA methylation occurs at CpG sites, which may cluster in regions called CpG islands. DNA methylation sites may overlap with or be close to promoter regions and transcription start sites. DNA methylation can repress gene expression by either interfering with the binding of transcription factors or modifying chromatin structure to a repressive state. Methylated DNA regions may therefore be useful in the methods of the present invention.
[0074] MeDIP can be performed as part of the enrichment step of the methods of the invention. MeDIP involves isolating methylated DNA fragments by immunoprecipitation using means for specifically binding 5-methylcytosine (5mC) e.g., one or more antibodies that specifically bind to a methylated DNA. The DNA can be then released from its associated proteins and analyzed using various methods.
[0075] Typically, the DNA fragments are denatured to produce single-stranded DNA, before an immunoprecipitation step using antibodies that specifically bind to 5mC. Suitable immunoprecipitation techniques are described in Pomraning KR, el al. Methods. 47 (3): 142- 50. After immunoprecipitation, proteinase K is added to digest the antibodies and release the DNA, which can be collected and prepared for DNA detection.
[0076] Other suitable methods for MeDIP are known in the art, such as MagMeDIP Kit (Diagenode), and Abeam Methylated DNA Immunoprecipitation (MeDIP) Kit - DNA (abl 17133). Suitable anti-5mC antibodies for use in MeDIP include Invitrogen MA5-24694, Invitrogen MA5-31475, and Abeam RM231.
[0077] Once isolated, the methylated DNA may be further purified and / or amplified. The amplification may be by polymerase chain reaction (PCR) or isothermal amplification. Purification can be by any means known in the art, including organic extraction, silica-column- based techniques, ethanol precipitation, and anion-exchange chromatography.Nucleotide sequences to be detected
[0078] By isolating cfDNA from a sample (e.g., blood, plasma, or serum) obtained from a subject and enriching the isolated cfDNA for H3K27ac-associated DNA and / or H3K4me3- associated DNA, nucleotide sequences comprising gene regulatory elements specifically associated with a cancer comprising a transcription factor gene fusion (e.g., tRCC, or certain subtypes of prostate cancer or synovial sarcoma) can be reliably detected. For example, the methods described herein identify genomic loci that are overrepresented, and as such detectable at a higher frequency, in cfDNA isolated by ChIP and / or MeDIP from cancer cells comprising the transcription factor gene fusion (e.g., tRCC cells) as compared to cells that do not, (e.g., ccRCC cells). Nucleotide sequences comprising gene regulatory elements specifically associated with a cancer comprising a transcription factor gene fusion (e.g., tRCC) can be identified as described herein, for instance as set out in the “Peak Calling” methods described in Examples 1 and 5.
[0079] Gene regulatory elements may include, for example, promoter regions, enhancer regions, silencer regions, transcription factor binding sites, polymerase binding sites, or any other nucleotide sequence that interacts with proteins involved in transcription. For example, H3K27ac-associated DNA commonly includes active promoters and enhancers, while H3K4me3 -associated DNA commonly includes active promoters. Nucleotide sequences associated with transcription factor binding sites, gene fusions and / or gene regulatory elements within differentially expressed genes associated with a cancer comprising a transcription factor gene fusion (e.g., tRCC) can be obtained from the GTRD database (Yevshin et al Nucleic Acids Research, Volume 45, Issue DI, January 2017, Pages D61-D67).
[0080] Typically, in the methods disclosed herein, the nucleotide sequences comprising gene regulatory elements specifically associated with a cancer comprising a transcription factor gene fusion (e.g., tRCC) are nucleosome-sized DNA fragments. For example, the one or more nucleotide sequences comprising gene regulatory elements specifically associated with the cancer comprising the transcription factor gene fusion (e.g., tRCC) may each be 50-300bp, or 100-250bp, or 150-200bp. In some instances, the one or more nucleotide sequences comprising gene regulatory elements specifically associated with the cancer comprising the transcriptionfactor gene fusion (e.g., tRCC) may each have a length of 100-250bp. In one instance, the nucleic acid sequences comprising gene regulatory elements specifically associated with the cancer comprising the transcription factor gene fusion (e.g., tRCC) are about 150bp in length.
[0081] To account for the variability between individual patient samples, the various methods disclosed herein (e.g., methods of diagnosing or monitoring a cancer comprising a transcription factor gene fusion, such as tRCC) typically detect the presence, or level, of more than one nucleotide sequences comprising gene regulatory elements specifically associated with a cancer comprising a transcription factor gene fusion (e.g., tRCC). For example, described herein are methods that use at least 1,000 (e.g., at least 5,000 or at least 10,000) nucleotide sequences comprising gene regulatory elements specifically associated with the cancer comprising the transcription factor gene fusion (e.g., tRCC). In some instances, the methods disclosed herein may be used for the simultaneous detection of up to about 17,000 nucleotide sequences comprising gene regulatory elements specifically associated with the cancer comprising the transcription factor gene fusion (e.g., tRCC), for example about 1,000-6,000, about 1,000-8,000, about 1,000-12,000, about 1,000-17,000, or 5,000-6,000, about 5,000-8,000, about 5,000-12,000, about 5,000-17,000.
[0082] For example, the methods disclosed here may be employed to detect about 5, GOO- 12, 000 (e.g., about 5,500 or about 11,000) nucleotide sequences comprising gene regulatory elements specifically associated with a cancer comprising a transcription factor gene fusion (e.g., tRCC). As shown in FIG. 4A, 11,435 nucleotide sequences comprising gene regulatory elements specifically associated with tRCC were used successfully to differentiate tRCC from ccRCC and healthy plasma samples, i.e., tRCC-specific H3K27ac sites. Similarly, FIG. 4B illustrates the use of 5,128 nucleotide sequences associated with a TFE3 gene fusion reliably differentiate tRCC from ccRCC and healthy plasma samples.
[0083] In some instances, the methods disclosed here may be employed to detect about 2,000-21,000 (e.g., about 2,500, about 6,500, about 9,000, about 11,500, about 14,000, or about 20,500) nucleotide sequences comprising gene regulatory elements specifically associated with a cancer comprising a transcription factor gene fusion (e.g. , tRCC). As shown in FIG. 10A, 2,450 nucleotide sequences comprising gene regulatory elements specifically associated with tRCCwere used successfully to differentiate tRCC from ccRCC and healthy plasma samples, i.e., tRCC-specific H3K4me3 sites. Similarly in FIG. 10B, 11,443 nucleotide sequences comprising gene regulatory elements specifically associated with tRCC were used successfully to differentiate tRCC from ccRCC and healthy plasma samples, i.e., tRCC-specific H3K27ac sites. Further, FIG. 4B illustrates the use of 6,540 nucleotide sequences associated with a TFE3 gene fusion reliably differentiate tRCC from ccRCC and healthy plasma samples.Nucleotide sequences associated with a transcription factor gene fusion
[0084] Chromosomal rearrangements that result in transcription factor gene fusions are associated with various different cancers. For example, tRCC is an aggressive subtype of RCC driven by gene fusions involving an MiT / TFE family transcription factor, most commonly TFE3. The TFE3 fusions result in the production of a neomorphic transcription factor with the potential to drive a distinctive epigenomic signature (Damayanti, N. P. el al. Clin Cancer Res 24, 5977- 5989 (2018)). Similarly, there are also subtypes of prostate cancer characterized by a transcription factor gene fusion, e.g., TMPRSS2-ERG, that yields a neomorphic transcription factor resulting in a distinctive epigenomic signature. Another example is synovial sarcoma which is often characterized by a translocation between chromosomes X and 18, leading to the formation of SS18:SSX fusion oncogenes, including SSX2 gene fusions.
[0085] Nucleotide sequences associated with a transcription factor gene fusion (e.g., TFE3, TMPRSS2-ERG or SSX2) include gene regulatory elements that are bound by the chimeric transcription factor fusion protein encoded by the transcription factor gene fusion. The chimeric transcription factor fusion protein encoded by the transcription factor gene fusion is also referred to as a neomorphic transcription factor. As demonstrated herein, such gene regulatory elements are often distinct from gene regulatory elements bound by the corresponding native transcription factor.
[0086] Association of a chimeric transcription factor fusion protein with gene regulatory elements such as transcription factor binding sites can result in the recruitment of histone acetylases and histone methylases. Nucleotide sequences associated with a transcription factor gene fusion that are associated with histones that are differentially acetylated and / or methylatedin a cancer genome comprising a transcription factor gene fusion relative to a cancer genome that does not comprise the transcription factor gene fusion can then be enriched and detected as described herein. For example, nucleotide sequences associated with a TFE3 gene fusion that are associated with histones that are differentially acetylated and / or methylated in the tRCC cancer genome relative to other RCC genomes can then be enriched and detected as described herein.
[0087] Association of a chimeric transcription factor fusion protein with gene regulatory elements such as transcription factor binding sites may also result in the recruitment of DNA methylases. Nucleotide sequences associated with a transcription factor gene fusion that are differentially methylated in a cancer genome comprising a transcription factor gene fusion relative to a cancer genome that does not comprise a transcription factor gene fusion can likewise be enriched and detected as described herein.
[0088] As shown herein, the described methods are capable of detecting the epigenetic signature of a chimeric transcription factor fusion protein in a cancer genome comprising a transcription factor gene fusion. The methods are highly sensitive and are directly linked to the underlying disease biology of the cancer comprising the transcription factor gene fusion. The methods are not limited to the detection of any particular translocation type or event, or any specific transcription factor gene fusion. Accordingly, the methods described herein may be used for the detection of any cancer comprising a transcription factor gene fusion.
[0089] This is especially important in the detection of tRCC, because over twenty different MiT / TFE fusion partners have been reported to date and genomic breakpoint locations can vary substantially between patients, which makes detecting the translocation per se an unreliable method of detecting tRCC. In contrast, as shown herein, the epigenomic signatures of various TFE3 gene fusions resulting from different genomic break points overlap between different types of tRCC and can be detected by the methods disclosed herein. In each instance, the TFE3 gene fusions may be associated with substantially the same epigenomic signature.Nucleotide sequences that are not associated with a transcription factor gene fusion
[0090] The methods disclosed herein may also include detecting the presence of or level of one or more nucleotide sequences that are not associated with the cancer (or cancer subtype) comprising the transcription factor gene fusion (e.g., tRCC). For example, the method may further comprise detecting the presence or level of one or more nucleotide sequences associated with differentially methylated regions (DMRs) associated with a cancer subtype affecting the same tissue or organ that does not comprise the transcription factor gene fusion (e.g., a non- tRCC subtype of RCC, e.g., ccRCC). In this instance, a cancer comprising a transcription factor gene fusion is characterized by the absence of such DMRs. For example, tRCC is characterized by the absence of DMRs associated with ccRCC.Exemplary nucleotide sequences
[0091] Provided herein are nucleotide sequences comprising gene regulatory elements specifically associated with tRCC which can be used to differentiate tRCC from other common subtypes of RCC, such as ccRCC. These include nucleotide sequences associated with a TFE3 gene fusion.
[0092] Exemplary nucleotide sequences comprising gene regulatory elements specifically associated with tRCC are set forth in SEQ ID NOs: 1-16,563. The nucleotide sequences set forth in SEQ ID NOs: 1-11,435 comprise gene regulatory elements specifically associated with tRCC that do not overlap with the nucleotide sequences associated with TFE3 gene fusions identified herein. Nucleotide sequences that are specifically associated with TFE3 gene fusions as identified herein, specifically TFE3 binding sites, are set forth in SEQ ID NOs: 11,436-16,563.
[0093] Exemplary nucleotide sequences that are differentially methylated in tRCC include those set forth in SEQ ID NOs: 16,564-17,069. Exemplary nucleotide sequences that are differentially methylated in ccRCC include those set forth in SEQ ID NOs: 17,070-17,184. The presence of nucleotide sequences that are not associated with TFE3 gene fusion may suggest that the subject does not have tRCC. For example, the presence of or increased level of one or more nucleotide sequences as set forth in SEQ ID NOs: 17,070-17,184 is indicative that the subject has ccRCC.
[0094] Due to the variability associated with the sites of translocation events and / or single nucleotide polymorphisms (SNPs) present in individual cancer genomes, nucleotide sequences comprising gene regulatory elements specifically associated with tRCC (e.g., nucleotide sequences associated with a TFE3 gene fusion), nucleotide sequences that are differentially methylated in tRCC, and nucleotide sequences that are differentially methylated in ccRCC may be at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or identical to the sequences set forth in SEQ ID NOs: 1-17,184.
[0095] Also provided herein are nucleotide sequences comprising gene regulatory elements specifically associated with prostate cancer which can be used to differentiate prostate cancer comprising a TMPRSS2-ERG gene fusion from other common subtypes of prostate cancer. These include nucleotide sequences associated with a TMPRSS2-ERG gene fusion.
[0096] Exemplary nucleotide sequences comprising gene regulatory elements specifically associated with a prostate cancer comprising a TMPRSS2-ERG gene fusion include those set forth in SEQ ID NOs: 17,185-24,715. Translocation events and / or SNPs present in individual cancer genomes comprising a TMPRSS2-ERG fusion may vary. Nucleotide sequences associated with a prostate cancer comprising a TMPRSS2-ERG gene fusion may be at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or identical to the sequences set forth in SEQ ID NOs: 17,185-24,715.Libraries
[0097] Also provided herein is a library of isolated chromatin-associated DNA enriched for one or more nucleotide sequences comprising gene regulatory elements specifically associated with a cancer comprising a transcription factor gene fusion (e.g., tRCC,). The library may be obtained by the methods described herein.
[0098] For example, also provided herein is a method of preparing a library of isolated chromatin-associated DNA comprising gene regulatory elements specifically associated with a cancer comprising a transcription factor gene fusion (e.g., tRCC), comprising isolating cell-free DNA (cfDNA) from a sample obtained from a subject suffering from the cancer comprising the transcription factor gene fusion (e.g., tRCC), enriching the cfDNA with a means for specificallybinding histone H3 protein acetylated at the lysine at residue 27 (H3K27ac) for H3K27ac- associated DNA and / or with a means for specifically binding histone H3 protein methylated at the lysine at residue 4 (H3K4me3) for H3K4me3 -associated DNA and isolating nucleotide sequences comprising gene regulatory elements specifically associated with the cancer comprising the transcription factor gene fusion (e.g., tRCC) from the cfDNA enriched for H3K27ac-associated DNA and / or H3K4me3 -associated DNA. In some instances, the method further comprises a step of sequencing the isolated nucleotide sequences.
[0099] Accordingly, also provided herein is a library comprising one or more of the nucleotide sequences comprising gene regulatory elements specifically associated with a cancer comprising a transcription factor gene fusion (e.g., tRCC). Such a nucleotide sequence library comprising an epigenomic signature of a cancer comprising a transcription factor gene fusion (e.g., tRCC) may be used in the computational methods described herein that can be applied in the diagnosis and monitoring of this cancer in a subject. The library may comprise a set of nucleotide sequences comprising gene regulatory elements specifically associated with a cancer comprising a transcription factor gene fusion, for example, about 1,000-6,000, about 1,000- 8,000, about 1,000-12,000, or about 1,000-17,000, or about 5,000-6,000, about 5,000-8,000, about 5,000-12,000, or about 5,000-17,000 nucleotide sequences comprising gene regulatory elements specifically associated with the cancer comprising the transcription factor gene fusion.[000100] In some instances, the library may comprise one or more additional nucleotide sequences that can be used in the differential diagnosis of cancer subtypes affecting the same organ or tissue. For example, the library may further comprise one or more nucleotide sequences associated with differentially methylated regions (DMRs) associated with a cancer subtype affecting the same tissue or organ that does not comprise the transcription factor gene fusion. For example, a library for use in diagnosing and / or monitoring tRCC may comprise one or more of SEQ ID NOs: 1-11,435, SEQ ID NOs: 11,436-16,563, and / or SEQ ID NOs: 16,564-17,069, e.g, about 1,000, about 5,000, about 10,000, or about 17,000 (or more) of the nucleotide sequences set forth in SEQ ID NOs: 1-17,069, e.g., all nucleotide sequences set for in SEQ ID NOs: 1-11,435, SEQ ID NOs: 11,436-16,563, and / or SEQ ID NOs: 16,564-17,069. In some instances, the library for tRCC may comprise a subset of the nucleotide sequences comprisinggene regulatory elements specifically associated with tRCC described herein, for example, about 1,000-6,000, about 1,000-8,000, about 1,000-12,000, or about 1,000-17,000, or about 5,000- 6,000, about 5,000-8,000, about 5,000-12,000, or about 5,000-17,000 nucleotide sequences comprising gene regulatory elements specifically associated with tRCC. The library may further comprise one or more nucleotide sequences not associated with tRCC. For example, the library may comprise one or more nucleotide sequences associated with differentially methylated regions (DMRs) associated with a non-tRCC subtype, e.g. such as ccRCC, e.g., one or more (or all) of the nucleotide sequences set forth in SEQ ID NOs: 17,070-17,184.[000101] Also provided is a nucleotide sequence library comprising an epigenomic signature of a prostate cancer comprising a TMPRSS2-ERG gene fusion that can be applied in the diagnosis and monitoring of this cancer. For example, such a library may comprise one or more of SEQ ID NOs: 17,185-24,715, e.g., about 500, about 1,000, about 5,000, or about 7,000 (or more) of the nucleotide sequences set forth in SEQ ID NOs: 17,185-24,715.DetectionSequencing[000102] The one or more nucleotide sequences comprising gene regulatory elements specifically associated with a cancer comprising a transcription factor gene fusion such as tRCC (e.g., one or more nucleotide sequences associated with a TFE3 gene fusion) can be detected by sequencing after ChIP or MeDIP as described herein. Suitable ChlP-Seq and MeDIP-seq methods are known in the art, for example, as described in Baca, et al. Nature Medicine vol. 29,11 (2023): 2737-2741.[000103] When employed in the present methods, the ChlP-Seq and optionally MeDIP- Seq methods may be used for the simultaneous detection of up to about 17,000 (or more) nucleotide sequences comprising gene regulatory elements specifically associated with a cancer comprising a transcription factor gene fusion (e.g., tRCC), for example, about 1,000-6,000, about 1,000-8,000, about 1,000-12,000, or about 1,000-17,000, or about 5,000-6,000, about 5,000- 8,000, about 5,000-12,000, or about 5,000-17,000 nucleotide sequences.[000104] ChlP-Seq and / or MeDIP-Seq can also be used to detect increased levels of the one or more nucleotide sequences comprising gene regulatory elements specifically associated with a cancer comprising a transcription factor gene fusion (e.g., tRCC) as compared to the levels of the one or more nucleotide sequences in a control sample (for instance, from a healthy subject or a subject suffering from another cancer affecting the same organ or tissue, e.g., an RCC other than tRCC). In ChlP-Seq and optionally MeDIP-Seq, the increased level corresponds to an increased read density of the nucleotide sequences comprising gene regulatory elements specifically associated with a cancer comprising a transcription factor gene fusion (e.g., tRCC), as such nucleotide sequences are overrepresented in the cfDNA of samples when the cancer comprising the transcription factor gene fusion (e.g., tRCC) is present as compared to a control sample (e.g., from a healthy individual).[000105] In the described methods, the presence of, or an increased level of, one or more nucleotide sequences comprising gene regulatory elements specifically associated with a cancer comprising a transcription factor gene fusion such as tRCC (e.g., one or more nucleotide sequences associated with a TFE3 gene fusion) in a sample are indicative of a subject suffering from the cancer comprising a transcription factor gene fusion. The level of one or more nucleotide sequences comprising gene regulatory elements specifically associated with a cancer comprising a transcription factor gene fusion such as tRCC (e.g., one or more nucleotide sequences associated with a TFE3 gene fusion) may change in the presence of the cancer comprising the transcription factor gene fusion (e.g., as result of the presence of tRCC and / or as a result of tumor progression or remission). The level of one or more nucleotide sequences comprising gene regulatory elements specifically associated with a cancer comprising a transcription factor gene fusion (e.g., tRCC) may be compared to the corresponding level of the nucleotide sequences comprising gene regulatory elements specifically associated with the cancer comprising the transcription factor gene fusion in a control sample. The control sample may be from a subject without cancer. The control sample may be from a subject diagnosed with a subtype of a cancer affecting the same organ or tissue (e.g., an RCC other than tRCC).[000106] The level of the one or more nucleotide sequences comprising gene regulatory elements specifically associated with a cancer comprising a transcription factor gene fusion suchas tRCC (e.g., one or more nucleotide sequences associated with a TFE3 gene fusion) may be more than 2 times greater in the test sample (e.g., at least 3 times, at least 5 times, at least 10 times, or at least 100 times greater) than the level of the corresponding nucleotide sequence in the control sample. For example, the level of the nucleotide sequence comprising gene regulatory elements specifically associated with a cancer comprising a transcription factor gene fusion such as tRCC (e.g., one or more nucleotide sequences associated with a TFE3 gene fusion) may be at least 10 times greater (e.g., at least 20 times, at least 50 times, at least 100 times, at least 1000 times greater) than the level of the corresponding nucleotide sequence in the control sample.[000107] In some instances, the increase in the level of the one or more nucleotide sequences comprising gene regulatory elements specifically associated with a cancer comprising a transcription factor gene fusion (e.g., tRCC) is determined relative to a pre-determined threshold value and thus provides a relative measure. For example, as shown in Examples 3 and 4, pre-determined threshold values for signal density may be used to determine whether the level of one or more nucleotide sequences comprising gene regulatory elements specifically associated with a cancer comprising a transcription factor gene fusion (e.g., tRCC) are increased or decreased, e.g, over the course of a treatment for that cancer. The levels can be determined by a ChlP-Seq and MeDIP-Seq method as described herein. In other examples, an absolute number of sequence reads or an area can be used instead of signal density, as explained below.[000108] Typically, the presence of, or increased levels of, a plurality of nucleotide sequences comprising gene regulatory elements specifically associated with a cancer comprising a transcription factor gene fusion (e.g, 1,000 or more, e.g, 5,000 or more, or 10,000 or more) is indicative of a subject suffering from that cancer. The plurality of nucleotide sequences comprising gene regulatory elements specifically associated with a cancer comprising a transcription factor gene fusion may be referred to as an “epigenomic signature”. Generating epigenomic signatures for cancer subtyping cancer is described in Baca, et al. Nature Medicine vol. 29,11 (2023): 2737-2741. For example, an epigenomic signature for tRCC may include one or more of SEQ ID NOs: 1-11,435 SEQ ID NOs: 11,436-16,563, and / or SEQ ID NOs: 16,564- 17,069. An epigenomic signature for prostate cancer comprising a TMPRSS2-ERG gene fusion may include one or more of SEQ ID NOs: 17,185-24,715.[000109] As shown herein, an epigenomic signature of a cancer comprising a transcription factor gene fusion (e.g., tRCC) is distinct from an epigenomic signature of a cancer that does not comprise the transcription factor gene fusion (e.g., other types of RCC), and an epigenomic signature seen in healthy control samples. The plurality of nucleotide sequences from the control sample may contain fewer nucleotide sequences comprising gene regulatory elements specifically associated with a cancer comprising a transcription factor gene fusion such as tRCC (e.g., one or more nucleotide sequences associated with a TFE3 gene fusion) than the plurality of nucleotide sequences in the test sample. Alternatively, the plurality of nucleotide sequences from the control sample may contain different nucleotide sequences, i.e., other nucleotides sequence than those comprising gene regulatory elements specifically associated with a cancer comprising a transcription factor gene fusion. For example, when monitoring tumor progression in a cancer comprising a transcription factor gene fusion (e.g., tRCC), the number and / or identity of the nucleotide sequences comprising gene regulatory elements specifically associated with the cancer may change over time. For instance, the level of the one or more nucleotide sequences comprising gene regulatory elements specifically associated with the cancer comprising the transcription factor gene fusion may increase as the tumor progresses.[000110] In one aspect, a method for detecting (e.g., diagnosing or monitoring) a cancer comprising a transcription factor gene fusion (e.g., tRCC) in a subject is provided that comprises: (a) isolating cell-free (cfDNA) from a sample obtained from the subject; (b) enriching the cfDNA with a means for specifically binding histone H3 protein acetylated at the lysine at residue 27 (H3K27ac) for H3K27ac-associated DNA and / or with a means for specifically binding histone H3 protein methylated at the lysine at residue 4 (H3K4me3) for H3K4me3 -associated DNA; and (c) detecting the presence of or level of one or more nucleotide sequences comprising gene regulatory elements specifically associated with the cancer comprising the transcription factor gene fusion in the cfDNA enriched for H3K27ac-associated DNA and / or H3K4me3 -associated DNA, wherein the presence or level of the one or more nucleotide sequences comprising gene regulatory elements specifically associated with the cancer comprising the transcription factor gene fusion in the enriched cfDNA is indicative of the subject suffering from the cancer comprising the transcription factor gene fusion.[000111] To increase assay sensitivity, more than one set of nucleotide sequences comprising gene regulatory elements specifically associated with a cancer comprising a transcription factor gene fusion may be analyzed. For example, assay sensitivity can be increased by separately analyzing ChlP-seq reads obtained from H3K27ac-associated DNA in a cfDNA sample obtained from a subject to determine the presence and / or level of (a) a pre-determined first set of nucleotide sequences comprising gene regulatory sequences specifically associated with cancer comprising a transcription factor gene fusion (e.g., tRCC), and (b) a second predetermined set of nucleotide sequences comprising transcription factor binding sites of the fusion transcription factor encoded by the transcription factor gene fusion (e.g., a TFE3 gene fusion). Sensitivity may be further boosted by separately analyzing ChlP-seq reads obtained from H3K4me3 -associated DNA in the cfDNA sample to determine the presence and / or level of a pre-determined third set of nucleotide sequences comprising gene regulatory sequences specifically associated with the cancer comprising the transcription factor gene fusion. Typically, there is no overlap between the nucleotide sequences comprised in the first and second predetermined sets. In some instances, there is no overlap between the nucleotide sequences comprised in the first, second, and third pre-determined sets. The levels (or number of sequencing reads) of each nucleotide sequence may be aggregated separately for each predetermined list. The aggregate numbers can be used to calculate separate aggregate scores. As shown herein, combining two or three aggregate scores increases the sensitivity of the methods described herein.[000112] In one aspect, the plurality of nucleotide sequences comprising gene regulatory elements specifically associated with a cancer comprising a transcription factor gene fusion such as tRCC (e.g., one or more nucleotide sequences associated with a TFE3 gene fusion) of the test sample may be compared to a pre-determined set or plurality of nucleotide sequences comprising gene regulatory elements specifically associated with the same cancer comprising the same transcription factor gene fusion that is indicative of the presence of that cancer, e.g, a library of one or more of the nucleotide sequences comprising gene regulatory elements specifically associated with said cancer.[000113] For example, the pre-determined plurality of sequences comprising gene regulatory elements specifically associated with tRCC may include one or more of any one of SEQ ID NOs: 1-11,435, SEQ ID NOs: 11,436-16,563, and / or SEQ ID NOs: 16,564-17,069 e.g, one or more of SEQ ID NOs: 1-11,435, SEQ ID NOs: 11,436-16,563, and / or SEQ ID NOs: 16,564-17,069, e.g., about 1,000, about 5,000, about 10,000, or about 17,000 (or more) of the nucleotide sequences set forth in SEQ ID NOs: 1-17,069, e.g., all nucleotide sequences set for in SEQ ID NOs: 1-11,435, SEQ ID NOs: 11,436-16,563, and / or SEQ ID NOs: 16,564-17,069. The pre-determined plurality of sequences comprising gene regulatory elements specifically associated with a prostate cancer comprising a TMPRSS2-ERG gene fusion may include one or more of any one of SEQ ID NOs: 17,185-24,715, e.g., about 500, about 1,000, about 5,000, or about 7,000 (or more) of the nucleotide sequences set forth in SEQ ID NOs: 17,185-24,715.[000114] The pre-determined plurality of nucleotide sequences comprising gene regulatory elements specifically associated with a cancer comprising a transcription factor gene fusion such as tRCC (e.g., one or more nucleotide sequences associated with a TFE3 gene fusion) that is indicative of the presence of said cancer may be referred to as a “reference epigenomic signature”. Substantial overlap between the nucleotide sequences comprising gene regulatory elements specifically associated with a cancer comprising a transcription factor gene fusion present in the test sample with the pre-determined plurality of nucleotide sequences comprising gene regulatory elements specifically associated with said cancer is indicative of the presence of said cancer in the test sample. For example, 70%, 80%, 90% or more overlap in the nucleotide sequences in the test and reference samples is indicative of the presence of said cancer in the test sample.[000115] For example, using the methods described herein, signal density and / or area under the curve and / or sequence read count may be obtained for peaks of sequencing reads that correspond to nucleotide sequences comprising gene regulatory elements specifically associated with a cancer comprising a transcription factor gene fusion (e.g., tRCC). The values obtained for the signal density and / or area under the curve and / or sequence read count can be compared to pre-determined threshold values for signal density and / or area under the curve and / or sequence read count. The pre- determined threshold values can be obtained by comparing correspondingsignal density and / or area under the curve and / or sequence read count values from a plurality of samples that are confirmed to be the cancer comprising the transcription factor gene fusion (e.g., tRCC) and a plurality of control samples. For example, the threshold value for the signal density can be set to return an area under the curve >0.8 for a receiver operating characteristic curve of a classifier relying on the threshold value. The threshold value for the signal density varies depending on the specific epigenomic signature, i.e., the pre-determined set of nucleotide sequences comprising gene regulatory elements specifically associated with the cancer comprising the transcription factor gene fusion (e.g., tRCC). Typically, the threshold value is higher the larger the number of nucleotide sequences in the pre-determined set.[000116] In some instances, to increase assay sensitivity, signal density and / or area under the curve (AUC) and / or sequence read count values may be obtained for (a) peaks of ChlP-seq reads obtained from H3K27ac-associated DNA for a pre-determined first set of nucleotide sequences comprising gene regulatory sequences specifically associated with a cancer comprising a transcription factor gene fusion (e.g., tRCC) and a second pre-determined set of nucleotide sequences comprising transcription factor binding sites of the fusion transcription factor encoded by the transcription factor gene fusion (e.g., a TFE3 gene fusion), and optionally (b) peaks of ChlP-seq reads obtained from H3K4me3 -associated DNA for a pre-determined third set of nucleotide sequences comprising gene regulatory sequences specifically associated with the cancer comprising the transcription factor gene fusion. As demonstrated in Examples 6 and 7, combining these values (e.g., sequence read count values for each of the first, second, and optionally third pre-determined sets of nucleotide sequences) to generate an integrated epigenomic score can boost sensitivity of the disclosed detection methods.Next seneration sequencing and other sequencing methods[000117] Isolated cfDNA enriched for H3K27ac-associated DNA and / or H3K4me3- associated DNA can be subjected to next generation sequencing (NGS) to detect one or more nucleotide sequences comprising gene regulatory elements specifically associated with a cancer comprising a transcription factor gene fusion (e.g., tRCC). Multiple different NGS platform technologies are known to the person skilled in the art such as systems provided by Illumina, Sequenom, ION Torrent Systems, Halcyon Molecular, NABsys, IBM, and GE Global.Alternative sequencing platforms are known to the person skilled in the art such as Nanopore sequencing (Oxford Nanopore Technologies) and single molecule real-time (SMRT) sequencing, e.g., as provided by Pacific Biosciences Hifi.Whole genome sequencing[000118] In some instances, whole genome sequencing (WGS) may be used for the detection of one or more nucleotide sequences comprising gene regulatory elements specifically associated with a cancer comprising a transcription factor gene fusion such as tRCC (e.g, one or more nucleotide sequences associated with a TFE3 gene fusion). WSG is a comprehensive method for analyzing genomic DNA sequences within a data set at a single time. This technique typically involves breaking down the DNA into smaller fragments, sequencing these fragments, e.g, by next generation sequencing (NGS), and then using computational methods to reassemble the sequences to generate a complete picture of the genetic data set.[000119] WGS provides detailed information on genetic variations, including single nucleotide polymorphisms (SNPs), insertions, deletions, and structural variants. WGS, including low pass WGS (IpWGS), allows for the identification of nucleotide sequences comprising gene regulatory elements specifically associated with a cancer comprising a transcription factor gene fusion such as tRCC (e.g., nucleotide sequences associated with a TFE3 gene fusion) across an entire cancer genome.[000120] In some instances, the methods disclosed herein may employ ultra-low pass whole genome sequencing (ulpWGS) as described in Christodoulou, E., etal. Precis. One. 7, 21 (2023). ulpWGS utilizes NGS technologies, such as Illumina or Ion Torrent, to generate millions of short DNA reads. These reads are aligned to a reference genome to identify genetic information.[000121] While traditional WGS involves deep coverage to capture every base pair of an individual's genome, ulpWGS sequences a genome at a lower depth, typically around O.lx to 5x coverage, as compared to the standard depth of 30x to 50x in traditional WGS. While the lower coverage may result in missing some rare variants, it still allows for the identification of common genetic sequences across the data set. On the other hand, this reduction in coverage significantlydecreases the cost and time required for analysis while still providing valuable genomic information. As demonstrated herein, WGS, IpWGS and ulpWGS can be used to infer copy number profiles and cfDNA tumor content.[000122] All of WGS, IpWGS and ulpWGS, allow for the simultaneous detection of a very large number of nucleotide sequences comprising gene regulatory elements specifically associated with a cancer comprising a transcription factor gene fusion such as tRCC (e.g., nucleotide sequences associated with a TFE3 gene fusion) in a single assay. WGS allows for the simultaneous detection of a plurality of nucleotide sequences comprising gene regulatory elements specifically associated with a cancer comprising a transcription factor gene fusion (e.g., tRCC) that make up an epigenomic signature of said cancer.[000123] In some instances, WGS (e.g., LPWGS) is performed on the isolated cfDNA to determine the tumor fraction. The tumor fraction of cfDNA is commonly referred to as circulating tumor DNA (ctDNA). An exemplary method to determine the tumor fraction in isolated cfDNA (i.e., the cfDNA fraction) from read abundance across bins spanning the genome of a subject is described in Example 2.[000124] The methods for diagnosing and / or monitoring a cancer comprising a transcription factor gene fusion in a subject described herein are especially sensitive, in particular when compared to methods that rely on detecting copy number alterations in cfDNA or changes in the transcriptional profile of circulating tumor cells (CTCs). Specifically, the disclosed methods are sensitive enough that they can be used in subjects having a ctDNA fraction of about 1% or more.Comyuter-imylemented methods[000125] International patent applications PCT / US2024 / 051147 and PCT / US2024 / 056491 disclose computer-implemented methods similar to those described below, and can be adapted for use in a method of analysis as described in the present application.[000126] In another aspect, provided herein is a computer program comprising a computer program code configured to cause one or more physical computing devices to perform a methodof analysis of sample results for a subject when the code is run on the one or more physical computing devices, the method of analysis comprising: a) Receiving sequencing read data obtained from cfDNA isolated from a sample obtained from the subject and enriched for H3K27ac-associated DNA and optionally sequencing read data obtained from cfDNA isolated from the sample and enriched for H3K4me3- associated DNA; b) Aligning the sequencing read data obtained in (a) with whole human genome read data; c) Receiving a first predetermined list of transcription factor fusion binding sites, a second predetermined list of gene regulatory elements associated with H3K27ac, and optionally a third predetermined list of gene regulatory elements associated with H3K4me3, wherein the gene regulatory elements in the second predetermined list and optionally the third predetermined list do not comprise the transcription factor fusion binding sites in the first predetermined list; d) For each site or element present in one of the lists received in (c) determining the number of aligned sequencing reads obtained in (b); e) Aggregating the number of sequencing reads for each element in each list; f) Receiving a reference value; g) For each aggregate, normalizing the aggregate based on the reference value; h) Calculating a combined score based on the two or more aggregates; and i) Comparing the combined score to a predetermined score to determine whether the cancer comprises the transcription factor gene fusion.Aligning the sequencing read data[000127] The sequencing read data can be aligned with whole human genome read data using Burrows- Wheeler Aligner (e.g., version 0.7.1740), or any other suitable software known in the art. By aligning the data, portions of the data corresponding to the transcription factor fusion binding sites in the first predetermined list, gene regulatory elements associated with H3K27ac in the second predetermined list, and optionally gene regulatory elements associated with H3K4me3 in the third predetermined list can be identified.Determining the number of aligned sequencing reads[000128] Peaks in the data at the transcription factor fusion binding sites in the first predetermined list and gene regulatory elements in the second and third predetermined lists can be identified using, for example, the ChiLin computational pipeline or any other suitable software known in the art. In some examples, the peak identification may be configured to identify peaks having a width less than a threshold. The threshold may be configurable, and may be about 8kb (that is, any peaks extending by more than 4kb from the center of the peak). Filtering out peaks having widths greater than this threshold can improve the final classification by increasing the likelihood that the peaks being considered to represent specific transcription factor (TF) binding events or localized histone modifications, and reducing the likelihood that the peaks are chromatin accessibility sites that are non-specific for the cancer comprises the transcription factor gene fusion.[000129] To facilitate a standardized analysis of the peaks, the peaks can be resized to a set width (e.g., to a standard width of about 2-8kb, i.e., ± l-4kb extending from the center of the peak). For example, the peaks may be resized to have a standard width of 6kb (that is, extending from the center of the peak by 3kb in each direction). For peaks wider than 6kb, this can be achieved by discarding the tails of the peaks to reduce the peak width to 6kb. For peaks narrower than 6kb, this can be achieved by expanding what is considered to be the peak by including ‘background’ data surrounding the peak. The resizing of the peaks is not essential; however, it can assist in improving the accuracy of the final classification.[000130] For a given list of transcription factor binding sites or gene regulatory elements, the number of sequencing reads at each corresponding peak is determined. For example, a first peak may contain 1000 reads while a second peak may contain 1500 reads. In some examples, this can be expressed as a signal density or an area under the curve defined by the peak. Where an example makes use of one of these measures, it should be understood that either of the other measures could be used instead.[000131] In some examples, each peak may be divided into a plurality of bins, with each bin containing a subset of total number of sequence reads. For NGS, the average sequence readlength typically ranges from about 75-15Obp. Each bin may have a width of about 30-1 OObp (e.g., about 35bp, 40bp, 45bp, 50bp, 55bp, 60bp, 65bp, 70bp, 75bp, 80bp, 85bp, 90bp, or 95bp). For example, each 6kb wide peak may be divided into 150 bins having a width of 40bp. The width of the bins is configurable, and in some examples might be larger than or less than 40bp. Dividing the peaks into bins is not essential, and in some examples this step may be omitted.Aggregating the sequencing reads[000132] For each list, the peaks corresponding to the sites or elements of that list are aggregated. In a simple case where the peaks are not divided into bins, the peaks can be aggregated by summing the number of sequencing reads (or area or signal density) of the peaks. This results in an aggregate value for each list.[000133] In cases where each peak has been divided into bins, the aggregation can be performed on a per-bin basis. For example, for each list, the number of sequence reads in the first bin of each peak can be summed together to create a first aggregated bin for an aggregated peak. Each subsequent bin can be aggregated in the same way, to produce an aggregated peak made up of the aggregated bins.Receiving one or more reference values[000134] One or more reference values can be obtained for normalizing the aggregates and calculating scores, based on a pre-determined set of DNAse hypersensitivity sites (DHS). In examples where the aggregates are normalized by division-based scaling, the reference values can be an aggregation of peaks at the DHS, obtained using the same method as described above for the predetermined lists, e.g., the gene regulatory elements associated with H3K27ac, and optionally H3K4me3. Different reference values may be used for normalizing the aggregate peaks obtained from the sample enriched for H3K27ac-associated DNA and the aggregate peaks enriched for H3K4me3 -associated DNA.[000135] The pre-determined set of DHS may be a subset of previously identified DHS, e.g., as identified in Baca, et al. supra. In some examples, between 1000 and 20,000 DHS (e.g., about 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10,000, 11,000, 12,000, 13,000, 14,000,15,000, 16,000, 17,000, 18,000, or 19,000) may be used to generate the one or more reference values. For instance, between 5,000 and 15,000 DHS may be used to generate the one ore more reference values. Typically, the reference values may be an aggregation of sequence reads for a pre-determined set of about 10,000 DHS.[000136] The reference values can be obtained from the aligned sequence read data in the same way that the scores for each list were obtained.Normalizing the aggregates[000137] Normalization of the aggregates can, in some examples, comprise two or more normalization steps.[000138] In a first (optional) step, applicable, e.g., where the peaks have been divided into bins, the aggregated peaks can undergo shoulder normalization. A background value for each aggregated peak can be determined based on an aggregation of the sequence reads at the tail end portions of the peak (e.g., 5-10% of the peak width at each tail end of the peak). For example, the sequence reads in the locations -3kb to -2.8kb and +2.8kb to +3 kb (relative to the center of the peak) can be aggregated to obtain a background value. In other examples, other locations for the tails can be used. For example, a range of ±0.05-0.4kb at each tail end may be used for shoulder normalization.[000139] The background value can then be subtracted from the aggregate, to effectively set the value of the shoulders of the aggregated peak to 0. In some examples, the aggregated peaks might not undergo shoulder normalization.[000140] Each of the aggregated peaks is normalized relative to a reference value obtained using the predetermined set of DHS, by calculating the ratio between the number of sequence reads in an aggregated peak and the reference value. This ratio is the score for that aggregate. The aggregated peaks obtained from the sequencing read data obtained from a sample of cfDNA enriched for H3K27ac-associated DNA are normalized relative to a reference value obtained from the sequence read data obtained for the predetermined set of DHS obtained from same sample. The aggregated peak obtained from the sequencing read data obtained from a sample ofcfDNA enriched for H3K4me3-associated DNA is normalized relative to a reference value obtained from sequence read data obtained using the predetermined set of DHS obtained from the same sample. This means that the two reference values may not be the same.[000141] Where signal density or area under the peaks is used in place of the number of sequence reads, normalization can be performed in the same way, with the reference values being corresponding signal densities or areas calculated using the sequence read data obtained using the predetermined set of DHS.Calculating a combined score[000142] In some examples, a combined score is calculated based on the individual scores for the aggregates. The combined score can be calculated by summing the individual scores for the aggregates. In other examples where only one reference value is required (e.g., where data is obtained only from a sample enriched for H3K27ac-associated DNA), a combined score may be calculated by first summing the number of sequences reads in the aggregates to create a superaggregate, and then calculating the ratio of the super-aggregate and the reference value obtained using the pre-determined set of DHS.[000143] In other examples, as explained below, a combined score can be determined using a statistical model that has been trained to generate scores for each of the pre-determined first, second and optionally third lists based on sequencing read data.Comparing the combined score to a predetermined score[000144] The combined score can be compared to a predetermined score to determine whether the subject has the transcription factor gene fusion. The exact value of the predetermined score depends on the choices of the threshold peak width, the width to which the peaks are standardized, the width of the bins, and the size of the tail end portion chosen for shoulder normalization.[000145] For instance, in the provided example involving tRCC, the threshold peak width was originally chosen to be 8kb (4kb in each direction from the peak center). The peaks were then resized to a standard width of 6kb (3kb in each direction from the peak center), the peakswere divided into bins having a width of 40b, and shoulder normalization was performed based on the sequence reads in the locations -3000 to -2800 and +2800 to +3000 about the peak center. In this example, the method was able to distinguish between subjects having tRCC and subjects having ccRCC with an AUROC (area under the receiver operating characteristic curve) of 0.87. The method was able to distinguish between subjects having tRCC and healthy control subjects (that is, subjects having neither tRCC nor ccRCC) with an AUROC of 0.91.[000146] The combined score can be compared to the predetermined score using any classifier known to be suitable for binary classification tasks, such as logistic regression.Determining the predetermined score[000147] The predetermined score may be determined by any known method of obtaining an optimal threshold score.[000148] For example, the predetermined score may be determined using the ROC (receiving operation characteristic) curve and Youden’s J statistic (J=True positive rate - False positive rate). When provided with a set of scores obtained from subjects that have been preclassified as having or not having a cancer comprising transcription factor gene fusion, an initial threshold score can be arbitrarily selected and incrementally varied until Youden’s J statistic is maximized.[000149] In another example, a cost-based method can be used. For example, a custom loss function can be defined as Cost=(False Positive Rate*Cl )+(False Negative Rate*C2), where Cl and C2 are configurable variables that weight the cost of a false positive and a false negative. By minimizing the cost, an optimal predetermined score can be determined while allowing the classification to be intentionally biased towards false positives or false negatives to account for the difference in real-world significance between a false positive result and a false negative result.[000150] The method can be implemented as a computer program comprising a computer program code configured to cause one or more physical computing devices to perform a method of determining a threshold score when the code is run on the one or more physical computing devices.Method using a statistical model[000151] In another aspect, provided herein is a computer program comprising a computer program code configured to cause one or more physical computing devices to perform a method of analysis of sample results for a subject when the code is run on the one or more physical computing devices, the method of analysis comprising:(a) receiving data indicative of the presence or level of one or more nucleotide sequences comprising gene regulatory elements specifically associated with a cancer comprising a transcription factor gene fusion such as tRCC (e.g., one or more nucleotide sequences associated with a TFE3 gene fusion) in a sample collected from the subject; (b) obtaining a statistical model; (c) comparing the data and to the statistical model to generate a score; and (d) outputting the score from the statistical model.[000152] The statistical model may be configured such that the score is suitable for at least one of (a) identifying whether or not the subject has a cancer comprising a transcription factor gene fusion in an organ or tissue (such as a kidney, e.g., tRCC or ccRCC) and (b) identifying whether the subject is responsive to a standard therapy for a cancer affecting the same organ or tissue (e.g., RCC).[000153] Data obtained from the methods disclosed herein may be used to generate a statistical model or machine learning algorithm to classify a cancer comprising a transcription factor gene fusion, such as renal cell carcinoma (RCC).Obtaining a list of transcription factor fusion binding sites[000154] The above-mentioned methods make use of a predetermined list of transcription factor fusion binding sites. Accordingly, in another aspect, provided herein is a computer program comprising a computer program code configured to cause one or more physical computing devices to perform a method of obtaining a list of transcription factor fusion binding sites when the code is run on the one or more physical computing devices. The method of obtaining a list of transcription factor fusion binding sites may comprise:(a) obtaining a first reference dataset and a second reference dataset, wherein the first reference dataset comprises a plurality of reference sequencing reads of DNA fragments associated with a fusion transcription factor encoded by a transcription factor gene fusion and the second reference dataset comprises a plurality of reference sequencing reads of DNA fragments associated with a corresponding wildtype transcription factor;(b) carrying out a peak analysis on the first reference dataset and the second reference dataset to identify a plurality of sites that overlap between the first reference dataset and the second reference data set;(c) determining significant overlap sites; and(d) selecting, from the significant overlap sites, the list of transcription factor fusion binding sites by selecting sites having a log2fold change in either direction greater than 1.[000155] The plurality of reference sequencing reads of the first reference dataset may be obtained by isolating DNA from a cell comprising the transcription factor gene fusion (e.g., an established cell line comprising the transcription factor gene fusion). For example, DNA may be isolated from the cell comprising the transcription factor gene fusion, and the isolated DNA may be fragmented. The resulting DNA fragments may be contacted with an antibody that specifically binds to the fusion transcription factor to enrich for DNA fragments that are associated with the fusion transcription factor. The plurality of reference sequencing reads of the second reference dataset may be obtained by isolating DNA from a cell comprising the wild-type transcription fusion. For example, DNA may be isolated from the cell comprising the wild-type transcription factor, and the isolated DNA may be fragmented. The resulting DNA fragments may be contacted with an antibody that specifically binds to the wild-type transcription factor to enrich for DNA fragments that are associated with the wild-type transcription factor. The enriched DNA is then subjected to sequencing to generate the first and second reference datasets.[000156] The peak analysis may be implemented using Diffbind, as is explained in more detail in the examples below. In some examples, overlap sites were determined to be significantif they had a false discovery rate ‘q’ of less than or equal to 0.01. In other examples, a different threshold q value may be used.Obtaining a list of gene regulatory elements associated with H3K27ac[000157] The above-mentioned methods make use of a predetermined list of gene regulatory elements associated with H3K27ac. In another aspect, provided herein is a computer program comprising a computer program code configured to cause one or more physical computing devices to perform a method of obtaining a list of gene regulatory elements associated with H3K27ac when the code is run on the one or more physical computing devices. The method of obtaining a list of gene regulatory elements associated with H3K27ac may comprise:(a) obtaining a first reference dataset and a second reference data set, wherein the first reference dataset comprises plurality of reference sequencing reads of DNA fragments associated with H3K27ac of a first cancer cell population comprising a transcription gene fusion and the second reference dataset comprises a plurality of reference sequencing reads of DNA fragments associated with H3K27ac in a second cancer cell population of the same tissue type as the cancer comprising the transcription factor gene fusion, wherein the second cancer cell population does not comprise the transcription factor gene fusion;(b) carrying out a differential peak analysis on the first reference dataset and the second reference dataset to identify a plurality of differentially marked sites present only in the first cancer cell population comprising the transcription gene fusion;(c) determining significant differentially marked sites of the differentially marked sites; and(d) selecting, from the significant differentially marked sites, the list of gene regulatory elements associated with H3K27ac by selecting sites having a logifold change in either direction greater than 1.[000158] The method may comprise comparing the list of gene regulatory elements associated with H3K27ac with the list of transcription factor fusion binding sites identified above. The list of gene regulatory elements associated with H3K27ac may be modified byremoving any significant differentially marked sites present in both lists. In other words, the gene regulatory elements in the second predetermined list (and optionally the third predetermined list) might not comprise the transcription factor fusion binding sites in the first predetermined list. Alternatively, the list of transcription factor fusion binding sites may be modified by removing any transcription factor fusion binding sites present in both lists. Alternatively, sites present in both lists might not be deleted and instead may be discounted when aggregating the peaks corresponding to one of the first and second predetermined lists.[000159] The first and second reference data sets may be obtained by isolating DNA from the first and second cancer populations. For example, DNA may be isolated from the first and second cancer cell populations, and the isolated DNA may be fragmented. The resulting DNA fragments may be contacted with an antibody that specifically binds to H3K27ac to enrich for DNA fragments that are associated with H3K27ac. The enriched DNA is then subjected to sequencing to generate the plurality of reference sequencing reads of DNA fragments associated with H3K27ac in the first and second cancer cell populations.[000160] The peak analysis may be implemented using Diffbind, as is explained in more detail in the examples below. In some examples, peaks were determined to be significant (that is, corresponding to enriched regions) if they had a false discovery rate ‘q’ of less than or equal to 0.01. In other examples, a different threshold q value may be used.Obtaining a list of gene regulatory elements associated with H3K4me3[000161] The above-mentioned methods optionally make use of a predetermined list of gene regulatory elements associated with H3K4me3. Accordingly, in another aspect, provided herein is a computer program comprising a computer program code configured to cause one or more physical computing devices to perform a method of obtaining a list of gene regulatory elements associated with H3K4me3when the code is run on the one or more physical computing devices. The method of obtaining a list of gene regulatory elements associated with H3K4me3 may comprise:(a) obtaining a first reference dataset and a second reference data set, wherein the first reference dataset comprises a plurality of reference sequencing reads of DNAfragments associated with H3K4me3 of a first cancer cell population comprising a transcription gene fusion and the second reference dataset comprises a plurality of reference sequencing reads of DNA fragments associated with H3K4me3 in a second cancer cell population of the same tissue type as the cancer comprising the transcription factor gene fusion, wherein the second cancer cell population does not comprise the transcription factor gene fusion;(b) carrying out a differential peak analysis on the first reference dataset and the second reference dataset to identify a plurality of differentially marked sites present only in the first cancer cell population comprising the transcription gene fusion;(c) determining significant differentially marked sites of the differentially marked sites; and(d) selecting, from the significant differentially marked sites, the list of gene regulatory elements associated with H3K4me3 by selecting sites having a logifold change in either direction greater than 2.[000162] The first and second reference data sets may be obtained by isolating DNA from the first and second cancer populations. For example, DNA may be isolated from the first and second cancer cell populations, and the isolated DNA may be fragmented. The resulting DNA fragments may be contacted with an antibody that specifically binds to H3K4me3 to enrich for DNA fragments that are associated with H3K27ac. The enriched DNA is then subjected to sequencing to generate the plurality of reference sequencing reads of DNA fragments associated with H3K4me3 in the first and second cancer cell populations.[000163] The peak analysis may be implemented using Diffbind, as is explained in more detail in the examples below. In some examples, peaks were determined to be significant (that is, corresponding to enriched regions) if they had a false discovery rate ‘q’ of less than or equal to 0.01. In other examples, a different threshold q value may be used.Exemplary methods[000164] An exemplary method for detecting a cancer comprising a transcription factor gene fusion in a liquid biopsy sample (e.g., blood or plasma) obtained from a subject comprises:(a) isolating cell-free DNA (cfDNA) from the sample;(b) enriching a first aliquot of the cfDNA with an antibody that specifically binds histone H3 protein acetylated at the lysine at residue 27 (H3K27ac) for H3K27ac- associated DNA and optionally enriching a second aliquot of the cfDNA with an antibody that specifically binds histone H3 protein methylated at the lysine at residue 4 (H3K4me3) for H3K4me3-associated DNA;(c) sequencing the enriched cfDNA aliquots obtained in step (b) to obtain a plurality of sequencing reads for H3K27ac-associated DNA and optionally a plurality of sequencing reads for H3K4me3-associated DNA; and(d) identifying (i) within the plurality of sequencing reads for the H3K27ac- associated DNA a pre-determined first set of nucleotide sequences comprising gene regulatory elements specifically associated with the cancer comprising the transcription factor gene fusion and a second pre-determined set of nucleotide sequences comprising transcription factor binding sites of the fusion transcription factor encoded by the transcription factor gene fusion (e.g., a TFE3 gene fusion), and optionally (ii) within the plurality of sequencing reads for the H3K4me3 -associated DNA a pre-determined third set of nucleotide sequences comprising gene regulatory elements specifically associated with the cancer comprising the transcription factor gene fusion.[000165] An exemplary method for detecting translocation renal cell carcinoma (tRCC) in a liquid biopsy sample (e.g., blood or plasma) obtained from a subject comprises:(a) isolating cell-free DNA (cfDNA) from the sample;(b) enriching the cfDNA with an antibody that specifically binds histone H3 protein acetylated at the lysine at residue 27 (H3K27ac) for H3K27ac-associated DNA;(c) sequencing the enriched cfDNA obtained in step (b) to obtain a plurality of sequencing reads; and(d) identifying within the plurality of sequencing reads a pre-determined set of nucleotide sequences comprising gene regulatory elements specifically associated with tRCC.[000166] An exemplary method for detecting a prostate cancer comprising a transcription factor gene fusion in a liquid biopsy sample (e.g., blood or plasma) obtained from a subject comprises:(a) isolating cell-free DNA (cfDNA) from the sample;(b) enriching the cfDNA with an antibody that specifically binds histone H3 protein acetylated at the lysine at residue 27 (H3K27ac) for H3K27ac-associated DNA;(c) sequencing the enriched cfDNA obtained in step (b) to obtain a plurality of sequencing reads; and(d) identifying within the plurality of sequencing reads a pre-determined set of nucleotide sequences comprising gene regulatory elements specifically associated with a prostate cancer comprising a transcription factor gene fusion.[000167] An exemplary method for detecting translocation renal cell carcinoma (tRCC) in a liquid biopsy sample (e.g., blood or plasma) obtained from a subject comprises:(a) isolating cell-free DNA (cfDNA) from the sample;(b) enriching the cfDNA with an antibody that specifically binds histone H3 protein methylated at the lysine at residue 4 (H3K4me3) for H3K4me3 -associated DNA;(c) sequencing the enriched cfDNA obtained in step (b) to obtain a plurality of sequencing reads; and(d) identifying within the plurality of sequencing reads a pre-determined set of nucleotide sequences comprising gene regulatory elements specifically associated with tRCC.[000168] An exemplary method for detecting a prostate cancer comprising a transcription factor gene fusion in a liquid biopsy sample (e.g., blood or plasma) obtained from a subject comprises:(a) isolating cell-free DNA (cfDNA) from the sample;(b) enriching the cfDNA with an antibody that specifically binds histone H3 protein methylated at the lysine at residue 4 (H3K4me3) for H3K4me3 -associated DNA;(c) sequencing the enriched cfDNA obtained in step (b) to obtain a plurality of sequencing reads; and(d) identifying within the plurality of sequencing reads a pre-determined set of nucleotide sequences comprising gene regulatory elements specifically associated with a prostate cancer comprising a transcription factor gene fusion.[000169] Typically, the method further comprises determining signal density and area under the curve (AUC) for peaks of the sequencing reads that correspond to the identified nucleotide sequences comprising gene regulatory elements specifically associated with a cancer comprising a transcription factor gene fusion (e.g., tRCC) and comparing the values obtained for the signal density and area under the curve to pre-determined threshold values for signal density and area under the curve. The pre-determined threshold values can be obtained by comparing corresponding signal density and area under the curve (AUC) values from a plurality of samples that are confirmed to be a cancer comprising a transcription factor gene fusion (e.g., tRCC) and a plurality of control samples. The control samples can be samples obtained from healthy subjects, or from a subject suffering from a cancer that does not comprise the transcription factor gene fusion (e.g., clear cell renal cell carcinoma (ccRCC). A cancer comprising a transcription factor gene fusion (e.g., tRCC) is identified when the signal density and area under the curve of the identified nucleotide sequences comprising gene regulatory elements specifically associated with the cancer comprising the transcription factor gene fusion are above the pre-determined threshold values.Applications[000170] The disclosed methods for diagnosing a cancer comprising a transcription factor gene fusion (e.g., tRCC) may be used in subjects that are suspected of suffering from a cancer comprising a transcription factor gene fusion after an initial diagnosis of cancer (e.g., renal cell carcinoma (RCC). Confirming whether a subject suffering from RCC is suffering from tRCC is useful because not all therapies effective against certain types of RCC such clear cell RCC (ccRCC) are effective in subjects suffering from tRCC due to differences in the mutational makeup of tRCC relative to ccRCC.[000171] The methods described herein can be used to monitor disease burden in patients with a cancer comprising a transcription factor gene fusion (e.g., tRCC) even in the absence of DNA alterations detectable by conventional methods and at very low tumor fraction. Furthermore, the methods disclosed herein are capable of discriminating a cancer comprising a transcription factor gene fusion (e.g., tRCC samples from other RCC and healthy samples) even in cases of a low tumor fraction in cfDNA. Moreover, the disclosed methods are more sensitive than conventional methods of detecting a cancer comprising a transcription factor gene fusion (e.g., tRCC), especially when a patient’s tumor burden is low. For example, using the disclosed methods, it is possible to detect an increase in one or more nucleotide sequences comprising gene regulatory elements specifically associated with the cancer comprising the transcription factor gene fusion (e.g., tRCC) in cfDNA enriched for H3K27ac-associated DNA and / or H3K4me3 -associated DNA, in advance of clinical progression or tumor detection by standard imaging technologies.[000172] Accordingly, the methods disclosed herein may be used to support the diagnosis of a cancer comprising a transcription factor gene fusion (e.g., tRCC), and due to their sensitivity and specificity may allow an earlier diagnosis and / or detection. The disclosed methods may also be used for the monitoring of disease burden, allow earlier detection of disease progression and more sensitive and accurate monitoring of treatment efficacy to guide therapy selection.Monitoring cancer progression and / or treatment response[000173] The present methods are useful for the monitoring and / or detection of a cancer comprising a transcription factor gene fusion (e.g., tRCC) including new occurrences or recurrences. In particular, the present methods allow for early detection of new occurrences or recurrences of cancer from a liquid biopsy without the need to identify a new or recurrent tumor location.[000174] Provided herein is a method for monitoring the tumor progression in a subject suffering from a cancer comprising a transcription factor gene fusion (e.g., tRCC) which comprises measuring, from a first sample obtained from the subject at a first time point, the presence and / or level of one or more nucleotide sequences comprising gene regulatory elements specifically associated with the cancer comprising the transcription factor gene fusion; measuring, from a second sample obtained from the subject at a second or subsequent time point, the presence and / or level of the one or more nucleotide sequences comprising gene regulatory elements specifically associated with the cancer comprising the transcription factor gene fusion; and comparing the measurements from the first and second or subsequent time points, thereby monitoring the tumor progression. The measuring steps are carried out according to a method described herein.[000175] Any difference in the level of the one or more nucleotide sequences comprising gene regulatory elements specifically associated with the cancer comprising the transcription factor gene fusion (e.g., tRCC) between the first time point to the second or subsequent (or between subsequent time points can indicate the progression of the cancer comprising the transcription factor gene fusion (e.g., tRCC). For example, an increased level between the first and second or subsequent time point may indicate tumor progression e.g., tumor growth, and a decreased level between the first and subsequent time point may indicate that the subject is in remission e.g., tumor shrinkage.[000176] The first time point can be before the start of treatment for the cancer comprising the transcription factor gene fusion and the second or subsequent time point can be during or after treatment for the cancer comprising the transcription factor gene fusion. The first time pointcan be after treatment for the cancer comprising the transcription factor gene fusion , and the second time point can be up to one month, 2 months, 3 months, 4 months, 5 months, 6 months, 9 months, 12 months or more, 1 year, 2 years, 3 years, 4 years, 5 years, 10 years or more thereafter. The threshold for detecting the increase in the level of the nucleotide sequence comprising gene regulatory elements specifically associated with the cancer comprising the transcription factor gene fusion may be a baseline level determined by the first time point.[000177] An increase in the level of the one or more nucleotide sequences comprising gene regulatory elements specifically associated with a cancer comprising a transcription factor gene fusion (e.g., tRCC) from the first time point to the second or subsequent time points indicates that the therapeutic agent is ineffective for the treatment of the cancer comprising the transcription factor gene fusion. Accordingly, the treatment with the therapeutic agent may be discontinued and / or the subject may receive a different therapy.[000178] Accordingly, also provided herein is a method for monitoring a response to cancer therapy in a subject, the method comprising: measuring, from a sample derived from a subject at a first time point prior to the cancer therapy, the presence and / or level of one or more nucleotide sequences comprising gene regulatory elements specifically associated with a cancer comprising a transcription factor gene fusion (e.g., tRCC); measuring, from a sample derived from a subject at a second time point after the onset of cancer therapy, the presence and / or level of the one or more nucleotide sequences comprising gene regulatory elements specifically associated with the cancer comprising the transcription factor gene fusion; and comparing the measurements from the first and second time points, thereby monitoring the response to the cancer therapy. Typically, both of the measuring steps are carried out according to a method described herein.[000179] Also provided herein is a method for monitoring the effectiveness of a therapy in a subject with a cancer comprising a transcription factor gene fusion (e.g., tRCC) which comprises (a) providing a first sample obtained from the subject at a first time point; (b) isolating cell-free DNA (cfDNA) from the first sample; (c) enriching the cfDNA with a means for specifically binding histone H3 protein acetylated at the lysine at residue 27 (H3K27ac) for H3K27ac-associated DNA and / or with a means for specifically binding histone H3 proteinmethylated at the lysine at residue 4 (H3K4me3) for H3K4me3 -associated DNA; and (d) determining levels of one or more nucleotide sequences comprising gene regulatory elements specifically associated with the cancer comprising the transcription factor gene fusion in the cfDNA enriched for H3K27ac-associated DNA and / or H3K4me3 -associated DNA; (e) providing a second or subsequent sample obtained from the subject and repeating steps (b)-(d) with the second sample; and (f) comparing the levels of the one or more nucleotide sequences obtained in steps (d) and (e), wherein an increase in the level of the one or more nucleotide sequences determined in step (e) relative to step (d) indicates that the therapy is not effective against the subject’s cancer comprising the transcription factor gene fusion and no change or a decrease in the level of the one or more nucleotide sequences determined in step (e) relative to step (d) indicates that the therapy is effective against the subject’s cancer comprising the transcription factor gene fusion.[000180] Any difference in the level of the one or more nucleotide sequences comprising gene regulatory elements specifically associated with the cancer comprising the transcription factor gene fusion (e.g., tRCC) between the first time point to the second or subsequent (or between subsequent time points can indicate the response to therapy and its effectiveness. For example, an increased level between the first and second or subsequent time point may indicate tumor progression and that the therapeutic agent is ineffective for the treatment of the cancer comprising the transcription factor gene fusion, and a decreased level between the first and subsequent time point may indicate tumor regression and that the therapeutic agent is effective for the treatment of the cancer comprising the transcription factor gene fusion. If a therapeutic agent is shown to be ineffective, the therapeutic agent may be discontinued and / or the subject may receive a different therapy.[000181] The first time point can be before the start of the therapy and the second or subsequent time point can be during or after the therapy. Alternatively, the first time point can be after the start of the therapy, and the second time point can be up to one month, 2 months, 3 months, 4 months, 5 months, 6 months, 9 months, 12 months or more, 1 year, 2 years, 3 years, 4 years, 5 years, 10 years or more thereafter, e.g. while the therapy is ongoing. The threshold for detecting the increase in the level of the nucleotide sequence comprising gene regulatoryelements specifically associated with the cancer comprising the transcription factor gene fusion present may be a baseline level determined by the first time point.Classification of cancers by subtype[000182] Also provided herein is a method for classifying a cancer subtype which comprises detecting, from a sample derived from a subject suffering from a cancer (e.g., RCC), the presence and / or level of one or more nucleotide sequences comprising gene regulatory elements specifically associated with the cancer comprising the transcription factor gene fusion (e.g., nucleotide sequences associated with tRCC), and identifying the type or subtype of the cancer as the cancer comprising the transcription factor gene fusion based on the presence of the one or more nucleotides sequences comprising gene regulatory elements specifically associated with the cancer comprising the transcription factor gene fusion (e.g., tRCC), wherein the detecting step is carried out according to a method disclosed herein.[000183] For example, RCCs can be classified by subtype. RCC can include clear cell RCC (ccRCC), papillary RCC, chromophobe RCC, collecting duct carcinoma, renal medullary carcinoma (RMC), renal sarcoma, nephroblastoma, transitional cell carcinoma, and tRCC. The methods disclosed herein can be used to differentiate between tRCCs and non-tRCCs.[000184] Similarly, prostate cancer can be classified by whether or not the prostate cancer comprises a TMPRSS2-ERG gene fusion. The methods disclosed herein can be used to differentiate between a prostate cancer that comprises a TMPRSS2-ERG gene fusion and other common subtypes of prostate cancer that do not.Selecting therapy and treatment[000185] While therapies developed for ccRCC are often deployed off label to subjects suffering from tRCC, these therapies typically have lower response rates in tRCC, owing to its distinct biology. In addition, certain therapies such as treatment with a Hypoxia-Inducible Factor 2a (HIF2a) inhibitor such as belzutifan have limited mechanistic rationale in tRCC, given that tRCC does not harbor alterations in the von Hippel-Lindau (VHL) gene, and therefore may be unsuitable as a therapy for tRCC. It is therefore beneficial to limit therapeutic interventions in subjects suffering from tRCC to treatments that are likely to be effective.[000186] Accordingly, a method of selecting a therapy for a subject suffering from a cancer (e.g., tRCC) is provided that comprises: (a) isolating cell-free DNA (cfDNA) from a sample obtained from the subject; (b) enriching the cfDNA with a means for specifically binding histone H3 protein acetylated at the lysine at residue 27 (H3K27ac) for H3K27ac-associated DNA and / or with a means for specifically binding histone H3 protein methylated at the lysine at residue 4 (H3K4me3) for H3K4me3 -associated DNA; and (c) detecting the presence (and / or level) of one or more nucleotide sequences comprising gene regulatory elements specifically associated with the cancer comprising the transcription factor gene fusion in the cfDNA enriched for H3K27ac- associated DNA and / or H3K4me3 -associated DNA, wherein the presence (or level) of the one or more nucleotide sequences comprising gene regulatory elements specifically associated with the cancer comprising the transcription factor gene fusion in the enriched cfDNA indicates that the subject’s cancer is the cancer comprising the transcription factor gene fusion; and (d) selecting a therapy that is effective against the cancer comprising the transcription factor gene fusion.[000187] For example, the presence or level of the one or more nucleotide sequences comprising gene regulatory elements specifically associated with tRCC in the enriched cfDNA may indicate that therapy with a Hypoxia-Inducible Factor 2a (HIF2a) inhibitor is unsuitable or unlikely to be effective against the subject’s RCC; and lead to the selection of a therapy that has a greater likelihood of being effective against the subject’s RCC (e.g., an alternative therapy that does not include treatment with a HIF2a inhibitor). The alternative therapy may be selected from any known therapies (such as the exemplary therapies described herein) that may have a greater likelihood of being effective against the subject’s tRCC than HIF2a inhibitors.[000188] Furthermore, a method of treating a cancer (e.g., tRCC) in a subject is provided that comprises: (a) isolating cell-free DNA (cfDNA) from a sample obtained from the subject; (b) enriching the cfDNA with a means for specifically binding histone H3 protein acetylated at the lysine at residue 27 (H3K27ac) for H3K27ac-associated DNA and / or with a means for specifically binding histone H3 protein methylated at the lysine at residue 4 (H3K4me3) for H3K4me3 -associated DNA; and (c) detecting the presence (or level) of one or more nucleotide sequences comprising gene regulatory elements specifically associated with the cancercomprising the transcription factor gene fusion in the cfDNA enriched for H3K27ac-associated DNA and / or H3K4me3-associated DNA, wherein the presence (or level) of the one or more nucleotide sequences comprising gene regulatory elements specifically associated with the cancer comprising the transcription factor gene fusion in the enriched cfDNA is indicative of the subject suffering from the cancer comprising the transcription factor gene fusion and (d) administering to the subject a therapeutically effective amount of a therapeutic agent effective in the treatment of the cancer comprising the transcription factor gene fusion.[000189] Certain therapies have been shown to provide a therapeutic benefit in tRCC. A suitable therapeutic agent for use in the methods disclosed herein may be a tyrosine kinase inhibitor (TKI). Exemplary TKIs that can be used for the methods disclosed include sunitinib, cabozantinib, lenvatinib, tivozanib, axitinib, and pazopanib. Alternatively, or additionally, a suitable therapeutic agent can be an immunotherapeutic agent. Exemplary immunotherapeutic agents include pembrolizumab, ipilimumab, and nivolumab.[000190] A combination of therapeutic agents for use in the methods disclosed herein may be suitable. For example, a subject may be administered one or more TKIs and / or one or more immunotherapeutic agents (e.g., nivolumab and ipilimumab). Combinatorial therapies may be administered simultaneously or sequentially.EXAMPLES[000191] Although methods and materials similar or equivalent to those described herein can be used, suitable methods and materials are described below. The following examples are included for illustrative purposes only and are not intended to be limiting.Example 1. Determining an epigenomic signature for tRCC[000192] Generating epigenomic profiles for cancer subtyping cancer was previously described in Baca, et al. supra. An epigenomic signature driven by the TFE3 gene fusion was generated by profiling 25 epigenomic libraries from 4 tRCC and 6 ccRCC cell lines using chromatin immunoprecipitation sequencing (ChlP-seq) for post-translational histone modifications (H3K27ac and H3K4me3) and methylated CpG dinucleotide sequencing (MeDIP- seq) on each cell line. H3K27ac ChlP-seq captures active promoters and enhancers, whileH3K4me3 is associated with active promoters. MeDIP-seq captures differentially methylated DNA regions. Across all 10 RCC cell lines, a total of 119,504 and about 50,000 regulatory elements (REs) were captured using H3K27ac and H3K4me3 ChlP-seq, respectively, while a total of 427,458 methylated regions were captured using MeDIP-seq.[000193] To identify TFE3 binding sites, CUT&RUN was performed on the cell line UOK109 (comprising a NONO-TFE3 fusion) using the CUT&RUN Assay Kit (Cell Signaling) following manufacturer protocol. Briefly, for each reaction 250,000 cells were immobilized on activated Concanavalin A-coated magnetic beads. Immobilized cells were then permeabilized prior to an overnight incubation at 4°C with anti-TFE3 antibody (ZRB1272) in binding buffer. The complex was then digested by Protein AG-MNase, and DNA fragment extracted. Libraries were prepared with KAPA Hyper Prep Kit (Roche cat# 07962363001) following the manufacturer's instructions, starting from 0.5 ng of DNA. After library amplification, the DNA was purified by KAPA HyperPure Beads (Roche cat# 09006583001). The size distribution of the purified libraries was examined using Agilent 4200 TapeStation with a high sensitivity DNA Chip (Agilent, cat# 5067-5584). The library was submitted for the 150 base-pair paired end sequencing on an Illumina NovaSeq6000 system (Novogene Corporation, CA).Peak Calling[000194] MeDIP-seq, ChlP-seq, and CUT&RUN peak calling were performed using the ChiLin computational pipeline that automates the quality control and data analyses of ChlP-seq, with the parameters -p narrow and -r histone and hg38 as reference genome. For differential peak analysis between tRCC and ccRCC cell lines, a consensus peaks file was generated using DiffBind (v. 3.10.1) and a DeSeq2-based differential peak calling was performed (log2FC >1 or < -1 and FDR<0.05). Peak calling was used to identify gene regulatory elements specifically associated with tRCC.[000195] Inferred regulatory elements activity at sites of interest based on H3K27ac and H3K4me3 were determined as follows. Sites of width <8 kb (± 4kb from the center of a peak) were included and then resized to a 6-kb interval centered on the original peak (± 3kb from the center of the peak). Peaks were separated into 40-bp windows, and fragment counts wereaggregated across a given window for all peaks to obtain aggregate profiles for a sample. Subsequently, there were two normalization steps. First, a ‘shoulder normalization’ step to account for variation in background signal across samples. The region between -3,000, -2,800 bp and 2,800, 3,000 bp around the center of each transcription factor binding site (TFBS) and aggregated counts at these sites for each sample was considered. This value was subtracted from the aggregate counts to set the ‘shoulder’ of peaks to zero. Following the shoulder normalization step, signals in each bin were normalized to the aggregated signal at the common 10,000 DNase hypersensitivity sites that are expected to be active across most tissue types and defined across the largest number of samples in Meuleman W, et al. (Nature. 2020;584(7820):244-251).[000196] Principal component analyses (PC A) and unsupervised hierarchical clustering of H3K27ac and H3K4me3 peaks revealed distinct segregation of sequence read peaks associated with tRCC and ccRCC cell lines (FIG. 1). As shown in FIG. 1, eight cell lines were used for the analysis comprising four ccRCC cell lines (CAK1, A498, 7860 and RFX393; depicted by dark grey filled circles) and four tRCC cell lines (sTFE, FUUR1, UOK146 and UOK109; depicted by light grey filled circles). As illustrated by the dark grey and light grey circles, there is no overlap between the sequence reads for ccRCC and tRCC cell lines, respectively. Active histone marks reflect inter alia differences driven by the TFE3 gene fusion in tRCC. Analysis of differential sequence read peaks between tRCC and ccRCC cell lines identified 11,435 H3K27ac and 148 MeDIP sequencing read peaks that were selectively enriched in tRCC (q < 0.05, FIG. 2A-B). As shown in FIG. 2, each filled circle represents a peak (A: H3K27ac; B: MeDIP) that is upregulated in ccRCC (dark grey) and tRCC (light grey), respectively. FIG. 2A illustrates the differential H3K27ac sequence read peaks found in tRCC compared to ccRCC. FIG. 2B illustrates the differential MeDIP sequence read peaks found in tRCC compared to ccRCC. Both FIG. 2A and FIG. 2B highlight the clear differences for distinguishing tRCC from ccRCC. FIG. 3 illustrates upregulated peaks between tRCC cell lines (light grey bar) and ccRCC cell lines (dark grey bar). FIG. 3A (H3K27ac) and FIG. 3B (MeDIP) similarly illustrates the clear differences in upregulated sequence read peaks between tRCC cell lines (light grey bar) and ccRCC cell lines (dark grey bar). These differential peaks provide an epigenomic signature of nucleotidesequences comprising gene regulatory elements specifically associated with tRCC that can be used for diagnosing and / or monitoring tRCC.[000197] Comparison of peaks identified from tRCC and ccRCC cell lines identified the nucleotide sequences set forth in SEQ ID NOs: 1 - 11 ,435 as gene regulatory elements specifically associated with tRCC that do not overlap with the nucleotide sequences associated with TFE3 gene fusions identified herein (which include transcription factor binding sites bound by neomorphic transcription factors resulting from such gene fusions).Intersection between GTRD and TFE3 gene fusions[000198] Known ' / ' / ' / G-nai've transcription factor binding sites (TFBS) were downloaded from the GTRD database (Yevshin el al. supra). This database contains a compilation of ChlP- seq data from various sources. The 24,002 peaks available in GTRD database were overlapped with the 25,979 peaks identified by CUT&RUN in UOK109 to determine a set of 5,128 TFBS for which the fusion- TFE3 has a particular affinity. Overlap of peaks was assessed using BEDTools v2.27.1. Peaks were considered overlapping if they shared one or more base pairs. The intersection between the GTRD database and the peaks identified herein by CUT&RUN in UOK109 resulted in the identification of the nucleotide sequences set forth in SEQ ID NOs: 11,436-16,563 as nucleotide sequences associated with TFE3 gene fusions. These nucleotide sequences represent transcription factor binding sites that are specifically associated with neomorphic transcription factors caused by TFE3 gene fusions.Example 2. Detection of epigenomic signatures in cjDNA isolated from plasma[000199] The profile of 149 epigenomic libraries from 51 plasma samples from patients with tRCC, ccRCC, and healthy individuals was determined.[000200] Plasma samples were collected from patients with tRCC and ccRCC diagnosed and treated at the Dana-Farber Cancer Institute (DFCI) between 2005 and 2022. All patients provided written informed consent. The use of samples was approved by the DFCI (01-045 and 09-171) IRB protocols. Studies were conducted in accordance with recognized ethical guidelines.[000201] Plasma samples from healthy individuals without a history of diabetes, cancer, or major medical illnesses were obtained from the Mass General Brigham Biobank. Written informed consent was obtained from all healthy donors, and sample collection was approved by the Brigham and Women’s Hospital IRB (2009P002312), following ethical regulations. cfDNA processing[000202] cfDNA samples were processed by the following method. Peripheral blood was collected in EDTA Vacutainer tubes (BD) and processed within 3 hours of collection. Plasma was separated by centrifugation at 2,500 x g for 10 minutes, transferred to microcentrifuge tubes, and centrifuged at 2,500 x g at room temperature for 10 minutes to remove cellular debris. The supernatant was aliquoted into 1 to 2 mb aliquots and stored at -80°C until DNA extraction. cfDNA was isolated from 1 mL of plasma using the QIAGEN Circulating Nucleic Acids Kit (QIAGEN), eluted in AE buffer, and stored at -80°C.Tumor fraction calculation[000203] Low-pass whole-genome sequencing (LPWGS) was performed on all cfDNA samples. The ichorCNA R package was used to infer copy-number profiles and cfDNA tumor (ctDNA) fraction from read abundance across bins spanning the genome using default parameters from Adalsteinsson, V. A. et al. Nat. Commun. 8, 1324 (2017)Chromatin immunoprecipitation and sequencing[000204] The method used for cfChlP-seq was previously published in Baca, S. C. et al. (supra). 1 pg of anti-histone antibody was coupled with 10 pL protein A (Invitrogen, cat# 10002D) and 10 pL protein G (Invitrogen, cat# 10004D) for at least 6 hours at 4°C with rotation in 0.5% BSA (Jackson Immunology, cat# 001-000-161) in PBS (Gibco, cat# 14190250), followed by blocking with 1% BSA in PBS for 1 hour at 4°C with rotation. The following anti- histone antibodies were used for immunoprecipitation: H3K4me3, ThermoFisher # PA5-27029; H3K27ac, Abeam # ab4729.[000205] Thawed plasma was centrifuged at 3,000 x g for 15 minutes at 4°C. The supernatant was precleared with the magnetic beads with 20 pL protein A and 20 pL protein Gfor 2 hours at 4°C. The precleared and conditioned plasma samples were subjected to antibody- coupled magnetic beads overnight with rotation at 4°C. The reclaimed magnetic beads were washed with 1 mb of each washing buffer twice. Three washing buffers were used in following order: low salt washing buffer (0.1% SDS, 1% Triton X-100, 2 mM EDTA, 150mM NaCl, 20 mM Tns-HCl pH 7.5), high salt buffer (0.1 % SDS, 1 % Triton X-100, 2 mM EDTA, 500 mM NaCl, 20 mM Tns-HCl pH 7.5), and LiCl washing buffer (250 mM LiCl, 1%NP-4O, 1% Na Deoxy cholate, 1 mM EDTA, 10 mM Tris-HCl pH 7.5). Subsequently, the beads were rinsed with TE buffer (Fisher Sci, cat# BP2473500), and resuspended and incubated in 100 pL of DNA extraction buffer containing 0.1 M NaHCO3, 1% SDS and 0.6 mg / mL Proteinase K (Qiagen, cat#19131) and 0.4 mg / mL RNaseA (ThermoFisher, cat#12091021) for 10 minutes at 37°C, for 1 hour at 50°C, and for 90 minutes at 65°C. DNA was purified through Phenol extraction (Invitrogen, cat# 15593031) and Ethanol precipitation was performed with 3M sodium acetate (NaOAc) (Ambion, cat# AM9740) and glycogen (Ambion, cat# AM9510).[000206] CfChlP-seq libraries were prepared with ThruPLEX DNA-Seq Kit (Takara Bio, cat# R400675) following the manufacturer's instructions. After library amplification, the DNA was purified by AMPure XP (Beckman coulter, cat# A63880). The size distribution of the purified libraries was examined using Agilent 2100 Bioanalyzer with a high sensitivity DNA Chip (Agilent, cat# 5067-4626). The library was submitted for the 150 base-pair paired end sequencing on an Illumina NovaSeq6000 system (Novogene Corporation).CfMeDIP-seg Assay[000207] cfMeDIP-seq was performed on plasma samples using previously published methods ofBaca S. C. etal. (supra). cfDNA library preparation was performed on 10 ng of DNA using the KAPA HyperPrep Kit (KAPA Biosystems) according to the manufacturer's protocol. End-repair, A-tailing, and ligation of NEBNext adaptors was then performed (NEBNext Multiplex Oligos for Illumina kit, New England BioLabs). Libraries were digested using the USER enzyme (New England BioLabs). A. DNA, consisting of unmethylated and in vitro methylated DNA, was added to prepared libraries to achieve a total amount of 100 ng DNA. Methylated and unmethylated Arabidopsis thaliana DNA (Diagenode) was added for quality control.[000208] MeDIP was performed using the MagMeDIP Kit (Diagenode) following the manufacturer's protocol. Samples were purified using the iPure Kit v2 (Diagenode). Success of the immunoprecipitation was confirmed using qPCR to detect recovery of the spiked-in Arabidopsis thaliana methylated and unmethylated DNA. KAPA HiFi Hotstart Ready Mix (KAPA Biosystems) and NEBNext Multiplex Oligos for Illumina (New England Biolabs) were added to a final concentration of 0.3 pmol / L.[000209] Libraries were amplified as follows: activation at 95 °C for 3 minutes, amplification cycles at 98°C for 20 seconds, 65°C for 15 seconds, 72°C for 30 seconds, and a final extension at 72°C for 1 minute. Samples were pooled and sequenced (Novogene Corporation) on Illumina HiSeq 4000 to generate 150 bp paired-end reads.Cell-free DNA sequencing data processing[000210] CfChlP-seq / cfMeDIP-seq reads were aligned to the hgl9 human genome build using Burrows-Wheeler Aligner version 0.7.1740. Non-uniquely mapping and redundant reads were discarded. MACS version 2.1.1.2014061641 was used for ChlP-seq peak calling with a q value (false discovery rate (FDR)) threshold of 0.01. Fragment locations were converted to BED files using BEDTools (version 2.29.2) bamtobed with the -bedpe flag set.[000211] For analyses involving overlap with genomic regions such as differentially marked regions or TFE3 binding sites, fragments were imported as GRanges objects and collapsed to 1 bp at the center of the fragment location to ensure that a fragment can map to only one site.Identifying epigenomic signatures in cfChlP-seq and cfMeEP-seq reads[000212] Two epigenomic signatures, derived from peak calling and the intersection between GTRD and TFE3 gene fusions as described above in Example 1, both successfully identified tRCC patient plasma samples, as distinct from ccRCC and healthy plasma samples, with a high degree of specificity and sensitivity. Combining both epigenomic profiles also provided a robust method to identify tRCC from plasma samples.[000213] 11,435 tRCC-specific H3K27ac-associated regulatory elements (SEQ ID NOs 1-11,435), as determined from cell line epigenomic profiling (see Example 1), were used as the first epigenomic signature. Aggregated coverage at these sites in the H3K27ac cfChIP revealed an increase in signal in plasma from patients with tRCC as compared with patients with ccRCC (p=0.0029) and with healthy patients (p<10‘5) (FIG. 4). As shown in FIG. 4A, there was no overlap between the tRCC samples and healthy patient samples and only a small overlap between tRCC and ccRCC.[000214] To refine this discrimination, a second epigenomic signature comprising 5,128 H3K27ac-associated TFE3 genomic binding sites (SEQ ID NOs: 11,436-16,563) was identified based on tRCC cell line profiling and integration with public data as set out above in Example 1. TFE3 gene fusion binding sites showed increased discrimination versus ccRCC (p=0.0005) but was less discriminating versus healthy patients (p=1.76e-5). This may be due to TFE3 activity in white blood cells, particularly macrophages, in which MiT / TFE genes are known to be active (FIG. 4B).[000215] The signal density threshold was set at > 7.1 for the 11,435 tRCC-specific H3K27ac-associated regulatory elements. For 5,128 H3K27ac-associated TFE3 genomic binding sites, the signal density threshold was set at > 5.9. For both data sets together, the signal density threshold was set at > 13. An area under the receiver operating characteristic curve (AUROC) for the ability of all tRCC-specific H3K27ac sites (diff-K27ac sites) or H3K27ac sites at TFE3 binding sites (diff-77'7G / K27ac sites) to discriminate tRCC from ccRCC and healthy plasma was calculated. This provided excellent performance for both diff-K27ac (0.80) and diff- T E3i 27ac sites (0.87) in discriminating tRCC from ccRCC. The integration of both scores increased the AUC to 0.88 (FIG. 5 A). For discriminating tRCC from healthy plasma, AUCs were 0.99, 0.925, and 0.98, respectively (FIG. 5B).[000216] The methods as described above were able to discriminate tRCC samples even in cases of low calculated tumor fraction (as determined by ichorCNA).Example 3. Monitoring tRCC tumor progression[000217] tRCC tumor progression was monitored in three patients via cfH3K27ac ChlP- seq as described in Example 2. Plasma was collected at multiple time-points during treatment, indicated in weeks on the x axis of the graphs in FIG. 6. The results of ichorCNA (to determine the tumor fraction in the cfDNA) and ChlP-Seq (to determine the presence and level of nucleotide sequences comprising gene regulatory elements specifically associated with tRCC) are shown in FIG. 6. In all three patients, there was an observed increase in gene regulatory elements specifically associated with tRCC, which preceded clinical progression by weeks to months.[000218] The top panel in FIG. 6 shows tumor fraction determined by ichorCNA as described in Example 2. The threshold cut-off for healthy plasma vs tRCC plasma is shown as a shaded grey box with a dashed upper border. The dashed upper border represents the cut off indicating the presence of tRCC. ichorCNA values above this threshold are indicative of the presence of tRCC and / or tumor progression.[000219] The middle panel of FIG. 6 shows the signal density for tRCC-specific H3K27ac- associated regulatory elements (H3K27ac tRCC up peaks) described in Example 2. The grey box bounded by dashed lines indicates the signal density threshold minimum and maximum for healthy plasma, set at a cut-off whereby a signal density above the threshold maximum is indicative of the presence of tRCC and / or tumor progression.[000220] The bottom panel of FIG. 6 shows the signal density for tRCC-specific H3K27ac- associated TFE3 genomic binding sites (H3K27ac TFE3) described in Example 2. The grey box bordered by dashed lines indicates the signal density threshold minimum and maximum for healthy plasma, set at a cut-off whereby a signal density above the threshold maximum is indicative of the presence of tRCC and / or tumor progression.[000221] Notably, in all three cases (FIG. 6, left / middle / right), an increase in gene regulatory elements specifically associated with tRCC (tRCC-specific H3K27ac-associated regulatory elements and non-overlapping tRCC-specific H3K27ac-associated TFE3 genomic binding sites) was apparent even when the tumor fraction was undetectable using knownmethods that detect tRCC tumors by large-scale copy number alterations. This demonstrates the increased sensitivity of the methods described herein in detecting the presence of and monitoring the progression of tRCC.Example 4. Optimizing Therapy Selection[000222] Accurate detection of tRCC can be used in optimal therapy selection. While therapies developed for ccRCC are often deployed off-label to patients with tRCC, these therapies typically have lower response rates in tRCC owing to its distinct biology. In addition, certain therapies such as belzutifan (HIF2a inhibitor) have limited mechanistic rationale in tRCC, given that tRCC does not harbor alterations in VHL.[000223] In the left-hand panel, the subject was originally treated with nivolumab (“nivo” in FIG. 6). Treatment was changed to cabozantinib (“cabo” in FIG. 6) as indicated by the vertical line shortly after week 12. Treatment was changed again after week 48, when tumor progression was detected.[000224] As shown in the left-hand panel of FIG. 6, after administration of nivolumab and cabozantinib, the level of gene regulatory elements specifically associated with tRCC increased. Following a change of treatment regimen (commenced at the time point indicated by the arrow labeled “progression”) to the combination of sunitinib and gemcitabine, clinical and molecular signs of tRCC (including the level of gene regulatory elements specifically associated with tRCC) diminished, showing the efficacy of this treatment protocol for this patient.[000225] In the right-hand panel, the subject was originally treated with sunitinib and gemcitabine (“Sunit + Gem” in FIG. 6). Treatment was changed to cabozantinib, nivolumab, and ipilimumab (“cabo nivo ipi” in FIG. 6) as indicated by the vertical line shortly at approximately week 18. Tumor progression was detected after week 24 by ichorCNA. It was preceded by an increase in gene regulatory elements specifically associated with tRCC (both tRCC-specific H3K27ac-associated regulatory elements and non-overlapping tRCC-specific H3K27ac- associated TFE3 genomic binding sites).[000226] This example demonstrates the suitability of the methods described herein for monitoring tumor progression in subjects suffering from tRCC and for assisting in therapeuticdecision making such as selecting an effective treatment or ceasing a treatment that is ineffective, without having to wait for clinical disease progression.Example 5. Refinement of epigenomic signatures for tRCC identification[000227] An epigenomic signature driven by the TFE3 gene fusion was generated by reanalyzing the 25 epigenomic libraries described in Example 1 and two additional epigenomic libraries using the same profiling method. H3K27ac and TFE3 ChlP-seq data for tRCC cell lines that had previously been generated in-house were obtained from GEO under accession number GSE266530. H3K4me3 and H3K27ac ChlP-seq for ccRCC cell lines (Caki-1, A-498, RFX393, and 786-0) were obtained from GEO under accession number GSE143653 (Gopi and Kidder, Nat Commun. 2021 Mar 3; 12(1): 1419).[000228] H3K4me3 ChlP-seq was performed on three tRCC cell lines (UOK146, s-TFE, and FU-UR-1). Briefly, 3x106cells per reaction were collected and crosslinked with 1% formaldehyde. The crosslinking reaction was quenched with 0.125 M glycine. After washing with ice-cold PBS, the pellet was resuspended in 130 pL SDS lysis buffer. The lysate was sonicated to 200-500 bp (Covaris E220 sonicator). Following sample precipitation and centrifugation, 100 pL of each sample were diluted 10-fold in ChIP dilution buffer and incubated with protein A and Dynabeads protein G (1: 1; 10 pL each) and 10 pL of H3K4me3 antibody (Rabbit mAb #9751, Cell Signaling, RRID: AB 2616028). Following an overnight incubation at 4°C, antibody-bound DNA was washed with Low Salt Immune Complex Wash Buffer for 5 minutes, High Salt Immune Complex Wash for 5 minutes, LiCl Immune Complex Wash Buffer for 5 minutes followed by two washes in TE buffer. ChIP DNA was reverse cross-linked and purified for DNA library construction using the KAPA HyperPrep Kit (KAPA Biosystems, KR0961).[000229] MeDIP-seq was performed on eight cell lines (UOK146, s-TFE, FU-UR-1, Caki- 1, A-498, 786-0, 786-P, and KMRC-1) by adapting the method described by Nuzzo et al., Nat Med. 2020 Jul;26(7): 1041-1043.Peak calling[000230] Peak calling was performed as described in Example 1. Differentially marked sites were deemed significant if the FDR q-value was <0.01. Sites exhibiting a logifold change in either direction greater than 1 for H3K27ac and greater than 2 for H3K4me3 were focused on. Across 10 RCC cell lines (i.e., 4 tRCC, 6 ccRCC), a median of 29,588 peaks (range 25,314- 30,076) were captured by H3K4me3 ChlP-seq, a median of 57,226 peaks (44,470-73,400) by H3K27ac ChlP-seq, and a median of 229,624 peaks (111,904-297,344) by MeDIP-seq.[000231] Next, inferred regulatory elements activity at sites of interest based on H3K27ac and H3K4me3 were determined. Unsupervised hierarchical clustering and PCA of the H3K27ac and H3K4me3 peaks revealed clear segregation of tRCC and ccRCC cell lines. Overlapping peaks across samples were merged for each mark, creating a consensus set of 26,529, 63,322, and 342,285 peaks for the H3K4me3, H3K27ac, and MeDIP profiles, respectively.[000232] Differential peak analysis of ChlP-seq data identified 2,860 differential H3K4me3 peaks (of which 2,450 were enriched in tRCC compared to ccRCC; “tRCC-up”; FDR- q<0.01 and Log2fold-change (FC) > 2), 21,325 differential H3K27ac peaks (of which 11,443 were enriched in tRCC compared to ccRCC; “tRCC-up”; FDR-q<0.01 and Log2FC > 1), and 627 differentially methylated regions (DMRs) enriched tRCC compared to ccRCC (FDR-q<0.01 and Log2FC > 1). The results are summarized in FIG. 7. Each filled circle represents a peak (FIG, 7A: H3K4me3; FIG. 7B: H3K27ac; FIG. 7C: MeDIP) that was upregulated in ccRCC (light grey) and tRCC (dark grey), respectively. Consistent with the data in Example 1, FIG. 7A, FIG. 7B, and FIG. 7C highlight clear differences in peak distribution that make it possible to distinguish tRCC from ccRCC. This is further illustrated in FIG. 8 which visualizes the differential distribution of H3K4me3 -associated peaks (FIG. 8A) and H3K27ac-associated peaks (FIG. 8B) between tRCC cell lines (dark grey bar) and ccRCC cell lines (light grey bar).[000233] Motif analysis of the 11,443 tRCC-up H3K27ac peaks identified significant enrichment for sequences bound by TFE3 and its paralog MITF, which share consensus binding sites. This supports the hypothesis that the identified tRCC-up H3K27ac-associated sites include gene regulatory elements activated by the direct binding of TFE3 gene fusion transcriptionfactors. This finding is also consistent with recent reports that the binding of TFE3 gene fusion transcription factors to chromosomal DNA may facilitate the organization of enhancer loops. In short, the identified differential peaks of H3K27ac-associated DNA provide an epigenomic signature of nucleotide sequences comprising gene regulatory elements specifically associated with tRCC that can be used for diagnosing and / or monitoring tRCC.Intersection between GTRD and TFE3 gene fusions[000234] To identify differentially activated TFE3 gene fusion transcription factor binding sites (TFBS), known TFE3 TFBS were downloaded from the GTRD database as described in Example 1. The 24,050 peaks available in the GTRD database (originating from two non-RCC cell lines) were overlapped with 29,785 peaks identified by the union of TFE3 ChlP-seq in three tRCC cell lines (FU-UR-1, UOK109 and s-TFE), resulting in the identification of a set of 6,540 fusion-occupied TFBS. This is illustrated in FIG. 9. The two non-RCC cell lines in the GTRD database were LoVo (colorectal cancer cell line) and HepG2 (hepatocellular carcinoma cell line) and provided wild-type TFE3 TFBS. Overlap of peaks was assessed as described in Example 1.[000235] Assessing aggregated H3K27ac signals across these 6,540 fusion-occupied TFBS revealed a higher signal in all tRCC cell lines as compared to ccRCC cell lines. This difference in signal intensity was less pronounced when considering all TFBS from GTRD (n=24,050), or the fusion non-occupied binding sites which did not overlap with TFE3 fusion ChlP-seq peaks (n=17,510), underscoring the importance of building a robust consensus set of TFE3 fusion- occupied TFBS. Furthermore, only 853 of the 6,540 fusion-occupied TFBS overlapped with the H3K27ac-tRCC-up sites (n=l 1,443) identified using DiffBind, suggesting that these two methods identify partly non-overlapping regulatory sites associated with tRCC.Example 6. Detection of epigenomic signatures in cjDNA isolated from plasma[000236] The epigenomic signatures determined in Example 5 were then used to analyze 141 epigenomic libraries that had been prepared from 51 plasma samples from patients with tRCC (n=30), ccRCC (n=12), and healthy individuals (n=9).[000237] Plasma samples were collected from patients with tRCC and ccRCC diagnosed and treated at the Dana-Farber Cancer Institute (DFCI) between 2005 and 2024. All patientsprovided writen informed consent. The collection and use of samples was conducted under an IRB-approved protocol at DFCI. Studies were conducted in accordance with recognized ethical guidelines. Plasma samples from healthy individuals, which had been obtained as described in Example 2, were reanalyzed.[000238] Analysis of cfChIP data revealed increased signals at tRCC-up or ccRCC-up sites in plasma from RCC patients that were absent in plasma from healthy volunteers. For example, in tRCC plasma, H3K4me3 and H3K27ac signals were elevated at the GPR143 gene locus, a tRCC-specific site and TFE3 -fusion target gene. Conversely, an increased H3K4me3 and H3K27ac signal at C1QL1 gene locus in ccRCC plasma samples was observed. These findings were concordant with the published RNA-seq data in ccRCC and tRCC cell lines, as well as in tRCC and ccRCC tumor samples from a cohort of patients with metastatic ccRCC (Motzer et al. Cancer Cell. 2020 Dec 14; 38(6):803-817). Consistent with the findings described in Example 2, epigenomic signatures derived from peak calling as well as the intersection between GTRD and TFE3 gene fusions as described in Example 5 could successfully identify plasma samples from tRCC patients as distinct from ccRCC and healthy plasma samples.[000239] To develop an epigenomic- wide cfChIP signature for detecting tRCC, H3K4me3 and H3K27ac coverage was compared in patient plasma at the 2,450 H3K4me3 -associated tRCC-up sites and 11,443 H3K27ac-associated tRCC-up sites, which have been identified by cell line profiling as described in Example 5. As shown in FIG. 10, aggregated coverage at both H3K4me3 and H3K27ac tRCC-up sites was increased in plasma from patients with tRCC compared to patients with ccRCC (p=0.0027 and p=0.0029, respectively) or healthy controls (p=0.079 and p=0.00017, respectively). Further, as shown in FIG. 10B, H3K27ac coverage at the 6,540 TFE3 fusion-occupied binding sites (as described in Example 6) showed higher discriminating power versus ccRCC (p=0.0003) and versus healthy patients (p=0.00042), consistent with the finding in Example 2.[000240] The three sets of epigenomic data were used to build a robust cfChIP classifier for tRCC. This tRCC classifier utilized three distinct cell line-informed epigenomic signatures: H3K4me3 -associated tRCC-up sites, H3K27ac-associated tRCC-up sites, and TFE3 fusion- occupied binding sites. Aggregating plasma H3K4me3 signals in the cfChlP-seq data acrossH3K4me3 -associated tRCC-up sites (n=2,450) distinguished tRCC from ccRCC plasma samples (n=27 and n=l 1, respectively) with an area under the curve (AUC) of 0.78. Aggregating plasma H3K27ac signals in the cfChlP-seq data across H3K27ac tRCC-up sites (n=l 1,443) and TFE3 fusion-occupied binding sites (n=6,540) achieved AUCs of 0.8 and 0.84, respectively. This finding was consistent with the data described in Example 2.[000241] To enhance the performance of the classifier, the three distinct scores for each sample were combined to create a tRCC integrated epigenomic score (TIES). In the present example, the TIES was created by summing the three distinct scores. Signals at overlapping sites between H3K27ac-associated tRCC-up sites and TFE3 fusion-occupied binding sites were not double-counted. This approach achieved an AUC of 0.87 for the discrimination of tRCC from ccRCC (FIG. 11 A). For discriminating tRCC from healthy plasma samples (n=27 and n=9, respectively), aggregated H3K4me3 signals at H3K4me3 -associated tRCC-up sites and aggregated H3K27ac signals at H3K27ac-associated tRCC-up sites and TFE3 fusion-occupied binding sites achieved AUCs of 0.7, 0.89 and 0.87, respectively, with an integrated AUC of 0.910 (FIG. 11B). These findings using TIES demonstrate a significant improvement in discriminating tRCC from ccRCC and healthy plasma compared to the individual aggregated H3K4me3 signals at H3K4me3 -associated tRCC-up sites, aggregated H3K27ac signals at H3K27ac-associated tRCC-up sites, and TFE3 fusion-occupied binding sites, respectively .[000242] Among the tRCC plasma samples, two samples with < 3% ctDNA fraction as detected by ichorCNA (see Example 2) had amongst the highest signal by TIES. These samples had many copy number alterations (CNAs) that could be detected in the DNA from a tumor sample, whereas no CNAs were evident in the corresponding cfDNA from plasma. This highlights the limitation of CNA-centered methods for sensitive detection of tRCC via liquid biopsy.[000243] Given this observation, the limit of detection of the TIES was estimated using in silico dilution. The 11 tRCC samples with ctDNA fraction > 3% (as estimated by ichorCNA) were combined with each of the 8 healthy plasma samples at 11 dilution ratios (88 combinations for each of 11 dilution levels ranging from 0.9 to 0.01 ratio of tumor: healthy plasma by number of reads). Diluted samples were binned into intervals of 0.4% expected tumor fraction (rangingfrom <0.4% to >3.2 %). In each bin, the TIES score for the tumor dilutions was compared against the TIES score for the 8 healthy plasma samples via Wilcoxon test. The last interval with a significant difference is 0.8-1.2% (p<0.05). This suggests a sensitivity for detecting tRCC via cfChIP in the sub-1% TF range. Thus, using the TIES-based analysis method also provides a significant advantage over CNA-centered methods.Example 7. Monitoring tRCC tumor progression[000244] Using the TIES-based method described in Example 6, disease burden was measured in three patients with metastatic tRCC (Patient-I, -II, and -III). Plasma was collected at multiple time points during treatment, and cfDNA was isolated and enriched for H3K4me3- associated and H3K27ac-associated DNA as described in Example 2. Following sequencing, the resulting reads were analyzed using the method described in Example 6. The results are summarized in FIG. 12, which shows the longitudinal tracking of the TIES (light grey) and ctDNA fraction in the isolated cfDNA (dark grey) for each patient. The ctDNA fraction was determined as described in Example 2. FIG. 12 A, FIG. 12B, and FIG. 12C further depict radiographic changes in an index lesion (pleural metastasis) and the timings and doses of administered systemic therapies. Therapeutic interventions and plasma sample collection time points are schematically illustrated at the bottom of each graph. Disease burden was categorized as stable disease (SD), progressive disease (PD), or no evidence of disease (NED). Partial response (PR) is also annotated on the graphs, where appropriate.[000245] As can be seen from FIG. 12A, FIG. 12B, and FIG. 12C, variations in the TIES were concordant with the clinical course of response and progression to systemic therapy in all three patients. For example, as illustrated in FIG. 12A, an increase of the TIES at three time points (13-, 26, and 52-months post-diagnosis) was observed in a Patient-I corresponding to radiographic progression. A decrease of the TIES was observed at three time points corresponding to disease control or response (18 months and 39 months post-diagnosis) Patient- II was initially misdiagnosed as ccRCC on pathology. In this patient, an increase in TIES was observed at disease recurrence, followed by a subsequent decrease after a change in systemic therapy that resulted in radiographic disease control (FIG. 12B). Similarly, in Patient-III, an initial decrease in TIES following curative nephrectomy was observed, followed by an increaseof the TIES 16 months later, aligned with disease recurrence and metastasis to the liver (FIG. 12C). Moreover, TIES was detectable and dynamic even when the tumor fraction was in the undetectable range (<3%) by a method that estimates ctDNA using copy number alterations (Adalsteinsson et al. Nat Commun. 2017 Nov 6;8(1):1324).[000246] To evaluate cfChIP for monitoring tRCC treatment response, changes in the TIES and ctDNA fraction for each pair of consecutive plasma draws were calculated. As shown in FIG. 13 A, a comparison of these changes during intervals of disease progression, stability, or response revealed that TIES between consecutive draws increased at times of disease progression and decreased during response or disease stability (p=0.027). Conversely, changes in tumor fraction were less pronounced (p=0.63), as shown in FIG. 13B.[000247] This example indicates that the described liquid biopsy epigenomic assay is more effective at tracking disease evolution and response compared to other methods that rely solely on CNAs, likely due to the low frequency of such alterations in tRCC.Example 8. Detection of epigenomic signatures in cjDNA is more sensitive than CTC profiling[000248] The profiling of the epigenomic libraries of Example 5 revealed that H3K4me3 and H3K27ac signals were elevated at different gene loci in tRCC plasma compared to ccRCC plasma. For example, in tRCC plasma, H3K4me3 and H3K27ac signals were found to be elevated at the GPR143 gene locus. GPR143 is a tRCC-specific peak and target gene of the neomorphic TFE3 gene fusion transcription factor. Conversely, an increased H3K4me3 and H3K27ac signal was observed at C1QL1 gene locus in ccRCC plasma samples. Published RNA- seq data from tRCC and ccRCC cell lines and tumor samples validated these observations.[000249] The suitability of profiling circulating tumor cells (CTCs) for discriminating tRCC and ccRCC samples was assessed. As GPR143 and Cl QL1 appeared to be highly selective marker genes for tRCC and ccRCC, respectively, these two markers were investigated for expression in CTCs. A cohort of 10 patients with metastatic ccRCC, 7 patients with metastatic tRCC, and one healthy control was sampled for CTCs.[000250] CTCs were isolated from whole blood samples within four hours of collection to ensure high recovery of intact CTCs with quality RNA. Enrichment was performed using a CTC-iChip (TellBioDX) prior to expression analysis. Briefly, leukocytes were depleted using the microfluidic CTC-iChip system using biotinylated antibodies targeting CD45 (clone HI30), CD66b (clone 80H3), and CD 16 (clone 3G8), respectively, followed by incubation with Dynabeads MyOne Streptavidin T1 for magnetic labelling. RNA extraction was performed on the CTC samples followed by whole transcriptome amplification. Digital droplet PCR (ddPCR) was performed to detect tRCC (GPR143 and TRIM63) or ccRCC (ClQLl)-specific transcripts.[000251] As a positive control, healthy blood spiked with RNA derived from tRCC or ccRCC cell lines was included. Transcripts could be detected in the tRCC and ccRCC positive control samples. However, no signal was detectable for these transcripts in the patient samples. Four patients were sampled for both cfDNA and CTC isolation, with two of them sampled from the same blood draw.[000252] This example illustrates that detection of the epigenomic signatures of a neomorphic gene fusion transcription factor in cfDNA using the ChlP-based methods described herein is more sensitive than standard methods for profiling CTCs.Example 9. Detection of fusion-specific epigenomic signatures in prostate cancer[000253] To evaluate the extensibility of the approach, it was investigated whether epigenomic signatures of neomorphic gene fusion transcription factors in plasma samples could be detected in other cancers, using methods similar to those described in Examples 2 and 5 for tRCC. Plasma samples of patients with prostate cancer with (n=5) or without (n=8) a TMPRSS2- ERG gene fusion were compared using cfChlP. The TMPRSS2-ERG gene fusion places the ETS- family transcription factor ERG under the control of the androgen-responsive gene TMPRSS2. This gene fusion is found in approximately 50% of prostate cancer cases and is associated with a distinctive transcriptional signature.[000254] As shown in FIG. 14, the plasma from patients with the TMPRSS2-ERG fusion exhibited significantly higher H3K27ac signals at fusion-specific H3K27ac sites (n=7,531) compared with samples from patients with fusion-negative cancers (p=0.006). As shown in FIG. 15, samples with and without the TMPRSS2-ERG fusion could be distinguished with an AUC of 0.95.1[000255] This example illustrates the applicability of the methods described herein to identify epigenomic signatures in other cancers comprising a transcription factor gene fusion.Example 10. Identifying “mutationally quiet” cancers[000256] Cancers comprising a transcription factor gene fusion can be difficult to detect by transcriptional profiling. To identify additional such cancers that may be particularly amenable to profiling via cfChIP, a pan-cancer analysis of the fraction of genome altered (FGA) was performed. FGA is a metric of the proportion of the genome affected by copy number alteration. A significant variation was observed in FGA both between and within cancer lineages.[000257] Tumor samples were then categorized as to whether they harbored transcription factor gene fusions or not (analogous to a TFE3 gene fusion in tRCC and a TMPRSS2-ERG gene fusion in prostate cancer). It was then determined how FGA aligned with fusion status across different cancer types. As shown in FIG. 16, it was observed that cancers comprising a transcription factor gene fusion had significantly lower FGA than cancers that did not comprise a transcription factor gene fusion (median 0.06 vs. 0.20, respectively; p<2.2xl0'16). For example, FIG. 17 shows that FGA in renal cell carcinomas comprising a TFE3 gene fusion was significantly lower than in other RCCs that do not comprise a transcription factor gene fusion (median 0.08 vs 0.15, respectively; p=0.038). Similarly, FIG. 18 shows that FGA in synovial sarcomas comprising an SSX2 gene fusion was significantly lower than in other sarcomas that do not comprise a transcription factor gene fusion (median 0.08 vs 0.34, respectively; p=0.011). As shown in FIGs 16-18, there is an inverse correlation between the FGA of a cancer and the presence of a transcription factor gene fusion in that cancer.[000258] This example demonstrates that low FGA can identify a subgroup of cancers that may comprise a transcription factor gene fusion and may be amenable to profiling by cfChIP.[000259] It should be understood that the details provided herein are given by way of illustration only, not limitation. Other features, objects, and advantages are apparent from the above detailed description, drawings and examples. Various changes and modifications will be apparent to those skilled in the art.[000260] All patents, patent publications and non-patent publications referenced herein are indicative of the level of skill of those skilled in the art to which this invention pertains. All these publications are herein incorporated by reference to the same extent as if each individual publication were specifically and individually indicated as being incorporated by reference.
Claims
AMENDED CLAIMS received by the International Bureau on 19 December 2025 (19.12.2025)1. A method for diagnosing a cancer comprising a transcription factor gene fusion in a subject, the method comprising:(a) isolating cell-free DNA (cfDNA) from a sample obtained from the subject;(b) enriching the cfDNA: i.with a means for specifically binding histone H3 protein acetylated at the lysine at residue 27 (H3K27ac) for H3K27ac-associated DNA; and / or ii.with a means for specifically binding histone H3 protein methylated at the lysine at residue 4 (H3K4me3) for H3K4me3 -associated DNA; and(c) detecting the presence of or level of one or more nucleotide sequences comprising gene regulatory elements specifically associated with the cancer comprising the transcription factor gene fusion in the cfDNA enriched for H3K27ac-associated DNA and / or H3K4me3 -associated DNA; wherein the presence or level of the one or more nucleotide sequences comprising gene regulatory elements specifically associated with the cancer comprising a transcription factor gene fusion in the enriched cfDNA is indicative of the subject suffering from the cancer comprising a transcription factor gene fusion.
2. The method of claim 1, wherein the cancer comprising a transcription factor gene fusion is selected from the group consisting of renal cell carcinoma (RCC), prostate cancer, and synovial sarcoma.
3. The method of claim 2, wherein:(a) the cancer is RCC and the one or more nucleotide sequences comprising gene regulatory elements specifically associated with RCC comprise one or more nucleotide sequences associated with a TFE3 gene fusion; or(b) the cancer is prostate cancer and the one or more nucleotide sequences comprising gene regulatory elements specifically associated with prostate cancer comprise one or more nucleotide sequences associated with a TMPRSS2-ERG gene fusion; or(c) the cancer is synovial sarcoma and the one or more nucleotide sequences comprising gene regulatory elements specifically associated with synovial sarcoma comprise one or more nucleotide sequences associated with an SSX2 gene fusion.
4. The method of any one of the preceding claims, wherein each of the one or more nucleotide sequences has a length of 100-250 base pairs.
5. The method of any one of the preceding claims, wherein the sample is blood, plasma, or serum.
6. The method of any one of the preceding claims, wherein the means for specifically binding H3K27ac is an antibody.
7. The method of any one of the preceding claims, wherein the means for specifically binding H3K4me3 is an antibody.
8. The method of any one of the preceding claims, wherein the step of enriching the cfDNA comprises chromatin immunoprecipitation (ChIP).
9. The method of any one of the preceding claims, wherein the detecting step comprises sequencing.
10. The method of claim 9, wherein the sequencing comprises next-generation sequencing (NGS).
11. The method of claim 9 or 10, wherein the sequencing comprises ChlP-Seq and optionally MeDIP-Seq.
12. The method of any one of the preceding claims, wherein the one or more nucleotide sequences comprise one or more of: a promoter, an enhancer, a CpG island, a transcription factor binding site, and a differentially methylated region (DMR).
13. The method of any one of the preceding claims, wherein the method further comprises detecting the presence of or level of one or more nucleotide sequences associated with differentially methylated regions (DMRs) associated with a cancer that does not comprise a transcription factor gene fusion.
14. The method of claim 13, wherein the cancer that does not comprise a transcription factor gene fusion is clear cell RCC (ccRCC).
15. A method for monitoring the tumor progression in a subject suffering from a cancer comprising a transcription factor gene fusion, the method comprising:(a) measuring, from a first sample obtained from the subject at a first time point, the presence and / or level of one or more nucleotide sequences comprising gene regulatory elements specifically associated with the cancer comprising the transcription factor gene fusion;(b) measuring, from a second sample obtained from the subject at a second or subsequent time point, the presence and / or level of the one or more nucleotide sequences comprising gene regulatory elements specifically associated with the cancer comprising the transcription factor gene fusion; and(c) comparing the measurements from steps (a) and (b), thereby monitoring the tumor progression, wherein the measuring steps are carried out according to the method of any one of claims 1-14.
16. The method of claim 15, wherein the subject is administered a therapeutic agent to treat the cancer.
17. The method of claim 16, wherein the first time point is before administering the therapeutic agent, and the second or subsequent time point is after administering the therapeutic agent.
18. The method of claim 16 or 17, wherein an increase in the level of the one or more nucleotide sequences from the first time point to the second or subsequent time points indicates that the therapeutic agent is ineffective for the treatment of the cancer.
19. The method of claim 18, wherein the treatment with the therapeutic agent is discontinued and the subject receives a different therapy.
20. The method of claim 15, wherein an increase in the level of the one or more nucleotide sequences from the first time point to the second or subsequent time points indicates tumor progression, and wherein a decrease in the level of the one or more nucleotide sequences, or the absence of a detectable level of the one or more nucleotides sequences indicates that the subject is in remission.
21. A method for classifying a cancer subtype, the method comprising:(a) detecting, from a sample derived from a subject suffering from a cancer, the presence and / or level of the one or more nucleotide sequences comprising gene regulatory elements specifically associated with a cancer comprising a transcription factor gene fusion, and(b) identifying the cancer subtype as the cancer comprising the transcription factor gene fusion based on the presence of the one or more nucleotide sequences; wherein the detecting step is carried out according to the method of any one of claims 1-14.
22. A method for classifying a cancer subtype according to claim 21, wherein the cancer comprising a transcription factor gene fusion is:(a) RCC; or(b) prostate cancer; or(c) sarcoma; and the cancer subtype is identified as: i.RCC comprising a TFE3 gene fusion; or ii. prostate cancer comprising a TMPRSS2-ERG gene fusion; or iii. synovial sarcoma comprising a SSX2 gene fusion.
23. A method of selecting a therapy for a subject suffering from a cancer comprising a transcription factor gene fusion, comprising:(a) isolating cell-free DNA (cfDNA) from a sample obtained from the subject;(b) enriching the cfDNA:(i) with a means for specifically binding histone H3 protein acetylated at the lysine at residue 27 (H3K27ac) for H3K27ac-associated DNA; and / or(ii) with a means for specifically binding histone H3 protein methylated at the lysine at residue 4 (H3K4me3) for H3K4me3-associated DNA; and(c) detecting the presence of one or more nucleotide sequences comprising gene regulatory elements specifically associated with the cancer comprising the transcription factor gene fusion in the cfDNA enriched for H3K27ac-associated DNA and / or H3K4me3 -associated DNA, wherein the presence of the one or more nucleotide sequences comprising gene regulatory elements specifically associated with the cancer comprising the transcription factor gene fusion in the enriched cfDNA indicates that the subject’s cancer is the cancer comprising a transcription factor gene fusion; and(d) selecting a therapy that is effective against the cancer comprising the transcription factor gene fusion.
24. A method for selecting a therapy according to claim 23, wherein the cancer comprising the transcription factor gene fusion is:(a) RCC comprising a TFE3 gene fusion; or(b) a prostate cancer comprising a TMPRSS2-ERG gene fusion; or(c) a synovial sarcoma comprising an SSX2 gene fusion.
25. A method of treating a cancer comprising a transcription factor gene fusion in a subject, the method comprising:(a) isolating cell-free DNA (cfDNA) from a sample obtained from the subject;(b) enriching the cfDNA:(i) with a means for specifically binding histone H3 protein acetylated at the lysine at residue 27 (H3K27ac) for H3K27ac-associated DNA; and / or(ii) with a means for specifically binding histone H3 protein methylated at the lysine at residue 4 (H3K4me3) for H3K4me3-associated DNA; and(c) detecting the presence of one or more nucleotide sequences comprising gene regulatory elements specifically associated with the cancer comprising the transcription factor gene fusion in the cfDNA enriched for H3K27ac-associated DNA and / or H3K4me3 -associated DNA, wherein the presence of the one or more nucleotide sequences comprising gene regulatory elements specifically associated with the cancer comprising the transcription factor gene fusion in the enriched cfDNA is indicative of the subject suffering from the cancer comprising the transcription factor gene fusion, and(d) administering to the subject a therapeutically effective amount of a therapeutic agent effective in the treatment of the cancer comprising the transcription factor gene fusion.
26. The method of claim 25, wherein the cancer is translocation renal cell carcinoma (tRCC) and optionally wherein the therapeutic agent is selected from the group consisting of:(a) a tyrosine kinase inhibitor (TKI), optionally sunitinib, cabozantinib, lenvatinib, tivozanib, axitinib, and / or pazopanib; and / or(b) an immunotherapy, optionally pembrolizumab and / or nivolumab.
27. A method of monitoring the effectiveness of a therapy in a subject with a cancer comprising a transcription factor gene fusion, the method comprising:(a) providing a first sample obtained from the subject at a first time point;(b) isolating cell-free DNA (cfDNA) from the first sample;(c) enriching the cfDNA:(i) with a means for specifically binding histone H3 protein acetylated at the lysine at residue 27 (H3K27ac) for H3K27ac-associated DNA; and / or(ii) with a means for specifically binding histone H3 protein methylated at the lysine at residue 4 (H3K4me3) for H3K4me3-associated DNA; and(d) determining levels of one or more nucleotide sequences comprising gene regulatory elements specifically associated with the cancer comprising the transcription factor gene fusion in the cfDNA enriched for H3K27ac-associated DNA and / or H3K4me3- associated DNA;(e) providing a second or subsequent sample obtained from the subject and repeating steps (b)-(d) with the second sample; and(f) comparing the levels of the one or more nucleotide sequences obtained in steps (d) and (e), wherein:(i) an increase in the level of the one or more nucleotide sequences determined in step (e) relative to step (d) indicates that the therapy is not effective against the subject’s cancer; and(ii) no change or a decrease in the level of the one or more nucleotide sequences determined in step (e) relative to step (d) indicates that the therapy is effective against the subject’s cancer.
28. The method of claim 27, wherein the first sample is obtained from the subject prior to the commencement of therapy and the second or subsequent samples are obtained from the subject after commencement of therapy.
29. A computer program comprising a computer program code configured to cause one or more physical computing devices to perform a method of analysis of sample results for a subject when the code is run on the one or more physical computing devices, the method of analysis comprising:(a) receiving data indicative of the presence or level of one or more nucleotide sequences comprising gene regulatory elements specifically associated with a cancer comprising a transcription factor gene fusion in a sample collected from the subject;(b) obtaining a statistical model;(c) comparing the data to the statistical model to generate a score; and(d) outputting the score from the statistical model.
30. The computer program of claim 29, wherein the sample was processed to:(a) isolate cell-free DNA (cfDNA) from a sample obtained from the subject;(b) enrich the cfDNA:(i) with a means for specifically binding histone H3 protein acetylated at the lysine at residue 27 (H3K27ac) for H3K27ac-associated DNA; and / or(ii) with a means for specifically binding histone H3 protein methylated at the lysine at residue 4 (H3K4me3) for H3K4me3 -associated DNA; and(c) detect the level of one or more nucleotide sequences comprising gene regulatory elements specifically associated with the cancer comprising the transcription factor gene fusion in the cfDNA enriched for H3K27ac-associated DNA and / or H3K4me3- associated DNA.
31. The method of claim 27 or 28 or the computer program of claim 29 or 30, wherein the cancer comprising a transcription factor gene fusion is:(a) RCC comprising a TFE3 gene fusion; or(b) prostate cancer comprising a TMPRSS2-ERG gene fusion; or(c) synovial sarcoma comprising an SSX2 gene fusion.
32. The computer program of claim 29 or 30, wherein the statistical model is configured such that the score is suitable for at least one of:(a) identifying whether or not the subject has a cancer comprising a transcription factor gene fusion in an organ or tissue; and(b) identifying whether the subject is responsive to a standard therapy for a cancer affecting the same organ or tissue.
33. A computer-implemented method for determining whether a subject’s cancer comprises a transcription factor gene fusion, the method comprising:(a) Receiving sequencing read data obtained from cfDNA isolated from a sample obtained from the subject and enriched for H3K27ac-associated DNA and optionally sequencing read data obtained from cfDNA isolated from the sample and enriched for H3K4me3 -associated DNA;(b) Aligning the sequencing read data obtained in (a) with whole human genome read data;(c) Receiving a first predetermined list of transcription factor fusion binding sites, a second predetermined list of gene regulatory elements associated with H3K27ac, and optionally a third predetermined list of gene regulatory elements associated with H3K4me3;(d) For each site or element present in one of the lists received in (c) determining the number of aligned sequencing reads obtained in (b);(e) Aggregating the number of sequencing reads for each element in each list;(f) Receiving a reference value;(g) For each aggregate, normalizing the aggregate based on the reference value;(h) Calculating a combined score based on the two or more aggregates; and(i) Comparing the combined score to a predetermined score to determine whether the cancer comprises the transcription factor gene fusion.
34. The computer implemented method of claim 33, wherein the gene regulatory elements in the second predetermined list and optionally the third predetermined list do not comprise the transcription factor fusion binding sites in the first predetermined list.
35. The computer implemented method of claim 33 or 34, wherein calculating a combined score comprises:(a) calculating a score for each of the two or more aggregates; and(b) combining the scores.
36. The computer implemented method of any one of claim 33-35, wherein normalizing each aggregate comprises determining a ratio of the aggregate to the reference value, and wherein the ratio is the score for the aggregate.
37. The computer implemented method of any one of claims 33-36, wherein aggregating the number of sequencing reads comprises:(a) for each list;(b) for each site or element in the list identifying a corresponding peak, resizing the corresponding peaks to obtain a plurality of resized peaks having a predetermined peak width, and centering each resized peak on its corresponding peak;(c) dividing each resized peak into a plurality of bins:(d) obtaining, for each resized peak, a plurality of bin counts by counting the number of sequencing reads for each bin; and(e) for each bin count, aggregating the plurality of bin counts across each resized peak to obtain an aggregate.
38. The method of claim 37, wherein:(a) the predetermined peak width is about 2-6kb; and(b) each bin has a width of about 10-100bp.
39. The computer implemented method of any one of claims 33-38, wherein normalizing the aggregates comprises carrying out a shoulder normalization process.
40. The method of any one of claims 33-39, wherein the reference value represents the number of sequencing reads for a pre-determined set of at least 1,000 DNAse hypersensitivity sites in the sequencing read data obtained in (a).
41. The computer implemented method of any one of claims 33-40, wherein aligning the sequencing read data comprises discarding non-uniquely mapping sequencing reads.
42. The computer implemented method of any one of claims 33-41, wherein calculating the combined score comprises:(a) providing the two or more aggregates as inputs to a statistical model; and(b) receiving a combined score as an output of the statistical model.
43. The computer implemented method of claim 42, wherein the statistical model is a logistic regression model, a Cox Proportional Hazards model, or a machine learning model.
44. A computer implemented method for obtaining a list of transcription factor fusion binding sites for use in the method of any one of claim 33-43, the method comprising:(a) obtaining a first reference dataset and a second reference data set, wherein the first reference dataset comprises a plurality of reference sequencing reads of DNA fragments associated with a fusion transcription factor encoded by a transcription factor gene fusion and the second reference dataset comprises a plurality of reference sequencing reads of DNA fragments associated with a corresponding wildtype transcription factor;(b) carrying out a peak analysis on the first reference dataset and the second reference dataset to identify a plurality of sites that overlap between the first reference dataset and the second reference data set;(c) determining significant overlap sites; and(d) selecting, from the significant overlap sites, the list of transcription factor fusion binding sites by selecting sites having a logifold change in either direction greater than 1.
45. The method of claim 44, wherein determining significant overlap sites comprises identifying overlap sites having a false discovery rate q-value equal to or less than 0.01.
46. A computer implemented method for obtaining a list of gene regulatory elements associated with H3K27ac for use in the method of any one of claims 33-43, the method comprising:(a) obtaining a first reference dataset and a second reference data set, wherein the first reference dataset comprises reference sequencing reads of DNA fragments associated with H3K27ac of a first cancer cell population comprising a transcription gene fusionand the second reference dataset comprising a plurality of reference sequencing reads of DNA fragments associated with H3K27ac in a second cancer cell population of the same tissue type as the cancer comprising the transcription factor gene fusion, wherein the second cancer cell population does not comprise the transcription factor gene fusion;(b) carrying out a differential peak analysis on the first reference dataset and the second reference dataset to identify a plurality of differentially marked sites present only in the first cancer cell population comprising the transcription gene fusion;(c) determining significant differentially marked sites of the differentially marked sites; and(d) selecting, from the significant differentially marked sites, the list of gene regulatory elements associated with H3K27ac by selecting sites having a logifold change in either direction greater than 1.
47. A computer implemented method for obtaining a list of gene regulatory elements associated with H3K4me3 for use in the method of any one of claim 33-43, the method comprising:(a) obtaining a first reference dataset and a second reference data set, wherein the first reference dataset comprises a plurality of reference sequencing reads of DNA fragments associated with H3K4me3 of a first cancer cell population comprising a transcription gene fusion and the second reference dataset comprising a plurality of reference sequencing reads of DNA fragments associated with H3K4me3 in a second cancer cell population of the same tissue type as the cancer comprising the transcription factor gene fusion, wherein the second cancer cell population does not comprise the transcription factor gene fusion;(b) carrying out a differential peak analysis on the first reference dataset and the second reference dataset to identify a plurality of differentially marked sites present only in the first cancer cell population comprising the transcription gene fusion;(c) determining significant differentially marked sites of the differentially marked sites; and(d) selecting, from the significant differentially marked sites, the list of gene regulatory elements associated with H3K4me3 by selecting sites having a logifold change in either direction greater than 2.
48. The method of claim 46 or 47, wherein determining the significantly differentially marked sites comprises identifying differentially marked sites having a false discovery rate q- value equal to or less than 0.01.