Methods of detecting, diagnosing, classifying and / or monitoring cancers using cell-free DNA and kits or devices implementing the same
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2026-04-09
AI Technical Summary
Traditional cancer detection and monitoring methods are invasive, time-consuming, and lack specificity and sensitivity, especially for cancers without accessible tumor biopsies, and personalized molecular residual disease testing is costly and not suitable for widespread screening.
A multi-step method involving isolating cell-free DNA, enriching for chromatin-associated DNA using immunoprecipitation and probe capture, followed by quantitative PCR to detect differentially activated regulatory regions of cancer-associated genes, enhancing sensitivity and specificity in liquid biopsies.
Improves the sensitivity and accuracy of cancer detection, monitoring, and classification through non-invasive means, enabling early detection and tracking of cancer progression and response to therapy.
Smart Images

Figure IB2025058005_09042026_PF_FP_ABST
Abstract
Description
METHODS OF DETECTING, DIAGNOSING, CLASSIFYING AND / OR MONITORING CANCERS USING CELL-FREE DNA AND KITS OR DEVICES IMPLEMENTING THE SAMECROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims benefit of U.S. Provisional Application No. 63 / 679,845 filed August 6, 2024, which is incorporated herein by reference.SEQUENCE LISTING
[0002] The present specification makes reference to a Sequence Listing, submitted electronically as an .xml file name “2025-07-31 DFCI 3358 Sequence Listing” on August, 6 2025. The .xml file was generated on 31 July 2025 and is 12.7 in size. The entire contents of the Sequence Listing are herein incorporated by reference.FIELD
[0003] Provided herein are methods of detecting, diagnosing, classifying and / or monitoring cancer in a subject by isolating and enriching cell-free, chromatin-associated DNA and performing quantitative PCR (qPCR) to detect one or more nucleotide sequences associated with the presence of cancer. Also provided are methods of detecting, diagnosing, and / or monitoring bladder cancer using specific gene regulatory regions and associated genes disclosed herein. Kits or devices for implementing the methods are also provided, as are libraries of enriched cell-free, chromatin-associated DNA obtained by the methods.BACKGROUND
[0004] It is estimated that 1 in 2 people will develop cancer at some point in their lifetime. Traditional methods of the detection, identification and / or monitoring of cancer can be invasive e.g., requiring a tumor sample or solid biopsy, time consuming, and may require specialist equipment. Accordingly, such methods are not readily accessible to many patients. They also may not be useful in cancers for which tumor biopsies are not available. Traditional techniques can also lack specificity and sensitivity.
[0005] Newer techniques for monitoring cancers involving personalized medicine, such as personalized molecular residual disease (MRD) testing, require a tumor sample to prepare a unique tumor signature that can later be used for monitoring the recurrence of cancer. One example of a commercially available MRD system is the Natera system. Although highly specialized, such systems may be too costly to be used to screen for, detect and monitor cancer in the general population.
[0006] Effective non-invasive detection is currently available only for a small number of cancers. Therefore, there is an urgent need to develop simple, accurate, and non-invasive methods for the early detection of cancer.SUMMARY
[0007] Methods for identifying cancer biomarkers commonly involve the detection of cancer-associated gene expression products, such as mRNA or protein. In contrast, the present methods focus on the detection of epigenetic markers of cancer in the genome. Epigenetic alterations in tumor-derived cell-free DNA (cfDNA) can be used to improve detection, diagnosis, and monitoring subjects with cancer. In particular, the present methods relate to the detection of chromatin-associated DNA (i.e., genomic DNA) and more specifically differentially activated regulatory regions of cancer-associated genes (e.g., highly expressed genes that can drive cell de-differentiation and / or cancer progression).
[0008] It has been discovered that a multi-step method provides a highly sensitive and accurate method of detecting the presence or increased level of such cancer-associated differentially activated regulatory regions in a liquid biopsy (e.g., a blood or urine sample). This multi-step method involves isolating cell-free DNA (cfDNA) and enriching for chromatin- associated DNA comprising one or more nucleotide sequences associated with the presence of cancer (e.g., differentially activated regulatory regions of genes associated with one or more cancers) followed by quantitative PCR (qPCR) to amplify and detect a portion of the one or more nucleotide sequences associated with the presence of cancer. Specifically, by combining the steps of enriching cfDNA first for chromatin-associated DNA (e.g., using immunoprecipitation) and then for known gene regulatory sequences associated with the presence of cancer (usingprobe capture), it has been surprisingly found that sensitivity and specificity of the detection of such sequences in a liquid biopsy (e.g., a blood or urine sample) can be dramatically improved.
[0009] Accordingly, in one aspect, there is provided a method for diagnosing or detecting a cancer in a subject, the method comprising providing a sample obtained from the subject; isolating cell-free DNA (cfDNA) from the sample; contacting the cfDNA with a means for specifically binding chromatin to isolate chromatin-associated DNA; enriching the isolated chromatin-associated DNA using oligonucleotide probes that specifically hybridize to one or more nucleotide sequences associated with the presence of the cancer, wherein each oligonucleotide probe binds to a distinct region of the one or more nucleotide sequences; and performing quantitative PCR (qPCR) by contacting the enriched isolated chromatin-associated DNA with one or more primer pairs to amplify and detect a portion of the one or more nucleotide sequences associated with the presence of the cancer, wherein the presence or increased level of the one or more nucleotide sequences associated with the presence of the cancer in the sample indicates that the subject has the cancer.
[0010] In another aspect, there is provided a method for monitoring the progression of a cancer in a subject that employs the same steps as the methods described above to detect, in a first sample obtained from the subject at a first time point, the presence and / or level of one or more nucleotide sequences associated with the presence of the cancer; detecting, in a second sample obtained from the subject at a second time point (e.g., one or more subsequent time points), the presence and / or level of the one or more nucleotide sequences associated with the presence of the cancer; and comparing the results obtained for the first and second samples, thereby monitoring the progression of the cancer. In some embodiments, the monitoring is performed to assess the subject’s response to an anti-cancer therapy.
[0011] The present methods can also be used to classify a cancer type and / or subtype. Accordingly, in a further aspect, there is provided a method for classifying a type and / or subtype of cancer in a subject, the method comprising detecting, in a sample obtained from the subject, the presence and / or level of one or more nucleotide sequences associated with the type and / or subtype of the cancer; and optionally comparing the results obtained for the sample to a reference data set for the cancer type and / or subtype.
[0012] Also provided are kits and devices for carrying out the methods described herein. The kit or device comprises reagents for isolating cfDNA; reagents for specifically binding chromatin to isolate chromatin-associated DNA; reagents (e.g., oligonucleotide probes) for enriching the chromatin-associated DNA for one or more nucleotide sequences associated with the presence of a cancer; and optionally reagents (e.g., one or more primer pairs) for carrying out qPCR to amplify and / or detect the presence or level of a portion of the one or more nucleotide sequences associated with the presence of the cancer.
[0013] In a further aspect, a method of preparing a library of isolated cell-free chromatin- associated DNA comprising nucleotide sequences associated with the presence of a cancer is provided. The method comprises providing a sample obtained from a subject with the cancer; isolating cell-free DNA (cfDNA) from the sample; contacting the cfDNA with a means for specifically binding chromatin to isolate chromatin-associated DNA; enriching the isolated chromatin-associated DNA using oligonucleotide probes that specifically hybridize to one or more nucleotide sequences associated with the presence of the cancer, wherein each oligonucleotide probe binds to a distinct region of the one or more nucleotide sequences; and optionally amplifying the chromatin-associated DNA, e.g, using an isothermal amplification reaction.
[0014] In the course of validating the methods described above, gene regulatory sequences and associated genes were discovered that are suitable to classify bladder cancer as luminal-like or basal-like. Accordingly, in a further aspect, provided herein are methods for classifying or diagnosing a bladder cancer in a subject. A method for classifying a bladder cancer as luminal-like or basal-like may comprise providing a sample obtained from a subject suffering from bladder cancer; and determining in the sample expression of, and / or activation state of a gene regulatory region of, one or more of BMP7, ERN2, SRPX2, UGT1A10, CLCA2, SPOCD1, SERPINB5, SEMA4B, CLCA4, CLU, UGT1A7, SH3PXD2A, AQP3, ANXA10 and SLITRK6; and expression of, and / or activation state of a gene regulatory region of, one or more of FBN2, SLC4A4, CTTNBP2, KALRN, SELL, PMP22, UGT2B7, COL12A1, PTH2R, ANKRD36, DCDC2, PRTG, BPGM, PDE9A, and ANXA3, wherein overexpression of, or an active state of the regulatory region of, one or more of BMP7, ERN2, SRPX2, UGT1A10, CLCA2, SPOCD1, SERPINB5, SEMA4B, CLCA4, CLU, UGT1A7, SH3PXD2A, AQP3, ANXA10 and SLITRK6indicates that the bladder cancer is basal-like; and wherein overexpression of, or an active state of the gene regulatory region of, one or more of FBN2, SLC4A4, CTTNBP2, KALRN, SELL, PMP22, UGT2B7, COL12A1, PTH2R, ANKRD36, DCDC2, PRTG, BPGM, PDE9A, m<3ANXA3 indicates that the bladder cancer is luminal-like.
[0015] For example, a method for classifying a bladder cancer as luminal-like or basal- like may comprise providing a sample obtained from a subject suffering from bladder cancer; and determining in the sample expression of, and / or activation state of a gene regulatory region of, one or more of BMP7, ERN2, SRPX2, UGT1A10, CLCA2, SERPINB5, CLCA4, CLU, SH3PXD2A, AQP3, ANXA10, and SLITRK6; and expression of, and / or activation state of a gene regulatory region of, one or more of CTTNBP2, KALRN, SELL, PMP22, COL12A1, PRTG, PDE9A, and ANXA3, wherein overexpression of, or an active state of the regulatory region of, one or more of BMP7, ERN2, SRPX2, UGT1A10, CLCA2, SERPINB5, CLCA4, CLU, SH3PXD2A, AQP3, ANXA10, and SLITRK6 indicates that the bladder cancer is basal-like; and wherein overexpression of, or an active state of the gene regulatory region of, one or more of CTTNBP2, KALRN, SELL, PMP22, COL12A1, PRTG, PDE9A, and ANXA3 indicates that the bladder cancer is luminal-like.
[0016] Additionally, the method may comprise determining in the sample expression of, and / or activation state of a gene regulatory region selected from the group consisting of UGT1A1, TP63- and the group consisting of CASTOR3, CHKA, UPK1A, and PLEKHF1, wherein overexpression of, or an active state of the regulatory region of, one or more of UGT1A1 and TP63 indicates that the bladder cancer is basal-like, and wherein overexpression of, or an active state of the gene regulatory region of, one or more of CASTOR3, CHKA, UPK1A, and PLEKHF1 indicates that the bladder cancer is luminal-like.
[0017] Gene regulatory sequences and associated genes have also been discovered that are suitable to classify bladder cancer as micropapillary. Accordingly, in a further aspect, provided herein is a method for classifying a bladder cancer as micropapillary, wherein the method comprises providing a sample obtained from a subject suffering from bladder cancer; and determining in the sample expression of, and / or activation state of a gene regulatory region of, one or more of ANPEP, BPGM, SPNS2, MAL, ASS1, COL4A4, MUC4, TMC5, MMP7, PLCXD3, KRT80, PROMI, DAB2, KRT23, anARHPN2, wherein overexpression of, or an activestate of the regulatory region of, one or more of ANPEP, BPGM, SPNS2, MAE, ASS1, COL4A4, MUC4, TMC5, MMP7, PLCXD3, KRT80, PROMI, DAB2, KRT23, and RHPN2 indicates that the bladder cancer is micropapillary.
[0018] It has further been discovered that certain gene regulatory regions and associated genes are common to the investigated subtypes of bladder cancer and particularly suitable to determine whether a subject suffers from bladder cancer. Thus, in yet a further aspect, a method for diagnosing bladder cancer in a subject is provided, comprising providing a sample obtained from the subject; and determining in the sample the expression of, and / or activation state of one or more (e.g., all) of gene regulatory regions of, NECTIN4, SOX4, FGFR3, TP63, KDM6A, KRT5, SPP1, m G BPS, wherein overexpression of, or an active state of the one or more (e.g., all) of gene regulatory regions of, NECTIN4, SOX4, FGFR3, TP63, KDM6A, KRT5, SPP1 and IGFBP3 indicates that the subject suffers from bladder cancer. Additionally, the method may comprise determining in the sample the expression of, and / or activation state of one or more gene regulatory regions selected from the group consisting of UGT1A8, SEMA4B, KRT80, KRT7, ERBB2, TACSTD2, ELK3, PSCA, EPCAM, EHF, TBX3, and ELF3.
[0019] For example, a method for diagnosing bladder cancer in a subject is provided, comprising providing a sample obtained from the subject; and determining in the sample the expression of, and / or activation state of one or more (e.g., all) of gene regulatory regions of, UGT1A8, SEMA4B, NECTIN4, KRT80, KRT7, ERBB2, TACSTD2, KDM6A, ELK3, PSCA, EPCAM, EHF, TBX3, and ELF3, wherein overexpression of, or an active state of the one or more (e.g, all) of gene regulatory regions of, UGT1A8, SEMA4B, NECTIN4, KRT80, KRT7, ERBB2, TACSTD2, KDM6A, ELK3, PSCA, EPCAM, EHF, TBX3, and ELF 3 indicates that the subject suffers from bladder cancer.
[0020] Also provided is a computer-implemented method for classifying a bladder cancer in a subject as luminal-like or basal-like, the method comprising (a) receiving data indicative of expression of, and / or activation state of a gene regulatory region of: (i) one or more of FBN2, SLC4A4, CTTNBP2, KALRN, SELL, PMP22, UGT2B7, COL12A1, PTH2R, ANKRD36, DCDC2, PRTG, BPGM, PDE9A, tm<3ANXA3, and (n) one or more of BMP7, ERN2, SRPX2, UGT1A10, CLCA2, SPOCD1, SERPINB5, SEMA4B, CLCA4, CLU, UGT1A7, optionally (hi) one or more of NECTIN4,SOX4, FGFR3, TP63, KDM6A, KRT5, SPP1, IGFBP3, KMT2D, RBI, MYC, and BRCA2, in a sample collected from the subject; (b) obtaining a statistical model; (c) comparing the data to the statistical model to generate a score; and (d) outputting the score from the statistical model, wherein the score indicates the probability of the bladder cancer being luminal-like or basal-like. Additionally, the method may comprise receiving data indicative of expression of, and / or activation state of one or more gene regulatory regions selected from the group consisting of (i) CASTOR3, CHKA, UPK1A, and PLEKHF1, and (h) UGT1A1, and TP63, and (111) UGT1A8, SEMA4B, KRT80, KRT7, ERBB2, TACSTD2, ELK3, PSCA, EPCAM, EHF, TBX3, and ELF3.
[0021] For example, a computer-implemented method for classifying a bladder cancer in a subject as luminal-like or basal-like, the method comprising (a) receiving data indicative of expression of, and / or activation state of a gene regulatory region of: (i) one or more of CTTNBP2, KALRN, SELL, PMP22, COL12A1, PRTG, PDE9A, ANXA3, CASTOR3, CHKA, UPK1A, PLEKHF1, and (h) one or more of BMP7, ERN2, SRPX2, UGT1A10, CLCA2, SERPINB5, CLCA4, CLU, UGT1A1, SH3PXD2A, AQP3, ANXA10, SLITRK6, and TP63, and optionally (111) one or more of UGT1A8, SEMA4B, NECTIN4, KRT80, KRT7, ERBB2, TACSTD2, KDM6A, ELK3, PSCA, EPCAM, EHF, TBX3, ELF3, in a sample collected from the subject; (b) obtaining a statistical model; (c) comparing the data to the statistical model to generate a score; and (d) outputting the score from the statistical model, wherein the score indicates the probability of the bladder cancer being luminal-like or basal-like.
[0022] Also provided is a computer-implemented method for classifying a bladder cancer in a subject as micropapillary bladder cancer, the method comprising (a) receiving data indicative of expression of, and / or activation state of a gene regulatory region of: (i) one or more oiANPEP, BPGM, SPNS2, MAL, ASS1, COL4A4, MI / C4, TMC5, MMP7, PLCXD3, KRT80, PROMI, DAB2, KRT23, RHPN2, and optionally (11) one or more of UGT1A8, SEMA4B, NECTIN4, KRT80, KRT7, ERBB2, TACSTD2, KDM6A, ELK3, PSCA, EPCAM, EHF, TBX3, ELF3, in a sample collected from the subject; (b) obtaining a statistical model; (c) comparing the data to the statistical model to generate a score; and (d) outputting the score from the statistical model, wherein the score indicates the probability of the bladder cancer being micropapillary bladder cancer.BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Aspects will be described, by way of example, with reference to the following drawings.
[0024] FIG. 1 illustrates that a probe capture step enhances detection of the bladder cancer-associated promoter region of RBI by qPCR of anti-H3K27ac-enriched cfDNA isolated from a urine sample.
[0025] FIG. 2 illustrates the identification of three different clusters of urothelial carcinomas (UC) using unsupervised principal component analysis of cancer-associated nucleotide sequences obtained by sequencing anti-H3K27ac-enriched DNA isolated from tissue biopsies of patients suffering from UC. Tissue samples of UC with luminal like-inflammatory (LLI) characteristics (UC-LLI) were more closely associated with UC having a micropapillary component (MP) and could be distinguished from UC with basal-like (BL) characteristics (UC- BL).
[0026] FIG. 3 illustrates the sequencing peaks of nucleotide sequences of differentially activated gene regulatory sequences associated with UC-LLI or UC-BL. The nucleotide sequences were obtained by sequencing anti-H3K27ac-enriched DNA isolated from tissue biopsies of patients suffering from UC.
[0027] FIG. 4 illustrates that a probe capture step enhances detection of bladder cancer- associated gene regulatory regions of KDM6A (panel B), KMT2D (panel C), and KRT5 (panel D) by qPCR of anti-H3K27ac-enriched cfDNA isolated from a urine sample. Panel A shows detection of a housekeeping gene GAPDH.
[0028] FIG. 5 shows a volcano plot of differentially activated gene regulatory regions and associated genes for luminal bladder cancer (FIG. 5A) and basal bladder cancer (FIG. 5B). Both FIG. 5A and FIG. 5B are based on the same data set.
[0029] FIG. 6 shows a representative window of integrative genomics viewer (IGV) tracks from ChlP-seq data from bladder cancer tissues identified as luminal-like samples, micropapillary samples, basal-like samples, and healthy control samples. Chromatin signals are shown for the chromosomal region around the ERBB2 gene. Signals at the promoter areindicated as “A”, and signals at a nearby site exclusive to the bladder cancer samples is indicated as “B”.
[0030] FIG. 7 shows a heatmap of sequencing peaks from 8000 bladder cancer- associated gene regulatory regions detected in chromatin-associated DNA obtained from anti- H3K27ac-enriched cfDNA isolated from plasma (P), serum (S), and urine (U) pre- and postcapture.DETAILED DESCRIPTION
[0031] In order for the following description to be more readily understood, certain terms are first defined below. Additional definitions may be set forth throughout the specification.
[0032] Unless otherwise required by context, singular terms shall include pluralities and plural terms shall include the singular. Thus, as used in this specification and the appended claims, the singular forms “a”, “an” and “the” include plural referents unless the context clearly dictates otherwise. For example, “a nucleotide sequence” is understood to represent one or more nucleotide sequences. As such, the terms “a” (or “an”), “one or more,” and “at least one” can be used interchangeably herein.
[0033] Throughout this specification and aspects, the words “have” and “comprise”, or variations such as “has”, “having”, “comprises”, or “comprising”, will be understood to imply the inclusion of a stated integer or group of integers but not the exclusion of any other integer or group of integers.
[0034] “And / or” where used herein is to be taken as specific disclosure of each of the two specified features or components with or without the other. Thus, the term “and / or” as used in a phrase such as “A and / or B” herein is intended to include “A and B”, “A or B”, “A” (alone), and “B” (alone). Likewise, the term “and / or” as used in a phrase such as “A, B, and / or C” is intended to include “A and / or B and / or C” and to thus encompass each of the following aspects: A, B, and C; A, B, or C; A or C; A or B; B or C; A and C; A and B; B and C; A (alone); B (alone); and C (alone).
[0035] It is understood that wherever aspects are described herein with the language “comprising”, otherwise analogous aspects described in terms of “consisting of’ and / or“consisting essentially of’ are also provided. In other words, if an aspect is described as comprising A, B and C, aspects consisting essentially of A, B and C are also contemplated, as are aspects consisting of A, B and C.
[0036] The term “about” refers to an interval of accuracy that a person skilled in the art will understand to still ensure the technical effect of the feature in question. The term indicates a deviation from the indicated numerical value of ±10%. In some embodiments, the deviation is ±5% of the indicated numerical value. In certain embodiments, the deviation is ±1% of the indicated numerical value.
[0037] The terms “cancer”, “malignancy”, “neoplasm”, “tumor”, and “carcinoma” are used interchangeably to refer to a disease, disorder, or condition in which cells exhibit or exhibited relatively abnormal, uncontrolled, and / or autonomous growth, so that they display or displayed an abnormally elevated proliferation rate and / or aberrant growth phenotype.
[0038] The term “differentially activated regulatory regions” refers to chromatin- associated nucleotide sequences comprising at least one post-translational modification e.g., acetylation or methylation to histone proteins that drive the expression of downstream genes. Certain markers e.g., acetylation markers and methylation markers are associated with enhanced gene expression. Expression of certain genes is activated in some types of cancers. Such enhanced gene expression can be the result of post-translational modifications to histone proteins e.g., acetylation or methylation which renders the chromatin more accessible to the transcriptional machinery required for gene expression. The presence or increased level of such nucleotide sequences in enriched chromatin-associated DNA therefore is indicative of the presence of a cancer. Conversely, expression of certain genes is inactivated in some types of cancers. Such decreased gene expression can be the result of post-translational modifications to histone proteins e.g., deacetylation or demethylation which renders the chromatin less accessible to the transcriptional machinery required for gene expression. The absence of such nucleotide sequences in enriched chromatin-associated DNA therefore may also be indicative of the presence of a cancer.
[0039] The term “specifically hybridizes” refers to a process in which a nucleic acid strand anneals to and forms a stable duplex under suitable hybridization conditions with a secondcomplementary nucleic acid strand, and does not form a stable duplex with unrelated nucleic acid molecules under the same conditions. The formation of a duplex is accomplished by annealing two complementary nucleic acid strands in a hybridization reaction. The hybridization reaction can be made to be highly specific by adjustment of the hybridization conditions, such that hybridization between two nucleic acid strands will not form a stable duplex, unless the two nucleic acid strands contain sequences which are substantially or completely complementary to each other. Suitable hybridization conditions are readily determined for any given hybridization reaction. See, for example, Ausubel et al., Current Protocols in Molecular Biology, John Wiley & Sons, Inc., New York. Oligonucleotide probes and primers described herein may specifically hybridize to either strand of a cancer-associated nucleotide sequence. In the context of the present disclosure, each oligonucleotide probe binds to a distinct complementary region of the one or more cancer-associated nucleotide sequences disclosed herein. Likewise, each primer binds to a complementary region of the one or more cancer-associated nucleotide sequences disclosed herein. Typically, with respect to a specific cancer-associated nucleotide sequence, the complementary region bound by an oligonucleotide probe does not overlap with the regions bound by the primers that are used to amplify a portion of that specific cancer-associated nucleotide sequence.
[0040] The term “progression” is used to refer to monitoring changes in the presence and / or level of one or more cancer-associated nucleotide sequences disclosed herein. Such change can be used to describe an appearance or disappearance, or an increased or decreased level, of one or more cancer-associated nucleotide sequences disclosed herein.
[0041] The terms “diagnosis”, “detection”, or “screening” of / for a cancer includes determining whether a subject has or likely has cancer. For example, in some instances, further diagnostic tests (e.g., imaging or tissue biopsy) may be required to confirm the presence of a cancer. Diagnosis can include a determination relating to the risk, type, stage, malignancy, or other classification of the cancer. In some instances, a diagnosis can be or include a determination relating to prognosis and / or the probability of a response to one or more general or particular therapeutic agents or regimens.
[0042] Unless otherwise defined herein, technical and scientific terms used herein have the same meaning as commonly used and / or understood by one of ordinary skill in the art towhich this application belongs. In case of conflict, the present specification, including definitions, will control.
[0043] Generally, nomenclature used in connection with, and techniques of, cell and tissue culture, molecular biology, virology, immunology, microbiology, genetics, analytical chemistry, synthetic organic chemistry, medicinal and pharmaceutical chemistry, and protein and nucleic acid chemistry and hybridization described herein are those well-known and commonly used in the art. Enzymatic reactions and purification techniques are performed according to manufacturer’s specifications, as commonly accomplished in the art or as described herein.Method workflow
[0044] Methods are provided for the detection and / or monitoring of cancer that comprise the workflow set out in the following. Additionally, kits and devices are provided that can be used to implement the described workflow.
[0045] Briefly, the method workflow comprises a first step of isolating cfDNA (free and / or in extracellular vesicles such as exosomes). The isolation step is followed by a first enrichment step comprising immunoprecipitation to enrich the cfDNA for chromatin-associated DNA. Probe capture is then performed on the chromatin-associated DNA as a second enrichment step to further enrich for nucleotide sequences that are associated with presence of cancer. For example, probe capture may be performed with oligonucleotide probes that specifically hybridize to differentially activated regulatory regions of cancer-associated genes, wherein each oligonucleotide probe binds to a distinct region of the one or more nucleotide sequences. A detection step follows probe capture. The further enriched sample is contacted with one or more primer pairs to amplify and detect a portion of the one or more nucleotide sequences associated with the presence of cancer using qPCR. For instance, the primer pairs may bind to nucleotide sequences within the differentially activated regulatory regions of cancer-associated genes that have been enriched by the probe capture step.
[0046] The workflow can be adapted as needed for the desired assay readout. For example, the number of oligonucleotide probes and primer pairs can be adjusted for particular applications, e.g., to detect a specific cancer type or subtype. In some embodiments, a largenumber (e.g., 10,000 or more, 50,000, or 100,000 or more) of oligonucleotide probes is used to enrich a multitude of nucleotide sequences associated with the presence of many different types of cancers. Such a generic pan-cancer enrichment step (see Example 1) makes it possible to use the same workflow to detect multiple cancer types and subtypes. In other embodiments, oligonucleotide probe sets may be provided that hybridize specifically to nucleotide sequences that are associated with a specific type of cancer (e.g., bladder cancer).
[0047] The qPCR step is typically designed to detect a specific type and / or subtype of cancer. For example, a set of primer pairs may be provided to amplify and detect a portion of one or more nucleotide sequences associated with the presence of specific type or subtype of cancer (e.g., basal-like or luminal-like bladder cancer, or micropapillary bladder cancer).Sample
[0048] A sample obtained from a subject can be or may include cells, tissue, or bodily fluids. Typically, the sample is a liquid such as a bodily fluid. Liquid biopsies can often be obtained by non-invasive or minimally invasive means and therefore are easier to collect than solid tumor biopsies. “Bodily fluids” refer to fluids that are excreted or secreted from the body as well as fluids that are normally not (e.g., blood and blood plasma, Cowper’s fluid or pre- ejaculatory fluid, chyle, chyme, stool, interstitial fluid, intracellular fluid, lymph, menses, saliva, sebum, semen, serum, sweat, synovial fluid, tears, urine, vitreous humor, vomit).
[0049] A sample may be obtained by collecting blood, ascites, sputum, saliva, feces, urine, cerebrospinal fluid, peritoneal fluid, pleural fluid, lymph, gynecological fluids, secretions, or excretions. In some instances, bodily fluids may be obtained using vaginal swabs, oral swabs, nasal swabs, washings or lavages such as a ductal or bronchioalveolar lavages, aspirates etc. In some embodiments, a sample is or includes cells obtained from the subject.
[0050] The bodily fluid can be blood, plasma, serum, or urine. As it may be difficult to identify or biopsy the primary or original tumor in metastatic cancer, liquid biopsies can be particularly useful in such instances to support the diagnosis and / or monitoring of the cancer. The present methods, kits, and devices are therefore particularly advantageous in the detection of both solid tumors and metastatic cancer.
[0051] When cancer cells die and undergo apoptosis, DNA and / or chromatin is released into the blood and other bodily fluids as cell-free DNA (cfDNA). Using cfDNA for cancer detection, diagnosis, and / or monitoring is particularly advantageous for solid tumors because it avoids the need for an invasive tissue biopsy. cfDNA includes all extracellular DNA that may be freely circulating and / or present in exosomes. cfDNA associated with the presence of tumors may be present as circulating mono-nucleosomes (Thierry, Cell Genom. 2023; 3(1): 100242). Typically, the isolated chromatin associated DNA is mononucleosomic.
[0052] cfDNA can be isolated from a sample such as blood, urine or other bodily fluids by any means known in the art. Kits for the isolation of cfDNA are readily available commercially (e.g., QIAGEN’s cell-free DNA Extraction kits). Suitable methods for isolating cfDNA are described in, e.g., the QIAseq cfDNA Extraction Kit Quick-Start Protocol or the QIAseq cfDNA All-in-One Kit Handbook. Cells and cellular debris may be removed by centrifugation prior to cfDNA isolation. cfDNA may also be isolated from extracellular vesicles such as exosomes, e.g., by employing positive selection using microbeads recognizing the tetraspanin proteins CD9, CD63, and CD81 (as described, e.g., in instructions for the Miltenyi Biotec EV Isolation Kit). Isolating chromatin-associated cfDNA from extracellular vesicles such as exosomes and making this cfDNA available for chromatin immunoprecipitation is particularly advantageous in the context of the present methods, kits, and devices.Isolating chromatin-associated DNA
[0053] The methods described herein involve the step of enriching the cfDNA for chromatin-associated DNA. Chromatin immunoprecipitation (ChIP) can be performed as part of this enrichment step. ChIP usually involves cross-linking of chromatin-bound proteins (e.g., histones), e.g, using a crosslinking agent such as formaldehyde, followed by a fragmentation step (e.g, sonication or nuclease treatment) to obtain DNA fragments. Immunoprecipitation can be then carried out using a means for specifically binding chromatin, e.g., one or more antibodies that specifically bind to a chromatin-bound protein of interest. The DNA can be then released from the proteins and analyzed using various methods. X-ChIP methods utilize fixed chromatin fragmented by sonication, while the N-ChIP methods utilize native chromatin, which can be unfixed and nuclease digested.
[0054] Formaldehyde is a commonly used cross-linking agent. One advantage of using formaldehyde can be the ease of reversibility of the cross-links and its ability to form bonds that span a distance of approximately 2 angstroms. This means that formaldehyde can bind molecules in close association with each other. Other cross-linking agents that have been used include methylene blue, acridine orange, cisplatin, dimethylarsinic acid, potassium chromate, and ultraviolet (UV) light and lasers. The methods, kits, and devices can use any cross-linking agent known in the art.
[0055] The chromatin-associated DNA can be fragmented, e.g., using one or more sonication cycle(s), to obtain DNA fragments of 100-400 base pairs in size, e.g., 100-300 base pairs in size. An alternative fragment method to sonication can be nuclease digestion of the chromatin, e.g., in N-ChIP methods. The methods, kits, and devices can use any method known in the art for fragmenting DNA. In some embodiments, fragmentation (e.g., sonication or nuclease digestion) of the chromatic-associated DNA is not required.
[0056] Chromatin architecture, nucleosomal positioning, and access to DNA for gene transcription is largely controlled by histone proteins. A nucleosome comprises two identical subunits, each of which contains four histones (H2A, H2B, H3 and H4). The methods, kits, and devices comprise the immunoprecipitation of chromatin-associated DNA that is associated with modified histone proteins. Histone proteins may be modified in different ways which influence DNA interactions. Certain modifications correspond to the disruption of histone-DNA interaction, causing nucleosomes to unwind. In this “open” conformation, known as euchromatin, DNA is accessible for the binding of transcriptional machinery and subsequent gene activation. In contrast, other modifications can strengthen histone-DNA interactions which create a tightly packed chromatin structure called heterochromatin. In this compact form, transcriptional machinery cannot access DNA, resulting in gene silencing. Consequently, histone modification influences chromatin architecture and gene activation / inactivation. Histone proteins may be modified by methylation, acetylation and phosphorylation among other modifications.
[0057] Accordingly, a means for specifically binding chromatin for use in the methods, kits and devices can include binding a modified histone, such as an acetylated histone or a methylated histone e.g., histone H3. The means for specifically binding chromatin can be anantibody. For example, the methods may comprise contacting a sample with an antibody that specifically binds to an acetylated histone (e.g., acetylated histone H3). In some embodiments, the methods comprise contacting a sample with an antibody that specifically binds to histone H3 protein acetylated at the lysine at residue 27 (H3K27ac), an antibody that specifically binds to histone H3 protein acetylated at the lysine at residue 9 (H3K9ac), an antibody that specifically binds to histone H3 protein acetylated at the lysine at residue 14 (H3K14ac), an antibody that specifically binds to histone H3 protein acetylated at the lysine at residue 18 (H3K18ac), or an antibody that specifically binds to histone H3 protein acetylated at the lysine at residue 23 (H3K23ac), or combinations of two or more of such antibodies. In a specific embodiment, the means for specifically binding chromatin is an antibody that specifically binds to histone H3 protein acetylated at the lysine at residue 27 (H3K27ac).
[0058] The methods may also or alternatively comprise contacting a sample with an antibody that specifically binds to a methylated histone (e.g., methylated histone H3). For instance, the methods may comprise contacting a sample with an antibody that specifically binds to a histone methylated at the lysine at residue 4 (e.g., H3K4mel, H3K4me2, and / or H3K4me3), an antibody that specifically binds to a histone methylated at the lysine at residue 27 (H3K27me3), or an antibody that specifically binds to a histone methylated at the lysine at residue 9 (H3K9me3), or combinations of two or more of such antibodies. In a specific embodiment, the means for specifically binding chromatin is an antibody that specifically binds to histone H3 protein methylated at the lysine at residue 4 (H3K4me3).
[0059] Commercially available antibodies for use with the methods, kits, and devices include, but are not limited to, the anti-H3K27ac antibody C15410196, the anti-H3K4mel antibody ab8895, the anti-H3K4me2 antibody 07030MI, and the anti-H3K4me3 antibody ab8580.
[0060] Multiple antibodies may be used sequentially and / or simultaneously to isolate chromatin-associated DNA that is associated with a combination of histone modifications. For example, a first antibody used in a first ChIP may specifically bind to an acetylated histone, and a second antibody used in a second ChIP may specifically bind to a methylated histone. Alternatively, a first antibody used in a first ChIP may specifically bind to a histone acetylated a first position, and a second antibody used in a second ChIP may specifically bind to a histoneacetylated at a second position, wherein the first and second positions are different. For instance, a first ChIP with a first antibody that specifically binds to histone H3 protein acetylated at the lysine at residue 27 (H3K27ac), histone H3 protein acetylated at the lysine at residue 9 (H3K9ac), histone H3 protein acetylated at the lysine at residue 14 (H3K14ac), histone H3 protein acetylated at the lysine at residue 18 (H3K18ac), histone H3 protein acetylated at the lysine at residue 23 (H3K23ac), histone H3 protein methylated at the lysine at residue 4 (e.g., H3K4mel, H3K4me2, and / or H3K4me3), histone H3 protein methylated at the lysine at residue 27 (H3K27me3), or histone H3 protein methylated at the lysine at residue 9 (H3K9me3) may be followed by one or more subsequent ChIP using one or more of the antibodies not used for the first ChIP. Thus, the chromatin-associated DNA may comprise a combination of methylated and acetylated histone H3 proteins and / or histone H3 proteins that are acetylated at different positions.
[0061] A means for specifically binding chromatin for use in the present methods, kits and devices can be by binding a modified histone, such as an acetylated histone or a methylated histone e.g., histone H3. The means for specifically binding chromatin can be an antibody. The antibody can bind a group of histone modifications associated with increased transcription activation and / or an antibody that binds a group of histone modifications associated with increased transcription repression. For example, a pan-acetylation antibody, i.e., an antibody that recognizes acetylation across many histone proteins, and / or a pan-methylation antibody, i.e., an antibody that recognizes methylation across many histone proteins, may be used. In some embodiments, a combination of a pan-acetylation or pan-methylation antibody and a histone modification-specific antibody is used. For example, a pan-methylation antibody may be combined with an antibody that specifically binds to histone protein H3K27ac. Alternatively, a pan-acetylation antibody and a histone modification-specific antibody is used. For example, a pan-acetylation antibody may be combined with an antibody that specifically binds to histone H3K4me3. Or, histone modification specific antibodies that each bind to histone H3K27ac and H3K4me3, respectively, may be used in combination.
[0062] The chromatin-associated DNA can be isolated by immunoprecipitation. Immunosorbents commonly used to separate the antigen-antibody complex from a sampleinclude salmon sperm DNA-protein A-Sepharose®, protein G, magnetic beads, and other engineered immunoprecipitation systems known to those of skill in the art.
[0063] Once isolated, the chromatin associated DNA may be further purified and / or amplified. The amplification may be an isothermal amplification. Purification can be by any means known in the art, including organic extraction, silica-column-based techniques, ethanol precipitation, and anion-exchange chromatography.Identification of nucleotide sequences associated with cancer
[0064] Nucleotide sequences (e.g., gene regulatory regions) that are differentially activated in cancer can be identified by isolating chromatin-associated DNA from cancer tissue samples using the present methods, kits or devices. For example, soluble chromatin can be isolated and immunoprecipitated with an anti-H3K27ac antibody. A sequencing library of the immunoprecipitated chromatin-associated DNA (ChlP-seq library) can be constructed as described herein using methods known in the art. The ChlP-seq library can be sequenced using methods well known in the art, such as short-read sequencing technologies. Typically, to identify gene regulatory regions (e.g., promoters and enhancers) that are differentially activated in a cancer, ChlP-seq libraries of soluble chromatin isolated and immunoprecipitated with an anti- H3K27ac antibody are prepared from both a sample of cancer cells and a sample of healthy cells. Differentially activated gene regulatory regions can be identified using computational methods as described herein, for instance, in Example 3. Gene regulatory regions can include promoter regions, enhancer regions, transcription factor binding sites, polymerase binding sites, or any other nucleotide sequence that interacts with proteins involved in transcription.
[0065] Sequencing peaks that may represent differentially activated gene regulatory regions can be validated using multiple quality control criteria. For instance, a sequencing peak may be selected for further analysis if it represents a uniquely mappable read, i.e., a read that can only map to one location in the genome). Similarly, a sequencing peak may be selected for further analysis if it represents a uniquely mappable location, i.e., a location that can be mapped only by at least one read. A sequencing peak may be discarded if it overlaps with a Velcro region, i.e., a comprehensive set of consensus signal artifact regions in the human genome that have anomalous, unstructured high signal or read counts in next-generation sequencing experimentsindependent of cell line and of type of experiment. Typically, a sequencing peak is selected only if it meets certain threshold values, e.g. , comprises at least 5,000, at least 8,000, or at least 10,000 sequencing reads and is at least 5 -fold or at least 10-fold higher than background.
[0066] Nucleotide sequences that pass the quality control criteria may be further validated. For example, identified nucleotide sequences that are selected based on their association with the presence of cancer can be validated using annotated nucleotide sequence data provided in one or more curated databases, e.g., in the Encyclopedia of DNA Elements (ENCODE) Consortium database (encodeproject.org). For example, a nucleotide sequence may be selected if it corresponds to an evolutionarily conserved, DNase I hypersensitive region of the human genome found in one or more tumors (e.g., one or more solid tumors). A filtering step of the nucleotide sequences in ENCODE can be performed to select nucleotide sequences that overlap with a DNase I hypersensitive region of the human genome, e.g., by at least 70% or at least 80%. During this step, nucleotide sequences associated with solid tumors may be selected. Nucleotide sequences associated with hematological malignancies may be discarded.
[0067] To further refine the sequence selection, an evolutionary conservation analysis may be performed, e.g, by calculating conservation scores for each selected nucleotide sequence. It is believed that nucleotide sequences with similarity across species are more likely to be functional. A conservation score can be determined by using genome wide conservation scores provided by the Vertebrate Multiz Alignment and Conservation track software. The obtained conservation scores represent a probability that the base is conserved in the multialignment of a set of species. For example, the reference human genome hgl9 (GRCh37 Genome Reference Consortium Human Build 37; NCBI RefSeq assembly: GCF 000001405.13; GenBank assembly: GCA 000001405.1) can be used to identify an open reading frame and align the open reading frame with a corresponding open reading frame of one or more species (e.g, at least two, three, or four of rhesus monkey, mouse, dog, elephant, chicken, xenopus tropicalis, and zebrafish). Alternatively, the reference human genome hg38 (GRCh38.14 Genome Reference Consortium Human Build 38.14; NCBI RefSeq assembly: GCF 000001405.40; GenBank assembly: GCA_000001405.29) can be used.
[0068] For example, phastCons scores for the set of mammalian species averaging over the regions may be used where the top-most conserved regions are selected. Nucleotidesequences may be assigned a “low” phastCons score if the value is below a threshold value, e.g., below about 0.95 or less, about 0.9 or less, and about 0.85 or less may, and removed from further analysis. The threshold value can be selected to give a size for the top set of about 100,000 nucleotide sequences. Accordingly, a phastCons score threshold value of about 0.8 or less, about 0.7 or less, or about 0.6 or less may be acceptable.
[0069] In some instances, conservation scores are merely used to rank evolutionarily conserved nucleotide sequences to select the most conserved nucleotide sequences. For example, only the top 10,000, top 50,000, or top 100,000 most conserved nucleotide sequences may be selected for subsequent analysis. A similarity-based analysis may be performed to remove any repetitive (or near-repetitive) nucleotide sequences.
[0070] Regulatory regions of cancer-associated genes that can be used in the present methods, kits or devices can be c / .s-regulatory elements, as identified by ENCODE. Chromosomal distribution and evolutionary conservation of DNase I hypersensitive regions and c / .s-regulatory elements are similar. In particular, it has been shown that c / .s-regulatory elements are anchored to DNase I hypersensitive regions and can be 150-350 base pairs in length. Cis- regulatory elements can be classified based on their function-associated signatures (ENCODE Project Consortium et al., Nature 2020; 583(7818):699-710). A c / .s-regulatory element (i.e., a regulatory region of a cancer-associated gene) of interest can be an element with an enhancerlike signature that is in a DNase I hypersensitive region and associated with H3K27ac (e.g., as provided by ENCODE). Regulatory regions of cancer-associated genes may contain a proximal enhancer-like signature or a distal enhancer-like signature, i.e., the regions may be within 2 kb of a transcription start site (TSS) or more than 2 kb removed from the closest TSS, respectively. In some instances, a regulatory region of a cancer-associated gene is excluded for use with the present methods, kits or devices if it is located within 2000 bp of a transcription start site (TSS) and has a high H3K4me3 signal.
[0071] Once candidate nucleotide sequences associated with the presence of cancer or a cancer subtype (e.g., regulatory regions of cancer-associated genes) have been identified, the nucleotide sequences may be screened for use with a particular sample type, e.g., cfDNA obtained from a liquid biopsy such as a blood sample. In some instances, candidate nucleotide sequences that yield a low signal-to-noise ratio against the background of cfDNA derived fromhealthy cells may be removed. For example, as described herein, a comparison of signals at gene regulatory regions of the top differentially expressed genes obtained from ChlP-seq data can be used to identify nucleotide sequences that are present in the cancer type (and / or subtype) of interest but not in a control sample, such as plasma samples from healthy subjects.
[0072] Typically, in the methods, kits or devices, the nucleic acid sequences associated with the presence of cancer are nucleosome-sized DNA fragments. For example, the one or more nucleotide sequences associated with the presence of cancer may each be 50-300bp, 100-250bp, 150-200bp. The one or more nucleotide sequences associated with the presence of cancer may each have a length of 100-250bp. In a specific embodiment, the nucleic acid sequences associated with the presence of cancer are about 150bp in length.
[0073] Using this approach, chromosomal regions associated with different cancers were identified that can be used with the present methods, kits or devices. Suitable chromosomal regions associated with bladder cancer, include those set forth in SEQ ID NOs: 1-6372 and 9792- 9808. Suitable chromosomal regions associated with bladder cancer having a luminal phenotype include those set forth in SEQ ID NOs: 6373-7634 and 9809-9820. Suitable chromosomal regions associated with bladder cancer having a basal phenotype include those set forth in SEQ ID NOs: 7635-9733 and 9821-9832. Suitable chromosomal regions associated with bladder cancer having a micropapillary phenotype include those set forth in SEQ ID NOs: 7172 and 9833-9847. The corresponding chromosome number can be identified by aligning any one of SEQ ID NOs: 1-9733 and 9792-9847 against the reference human genome, hgl9 (GRCh37 Genome Reference Consortium Human Build 37; NCBI RefSeq assembly: GCF 000001405.13; GenBank assembly: GCA_000001405.1).Oligonucleotide probe-based enrichment
[0074] The present methods, kits or devices described herein involve the step of enriching the isolated chromatin-associated DNA using oligonucleotide probes (DNA or RNA) that specifically hybridize to one or more nucleotide sequences associated with the presence of cancer, wherein each oligonucleotide probe binds to a distinct region of the one or more nucleotide sequences. For example, the oligonucleotide probes may be DNA oligonucleotideprobes. The use of RNA oligonucleotide probes may be advantageous because RNA oligonucleotides can easily be removed by nuclease digestion post-capture.
[0075] The nucleotide sequences may be associated with modified histone proteins (e.g. , acetylated or methylated histone proteins), as described herein. The association of the enriched isolated DNA with modified histone proteins such acetylated histone H3 may indicate that the isolated chromatin-associated DNA comprises gene regulatory regions which are differentially activated and that the genes downstream of the regulatory regions are expressed. Accordingly, the oligonucleotide probes can be designed to specifically hybridize within regulatory regions of cancer-associated genes. For example, one or more oligonucleotide probes may specifically hybridize to a regulatory region of a gene that is differentially activated in a cancer cell.
[0076] A set of oligonucleotide probes may specifically hybridize to regulatory regions of cancer-associated genes that are associated with a specific cancer type or subtype. The present methods, kits or devices are applicable to any type of cancer (especially solid cancers), including, e.g., cancers of the bladder, breast, prostate, liver, kidney, lung, pancreas, gallbladder, skin, bone, adrenal gland, brain, heart, intestine, ovary, cervix, uterus, testicles, and colon, as well as their associated subtypes. For example, the regulatory nucleotide sequences associated with the presence of cancer may be regulatory regions of cancer-associated genes that are associated with the presence of a specific cancer type such as bladder cancer. Exemplary nucleotide sequences associated with the presence of bladder cancer are set forth in SEQ ID NOs: 1-6372. Accordingly, in some embodiments, each of the oligonucleotides probes specifically hybridize to a distinct region within one or more of the nucleotide sequences set forth in SEQ ID NOs: 1- 6372, or one or more complementary nucleotide sequences thereof. A set of oligonucleotide probes for use in the methods, kits, or devices described herein may specifically hybridize to distinct regions of one or more (e.g., two, three, four, five, six, seven, eight, or all) nucleotide sequences as set forth in Table 1.Table 1A. Exemplary chromosomal regions associated with bladder cancer
[0077] Further genes associated with bladder cancer include KMT2D, RBI, MYC and BRCA2. Accordingly, additional non-limiting examples of cancer-associated nucleotide sequences for use with the present methods, kits or devices include the gene regulatory regions of KMT2D (e.g., the chromosomal region of SEQ ID NO: 6369), RBI (e.g., the chromosomal region of SEQ ID NO: 6370), MYC (e.g, the chromosomal region of SEQ ID NO: 6371) and BRCA2 (e.g, the chromosomal region of SEQ ID NO: 6372).
[0078] Further exemplary nucleotide sequences associated with the presence of bladder cancer are set forth in SEQ ID NOs: 9792-9808. Accordingly, in some embodiments, each of the oligonucleotides probes specifically hybridize to a distinct region within one or more of the nucleotide sequences set forth in SEQ ID NOs: 9792-9808, or one or more complementary nucleotide sequences thereof. A set of oligonucleotide probes for use in the methods, kits, or devices described herein may specifically hybridize to distinct regions of one or more (e.g., two, three, four, five, six, seven, eight, or all) nucleotide sequences as set forth in Tables IB. Some of the nucleotide sequences associated with the presence of bladder cancer may be within a promoter region or an enhancer region. For example, the nucleotide sequences within the gene regulatory regions of SEMA4B, NECTIN4, ERBB2, TACSTD2, ELK3, PSCA, and TBX3 may be within an enhancer region. Alternatively, the nucleotide sequences within the gene regulatory regions of NECTIN4 may be within a promoter region. A nucleotide sequence associated withthe presence of bladder cancer may be within an intron. For example, the nucleotide sequence within the gene regulatory region of EPCAM may be within an intron.Table IB. Further exemplary chromosomal regions associated with bladder cancer
[0079] The nucleotide sequences associated with the presence of cancer may be regulatory regions of cancer-associated genes that are associated with the presence of a specific subtype of a cancer, e.g., luminal or basal subtypes of cancer, such as luminal or basal bladdercancer. For example, gene regulatory sequences associated with luminal-like bladder cancer include nucleotide sequences comprising an IRF2 DNA-binding motif. Gene regulatory sequences associated with luminal-like bladder cancers with a micropapillary component comprise a GRHL2 DNA-binding motif. Gene regulatory sequences associated with basal-like bladder cancer include nucleotide sequences comprising a TP63 DNA-binding motif. Such motifs can be identified within differentially activated gene regulatory sequences associated with bladder cancer using readily available computational tools and databases (see, e.g., Example 4).
[0080] Exemplary nucleotide sequences associated with the presence of luminal bladder cancer are set forth in SEQ ID NOs: 6373-7634. Exemplary nucleotide sequences associated with the presence of basal bladder cancer are set forth in SEQ ID NOs: 7635-9733. Accordingly, in specific embodiments, each of the oligonucleotide probes specifically hybridize to a distinct region within one or more of the nucleotide sequences set forth in SEQ ID NOs: 6373-7634 or 7635-9733, or one or more complementary nucleotide sequences thereof. A set of oligonucleotide probes for use in the methods, kits, or devices described herein may specifically hybridize to distinct regions of one or more (e.g., two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, fourteen, or all) nucleotide sequences as set forth in Tables 2A and 3 A.
[0081] Further exemplary nucleotide sequences associated with the presence of luminal bladder cancer are set forth in SEQ ID NOs: 9809-9820. Further exemplary nucleotide sequences associated with the presence of basal bladder cancer are set forth in SEQ ID NOs: 9821-9832. Accordingly, in specific embodiments, each of the oligonucleotide probes specifically hybridize to a distinct region within one or more of the nucleotide sequences set forth in SEQ ID NOs: 9809-9820 or 9821-9832, or one or more complementary nucleotide sequences thereof. A set of oligonucleotide probes for use in the methods, kits, or devices described herein may specifically hybridize to distinct regions of one or more (e.g., two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, fourteen, or all) nucleotide sequences as set forth in Tables 2B and 3B. Some of the nucleotide sequences associated with the presence of luminal and basal bladder cancer may be within a promoter region or an enhancer region. For example, thenucleotide sequences within the gene regulatory regions of SELL, UPK1A, PLEKHF1 and TP63 may be within an enhancer region.Table 2A. Exemplary chromosomal regions associated with luminal bladder cancerTable 2B. Further exemplary chromosomal regions associated with luminal bladder cancerTable 3A. Exemplary chromosomal regions associated with basal bladder cancerTable 3B. Further exemplary chromosomal regions associated with basal bladder cancer
[0082] Exemplary nucleotide sequences associated with the presence of micropapillary bladder cancer are set forth in SEQ ID NOs: 7172 and 9833-9847. Accordingly, in specific embodiments, each of the oligonucleotide probes specifically hybridize to a distinct region within one or more of the nucleotide sequences set forth in SEQ ID NOs: 7172 and 9833-9847, or one or more complementary nucleotide sequences thereof. A set of oligonucleotide probes for use in the methods, kits, or devices described herein may specifically hybridize to distinct regions of one or more (e.g., two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, fourteen, or all) nucleotide sequences as set forth in Table 4. Some of the nucleotide sequences associated with the presence of micropapillary bladder cancer may be within a promoter region or an enhancer region. For example, the nucleotide sequences within the gene regulatory regions (AB PGM. SPNS2, MAL, MMP7, KRT80, PROMI, DAB 2, and KRT23 may be within an enhancer region.Table 4. Exemplary chromosomal regions associated with micropapillary bladder cancer
[0083] The oligonucleotide probes may specifically hybridize to regulatory regions of cancer-associated genes that are associated with multiple different cancer types or subtypes, including those set forth in SEQ ID NOs: 1-9733, or one or more complementary nucleotide sequences thereof. Accordingly, in some embodiments, each of the oligonucleotides probes specifically hybridize to a region within the nucleotide sequences set forth in SEQ ID NOs: 1- 9733. The methods, kits, and devices may use more than one oligonucleotide probes that specifically hybridize to different regions within the same nucleotide sequences as set forth in SEQ ID NOs: 1-9733. Alternatively, the methods, kits, and devices may use oligonucleotide probes that each specifically hybridize to regions within different nucleotide sequences as set forth in SEQ ID NOs: 1-9733.
[0084] Alternatively or additionally, the oligonucleotide probes may specifically hybridize to regulatory regions of cancer-associated genes that are associated with multiple different cancer types or subtypes, including those set forth in SEQ ID NOs: 7172 and 9792- 9847, or one or more complementary nucleotide sequences thereof. Accordingly, in some embodiments, each of the oligonucleotides probes specifically hybridize to a region within the nucleotide sequences set forth in SEQ ID NOs: 7172 and 9792-9847. The methods, kits, and devices may use more than one oligonucleotide probes that specifically hybridize to different regions within the same nucleotide sequences as set forth in SEQ ID NOs: 7172 and 9792-9847. Alternatively, the methods, kits, and devices may use oligonucleotide probes that each specifically hybridize to regions within different nucleotide sequences as set forth in SEQ ID NOs: 7172 and 9792-9847.
[0085] Oligonucleotide probe design parameters are well known in the art, as are means to optimize oligonucleotide probe design. For instance, each of the oligonucleotide probes may have a length of less than about 50 nucleotides, less than about 40 nucleotides, less than about30 nucleotides, less than about 20 nucleotides. Each of the oligonucleotide probes may have a length of about 20-30 nucleotides. Each of the oligonucleotide probes may have a melting temperature (Tm) of 40-60°C.
[0086] Assay sensitivity and / or specificity can be improved by using a multitude of oligonucleotide probes. Moreover, using a multitude of probes has the additional advantage that the enriched sample can be interrogated for many nucleotide sequences associated with cancer at the same time (multiplexing). Accordingly, a pool of oligonucleotide probes may be used that specifically hybridize to a plurality of nucleotide sequences associated with cancer (e.g., regulatory regions of a plurality of cancer-associated genes differentially expressed in cancer). A pool of oligonucleotide probes for use with the present methods, kits or devices may specifically hybridize to up to about 100,000 individual nucleotide sequences that are associated with cancer. For example, more than one (e.g., 5, 10, 50, or 100) oligonucleotide probes may be provided that specifically hybridizes to more than one location in the regulatory region of a gene differentially expressed in a cancer cell. A pool of up to about 100,000 oligonucleotide probes that specifically hybridize to up to about 100,000 individual nucleotide sequences can comprehensively enrich regulatory regions of cancer-associated genes associated with solid cancers characterized to date with high specificity and sensitivity.
[0087] A skilled artisan will appreciate that pools with fewer oligonucleotide probes can be used for more specific applications, e.g, to detect, diagnose, and / or monitor specific types and / or subtypes of cancer. For example, for specific applications, a pool of about 10-100 oligonucleotide probes, for example 10-50, 10-40, 10-30 or 10-20 that specifically hybridize to individual nucleotides sequence may be adequate. Accordingly, a pool of about 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 of oligonucleotide probes may be used, depending on the specific type of assay. For example, a pool of 10-20 oligonucleotide probes (e.g., 10, 15 or 20 oligonucleotide probes) may be a suitable number for an assay using the method described herein to determine a cancer type or subtype in a liquid biopsy sample such as serum or urine. In some instances, the pool of oligonucleotides probes specifically hybridize to 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 of the nucleotide sequences set forth in SEQ ID NOs: 1-9733. Alternatively or additionally, the pool of oligonucleotides probes may specifically hybridize to 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 of the nucleotide sequences set forth in SEQ ID NOs: 7172 and 9792-9847, or one ormore complementary nucleotide sequences thereof and can be used to diagnose and / or classify bladder cancer.
[0088] The number of oligonucleotide probes that specifically hybridize to individual cancer-associated nucleotides sequences in the pool is determined by the number of regulatory regions of cancer-associated genes that are detected by qPCR. For example, up to about 100 qPCRs may be performed in parallel in a single run in a microfluidic device. Some commercial applications are designed to perform about 5-20 qPCRs (e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20) in parallel in a single run. Accordingly, a pool may comprise about 5-20 oligonucleotide probes that specifically hybridize to individual nucleotides sequences associated with the presence of cancer. It may be desirable to provide more than one oligonucleotide probe per gene regulatory region, e.g., at least 2, 3, 4, 5, 6, 7, 8, 9, or 10 oligonucleotide probes per gene regulatory region. Therefore, a pool may comprise about 10-200 oligonucleotide probes that specifically hybridize to individual nucleotides sequences associated with the presence of cancer.
[0089] As noted above, it may be desirable to perform the enrichment step with a pool of oligonucleotide probes that enriches regulatory regions of cancer-associated genes associated with a multitude of different cancers, e.g., a multitude of different solid cancers. This has the advantage that the same set of oligonucleotide probes can be used for sample preparation, even though the specific primer sets used to perform the qPCR detect only nucleotide sequences associated with one specific type or subtype of cancer (e.g., bladder cancer, or subtypes thereof such a basal or luminal bladder cancer, or micropapillary bladder cancer). This allows the user to fine tune the set of primer pairs used for qPCR without any change to the enrichment platform used for sample preparation. Accordingly, the pool of oligonucleotide probes may comprise up to about 500, 1000, 5000, 10,000, 20,000, 50,000 or 100,000 probes. For example, the pool of oligonucleotide probes may comprise 500-100,000, 1,000-50,000, or 5,000-20,000 oligonucleotide probes. Each pool of oligonucleotide probes may specifically hybridize to individual nucleotide sequences associated with multiple cancers (e.g., multiple cancer types and / or subtypes). For example, the pool of oligonucleotides probes specifically hybridize to 500, 1000, 5000, or all of the nucleotide sequences set forth in SEQ ID NOs: 1-9733, or one or more complementary nucleotide sequences thereof. Alternatively or additionally, the pool ofoligonucleotides probes specifically hybridize to 500, 1000, 5000, or all of the nucleotide sequences set forth in SEQ ID NOs: 7172 and 9792-9847, or one or more complementary nucleotide sequences thereof
[0090] Kits for enriching isolated chromatin-associated DNA using oligonucleotide probes are commercially available from Agilent (e.g., SureSelect), Twist Bioscience, and Integrated DNA Technologies. The optimization of hybridization conditions is conventional and known in the art and can be achieved by following manufacturer instructions. In a specific example, each custom Sure-Select target enrichment kit contains a mixture of custom Sure- Select RNA oligonucleotide probes that are biotinylated for easy capture using streptavidin - conjugated magnetic beads. RNA oligonucleotide probes are hybridized with a DNA library prepared using Illumina’s Genomic DNA Sample Prep Kit. RNA-oligonucleotide probe-DNA hybrids are then isolated from the complex mixture using streptavidin-conjugated magnetic beads. The RNA oligonucleotide probe is then digested before the next analysis step (e.g. , qPCR) of the probe capture-enriched DNA.
[0091] Once enriched for the one or more nucleotide sequences associated with the presence of cancer, the enriched chromatin-associated DNA may be purified. Purification can be by any means known in the art, including organic extraction, silica-column-based techniques, ethanol precipitation, and anion-exchange chromatography.
[0092] The enriched chromatin-associated DNA may be stored as a library and / or subject to further analysis.Libraries
[0093] Also provided herein is a library of isolated cell-free chromatin-associated DNA comprising nucleotide sequences associated with the presence of cancer, wherein the nucleotide sequences are enriched for regulatory regions of cancer-associated genes. The library may be obtained by the methods described herein.
[0094] Also provided herein are methods of preparing a library of isolated cell-free chromatin-associated DNA comprising nucleotide sequences associated with the presence of cancer, said method comprising: providing a sample obtained from a subject; isolating cell-free DNA (cfDNA) from the sample; contacting the cfDNA with a means for specifically bindingchromatin (e.g., an antibody that specifically binds acetylated or methylated histones such as an anti-H3K23ac antibody) to isolate chromatin-associated DNA; enriching the isolated chromatin- associated DNA using oligonucleotide probes that specifically hybridize to one or more nucleotide sequences associated with the presence of cancer (e.g. , one or more regulatory regions of cancer-associated genes), wherein each oligonucleotide probe binds to a distinct region of the one or more nucleotide sequences; and optionally amplifying the chromatin-associated DNA, e.g., using an isothermal amplification reaction. qPCR-based detection
[0095] The one or more nucleotide sequences associated with the presence of cancer are amplified and detected by quantitative PCR (qPCR). To this end, the isolated chromatin- associated DNA enriched for one or more nucleotide sequences associated with the presence of cancer is contacted with one or more primer pairs. In some instances, the enriched isolated chromatin-associated DNA may be divided into several aliquots. For each individual aliquot, qPCR is then performed using one or more cancer type-specific primer pairs. Both aliquoting of the sample and adding of primers can be done on an integrated microfluidic device as is known in the art. The aliquots can be contacted with primer pairs to detect nucleotides sequences associated with different cancer types or subtypes. For example, a first aliquot may be contacted with one or more primer pairs to detect nucleotide sequences associated with a first cancer, a second aliquot may be contacted with one or more primer pairs to detect a second cancer, a third aliquot may contacted with one or more primer pairs to detect a nucleotide sequence associated with a third cancer, etc. Accordingly, the present methods, kits or devices can be used as part of a pan-cancer detection strategy.
[0096] Alternatively, a first aliquot may be contacted with one or more primer pairs to detect a nucleotide sequence associated with a type of cancer, e.g, bladder cancer, a second aliquot may be contacted with one or more primer pairs to detect a subtype of the cancer, e.g, luminal bladder cancer. Further, a third aliquot may be contacted with one or more primer pairs to detect a different subtype of the cancer, e.g., basal bladder cancer. In some instances, a subject may have received a cancer diagnosis e.g., bladder cancer. A first aliquot may be contacted with one or more primer pairs to detect a nucleotide sequence associated with a first subtype (e.g., luminal bladder cancer) of the diagnosed cancer, a second aliquot may be contacted with one ormore primer pairs to detect a second subtype (e.g., basal bladder cancer) of the diagnosed cancer, a third aliquot may be contacted with one or more primer pairs to detect a third subtype (e.g., micropapillary bladder cancer) of the diagnosed cancer, etc. Accordingly, the present methods, kits or devices can be used as part of a cancer typing strategy.
[0097] qPCR can be carried out by any method known in the art, using any suitable conditions that result in a PCR product. Suitable protocols are disclosed in Heid et al. Genome Res. 1996; 6(10): 986-94 and the commercially available TaqMan® Gene Expression Assays. The PCR product can be detected by any means known in the art. Typically, a fluorescent reporter dye is used as an indirect measure of the amount of nucleic acid present during each amplification cycle of the qPCR reaction. The increase in fluorescent signal is directly proportional to the quantity of exponentially accumulating PCR product molecules produced during the amplification reaction.
[0098] Suitable qPCR systems for diagnostic applications include, e.g., the Cepheid Gene Xpert and Taqman Gene Expression Array cards and plates from ThermoFisher Scientific. Using qPCR for detection of genomic DNA sequences associated with the presence of cancer provides advantages in terms of making the present methods, kits or devices accessible and rapid. For example, avoiding the need for sequencing and / or RT-PCR allows the present methods kits or devices to be widely used in a variety of settings without a centralized processing step carried out off site or in a specialist facility.
[0099] The present methods, kits or devices may be used for the detection of multiple nucleotide sequences that are associated with presence of cancer. The present methods, kits, and devices may be used for the simultaneous detection of up to 1000 nucleotide sequences that are associated with the presence of cancer, for example 1-6, 1-8, 1-12, 1-24, 1-48, 1-100, 1-256 or more nucleotide sequences that are associated with the presence of cancer. Using, e.g., a microfluidic device, the methods, kits or devices may be used for the simultaneous qPCR detection of up to 100 nucleotide sequences that are associated with the presence of cancer. Or, up to 12 (e.g., 6-10) nucleotide sequences that are associated with the presence of cancer are detected by qPCR simultaneously, e.g, by multiplexing.
[0100] The primer pairs may be specific for one or more nucleotide sequences associated with a specific type of cancer such as bladder cancer. For example, the primer pairs may be specific for one or more gene sequences or regulatory sequences associated with a specific type or subtype of cancer. The primer pairs may be specific for one or more nucleotide sequences associated with bladder cancer. The primer pairs may be specific for one or more gene sequences or regulatory sequences associated with bladder cancer.
[0101] Depending on the format, 6, 9, 12, or 16 individual primer pairs may be used in a single run. Each of the primers may have a length of between 15-25 nucleotides. Each of the primers may have a G / C content of at least 40%. As described herein, oligonucleotides for use as primers are typically 18-24 bases long; a 40-60% G / C content; optionally start and end with 1-2 G / C pairs; have a Tm of 50-60°C and primer pairs should have a Tm within 5°C of each other.
[0102] The primer pair is typically designed to yield an amplicon of 100-250 base pairs in length. Accordingly, the amplified portion of the one or more nucleotide sequences associated with cancer may be no more than 250 base pairs in length. For example, the amplified portion may be between 100 and 250 (e.g., 100 and 200) base pairs in length.
[0103] Specific sets of primer pairs can be used to amplify (a) portion(s) of the one or more nucleotide sequences associated with the presence of specific cancers e.g., bladder cancer. Exemplary primer pairs for the detection of nucleotide sequences in the regulatory regions of NECTIN4, SOX4, FGFR3, TP63, KDM6A, KRT5, SPP1 and IGFBP3 are provided in Table 5.Table 5. Exemplary primer pairs for bladder cancer
[0104] Further genes associated with bladder cancer include KMT2D, RBI, MYC and BRCA2. Accordingly, additional non-limiting examples of primer pairs include those directed to bladder cancer-associated genes such as KMT2D (SEQ ID NOs: 9752 and 9753), RBI (SEQ ID NOs: 9754 and 9755), MYC (SEQ ID NOs: 9756 and 9757) and BRCA2 (SEQ ID NOs: 9758 and 9759).Analysis of qPCR data
[0105] In the methods, kits or devices, the presence or increased level of one or more nucleotide sequences associated with the presence of cancer in a sample are indicative of the subject suffering from cancer.
[0106] The presence of one or more nucleotide sequences associated with the presence of cancer will therefore be indicative of the presence of cancer. For example, the presence of one or more nucleotide sequences associated with cancer may be determined by a threshold cycle (Ct) value. For example, the detection of an amplification product for a nucleotide sequence associated with the presence of cancer at a Ct value of less than 40 indicates the presence of the nucleotide sequence in the sample. In some instances, the Ct value may be 35 or less, e.g., 30 or less, 25 or less, or 20 or less. In the absence of cancer, the level of the one or more nucleotidesequences associated with the presence of cancer is below the detection limit for qPCR at the chosen Ct value.
[0107] To enable comparison between different samples (e.g., a patient sample and a control sample from a healthy individual), the levels of the one or more nucleotide sequences associated with the presence of cancer are typically normalized to one or more nucleotide sequences whose levels do not significantly fluctuate. For example, regulatory regions of housekeeping genes are typically associated with acetylated histones as they are constantly expressed. Commonly used housekeeping genes in qPCR include beta actin (ACTB), glyceraldeyde-3 -phosphate dehydrogenase (GAPDH), ribosome small subunit (18S) ribosomal RNA (rRNA), Ubiquitin C (UBC), hypoxanthine guanine phosphoribosyl transferase (HPRT), succinate dehydrogenase complex, subunit A (SDHA) and Tyrosine 3- monooxygenase / tryptophan 5-monooxygenase activation protein, zeta polypeptide (YWHAZ) (Silver et al. BMC Mol Biol. 2008: 9:64). Alternatively, where the presence of one or more nucleotide sequences associated with the presence of cancer is determined by a threshold Ct value, the amplification of regulatory regions of one or more housekeeping genes can serve as an internal control (i.e., to confirm that the qPCR was successful).
[0108] In some circumstances, the level of one or more nucleotide sequences associated with the presence of cancer may change in the presence of cancer. The control sample may be from a subject without cancer. Or, the control sample is a sample from a subject obtained at an earlier time point (e.g., prior to progression such as the development of metastatic cancer). Or, the level of a nucleotide sequence associated with the presence of cancer is compared to the corresponding level of the nucleotide sequence in a control sample (e.g., from a healthy individual). An increased level of the nucleotide sequence associated with the presence of cancer relative to the corresponding nucleotide sequence in the control sample may indicative of the presence of cancer. Or, a decreased level of the nucleotide sequence associated with the presence of cancer relative to the corresponding nucleotide sequence in the control sample may indicative of the presence of cancer.
[0109] The level of the nucleotide sequence associated with the presence of cancer may be at least 3 times greater (e.g., at least 4 times, at least 5 times, at least 10 times, or at least 100 times greater) than the level of the corresponding nucleotide sequence in the control sample. Or,the level of the nucleotide sequence associated with the presence of cancer may be at least 10 times greater (e.g., at least 20 times, at least 50 times, at least 100 times, at least 1000 times greater) than the level of the corresponding nucleotide sequence in the control sample.
[0110] Alternatively, the level of the nucleotide sequence associated with the presence of cancer may be at least 3 times lower (e.g., at least 4 times, at least 5 times, at least 10 times, or at least 100 times lower) than the level of the corresponding nucleotide sequence in the control sample. Or, the level of the nucleotide sequence associated with the presence of cancer may be at least 10 times lower (e.g., at least 20 times, at least 50 times, at least 100 times, at least 1000 times lower) than the level of the corresponding nucleotide sequence in the control sample.Applications
[0111] The present methods, kits or devices can be used to detect, diagnose, and / or monitor a cancer, or a specific type / subtype of a cancer, in a subject. Cancer is a neoplastic disease that can be or include cells that are benign, malignant, pre-metastatic, metastatic, and / or non-metastatic. A cancer can include one or more tumors. Typically, the cancer detected or diagnosed by the present methods, kits or devices is a solid tumor. Different types of cancers that form solid tumors are known in the art and include, for example, colorectal cancer, sarcomas, melanomas, adenomas, solid tissue carcinomas, squamous cell carcinomas (e.g., of the mouth, throat, larynx, or lung), liver cancer, genitourinary cancers such as prostate, cervical, bladder, uterine, and endometrial cancer and renal cell carcinomas, bone cancer, pancreatic cancer, skin cancer, cutaneous or intraocular melanoma, cancer of the endocrine system, cancer of the thyroid gland, cancer of the parathyroid gland, head and neck cancers, breast cancer, gastro-intestinal cancers and nervous system cancers, benign lesions such as a papilloma, and the like. The present methods, kits or devices have been illustrated using bladder cancer, and luminal, basal and micropapillary subtypes thereof, as an example.
[0112] The presence of cancer, especially an at early stage, may be associated with vague and non-discriminatory symptoms e.g., weight loss and fatigue, which can delay a timely diagnosis. The present methods, kits, and devices may be used to support the diagnosis of cancer, and due to their sensitivity and specificity may allow an earlier diagnosis of the type and / or subtype of cancer. In some instances, the methods, kits, and devices may support identificationof the primary tumor in a subject with metastatic disease to determine a suitable course of treatment.
[0113] The present methods, kits and devices are useful for the monitoring and / or detection of cancer including new occurrences or recurrences. The methods, kits or devices can allow for early detection of new occurrences or recurrences of cancer from a liquid biopsy without the need to identify a new or recurrent tumor location for a solid biopsy.Monitorins cancer progression and / or treatment response
[0114] In another aspect, there is provided a method for monitoring the progression of a cancer, the method comprising: detecting, from a sample (e.g., plasma, urine, or serum) obtained from a subject at a first time point, the presence and / or level of one or more nucleotide sequences associated with the presence of cancer; detecting, from a sample obtained from a subject at a second or subsequent time point, the presence and / or level of the one or more nucleotide sequences associated with the presence of cancer; and comparing the results from the first and second time points, thereby monitoring the progression of the cancer, wherein one or both of the measuring steps are carried out according to a method described herein. A difference in the level of the one or more nucleotide sequences associated with the presence of cancer between the subsequent time points can indicate the progression of the tumor. For example, an increased level between the first and second or subsequent time point may indicate tumor growth and a decreased level between the first and subsequent time point may indicate tumor shrinkage.
[0115] The first time point can be before the start of treatment for the cancer and the second time point can be during or after treatment for the cancer. The first time point can be after treatment for cancer, and the second time point can be up to one month, 2 months, 3 months, 4 months, 5 months, 6 months, 9 months, 12 months or more, 1 year, 2 years, 3 years, 4 years, 5 years, 10 years or more after the treatment for cancer. The threshold for detecting the increase in the level of the nucleotide sequences associated with the cancer may be a baseline determined by the first time point.
[0116] Accordingly, in a further aspect, provided is a method for monitoring a response to cancer therapy in a subject, the method comprising: detecting, from a sample derived from a subject at a first time point prior to the cancer therapy, the presence and / or level of one or morenucleotide sequences associated with the presence of cancer; detecting, from a sample derived from a subject at a second time point after the onset of cancer therapy, the presence and / or level of the one or more nucleotide sequences associated with the presence of cancer; and comparing the results from the first and second time points, thereby monitoring the response to the cancer therapy, wherein one or both of the measuring steps are carried out according to a method described herein.Classifying cancer types or subtypes
[0117] Cancers can be classified by type and / or by subtype. The present methods, kits or devices can be used for the identification of a cancer type and / or subtype. Accordingly, in another aspect, there is provided a method for classifying a cancer type and / or subtype, the method comprising: detecting, from a sample derived from a subject suffering from cancer, the presence and / or level of the one or more nucleotide sequences associated with a specific type or subtype of cancer; and comparing the results to a reference data set for said cancer type and / or subtype, thereby identifying the type or subtype of the cancer; wherein the detecting step is carried out according to a method described herein.
[0118] The methods, kits or devices can be used as part of a pan-cancer detection strategy. Accordingly, a cancer type can be identified by detecting the presence of one or more nucleotide sequences associated with the presence of cancer using different aliquots obtained from the same sample from a subject. Accordingly, the enriched isolated chromatin-associated DNA in each of the aliquots is contacted with one or more primer pairs to amplify and detect (a) portion(s) of the one or more nucleotide sequences associated with the presence of different cancers. For example, a first aliquot can comprise primer pairs to amplify a portion of the one or more nucleotide sequences associated with the presence of a first cancer and a second aliquot can comprise primer pairs to amplify a portion of the one or more nucleotide sequences associated with the presence of a second, different type of cancer.
[0119] The methods, kits or devices can also be used to detect cancer in a subject that has previously been diagnosed with cancer to support the early identification of relapse. For example, the methods may include providing a sample obtained from a subject in remission from cancer.
[0120] The methods, kits or devices can also be used for the detection of metastases, e.g. , by detecting cancer-associated nucleotide sequences in cfDNA in a blood or serum sample. For example, a sample can be obtained from a subject that has been diagnosed with cancer that can be used in the methods, kits or devices to monitor the progression of the cancer and also detect the presence of one or more nucleotide sequences associated with cancer at distant sites.
[0121] It is well understood that different cancer subtypes have different prognoses and treatment options. The identification of a cancer sub-type with the methods can be used to select the most suitable therapeutic approach.Kits or devices
[0122] Also provided herein is a kit or device comprising reagents for isolating cell free DNA (cfDNA); reagents for specifically binding chromatin to isolate chromatin-associated DNA; reagents (e.g., oligonucleotide probes) for enriching the chromatin-associated DNA for one or more nucleotide sequences associated with the presence of the cancer; and optionally reagents (e.g., one or more primer pairs) for carrying out qPCR to amplify and / or detect the presence of a portion of the one or more nucleotide sequences associated with the presence of cancer. Kits or devices may also contain instructions for use, and may be used in conjunction with or to implement the present methods.
[0123] Reagents for carrying out qPCR to determine the presence of the one or more nucleotide sequences associated with the presence of cancer may be provided in the same or a separate kit or device from the reagents for the isolation of cfDNA, reagents for specifically binding chromatin, reagents for isolating chromatin-associated DNA and reagents for enriching the chromatin-associated DNA. Kits or devices for carrying out qPCR to determine the presence of the one or more nucleotide sequences associated with the presence of cancer may be for the detection of more than one nucleotide sequence associated with the presence of the cancer. For example, the kit may be for the detection of 1 to 256 nucleotide sequences associated with the presence of the cancer, e.g., 1-6, 1-8, 1-12, 1-24, 1-48, 1-100, 1-256 different nucleotide sequences associated with the presence of the cancer. Reagents may include at least primer pairs for the detection of the nucleotide sequences associated with the presence of the cancer and / orreagents for carrying out a qPCR reaction. Kits and devices for carrying out qPCR are conventional in the art.
[0124] The reagents for carrying out qPCR comprise one or more primer pairs for amplifying a portion of the one or more nucleotide sequences associated with the presence of cancer. Each of the primer pairs can specifically hybridize to a portion of a nucleotide sequence associated with the presence of cancer. In a typical embodiment, the amplified portion of the one or more nucleotide sequences associated with cancer is no more than 250 base pairs in length. More typically, the amplified portion may be between 100 and 200 base pairs in length.
[0125] The person skilled in the art will be aware of existing kits that may be compatible with the methods described herein, for example, the Cepheid qPCR kits, SybrGreen kits and Taqman Cards kits known in the art. The numbers of primer pairs used may be 12, 24 or up to 100. Several primer design tools are available online e.g., PrimerQuest (IDT) or GenScript, where the user imports a target sequence and other appropriate criteria e.g., primer length, GC content and melting temperature (Tm) and suitable primer sequences are provided as an output.Bladder cancer
[0126] Bladder cancer encompasses a highly heterogeneous group of cancers arising from tissues of the urinary bladder. It typically occurs when epithelial cells that line the bladder become malignant. Common symptoms include blood in the urine, frequent urination, urgent urination, pain with urination, and lower back pain.
[0127] Bladder cancer occurs primarily in patients aged 55 years and older. Most bladder cancers are diagnosed at an early stage, when the cancer is highly treatable. While initial treatment is often successful, new occurrences or recurrences of bladder cancer are common, and subjects having bladder cancer often need follow-up tests for years after treatment. The present methods, kits and devices are useful for the monitoring and / or detection of bladder cancer including new occurrences or recurrences. The present methods allow for early detection of new occurrences or recurrences of a bladder cancer from a liquid biopsy without the need to identify a new or recurrent tumor location for a solid biopsy.
[0128] Typically, the symptoms associated with bladder cancer (e.g., blood in urine, frequent urination, etc. lead to the initial detection of bladder cancer. Blood will cause urine tochange color, but urinalysis may be appropriate when there are only trace amounts of blood in the urine. Risk factors associated with bladder cancer include smoking, chemical exposure, race (Caucasians are twice as likely to develop bladder cancer), age (risk increases with age), gender (bladder cancer is more often found in men), and family history. The present methods, kits and devices described herein can be used to screen for cancer in a subject that is considered high- risk for bladder cancer. Accordingly, a sample obtained from a high-risk subject can be tested for the presence of bladder cancer using the present methods, kits or devices in the absence of symptom development. Such preventative screening can help diagnose subjects suffering from bladder cancer at an early stage of disease.
[0129] There are multiple different classifications of bladder cancer, including urothelial cell carcinoma (UCC), squamous cell carcinoma (SCC) and adenocarcinoma. Of the bladder cancer classifications, UCC is the most common making up approximately 90% of all bladder cancers. The method of detecting, diagnosing, and monitoring bladder cancer described herein are especially applicable to UCC and its subtypes. UCC develops from the transitional or urothelial cells of the bladder lining (urothelium). Approximately 5% and 1-2% of bladder cancers are SCC and adenocarcinoma, respectively.
[0130] Bladder cancers can also be classified (or graded) as non-muscle invasive bladder cancer (NMLBCs) or muscle-invasive bladder cancer (MIBC). The majority of bladder cancers (approximately 80%) are diagnosed as NMIBC, with limited metastatic and / or lethal potential. NMIBCs have high rates of local recurrence, and a significant fraction (25-30%) progress to MIBC. MIBCs are characterized by rapid progression to metastatic disease and are associated with higher mortality. While the distinction between NMIBC and MIBC is clinically useful, each is highly heterogeneous, exhibiting a wide range of clinical outcomes and treatment responses. NMIBCs and MIBCs exhibit distinct gene expression profiles. They can also each be further subdivided into further expression subtypes exhibiting distinct biological, and possibly clinical properties.
[0131] The methods, kits or devices can be used to identify a particular bladder cancer type and / or subtype. As demonstrated herein, the methods, kits or devices are capable of distinguishing these different types / subtypes of bladder cancer. There are two main subtypes of bladder cancer, referred to as luminal or basal. Bladder cancers may be described as “luminal-like” (e.g., luminal-infiltrated bladder cancer, or luminal-papillary bladder cancer) or “basal- like” (e.g., basal-squamous cancer bladder cancer). The methods, kits or devices can be used to classify a cancer type and / or subtype e.g., basal-like or luminal-like type, and UCC or non-UCC bladder cancer. The classification of a bladder cancer as luminal-like can indicate that the cancer is progressing towards MIBC. Another subtype of bladder cancer is micropapillary bladder cancer (MPBC). This is a rare and aggressive subtype of bladder cancer that is associated with poor outcome.
[0132] For example, the methods, kits or devices may be used to diagnose or monitor the progression of bladder cancer, or to monitor the response to a cancer therapy that is used for treating bladder cancer. The methods may comprise determining in one or more samples (e.g., in urine and / or serum) the level, expression, and / or activation state of a gene or gene regulatory region selected from the group consisting of NECTIN4, SOX4, FGFR3, TP63, KDM6A, KRT5, SPP1, IGFBP3, KMT2D, RBI, MYC and BRCA2. For instance, it may particularly useful to determine the level, expression, and / or activation state of a gene or gene regulatory region selected from the group consisting of NECTIN4, SOX4, FGFR3, TP63, KDM6A, KRT5, SPP1, and IGFBP3. Additionally, the methods may comprise determining in one or more samples the level, expression, and / or activation state of a gene or gene regulatory region selected from the group consisting of UGT1A8, SEMA4B, KRT80, KRT7, ERBB2, TACSTD2, ELK3, PSCA, EPCAM, EHF, TBX3, and ELF3. In some instances, the level, expression, and / or activation state of at least 5 (e.g., at least the first 5 in the list) or all of the gene or gene regulatory regions selected from the group consisting of NECTIN4, SOX4, FGFR3, TP63, KDM6A, KRT5, SPP1, and IGFBP3 is determined. For example, oligonucleotide probes and / or primer pairs may be used to determine the presence and / or level of one or more (e.g., at least 5, at least 10, or all) nucleotide sequences of SEQ ID NOs: 2273 (NECTIN4), 6365 (NECTIN4), 6366 (SOX4), 1211 (FGFR3), 3187 (TP63), 6367 (KDM6A), 6368 (KRT5), 3314 (SPP1) and 4037 (IGFBP3). Alternatively, the level, expression, and / or activation state of at least 5 (e.g., at least the first 5 in the list) or all of the gene or gene regulatory regions selected from the group consisting of UGT1A8, SEMA4B, NECTIN4, KRT80, KRT7, ERBB2, TACSTD2, KDM6A, ELK3, PSCA, EPCAM, EHF, TBX3, ELF 3 is determined. For example, oligonucleotide probes and / or primer pairs may be used to determine the presence and / or level of one or more (e.g., at least 5, at least10, or all) nucleotide sequences of SEQ ID NOs: 9792 (UGT1A8), 9793 (SEMA4B), 9794 (NECTIN4), 9795 (NECTIN4), 9796 (KRT80), 9797 (KRT7), 9798 (ERBB2), 9799 (TACSTD2), 9800 (KDM6A), 9801 (ELK3), 9802 (PSCA), 9803 (PSCA , 9804 (EPCAM), 9805 (EHF), 9806 (TBX3), 9807 (TBX3), and 9808 (ELF3).
[0133] In addition, or alternatively, the methods may comprise determining in one or more samples (e.g., in urine and / or serum) the level, expression, and / or activation state of a gene or gene regulatory region selected from the group consisting of FBN2, SLC4A4, CTTNBP2, KALRN, SELL, PMP22, UGT2B7, COL12A1, PTH2R, ANKRD36, DCDC2, PRTG, BPGM, PDE9A anAANXA3 (e.g., to determine whether the bladder cancer is luminal-like). Additionally, the methods may comprise determining in one or more samples the level, expression, and / or activation state of a gene or gene regulatory region selected from the group consisting of CASTOR3, CHKA, UPK1A, and PLEKHF1. In some instances, the level, expression, and / or activation state of at least 5 (e.g., at least the first 5 in the list), at least 10 (e.g., at least the first 10 in the list), or all of the gene or gene regulatory regions selected from the group consisting of FBN2, SLC4A4, CTTNBP2, KALRN, SELL, PMP22, UGT2B7, COL12A1, PTH2R, ANKRD36, DCDC2, PRTG, BPGM, PDE9A anAANXA3 is determined. For example, oligonucleotide probes and / or primer pairs may be used to determine the presence and / or level of one or more (e.g., at least 5, at least 10, or all) nucleotide sequences of SEQ ID NOs: 7620-7634. Alternatively, the level, expression, and / or activation state of at least 5 (e.g., at least the first 5 in the list), at least 6 (e.g., at least the first 6 in the list), at least 7 (e.g., at least the first 7 in the list), or all of the gene or gene regulatory regions selected from the group consisting of CTTNBP2, KALRN, SELL, PMP22, COL12A1, PRTG, PDE9A, ANXA3 is determined. For example, oligonucleotide probes and / or primer pairs may be used to determine the presence and / or level of one or more (e.g., at least 5, at least 10, or all) nucleotide sequences of SEQ ID NOs: 6664, 7622-7625, 7627, 7631, 7633-7634, and 9809-9815. Alternatively, the level, expression, and / or activation state of at least 5 (e.g., at least the first 5 in the list), at least 10 (e.g., at least the first 10 in the list), or all of the gene or gene regulatory regions selected from the group consisting of CTTNBP2, KALRN, SELL, PMP22, COL12A1, PRTG, PDE9A, ANXA3, CASTOR3, CHKA, UPK1A, and PLEKHF1 is determined. For example, oligonucleotide probes and / or primer pairs may be used to determinethe presence and / or level of one or more (e.g., at least 5, at least 10, or all) nucleotide sequences of SEQ ID NOs: 9809-9820.
[0134] In addition, or alternatively, the methods may comprise determining in one or more samples (e.g., in urine and / or serum) the level, expression, and / or activation state of a gene or gene regulatory region selected from the group consisting of BMP 7, ERN2, SRPX2, UGT1A10, CLCA2, SPOCD1, SERPINB5, SEMA4B, CLCA4, CLU, UGT1A7, SH3PXD2A, AQP3, ANXA10, and SLITRK6 (e.g, to determine whether the bladder cancer is basal-like). Additionally, the methods may comprise determining in one or more samples the level, expression, and / or activation state of a gene or gene regulatory region selected from the group consisting of UGT1A1 and TP63. In some instances, the level, expression, and / or activation state of at least 5 (e.g, at least the first 5 in the list), at least 10 (e.g., at least the first 10 in the list), or all of the gene or gene regulatory regions selected from the group consisting of BMP7, ERN2, SRPX2, UGT1A10, CLCA2, SPOCD1, SERPINB5, SEMA4B, CLCA4, CLU, UGT1A7, SH3PXD2A, AQP3, ANXA10 and SLITRK6 is determined. For example, oligonucleotide probes and / or primer pairs may be used to determine the presence and / or level of one or more (e.g., at least 5, at least 10, or all) nucleotide sequences of SEQ ID NOs: 7635-9733. Alternatively, the level, expression, and / or activation state of at least 5 (e.g., at least the first 5 in the list), at least 10 (e.g. , at least the first 10 in the list), or all of the gene or gene regulatory regions selected from the group consisting of BMP7, ERN2, SRPX2, UGT1A10, CLCA2, SERPINB5, CLCA4, CLU, SH3PXD2A, AQP3, ANXA10, and SLITRK6 is determined. For example, oligonucleotide probes and / or primer pairs may be used to determine the presence and / or level of one or more (e.g., at least 5, at least 10, or all) nucleotide sequences of SEQ ID NOs: 7699, 8953, 8158, 9401, 9719- 9723, 9725, 9727-9728, 9730-9733, 9821-9828, and 9830-9831. Alternatively, the level, expression, and / or activation state of at least 5 (e.g., at least the first 5 in the list), at least 10 (e.g., at least the first 10 in the list), or all of the gene or gene regulatory regions selected from the group consisting of BMP7, ERN2, SRPX2, UGT1A10, CLCA2, SERPINB5, CLCA4, CLU, UGT1A1, SH3PXD2A, AQP3, ANXA10, SLITRK6, and TP63 is determined. For example, oligonucleotide probes and / or primer pairs may be used to determine the presence and / or level of one or more (e.g., at least 5, at least 10, or all) nucleotide sequences of SEQ ID NOs: 9821- 9832.
[0135] In addition, or alternatively, the methods may comprise determining in one or more samples (e.g., in urine and / or serum) the level, expression, and / or activation state of a gene or gene regulatory region selected from the group consisting of ANPEP, BPGM, SPNS2, MAL, ASS1, COL4A4, MUC4, TMC5, MMP7, PLCXD3, KRT80, PROMI, DAB2, KRT23, and RHPN2 (e.g., to determine whether the bladder cancer is a micropapillary bladder cancer). In some instances, the level, expression, and / or activation state of at least 5 (e.g., at least the first 5 in the list), at least 10 (e.g., at least the first 10 in the list), or all of the gene or gene regulatory regions selected from the group consisting of ANPEP, BPGM, SPNS2, MAL, ASS1, COL4A4, MUC4, TMC5, MMP7, PLCXD3, KRT80, PROMI, DAB2, KRT23, and RHPN2 is determined. For example, oligonucleotide probes and / or primer pairs may be used to determine the presence and / or level of one or more (e.g., at least 5, at least 10, or all) nucleotide sequences of SEQ ID NOs: 7172 and 9833-9847.
[0136] The detection of the presence and / or level of the one or more nucleotide sequences associated with the presence of cancer can be performed using any suitable method. For instance, liquid biopsies (e.g., urine or serum) may be assayed using the workflow as described herein. Other exemplary methods include RNA sequencing (RNAseq), qPCR, immunohistochemistry (IHC), immunofluorescence (IF) and RNAscope. For example, qPCR, RNAseq, or immunohistochemistry (IHC) may be used on tissue sample.
[0137] Traditionally, classification of bladder cancer has relied on histopathologic status, such as tumor size, nodal status, and distance metastasis (TNM classification). While the TNM classification system is well established, clinical diagnosis of bladder cancer using the level, expression, and / or activation state of a gene or gene regulatory region described herein offers several significant advantages as compared to standard histological diagnosis. These advantages include, among others: (i) the provision of an objective scoring system (standard histological approaches are inherently subjective), which can result in more consistent diagnoses, (ii) the identification of bladder cancer heterogeneity and features that cannot be detected via histological analysis, and (iii) the identification of therapeutic targets that cannot be determined using solely DNA mutation status.
[0138] For example, luminal-like, basal-like and micropapillary bladder cancers show different prognoses and are susceptible to different targeted therapies as is known in the art.Accordingly, identifying whether a bladder cancer e.g., a resected cancer or a tissue sample thereof) has a luminal-like or basal-like subtype or is a micropapillary bladder cancer can support selection of the most suitable therapeutic approach. The present methods, kits, and devices can be used to identify whether a bladder cancer is of a basal-like or luminal-like subtype, or a micropapillary bladder cancer.
[0139] Common sites of bladder cancer metastasis include the liver and lung among other organs. The methods, kits or devices can be used for the detection of metastases, e.g., by detecting cancer-associated nucleotide sequences in cfDNA in a blood sample (e.g., serum or plasma). For example, a sample can be obtained from a subject that has been diagnosed with bladder cancer to monitor the progression of the bladder cancer and also detect the presence of one or more nucleotide sequences associated with bladder cancer at sites distant to the original tumor.Kit and device for detecting, monitoring, or classifyin bladder cancer
[0140] Kits for the detection of bladder cancer may comprise oligonucleotide probes that specifically hybridize to one or more of the bladder cancer-associated nucleotide sequences as set forth in SEQ ID NOs: 1-6372, or complementary nucleotide sequences thereof, wherein each oligonucleotide probe binds to a distinct region of the one or more bladder cancer-associated nucleotide sequences. In addition, such kits may comprise one or more primer pairs wherein each primer pair amplifies a portion of a nucleotide sequence within the one or more bladder cancer-associated nucleotide sequences as set forth in SEQ ID NOs: 1-6372, or complementary nucleotide sequences thereof. Each of the oligonucleotide probe sequences may specifically hybridize to a distinct region of one or more (or all) nucleotide sequences as set forth in SEQ ID NOs: 2273 (NECTIN4), 6365 (NECTIN4), 6366 (SOX4), 1211 (FGFR3), 3187 (TP63), 6367 (KDM6A), 6368 (KRT5), 3314 (SPP1) and 4037 (IGFBP3), and each of the primer pairs may amplify a portion of a nucleotide sequence of one or more (or all) of the nucleotide sequences as set forth in SEQ ID NOs: 2273 (NECTIN4), 6365 (NECTIN4), 6366 (SOX4), 1211 (FGFR3), 3187 (TP63), 6367 (KDM6A), 6368 (KRT5), 3314 (SPP1) and 4037 (IGFBP3). For the detection of bladder cancer, the primer pairs may have the sequences as set forth in SEQ ID NOs: 9734- 9759.
[0141] Alternatively or additionally, the kit may comprise oligonucleotide probes that specifically hybridize to one or more of the bladder cancer-associated nucleotide sequences as set forth in SEQ ID NOs: 9792-9808 or complementary nucleotide sequences thereof, wherein each oligonucleotide probe binds to a distinct region of the one or more bladder cancer- associated nucleotide sequences. In addition, such kits may comprise one or more primer pairs wherein each primer pair amplifies a portion of a nucleotide sequence within the one or more bladder cancer-associated nucleotide sequences as set forth in SEQ ID NOs: 9792-9808, or complementary nucleotide sequences thereof. Each of the oligonucleotide probe sequences may specifically hybridize to a distinct region of one or more (or all) nucleotide sequences as set forth in SEQ ID NOs: 9792 (UGT1A8), 9793 (SEMA4B), 9794 (NECTIN4), 9795 (NECTIN4), 9796 (KRT80), 9797 (KRT7), 9798 (ERBB2), 9799 (TACSTD2), 9800 (KDM6A), 9801 (ELK3), 9802 (PSCA), 9803 (P G4), 9804 (EPCAM), 9805 (EHF), 9806 (TBX3), 9807 (TBX3), and 9808 (ELF3), and each of the primer pairs may amplify a portion of a nucleotide sequence of one or more (or all) of the nucleotide sequences as set forth in SEQ ID NOs: 9792 (UGT1A8), 9793 (SEMA4B), 9794 (NECTIN4), 9795 (NECTIN4), 9796 (KRT80), 9797 (KRT7), 9798 (ERBB2), 9799 (TACSTD2), 9800 (KDM6A), 9801 (ELK3), 9802 (PSCA), 9803 (PSCA), 9804 (EPCAM), 9805 (EHF), 9806 (TBX3), 9807 (TBX3), and 9808 (ELF3).
[0142] Kits for the detection of luminal-like bladder cancer may comprise oligonucleotide probes that specifically hybridize to one or more of the luminal-like bladder cancer-associated nucleotide sequences as set forth in SEQ ID NOs: 6373-7634, or complementary nucleotide sequences thereof, wherein each oligonucleotide probe binds to a distinct region of the one or more luminal-like bladder cancer-associated nucleotide sequences. In addition, such kits may comprise one or more primer pairs wherein each primer pair amplifies a portion of a nucleotide sequence within the one or more luminal-like bladder cancer-associated nucleotide sequences as set forth in SEQ ID NOs: 6373-7634, or complementary nucleotide sequences thereof. Each of the oligonucleotide probe sequences may specifically hybridize to a distinct region of one or more (or all) nucleotide sequences as set forth in SEQ ID NOs: 7620- 7634, and each of the primer pairs may amplify a portion of a nucleotide sequence of one or more (or all) of the nucleotide sequences as set forth in SEQ ID NOs: 7620-7634.
[0143] Alternatively or additionally, the kit may comprise oligonucleotide probes that specifically hybridize to one or more of the luminal-like bladder cancer-associated nucleotide sequences as set forth in SEQ ID NOs: 9809-9820 or complementary nucleotide sequences thereof, wherein each oligonucleotide probe binds to a distinct region of the one or more luminallike bladder cancer-associated nucleotide sequences. In addition, such kits may comprise one or more primer pairs wherein each primer pair amplifies a portion of a nucleotide sequence within the one or more luminal-like bladder cancer-associated nucleotide sequences as set forth in SEQ ID NOs: 9809-9820, or complementary nucleotide sequences thereof. Each of the oligonucleotide probe sequences may specifically hybridize to a distinct region of one or more (or all) nucleotide sequences as set forth in SEQ ID NOs: 9809-9820, and each of the primer pairs may amplify a portion of a nucleotide sequence of one or more (or all) of the nucleotide sequences as set forth in SEQ ID NOs: 9809-9820.
[0144] Kits for the detection of basal-like bladder cancer may comprise oligonucleotide probes that specifically hybridize to one or more of the basal-like bladder cancer-associated nucleotide sequences as set forth in SEQ ID NOs: 7635-9733, or complementary nucleotide sequences thereof, wherein each oligonucleotide probe binds to a distinct region of the one or more basal-like bladder cancer-associated nucleotide sequences. In addition, such kits may comprise one or more primer pairs wherein each primer pair amplifies a portion of a nucleotide sequence within the one or more basal-like bladder cancer-associated nucleotide sequences as set forth in SEQ ID NOs: 7635-9733, or complementary nucleotide sequences thereof. For example, each of the oligonucleotide probe sequences may specifically hybridize to a distinct region of one or more (or all) nucleotide sequences as set forth in SEQ ID NOs: 9719-9733, and each of the primer pairs may amplify a portion of a nucleotide sequence of one or more (or all) of the nucleotide sequences as set forth in SEQ ID NOs: 9719-9733.
[0145] Alternatively or additionally, the kit may comprise oligonucleotide probes that specifically hybridize to one or more of the basal-like bladder cancer-associated nucleotide sequences as set forth in SEQ ID NOs: 9821-9832 or complementary nucleotide sequences thereof, wherein each oligonucleotide probe binds to a distinct region of the one or more basal- like bladder cancer-associated nucleotide sequences. In addition, such kits may comprise one or more primer pairs wherein each primer pair amplifies a portion of a nucleotide sequence withinthe one or more basal-like bladder cancer-associated nucleotide sequences as set forth in SEQ ID NOs: 9821-9832, or complementary nucleotide sequences thereof. Each of the oligonucleotide probe sequences may specifically hybridize to a distinct region of one or more (or all) nucleotide sequences as set forth in SEQ ID NOs: 9821-9832, and each of the primer pairs may amplify a portion of a nucleotide sequence of one or more (or all) of the nucleotide sequences as set forth in SEQ ID NOs: 9821-9832.
[0146] Kits may be provided that can be used for the detection of luminal-like or basal- like bladder cancer. Such kits may comprise one or more oligonucleotide probes that specifically hybridize to one or more of the luminal-like bladder cancer-associated nucleotide sequences as set forth in SEQ ID NOs: 6373-7634, or complementary nucleotide sequences thereof, and one or more oligonucleotide probes that specifically hybridize to one or more of the basal-like bladder cancer-associated nucleotide sequences as set forth in SEQ ID NOs: 7635-9733, or complementary nucleotide sequences thereof, wherein each oligonucleotide probe binds to a distinct region of the one or more luminal-like or basal-like bladder cancer-associated nucleotide sequences. In addition, such kits may comprise two or more primers pairs wherein at least one primer pair amplifies a portion of a nucleotide sequence within the one or more luminal-like bladder cancer-associated nucleotide sequences as set forth in SEQ ID NOs: 6373-7634, or complementary nucleotide sequences thereof, and at least one primer pair amplifies a portion of a nucleotide sequence within the one or more basal-like bladder cancer-associated nucleotide sequences as set forth in SEQ ID NOs: 7635-9733, or complementary nucleotide sequences thereof. For example, each of the one or more oligonucleotide probe sequences may specifically hybridize to a distinct region of one or more (or all) nucleotide sequences as set forth in SEQ ID NOs: 7620-7634, each of the one or more the oligonucleotide probe sequences may specifically hybridize to a distinct region of one or more (or all) nucleotide sequences as set forth in SEQ ID NOs: 9719-9733, at least one primer pairs may amplify a portion of a nucleotide sequence of one or more (or all) of the nucleotide sequences as set forth in SEQ ID NOs: 7620-7634, and each of the primer pairs may amplify a portion of a nucleotide sequence of one or more (or all) of the nucleotide sequences as set forth in SEQ ID NOs: 9719-9733.
[0147] Alternatively or additionally, the kit may comprise one or more oligonucleotide probes that specifically hybridize to one or more of the luminal-like bladder cancer-associatednucleotide sequences as set forth in SEQ ID NOs: 9809-9820, or complementary nucleotide sequences thereof, and one or more oligonucleotide probes that specifically hybridize to one or more of the basal-like bladder cancer-associated nucleotide sequences as set forth in SEQ ID NOs: 9821-9832, or complementary nucleotide sequences thereof, wherein each oligonucleotide probe binds to a distinct region of the one or more luminal-like or basal-like bladder cancer- associated nucleotide sequences. In addition, such kits may comprise two or more primers pairs wherein at least one primer pair amplifies a portion of a nucleotide sequence within the one or more luminal-like bladder cancer-associated nucleotide sequences as set forth in SEQ ID NOs: 9809-9820, or complementary nucleotide sequences thereof, and at least one primer pair amplifies a portion of a nucleotide sequence within the one or more basal-like bladder cancer- associated nucleotide sequences as set forth in SEQ ID NOs: 9821-9832, or complementary nucleotide sequences thereof.
[0148] Such kits for the detection of luminal-like or basal-like cancer subtypes may also include the oligonucleotide probes that specifically hybridize to one or more of the bladder cancer-associated nucleotide sequences as set forth in SEQ ID NOs: 1-6372 (e.g., SEQ ID NOs: 2273, 6365, 6366, 1211, 3187, 6367, 6368, 3314 and 4037), as well as one or more primer pairs wherein each primer pair amplifies a portion of a nucleotide sequence within the one or more bladder cancer-associated nucleotide sequences as set forth in SEQ ID NOs: 1-6372 (e.g., SEQ ID NOs: 2273, 6365, 6366, 1211, 3187, 6367, 6368, 3314 and 4037), or complementary nucleotide sequences thereof. Alternatively or additionally, the kit may comprise oligonucleotide probes that specifically hybridize to one or more of the bladder cancer-associated nucleotide sequences as set forth in SEQ ID NOs: 9792-9808, as well as one or more primer pairs wherein each primer pair amplifies a portion of a nucleotide sequence within the one or more bladder cancer-associated nucleotide sequences as set forth in SEQ ID NOs: 9792-9808, or complementary nucleotide sequences thereof.
[0149] Kits for the detection of micropapillary bladder cancer may comprise oligonucleotide probes that specifically hybridize to one or more of the micropapillary bladder cancer-associated nucleotide sequences as set forth in SEQ ID NOs: 7172 and 9833-9847, or complementary nucleotide sequences thereof, wherein each oligonucleotide probe binds to a distinct region of the one or more micropapillary-like bladder cancer-associated nucleotidesequences. In addition, such kits may comprise one or more primer pairs wherein each primer pair amplifies a portion of a nucleotide sequence within the one or more micropapillary-like bladder cancer-associated nucleotide sequences as set forth in SEQ ID NOs: : 7172 and 9833- 9847, or complementary nucleotide sequences thereof. Each of the oligonucleotide probe sequences may specifically hybridize to a distinct region of one or more (or all) nucleotide sequences as set forth in SEQ ID NOs: : 7172 and 9833-9847, and each of the primer pairs may amplify a portion of a nucleotide sequence of one or more (or all) of the nucleotide sequences as set forth in SEQ ID NOs: : 7172 and 9833-9847.
[0150] Also provided are kits for classifying a bladder cancer comprising one or more reagents for determining in a sample (i) expression of, and / or activation state of a gene regulatory region of, one or more of (e.g., at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 of) BMP7, ERN2, SRPX2, UGT1A10, CLCA2, SPOCD1, SERPINB5, SEMA4B, CLCA4, CLU, UGT1A7, SH3PXD2A, AQP3, ANXA10 and SLITRK6; and (ii) expression of, and / or activation state of a gene regulatory region of, one or more of (e.g., at least 2, 3, 4, 5, 6, 7, 8„ 9, 10, 11, 12, 13, 14, or 15 of) FBN2, SLC4A4, CTTNBP2, KALRN, SELL, PMP22, UGT2B7, COL12A1, PTH2R, ANKRD36, DCDC2, PRTG, BPGM, PDE9A, and ANXA3. Such kits may further comprise one or more reagents for determining in a sample expression of, and / or activation state of a gene regulatory region of, one or more of (e.g, at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 or 12 of) NECTIN4, SOX4, FGFR3, TP63, KDM6A, KRT5, SPP1, IGFBP3, KMT2D, RBI, MYC and BRCA2.
[0151] Also provided are kits for classifying a bladder cancer comprising one or more reagents for determining in a sample (i) expression of, and / or activation state of a gene regulatory region of at least the first 3, 5, or 10 of the list of BMP7, ERN2, SRPX2, UGT1A10, CLCA2, SPOCD1, SERPINB5, SEMA4B, CLCA4, CLU, UGT1A7, SH3PXD2A, AQP3, ANXA10 and SLITRK6; and (ii) expression of, and / or activation state of a gene regulatory region of at least the first 3, 5, or 10 of the list of FBN2, SLC4A4, CTTNBP2, KALRN, SELL, PMP22, UGT2B7, COL12A1, PTH2R, ANKRD36, DCDC2, PRTG, BPGM, PDE9A, and ANXA3. Such kits may further comprise one or more reagents for determining in a sample expression of, and / or activation state of a gene regulatory region of at least the first 3, 5, or 8 of the list of NECTIN4, SOX4, FGFR3, TP63, KDM6A, KRT5, SPP1, IGFBP3, KMT2D, RBI, MYC and BRCA2.
[0152] Additionally, the kit may comprise one or more reagents for determining in a sample (i) expression of, and / or activation state of a gene regulatory region selected from the group consisting of UGT1A1 and TP63 and expression of, and / or activation state of a gene regulatory region selected from the group consisting of CASTOR3, CHKA, UPK1A, PLEKHF1, and optionally (iii) expression of, and / or activation state of a gene regulatory region selected from the group consisting of UGT1A8, SEMA4B, KRT80, KRT7, ERBB2, TACSTD2, ELK3, PSCA, EPCAM, EHF, TBX3, ELF3.
[0153] For example, a kit for classifying a bladder cancer may comprise one or more reagents for determining in a sample (i) expression of, and / or activation state of a gene regulatory region of at least the first 3, 5, or 10 of the list of BMP 7, ERN2, SRPX2, UGT1A10, CLCA2, SERPINB5, CLCA4, CLU, SH3PXD2A, AQP3, ANXA10, SLITRK6; and (n) expression of, and / or activation state of a gene regulatory region of at least the first 3, 5, or 7 of the list of CTTNBP2, KALRN, SELL, PMP22, COL12A1, PRTG, PDE9A, ANXA3 and optionally (hi) expression of, and / or activation state of a gene regulatory region of at least one of NECTIN4, and KDM6A.
[0154] Alternatively, a kit for classifying a bladder cancer may comprise one or more reagents for determining in a sample (i) expression of, and / or activation state of a gene regulatory region of at least the first 3, 5, or 10 of the list of BMP7, ERN2, SRPX2, UGT1A10, CLCA2, SERPINB5, CLCA4, CLU, UGT1A1, SH3PXD2A, AQP3, ANXA10, SLITRK6, TP63; and (h) expression of, and / or activation state of a gene regulatory region of at least t the first 3, 5, or 10 of the list (ACTTNBP2, KALRN, SELL, PMP22, COL12A1, PRTG, PDE9A,ANXA3, CASTOR3, CHKA, UPK1A, PLEKHFL, and optionally (iii) expression of, and / or activation state of a gene regulatory region of at least the first 3, 5, or 10 of the list of UGT1A8, SEMA4B, NECTIN4, KRT80, KRT7, ERBB2, TACSTD2, KDM6A, ELK3, PSCA, EPCAM, EHF, TBX3, ELF3.
[0155] Also provided are kits for classifying a bladder cancer comprising one or more reagents for determining in a sample expression of, and / or activation state of a gene regulatory region of at least the first 3, 5, or 10 of the list oiANPEP, BPGM, SPNS2, MAL, ASS1, COL4A4, MUC4, TMC5, MMP7, PLCXD3, KRT80, PROMI, DAB2, KRT23, RHPN2. Such kits may further comprise one or more reagents for determining in a sample expression of, and / or activation state of a gene regulatory region of at least the first 3, 5, or 8 of the list of UGT1A8,SEMA4B, NECTIN4, KRT80, KRT7, ERBB2, TACSTD2, KDM6A, ELK3, PSCA, EPCAM, EHF, TBX3, ELF3.
[0156] The expression of the one or more genes, and / or the activation state of one or more gene regulatory regions may be determined using one or more oligonucleotide probes that specifically hybridize to the one or more genes or gene regulatory regions, or one or more primer pairs that can be used in a PCR reaction to specifically amplify the one or more genes or gene regulatory regions. Alternatively, expression of the one or more genes may be determined using one or more antibodies each of which specifically binds to one of the proteins encoded by the one or more genes. The antibody may be linked to a detectable moiety, or the kit may comprise one or more secondary antibodies to detect the one or more antibodies that specifically bind to the proteins encoded by the one or more genes.Computational methods for classifying bladder cancer
[0157] Also provided is a computer-implemented method for classifying a bladder cancer in a subject as luminal-like or basal-like, the method comprising (a) receiving data indicative of expression of, and / or activation state of a gene regulatory region of, (i) one or more of FBN2, SLC4A4, CTTNBP2, KALRN, SELL, PMP22, UGT2B7, COL12A1, PTH2R, ANKRD36, DCDC2, PRTG, BPGM, PDE9A, tm<3ANXA3, and (11) one or more of BMP7, ERN2, SRPX2, UGT1A10, CLCA2, SPOCD1, SERPINB5, SEMA4B, CLCA4, CLU, UGT1A7, SH3PXD2A, AQP3, ANXA10, and SLITRK6, and optionally (hi) one or more of NECTIN4, SOX4, FGFR3, TP63, KDM6A, KRT5, SPP1, IGFBP3, KMT2D, RBI, MYC, and BRCA2, in a sample collected from the subject; (b) obtaining a statistical model; (c) comparing the data to the statistical model to generate a score; and (d) outputting the score from the statistical model, wherein the score indicates the probability of the bladder cancer being luminal-like or basal-like.
[0158] Additionally, the method may comprise receiving data indicative of expression of, and / or activation state of one or more gene regulatory regions selected from the group consisting of (i) CASTOR3, CHKA, UPK1A, and PLEKHF1, and (h) UGT1A1, and TP63, and (hi) UGT1A8, SEMA4B, KRT80, KRT7, ERBB2, TACSTD2, ELK3, PSCA, EPCAM, EHF, TBX3, and ELF3.
[0159] For example, a computer-implemented method for classifying a bladder cancer in a subject as luminal-like or basal-like, the method comprising (a) receiving data indicative ofexpression of, and / or activation state of a gene regulatory region of: (i) one or more of c, and (11) one or more of BMP7, ERN2, SRPX2, UGT1A10, CLCA2, SERPINB5, CLCA4, CLU, SH3PXD2A, AQP3, ANXA10, SLITRK6, and optionally (in) one or more of NECTIN4, and KDM6A, in a sample collected from the subject; (b) obtaining a statistical model; (c) comparing the data to the statistical model to generate a score; and (d) outputting the score from the statistical model, wherein the score indicates the probability of the bladder cancer being luminallike or basal-like.
[0160] For example, a computer-implemented method for classifying a bladder cancer in a subject as luminal-like or basal-like, the method comprising (a) receiving data indicative of expression of, and / or activation state of a gene regulatory region of: (i) one or more of CTTNBP2, KALRN, SELL, PMP22, COL12A1, PRTG, PDE9A, ANXA3, CASTOR3, CHKA, UPK1A, PLEKHF1, and (h) one or more of BMP7, ERN2, SRPX2, UGT1A10, CLCA2, SERPINB5, CLCA4, CLU, UGT1A1, SH3PXD2A, AQP3, ANXA10, SLITRK6, and TP63, and optionally (111) one or more of UGT1A8, SEMA4B, NECTIN4, KRT80, KRT7, ERBB2, TACSTD2, KDM6A, ELK3, PSCA, EPCAM, EHF, TBX3, ELF3, in a sample collected from the subject; (b) obtaining a statistical model; (c) comparing the data to the statistical model to generate a score; and (d) outputting the score from the statistical model, wherein the score indicates the probability of the bladder cancer being luminal-like or basal-like.
[0161] In some instances, step (a) may comprise receiving data indicative of the expression of, and / or the activation state of the gene regulatory region of (i) at least 2, 5, or 10, (e.g., 2, 3, 4, 5 6, 7, 8, 9, 10, 11, 12, 13, 14 or all) of the gene or gene regulatory region of FBN2, SLC4A4, CTTNBP2, KALRN, SELL, PMP22, UGT2B7, COL12A1, PTH2R, ANKRD36, DCDC2, PRTG, BPGM, PDE9A, avAANXA3,’ and (h) at least 2, 5, or 10, (e.g., 2, 3, 4, 5 6, 7, 8,9, 10, 11, 12, 13, 14 or all) of the gene or gene regulatory region of BMP7, ERN2, SRPX2, UGT1A10, CLCA2, SPOCD1, SERPINB5, SEMA4B, CLCA4, CLU, UGT1A7, SH3PXD2A, AQP3, ANXA10, and SLITRK6,' and optionally (hi) at least 2, 5, or 10, (e.g., 2, 3, 4, 5 6, 7, 8, 9,10, 11, 12, 13, 14 or all) of NECTIN4, SOX4, FGFR3, TP63, KDM6A, KRT5, SPP1, IGFBP3, KMT2D, RBI, MYC, and BRCA2.
[0162] Additionally, the method may comprise receiving data indicative of expression of, and / or activation state of one or more gene regulatory regions selected from the group consisting(in) UGT1A8, SEMA4B, KRT80, KRT7, ERBB2, TACSTD2, ELK3, PSCA, EPCAM, EHF, TBX3, and ELF3.
[0163] For example, step (a) may comprise receiving data indicative of the expression of, and / or the activation state of the gene regulatory region of (i) at least 2, 5, or 7, (e.g., 2, 3, 4, 5 6, 7, or all) of the gene or gene regulatory region of CTTNBP2, KALRN, SELL, PMP22, COL12A1, PRTG, PDE9A, ANXA3' and (n) at least 2, 5, or 10, (e.g., 2, 3, 4, 5 6, 7, 8, 9, 10, 11, or all) of the gene or gene regulatory region of BMP7, ERN2, SRPX2, UGT1A10, CLCA2, SERPINB5, CLCA4, CLU, SH3PXD2A, AQP3, ANXA10, SI.ITRK6, and optionally (m) at least one of the gene or gene regulatory region of NECTIN4, and KDM6A.For example, step (a) may comprise receiving data indicative of the expression of, and / or the activation state of the gene regulatory region of (i) at least 2, 5, or 10, (e.g., 2, 3, 4, 5 6, 7, 8, 9, 10, 11, or all) of the gene or gene regulatory region of CTTNBP2, KALRN, SELL, PMP22, COL12A1, PRTG, PDE9A, ANXA3, CASTOR3, CHKA, UPK1A, PLEKHFF, and (n) at least 2, 5, or 10, (e.g., 2, 3, 4, 5 6, 7, 8, 9, 10, 11, 12, 13, or all) of the gene or gene regulatory region of BMP7, ERN2, SRPX2, UGT1A10, CLCA2, SERPINB5, CLCA4, CLU, UGT1A1, SH3PXD2A, AQP3, ANXA10, SLITRK6, TP63' and optionally (m) at least 2, 5, or 10 (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or all) of the gene or gene regulatory region of UGT1A8, SEMA4B, NECTIN4, KRT80, KRT7, ERBB2, TACSTD2, KDM6A, ELK3, PSCA, EPCAM, EHF, TBX3, ELF3.
[0164] In some instances, step (a) may comprise receiving data indicative of the expression of, and / or the activation state of the gene regulatory region of, (i) at least the first 1, 2, 3, 4, 5 6, 7, 8, 9, or 10 of the gene or gene regulatory region of FBN2, SLC4A4, CTTNBP2, KALRN, SELL, PMP22, UGT2B7, COL12A1, PTH2R, ANKRD36, DCDC2, PRTG, BPGM, PDE9A, and ANXA3,' and (ii) at least the first 1, 2, 3, 4, 5 6, 7, 8, 9, or 10 of the gene or gene regulatory region of BMP7, ERN2, SRPX2, UGT1A10, CLCA2, SPOCD1, SERPINB5, SEMA4B, CLCA4, CLU, UGT1A7, SH3PXD2A,AQP3,ANXA10, and SLITRK6,' and optionally (hi) at least the first 1, 2, 3, 4, 5 6, 7, 8, 9, or 10 of the gene or gene regulatory region of NECTIN4, SOX4, FGFR3, TP63, KDM6A, KRT5, SPP1, IGFBP3, KMT2D, RBI, MYC, and BRCA2.
[0165] In some instances, step (a) may comprise receiving data indicative of the expression of, and / or the activation state of the gene regulatory region of, (i) at least the first 1, 2, 3, 4, 5 6, 7, or 8, of the gene or gene regulatory region of CTTNBP2, KALRN, SELL, PMP22, COL12A1, PRTG, PDE9A, ANXA3- and (ii) at least the first 1, 2, 3, 4, 5 6, 7, 8, 9, or 10 of thegene or gene regulatory region of BMP7, ERN2, SRPX2, UGT1A10, CLCA2, SERPINB5, CLCA4, CLU, SH3PXD2A, AQP3, ANXA10, SLITRK6-, and optionally (in) at least the first 1, or 2 of the gene or gene regulatory region of NECTIN4 and KDM6A.
[0166] In some instances, step (a) may comprise receiving data indicative of the expression of, and / or the activation state of the gene regulatory region of, (i) at least the first 1, 2, 3, 4, 5 6, 7, 8, 9 or 10 of the gene or gene regulatory region of CTTNBP2, KALRN, SELL, PMP22, COL12A1, PRTG, PDE9A, ANXA3, CASTOR3, CHKA, UPK1A, PLEKHFL, and (n) at least the first 1, 2, 3, 4, 5 6, 7, 8, 9, or 10 of the gene or gene regulatory region of BMP7, ERN2, SRPX2, UGT1A10, CLCA2, SERPINB5, CLCA4, CLU, UGT1A1, SH3PXD2A,AQP3,ANXA10, SLITRK6, TP63,' and optionally (iii) at least the first 1, 2, 3, 4, 5 6, 7, 8, 9, or 10 of the gene or gene regulatory region of UGT1A8, SEMA4B, NECTIN4, KRT80, KRT7, ERBB2, TACSTD2, KDM6A, ELK3, PSCA, EPCAM, EHF, TBX3, ELF3.
[0167] Also provided is a computer-implemented method for classifying a bladder cancer in a subject as micropapillary bladder cancer, the method comprising (a) receiving data indicative of expression of, and / or activation state of a gene regulatory region of: (i) one or more oiANPEP, BPGM, SPNS2, MAL, ASS1, COL4A4, MUC4, TMC5, MMP7, PLCXD3, KRT80, PROMI, DAB2, KRT23, RHPN2, and optionally (n) one or more of UGT1A8, SEMA4B, NECTIN4, KRT80, KRT7, ERBB2, TACSTD2, KDM6A, ELK3, PSCA, EPCAM, EHF, TBX3, ELF3, in a sample collected from the subject; (b) obtaining a statistical model; (c) comparing the data to the statistical model to generate a score; and (d) outputting the score from the statistical model, wherein the score indicates the probability of the bladder cancer being micropapillary bladder cancer.
[0168] Also provided is a computer program comprising a computer program code configured to cause one or more physical computing devices to perform the method when the code is run on the one or more physical computing devices, as well as a computing or storage devices comprising the computer program or code.EXAMPLES
[0169] Although methods and materials similar or equivalent to those described herein can be used, suitable methods and materials are described below. The following examples are included for illustrative purposes only and are not intended to be limiting.Example 1. Method workflow
[0170] This example outlines a multi-step method for detecting the presence or increased level of cancer-associated differentially activated regulatory regions in a liquid biopsy (e.g., a blood or urine sample). This method combines the steps of enriching cfDNA first for chromatin- associated DNA (e.g., using chromatin immunoprecipitation) and then for known gene regulatory sequences associated with the presence of cancer (using probe capture). Post-capture qPCR is used to detect gene regulatory sequences of interest.Chromatin immunoprecipitation (first enrichment step)
[0171] A first enrichment step comprises the isolation of circulating DNA (free and exosome) followed by chromatin immunoprecipitation. Chromatin immunoprecipitation was performed as previously described in Cejas et al. Nat Commun. 2021; 12(1):5775. Matched peripheral clinical samples i.e., blood (plasma and serum), tumor tissue and urine were obtained from bladder cancer patients (n=13). Urine, blood, plasma or serum samples were thawed and centrifuged for 2 minutes at 4°C to remove debris. Samples were fixed with 1% formaldehyde for 10 minutes followed by a quenching step with 0.125M glycine for 5 minutes, with both steps performed at room temperature. Samples were then prepared by incubation with an ice-cold lysis buffer (50 mM Tris-HCl, lOmM, EDTA, 1% SDS, 5mM sodium butyrate and IX protease inhibitor, pH 8) on ice for 10-15 minutes. Lysis makes cfDNA present in exosomes and other extracellular vesicles available for ChlP. The chromatin was then sheared to 100-300 nucleotides using a Covaris E220 sonicator (140- watt peak incident power (PIP), 5% duty factor (DF), 200 cycles / burst). The samples were then centrifuged for 15 minutes at 4°C, and the supernatant transferred to a new tube and a portion taken for DNA preparation. Supernatants were diluted in ChlP dilution buffer (1% triton, 2mM EDTA, 150mM NaCl and 20mM Tris-HCl, pH 8). 5-20 pg DNA equivalent was then incubated with anti-H3K27ac (Diagenode, Cl 5410196), along with protein A and protein G beads (Life Technologies) in PBS supplemented with 0.5% BSA under constant 360° rotation at 4°C overnight followed by reverse crosslinking of the DNA for whole genome sequencing.
[0172] Elution of the chromatin was carried out in elution buffer (1% SDS, 0.1M NaHCCh), using magnetic bead collection as described in Cejas et al. (supra). After reversecrosslinking the DNA for 6-16 hours at 65°C with agitation at 800-1000 rpm, the supernatant was collected using magnetic bead collection, and the Supernatant using Qiagen MinElute Purification Kit, in pre-warmed EB at (37°C). Libraries were prepared using a ThruPLEX®-FD Prep Kit (Rubicon Genomics) with 15 cycles of amplification, in accordance with manufacturer’s instructions.Probe capture (second enrichment step)
[0173] Following the first enrichment step, a pan-cancer enrichment was performed using probe capture. Probe capture was carried out using the Agilent SureS elect protocol as described in the user manual (Version F0, September 2022). 50-500 ng of DNA was diluted in SureSelect XT HS2 blocker mix and a hybridization mix was prepared with RNase, probes and SureSelect fast hybridization buffer according to manufacturer’s instructions. Amplification was performed using a thermocycler using the following protocol settings: denaturation step of 5 min at 95°C, blocking of repetitive regions for 10 min at 65°C and 60 cycles of hybridization with 1 minute at 65°C and 3s at 37°C followed by a 65°C hold. Capture of the hybridized libraries followed the SureSelect protocol as described in the user manual, (Version F0, September 2022). 50 pL beads per sample were washed and resuspended in 200 pL of binding buffer. Amplified DNA was then combined with the beads for 30 minutes at room temperature followed by a wash step using prewarmed (70°C) SureSelect wash. DNA was eluted in nuclease-free water.
[0174] After capture, the eluted DNA was amplified by isothermal amplification using Herculease II Fusion DNA polymerase, SureSelect Post-Capture primer mix and 14 cycles of amplification (consisting of 30 seconds at 98°C, 30 seconds at 60°C and 1 minute at 72°C), a final elongation step of 5 minutes at 72°C followed by magnetic bead purification using AMPure XP beads and elution in IE buffer.
[0175] Following the pan-cancer enrichment step, a cancer- type specific qPCR detection step was performed. Post capture qPCR can be carried out by any known method.
[0176] DNA was amplified in PCR master mix (NEB #M0541) and SYBR Green I (Invitrogen #S-7563) and specific primer pairs for amplifying the cancer regions of interest. Each qPCR reaction was performed in triplicate with the same starting amount of DNA (0.75-1ng / reaction). Each qPCR experiment included a water and a “stable” primer pair as negative controls. The negative control primer pair (referred to as NEG2) was directed to a gene desert region (chromosome 4; start position: 137956850; end position: 137957300). The sequences of the NEG2 primer pair are as set forth in SEQ ID NOs: 9760 and 9761.
[0177] Samples were diluted in elution buffer to reach the desired concentration using the following calculation:DNA amount per reaction Desired concentration = - - — -Vf = 5 pL
[0178] Suitable PCR conditions can be determined for each primer pair using an annealing temperature based on the primers’ melting temperature (5 °C below the lowest Tmfor primers <20 bp). Standard PCR conditions are set out in Table 6. The annealing temperature is denoted by **, as this is determined by the specific primer pairs included for the reaction.Table 6. PCR conditions
[0179] Ct values were determined from the qPCR output using the appropriate software.
[0180] The “normalized” Ct for each Target and Sample combination were calculated as follows:
[0181] The difference between each sample and the control, for each target was calculated as follows:(sample, target) C t (samp le,targ et) Ct control, tar get)
[0182] The fold change value (2A-) was determined by the following calculation:Example 2. Probe capture significantly increases assay sensitivity
[0183] This example demonstrates that combining the steps of enriching cfDNA for (i) chromatin-associated DNA using chromatin immunoprecipitation (ChIP) and (ii) known gene regulatory sequences associated with the presence of cancer using probe capture dramatically increases the sensitivity and specificity of the detection of such sequences in a liquid biopsy.
[0184] As proof of principle, the method workflow set out in Example 1 was tested with a nucleotide sequence associated with the presence of bladder cancer (the RBI promoter located on chromosome 13 at location 48878039-48878195). Urine samples were obtained from subjects suffering from high-grade muscle-invasive bladder cancer (MIBC). cfDNA was isolated from the urine samples. An anti-H3K27ac antibody (Diagenode, C15410196) was used to isolate chromatin-associated DNA via ChIP. A 157 nucleotide-long capture probe for the / / U promoter was used for probe capture-based enrichment. The oligonucleotide probe specifically hybridized to the nucleotide sequence set forth in SEQ ID NO: 6370. Following probe capture, qPCR was performed using primers with the sequences as set forth in SEQ ID NOs: 9754 and 9755.
[0185] Samples were processed at different stages of the method workflow as described in Example 1. FIG. 1 illustrates the fold change (left) or log2 scale (right) of the RBI promoter sequence determined by qPCR. “Input” refers to isolated cfDNA, “IP” refers to chromatin- associated DNA obtained by chromatin immunoprecipitation (ChIP), and “Capture” relates to the enriched DNA obtained by probe capture. As shown in FIG. 1, there was no detectable qPCR signal for the RBI promoter sequence in the input cfDNA. After ChIP, but without probe capture, the detectable signal was negligible. Use of ChIP in combination with capture probe dramatically increased sensitivity of the qPCR assay.
[0186] This example illustrates that the method workflow of Example 1 provides a sensitive method for detecting a nucleotide sequence associated with the presence of cancer in a sample obtained from a subject. In particular, this example demonstrates the successful qPCR-based detection of an RBI promoter sequence in cfDNA isolated from urine samples of subjects suffering from bladder cancer and enriched by ChIP and probe capture. The method described above provides a sensitive and specific qPCR-based test for detecting and diagnosing cancer using cfDNA isolated from a liquid biopsy.Example 3. Chromatin analysis of bladder cancer sample reveals three clusters
[0187] This example describes the identification of three distinguished clusters of H3K27ac-associated (transcriptionally active) gene regulatory regions (promoters and enhancers) in bladder cancer samples following chromatin analysis.
[0188] Formalin-fixed paraffin-embedded (FFPE) archived clinical tissues from patients with non-muscle invasive bladder cancer (NMIBC) were obtained from collections at Hospital del Mar-Parc de Salut Mar-Biobank, Barcelona, Spain. Of the 17 selected NMIBC FFPE samples (labelled “UCC1” - “UCC17”), 6 had a micropapillary component (MP), which is associated with a worse clinical outcome. To increase the enrichment in cancer cells, FFPE sections were macro-dissected to obtain >80% tumor cells. The status of activation of enhancers and superenhancers (SE) was determined by H3K27ac profiling using fixed-tissue ChlP-seq for H3K27ac (FiTAc-seq) analysis on FFPE tissues, as previously described (Font-Tello et al. Nat Protoc. 2020; 15(8):2503-2518). Samples were sectioned at a thickness of 10 pm (n=10), incubated in xylene to remove paraffin and rehydrated in an ascending ethanol series.
[0189] The tissue was resuspended in lysis buffer and sonicated for 5 minutes using a Covaris E220 instrument (setting: 140 peak incident power, 5% duty factor, and 200 cycles per burst) in 1 ml adaptive focused acoustics (AFA) fiber millitubes. Soluble chromatin (5 pg) was immunoprecipitated with 10 pg of anti-H3K27ac antibody (Diagenode catalog number C15410196). ChlP-seq libraries were constructed using ThruPLEX-FD kits (Rubicon Genomics) following the manufacturer’s protocols. 75 -bp single-end reads were sequenced on a Nextseq instrument (Illumina).
[0190] All samples were processed through the computational pipeline developed at the Dana-Farber Cancer Institute Center for Functional Cancer Epigenetics using primarily open- source programs (github.com / liulab-dfci / CHIPS; Qiu et al. Genomics Proteomics Bioinformatics. 2021; 19(4): 652-661). Sequence reads were aligned with Burrows-WheelerAligner (BWA; Li and Durbin. Bioinformatics. 2009; 25(14): 1754-60) to build hgl9 and uniquely mapped, non-redundant reads were retained. These reads were used to generate binding sites with Model-Based Analysis of ChlP-Seq 2 (MACS v2.1.1.20160309), with a q-value (FDR) threshold of 0.01 (Zhang et al. Genome Biol. 2008;9(9):R137). Multiple quality control criteria were evaluated based on alignment information and peak quality: (i) sequence quality score; (ii) uniquely mappable reads (reads that can only map to one location in the genome); (iii) uniquely mappable locations (locations that can only be mapped by at least one read); (iv) peak overlap with Velcro regions, a comprehensive set of locations - also called consensus signal artifact regions - in the genome that have anomalous, unstructured high signal or read counts in next-generation sequencing (NGS) experiments independent of cell line and of type of experiment; (v) number of total peaks (the minimum required was 8,000); (vi) high-confidence peaks (the number of peaks that are tenfold enriched over background); (vii) overlap with known DNase I hypersensitive regions derived from the ENCODE Project (the minimum required was an 80% overlap); and (viii) peak conservation (a measure of sequence similarity across species based on the hypothesis that conserved sequences are more likely to be functional). Genome tracks were visualized by IGV (v2.14.1; Robinson etal. Cancer Res. 2017; 77(21):e31-e34).
[0191] The analysis produced high-quality results in terms of identified peaks and fraction of reads in peaks. Activation of enhancers at previously known genes associated with bladder cancer such as NECTIN4 and SOX4 was observed, validating the specificity and quality of the results. Unsupervised principal component analysis of the H3K27ac-bound DNA comprising (transcriptionally active) gene regulatory regions (promoter and enhancers) distinguished three clusters among the 17 NMIBC tissue samples, as shown in FIG. 2. One cluster corresponded to the 6 tissue samples with MP histology, and the other two subdivided the urothelial cancers (UC). Of the two UC clusters, the cluster labelled as UC-LLI comprising 5 of the test samples was closer to the micropapillary (MP) than to the other UC cluster labelled as UC-BL comprising 6 of the test samples.
[0192] As shown in FIG. 3, analysis of the differentially activated gene regulatory regions that distinguished UC-LLI from UC-BL revealed a set of nucleotide sequence peaks characteristic of UC-LLI and a set of nucleotide sequence peaks characteristic of UC-BL. The set of peaks characteristic of UC-BL is indicated in the set of panels in row A of FIG. 3 as thetop trace running above the other two traces in samples “UCC1”, UCC8”, “UCC9”, “UCC10”, “UCCH”, and “UCC13”, and as the bottom trace in samples “UCC3”, “UCC4”, “UCC5”, “UCC7”, and “UCC12”. The set of peaks characteristic of UC-LL1 is indicated as the bottom trace in black in samples “UCC1”, UCC8”, “UCC9”, “UCC10”, “UCCH”, and “UCC13” and as the top trace (also in black) running above the other two traces in samples “UCC3”, “UCC4”, “UCC5”, “UCC7”, and “UCC12”. The middle trace shows the set of peaks common to all 13 samples represented in FIG. 3.
[0193] That the set of peaks characteristic of UC-BL and the set of peaks characteristic of UC-LLI can be clearly distinguished can be seen from a comparison of the panels in rows B and C of FIG. 3. The row B panels represent the nucleotide sequence peaks characteristic of the UC-LLI cluster. The row C panels represent the nucleotide sequence peaks characteristic of the UC-BL cluster. In particular, a comparison of the panels in rows B and C reveals that samples “UCC1”, UCC8”, “UCC9”, “UCC10”, “UCCH”, and “UCC13” had virtually no overlap with the peak set characteristic of the UC-LLI cluster. Similarly, there was virtually no overlap for samples “UCC3” and “UCC12” with the peak set characteristic of UC-BL cluster.
[0194] The panels in row D of FIG. 3 show the remaining nucleotide sequence peaks. These include the set of nucleotide sequence peaks common to all 13 samples. These include nucleotide sequence peaks associated with gene regulatory regions for NECTIN4, SOX4, FGFR3, TP63, KDM6A, KRT5, SPP1, and IGFBP3 which are well-documented genes associated with bladder cancer. The identified nucleotide sequences associated with NECTIN4 (SEQ ID NOs: 2273 and 6365), SOX4 (SEQ ID NO: 6366), FGFR3 (SEQ ID NO: 1211), TP63 (SEQ ID NO: 3187), KDM6A (SEQ ID NO: 6367), KRT5 (SEQ ID NO: 6368), SPP1 (SEQ ID NO: 3314) and IGFBP3 (SEQ ID NO: 4037) can be used to design oligonucleotide probes and primer pairs that can be employed in the method described in Example 1 to determine whether a subject is likely suffering from bladder cancer. The presence or increased levels of these bladder cancer-associated nucleotide sequences in H3K27ac-enriched cfDNA isolated from a liquid biopsy such as blood or urine can indicate bladder cancer.
[0195] This example demonstrates that distinguishable clusters of transcriptionally active gene regulatory regions in bladder cancer samples can be identified based on the analysis of H3K27ac-enriched DNA.Example 4. Identification of basal-like and luminal-like bladder cancer subtypes
[0196] This example describes the identification of basal-like and luminal-like subtypes of bladder cancer.
[0197] Peaks from all samples were merged to create a union set of sites for each transcription factor and histone mark using BEDOPS (Neph etal. Bioinformatics. 2012; 28(14): 1919-20). Sample-sample correlation and differential peaks analysis were performed by the CoBRA pipeline (Qiu el al. (supra). To investigate the transcriptional mechanisms underlying the chromatin differences, CoBRA uses HomER (Huppert el al. Appl Opt. 2009; 48(10): D280- 98) to assess the enrichment of transcription factors (TF) DNA-binding motifs at the differential regions.
[0198] Read densities were calculated for each peak for each sample and used for the comparison of cistromes (gene regulatory elements including transcription factor binding sites and sites of histone modification) across samples. Sample similarity was determined by hierarchical clustering using the Spearman correlation between samples. Differential peaks were identified by DEseq2 with adjusted P < 0.05 and |log2FoldChange| > 0.5. A total number of reads in each sample was applied to the size factor in DEseq2, which can normalize the sequencing depth between samples. Peaks from each group were used for motif analysis by the motif search findMotifsGenome.pl in HomER2 (v3.0.0; Huppert el al. (supra)), with cutoff q-value < le-10. The signals of each sample on differential binding sites were visualized by deepTools (Ramirez etal. Nucleic Acids Res. 2014; 42(Web Server issue): W187-91).
[0199] The differential gene regulatory regions for the UC-BL cluster identified in Example 3 comprising “UCC1”, UCC8”, “UCC9”, “UCC10”, “UCC11”, and “UCC13” showed TP63 as the top enriched motif, suggesting more basal-like characteristics for this cluster. Closer to the MP than to UC-BL, the UC-LLI cluster identified in Example 3 showed enrichment in motifs for inflammatory related TFs such as IRF2 supporting a more luminal-like-inflammatory phenotype. The analysis also identified a luminal origin of the MP cluster identified in Example 3 by showing enrichment in GRHL2 as the top motif of the differential gene regulatory regions for the MP cluster. In summary, the chromatin analysis described in Example 3 identified threedistinct clusters of which MP and UC-LLI have a luminal phenotype and UC-BL has a basal phenotype.
[0200] To investigate the subtypes without the potential confounding parameter of a distinct histology, the differences between the two UC clusters was assessed. The comparison of UC-LLI with UC-BL resulted in the identification of 3726 gene regulatory regions differentially H3K27ac-activated in UC-LLI (visually represented in the row B panels shown in FIG. 3) and 9872 differentially H3K27ac-activated in UC-BL (visually represented the row C panels shown in FIG. 3). Nucleotide sequences associated with the presence of bladder cancer that were common to both the UC-LLI and UC-BL clusters are set forth in SEQ ID NOs: 1-6372. Nucleotide sequences specifically associated with luminal-like bladder cancer are set forth in SEQ ID NOs: 6373-7363. Nucleotide sequences specifically associated with basal-like bladder cancer are set forth in SEQ ID NOs: 7635-9733. Motif analysis of these two sets of gene regulatory regions showed similar results to what was observed in the previous analysis.
[0201] The differentially active gene regulatory regions were compared to published chromatin immunoprecipitation sequencing (ChlP-seq) profiles compiled in CistromeDB (Taing et al. Nucleic Acids Res. 2024; 52(D1):D61-D66). The results showed that top UC-LLI-peaks overlap with STAT4 and CD74, validating the luminal-like inflammatory phenotype revealed by the motif analysis. The UC-BL-activated regions showed the highest overlap with TP63 analyzed in keratinocytes validating the basal characteristics. Therefore, the analysis distinguishes a distinct phenotype for UC-LLI and UC-BL, with UC-BL featuring more basal-like (BL) and UC-LLI more luminal like-inflammatory (LLI) chromatin characteristics.
[0202] This example illustrates that subtypes of bladder cancer can be identified and distinguished from one another by detecting differentially active gene regulatory regions.Example 5. Bladder cancer detection
[0203] This example demonstrates that the bladder cancer-associated gene regulatory sequences can be used to detect bladder cancer in a liquid biopsy sample using the assay described in Example 1.
[0204] To further validate the method described in Example 1, the experiment described in Example 2 was repeated using a different set of oligonucleotide probes and primers. Theprimers had a length of 18-24 bases, a G / C content of 40-60%, and a melting temperature (Tm) of 50-60°C. Each primer pair was designed to have a Tmwithin 5°C of each other. In addition, the primers were designed to start and end with 1-2 G / C pairs.
[0205] An important characteristic of bladder cancer across classification stages is the occurrence of recurring mutations. Some of these mutations affect genes involved in chromatin structure, such as KMT2D and KDM6A. Therefore, oligonucleotide probes and primer pairs were designed to target gene regulatory regions of these bladder cancer-associated genes. In addition, oligonucleotide probes and primer pairs were designed to target the gene regulatory region of the bladder cancer-associated gene 7 / 75. The gene regulatory region of the constitutively active housekeeping gene GAPDH served as a control. The GAPDH primer pair amplifies a 76-bp fragment from intron 7 of GAPDH (Active Motif, Cat. No.: 71004). The oligonucleotide probes were designed to specifically hybridize to the nucleotide sequences as set forth in SEQ ID NOs: 6367-6369. The nucleotide sequences of the primers are shown in Table 7.Table 7. Primer pairs
[0206] The results of the qPCR amplification step are summarized in FIG. 4. GAPDH is shown in panel A, KDM6A in panel B, KMT2D in panel C, and KRT5 in panel D. “IP” refers to chromatin-associated DNA obtained by chromatin immunoprecipitation (ChIP), and “Capture” relates to the enriched DNA obtained by probe capture.
[0207] As can be seen from FIG. 4, using primers specific for gene regulatory regions of bladder cancer-associated genes KDMA6 (panel B), KMT2D (panel C), and KRT5 (panel D)provided a highly sensitive assay for detecting the presence of bladder cancer in liquid biopsy samples. As expected, the increase in sensitivity was specific to nucleotide sequences associated with the presence of bladder cancer. In comparison, the control sequence of the housekeeping gene GAPDH (panel A) did not show the same increase in sensitivity when steps 1 -3 were used.
[0208] This example demonstrates that the method of Example 1 can be used to detect nucleotide sequences associated with bladder cancer to provide a sensitive and specific assay for determining the presence of bladder cancer in isolated cfDNA from a liquid biopsy sample.Example 6. Differentially expressed genes associated with bladder cancer subtypes
[0209] This example demonstrates that the top differential peaks identified in the H3K27ac ChlP-seq analysis can be correlated with overexpressed genes. These overexpressed genes therefore can serve as surrogates to the gene regulatory regions in the classification of bladder cancer.
[0210] The differentially activated gene regulatory sequences (promoters and enhancers) identified in Example 3 were integrated with gene expression data from a subset of the initial tissue sample cohort by combining the differential peaks between UC-LLI and UC-BL with the corresponding differential gene expression.
[0211] Read alignment, quality control, and data analysis were performed using the Visualization Pipeline for RNA-seq (VIPER; Cornwell et al. BMC Bioinformatics. 2018; 19(1): 135). Alignment to the hgl9 human genome was performed using STAR v2.7.0f (Dobin etal. Bioinformatics. 2013; 29(1): 15-21) followed by transcript assembly using cufflinks v2.2.1 (Dobin et al. supra),' Trapnell et al. Nat Biotechnol. 2010; 28(5): 511-5) and RseQC v2.6.2 (Wang and Li, Bioinformatics. 2012; 28(16):2184-5). Differential gene expression analyses were performed comparing UC-LLI to UC-BL using DESeq2 vl.18.1 (Wang and Li supra),' Love et al. (Genome Biol. 2014; 15(12): 550), utilizing absolute gene counts for RNA-Seq data and raw read counts for transcriptomic profiling data.
[0212] The samples in the UC-LLI and UC-BL groups were matched with samples from H3K27ac ChlP-seq analysis. Specifically, in the H3K27ac dataset, there were 5 samples for UC- LLI and 6 for UC-BL, while in the RNA-seq cohort, there were 3 samples for UC-LLI and 4 for UC-BL. Gene Set Enrichment Analysis (GSEA) was conducted using the GSEA software(GSEA Java; v4.1.0) with Hallmark gene sets. Genes were pre-ranked based on Log2FC for the BL versus LLI comparison, and enrichment scores were computed (p = 1, weighted).
[0213] The differentially expressed genes from the RNA-seq analysis were visualized using a volcano plot as shown in FIG. 5A (luminal-like) and FIG. 5B (basal-like). These genes were chosen based on their correlation with the top differential peaks (ranked by adjusted p- value) identified in the corresponding H3K27ac ChlP-seq analysis (see Example 3). The resulting gene signatures for luminal-like and basal-like bladder cancer subtypes are summarized in Table 8.Table 8. Gene signatures for luminal-like and basal-like bladder cancer subtypes
[0214] As shown by the volcano plot in FIG. 5B, the top differentially expressed genes with corresponding differential enhancer activation in their vicinity, included genes involved in the basal phenotype like BMP7, ERN2, and SRPX2 for the BL cluster (FIG. 5B). The LLI side showed a more complex phenotype including genes associated with immune pathways like FBN2, SLC4A4, CTTNBP2, KALRN and SELL (FIG. 5A).
[0215] The relevance of the inflammatory genes for the UC-LLI was further revealed by the presence of Superenhancers (SE) at some of these genes as is the case for SELL and SELE along with SE at luminal genes like GATA3. Conversely, UC-BL showed activation of a transcriptional circuit for TP63 where the motif enrichment is accompanied by the activation of the super-enhancer at the TP63 locus potentially involved in the maintenance of the basal lineage for the UC-BL cluster (Saint-Andre et al. Genome Res. 2016;26(3):385-96), which is further reflected by the existence of a SE at KRT5.
[0216] This example demonstrates that the top differential peaks identified in the H3K27ac ChlP-seq analysis can be correlated with overexpressed genes which thus can serve as surrogates to the gene regulatory regions in classification of bladder cancer.Example 7. Subtype classification in bladder cancer patient cohorts
[0217] This example demonstrates a significant association between luminal-like cases and a higher risk of progression, in particular to MIBC.
[0218] To investigate the representation of the two subtypes (basal-like and luminal-like) in a larger HGT1 patient cohort, the chromatin derived score (CDS) was applied to score a bulk RNA-seq dataset composed of 62 cases (Bowden etal. Sci Rep. 2020;10(l):20135). The results showed a spectrum of scores ranging from cases with a pure basal-like score to cases with a pure luminal-like score, separated by a continuum of cases with mixed characteristics. A similar continuum distribution of scores was observed when evaluating the UROMOL cohort (Lindskrog et al. Nat Commun. 2021;12(l):2301), an independent dataset composed of 438 NMIBC cases that included 78 HGT1 tumors. For both cohorts, luminal-like cases significantly over expressed KRT20, while basal-like cases showed significant overexpression of KRT5 and TP63. Other previously described luminal markers, including PPARG and GAT A 3 and CDH3, did not show association with any of the subtypes.
[0219] Lindskrog et al. (supra) described a scoring system by the analysis of this UROMOL cohort (Hedegaard et al. Cancer Cell. 2016;30(l):27-42; Lindskrog et al. (supra)). The CDS score classification was compared with the Lindskrog system on both cohorts (62HGT1 and UROMOL). The Lindskrog classification divides NMIBC cases into 4 classes with different degrees of luminal and basal characteristics (Hedegaard et al. (supra)). While distinct, the two classification systems shared a number of features.
[0220] First, both the basal-like subtype and classes 1 and 3 display a higher expression of TP63, while luminal-like and Class 2a are both associated with higher KRT20 and ERBB2 expression. This result shows that, despite the difference in number of subtypes, the observed commonalities validate the existence of cases with distinct molecular characteristics within NMIBC.
[0221] Second, Class 2a in the Lindskrog classification has been associated with a higher risk of progression. In accordance with the luminal-like and Class 2a similarities (p value: 9 x 10-8 in 62 HGT1; and p- value: 1 xlO-4 for the UROMOL), a significant association (p-value: 5.39 x 10-7) was observed between luminal-like cases and the higher risk HGT1 cases in the UROMOL cohort. Moreover, when the UROMOL cases were stratified as either basal-like or luminal-like (based on the median CDS score) and survival analysis was performed, a significant association between luminal-like cases and a higher risk of progression was observed. Despitethe limited number of cases, the 62HGT1 cohort also showed an association between the luminal-like subtype and progression to muscle-invasive bladder cancer (MIBC; p-value: 0.02).
[0222] This example demonstrates a significant association between luminal-like cases and a higher risk of progression, in particular to MIBC.Example 8. Intratumor heterogeneity
[0223] This example shows significant intratumor heterogeneity, with different cancer cell populations within the same tumor showing either basal-like or luminal-like phenotypes.
[0224] Immunohistochemistry (IHC) on tissue microarrays was performed. The tissue arrays were made from a collection of 162 bladder tumor samples, composed mainly of HGT1 (102 cases) but also containing 41 low-grade tumors and 19 MIBC (Lloreta et al. Hum Pathol. 2017:62:222-231). Tissue sections were deparaffinized in xylene and rehydrated using an ascending ethanol series. After antigen retrieval, slides were treated with 3% H2O2 in PBS for 10 min to quench endogenous peroxidases, washed, and incubated in blocking solution (PBS containing 1% BSA and 1% Tween-20) for Ih at ambient temperature. Slides were incubated with antibodies targeting TP63 (FLEX Monoclonal Mouse Anti-Human p63 Protein, VENTANA anti-p63 (4A4), Part Number: GA66261-2), KRT5 (KRT 5 VENTANA anti- Cytokeratin 5 / 6 (D5 / 16B4) Mouse Monoclonal Primary Antibody), and KRT20 (KRT 20 VENTANA CONFIRM anti-Cytokeratin 20 (SP33) Rabbit Monoclonal Primary Antibody). Antibodies were diluted in blocking solution for 1 h. Slides were washed in PBS and incubated with the peroxidase-based EnVision Kit (Dako).
[0225] Due to the high association between KRT20 and the luminal-like subtype and between KRT5 and TP63 and the basal-like subtype, these markers were used as surrogates to track the chromatin subtypes. The results revealed that 53 (44%) tumors showed intratumor heterogeneity in terms of the coexistence of KRT5- and KRT20-expression in the same specimen. At the cellular level, the expression of KRT5 and KRT20 in these “mixed” cases showed anti correlation, with different cancer cells expressing KRT5 and KRT20. In contrast, 33 (27%) tumors showed homogenous expression of KRT5 cancer cells (pure BL cases) and 26 (21%) showed homogenous expression of KRT20 expression (pure LLI cases). Ten additional tumors (8%) were negative for both markers. The distribution of mixed cases is in goodagreement with the CDS scoring of the RNA-seq cohorts described in Example 7, suggesting that the approximately 50% of CDS scored cases with mixed characteristic could be attributed to subtypes coexistence.
[0226] Interestingly, the subtype localization showed a remarkable spatial feature in mixed cases. A consistent pattern was observed where KRT5 expressing cells were located in close proximity to the vascular stroma, while the KRT20 cells showed to be located toward the interior of the tumor. In general, the KRT5 cells constituted the first layer of cancer cells paving the contact with the vascular stroma.
[0227] The expression of TP63 showed high correlation with KRT5, which showed to be the subset of the stronger expressing TP63 cells. The spatial location of the KRT5 expressing cells is consistent with the lower association with angiogenesis revealed by the cell-cell interactions and further emphasizes the biological differences between basal-like and luminallike phenotypes.
[0228] This example shows that, although some bladder cancers can be clearly assigned to either a basal-like or a luminal-like subtype, significant intratumor heterogeneity can exist, with different cancer cell populations within the same tumor showing either basal-like or luminal-like phenotypes.Example 9. Identification of gene regulatory regions for blood-based diagnostic assay
[0229] This example illustrates the selection of nucleotide sequences of gene regulatory regions associated with a cancer or cancer subtype for use in a blood-based diagnostic assay.
[0230] A comparison of the signals at gene regulatory regions of the top differentially expressed genes obtained from ChlP-seq data was performed to select gene regulatory regions for a diagnostic assay. This initial analysis identified nucleotide sequences associated with bladder cancer (see Table 1A), and nucleotide sequences associated with bladder cancer subtypes, specifically, luminal-like bladder cancer (see Table 2A), and basal-like bladder cancer (see Table 3A).
[0231] To further refine the assay for use with cfDNA, ChlP-seq data from luminal-like bladder cancer samples, micropapillary bladder cancer samples, and basal-like bladder cancersamples were compared to plasma samples isolated from healthy (non-cancerous) control patients. This analysis was performed to ensure that the signals detected in a liquid biopsy sample, such as a blood sample, do not come from non-cancer origins, e.g., cfDNA of healthy cells found in the blood, and thus can reliably be attributed to the cancer cells.
[0232] FIG. 6 illustrates the chromatin signal around the ERBB2 gene, where a peak corresponds to a positive signal. A strong signal was observed at this locus across all bladder cancer samples, regardless of subtype. Accordingly, this gene can be considered a differentially expressed candidate gene for the detecting the presence of bladder cancer.
[0233] Despite the differential expression of the ERBB2 gene, a strong chromatin signal was observed at the promoter region marked as “A” in FIG. 6 in both the healthy control sample and the bladder cancer samples. Accordingly, the use of a nucleotide sequence from this part of the gene regulatory region would confound the results of a diagnostic assay. In contrast, strong signals were observed in the bladder cancer samples, but not the healthy control sample at a nearby site marked as “B” in FIG. 6. Accordingly, a nucleotide sequence from this part of the gene regulatory region is associated with the presence of bladder cancer. Corresponding analyses were performed to arrive at nucleotide sequences of gene regulatory regions associated with bladder cancer (see Table IB), and nucleotide sequences of gene regulatory regions associated with bladder cancer subtypes, specifically luminal-like bladder cancer (see Table 2B), basal-like bladder cancer (see Table 3B), and micropapillary bladder cancer (see Table 4).
[0234] This example illustrates the selection of nucleotide sequences of gene regulatory regions associated with a cancer or cancer subtype for use in a blood-based diagnostic assay.Example 10. Probe capture increases assay sensitivity
[0235] This example demonstrates that oligonucleotide probes can be used to enrich chromatin-associated DNA obtained from cfDNA isolated from plasma, serum, or urine, thereby increasing assay sensitivity.
[0236] CfDNA was isolated from plasma, serum, and urine obtained from two different patients and enriched for chromatin-associated DNA as described in Example 1. The sample was then divided up. The first aliquot was used to construct a first pre-capture library as described in Example 3. The second aliquot was contacted with a set of 8000 oligonucleotide probes, eachspecifically hybridizing to a distinct gene regulatory region, to further enrich the isolated chromatin-associated DNA for nucleotide sequences associated with the presence of bladder cancer, using the method of probe capture method described in Example 1. The 8000 oligonucleotide probes specifically hybridized to a subset of the chromosomal regions set forth in SEQ ID NOs: 1-9733, identified in Example 4. The further enriched isolated chromatin- associated DNA was used to construct a second post-capture library. Both libraries were then sequenced as described in Example 3.
[0237] FIG. 7 illustrates heatmaps of bladder cancer-associated gene regulatory regions at pre-capture (left panel) and post-capture stages (right panel) from plasma (P), serum (S), and urine (U) samples obtained from two bladder cancer patients. An aggregated NGS signal across 8000 capture regions is provided. Region representation is visualized by the black signal. Each row of the heatmap indicates one region out of the 8000 chromosomal regions enriched for by the 8000 oligonucleotide probes. As can be seen from FIG. 7, the probe capture process increased assay sensitivity for each of the three sample types. Most enrichment was observed in urine and serum.
[0238] It should be understood that the particular embodiments described herein are given by way of illustration only, not limitation. Other features, objects, and advantages are apparent from the above detailed description, drawings and examples. Various changes and modifications will be apparent to those skilled in the art.
[0239] All patents, patent publications and non-patent publications referenced herein are indicative of the level of skill of those skilled in the art to which this invention pertains. All these publications are herein incorporated by reference to the same extent as if each individual publication were specifically and individually indicated as being incorporated by reference.
Claims
AMENDED CLAIMS received by the International Bureau on 20 February 2026 (20.02.2026)1. A method for diagnosing a cancer in a subject, the method comprising: a) providing a sample obtained from the subject; b) isolating cell-free DNA (cfDNA) from the sample; c) contacting the cfDNA with a means for specifically binding chromatin to isolate chromatin-associated DNA; d) enriching the isolated chromatin-associated DNA using oligonucleotide probes that specifically hybridize to one or more nucleotide sequences associated with the presence of the cancer, wherein each oligonucleotide probe binds to a distinct region of the one or more nucleotide sequences; and e) performing quantitative PCR (qPCR) by contacting the enriched isolated chromatin- associated DNA from step (d) with one or more primer pairs to amplify and detect a portion of the one or more nucleotide sequences associated with the presence of the cancer; wherein the presence of or an increased level of the one or more nucleotide sequences associated with the presence of the cancer in the sample indicates that the subject has the cancer.
2. The method of claim 1, wherein each of the one or more nucleotide sequences associated with the presence of cancer has a length of 100-250 base pairs.
3. The method of claim 1 or 2, wherein each of the oligonucleotide probes has a length of about 20-30 nucleotides.
4. The method of any one of the preceding claims, wherein the amplified portion of the one or more nucleotide sequences associated with cancer is no more than 250 base pairs in length, optionally wherein the amplified portion is between 100 and 200 base pairs in length.
5. The method of any one of the preceding claims, further comprising a step of amplifying the enriched chromatin-associated DNA prior to performing qPCR, optionally wherein the amplification step is an isothermal amplification step.
856. The method of any one of the preceding claims, further comprising one or more step(s) of purifying: i) the isolated cfDNA; ii) the enriched chromatin-associated DNA; and / or iii) the amplified chromatin-associated DNA.
7. The method of any one of the preceding claims, wherein the means for specifically binding chromatin binds an acetylated or methylated histone protein, e.g., histone H3.
8. The method of any one of the preceding claims, wherein the means for specifically binding chromatin is an antibody.
9. The method of claim 8, wherein the antibody specifically binds to histone H3 protein i) acetylated at the lysine at residue 27 (H3K27ac), acetylated at the lysine at residue 9 (H3K9ac), protein acetylated at the lysine at residue 14 (H3K14ac), acetylated at the lysine at residue 18 (H3K18ac), or acetylated at the lysine at residue 23 (H3K23ac); or ii) methylated at the lysine at residue 4 (H3K4mel), methylated at the lysine at residue 4 (H3K4me3), methylated at the lysine at residue 27 (H3K27me3), methylated at the lysine at residue 9 (H3K9me3).
10. The method of any one of the preceding claims, wherein the sample is blood, plasma, serum, or urine.
11. The method of claim 10, wherein the sample is blood and the cfDNA comprises circulating tumor DNA (ctDNA).
12. The method of claim 11, wherein the ctDNA is comprised in exosomes.
13. The method of any one of the preceding claims, wherein the cancer is a solid cancer.8614. The method of claim 13, wherein the cancer is bladder cancer, e.g., basal-squamous bladder cancer, luminal bladder cancer, luminal-infiltrated bladder cancer, luminal-papillary bladder cancer, or micropapillary bladder cancer.
15. The method of claim 14, wherein the one or more nucleotide sequences associated with the presence of cancer are comprised with the regulatory region of one or more cancer-associated genes, optionally wherein the one or more cancer-associated genes are selected from the list consisting of NECTIN4, and KDM6A, and / or from the list consisting of CTTNBP2, KALRN, SELL, PMP22, COL12A1, PRTG, PDE9A, ANXA3, BMP7, ERN2, SRPX2, UGT1A10, CLCA2, SERPINB5, CLCA4, CLU, SH3PXD2A, AQP3, ANXA10, and SLITRK6.
16. The method of any one of the preceding claims, further comprising a step of treating the cancer of the subject, if required.
17. A method for monitoring the progression of a cancer, the method comprising: a) detecting, from a sample derived from a subject at a first time point, the presence and / or level of one or more nucleotide sequences associated with the presence of the cancer; b) detecting, from a sample derived from a subject at a second or subsequent time point, the presence and / or level of one or more nucleotide sequences associated with the presence of the cancer; and c) comparing the results from steps a) and b), thereby monitoring the progression of the cancer, wherein the measuring steps are carried out according to the method of any one of claims 1-16.
18. A method for classifying a type and / or subtype of a cancer in a subject, the method comprising: a) detecting, in a sample derived from the subject, the presence and / or level of one or more nucleotide sequences associated with the type and / or subtype of the cancer; and optionally87b) comparing the results from step a) to a reference data set for the cancer type and / or subtype; wherein the detecting step is carried out according to the method of any one of claims 1-16.
19. A kit or device comprising: a) reagents for isolating cfDNA; b) a means for specifically binding chromatin to isolate chromatin-associated DNA; and c) oligonucleotide probes that specifically hybridize to one or more nucleotide sequences associated with the presence of cancer in isolated chromatin-associated DNA, wherein each oligonucleotide probe binds to a distinct region of the one or more nucleotide sequences; and optionally d) one or more primer pairs for amplifying a portion of the one or more nucleotide sequences associated with the presence of the cancer.
20. The kit or device of claim 19, wherein the one or more nucleotide sequences associated with cancer are comprised within the regulatory region of one or more cancer-associated genes, optionally wherein the one or more cancer-associated genes are selected from the list consisting e NECTIN4, and KDM6A, and / or from the list consisting o CTTNBP2, KALRN, SELL, PMP22, COL12A1, PRTG, PDE9A, ANXA3, BMP7, ERN2, SRPX2, UGT1A10, CLCA2, SERPINB5, CLCA4, CLU, SH3PXD2A, AQP3, ANXA10, and SLITRK6.
21. The kit or device of any one of claims 19 or 20 for use with the method of any one of claims 1-18.
22. A method of preparing a library of isolated cell-free chromatin-associated DNA comprising nucleotide sequences associated with the presence of cancer, said method comprising: a) providing a sample obtained from a subject with cancer; b) isolating cell-free DNA (cfDNA) from the sample; c) contacting the cfDNA with a means for specifically binding chromatin to isolate chromatin-associated DNA;88d) enriching the isolated chromatin-associated DNA using oligonucleotide probes that specifically hybridize to one or more nucleotide sequences associated with the presence of the cancer, wherein each oligonucleotide probe binds to a distinct region of the one or more nucleotide sequences; and optionally e) amplifying the chromatin-associated DNA.
23. The method of claim 22, wherein amplifying the chromatin-associated DNA is performed using an isothermal amplification reaction.
24. The method of claim 22 or 23, wherein the nucleotide sequences associated with the presence of the cancer are regulatory regions of cancer-associated genes, optionally wherein the cancer-associated genes are selected from the group consisting of NECTIN4, and KDM6A, and / or the group consisting of CTTNBP2, KALRN, SELL, PMP22, COL12A1, PRTG, PDE9A, ANXA3, BMP7, ERN2, SRPX2, UGT1A10, CLCA2, SERPINB5, CLCA4, CLU, SH3PXD2A, AQP3, ANXA10, and SLITRK6.
25. The method of any one of claims 22-24, wherein the means for specifically binding chromatin is an antibody.
26. The method of claim 25, wherein the antibody specifically binds to histone H3 protein i) acetylated at the lysine at residue 27 (H3K27ac), acetylated at the lysine at residue 9 (H3K9ac), protein acetylated at the lysine at residue 14 (H3K14ac), acetylated at the lysine at residue 18 (H3K18ac), or acetylated at the lysine at residue 23 (H3K23ac); or ii) methylated at the lysine at residue 4 (H3K4mel), methylated at the lysine at residue 4 (H3K4me3), methylated at the lysine at residue 27 (H3K27me3), methylated at the lysine at residue 9 (H3K9me3).
27. A method for classifying a bladder cancer as luminal-like or basal-like, the method comprising:(a) providing a sample obtained from a subject suffering from bladder cancer; and89(b) determining in the sample:(i) expression of, and / or activation state of a gene regulatory region of, one or more of BMP7, ERN2, SRPX2, UGT1A10, CLCA2, SERPINB5, CLCA4, CLU, SH3PXD2A, AQP3, ANXA10 and SLITRK6; and(ii) expression of, and / or activation state of a gene regulatory region of, one or more of CTTNBP2, KALRN, SELL, PMP22, COL12A1, PRTG, PDE9A, and ANXA3, wherein overexpression of, or an active state of the gene regulatory region of, one or more of BMP7, ERN2, SRPX2, UGT1A10, CLCA2, SERPINB5, CLCA4, CLU, SH3PXD2A, AQP3, ANXA10 and SLITRK6 indicates that the bladder cancer is basal-like; and wherein overexpression of, or an active state of the gene regulatory region of, one or more <XCTTNBP2, KALRN, SELL, PMP22, COL12A1, PRTG, PDE9A, and / or ANXA3 indicates that the bladder cancer is luminal-like.
28. The method of claim 27, wherein the method comprises determining in the sample: a. expression of, and / or activation state of the gene regulatory regions of, at least 5 or more, at least 10 or more, or all of BMP7, ERN2, SRPX2, UGT1A10, CLCA2, SERPINB5, CLCA4, CLU, SH3PXD2A, AQP3, ANXA10 and SLITRK6; and b. expression of, and / or activation state of the gene regulatory regions of, at least 5 or more, at least 7 or more, or all of CTTNBP2, KALRN, SELL, PMP22, COL12A1, PRTG, PDE9A, m<3 ANXA3.
29. A computer-implemented method for classifying a bladder cancer in a subject as luminallike or basal-like, the method comprising:(a) receiving data indicative of expression of, and / or activation state of a gene regulatory region of:(i) one or more of CTTNBP2, KALRN, SELL, PMP22, COL12A1, PRTG, PDE9A, anA ANXA3, and(ii) one or more of BMP7, ERN2, SRPX2, UGT1A10, CLCA2, SERPINB5, CLCA4, CLU, SH3PXD2A, AQP3, ANXA10, and SLITRK6, and optionally(iii) one or both of NECTIN4, and KDM6A,90in a sample collected from the subject;(b) obtaining a statistical model;(c) comparing the data to the statistical model to generate a score; and(d) outputting the score from the statistical model, wherein the score indicates the probability of the bladder cancer being luminal-like or basal-like.
30. The method of claim 29, wherein step (a) comprises receiving data indicative of the expression of, and / or the activation state of the gene regulatory region of:(i) at least 5 or more, at least 7 or more, or all of CTTNBP2, KALRN, SELL, PMP22, COL12A1, PRTG, PDE9A, anA ANXA3’ and(ii) at least 5 or more, at least 10 or more, or all of BMP7, ERN2, SRPX2, UGT1A10, CLCA2, SERPINB5, CLCA4, CLU, SH3PXD2A,AQP3,ANXA10, and SLITRK6,' and optionally(iii) at least one, or both of NECTIN4 and KDM6A.91