Pan-cancer early detection and MRD CFDNA methylation
By identifying common DNA methylation alterations in pediatric cancers and using a machine learning model, the method enhances the sensitivity and specificity of cancer detection and monitoring, addressing the limitations of current diagnostic methods.
Patent Information
- Application Number
- JP2025547992
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-02-22
- Filing Date
- 2024-02-22
- Publication Date
- 2026-02-27
AI Technical Summary
Current methods for distinguishing between childhood cancers, adult cancers, and their recurrences are not sensitive and specific enough, and the rarity of pediatric cancers hinders the development of cancer-type-specific diagnostic, prognostic, or predictive biomarkers.
Identify minimal, differentially methylated regions (mDMRs) common to multiple pediatric cancers through whole-genome bisulfite sequencing, validate these regions in plasma cfDNA, and use a machine learning model to detect pediatric cancer, adult cancer, or minimal residual disease (MRD) based on methylation patterns.
Provides a basis for cfDNA methylation biomarkers for early detection, treatment response, and MRD monitoring in pediatric cancers, offering both scientific and clinical value.
Smart Images

Figure 2026506978000001_ABST
Abstract
Description
[Technical Field]
[0001] Related Applications This application claims priority under 35 U.S.C. § 119(e) to U.S. Provisional Patent Application No. 63 / 486,379, filed February 22, 2023, which is incorporated herein by reference. [Background technology]
[0002] Cancer is the second leading cause of death among children aged 1 to 14 years in the United States, with approximately 11,000 new cases and 1,200 deaths per year. The 5-year overall survival rate for childhood cancers has risen dramatically in recent years, from 58% in 1975 to nearly 85% in 2020. However, neuroblastoma—the third most common form of childhood cancer after leukemia and central nervous system (CNS) tumors—has a 75% 5-year overall survival rate, which plummets to 20% after the first disease recurrence. Improvements in overall survival for metastatic childhood cancers have been mixed over the past few decades, with neuroblastoma showing improved prognosis due to advances in treatment, while others, such as rhabdomyosarcoma and Ewing's sarcoma, have seen little improvement. A major factor limiting progress here is the rarity of many of these cancers, which limits research opportunities and hinders clinical trials.
[0003] Compared to adult-onset cancers, pediatric cancers generally have a much lower mutational burden. While the genomes of pediatric cancers are characteristically "quiet" and have a lower mutational burden, the epigenomes of many pediatric cancers appear "noisy," harboring widespread DNA methylation and histone marker alterations, as well as driver mutations in chromatin modifiers, such as SMARCB1 in malignant rhabdoid tumors. In this regard, targeting epigenetic regulators is being explored as a potential therapeutic option for several pediatric cancers.
[0004] DNA methylation alterations are widespread and ubiquitous in cancer and are one of the earliest abnormalities in tumorigenesis. Cancer methylomes typically exhibit genome-wide hypomethylation and focal hypermethylation, particularly at promoter sites. The efficacy of targeting DNA demethylase inhibitors has been explored in several cancers, particularly neuroblastoma, Ewing's sarcoma, and AML.
[0005] Therefore, there is a need for new methods for distinguishing between childhood cancers, adult cancers, and their recurrences that are more sensitive and more specific than current methods. The present disclosure meets these needs. Summary of the Invention
[0006] The rarity of pediatric cancers poses significant challenges for the development of cancer-type-specific diagnostic, prognostic, or predictive biomarkers. In this context, it is appealing to explore molecular signatures common to multiple cancer types, an approach that has also shown promise in multi-cancer early detection testing. Specifically, we aimed to identify DNA methylation alterations common to multiple pediatric cancers. Compared to adult-onset cancers, the genomes of pediatric cancers are characteristically "quiet" and have a low mutational load, whereas the epigenomes of many pediatric cancers appear "noisy" and harbor numerous driver mutations in chromatin modifiers, along with extensive DNA methylation and histone marker alterations. DNA methylation alterations are widespread in cancer and are one of the earliest aberrations in tumorigenesis.
[0007] In this study, we aimed to identify DNA methylation alterations common to multiple non-CNS pediatric solid tumors. We performed whole-genome bisulfite sequencing (WGBS) of pediatric cancers, including 31 tumor tissues, 13 normal tissues, and 20 plasma cfDNA samples from 27 individuals representing 11 different pediatric cancer subtypes. By integrating data across tumor types, we identified minimal, localized regions that were differentially methylated in the most samples across cancer types, which we named "minimally differentially methylated regions" (mDMRs). We also identified mDMR methylation differences in 518 pediatric cancer samples from four cancer types obtained from the Therapeutically Applicable Research to Generate Effective Treatments (TARGET) study and 6,426 adult cancer samples from 14 cancer types obtained from The Cancer Genome Atlas (TCGA). We further performed validation using a targeted hybridization probe capture assay in an independent set of 44 pediatric cancer tissue samples from six tumor types. Finally, we found that these methylation changes were detectable in cell-free (cf)DNA and could serve as potential cfDNA methylation biomarkers.
[0008] Identifying such signatures in pediatric cancers may provide the basis for cfDNA methylation liquid biopsies for early detection, treatment response, or minimal residual disease monitoring. Therefore, our research efforts to elucidate such methylation patterns have both scientific and clinical value, with the aim of translating these molecular insights into concrete advances in pediatric cancer treatment.
[0009] Thus, in some embodiments, a method for determining whether a subject has or is likely to develop pediatric cancer, adult cancer, or minimal residual disease (MRD) comprises the following steps: a) training a machine learning model to detect pediatric cancer, adult cancer, or MRD, wherein the machine learning model is trained using target regions from a plurality of cancer samples and corresponding target regions from non-cancer samples, the cancer samples comprising at least two different cancer types, and the machine learning model is configured to identify pediatric cancer, adult cancer, or MRD based on a comparison of the methylation pattern of the target region in the cancer sample compared to the methylation pattern of the corresponding target region in the non-cancer sample; b) determining a methylation pattern of the target region in a deoxyribonucleic acid (DNA) sample obtained from the subject; c) applying the trained machine learning model to the methylation pattern of the target region in the DNA obtained from the subject; and d) determining whether the subject has pediatric cancer, adult cancer, or MRD based on the output of the machine learning model.
[0010] In some embodiments, the methylation pattern of each of the plurality of target regions is determined using DNA methylation analysis, which includes one or more of whole-genome bisulfite sequencing (WGBS), reduced-representation bisulfite sequencing (RRBS), targeted bisulfite sequencing, hybridization probe capture, methylation bead array, and enzymatic methyl-sequence conversion. In some embodiments, the methylation pattern of each of the plurality of target regions is determined using hybridization probe capture after whole-genome bisulfite sequencing. In some embodiments, the hybridization probe capture includes one or more probes that hybridize to one or more target genomic regions, each of the one or more probes comprising a ribonucleic acid or a deoxyribonucleic acid, and optionally, each of the one or more probes comprising an affinity tag selected from the group consisting of biotin and streptavidin.
[0011] In some embodiments, the target region of steps a-c comprises between about 30% and about 50% of the target region of Table 1, between about 50% and about 70% of the target region of Table 1, between about 70% and about 90% of the target region of Table 1, between about 90% and about 95% of the target region of Table 1, or the plurality of target regions comprises greater than about 95% of the target region of Table 1.
[0012] These and other features and advantages of the present invention will be more fully understood from the following detailed description of the invention taken in conjunction with the appended claims, the scope of which is defined by the detailed description, and not by the specific descriptions of features and advantages set forth herein.
[0013] The following drawings form part of the present specification and are included to further demonstrate certain embodiments or various aspects of the present invention. In some cases, embodiments of the present invention may be best understood by reference to the accompanying drawings in combination with the detailed description set forth herein. The specification and accompanying drawings may emphasize particular examples or particular aspects of the invention. However, one skilled in the art will appreciate that some of the examples or aspects may be used in combination with other examples or aspects of the invention. [Brief explanation of the drawings]
[0014] [Figure 1A-E] Analysis of differential DNA methylation patterns in pediatric cancers. (A) Hierarchical clustering of beta values across 183 highly divergent DMRs in 31 tumor samples. Horizontal annotation bars indicate diagnosis, and vertical annotation bars indicate DMR clusters determined by k-means clustering. (B) Uniform manifold approximation and projection (UMAP) for each sample based on the 183 most divergent DMRs, as in C. (C) Heatmap of beta values for 166 / 183 regions in C from the 526 TARGET and 17 POETIC samples used (17 regions from C were omitted due to the 450k array limit). (D) UMAP of TARGET and POETIC samples across the regions in E by diagnosis and (E) origin. [Figure 2A-C] WGBS analysis reveals DNA methylation profiles across pediatric cancers. (A) Number of overlapping mDMRs (y-axis) in a given number of samples (x-axis) by directionality (light blue = hypomethylated regions; red = hypermethylated regions). The dotted red line indicates the selected cutoff value (≥22 samples; 70% of tumor samples) that resulted in 402 hypomethylated and 503 hypermethylated regions. (B) Box plot of mean beta values for hypomethylated and hypermethylated mDMRs. The tumor / normal comparison is statistically significant (p<0.0001; Wilcoxon rank-sum test). (C) Receiver operating characteristic (ROC) curve of a random forest classifier model constructed using mDMRs. The ROC curve was annotated by area under the curve (AUC). Significance annotation: ns: p>0.05, *: p≤0.05, **: p<0.01, ***: p<0.001, ****: p<0.0001. [Figure 3] Differential methylation of mDMRs in the TARGET and ENCODE datasets. Plots show the average beta values across hypermethylated (left) and hypomethylated (right) mDMRs in MRT, NBL, OS, and WT tumor samples from the TARGET database. Normal tissues from POETIC and ENCODE were used as controls. All comparisons were statistically significant (Wilcox p-value ≤ 0.0001) and consistent with the direction observed in the POETIC samples. Significance annotation: ns: p > 0.05, *: p ≤ 0.05, **: p ≤ 0.01, ***: p ≤ 0.001, ****: p ≤ 0.0001. [Figure 4A-C]mDMRs detected in multiple adult cancers from TCGA. (A) ROC curves from TCGA 450K DNA methylation data using 422 of the 905 pediatric cancer mDMRs derived by WGBS trained on a random forest model (a subset of 422 regions was used due to limitations of the 450K array). Plots are annotated by TCGA cancer code and AUC. (B) Graphical representation of the AUC in A. 95% CIs are indicated by error bars. See the GDC website for clarification of study abbreviations (gdc.cancer.gov / resources-tcga-users / tcga-code-tables / tcga-study-abbreviations). (C) Mean tumor and normal beta values across hypermethylated (top) and hypomethylated (bottom) mDMRs across all 14 adult cancer types. All comparisons were statistically significant (Wilcox p-value ≤ 0.0001) and consistent with the direction observed in the POETIC sample, except for the PAAD hypomethylation comparison. Significance annotation: ns: p > 0.05, *: p ≤ 0.05, **: p ≤ 0.01, ***: p ≤ 0.001, ****: p ≤ 0.0001. [Figure 5A-C] Overlapping DMRs between plasma and tissue. (A) Total DMRs called in cell-free (cf) DNA. (B) Stacked histogram showing the number of DMRs per sample or the percentage of DMRs called in cfDNA that overlap with at least one DMR called in patient-matched tumor tissue. Light gray represents hypermethylated regions, and dark gray represents hypomethylated regions. (C) Correlation plot of differential methylation of intersecting DMRs between patient-matched tissue and cfDNA. Each point represents the delta-beta in the intersecting region. The plot is annotated by R2, Spearman correlation coefficient, and best-fit line. Color represents point density (yellow = high, purple = low). [Figure 6A-B] Common features identified in solid tumors serve as biomarkers in plasma. (A / B) Average beta values from cfDNA across 402 hypomethylated and 905 hypermethylated mDMRs identified in tumor samples. All regions across all healthy and diseased plasma samples are shown. [Figure 7A-D]Targeted methylation analysis of hypomethylated mDMRs in an independent cohort. (A) Mean beta values across 402 hypomethylated mDMRs in 44 additional tissue samples from six tumor types obtained from CHLA compared to normal tissue samples. All comparisons are statistically significant (Wilcoxon rank-sum test, significance annotation: ns: p > 0.05, *: p ≤ 0.05, **: p ≤ 0.01, ***: p ≤ 0.001, ****: p ≤ 0.0001). (B) Heatmap of beta values within mDMRs in the POETIC and CHLA cohorts. Annotation bars indicate the diagnosis, tumor / normal status, and origin for each sample. A UMAP generated using the mean beta values across all 402 mDMRs is color-coded by origin (C) and cancer type (D). [Figure 8A-B] Number of DMRs called by WGBS by sample and cancer type. (A) Number of DMRs per sample. Stacked histogram showing the number of DMRs per sample. Light gray represents hypermethylated regions, and dark gray represents hypomethylated regions. The figure is subdivided by tumor type: embryonal rhabdomyosarcoma (ERMS), neuroblastoma (NBL), osteosarcoma (OS), hepatoblastoma (HB), malignant rhabdoid tumor (MRT), and fibrolamellar hepatocellular carcinoma (FHC); all cancer types are grouped as "other." (B) Stacked histogram showing the average number of DMRs per cancer type. n represents the number of samples per tumor type. [Figure 9A-B] Methylation profile of MRT. (A) Number of DMRs called in P01-019 (MRT). Shading represents hypermethylated (light gray) and hypomethylated (dark gray) regions. Each sample is annotated by the proportion of DMR calls that are hypermethylated. (B) Genome-wide beta scores for TARGET samples show hypermethylation of MRT. [Figure 10] UMAP of DMRs identified by k-means clustering in Figure 1c. Each point represents an individual DMR detailed in Figure 1c. The clusters in the vertical annotations in Figure 1c are "K-Means clusters." [Figure 11A-B]mDMRs detected in multiple stage I adult cancers from TCGA. (A) ROC curve from TCGA 450K DNA methylation data using 422 of the 905 pediatric cancer mDMRs derived by WGBS trained with a random forest model (a subset of 422 regions was used due to limitations of the 450K array). Plots are annotated by TCGA cancer code and AUC. (B) Graphical representation of AUC in A. 95% CIs are shown as error bars. Please refer to the GDC website for clarification of study abbreviations (gdc.cancer.gov / resources-tcga-users / tcga-code-tables / tcga-study-abbreviations). [Figure 12A-B] mDMRs detected in multiple stage II adult cancers from TCGA. (A) ROC curve from TCGA 450K DNA methylation data using 422 of the 905 pediatric cancer mDMRs derived by WGBS trained with a random forest model (a subset of 422 regions was used due to limitations of the 450K array). Plots are annotated by TCGA cancer code and AUC. (B) Graphical representation of AUC in A. 95% CIs are shown as error bars. Please refer to the GDC website for clarification of study abbreviations (gdc.cancer.gov / resources-tcga-users / tcga-code-tables / tcga-study-abbreviations). [Figure 13A-B]mDMRs detected in multiple stage III adult cancers from TCGA. (A) ROC curve from TCGA 450K DNA methylation data using 422 of the 905 pediatric cancer mDMRs derived by WGBS trained with a random forest model (a subset of 422 regions was used due to limitations of the 450K array). Plots are annotated by TCGA cancer code and AUC. (B) Graphical representation of AUC in A. 95% CIs are shown as error bars. Please refer to the GDC website for clarification of study abbreviations (gdc.cancer.gov / resources-tcga-users / tcga-code-tables / tcga-study-abbreviations). [Figure 14A-B] mDMRs detected in multiple stage IV adult cancers from TCGA. (A) ROC curve from TCGA 450K DNA methylation data using 422 of the 905 pediatric cancer mDMRs derived by WGBS trained with a random forest model (a subset of 422 regions was used due to limitations of the 450K array). Plots are annotated by TCGA cancer code and AUC. (B) Graphical representation of AUC in A. 95% CIs are shown as error bars. Please refer to the GDC website for clarification of study abbreviations (gdc.cancer.gov / resources-tcga-users / tcga-code-tables / tcga-study-abbreviations). [Figure 15] Curves from mDMRs in CNS tumors (Capper et al.) dataset. ROC curves from 450K DNA methylation data by Capper et al. using 422 of 905 pediatric cancer mDMRs derived by WGBS trained with a random forest model (a subset of 422 regions was used due to limitations of the 450K array). Each plot represents one methylation class from Capper et al. [Figure 16] Cell-free DNA yield from plasma. DNA yield in pg / ul per plasma sample. Vertically subdivided by diagnosis. DETAILED DESCRIPTION OF THE INVENTION
[0015] Although pediatric cancers typically have a lower mutational burden than adult-onset cancers, the epigenome in pediatric cancers is significantly altered and harbors widespread DNA methylation changes. The rarity of pediatric cancers poses significant challenges to the development of cancer-type-specific biomarkers for diagnosis, prognosis, and treatment monitoring. In this study, we explored the possibility of common molecular signatures across diverse pediatric cancers, focusing specifically on DNA methylation changes. To do so, we performed whole-genome bisulfite sequencing (WGBS) on 31 pediatric tumor tissues, 13 normal tissues, and 20 plasma cell-free (cf) DNA samples representing 11 different pediatric cancer types. We identified minimal local regions of differential methylation across samples from multiple cancer types, which we termed minimal differentially methylated regions (mDMRs). These methylation changes were also observed in 518 pediatric cancer samples and 6,426 adult cancer samples accessed from public databases, as well as in 44 pediatric cancer samples analyzed using a targeted hybridization probe capture assay. Finally, we found that these methylation changes were detectable in cfDNA and could serve as potential cfDNA methylation biomarkers.
[0016] Definition. The following definitions are included to provide a clear and consistent understanding of the specification and claims. As used herein, the described terms have the following meanings. Other terms and phrases used herein have their ordinary meanings to those of ordinary skill in the art. Such ordinary meanings can be obtained by consulting technical dictionaries such as R.J. Lewis, "Hawley's Condensed Chemical Dictionary, 14th Edition" (John Wiley & Sons, New York, 2001); Singleton, et al., "Dictionary of Microbiology and Molecular Biology, 2nd ed., John Wiley & Sons, New York (1994); and Hale & Markham, "The Harper Collins Dictionary of Biology," Harper Perennial, NY (1991). Common laboratory techniques (such as DNA extraction, RNA extraction, cloning, cell culture, etc.) are well known in the art and are described, for example, in "Molecular Cloning: A Laboratory Manual," J. Sambrook et al., 4th edition, Cold Spring Harbor Laboratory Press, 2012.
[0017] References in the specification to "one embodiment," "an embodiment," etc. indicate that the described embodiment includes a particular aspect, feature, structure, moiety, or characteristic, but not all embodiments necessarily include that aspect, feature, structure, moiety, or characteristic. Moreover, such phrases may, but do not necessarily, refer to the same embodiment referenced in other parts of the specification. Furthermore, when a particular aspect, feature, structure, moiety, or characteristic is described in connection with an embodiment, it is well known to those skilled in the art that that aspect, feature, structure, moiety, or characteristic also affects or relates to other embodiments, whether or not explicitly described.
[0018] When the term "comprising" is used herein, the alternatives of "consisting of" or "consisting essentially of" are contemplated. As used herein, "comprising" is synonymous with "including," "containing," or "characterized by," and is inclusive or open-ended, and does not exclude additional, unrecited elements or method steps. As used herein, "consisting of" excludes elements, steps, or ingredients not specified in the embodiment. As used herein, "consisting essentially of" does not exclude materials or steps that do not materially affect the basic and novel characteristics of the embodiment. In each instance herein, the terms "comprising," "consisting essentially of," and "consisting of" may be replaced with either of the other two terms. The disclosure illustratively described herein may suitably be practiced in the absence of elements, limitations, or components not specifically disclosed herein.
[0019] The singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. Thus, for example, a reference to "a compound" includes a plural of such compound, so that compound X includes multiple compound X. It is further noted that the claims may be drafted to exclude any element. Accordingly, this statement is intended to serve as antecedent basis for the use of exclusive terminology (e.g., "solely," "only," etc.) in connection with the recitation or use of "negative" limitations of any element described herein and / or claim element.
[0020] The term "and / or" means any one of the items associated with the term, any combination of those items, or all of those items. The phrases "one or more" and "at least one" are readily understood by those of ordinary skill in the art, particularly when read in the context of their use. For example, the phrase can mean 1, 2, 3, 4, 5, 6, 10, 100, or any upper limit that is about 10-fold, 100-fold, or 1000-fold higher than the stated lower limit. For example, one or more substituents on a phenyl ring refers to 1 to 5 substituents on the ring.
[0021] As will be understood by one of ordinary skill in the art, all numerical values, including numerical values (such as amounts of ingredients, properties such as molecular weight, reaction conditions, etc.), are approximate and are understood to be optionally modified in all cases by the term "about." These values may vary depending on the desired properties sought by one of ordinary skill in the art using the teachings herein. It is also understood that these values inherently contain variations necessarily resulting from the standard deviation in their respective testing measurements. When values are expressed as approximations, by use of the antecedent "about," it will be understood that the particular value without the modifier "about" also constitutes a further embodiment.
[0022] The terms "about" and "approximately" are used interchangeably. Both terms can refer to a variation of ±5%, ±10%, ±20%, or ±25% of the specified value. For example, "about 50" percent can, in some embodiments, include a variation of 45 to 55 percent, or other variation defined in a particular claim. For integer ranges, the term "about" can include one or two integers greater than and / or less than the stated integers at both ends of the range. Unless otherwise specified herein, "about" and "approximately" are intended to include values (e.g., weight percent) near the stated range that are equivalent with respect to the function of an individual component, composition, or embodiment. The terms "about" and "approximately" can also modify the endpoints of a stated range, as explained in the paragraph above.
[0023] As will be understood by those skilled in the art, for all purposes, particularly for the purpose of providing a written description, all ranges described herein also encompass all possible subranges and combinations of subranges, as well as the individual values, particularly integers, that make up the range. Thus, each unit between two specified units is also understood to be disclosed. For example, if 10 to 15 is disclosed, 11, 12, 13, and 14 are also disclosed individually and as part of a range. A recited range (e.g., weight percentage or carbon group) includes each specific value, integer, decimal, or identity within the range. Any recited range can be readily recognized as fully describing and enabling the same range to be divided into at least equal halves, thirds, quarters, fifths, or tenths. As a non-limiting example, each range discussed herein can be readily divided into a lower third, middle third, upper third, etc. As will be understood by those skilled in the art, terms such as "up to," "at least," "greater than," "less than," "over," "greater than or equal to," and the like, are inclusive of the recited numerical values, and these terms refer to ranges that can be further divided into subranges as discussed above. Similarly, all ratios recited herein include all subratios that fall within the broader ratio. Accordingly, specific values recited for radicals, substituents, and ranges are merely illustrative, and they do not exclude other defined values or other values within the defined ranges for radicals and substituents. It will be further understood that the endpoints of each range are significant not only in relation to the other endpoint, but independently of the other endpoint.
[0024] The present disclosure provides ranges, limits, and deviations for variables such as volume, mass, percentages, and ratios. Those skilled in the art will understand that a range from "number 1" to "number 2" refers to a continuous range of numbers, including integers and fractions. For example, 1 to 10 refers to 1, 2, 3, 4, 5, ..., 9, 10. It also refers to 1.0, 1.1, 1.2, 1.3, ..., 9.8, 9.9, 10.0, and further refers to 1.01, 1.02, 1.03, .... When a disclosed variable is a number less than "number 10," it refers to a continuous range, including integers and fractions less than "number 10," as discussed above. Similarly, when a disclosed variable is a number greater than "number 10," it refers to a continuous range, including integers and fractions greater than "number 10." These ranges may be qualified by the term "about," the meaning of which is explained above.
[0025] As used herein, the term "substantially" is a broad term and is used in its ordinary sense, including primarily, but not necessarily entirely, the meaning designated. For example, the term can refer to a numerical value that may not be 100% complete. The complete numerical value may be about 1%, about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9%, about 10%, about 15%, or about 20% less.
[0026] As used herein, the term "portion" or "part thereof" refers to consecutive nucleotides of the sequence of the specified region. A portion according to the present invention may comprise or consist of at least 15 or 20 consecutive nucleotides of the specified region, preferably at least 100, 200, 300, 500, or 700 consecutive nucleotides, and more preferably at least 1, 2, 3, 4, or 5 consecutive kb. For example, a portion may comprise or consist of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 consecutive kb of the specified region.
[0027] Those skilled in the art will also readily recognize that when members are grouped together in a common manner (such as in a Markush group), the invention includes not only the entire group recited as a whole, but also each member of the group individually and all possible subgroups of the main group. Moreover, for all purposes, the invention encompasses not only the main group, but also the main group absent one or more group members. The invention therefore contemplates the explicit exclusion of one or more of the described group members. Thus, provisos may be applied to any of the disclosed categories or embodiments, whereby any one or more described elements, species, or embodiments may be excluded from such category or embodiment, for example, by use in an explicit negative limitation.
[0028] The term "contacting" refers to the act of touching, making contact, or bringing into close proximity or proximity, and includes causing, for example, a physiological response, a chemical response, or a physical change at the cellular or molecular level, e.g., in a solution, in a reaction mixture, in vitro, or in vivo.
[0029] An "effective amount" refers to an amount effective to treat a disease, disorder, or condition or to cause a described effect. For example, an effective amount can be an amount effective to reduce the progression or severity of the disease, disorder, or symptom being treated. Determining a therapeutically effective amount is within the capabilities of one skilled in the art. The term "effective amount" is intended to include, for example, an amount of a compound described herein, or an amount of a combination of compounds described herein, that is effective to treat or prevent a disease or disorder, or to treat a symptom of a disease or disorder, in a host. Thus, an "effective amount" generally refers to an amount that provides a desired effect.
[0030] Alternatively, the term "effective amount" or "therapeutically effective amount" as used herein refers to a sufficient quantity of an agent, composition, or combination of compositions administered that will relieve to some extent one or more of the symptoms of the disease or condition being treated. The result can be reduction and / or alleviation of the signs, symptoms, or causes of the disease, or any other desired alteration of a biological system. For example, an "effective amount" for therapeutic use is the amount of a composition containing a compound disclosed herein required to provide a clinically significant reduction in disease symptoms. An appropriate "effective" amount in an individual case can be determined using techniques such as a dose escalation study. Doses can be administered in one or more administrations. However, the precise determination of what constitutes an effective dose can be based on factors specific to each patient, including, but not limited to, the patient's age, size, type or extent of disease, stage of disease, route of administration of the composition, type or extent of concurrent adjunctive therapy, ongoing disease process, and the type of treatment desired (e.g., active versus conventional treatment).
[0031] The terms "treating," "treat," and "treatment" include (i) preventing the occurrence of a disease, pathological, or medical condition (e.g., prophylaxis), (ii) inhibiting or arresting the occurrence of a disease, pathological, or medical condition, (iii) alleviating a disease, pathological, or medical condition; and / or (iv) alleviating symptoms associated with a disease, pathological, or medical condition. Thus, the terms "treat," "treatment," and "treating" can extend to prophylaxis and can include prevention, prophylactic measures, prophylaxis, or hindering the progression or severity of the condition or symptom being treated. Thus, the term "treatment" can include medical, therapeutic, and / or prophylactic administration, as appropriate.
[0032] As used herein, "subject" or "patient" refers to an individual who has symptoms of or is at risk for a disease or other malignancy. A patient may be human or non-human and may include animal breeds or species used as "model systems" for research purposes, such as the mouse model described herein. Similarly, a patient may include an adult or minor (e.g., a child). Furthermore, a patient may refer to any organism, preferably a mammal (e.g., human or non-human), that may benefit from the administration of the compositions discussed herein. Examples of mammals include, but are not limited to, any member of the mammalian class, such as humans, non-human primates such as chimpanzees, and other ape and monkey species; livestock animals such as cows, horses, sheep, goats, and pigs; pets such as rabbits, dogs, and cats; and laboratory animals, including rodents such as rats, mice, and guinea pigs. Examples of non-mammals include, but are not limited to, birds, fish, and the like. In one embodiment of the methods provided herein, the mammal is a human.
[0033] As used herein, the terms "providing," "administering," and "introducing" are used interchangeably herein and refer to the placement of a compound of the present disclosure into a subject by a method or route that results in at least partial localization of the compound to a desired site. The compound may be administered by any suitable route that results in delivery to the desired site in the subject.
[0034] The terms "inhibit," "inhibiting," and "inhibition" refer to slowing, stopping, or reversing the growth or progression of a disease, infection, condition, or group of cells. Inhibition can be, for example, greater than about 20%, 40%, 60%, 80%, 90%, 95%, or 99% compared to growth or progression that occurs in the absence of treatment or contact.
[0035] The term "amplicon" refers to a nucleic acid product resulting from amplification of a target nucleic acid sequence. Amplification is typically performed by PCR. Amplicons range in size from 20 base pairs to 15,000 base pairs for long-range PCR, but are more commonly 100 to 1,000 base pairs for bisulfite-treated DNA used in methylation analysis.
[0036] The term "amplification" refers to an increase in the copy number of a nucleic acid molecule. The resulting amplification product is called an "amplicon." Amplification of nucleic acid molecules (such as DNA or RNA molecules) refers to the use of techniques to increase the copy number of nucleic acid molecules in a sample. An example of amplification is polymerase chain reaction (PCR), in which a sample is contacted with a pair of oligonucleotide primers under conditions that allow hybridization of the primers to a nucleic acid template in the sample. Products of amplification can be characterized by techniques such as electrophoresis, restriction enzyme cleavage patterns, oligonucleotide hybridization or ligation, and / or nucleic acid sequencing. In some embodiments, the methods provided herein can include generating amplified nucleic acids under isothermal or thermovariable conditions.
[0037] The term "biological sample" refers to a sample taken from an individual. As used herein, biological samples include all clinical samples containing genomic DNA (such as cell-free genomic DNA) useful for cancer diagnosis and prognosis, including, but not limited to, cells, tissues, body fluids (blood, blood derivatives and fractions (such as serum and plasma), oral epithelium, saliva, urine, stool, bronchial aspirate, sputum, biopsies (e.g., tumor biopsies), and CVS samples. A "biological sample" obtained from or derived from an individual includes any such sample that has been processed in any suitable manner after being obtained from the individual (e.g., processed to isolate genomic DNA for bisulfite treatment).
[0038] The term "bisulfite treatment" refers to the treatment of DNA with bisulfite or its salts, such as sodium bisulfite (NaHSO). Bisulfite reacts readily with the 5,6-double bond of cytosine but reacts poorly with methylated cytosine. Cytosine reacts with the bisulfite ion to form a sulfonated cytosine reaction intermediate that is susceptible to deamination, generating sulfonated uracil. The sulfonate group is removed under alkaline conditions, resulting in the formation of uracil. Uracil is recognized as thymine by polymerases, and amplification generates adenine-thymine base pairs rather than cytosine-guanine base pairs.
[0039] The term "cancer" refers to a biological state in which a malignant tumor or other neoplasm has undergone characteristic anaplasia characterized by loss of differentiation, increased rate of proliferation, invasion of surrounding tissues, and which is capable of metastasis. A neoplasm is a new and abnormal growth of tissue or cells, especially one in which the growth is uncontrolled and progressive. A tumor is one example of a neoplasm. Non-limiting examples of types of cancer include lung cancer, stomach cancer, colon cancer, breast cancer, uterine cancer, bladder cancer, head and neck cancer, kidney cancer, liver cancer, ovarian cancer, pancreatic cancer, prostate cancer, and rectal cancer.
[0040] The terms "polynucleotide" and "nucleotide" are used interchangeably and refer to at least two or more ribo- or deoxyribonucleotide base pairs (nucleotides) linked via a phosphate ester bond or equivalent. Nucleic acids include polynucleotides and polynucleosides. Nucleic acids include single molecules, double molecules, triplex molecules, circular molecules, or linear molecules. Examples of nucleic acids include, but are not limited to, RNA, DNA, cDNA, genomic nucleic acids, naturally occurring nucleic acids, and non-natural nucleic acids such as synthetic nucleic acids. Short nucleic acids and polynucleotides (e.g., 10-20, 20-30, 30-50, 50-100 nucleotides) are commonly referred to as single- or double-stranded DNA "oligonucleotides" or "probes."
[0041] The term "DNA (deoxyribonucleic acid)" refers to the long-chain polymer that comprises the genetic material of most living organisms. The repeating units in a DNA polymer are four different types of nucleotides, each of which contains one of four bases—adenine, guanine, cytosine, or thymine—linked to a deoxyribose sugar, to which a phosphate group is attached. Triplets of nucleotides (called codons) code for each amino acid in a polypeptide, or a stop signal. The term codon is also used for the corresponding (and complementary) sequence of three nucleotides in mRNA, into which the DNA sequence is transcribed.
[0042] The term "cell-free DNA" refers to DNA that is no longer entirely contained within intact cells, such as the DNA found in plasma or serum.
[0043] The term "target nucleic acid molecule" refers to a nucleic acid molecule intended for detection, amplification, quantification, qualitative detection, or a combination thereof. The nucleic acid molecule does not need to be in a purified form. Various other nucleic acid molecules may be present together with the target nucleic acid molecule. For example, the target nucleic acid molecule may be a specific nucleic acid molecule intended for amplification and / or evaluation of its methylation status. If necessary, purification or isolation of the target nucleic acid molecule may be performed by methods well known to those skilled in the art, such as by using a commercially available purification kit.
[0044] The term "methylation level" refers to the methylation status (methylated or unmethylated) of cytosine nucleotides at one or more CpG sites within a genomic sequence.
[0045] The terms "hypomethylated" or "hypermethylated" refer to the methylation state of a DNA molecule containing multiple CpG sites (e.g., more than 3, 4, 5, 6, 7, 8, 9, 10, etc.), in which a high proportion of the CpG sites (e.g., greater than 80%, 85%, 90%, or 95%, or any other proportion within the range of 50%-100%) are unmethylated or methylated, respectively.
[0046] The term "CpG site" refers to a dinucleotide DNA sequence containing a cytosine followed by a guanine in the 5' to 3' direction. The cytosine nucleotide at a CpG site in genomic DNA is a target of intracellular methyltransferases and can have a methylation state of either methylated or unmethylated. Reference to a "methylated CpG site" or similar term refers to a CpG site in genomic DNA that has a 5-methylcytosine nucleotide.
[0047] As used herein, "sequence identity" or "identity" in the context of two nucleic acid or polypeptide sequences refers to a specified percentage of residues in the two sequences that are the same when aligned for maximum correspondence over a specified comparison window, as measured by a sequence comparison algorithm or by visual inspection. When percentage sequence identity is used in the context of proteins, it is recognized that non-identical residue positions often differ by conservative amino acid substitutions, in which an amino acid residue is replaced with another amino acid residue having similar chemical properties (e.g., charge or hydrophilicity), thus not altering the functional properties of the molecule. When sequences differ by conservative substitutions, the percentage of sequence identity may be adjusted upward to correct for the conservative nature of the substitution. Sequences that differ by such conservative substitutions are said to have "sequence similarity" or "similarity." Methods for making this adjustment are well known to those skilled in the art. Typically, this involves scoring conservative substitutions as partial rather than complete mismatches, thereby increasing the percentage of sequence identity. Thus, for example, where identical amino acids are assigned a score of 1 and non-conservative substitutions are assigned a score of 0, conservative substitutions are assigned a score between 0 and 1. Scoring of conservative substitutions is calculated, for example, as implemented in the program PC / GENE (Intelligenetics, Mountain View, Calif.).
[0048] As used herein, "percent sequence identity" refers to a value determined by comparing two optimally aligned sequences over a comparison window, where the portion of the polynucleotide sequence in the comparison window may contain additions or deletions (i.e., gaps) compared to the reference sequence (which does not contain additions or deletions) due to optimal alignment of the two sequences. The percentage is calculated by determining the number of positions at which the same nucleotide base or amino acid residue occurs in both sequences to obtain the number of matching positions, dividing the number of matching positions by the total number of positions in the comparison window, and multiplying the result by 100 to obtain the percentage of sequence identity.
[0049] The term "substantial identity" in the context of peptides indicates that the peptide has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, or 94%, or even 95%, 96%, 97%, 98%, or 99% sequence identity with the reference sequence over a specified comparison window. In certain embodiments, optimal alignment is performed using the Needleman and Wunsch homology alignment algorithm (Needleman and Wunsch, JMB, 48, 443 (1970)). An indication that two peptide sequences are substantially identical is that one peptide immunologically reacts with an antibody raised against the other peptide. Thus, a peptide is substantially identical to a second peptide if the two peptides differ only by conservative substitutions. Accordingly, embodiments of the present invention also provide nucleic acid molecules and peptides that are substantially identical to the nucleic acid molecules and peptides presented herein.
[0050] For sequence comparison, one sequence usually serves as a reference sequence, to which the test sequence is compared. When using a sequence comparison algorithm, the test sequence and the reference sequence are input into a computer, subsequence coordinates are designated as needed, and sequence algorithm program parameters are designated. The sequence comparison algorithm then calculates the percent sequence identity for the test sequence relative to the reference sequence based on the designated program parameters.
[0051] The term "multiplexing" refers to the use of two or more pairs of primers intended to amplify multiple target gene segments simultaneously in a single tube. In this format, all primers can be contained in a single tube into which the sample is introduced or placed. All desired influenza virus and control gene segments are then amplified by multiple forward and reverse primers within the tube.
[0052] As used herein, the term "complement" refers to a sequence complementary to a nucleic acid according to standard Watson-Crick base pairing rules. A complementary sequence may also be a sequence of RNA complementary to a DNA sequence or its complementary sequence, and may also be cDNA. As used herein, the term "substantially complementary" means that two sequences hybridize under stringent hybridization conditions. Those skilled in the art will understand that substantially complementary sequences do not need to hybridize along their entire length. In particular, a substantially complementary sequence includes a contiguous sequence of bases that does not hybridize to a target or marker sequence, located 3' or 5' to a contiguous sequence of bases that hybridize to a target or marker sequence under stringent hybridization conditions.
[0053] "Hybridization" refers to a reaction in which one or more polynucleotides react to form a complex stabilized by hydrogen bonds between the bases of the nucleotide residues. Hydrogen bonding can occur through Watson-Crick base pairing, Hoogstein binding, or other sequence-specific manners. The complex may contain two strands forming a duplex structure, three or more strands forming a multistranded complex, a single self-hybridizing strand, or any combination of these. A hybridization reaction may constitute a step in a more extensive process, such as the initiation of a PCR reaction or the enzymatic cleavage of a polynucleotide by a ribozyme.
[0054] Examples of stringent hybridization conditions include an incubation temperature of about 25°C to about 37°C, a hybridization buffer concentration of about 6xSSC to about 10xSSC, a formamide concentration of about 0% to about 25%, and a wash solution of about 4xSSC to about 8xSSC. Examples of moderate hybridization conditions include an incubation temperature of about 40°C to about 50°C, a buffer concentration of about 9xSSC to about 2xSSC, a formamide concentration of about 30% to about 50%, and a wash solution of about 5xSSC to about 2xSSC. Examples of highly stringent hybridization conditions include an incubation temperature of about 55°C to about 68°C, a buffer concentration of about 1xSSC to about 0.1xSSC, a formamide concentration of about 55% to about 75%, and a wash solution of about 1xSSC, 0.1xSSC, or deionized water. Generally, hybridization incubation times range from 5 minutes to 24 hours, including one, two, or more wash steps, with wash incubation times of approximately 1, 2, or 15 minutes. SSC is a 0.15 M NaCl, 15 mM citrate buffer. It is understood that equivalents of SSC using other buffer systems can be used.
[0055] As used herein, the term "reference genome" refers to any particular known, partial, or complete, sequenced, or characterized genome that can be used to reference sequences identified from a subject. Exemplary reference genomes used for human subjects and many other organisms are provided in online genome browsers operated by the National Center for Biotechnology Information (NCBI) or the University of California, Santa Cruz (UCSC). "Genome" refers to the complete genetic information of an organism or virus, represented by nucleic acid sequence. As used herein, a reference sequence or reference genome is often a genome sequence assembled or partially assembled from an individual or multiple individuals. In some embodiments, a reference genome is a genome sequence assembled or partially assembled from one or more human individuals. A reference genome can be considered a representative set of genes for a species. In some embodiments, a reference genome includes sequences assigned to chromosomes. An example of a human reference genome is GRCh37 (UCSC equivalent: hg19).
[0056] As used herein, the term "normal reference standard" refers to the control level, degree, or range of DNA methylation at a particular genomic region or gene in a sample not associated with cancer. The term "normal reference cutoff value" refers to the control threshold level of DNA methylation at a particular genomic region or gene, or differential methylation value (DMV). In some embodiments, enriched DNA methylation levels above the normal reference cutoff value are associated with having or developing cancer. In some embodiments, DNA methylation levels at or below the normal reference cutoff value are associated with not having or developing cancer.
[0057] As used herein, "detecting" refers to determining the presence and / or degree of methylation in a nucleic acid of interest in a sample. Detection does not require a method that provides 100% sensitivity and / or 100% specificity.
[0058] "RT-PCR" refers to reverse transcription polymerase chain reaction, which is used to detect specific RNA (in this case, a specific gene segment of the influenza virus genome) by transcribing the RNA of interest into its DNA complement using reverse transcriptase. The newly synthesized cDNA can be amplified using conventional PCR. In one embodiment, the RT-PCR provided herein is a one-step approach, in which the entire reaction, from cDNA synthesis to PCR amplification, occurs in a single tube. Alternatively, the method described herein is compatible with two-step reactions, in which the reverse transcriptase reaction and PCR amplification are performed in separate tubes. See Real-Time PCR: Current Technology and Applications, Logan, Edwards, and Saunders eds., Caister Academic Press, 2009; Bustin AZ of Quantitative PCR (IUL Biotechnology, No. 5).
[0059] As used herein, a "fragment" of DNA refers to a fragment of DNA having a length of about 10 bp, about 20 bp, about 30 bp, about 40 bp, about 50 bp, about 60 bp, about 70 bp, about 80 bp, about 90 bp, about 100 bp, about 110 bp, about 120 bp, about 130 bp, about 140 bp, about 150 bp, about 160 bp, 170 bp, about 180 bp, about 190 bp, about 200 bp, This refers to cell-free DNA fragments that are about 210 bp, about 220 bp, about 230 bp, about 240 bp, about 250 bp, about 260 bp, about 270 bp, 280 bp, about 290 bp, about 300 bp, about 310 bp, about 320 bp, about 330 bp, about 340 bp, about 350 bp, about 360 bp, about 370 bp, about 380 bp, about 390 bp, or about 400 bp. Typically, DNA fragments are about 100 bp to about 200 bp, about 120 bp to about 180 bp, or about 140 bp to about 160 bp.
[0060] The term "neoadjuvant therapy" refers to treatment (such as chemotherapy or hormone therapy) given before primary cancer treatment (such as surgery) to improve the outcome of the primary treatment.
[0061] The term "chemotherapy" refers to the treatment of cancer using anti-tumor or chemotherapeutic agents as part of a standard regimen. Chemotherapy may be administered with curative intent, or it may be administered to prolong survival or alleviate symptoms. It may also be administered in combination with other cancer treatments, such as radiation therapy or surgery.
[0062] The term "methylation" refers to the addition of a methyl group to the 5' carbon of cytosine bases in CpG deoxyribonucleic acid sequences within the genome.
[0063] The term "adjacent CpG sites" refers to a collection of CpG sites within a genomic feature or spanning a short genetic distance. The genomic feature may be a promoter, enhancer, exon, intron, 5'-untranslated region (UTR), 3'-UTR, gene body, stem cell-associated region, CpG island, CpG shelf, CpG shore, LINE, SINE, or LTR. Short genetic distances are 10bp, 11bp, 12bp, 13bp, 14bp, 15bp, 16bp, 17bp, 18bp, 19bp, 20bp, 21bp, 22bp, 23bp, 24bp, 25bp, 26bp, 27bp, 28bp, 29bp, 30bp, 31bp, 32bp, 33bp, 34bp, 35bp, 36bp, 37bp, 38bp, 39bp, 40bp, 41bp, 42bp, 43bp, 44bp, 45bp, 46bp, 47bp, 48bp, 49bp, 50bp, 51bp, 52bp, 53bp, 54bp, 55bp, 56bp, 57bp, The length may be 58bp, 59bp, 60bp, 61bp, 62bp, 63bp, 64bp, 65bp, 66bp, 67bp, 68bp, 69bp, 70bp, 71bp, 72bp, 73bp, 74bp, 75bp, 76bp, 77bp, 78bp, 79bp, 80bp, 81bp, 82bp, 83bp, 84bp, 85bp, 86bp, 87bp, 88bp, 89bp, 90bp, 91bp, 92bp, 93bp, 94bp, 95bp, 96bp, 97bp, 98bp, 99bp, 100bp, 250bp, 500bp, 750bp, or 1,000bp. In some cases, adjacent CpG sites may occur within the sequencing read.
[0064] The term "minimal residual disease" or "MRD" refers to cancer cells (e.g., breast cancer cells) that remain after treatment and cannot be detected by scans or tests to identify a state of remission (i.e., cancer-free). Treatment of any of the cancers listed herein can result in MRD.
[0065] Embodiments of the present invention The present disclosure provides assays and various methods for detecting differential methylation patterns of target regions of DNA. Differences in the differential methylation patterns of target regions in samples (e.g., cfDNA, genomic DNA isolated from tissue) can indicate, for example, the presence or absence of childhood cancer, the presence or absence of adult cancer, and / or the recurrence of childhood or adult cancer (such as minimal residual disease). The methylation patterns of target regions of DNA in samples can be analyzed using machine learning models that are trained using DNA sets of defined target regions of cancer and non-cancer control samples.
[0066] Generally, a test sample from a subject may be processed to determine the presence or absence of a pediatric cancer, adult cancer, or MRD according to the following steps: i) extracting nucleic acid (e.g., DNA) from the test sample from a subject suspected of having a pediatric cancer, adult cancer, or MRD; ii) converting unmethylated cytosines in the extracted nucleic acid to uracil (e.g., by bisulfite conversion); iii) preparing a library of bisulfite-converted DNA; iv) enriching for target nucleic acids by hybridizing the bisulfite-converted DNA with a hybridization probe; v) sequencing the enriched nucleic acids containing the target region. vi) aligning sequence reads of the target regions with corresponding target regions in a reference genome (e.g., using Bismark); vii) performing differential methylation region analysis (e.g., using Metilene) of the test sample compared to a normal sample or pool of normal samples to identify regions differentially methylated between the test (cancer) and normal samples; viii) defining minimally differentially methylated regions (mDMRs) across the samples; ix) using machine learning tools to build a classifier model using the mDMRs, wherein an AUC value of 0.8 or greater indicates the presence of pediatric cancer, adult cancer, or MRD.
[0067] In one embodiment, a method for determining whether a subject has or is likely to develop pediatric cancer, adult cancer, and / or minimal residual disease (MRD) comprises the following steps: training a machine learning model to detect pediatric cancer, adult cancer, or MRD, wherein the machine learning model is trained using target regions from a plurality of cancer samples and corresponding target regions from non-cancer samples, the cancer samples comprising at least two different cancer types, and the machine learning model is configured to identify pediatric cancer, adult cancer, or MRD based on a comparison of the methylation patterns of the target regions in the cancer samples compared to the methylation patterns of the corresponding target regions in the non-cancer samples; determining a methylation pattern of the target regions in a deoxyribonucleic acid (DNA) sample obtained from the subject; applying the trained machine learning model to the methylation pattern of the target regions in the DNA obtained from the subject; and determining whether the subject has pediatric cancer, adult cancer, or MRD based on the output of the machine learning model.
[0068] Embodiments of the present disclosure may include, for example, bisulfite converting nucleic acids from a subject's DNA sample using whole genome bisulfite sequencing (WGBS) or hybrid probe capture; next-generation sequencing the converted and / or enriched nucleic acids; collecting methylation data from targeted regions (e.g., target regions listed in Table 1); and using a trained machine learning model to determine, for example, the presence or absence of childhood cancer, recurrence of childhood cancer, and / or the presence or absence of adult cancer or recurrence of adult cancer.
[0069] In some embodiments, a method for determining the methylation pattern of one or more target nucleic acids includes methylation sequencing. For example, the methylation patterns of CpG sites within the target regions listed in Table 1 can be detected using DNA methylation sequencing. DNA methylation sequencing can involve, for example, treating DNA from a sample with bisulfite to convert unmethylated cytosines to uracil, followed by amplification (e.g., PCR amplification) of the target nucleic acids within the treated genomic DNA, and sequencing the resulting amplicons. Sequencing generates nucleotide reads, which can be aligned to a genomic reference sequence and used to quantify the methylation levels of all CpGs within the amplicon. Cytosines in non-CpG contexts can be used to track bisulfite conversion efficiency for each sample. This procedure is efficient in both time and cost, and multiple samples can be sequenced in parallel using 96-well plates, generating reproducible measurements of methylation when assayed in independent experiments.
[0070] Nucleic acid molecules can also be subjected to conditions sufficient to convert unmethylated cytosines in the nucleic acid molecule to uracil (e.g., after extraction from a sample). For example, to detect DNA methylation, certain embodiments provide for first converting the DNA to be analyzed, thereby converting unmethylated cytosines to uracil. In one embodiment, a chemical reagent can be used that selectively modifies either the methylated or unmethylated form of a CpG dinucleotide motif. Suitable chemical reagents include hydrazine, bisulfite ions, and the like. Preferably, the separated DNA is treated with sodium bisulfite (NaHSO), which converts unmethylated cytosines to uracil while maintaining methylated cytosines. Without wishing to be bound by theory, it is understood that sodium bisulfite readily reacts with the 5,6-double bond of cytosine but hardly reacts with methylated cytosine. Cytosine reacts with bisulfite ions to form a sulfonated cytosine reaction intermediate that is susceptible to deamination, generating sulfonated uracil. The sulfonated group can be removed under alkaline conditions, resulting in the formation of uracil. The conversion of the nucleotide results in a change in the sequence of the original DNA. It is well known that the resulting uracil has the base-pairing behavior of thymine, which differs from that of cytosine. For this reason, uracil is recognized as thymine by DNA polymerase. Therefore, after PCR or sequencing, the resulting product contains cytosine only at positions where 5-methylcytosine occurs in the starting template DNA. This distinguishes between unmethylated and methylated cytosine.
[0071] Nucleic acid molecules may also be subjected to further processing, including other derivatization processes (e.g., to incorporate, modify, and / or delete one or more sequences, tags, or labels). In some cases, functional sequences (e.g., sequencing adapters, flow cell adapters, sequencing primers, etc.) may be added to nucleic acid molecules to facilitate nucleic acid sequencing. Thus, derivatives of nucleic acid molecules from a sample may include processed nucleic acid molecules, including bisulfite-modified nucleic acid molecules, reverse-transcribed nucleic acid molecules, tagged nucleic acid molecules, barcoded nucleic acid molecules, and other modified nucleic acid molecules.
[0072] In some embodiments, the methylation pattern of the target region may be determined using one or more of hybrid probe capture, targeted bisulfite amplicon sequencing, bisulfite DNA treatment, WGBS, combined bisulfite conversion and bisulfite restriction analysis (COBRA), bisulfite PCR, bisulfite modification, bisulfite pyrosequencing, methylated CpG island amplification, CpG-binding column-based CpG island isolation, CpG island array using differential methylation hybridization, high-performance liquid chromatography, DNA methyltransferase assay, methylation-sensitive PCR, cloning of differentially methylated sequences, post-restriction methylation detection, restriction landmark genomic scanning, methylation-sensitive restriction fingerprinting, or Southern blot analysis.
[0073] In one embodiment, the method used to determine the methylation level of one or more target regions in DNA is WGBS (Cokus, et al. 2008. Nature 452(7184):215-219; Lister, et al. 2009. Nature 462(7271):315-322; Harris, et al. 2010. Nat Biotechnol 28(10):1097-1105).
[0074] Other methods for assaying the methylation status of CpG sites can also be used. Numerous DNA methylation detection methods are known in the art, including, but not limited to, hybrid probe capture (REF), methylation-specific enzyme digestion (Singer-Sam et al., Nucleic Acids Res. 18(3):687, 1990; Taylor et al., Leukemia 15(4):583-9, 2001), methylation-specific PCR (MSP or MSPCR) (Herman et al., Proc Natl Acad Sci USA 93(18):9821-6, 1996), methylation-sensitive single-nucleotide primer extension (MS-SnuPE) (Gonzalgo et al., Nucleic Acids Res. 25(12):2529-31, 1997), restriction enzyme landmark genomic scanning (RLGS) (Kawai, Mol Cell Biol. 14(11):7421-7, 1994; Akama et al., Cancer Res. 57(15):3294-9, 1997), and differential methylation hybridization (DMH) (Huang et al., Hum Mol Genet. 8(3):459-70, 1999). In some embodiments, methylation levels can be determined using one or more DNA methylation sequencing assays, with or without bisulfite treatment of the DNA.
[0075] In one embodiment, reduced representation bisulfite sequencing (RRBS) is used to measure the methylation level of a target region. Generally, RRBS begins by treating nucleic acids with bisulfite to convert all unmethylated cytosines to uracil, followed by restriction enzyme digestion (e.g., with an enzyme that recognizes sites containing CG sequences, such as MspI), and sequencing of the complete fragments after coupling with an adaptor ligand. The selection of the restriction enzyme enriches fragments in CpG-dense regions, reducing the number of redundant sequences that may map to multiple gene locations during analysis. Therefore, RRBS reduces the sample complexity of a nucleic acid sample by selecting a subset of restriction fragments for sequencing (e.g., by size selection using preparative gel electrophoresis). In contrast to bisulfite whole genome sequencing, each fragment generated by restriction enzyme digestion contains DNA methylation information for at least one CpG dinucleotide. Thus, RRBS enriches samples in promoters, CpG islands, and other genomic features with high frequency of restriction enzyme cleavage sites in these regions, thus providing an assay for assessing the methylation status of one or more genomic loci.
[0076] A typical protocol for RRBS involves digesting sample nucleic acids with a restriction enzyme such as Mspl, projecting and filling with A-tails, ligating adapters, bisulfite conversion, and PCR (see, for example, Gu et al., (2010), Nat Methods 7:133-6; Meissner et al., (2005), Nucleic Acids Res. 33:5868-77).
[0077] In some embodiments, identifying the presence of childhood cancer, adult cancer, or MRD in a subject can include using hybrid capture probes configured to selectively enrich for nucleic acid molecules (e.g., DNA or RNA molecules) or their sequences. Such probes can be pull-down probes (e.g., bait sets). The selectively enriched nucleic acid molecules or their sequences can correspond to one or more target regions in the methylation profile of the dataset. The presence of specific sequences, modifications (e.g., methylation status), deletions, additions, single nucleotide polymorphisms, copy number variations, or other features in the selectively enriched nucleic acid molecules or their sequences can be indicative of the presence and / or recurrence of childhood cancer. The probes can be selective for a subset of specific target regions in Table 1 and / or for differentially methylated regions (e.g., CpG sites, CpA sites, CpT sites, and / or CpC sites) in the DNA sample. The probes can be configured to selectively enrich nucleic acid molecules (e.g., DNA or RNA molecules) or sequences thereof corresponding to multiple target nucleic acids of a target genomic sequence, such as a subset of one or more genomic regions and / or a subset of differentially methylated regions (e.g., CpG sites, CpA sites, CpT sites, and / or CpC sites) in a cell-free biological sample. The probes can be nucleic acid molecules (e.g., DNA or RNA molecules) having sequence complementarity with the target nucleic acid sequences. These nucleic acid molecules can be primers or enrichment sequences. Assaying nucleic acid molecules of a sample (e.g., a cell-free biological sample) using probes selected for target nucleic acid sequences can include the use of array hybridization, polymerase chain reaction (PCR), or nucleic acid sequencing (e.g., DNA sequencing or RNA sequencing).The number of target nucleic acid sequences selectively enriched using such a scheme can include at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 50, at least 100, at least 150, at least 200, at least 300, at least 500, or more than 500 target nucleic acid sequences in a target genomic region. The use of such probes for enrichment of target nucleic acids can be referred to as "hybrid capture." The use of such hybrid capture probes can occur before or after bisulfite conversion (if applicable). Exemplary target nucleic acid sequences include those associated with the target regions contained in Table 1. In some embodiments, the target region includes all of the sequences in Table 1. In some embodiments, the hybridization probes comprise an assay panel and may be complementary to one or more target regions of Table 1 and may include at least 50, 60, 70, 80, 90, 100, 120, 150, or 200 pairs of probes. In some embodiments, the hybridization probes comprise an assay panel and may include at least 1,000, 2,000, 2,500, 5,000, 6,000, 7,500, 10,000, 15,000, 20,000, 25,000, or 50,000 different probe pairs. In some embodiments, the assay panel may include at least 100, 120, 140, 160, 180, 200, 240, 300, or 400 different probes. In other embodiments, the assay panel may include at least 1,000, 2,000, 5,000, 10,000, 12,000, 15,000, 20,000, 30,000, 40,000, 50,000, or 100,000 different probes. Preferably, the number of probes is sufficient to overlap substantially all of the target regions of interest.
[0078] In some embodiments, one or more probes comprise deoxyribonucleic acid and / or ribonucleic acid, hi some embodiments, one or more probes comprise an affinity tag selected from the group consisting of biotin and streptavidin.
[0079] Thus, in some embodiments, methylation sequencing of multiple target regions uses one or more of whole genome sequencing, including one or more of whole genome bisulfite sequencing (WGBS), reduced representation bisulfite sequencing (RRBS), targeted bisulfite sequencing, hybridization probe capture, methylation bead array, and enzymatic methyl sequence conversion.
[0080] Nucleic acid molecules (e.g., extracted cfDNA) or their derivatives can be subjected to sequencing to provide multiple sequencing reads. The sequencing reads can be aligned with and / or analyzed relative to a reference genome. Based at least in part on the sequencing reads, the absolute or relative amounts of nucleic acid molecules corresponding to one or more genomic regions (including the absolute or relative levels of methylation within the molecules) can be measured. Alternatively, the sequencing reads can be used to determine the amount or relative amount of nucleic acid molecules. A dataset including a genomic profile (e.g., a methylation profile) of one or more genomic regions of the sample can be generated at least in part on the sequencing reads. The sequencing reads can be processed to identify the methylation pattern of target regions of DNA in the sample.
[0081] Sequence identification can be carried out, for example, by sequencing, array hybridization (for example: Affymetrix), or nucleic acid amplification (for example: PCR). Sequencing can be carried out by any suitable sequencing method, such as massively parallel sequencing (MPS), paired-end sequencing, high-throughput sequencing, next-generation sequencing (NGS), shotgun sequencing, single molecule sequencing, nanopore sequencing, nanopore sequencing (including direct detection or inference of methylation status), semiconductor sequencing, pyrosequencing, sequencing by synthesis (SBS), sequencing by ligation, sequencing by hybridization, and RNA-Seq (Illumina).
[0082] Sequencing and / or preparing a nucleic acid sample for sequencing may involve performing one or more nucleic acid reactions, such as one or more nucleic acid amplification processes (e.g., of DNA or RNA molecules). Nucleic acid amplification may include, for example, reverse transcription, primer extension, asymmetric amplification, rolling circle amplification, ligase chain reaction, polymerase chain reaction (PCR), and multiple displacement amplification. Examples of PCR methods include digital PCR (dPCR), emulsion PCR (ePCR), quantitative PCR (qPCR), real-time PCR (RT-PCR), hot-start PCR, multiplex PCR, asymmetric PCR, nested PCR, and assembly PCR. An appropriate number of rounds of nucleic acid amplification (e.g., PCR such as qPCR, RT-PCR, dPCR, etc.) may be performed to sufficiently amplify the initial amount of nucleic acid (e.g., DNA molecules) or its derivatives to an input amount appropriate for subsequent sequencing analysis. In some cases, PCR may be used for global amplification of nucleic acid molecules. This may involve first using adapter sequences ligated to different molecules, followed by PCR amplification using universal primers. PCR can be performed using any of a number of commercially available kits provided by Life Technologies, Affymetrix, Promega, Qiagen, etc. In other cases, only certain target nucleic acids within a population of nucleic acids may be amplified. Specific primers (optionally in combination with adapter ligation) may be used to selectively amplify specific targets for downstream sequencing. In some cases, nested primers may be used to target specific genomic regions. Nucleic acid amplification may include targeted amplification of one or more gene loci, genomic regions, cfDNA target regions, or differentially methylated regions (e.g., CpG sites, CpA sites, CpT sites, and / or CpC sites), and particularly the target regions listed in Table 1. In some cases, nucleic acid amplification is performed after bisulfite conversion. Such a procedure may be referred to as targeted bisulfite amplicon sequencing (TBAS).Nucleic acid amplification can involve the use of one or more of primers, probes, enzymes (e.g., polymerases), buffers, and deoxyribonucleotides. Nucleic acid amplification can be isothermal or involve thermal cycling. Thermal cycling can involve varying the temperature associated with various processes of nucleic acid amplification, including, for example, initialization, denaturation, annealing, and extension. Sequencing can involve the simultaneous use of reverse transcription (RT) and PCR, such as the OneStep RT-PCR kit protocol from Qiagen, NEB, Thermo Fisher Scientific, or Bio-Rad.
[0083] Nucleic acid molecules (e.g., DNA or RNA molecules) or their derivatives can be labeled or tagged with distinguishable tags to allow multiple samples to be multiplexed. For example, all nucleic acid molecules or their derivatives associated with a particular sample or subject can be tagged or labeled (e.g., with a barcode such as a nucleic acid barcode sequence or a fluorescent label). Nucleic acid molecules or their derivatives associated with other samples or subjects can be tagged or labeled with different tags or labels, thereby allowing the nucleic acid molecules or their derivatives to be associated with the samples or subjects from which they originate. Such tagging or labeling also facilitates multiplexing, allowing nucleic acid molecules or their derivatives from multiple samples and / or subjects to be analyzed (e.g., sequenced) simultaneously. Any number of samples can be multiplexed. For example, a multiplexed reaction can include at least about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, or more than 100 nucleic acid molecules or their derivatives from initial samples. These samples can be from the same or different subjects. For example, multiple samples can be tagged with sample barcodes (e.g., nucleic acid barcode sequences), so that each nucleic acid molecule (e.g., DNA molecule) or its derivative can be traced back to the sample (and / or subject) from which it originated. Sample barcodes allow samples from multiple subjects to be distinguished from one another, which allows sequences in such samples, such as in a pool, to be identified simultaneously. Tags, labels, and / or barcodes can be attached to nucleic acid molecules or derivatives thereof by ligation, primer extension, nucleic acid amplification, or other methods.In some cases, nucleic acid molecules or their derivatives of a particular sample are tagged, labeled, or barcoded with different tags, labels, or barcodes (e.g., unique molecular identifiers), such that different nucleic acid molecules or their derivatives from the same sample can be differently tagged, labeled, or barcoded. In some cases, nucleic acid molecules or their derivatives from a given sample can be labeled with both different and identical labels, such that each nucleic acid molecule or their derivative associated with the sample contains both a unique label and a common label.
[0084] In some embodiments, sequencing reads of target regions are aligned (e.g., using Bismark software) to the corresponding target regions of a reference genome (e.g., GRCh37). The methylation levels of the target regions may be compared to the methylation levels of the target regions in a pool of unusual or normal samples to identify hypomethylated and hypermethylated target regions (i.e., differentially methylated regions). This process may be facilitated by the use of a methylation calling program such as Metilene. Each of the hypomethylated and hypermethylated regions may be individually tested for the presence or absence of cancer using a machine learning model. [Table 1] TIFF2026506978000003.tif236159TIFF2026506978000004.tif237159TIFF2026506978000005.tif236159TIFF2026506978000006.tif238159TIFF2026506978000007.tif236159TIFF2026506978000008.tif240159TIFF2026506978000009.tif238159TIFF2026506978000010.tif237159TIFF2026506978000011.tif236159TIFF2026506978000012.tif236159TIFF2026506978000013.tif238159TIFF2026506978000014.tif236159TIFF2026506978000015.tif239159TIFF2026506978000016.tif236159TIFF2026506978000017.tif238159TIFF2026506978000018.tif241159TIFF2026506978000019.tif240159TIFF2026506978000020.tif237159TIFF2026506978000021.tif238159TIFF2026506978000022.tif240159TIFF2026506978000023.tif240159TIFF2026506978000024.tif241159TIFF2026506978000025.tif241159TIFF2026506978000026.tif240159TIFF2026506978000027.tif241159TIFF2026506978000028.tif240159TIFF2026506978000029.tif239159TIFF2026506978000030.tif56159
[0085] After nucleic acid molecules or their derivatives are subjected to sequencing, appropriate bioinformatics processing can be performed on sequence reads to generate a dataset comprising the methylation pattern of one or more target regions of DNA sample.For example, sequence reads can be aligned to one or more reference genomes (e.g., human genome).The aligned sequence reads can be quantified at one or more genome loci or target regions to generate a dataset comprising the methylation pattern profile of one or more target regions of acellular biological sample.The quantification of sequence can be expressed as unnormalized value or normalized value.
[0086] In some embodiments, alignment of bisulfite converted DNA is performed using a software program such as Bismark (Krueger et al. (2011) Bioinformatics, 27(11):157171). Bismark performs both read mapping and methylation calling in a single step, and its output identifies cytosines in CpG, CHG, and CHH contexts. Bismark is released under the GNU GPLv3+ license. The source code is freely available at bioinformatics.bbsrc.ac.uk / projects / bismark / . In some embodiments, methylation differentials are calculated for specific loci / regions using, for example, one or more publicly available programs for analyzing and / or determining methylation levels or target polynucleotide regions. In some embodiments, methods used to analyze and / or determine the methylation level of a target polynucleotide region include Metilene (Juhling et al., Genome Res., 2016;26(2):256-262) or GenomeStudio Software available online from Illumina, Inc. Other methods for determining differential methylation of a target polynucleotide region are described in Hovestadt et al., 2014; Nature, 510(7506), 537-541.
[0087] In some embodiments, the target regions examined to determine the presence or absence of childhood cancer in a subject comprise at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% of the target regions listed in Table 1.
[0088] In some embodiments, the target regions examined to determine the severity of a subject with pediatric cancer (i.e., stage I, stage II, stage III, or stage IV cancer) comprise at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% of the target regions listed in Table 1.
[0089] In some embodiments, the target genomic regions tested to determine the presence or absence of adult cancer in a subject comprise at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% of the target regions listed in Table 1.
[0090] In some embodiments, the target regions examined to determine the severity of a subject with adult cancer (i.e., stage I, stage II, stage III, or stage IV cancer) comprise at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% of the target regions listed in Table 1.
[0091] The target genomic regions tested to determine the presence of MRD in a subject include at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% of the target regions listed in Table 1.
[0092] In some embodiments, the biological sample can be collected by standard biopsy or liquid biopsy. Preferably, the biopsy is a liquid biopsy, and cfDNA can be collected from whole blood, plasma, serum, or urine. In some embodiments, the volume of the sample, such as whole blood, is about 50 μL to about 5 mL, about 100 μL to about 5 mL, about 150 μL to about 5 mL, about 200 μL to about 5 mL, about 250 μL to about 5 mL, about 300 μL to about 5 mL, about 350 μL to about 5 mL, about 400 μL to about 5 mL, about 450 μL to about 5 mL, about 500 μL to about 5 mL, or about 550 μL to about 5 mL. mL, about 600 μL to about 5 mL, about 700 μL to about 5 mL, about 750 μL to about 5 mL, about 800 μL to about 5 mL, about 850 μL to about 5 mL, about 900 μL to about 5 mL, about 950 μL to about 5 mL, about 1 mL to about 5 mL, about 1.5 mL to about 5 mL, about 2 mL to about 5 mL, about 2.5 mL to about 5 mL, or about 3 mL to about 5 mL. In another embodiment, the volume of the sample, such as whole blood, may comprise about 5 mL to about 10 mL.
[0093] The separation and extraction of DNA, particularly cfDNA, can be carried out by collecting bodily fluid using various techniques.In some cases, collecting can involve using syringe to draw bodily fluid from subject.In other cases, collecting can involve pipetting or directly collecting bodily fluid into collecting container.The method for separating DNA or other nucleic acids from tissue sample is well known in the art.
[0094] After collecting the body fluid, cfDNA can be separated and extracted using various techniques known to those skilled in the art. In some cases, cell-free nucleic acids can be separated, extracted, and prepared using a commercially available kit, such as the Qiagen Qiamp® Circulating Nucleic Acid Kit protocol. Other examples include the Qiagen Qubit™ dsDNA HS Assay Kit protocol, the Agilent™ DNA 1000 Kit, or the TruSeq™ Sequencing Library Preparation; Low-Throughput (LT) protocol.
[0095] Alternatively, cfDNA can be extracted and separated from body fluid through a partitioning step, which separates the cfDNA present in the solution of body fluid from cells and other insoluble components of body fluid.Partitioning includes, but is not limited to, techniques such as centrifugation or filtration.In other cases, cells can be dissolved instead of being first partitioned from cfDNA.For example, the genomic DNA of intact cells can be partitioned by selective precipitation.
[0096] In some embodiments, the pediatric or adult cancer is osteosarcoma, medulloblastoma, ependymoma, optic glioma, brainstem glioma, oligostoma, ganglioglioma, pineal tumor, hepatoblastoma, fibrolamellar hepatocellular carcinoma, Hodgkin's lymphoma, non-Hodgkin's lymphoma, acute lymphocytic leukemia, acute myeloid leukemia, juvenile myelomonocytic leukemia, acute promyelocytic leukemia, chronic lymphocytic leukemia, chronic myeloid leukemia, diffuse pontine glioma, neuroblastoma, retinoblastoma, rhabdoid tumor, Ewing's sarcoma, rhabdoid tumor, embryonal rhabdomyosarcoma, fibrosarcoma, mesenchymoma, synovial sarcoma, teratoma, liposarcoma, spinal tumor, ovarian cancer, or Wilms' tumor. In some embodiments, the cancer is recurrent cancer (e.g., recurrent ovarian cancer, recurrent teratoma, recurrent Wilms' tumor, etc.).
[0097] In some embodiments, the pediatric or adult cancer is selected from the group consisting of anaplastic pilocytic astrocytoma, atypical teratoid / rhabdoid tumor, subclass MYC; atypical teratoid / rhabdoid tumor, subclass SHH; atypical teratoid / rhabdoid tumor, subclass TYR; cerebellar liponeurocytoma; CNS Ewing's sarcoma family of tumors with CIC alterations; CNS high-grade neuroepithelial tumor with BCOR alterations, CNS high-grade neuroepithelial tumor with MN1 alterations, CNS neuroblastoma with FOXR2 activation; diffuse meningeal glioneuronal tumor; diffuse midline glioma H3K27M mutant, embryonal tumor with multilayered rosettes; posterior fossa ependymoma group A; ependymoma, RELA fusions; esthesioneuroblastoma, subclass A; glioblastoma, IDH wild-type, H3.3 G34 mutant; glioblastoma, IDH wild-type, subclass mesenchymal; glioblastoma, IDH wild-type, subclass median; glioblastoma, IDH wild-type, subclass MYCN; glioblastoma, IDH wild-type, subclass RTK; glioblastoma, IDH wild-type, subclass RTKII; glioblastoma, IDH wild-type, subclass RTKIII; IDH glioma, subclass 1p / 19q codeleted oligodendroglioma; IDH glioma, subclass astrocytoma; IDH glioma, subclass high-grade astrocytoma; low-grade glioma, MYB / MYBL1; low-grade glioma, rosette-forming glioneuronal tumor; lymphoma; medulloblastoma, subclass group 3; medulloblastoma, subclass group 4; medulloblastoma, subclass SHH A (pediatric and adult); medulloblastoma, subclass SHH B (infant), medulloblastoma, WNT; melanoma; pineal papillary tumor Group B; pineal parenchymal tumor; pineoblastoma Group A / intracranial retinoblastoma, pineoblastoma Group B; plexus tumor, subclass childhood B, retinoblastoma, or spinal subependymoma. In some embodiments, the central nervous system cancer is a recurrent cancer.
[0098] In some embodiments, the pediatric or adult cancer is bladder urothelial carcinoma, invasive breast cancer, colon adenocarcinoma, esophageal cancer, head and neck squamous cell carcinoma, renal clear cell carcinoma, papillary renal cell carcinoma, liver hepatocellular carcinoma, lung adenocarcinoma, lung squamous cell carcinoma, pancreatic adenocarcinoma, prostate cancer, thyroid cancer, or uterine endometrioid carcinoma. In some embodiments, the cancer is breast cancer. In some embodiments, the cancer is recurrent cancer (e.g., recurrent colon adenocarcinoma, recurrent esophageal cancer, recurrent lung adenocarcinoma, etc.).
[0099] Once a subject is identified as having, for example, a childhood cancer, an adult cancer, or an MRD, a clinical procedure or cancer treatment can be administered to the subject. Exemplary therapies or procedures include, but are not limited to, surgery, radiation therapy, chemotherapy, hormone therapy, targeted therapy, and / or administration of an effective amount of one or more of the following therapeutic agents: angiogenesis inhibitors, such as angiostatin K1-3, DL-α-difluoromethylornithine, endostatin, fumagillin, genistein, minocycline, staurosporine, and (±)-thalidomide; DNA intercalators / crosslinkers, such as bleomycin, carboplatin, carmustine, chlorambucil, cyclophosphamide, cis-diammineplatinum(II) dichloride (cisplatin), melphalan, mitoxantrone, oxaliplatin; DNA synthesis inhibitors, such as (±)-amethopterin (methotrexate), 3-amino-1,2,4-benzotriazine, 1,4-dioxide, aminopterin, cytosine β-D-arabinofuranosides, 5-fluoro-5′-deoxyuridine, 5-fluorouracil, ganciclovir, hydroxyurea, and mitomycin C; DNA-RNA transcription regulators, such as actinomycin D, daunorubicin, doxorubicin, homoharringtonine, and idarubicin; enzyme inhibitors, S(+)-camptothecin, curcumin, (-)-deguelin, and 5,6-dichlorobenzimidazole 1-β-D-ribofuranoside, etoposide, formestane, fostriecin, hispidin, 2-imino-1-imidazolidineacetic acid (cyclocreatine), mevinolin, trichostatin A, tyrophostin AG34, and tyrophostin AG879; gene modulators, such as 5-aza-2′-deoxycytidine, 5-azacytidine, cholecalciferol (vitamin D3), 4-hydroxytamoxifen, melatonin, mifepristone, raloxifene, all-trans-retinal (vitamin A aldehyde), retinoic acid, all-trans (vitamin A acid), 9-cis-retinoic acid, 13-cis-retinoic acid, retinol (vitamin A), tamoxifen, and troglitazone;Microtubule inhibitors, such as colchicine, dolastatin 15, nocodazole, paclitaxel, podophyllotoxin, rhizoxin, vinblastine, vincristine, vindesine, and vinorelbine (navelbine); and unclassified antitumor agents, such as 17-(allylamino)-17-demethoxygeldanamycin, 4-amino-1,8-naphthalimide, apigenin, brefeldin A, cimetidine, dichloromethylenediphosphonic acid, leuprolide (leuprorelin), luteinizing hormone-releasing hormone, pifithrin-α, rapamycin, sex hormone-binding globulin, thapsigallin, and urinary trypsin inhibitor fragment (bikunin). The antitumor agent may be a neoantigen. Neoantigens are tumor-associated peptides that function as active pharmaceutical ingredients in vaccine compositions that stimulate anti-tumor responses, and are described in U.S. Patent Application Publication No. 2011 / 0293637, which is incorporated herein by reference in its entirety. Anti-tumor agents include monoclonal antibodies such as rituximab, alemtuzumab, ipilimumab, bevacizumab, cetuximab, panitumumab, and trastuzumab, vemurafenib, imatinib mesylate, erlotinib, gefitinib, vismodegib; 90 Y-ibritumomab tiuxetan, 131 l-tositumomab, ado-trastuzumab emtansine, lapatinib, pertuzumab, ado-trastuzumab emtansine, regorafenib, sunitinib, denosumab, sorafenib, pazopanib, axitinib, dasatinib, nilotinib, bosutinib, ofatumumab, obinutuzumab, ibrutinib, idelalisib, crizotinib, erlotinib (Tarceva®), afatinib dimaleate, ceritinib, tositumomab, and 131The anti-tumor agent may be I-tositumomab, ibritumomab tiucetan, brentuximab vedotin, bortezomib, siltuximab, trametinib, dabrafenib, pembrolizumab, carfilzomib, ramucirumab, cabozantinib, or vandetanib. The anti-tumor agent may be a cytokine such as interferon (INF), interleukin (IL), or hematopoietic growth factor. The anti-tumor agent may be INF-α, IL-2, aldesleukin, IL-2, erythropoietin, granulocyte-macrophage colony-stimulating factor (GM-CSF), or granulocyte colony-stimulating factor. Antitumor agents include toremifene, fulvestrant, anastrozole, exemestane, letrozole, dib-aflibercept, alitretinoin, temsirolimus, tretinoin, denileukin-diftitox, vorinostat, romidepsin, bexarotene, pralatrexate, lenalidomide, belinstat, pomalidomide, cabazitaxel, enzalutamide, abiraterone acetate, 223 The anti-tumor agent may be a targeted therapy agent such as radium chloride or everolimus. The anti-tumor agent may be a checkpoint inhibitor such as an inhibitor of the programmed cell death-1 (PD-1) pathway, e.g., an anti-PD1 antibody (nivolumab). The inhibitor may be an anti-cytotoxic T-lymphocyte-associated antigen (CTLA-4) antibody. The inhibitor may target another member of the CD28 CTLA4 Ig superfamily, such as BTLA, LAG3, ICOS, PDL1, or KIR. The checkpoint inhibitor may target a member of the TNFR superfamily, such as CD40, OX40, CD137, GITR, CD27, or TIM-3. Furthermore, the anti-tumor agent may be an epigenetic targeting drug such as an HDAC inhibitor, a kinase inhibitor, a DNA methyltransferase inhibitor, a histone demethylase inhibitor, or a histone methylation inhibitor. The epigenetic drug can be azacitidine, decitabine, vorinostat, romidepsin, or ruxolitinib.
[0100] In some embodiments, methods for treating pediatric cancer, adult cancer, or MRD may comprise administering an effective amount of a suitable agent capable of targeting intracellular proteins, small molecules, or nucleic acid molecules, alone or in combination with a suitable carrier or vehicle, including, but not limited to, antibodies or functional fragments thereof (e.g., Fab′, F(ab′)2, Fab, Fv, rlgG, and scFv fragments, and genetically engineered or otherwise modified forms of immunoglobulins, e.g., intracellular antibodies and chimeric antibodies), small molecule inhibitors of proteins, chimeric proteins, or peptides, gene therapy agents for transcription inhibition, or RNA interference (RNAi)-related molecules or morpholino molecules that inhibit gene expression and / or translation. In one embodiment, the inhibitor is an RNAi-related molecule, such as siRNA or shRNA, for the inhibition of translation. RNA interference (RNAi) molecules are small nucleic acid molecules, such as small interfering RNA (siRNA), double-stranded RNA (dsRNA), microRNA (miRNA), or short hairpin RNA (shRNA) molecules, that bind complementary to a portion of a target gene or mRNA so as to reduce the expression level of the target.
[0101] Appropriate pharmaceutical compositions containing one or more agents described herein are administered and dosed according to good medical practice, taking into account the individual patient's clinical condition, the site and method of administration, the administration schedule, the patient's age, sex, and weight, and other factors known to medical professionals. The therapeutically effective amount for purposes herein is therefore determined by considerations known to those skilled in the art. For example, an effective amount of a pharmaceutical composition is the amount necessary to produce a therapeutically effective decrease in the expression of a target gene. The amount of the pharmaceutical composition should be effective to achieve improvement, including, but not limited to, complete prevention, improved survival or more rapid recovery, or improvement or elimination of symptoms associated with the chronic inflammatory condition being treated, and other indicators selected as appropriate measures by those skilled in the art. In accordance with the technology of the present invention, an appropriate single dose size is a dose that, when administered one or more times over an appropriate period of time, can prevent or alleviate (reduce or eliminate) symptoms in a patient. Those skilled in the art can easily determine the appropriate single dose size for systemic administration based on the patient's size and the route of administration.
[0102] Pharmaceutical compositions can be formulated according to known methods for preparing pharmaceutically useful compositions. Furthermore, as used herein, "pharmaceutically acceptable carrier" refers to any standard pharmaceutically acceptable carrier. Pharmaceutically acceptable carriers can include diluents, adjuvants, and vehicles, as well as implant carriers, and inert, non-toxic solid or liquid fillers, diluents, or encapsulating materials that do not react with the active ingredients of the technology. Examples include, but are not limited to, phosphate-buffered saline, saline, water, and emulsions such as oil / water emulsions. Carriers can be solvents or dispersion media containing, for example, ethanol, polyols (e.g., glycerol, propylene glycol, liquid polyethylene glycol, etc.), suitable mixtures thereof, and vegetable oils.
[0103] Pharmaceutically acceptable carrier compositions are described in numerous references well known and readily available to those skilled in the art. For example, Remington: The Science and Practice of Pharmacy (Gerbino, PP2005) Philadelphia, Pa., Lippincott Williams & Wilkins, 21st ed.) describes pharmaceutical formulations that can be used in connection with this technology. Suitable formulations for parenteral administration include sterile injectable solutions, which may contain, for example, antioxidants, buffers, bacteriostats, and solutes that render the formulation isotonic with the blood of the intended recipient, as well as aqueous and non-aqueous sterile suspensions, which may contain suspending agents and thickening agents. The formulations may be in single-dose or multi-dose containers (e.g., sealed ampoules and vials) and may be stored in a freeze-dried (lyophilized) state, requiring only a sterile liquid carrier (e.g., water for injection) prior to use. Extemporaneous injection solutions and suspensions may be prepared from sterile powders, granules, tablets, etc. In addition to the ingredients particularly mentioned above, the formulations of the present technology may include other agents conventional in the art having regard to the type of formulation in question.
[0104] The present disclosure also provides assay panels (nucleic acid hybridization probes or sets of probes) comprising a plurality of polynucleotide probes, each of which is configured to hybridize to bisulfite converted fragments obtained from processing of DNA from a subject, or more preferably, cfDNA molecules, each of which corresponds to, is derived from, or comprises one or more target regions selected from Table 1.
[0105] In some embodiments, the methods described herein can also be implemented using a computer system. For example, any of the above steps for evaluating sequence reads to determine the methylation status of CpG sites can be performed by software components loaded onto a computer or other information appliance or digital device. When so enabled, the computer, appliance, or device can then perform all or part of the above steps to assist in the analysis of values associated with the methylation of one or more CpG sites or to compare such associated values. The above features implemented in one or more computer programs can be performed by one or more computers executing such programs.
[0106] Additionally, various aspects of the methods disclosed herein may be implemented using computer-based computation, machine learning (e.g., support vector machines (SVM), lasso, generalized linear models (GLM), gradient boosting models (GBM), extreme gradient boosting (XGB), elastic net regularized generalized linear models (Glmnet), random forests, gradient boosting (on random forests, C5.0 decision trees), and other software tools, or combinations thereof. For example, methylation states for CpG sites may be assigned computationally based on the underlying sequence reads of amplicons from a sequencing assay. In another example, methylation values for DNA regions or portions thereof may be compared computationally to thresholds, as described herein. These tools are advantageously provided in the form of computer programs executable by conventional general-purpose computer systems.
[0107] In some embodiments, methods used to analyze and / or determine the methylation level of a target polynucleotide region include Metilene (Juhling et al., Genome Res., 2016;26(2):256-262) or GenomeStudio Software available online from Illumina, Inc., or the methods described in Hovestadt et al., 2014; Nature, 510(7506), 537-541. In some embodiments, the methylation data may be further processed by algorithms and / or software to determine difference values (i.e., differential methylation values) and identify differentially methylated regions (DMRs). Differential methylation values may be calculated by methods known in the art (e.g., Hovestadt et al. (2014). Nature, 510(7506), 537-541). In some embodiments, Metilene, a software program for calling differentially methylated regions, may be used to identify differentially methylated regions in whole-genome and targeted sequencing data. In some embodiments, the methylation data may be classified into hypomethylated target regions, hypermethylated target regions, and a combination of hypomethylated and hypermethylated regions. In some embodiments, a machine learning model may be trained using hypomethylated target regions, hypermethylated target regions, or both.
[0108] In some embodiments, a method for identifying pediatric cancer, adult cancer, or MRD in a subject may include the use of a machine learning model. The machine learning model may be a trained algorithm. The machine learning model is trained on one or more features and used to process a dataset generated by assaying nucleic acid molecules in a sample (e.g., an acellular biological sample). The dataset includes a methylation profile of one or more target genomic regions of the biological sample. Examples of the use and training of machine learning models are described, for example, in International Publication Nos. WO 2022 / 178108 (Salhia et al.); WO 2019 / 178277 (Gross et al.); and U.S. Patent Application Publication Nos. US 2021 / 00250 (Gross et al.); and US 2019 / 0287652 (Gross et al.). In some embodiments, the machine learning model is trained using a sample of DNA containing a target region in Table 1 obtained from a subject with known cancer. In some embodiments, the machine learning model is trained using 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more cancer samples.
[0109] In some embodiments, target regions of DNA isolated from a subject may be partitioned into hypomethylated target regions, hypermethylated target regions, or a combination thereof, and analyzed separately, e.g., using a machine learning model. As another example, target regions of a subject's DNA may be analyzed using a machine model trained only on hypermethylated or hypomethylated regions.
[0110] In some embodiments, a computer including at least one processor can be configured to receive multiple sequencing results from a DNA methylation sequencing reaction (e.g., after WGBS) from a patient with or suspected of having a tumor or other tumor (e.g., DNA isolated from a tumor or tumor, or DNA isolated from cfDNA from a person with a tumor or tumor). The results can include methylation patterns of one or more target regions (e.g., Table 1) disclosed herein. The methylation patterns of the received sequence reads can be determined by sequence alignment with a reference genome, for example, using Metilene or other commercially available products to identify regions that are differentially methylated between the sample and a normal sample or pool of normal samples. A trained machine learning model can be applied to these results to output a result (e.g., pediatric or adult cancer, presence or absence of MRD, etc.). As used herein, a "processor" may be any type of processor, such as, for example, a general-purpose microprocessor or microcontroller (e.g., Intel™ x86, PowerPC™, ARM™ processor, etc.), a digital signal processor (DSP), an integrated circuit, a field programmable gate array (FPGA), or a combination thereof.
[0111] In some embodiments, the machine learning model used to detect pediatric cancer, adult cancer, or MRD comprises analyzing the methylation pattern of a plurality of target regions in a cancer sample compared to the methylation pattern of a plurality of target regions in a non-cancer sample. In some embodiments, a signature of pediatric cancer, adult cancer, and / or MRD is detected by determining and analyzing the methylation pattern of a plurality of target regions in both cancer and non-cancer samples, wherein the plurality of target regions comprises at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% of the target regions listed in Table 1. In some embodiments, the signature is developed by analyzing and comparing all of the target regions in Table 1.
[0112] In some embodiments, the machine learning model is trained using target regions from a plurality of cancer samples and corresponding target regions from non-cancer samples, wherein the cancer samples comprise at least two different cancer types, and wherein detection of pediatric cancer, adult cancer, or MRD is based on a comparison of the methylation pattern of the target regions in the cancer samples compared to the methylation pattern of the corresponding target regions in the non-cancer samples.
[0113] In some embodiments, the machine learning model is trained by the following procedure: Aligning the sequence reads of target regions of multiple cancer samples to corresponding regions of a reference genome; Then, performing methylation analysis of the target regions of multiple cancer samples against a pool of normal samples or normal samples to identify differentially methylated target regions; Identifying differentially methylated regions that are common among cancer samples; Using the common or shared differentially methylated regions, train the machine learning model to distinguish between cancer samples and non-cancer samples. In some embodiments, the target regions of the samples that have a methylation level different from that of the target sequence of the normal sample are differentially methylated regions (DMRs). In some embodiments, the sequence reads are aligned after whole genome sequencing (e.g., WGBS, RRBS).
[0114] In some embodiments, training the machine learning model includes identifying a minimum number of differentially methylated regions that are common among multiple cancer samples. In some embodiments, the minimal differentially methylated regions (mDMRs) are common to about 40% of the cancer samples, about 45% of the cancer samples, about 50% of the cancer samples, about 55% of the cancer samples, about 60% of the cancer samples, about 65% of the cancer samples, about 70% of the cancer samples, about 75% of the cancer samples, about 80% of the cancer samples, about 85% of the cancer samples, about 90% of the cancer samples, about 95% of the cancer samples, or more than 95% of the cancer samples. In some embodiments, the mDMR is common or shared among about 70% of the cancers from which the target sample is derived, about 75% of the cancers from which the target sample is derived, about 80% of the cancers from which the target sample is derived, about 85% of the cancers from which the target sample is derived, about 90% of the cancers from which the target sample is derived, about 95% of the cancers from which the target sample is derived, or about 100% of the cancers from which the target sample is derived.
[0115] In some embodiments, training the machine learning model comprises the following steps: i) receiving methylation sequencing reads of a plurality of test samples, including cancer samples and non-cancer samples, to obtain methylation patterns of target regions in the plurality of test samples, wherein the cancer samples comprise at least two different cancer types; ii) aligning the target regions of the plurality of test samples to a reference genome, wherein each target region of the plurality of test samples is aligned with a corresponding target region in the reference genome; iii) performing differential methylation region analysis (e.g., using Metilene) of the test sample compared to a normal sample or pool of normal samples to identify regions that are differentially methylated between the test (cancer) sample and the normal sample; vi) defining minimal differentially methylated regions (mDMRs) between the test samples; and vii) using a machine learning tool to build a classifier model based on the mDMRs to distinguish between cancer and non-cancer samples, wherein an AUC value of 0.8 or greater indicates the presence of pediatric cancer, adult cancer, or MRD. In some embodiments, step iii) further comprises determining hypomethylated and hypermethylated target regions, for example, using a methylation caller (e.g., Metilene).
[0116] In some embodiments, target regions may be assigned a methylation value based on, for example, a positive number if the target region has CpGs that are methylated compared to corresponding positions in the reference genome (i.e., the CpGs in the reference genome are unmethylated), and a negative number if the target region has CpGs that are unmethylated compared to corresponding positions in the reference genome (i.e., the CpGs in the reference genome are methylated). Thus, a hypermethylated target region may have a positive overall score, and a hypomethylated target region may have a negative overall score.
[0117] In some embodiments, the output of the machine learning model comprises an area under the curve (AUC) value for the test sample. In some embodiments, an AUC of 0.8 or greater, 0.85 or greater, 0.9 or greater, 0.95 or greater, 0.96 or greater, 0.97 or greater, 0.98 or greater, or 0.99 or greater indicates the presence of pediatric cancer, adult cancer, or MRD.
[0118] In some embodiments, the first set of target regions, the second set of target regions, and the corresponding target regions comprise the same target regions of Table 1.
[0119] In some embodiments, the cancer samples used to train the machine learning model include stage I-IV cancer samples, such as metastatic breast cancer. In other embodiments, the cancer samples used to train the machine learning model are from stage I or stage II cancer samples. The methylation patterns of the cancer samples may then be compared to the methylation patterns of non-cancerous samples.
[0120] Any of the computer-readable media described herein can be non-transitory (e.g., volatile memory such as DRAM or SRAM; non-volatile memory such as magnetic storage, optical storage, etc.) and / or tangible memory. Any of the save operations described herein can be performed by storing in one or more computer-readable media (e.g., computer-readable storage media or other tangible media). Anything described as "stored" herein (e.g., data created and used during performance) can be stored in one or more computer-readable media (e.g., computer-readable storage media or other tangible media). Computer-readable media may be limited to embodiments that do not consist of signals.
[0121] In some embodiments, kits for detecting cancers such as pediatric cancers, adult cancers, MRD, etc. are provided, comprising reagents for carrying out the aforementioned methods and instructions for detecting cancer signals. The reagents may include, for example, primer sets, PCR reaction components, multiple probe sets complementary to target regions in Table 1, sequencing reagents, and, optionally, solid supports for the probes (e.g., glass slides or chips, surfaces of beads, surfaces of matrices, etc.).
[0122] The following examples are intended as illustrations to explain the above invention and should not be construed as limiting its scope. Those skilled in the art will readily recognize that the examples suggest many other ways of practicing the invention. It should be understood that numerous variations and modifications can be made within the scope of the invention.
[0123] Example Example 1. Analysis of differential DNA methylation patterns in childhood cancer. We performed WGBS on 31 tumor samples and 13 patient-matched adjacent normal tissue samples (Table 2), representing 11 different pediatric cancer types. We first performed differential methylation analysis using Metilene to identify differentially methylated regions (DMRs). Tumor samples were compared with their patient-matched adjacent normal samples, if available. For tumor samples lacking matched normal tissue, a pool of normal samples from other patients with the same diagnosis was used. Differential DNA methylation analysis revealed a variable number of DMRs both within and between tumor types (Figure 8A-B, Table 3). On average, 74% of DMRs across all tumors were hypomethylated, compared with 26% hypermethylated. However, malignant rhabdoid tumors (MRTs) displayed a hypermethylator phenotype, with 90% of DMRs hypermethylated (Figure 9A). Genome-wide hypermethylation in MRTs was also observed in 68 TARGET samples (Figure 9B). [Table 2]
[0124] Hierarchical clustering using the 2.5% most variable and statistically significant DMR calls (183 unique regions after merging overlapping regions) demonstrated separation by tumor type and identified four distinct DMR subgroups (Figures 1A-B, 10). These DMRs were primarily located in CpG islands (n = 151), with the remaining DMRs located in CpG shores (n = 6), shelves (n = 4), and open seas (n = 22). Ingenuity Pathway Analysis (IPA) of genes in Cluster 1 revealed enrichment for netrin signaling and GABA receptor signaling (Table 4). Cluster 2 showed enrichment for embryonic stem cell differentiation and sonic hedgehog signaling (Table 5). Cluster 3 (Figure 1A, Table 6, 10) was primarily hypermethylated in neuroblastoma (NBL) and associated with ERK / MAPK signaling. Cluster 4 was hypomethylated in neuroblastoma and hypermethylated in all other tumor types. 20 of the 32 DMR calls in this cluster were within known genes. Genes associated with this cluster included multiple tumor suppressor genes (BRCA, KANK1, ASB3, NBAT1, PIP4K2A, NFATC1, and ZNRF3) and oncogenes (HOXA3 and HOXB-AS3). IPA identified genes with DMRs in this cluster associated with PI3K signaling, DNA double-strand break repair, and G2 / M checkpoint regulation (Table 7). [Table 3] [Table 4] [Table 5] [Table 6] Methylation beta values were also extracted from 518 TARGET samples (Table 8) across 166 of the 183 highly variable DMRs that overlapped with at least one probe on the HM450 methylation array. Analysis of the TARGET data also demonstrated strong separation by tumor type (Figure 1C-D). POETIC samples clustered with TARGET samples according to tumor type, as seen by hierarchical clustering and UMAP analysis (Figure 1E). It is also important to note that TARGET samples were primarily collected from primary tumors (Table 8), whereas all POETIC cases were collected from patients with recurrent metastatic disease, indicating that the methylation profiles of recurrent tumors are more similar than distinct from those of primary tumors. Coclustering of POETIC recurrent and TARGET samples was also observed when each tumor type was analyzed separately (data not shown). [Table 7]
[0125] DNA methylation changes shared across cancer types. We identified a set of minimally differentially methylated regions (mDMRs) as subregions of each DMR that were shared (in the same direction; hypomethylated or hypermethylated) across multiple samples (and cancer types). Briefly, each CpG site was scored based on the number of samples with a DMR call containing that CpG (separately for hypomethylated and hypermethylated DMR calls). Clusters of adjacent CpG sites, each with a DMR present in at least N samples, were merged to create contiguous mDMR regions. Although the number of shared mDMRs decreased as N increased for both hypomethylated and hypermethylated regions (Figure 2A), it was possible to identify a set of mDMRs shared between at least 70% (n = 22 / 31) of samples. The mDMRs in this subset included 402 hypomethylated regions with a median width of 276 bp and 503 hypermethylated regions with a median width of 230 bp (905 in total, Figure 2A). The average beta value across hypomethylated mDMRs was 0.406 in tumor tissues compared with 0.732 in normal tissues. The beta value across hypermethylated mDMRs was 0.643 in tumor tissues compared with 0.308 in normal tissues (Figure 2B).
[0126] Based on the methylation patterns of the mDMR set, we constructed a random forest classifier to evaluate the utility of using the data for tumor detection. The cross-validated receiver operating characteristic (ROC) curve for methylation in mDMRs had an area under the curve (AUC) of 0.95, indicating that mDMRs can distinguish tumors from normal tissue (Figure 2C).
[0127] To further validate the mDMR signature, we analyzed 518 samples from the TARGET database (Table 8) for Wilms tumor (WT), MRT, osteosarcoma (OS), and NBL samples. We also analyzed 90 normal tissue samples from multiple tissue types from "The Encyclopedia of DNA Elements" (ENCODE) (Table 9). Both of these datasets were analyzed using Infinium methylation HM450 and EPIC arrays (Illumina). We were able to estimate the average beta values for 344 hypermethylated and 71 hypomethylated mDMRs that overlapped with at least one probe on the HM450 and EPIC beadchips. We found that the methylation patterns of mDMRs in the TARGET data were similar to those in the POETIC WGBS data and were more hypermethylated or hypomethylated in tumors compared with normal controls (Figure 3; Wilcoxon p-values <0.001 for all comparisons). Taken together, these results demonstrate that the identified mDMRs are generalizable across multiple pediatric cancer types. [Table 8] Hypomethylated mDMRs were located throughout the genome, with 152 mDMRs within known genes, 30 mDMRs within 2000 bp of gene transcription start sites, and 220 mDMRs in intergenic regions. The majority of these regions were located in open seas (n=328), with 16 regions in CpG islands, 39 regions in CpG shores, and 18 regions in CpG shelves. IPA analysis indicated that genes associated with hypomethylated mDMRs are involved in multiple immune signaling pathways, including natural killer cell signaling, WNT / β-catenin signaling, and TREM1 signaling (Table 10). [Table 9] Hypermethylated mDMRs were more likely to be associated with genes than hypomethylated mDMRs, with 316 of the 503 regions located within known genes (63% compared with 38% in hypomethylated regions). Most hypermethylated mDMRs were located in CpG islands (n = 387), with the remaining mDMRs located in CpG shores (n = 6), shelves (n = 4), and open seas (n = 22). Hypermethylated mDMRs were disproportionately associated with the protocadherin genes PCDHGA8, PCDHGA1, PCDHA1, and PCDHA9. IPA analysis found that genes associated with selected mDMRs were significantly associated with epithelial-mesenchymal transition (EMT), NANOG signaling, TGF signaling, and TREM1 signaling (Table 11). Together, these results indicate that the selected mDMRs may represent a pan-pediatric cancer signature associated with a wide range of cancer-specific pathways. [Table 10] TIFF2026506978000040.tif234159TIFF2026506978000041.tif17159 [Table 11]
[0128] mDMRs in pediatric cancers are also detected in adult cancers. To determine the generalizability of the 905 mDMRs across a broad set of adult tumor types, we evaluated a tumor / normal classifier using HM450 data from TCGA. 422 of the 905 mDMRs overlapped with at least one probe on the HM450 array, and we constructed a separate random forest classifier based on this reduced set of regions. ROC curves (and AUCs) were calculated for 14 adult solid tumors (6,426 samples, Table 14) (Figure 4A-B). Our model achieved an average AUC of 0.95, demonstrating that the reduced CpG mDMR set can distinguish tumors from normals in samples obtained from TCGA. All cancer types except two (THCA - thyroid cancer and PRAD - prostate adenocarcinoma) achieved an AUC > 0.9 (Figure 4C). Performance was consistent across all tumor stages in TCGA (Figures 11-14). This indicates that the identified panel of CpG mDMRs could potentially serve as pan-cancer methylation biomarkers applicable to both pediatric and adult cancers.
[0129] mDMRs distinguish some CNS tumors from normal tissues. Although mDMR signatures have been identified in non-CNS tumors, we sought to determine whether these regions are generalizable to pediatric CNS tumors, as CNS tumors are the second most common tumor type and a high-risk tumor in children. To do this, we downloaded CNS cancer samples profiled by the HM450 array generated by Capper et al. (Nature 555, 469-474 (2018)) (hereafter referred to as "DKFZ samples"). This dataset contains 2,682 cancer samples and 119 normal control CNS tissues. This dataset includes 91 unique methylation classes (Capper et al., Table 3)—82 cancer and 9 normal tissue types. We calculated the average beta across 422 mDMRs that overlap with at least one probe on the HM450 array. We found 44 of the 82 tumor types in which beta values differed significantly between tumor and control tissues (Wilcoxon p-value <0.05) with consistent directionality for both hypermethylated and hypomethylated mDMRs. We then input the beta values into the same random forest model used to classify the TCGA samples, and ROC curves were calculated for each methylation class using 119 control CNS tissues as controls. The classifier was able to distinguish tumor from normal with high sensitivity and specificity (AUC >0.9) in 41 of the 82 CNS tumor types (Figure 15). Notably, some tumor types in the DKFZ samples are histologically similar to tumor types in the POETIC cohort. For example, three methylation classes—esthesioneuroblastoma (ENB) A, ENB B, and CN NBL—represent types of NBL that arise in olfactory neurons and other neural crest cells within the CNS, respectively. All CN NBL cases were classified as tumors (AUC=1), with excellent performance in ENB A (AUC=0.97) and moderate performance in ENB B (AUC=0.67). Similarly, all ATRT DKFZ samples, histologically similar to the POETIC MRT samples, were classified as cancers by the random forest model (AUC=1).Taken together, these results demonstrate that mDMR is generalizable to several CNS tumor types and is highly generalizable to similar tumor types.
[0130] Cell-free DNA methylation reflects the methylation patterns in tumor tissue. Cell-free DNA methylation is gaining widespread acceptance as a novel biomarker for liquid biopsies. In this study, we performed WGBS on 17 individual patient-matched cfDNA samples and three normal samples to determine the extent to which DNA methylation patterns in genomic DNA (gDNA) are found in cfDNA (Table 2). cfDNA extracted from these samples had a mean yield of 20.4 ng / ml (3-2-87.5 ng) (Figure 16). DMRs were called between the cfDNA from cancer patients and three healthy cfDNA samples and filtered in the same manner as described above for tissue-derived gDNA. The median number of DMRs per sample was 42,935 (Figure 5A) (Table 13). [Table 12]
[0131] We determined the number of overlapping DMRs between cfDNA and gDNA in patient-matched samples. To count as an overlapping DMR, at least three methylated CpGs in the same direction had to be present in both sample types. The average percentage of overlapping DMRs between cfDNA and gDNA (Figure 5B) was 24.9% (Bettegowda et al., Sci Transl Med 6, 224ra224 (2014)). Some notable exceptions included the following samples, which had exceptional overlap between cfDNA and gDNA DMRs: P01-036 (NBL, 76.0%), P0-021 (hepatoblastoma (HB), 45.6%), P01-028 (teratoma, 35.8%), and P01-029 (OS, 37.4%). Overlapping DMRs between gDNA and cfDNA had an average R of 0.45. 2 and were positively correlated (Figure 5C).
[0132] Pan-pediatric cancer mDMRs detectable in cfDNA. To determine whether the 905 mDMRs (402 hypomethylated and 503 hypermethylated) across tumor types (Figure 2) could distinguish tumors from normals in cfDNA, we performed WGBS in 17 pediatric cancer patients and 15 healthy controls and calculated the average beta across all hypermethylated / hypomethylated mDMRs identified in the patient's matched tumor tissue for each WGBS cfDNA sample. The average methylation for tissue-identified hypomethylated mDMRs was significantly lower in cfDNA cancer samples compared to normals, with an average methylation difference of 0.10 (p<0.0001) (Figure 6A). However, the average methylation for tissue-identified hypermethylated regions was also unexpectedly lower in cfDNA cancer samples compared to normals (Figure 6B). Therefore, we focused on further validation of mDMRs in only hypomethylated regions and used a targeted hybridization probe capture assay designed against 402 hypomethylated mDMRs.
[0133] We obtained an additional cohort of 44 pediatric cancer samples from Children's Hospital Los Angeles (CHLA), for which cfDNA was not available. This cohort consisted of six cancer types: NBL, OS, WT, anaplastic small round cell tumor (DSRCT), embryonal rhabdomyosarcoma (ERMS), and teratoma (Table 14). We found that each of these tumor types was also significantly hypomethylated relative to normal tissue (Figure 7A). Hierarchical clustering showed that CHLA samples clustered with POETIC samples and separated from normal tissue samples (Figure 7B). Furthermore, UMAP analysis showed no origin effect (Figure 7C), indicating that tumors of the same type clustered together regardless of origin (Figure 7D). [Table 13] TIFF2026506978000045.tif27159
[0134] Identification of DNA methylation deserts in neuroblastoma. After identifying and validating common focal epigenetic alterations, we shifted our focus to large-scale genomic and epigenomic alterations, such as partially methylated regions (PMDs; large regions of hypomethylation spanning hundreds of kilobases to several megabases, which have been described as common epigenetic alterations in cancer) and copy number variations. PMD-positive samples had an average of 98 PMDs greater than 1 megabase. We also identified significantly hypomethylated, multi-kilobase regions, significantly below the levels previously reported for PMD, which we term DNA methylation deserts. These deserts were found in 5 of 9 neuroblastoma cases but not in other tumor types. Methylation deserts were characterized by an average DNA methylation beta value of less than 0.2, which is significantly lower than that seen in conventional PMDs, which typically have a beta value of 0.7 (Lister et al., Nature 462, 315-322 (2009)). PMD and methylation desert were observed in cfDNA from P01-010, P01-029, and P01-036.
[0135] Interestingly, the forkhead box genes FOXA1 and FOXP2 were found within the methylation deserts of five NBL samples. Conventional PMDs were also detected in multiple other tumors, including three additional NBL samples and 11 other samples from one patient, including five EMRS, three OS, and two HB samples. Using permutation tests (Gel et al., Bioinformatics 32, 289-291 (2016)), we assessed whether these PMDs and methylation deserts were associated with copy number variation, but found no significant associations. We identified 178 regions that contained PMD calls in more than 50% of the samples; these regions ranged from 5 kb to 2.1 Mb, had an average width of 308 kb, and an average of 2,699 CpGs per region. Hierarchical clustering of samples based on overall methylation values within these 178 consensus regions separated the 5 NBL samples with methylation deserts from the 14 samples with PMD and from the remaining samples (in which PMD was not significantly called).
[0136] Copy number estimation by WGBS in gDNA from tumor tissue. WGBS is primarily used to measure DNA methylation across the genome, but it can also be reliably used to determine copy number variations (CNVs). In this study, we performed copy number analysis from WGBS data, which revealed CNVs in recurrent tumors consistent with known copy number variations in the various primary tumor types studied. Specifically, we identified 1q gain, a structural mutation associated with poor prognosis, in 18 of 31 solid tumor samples (1 / 1 anaplastic ependymoma (AE), 1 / 1 DSRCT, 7 / 10 NBL, 5 / 5 ERMS, and 3 / 3 HB). Strong 8q gain in two hepatoblastoma tumor samples was also associated with poor prognosis. 8q was amplified in P01-020, with a mean copy number of 3.7 in the first sample collected (P01-020-T-1) and 4.8 in the second sample collected 6 months later (P01-020-T-2). The copy number profiles in these two samples were generally consistent, except for a loss of 13q and a gain of 18q observed at the second but not the first time point. In the NBL samples, we observed a loss of 1p in 4 of 10 samples, a gain of 17q in all samples, a loss of 11q in 6 of 10 samples, and a loss of 3p in 5 of 10 samples. All of these structural variants are relatively common in NBL. In OS samples (P01-012, P01-016, P01-029, and P01-030), copy number alterations were widespread and chaotic, often associated with numerous structural alterations and chromosomal breakages, characteristic of OS. Four samples with the excess methylator phenotype did not have large chromosomal abnormalities: P01-019-T-1 (MRT), P01-019-T-2 (MRT), P01-027-T1 (Hodgkin lymphoma), and P01-024-T-1 (WT).
[0137] We also detected numerous focal deletions and amplifications known to drive specific tumor types in this study. N-MYC amplification was observed in two of 10 NBL samples. N-MYC amplification in NBL occurs in 20–25% of NBL cases and is associated with poor prognosis. We identified homozygous loss of SMARCB1 in two MRT samples, a known driver mutation for that tumor type. We observed homozygous loss of MMP11 in MRT. While MMP11 CNVs have been reported in NBL and WT in the Catalog of Somatic Mutations in Cancer (COSMIC), no MMP11 CNVs have been reported in MRT. All samples from P01-022 (ERMS) had loss of PTEN, a CNV previously reported and associated with an aggressive cancer phenotype. This homozygous loss in P01-022 also involved ATAD1, which is associated with cancer cell progression. PTEN loss in these ERMS tumors is likely a driver tumor event present in all cells.
[0138] Detection of CNVs in cfDNA is less sensitive than cfDNA methylation. Next, we analyzed CNVs in cfDNA using WGBS data. Unlike many known CNVs found in tissue gDNA, CNVs were rarely detected in cfDNA. Five samples from three patients were notable exceptions, with CNV calls broadly concordant with CNV detection from tissue gDNA. In one case (P01-020), gains of 1q and 8q observed in two tissue samples were observed in cfDNA. Blood from this patient was collected at the same time as the first tumor sample (P01-020-T1), and both samples had similar CNV profiles. A second tumor sample (P01-020-T2), collected at a later date, showed a loss of 13q that was not seen in the first sample (T1) or its cfDNA. In P01-036 (NBL), cfDNA CNV calls were also highly concordant with tumor sample CNV data, with gains of 1q, 2p, 7, 9q, 12q, 13q, and 17q, and losses of 1p, 3p, 4q / p, 11q, 17q, and 19q. Interestingly, cfDNA was able to alter some CNVs not seen in tissue, namely losses of 10p and 15q. This was also true for P01-029, where plasma CNVs were not seen in gDNA. Furthermore, we were able to identify focal N-MYC amplification in P01-026 but not in P01-023, as previously identified in gDNA. The reason why cfDNA CNV calls were concordant with gDNA in only some samples is unclear, but it was unrelated to tumor cfDNA yield (an indicator of tumor burden) (Figure 16). Overall, this analysis showed that cfDNA methylation is more sensitive for detecting tissue-associated mutations than CNV analysis in plasma.
[0139] In this study, we aimed to determine the set of DMRs common to multiple non-CNS pediatric solid tumors. To this end, we present WGBS data for 31 tumor samples and 13 normal tissue samples from recurrent non-CNS pediatric solid tumors representing 11 different tumor types. We further validated the DNA methylation results in 518 pediatric cancer samples from TARGET and 6,426 adult cancer samples from TCGA, along with an additional 44 pediatric cancer tissue samples examined by targeted hybridization probe capture assay. Furthermore, we analyzed corresponding cfDNA methylation from 17 patient-matched plasma samples. Previous studies have identified DNA methylation alterations in many of the pediatric cancers evaluated in this study, including NBL, OS, WT, AE, HB, ERMS, and fibrolamellar hepatocellular carcinoma, but methylation analysis of recurrent pediatric cancer cases has been limited. A novel finding of this study is that in addition to tumor-specific alterations, DNA methylation patterns are also shared among the diverse tumor types analyzed, regardless of whether the tumors originated from primary or recurrent tumors. Furthermore, DNA methylation alterations have been shown to be common across multiple adult cancer types, frequently associated with tumor suppressor genes and often associated with survival. However, similar findings in pediatric cancers have been rare. Given the rarity of many pediatric cancers, identifying DNA methylation patterns across multiple pediatric cancer types may be of great value for developing biomarker approaches for early detection and / or disease recurrence through MRD detection.
[0140] To identify commonalities in DNA methylation across cancer types, we identified minimal regions of differential methylation common to multiple samples across cancer types, termed minimally differentially methylated regions (mDMRs). These mDMRs were significantly differentially methylated in multiple cancer types, including rare tumors such as DSRCT, when compared with normal tissue. We also evaluated 518 pediatric cancer samples from the TARGET database and detected methylation changes in mDMR regions consistent with those observed in POETIC samples, demonstrating the generalizability of this signature. Furthermore, because this signature was derived from relapsed patients, it may reflect a method for monitoring recurrence by MRD detection in pediatric cancers. Finally, this signature also showed high sensitivity and specificity when tested in adult cancer samples from TCGA, suggesting that these mDMRs may also function as pan-cancer detection markers in adult cancers. Another notable finding of this study is that the methylation profiles of relapsed POETIC samples were highly similar to those of primary tumors obtained from TARGET. Previous studies have shown that primary and recurrent tumors have similar methylation profiles, which has not been previously reported in pediatric cancers, but the TARGET data also suggest this may be true.
[0141] Among the tumor-specific DNA methylation changes, intriguing was the identification of DNA methylation deserts in a subset of NBLs, where large regions of DNA methylation were virtually eliminated. The functional significance of these DNA methylation deserts remains unclear, but the observed degree of loss of DNA methylation differs from previously reported PMDs. Previously, PMDs have been associated with CpG island methylation in breast cancer, and PMDs are also associated with lamina-associated regions (LADs). Increased expression of FOXA1, one of the few genes found in DNA methylation deserts, has previously been associated with late recurrence. Further research is needed to elucidate the possible mechanisms underlying these desert regions in NBLs. Another tumor-specific methylation change we observed was strong global hypermethylation in MRTs, an atypical cancer phenotype in pediatric and adult tumors. Hypermethylation in cancers typically occurs in promoter regions of tumor suppressor genes and in CpG islands, where it is associated with poor prognosis in many adult and pediatric cancers and may be associated with CIMP. However, global hypermethylation has only rarely been reported in cancer. While MRT cells have previously been reported to have focal hypermethylation, the global hypermethylator phenotype we observed has not been reported previously. Furthermore, restoration of SMARCB1 in MRT cell lines has been found to cause widespread chromatin activation. Conversely, this hypermethylator phenotype may be associated with genome-wide chromatin inactivation, although further studies are needed to assess whether this is true.
[0142] Unlike adult cancers, most pediatric cancers stem from a single genetic driver, and these genetic mutations tend to be highly tumor-specific. For example, loss of SMARCB1 specifically leads to the development of MTRT. Specific gene fusions cause alveolar rhabdomyosarcoma (PAX3 / 7-FOXO1), Ewing's sarcoma (EWS-FLI1), and CML (BCR-ABL1). Mutations in certain genes, such as the RB gene, are specific to retinoblastoma and OS, while mutations in TP53 are extremely common in multiple pediatric cancers. WGBS is the gold standard for methylation assessment because it assesses every CpG in the genome, but we also used it for copy number estimation. This maximizes data generation from each sample and enables multi-omics analysis from the same sample and aliquot, which also minimizes sampling bias and reduces costs. This is particularly useful when working with rare tumor types or sample types with very limited DNA availability (e.g., cfDNA). CNV analysis from WGBS identified large-scale mutations and focal gains / losses, such as gain of MYCN in NBL, loss of SMARCB1 in MRT, and loss of PTEN in ERMS.
[0143] A key element of this study was to assess the extent to which genomic alterations in solid tumors are reflected in cfDNA, as cfDNA analysis is rapidly becoming a primary tool for noninvasive screening, diagnosis, treatment, and monitoring of human tumors. In renal cell carcinoma, cfDNA methylation (using MeDIP-seq) was found to be much more sensitive and specific than cfDNA SNV markers. In this study, we found that detecting DNA methylation from tumor tissue using WGBS was superior to detecting CNVs in cfDNA. In this study, only three of 17 cfDNA samples reflected the tumor copy number profile, and the remaining samples lacked much of the signal found in tissue. On the other hand, cfDNA methylation more strongly resembled tissue DNA methylation, with an average of approximately 25% DMRs shared between tissue and plasma in each sample. We also found that hypermethylated regions were likely less retained in plasma and were often hypomethylated. While the reasons for this are unclear, the implications for biomarker development strategies are important.
[0144] In summary, this disclosure provides a comprehensive analysis of multiple pediatric cancers using WGBS. We identified a pan-cancer methylation signature detectable in cfDNA that is common to multiple pediatric cancer types, including extremely rare neoplasms such as DSRCT and MRT. We also detected CNVs using WGBS to directly compare CNV and methylation detection from the same sample and aliquot. We found that DNA methylation was superior to CNVs in detecting tumor-specific signals in cfDNA. The pan-cancer cfDNA methylation signature in this study has potential utility for monitoring and early detection of minimal residual disease and warrants further investigation in pediatric and adult cancers.
[0145] Example 2. Materials and Methods. Sample collection. Samples were obtained from patients with various recurrent childhood cancers with written parental consent from the Pediatric Oncology Experimental Therapeutics Investigators Consortium (POETIC) at Memorial Sloan Kettering Cancer Center (New York, USA) (Table 2). Tissue samples were snap-frozen after resection. Peripheral blood samples were collected in EDTA purple-top tubes before surgery, and plasma was collected. [Table 14] TIFF2026506978000047.tif214159
[0146] All participants in this study were patients with recurrent pediatric cancer. In total, 44 tissue samples were collected from 24 patients, including 31 tumor samples and 13 matched adjacent normal samples. All tissue samples were collected during surgical resection of the recurrent tumor. Seventeen patient-matched plasma samples were collected, along with three additional plasma samples from young healthy individuals purchased from Conversant Bio (Table 2). Twelve additional cfDNA healthy controls from adults were included for model building in cfDNA, as described in detail below (Conversant Bio). Patient ages ranged from 18 months to 35 years. In total, 11 different cancer types were included in this study: neuroblastoma (NBL, n = 9), osteosarcoma (OS, n = 4; 3 from lung metastases and 1 from retroperitoneal metastasis), fibrolamellar hepatocellular carcinoma (FHC, n = 2), hepatoblastoma (HB, n = 2; both resected from the liver with one additional specimen from a lung metastasis), anaplastic ependymoma (n = 1), desmoplastic small round cell tumor (n = 1; from a gastric lesion), malignant rhabdoid tumor (MRT, n = 1; 2 specimens from abdominal lesions), Hodgkin lymphoma (n = 1; from a lung biopsy), embryonal rhabdomyosarcoma (ERMS, n = 1; with 5 specimens from omental, diaphragmatic, venous sinus, pelvic, and sigmoid colon lesions), immature teratoma (n = 1; with 2 specimens from conjunctival sac and ileal lesions), and Wilms' tumor (n = 1). [Table 15]
[0147] We also obtained a separate cohort of 44 pediatric cancer tissue samples for validation from the Children's Hospital Los Angeles Pathology Research Center (CHLA). These samples included six cancer types (N=10 NBL, N=10 OS, N=10 WT, N=5 DSRCT, N=5 teratoma, N=5 ERMS). They were collected with written consent and stored in OCT compound (Table 14). Additionally, we included 68 MRT, 221 NBL, 86 OS, 131 WT, and 12 adjacent normal tissue samples (518 total) from TARGET, 90 normal tissue samples from ENCODE, and 6,426 adult cancer samples from TCGA derived from 14 tumor types. All three of these cohorts were analyzed using the HM450 array (Illumina) (Tables 8, 9, and 12).
[0148] Sample extraction and library preparation for whole-genome bisulfite sequencing. Genomic (g) DNA and total RNA were extracted from flash-frozen normal or tumor tissues using the AllPrep DNA / RNA Mini Kit (Qiagen) according to the manufacturer's instructions. Briefly, tissues were agitated with a mixture of 0.9-2.0 mm RNase-free stainless steel beads using a Bullet Blender homogenizer (Next Advance) at maximum speed for 5 minutes. The homogenate was passed through a QIAshredder (Qiagen) to remove any remaining particulate matter. Plasma was separated from whole blood by spinning it at 300 g for 20 minutes. Cell-free (cf) DNA was extracted from plasma using the QIAamp DNA Blood Maxi Kit (Qiagen) according to the manufacturer's recommendations.
[0149] The quantity and purity of isolated gDNA were determined using the Qubit dsDNA High Sensitivity Fluorescence Assay (Invitrogen). cfDNA was quantified using the TapeStation High Sensitivity D1000 assay according to the manufacturer's protocol. Extracted gDNA and cfDNA were used for whole-genome bisulfite sequencing (WGBS) analysis, as described, for example, in Legendre et al., Clin Epigenetics 7, 100 (2015). Directional, bisulfite-converted libraries for paired-end sequencing were prepared using the Ovation Ultralow Methyl-Seq Library System (NuGen) according to the manufacturer's recommended protocol. Bisulfite conversion was performed using the EpiTect Fast DNA Bisulfite Kit (Qiagen). Post-library QC was performed on the 4200 Tapestation using High Sensitivity D1000 ScreenTapes (Agilent). Paired-end sequencing was performed on an Illumina NovaSeq 6000 platform using an S2 or S4 flow cell for a total read length of 2 x 150 bp. Paired-end sequencing was performed on bisulfite-treated gDNA and cfDNA. Tissue samples were sequenced by Macrogen on a HiSeq X (Illumina). Read pairs were processed through our alignment and methylation calling pipeline, using Brabham Bioinformatics' Bismark alignment software (Krueger et al., Bioinformatics 27, 1571-1572 (2011)) for read mapping and methylation assessment. Sequencing of cfDNA was performed at USC on an Illumina NovaSeq 6000 using an S2 chip. All reads were mapped to hg19.
[0150] DMR calling. Differentially methylated regions (DMRs) were assessed using the DMR caller Metilene (Juhling et al., Genome Res 26, 256-262 (2016)). Tumor samples were compared with their patient-matched adjacent normal samples when possible. For tumor samples without matched normal tissue, a pool of normal samples from other patients with the same diagnosis was used. For unique cancer types without matched normal samples, a pool of all normal samples was used. DMRs were filtered to include those with a Mann-Whitney p-value <0.05 and |Δβ| >0.15. Regions derived from chromosomes X and Y were removed. DMRs were annotated by the name of the nearest gene. Overlapping DMRs were defined as regions in which DMRs from two or more samples shared at least three CpGs with the same orientation (hypermethylated vs. hypomethylated).
[0151] Analysis of partially methylated / hypomethylated regions. Partially methylated regions (PMDs) are hypomethylated regions with intermediate levels of methylation spanning several kilobases to several megabases. PMDs were called using MethPipe (Song et al. PLoS One 8, e81148 (2013)) based on methylation and coverage information from Bismark. The MethPipe code was modified to generate additional metrics: the average methylation level among each PMD and the standard deviation of beta values within each PMD call. Analysis and plotting of PMD regions were performed in R.
[0152] CpG minimal differentially methylated regions. We identified a set of consensus DMRs with significant enrichment of hypomethylated and hypermethylated DMRs across all samples. Each CpG position was scored by the number of samples with DMRs overlapping this locus (separately for hypermethylated and hypomethylated DMRs). These CpGs were filtered to include those with a count of 22 or more samples (70% of samples), and filtered CpGs within 500 bases of each other were combined to form distinct regions called minimally differentially methylated regions (mDMRs).
[0153] Copy number analysis. Copy number variations (CNVs) were called for WGBS data using the R package QDNAseq (Scheinin et al., Genome Res 24, 2022-2032 (2014)). This method uses read counts within fixed-size bins. We chose a 30 kb bin size for genome-wide copy number variation assessment and used 1 kb, 5 kb, or 10 kb bin sizes for focal amplifications, such as MYCN upregulation in neuroblastoma.
[0154] Ingenuity Pathway Analysis. To investigate the biological pathways associated with DMRs, we used Ingenuity Pathway Analysis (IPA, Qiagen). DMRs were annotated by the nearest gene and the distance to the nearest gene. The resulting gene list was used as input for IPA. We performed pathway enrichment analysis using default IPA settings.
[0155] Machine learning classifier. The mean beta value matrix was used to build a random forest model using a ranger. We used repeated (n=10) 3-fold cross-validation to estimate the performance of all models. For classification of TCGA solid tumors, we used a single model built from the beta values of all pediatric solid tissues.
[0156] Hybridization probe capture. Hybridization probe capture, assay design, sequencing, and bioinformatics analysis were performed as previously reported (Buckley et al., NAR Genom Bioinform 4, lqac099(2022); Buckley et al. Clin Cancer Res, 2023 Dec 15;29(24):5196-5206).
[0157] While the above specific embodiments have been described with reference to disclosed embodiments and examples, these embodiments are illustrative only and do not limit the scope of the present invention. Changes and modifications can be made within the ordinary skill of one of ordinary skill in the art without departing from the broader aspects of the present invention as defined in the following claims.
[0158] All publications, patents, and patent documents are incorporated herein by reference, as if individually incorporated by reference. No limitations inconsistent with the present disclosure should be construed therefrom. The present invention has been described with reference to various specific and preferred embodiments and techniques. However, it should be understood that many variations and modifications are possible within the spirit and scope of the invention.
Claims
1. 1. A method for determining whether a subject has childhood cancer, adult cancer, or minimal residual disease (MRD), comprising: a) training a machine learning model to detect the pediatric cancer, the adult cancer, or the MRD, wherein the machine learning model is trained using a first set of target regions from a plurality of cancer samples and corresponding target regions from non-cancerous samples, the plurality of cancer samples comprising at least two different cancer types, and the machine learning model is configured to identify the pediatric cancer, the adult cancer, or the MRD based on methylation patterns of the first set of target regions in the plurality of cancer samples compared to methylation patterns of the corresponding target regions in the non-cancerous samples; b) determining the methylation pattern of a second set of target regions of a deoxyribonucleic acid (DNA) sample obtained from said subject; c) applying the trained machine learning model to the second set of methylation patterns of target regions of the DNA obtained from the subject; and d) determining whether the subject has pediatric cancer, adult cancer, or MRD based on the output of the machine learning model; A method comprising:
2. 2. The method of claim 1, wherein the methylation patterns of the first set of target regions, the second set of target regions, and the corresponding target regions are determined using DNA methylation analysis, wherein the DNA methylation analysis comprises one or more of whole genome bisulfite sequencing (WGBS), reduced representation bisulfite sequencing (RRBS), targeted bisulfite sequencing, hybridization probe capture, methylation bead array, and enzymatic methyl sequence conversion.
3. 2. The method of claim 1, wherein the methylation patterns of the first set of target regions, the second set of target regions, and the corresponding target regions are determined using hybridization probe capture after whole genome bisulfite sequencing.
4. 4. The method of claim 3, wherein the hybridization probe capture comprises one or more probes that hybridize to the one or more target genomic regions, each of the one or more probes comprising ribonucleic acid or deoxyribonucleic acid, and optionally, each of the one or more probes comprising an affinity tag selected from the group consisting of biotin and streptavidin.
5. The method of claim 4 , wherein the one or more probes are immobilized on a solid support.
6. 10. The method of claim 1, wherein the DNA sample comprises cell-free DNA (cfDNA) or genomic DNA isolated from tissues or cells.
7. 7. The method of claim 6, wherein the cfDNA is extracted from one or more of whole blood, plasma, serum, urine, peritoneal fluid, cerebrospinal fluid, and aspirate.
8. 2. The method of claim 1, wherein the pediatric cancer or the adult cancer is osteosarcoma, medulloblastoma, meningioma, optic glioma, brainstem glioma, oligodendroglioma, ganglioglioma, pineal tumor, hepatoblastoma, fibrolamellar hepatocellular carcinoma, Hodgkin's lymphoma, non-Hodgkin's lymphoma, acute lymphocytic leukemia, acute myeloid leukemia, juvenile myelomonocytic leukemia, acute promyelocytic leukemia, chronic lymphocytic leukemia, chronic myeloid leukemia, diffuse pontine glioma, neuroblastoma, retinoblastoma, rhabdoid tumor, Ewing's sarcoma, rhabdomyosarcoma, embryonal rhabdomyosarcoma, fibrosarcoma, mesenchymoma, synovial sarcoma, teratoma, liposarcoma, spinal tumor, ovarian cancer, or Wilms' tumor.
9. 9. The method of claim 8, wherein the childhood cancer or the adult cancer is stage I, stage II, stage III, or stage IV cancer.
10. 2. The method of claim 1, wherein the first set of target regions, the second set of target regions, and the corresponding target regions comprise about 30% to about 50% of the target regions of Table 1.
11. 2. The method of claim 1, wherein the first set of target regions, the second set of target regions, and the corresponding target regions comprise about 50% to about 70% of the target regions of Table 1.
12. 2. The method of claim 1, wherein the first set of target regions, the second set of target regions, and the corresponding target regions comprise about 70% to about 90% of the target regions of Table 1.
13. 2. The method of claim 1, wherein the first set of target regions, the second set of target regions, and the corresponding target regions comprise about 90% to about 95% of the target regions in Table 1.
14. 2. The method of claim 1, wherein the first set of target regions, the second set of target regions, and the corresponding target regions comprise greater than about 95% of the target regions in Table 1.
15. 10. The method of claim 1, further comprising treating the pediatric or adult cancer, or the MRD in the subject, wherein the treatment comprises one or more of radiation therapy, surgery to remove the cancer, and administering a therapeutic agent to the subject.
16. 10. The method of claim 1, wherein the machine learning model comprises one or more of a random forest, a support vector machine (SVM), a neural network, a generalized linear model (GLM), a gradient boosting model (GBM), an extreme gradient boosting (XGB), and a deep learning algorithm.
17. 10. The method of claim 1, wherein the first set of target regions comprises the target regions of Table 1.
18. 18. The method of claim 17, further comprising comparing methylation levels of the first set of sequence reads of target regions of the plurality of cancer samples to methylation levels of corresponding target regions of the non-cancer samples to identify differentially methylated regions (DMRs) between the plurality of cancer samples and the non-cancer samples.
19. 20. The method of claim 18, further comprising identifying DMRs common to the plurality of cancer samples to define a set of minimally differentially methylated regions (mDMRs), wherein the mDMRs are used to train the machine learning model.
20. A kit comprising a plurality of polynucleotide probes complementary to said plurality of target regions listed in Table 1.