Method for the detection of a premalignant lesion in a subject

WO2026202413A1PCT designated stage Publication Date: 2026-10-01DEUTES KREBSFORSCHUNGSZENT STIFTUNG DES OFFENTLICHEN RECHTS +3
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2026/059170
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-28
Filing Date
2026-03-30
Publication Date
2026-10-01

Smart Images

  • Figure IMGF000017_0001_TABLE
    Figure IMGF000017_0001_TABLE
  • Figure IMGF000018_0001_TABLE
    Figure IMGF000018_0001_TABLE
  • Figure IMGF000019_0001_TABLE
    Figure IMGF000019_0001_TABLE
Patent Text Reader

Abstract

The present invention relates to an in vitro or ex vivo method for the detection of a premalignant lesion in a subject, preferably a newborn or infant or for the identification of a subject, preferably a newborn or infant being at risk of developing cancer, wherein the method comprises (a) providing a sample comprising cell-free DNA (cfDNA) of a subject, preferably a newborn or infant, (b1) determining the DNA methylation status of a multitude of genomic CpG positions in the cfDNA of the sample of (a) by methylation conversion followed by whole genome sequencing or a methylation array platform, or by direct methylation sequencing (e.g. Nanopore Sequencing), (b2) determining the copy number variation (CNV) pattern of a multitude of independent DNA segments in the cfDNA of the sample of (a) by whole genome sequencing and / or chromosome translocations at a multitude of genomic loci in the cfDNA of the sample of (a) by capture-enriched DNA sequencing and / or providing a sample comprising cell-free RNA (cfRNA) of the subject, preferably a newborn or infant and determining chromosome translocations at a multitude of loci in the cfRNA by capture-enriched RNA sequencing, (c) classifying the subject, preferably newborn or infant as having or not having a premalignant lesion or as being or as not being at risk of developing cancer based on (i) a comparison of the DNA methylation status of (b1) with the DNA methylation status of the multitude of genomic CpG positions in the DNA of a plurality of reference tumor and / or non-tumor samples that are or have been obtained by methylation conversion followed by sequencing or a methylation array platform, or by direct methylation sequencing (e.g. Nanopore Sequencing), and / or the DNA methylation status of (b1), and (ii) a comparison of the CNV pattern of (b2) with the CNV pattern of a human reference genome, and / or a plurality of reference tumor and / or non-tumor samples, and / or the chromosome translocations, and / or the CNV pattern of (b2).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] New PCT-Patent Application

[0002] Deutsches Krebsforschungszentrum; et al.

[0003] Vossius Ref.: AJ2044 PCT

[0004] Method for the detection of a premalignant lesion in a subject

[0005] The present invention relates to an in vitro or ex vivo method for the detection of a premalignant lesion in a subject, preferably a newborn or infant or for the identification of a subject, preferably a newborn or infant being at risk of developing cancer, wherein the method comprises (a) providing a sample comprising cell-free DNA (cfDNA) of a subject, preferably a newborn or infant, (bl) determining the DNA methylation status of a multitude of genomic CpG positions in the cfDNA of the sample of (a) by methylation conversion followed by whole genome sequencing or a methylation array platform, or by direct methylation sequencing (e.g. Nanopore Sequencing), (b2) determining the copy number variation (CNV) pattern of a multitude of independent DNA segments in the cfDNA of the sample of (a) by whole genome sequencing and / or chromosome translocations at a multitude of genomic loci in the cfDNA of the sample of (a) by capture-enriched DNA sequencing and / or providing a sample comprising cell-free RNA (cfRNA) of the subject, preferably a newborn or infant and determining chromosome translocations at a multitude of loci in the cfRNA by capture-enriched RNA sequencing, (c) classifying the subject, preferably newborn or infant as having or not having a premalignant lesion or as being or as not being at risk of developing cancer based on (i) a comparison of the DNA methylation status of (bl) with the DNA methylation status of the multitude of genomic CpG positions in the DNA of a plurality of reference tumor and / or non-tumor samples that are or have been obtained by methylation conversion followed by sequencing or a methylation array platform, or by direct methylation sequencing (e.g. Nanopore Sequencing), and / or the DNA methylation status of (bl), and (ii) a comparison of the CNV pattern of (b2) with the CNV pattern of a human reference genome, and / or a plurality of reference tumor and / or non-tumor samples, and / or the chromosome translocations, and / or the CNV pattern of (b2).

[0006] In this specification, a number of documents including patent applications and manufacturer's manuals are cited. The disclosure of these documents, while not considered relevant for the patentability of this invention, is herewith incorporated by reference in its entirety. More specifically, all documents are incorporated by reference to the same extent as if each individual document was specifically and individually indicated to be incorporated by reference.Cell-free fetal DNA has been found in the plasma of pregnant women. Its detection has made non-invasive prenatal testing, most notably for chromosomal aneuploidies, a clinical reality. Following organ transplantation, donor-derived DNA from the transplanted organs has been detected in the plasma of the recipients and has been used for monitoring graft rejection. Tumor-derived DNA has been found in the plasma of cancer patients, offering the possibility of performing "liquid biopsy" for cancer assessment and monitoring (Sun eat al. (2015), PNAS, 112 (40) E5503-E5512).

[0007] This circulating tumor DNA analysis holds promise for early detection of cancer. However, costs and complexity of genomic methods, and logistical requirements for blood collection and processing limit large scale implementation (Stephens et al. (2024) medRxiv, 12.22.24319522 doi: https: / / doi.org / 10.1101 / 2024.12.22.24319522).

[0008] Circulating tumor DNA analysis has shown promising results for cancer detection and response monitoring during treatment. Early work in this field relied on deep molecular analysis of a few genomic loci with known or recurrent somatic genomic alterations and methylation changes. More recently, new approaches have traded high sequencing depth at a few loci with wide breadth of genomic coverage at low depth. Low-depth whole genome sequencing methods analyze variation in genome-wide features such as sequencing coverage (driven by copy number alterations), mutation signatures, fragmentomic features, and nucleotide frequencies surrounding fragment ends. Incorporating several thousand or more loci from across the genome into these features enables high accuracy while reducing requirements of sequencing depth, amount of input DNA for sequencing library preparation, and sequencing costs. However, unlike somatic single nucleotide variants, differences in fragmentation features are not inherently tumor specific, making them more susceptible to pre-analytical sources of variation. Wide adoption of plasma DNA fragmentation analysis for cancer detection and monitoring is hindered by stringent upstream requirements for collection, processing, and storage of blood samples. In resource-limited environments, particularly low- and middle-income countries (LMICs), the costs associated with phlebotomy, sample processing to isolate plasma, storage, and shipping while maintaining an effective cold chain can be prohibitive. These challenges reduce the likelihood that patients will receive molecular testing and perpetuate disparities in cancer detection and outcomes. An inexpensive, and logistically simpler alternative has immense potential to bring blood-based early detection of cancer to resource limited environments (Stephens et al. (2024) medRxiv, 12.22.24319522 doi: https: / / doi.org / 10.1101 / 2024.12.22.24319522).However, whether the circulating tumor DNA analysis holds promise for the detection of premalignant stages of a cancer or the detection of a risk of developing a cancer and, if yes, what technical steps are required for successful practice is a largely unexplored field that is now specifically addressed herein.

[0009] Accordingly, the present invention relates in a first aspect to a method, preferably an in vitro or ex vivo method for the detection of a premalignant lesion in a subject, preferably a newborn or infant or for the identification of a subject, preferably a newborn or infant being at risk of developing cancer, wherein the method comprises (a) providing a sample comprising cell-free DNA (cfDNA) of a subject, preferably a newborn or infant, (bl) determining the DNA methylation status of a multitude of genomic CpG positions in the cfDNA of the sample of (a) by methylation conversion followed by whole genome sequencing or a methylation array platform, or by direct methylation sequencing (e.g. Nanopore Sequencing), (b2) determining the copy number variation (CNV) pattern of a multitude of independent DNA segments in the cfDNA of the sample of (a) by whole genome sequencing and / or chromosome translocations at a multitude of genomic loci in the cfDNA of the sample of (a) by capture-enriched DNA sequencing and / or providing a sample comprising cell-free RNA (cfRNA) of the subject, preferably a newborn or infant and determining chromosome translocations at a multitude of loci in the cfRNA by capture-enriched RNA sequencing, (c) classifying the subject, preferably newborn or infant as having or not having a premalignant lesion or as being or as not being at risk of developing cancer based on (i) a comparison of the DNA methylation status of (bl) with the DNA methylation status of the multitude of genomic CpG positions in the DNA of a plurality of reference tumor and / or non-tumor samples that are or have been obtained by methylation conversion followed by sequencing or a methylation array platform, or by direct methylation sequencing (e.g. Nanopore Sequencing), and / or the DNA methylation status of (bl), and (ii) a comparison of the CNV pattern of (b2) with the CNV pattern of a human reference genome, and / or a plurality of reference tumor and / or non-tumor samples, and / or the chromosome translocations, and / or the CNV pattern of (b2).

[0010] The first aspect also relates to a method, preferably an in vitro or ex vivo method for the detection of a premalignant lesion in a subject, preferably a newborn or infant or for the identification of a subject, preferably a newborn or infant being at risk of developing cancer, wherein the method comprises (a) providing a sample comprising cell-free DNA (cfDNA) of a subject, preferably a newborn or infant, (bl) determining the DNA methylation status of a multitude of genomic CpG positions in the cfDNA of the sample of (a) by methylation conversion followed by whole genome sequencing or a methylation array platform, or by direct methylation sequencing (e.g. Nanopore Sequencing), (b2) determining the copy number variation (CNV) pattern of a multitude of independent DNA segments in the cfDNA of the sample of (a) by whole genome sequencing and / or chromosome translocations at a multitude ofgenomic loci in the cfDNA of the sample of (a) by capture-enriched DNA sequencing, (c) classifying the subject, preferably newborn or infant as having or not having a premalignant lesion or as being or as not being at risk of developing cancer based on (i) a comparison of the DNA methylation status of (bl) with the DNA methylation status of the multitude of genomic CpG positions in the DNA of a plurality of reference tumor and / or non-tumor samples that are or have been obtained by methylation conversion followed by sequencing or a methylation array platform, or by direct methylation sequencing (e.g. Nanopore Sequencing), and (ii) a comparison of the CNV pattern of (b2) with the CNV pattern of a human reference genome, and / or a plurality of reference tumor and / or non-tumor samples, and / or the chromosome translocations.

[0011] With regard to steps (a), (bl), (b2), (c)(i) and (c)(ii) of the method of the first aspect it is to be understood that (a) is before (bl) and (bl) is before (c)(i), and similarly that (a) before (b2) and (b2) is before (c) ( ii ) . The order of (bl) and (b2) can be changed or they can be carried out in parallel. The same applies to (c)(i) and (c)(ii).

[0012] An ex vivo method is a method being carried outside of a living subject. In vitro methods are performed with microorganisms, cells, or biological molecules outside their normal biological context.

[0013] The first aspect relates to a method for the detection of premalignant lesion in a subject, or for the identification of a subject being at risk of developing cancer.

[0014] A lesion is generally any damage or abnormal change in the tissue of a subject. A premalignant lesion is a lesion that is likely to become a cancer. A premalignant lesion comprises abnormal cells which are associated with an increased risk of developing into cancer. Hence, a subject having premalignant lesion has a high risk of developing a cancer in the future. For this reason, the method of the first aspect is also a method for the identification of a subject being at risk of developing cancer.

[0015] Premalignant lesions are often characterized by driver mutations (SNV or CNV) that provide a proliferative advantage for the cells but do not make the cells cancer or cancerous cells. A driver mutation can be acquired in earlier generations and become cancerous in the latter. Premalignant lesions can be acquired as early as in fetal development and drive a pre-malignant clone. The transformation of the pre-malignant clone into a malignant clone can take years and the cancer may, for example, develop at any time within the first decade of life.The subject is preferably a mammalian subject and most preferably a human subject. Also, the subject is preferably a newborn or infant and accordingly more preferably a mammalian newborn or infant and most preferably a human newborn or infant. Among the options newborn and infant a newborn is preferred, noting that in the appended examples dried blood samples of newborns (Guthrie cards) are examined as test sample.

[0016] The term "newborn" (or neonate) as used herein preferably defines a subject under 28 days of age. The term "infant" (or toddler) as used herein preferably defines a subject being older than 29 days but younger then 1 year. These preferred ages in particular apply to human newborns and infants.

[0017] Cell-free nucleic acid (cfNA) are DNA or RNA fragments (typically 50-200 bp) circulating in extracellular body fluids, such as blood, urine, and saliva. cfNA is primarily released through cell death (apoptosis / necrosis) or active secretion. cfNA can serve as crucial, non-invasive biomarkers for cancer (liquid biopsy), fetal screening, and organ transplant monitoring (see for review, for example, Kubiritova etal. (2019), Int J Mol Sci, 20(15):3662).

[0018] For the detection of chromosome translocations cfNA can be cell-free DNA or cell-free RNA. This is because chromosome translocations may result in expressed fusion proteins that can also be detected at the RNA level as RNA transcripts encoding the fusion proteins. Unless for the detection of chromosome translocations cfNA herein is cell-free DNA. Cell-free DNA (cfDNA) (also known as circulating free DNA) are degraded DNA fragments released to body fluids such as blood, urine, cerebrospinal fluid, etc. Typical sizes of cfDNA fragments reflect chromatosome particles (average size about 166bp), as well as multiples of nucleosomes, which protect DNA from digestion by nucleases. As used herein the term "about" means with increasing preference ±20%, ±10% and ±5%. The DNA from cancer cells and - as shown herein - also from premalignant lesions cell gets released by cell-death, secretion or other mechanisms still not known into body fluids in the form of cfDNA, so that cfDNA is suitable sample for the detection of not only cancer but - as show herein - also for the detection of premalignant lesions and therefore also the detection of the risk for developing cancer. It is to be understood that the cfDNA sample comprises DNA from the premalignant lesion as well as DNA from normal tissue and cells. In general, only a small portion of the cfDNA is derived from the premalignant lesion and the vast majority of the cfDNA is derived from normal tissue and cells. The sample comprising cfDNA comprises preferably a cfDNA fraction of at least 5%. This applies mutatis mutandis to the The sample comprising cfRNA. The ctDNA fraction (ctDNA%) is the proportion of tumor-derived cfDNA of the total cDNA in the sample.The cfDNA as used herein is preferably blood cfDNA. Means and methods for the preparation of plasma as well as plasma cfDNA from plasma are known in the art. Protocols and reagents for the preparation of plasma and plasma cfDNA are available, for example, from Qiagen and Roche.

[0019] DNA methylation in cancer plays a variety of roles and DNA methylation contributes to change a healthy cell by misregulation of normal gene expression into a cancer cells. One of the most widely studied DNA methylation dysregulation is the promoter hypermethylation, where the CPGs islands in the promoter region are methylated contributing or causing gene expression to be reduced or silenced. Even if silencing of a gene is initiated by another mechanism, this often is followed by methylation of CpG sites in the promoter CpG island to stabilize the silencing of the gene. On the other hand, also hypomethylation of CpG islands in promoters can result in gene over-expression. Hypomethylation of CpG islands may also drive cancer in the case of tumor oncogenes. In cancers, loss of expression of genes occurs about 10 times more frequently by hypermethylation of promoter CpG islands than by mutations.

[0020] In view of this important role of DNA methylation of CpG islands in the development of cancer in step (bl) of the method of the first aspect the DNA methylation status of a multitude of genomic CpG positions in the cfDNA of the sample is determined.

[0021] The DNA methylation status is determined by methylation conversion followed by whole genome sequencing or a methylation array platform, or by direct methylation sequencing (e.g. Nanopore Sequencing).

[0022] The methylation conversion is preferably by enzymatic conversion as will be described herein below (utilizing the enzymes Tet-methylcytosin-dioxygenase 2 (TET2) and apolipoprotein B mRNA editing enzyme, catalytic polypeptide (APOBEC)) but can also be by bisulfite conversion. Treatment of DNA with sodium bisulfite converts unmethylated cytosine to uracil, which is subsequently converted to thymine during PCR amplification, while methylcytosine remains unchanged. Whole genome sequencing (WGS) is the process of determining the entirety, or nearly the entirety, of the DNA sequence of an organism's genome at a single time. Whole-genome sequencing can detect single nucleotide variants, insertions / deletions, copy number changes, and large structural variants.

[0023] Whole genome bisulfite sequencing (WGBS) has been the gold standard technique for base resolution analysis of DNA methylation for the last 15 years (Olova and Andrews (2025), Methods Mol Biol, 2866:73-98).Instead of bisulfite conversion followed by whole genome sequencing also bisulfite conversion followed by a methylation array platform can be used. Illumina first provided the HumanMethylation27 (27k) array, which measured methylation at approximately 27,000 CpGs, primarily in gene promoters. Meanwhile also a 450k array and even 850k array are available.

[0024] Direct methylation sequencing is a bisulfite-free technique using enzymatic modification of DNA for direct and accurate methylation mapping (Abakir and Reis, Nature Chemical Biology, 19:932-933). The direct methylation sequencing is preferably based on the nanopore technology. Nanopore technology can detect methylation directly from native DNA — without the need for a conversion step. This capability means that DNA methylation and canonical base sequence data can be produced simultaneously (Si eat al. (2023), J. Neuroimmunol., 381:578134).

[0025] The plurality of reference tumor and / or non-tumor samples is preferably a plurality of reference tumor and non-tumor samples. As reference samples for the comparison of the DNA methylation status preferably publicly available methylation EPIC array data Geo accession: GSE314261, preferably in the version of Mar 12, 2026 (https: / / www.ncbi. nlm.nih.gov / geo / query / acc.cgi?acc=GSE314261) from 5013 cancer survivor and / or the CSF methylation sequencing data from 60 patients as published in Smith et al. (2026), Nature Cancer, https: / / doi.org / 10.1038 / s43018-026-01115-4 is / are used. Accordingly, the plurality of reference tumor and / or non-tumor samples is with increasing preference at least 10, at least 20, alt least 30, at least 50 and at least 60 reference tumor and / or non-tumor samples (i.e. tumor samples, non-tumor samples or both).

[0026] A copy number variation (abbreviated CNV) refers to a circumstance in which the number of copies of a specific segment of DNA varies among different individuals' genomes. These structural differences may have come about through duplications, deletions or other changes and can affect long stretches of DNA. CNVs are preferably deletions, duplications or insertions larger than 50 base pairs. CNVs are a key characteristic of tumor development and progression. The accumulation of various CNVs during tumor development plays a critical role in driving tumor evolution (Wang et al. (2023), Brief Bioinform, 24(6):bbad341. doi: 10.1093 / bib / bbad341). Patterns (or signatures) of CNV alterations are indicative for cancer and may even predict certain cancer types or cancer stages (Steele et al. (2022), Nature, 606:984-991).

[0027] A chromosomal translocation is defined as a genome abnormality in which a chromosome breaks and either the whole or a portion of it reattaches to a different chromosome. Chromosomal translocations are considered as the primary cause for many cancers including lymphoma, leukemia and some solidtumors; see, for review, Nambiar et al. (2018), Biochim Biophys Acta, 1786(2):139-52. For example, the t(9;14)(pl3;q32) chromosomal translocation is associated with lymphoplasmacytoid lymphoma. Non-limiting but preferred examples of chromosomal translocations and their associated cancer or tumor type are provided in the following table.

[0028] Chromosomal translocations

[0029] Cancer or tumor type (Gene Fusions)

[0030] B-ALL MYB-PLAGL1

[0031] B-ALL NUP214-ABL1

[0032] B-ALL PCM1-JAK2

[0033] B-ALL TCF7-SPI1

[0034] B-ALL TCF3-PBX

[0035] B-ALL MLL-AF4

[0036] B-ALL PAX5-ETV6 / PML / JAK2

[0037] B-ALL BCR / RANBP2 / SNX2 / NUP24 / ZMIZ / ETV6 / RCSD-ABL

[0038] B-ALL TCF3-HLF

[0039] B-ALL E2A-PRL

[0040] B-ALL ABCD3 / KMT2E

[0041] B-ALL ACINl-NUTMl

[0042] B-ALL AFF1-KMT2A / NUTM1 / RAD51B

[0043] B-ALL APBB1IP-SCARB1

[0044] B-ALL ARHGDIB-ETV6

[0045] B-ALL ARL11-RNASEH2B

[0046] B-ALL ATF7IP-PDGFRB

[0047] B-ALL ATP5I-SLK

[0048] B-ALL BCR-ABL1 / JAK2

[0049] B-ALL CBFA2T3-GALNS

[0050] B-ALL CCL4L1,CCL4L2-MYO19

[0051] B-ALL CDKN2A-MTAP

[0052] B-ALL CENPC-ABL1

[0053] B-ALL CEP192-LDHB

[0054] B-ALL CLIC5-SUPT3H

[0055] B-ALL CRAMP1-CLCN7

[0056] B-ALL CRCP-PIK3AP1

[0057] B-ALL CREBBP-DEXI / MEFV

[0058] B-ALL CUX1-AFP

[0059] B-ALL CUX1-NUTM1 / ZNF398

[0060] B-ALL CXCR4-TSC22D3

[0061] B-ALL DBX1-PAX5

[0062] B-ALL DCP2-HIF1A

[0063] B-ALL DDX6-FOXR1

[0064] B-ALL DENND1B-ZCCHC7

[0065]

[0066] B-ALL DHX9-NPLB-ALL E2F4-RPL14

[0067] B-ALL ELM01-NKX2-1

[0068] B-ALL ENTPD4-PSD3

[0069] B-ALL EP300-ZNF384

[0070] B-ALL EPB41L2-OSTM1

[0071] B-ALL ERG-DUX4 / DYRK1A

[0072] B-ALL ESYT2-ATRN

[0073] B-ALL ETS2-ETV6

[0074] ETV6-ABL1 / BCL2L14 / CDK2AP1 / ELMO1 / NA / PDGFRB / QSOX1 / RNU6- B-ALL 19P / RUNX1 / EBF1 / BORCS5 / CDKN1B / CREBBP / NID1 / PMEL / PTPRO B-ALL FAMlllA-MYB

[0075] B-ALL FBRSL1-PAX5

[0076] B-ALL FIGNL1-IKZF1

[0077] B-ALL FLT1-STAU1

[0078] B-ALL FOXJ2-MEF2D

[0079] B-ALL FOXK2-CLIP2

[0080] B-ALL GATAD2A-LYN

[0081] B-ALL GGNBP2-MYO19

[0082] B-ALL HBS1L-AHI1

[0083] B-ALL HM13-TSPAN31

[0084] B-ALL H00K3-FGFR1

[0085] B-ALL IGH-CEBPD / CRLF2 / DCP2 / EPOR / DUX4

[0086] B-ALL IGK-ETV6

[0087] B-ALL IKZF1-CDK2 / SETD5 / TRPV2

[0088] B-ALL IKZF1-ETV6 / NUTM1 / ZEB2

[0089] B-ALL INO80-EXD1

[0090] B-ALL ITM2B-RB1

[0091] B-ALL ITSN1-KCNE2

[0092] B-ALL KMT2A-AFF1 / MLLT1

[0093] B-ALL KMT2D-SEC31A

[0094] B-ALL LDLRAD4-PHACTR3

[0095] B-ALL LPAR6-ITM2B

[0096] B-ALL MAML3-CCNG2

[0097] B-ALL MBNL1-PAX5

[0098] B-ALL MED12-HOXA9

[0099] B-ALL MED17-LM01

[0100] B-ALL MEF2D-BCL9 / HNRNPUL1 / SS18

[0101] B-ALL MEF2D-FOXJ2

[0102] B-ALL MLLT10-EED

[0103] B-ALL MSH6-ETV6

[0104] B-ALL NCOA6-PRDM14

[0105] B-ALL NDUFA9-DYRK4

[0106] B-ALL NGEF-INPP5D

[0107] B-ALL NONO-TFE3

[0108]

[0109] B-ALL NUP93-SUGCTB-ALL NXF1-GANAB

[0110] B-ALL OGA-RNLS

[0111] B-ALL P2RY8-CRLF2

[0112] PAX5-CBFA2T3 / CBFA2T2 / DACH1 / DACH2 / DMRTA2 / ETV6 / FKBP15 / JAK2 / MBNL1 / NOL4L / RHOXF2 / TAF3 / ZCCHC7 / ZNF276 / C20orfll2 / BCOR / TIV1PRSS9 / B-ALL FBRSL1 / GRM3 / NCOA5 / NOL4L / AUTS2 / P2RY8 / RNF38 / TMEIV114B / ZNF521 / ESRRB B-ALL PDZD7-ETV6

[0113] B-ALL PIM3-SCO2

[0114] B-ALL POLB-LRRC6

[0115] B-ALL PRRC2C-PRLR

[0116] B-ALL PTEN-EXTL2

[0117] B-ALL RAF1-TMEM40

[0118] B-ALL RAG1-HAND1

[0119] B-ALL RFX3-JAK2

[0120] B-ALL RIPOR2-VPS54

[0121] B-ALL RNF138-RNF125

[0122] B-ALL RNF168-SMCO1

[0123] B-ALL RSCD1-ABL2

[0124] B-ALL RSRP1-TMEM50A

[0125] RUNX1-ARHGDIB / ATF7IP / BAZ2A / ETS2 / MRPL47 / PCNT / DDX47 / ETV6 /

[0126] B-ALL FAM117A / KRR1 / LIPA / MLYCD / PDZD7 / PTHLH / UGGT2 / ZADH2DSCAM / DYRK1A B-ALL SENP7-ETV6

[0127] B-ALL SETD2-CCDC12

[0128] B-ALL SETD5-IKZF1

[0129] B-ALL SF1-SF3B1

[0130] B-ALL SF3A2-FNBP1

[0131] B-ALL SIN3A-ANKRD18B

[0132] B-ALL SLC12A6-NUTM1

[0133] B-ALL SLC12A7-PDCD6

[0134] B-ALL SMAD2-ZBTB7C

[0135] B-ALL SMARCA2-VLDLR / ZNF384

[0136] B-ALL SMARCA4-DHPS

[0137] B-ALL SSBP2-CSF1R / JAK2

[0138] B-ALL STIM2-IKZF1

[0139] B-ALL TAF15-ZNF384

[0140] B-ALL TBC1D15-RAB21

[0141] B-ALL TBL1XR1-CSF1R

[0142] B-ALL TCF3-HLF / PBX1 / ZNF384

[0143] B-ALL TFE3-NONO

[0144] B-ALL TLR1-BICD1

[0145] B-ALL UBASH3B-TRA2A

[0146] B-ALL WDFY1-SERPINE2

[0147] B-ALL XPO1-TNRC18

[0148] B-ALL ZCCHC7-PAX5

[0149]

[0150] B-ALL ZEB2-CXCR4B-ALL ZMYM5-PSPC1

[0151] B-ALL ZNF274-JAK2

[0152] B-ALL ZNF608-ARL14EPL

[0153] B-ALL ZNF618-NUTM1

[0154] B-ALL ZNF720-IGK

[0155] B-ALL ATF7IP / BCR-JAK2

[0156] B-ALL, AML ETV6-RUNX1 / NTRK3

[0157] B-ALL / Ph-like ALL EBF-PDGFRB

[0158] Alveolar

[0159] rhabdomyosarcoma ARHGEF7-ATP11A

[0160] Alveolar

[0161] rhabdomyosarcoma CCDC32-CBX3

[0162] Alveolar

[0163] rhabdomyosarcoma CIAO2A-CSNK1G1

[0164] Alveolar

[0165] rhabdomyosarcoma XRCC5-FRMD4A

[0166] Alveolar

[0167] rhabdomyosarcoma YLPM1-LIN52

[0168] Alveolar

[0169] rhabdomyosarcoma ZNF160-ZNF415

[0170] Alveolar

[0171] rhabdomyosarcoma PAX3-FOXO4

[0172] Alveolar

[0173] rhabdomyosarcoma FOXO1-FGFR1

[0174] Alveolar

[0175] rhabdomyosarcoma PAX3-NCOA2

[0176] Alveolar

[0177] rhabdomyosarcoma STAG1-ATF7

[0178] Alveolar

[0179] rhabdosarcoma PAX3 / 7-FOXO1

[0180] Alveolar soft part

[0181] sarcoma ASPSCR1-TFE3

[0182] AML DEK-NUP214

[0183] AML KMT2A (MLL) gene rearrangements AML RUNX1-RUNXT1

[0184] AML CBFB-MYH11

[0185] AML PML-RARA

[0186] AML MNX1-ETV6

[0187] AML KAT6A-CREBBP

[0188] AML NUP98-NSD1

[0189] AML RBM15-MKL1

[0190] AML CBFA2T3-GLIS2

[0191] anaplastic

[0192] astrocytoma SPECC1L-NTRK2

[0193] anaplastic

[0194]

[0195] oligodendroglioma PTPRZ1-ETV1anaplastic

[0196] pleomorphic

[0197] astrocytoma CIC-LEUTX angiocentric glioma MYB-ESR1 angiocentric glioma MYB-QKI astroblastoma EWSR1-PATZ1 astroblastoma MN1-GTSE1 B-AII(Ph-like ALL) STRN3-JAK2

[0198] B-AII(Ph-like ALL) IGH-EPOR / CRLF2 CNS EFT-CIC CIC-NUTMl CNS EFT-CIC ATXN1-NUTM1 CNS embryonal tumor CIC-LEUTX CNS

[0199] ganglioneuroblastoma MYO5A-NTRK3 CNS HGNET-MN1 MN1-BEND2 CNS HGNET-MN1 MN1-CXXC5 diffuse astrocytoma FGFR1-TACC1 diffuse astrocytoma FGFR3-TACC3 diffuse astrocytoma KIAA1549-BRAF diffuse astrocytoma CEP85L-ROS1 diffuse astrocytoma MYB-MMP16 diffuse astrocytoma MYB-PCDHGAl diffuse

[0200] leptomengineal

[0201] glioneural tumor KIAA1549-BRAF diffuse

[0202] leptomengineal

[0203] glioneural tumor TRIM33-RAF1 diffuse

[0204] oligodendroglioma FGFR1-TACC1 DIPG ETV6-NTRK3 DIPG VCL-NTRK2 DIPG BTBD1-NRTK3 DNET FGFR1-TACC1 DNET FGFR2-INA DSRCT EWSR1-WT1 ependymoma CCDC88A-ALK Ependymoma EWSR1-PLAGL1 Ependymoma GOPC-ROS1 Ependymoma YAP1-FAM118B Ependymoma ZFTA-NCOA1 Ependymoma ZFTA-NCOA2 Ependymoma ZFTA-RELA Ependymoma ZFTA-MAML2 Ependymoma ZFTA-MAML3

[0205]

[0206] Ependymoma ZFTA-YAP1Ependymoma YAP1-MAMLD1 Ependymoma PATZ1-MN1 Ependymoma PLAGL1-FOXO1 ependymoma like

[0207] tumors with

[0208] mesenchymal

[0209] differentiation

[0210] (ELTMD) ZFTA-NCOA1 ependymoma like

[0211] tumors with

[0212] mesenchymal

[0213] differentiation

[0214] (ELTMD) ZFTA-NCOA2 ependymoma like

[0215] tumors with

[0216] mesenchymal

[0217] differentiation

[0218] (ELTMD) ZFTA-RELA

[0219] ewing

[0220] sarcoma / primitive

[0221] neuroectodermal

[0222] tumor EWSR1-SMARCA5 Ewings like sarcoma CIC-DUX4

[0223] Ewings like sarcoma EWSR1- NFATc2 Ewings like sarcoma BCOR-CCNB3

[0224] Ewings sarcoma FUS-ERG

[0225] Ewings sarcoma EWSR1 / SMARCA5 Ewings sarcoma CTSB / BACE2

[0226] Ewings sarcoma EWSR1 / FUS-FL1 / ETS / ERG ganglioglioma AGK-BRAF ganglioglioma FXR1-BRAF ganglioglioma EWSR1-PATZ1 ganglioglioma FGFR2-SHTN1 ganglioglioma GNAI1-BRAF ganglioglioma KCTD16-NTRK2 ganglioglioma MACF1-BRAF ganglioglioma SETD2-ROBO1 ganglioglioma TMEM106B-braf Ganglioglioma TRIM24-NTRK2 ganglioglioma SLMAP-NTRK2 ganglioglioma TLE4-NTRK2

[0227] glial neuronal ARHGEF2-NTRK1

[0228] glial neuronal EWSR1-PATZ1

[0229] glial neuronal EWSR1-PLAGL1

[0230] glial neuronal FGFR2-CTNNA3

[0231] glial neuronal KIAA1549-BRAF

[0232]

[0233] glioblastoma ZCCHC8-ROS1glioblastoma CHTOP-NTRK1 glioblastoma TFG-MET

[0234] glioblastoma DGKB-ETV1

[0235] glioblastoma BCOR-CREBBP glioblastoma RNF213-SLC26A11

[0236] glioma BCOR-EP300

[0237] HGG EGFR-SEPT14

[0238] HGG EGFR-PSPH

[0239] HGG EGFR-HMGA2

[0240] HGG EGFR-VWC2

[0241] HGG FGFR1FGRR3-TACC1

[0242] HGG FGFR1FGRR3-TACC 3

[0243] HGG FGFR1FGRR3-BRAP

[0244] HGG FGFR1FGRR3-NBR1

[0245] HGG NTRK1NTRK2NTRK3-NFASC HGG NTRK1NTRK2NTRK3-BCAN HGG NTRK1NTRK2NTRK3-CHTOP HGG NTRK1NTRK2NTRK3-ARHGEF2 HGG NTRK1NTRK2NTRK3-GKAP HGG NTRK1NTRK2NTRK3-KCTD8 HGG NTRK1NTRK2NTRK3-NOS1AP HGG NTRK1NTRK2NTRK3-SQSTM 1 HGG NTRK1NTRK2NTRK3-TBC1D2 HGG NTRK1NTRK2NTRK3-VCAN HGG NTRK1NTRK2NTRK3-EML4 HGG MET-PTPRZ1

[0246] HGG MET-TGF

[0247] HGG MET-CLIP2

[0248] HGG MET-CAPZA2

[0249] HGG MET-ST7

[0250] HGG MET-TPR

[0251] HGG ROSl-FIG(GOPC)

[0252] HGG AGBL4-NTRK2

[0253] HGG ARHGEF2-NTRK1

[0254] HGG ATG7-RAF1

[0255] HGG CLIP2-MET

[0256] HGG KTN1-ALK

[0257] HGG EML4-ALK

[0258] HGG ETV6-NTRK3

[0259] HGG FGFR1-TACC1

[0260] HGG FGFR3-TACC3

[0261] HGG GOPC-ROS1

[0262] HGG NAB2-STAT6

[0263] HGG NOS1AP-NTRK1

[0264]

[0265] HGG PPP1CB-ALKHGG PTPRZ1-ETV1 HGG PTPRZ1-MET HGG TRIM24-BRAF HGG EML4-NTRK3 highly infiltrative

[0266] fibroblastic spindle

[0267] cell neoplasm RBPMS-NRTK3 Homogeneous type of

[0268] well-differentiated rhabdomyosarcoma SRF-FOXO HPC NAB2-STAT6 IHG HIP1-ALK

[0269] IHG MAD1L1-ALK IHG MAP2-ALK IHG MSI2-ALK

[0270] IHG PRKAR2A-ALK IHG SPECC1L-ALK IHG ZC3H7A-ALK IHG CHCHD3-ROS1 IHG AGBL4-NTRK2 IHG CCDC88A-ALK IHG CLIP2-MET IHG EML4-ALK

[0271] IHG ETV6-NTRK3 IHG GOPC-ROS1 IHG KIF5B-NTRK2 IHG KCTD16-NTRK2 IHG NOS1AP-NTRK1 IHG PPP1CB-ALK IHG TPM3-NTRK1 IHG MN1-PATZ1 inflammatory

[0272] myofibroblastic tumor KIF5B-ALK intracranial

[0273] angiomatoid Fibrous

[0274] Histiocytoma EWSR1-CREB intracranial myxoid

[0275] mesenchymal tumor EWSR1-CREB intracranial myxoid

[0276] mesenchymal tumor EWSR1-CREM intracranial

[0277] Rhabdomyosarcoma PAX3-NCOA1 LGG SPTBN1-ALK LGG KLC1-ROS1 LGG ETV6-NTRK3 LGG FGFR3-TACC3

[0278]

[0279] LGG GOPC-ROS1LGG PPP1CB-ALK

[0280] LGG TPM3-NTRK1

[0281] LGG TRIM24-BRAF

[0282] Low-grade

[0283] fibromyxoid sarcoma -Renal EWSR1-CREB3L Medulloblastoma DNAJB6 / SHH Medulloblastoma LCLAT1 / ERBB4 Medulloblastoma MLLT6 / MRPL45 Medulloblastoma ACOT7-ZFP69 Medulloblastoma ASAP2-MGMT Medulloblastoma BRK1-THADA Medulloblastoma CTPS1-NFYC Medulloblastoma DDX58-TOPORS-AS1 Medulloblastoma EBNA1BP2-RRAGC Medulloblastoma EIF2AK1-COL28A1 Medulloblastoma EIF2AK1-OGDH Medulloblastoma ERRFI1-CUBN Medulloblastoma ESCO1-CNDP2 Medulloblastoma FAM171A1-FAM188A Medulloblastoma FAM188A-SIL1 Medulloblastoma FAM208B-MACF1 Medulloblastoma HNRNPA2B1-SLCO1A2 Medulloblastoma KIAA1324L-TMEM243 Medulloblastoma MGMT-C8orf34 Medulloblastoma NDUFS1-SLC25A51 Medulloblastoma NOM1-LOC100128264 Medulloblastoma PRKCZ-CUBN Medulloblastoma PRKCZ-FAM171A1 Medulloblastoma RERE-ERI3 Medulloblastoma RSU1-CAMTA1 Medulloblastoma VPS53-NXN Medulloblastoma ZNF208-NRG2 Medulloblastoma CARKD-ADPRHL1 Medulloblastoma MYO16-IRS2 Medulloblastoma CCDC32-CDK4 Medulloblastoma CDH4-MACROD2 Medulloblastoma DDX3X-SULTC2P Medulloblastoma DNAL-ZNF385D Medulloblastoma GLI2-NYAP2 / EPB4L5 Medulloblastoma GMEB2-CDH4 Medulloblastoma GPATCH8-NF Medulloblastoma KLHL29-LMAN2L / LOC0929237 Medulloblastoma LCLAT-ERBB4

[0284]

[0285] Medulloblastoma MARCKSL-PIK3CDMedulloblastoma PHC-ING4

[0286] Medulloblastoma PTEN-THAP9 / NR2F-AS Medulloblastoma PVT-LI NC00964 / LOC09274 / ZC3H3 Medulloblastoma RAB7A-DCC

[0287] Medulloblastoma RUFY2-TACR2 Medulloblastoma SUFU-CYP7AOS Medulloblastoma TCF4-ROCK

[0288] Medulloblastoma TP53-BCL6B

[0289] Medulloblastoma UNC9-MYO8A Medulloblastoma ZNF365-CCFC02B

[0290] meningioma TFG-ROS1

[0291] meningioma NAB2-STAT6

[0292] meningioma YAP1-FAM118B

[0293] meningioma YAP1-MAML2

[0294] NBS-HGG TPM3-NTRK1

[0295] neuro epithelial

[0296] tumor SPECC1L-NTRK2 Neuroblastoma ACER3-CALCOCO2 Neuroblastoma ATPA3-RABAC

[0297] Neuroblastoma CACNG3-PRKCB Neuroblastoma CRY2-MAPK8IP

[0298] Neuroblastoma CYB5R-ADIPOR

[0299] Neuroblastoma DCTN2-DDIT3

[0300] Neuroblastoma DDX-MYCNUT

[0301] Neuroblastoma E2F-NECAB3

[0302] Neuroblastoma ENO-EDARADD

[0303] Neuroblastoma FAIM2-BCDIN3D Neuroblastoma HDAC3-DIAPH

[0304] Neuroblastoma HDAC6-ERAS

[0305] Neuroblastoma KDM5A-SLC6A3 Neuroblastoma MAML-SQSTM

[0306] Neuroblastoma MYCNUT-MYCN Neuroblastoma NRL-DHRS4

[0307] Neuroblastoma NTRK-PEAR

[0308] Neuroblastoma PTRF-STAT3

[0309] Neuroblastoma PXMP4-E2F

[0310] Neuroblastoma RASA4-POLR2J4 Neuroblastoma RHBDD2-POR

[0311] Neuroblastoma RIPK-SERPINB9

[0312] Neuroblastoma SLC29A-HSP90AB Neuroblastoma ST3GAL-NDRG

[0313] Neuroblastoma SEMA3D-ABCB1 Neuroblastoma SORD-NT5DC1

[0314] Neuroblastoma SSBP3-ACOT11

[0315]

[0316] Neuroblastoma ATRX-PASD1Neuroblastoma EWSR1-FLI1 Neuroblastoma RNF121-TRIM37 Neuroblastoma GLIS2-PHF21A Neuroblastoma LSAMP-STAG1 Neuroblastoma PRELID2-MAPK9 Neuroblastoma RETREG1-CDH18 Neuroblastoma CTNND2-TRIO Neuroblastoma ANKH-TENM2 Neuroblastoma EXOC3-ERGIC1 Neuroblastoma FAM104A-C170RF80 Neuroblastoma WNK4-ZNF787 Neuroblastoma MYCN-GULP1 Neuroblastoma NBAS-BAZ2B / CCNT2 Neuroblastoma CYRIA-RIF1 Neuroblastoma LRP1B-GULP1 Neuroblastoma MYLK-TXK Neuroblastoma NDUFA12-USP44 Neuroblastoma KMT2A-FOXR1 Neuroblastoma PAFAH1B2-FOXR1 Neuroblastoma TCF25-TUBB3 Neuroblastoma TNK2-AC24944.5 Neuroblastoma TOMM40-APOE Neuroblastoma VEGFB-FKBP2 Neuroblastoma RASA4B-POLR2J 2 / 3 / 4 neurofibroma SETD2-ROBO1 oligoastrocytoma NAV1-NTRK2 oligodendroglioma FAM131B-BRAF oligodendroglioma FGFR2-CTNNA3 oligodendroglioma FGFR2-INA oligodendroglioma FGFR3-TACC3 oligodendroglioma KIAA1549-BRAF oligodendroglioma FGFR3-PHGDH oligodendroglioma QKI-RAF1 oligodendroglioma MYB-MAML2 Osteosarcoma ANOO-GPDL Osteosarcoma ASCC3-LRPPRC Osteosarcoma NSUN6 / PITPNC1 Osteosarcoma PGS1 / DNAH17 Osteosarcoma PPP1R18 / FXR2 Osteosarcoma PXMP2 / GPR133 Osteosarcoma SNUPN / KIF1C Osteosarcoma TAF2 / COL14A1 Osteosarcoma TP53 / ZNF565 Osteosarcoma ZNF565 / CATSPERD

[0317]

[0318] Osteosarcoma COLEC-FAM20A-ADARB2Osteosarcoma FAM05B-EIF Osteosarcoma FBXL4-AFF3 Osteosarcoma FRYL-TEC Osteosarcoma GYPB-PTPRM Osteosarcoma VSTM4-DNAH2 Osteosarcoma COLA2-MMP4 Osteosarcoma FUS-LGMN Osteosarcoma DHFR-PABPC Osteosarcoma PTMA-NPM Osteosarcoma PMP22-EVOVL5 Osteosarcoma TP53-PPRAD Osteosarcoma PLZNA2-RUNX Osteosarcoma EEFA-VIM Osteosarcoma COLA-HSP90AB Osteosarcoma CREB3L-HNRPA papillary glioneuronal

[0319] tumor SLC44A1-PRKCA pilocytic astrocytoma CLCN6-BRAF pilocytic astrocytoma GIT2-BRAF pilocytic astrocytoma GTF2I-BRAF pilocytic astrocytoma MKRN1-BRAF pilocytic astrocytoma RNF130-BRAF pilocytic astrocytoma FYCO1-RAF1 pilocytic astrocytoma NFIA-RAF1 pilocytic astrocytoma SRGAP3-RAF1 pilocytic astrocytoma KANK1-NTRK2 pilocytic astrocytoma NACC2-NTRK2 pilocytic astrocytoma QKI-NTRK2 pilocytic astrocytoma FAM131B-BRAF pilocytic astrocytoma FGFR1-TACC1 pilocytic astrocytoma GNAI1-BRAF pilocytic astrocytoma KIAA1549-BRAF pilocytic astrocytoma PTPRZ1-ETV1 pilocytic astrocytoma PTPRZ1-MET pilocytic astrocytoma QKI-RAF1 pilocytic astrocytoma TRIM24-NTRK2 pleomorphic

[0320] xanthoastrocytoma NRF1-BRAF Pleomorphic

[0321] xanthoastrocytoma ETV6-NTRK2 pleomorphic

[0322] xanthoastrocytoma FGFR2-CLIP2 pleomorphic

[0323] xanthoastrocytoma ATG7-RAF1 pleomorphic

[0324]

[0325] xanthoastrocytoma ETV6-NTRK3pleomorphic

[0326] xanthoastrocytoma TMEM106B-braf pleomorphic

[0327] xanthoastrocytoma TPM3-NTRK1

[0328] PLNTY FGFR2-CTNNA3

[0329] PLNTY FGFR2-SHTN1

[0330] PLNTY FGFR3-TACC3

[0331] primary epidural

[0332] spinal sarcoma CIC-DUX4

[0333] Renal carcinoma EWSR1-KLF5

[0334] Renal carcinoma RBMX-TFE3

[0335] Renal carcinoma SMARCBl fusions Retinoblastoma LSG-TMEM44 Retinoblastoma CIRBP-C9orf24

[0336] SFT NAB2-STAT6

[0337] Synovial sarcoma SS18-SSX1

[0338] T-ALL BCLB-TLX3

[0339] T-ALL BG20338-ETV6

[0340] T-ALL CCND3-STIL

[0341] T-ALL GNPTAB-ATF7IP

[0342] T-ALL IKZF-NOTCH

[0343] T-ALL KMT2A (MLL)-multiple partners T-ALL PICALM-MLLTO

[0344] T-ALL SLC38A2-ABL2

[0345] T-ALL SPI-BCLB

[0346] T-ALL TAL2-TCRB

[0347] T-ALL TCRA-TLX

[0348] T-ALL TRB-AHI

[0349] T-ALL ZBTB6-ABL

[0350] T-ALL ETV6-NCOA2 / INO80D

[0351] T-ALL MYB-TCRB / AHI

[0352] T-ALL NKX2-1-TCRA / TCRD

[0353] T-ALL NUP214-ABL / SQSTM

[0354] T-ALL AHI1-TRB

[0355] T-ALL CBL-KMT2A

[0356] T-ALL CCND2-TRA

[0357] T-ALL CREBBP-TRAP1

[0358] T-ALL EVL-NKX2-1

[0359] T-ALL EVL-SFTA3

[0360] T-ALL IKZF1-NOTCH1

[0361] T-ALL KMT2A-CBL / M LLTl / M LLT4

[0362] T-ALL LMOl-TRAC

[0363] T-ALL LMO2-TRAC

[0364] T-ALL MBNL1-ANXA3

[0365]

[0366] T-ALL MLLT10-PICALMT-ALL NKX2-5-BCL11B

[0367] T-ALL NUP98-VRK1

[0368] T-ALL PICALM-MLLT10

[0369] T-ALL SET-NUP214

[0370] T-ALL STIL-TAL1

[0371] T-ALL TRB-AHI1

[0372] T-ALL ZBTB16-ABL1

[0373] T-ALL ACTR1O-GALC

[0374] T-ALL ASXL2-PFN4

[0375] T-ALL BCL2L1-TTLL9

[0376] T-ALL BCL7A-DAZAP1

[0377] T-ALL BPTF-NUTM1

[0378] T-ALL BRAP-BLMH

[0379] T-ALL CD99-JAK2

[0380] T-ALL CDK6-NKX2-1 / TLX3

[0381] T-ALL CDKN2A-TRA

[0382] T-ALL CEP164-EMP3

[0383] T-ALL CSTF3-LMO2

[0384] T-ALL DHX9-TAL1

[0385] T-ALL DNM2-HEPACAM2

[0386] T-ALL EDF1-NOTCH1

[0387] T-ALL ENTPD5-AHI1

[0388] T-ALL ERCC6L2-SF3B1 / TGFB1

[0389] T-ALL ETV6-CTNNB1 / ABL1

[0390] T-ALL FAM117B-SEC62

[0391] T-ALL FOXJ3-LMO2

[0392] T-ALL GET4-TNRC18

[0393] T-ALL GN PAT-TALI

[0394] T-ALL GTPBP1O-BARD1

[0395] T-ALL HELLS-CSMDl

[0396] T-ALL HOXA10-TRB / POLR2E

[0397] T-ALL IL9R-VAMP7

[0398] T-ALL KCTD1-SPTBN1

[0399] T-ALL KIF3B-TRB

[0400] T-ALL KMT2A-M LLT10 / CT45A3 / ELL T-ALL LEF1-TAF8

[0401] T-ALL LMAN2-PAPOLA / NSD1

[0402] T-ALL LMOl-TRA

[0403] T-ALL LYL1-TRB

[0404] T-ALL MAP1A-NUP188

[0405] T-ALL MBNL1-ABL1

[0406] T-ALL M LLT1O-CAPS2 / FAM 171A1 T-ALL MLLT6-KMT2A

[0407] T-ALL MPLKIP-H1-5

[0408]

[0409] T-ALL MTREX-SIK3T-ALL MYB-AHIl / BDPl / CHMPlA / PLAGLl T-ALL NAP1L1-MLLT10

[0410] T-ALL NCOR1-NFIA

[0411] T-ALL NIN-UNC79

[0412] T-ALL NKX2-1-TRA / BCL11B / DIO2

[0413] T-ALL NONO-ZFP36L2

[0414] T-ALL NRBP2-TMEM245

[0415] T-ALL NUP98-CCDC28A / LNP1 / PSIP1

[0416] T-ALL PARG-BMSl

[0417] T-ALL PCBP2-PRRC2B

[0418] T-ALL PELI2-C14ORF39

[0419] T-ALL PIAS4-ZBTB7A

[0420] T-ALL POFUT1-HCK

[0421] T-ALL PPP1R12C-CIC

[0422] T-ALL PPP1R37-PPP6R1

[0423] T-ALL PPP2R5C-SLC38A6

[0424] T-ALL PPP6R3-INTS4

[0425] T-ALL PRKCH-ARMH4

[0426] T-ALL RBM11-LIPI

[0427] T-ALL RNF144A-LAPTM4A

[0428] T-ALL RTRAF-STXBP6

[0429] T-ALL SEC24B-MANBA

[0430] T-ALL SETD1A-NLRC3

[0431] T-ALL SFPQ-ZFP36L2

[0432] T-ALL SLC12A9-MYB

[0433] T-ALL STMN1-SPI1

[0434] T-ALL SUSD6-RBM25

[0435] T-ALL TAL1-TRA / STIL

[0436] T-ALL TAL2-TRB

[0437] T-ALL TBL1XR1-COPB2

[0438] T-ALL TFG-ADGRG7

[0439] T-ALL TGFB1-TMEM91

[0440] T-ALL TLX1-TRA / TRB

[0441] T-ALL TLX3-BCL11B / CDK6

[0442] T-ALL TMEM18-ASXL1

[0443] T-ALL TRA-CEP85L / NKX2-1

[0444] T-ALL TRA-LMO2

[0445] T-ALL TRA-TLX1

[0446] T-ALL TRB-HOXA10 / IL2RB

[0447] T-ALL TRB-MYC

[0448] T-ALL TTC7B-SYNE2

[0449] T-ALL USP7-TRA

[0450] T-ALL UTRN-ENTPD5 / POLE4

[0451] T-ALL VILL-ABCA2

[0452]

[0453] T-ALL WDR76-SERF2T-ALL ZC3HAV1-ABL2 / AKAP11 T-ALL ZCCHC8-RSRC2

[0454] T-ALL ZFP36L2-ZFP36L2

[0455] T-ALL CBFA2T3-ACSF3

[0456] T-ALL CEP43-FGFR1

[0457] T-ALL CREBRF-AK7

[0458] T-ALL CTCF-SLC7A6

[0459] T-ALL HDAC4-GHR

[0460] T-ALL KMT2A-MLLT6

[0461] T-ALL KMT2A-TNRC18

[0462] T-ALL NF1-PSMD11

[0463] T-ALL NIPBL-ETV6

[0464] T-ALL PUM1-H0XA11

[0465] T-ALL RUNX1-AFF3

[0466] T-ALL RUNX1-EVX1

[0467] T-ALL SEC16A-NOTCH1

[0468] T-ALL SFPQ-ZFP36L1

[0469] T-ALL VANGL2-UFC1

[0470] T-ALL NUP98-VRK / RAPGDS T-ALL RUNX-AFF3 / EVX

[0471] T-ALL TAL1-TCRA / TCRD

[0472] T-ALL TCRB-TAL2 / TLX

[0473] T-ALL TLX-TCRA / TCRB

[0474] T-ALL TLX3-TCRD / BCLB

[0475] T-ALL DDX3X / NAPL-MLLT0 T-ALL EVL / TRD-NKX2- T-ALL LMO-TCRA / TCRD

[0476] T-ALL LMO2-TCRA / TCRB

[0477] T-ALL MBNL / STIL-TALl

[0478] T-ALL MLLTO-NAPL

[0479] T-ALL NDST2-RUNX

[0480] T-ALL SET / SQSTM-NUP24 T-ALL EEFSEC-PDHX

[0481] T-ALL EEFSEC-PDHX

[0482] T-ALL MELK-SIK3

[0483] T-ALL MELK-SIK3

[0484] T-ALL MYB-PLAGL1

[0485] T-ALL MYH9-JAK2

[0486] T-ALL MYH9-JAK2

[0487] T-ALL NUP214-ABL1

[0488] T-ALL PCM1-JAK2

[0489] T-ALL PPP4R3A-IGH

[0490] T-ALL PPP4R3A-IGH

[0491] T-ALL TCF7-SPI1

[0492]

[0493] T-ALL TRA-SALL2T-ALL TRA-SALL2

[0494] T-ALL TRD-NKX2-1

[0495] T-ALL TRD-NKX2-1

[0496] T-ALL ZFP36L2-TRA

[0497]

[0498] T-ALL ZFP36L2-TRA

[0499] The chromosomal translocations and / or the cancer or tumor type of the premalignant lesion or the cancer or tumor of the risk of developing cancer are preferably in accordance with the above table. For example, "PAX5-ETV6 / PI\ / IL / JAK2" mean a translocation between (or fusion gene generated from) the gene PAX5 and the gene ETV6 or PML or JAK2.

[0500] The occurrence of a specific CNV pattern and / or a specific chromosomal translocation may also indicate the kind of cancer a subject being at risk for cancer will likely develop or can predict the kind of cancer into which a premalignant lesion is likely to develop.

[0501] For this reason, in step (b2) of the first aspect the CNV pattern of a multitude of independent DNA segments in the cfDNA of the sample of (a) and / or chromosome translocations at a multitude of genomic loci in the cfDNA of the sample of (a) or at a multitude of loci in the cfRNA of the subject are determined. The cfDNA and the cfRNA are from the same subject; i.e. the subject in which a premalignant lesion or at risk of developing cancer is to be detected by the method of the first aspect of the invention.

[0502] The CNV pattern is determined by whole genome sequencing. Whole genome sequencing has been described herein above in connection with step (bl). It is important to understand that for step (b2) no methylation conversion is carried out but only whole genome sequencing to identify the CNV pattern. With regard to the order of steps (bl) and (b2) it is to be understood that they can be done consecutively ((bl) before (b2), or (b2) before (bl)) or in parallel. For instance, the CNV pattern can be identified after methylation conversion, but methylation conversion is not necessary.

[0503] The chromosome translocations are determined by capture-enriched DNA sequencing in cfDNA sample and by capture-enriched RNA sequencing in cfRNA sample. Capture-enriched DNA sequencing is a method for the high throughput mapping of chromosomal rearrangements (Olivera et al. (2011), Immunol Methods.; 375(1-2):176-181). Target enrichment works, for example, by capturing genomic regions of interest by hybridization to target-specific biotinylated probes, which are then isolated by magnetic pulldown. Thereby a subset of genes or regions of the genome being of interest can be enriched and sequenced. These genes or regions of the genome being of interest are the genes orregions of the genome being involved in the chromosome translocations that are known to be associated with certain cancer types.

[0504] In step (c) of the first aspect the subject is finally classified as having or not having a premalignant lesion or as being or as not being at risk of developing cancer.

[0505] The classification may be based on (i) the DNA methylation status of (bl) and / or a comparison of the DNA methylation status of (bl) with the DNA methylation status of the multitude of genomic CpG positions in the DNA of a plurality of reference tumor and / or non-tumor samples that are or have been obtained by methylation conversion followed by sequencing or a methylation array platform, or by direct methylation sequencing.

[0506] It is to be understood that if the DNA methylation status of (bl) was determined by methylation conversion followed by sequencing then also in step (c)(i) the DNA methylation status of the multitude of genomic CpG positions in the DNA of a plurality of reference tumor and / or non-tumor samples that are or have been obtained by methylation conversion followed by sequencing. The same applies mutatis mutandis to methylation conversion followed by a methylation array platform, and direct methylation sequencing.

[0507] The classification may also be based on (ii) the CNV pattern of (b2) and / or a comparison of the CNV pattern of (b2) with the CNV pattern of a human reference genome, and / or a plurality of reference tumor and / or non-tumor samples, and / or the chromosome translocations.

[0508] In the case of a CNV pattern of a human reference genome it is to be understood thar reference genome is from a healthy or diseased subject. CNV data on human reference genome can be generated de novo but can as be taken from a public data base, e.g., the ensembl genome browser. The plurality of reference tumor and / or non-tumor samples can be taken from tissue or plasma sample libraries. Preferred source will be discussed provided herein below.

[0509] It can be taken from the appended examples herein below that it was surprisingly found that cfDNA can not only be used as in the prior art to detect cancer in a subject but can also be used to detect a premalignant lesion or risk of developing cancer in the subject. It was also found that this can only be done with high accuracy based on the (1) DNA methylation of CpGs as first diagnostic marker, and (2) CNV pattern and / or chromosome translocations as second diagnostic marker. It is of note that the first diagnostic marker is an epigenetic marker and the second diagnostic marker is a genetic markerinvolving chromosomal aberrations. Only the combination of both diagnostic markers provides high accuracy. The specificity is increased, and the number of true positives is increased. This can be taken from Figure 4, wherein it is shown that a combined CNV + methylation model provides superior sensitivity for cancer detection in particular in low ctDNA contexts without requiring a matched tumor tissue methylome. It is yet further of note that the comparison in step (c) may be made based on a plurality of reference tumor and / or non-tumor samples and not based on the DNA methylation status of the multitude of genomic CpG positions in the DNA of a plurality of reference premalignant lesion and / or non- premalignant lesion samples. The combination of both diagnostic markers was previously used for cancer detection (e.g. Wei et al. (2022), Briefings in Bioinformatics, 23(4):1-11 and Chan et al. (2013), PNAS, 110 (47) 18761-18768) but not for the detection of premalignant lesions. In addition, it was found that the diagnosis of a disease state (premalignant lesion) is possible based on reference subjects that do not have the same disease state (cancer / non-cancer). It was surprising and could not have been expected beforehand that the method of the first aspect works in the claimed particular format and is able to identify with high accuracy a subject having a premalignant lesion or being at risk of developing cancer.

[0510] In accordance with a preferred embodiment of the first aspect in step (c)(i) the subject, preferably newborn or infant is classified as being or as not being at risk of developing cancer based on a methylation-based predictive algorithm for tumors that is a deep neural network-based model that has been generated from a plurality of reference tumor samples with varying degrees of tumor purity and CpG sparsity by

[0511] (1) simulating the tumor purity in cfDNA in the reference tumor samples by spiking in methylation profiles from non-tumor sources of cfDNA,

[0512] (2) addressing CpG sparsity in the deep neural network-based model by a network-based regression diffusion model to impute methylation values uncaptured by sequencing, thereby obtaining an imputation network for the deep neural network-based model,

[0513] (3) training the imputation network on the multitude of genomic CpG positions of step (bl), and (4) enhancing the tumor signal and classification probabilities of the deep neural network-based model by beta regressing methylation signatures from non-malignant cell types to reduce normal DNA methylation contamination from tumor-specific DNA methylation.

[0514] As explained herein above the cfDNA for the subject as provided in step (a) comprises (if any) only small amounts of DNA from the premalignant lesion and large amounts of DNA from normal tissue and cells.Step (1) takes into account and balances for this composition of the sample comprising cfDNA from the subject by spiking in methylation profiles from non-tumor sources of cfDNA into the reference tumor samples.

[0515] Step (2) preferably comprises the following imputation. Observed methylation beta values per CpG are input into a pre-learned network of CpG correlations. Missing CpG beta values are predicted based on the proximity to observed values within the network as implemented in a heat diffusion regression with boundary constraints between 0 and 1.

[0516] In step (3) the imputation network is trained on the multitude of genomic CpG positions of step (bl). The training process of a neural network is generally an iterative process in which the calculations are carried out forward and backward through each layer in the network until the loss function is minimized. The trained network can generate a response for the testing data depending upon the trained network parameters.

[0517] The beta regressing in step (4) preferably comprises to logit transform imputed methylation beta values per CpG to inverse the beta distribution of methylation values and then to fit them to a linear regression model. Transformed beta values are then modeled with a panel of non-malignant reference methylation profiles. Resulting residuals are subtracted from the observed logit transformed beta values before using a logistic sigmoid function to inverse the logit transformation resulting in purified beta values per CpG. This reduces normal DNA methylation contamination from tumor-specific DNA methylation.

[0518] The approach according to the preferred embodiment of the first aspect is illustrated by the appended examples. In particular reference is made to the M-PACT procedure as outlined in Figure 3.

[0519] In accordance with a more preferred embodiment of the first aspect steps (3) and (4) are carried out for three independent sets of a multitude of genomic CpG positions thereby obtaining three pre-deep neural network-based model and layering the three models together into the final deep neural network-based model.

[0520] The classification in accordance with the above more preferred embodiment is preferably done by generating from both, the network diffusion imputation and beta regression purification methods, a set of CpG methylation beta values and multiply them by the weights of the three independent multilayer perceptron deep neural networks the compose an ensemble model (referred to as M-PACTin the examples herein below) to calculate probability scores for each tumor class. The highest probability score amongst all three models represents the overall classification of the subject.

[0521] In accordance with a preferred embodiment of the first aspect, the DNA methylation status of a multitude of genomic CpG positions is determined by enzymatic methyl sequencing, wherein nondestructive enzymatic reactions, utilizing the enzymes Tet-methylcytosin-dioxygenase 2 (TET2) and apolipoprotein B mRNA editing enzyme, catalytic polypeptide (APOBEC) convert unmethylated cytosines to uracils, followed by CpG methylation profiling at single-CpG-site level based on a methylation array platform.

[0522] TET2 then oxidizes 5-mC and 5-hmC, providing protection from deamination by APOBEC in the next step. In contrast, unmodified cytosines are deaminated to uracils. Hence, the use of TET2 of APOBEC is an alternative approach for the above-described bisulfite treatment. Methylation array platforms from Illumina that allow CpG methylation profiling at single-CpG-site level have been described herein above. These methylation array platforms may also used in connection with the above preferred embodiment.

[0523] In accordance with a further preferred embodiment of the first aspect, wherein the multitude of genomic CpG positions comprises at least 1,000,000 CpG positions.

[0524] The implementation of at least 1,000,000 CpG positions ensures genome wide coverage and therefore the very reliable identification of hyper- of hypomethylated CpG islands.

[0525] In accordance with another preferred embodiment of the first aspect, in (b2) the copy number variation (CNV) pattern is determined by low-coverage whole genome sequencing and identification of the CNV pattern in the sequencing data by aligning the sequence reads to the human reference genome (preferably GRCh37).

[0526] Compared with deep sequencing strategies, low-coverage whole-genome sequencing (WGS) produces a mere fraction of the data per sample and relies on computational methods to fill in the missing information. Low-coverage whole-genome sequencing is therefore cheaper and faster as compared to deep sequencing strategies. CNV analysis using low-coverage WGS is still very efficient and was shown to outperform array-based analysis (Zhou et al. (2018), J Med Genet; 55(ll):735-743).In accordance with a more preferred embodiment of the first aspect, the method further comprises determining from the sequencing data the cfDNA fragment lengths of genomic segments in the cfDNA of the sample of (a) and preferably the ratio of short fragments of 100-150 bps to long fragments of 151-250 bp, thereby estimating the amount of circulating tumor DNA in the sample.

[0527] Cristiano et al. (2019), Nature. 2019, 570:385-389 established fragmentation profiles being defined as the ratio of short to long fragments as means for estimating the amount of circulating tumor DNA in the sample. It is assumed for the above more preferred embodiment of the first aspect that in silica size selection can enhance the signal of circulating tumor DNA in the sample or the signal of DNA from a premalignant lesion.

[0528] In accordance with a preferred embodiment of the first aspect, the cancer is any cancer arising in young people until the age of 30 years or tumor types that are not associated with carcinogenesis through chronic mutagen exposure and is preferably selected from hematopoietic and lymphoid tumors (such as leukemias, acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), juvenile myelomonocytic leukemia (JMML), chronic myeloid leukemia (CML), other specified leukemias, lymphomas, Hodgkin lymphoma, Non-Hodgkin lymphomas (e.g., Burkitt lymphoma, diffuse large B-cell lymphoma, lymphoblastic lymphoma, anaplastic large cell lymphoma), primary central nervous system (CNS) lymphomas); central nervous system tumors (such as embryonal CNS tumors, edulloblastoma, atypical teratoid / rhabdoid tumor (AT / RT), CNS embryonal tumor NOS (not otherwise specified), gliomas, pilocytic astrocytoma, diffuse midline glioma, diffuse intrinsic pontine glioma (DIPG), glioblastoma, other CNS tumors, ependymoma, craniopharyngioma, choroid plexus tumors, neuronal and mixed neuronal-glial tumors (e.g., ganglioglioma, dysembryoplastic neuroepithelial tumor); embryonal tumors (such as neuroblastoma, nephroblastoma (Wilms tumor), hepatoblastoma, retinoblastoma) soft tissue and bone tumors (soft tissue sarcomas, rhabdomyosarcoma, nonrhabdomyosarcoma soft tissue sarcomas (NRSTS) (e.g., synovial sarcoma, fibrosarcoma, malignant peripheral nerve sheath tumor), infantile fibrosarcoma, bone tumors, osteosarcoma, Ewing sarcoma, chondrosarcoma); germ cell tumors (such as gonadal germ cell tumors, testicular germ cell tumors, ovarian germ cell tumors, extragonadal germ cell tumors, sacrococcygeal teratomas, mediastinal germ cell tumors, intracranial germ cell tumors); and other pediatric tumors (such as Langerhans cell histiocytosis, hemophagocytic lymphohistiocytosis, melanotic neuroectodermal tumor of infancy, infantile hemangioma, plexiform neurofibroma).The preferred examples of the cancer are non-limiting, are preferred example of cancer arising in young people until the age of 30 years or tumor types that are not associated with carcinogenesis through chronic mutagen exposure.

[0529] In accordance with a further preferred embodiment of the first aspect, the sample is a blood, plasma, or serum sample of the subject, preferably newborn or infant.

[0530] As discussed above, cfDNA is usually isolated from plasma. Plasma can in turn be isolated from blood and serum.

[0531] In accordance with a more preferred embodiment of the first aspect, the sample is a dried blood sample, preferably a dried blood spot on filter paper and most preferably dried blood of a newborn on a Guthrie card.

[0532] Dried blood spot (DBS) testing is a form of bio-sampling where blood samples are blotted and dried on filter paper. DBS are advantageous for several reasons. For example, the dried samples can be obtained at low costs, without highly qualified staff, and can easily be shipped without special conditions for storage or transportation to an analytical laboratory and analysed using various methods such as DNA amplification or HPLC. Subjects can even collect their sample at home at the time that suits them. DBS testing is exemplified herein by the data in Figure 5. Here, the so-called "Ewing fusion" (EWS-FLI1 fusion protein i.e. a reciprocal translocation between chromosomes 11 and 22, t(ll,22)) that can be found in about 85% of the cases of the pediatric cancer Ewing sarcoma is detected. When the sample is a dried blood sample, preferably a dried blood spot on filter paper and most preferably dried blood of a newborn on a Guthrie card step (b2) preferably comprises of consists of determining chromosome translocations at a multitude of genomic loci in the cfDNA of the sample of (a) by capture-enriched DNA sequencing.

[0533] In the examples herein below the samples for the subject to be diagnosed are dried blood of a newborn on a Guthrie card. The Guthrie card is generally done for newborns in clinical practice in the fist days of life because it is a simple way to check if the newborn has one of 9 or more rare but serious conditions (cystic fibrosis (CF), sickle cell disease (SCD), congenital hypothyroidism (CHT), phenylketonuria (PKU), medium-chain acyl-CoA dehydrogenase deficiency (MCADD), maple syrup urine disease (MSUD), isovaleric acidaemia (IVA), glutaric aciduria type 1 (GAI), homocystinuria (HCU)). Hence for the vast majority of newborns blood on a Guthrie card is available.In line with the use of newborns blood on a Guthrie card as test sample also as the reference samples were dried blood spots on filter paper from pediatric oncology patients and healthy volunteers. Hence, in case the test sample is a dried blood spot on filter paper and most preferably dried blood of a newborn on a Guthrie card also the reference samples are preferably based on dried blood spots on filter paper from tumor and / or non-tumor subjects.

[0534] In accordance with an even more preferred embodiment of the first aspect, the method comprises in step (a) isolating cfDNA from a dried blood sample by

[0535] (i) recovering the cfDNA of the dried blood sample under denaturing conditions by a protease K-containing buffer, wherein the buffer preferably comprises 2.5-10% sodium dodecyl sulfate and about 20 pl proteinase K,

[0536] (ii) adding a buffer comprising a chaotropic salt, preferably 25-50% guanidinium chloride to the lysate of (i) comprising a carrier RNA,

[0537] (iii) adding ethanol to the lysate of (ii),

[0538] (iv) applying the lysate of (iii) to a silica-membrane-based anion exchange resin, wherein the DNA binds to the membrane,

[0539] (v) optionally washing the membrane to residual contaminants, such as RNA, proteins, and low-molecular-weight impurities by one or more washing steps, wherein in the one or more washing steps preferably a solution of 50-100% guanidinium chloride and a solution of 0.1% sodium azide are used as washing buffers,

[0540] (vi) optionally washing the membrane with ethanol,

[0541] (vii) eluting the cfDNA from the membrane by distilled water or a buffer comprising 2.5-10% sodium dodecyl sulfate.

[0542] The above steps for isolating cfDNA from a dried blood sample are based on the based on the "Protocol: Isolation of Total DNA from FTA and Guthrie Cards" from the Qiagen handbook (version January 2020; file: / / / C: / downloads / HB-0355-005_HB_QA_DNA_lnvestigator_0120_WW%20(l).pdf)

[0543] The buffer of the above step (i) is the Buffer ATL, of step (ii) the Buffer AL and of step (v) the buffers AW1 and AW2. The buffer ingredients have been made public by the document "LABORATORY SAFETY GUIDELINE Qiagen Kits", Revision Date: 2 / 28 / 2020; Copyright © 2018-2019 The President and Fellows of Harvard College (https: / / www.ehs.harvard.edu / sites / default / files / lab_safety_guideline_qiagen_kits.pdf). The buffer ingredients are indicated in the above even more preferred embodiment of the first aspect.In the above step (i) is also possible to use or 0.1%-10% sodium dodecyl sulfate and about 20 pl proteinase K (so called "KiTZ-SP-Buffer", and then to use in step (ii) 25-50% guanidinium thiocyanate chloride to the lysate of (i) comprising a carrier RNA (so called "KiTZ-GT-Buffer"). Both, guanidinium thiocyanate and SDS improve the isolation capacities of the KITZ buffer. The KiTZ buffer preferably comprises about 100 mM Tris-HCI, about 200 mM NaCI, about 5mM EDTA and about 1% Triton.

[0544] While also other methods for the isolating cf NA from a dried blood sample exist and work in connection with the method of the first aspect, the above steps for isolating cfNA from a dried blood sample (Qiagen or KiTZ method) are preferred since they resulted in superior cfNA quality for the method of the first aspect as compared to other tested methods.

[0545] In accordance with a preferred embodiment of the first aspect, the subject is a newborn or infant and the sample has been obtained from the newborn or infant within 1 year, preferably within 1 month, more preferably within 20 days and most preferably within about 10 days after birth.

[0546] As discussed herein above newborns and infants are preferably 1 year or below. For the case where the sample been obtained within 20 days or about 10 days after birth the sample been obtained for a newborn.

[0547] In accordance with another preferred embodiment of the first aspect, determining the CNV pattern comprises determining CNVs that involve gains, deletions, and / or rearrangements of chromosomes, chromosome arms, or focal genome segments.

[0548] According to this preferred embodiment the most important classes of CNVs are analyzed.

[0549] In accordance with another preferred embodiment of the first aspect, in step (c) a classification rule is used for the classification.

[0550] A classification rule or classifier is generally a procedure by which the elements of the population set are each predicted to belong to one of the classes. In the simplest form the classification rule as used for the first aspect the classification rule determines that the subject belongs to the class of subjects having a premalignant lesion or being a risk for developing cancer or to the class of subjects not having a premalignant lesion and not being at risk for developing cancer.In accordance with a more preferred embodiment of the first aspect, the classification rule has been obtained based on random forest analysis, neuronal network, support vector machine, K-nearest neighbors, XGBoost or similar machine learning approaches.

[0551] Random forest analysis is a way of averaging multiple deep decision trees, trained on different parts of the same training set, with the goal of reducing the variance. A neural network consists of connected units or nodes called artificial neurons, which loosely model the neurons in the brain.

[0552] A neuronal network (also artificial neural network or neural net, abbreviated ANN or NN) is a model inspired by the structure and function of biological neural networks in animal brains.

[0553] Support vector machines (SVMs, also support vector networks) are supervised max-margin models with associated learning algorithms that analyze data for classification and regression analysis. Developed at AT&T Bell Laboratories, SVMs are one of the most studied models, being based on statistical learning frameworks of VC (Vapnik and Chervonenkis) theory.

[0554] The k-nearest neighbors (KNN) algorithm is a non-parametric, supervised learning classifier, which uses proximity to make classifications or predictions about the grouping of an individual data point. It is a popular and simple classification and regression classifier used in machine learning.

[0555] XGBoost (Extreme Gradient Boosting) is based on decision trees and is an improvement of other machine learning approaches, such as random forest and gradient boosting.

[0556] A machine learning approach generally refers to the algorithms, techniques, and methods used to create models that solve real-world problems by processing data.

[0557] In accordance with a preferred embodiment of the first aspect, the reference tumor and / or non-tumor samples have been obtained from subjects that developed a pediatric cancer within the first 10 years of life and / or did not develop a pediatric cancer within the first 10 years of life.

[0558] The "and / or" is preferable "and".

[0559] This preferred embodiment is in particular preferred in combination with the other preferred embodiment requiring that the subject is a newborn or infant and the sample has been obtained from the newborn or infant within 1 year, preferably within 1 month, more preferably within 20 days andmost preferably within about 10 days after birth.

[0560] As shown in the appended examples, it was surprisingly found based on a comparison of samples from newborn with reference tumor and non-tumor samples from subjects that developed a pediatric cancer within the first 10 years of life and did not develop a pediatric cancer within the first 10 years of life, respectively, a premalignant lesion or risk of developing cancer can be detected in the newborn. Is emphasized that it was surprisingly found that a premalignant stage can be identified based on samples from malignant stages and non-malignant stages.

[0561] In accordance with a preferred embodiment of the first aspect, the reference tumor and / or non-tumor samples are from different cancer types relevant to the age group of children, adolescents and young adults or cancers that are not associated with carcinogenesis through chronic mutagen exposure, wherein the different cancer types are preferably selected from hematopoietic and lymphoid tumors (such as leukemias, acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), juvenile myelomonocytic leukemia (JMML), chronic myeloid leukemia (CML), other specified leukemias, lymphomas, Hodgkin lymphoma, Non-Hodgkin lymphomas (e.g., Burkitt lymphoma, diffuse large B-cell lymphoma, lymphoblastic lymphoma, anaplastic large cell lymphoma), primary central nervous system (CNS) lymphomas); central nervous system tumors (such as embryonal CNS tumors, edulloblastoma, atypical teratoid / rhabdoid tumor (AT / RT), CNS embryonal tumor NOS (not otherwise specified), gliomas, pilocytic astrocytoma, diffuse midline glioma, diffuse intrinsic pontine glioma (DIPG), glioblastoma, other CNS tumors, ependymoma, craniopharyngioma, choroid plexus tumors, neuronal and mixed neuronal-glial tumors (e.g., ganglioglioma, dysembryoplastic neuroepithelial tumor); embryonal tumors (such as neuroblastoma, nephroblastoma (Wilms tumor), hepatoblastoma, retinoblastoma) soft tissue and bone tumors (soft tissue sarcomas, rhabdomyosarcoma, nonrhabdomyosarcoma soft tissue sarcomas (NRSTS) (e.g., synovial sarcoma, fibrosarcoma, malignant peripheral nerve sheath tumor), infantile fibrosarcoma, bone tumors, osteosarcoma, Ewing sarcoma, chondrosarcoma); germ cell tumors (such as gonadal germ cell tumors, testicular germ cell tumors, ovarian germ cell tumors, extragonadal germ cell tumors, sacrococcygeal teratomas, mediastinal germ cell tumors, intracranial germ cell tumors); and other pediatric tumors (such as Langerhans cell histiocytosis, hemophagocytic lymphohistiocytosis, melanotic neuroectodermal tumor of infancy, infantile hemangioma, plexiform neurofibroma).

[0562] As mentioned above, the preferred examples of the cancer are non-limiting are preferred example of cancer arising in young people until the age of 30 years or tumor types that are not associated with carcinogenesis through chronic mutagen exposure.In accordance with a further preferred embodiment of the first aspect, the method is a computer-implemented method.

[0563] A computer-implemented method is a method which involves the use of a computer, computer network or other programmable apparatus, where one or more features are realised wholly or partly by means of a computer program.

[0564] In this connection it is to be understood that step (a) of the method of the invention (a) "providing a sample comprising cell-free nucleic acid (cfNA), preferably cell-free DNA (cfDNA) of a subject, preferably a newborn or infant" does generally not involve the use of a computer, computer network or other programmable apparatus, where one or more features are realised wholly or partly by means of a computer program.

[0565] On the other hand, the classifying step (c) and preferably also the determination steps (bl) and / or (b2) involve the use of a computer, computer network or other programmable apparatus, where one or more features are realised wholly or partly by means of a computer program.

[0566] In particular, the various sequencing technologies as described herein above all have in common that they are next-generation sequencing (NGS) technologies. With the advantages of next-generation sequencing (NGS), medical researchers have access to identify massive amounts of data around genomic profiles of cancers. In order to store, process and evaluate the massive amounts of data generally computers are used. For this reason, the method of the first aspect is preferably a computer-implemented method.

[0567] In accordance with another preferred embodiment of the first aspect, the multitude of genomic CpG positions in the cfDNA comprise genomic CpG positions from human chromosome llpl5.5.

[0568] Loss of methylation at chromosome llpl5.5 is common in human adult tumors (Scelfo eat al (2022), Oncogene; 21(16):2564-72) and for this reason this genomic region is of particular interest. Human chromosome band llpl5.5 houses a large cluster of genes that are imprinted. Dysregulation of this gene cluster is associated with the overgrowth and tumor predisposition syndrome, Beckwith-Wiedemann syndrome.

[0569] In accordance with a more preferred embodiment of the first aspect, the genomic CpG positions comprise genomic CpG positions in the genomic region encoding one or more of the long-coding RNAH19, the gene IGF2, the gene KCNQ1, the gene CDKN1C, the non-coding RNA KCNQ1OT1, the noncoding RNAs KCNQ1-AS1, and the non-coding RNA KCNQ1DN, wherein the long-coding RNA H19 is preferred.

[0570] The large cluster of genes on human chromosome band llpl5.5 that are imprinted comprises H19, IGF2, KCNQ1 and CDKN1C. H19 is preferred because it acts as a tumor suppressor (Yoshimizu et al. (2008), PNAS, 105(34):12417-12422).

[0571] The present invention relates in a second aspect to a computer-readable storage medium having computer-executable instructions stored or a computer program, that, when executed, causes a computer to perform a method according to the first aspect.

[0572] Non-limiting but preferred examples of computer-readable storage media are pendrives, CDs, DVDs Blue rays, flash storages, cloud storages, SSDs, floppy discs and magnetic drums.

[0573] The computer may comprise one or more of a storage; a display; an operation section; a first type interface section configured to perform a wireless communication in compliance with a first communication protocol; a second type interface section configured to perform a wireless communication in compliance with a second communication protocol being different from the first communication protocol; and a processor.

[0574] The computer-executable instructions may cause the processor of a computer to execute perform a method according to the first aspect and to store in a storage area of the storage of the computer whether the subject has a premalignant lesion or is at risk of developing cancer.

[0575] Also described herein is a computer program causing the system of the below described third aspect of the invention to execute the steps of the method of the fist aspect of the invention. Furthermore, described herein is a computer-readable storage medium having stored thereon this computer program.

[0576] In this connection it is again noted that that step (a) of the method of the invention (a) generally does not involve the use of a computer, computer network or other programmable apparatus, where one or more features are realised wholly or partly by means of a computer program, whereas the classifying step (c) and preferably also the determination steps (bl) and / or (b2) involve the use of a computer, computer network or other programmable apparatus, where one or more features arerealised wholly or partly by means of a computer program.

[0577] Accordingly, the second aspect relates in a preferred embodiment to a computer-readable storage medium or computer program for the detection of a premalignant lesion in a subject, preferably a newborn or infant or for the identification of a subject, preferably a newborn or infant being at risk of developing cancer,

[0578] wherein the computer-readable storage medium or the or computer program has stored or is configured to store,

[0579] (a) the DNA methylation status of a multitude of genomic CpG positions in cell-free DNA (cfDNA) of a sample comprising cfDNA of the subject as was obtained by methylation conversion followed by whole genome sequencing or a methylation array platform, or by direct methylation sequencing,

[0580] (b) the copy number variation (CNV) pattern of a multitude of independent DNA segments in the cfDNA of the sample of (a) as was obtained by whole genome sequencing and / or chromosome translocations at a multitude of genomic loci in the cfDNA of the sample of (a) by capture-enriched DNA sequencing, (c) the DNA methylation status of the multitude of genomic CpG positions in the DNA of a plurality of reference tumor and / or non-tumor samples that have been obtained by methylation conversion followed by sequencing or a methylation array platform, or by direct methylation sequencing, and (d) the CNV pattern of a human reference genome, and / or a plurality of reference tumor and / or nontumor samples, and

[0581] wherein the computer-readable storage has computer-executable instructions stored or the computer program, that, when executed, causes a computer to

[0582] (e) classify the subject, preferably newborn or infant as having or not having a premalignant lesion or as being or as not being at risk of developing cancer based on

[0583] (i) the DNA methylation status of (a) and / or a comparison of the DNA methylation status of (a) with the DNA methylation status of (c), and

[0584] (ii) the CNV pattern of (b) and / or a comparison of the CNV pattern of (b) with the CNV pattern of (d), and / or the chromosome translocations (b).

[0585] The second aspect relates in an alternative preferred embodiment to a computer-readable storage medium having computer-executable instructions stored or computer program for the detection of a premalignant lesion in a subject, preferably a newborn or infant or for the identification of a subject, preferably a newborn or infant being at risk of developing cancer,

[0586] wherein the computer-readable storage medium having computer-executable instructions stored or computer program, when executed, causes a computer to perform a method for the detection of a premalignant lesion in the subject, preferably a newborn or infant or for the identification of thesubject, preferably a newborn or infant being at risk of developing cancer, said method comprising: classifying the subject, preferably newborn or infant as having or not having a premalignant lesion or as being or as not being at risk of developing cancer based on

[0587] (i) a comparison of a DNA methylation status of a multitude of genomic CpG positions in cell-free DNA (cfDNA) of a sample comprising cfDNA of the subject, preferably a newborn or infant that are or have been obtained by methylation conversion followed by sequencing or a methylation array platform, or by direct methylation sequencing with a DNA methylation status of a multitude of genomic CpG positions in the DNA of a plurality of reference tumor and / or non-tumor samples that are or have been obtained by methylation conversion followed by sequencing or a methylation array platform, or by direct methylation sequencing, and / or

[0588] said DNA methylation status of the multitude of genomic CpG positions in the cfDNA of the sample, and

[0589] (ii) a comparison of a copy number variation (CNV) pattern of a multitude of independent DNA segments in the cfDNA of the sample of (i) that is or has been obtained by whole genome sequencing with the CNV pattern of a human reference genome, and / or a plurality of reference tumor and / or nontumor samples, and / or

[0590] chromosome translocations at a multitude of genomic loci in the cfDNA of the sample comprising cfDNA of the subject that are or have been obtained by capture-enriched DNA sequencing, and / or said CNV pattern of the multitude of independent DNA segments in the cfDNA of the sample of (i).

[0591] The present invention relates in a third aspect to a system for the identification of a subject, preferably a newborn or infant being at risk of developing cancer, comprising: (a) one or more processors; and (b) memory coupled to the one or more processors and comprising instructions executable by the one or more processors to perform a method according to the first aspect. Similarly, the present invention relates to a system for the identification of a subject, preferably a newborn or infant being at risk of developing cancer, wherein the system comprises one or more processors adapted to or configured to perform a method according to the first aspect.

[0592] The system is preferably a computer and / or data processing system. As the system preferably comprises one or more of a storage; a display; an operation section; a first type interface section configured to perform a wireless communication in compliance with a first communication protocol; a second type interface section configured to perform a wireless communication in compliance with a second communication protocol being different from the first communication protocol; and a processor.The third aspect relates in a preferred embodiment to a system for the identification of a subject, preferably a newborn or infant being at risk of developing cancer, comprising:

[0593] (a) one or more processors; and

[0594] (b) memory coupled to the one or more processors and comprising instructions executable by the one or more processors,

[0595] wherein the memory has stored or is configured to store,

[0596] (a) the DNA methylation status of a multitude of genomic CpG positions in the cell-free DNA (cfDNA) of a sample comprising cfDNA of the subject as was obtained by methylation conversion followed by whole genome sequencing or a methylation array platform, or by direct methylation sequencing, (b) the copy number variation (CNV) pattern of a multitude of independent DNA segments in the cfDNA of the sample of (a) as was obtained by whole genome sequencing and / or chromosome translocations at a multitude of genomic loci in the cfDNA of the sample of (a) by capture-enriched DNA sequencing, (c) the DNA methylation status of the multitude of genomic CpG positions in the DNA of a plurality of reference tumor and / or non-tumor samples that have been obtained by methylation conversion followed by sequencing or a methylation array platform, or by direct methylation sequencing, and (d) the CNV pattern of a human reference genome, and / or a plurality of reference tumor and / or nontumor samples, and

[0597] (e) classify the subject, preferably newborn or infant as having or not having a premalignant lesion or as being or as not being at risk of developing cancer based on

[0598] (i) the DNA methylation status of (a) and / or a comparison of the DNA methylation status of (a) with the DNA methylation status of (c), and

[0599] (ii) the CNV pattern of (b) and / or a comparison of the CNV pattern of (b) with the CNV pattern of (d), and / or the chromosome translocations of (b).

[0600] The third aspect relates in an alternative preferred embodiment to a system for the identification of a subject, preferably a newborn or infant being at risk of developing cancer, comprising:

[0601] (a) one or more processors; and

[0602] (b) memory coupled to the one or more processors and comprising instructions executable by the one or more processors,

[0603] wherein the system, when executed, causes a computer to perform a method for the detection of a premalignant lesion in the subject, preferably a newborn or infant or for the identification of the subject, preferably a newborn or infant being at risk of developing cancer, said method comprising: classifying the subject, preferably newborn or infant as having or not having a premalignant lesion or as being or as not being at risk of developing cancer based on

[0604] (i) a comparison of a DNA methylation status of a multitude of genomic CpG positions in cell-free DNA(cfDNA) of a sample comprising cfDNA of the subject, preferably a newborn or infant that are or have been obtained by methylation conversion followed by sequencing or a methylation array platform, or by direct methylation sequencing with a DNA methylation status of a multitude of genomic CpG positions in the DNA of a plurality of reference tumor and / or non-tumor samples that are or have been obtained by methylation conversion followed by sequencing or a methylation array platform, or by direct methylation sequencing, and / or

[0605] said DNA methylation status of the multitude of genomic CpG positions in the cfDNA of the sample, and

[0606] (ii) a comparison of a copy number variation (CNV) pattern of a multitude of independent DNA segments in the cfDNA of the sample of (i) that is or has been obtained by whole genome sequencing with the CNV pattern of a human reference genome, and / or a plurality of reference tumor and / or nontumor samples, and / or

[0607] chromosome translocations at a multitude of genomic loci in the cfDNA of the sample comprising cfDNA of the subject that are or have been obtained by capture-enriched DNA sequencing, and / or said CNV pattern of the multitude of independent DNA segments in the cfDNA of the sample of (i).

[0608] The present invention relates in another (fourth) aspect to a system for the identification of a subject, preferably a newborn or infant being at risk of developing cancer, comprising (a) one or more processors; and (b) memory coupled to the one or more processors and comprising instructions executable by the one or more processors to perform a method according to the first aspect or at least step (c) of the method, and / or, to perform a classification of the subject, preferably newborn or infant as having or not having a premalignant lesion or as being or as not being at risk of developing cancer based on

[0609] (i) a comparison of a DNA methylation status of a multitude of genomic CpG positions in a cell-free DNA (cfDNA) of the subject with the DNA methylation status of the multitude of genomic CpG positions in the DNA of a plurality of reference tumor and / or non-tumor samples that are or have been obtained by methylation conversion followed by sequencing or a methylation array platform, or by direct methylation sequencing, and / or said DNA methylation status of the multitude of genomic CpG positions in the cfDNA of the subject, and

[0610] (ii) a comparison of a copy number variation (CNV) pattern of a multitude of independent DNA segments in the cfDNA of the subject with the CNV pattern of a human reference genome, and / or a plurality of reference tumor and / or non-tumor samples, and / or the chromosome translocations ata multitude of genomic loci in the cfDNA of the subject or a multitude of loci in the cfRNA of the subject, and / orsaid CNV pattern of the multitude of independent DNA segments in the cfDNA of the subject.

[0611] The DNA methylation status of the multitude of genomic CpG positions in the cfDNA of the subject may have been determined by methylation conversion followed by whole genome sequencing or a methylation array platform, or by direct methylation sequencing, particularly preferably by step (bl) of the method according to the first aspect.

[0612] The CNV pattern of the multitude of independent DNA segments in the cfDNA of the subject may have been determined by whole genome sequencing.

[0613] The chromosome translocations at a multitude of genomic loci in the cfDNA of the subject may have been determined by capture-enriched DNA sequencing and / or chromosome translocations at a multitude of loci in the cfRNA of the subject by capture-enriched RNA sequencing, particularly preferably by step (b2) of the method according to the first aspect.

[0614] The present invention relates in another (fifth) aspect to a computer-readable storage medium having computer-executable instructions stored and / or to a computer program, that, when executed, causes a computer or the system according to the fourth aspect to perform a method according to the first aspect or to perform at least step (c) of the method according to the first aspect, and / or that particularly, when executed, causes a (or the) computer to perform a classification of a subject, preferably newborn or infant as having or not having a premalignant lesion or as being or as not being at risk of developing cancer based on

[0615] (i) a comparison of a DNA methylation status of a multitude of genomic CpG positions in cell-free DNA (cfDNA) of the subject with the DNA methylation status of the multitude of genomic CpG positions in the of a plurality of reference tumor and / or non-tumor samples that are or have been obtained by methylation conversion followed by sequencing or a methylation array platform, or by direct methylation sequencing, and / or

[0616] said DNA methylation status of the multitude of genomic CpG positions in the cfDNA of the subject, and

[0617] (ii) a comparison of a copy number variation (CNV) pattern of a multitude of independent DNA segments in the cfDNA of the subject with the CNV pattern of a human reference genome, and / or a plurality of reference tumor and / or non-tumor samples, and / or the chromosome translocations at a multitude of genomic loci in the cfDNA of the subject or a multitude of loci in the cfRNA of the subject, and / orsaid CNV pattern of the multitude of independent DNA segments in the cfDNA of the subject.

[0618] The DNA methylation status of the multitude of genomic CpG positions in the cfDNA of the subject may have been determined by methylation conversion followed by whole genome sequencing or a methylation array platform, or by direct methylation sequencing, particularly preferably by step (bl) of the method according to the first aspect.

[0619] The CNV pattern of the multitude of independent DNA segments in the cfNA of the subject may have been determined by whole genome sequencing.

[0620] The chromosome translocations at a multitude of genomic loci in the cfDNA of the subject may have been determined by capture-enriched DNA sequencing and / or chromosome translocations at a multitude of loci in the cfRNA of the subject by capture-enriched RNA sequencing, particularly preferably by step (b2) of the method according to the first aspect.

[0621] The present invention relates in another (sixth) aspect to a methylation-based predictive algorithm for tumors, particularly a deep neural network-based model, for the classification of the subject, preferably newborn or infant as being or as not being at risk of developing cancer, particularly according to the first aspect in step (c)(i). In other words, in accordance with a preferred embodiment of the first aspect in step (c)(i) the subject, preferably newborn or infant is classified as being or as not being at risk of developing cancer based on the methylation-based predictive algorithm for tumors that may be a deep neural network-based model.

[0622] The deep neural network-based model may have been generated from a plurality of reference tumor samples with varying degrees of tumor purity and CpG sparsity by (1) simulating the tumor purity in cfNA, preferably cfDNA in the reference tumor samples by spiking in methylation profiles from nontumor sources of cfNA, preferably cfDNA, (2) addressing CpG sparsity in the deep neural networkbased model by a network-based regression diffusion model to impute methylation values uncaptured by sequencing, thereby obtaining an imputation network for the deep neural network-based model, (3) training the imputation network on the multitude of genomic CpG positions of step (bl) of the method according to the first aspect, and (4) enhancing the tumor signal and classification probabilities of the deep neural network-based model by beta regressing methylation signatures from non-malignant cell types to reduce normal DNA methylation contamination from tumor-specific DNA methylation.Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skilled in the art to which this invention belongs. In case of conflict, the patent specification including definitions, will prevail.

[0623] Regarding the embodiments characterized in this specification, in particular in the claims, it is intended that each embodiment mentioned in a dependent claim is combined with each embodiment of each claim (independent or dependent) said dependent claim depends from. For example, in case of an independent claim 1 reciting 3 alternatives A, B and C, a dependent claim 2 reciting 3 alternatives D, E and F and a claim 3 depending from claims 1 and 2 and reciting 3 alternatives G, H and I, it is to be understood that the specification unambiguously discloses embodiments corresponding to combinations A, D, G; A, D, H; A, D, I; A, E, G; A, E, H; A, E, I; A, F, G; A, F, H; A, F, I; B, D, G; B, D, H; B, D, I; B, E, G; B, E, H; B, E, I; B, F, G; B, F, H; B, F, I; C, D, G; C, D, H; C, D, I; C, E, G; C, E, H; C, E, I; C, F, G; C, F, H; C, F, I, unless specifically mentioned otherwise.

[0624] Similarly, and also in those cases where independent and / or dependent claims do not recite alternatives, it is understood that if dependent claims refer back to a plurality of preceding claims, any combination of subject-matter covered thereby is considered to be explicitly disclosed. For example, in case of an independent claim 1, a dependent claim 2 referring back to claim 1, and a dependent claim 3 referring back to both claims 2 and 1, it follows that the combination of the subject-matter of claims 3 and 1 is clearly and unambiguously disclosed as is the combination of the subject-matter of claims 3, 2 and 1. In case a further dependent claim 4 is present which refers to any one of claims 1 to 3, it follows that the combination of the subject-matter of claims 4 and 1, of claims 4, 2 and 1, of claims 4, 3 and 1, as well as of claims 4, 3, 2 and 1 is clearly and unambiguously disclosed.

[0625] This also holds true for alternatives in different claims that depend from each other. Thus, if claim 1 recites three alternatives of the same category and claim 2 recites three alternatives of a different category as recited in claim 1, and refers back to claim 1, all combinations of the alternatives as recited in claims 1 and 2 are explicitly disclosed herein.

[0626] The above considerations apply mutatis mutandis to all appended claims.The figures show.

[0627] Figure 1 - Expansion of liquid biopsy technologies to DBS. A Library yields of the different cfDNA isolation protocols. B Comparison of the effect of size selection on library yield and average size postlibrary preparation. Samples were isolated with Investigator Kit. C Exemplary CNV profile and D Fragment length analysis of sarcoma bearing PDX mouse. cfDNA was isolated using Investigator Kit followed by 7:1 size selection. Statistical analysis was performed using unpaired t-test and one-way ANOVA. ****=p<0.0001

[0628] Figure 2 - Optimization of cfDNA isolation from DBS. A DNA amounts isolated from same sized DBS punches and library yield using different combinations of isolation buffers and purification columns. B Effect of storage time on DNA amount isolated with Investigator Kit. Statistical analysis was performed usingtwo-way ANOVA with Sidak's multiple comparisons test and unpaired t-test. *=p<0.05, **=p<0.01

[0629] Figure 3 - Development of a DNA methylation-based classifier for accurate tumor detection and entity classification from liquid biopsies. A Schematic workflow of the M-PACT (Methylation-based Predictive Algorithm for Tumors) pipeline. B Boxplot summarizing difference between observed and predicted p-values for a range of 200,000 to 50,000 CpGs used to impute the 850,000 CpGs which are included in the Illumina EPICvl. C UMAP projecting the n=71 samples of the non-malignant background classes, including n=18 CSF samples from non-oncological donors. D Balanced accuracy of classification performance against the number of training epochs for one of the three neural networks, which were combined for a three-layer ensemble model. E Box plot comparing simulated purity for validation in silico samples successfully and non-successfully classified by M-PACT.

[0630] Figure 4 - Performance of methylation- and CNV-based approaches across decreasing ctDNA fractions in cell-free DNA EM-seq samples. Receiver operating characteristic (ROC) curves for cancer detection at ~50%, 10%, and 5% ctDNA fractions generated from an in silico titration series of EM-seq samples. High-ctDNA (>60%) samples were computationally diluted with a pool of normal controls to the indicated fractions; all diluted samples were therefore treated as true positives. A combined CNV + methylation model provides gains at higher fractions. These results demonstrate superior sensitivity of combined CNV and methylation-based cancer detection in low ctDNA contexts without requiring a matched tumor tissue methylome.Figure 5 - Fusion Detection from Dried Blood Spot Cards a) Representative electropherogram of RNA isolated from a dried blood spot (DBS) sample demonstrating fragment distribution and RNA integrity prior to downstream analysis, b) RNA yield (ng / pL) obtained from healthy controls and pediatric tumor-derived DBS samples. Bars indicate mean ± SD. c) Amplification curves for GAPDH (housekeeping control) and Ewing fusion transcript in A673 positive control and representative Ewing PDX DBS sample. Threshold line indicates Cq determination. Corresponding mean Ct values are shown below, d) ID amplitude plot showing separation of positive and negative droplets for Ewing fusion detection in DBS samples. Positive droplets are clearly distinguishable from background signal, e) Absolute copy number per pL reaction mix for each DBS sample. Positive droplets were detected in selected PDX-derived DBS samples, whereas healthy control and NTC showed no signal.

[0631] The examples illustrate the claimed invention.

[0632] Example 1 - Material and methods

[0633] DBS Collection

[0634] DBS were collected from both pediatric oncology patients and healthy volunteers through standard venipuncture using a BD Vacutainer® Safety-Lok™ blood collection set (Becton Dickinson, Franklin Lakes, NJ, USA). Following the blood draw, the sample was carefully applied to Whatman™ filter paper (Cytiva, Buckinghamshire, UK), ensuring an even distribution across the designated collection areas. DBS provided by the Heidelberg Newborn Screening Unit were collected via heel prick with a safety lancet according to standard operating procedure for the Newborn Screening Program. DBS were stored at 4°C until further use. For DBS of PDX mice, after euthanasia up to 500 pl blood was collected via final cardiac puncture into EDTA tubes (Sarstedt, Numbrecht, Germany). The blood was then immediately applied to the filter paper aiming for an even distribution across the designated collection area. To validate the DBS analyses, blood samples were simultaneously drawn from patients using a BD Vacutainer® Safety-Lok™ blood collection set (Becton Dickinson, Franklin Lakes, NJ, USA) and collected into PAXgene Blood ccfDNA tubes (Qiagen, Hilden, Germany). Each tube was filled to capacity, inverted after collection and processed to plasma in accordance with the manufacturers' guidelines. The plasma was stored at -20°C until further use.

[0635] cfDNA Isolation

[0636] CfDNA Isolation from DSB was carried out using the QIAamp DNA Investigator Kit (Investigator Kit, Qiagen, Hilden, Germany). The manufacturer's protocol for isolation from FTA and Guthrie Cards wasfollowed including the addition of carrier RNA as per the recommendation for small amounts of starting material. CfDNA was eluted twice in 25 pl of the provided elution buffer.

[0637] For method comparison DBS cfDNA was also isolated with a buffer from the Heidelberg Newborn Screening Unit (HD Screening buffer) and a DNA extraction buffer with varying ratios of sodium dodecyl sulfate (0.1%-10%) and guanidinium thiocyanate (25-50%) (KiTZ Buffer). Samples processed with the HD Screening buffer followed the SOP for DNA isolation from the Heidelberg Newborn Screening Unit and cfDNA was eluted in 50 pl nuclease-free water.

[0638] The KiTZ buffer consists of 100 mM Tris-HCI, 200 mM NaCI, 5mM EDTA and 1% Triton and is originally established for DNA extraction from tissue. 300 pl of this buffer was used for isolation with 20 pl Proteinase K (part of Investigator Kit) and varying ratios of sodium dodecyl sulfate (0.1%-10%) and guanidinium thiocyanate (25-50%). The samples were then incubated overnight at 55°C followed by one hour at 85°C.

[0639] To further optimize cfDNA isolation from DBS, all mentioned isolation methods were also tested with additional clean up steps via the Nucleosnap (Macherey-Nagel, Duren, Germany) as well as the MinElute columns included in the Investigator Kit. For this comparative analysis, eluates were introduced in the Investigator Kit protocol at step nine by adding Ethanol for efficient column binding. To explore added value of reduced gDNA contamination, a right-side size selection with AMPure XP magnetic beads (Beckman Coulter, Brea, CA, USA) was investigated. Size selection was carried out as described previously by adding AMPure XP beads at a 1 x beads-to-sample ratio (Heider et al. 2020). The supernatant, as is necessary for right-side size selection, was transferred to a sterile 2 ml Eppendorf tube, and fresh beads were added at a 7.0 x beads-to-sample ratio. After discarding the supernatant, the size-selected cfDNA was eluted in 25 pl of nuclease-free water.

[0640] To accurately quantify and assess the quality of the extracted cfDNA, two analytical techniques were employed. The Qubit 3 fluorometer, along with the Qubit dsDNA high sensitivity assay kit (Thermo Fisher Scientific, Waltham, MA, USA), was used to determine the concentration of cfDNA by fluorescent dye binding. Prior to each analysis, the fluorometer was calibrated with two standards, and then all samples were measured to establish absolute DNA concentrations. To assess cfDNA purity and potential contamination with genomic DNA, the Agilent 2100 Bioanalyzer (Bioanalyzer, Agilent Technologies, Santa Clara, CA, USA) was used with the matching High Sensitivity DNA Kit according to the manufacturer's instructions. This system performs on-chip automated electrophoresis, distinguishing between different-sized DNA fragments. As an alternative for Bioanalyzer the Agilent 4150 or 4200 Tapestation (Tapestation, Agilent Technologies, Santa Clara, CA, USA) was used with the High Sensitivty D1000 reagents. The Tapestation is a Screen Tape-based system for the discrimination of DNA fragments of different lengths. For cfDNA quantification, fragments ranging from 150 to 500bp were considered, while genomic DNA contamination was evaluated by measuring fragments from 50 to 7.000 bp.

[0641] CfRNA Isolation and downstream analysis

[0642] Cell-free RNA (cfRNA) was isolated from dried blood spots (DBS) using the miRNeasy Micro Kit (Qiagen, Hilden, Germany) according to the manufacturer's instructions, with modifications to accommodate DBS as input material. Briefly, up to eight punches (3.2 mm diameter each) were obtained from DBS filter paper and incubated in 700 pL QIAzol lysis reagent (Qiagen, Hilden, Germany) for 60 minutes at room temperature with intermittent vortexing every 15 minutes to facilitate efficient elution of blood components from the matrix. Following lysis, chloroform was added, phase separation was performed, and the aqueous phase containing RNA was further processed according to the manufacturer's instructions. RNA was eluted in 14 pL of RNase-free water.

[0643] RNA quality control and quantification were performed using the Qubit RNA High Sensitivity (HS) Assay Kit on a Qubit 3 Fluorometer (Thermo Fisher Scientific, Waltham, MA, USA), and RNA integrity was assessed using High Sensitivity RNA ScreenTape on a Tapestation system (Agilent Technologies, Santa Clara, CA, USA).

[0644] For RNA sequencing applications, eluates were immediately frozen and stored at -80 °C until further processing. For downstream PCR-based analyses, RNA was immediately reverse transcribed using ProtoScript® II Reverse Transcriptase (New England Biolabs, Ipswich, MA, USA) according to the manufacturer's instructions to generate cDNA, which was subsequently diluted 1:5 prior to amplification.

[0645] Quantitative PCR (qPCR) was performed using SYBR Green-based chemistry (primaQUANT CYBR-Green 2x PCR Master Mix, Steinbrenner Laborsysteme GmbH, Wiesenbach, Germany) on a Quantstudio 5 Real-Time PCR System (Applied Biosystems, Thermo Fisher Scientific, Waltham, MA, USA). Amplification was carried out under the following cycling conditions: initial denaturation at 95 °C for 10 minutes, followed by 50 cycles of denaturation at 95 °C for 15 seconds and combined annealing / extension at 60 °C for 1 minute. A melt curve analysis was performed at the end of the amplification to assess product specificity. Data were analyzed using Quantstudio Design & Analysis Software (vl.6.1; Applied Biosystems, Thermo Fisher Scientific, Waltham, MA, USA).

[0646] Droplet digital PCR (ddPCR) was performed using the QX200 Droplet Digital PCR System (Bio-Rad Laboratories, Hercules, CA, USA) with QX200 ddPCR EvaGreen Supermix and Droplet Generation Oil for EvaGreen (Bio-Rad Laboratories, Hercules, CA, USA). Amplification was carried out under the following cycling conditions: initial enzyme activation at 95 °C for 5 minutes, followed by 40 cycles of denaturation at 95 °C for 30 seconds and annealing / extension at 60 °C for 60 seconds with a ramp rate of 2 °C / s, followed by 4 °C for 5 minutes and 90 °C for 5 minutes, and a final hold at 4 °C. Droplets weregenerated, thermocycled, and analyzed using QX Manager Standard Edition software (v2.3.2; Bio-Rad Laboratories, Hercules, CA, USA).

[0647] For both qPCR and ddPCR analyses, primers targeting the EWSR1::FLI1 fusion transcript were used. Primer sequences were as follows (5'— >3'):

[0648] EF1_F (EWSR1): GCCAAGCTCCAAGTCAATATAGC (SEQ ID NO: 1)

[0649] EF2_R (FLU): GAGGCCAGAATTCATGTTATTGC (SEQ ID NO: 2)

[0650] The assay was validated using DBS derived from patient-derived xenograft (PDX) samples of the A-673 cell line (RRID:CVCL_0080).

[0651] Low coverage Whole Genome Sequencing (IcWGS)

[0652] Libraries for IcWGS were prepared using the 2S™ Hyb DNA Library Kit and the 2S™ MID Adapter SI- S4 Indexing Kit (both Integrated DNATechnologies (IDT), Coralville, IA, USA), following the manufacturer's instructions. The target input amount was 500-4000 pg cfDNA per sample, as quantified by Bioanalyzer or Tapestation. If no or not sufficient cfDNA amounts were measurable, 20 pl were used as maximal input volume. Briefly, Illumina indices were ligated to the 5' and 3' ends of fragmented DNA, and smaller fragments were removed through size selection using magnetic beads. Library preparation involved 12 to 14 PCR cycles, with no template controls included in the runs. After preparation, libraries were quantified using Bioanalyzer or Tapestation and pooled in equimolar ratios for sequencing on a NovaSeq6000 or NovaSeqX (both Illumina, San Diego, California, USA) at the NGS Core Facility, DKFZ, to achieve 100 bp paired-end or 150 bp paired-end reads respectively with a 10x target coverage.

[0653] For downstream analysis, the raw sequencing data were transferred to the Omics IT and Data Management Core Facility (ODCF) at DFKZ. Using ODCF's in-house AlignmentAndQCWorkflows (vl.2.73-1) pre-processing was conducted. This workflow included adapter trimming, aligning sequences to the human reference genome (GRCh37), marking duplicate reads, sorting, and extracting the sequence quality matrix. Unique molecular index (UMI) processing was carried out with the fgbio workflow (vl.1.0, Fulcrum Genomics) and the Picard toolkit, enabling BAM file collapse and identification of high-quality, de-duplicated molecular consensus reads. These reads were subsequently re-aligned using the Burrows-Wheeler Aligner. For CNV analysis, ichorCNA (vO.3.2) was used to perform segmentation and estimate ctDN A fraction by calculating normalized Iog2 ratios based on read counts within each 1MB genomic window. The CNV profiles of tumor and plasma samples were visualized and compared in R (v3.6.0). The fragment lengths of cfDNA templates were determined from paired-end sequencing reads by utilizing the genomic locations of both ends after aligning them to the reference genome. Samtools (Dawson et al. 2013) was employed to select paired reads with fragment sizes between 50 and 150 bp. CNVs were called from size-selected samples usingichorCNA, and a Panel-of-Normal (PoN) was generated from the size-selection process. An in-house R package, cfdnakit, was developed to analyze fragment length information from low-coverage wholegenome sequencing (IcWGS) data. Reads overlapping with DUKE and DAC blacklisted regions (Amemiya et al. 2019), as well as centromeric regions, were excluded from further analysis. The short-to-long-fragment ratio (S / L Ratio) was defined as the ratio of fragments within the 100-150 bp range to those in the 151-250 bp range. The number of short and long reads was calculated per nonoverlapping genomic regions of 1 MB each. For each region, the counts of short and long fragments were corrected for GC content and mappability. The S / L Ratio for each bin was computed based on these corrected read counts, and the bin's S / L Ratio was normalized using the genome-wide median S / L Ratio (z-score). Test sample z-scores were further normalized by comparing them to process-matched PoN sample z-scores for each bin. Lastly, the bins were segmented using Circular Binary Segmentation (Madanat-Harjuoja et al. 2022), and median z-scores per segment were calculated. This package also provides visualization of fragment length distribution and z-scores with segmentation for non-overlapping windows. The R package is available at [github.com / Pitithat-pu / cfdnakit], For the quantification of tumor burden the ctDNA estimation score (CES) was calculated based on CNV status as well as cfDNA segmentation patterns taking into consideration the z-score. The CES is a further development of the CPA score (Raman et al. 2020) and was improved within the INFORM liquid biopsy study (Maass et al.).

[0654] DNA Methylation Profiling

[0655] Library preparation was performed using the NEBNext* Enzymatic Methyl-seq library Kit (EM-seq, New England Biolabs, Ipswich, MA, USA) according to manufacturer's recommendation. The target input amount was 500-2000 pg cfDNA per sample, as determined by the Bioanalyzer or Tapestation. If no or not enough cfDNA was measurable 20 pl were used as maximal input volume. EM-seq is based on enzymatic conversion of unmethylated cytosin to uracil. Briefly, libraries are prepared, followed by oxidation of 5-mC and 5-hmC protecting from deamination. In the next step the unmodified cytosines are converted to uracils, followed by amplification of the libraries. Library preparation involved 12 to 14 PCR cycles, with no template controls included in the runs. After preparation, libraries were quantified using Qubit and Bioanalyzer or Tapestation. Libraries were pooled in equimolar ratios for sequencing on NovaSeqX at the NGS Core Facility, DKFZ, to achieve 150 bp paired-end reads with a target coverage of 10x.

[0656] Genome-wide DNA methylation profiles were generated for all available tumor and liquid biopsy samples using the Infinium MethylationEPIC v2.0 Beadchip arrays, as previously described (Capper et al. 2018; Koelsche et al. 2021) at the DKFZ Array Core Facility. To reduce dimensionality and cluster tumor tissue and liquid biopsy samples, a t-distributed Stochastic Neighborhood Embedding (t-SNE)approach was applied. The t-SNE analysis was performed using the Rtsne R package and visualized with ggplot2. The conumee package (Bioconductor) was utilized to identify segmentation patterns and genome-wide CNVs (Hovestadt and Zapatka 2017). Genomic segments were classified as amplifications when the Iog2 ratio exceeded 0.07, and as deletions when the Iog2 ratio dropped below -0.074.

[0657] For unsupervised heatmap clustering of differentially methylated sites the DNA methylation data were preprocessed using the "Noob" function from the minfi package. Subsequently, the standard deviations for each beta value were computed, and the 10.000 CpGs with the highest standard deviations were chosen. These selected CpGs were visualized using the "heatmap.2" function from the gplots package.

[0658] Methylation imputation

[0659] An EPIC methylation array reference cohort of 914 samples was constructed from methylation profiles associated with CNS tumors (GSE215240) and immune cells (GSE167998) (PMID:36928815, PMID:35140201). A CpG correlation network based on the Euclidean distance between CpG beta values across all samples was developed using the scikit-network (vO.33.1) python package. The resulting network served as the foundation for a network diffusion regression model to predict beta values of missing CpGs from sparse methylomes. Imputation was validated on a cohort of 950 EPIC arrays (GSE276299) (PMID:39358389).

[0660] Neural network-based ctDNA classification

[0661] A reference methylation cohort of 5,014 methylation profiles encompassing 4,996 methylation arrays from CNS tumors (GSE90496, GSE109379, GSE215240, GSE183798, and GSE175543), control blood (GSE89278), and immune cells (GSE167998), in addition to 18 non-malignant CSF (current study), was assembled (PMID: 29539639, PMID: 36964296, PMID: 36928815, PMID: 34545084, PMID: 27822319, PM ID: 35140201). Methylation profiles were divided into 97 categories for classification. A series of Multi-Layer Perceptron (MLP) deep neural networks were trained and integrated into an ensemble classifier to predict these categories from sparse, low tumor purity methylomes. HM450 methylation array probes were ordered by importance scores via implementation of a random forest classifier using 1000 decision trees. These scores were used to partition the 450,000 CpGs into their most informative features, resulting in three sets of 100,000 CpGs with predictive value and removal of non-predictive CpGs.

[0662] The first MLP classifier in the ensemble model was trained directly on the dense methylation cohort using scikit-learn (vl.5.0). Specifically, three neural networks were trained on separate folds of CpGs, each using 100,000 CpGs utilizing two hidden layers (256 and 128) with sigmoid activation. The highestprobability score across all three models was used for subsequent class prediction. The second and third classifiers in the ensemble model were trained using a similar approach to Sturgeon29. In brief, the reference methylation cohort was used to create a simulated cohort for training. Each simulation was created by first generating a novel methylation profile from a randomly weighted average of all samples belonging to the same class (e.g. ETMR). Next, a novel non-malignant methylation profile was generated using a randomly weighted average of all samples categorized as blood, control CSF, immune cells, plasma, control reactive tissue, and control inflammatory tissue. Averaging non-malignant and tumor profiles at different proportions of weighted averages, we simulated samples harboring a range of tumor purities. Lastly, CpGs were randomly removed from simulated samples to mimic sparse methylomes. The second and third MLP classifiers were each trained on >3 million total simulations comprised of 36,000 per class utilizing a two-layer architecture with sigmoid activation. Like the first classifier, these two classifiers were comprised of three individual neural networks trained on distinct folds of 100,000 CpGs. The weighted average of each model was then scaled to sum to 1, yielding a total probability score for class prediction. This approach reduces the likelihood of confidently called false positives due to the concordance required between models. Layering the three models together, an ensemble classifier was created, with class prediction reported by the model with the highest prediction probability score.

[0663] Methylation deconvolution and regression

[0664] A reference whole genome bisulfite sequencing (WGBS) cohort of non-malignant background signatures spanning 17 different cell types (GSE186458) and non-malignant CSF (this study) was curated for methylation-based deconvolution (PMID: 36599988). A signature matrix was created by calculating the average methylation profile for each non-malignant class. Additionally, a tumor methylation profile was added to the signature matrix by averaging reference methylation array profiles. A Huber regression was implemented on logit transformed beta values from the 4,000 most variable CpGs amongst the signature matrix and those captured in the sample. Regression coefficients were scaled to sum to 1 and used as the predicted cell type proportion. The regression was conducted serially by removing all signature cell types with a predicted proportion of 0 and re-run iteratively until all cell types contained non-zero proportion predictions. The same signature matrix was used to regress out non-malignant background. A beta regression was implemented, with residuals subtracted from the sample methylation profile. All beta values below 0 were set to 0, and those above 1 were set to 1.M-PACT Methodology

[0665] Imputation. Observed methylation beta values per CpG are input into a pre-learned network of CpG correlations as determined by Euclidean distance. Missing CpG beta values are predicted based on the proximity to observed values within the network as implemented in a heat diffusion regression with boundary constraints between 0 and 1.

[0666] Purification. Imputed methylation beta values per CpG are logit transformed to inverse the beta distribution of methylation values and then fit to a linear regression model. Transformed beta values are modeled with a panel of non-malignant reference methylation profiles. Resulting residuals are subtracted from the observed logit transformed beta values before using a logistic sigmoid function to inverse the logit transformation resulting in purified beta values per CpG.

[0667] M-PACT classification. From both the network diffusion imputation and beta regression purification methods, a set of CpG methylation beta values are generated and multiplied by the weights of the three independent multilayer perceptron deep neural networks the compose the ensemble model referred to as M-PACT to calculate probability scores for each tumor class. The highest probability score amongst all three models represents the overall classification.

[0668] Example 2 - Results

[0669] Newborn screening (NBS) programs have been routinely collecting heel prick blood on filter paper, so called dried blood spots (DBS), for more than 60 years to detect a variety of congenital, treatable diseases affecting the metabolic, endocrine, blood, immune and neuromuscular systems in neonates allowing for early detection and intervention to reduce morbidity and mortality. Classically enzymes, metabolites and antibodies are detected through mass spectrometry, but also molecular genetic analyses are now used to detect, e.g., spinal muscle atrophy (Czibere et al. 2020). Recently next generation sequencing (NGS) has made genomic screening of a large number of monogenic diseases possible, that are at the moment not covered by the standard NBS, which led to several genomic NBS (gNBS) studies showing the feasibility and potential of gNBS (Holm et al. 2018; Kingsmore et al. 2022; Roman et al. 2020). Within the BabySeq project 9.4% of newborns were found to harbor genetic abnormalities associated with disease manifesting in childhood and 3.4% with adult-onset disease including cancer predispositions (Ceyhan-Birsoy et al. 2019). Besides genomic DNA, it is also possible to isolate cfDNA from DBSs exposing the potential for new approaches to liquid biopsy sample collection on the one hand (Heider et al. 2020), but also to potentially allow screening of newborns for malignancies or premalignant lesions at birth on the other hand. It has been shown that, e.g., high-riskRBI mutations can be detected through screening programs leading to diagnosis in an earlier stage and less aggressive therapies (Gerrish et al. 2020).

[0670] DBSs provide a cost-effective and easily accessible method for samples acquisition and coupled with easy long-term storage they offer the unique opportunity to establish population-based screening programs on a global scale.

[0671] Addressing the problems of limited sample volumes and decentralized sample collection in pediatric oncology, this work expanded liquid biopsy methodology to dried blood spots (DBS). By requiring minimal blood volumes and allowing for easy long-term storage and transport without refrigeration, DBS provide a scalable and practical alternative for integrating liquid biopsy techniques into pediatric cancer management. Building on their established role in newborn screening programs, DBS also offer the potential to expand tumor surveillance and early detection programs. For feasibility studies on the potential of cfDNA isolations from DBS for downstream genomic and epigenomic analyses, we explored the compatibility of cfDNA from DBSs with IcWGS and EM-Seq with subsequent assessment of sensitivity and specificity in preclinical models of pediatric cancer. Establishment of methodology using DBS was performed in a cohort of PDX-mouse models. Whole-genome analyses of cfDNA from DBS of mice carrying various pediatric tumors enabled differentiation between human tumor-derived cfDNA and mouse host-derived cfDNA.

[0672] Firstly, two cfDNA isolation protocols were compared: the QIAamp DNA Investigator Kit (IK) and a protocol from the Heidelberg Newborn Screening unit (HNS). The IK protocol produced significantly higher DNA yields post-library preparation (Figure 1A, p<0.0001), whereas the HNS protocol only provided cfDNA of inferior quality for downstream analysis. Secondly, the effect of magnetic beadbased size selection prior to library preparation was investigated. The relation of beads to sample volume determines the length of DNA fragments that are selected with 1:1 ratio favoring genomic DNA and 7:1 ratio expected to favor cfDNA. Contrary to expectations, size selection did not significantly impact the yield, and non-size-selected samples generated highest DNA amounts (Figure IB, p=0.4736). Library preparations from all tested conditions generated DNA fragments of expected postlibrary cfDNA length (Figure IB, p=0.0620). CNV analysis of a DBS of Mouse 4 bearing a synovial sarcoma showed clear copy number alterations in several chromosomes and an estimated ctDNA fraction of 0.0901 further supporting the feasibility of ctDNA analysis from DBSs. Due to the possibility of differentiating between host murine DNA and tumor human DNA, fragment length analysis was able to show that tumor derived DNA fragments were distinctly shorter with an average length of 142 bp compared to 167bp (Figure ID).Further optimization of the isolation protocol was performed using DBS samples from healthy human subjects. A range of isolation buffers, including the Investigator Kit buffer ( I K-B), HNS buffer (HNS-B), and KiTZ buffer (optimized for tissue DNA extraction), were combined with different purification columns (MinElute and Nucleosnap). The HNS buffer repeatedly failed to isolate sufficient cfDNA for analysis, while IK-B significantly outperformed all other buffers, regardless of the purification column used. (Figure 2A). The IK-B with MinElute column combination, which is part of the standard protocol of the Investigator Kit, isolated the highest DNA amount. After library preparation IK-B samples outperformed KiTZ buffer samples significantly (p=0.0191). To assess the impact of storage on cfDNA recovery, the yield from five-year-old archived DBSs from newborns was compared to freshly collected DBS from healthy adults. Interestingly, the archived DBSs produced higher cfDNA yields than the fresh samples (Figure 2B, p = 0.0421).

[0673] Prior efforts to implement DNA methylation-based tumor classification in liquid biopsies worked but have been challenged by i) low cfDNA inputs, ii) sparse CpG capture, and iii) low tumor fraction 21,25-27. To improve obstacles, we established a workflow that applies enzymatic-based methylation conversion (EM-seq), a recently introduced bisulfite-free platform for methylation sequencing 28, to sub-nanogram quantities of cfDNA and built a deep neural network-based model, M-PACT, to classify sparse, low tumor fraction samples (Figure 3A).

[0674] For classifier development, a large set of n=3, 195,000 in silico samples with varying degrees of tumor purity and CpG sparsity (ranging from 0.1 to 1.0) was generated from an extensive reference data set representing 87 CNS tumor and 10 non-tumor classes (n=5,014, Figure 3A). Tumor purity was simulated by computationally spiking in methylation profiles from non-tumor sources of physiologically expected ambient cfDNA. To address CpG sparsity, we created a network-based regression diffusion model to impute methylation values uncaptured by sequencing (Figure 3A). Our imputation network was trained on a subset of the array reference cohort (n=914), achieving high accuracy even when provided values for only 100,000 CpGs (Figure 3B). Tumor signal and classification probabilities were enhanced by linearly regressing methylation signatures from 16 non-malignant cell types to subtract normal background content from bulk methylation profiles (Figure 3C). Three neural networks which attained a balanced accuracy of >0.9 after n=100 epochs were subsequently integrated into a three-layer ensemble model, yielding M-PACT (Figure 3A, 3D). Overall, M-PACT achieved an average Fl score of 0.93 and 0.89 for the in silico validation cohort with tumor purities greater than 0.5 and less than 0.5, respectively. Simulated samples with lower tumor fractions were more likely to be misclassified (median=0.13, Figure 3E). For samples with tumor classification scoresabove 0.7, tumor burden was estimated by deconvoluting methylation profiles into malignant and non-malignant fractions.

[0675] The methylation- and CNV-based approaches as provided herein work across various and decreasing ctDNA fractions in cell-free DNA EM-seq samples. Even samples with only 5% ctDNA were considered true positives which demonstrates the excellent sensitivity of the methylation- and CNV-based approaches as provided herein (Figure 4).

[0676] The results demonstrate the feasibility of using DBSs for cfDNA analysis in oncology. Successful isolation and profiling of tumor-derived cfDNA, including the identification of characteristic fragment lengths and CNVs, highlight the potential of DBSs for non-invasive tumor monitoring. The stability of cfDNA in archived samples further supports the suitability of DBSs for retrospective studies and large-scale biobanking. These findings establish DBSs as a robust and practical tool for molecular diagnostics and precision medicine in resource-limited and pediatric settings.

[0677] Example 3 - Discussion

[0678] Previous studies have shown that a significantly larger proportion of pediatric cancers than previously assumed is initiated by large-scale copy number variations (CNVs) rather than single-gene mutations or alterations. These premalignant clones often originate embryonically in utero, as observed in most cases of neuroblastoma, medulloblastoma, and high-grade gliomas, and likely other entities as well. The "second hit," which drives malignant transformation, typically occurs years before symptoms emerge. Additionally, the patterns of the initial CNVs are highly specific to the tumor and its cell of origin, frequently recurring in fully developed tumors. This suggests their potential for predicting the malignant transformation of premalignant clones.

[0679] Previous urine-based catecholamine screenings for neuroblastoma in infants increased the detection of early-stage tumors but failed to reliably detect tumors in advanced stages. Overdiagnosis of favorable-outcome neuroblastomas often led to unnecessary interventions, posing physical and psychological risks. To address these issues, we identified predictive CNVs from large tumor datasets as specific and sensitive biomarkers. Additionally, we developed an enzymatic methylation conversion protocol for cfDNA with improved recovery rates to provide detailed methylation profiles. This approach is widely applied in diagnostics for childhood brain tumors and peripheral neuroblastomas and shows promising diagnostic and prognostic value for sarcomas.Methylation-based tumor classification from DBS could revolutionize pediatric cancer screening, offering precise tumor detection and organ-specific identification at early stages. This groundbreaking technology enables tailored surveillance programs and avoids unnecessary medical interventions, setting the stage for targeted therapeutic measures. It holds the potential to fundamentally transform pediatric oncology, paving the way for a new era in cancer prevention, diagnosis, and treatment.

[0680] Example 4 - A further exemplary embodiment of the invention

[0681] The following description relates to one or more solely optional variants and implementations of the invention:

[0682] According to exemplary embodiments, the invention may be used for the detection of a premalignant lesion in a subject or for the identification of a subject being at risk of developing cancer. The subject may be a newborn or infant. Preferably, to this end, a classification and preferably a binary classification of the subject is carried out. The result of the binary classification may indicate the presence or absence of a tumor.

[0683] According to exemplary embodiments of the invention, the sample may comprise cell-free DNA (cfDNA) of the subject. It may be provided that cfDNA is isolated from a sample of the subject, preferably from cerebrospinal fluid (CSF) or a dried blood sample.

[0684] It is possible that a DNA methylation status of a multitude of genomic CpG positions in the cfDNA of the sample is determined, for example by methylation conversion followed by whole genome sequencing or a methylation array platform, or by direct methylation sequencing. The determined DNA methylation status may then be provided in the form of and / or represented by digital data, also referred to as methylation data. The methylation data may comprise, particularly genome-wide, CpG methylation beta values. Particularly, the methylation data may comprise a vector of beta values, and each beta value may represent the methylation level at a specific genomic CpG Position as a continuous value between 0 (fully unmethylated) and 1 (fully methylated). The methylation data may be indexed by CpG identifier and may be stored on a computer-readable storage medium, for example in a tabular data format.

[0685] It is also possible that a copy number variation (CNV) pattern of a multitude of independent DNA segments in the cfDNA of the sample is determined by whole genome sequencing and / or chromosome translocations at a multitude of genomic loci in the cfDNA of the sample by capture-enriched DNAsequencing. The determined CNV pattern may then be provided in the form of and / or represented by digital data, also referred to as CNV data. The CNV data may comprise a set of genomic segments, and each segment may be defined by chromosomal coordinates and / or associated with a Iog2 ratio value representing relative copy number gain or loss. The CNV may be determined from the same sequencing data used for methylation analysis, without requiring a separate assay.

[0686] It is also possible that a sample comprising cell-free RNA (cfRNA) of the subject is provided and chromosome translocations are determined at a multitude of loci in the cfRNA by capture-enriched RNA sequencing. The determined chromosome translocations may then be provided in the form of and / or represented by digital data, also referred to as translocation data.

[0687] According to exemplary embodiments of the invention, at least a part of the digital data, particularly the methylation data and / or the CNV data and / or the translocation data, may be used for or in the context of the classification of the subject and particularly as input(s) or parameters for a classification framework, also referred to as classification workflow or classification algorithm.

[0688] The classification framework may be used to classifying the subject as having or not having a premalignant lesion or as being or as not being at risk of developing cancer. The classification may be carried out based on (i) the DNA methylation status and particularly a comparison of the DNA methylation status (represented by the methylation data) with the DNA methylation status of the multitude of genomic CpG positions in the DNA of a plurality of reference tumor and / or non-tumor samples that are or have been obtained by methylation conversion followed by sequencing or a methylation array platform, or by direct methylation sequencing, and / or (ii) the CNV pattern and particularly a comparison of the CNV pattern (represented by the CNV data) with the CNV pattern of a human reference genome, and / or a plurality of reference tumor and / or non-tumor samples, and / or the chromosome translocations. This is described in more detail below by way of example. It should be noted that the classification framework may be stored as computer-executable instructions or may be provided as a computer program according to aspects of the invention.

[0689] With regards to (i), it is possible that the classification framework is provided as a statistical framework for binary classification that quantifies the deviation of the sample's DNA methylation status, particularly methylation profile, as represented by the methylation data, from a non-tumor control baseline, particularly forming the abovementioned plurality of reference tumor and / or non-tumor samples. For example, the (genome-wide) CpG methylation beta values may be aggregated into 50,000 variably methylated genomic blocks for four sample cohorts: a peripheral blood reference panel, acontrol neuron cohort, a non-malignant CSF cohort, and a set of tumor-bearing cfDNA samples. The set of tumor-bearing cfDNA samples may be down sampled to varying tumor purities.

[0690] For each sample of the control neuron cohort, the non-malignant CSF cohort and the set of tumorbearing cfDNA samples, a pairwise weighted cosine similarity may be computed against every sample in the peripheral blood reference panel. This may produce a matrix of similarity scores for each of the control neuron cohort, the non-malignant CSF cohort and the set of tumor-bearing cfDNA samples.

[0691] The weighting of the pairwise weighted cosine similarity may be performed using a weighting scheme. The weighting scheme may assign a weight to each of the 50,000 variably methylated genomic blocks. The assigned weights may be increased in a monotonic manner, preferably such that the assigned weight is increased by 1 for each block (i.e., wi = i + 1). This may emphasize higher-ranked blocks, particularly to amplify signal from the most informative loci. The control neuron cohort and / or the non-malignant CSF cohort may then be split into two partitions. The first 30 samples of the respective cohort may serve as reference samples and the remaining samples of the respective cohort may serve as test samples. For each of the reference samples, the 5th percentile of its similarity vector across the peripheral blood reference panel may be extracted, producing a 30-sample null distribution representing baseline non-tumor affinity.

[0692] For each sample of the set of tumor-bearing cfDNA samples, the z-score may be calculated by extracting the 95th percentile of its similarity vector against the peripheral blood reference panel and standardizing against the 30-sample null distribution:

[0693] Z — (observedgsth—Pcontrol-ref) / Ocontrol-ref

[0694] Here, pcontroi-ref and oCOntroi-ref may be the mean and standard deviation of the 30-sample null distribution. The z-score may be calculated in the same way for each of the test samples, particularly to establish an empirical false-positive rate under the null. A z-score cutoff of -1.2 may be applied for the binary classification, less than (or possibly equal to) -1.2 may indicate the presence of a tumor and a z-score greater than (or possibly equal to) -1.2 may indicate the absence of a tumor. This may distinguish samples with detectable tumor signal from those consistent with a non-malignant background.

[0695] With regards to (ii), i.e. the CNV pattern and particularly the comparison of the CNV pattern (represented by the CNV data) with the CNV pattern of a human reference genome, and / or a pluralityof reference tumor and / or non-tumor samples, CNV-based ctDNA purity predictions may be determined, described exemplarily as follows:

[0696] CNVs may be detected from sequencing data by partitioning the genome into non-overlapping bins, correcting for GC content and mappability bias, and normalizing coverage as Iog2 ratios based on GC and mappability calculated from the human reference genome. Additional noise reduction techniques can be implemented via normalization with a panel of normal CNV profiles utilizing a robust principal component analysis. Comparison to the non-tumor samples facilitates the removal of background biases and enriching any tumor enriched foreground. Once normalized, contiguous segments of consistent copy number may be identified using a bayesian changepoint detection algorithm. A probabilistic model, such as a Hidden Markov Model (HMM) model, may then be fitted to predict tumor fraction based on large segments of the genome that have an altered copy number state, thereby obtaining CNV-based ctDNA purity predictions.

[0697] The calculated z-scores, which are particularly based on the methylation, may be merged with the CNV-based ctDNA purity predictions. This may be done using a weighted blending metric:

[0698] S = a x Norm(-z) + (1 - a) x PuritycNv-based

[0699] Here, z may represent a calculated z-score and PuritycNv-based may represent a CNV-based ctDNA purity prediction. Furthermore, a may represent a weighting between the z-scores and the purity. The z-scores may be normalized using a sigmoid function to scale the z-scores between 0 and 1. Based on S, the binary classification may be performed.

[0700] The methylation-based classifier and CNV merging method may be expanded to incorporate additional measurements, such as fragment length distributions and nucleosome positioning for greater sensitivity and specificity.

[0701] It is also possible that the methylation data and particularly the beta values are classified by a machinelearning model such as a deep neural network-based model that assigns the sample to one of a plurality of tumor or non-tumor classes. For samples where the machine-learning model yields a classification probability below the threshold, or as a complementary assay to increase sensitivity, the described binary classification (algorithm) may determine whether the cfDNA sample is tumor-positive or tumor-negative.References

[0702] Amemiya, H. M., Kundaje, A. and Boyle, A. P. (2019). The ENCODE Blacklist: Identification of Problematic Regions of the Genome. Scientific Reports 9 (1), 9354, doi: 10.1038 / s41598-019-45839-z.

[0703] Capper, D., Jones, D. T. W., Sill, M., Hovestadt, V., Schrimpf, D., Sturm, D., Koelsche, C., Sahm, F., Chavez, L., Reuss, D. E., Kratz, A., Wefers, A. K., Huang, K., Pajtler, K. W., Schweizer, L., Stichel, D., Olar, A., Engel, N. W., Lindenberg, K., Harter, P. N., Braczynski, A. K., Plate, K. H., Dohmen, H., Garvalov, B. K., Coras, R., Holsken, A., Hewer, E., Bewerunge-Hudler, M., Schick, M., Fischer, R., Beschorner, R., Schittenhelm, J., Staszewski, O., Wani, K., Varlet, P., Pages, M., Temming, P., Lohmann, D., Self, F., Witt, H., Milde, T., Witt, O., Aronica, E., Giangaspero, F., Rushing, E., Scheurlen, W., Geisenberger, C., Rodriguez, F. J., Becker, A., Preusser, M., Haberler, C., Bjerkvig, R., Cryan, J., Farrell, M., Deckert, M., Hench, J., Frank, S., Serrano, J., Kannan, K., Tsirigos, A., Bruck, W., Hofer, S., Brehmer, S., Seiz-Rosenhagen, M., Hanggi, D., Hans, V., Rozsnoki, S., Hansford, J. R., Kohlhof, P., Kristensen, B. W., Lechner, M., Lopes, B., Mawrin, C., Ketter, R., Kulozik, A., Khatib, Z., Heppner, F., Koch, A., Jouvet, A., Keohane, C., Muhleisen, H., Mueller, W., Pohl, LL, Prinz, M., Benner, A., Zapatka, M., Gottardo, N. G., Driever, P. H., Kramm, C. M., Muller, H. L., Rutkowski, S., von Hoff, K., Fruhwald, M. C., Gnekow, A., Fleischhack, G., Tippelt, S., Calaminus, G., Monoranu, C. M., Perry, A., Jones, C., Jacques, T. S., Radlwimmer, B., Gessi, M., Pietsch, T., Schramm, J., Schackert, G., Westphal, M., Reifenberger, G., Wesseling, P., Weller, M., Collins, V. P., Blumcke, I., Bendszus, M., Debus, J., Huang, A., Jabado, N., Northcott, P. A., Paulus, W., Gajjar, A., Robinson, G. W., Taylor, M. D., Jaunmuktane, Z., Ryzhova, M., Platten, M., Unterberg, A., Wick, W., Karajannis, M. A., Mittelbronn, M., Acker, T., Hartmann, C., Aidape, K., Schuller, LL, Buslei, R., Lichter, P., Kool, M., Herold-Mende, C., Ellison, D. W., Hasselblatt, M., Snuderl, M., Brandner, S., Korshunov, A., von Deimling, A. and Pfister, S. M. (2018). DNA methylation-based classification of central nervous system tumours. Nature 555 (7697), 469-474, doi: 10.1038 / nature26000.

[0704] Ceyhan-Birsoy, O., Murry, J. B., Machini, K., Lebo, M. S., Yu, T. W., Fayer, S., Genetti, C. A., Schwartz, T. S., Agrawal, P. B., Parad, R. B., Holm, I. A., McGuire, A. L., Green, R. C., Rehm, H. L., Beggs, A. H., Agrawal, P. B., Beggs, A. H., Betting, W. N., Ceyhan-Birsoy, O., Christensen, K. D., Dukhovny, D., Fayer, S., Frankel, L. A., Genetti, C. A., Graham, C., Green, R. C., Guiterrez, A. M., Harden, M., Holm, I. A., Krier, J. B., Lebo, M. S., Levy, H. L., Lu, X., Machini, K., McGuire, A. L., Murry, J. B., Naik, M., Nguyen, T. T., Parad, R. B., Peoples, H. A., Pereira, S., Petersen, D., Ramamurthy, U., Ramanathan, V., Rehm, H. L., Roberts, A., Robinson, J. O., Roumiantsev, S., Schwartz, T. S., Truong, T. K., VanNoy, G. E., Waisbren, S. E. and Yu, T. W. (2019). Interpretation of Genomic Sequencing Results in Healthy and III Newborns: Results from the BabySeq Project. The American Journal of Human Genetics 104 (1), 76-93, doi: https: / / doi.Org / 10.1016 / j.ajhg.2018.ll.016.

[0705] Czibere, L., Burggraf, S., Fleige, T., Gluck, B., Keitel, L. M., Landt, O., Durner, J., Rbschinger, W., Hohenfellner, K., Wirth, B., Muller-Felber, W., Vill, K. and Becker, M. (2020). High-throughput genetic newborn screening for spinal muscular atrophy by rapid nucleic acid extraction from dried blood spots and 384-well qPCR. European Journal of Human Genetics 28 (1), 23-30, doi: 10.1038 / s41431-019-0476-4.Dawson, S.-J., Tsui, D. W. Y., Murtaza, M., Biggs, H., Rueda, O. M., Chin, S.-F., Dunning, M. J., Gale, D., Forshew, T., Mahler-Araujo, B., Rajan, S., Humphray, S., Becq, J., Halsall, D., Wallis, M., Bentley, D., Caldas, C. and Rosenfeld, N. (2013). Analysis of Circulating Tumor DNA to Monitor Metastatic Breast Cancer. New England Journal of Medicine 368 (13), 1199-1209, doi: doi:10.1056 / NEJMoal213261.

[0706] Gerrish, A., Bowns, B., Mashayamombe-Wolfgarten, C., Young, E., Court, S., Bott, J., McCalla, M., Ramsden, S., Parks, M., Goudie, D., Carless, S., Clokie, S., Cole, T. and Allen, S. (2020). Non-lnvasive Prenatal Diagnosis of Retinoblastoma Inheritance by Combined Targeted Sequencing Strategies. J Clin Med 9 (11), doi: 10.3390 / jcm9113517.

[0707] Heider, K., Wan, J. C. M., Hall, J., Belie, J., Boyle, S., Hudecova, I., Gale, D., Cooper, W. N., Corrie, P. G., Brenton, J. D., Smith, C. G. and Rosenfeld, N. (2020). Detection of ctDNA from Dried Blood Spots after DNA Size Selection. Clin Chem 66 (5), 697-705, doi: 10.1093 / clinchem / hvaa050.

[0708] Holm, I. A., Agrawal, P. B., Ceyhan-Birsoy, O., Christensen, K. D., Fayer, S., Frankel, L. A., Genetti, C. A., Krier, J. B., LaMay, R. C., Levy, H. L., McGuire, A. L., Parad, R. B., Park, P. J., Pereira, S., Rehm, H. L., Schwartz, T. S., Waisbren, S. E., Yu, T. W., Agrawal, P. B., Beggs, A. H., Betting, W. N., Blout, C. L., Ceyhan-Birsoy, O., Christensen, K. D., Diamond, P., Dukhovny, D., Dunn, K. E., Fayer, S., Frankel, L. A., Genetti, C. A., Graham, C., Green, R. C., Gutierrez, A. M., Harden, M., Helm, M. H., Hoffman-Andrews, L., Holm, I. A., Krier, J. B., Lebo, M. S., Lee, K. B., Levy, H. L., Lu, X., Kalia, S. S., Machini, K., McGuire, A. L., Murry, J. B., Naik, M., Nguyen, T., Parad, R. B., Peoples, H. A., Pereira, S., Petersen, D., Ramamurthy, LL, Ramanathan, V., Rehm, H. L., Roberts, A., Robinson, J. O., Roumiantsev, S., Schwartz, T. S., Steffens, E. B., Towne, M. C., Truong, T. K., VanNoy, G. E., Waisbren, S. E., Weipert, C. M., Yu, T. W., Green, R. C., Beggs, A. H. and The BabySeq Project, T. (2018). The BabySeq project: implementing genomic sequencing in newborns. BMC Pediatrics 18 (1), 225, doi: 10.1186 / sl2887-018-1200-l.

[0709] Hovestadt, V. and Zapatka, M. (2017). conumee: Enhanced copy-number variation analysis using Illumina DNA methylation arrays. URL: https: / / bioconductor.org / packages / release / bioc / html / conumee.html [as of06 / 09 / 2024],

[0710] Kingsmore, S. F., Smith, L. D., Kunard, C. M., Bainbridge, M., Batalov, S., Benson, W., Blincow, E., Caylor, S., Chambers, C., Del Angel, G., Dimmock, D. P., Ding, Y., Ellsworth, K., Feigenbaum, A., Frise, E., Green, R. C., Guidugli, L., Hall, K. P., Hansen, C., Hobbs, C. A., Kahn, S. D., Kiel, M., Van Der Kraan, L., Krilow, C., Kwon, Y. H., Madhavrao, L., Le, J., Lefebvre, S., Mardach, R., Mowrey, W. R., Oh, D., Owen, M. J., Powley, G., Scharer, G., Shelnutt, S., Tokita, M., Mehtalia, S. S., Oriol, A., Papadopoulos, S., Perry, J., Rosales, E., Sanford, E., Schwartz, S., Tran, D., Reese, M. G., Wright, M., Veeraraghavan, N., Wigby, K., Willis, M. J., Wolen, A. R. and Defay, T. (2022). A genome sequencing system for universal newborn screening, diagnosis, and precision medicine for severe genetic diseases. The American Journal of Human Genetics 109 (9), 1605-1619, doi: https: / / doi.Org / 10.1016 / j.ajhg.2022.08.003.

[0711] Koelsche, C., Schrimpf, D., Stichel, D., Sill, M., Sahm, F., Reuss, D. E., Blattner, M., Worst, B., Heilig, C. E., Beck, K., Horak, P., Kreutzfeldt, S., Paff, E., Stark, S., Johann, P., Selt, F., Ecker, J., Sturm, D., Pajtler, K. W., Reinhardt, A., Wefers, A. K., Sievers, P., Ebrahimi, A., Suwala, A., Fernandez-Klett, F., Casalini, B., Korshunov, A., Hovestadt, V., Kommoss, F. K. F., Kriegsmann, M., Schick, M., Bewerunge-Hudler, M., Milde, T., Witt, O., Kulozik, A. E., Kool, M., Romero-Perez, L., Grunewald, T. G. P., Kirchner, T., Wick, W.,Platten, M., Unterberg, A., Uhl, M., Abdollahi, A., Debus, J., Lehner, B., Thomas, C., Hasselblatt, M., Paulus, W., Hartmann, C., Staszewski, O., Prinz, M., Hench, J., Frank, S., Versleijen-Jonkers, Y. M. H., Weidema, M. E., Mentzel, T., Griewank, K., de Alava, E., Martin, J. D., Gastearena, M. A. I., Chang, K. T., Low, S. Y. Y, Cuevas-Bourdier, A., Mittelbronn, M., Mynarek, M., Rutkowski, S., Schuller, U., Mautner, V. E, Schittenhelm, J., Serrano, J., Snuderl, M., Buttner, R., Klingebiel, T., Buslei, R., Gessler, M., Wesseling, R, Dinjens, W. N. M., Brandner, S., Jaunmuktane, Z., Lyskjaer, I., Schirmacher, P., Stenzinger, A., Brors, B., Glimm, H., Heining, C., Tirado, O. M., Sainz-Jaspeado, M., Mora, J., Alonso, J., Del Muro, X. G., Moran, S., Esteller, M., Benhamida, J. K., Ladanyi, M., Wardelmann, E., Antonescu, C., Flanagan, A., Dirksen, U., Hohenberger, P., Baumhoer, D., Hartmann, W., Vokuhl, C., Flucke, U., Petersen, I., Mechtersheimer, G., Capper, D., Jones, D. T. W., Frbhling, S., Pfister, S. M. and von Deimling, A. (2021). Sarcoma classification by DNA methylation profiling. Nat Commun 12 (1), 498, doi: 10.1038 / s41467-020-20603-4.

[0712] Madanat-Harjuoja, L. M., Klega, K., Lu, Y., Shulman, D. S., Thorner, A. R., Nag, A., Tap, W. D., Reinke, D. K., Diller, L., Ballman, K. V., George, S. and Crompton, B. D. (2022). Circulating Tumor DNA Is Associated with Response and Survival in Patients with Advanced Leiomyosarcoma. Clin Cancer Res 28 (12), 2579-2586, doi: 10.1158 / 1078-0432.Ccr-21-3951.

[0713] Raman, L., Van der Linden, M., Van der Eecken, K., Vermaelen, K., Demedts, I., Surmont, V, Himpe, U., Dedeurwaerdere, E, Ferdinande, L., Lievens, Y., Claes, K., Menten, B. and Van Dorpe, J. (2020). Shallow whole-genome sequencing of plasma cell-free DNA accurately differentiates small from non-small cell lung carcinoma. Genome Med 12 (1), 35, doi: 10.1186 / sl3073-020-00735-4.

[0714] Roman, T. S., Crowley, S. B., Roche, M. L, Foreman, A. K. M., O'Daniel, J. M., Seifert, B. A., Lee, K., Brandt, A., Gustafson, C., DeCristo, D. M., Strande, N. T., Ramkissoon, L., Milko, L. V., Owen, P., Roy, S., Xiong, M., Paquin, R. S., Butterfield, R. M., Lewis, M. A., Souris, K. J., Bailey, D. B., Rini, C., Booker, J. K., Powell, B. C., Week, K. E., Powell, C. M. and Berg, J. S. (2020). Genomic Sequencing for Newborn Screening: Results of the NC NEXUS Project. The American Journal of Human Genetics 107 (4), 596-611, doi: https: / / doi.Org / 10.1016 / j.ajhg.2020.08.001.

Claims

CLAIMS1. An in vitro or ex vivo method for the detection of a premalignant lesion in a subject, preferably a newborn or infant or for the identification of a subject, preferably a newborn or infant being at risk of developing cancer, wherein the method comprises(a) providing a sample comprising cell-free DNA (cfDNA) of a subject, preferably a newborn or infant,(bl) determining the DNA methylation status of a multitude of genomic CpG positions in the cfDNA of the sample of (a) by methylation conversion followed by whole genome sequencing or a methylation array platform, or by direct methylation sequencing, (b2) determining the copy number variation (CNV) pattern of a multitude of independent DNA segments in the cfDNA of the sample of (a) by whole genome sequencing and / or chromosome translocations at a multitude of genomic loci in the cfDNA of the sample of (a) by capture-enriched DNA sequencing and / orproviding a sample comprising cell-free RNA (cfRNA) of the subject, preferably a newborn or infant and determining chromosome translocations at a multitude of loci in the cfRNA by capture-enriched RNA sequencing,(c) classifying the subject, preferably newborn or infant as having or not having a premalignant lesion or as being or as not being at risk of developing cancer based on (i) a comparison of the DNA methylation status of (bl) with the DNA methylation status of the multitude of genomic CpG positions in the DNA of a plurality of reference tumor and / or non-tumor samples that are or have been obtained by methylation conversion followed by sequencing or a methylation array platform, or by direct methylation sequencing, and / or the DNA methylation status of (bl), and(ii) a comparison of the CNV pattern of (b2) with the CNV pattern of a human reference genome, and / or a plurality of reference tumor and / or non-tumor samples, and / or the chromosome translocations, and / or the CNV pattern of (b2).

2. The method of claim 1, wherein in step (c)(i) the subject, preferably newborn or infant is classified as being or as not being at risk of developing cancer based on a methylation-based predictive algorithm for tumors that is a deep neural network-based model that has been generated from a plurality of reference tumor samples with varying degrees of tumor purity and CpG sparsity by(1) simulating the tumor purity in cfDNA in the reference tumor samples by spiking in64methylation profiles from non-tumor sources of cfDNA,(2) addressing CpG sparsity in the deep neural network-based model by a network-based regression diffusion model to impute methylation values uncaptured by sequencing, thereby obtaining an imputation network for the deep neural network-based model,(3) training the imputation network on the multitude of genomic CpG positions of step (bl), and (4) enhancing the tumor signal and classification probabilities of the deep neural network-based model by beta regressing methylation signatures from non-malignant cell types to reduce normal DNA methylation contamination from tumor-specific DNA methylation.

3. The method of claim 2, wherein steps (3) and (4) are carried out for three independent sets of a multitude of genomic CpG positions thereby obtaining three pre-deep neural network-based model and layering the three models together into the final deep neural network-based model.

4. The method of any one of claims 1 to 3, wherein the DNA methylation status of a multitude of genomic CpG positions is determined by enzymatic methyl sequencing, wherein non-destructive enzymatic reactions, utilizing the enzymes Tet-methylcytosin-dioxygenase 2 (TET2) and apolipoprotein B mRNA editing enzyme, catalytic polypeptide (APOBEC) convert unmethylated cytosines to uracils, followed by CpG methylation profiling at single-CpG-site level based on a methylation array platform.

5. The method of claim 1 or 2, wherein in (b2) the copy number variation (CNV) pattern is determined by low-coverage whole genome sequencing and identification of the CNV pattern in the sequencing data by aligning the sequence reads to the human reference genome.

6. The method of claim 5 further comprising determining from the sequencing data the cfDNA fragment lengths of genomic segments in the cfDNA of the sample of (a) and preferably the ratio of short fragments of 100-150 bps to long fragments of 151-250 bp, thereby estimating the amount of circulating tumor DNA in the sample.

7. The method of any one of claims 1 to 6, wherein the sample is a blood, plasma, or serum sample of the subject, preferably newborn or infant.

8. The method of claim 7, wherein the sample is a dried blood sample, preferably a dried blood spot on filter paper and most preferably dried blood of a newborn on a Guthrie card.

9. The method of claim 8, wherein the method comprises in step (a) isolating cfDNA from a dried blood sample by(i) recovering the cfDNA of the dried blood sample under denaturing conditions by a protease K-containing buffer, wherein the buffer preferably comprises 2.5-10% sodium dodecyl sulfate and about 20 pl proteinase K,(ii) adding a buffer comprising a chaotropic salt, preferably 25-50% guanidinium chloride to the lysate of (i) comprising a carrier RNA,(iii) adding ethanol to the lysate of (ii),(iv) applying the lysate of (iii) to a silica-membrane-based anion exchange resin, wherein the DNA binds to the membrane,(v) optionally washing the membrane to residual contaminants, such as RNA, proteins, and low- molecular-weight impurities by one or more washing steps, wherein in the one or more washing steps preferably a solution of 50-100% guanidinium chloride and a solution of 0.1% sodium azide are used as washing buffers,(vi) optionally washing the membrane with ethanol,(vii) eluting the cfDNA from the membrane by distilled water or a buffer comprising 2.5-10% sodium dodecyl sulfate.

10. The method of any one of claims 1 to 9, wherein subject is a newborn or infant and the sample has been obtained from the newborn or infant within 1 year, preferably within 1 month, more preferably within 20 days and most preferably within about 10 days after birth.

11. The method of any one of claims 1 to 10, wherein in step (c) a classification rule is used for the classification.

12. The method of any one of claim 11, wherein the classification rule has been obtained based on random forest analysis, neuronal network, support vector machine, K-nearest neighbors, XGBoost or similar machine learning approaches.

13. The method of any one of claims 1 to 12, wherein the reference tumor and / or non-tumor samples have been obtained from subjects that developed a pediatric cancer within the first 10 years of life and / or did not develop a pediatric cancer within the first 10 years of life.6614. A system for the identification of a subject, preferably a newborn or infant being at risk of developing cancer, comprising:(a) one or more processors; and(b) memory coupled to the one or more processors and comprising instructions executable by the one or more processors to perform at least step (c) of a method according to any one of claims 1 to 13, and particularly, to perform a classification of the subject, preferably newborn or infant as having or not having a premalignant lesion or as being or as not being at risk of developing cancer based on(i) a DNA methylation status of a multitude of genomic CpG positions in a cell-free DNA (cfDNA) of the subject, and / ora comparison of said DNA methylation status with a DNA methylation status of a multitude of genomic CpG positions in the DNA of a plurality of reference tumor and / or non-tumor samples that are or have been obtained by methylation conversion followed by sequencing or a methylation array platform, or by direct methylation sequencing, and / or (ii) a copy number variation (CNV) pattern of a multitude of independent DNA segments in the cfDNA of the subject, and / ora comparison of said CNV pattern with a CNV pattern of a human reference genome, and / or a plurality of reference tumor and / or nontumor samples, and / orchromosome translocations at a multitude of genomic loci in the cfDNA of the subject or a multitude of loci in the cfRNA of the subject, wherein preferably the DNA methylation status of the multitude of genomic CpG positions in the cfDNA of the subject has been determined by methylation conversion followed by whole genome sequencing or a methylation array platform, or by direct methylation sequencing, particularly preferably by step (bl) of the method according to any one of claims 1 to 13,wherein preferably the CNV pattern of the multitude of independent DNA segments in the cfDNA of the subject has been determined by whole genome sequencing, and / or the chromosome translocations at the multitude of genomic loci in the cfDNA of the subject have been determined by capture-enriched DNA sequencing, and / or the chromosome translocations at the multitude of loci in the cfRNA of the subject have been determined by capture-enriched RNA sequencing, each one particularly preferably by step (b2) of the method according to any one of claims 1 to 13.6715. A computer-readable storage medium having computer-executable instructions stored or a computer program, that, when executed, causes a computer or the system of claim 14 to perform a method according to any one of claims 1 to 13 or to perform at least step (c) of the method according to any one of claims 1 to 13, and that particularly, when executed, causes a computer to perform a classification of a subject, preferably newborn or infant as having or not having a premalignant lesion or as being or as not being at risk of developing cancer based on (i) a DNA methylation status of a multitude of genomic CpG positions in a cell-free DNA (cfDNA) of the subject, and / ora comparison of said DNA methylation status with a DNA methylation status of a multitude of genomic CpG positions in the DNA of a plurality of reference tumor and / or non-tumor samples that are or have been obtained by methylation conversion followed by sequencing or a methylation array platform, or by direct methylation sequencing, and / or(ii) a copy number variation (CNV) pattern of a multitude of independent DNA segments in the cfDNA of the subject, and / ora comparison of said CNV pattern with a CNV pattern of a human reference genome, and / or a plurality of reference tumor and / or non-tumor samples, and / or chromosome translocations at a multitude of genomic loci in the cfDNA of the subject or a multitude of loci in the cfRNA of the subject,wherein preferably the DNA methylation status of the multitude of genomic CpG positions in the cfDNA of the subject has been determined by methylation conversion followed by whole genome sequencing or a methylation array platform, or by direct methylation sequencing, particularly preferably by step (bl) of the method according to any one of claims 1 to 13. wherein preferably the CNV pattern of the multitude of independent DNA segments in the cfDNA of the subject has been determined by whole genome sequencing, and / or the chromosome translocations at the multitude of genomic loci in the cfDNA of the subject have been determined by capture-enriched DNA sequencing, and / or the chromosome translocations at the multitude of loci in the cfRNA of the subject have been determined by capture-enriched RNA sequencing, each one particularly preferably by step (b2) of the method according to any one of claims 1 to 13.