Prediction of tumour growth and prognosis
By utilizing DNA methylation data from fluctuating CpG loci to predict the historical growth rate of a cancer, this method addresses the limitations of existing tumor growth prediction methods, offering a cost-effective and clinically scalable solution for predicting cancer progression risk.
Patent Information
- Application Number
- PCT/GB2024/052831
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-08
- Filing Date
- 2024-11-07
- Publication Date
- 2025-05-15
AI Technical Summary
Current methods for predicting the growth rate and progression of tumors are limited by the need for multi-region high-depth whole genome sequencing, which is labor-intensive and costly, making it unsuitable for clinical scalability.
A method using DNA methylation data from fluctuating CpG loci to predict the historical growth rate of a cancer, allowing for the prediction of cancer progression risk without the need for longitudinal sampling or high-cost sequencing.
Enables the prediction of cancer progression risk and patient mortality based on standard DNA methylation data, providing valuable prognostic information for informing treatment decisions.
Smart Images

Figure GB2024052831_15052025_PF_FP_ABST
Abstract
Description
[0001]P603125PC001 PREDICTION OF TUMOUR GROWTH AND PROGNOSISTechnical Field Embodiments of the present invention described herein relate to methods ofpredicting the risk of progression of a cancer of a subject. Background to the Invention The clonal history of a cell is recorded within its (epi)genome via the accumulation of heritable changes. Studying the patterns of these heritable changes allows for the reconstruction of a tissue’s clonal architecture and the dynamics of clone replacements.However, existing markers, such as the accumulation of single nucleotide variants (SNVs),occur infrequently and therefore are only suitable for recording clonal dynamics that occur over long timescales. The inventors previously developed a new method that can measure contemporarycell dynamics with fluctuating methylation clocks (FMCs) (Gabbutt et al. 2022). Theinventors identified a set of fluctuating CpG (fCpG) loci which fluctuate between homozygous methylated, heterozygous, and homozygous demethylated states within individual cells. In a polyclonal population, these FMCs are effectively desynchronized, leading to the average methylation being randomly distributed around the steady-state methylation level. Following a clonal expansion, however, the methylation of the entire population resembles that of the progenitor cell, leaving a clear signal on the bulk methylation patterns. The inventors have previously employed these FMCs to quantitatively measure the human adult stem cell dynamics of healthy intestinal glands. Measuring the historical growth dynamics of an individual’s tumour is alongstanding problem, as longitudinal samples from non-treated disease are typically notavailable, because treatment is generally initiated soon after a cancer is detected.Attempts have been made to infer the growth rate of individual tumours using phylogeneticmethods applied to SNV data from multiple regions within the tumour, but such methodsrequire multi-region high-depth whole genome sequencing. These methods are thus notscalable to the clinic due to their labour-intensiveness and prohibitively high cost. Variousmodels of tumour growth have been proposed, but there is no consensus about the growth patterns that tumours exhibit. This is an important problem because there is a need to accurately assess an individual’s tumour growth in order to inform treatment decisions. This problem has been addressed by the present invention.P603125PC002 Summary of the Disclosure The present disclosure addresses the above, by providing methods of predictingthe risk of progression of a cancer of a subject. In view of the above, the present disclosure relates to a method of predicting therisk of progression of a cancer of a subject, wherein the method comprises: obtaining DNAmethylation data of fluctuating CpG loci of DNA derived from the cancer; using the DNAmethylation data to generate a methylation distribution from which a historical growth rate of the cancer can be determined; and predicting the risk of progression of the cancer based on the historical growth rate of the cancer. Advantageously, this method enables the historical growth rate of an individual’scancer to be predicted based on standard DNA methylation data, and without the need for longitudinal sampling. The inventors have demonstrated that measurements of the dynamics of historical cancer growth predicts the risk of cancer progression and patient death. The inventors have developed a model which can be applied to methylation data on a patient-by-patient basis to determine key parameters of historical cancer growth, and use this to predict the individual patient’s risk of cancer progression and death. This information can be used to inform treatment, such as whether more aggressive or palliative treatment is required, or whether treatment should be deferred (i.e., a “watch and wait” approach). Accordingly, the method holds valuable prognostic information enabling patient stratification beyond that of standard clinical practice. The term “predicting the risk of progression of a cancer of subject”, as used herein, is intended to mean determining the likelihood that the cancer will progress to the next stage of the disease (e.g., from stage 0 to stage I; from stage I to stage II; from stage II to stage III; from stage III to stage IV) or to a later stage of the disease (e.g., from stageI to stage IV; from stage 0 to stage II; etc.) or to a new form that is unresponsive to aparticular treatment. In some embodiments, the risk of progression of the cancer to the next stage or a later stage within a certain time window is predicted. In some embodiments, the time window is at least 1 month, at least 3 months, at least 6 months, at least 9 months, at least 1 year, at least 2 years, at least 3 years, at least 4 years, at least 5 years, or at least 10 years. In some embodiments, the time window is at least 6 months. In some embodiments, the time window is 6 months. In some embodiments, the time window is no more than 20 years, no more than 10 years, no more than 5 years, no more than 4 years, no more than 3 years, no more than 2 years, no more than 1 year, or no more than 9 months. In some embodiments, the time window is the entire lifespan of the subject. Typically, the time window is calculated from the time point at which the DNA methylation data was obtained.P603125PC003 The term “DNA derived from the cancer”, as used herein, is intended to mean DNAobtained from a cancer cell of the cancer or DNA that was released from a cancer cell ofthe cancer (i.e., cell-free DNA). This term may be replaced with “DNA obtained from thecancer” or “cancer DNA”.Cancer The cancer may be any cancer. For example, the cancer may be a blood cancer ora solid tumour. In some embodiments, the cancer is lymphoma, leukemia, myeloma,carcinoma, sarcoma, or melanoma. In some embodiments, the cancer is a blood cancer.In some embodiments, the cancer is a hematologic malignancy. In some embodiments,the cancer is leukemia (e.g., chronic lymphocytic leukemia), lymphoma (e.g., non-Hodgkin's lymphoma), or multiple myeloma. In some embodiments, the cancer is chroniclymphocytic leukemia (CLL).In some embodiments, the cancer is chronic lymphocytic leukemia, small lymphocytic lymphoma, non-Hodgkin's lymphoma, indolent non-Hodgkin's lymphoma (iNHL), refractory iNHL, mantle cell lymphoma, follicular lymphoma, lymphoplasmacytic lymphoma, marginal zone lymphoma, cell lymphoma large immunoblastic, lymphoblasticlymphoma, splenic marginal zone B-cell lymphoma (+ / - hairy lymphocytes), nodalmarginal zone lymphoma (+ / - monocytoid B cells), extranodal marginal zone B-celllymphoma of the mucosa-associated lymphoid tissue, cutaneous T-cell lymphoma, extranodal T-cell lymphoma, anaplastic large cell lymphoma, angioimmunoblastic T-cell lymphoma, mycosis fungoides, B-cell lymphoma, diffuse large B-cell lymphoma, mediastinal cell lymphoma B large, intravascular large B cell lymphoma, primary effusion lymphoma, uncleaved small cell lymphoma, Burkitt's lymphoma, multiple myeloma, plasmoc itoma, acute lymphocytic leukemia, T-cell acute lymphoblastic leukemia, B-cell acute lymphoblastic leukemia, B-cell cytic prolymphoblastic leukemia, acute myeloid leukemia, juvenile myelomonocytic leukemia, minimal residual disease, hairy cell leukemia, primary myelofibrosis, secondary myelofibrosis, chronic myeloid leukemia, myelodysplastic syndrome, myeloproliferative disease, or Waldestrom's macroglobulinemia. The term “cancer”, as used herein, is also intended to includepreneoplastic disorders such as monoclonal gammopathy of undetermined significance(MGUS). In some embodiments, the cancer is pancreatic cancer, urologic cancer, bladdercancer, colorectal cancer, colon cancer, breast cancer, prostate cancer, kidney cancer,hepatocellular cancer, thyroid cancer, gallbladder cancer, lung (e.g., cancer, small celllung cancer), ovarian cancer, cervical cancer, gastric cancer, endometrial cancer, esophageal cancer, head and neck cancer, melanoma, neuroendocrine cancer, CNSP603125PC004 cancer, tumours brain (e.g., glioma, anaplastic oligodendroglioma, adult glioblastoma multiforme, and adult anaplastic astrocytoma), bone cancer, soft tissue sarcoma, retinoblastomas, neuroblastomas, peritoneal effusions, malignant pleural effusions, mesotheliomas, Wilms tumours, trophoblastic neoplasms, hemangiopericytomas, Kaposi's sarcomas, myxoid carcinoma, round cell carcinoma, round cell carcinoma, squamous cell carcinomas, oral carcinomas, cancers of the adrenal cortex or ACTH-producing tumours. Subject Typically, the method determines the historical growth rate of a cancer in a subject(e.g., an endogenous cancer). In alternative embodiments, the method determines thehistorical growth rate of a cancer in vitro or ex vivo.The subject may be any subject with cancer. The term "subject" is intended to include living organisms in which an immune response can be elicited (e.g., mammals). In some embodiments, the subject is non-human. Examples of non-human subjectsinclude dogs, cats, mice, rats, and transgenic species thereof. In some embodiments, thesubject is human. In some embodiments, the method further comprises the step of obtaining asample comprising cancer-derived DNA or cancer cells from the subject. Any suitablesample may be obtained provided it comprises cancer-derived DNA. In someembodiments, the sample is a tumour biopsy or primary tumour sample. In someembodiments, the sample is a liquid sample, such as a blood sample, a serum sample, a plasma sample, a bone marrow sample, a cerebrospinal fluid sample, a urine sample or an ascites sample. In some embodiments, the sample is a blood sample. In some embodiments, the method comprises the step of obtaining a sample comprising cancer cells from the subject. In some embodiments, the method further comprises the step of processing thesample comprising cancer-derived DNA or cancer cells. Relevant processing techniquesare known and routine in the art. For example, the sample may be purified (e.g., centrifugation, filtration, sorting). As a further example, DNA may be extracted (e.g., using QIAamp Blood Mini Kit). As a further example, an FFPE Restore Protocol may be applied (e.g., Infinium HD FFPE Workflow). In some embodiments, the method further comprises the step of treating the sample with sodium bisulfite (i.e., sodium bisulfite conversion). In alternative embodiments, DNA methylation data can be obtained without bisulfite conversion, forexample using the method described in Füllgrabe et al. (Simultaneous sequencing ofgenetic and epigenetic bases in DNA. Nat Biotechnol 41, 1457–1464 (2023)) or in SimpsonP603125PC005 et al. (Detecting DNA cytosine methylation using nanopore sequencing. Nat Methods14(4):407-410 (2017).DNA methylation data The term “obtaining DNA methylation data”, as used herein, is intended to coveractively collecting the DNA methylation data (e.g., carrying out a DNA methylation experimental step) and / or retrieving DNA methylation data that has been previously collected (e.g., stored in a database). In some embodiments, the step of obtaining DNA methylation data comprises retrieving DNA methylation data from a database. In some embodiments, the step of obtaining DNA methylation data comprisesperforming a DNA methylation analysis (e.g., by microarray and / or next-generationsequencing and / or nanopore long read sequencing). In some embodiments, the step ofobtaining DNA methylation data comprises performing a DNA methylation microarray. DNAmethylation microarrays are routinely used in the art (e.g., Illumina InfiniumHumanMethylation450 and MethylationEPIC BeadChip). In some embodiments, the stepof obtaining DNA methylation data comprises performing next-generation sequencing, such as bisulfite sequencing, or tet-assisted bisulfite sequencing. In some embodiments, the method further comprises a quality control step of the DNA methylation data. In some embodiments, the method further comprises a step of normalising the DNA methylation data. Normalisation methods are known to one skilled in the art and include, for example, single-sample Noob normalisation, quantile normalisation, all sample mean normalisation, peak-based correction, subset-quantile within array normalisation,subset quantile normalisation, and β-mixture quantile normalization. In someembodiments, the DNA methylation data is normalised by single-sample Noobnormalisation. This provides the advantage of simply correcting for known technical noise(unlike other methods which risk overcorrecting the fCpG signal), and works on a singlesample basis, unlike many other batch correction methods. In certain embodiments, thesingle-sample Noob normalisation combines a background correction step and a dye bias correction step, using the control probes on the methylation microarray. The term “fluctuating CpG loci”, as used herein, is intended to mean CpG loci inwhich each allele stochastically methylates and demethylates at a similar rate withinindividual cells in a selectively neutral manner and not subject to direct biologicalregulation. This term covers all of the fluctuating CpG loci, as well as a subset of the fluctuating CpG loci. In some embodiments, DNA methylation data of a subset offluctuating CpG loci of DNA derived from the cancer is obtained. The term “derived from”,as used herein, may be replaced with “from”.P603125PC006 In some embodiments, the fluctuating CpG loci comprise at least 50% of the CpG loci listed in Table 1 below. In some embodiments, the fluctuating CpG loci comprise at least 60% of the CpG loci listed in Table 1 below. In some embodiments, the fluctuating CpG loci comprise at least 70% of the CpG loci listed in Table 1 below. In some embodiments, the fluctuating CpG loci comprise at least 80% of the CpG loci listed in Table 1 below. In some embodiments, the fluctuating CpG loci comprise at least 90% of the CpG loci listed in Table 1 below. In some embodiments, the fluctuating CpG loci comprise at least 95% of the CpG loci listed in Table 1 below. In some embodiments, the fluctuating CpG loci comprise at least 98% of the CpG loci listed in Table 1 below. In some embodiments, the fluctuating CpG loci comprise at least 99% of the CpG loci listed in Table 1 below. In some embodiments, the fluctuating CpG loci comprise 100% of the CpG loci listed in Table 1 below. Table 1: Exemplary fCpG loci Name cg18174678 cg18438591 cg05821046 cg20433858 cg01768082 cg08706567 cg26354017 cg26013992 cg20510474 cg16526047 cg16342298 cg14159672 cg00597687 cg03429644 cg10126324 cg17428085 cg18086868 cg24049468 cg01621716 cg06036677 cg06320401 cg13986762 cg05694621 cg19584871 cg17132079 cg17742781 cg20704450 cg04896168 cg01931792 cg13354934 cg15546227 cg01730064 cg18143535 cg04082016 cg09368716 cg25189904 cg26536593 cg26537280 cg03350299 cg12208770 cg17083209 cg12726839 cg00729875 cg25071744 cg26863172 cg24724587 cg24598973 cg02605776 cg08965143 cg23288755 cg10593400 cg18224942 cg22697853 cg02716646 cg06770735 cg17108141 cg24557048 cg08477332 cg22488717 cg20312132 cg23627354 cg01894875 cg22941637 cg00956987 cg26220594 cg26523099 cg23656015 cg17935233 cg24209528 cg05129477 cg22374525 cg08808724 cg24058386 cg10588622 cg07981328 cg09815769 cg13885623 cg23640929 cg14223654 cg24792360 cg26050734 cg22698629 cg24329783 cg14758072 cg16585234 cg01963297 cg24876897 cg07580762 cg24686644 cg01811796 cg09037777 cg24730688 cg21691116 cg14339650 cg08134856 cg19321979 cg22077361 cg10092377 cg24220031 cg01815720 cg04843111 cg25593625 cg03948781 cg00058515 cg21151355 cg03390245 cg08133631 cg17178900 cg24222817 cg18931815 cg04415780 cg19633205 cg14893161 cg22111694 cg14076729 cg10963061 cg04673462 cg23615892 cg26874229 cg14523898 cg05843457 cg17434008 cg06939851 cg25924274 cg02851558 cg24509168 cg12410921 cg24860534 cg11942181 cg09847153 cg13229857 cg13393721 cg02012159 cg24980657 cg12051614 cg06775570 cg22345063 cg14099468 cg27664689 cg15986668 cg16029875 cg00405069 cg21397540 cg26817877P603125PC007 cg21361856 cg17658717 cg12494515 cg23521468 cg27547543 cg21627775 cg03855994 cg25932599 cg06060754 cg23906067 cg01217071 cg20540428 cg10432947 cg23248424 cg00872984 cg10634136 cg20232291 cg08644045 cg00619978 cg07885132 cg22308949 cg11874323 cg23139982 cg16035780 cg04749507 cg14157435 cg04865442 cg17814781 cg09075844 cg20980321 cg27369013 cg20988073 cg02136252 cg06836380 cg09459982 cg19075252 cg18026197 cg01096266 cg07401067 cg22297966 cg01898867 cg15709065 cg03315557 cg26875852 cg00116315 cg26302230 cg09373148 cg06634862 cg18394854 cg02377690 cg00680696 cg18005180 cg02170478 cg27586797 cg13602274 cg00899350 cg00472710 cg06483432 cg01809941 cg09140531 cg03727500 cg00010946 cg13210763 cg09560636 cg11611580 cg15371801 cg15180789 cg08572214 cg06837426 cg15234946 cg27661104 cg21030598 cg27242132 cg20451680 cg08313539 cg13716566 cg17036624 cg09511421 cg08336234 cg21855021 cg14853772 cg04207166 cg25654705 cg17503389 cg07571531 cg23531340 cg17546649 cg25886683 cg14996985 cg16699385 cg04131969 cg16930811 cg22025206 cg17253517 cg06864789 cg01583753 cg08072716 cg04156016 cg02250764 cg18136963 cg20097593 cg22158068 cg12518844 cg11620475 cg16073143 cg05414442 cg17057719 cg26690304 cg11809091 cg18322025 cg03813688 cg24576535 cg25102726 cg17763019 cg01238672 cg08201735 cg14218447 cg23522872 cg07017875 cg15279541 cg06383022 cg13511885 cg16079430 cg17303540 cg14115346 cg05723825 cg06726820 cg05530317 cg09546802 cg04767753 cg22022881 cg01324343 cg01784614 cg11609571 cg08660295 cg25079915 cg24960291 cg05465935 cg12352302 cg19579217 cg00280895 cg14463292 cg14789272 cg16006841 cg18278486 cg25223285 cg24793722 cg19590578 cg07763376 cg04612667 cg05353292 cg08222618 cg12980795 cg17974424 cg12623302 cg08719380 cg23855319 cg13003311 cg16935061 cg26968378 cg25487775 cg02754929 cg18665594 cg00101728 cg12584458 cg15742848 cg09573795 cg18586823 cg01474544 cg07734926 cg15161959 cg12948920 cg16988194 cg17204113 cg09012858 cg26853787 cg11266682 cg25218470 cg00942918 cg14132236 cg24715680 cg15175162 cg15574301 cg25063707 cg13143120 cg18887096 cg03238216 cg08470180 cg02631126 cg15070894 cg03817727 cg03313172 cg06536614 cg03350138 cg08497487 cg24330386 cg00236919 cg25340688 cg00219626 cg15946590 cg05868531 cg03463411 cg26896946 cg04522432 cg09876447 cg11559198 cg11388320 cg15548198 cg11406274 cg23252259 cg21054919 cg01883540 cg19727817 cg03228974 cg14036627 cg10883069 cg11991617 cg02147194 cg08710841 cg11811828 cg02555923 cg08125755 cg16391973 cg17931227 cg01869058 cg06234584 cg17069313 cg07751793 cg26668675 cg10223234 cg27364650 cg14016257 cg22675956 cg26818629 cg04142955P603125PC008 cg13069441 cg03454541 cg04966191 cg16351002 cg08355456 cg12672189 cg17264513 cg03894796 cg04309849 cg21570702 cg18816397 cg15575538 cg19682134 cg13221347 cg05353869 cg14590011 cg01499522 cg01319235 cg25258098 cg19624354 cg06386482 cg15881597 cg14182974 cg01794156 cg23149852 cg25716013 cg00506049 cg21891967 cg06974175 cg03222066 cg14989243 cg15285764 cg13891189 cg11976911 cg20118840 cg12602851 cg19435720 cg00220243 cg07238401 cg10961484 cg05237015 cg02984009 cg03191962 cg05061471 cg10262710 cg05567646 cg09669835 cg01126560 cg07934552 cg01094687 cg02270332 cg27203184 cg13573611 cg04862679 cg04468568 cg01883195 cg21406967 cg04824555 cg01421380 cg15235707 cg25399239 cg20665157 cg03576809 cg03101664 cg24829292 cg06388363 cg02193461 cg07881032 cg15887846 cg13147090 cg26850117 cg23117316 cg21177396 cg20051530 cg12675583 cg07257824 cg23305129 cg19519624 cg23188684 cg09328283 cg05779406 cg21724654 cg13518232 cg00864171 cg09601159 cg14662585 cg26967842 cg00219169 cg07716408 cg17570948 cg26012061 cg04671734 cg09955645 cg10283695 cg22566355 cg06955417 cg08242633 cg20503657 cg25132257 cg06868955 cg22750906 cg20463151 cg06188367 cg08743050 cg19893439 cg22330875 cg26041569 cg12845268 cg14201467 cg08409225 cg13710556 cg13346869 cg06850159 cg05566455 cg10306450 cg02207386 cg21505509 cg21012874 cg24953078 cg00700412 cg06166490 cg08247852 cg02971322 cg17840408 cg08784247 cg05929675 cg23245289 cg22704845 cg12110980 cg23390920 cg21004899 cg11491002 cg24143137 cg10026995 cg24460369 cg23068772 cg19632760 cg00492979 cg01694451 cg06840837 cg03552937 cg03003919 cg07262395 cg19415743 cg25752163 cg25783361 cg27053337 cg12765123 cg12904004 cg23469448 cg13578160 cg20366110 cg24696332 cg23065768 cg03495030 cg01830073 cg04724477 cg00153041 cg10570158 cg03554925 cg16453056 cg01385052 cg16977735 cg18052274 cg17566325 cg23192873 cg06070970 cg11444332 cg04106006 cg15878909 cg18209323 cg05807444 cg00142642 cg06705930 cg03353445 cg05667818 cg20845050 cg12428160 cg25948180 cg10809165 cg18577231 cg00643296 cg03444077 cg00233420 cg16267202 cg03027739 cg09945801 cg04004590 cg03606954 cg25134647 cg12560447 cg19806221 cg24160569 cg06962549 cg01844176 cg10157098 cg11583762 cg12450708 cg26418434 cg17292337 cg06761584 cg10327428 cg00169964 cg00089550 cg10317314 cg22047901 cg02980621 cg15233961 cg11132272 cg14687298 cg09916174 cg18570392 cg14251734 cg14933468 cg04348558 cg20104432 cg21244580 cg18103548 cg07539927 cg03320170 cg15935721 cg24636969 cg06620993 cg10130811 cg20402783 cg11945929 cg10039928 cg10475928 cg24851651 cg08455772 cg04816699 cg09981361 cg23331156 cg03349769 cg17295489P603125PC009 cg07910460 cg10775991 cg16548362 cg07391831 cg14192542 cg21961771 cg23099211 cg16367168 cg16893174 cg09601923 cg10651121 cg01214012 cg26538942 cg10092265 cg14952312 cg15996769 cg05225684 cg00639984 cg06418907 cg12443604 cg16624888 cg09989251 cg00671386 cg12156887 cg08084655 cg20092036 cg26421140 cg06334495 cg18048405 cg06697425 cg11367159 cg01115857 cg09242293 cg21197958 cg05865011 cg05134054 cg24036523 cg09798061 cg17033854 cg07010633 cg08800396 cg04716394 cg20035459 cg10747531 cg11043990 cg05361406 cg18220816 cg02873098 cg20291162 cg00423030 cg20050828 cg05150608 cg02859934 cg25067162 cg24207009 cg26750617 cg10095084 cg06138239 cg08471713 cg13694343 cg23591611 cg22454744 cg16014085 cg13947929 cg17563271 cg07610777 cg25440893 cg24642468 cg02219949 cg07016075 cg20334079 cg03945400 cg17484472 cg11732492 cg03889540 cg12206359 cg17221584 cg27268835 cg12165772 cg11520030 cg09084477 cg23634401 cg26580413 cg01992590 cg01908551 cg13868001 cg20481312 cg02574464 cg09889948 cg21266975 cg12436019 cg26710563 cg27014354 cg13655803 cg10092257 cg03898952 cg24382249 cg00618711 cg02237906 cg19453250 cg22731513 cg26218577 cg07307142 cg12407791 cg00057722 cg16032841 cg16146501 cg03072681 cg12700283 cg05458220 cg06897282 cg08637403 cg04810063 cg04498014 cg13287964 cg25658983 cg20246113 cg00511475 cg10759591 cg21237861 cg03593908 cg18105749 cg08521396 cg20067780 cg06511678 cg07974719 cg25985504 cg00591660 cg20932150 cg18086635 cg19759502 cg07615383 cg03264209 cg23246911 cg05872758 cg07937631 cg00098799 cg01680674 cg22635541 cg24864576 cg00750481 cg13812291 cg08230268 cg08103988 cg02978168 cg21805940 cg03380198 cg05787209 cg08750459 cg14887613 cg25444002 cg17074431 cg05725404 cg15660077 cg15002904 cg09460553 cg03822622 cg02451853 cg16696476 cg05308639 cg22721998 cg12756504 cg09758204 cg15041662 cg08350509 cg08183074 cg12163955 cg08638180 cg24156181 cg15255455 cg18624544 cg06891458 cg10433128 cg18091275 cg13617776 cg14184890 cg17082938 cg06655190 cg05451094 cg27555894 cg11071193 cg08133919 cg06521653 cg19865134 cg11056604 cg13097993 cg12036633 cg04245248 cg19789473 cg20587168 cg23511824 cg25481157 cg06471402 cg09760387 cg04351156 cg18771300 cg00474218 cg13768055 cg17757575 cg10750464 cg06892679 cg25924688 cg12770741 cg02633371 cg11738485 cg07922154 cg24795825 cg02866639 cg02481789 cg01632288 cg13099429 cg19708986 cg00668150 cg26782881 cg26914865 cg18313317 cg02054724 cg11684897 cg01405107 cg16701007 cg08072202 cg07695058 cg10553748 cg09983216 cg15428479 cg18333511 cg24667603 cg24768135 cg06602723 cg19492423 cg04329125 cg10135717 cg21358336 cg10237252 cg17969123P603125PC0010 cg14265502 cg22603450 cg25786696 cg02544614 cg16501323 cg03657045 cg10503234 cg27331471 cg02519263 cg17356252 cg26643967 cg19870717 cg04382643 cg01735357 cg11113589 cg12216477 cg14111332 cg18741372 cg08389814 cg09417038 cg17069873 cg23671699 cg11237948 cg20964064 cg09374293 cg18458353 cg15043935 cg17716663 cg22463915 cg24435209 cg22376864 cg24084706 cg11773468 cg03173827 cg15498409 cg09438457 cg20467412 cg19851487 cg00237512 cg07331806 cg15173319 cg03489492 cg17071417 cg09990613 cg21401457 cg14061069 cg17490089 cg17729891 cg09123507 cg11141652 cg18861015 cg01272599 cg12289045 cg08424491 cg16210088 cg23448153 cg25506915 cg19014858 cg20102019 cg20548231 cg21198638 cg06189175 cg02592403 cg11133732 cg01704580 cg12068280 cg15050529 cg06878361 cg05389922 cg00276389 cg14780446 cg23899408 cg22331349 cg03539340 cg18982625 cg16194588 cg10384482 cg15721243 cg09607548 cg00610021 cg03630273 cg15010352 cg15061569 cg05139523 cg06961233 cg05696092 cg00706441 cg04358463 cg19348676 cg25163076 cg21723184 cg00003818 cg18282375 cg02174884 cg07459594 cg26651782 cg14021871 cg05386230 cg23303369 cg06942979 cg10635122 cg05516285 cg00216901 cg10797197 cg26730347 cg24794228 cg22548353 cg17083925 cg17029694 cg09696535 cg03153360 cg06942742 cg13928427 cg24818238 cg08178168 cg15636087 cg09007470 cg04307107 cg25421566 cg03688400 cg07786254 cg26920327 cg11598935 cg01481690 cg06781532 cg19907366 cg14881601 cg12751644 cg08431893 In some embodiments, the method further comprises a step of identifyingfluctuating CpG loci of DNA derived from the cancer. In some embodiments, the CpG lociare identified as fluctuating CpG loci by selecting CpG loci that: (i) are heterogeneousacross different subjects with the same disease, (ii) have an intermediate averagemethylation across different subjects with the same disease (i.e. similar methylation and demethylation rates), and (iii) are unlikely to be associated with specific cell-types, cancertype / subtype or epigenetic remodelling.For (i), the standard deviation of each CpG locus in each cancer type is typicallyemployed, averaged over all the cancer types in the training set as a proxy for heterogeneity. For (ii), the average absolute distance from 50% methylated in each locus is typically used, averaged across all the cancer types in the training set. To limit the effect of cell / cancer type specific methylation, as in (iii), the Laplacian Score is typically used to rank CpG loci in an unsupervised manner to identify which CpG loci encode are importantfor clustering similar subject samples (He et al., "Laplacian Score for Feature Selection."NIPS 2005). A number of different metrics could be employed in step (iii), including: theP603125PC0011 loading of the first few principal components in a principle component analysis (PCA), the feature importance score from a supervised machine learning classification method (e.g. random forest or XGBoost) or the F value in an ANOVA test. By taking CpG loci with a high intra-disease heterogeneity, an average methylation near 50% and those CpG sites with a low feature importance, a set of putative fCpG loci can be identified. Stochastic forward model The DNA methylation data is used to generate a methylation distribution from which a historical growth rate of the cancer can be determined. In some embodiments, the methylation distribution is generated by fitting a stochastic forward model to the methylation data using an inference method. Advantageously, the stochastic forward model maps from a set of biologically meaningful parameters to the methylation distribution of the fCpG loci. One of the biologically meaningful parameters is the historical growth rate of the cancer. The term “historical growth rate of a cancer”, as used herein, is intended to mean the growth rate of the cancer over the period of existence of the cancer (i.e., between day 0 of the cancer and the day on which the methylation data was collected). In some embodiments, the historical growth rate assumes an exponentially growing population of cancer cells. In some embodiments, the age of the subject when the most recent commonancestor (MRCA) of the sampling population emerged can be determined from themethylation distribution. In some embodiments, the cancer progenitor size can be determined from themethylation distribution. The cancer progenitor size is a composite parameter, equallingexp(growth rate * (age - t_MRCA)).In some embodiments, one or more further parameters can be determined fromthe methylation distribution, such as one or more epigenetic switching rates selected fromhomozygous to heterozygous methylation rate, heterozygous to homozygous methylation rate, homozygous to heterozygous demethylation rate, and heterozygous to homozygous demethylation rate. In some embodiments, the homozygous to heterozygous methylation rate, the heterozygous to homozygous methylation rate, the homozygous to heterozygousdemethylation rate, and / or the heterozygous to homozygous demethylation rate isdetermined from the methylation distribution. In some embodiments, the average methylation level of contaminating normal cells is determined from the methylation distribution.P603125PC0012 In some embodiments, the stochastic forward model is based on a discrete time scheme, whereby the number of homozygous methylated, heterozygous methylated cells and homozygous demethylated cells are updated at each time step, according to thegrowth rate and epigenetic switching rate parameters outlined above.In some embodiments, the methylation of the cancer population is based on thefollowing equation: where βc(t) is the methylation fraction of the cancer cells at time t at that locus, k(t) is thenumber of cells with one allele methylated at time t, m(t) is the number of cells with bothalleles methylated at time t, S(t) is the number of cancer cells at time t.In some embodiments, the simulated model output is based on the followingequation: ^(^) = ^^^(^) + (1 − ^)^^where β(t) is the measured methylation fraction at time t, ρ is the purity of the cancersample, βc(t) is the methylation fraction of the cancer cells at time t at that locus, β(n) isthe average methylation of the contaminating normal cells.Inference method The stochastic forward model may be fitted to the methylation data using an inference method. In some embodiments, the inference method is a Bayesian inference method. In some embodiment, the inference method is a pseudo-marginal Bayesian inference method. In some embodiments, the inference method estimates a likelihood of the data. In some embodiments, the inference method estimates a log-likelihood of the data. In some embodiments, the inference method estimates a log-likelihood of the data under a finite number of simulations. In some embodiments, the log-likelihood is calculated as follows: ^log ^P^^^|^^, Δ, ^, ^^^ + log ^P^^^|^, ^, ^, ^, ^, ^, ^^^^^ where ℒ is the likelihood of the data according to the model, ^^^ is the logsumexp function^^^^(^^, … , ^^) = log^∑ ^ ^^^ ^, P^^^|^^ , Δ, ^, ^^ is the probability of observing the ^^ data pointoriginating from the ^^ simulated value assuming beta distributed noise, is the probability of observing a methylation value ^^, ^ is the cancer growth rate, ^ is theage of the patient when the MRCA emerged, ^, ^, ^, ^ are the epigenetic switching rates, ^^is the average methylation of the normal contaminating cells, and ∆, ^, ^ are model fittingparameters encoding the beta distributed noise.P603125PC0013 In some embodiments, the inference method comprises a penalisation function to penalise sizes of cancer that are non-physical. In some embodiments, the penalisationfunction (ℒ^^^^^^^^) is calculated according to the following equation: where Smin is the minimum plausible cancer size, Smax is the maximum plausible cancersize…, θ is cancer growth rate, ^ is the cancer growth rate, ^ is the age of the patient whenthe MRCA emerged, and ^ is the age of the patient at sampling.In some embodiments, the inference method comprises a noise function to account for noise. In some embodiments, the noise function is calculated according to the following equation: ^^ = (^ − ∆)^^ + ∆where xj is the beta distribution mean, ∆ and ^ are the offsets from 0 and 1 respectivelydue to background noise. In some embodiments, the inference method comprises a correction function to account for finite sampling bias. In some embodiments, the correction function is as follows: +0.005^^^^ ^^^^where ℒ^^^^is the bias in the likelihood due to the stochastic estimation of the likelihood,where ^ is precision of the beta distributed noise, where Nsim is the number of simulationsin the stochastic forward model. In some embodiments, the log likelihood is penalised and corrected for bias according to the following equation: log(ℒ′) = log(ℒ) + where ℒ′ is the corrected likelihood, ℒ is the uncorrected likelihood of the data accordingto the stochastic forward model, ℒ^^^^^^^^is the penalisation function and ℒ^^^^is the bias in the likelihood due to the stochastic estimation of the likelihood. The method generates a methylation distribution from which a historical growthrate of the cancer can be determined. The term “methylation distribution”, as used herein,is intended to mean the overall representation of the inferred methylation state of each ofthe selected fluctuating CpG loci between 0% methylated and 100% methylated acrossthe sampled population. Typically, the methylation distribution is W-shaped in cancersamples. Advantageously, the inventors have demonstrated that the specific shape of the W-shaped distribution is sensitive to the parameters of the stochastic forward model and therefore encodes the historical growth dynamics of the cancer.P603125PC0014 Predicting risk of progression In some embodiments, the method is a method of predicting overall survival of a subject with cancer. In some embodiments, the method is a method of predicting progression-free survival of a subject with cancer. In some embodiments, the method is a method of survival prognosis of a subject with cancer. The step of predicting the risk of progression of the cancer is based on the historicalgrowth rate of the cancer. This is intended to mean that the quantitative parameter ofthe historical growth rate that has been obtained from the methylation distribution is usedto infer the risk of progression. In some embodiments, predicting the risk of progression is further based onadditional parameters, such as the age of the subject when the most recent commonancestor (MRCA) of the sampling population emerged.In some embodiments, the risk of progression of the cancer (or assessing the progression-free survival of the subject, or assessing the overall survival of the subject) based on the historical growth rate of the cancer is predicted using a survival analysismethod. Any suitable survival analysis method may be used, which are known to theskilled person. In some embodiments, the survival analysis method is a Cox regressionmodel. In some embodiments, the survival analysis method is a Cox proportional hazardsregression model. In some embodiments, the survival analysis method is a multivariate Cox proportional model. In some embodiments, the survival analysis method is themaxstat statistics. In some embodiments, the survival analysis method is a Kaplan-Meiermethod. In some embodiments, the survival analysis method is a log-rank test. Treatment In some embodiments, the method further comprises treating the subject with atherapeutic drug, a palliative drug, radiotherapy, surgical resection of the cancer, or a combination thereof. In some embodiments, the method further comprises treating the subject with an anti-neoplastic agent. Typically, any anti-neoplastic agent that has activity versus asusceptible cancer being treated may be utilised in the method. Typical anti-neoplasticagents useful in the present invention include, but are not limited to, anti-microtubule agents such as diterpenoids and vinca alkaloids; platinum coordination complexes; alkylating agents such as nitrogen mustards, oxazaphosphorines, alkylsulfonates, nitrosoureas, and triazenes; antibiotic agents such as anthracyclins, actinomycins and bleomycins; topoisomerase II inhibitors such as epipodophyllotoxins; antimetabolites such as purine and pyrimidine analogues and anti-folate compounds; topoisomerase I inhibitors such as camptothecins; hormones and hormonal analogues; signal transduction pathwayP603125PC0015 inhibitors; non-receptor tyrosine kinase angiogenesis inhibitors; immunotherapeutic agents; proapoptotic agents; and cell cycle signaling inhibitors. Identification of fCpG loci In some embodiments there is further provided a method of identifying fluctuating CpG (fCpG) loci that are prognostic of the progression of a cancer, the method comprising: receiving a dataset of fCpG loci from DNA derived from cancer cells obtained from a largeplurality of cancer patient subjects with different cancer presentations; calculating arespective Laplacian Score feature selection metric for the fCpG loci in the dataset; and selecting as prognostic of the progression of a cancer those fCpG loci with a Laplacian Score greater than a predetermined threshold score. In some embodiments the selected fCpG loci are then used as the fCpG loci in amethod of predicting the risk of progression of a cancer of a subject as described aboveand elsewhere herein. In some embodiments the calculating step of the Laplacian Score further comprisesconstructing the nearest neighbour graph of the methylation samples in the dataset. In some embodiments, the Laplacian Score is a measure of how well each feature (i.e. CpG) preserves the nearest neighbour graph. In some embodiments, the relationship between methylation samples may bevisualised by performing a principal component analysis (PCA) on methylation values ofthe received fCpG loci in the dataset. In this case, in some embodiments at least the firstn, where n=> 1, principal components reveal disease specific clustering in methylation.In more detail, in some embodiments the CpG loci with the most informativefeatures for clustering and thereby having the lowest Laplacian Scores show disease-specific methylation, whereas the CpG loci that are least informative for clustering andthereby have the highest Laplacian Scores have either low variability or high variabilitythat are well-dispersed across the PCA. It is these CpG loci with the highest Laplacianscores that are selected as encoding evolutionary dynamics. In particular, in some embodiments, the CpG loci with the at least 50% and preferably at least 10% and more preferably at least 5% highest Laplacian Scores are selected as predictive of cancer progression. In a further embodiment the fCpG selection further comprises selecting as prognostic of the progression of a cancer those fCpG loci that are heterogeneous acrossdifferent cancer patient subjects with the same disease. In this case, the selectingpreferably further comprises accepting CpG loci with a top 5% of standard deviation ofmethylation value within a cancer type. In some embodiments, the selecting comprisesP603125PC0016 accepting CpG loci with average intra-disease standard deviation of methylation above 0.15 across cancer types. In a yet further embodiment, the fCpG selection further comprises selecting asprognostic of the progression of a cancer those fCpG loci that are equally likely to bemethylated or unmethylated. In this case, the selecting preferably further comprisesaveraging across all cases in a cancer type and accepting only fCpG sites with average methylation of approximately 0.5 in the cancer type. In most embodiments the method of identifying fluctuating CpG (fCpG) loci thatare prognostic of the progression of a cancer is computer implemented. Therefore, afurther aspect also provides a computer system for identifying fluctuating CpG (fCpG) locithat are prognostic of the progression of a cancer, the system comprising: one or moreprocessor units; and a computer readable storage medium storing one or more computerprograms such that when performed by the one or more processor units cause the computer system to operate in accordance with the method a method of identifying fluctuating CpG (fCpG) loci that are prognostic of the progression of a cancer as described above. In some embodiments there is further provided a method of identifying fluctuating CpG (fCpG) loci, the method comprising: receiving a dataset of CpG methylation values from DNA derived from cancer cells obtained from a large plurality of cancer patient subjects with different cancer presentations; calculating a respective Laplacian Score feature selection metric for the CpG methylation values in the dataset; and identifying fCpG loci with a Laplacian Score greater than a predetermined threshold score. In some embodiments the identified fCpG loci are then used as the fCpG loci in amethod of predicting the risk of progression of a cancer of a subject as described aboveand elsewhere herein. In some embodiments the calculating step of the Laplacian Score further comprisesconstructing the nearest neighbour graph of the CpG methylation values in the dataset.In some embodiments, the Laplacian Score is a measure of how well each feature (i.e. CpG) preserves the nearest neighbour graph. In some embodiments, the relationship between CpG methylation values may bevisualised by performing a principal component analysis (PCA) on the CpG methylationvalues. In this case, in some embodiments at least the first n, where n=> 1, principal components reveal disease specific clustering in methylation. In more detail, in some embodiments the CpG methylation values with the mostinformative features for clustering and thereby having the lowest Laplacian Scores showdisease-specific methylation, whereas the CpG methylation values that are leastinformative for clustering and thereby have the highest Laplacian Scores have either lowP603125PC0017variability or high variability that are well-dispersed across the PCA. It is these CpGmethylation values with the highest Laplacian scores that may be selected as fluctuatingCpGs (fCpG). In particular, in some embodiments, the CpG methylation values with the at least50% and preferably at least 10% and more preferably at least 5% highest Laplacian Scores are selected. In a further embodiment the fCpG identification further comprises selecting thosefCpG loci that are heterogeneous across different cancer patient subjects with the samedisease. In this case, the selecting preferably further comprises accepting CpG loci with atop 5% of standard deviation of methylation value within a cancer type. In some embodiments, the selecting comprises accepting CpG loci with average intra-disease standard deviation of methylation above 0.15 across cancer types. In a yet further embodiment the fCpG identification further comprises selectingthose fCpG loci that are equally likely to be methylated or unmethylated. In this case, theidentifying preferably further comprises averaging across all cases in a cancer type andaccepting only fCpG sites with average methylation of approximately 0.5 in the cancer type. In most embodiments the method of identifying fluctuating CpG (fCpG) loci iscomputer implemented. Therefore, a further aspect also provides a computer system foridentifying fluctuating CpG (fCpG) loci, the system comprising: one or more processorunits; and a computer readable storage medium storing one or more computer programssuch that when performed by the one or more processor units cause the computer system to operate in accordance with the method a method of identifying fluctuating CpG (fCpG) loci as described above. Methods of diagnosis The present disclosure also relates to a method of diagnosing a cancer in a subject,wherein the method comprises: obtaining DNA methylation data of fluctuating CpG loci ofa sample obtained from the subject; using the DNA methylation data to generate amethylation distribution, identifying the methylation distribution as statistically differentrelative to a reference methylation distribution, wherein a statistically different methylation distribution relative to the reference methylation distribution is indicative of the subject having cancer. Normal, healthy cells are polyclonal and so the distribution of fCpGs will beunimodal, but as a rapidly expanding cancer population derives from a common ancestor, the methylation pattern of that first cancer cell will become over-represented in DNAobtained from the subject. Thus, the fCpG loci appear “synchronised” in the cancer-derivedP603125PC0018DNA, increasing the variance of the bulk population fCpG methylation distribution.Advantageously, this additional growth information may help distinguish clinically relevanttumours from benign, allowing simultaneously early detection and risk-stratification. The above description in relation to the methods of predicting the risk ofprogression of a cancer of a subject is equally applicable to this aspect. In particular, the fluctuating CpG loci and the generation of the methylation distribution may be as described above. Statistical difference is often determined by comparing two or more populationsand determining a confidence interval and / or a p value. See, e.g., Dowdy and Wearden, Statistics for Research, John Wiley & Sons, New York, 1983. Exemplary confidence intervals of the present subject matter are 90%, 95%, 97.5%, 98%, 99%, 99.5%, 99.9% and 99.99%, while exemplary p values are 0.1, 0.05, 0.025, 0.02, 0.01, 0.005, 0.001, and 0.0001. The “reference methylation distribution” as used herein, may be a methylation distribution generated from DNA methylation data of the same fluctuating CpG lociobtained from a sample or multiple samples from a subject or subjects that do not havecancer, or a sample or multiple samples from a normal or healthy tissue. The “referencemethylation distribution” as used herein, may be a data set comprising the methylation distribution of the same fluctuating CpG loci in a normal or healthy individual or a population of normal or healthy individuals. In some embodiments, the method further comprises treating the subject with a therapeutic drug, a palliative drug, radiotherapy, surgical resection of the cancer, or a combination thereof, as discussed above. Brief Description of the Drawings Embodiments of the invention will now be further described by way of example only and with reference to the accompanying drawings. Figure 1. A schematic showing steps of the method according to an embodimentof the present invention. Figure 2. A schematic showing steps of the method according to an embodimentof the present invention. Figure 3. A schematic showing steps of the method according to an embodimentof the present invention. Figure 4. A schematic showing steps of an exemplary Bayesian inference method.Figure 5. A block diagram of a computer system for use with the risk predictionmethod according to and embodiment of the present disclosure.P603125PC0019 Figure 6. Selection and properties of fCpG loci. A: Illustration of how fCpGs enabletracking of lineage dynamics, describing three limiting cases of a population evolutionarystructure: (top) polyclonal, in which the most recent common ancestor (MRCA) exists farin the past; (middle) clonal, in which the MRCA is very recent; (bottom) ongoing evolution following a population bottleneck. In a polyclonal population, the different fCpGs (white –unmethylated, grey – heterozygous methylated, black – methylated) are all different dueto fluctuations that occurred since the MRCA. Thus, the average methylation of the population as each fCpG is roughly the steady state methylation value. Following a clonal expansion, all the fCpGs in the population are the same, so the average methylation reflects that of the methylation of the founder cell. However, as the population continues to grow ongoing fluctuations alter the barcodes in individual cells from the ancestral barcode, changing the distribution of bulk methylation values. B: fCpGs were selected by combining three filters: (i) CpGs with low intra-disease heterogeneity were removed as sites that were likely not fluctuating, (ii) CpGs with a mean methylation far from 0.5 were removed as likely belonging to CpG sites with skewed methylation / demethylation rates, and (iii) CpGs which preserved the nearest-neighbour graph (quantified using the Laplacian Score) were removed as sites that were likely under selection / strict regulation.C: A hierarchically-clustered heatmap of the 978 fCpGs. D: (top) Example methylationdistributions of the fCpGs from healthy B cells, B-ALL cells and whole blood from a patient in remission from B-ALL. (bottom) Boxplots (whiskers extending to ±1.5×IQR) displaying the standard deviation of the methylation distributions separated by cell / disease type. E: A heatmap showing the log2 fold change in chromatin status at the site of fCpG loci compared to other CpG loci on the 450K Illumina array in healthy lymphoid cells, MCL,CLL and MM samples. White stars signify the degree of statistical significance (^^ tests, hscorrection). F: Expression analysis on the genes associated with fCpG and non-fCpG loci:(left) a Q-Q comparing non-fCpG and fCpG associated genes, demonstrating fCpG associated genes have a lower expression (in transcripts per million), (middle) a comparison between fCpG and non-fCpG associated genes in a single CLL sample (P=3.96e10-11, Wilcoxon test), (right) the expression of fCpG associated genes separated by discretised methylation status (all p>0.05, Wilcoxon test). Figure 7. A stochastic mathematical model well-describes the observed patternsin cancer fCpG data. A: An illustration of the mathematical model describing how the fCpGdistributions vary with the evolutionary parameters of a cancer. (top) The model is splitinto 2 phases, prior to the MRCA (^) whereby methylation changes occur in a single cell,and following the MRCA where the population grows exponentially (^). (bottom) At each time step, epigenetic switching is allowed to occur between the 3 possible states (rates ^,^, ^ and ^). B: An example simulation run, which has the characteristic “W-shape”P603125PC0020observed in lymphoid cancer data (e.g. fig. 1D). C: Illustrative limiting cases of: (top) avery old cancer in which the fCpGs have had a long time to desynchronise vs a youngcancer in which the fCpGs are largely synchronised, (bottom) a rapidly growing cancer inwhich the rate at which cells divide is high compared to the switching rate vs a slowlygrowing cancer in which the effect of stochasticity is amplified. D: The resulting posteriordistributions after fitting the simulated data in fig. 2B with the Bayesian inference method. The posterior (orange) displays tightening around the ground truth parameter values (red)compared to the prior (grey). E: The posterior predictive distribution (orange)superimposed upon the simulated data (blue), showing an excellent fit. Figure 8. EVOFLUx reveals the evolutionary dynamics of lymphoid cancers. A: Ascatter plot of the inferred growth rate (^) vs effective population size (^^, ^^(^^^)) per sample, using the posterior median as a descriptive summary statistic. Coloured accordingto cancer type. B: A scatter plot of the inferred time since the MRCA (^) vs the meanepigenetic switching rate (averaged across different rate parameters) per sample. C-E: 3example histograms samples in which the neutral (C), subclonal selection (D) andindependent clonal origins (E) models are preferred. F: A histogram showing the variantallele frequency spectrum of matched whole exome sequencing (WES) data, with the results of subclonal deconvolution (performed using Mobster (Caravagna et al. 2020)) overlaid. In this sample, there is strong evidence of an emerging subclone (C2), in linewith the inference from EVOFLUx. G: Boxplots comparing the distribution of subclonalweightings inferred by EVOFLUx in samples called as neutral vs under subclonal selectionvia the WES data. Figure 9. fCpGs allow for phylogenetic reconstruction of longitudinal lymphoidcancer samples. A: Timeline of two CLL patients with samples collected longitudinally, annotated with treatment received. Circles represent methylation array (blue: 450K, orange EPIC), squares represent whole genome sequencing, treatment is represented with a green vertical line and Richter transformation (RT) is represented with a vertical pink line. B: (left) The reconstructed phylogenies of the relationship between samples, annotated with the clinical classification of each sample. The black triangles represent the time that occurred since the most recent common ancestor, taken as the posterior medianof ^ − ^ from the single-sample EVOFLUx inferences. (right) A heatmap representing the978 fCpG loci, with the colour a representing the fraction methylated (0% blue, 100% red). Figure 10. The evolutionary history of a tumour varies between molecular subtypeand is predictive of clinical outcome. A-C: The inferred growth rate (top) and effectivepopulation size (Ne, bottom) of individual cancer samples separated by molecular subtype in B-acute lymphoblastic leukaemia (A, BCP-ALL), mantle cell lymphoma (B, MCL) andP603125PC0021 chronic lymphocytic leukaemia (C, CLL). Differences between subtypes were tested using Mann-Whitney U tests, with Holm-Bonferroni corrections applied. In A, the 11q23 / MLL subtype had significantly higher growth rate than each of the other B-ALL subtypes (all P<0.005), but for ease of presentation, only the comparison with t(1;19) is shown here.D: Univariate survival analysis of the time to first treatment (TTFT, blue) and overallsurvival (OS, red) in the discovery CLL cohort for evolutionary variables inferred viaEVOFLUx. E: Kaplan-Meier curves comparing the TTFT between patients with high vs lowinferred cancer growth rates, separated by IGHV mutational status. F: Multivariate Coxregression of the TTFT shows the cancer growth rate is significant when controlling for IGHV status, TP53 status and age. Figure 11. Heatmaps of putative fCpG loci by disease. Clustered heatmaps (hierarchical with average linkage and a Euclidean metric) of our set of 978 pan-lymphoidfCpGs with disease-specific subtypes annotated in: A chronic lymphocytic leukaemia(CLL), B B-cell acute lymphoblastic leukaemia (B-ALL), C mantle cell lymphoma (MCL),and D diffuse large B-cell lymphoma – not otherwise specified (DLBCL-NOS).Figure 12. Genomic properties of fCpG loci. A: fCpG loci are similarly distributedacross the genome when compared to non-fCpG loci also on the 450K Illumina array, except a significant enrichment on chromosome 19 (p=2.2e-5, pairwise ^^tests, holm-sidak (hs) correction). B: Paired bar charts demonstrating that fCpG loci are enriched onthe shores of CpG Islands but depleted on both the islands themselves and in the OpenSea. C: A comparison of the mutual overlap between different epigenetic clocks [refs] andfCpGs. D: fCpG loci are significantly less likely to be associated with a gene (annotationsprovided by UCSC, p<0.001, ^^test). E: fCpG loci are significantly less likely to be associated with a promotor, and much more likely to be either unclassified or unclassified but associated with a specific cell type (annotations provided by ENCODE (Dunham et al. 2012)) (all significant p values < 10-10, pairwise ^^tests, hs correction). Figure 13. Local genomic neighbourhood surrounding fCpG loci. A: Heatmaps ofthe methylation values of sorted healthy B- and T-cells (Loyfer et al. 2023) of fCpG loci,CpG loci within 1-20 bps and 21-40 bps. In polyclonal populations, measurements of bulk sorted population are expected to have intermediate methylations, whereas CpG loci understrict regulation are expected to be hypo- or hypermethylated. B: The average absolutedifference from 0.5 for methylation values from fCpG loci and CpG loci in the localneighbourhood, in different B- and T- cell populations. The fraction of hypo- and hyper-methylated CpG loci increases as a function of distance from the reference fCpG locus. Figure 14. Gene-set enrichment analysis of fCpG-associated genes. Gene-setdepletion (A) and enrichment (B) analysis of genes associated with fCpG loci according to UCSC [ref] using gProfiler, with the target gene-space corresponding to all genes with anP603125PC0022 associated CpG locus present on the 450K Illumina methylation bead array. The GeneOntology biological processes (Ashburner et al. 2000) and the Human Protein Atlas (Uhlénet al. 2015) databases were employed. fCpG associated genes were significantly enrichedin developmental pathways but depleted in lymph, testis, and tonsil tissues. Figure 15. Example EVOFLUx posterior and fit. A: An example of the posteriorresulting from running EVOFLUx on a CLL sample – a pairs plot showing the marginal(diagonal) and pairwise (off-diagonal) inferred posterior (orange) distributions, with the prior distributions overlaid (gray). The posteriors show marked tightening compared tothe priors, demonstrating the parameters are well informed by the data. B: A histogramof the fCpG methylation distribution (blue) with the posterior predictive of the model fit overlaid (orange). Figure 16. EVOFLUx captures lymphoid cancers’ evolutionary histories andmethylation epimutation dynamics. A-D: Boxplots (whiskers extending to ±1.5×IQR)showing the distribution of inferred growth rate (A), effective population size (B), time since the most recent common ancestor (C), and mean epigenetic switching rate (D) by disease. For interpretability, only a subset of the pairwise p values are annotated (Mann-Whitney U tests, hs correction). E: Linear regression between the growth and epigeneticswitching rates separated by cancer types. There is a positive association in B-ALL (p= 2.4e-98, R2=0.44) and T-ALL (p=5.9e-06, R2=0.22), a weak negative association in MM (p=1.6e-05, R2=0.18) and no association in CLL (p=0.060, R2=0.005) or the otherentities. F: A pairs plot comparing the four epigenetic switching rates (μ, υ, γ and ζ) ineach of the cancers, coloured by disease. Figure 17. Simulations of non-neutral evolutionary dynamics. A-B: Simulationscomparing the fCpG methylation distribution in the neutral (orange) case vs subclonal selection (left, green) and independent clonal origins (right, red) in the case with no noise(A) and with noise added (B). C: Illustrations of three alternative evolutionary models:neutral evolution, in which the cancer population with a MRCA emerging at time ^ growsexponentially at rate ^; subclonal selection, in which an initial population emerging at time ^^growing at rate ^^is outcompeted by a fitter subclonal emerging at rate ^^with growth rate ^^; and independent clonal origins, in which the fitter clone emerging at time ^^withgrowth rate ^^bears no clonal relationship to the initial clone. D: A heatmap of showinghow the subclonal model weight of a simulation varies with the fraction of the populationmade up of the fitter subclone and the growth advantage of the fitter subclone. E: Acomparison of the inferred relative growth advantage of a simulation vs the actual ground truth value, where the colour of the point reflects the fraction of the subclonal population. Figure 18. Inference and validation of subclonal selection using fCpG loci. A: Barchart comparing the fraction of cancer samples identified as subclonal by EVOFLUx acrossP603125PC0023disease (pairwise ^^ tests, holm-sidak (hs) correction). B: Logistic regression between theprobability of a cancer being identified as subclonal via running our subclonal deconvolution tool, MOBSTER, on WES data and the number of mutations detected in the sample. Regression run separately for samples also identified as subclonal / neutral via EVOFLUx. C: Logistic regression between the inferred subclonal weighting by EVOFLUx and the number of mutations detected in WES finds no association between the two (p>0.05,logistic regression). D: Boxplots comparing the distribution of independent modelweightings inferred by EVOFLUx in samples containing multiple IgHV rearrangements (likely independent cancers) vs those with only a single IgHV rearrangement (likely a single clonal origin). Figure 19. WGS SNV phylogenies validate fCpG phylogenies in CLL. Phylogeniesreconstructed on matched high-depth WGS SNV data have a similar topology and absolute timing of branch points as to those reconstructed using fCpGs (fig 4). Figure 20. fCpG phylogenies in ALL. (left) The reconstructed phylogenies of therelationship between samples collected longitudinally in individual ALL patients, annotated with the clinical classification of each sample. The black triangles represent the time thatoccurred since the most recent common ancestor, taken as the posterior median of ^ − ^from the single-sample EVOFLUx inferences. (right) A heatmap representing the 978 fCpG loci, with the colour a representing the fraction methylated (0% blue, 100% red). Figure 21. Additional subtype comparisons of the inferred evolutionaryparameters. The inferred growth rate (A&C) and effective population size (B&D) ofindividual cancer samples separated by molecular subtype in diffuse large B-cell lymphoma (A&B, DLBCL-NOS) and monoclonal B-cell lymphocytosis (C&D, MBL). Differences between subtypes were tested using Mann-Whitney U tests. Figure 22. Correlations in CLL evolutionary parameters and additional survivalanalysis. A: A pairs plot showing the marginal relationships between the observedevolutionary parameters in the discovery cohort as scatter plots (lower left triangle), histograms (leading diagonal) and summary correlation coefficients (stars indicate significance). Patients separated by IGHV mutational status (unmutated orange, mutatedpurple). B: Kaplan-Meier curves comparing the OS between patients with high vs lowinferred effective population sizes (Ne), separated by IGHV mutational status. C: Multivariate Cox regression of the OS shows the Neis significant when controlling for IGHV status, TP53 status and age in the discovery cohort. Figure 23. Survival analysis in validation cohort. A: Univariate survival analysis ofthe time to first treatment (TTFT, blue) and overall survival (OS, red) in the validation CLL cohort for evolutionary variables inferred via EVOFLUx. Note this cohort contains a mixture of treated and untreated samples, of which only the untreated samples were included inP603125PC0024 the TTFT analysis. B: Kaplan-Meier curves comparing the TTFT between patients with high vs low inferred cancer growth rates in the validation cohort, separated by IGHV mutationalstatus. C: Multivariate Cox regression of the effect of the cancer growth rate on the in thevalidation cohort, controlling for IGHV status, TP53 status and age. Figure 24. Allele-specific epigenetic switching. A: A histogram displaying the fCpGmethylation value distribution of a CLL sample (blue) with a posterior predictive of a beta mixture model (with 3 components) overlaid (gray). The vertical red lines indicate themode of the 3 inferred beta mixture components. B: A scatterplot comparing the distancebetween the inferred modes of the two homozygous peaks and the heterozygous peak in the CLL cohort. The distance between the homozygous unmethylated and theheterozygous peak is larger in the vast majority of samples. C: A schematic illustratingtwo competing models, one in which the methylation of a given allele depends on the methylation status of the other (top, 4 parameters) and a model in which the two allelesare independent (bottom, 2 parameters). D: (top) A histogram displaying the fCpGmethylation value distribution of a CLL sample (blue) with the model fits of the allele specific (orange) and allele invariant epigenetic (lilac) switching models overlaid. (bottom) A leave one out cross validation model comparison, showing the allele specific model isstrongly preferred. E: Scatterplots showing the association between tumour purity and thestandard deviation of fCpG methylation in lymphoid cancers. Figure 25. Loglikelihood correction. A: Repeated calculations (n=100) of theuncorrected synthetic loglikelihood (LL) value of simulated data calculated at the parameter values used to simulate the data, as a function of the number of simulations (nsim) used in the LL calculation. Using a finite nsimsystematically underestimates the trueLL. B: The bias corrected synthetic LL, the mean of which does not vary as a function ofnsim. Figure 26. A summary of methylation array samples. Abbreviations: Peripheralblood mononuclear cells (PBMCs); precursor T-acute lymphoblastic leukemias (T-ALL); precursor B-acute lymphoblastic leukemias (B-ALL); mantle cell lymphoma; monoclonal B-cell lymphocytosis (MBL); chronic lymphocytic leukemias (CLL); Richter transformation (RT); diffuse large B-cell lymphoma, not otherwise specified (DLBCL-NOS); monoclonal gammopathy of undetermined significance (MGUS); and multiple myeloma (MM). Female (F); Male (M); bone marrow (BM); peripheral blood (PB); lymph node (LN). Figure 27. Copy number alterations are rare but are consistent with fCpGsbehaving in an allele specific manner. A: Copy number profiles in the CLL (top) and MCL(bottom) cohorts. Rows represent individual patient samples; columns represent genomic locations. Diploid regions are shown in white, gains in red and losses in blue. B: An illustration of the possible combinations of fCpG states in monosomic, diploid and trisomicP603125PC0025regions, along with the corresponding fraction methylation value. C-D: Examplehistograms showing the fCpG methylation distribution of fCpGs located on just monosomic(C) or trisomic (D) regions within 2 patient samples. E-F: Heatmaps of fCpGs organisedby genomic location in the CLL (E) and MCL (F) cohorts, with fCpGs present on non-diploid regions coloured in black. Figure 28. Dynamics of fCpGs in normal lymphoid cells during aging. A: Mean and standard deviation of fCpGs as a function of age estimated by the Horvath clock in normallymphoid cells included in the study. B: Same as A but with real age of donors in 2independent validation studies. C: Mean and standard deviation of whole-blood samples in 4 independent studies covering the entire human lifespan. P-values are adjusted bycellular fractions calculated with EpiDISH package using 12 cell types. D: Boxplots ofstandard deviation of fCpGs of samples in C divided by age groups together with lymphoid tumors. P-values were derived using two-tailed t tests. Boxplots whiskers represent ±1.5 IQR. Figure 29. Testing the effect of non-exponential growth on EVOFLUx inferences.A: Plots showing the relationship between the population size and the carrying capacity inthe logistic growth model. B: Histograms showing the fCpG methylation distributions of10,000 fCpGs simulated under the logistic growth model with varying carrying capacity.C-F: The EVOFLUx posterior median and 95% credible interval when run on a subset of2,000 of the simulated fCpGs above as a function of the carrying capacity for the growth rate (C), most recent common ancestor age (D) and epigenetic switching rates (E&F). The dashed red line represents the ground truth parameter value. Figure 30. Testing the effect of variable epigenetic switching rates on EVOFLUxinferences. A: Histograms showing the relationship between the standard deviation of theepigenetic switching rates (assumed to be lognormal) and the simulated switching rate foreach fCpG. B: Histograms showing the fCpG methylation distributions of 10,000 fCpGssimulated under the heterogenous epigenetic switching model with varying switching ratestandard deviation. C-F: The EVOFLUx posterior median and 95% credible interval whenrun on a subset of 2,000 of the simulated fCpGs above as a function of the switching rate standard deviation for the growth rate (C), most recent common ancestor age (D) and epigenetic switching rates (E&F). The dashed red line represents the ground truth parameter value. Figure 31. Comparison of lymphocyte counts with EVOFLUx growth rates. A: Example time courses showing the longitudinal lymphocyte counts (a proxy for the numberof CLL cells in the blood) prior to treatment (green vertical line). B: Comparison of therate of change in lymphocyte count between U-CLL and M-CLL (p=2.2e-12). C: Linear regression between the EVOFLUx inferred early evolutionary growth rate and the rate ofP603125PC0026change in lymphocyte number (p=2e-5). D: Forest plot showing the result of a multivariateCox proportionate hazard regression on the time to first treatment for 229 CLL patients. Figure 32. The effect of different aspects of the fCpG selection process. A-B:Clustered heatmaps (hierarchical with average linkage and a Euclidean metric) showing CpGs that are excluded from the fCpG selection by the Laplacian Score filter (A) or the standard deviation filter (B). Figure 33. Selection of biased fCpGs. A-B: Heatmaps showing the fCpGmethylation distributions for fCpGs selected where the mean methylation between 0.15and 0.35, or 0.65 and 0.85 respectively. C-D: Example histograms showing thedistribution of biased hypo- / hyper-methylated fCpGs in a CLL sample (SCLL-006). E-F: Example histograms showing the distribution of biased hypo- / hyper-methylated fCpGs ina normal whole blood sample (WB-01). G-H: Scatterplots showing the relationshipbetween the fCpG standard deviation in between the original unbiased fCpGs and the biased hypo- / hyper-methylated fCpGs respectively. Figure 34. fCpGs and SNPs show distinct methylation dynamics. A: Commoncontrol SNP probes from Illumina 450k and EPIC arrays in normal and neoplastic lymphoid cells showing distinct patterns compared to fCpGs. B: Two B-ALL patients with longitudinal samples including diagnosis, remission and relapse. Pairwise comparisons between diagnosis and remission and remission and relapse are shown for 59 control SNP probes(top panels) and 59 randomly selected fCpGs (bottom panels). C: Same as in B but fortwo CLL patients in diagnosis, progression, relapse and Richter transformation (RT) in the last sample timepoints. Using matched WGS, we show that there are changes in the tumorsubclonal composition in T1 vs T2 and T3 vs T4 in SCLL-12, whereas in SCLL-019 thebiggest changes occur between T3 vs T4 and T4 vs T5 (Fig 5). Fig. 5 also show themethylation values of the full fCpGs set, and supplementary table 12 the full clinical annotation of these 2 patients. Figure 35. Concordance of bead array and nanopore methylation calls. A: A barplot showing the fraction of fCpGs that were called as unmodified cytosine (C), 5- Methylcytosine (5mC), 5-hydroxymethylcytosine (5hmC) or a non-canonical base (A, G orT). B: A scatterplot showing that the variant allele frequency (VAF) of non-canonical fCpGcalls is inversely proportional to the coverage at that locus. C: The proportion of pairwise phased read comparisons in which the same fCpG site is in the same methylation state in normal B cells, split by whether the pattern of variants in a 50bp window surrounding that fCpG site is the same. D: Histograms showing the characteristic “w-shaped” fCpG methylation distribution in matched methylation bead array (top) nanopore (bottom) data.E: A heatmap showing the fraction methylated (5mC or 5hmC) in 2 matched CLL and RTsamples, and 6 samples from normal B cells. F: A scatterplot showing the correlationP603125PC0027 between the fCpG fraction methylated in matched nanopore and bead array data. Points are coloured according to the absolute difference between the two measurements. Figure 36. Validation of bead array methylation measurement of fCpGs via wholegenome bisulfite sequencing. A: Heatmaps showing the fCpG methylation measured viawhole genome bisulfite sequencing (WGBS) of 19 cancer samples (left) and 19 normal B cell samples (right). Only a subset of 730 fCpGs which had a high average coverage wereincluded. Sites with <10 reads are coloured black. B-C: Example histograms of the fCpGmethylation distribution for a mantle cell lymphoma (MCL, B) and a normal B cell (C) sample. Only fCpGs with ≥10 reads are included. Figure 37. E: fCpG methylation of phased Oxford nanopore long-reads of 2 CLLand 2 matched RT samples, 3 sorted mature and 3 sorted naïve B cells samples. Long reads which cover the 4 fCpGs within the region chr7:25,854,120-25,856,220 are shown.F: (left) A box plot comparing the mean intra-haplotype Hamming distance of phasedreads from cancer and normal samples (p=0.038, MW-U test). (right) Paired comparisonof the mean intra- and inter-haplotype Hamming distances in 1 CLL and 2 RT samples(p=0.030, paired T-test). One of the CLL samples only had reads from one haplotype, andthus was excluded from this plot. Figure 38. Paired B-ALL and remission samples. A: An example histogram of alongitudinal pair of samples showing the fCpG methylation distributions of the diagnostic sample (W-shaped) and once the patient is in remission following treatment (unimodal).B: A paired boxplot showing the standard deviation of the fCpG methylation distributionsof B-ALL samples is greater than their matched remission sample ((p=9.6e-39, paired t-test). C: A comparison of the standard deviation of the fCpG methylation distributions ofB-ALL patients in remission (i.e. no cancer cells present in blood) vs normal whole blood (p=0.067, MW-U test). Figure 39. Local correlations between fCpGs. A: fCpG methylation of phasedOxford nanopore long-reads of 2 CLL and 2 matched RT samples, 3 sorted mature and 3 sorted naïve B cells samples within the region chr6:31,180,590-31,180,895 (~300bps)are shown. There is high correlation between the methylation status of neighbouring fCpGson the same read. B: A heatmap showing the same 6 fCpGs assessed via methylationbead array, showing strong correlation on the bulk methylation values at these sites. C: Aheatmap showing the pairwise correlation coefficient calculated from 2,204 bulk methylation bead array of every fCpG with every other fCpG on the same chromosome,arranged by genomic position. D: A scatterplot showing the relationship between thepairwise correlation coefficient of every fCpG with other fCpG on the same chromosome, and the genomic distance between them.P603125PC0028 Figure 40. (top) A Fisher plot showing the change in the subclonal composition within a simulated set of longitudinal samples. (bottom) Scatter plots showing the marginal fCpG methylation distribution between simulated pairwise longitudinal samples. Points are coloured according to the difference in methylation between the first and second timepoint. Figure 41. The effect of the number of fCpGs on the inferred EVOFLUx parameters. A: Example histograms showing the distribution of all 978 fCpGs (blue) in a CLL sample, with a randomly downsampled subset of 10% (left), 50% (middle) and 90% (right)superimposed (yellow). B-E: Plots showing the effect of the number of downsampledfCpGs included in the inference process against the inferred growth rate (B), time since the most recent common ancestor (C), mean epigenetic switching rate (D), and the effective population size (E). For each set of 10% increment, 10 replicate fCpG subsets were generated (grey dots) and the EVOFLUx inference repeated. The mean and standard error of the replicates are represented with a blue dot and error bars respectively. Figure 42. Correlations between the patient age and the fCpG epigenetic switchingrate. Linear regressions between the patient age at sampling and the mean inferredepigenetic switching rate in chronic lymphocytic leukaemia (CLL, A), mantle cell lymphoma (MCL, B) multiple myeloma (MM, C), diffuse large B cell lymphoma (DLBCL, D), B cell acute lymphoblastic leukaemia (B-ALL, E) and T cell acute lymphoblastic leukaemia (T- ALL, F). Figure 43. EVOFLUx captures lymphoid cancers’ evolutionary histories andmethylation epimutation dynamics. A-D: Boxplots (whiskers extending to ±1.5×IQR)showing the distribution of inferred growth rate (A), effective population size (B), time since the most recent common ancestor (C), and mean epigenetic switching rate (D) by disease. For interpretability, only a subset of the pairwise p values are annotated (Mann-Whitney U tests, hs correction). E: Linear regression between the growth and epigeneticswitching rates separated by cancer types. There is a positive association in B-ALL (p= 2.4e-98, R2=0.44) and T-ALL (p=5.9e-06, R2=0.22), a weak negative association in MM (p=1.6e-05, R2=0.18) and no association in CLL (p=0.060, R2=0.005) or the otherentities. F: A pairs plot comparing the four epigenetic switching rates (μ, υ, γ and ζ) ineach of the cancers, coloured by disease. Figure 44. EVOFLUx in robust to removal of fCpGs present on non-diploid. A:Regression plots between the parameters inferred by running EVOFLUx on all 978 fCpGs (x-axis) vs just those fCpGs present on diploid regions (y-axis) in the CLL (top) and MCL (bottom) cohorts. B: Regression plots between the number of fCpGs located on non-diploid regions (x-axis) vs the log-fold change in the parameter inferred from the inference runP603125PC0029 on fCpGs on diploid regions compared to the original inference (y-axis) in the CLL (top) and MCL (bottom) cohorts. Figure 45. Validation of EVOFLUx subclonality calling using WGS. A: Boxplotscomparing the distribution of subclonal weightings inferred by EVOFLUx in samples called as neutral vs under subclonal selection within WGS data. B: Logistic regression between the probability of a cancer being identified as subclonal via running our subclonal deconvolution tool, MOBSTER, on WGS data and the number of mutations detected in the sample. Figure 46. Receiver operating characteristic (ROC) curves showing the accuracy at predicting the subclonality of WGS data of competing logistic classifiers trained on just the EVOFLUx subclonality weighting, just the WES subclonality call and a model trained on both sources of data. Figure 47. fCpGs allow for phylogenetic reconstruction of longitudinal lymphoidcancer samples. A-B: (top) Timelines and Fisher plots derived from WGS of two CLLpatients with samples collected longitudinally, annotated with treatment received. The Richter transformed (RT) clone is shown coloured puce. (middle) Scatter plots showing the marginal fCpG methylation distribution between pairwise samples. Points are coloured according to the difference in methylation between the first and second timepoint. (bottom) The reconstructed phylogenies of the relationship between samples, annotated with the clinical classification of each sample. The black triangles represent the time thatoccurred since the most recent common ancestor, taken as the posterior median of ^ − ^from the single-sample EVOFLUx inferences. The methylation fraction of the 978 fCpG lociare presented as heatmaps (0% blue, 100% red). C-D: (left) Two example longitudinalseries showing the development of the fCpG distribution from diagnostic B-ALL, through remission and relapse. (right) Scatter plots showing the marginal fCpG methylation distribution between diagnosis and relapse. Figure 48. Phylogenetic trees based on random CpGs. Phylogenetic treesreconstructed from random subsets of 1,000 CpGs (10 replicates) for case 12 (A) and case 19 (B). Figure 49. WGS SNV phylogenies validate fCpG phylogenies (Case 12). A-E:Phylogenies reconstructed on matched WGS SNVs (A), fCpGs (B), Horvath’s CpGs (C)(Horvath 2013), epiCMIT (D) (Duran-Ferrer et al. 2020) and a random subset of 1,000CpGs (E). F: Comparison of the Robinson Fould (RF) distance between the WGS SNV treeand the methylation-based trees. G: A scatterplot showing the inferred tip coalescenttimes (posterior median and 95% CI) for the WGS SNV tree vs each of the methylation-based trees. H: Comparison of the mean error between the tip coalescent times betweenthe WGS SNV tree and the methylation-based trees.P603125PC0030 Figure 50. WGS SNV phylogenies validate fCpG phylogenies (Case 19). A-E:Phylogenies reconstructed on matched WGS SNVs (A), fCpGs (B), Horvath’s CpGs (C)(Horvath 2013), epiCMIT (D) (Duran-Ferrer et al. 2020) and a random subset of 1,000CpGs (E). F: Comparison of the Robinson Fould (RF) distance between the WGS SNV treeand the methylation-based trees. G: A scatterplot showing the inferred tip coalescenttimes (posterior median and 95% CI) for the WGS SNV tree vs each of the methylation-based trees. H: Comparison of the mean error between the tip coalescent times betweenthe WGS SNV tree and the methylation-based trees. Figure 51. fCpG phylogenies in ALL. (left) The reconstructed phylogenies of the relationship between samples collected longitudinally in individual ALL patients, annotated with the clinical classification of each sample. The black triangles represent the time thatoccurred since the most recent common ancestor, taken as the posterior median of ^ − ^from the single-sample EVOFLUx inferences. (right) A heatmap representing the 978 fCpG loci, with the colour a representing the fraction methylated (0% blue, 100% red). Figure 52. Genotype-phenotype driver mutation map. The inferred growth rate(A) and effective population size (B) of individual CLL samples separated by driver mutational status for common driver mutations (TP53, SF3B1, NOTCH1, ATM, POT1 and IGLV321.R110, del11q22.3, del13q14.3, del17p13.1 and trisomy12). Differences between genotypes were tested using Mann-Whitney U tests and Benjamini-Hochberg FDR corrected. Tests were performed on U-CLL and M-CLL patients separately to remove differences solely to do with IGHV status. Figure 53. D-E: The inferred growth rate (left) and effective population size (right)for patients with a TP53 driver mutation, controlling for IGHV status. Figure 54. Inferred effective population size is predictive of progression free survival in colorectal cancer. A Kaplan-Meier plot of the progression free survival of 114colorectal cancer patients, split by the median effective population size (ScancerLog – blueis low population size, orange high). 95% confidence intervals are shown as shaded regions and were calculated using the exponential Greenwood method (Kalbfleisch and Prentice 1980). Censored patients are indicated by a cross, and the number of patients at risk over time in each group are summarised below. Figure 55. Multivariate Cox regression of progression free survival in colorectalcancer. A forest plot showing the calculated coefficients of a multivariate Cox proportionalhazards regression of progression free survival against the inferred effective population size (ScancerLog) and tumour purity. 95% confidence intervals are shown as bars and p values are annotated. Description of the EmbodimentsP603125PC0031 Overview The inventors have developed a novel stochastic model and associated inference method to infer the historical growth dynamics of an individual’s cancer using fluctuating methylation clocks (FMCs). In the case of chronic lymphocytic leukaemia (CLL), the inferred growth rate holds valuable prognostic information, enabling patient stratification beyond that of standard clinical practice. The inventors previously identified that clonal expansions leave a clear signatureon the patterns of FMCs, which are specific CpG loci which fluctuate between homozygousmethylated, heterozygous, and homozygous demethylated states within individual cells(Gabbutt et al. ‘Fluctuating Methylation Clocks for Cell Lineage Tracing at High TemporalResolution in Human Tissues’. Nature Biotechnology 2022, January, 1–11.). In a polyclonalpopulation, these FMCs are effectively desynchronised, leading to the average methylation being randomly distributed around the steady-state methylation level. Immediately following a clonal expansion, however, the methylation of the entire population resembles that of the progenitor cell, leaving a distinctive “w-shaped” signature on the distribution of FMC methylation patterns of cancer samples. The hematopoietic system is maintained by ~105-106stem cells that share a very early common ancestor (Lee-Six, Henry, Nina Friesgaard Øbro, Mairi S Shepherd, Sebastian Grossmann, Kevin Dawson, Miriam Belmonte, Robert J Osborne, et al. 2018. ‘Population Dynamics of Normal Human Blood Inferred from Somatic Mutations’. Nature 561 (7724): 473–78.), hence in healthy blood the fluctuating CpG (fCpG) loci are largely desynchronised and the FMC distribution is generally unimodal about the steady-state methylation value. However, in exponentially growing blood cancers, all the malignant cells share a recent common ancestor and thus the descendent cells should largely share the FMC methylation patterns of the ancestor, leading to a characteristic “w-shaped” distribution. The inventors hypothesised that the specific shape of a given patient’s FMC distribution would encode the growth dynamics of that patient’s cancer. The methylome can be measured using standard Illumina 450K / EPIC arrays, which probe the methylation status of ~450,000 / 850,000 CpG loci respectively, returning the fraction of methylation at a given CpG locus (often termed the beta value). The inventors employed standard single-sample Noob normalisation to pre-process the array data (Fortin, Jean-Philippe, Timothy J Triche Jr, and Kasper D Hansen. 2017. ‘Preprocessing, Normalization and Integration of the Illumina HumanMethylationEPIC Array with Minfi’.Bioinformatics 33 (4): 558–60.).The inventors have refined how they select fCpG loci to apply in cases where they do not have multiple samples from the same individual and in cases in which there exist multiple molecular subtypes of disease. This involved identifying CpG loci that tended toP603125PC0032 be heterogeneous across different patients with the same disease, but which did not tend to have large average differences between cancers with different molecular subtypes. As predicted from theory, the resulting FMC methylation distributions had w-shaped distributions but the patterns of methylation in individual fCpG loci were uncorrelated and varied from cancer-to-cancer. The inventors developed a forward model to map from a set of biologically meaningful parameters to the FMC distribution. These parameters were: (i) the growth rate of the cancer (assuming an exponentially growing population); (ii) the age of the patient when the most recent common ancestor (MRCA) of the sampling population emerged; (iii) the four epigenetic switching rates, corresponding to the four possible transitions between homozygous unmethylated, heterozygous methylated and homozygous methylated; and (iv) the average methylation level of contaminating normal cells. The inventors further employed a number of fitting parameters that were not biologically meaningful, but rather related to the noise associated with the array. The forward model involves tracking the number of cells in the population that arein the homozygous (de)methylated states or heterozygous state for each of ^^^^loci, and how these populations change over time as both a result of the exponential growth of thepopulation and the ongoing epigenetic switching. It consists of effectively two phases, thetime prior to the MRCA of the cancer emerged and the time between the MRCA and the sampling time. Prior to the MRCA emerging, the ongoing fluctuations in the methylation of fCpG within a single cell can be calculated using standard matrix exponentiation, fixing the FMC distribution in the founder cell according to the epigenetic switching rate. Thesecond phase of the model is split into discrete time steps, in which the total populationof the cancer grows according to a deterministic exponential model, after which transitions between the different homozygous / heterozygous populations occur stochastically. Once the simulation has reached the sampling age, the methylation fraction of the population at a particular locus can be calculated by summing the number of methylated alleles in the population and dividing it by the total number of alleles. The inventorsaccounted for the fact the sample was not purely cancer cells by modifying the methylationfraction of the cancer cells by the average methylation of the contaminating cells, weightedby the known sample purity. Finally, to mimic the noise introduced by the methylationarray, the inventors added beta distributed errors centred upon the “true” methylationvalue in the population. This forward model qualitatively recapitulates the qualitative patterns observed in patient data, including the characteristic w-shape. To fit this model to data, the inventorsP603125PC0033employed a pseudo-marginal Bayesian approach, using their stochastic simulations toestimate the log-likelihood of the data under a finite number of simulations and applying a correction to account for the finite sampling. This log-likelihood estimate was inspired by the error model employed previously (Gabbutt et al. 2022), in which the probability of observing a datapoint given some known “true” methylation fraction is beta distributed, and the likelihood is then the sum of these probabilities over all the simulation runs. Theinventors also added a further term to the log-likelihood to penalise cancer populationsizes that were non-physical. Surprisingly, Bayesian methods applied to noisy estimators of the log-likelihood can still sample from the correct posterior distribution associated with the intractable exact log-likelihood as long as the estimator is unbiased (Andrieu and Roberts 2009). Hence,the inventors employed standard nested sampling using the python dynesty package(Speagle 2019), which allows both the posterior to be inferred and the associated modelevidence, which is useful for model comparison. The inventors further included abootstrapping and importance sampling step, in which they re-estimated the log-likelihooda number of times for every point in the nested sampling run and updated these pointswith the mean of the re-estimated log-likelihoods with bootstrapped errors. This stepslightly reduced the uncertainty in the posterior estimates and shifted the estimated modellikelihood. The inventors applied their inference method to simulated data and were ableto accurately recover the ground truth parameters. The inventors applied their inference method to a total of 718 CLL patient samples(plus a further 55 from its precursor condition called monoclonal B cell lymphocytosis (MBL) and 6 its high-grade transformation called Richter transformation (RT)) using 978 fCpG loci. A small fraction of the CLL cases (16 / 718) had unusually high inferred noise and were considered outliers. In the remaining cases, the posterior was significantly tighter than the prior, suggesting the data has provided good evidence informing the model parameters. CLL patients are often stratified using the IGHV mutational status, which reflects the cell of origin of the cancer, with unmutated patients having a significantly worseprognosis (Hamblin et al. 1999). Consistent with this, the inventors found that the inferredgrowth rate of unmutated CLL patients was significantly higher than unmutated patients(^ = 1 × 10^^^, Mann-Whitney U test). More clinically relevant, they also found that withinboth the unmutated and mutated patient groups, the growth rate was highly correlatedwith the time to first treatment (TTFT, ^ = 2 × 10^^^) in a univariate analysis. In amultivariate Cox regression model including IGHV status, likely cell of origin and otherprognostic markers, the growth rate remained significantly correlated with TTFT (^ =2 × 10^^^). Similarly, the inferred effective cancer population is prognostic of the overallP603125PC0034survival in both univariate (^ = 5 × 10^^) and analogous multivariate analysis (^ = 0.025).Together, these survival analyses suggest that the evolutionary history of a cancer holds valuable prognostic information with regard to the progression of disease and eventual treatment failure. Hence, the inventors have developed a method to infer the evolutionary history ofcancers by combining their novel inference method with relatively cheap, standard DNAmethylation arrays. This inferred history may be used to predict individual patient outcomes, allowing for stratification beyond that of standard-of-care practice. Figure 1 is a schematic showing steps of the method according to an embodiment of the present invention. Figure 1 particularly relates to the steps required to implement the overarching risk prediction method / program 100, but also discusses the use of the stochastic forward model / program, the inference method / program, and the survival analysis method / program. In step 102 of the risk prediction method 100, DNA methylation data of fluctuating CpG loci of cancer cells is obtained. In step 104, a stochastic forward model is fitted to the methylation data using an inference method to generate a methylation distribution. In step 106, the methylation distribution is used to infer the historical growth rate of the cancer. Finally, in step 108 the risk of progression of the subject is predicted. The prediction is achieved using a survival analysis method based on the historical growth rate of the cancer. Figure 2 is a schematic showing steps of the method according to a furtherembodiment of the present invention. Again, Figure 2 particularly relates to the steps required to implement the overarching risk prediction method / program 200, but also discusses the use of the stochastic forward model / program, the inference method / program, and the survival analysis method / program. Steps 202, 204 and 206 ofFigure 2 are the same as steps 102, 104 and 106 of Figure 1. Accordingly, step 202 of therisk prediction method 200, DNA methylation data of fluctuating CpG loci of cancer cells is obtained. In step 204, a stochastic forward model is fitted to the methylation data using an inference method to generate a methylation distribution. In step 206, the methylation distribution is used to infer the historical growth rate of the cancer. Finally, step 208 of Figure 2 differs to step 108 of Figure 1. In step 208, the likelihood of overall survival or progression-free survival of the subject is predicted. The prediction is formed using a survival analysis method and is based on the historical growth rate of the cancer. Figure 3 is a schematic showing steps of the method according to a furtherembodiment of the present invention. Again, Figure 3 particularly relates to the steps required to implement the overarching risk prediction method / program 300, but also discusses the use of the stochastic forward model / program, the inference method / program, and the survival analysis method / program. In step 302 of the riskP603125PC0035 progression method 300, a cancer sample from the subject is obtained. Next, in step 304 a set of fluctuating CpG loci of the cancer cells are identified from the cancer sample obtained in step 302. DNA methylation data corresponding to the set of fluctuating CpGloci is obtained in step 306. Next, in step 308 a stochastic forward model is fitted to themethylation data using an inference method to generate a methylation distribution. In step 310, the historical growth rate of the cancer is inferred from the methylation distribution and optionally one or more other parameters. These parameters may include,but are not limited to, the age of the subject when the most recent common ancestor(MRCA) of the sampling population emerged, the four epigenetic switching rates,corresponding to the four possible transitions between homozygous unmethylated, heterozygous methylated and homozygous methylated; and the average methylation levelof contaminating normal cells. Finally, in step 312 the risk of progression of the subject ispredicted using a survival analysis method. The prediction is based on the historical growthrate of the cancer and optionally one or more parameters that may include: the age of thesubject when the most recent common ancestor (MRCA) of the sampling populationemerged, the four epigenetic switching rates, corresponding to the four possible transitions between homozygous unmethylated, heterozygous methylated and homozygous methylated; and the average methylation level of contaminating normal cells. Figure 4 is a schematic showing steps of an exemplary Bayesian inference method400. The Bayesian inference method could be used in conjunction with the methodschematics 100, 200, 300 discussed above, specifically in steps 104, 204 and 308. In step 402 of the exemplary Bayesian inference method 400, the log-likelihood of the methylation data is estimated. Next, in step 404 a penalisation function is applied to penalise sizes of the cancer that are non-physical. In step 406, a correction function is applied in order to account for finite sampling bias. In step 408, the method completes bootstrapping and importance sampling. Typically, this step involves re-estimating the log-likelihood a number of times for every point in the nested sampling run and taking a mean for eachpoint with bootstrapped errors. Finally, in step 410 the methylation distribution is definedby a set of parameters, wherein the set of parameters comprises the historical growth rate of the cancer. Figure 5 is a block diagram of a typical general purpose computer system 500 that can form the processing platform for the risk prediction of progression of a subject withcancer, as described, according to an embodiment of the present invention. The computersystem 500 comprises a central processing unit (CPU) 502, random access memory (RAM) 504, and input / output ports (I / O) 506 into which data can be received and output therefrom as is well known in the art. Additionally included is a visual display unit (VDU) 530 for displaying necessary user data and outputs and network connection 532 which canP603125PC0036 communicate with other computer systems via the network I / C 508. Although it should be appreciated that many further appliances could be connected to the computer system 500 such as a mouse 526 and a keyboard 528 if necessary. The computer system 500 also includes some non-volatile storage 512, such as a hard disk drive, solid-state drive, or NVMe drive. Stored on the non-volatile storage 512 is a number of executable computer programs together with data and data structures required for their operation or training. Overall control of the computer system 500 is undertaken through the execution of the control program 514 which executed programs relating to the general operation of the computer system 500 such as initialisation or shut- down etc. However, the non-volatile storage 512 also includes programs which when executed perform the specific requirements of the present invention as detailed above. These programs include but are not limited to a risk prediction program 516 for predictingthe risk of progression of a subject with cancer, a survival analysis program 518 whichworks in tandem with the risk prediction program 516, the inference method program 520and the stochastic forward program 522. These relate to the methods / programs as discussed in relation to Figures 1 to 4. The other data contained in the non-volatile storage512 is the training data 524 which would be required to train the associated programs.The computer-implemented system is applicable with a plethora of devices i.e.,devices that are capable of processing large data files. The devices should be capable ofat least receiving, processing, analysing, and storing of data that is required for risk prediction of progression of a subject with cancer. Therefore, devices that may beapplicable for use in this system, but are not limited to, include:• desktop computers• laptop computers• mobile telephones• tablet computersVarious modifications whether by way of addition, deletion, or substitution offeatures may be made to above-described embodiments to provide further embodiments,any and all of which are intended to be encompassed by the appended claims. All patent and literature references cited in the present specification are hereby incorporated by reference in their entirety.Example 1An unpublished paper written by the inventors is included below and forms part ofthe description of the embodiments, thus forming part of the patent application. This paperdescribes experiments which use and explain embodiments of the present invention indetail.P603125PC0037 Introduction The growth and dissemination of cancer is fundamentally an evolutionary process(Nowell 1976; Merlo et al. 2006); consequently, the past evolutionary trajectory of acancer should strongly influence its future trajectory and allow inference of the clinical path of a patient (Greaves and Maley 2012). However, testing this hypothesis directly is challenging as a precise characterization of cancer evolutionary dynamics would require the analysis of multiple sequential patient samples. Instead, evolutionary trajectories must be inferred by means of a proxy; for example, somatic (epi)mutations are patterned in distinctive ways by differing evolutionary dynamics (Turajlic et al. 2019). In the haematological system, genome sequencing of single cells or single cell colonies are employed to infer the phylogenetic relationships among cells (Gaiti et al. 2019; Lee-Six et al. 2018; N. Williams et al. 2022). The expense of this approach has restricted analyses to small numbers of cases, which limits its scaling for clinical translation. DNA methylation can also be employed as a heritable lineage marker, recording the clonal architecture of a population of cells (Yatabe, Tavaré, and Shibata 2001; Honget al. 2010; Brocks et al. 2014; Hao et al. 2016; Siegmund et al. 2009) or the total cellproliferative history (Duran-Ferrer et al. 2020; Endicott et al. 2022). We have recentlyidentified CpGs whose methylation state in each allele is heritable, but stochasticallyfluctuate over time. These fluctuating CpGs (fCpGs) function as a methylation barcode and so offer a low-cost strategy to provide high temporal resolution lineage tracing in patientsamples (Gabbutt et al. 2022).In a single diploid region, an fCpG can take one of three states: neither allele methylated (0% methylated), one allele methylated (50% methylated) or both alleles methylated (100% methylated). If multiple fCpGs are considered together, collectively they represent a vast number of possibilities (i.e. 3 to the power of the number of fCpGs) where distantly related cells carry distinct fCpG barcodes. These barcodes are not stable, but stochastically methylate and demethylate over time at a timescale measured in years (Gabbutt et al. 2022). Therefore, they not only provide a means to trace the ancestor of a clonal expansion, but also an “evolving barcode” that tracks clonal evolution: two somatic cells with close ancestry will share a near-identical pattern of fCpG methylation, whereas distantly related cells will have divergent fCpG methylation patterns. In bulk populations of somatic cells, the dominant fCpG pattern in the sample represents the fCpG state of the founder cell of the population and intermediate methylation values that deviate from 0, 0.5 and 1 are caused by subclonal expansions in the population (Fig 1A). Consequently, the full bulk distribution of fCpG methylation is determined by the evolutionP603125PC0038 of the population and so provides a precise proxy-measurement of the evolutionary dynamics. Here, we use fCpGs to infer the evolutionary dynamics of cancer cells from clinical specimens, at scale. We focus on lymphoid neoplasms, which cover a broad spectrum of diseases and subtypes with highly variable biological features and clinical manifestations (de Leval et al. 2022; Arber et al. 2022). These tumours have been extensively profiled by DNA methylation arrays, which have provided important insights into the cellular origin, pathogenesis, and clinical behaviour of these tumours (Oakes and Martin-Subero 2018). While their temporal clonal dynamics has been analysed in part (Gruber et al. 2019; Gutierrez et al. 2021), their precise evolutionary dynamics remains poorly characterised. We construct a quantitative modelling framework called “EVOlutionary inference using FLUctuating methylation” (EVOFLUx) that enables precise quantitative estimation of cancer cell dynamics in lymphoid tumours from individual patients. In 1,976 individual lymphoid malignancies, we precisely measure evolutionary history and show that the patient-specific evolutionary dynamics is strongly associated with disease outcomes. Results Bulk Methylation Array Samples from 2430 samples We assembled and processed with a harmonized pipeline 2,430 bulk sample Illumina methylation array data of normal and neoplastic lymphoid cells (Kulis et al.2015; Nordlund et al. 2013; Reinius et al. 2012; Lee et al. 2015; Queirós et al. 2016; Nadeu et al. 2020; Duran-Ferrer et al. 2020; Nadeu et al. 2022; Oakes et al. 2016; Dietrich et al. 2018; Agirre et al. 2015). After quality control, we retained 2,204 samples from 2,054 patients (including 22 technical replicates, 3 synchronic and 125 longitudinal samples from the same patients; fig.26) and 389,180 CpGs for downstream analyses. As healthy control samples, this dataset contained sorted CD19+B cells (n=40), CD3+T cells (n=35),peripheral blood mononuclear cells (PBMCs, n=6) and whole blood samples (n=6). Astumor samples, we included precursor 797 B- and 90 T-acute lymphoblastic leukemias (B-and T-ALL, respectively) at diagnosis, 28 B and 2 T-ALL at relapse as well as 74 B and 12T-ALL at complete remission (i.e. normal blood); 149 mantle cell lymphomas (MCL); 722chronic lymphocytic leukemias (CLL), 55 of its precursor condition monoclonal B-cell lymphocytosis (MBL), and 6 samples from CLL patients undergoing a diffuse large B-celllymphoma (DLBCL) transformation called Richter transformation (RT); 62 primary DLBCL,not otherwise specified (NOS); and 104 multiple myeloma (MM) and 16 of its precursor condition monoclonal gammopathy of undetermined significance (MGUS). In addition, clinical follow-up was available for 1,667 patients and matched DNA methylation with genetic (whole genome / exome sequencing (WGS / WES)) or with RNA-seq data wasP603125PC0039available for 505 and 293 CLL / RT samples, respectively (fig. 26). Therefore, this exquisitedataset was unique from multiple perspectives, including tumors derived from lymphocytes at various maturation stages, tumors with a highly proliferative acutepresentation and more indolent chronic leukemias, tumors from patients ranging frominfants to old adults, and tumor samples from different niches and disease stages with available genetic, transcriptomic, and clinical data. Identification of fCpG Loci in Blood Malignancies We have previously observed that fCpG loci are tissue-specific (Gabbutt et al.2022). To identify fCpGs in the lymphoid system we first partitioned our dataset intodiscovery (n=1471) and validation sets (n=733, tumor samples n=536; fig. 26). Usingthe discovery set, we selected for CpG loci that fulfilled the following criteria:(i) Heterogeneous across different subjects with the same disease (by accepting CpG loci with top 5% of standard deviation of methylation value within a cancer type) (ii) Equally likely to be methylated or unmethylated (by averaging across all cases in a cancer type and accepting only CpG sites with average methylation of approximately 0.5 in the cancer type) (iii) Unlikely to be associated with specific cell or cancer types. This was done using anunsupervised Laplacian Score (LS) feature selection metric (He, Cai, and Niyogi 2005) thatranked CpG loci by their tendency to preserve the nearest-neighbour graph, where we accepted the 5% least informative CpG loci. To visualise how the LS quantified the informativeness of individual CpG loci, we performed a principal component analysis (PCA) on methylation values (fig. 1B). The first 2 principal components revealed disease specific clustering in methylation. The CpG loci with the most informative features for clustering (lowest LS) showed disease-specific methylation. In contrast, the CpG loci that were least informative for clustering (highest LS) tended to have either low variability or high variability that was well-dispersed across the PCA. Combining these filters yielded 978 pan-lymphoid cancer fCpGs. These fCpGs did not cluster the samples by either disease (fig. 1C) or disease-subtype in either discoveryor validation samples (fig. 11A-D), nor with the version of methylation array employed.We then examined the distribution of these 978 fCpG methylation values in individual samples. As we previously predicted (Gabbutt et al. 2022), in each cancer sample the fCpG loci followed a characteristic “W-shaped” distribution that depicts the fCpG methylation patterns of the founder cell of the cancer sample (fig 1D). In contrast,the healthy B cell subpopulations, which were not included in the discovery set, hadunimodal distributions with intermediate methylation consistent with these being polyclonal populations (fig. 1D). The fCpG methylation distributions of samples from ALLP603125PC0040 patients in remission (i.e. contain mostly healthy cells) similarly had unimodal distributions and clustered with the normal B cell subpopulations (fig 1C-D). Employing the standard deviation of each sample’s fCpG methylation distribution as a proxy for its “W-shape”, we observed that lymphoid cancers had significantly higher standard deviation than both peripheral blood cells and sorted B and T lymphocytes, as expected (fig 1D). fCpG loci are enriched in silent regions of the genome fCpGs represented less than 1% of measured CpGs and were generally evenly dispersed across all chromosomes, with the exception of an enrichment on chromosome 19 (P=2.2e-5, ^^test, holm-sidak (hs) correction, fig.12A). Despite the similar distribution across chromosomes, fCpG loci were enriched on the shores of CpG islands, whilst being depleted within CpG islands themselves and the open sea (all significant p values < 10-7,^^ test, hs correction, fig. 12B). Remarkably, fCpG were poorly overlapped with CpGpreviously used in a variety of “epigenetic clocks”, including mitotic age, chronological age, trait predictors and biological and mortality clocks (fig 12C). Next, to extend our analyses beyond Illumina array data, we looked at the methylation patterns of sorted bulk B-cell and T-cell populations using publicly available whole genome bisulfite sequencing data (Loyfer et al. 2023). Methylation at fCpG loci in these polyclonal normal samples was largely intermediate (fig. 13A), consistent with ourprevious findings in array data (fig. 1C-D). Over a small window of 100bp, as the distancefrom the fCpG increased, an increasing fraction of the neighbouring CpG loci were eitherhyper- or hypo-methylated (fig. 13A-B), suggesting that a particular 3D DNA structuremay allow the fluctuation of methylation at the specific sites where fCpGs are located. The chromatin configuration of fCpG loci in normal and neoplastic B cells showed that fCpG loci were enriched in weak promoters and enhancers as well as H3K27me3- marked regions, whereas in active promoters and H3K36me3-marked regions they weresignificantly underrepresented (fig. 1E). In line with this, RNA-seq analysis of 294 CLLsamples showed that the expression of genes associated with fCpGs was highly skewed towards lower expression compared to genes exclusively associated with non-fCpG loci(fig. 1F). Using a subset of CLL samples (n=224) having matched RNA-seq and DNAmethylation data, we observed similarly low gene expression levels between fCpG- associated genes regardless of the methylation allele states (fig. 1F). These results were further supported by a variety of statistical analyses and database annotations, includinga significantly lower likelihood of fCpGs being annotated as associated with a gene (fig.12C,D, p<0.001) or annotated as promoters by the ENCODE Methylation Consortium(Dunham et al. 2012) (fig. 12E, p<10-10, ^^ test, hs correction). fCpG associated genesP603125PC0041were underrepresented in pathways ubiquitously expressed across multiple tissue types(e.g. lymph node, tonsil and testis) by gene-set enrichment analysis (fig. 14A).Conversely, fCpGs were enriched in CpGs unclassified by ENCODE (p < 10-10, ^^ test, hscorrection, fig. 12E) and enriched in developmental pathways (fig. 14B). Together, these results indicate that fCpG loci are located in particular regions of the genome where it is likely that the DNA structure enables fluctuation, and further that fCpGs do not regulate transcription and thus are unlikely to experience clonal selection. EVOFLUx measures clonal evolution We developed a stochastic modelling and inference framework called EVOFLUx tosimulate how clonal evolution quantitatively determines fluctuating methylation values and enables inference of evolutionary dynamics on a case-by-case basis from patient data. The model simulates the ongoing gain and loss of methylation at fCpGs within a lineage from a patient’s birth until the beginning of a cancer-associated clonal expansion at somespecified time, and then continues to simulate methylation fluctuations within the growingpopulation of cancer cells until the cancer sample is collected at time ^ (fig 2A). The keyparameters in the model are: ^Cancer growth rate per year (^), assuming an exponentially growing population.^ Patient age at cancer emergence (^), measured in years.^ fCpG switching rates per allele per year. Four parameters corresponding to the fourpossible transitions between homozygous unmethylated, heterozygous methylated and homozygous methylated (^, ^, ^ and ^).Importantly, the time of cancer emergence should be thought of as the time of occurrence of the most recent common ancestor (MRCA) of the sampled extant cancer population, rather than the cancer age per se, as there may have been previous clonal expansions prior to the “final” clonal sweep that produced a clinically-detectable cancer, and the MRCA of the sample is inevitably the same or more recent than the MRCA of the whole cancer population. By combining the cancer growth rate and the time that has passed since the MRCA, the cancer effective population size (Ne) present within the samplecan be calculated as ^^ = ^^(^^^). Note that this is the ^^of the sample, rather than the cancer as a whole, and under the assumption the cancer is growing exponentially. EVOFLUx accounts for sample impurity (contaminating non-cancer cells in the sample) by modifying the methylation fraction of the cancer cells by the average methylation of the contaminating cells (assumed to be random), weighted by samplepurity (Methods). The framework also mimics the technical noise introduced by themethylation array, by adding beta-distributed noise to the methylation values centred upon the “true” methylation value in the population.P603125PC0042 Simulations showed that the model qualitatively recapitulates the patterns observed in patient data, including the characteristic W-shape (fig.2B). The specific shape of the resulting W-distribution is determined by the model parameters. For example, asthe cancer age is increased (i.e. the time ^ occurs earlier in a patient’s life, all otherparameters held constant), there is more time for fluctuations to “desynchronise” the methylation pattern of the founder cell, causing the flanking peaks of the W-shaped distribution to move towards the central peak (fig. 2C). Similarly, if we keep the time since the MRCA and the epimutation rate (per year) fixed but decrease the growth rate of the tumour population, each subsequent division has more time to accumulate epimutations, magnifying the stochasticity of the system and broadening the width of the peaks in the W-distribution (fig. 2C). EVOFLUx utilises an extensive Bayesian inference method to learn model parameters from input fCpG methylation distribution data (Methods). To evaluate the accuracy of the inference, we simulated fCpG data and tested the ability of EVOFLUx to recover the known input parameters (fig. 2C). These values were inferred with good confidence, with the 95% credible intervals encompassing the ground truth values (fig. 2D), significant narrowing of the posterior compared to the prior and with the posterior predictive distribution well-matching that of the original simulated dataset (fig.2E). Hence, these simulation experiments demonstrate that the parameters defining the evolutionary dynamics of growing cancers can be measured using bulk methylation arrays. The evolutionary dynamics of lymphoid malignancies We applied EVOFLUx to 1,976 samples from lymphoid cancer (including T-ALL, B- ALL, CLL, MCL, DLBCL and MM) and premalignant conditions (i.e. MBL and MGUS) where we had age and tumour cell purity information, and inferred the cancer growth rate, time to the MRCA and epigenetic switching rates of each cancer independently. Posteriordistributions were well-formed (fig. 15A), and posterior predictive distributionsrecapitulated the input data well (fig. 15B), emphasising the excellent fit of the model todata. Acute paediatric leukemias derived from precursor B / T cells and adult lymphoidneoplasms derived from mature B cells were observed to have markedly differentevolutionary histories (taking the median as a point estimate; fig. 3A). In particular, EVOFLUx inferred ALLs growing at a significantly higher rate (fig. 16A, P= 9.3e-306, MW- U, hs correction), having a smaller cancer effective population size (^^, fig. 16B, P= 8.1e- 25) and shorter time since the MRCA (fig. 16C, P= 6.0e-306) than all other lymphoidmalignancies. T-ALL were found to grow faster than B-ALL (P=0.0017, hs correction),although the variance on the B-ALL growth rates was larger than for T-ALL (P=0.00044,P603125PC0043 Levene test). In adult cancers, MBL, an asymptomatic precursor condition to CLL, was found to have lower growth rates than CLL (fig. 16A, P= 9.7e-10) and longer time since the MRCA (P=9.9e-13), although there was no statistically significant difference in the inferred ^^(P= 0.53). DLBCL was found to have the largest ^^, despite not having a particularly elevated growth rate compared to the other adult cancers (potentially due tothe lower purity of DLBCL samples).There was also significant variation in the epigenetic switching rates (epimutationrates) between diseases, with acute paediatric leukaemias having an order of magnitudefaster switching than adult mature B cell neoplasms (fig. 3B; fig. 16D, P=5.6e-301).Further, we observed a positive correlation between growth rate and methylation switching rate in both B-ALL and T-ALL (fig. 16E, P= 2.4e-98, R2=0.44 & P=5.9e-06, R2=0.22 respectively), a weak negative correlation in MM (P=1.6e-05, R2=0.18) and no correlation in the other adult malignancies, suggesting that fCpG switching rates are elevated during childhood, but fall to a steady rate during adulthood. Notably, when comparing the fourepigenetic switching parameters (^, ^, ^ and ^), the different cancer types occupieddifferent regions of the epimutation parameter space (fig. 16F). The Majority of CLLs Grow Neutrally Previous work has proposed that many cancers grow effectively neutrally with no subclonal selection detectable (Sottoriva et al. 2015; M. J. Williams et al. 2016; 2018; Ling et al. 2015). In CLL, two major molecular subtypes are defined based on the extent of somatic hypermutation in the heavy-chain variable region of the IG gene, namelyunmutated (U-CLL), and mutated (M-CLL). Nonetheless, a small fraction of CLL patientspresent with multiple IGHV rearrangements (Plevova et al. 2014), and likely consistent of two cancers with independent clonal origins. Subclonal, independent and neutral growth dynamics are distinguishable by characteristic patterns of mutations across the cancer genome (Caravagna et al. 2020; M. J. Williams et al. 2018). Simulations showed that in fCpG data the signature of subclonal selection is the presence of intermediate peaks between the peaks corresponding to fCpG loci that were in the homozygous unmethylated, heterozygous, and homozygous methylated states in the MRCA cell (fig. 17A). These additional intermediate peaks are formed by fCpG loci that have switched methylation state in the lineage that formed the subclone within the intervening time between the MRCA of the sample population and the emergence of the outgrowing subclone. However, direct detection of these fCpG methylation patterns in the data (e.g. via a peak detection method) was generally impractical due to the relatively small number of fCpG loci that are expected to change methylation status in the timebetween the MRCA and the emergence of the subclone (e.g. assuming ^ ≈ 0.01 per year,P603125PC0044as in CLL, 1 − ^^^^(^^^^^) ≈ 2% of fCpGs in the homozygous unmethylated state will haveundergone 1 fluctuation in a single year) and technical noise of fCpG methylation measured by the methylation array (fig. 17B). Therefore, we re-fitted each of our patient’s fCpG methylation distributions with modified EVOFLUx models that included subclonal selection and independent clonal origins, and used Bayesian model selection and a leave-one-out approach to determine which model best represented the data (fig. 3C-E, 17C). We found that in the majority of samples (1610 / 1976), there was no evidence of either subclonal or two independent cancers at 95% confidence. There was also a high variance in the fraction of cancers growing under subclonal selection between different cancer types (fig.18A, from over 30% (232 / 718) in CLL to less than 5% in DLBCL (1 / 57). However, these results need to be interpreted cautiously given the heterogeneity in purity across the sample groups, which likely limits the power to detect subclonal selection in lower purity samples. We note that only strong selection occurring neither “too early” or “too late” is detectable in the data (fig. 17D), analogously to detecting selection using point mutations in WGS (Heide et al. 2018). This is because late arising or weakly selected subclones will only be present at very low frequencies, whereas very strongly selected early subclones will tend to have swept prior to sampling, and very early arising subclones have not had sufficient time to evolve fCpG patterns dissimilar enough to be distinguishable from the ancestor clone. Furthermore, whilst our simulations demonstrated that EVOFLUx can correctly identify subclonal section, the absolute value of the inferred selection coefficients are not well informed due to the noisiness of the methylation array data (fig. 17E). To verify EVOFLUx based inferences of subclonal selection we used matched WESdata available for 425 CLL cases. Using MOBSTER, our subclonal deconvolution tool(Caravagna et al. 2020), to detect emerging subclone(s) from the variant allele frequency(VAF) distributions of somatic mutations (fig. 3F) we found 78 / 425 samples containedevidence of subclonality by WES. Samples which were classified as subclonal by MOBSTER had significantly higher subclonality weights inferred from our EVOFLUx inference (P=2.0e-4, MW-U test, fig. 3G). The paucity of point mutations in WES data limits the power to detect subclonality. MOBSTER was more likely to assign a sample as containing a subclone when the samplecontained more mutations (p=0.0023, logistic regression against log number of mutations,fig. 18B). Hence, we suggest that those samples with high subclonality weights from our EVOFLUx inference, despite being labelled as “neutral” by MOBSTER, may truly contain subclones. Conversely, in cases assigned as subclonal by MOBSTER, there was no association between the mutation burden and EVOFLUx subclonality weighting, suggestingP603125PC0045 these may be genuinely subclonal samples that EVOFLUx has missed (P=0.87, logistic regression, fig. 18D). Cancers with two independent clonal origins were detected in 22 / 718 CLL cases.We attempted to validate these inferences by comparing them to orthogonal measurementof the number of IG gene rearrangements based on WES / WGS and RNA-Seq (Knisbacher2022). Those patients with multiple IG gene rearrangements had elevated independentorigins model weightings compared to those without (fig. 18D, P= 0.028), although wenote that EVOFLUx was not as accurate at detecting independent cancers as detectingsubclonality.EVOFLUx detects early subclonal origins of Richter transformed cells in CLLA small fraction of CLL patients will undergo RT, the emergence of an aggressive phenotype with dismal prognosis. Recently, using single cell genomics and transcriptomics, we found that RT cells are present at low frequencies within the diagnostic CLL sample, preceding the clinical emergence of RT in some cases by two decades (Nadeuet al. 2022). For 2 same patients analysed previously by WGS and single cell approaches(Nadeu et al. 2022), we also had DNA methylation arrays of longitudinal samples fromdiagnosis to RT (fig. 4A). We used EVOFLUX to construct phylogenetic relationshipsbetween samples from differences in fCpG methylation, extracting valuable temporalinformation regarding the divergence of RT cells from the ancestral clone.We found the inferred phylogeny well matched the known clinical relationshipsbetween the different samples (fig. 4B). Notably, the evolutionary distance between twotechnical replicates of untreated samples (450k and EPIC data) was small in both cases (fig.4B). In both phylogenies, the RT clone diverged exceptionally early, roughly a decadebefore the MRCA of the non-transformed samples (9 and 12 years for case 12 and 19,respectively). This is consistent with our previous work (Nadeu et al. 2022) in which wedetected RT cells within the diagnostic sample, but suggests that the initial RT divergence occurred well before diagnosis, over 30 years before the clinical presentation of RT. To validate the fCpG phylogenetic inference, we built phylogenetic trees on SNVdata from matched WGS (Nadeu et al. 2022) (figs. 4A and 19). We restricted our analysisto clock-like SBS1 mutations (Nik-Zainal et al. 2012) which were clonal within the sample(Caravagna et al. 2020). The topology of the SNV phylogenies matched that of the fCpG phylogenies and the 95% credible intervals of the estimated MRCA overlapped. Hence, fCpGs have the potential to serve as a high temporal resolution phylogenetic character,enabling multi-sample phylogenetic reconstruction at a fraction of the cost of WGS.In addition, we applied our phylogenetic analysis to 3 longitudinal B-ALL samples, tracking the disease course from initial diagnosis to relapse using the same methodologyP603125PC0046 (fig.20). In each of the patients, the subsequent relapse samples formed a separate clade from the initial diagnostic samples, suggesting the cancer population has undergone a significant bottleneck through treatment. The branch lengths of the diagnostic samples were all very short, suggesting there was initially little evolutionary distance between the surviving resistant population and the sensitive cells. Evolutionary growth dynamics of CLL predicts clinical outcomes Since cancer development and progression is fundamentally an evolutionary process, the evolutionary trajectory of the disease should predict the future course of the disease and clinical outcome of the patients. To evaluate this hypothesis, we first compared the inferred growth rate and ^^between different molecular subtypes of the same cancer type. In B-ALL, cases having the translocation 11q23 / MLL had a significantly highergrowth rate (fig. 5A, P=1.3e-13, MW-U, 44.3±6.1 vs 11.7±0.2 per year [mean ± standarderror]), but lower ^^(P=3.7e-07, 1.8e5±0.2e5 vs 3.0e5±0.06e5 cells) than the other subtypes, consistent with its distinct clinical behaviour (Rice and Roy 2020). In MCL, thegenerally more indolent nnMCL (Nadeu et al. 2020) had a lower growth rate (fig. 5B,P=1.1e-03, 1.7±0.1 vs 2.1±0.1 per year) and Ne(P=7.4e-05, 4.7e5±1.4e5 vs 1.5e6±0.2e6 cells) than the more aggressive conventional MCL (cMCL). In DLBCLtranscriptomic subtypes (Alizadeh et al. 2000) there was no significant differences, likelydue to the smaller number of cases and the lower sample purity (fig. S11A&B). In CLL, the more aggressive U-CLL subtype (Hamblin et al. 1999; Damle et al.1999) showed a significantly higher growth rates (fig. 5C, P=1.3e-32, 2.3±0.04 vs1.8±0.02 per year) and ^^(P=2.1e-22, 7.2e5±0.3e5 vs 4.1e5±0.3e5 cells) compared to M-CLL. In its precursor condition MBL, although we had little power to detect a difference between U-CLL and M-CLL subtypes, we could observe a non-significant increase in growth rate in U-CLL cases mirroring the findings in CLL (fig. S11C&D, P=0.093, 1.7±0.1 vs 1.5±0.06 per year). This result is consistent with the known enrichment of MBLs in IGHV mutated cases, and suggestive that MBL patients with an unmutated IGHV gene generally progress more rapidly towards CLL and therefore the MBL phase is less frequently observed in the clinical setting. To further measure the influence the past evolutionary history of a cancer imparts upon its future clinical course, we exploited our access to a large clinically well-annotated series of CLL cases to evaluate the prognostic value of the inferred evolutionaryparameters. We used univariate Cox models to assess the impact of each parameter ontime to first treatment (TTFT), a commonly used variable that describes the natural history(biological behaviour) of the CLL cells, as well as overall survival (OS) which is a moreP603125PC0047 complex endpoint as it convolves disease biology with heterogeneous treatments administered to the patients. The univariate prognostic impact of the inferred growth rate of the cancer on both TTFT (fig. 5D, P=1.4e-30, hazard ratio (HR) HR=3.95) and OS (P=0.0053, HR=1.51) was stark. The ^^of a cancer did not have a strong impact on TTFT (P=0.058, HR=1.17), butits effect on OS was clear (P=1.3e-4, HR=1.41). The patient age at the time of the MRCAof the CLL population is highly corelated with the age of the patient (fig S12A), sounsurprisingly older patients had worse OS (P=2.3e-11, HR=1.79). The decrease in riskof progression with cancer age (P=3.8e-17, HR=0.65), measured by the time since theMRCA, was likely due to confounding with the growth rate, as these parameters werenegatively correlated (fig. S12A). The epigenetic switching rate parameters were largelyuninformative of prognosis.As the growth rate was different between U-CLL and M-CLL (fig. 5C), we nextanalysed its prognostic impact within each of them. Indeed, we observed that theestimated growth rate has a profound impact on TTFT in both CLL subtypes separately(fig. 5E, P=1.4e-5 for M-CLL, P=1.56e-7 for U-CLL, overall P=2.1e-53). Cases with higher growth rate consistently showed a shorter TTFT. Considering the growth rate as continuous variable in a multivariate Cox regression model, it maintained a strong independent prognostic impact (P=2.2e-10, HR=2.28) in the context of other well-established clinicalvariables, including the IGHV mutational status and TP53 aberrations (including bothmutations and deletions) and age at sampling (fig. 5F). The cancer ^^ was moresignificantly correlated with OS than the growth rate, and this effect was preserved in theU-CLL subtype (fig S12B, P=0.55 for M-CLL, P=1.0e-6 for U-CLL, overall P=1.7e-9), andin multivariate setting with IGHV status, TP53 aberrations and age at sampling (fig S12C,P=0.025, HR=1.33). The inference of the evolutionary parameters on our initial cohort was wholly blinded to the clinical outcomes. Nevertheless, we validated the prognostic value of the inferred evolutionary parameters using a second independent cohort of 210 CLL patients(136 untreated at sampling) (Oakes et al. 2016; Dietrich et al. 2018). We found that thegrowth rate was prognostically relevant as a continuous (fig. S13A, P=7.4e-4, HR=2.39)and dichotomic variable (fig. S13B, P=5.1e-5). In a multivariate Cox model with IGHV,TP53 status and age the growth rate approached significance (fig. S13C, P=0.11,HR=1.62). We note that the power of the study in this cohort was limited, as even thewell-validated IGHV status had borderline significance (P=0.053). In this cohort, thepredictive impact of ^^ on OS was not found to be significant (fig. S13A, P=0.25,HR=1.39), but we stress that TTFT is a better reflection of the natural history of disease, not confounded by heterogenous treatment decisions.P603125PC0048 These results demonstrate that the evolutionary parameters inferred from a cost- effective methylation array could have a direct clinical application by contributing to predict the clinical behaviour of CLL patients independently from well-established prognostic variables. Discussion Our study establishes a framework, called EVOFLUx, that enables quantitative measurement of the evolutionary dynamics of human malignancies. These dynamics represent the long-term evolutionary history of the neoplasm and are fundamentally distinct from characterisations of the contemporary phenotype of cancer cells, such as the fraction of proliferating cells. The framework only requires data obtained from single- timepoint bulk tumour samples that have been processed with widely-available and low- cost DNA methylation arrays as input. The methodology should also work identically for sequencing-based measurement of DNA methylation, should generalise across cancer types, and be applicable to proxy tumour measures such as tumour-derived cell free DNA (cfDNA) extracted from blood. Consequently, this strategy enables measurement of evolutionary dynamics at massive scale, which paves the way towards clinical translation. EVOFLUx utilises the clonal lineage identity signal that derives from stochastic fluctuations in methylated and unmethylated states at a subset of CpG sites across the genome, and outputs the evolutionary parameters defining the longitudinal dynamics of cancer growth, alongside precise measurement of DNA methylation and demethylation rates at these sites. Epimutation rates were found to differ by more than an order of magnitude between paediatric and adult malignancies. Thus, we suggest that future analyses of these quantitative data are likely to reveal new mechanistic insight into the somatic maintenance of DNA methylation. Applying our framework in lymphoid malignancies reveals that cancer evolutionary dynamics are strongly associated with patient clinical outcomes across disease types, and further that evolutionary dynamics add significant new prognostic value to state-of-the- art approaches currently applied to the management of CLL in the clinical setting. We consider this strong evidence that clonal evolution, the fundamental cellular process of disease development, underlies the clinical course of the disease. Further, as cancer development and progression are in essence evolutionary processes, we expect these results to generalise across all cancer types. We note that genome-wide DNA methylationanalyses also measure other important biological features of a cancer (for example in CLL:(Duran-Ferrer et al. 2020)), that could be combined with EVOFLUx-based inference of evolutionary dynamics to further improve the prognostic value of DNA methylation data.P603125PC0049 In summary, we present a cost-effective high-throughput platform for measuring cancer evolutionary dynamics in patient samples. These fundamental measurements of the disease biology hold significant prognostic value and represent an innovative asset inthe field of precision oncology.Methods Assembly and quality control of DNA methylation data We assembled and processed with a harmonized pipeline (Duran-Ferrer et al.2020) (version 4.1) 2,430 bulk sample Illumina methylation array data of normal and neoplastic lymphoid cells from previous publications (Kulis et al. 2015; Nordlund et al. 2013; Reinius et al. 2012; Lee et al. 2015; Queirós et al. 2016; Nadeu et al. 2020; Duran-Ferrer et al. 2020; Nadeu et al. 2022; Oakes et al. 2016; Dietrich et al. 2018; Agirre et al. 2015).Briefly, raw idat files were loaded and processed with R (version 4.3.1) using minfi package(Aryee et al. 2014; Fortin, Triche Jr, and Hansen 2017) (version 1.46.0) in batches. Briefly,the data was processed for each batch as follows. First, idats files were loaded into aRGChannelSet object, and minfi quality metrics using the qcReport function wereperformed, controlling for biased distributions of methylation values and low signal intensities of internal control probes for each sample. Next, further quality metrics werederived using the function minfiQC on unnormalized RGChannelSet obejct. Those sampleswith median signal intensities of unmethylated and methylated channels of at least 10.5 in log2 scale were considered as having good signal intensities (default value in minfi). Subsequently, detection P-values were calculated across all CpGs and samples using thedetectionP function for the unnormalized RGChannelSet object. Samples were consideredas good if having a mean detection P-value across all CpGs pf P≤0.01. On a CpG level, we retained CpGs with a detection P-value ≤1e-16in ≥90% of the samples. The RGChannelSetobject was normalized with the preprocessNoob function. We next retained only CpGs(excluding CH probes) that did not contain any single nucleotide polymorphism neither in the interrogated CpGs nor in the probe using the dropMethylationLoci anddropLociWithSnps functions. Furthermore, CpGs with any previous evidence of potentialcross-hybridization were excluded (Chen et al. 2013) and only CpGs mapping to autosomalchromosomes were subsequently retained for downstream analyses. Finally, we checked the distribution of normalized methylation values and performed principal component analyses separately for samples passing all quality checks as well as those considered to as bad samples. The final DNA methylation matrix contained 2,204 samples and 389,180 CpGs passing all the aforementioned quality controls, and included 2,054 patients (22 technical replicates, 3 synchronic and 125 longitudinal samples from the same patients).P603125PC0050 To determine to purity of samples, we used our previously deconvolution strategy to infer tumor cell content by DNA methylation (Duran-Ferrer et al.2020), which was used as a consensus purity in all the tumor samples except for DLBCL and MM. In these two tumor entities, we have previously identified a DNA methylation signature loss causing inaccurate tumor purity predictions using DNA methylation data, and therefore we used available genetic or flow cytometry data for DLBCL and MM, respectively. Statistical analysis Statistical tests performed throughout the study were performed as two-sided. Appropriate multiple test correction, such as the holm-sidak (hs) correction, is noted when applied. Characterisation of fCpGs To characterise the genomic and regulatory context of fCpGs we employed a series of statistical analyses and database annotations. We annotated fCpGs using Illumina manifest and other genomic annotationa available at Bioconductor packagesIlluminaHumanMethylation450kanno.ilmn12.hg19 (version 0.6.1) andIlluminaHumanMethylationEPICanno.ilm10b2.hg19 (version 0.6.0). We used Chi-squaredtests ^^to assess the enrichment of fCpGs in distinct genomic regions or elements. Wealso used CpG annotations of the regulatory features provided by the ENCODE MethylationConsortium (Dunham et al. 2012). We performed gene-set enrichment analysis upon thefCpG-associated genes using gProfiler (Raudvere et al. 2019), specifically focusing uponthe Gene Ontology biological processes (Ashburner et al. 2000) and the Human ProteinAtlas (Uhlén et al. 2015). The statistical domain space was limited to genes targeted by at least one CpG in the 389,180 candidate CpG set and significance was determined using the g:SCS algorithm (Reimand et al. 2007). Previous chromatin segmentation of normal and neoplastic B cells was used to assess the chromatin state enrichment of fCpG (Duran- Ferrer et al. 2020; Beekman et al. 2018). fCpG were checked for their overlap with previous “epigenetic clocks”, includingmitotic (Duran-Ferrer et al. 2020; Yang et al. 2016; Teschendorff 2020; Youn and Wang2018; Zhou et al. 2018), chronological age (Bocklandt et al. 2011; Garagnani et al. 2012; Hannum et al. 2013; Horvath 2013; Lin et al. 2016; Vidal-Bralo, Lopez-Golan, and Gonzalez 2016; Weidner et al. 2014; Zhang et al. 2019; Horvath et al. 2018; Shireby et al. 2020), gestational age (Bohlin et al. 2016; Knight et al. 2016; Y. Lee et al. 2019; Mayne et al. 2017; McEwen et al. 2020), biological age and mortality (Belsky et al. 2022; Levine et al. 2018; Lu et al. 2019), and trait predictors (McCartney et al. 2018; Liang et al. 2020). The package methylCIPHERP603125PC0051 (https: / / github.com / MorganLevineLab / methylCIPHER) was used to integrate the majority of the epigenetic clocks. CLL RNA-seq data Previously available RNA-seq data for 294 CLL patients was obtained (Puente et al.2015) and processed as previously described (Nadeu et al. 2020). Matched RNA-seq dataand DNA methylation data for the same patients at the same timepoint was available for 224 CLL patients. Transcript per million counts (tpms) were used to represent differential gene expression values across genes and samples. We used the gene annotation providedin the R Bioconductor package IlluminaHumanMethylationEPICanno.ilm10b2.hg19 toclassify genes associated with fCpGs. Genes targeted by any fCpG were considered as “fCpG genes”. In each methylation sample, the 978 fCpG were discretised as homozygous demethylated, heterozygous methylated or homozygous methylated (coded as [0,1,2] respectively). This was done by separately fitting a beta mixture model with 3 componentsto each sample using Stan (Carpenter et al. 2017) and extracting the component mixtureprobability. The gene expression value for genes classified as having and fCpG with 0, 1 or 2 alleles methylated were plotted as previously described. A Stochastic Model of how fCpGs vary in large, neutrally growing populations of cells We built a generative computational model of how the patterns of fCpGs vary over time (^) according to the evolutionary history of a cancer. Initially, our model focused on neutral evolution, before expanding to non-neutral modes of tumour evolution below. For the full explanation of the model, please see the supplementary information. Our model was parameterised in terms of the age of the patient at which the most recent common ancestor (MRCA) emerged (^), the exponential growth rate of the cancer (^), and the epigenetic switching rates of the fCpGs (^, ^, ^, ^). The model was partitionedinto 2 phases, prior and post the emergence of the MRCA. At time ^ = 0, the fCpGs wereassumed to be equally likely to be homozygously methylated or demethylated. The fCpGstatus of the MRCA at time ^ = ^ was calculated by applying matrix exponentiation.The second phase of the model consisted of a discrete time Markov process. The effective population size of the growing cancer was modelled as growing according to adeterministic exponential growth equation, ^^ = ^^(^^^). Each fCpG was consideredindependently – at each time step, ^ → ^ + ^^, the number of homozygous methylated (^),heterozygous methylated (^) and homozygous demethylated cells (^) at a specific fCpGwas updated according to the epigenetic switching rates.P603125PC0052 At the time of sample, ^, the fraction methylation of each simulated fCpG was calculated by summing the number of methylated alleles and normalising by the total number of alleles in the population: We further accounted for contaminating normal cells and the technical noise introduced by the methylation bead array. The methylation of the contaminated samples was assumed to be an average of the cancer methylation, weighted by the tumourpurity ^, and the average of the normal population, ^^, weighted by 1 − ^. Following ourprevious work, the bead array was assumed to saturate at extreme methylation values,shifting the minimum and maximum methylation by ^ and ^ respectively. The noise of thebead array was assumed to be beta distributed, with precision parameter ^. Non-neutral models of tumour evolution Alongside our model of neutral exponentially growing cancer populations, we devised two alternative models of cancer growth: 1. A subclonal selection model in which a single cell within the cancer develops aselective advantage and begins to grow at an increased growth rate. 2. An independent clonal origins model, whereby a patient has developed two distinctcancers concurrently. For the subclonal selection model, we replaced the growth rate (^) and the time of the MRCA (^) with the growth rates and time of the MRCA of the initial, slower-growing population (^^and ^^respectively), and that of the more recently emerging, faster-growingpopulation (^^ and ^^), constraining < ^^ and ^^ < ^^ (fig. 17C). We assumed that theinitial cancer population begins exponentially growing at as above, but at time ^ = ^^ weselect a single cell with a set of fCpG states drawn according to the cancer population and allow this second population to grow concurrently with a growth rate ^^. The independent-cancer model followed the same scheme as the nested subclonal selection model, except the methylation status of the emerging cancer was that of anindependent cell which experienced random fluctuations between ^ = 0 and ^ = ^^.If we let the number of cells in the less fit subclone in each methylation state be and in the fitter subclone be {^^, ^^, ^^}, following the convention above, then inboth cases the measured methylation patterns at the time of sample are: P603125PC0053 Prior functions For each methylation array blood sample, we had matched age (^) and purity (^) information. Hence, the parameters to be inferred are the growth rate (^), the age of the patient when the MRCA emerged (^), the epigenetic switching rates (^, ^, ^, ^; defined in fig 2A), the average fraction methylated of contaminating normal cells (^^), the betaoffsets from 0 and 1 due to the background noise on the methylation array (Δ and ^,respectively), and the precision of the beta distributed noise (^). These parameters are constrained either to be positive (^, ^, ^, ^, ^, ^ > 0) or to liewithin a specified range (0 <^ ^, ^^ , ^, ^ < 1), which we achieved using appropriate priordistributions. To better allow for priors to be set on a biologically meaningful scale, the priors for the lognormal distribution were set in terms of the real scale mean and standard deviation, rather than the standard log-scale. To reduce correlations in the posterior andmake sampling more efficient, the variables ^ and ^ were normalised by ^ and ^respectively. The priors are as follows: ^ ~lognormal(1 )^, 0.7^ ~lognormal(1, 0. )^7^^~beta(2, 2)^~beta(5, 95)^~beta(95,5) ^~halfnormal(100,30) When fitting non-neutral models of tumour growth, the inference is parameterised in terms of the relative growth of the fitter subclone, ^^^=^^ and the fraction of thepopulation consisting of the fitter subclone, ^ = ^^^(^^^^)^^^^(^^^^). The age at which the second clone emerges is then: ^− ^^ This parameterisation induces less correlations in the resulting posterior, which significantly improves the sampling efficiency. The priors on these additional parameters are:P603125PC0054 ^^ ~bet ( )^a 2, 2^^^~lognormal(1, 0.7)^~beta(2, 2)All the other priors were the same as in the neutral case. Bayesian Inference We developed a stochastic estimator of the loglikelihood function at a given set of parameters by simulating the fCpG methylation distribution a large number of times, correcting for the bias inherent with using a finite number of simulations and penalising the loglikelihood for extreme values of the Ne. The standard Bayesian algorithms developed to infer the posterior for a given set of data (e.g. MCMC, nested sampling) are typically used when the log-likelihood is analytically tractable and can be calculated exactly. Remarkably, it has been shown that, as long as the stochastic approximation of the log-likelihood is unbiased, MCMC methods can obtain an exact Bayesian inference of the true posterior, as in pseudo-marginal Metropolis–Hastings (Andrieu and Roberts 2009). Here, we employed a nested sampling approach using the dynesty package (Skilling 2004; 2006; Speagle 2019). Unlike pseudo-marginal Metropolis–Hastings, nested sampling is able to efficiently explore multi-modal posterior landscapes (which can occur under the subclonal and independent cancer models). Model comparison between non-neutral models of tumour evolution We employed an expected log pointwise predictive density (ELPD) (Vehtari,Gelman, and Gabry 2015) approach to compare our competing models of evolution foreach sample using the arviz Python package (Kumar et al. 2019), which uses PSIS-LOO- CV to compare the out-of-sample prediction accuracy between models whilst naturally penalising more complex models. This required the loglikelihood per datapoint and the posterior predictive for every point in the posterior. The weights of the respective models were calculated using pseudo-Bayesian Model averaging using Akaike-type weighting, stabilized using the Bayesian bootstrap (Yao et al. 2017). CLL and RT genomic analyses Previous mutated annotation files (MAF) from WES data (Knisbacher et al. 2022)and WGS (Nadeu et al. 2022) were used to further validate our distinct EVOFLUxevolutionary modes (i.e., neutral, subclonal and independent) and RT phylogenies.P603125PC0055 Subclonal deconvolution of WES and WGS data To detect subclones in bulk WES and WGS data, we employed MOBSTER (Caravagna et al. 2020), which fits the variant allele frequency (VAF) spectrum with a mixture model containing a Pareto distribution to account for the neutral tail (M. J. Williamset al. 2016) and a variable number of beta distributions to account for the clonal andsubclonal peaks. We ran MOBSTER using the default parameters, except employing a minimum 5% VAF threshold and lowering the minimum number of mutations to compose a cluster to 5 in WES samples due to the low number of mutations. We then manually QC’d all 377 WESsamples and 10 WGS, tuning the fitting parameters to better represent the data (forinstance, when the clonal peak had been called at a low frequency despite the median tumour purity being 95%). Phylogenetic inference of longitudinal methylation data A novel Bayesian phylogenetic method was used to reconstruct the evolutionary relationships and the time to most recent common ancestor (tMRCA) of longitudinal samples from the same patients. This was carried out in the BEAST v.1.8.4 framework(Drummond and Rambaut 2007; Drummond et al. 2012) using custom modelsimplemented in PISCA v1.1 (available from https: / / github.com / adamallo / PISCA) (Martinez et al. 2018). EVOFLUx provided an estimate of the age of the patient when the MRCA of each bulk sample emerged. To estimate the methylation status of each fCpG at the sample’s MRCA in each of our longitudinal samples, we discretised the fCpGs as described above (RNAseq Methods). We implemented a 4-parameter biallelic binary substitution model analogous to the pre-growth EVOFLUx model in PISCA. This plugin contains all the required statistical machinery to use this model for somatic phylogenetic estimation. The biallelic binary substitution model has three relative rate parameters 1) heterozygous methylation ^, 2) homozygous demethylation ^, and 3) heterozygous demethylation ^^, where homozygousmethylation ^^ was normalised to 1. For all relative transition rate parameters, a lognormalprior with mean 1 and standard deviation of 0.6 was used, with a half normal prior with mean 0 and standard deviation 0.13 for the molecular clock rate, using a strict clock model for the rate of evolution across the tree. Two demographic tree models, constantpopulation size (Kingman 1982) and exponential growth (Griffiths and Tavaré 1994), werecompared by marginal likelihood estimation using path-sampling (Baele et al. 2012) anda constant population model was deemed more appropriate.P603125PC0056 MCMC chains were run for 100 million generations sampled every 100,000 generations and convergence was assessed using Tracer v.1.7 (Rambaut et al. 2018), ensuring effective sample sizes greater than 500 for all parameters. Maximum clade credibility trees were then made using 10% burn-in and medium node heights. The resulting trees were plotted using ggtree (Yu et al. 2017). Phylogenetic inference of SNVs from WGS data Each bulk sample is represented by a set of clonal mutations found during the deconvolution of WGS data (see above). Where a mutation was deemed absent in the clonal peak, the reference nucleotide was used. Mutational signature assignment (Díaz-Gay et al. 2023) was used to select mutations in the clock-like SBS1 channel (Tate et al.2019). BEAST v.1.10 (Suchard et al. 2018) was then used with the simple binarysubstitution model (since SBS1 effectively represents just C to T substitutions), a strict clock model, a constant population size prior (Kingman 1982), and a flat prior on the age of most recent common ancestor (from zero to earliest patient sample), with ancestral state estimation at the root. Chains were run and ESS values assessed as described above. The distance between the ancestral state of the root at each MCMC state, and the clock rate were used to calculate the expected evolution distance between the root and the known germline. This was used to inform the length of the branch between germline (at birth) and the MRCA of the samples. Survival Analysis Clinical analyses were performed in CLL for time to first treatment (TTFT) andoverall survival (OS) from the time of sampling. Tumour growth rate (^), effectivepopulation size (^^) and epigenetic switching rates were analysed as continuous variables in univariate Cox regression models for both TTFT and OS. The effect size of hazard ratios (HRs) for each evolutionary variable were analysed considering different scaling factors.In particular, the growth rate was analysed in exponential growth per year (i.e., for ^ = 1,the population per year is e=2.71 times bigger) ; the ^^was considered per million cells; the cancer age or time from the MRCA was analysed for each 10 years; the homozygousto heterozygous (de)methylation rates (^ and ^, respectively) and heterozygous tohomozygous methylation rate (^) were analysed multiplied by a factor of 100 (per allele per year); and heterozygous to homozygous demethylation rate (^) was multiplied by a factor of 1,000. In addition, growth rate and effective population were analysed ascontinuous variables in multivariate Cox regression models together with TP53 aberrations(considering mutations and deletions together), IGHV gene mutational status and the age of patients at sampling. Kaplan-Meier (KM) curves were generated for low and high growthP603125PC0057 rate and effective population size within IGHV subtypes using maximally selected log-rankstatistic using maxstats package (version 0.7-25). P values from KM curves were derivedusing the Log-rank statistic. Survival (version 3.5-7), survminer (version 0.4.9) and ggsurvfit (version 0.3.1) packages were used under R version 4.3.1. Supplementary Information In our previous work (Gabbutt et al. 2022), an analytic hidden Markov model was derived to describe the distribution of fluctuating methylation clocks in a fixed effectivepopulation of ^^ cells, replacing each other at a rate ^ per stem cell and with each fCpGloci allowed to alter its methylation state at a methylation / demethylation rate ^ / ^ respectively. In principle, the mathematics developed there applies for all values possible cell numbers; however, the number of possible states (i.e., the different possible combinations of heterozygous and homozygous (de)methylated cells at a specific CpGlocus) grows as ∼ ^^^, hence this approach is inappropriate for large populations of cells. Furthermore, the previous approach assumed a fixed number of cells, rather than a growing population. Hence, to extend our analysis to large, growing cell populations, such as those found within a blood cancer, an alternative approach to modelling the probability distribution that does not scale with the number of stem cells must be taken. Instead, wetook a stochastic approach to approximate the fCpG methylation distribution in anarbitrarily large, exponentially growing population. Rather than considering the individual states and the possible flows between them, we constructed equivalent stochastic jump differential equations. Finally, whereas previously it was the stochastic replacement of one stem cell by another that counteracted the ongoing (de)methylation, effectively resynchronising the methylation clocks upon a random clonal expansion; in the case of a growing cancer this neutral drift is negligible when compared to the exponential growth of each lineage within the tumour. In summary, then, our model consists of a single cell that undergoes random(de)methylation until time ^, at which point it begins exponentially growing at rate ^ peryear until the patient’s age at the time of sampling, ^, is reached. We allow for the possibility of allele specific methylation, and therefore require four epigenetic switching parameters to describe the system (fig 2A): ^^ – Homozygous to heterozygous methylation rate (per year)^ ^ – Heterozygous to homozygous methylation rate (per year)^ ^ - Homozygous to heterozygous demethylation rate (per year)^ ^ – Heterozygous to homozygous demethylation rate (per year)P603125PC0058 We employed four separate methylation switching parameters as these were necessary to reproduce the observed data. In CLL patient data, the distance between the homozygous unmethylated peak and the heterozygous methylated peak was almost always lower than the distance between the homozygous methylated peak and the heterozygous methylated peak (fig. 24A&B). A model where the homozygous toheterozygous rate is the same as the heterozygous to homozygous rate (i.e. ^ = ^ and ^ =^, fig.24A) provided a significantly worse fit to the data (accounting for the higher degrees of freedom in the 4-parameter model) than one in which the parameters are free to vary (fig. 24D). We assume that each methylation site is equally likely to be homozygousmethylated or unmethylated at ^ = 0, as in our previous work. We shall use the notationthat there are ^ cells with both alleles methylated, ^ cells with one allele methylated and^ cells with neither allele methylated. To calculate the methylation probability distributionof the single transformed cell at ^ = ^, the system is evolved according to the following setof differential equations: To calculate the initial state of the transformed cell at a particular locus at ^ = ^,the state (^, ^, ^) is drawn from a categorical distribution with probabilities^^(^; ^ = ^), ^(^; ^ = ^), ^(^; ^ = ^)^, which may be calculated using matrix exponentiation.The subsequent evolution of the system for ^ > ^ is then modelled using discrete timestochastic jump equations with a time step, ^^, that is sufficiently small that the probability of double jumps is negligible: The exponential growth and the ongoing (de)methylation of the cancer cell population are treated in two steps, first the population of the cancer grows according to the deterministic exponential growth equation (rounded to the closest integer): ^^(^) = ^^(^−^) With the new cells assigned to each population according to the multinomial distribution: ^^^, ^^, ^^^~Multinomial P603125PC0059 The flow from the homozygous methylated to the heterozygous methylated population due to ongoing demethylation can then be calculated as: Similarly, the flow from the homozygous unmethylated to the heterozygous methylated population is: To calculate the flow from the heterozygous population to the homozygous unmethylated and methylated populations, we first calculate the total number of cells that change state and then work out whether they flow to the homozygous methylated or unmethylated population: Updating the system at time ^ + ^^ is then: Note that the reason for this formulation, rather than simply calculating the number of cells that have switched from one state to another via each process (e.g. following a Poisson process) separately, is to strictly ensure that the total cell number is conservedand each population is strictly positive. The methylation fraction at time ^ can be simplycalculated as: The above assumes a perfectly pure tumour sample, however, there will inevitably be a degree of contamination from normal cells. To account for this contamination, we assumed that the contaminating normal cells did not share a recent common ancestor, sothe measured methylation of a tumour sample with purity ^ would be the weighted averageof the cancer methylation at that locus, ^^(^), and the average methylation of the contamination normal cells, Λ: ^(^) = ^^^(^) + (1 − ^)ΛThis assumption was informed by the observation that the standard deviation of the different cancer sample’s fCpG methylation distributions correlated linearly with the tumour purity (fig. 24E).P603125PC0060 Using the recursive relationship above, a single stochastic path can be generated. The methylation fraction probability distribution can then be numerically estimated by generating a series of ^^^^independent samples from the above process, the distribution of which will approximate the underlying hidden Markov model distribution. The Likelihood Function for a Growing Population The stochastic model above allowed samples to be drawn from the distribution offCpG methylation values given a set of parameters ^^, ^, ^, ^, ^, ^, ^, ^^ ^. We can use these samples to form an estimate of the likelihood function for a given set of data, ^.⃑ First, wecan bin the ^ stochastically generat ( )1 ^^^ ed ^ ^ into ^ bins with width^, to estimate aprobability mass function (p.m.f) for bins with centres (^ =1, 2, … , ^).To account for the noise introduced by the methylation array, we follow a similarapproach as in (Gabbutt et al. 2022), first employing a linear transform to account for thebackground noise of the array: ^^ = (^ − ∆)^^ + ∆. Then, we can model the measuredmethylation fraction values ^^ as a mixture of beta distributed random variables, drawn from a beta distribution with mean ^^and precision ^, weighted by the p.m.f.P^^^|^, ^, ^, ^, ^, ^, Λ^. Assuming that each measured fCpG site is independent, such that thelikelihood of a set of fCpG loci is the product the per-fCpG likelihood, the likelihood is then: ^^^^^ ℒ(^, ^, ^, ^, ^, ^, ^^ , ∆, ^, ^|^)⃑ = ^ ^ P^^^|^^, Δ, ^, ^^P^^^|^, ^, ^, ^, ^, ^, Λ^^^^ ^^^ It is often convenient to work with the log-likelihood, rather than the likelihooddirectly. Employing the logsumexp function, ^^^^(^^, … , ^^) = log(^^^+, … , +^^^), the log-likelihood is: ^ log^ℒ(^, ^, ^, ^, ^, ^, ^^ , ∆, ^, ^|^)⃑^ = ^ ^^^^ ^log ^P^^^|^^ , Δ, ^, ^^^ + log ^P^^^|^, ^, ^, ^, ^, ^, Λ^^^^^^ Here, we have modelled the growth rate and the time at which the MRCA emerged as independent; however, as we are assuming exponential growth, if we vary these two parameters independently, we implicitly assume that the total cancer burden can take onP603125PC0061implausibly high values. To penalise combinations of ^ and ^ that yield such high values,we add a penalisation factor for cancer burdens outside the plausible range{^^^^^ , ^^^^^}.^ ℒ^^^^ ^^^ − ^^^^^^^^^^^ = ^^^ − ^− ^^^Note that this is equivalent to placing a joint dependent prior over ^ and ^.Stochastic likelihood calculation correction The log-likelihood calculation above differs from a classical likelihood calculation inthat log(ℒ) is not an exact value, but rather an approximation of the “true” log-likelihoodwhich should become more exact as ^^^^ → ∞. To test whether this estimator issystematically biased, we simulated a set of ^⃑ data with a set of known parameters, thencalculated the log(ℒ) with the same set of parameters for a range of ^^^^ values (fig. 25A).We find that at low ^^^^values the approximation tends to systematically underestimate the log-likelihood, whilst as ^^^^increases, the calculated log-likelihood tends to asymptote to a single value. The increase in the mean estimated log-likelihood with ^^^^can be rationalised if one considers that we are attempting to estimate a continuous pdf by sampling from it, and as ^^^^increases, the stochastic noise from sampling is reduced and the approximation improves. In this sense, the fit at the peak of the log-likelihood distribution truly is more accurate, on average, as ^^^^increases, hence the calculated log-likelihood is higher. Unsurprisingly, the standard deviation of the calculated log-likelihood values falls as the number of simulated paths increases. 1 We found that the bias associated with having a finite ^^^^varied as log(ℒ^^^^)~ ^^^^. By varying different parameters, we found the log-likelihood bias could be estimated as: 0.5 + 0.4^3( ) 005^log ℒ^^^^ ≈ −^^^^Correcting for the bias yields log-likelihood estimates with a mean that does notvary with ^^^^ (fig. 25B). Hence, the bias-corrected cancer burden penalised log-likelihoodis:log(ℒ)′ = log(ℒ) + Example 2 Further work was carried out by the inventors to supplement the experiments and analysis discussed in Example 1. In summary:P603125PC0062 1. The inventors generated matched WGS data of their longitudinally-followed CLLcases that undergo Richter Transformation. They observed that subclonaldynamics inferred from the WGS data were mirrored in the methylation fCpG data.2. The inventors generated long-read Nanopore data for 10 samples, matched to theirexisting bead array data. These data prove that changes in methylation status at fCpGs are not because of underlying genetic mutation. The new sequencing-based data shows excellent concordance with the inventors’ previous array data provingthe fCpG signal is not technical noise. They have phased methylation patterns alongindividual long-reads, validating the allele-specific methylation of fCpGs and the inventors’ contention that fCpGs function as an evolving barcode. The inventorsobserved greater barcode heterogeneity in (polyclonal) normal B cells than in theclonally expanded CLL and RT samples. 3. The inventors generated matched SNP arrays to detect CNAs in their CLL (492samples) and MCL (85 samples) cohorts. Most of the samples from CLL and MCL show few CNAs and the inventors show that EVOFLUx model parameters are robustto CNAs. In addition, overlaying the fCpG methylation with the CNA data provides additional evidence that fCpGs fluctuate in an allele-specific manner.4. The inventors analysed the genotype-phenotype mapping. They added acomparison of the evolutionary parameters between CLLs by driver gene mutation status, observing, for example, that M-CLL cases with a TP53 mutation had significantly higher growth rates and effective population sizes than non-mutated cases. 5. The inventors carried out a comprehensive assessment of EVOFLUx robustness.They used simulated data with various assumptions weakened to test thedependency of EVOFLUx on these assumptions. The inventors found that EVOFLUxwas robust to both non-exponential growth of the cancer and heterogeneity of the epigenetic switching rate across fCpGs. 6. The inventors carried out a comparison of phylogenetic trees built using differentsets of non-fCpGs (Horvath and epicMIT methylation clocks, and random subsets of 1,000 CpGs from the 450K Illumina array) to the gold-standard WGS SNV tree. Unlike the trees built using fCpGs, the non-fCpG trees do not recover the correct topology or infer the correct times at which lineages diverged. Further details on the methods and results of this work are provided below.Methods External DNA methylation data of sorted and whole-blood samplesP603125PC0063 External DNA methylation data was download from the Gene expression Omnibusdatabase using the GEOquery R package (version 2.72.0). For sorted immune cells, theseinclude GSE137594 and GSE184269. For whole-blood samples, these include GSE72773, GSE55763, GSE40279 and GSE36054. Data was analysed with the normalization procedure used in each study together with the metadata provided. Mean and standard deviation for fCpGs were calculated with fCpGs present in the provided normalized matrices. Generation of long-read Oxford Nanopore dataFor long-read methylation sequencing in CLL and RT samples, concentration was assessedusing the Qubit assay and DNA integrity was analysed either with the Femto Pulse System (Agilent) or the Fragment Analyzer (Agilent). When more than 6ug of material with good integrity was available, DNA was additionally treated with the Short Fragment Eliminator Kit XS (PacBio) and eluted in EB buffer. Approximately 4ug of DNA was used for library preparation according to the standard LSK114 kit and protocol from Oxford Nanopore. The time for DNA repair and end-prep was increased up to 30 min at 20°C and 30 min at 65°C. Adapter ligation was performed for 1 hour at room temperature. All elutions were performed at 37°C for 1.5 hours, and 550-600ng of DNA were loaded onto a FLO-PRO114M (CLL cells) flow cells. Flow cells were washed (EXP-WSH004) after 1-2 days, if pore count decreased to less than 30%. A total of 1-4 washes were performed for each flow cell. Flow cells were run for 100 (CLL cells) hours in total with the Fast model (MinKNOW 23.11.7, Dorado 7.2.13). The raw data was rebasecalled using dorado duplex (version 0.5.3) andapplying the SUP and modified call to detect 5mC and 5hmC, (modeldna_r10.4.1_e8.2_400bps_sup@v4.3.0_5mCG_5hmCG@v1). In normal B cell samples, 1-3ug of DNA was used for whole genome sequencing. Libraries were prepared with the DNA ligation kit LSK110 with no modifications. Libraries were loaded onto a flow cell version FLO-PRO002 (R9.4) and were run for 90-110h. The basecalling was performed on live mode with the Guppy basecaller (version.6.2.7), included in the MinKNOW (version 22.08.6), using the SUP model for base modification detection of 5mC and 5hmC (dna_r9.4.1_450bps_modbases_5hmc_5mc_cg_sup.cfg). In all samples, the generated unmapped BAM files after the basecalling wereconverted to FASTQ files using samtools fastq -T Mm, Ml command. The FASTQ files werethen mapped to BAM files using the command minimap2 -ax map-ont -y .. / GCA_000001405.15_GRCh38_no_alt_analysis_set.fna.mmi. The methylation values were extracted from the BAMs into bedMethyl files using the in-house tool bam2bedmethyl (v0.3.2) and compressed / indexed using bgzip / tabix. Reads from each strand wereP603125PC0064 combined to generate DNA matrices for each CpG and were used for obtaining the methylation values of all fCpGs. In addition, mini BAM files containing all reads from the 978 fCpGs were generated. Subsequently, long-reads were phased using variants called using Clair 3 (v1.0.9, modelr941_prom_hac_g360+g422) (Zheng et al. 2022) with the Longphase package (v1.7) (J.H. Lin et al. 2022). The methylation status of each CpG was called using the modcallfunction within the Longphase package. At fCpGs, only 2.7% of the reads were non-canonical bases (fig. 27A). The variant allele frequency (VAF) of these mutations tended to be low and was negatively correlated with the coverage at that site (fig. 27B). Hence, the majority of these non-canonical base pairs are likely due to errors in nucleotide assignment. There is also no association between the methylation status of different reads and the variants present within a 50bp window of each fCpG locus (fig. 27C). Hence, assessment of fCpG methylation via bead array is not majorly confounded by miscalled variants. The fCpG methylation patterns seen in the bead array data were replicated in the long-read data (fig. 27D-E) and the correlation between the fraction methylated measured via bead array and long-read sequencing at fCpGs was excellent (fig. 27F). The same correspondence was observed in WGBS data (fig. 28). Characterisation of fCpGs To characterise the genomic and regulatory context of fCpGs, the inventors employed a series of statistical analyses and database annotations. In addition to theannotation used in Example 1, the inventors used the packagesSNPlocs.Hsapiens.dbSNP155.GRCh38 (version 0.99.24) and MafH5.gnomAD.v4.0.GRCh38(version 3.19) to check any possible germline or somatic genetic confounding on theresulting 978 fCpGs. The inventors found ~60% of fCpGs reported in gnomAD v4 database(with array background having ~65%), but with a very low minor allele frequency (mAF,median 1e-5, mean 1e-3). In addition, the inventors used the Illumina 450k and EPICarray internal SNP probes and showed a dramatically distinct methylation dynamicscompared to fCpGs (fig. 34 in single-timepoint and longitudinal samples.Adaption of simulations to the longitudinal setting The inventors modified the simulations of how the fCpG methylation distributionchanges over time to allow for multiple sequential sample collections. These simulations allow for neutral, independent clones, a single subclonal expansion or two subclonal expansions, which can either be nested or emerge from the clonal trunk in parallel. This required pre-specification of sampling times, along with the emergence times of anyP603125PC0065subclones / independent clones, which the inventors collected to form a set of “landmarktimes”. The discrete time steps of the simulation were split into phases between the landmark times, which evolved according to the discrete time Markov process outlined above. At each sampling time, the fCpG methylation fraction was calculated as above and stored as a column in the output matrix. Testing the modelling assumptions of EVOFLUx To interrogate their modelling assumptions that growth rate was constant and thatthe epigenetic switching rate of each fCpG was the same, the inventors simulated datafrom models where these assumptions have been weakened: 1) The growth rate of the cancer following a logistic growth model, rather thanexponential. This was parameterised using the carrying capacity (^, the maximum population size), such that the population size followed the growth equation ^ =^ (fig. 29). In the limit ^ → ∞, this growth equation reduces to theexponential growth equation used in their original model, ^ = ^^(^^^).2) Heterogeneity in the fCpG switching rate by fCpG site. The inventors assumed thatrather than each fCpG having the same epigenetic switching rates {^, ^, ^, ^}, the jthfCpG had switching rates equal to the mean switching rates multiplied by a multiplicative factor drawn from a lognormal distribution with mean 1 and standard deviation sigma ^ (fig. 30). In the limit ^ → 0, this model reduces to the originalhomogenous switching rate model. The inventors ran their EVOFLUx inference on these simulated datasets and testedhow each parameter varied with the carrying capacity / fCpG rate heterogeneity respectively. In the case of logistic growth, there was no difference between thedistributions of fCpG methylation as the carrying capacity is varied (fig. 29A-B). This isbecause the shape of the fCpG distribution, bar any ongoing subclonal selection, is driven by the early clonal dynamics of the cancer and the ongoing desynchronisation of fCpGs.Unsurprisingly, the inferred evolutionary parameters do not vary as a function of thecarrying capacity (fig. 29C-F).The distribution of fCpG methylation does change as a function of fCpG switchingrate heterogeneity (fig. 30A-B). Specifically, as the heterogeneity increases, the flankingclonal peaks of fCpGs homozygous methylated and demethylated in the founder cell shift to be closer to 1 and 0 respectively. This is due to the skew inherent in the lognormaldistribution employed to model the heterogeneity – the majority of heterogenous fCpGshave a switching rate lower than the mean of all fCpGs and thus shift less towards theP603125PC0066 steady state methylation value. In this case, whilst there is a weak correlation between the inferred parameters and the degree of fCpG switching rate heterogeneity, the ground truth value of the biological parameters of primary interest (i.e. the growth rate and MRCAtime) lie within the 95% credible interval as one varies the fCpG rate heterogeneity overa plausible regime (fig. 30C-F). Determining the rate of change in lymphocyte counts Historical records of the absolute number of lymphocytes in blood were obtained for CLL patients over the whole disease course (i.e., an approximate of the number ofmalignant CLL cells in blood). In 231 CLL patients, the inventors could obtain at least 10sample timepoints (i.e., 10 at least 10 medical appointments, median N=27, mean N=34) before the first treatment, allowing us to track the natural history of the disease prior totreatment intervention (fig. 31). The inventors fitted a linear model to all 231 cases andobtained the slope of the observed log number of lymphocytes (i.e., the coefficient of the univariate linear model) and compared it with growth rate estimates derived from EVOFLUx. Data availability Previously published DNA methylation data re-analysed in this study can be foundunder accession codes: B cells, EGAS00001001196; ALL, GSE56602, GSE49032,GSE76585, GSE69229; MCL, EGAS00001001637, EGAS00001004165; CLL,EGAD00010000871, EGAD00010000948, EGAD00010001975; MM, EGAS00001000841;DLBCL, EGAD00010001974. External DNA methylation data for sorted immune cells,GSE137594 and GSE184269. For whole-blood samples, GSE72773, GSE55763, GSE40279 and GSE36054. CLL gene expression data is available EGAS00001000374 and EGAS00001001306. ChIP-seq datasets are available from Blueprint https: / / www.blueprint-epigenome.eu / under the accession EGAS00001000326. Matched WES and WGS are available under accessions EGAS00000000092 andEGAD00001008954 respectively.Results Two approaches were used to evaluate the impact of the different fCpG selectioncriteria. First, the inventors repeated the fCpG selection process without including the LSor standard deviation filters. The inventors found 5336 CpGs that passed the other filtersbut were excluded by the LS filter (fig. 32A). These CpGs show clear disease specific clustering and stable hypo- / hypermethylation in normal cells. The 96 CpGs excluded byP603125PC0067 the standard deviation filter had intermediate methylation of in every sample, cancer or normal (fig. 32B). These could either represent CpGs that are stably maintained at 0.5 methylation, or those that undergo epimutations exceptionally quickly. As a secondapproach, the inventors repeated the fCpG derivation using mean methylation filters of0.15≤x ̅≤0.35 and 0.65≤x ̅≤0.85 to produce hypo- and hyper-fCpGs (fig. 33A-B). Withinindividual cancer samples, the hypo- and hyper-fCpGs still gave rise to trimodaldistributions, but with the occupancy of the peaks biased towards 0% and 100%methylated respectively (fig. 33C-D). Normal blood samples had unimodal distributionscentred around 0.25 / 0.75 respectively (fig. 33E-F). The standard deviation of the hypo- and hyper-fCpG methylation distribution were highly correlated with the standard deviation of the original fCpG standard deviation (R2=0.84 and R2=0.82 respectively, fig. 33G-H). Hence, fCpGs are a subset of CpGs prone to stochastic epigenetic errors which happen to have similar methylation and demethylation rates. The inventors next explored whether fCpGs can be confounded by the presence ofSNPs, as SNP-containing probes can follow a W-shaped distribution (due to commonheterozygous SNP variants). The inventors compared the features of array probescontaining SNPs (which were removed during the initial array processing) versus fCpGs. Unlike fCpGs, SNP probes show the same distribution in all samples, rather than exclusively in clonally expanded samples (fig.34A). Furthermore, whereas fCpGs continue to undergo stochastic epigenetic changes methylation over time, SNP-confounded probesare stable longitudinally (fig. 34B-C). Furthermore, to directly rule out that the fCpGs werenot related to SNPs, the inventors performed long read sequencing on 6 normal B cellsand two pairs of matched CLL / RT samples, which enables the measurement of methylation on single molecules of length ~50kb without the need for bisulfite conversion (Clarke etal. 2009). The inventors indeed identified that fCpGs reflect the cytosine methylationstatus rather than genetic variation. Moreover, the concordance between DNA methylation levels as measured via bead array and with long-read sequencing at fCpGs was excellent (fig.35) and the same correspondence was observed in whole genome bisulfite sequencing (WGBS) data (fig. 36). The length of long-reads allows for multiple neighbouring fCpGs to be observed onthe same read and for the reads to be phased by haplotype, directly testing their “evolvingbarcode” hypothesis. The inventors present an example of a 2kbp region on chromosome7 which contained 4 fCpGs that were often captured on the same read (fig. 37A). The phased reads from the CLL and RT samples showed significantly less heterogeneity than the reads from the B cell samples (p=0.038, fig. 37B). The patterns of fCpG methylation in the CLL / RT reads were also markedly different between alleles (p=0.030, fig. 37B), demonstrating epimutations on fCpGs occur in an allele-specific manner.P603125PC0068 To validate fCpG patterns fluctuate in an allele specific manner, the inventors checked the methylation distributions of fCpGs located in copy number alterations (CNAs) using matched SNP array data in 492 CLL and 85 MCL samples (Puente et al.2015; Nadeu et al. 2020). As expected, CNAs were relatively rare in CLL samples, but more common inMCL samples (fig. 27A). In the case of fCpGs present on monosomic DNA, one shouldexpect to see peaks only at 0 and 1 (fig. 27B), whereas fCpGs present on fCpGs with copynumber 3 should have 4 peaks (0,2 ^ 3 , 1). Indeed, the fCpG methylation distributionsfor these regions exactly match these predictions (fig. 27C-D). Few fCpGs were located in copy-altered regions (mean 14.5 / 978 fCpGs in CLL, 45.7 / 978 in MCL, fig. 27E-F) As the DNA methylome is influenced by age (Seale et al. 2022), the inventors further explored whether fCpGs may be subjected to age-dependent epigeneticmodulation. The inventors did not find any association between age and mean fCpGmethylation in any of their normal blood samples or in independent cohorts of normalblood cells (N=4,099, p>0.05, fig. 28A-B). However, employing the standard deviation ofeach sample’s fCpG methylation distribution as a proxy for its “W-shape”, the inventors observed a significant correlation between the standard deviation of fCpGs and age in several normal cell subpopulations, which is consistent with the well-described aging- related clonal expansions in the hematopoietic system (p<0.05, fig. 28C-D) (Jaiswal and Ebert 2019; Boddicker et al. 2023). As expected, highly enriched clonal populations of lymphoid cancers had a much higher standard deviation than both peripheral blood cellsand sorted B and T lymphocytes regardless the age of the donor (fig 1D, fig 28D).Moreover, B-ALL samples from patients in remission show a marked decrease in the fCpGstandard deviation compared to the matched malignant B-ALL sample (p=9.6e-39, pairedt-test, fig. 38A), similar to that observed in normal blood samples (p=0.067, MW-U, fig.38B). The inventors also investigated whether the methylation status of nearby fCpGswere correlated. In their long-read data, some neighbouring fCpGs within 1kbp did showlocal correlation (fig. 39A), consistent with prior work (Affinito et al. 2020). Using theirbead array data, the inventors found that the pairwise correlation between fCpGs on thesame chromosome tended to be high for fCpGs within approximately 1 kbps, but rapidly dropped to 0 at greater distances (fig. 39B-D). Only a very small fraction of fCpGsdisplayed these correlations (74 / 978) – notably, the 4 nearby fCpGs on chr7 did not displaythese correlations (fig. 37A). As discussed in Example 1, the inventors developed a stochastic modelling and inference framework called EVOFLUx to simulate how clonal evolution quantitatively determines fluctuating methylation values and enables inference of evolutionary dynamicson a case-by-case basis from patient data. The inventors adapted their simulations toP603125PC0069 model the case of a single subclone gaining a selective advantage some time following the MRCA, sampling the population longitudinally as the subclone expands. Between the MRCAand the subclone acquiring a selective advantage, a small fraction of the fCpGs willundergo a fluctuation, marking the lineage of the outgrowing subclone. By comparing the fCpG methylation patterns at subsequent timepoints, the intervening evolutionary dynamics can be elucidated (fig. 40). Comparing two early timepoints (T1 vs T2) where the subclone makes up a small fraction of the total population (0.1% and 1% respectively), the fCpG methylation all lies on the x-y line. In contrast, by later timepoint (T3 and T4)the subclone makes up a significant fraction of the population (57% and 99%respectively), and comparison with T1 show a number of fCpGs which lie off the x-y line. These represent fCpGs loci that have switched methylation state in the lineage that formed the subclone within the intervening time between the MRCA of the sample population and the emergence of the outgrowing subclone. Notably, at timepoint T3 where the subclonalexpansion has not yet swept to fixation, these discordant fCpGs form additionalintermediate peaks between the three clonal peaks. However, once the subclone has swept to fixation, these additional peaks are subsumed into the clonal peaks, causing a pairwise comparison of T1 or T2 vs T4 to have a characteristic “checkerboard” pattern. EVOFLUx utilises an extensive Bayesian inference method to learn model parameters from input fCpG methylation distribution data. To evaluate the accuracy of theinference, the inventors simulated fCpG data and tested the ability of EVOFLUx to recoverthe known input parameters (fig. 7D). These values were inferred with good confidence, with the 95% credible intervals encompassing the ground truth values (fig.6F), significant narrowing of the posterior compared to the prior and with the posterior predictive distribution well-matching that of the original simulated dataset. The EVOFLUx inference is robust to weakening the assumptions, such as strict exponential growth (fig. 29) orhomogeneity in the fCpG fluctuating rate (fig. 30). Notably, varying the carrying capacityof the simulated tumour has no effect upon the fCpG methylation distribution, demonstrating that the patterns in the data are driven by the early clonal dynamics of the tumour. Furthermore, repeatedly downsampling the number of included fCpGs revealed that the majority of parameters were largely insensitive to the number of fCpGs included in the inference, particularly above ~600 (fig. 41). Hence, these simulation experiments demonstrate that the parameters defining the early evolutionary dynamics of growing cancers can be measured using bulk methylation arrays. The evolutionary dynamics of lymphoid malignancies In CLL, MCL, MM and DLBCL-NOS there is either no or a very weak correlation between the age of the patient and the mean epigenetic switching rate, but in both B-ALL and T- ALL there was a strong negative correlation (fig. 42, P= 2.4e-98, R2=0.44 & P=5.9e-06,P603125PC0070 R2=0.22 respectively). Together, these results suggests that fCpG switching rates are elevated during childhood, but fall to a steady rate during adulthood, although the difference in epigenetic switching rates could perhaps be due to disease specific alterations in methylation maintenance rather than a general feature of normal early development.Notably, when comparing the four epigenetic switching parameters (^, ^, ^ and ^), thedifferent cancer types occupied different regions of the epimutation parameter space (fig. 43F). The model underpinning EVOFLUx includes the assumption that fCpGs lie on a diploid region of the genome but copy number alterations (CNAs) are common in cancer.The inventors extracted all those fCpGs which were not affected by CNAs (fig. 27E-F) andre-ran their EVOFLUx inference method on just those diploid fCpGs. There was a very highcorrelation between the inferred growth rate, time since the MRCA and effective population size between the original EVOFLUx inference using all fCpGs and the subset of diploid fCpGs (fig. 44A). To ensure that this isn’t driven by the large number of samples with fewCNAs, the inventors correlated the log-fold change between the diploid vs originalEVOFLUx inference and the number of fCpGs not included in the inference. Both the growth rate and the time since the MRCA had no correlation in either cohort (fig.44B), suggesting that these parameters are not sensitive to CNAs. There was a correlation between the log- fold change in the effective population size and the number of non-diploid fCpGs, but this is likely the result of including fewer fCpGs in the inference (fig.41), rather than the effect of CNAs per se. The Majority of Lymphoid Cancers Grow Neutrally To validate that EVOFLUx could accurately detect subclonal selection, the inventorsran MOBSTER on a matched set of WGS data from 127 samples (depth ~30X). Similarly, samples identified as subclonal by WGS had a significantly higher EVOFLUx subclonality weighting (p=3.1e-4, fig. 45A). Unlike WES, there was no relationship between the number of mutations called and the probability of being called as subclonal by MOBSTER(p>0.05, fig. 45B). In a subset of 122 samples where there was also matched WES datathe inventors generated competing logistic regression classifiers to determine the accuracyof predicting the WGS call from WES alone, EVOFLUx alone, and the two combined (fig. 46). The EVOFLUx classifier (AUC=0.73) was notably better than the assessment from WES data (AUC=0.62) and only minimally improved by the inclusion of the WES data (AUC=0.74).Ongoing Clonal Evolution is Recorded by fCpGs in Longitudinal SamplesP603125PC0071 A small fraction of CLL patients will undergo Richter transformation (RT), theemergence of an aggressive phenotype with dismal prognosis. For two of the samepatients analysed previously by WGS (Nadeu et al. 2022), the inventors also had DNAmethylation arrays of longitudinal samples from diagnosis to RT (fig. 47A). Comparing theinferred subclonal relationships from the WGS with the pairwise fCpG patterns revealed aremarkable consistency. In their first case, in the 13.1 years between T1 and T2, a nested subclonalexpansion occurs, which in the pairwise comparison of fCpG methylation leads to an enrichment on the off-diagonal (fig. 47A). Conversely, between T2 and T3 the clonal composition of the samples is highly stable, so the pairwise fCpG methylation is strictly contained on the diagonal. At the final timepoint T4 the RT clone had expanded to a clonal fraction of 77%, leading to a further enrichment on the off-diagonal when compared to T3. In the second case, for the first three timepoints (T1-T4) the inventors observe theslow expansion of a subclone, which is consistent with the bulk of the fCpGs lying on the diagonal whilst a small fraction is discordant (fig. 47B). The large nested subclonal expansion at T5 shows a similar enrichment of fCpGs on the off-diagonal as in the first case. Finally, at timepoint T6 the RT clone sweeps through the population until it is almost fixed, which is reflected in the checkerboard fCpG methylation patterns. Hence, the patterns of fCpG methylation directly encode the ongoing clonal evolution that can be reconstructed via WGS. As discussed in Example 1, the inventors used EVOFLUX to construct phylogeneticrelationships between samples from differences in fCpG methylation, extracting valuable temporal information regarding the divergence of RT cells from the ancestral clone. Theinventors found the inferred phylogeny well matched the known clinical relationshipsbetween the different samples (fig. 47A-B). Notably, the evolutionary distance betweentwo technical replicates of untreated samples (450k and EPIC data) was small in both cases. In both phylogenies, the RT clone diverged exceptionally early, roughly a decade before the MRCA of the non-transformed samples (9 and 12 years for case 12 and 19,respectively). This is consistent with the inventor’s previous work (Nadeu et al. 2022) inwhich they detected RT cells at low frequencies within the diagnostic CLL sample, butsuggests that the initial RT divergence occurred well before diagnosis, over 30 years before the clinical presentation of RT. To validate the fCpG phylogenetic inference, the inventors built phylogenetic treeson alternative sets of methylation sites that have not been strictly filtered to ensure they encode lineage (i.e. random subsets of 1000 CpGs (fig. 48), and the Horvath (Horvath2013) and epiCMIT mitotic clock CpGs (Duran-Ferrer et al. 2020) (fig. 49-50)). TheP603125PC0072inventors then compared the concordance of those methylation-based trees with thosebuilt using SNV data from matched WGS (Nadeu et al. 2022) (fig. 49-50). The inventorsrestricted their SNV analysis to clock-like SBS1 mutations (Nik-Zainal et al. 2012) whichwere clonal within the sample (Caravagna et al. 2020). The topologies of the SNVphylogenies exactly matched that of the fCpG phylogenies but were different from the alternative methylation-based trees (fig. 49-50). In all cases the timings of the branch points (as measured by the mean pairwise error of branch coalescent times) was more consistent between the fCpG and SNV trees than any of the other methylation-based trees. Hence, fCpGs have the potential to serve as a high temporal resolution phylogenetic character, enabling multi-sample phylogenetic reconstruction at a fraction of the cost of WGS. As discussed in Example 1, the inventors applied their phylogenetic analysis to 3longitudinal B-ALL samples, tracking the disease course from initial diagnosis to relapse using the same methodology (fig. 51). In each of the patients, the subsequent relapse samples formed a separate clade from the initial diagnostic samples, suggesting the cancer population has undergone a significant bottleneck through treatment. For 2 B-ALL patients, the inventors have longitudinal samples covering the initialpresentation, remission and subsequent relapse. In these samples the unimodal distribution observed during remission is replaced by a W-shaped distribution, mimicking the initial clonal expansion (fig.47C-D). Comparing the fCpG distributions of the diagnostic vs relapse samples and comparing to the different scenarios outlined by simulation (fig. 40), the first case shows evidence of a strong bottleneck following treatment, whereas the second shows ongoing subclonal selection between the initial clone and an emerging subclone. Evolutionary growth dynamics of CLL predicts clinical outcomes Patients with mutations in specific driver genes, such as TP53, are well-known tohave a worse prognosis (Zenz et al. 2010). The inventors compared the inferred growthrates and effective population sizes accounting for IGHV status for the most prevalentdriver mutations in CLL: TP53, SF3B1, NOTCH1, ATM, POT1 and IGLV3-21R110, del11q22.3,del13q14.3, del17p13.1 and trisomy12 (fig. 52). The inventors found that M-CLL patientswith a TP53 mutations had a higher growth rate and effective population size (p=0.030and p=0.036 respectively, MW-U test, FDR corrected, fig. 53A-B). Contrastingly, U-CLL patients with an IGLV3-21R110mutation had a lower growth rate and effective population size (p=0.0032 and p=0.036 respectively). Interestingly, both U-CLL and M-CLL patients with trisomy12 had an increased effective population size (p=0.036 and p=0.036 respectively), but no difference in growth rate.P603125PC0073 Finally, the inventors found a moderate correlation between the EVOFLUx inferredevolutionary growth rate and observational estimates of the growth rate derived from longitudinal clinical measurements of lymphocyte count preceding treatment (p=2e-5, R=0.27, fig. 31). Importantly, in a multivariate regression including both these measures along with IGHV status, they are both independently predictive of TTFT (fig. 31D).Together with their previous findings from simulations that the pattern of fCpGs dependon the early growth dynamics of the tumours, this highlights the distinction between the contemporary growth of the tumour and the formerly inaccessible early evolutionarydynamics of the tumour revealed by EVOFLUx.Example 3 The inventors have demonstrated that EVOFLUx can be employed to stratify clinical outcomes pan-cancer. fCpGs fluctuate in a tissue / disease-specific context. The inventors therefore derived a novel set of colorectal cancer (CRC) fCpGs using the S-CORT cohort of samples (https: / / www.s-cort.org). This database contains Illumina EPIC bead array methylation data for: 985 primary tumours, 140 metastatic tumours and 10 normal samples. To derive the CRC-fCpGs, the metrics outlined in Example 1 were calculated for high purity cancer samples (n=288). Furthermore, the mean methylation in normal colon samples was calculated for each site, only selecting those with mean methylation between 0.4 and 0.6. Similarly, the inventors calculated the mean methylation in 7 normal blood cells where they had EPIC array data and selected only CpGs with mean methylation between 0.35 and 0.65. The distribution of the calculated Laplacian Score (i.e. informativeness of each CpG) was very different than in the lymphoid case, so a cutoff of the top 50% least informative CpGs was employed (this was a similar absolute cutoff as the lymphoid case, 0.706 vs 0.763 in the CRC and lymphoid cases respectively). All other thresholds were the same as in the lymphoid case (e.g. mean cancer methylation between 0.4 and 0.6, standard deviation across the cancer samples greater than 0.15). This process yielded an initial set of 1,454 CRC-specific fCpGs. In Examples 1 and 2, the inventors presented evidence that lymphoid fCpGs within ~1kbp can experience local correlations in methylation status. To remove these correlated fCpGs such that each site fluctuates independently, the inventors calculated the pairwise corelation between every fCpG on the same chromosome. The inventors then removed fCpGs that had a correlation coefficient of 0.5 or greater, leaving a set 1,388 CRC-specific fCpGs.P603125PC0074 The inventors ran EVOFLUx on 150 primary tumour samples, using the same inference process as described in Example 1. For 114 samples, the inventors had matched information of the progression free survival (PFS) of that patient. They found that theinferred effective population size was predictive of PFS (p=0.0384, log-rank test). SeeFigure 54. Furthermore, in a multivariate regression controlling for tumour purity, theeffective population size was still significantly predictive of PFS (p=0.0048, HR=2.31 (1.29-4.13, 95% CI), multivariate Cox proportional hazards). See Figure 55. Thus, the evolutionary history of a cancer, which can quantified via EVOFLUx, is predictive of clinical outcomes not just in the context of lymphoid cancers, but also solid tumours. References Affinito, Ornella, Domenico Palumbo, Annalisa Fierro, Mariella Cuomo, Giulia De Riso, Antonella Monticelli, Gennaro Miele, Lorenzo Chiariotti, and Sergio Cocozza. 2020. ‘Nucleotide Distance Influences Co-Methylation between Nearby CpG Sites’. Genomics 112 (1): 144–50. https: / / doi.org / 10.1016 / J.YGENO.2019.05.007. Agirre, Xabier, Giancarlo Castellano, Marien Pascual, Simon Heath, Marta Kulis, Victor Segura, Anke Bergmann, et al. 2015. ‘Whole-Epigenome Analysis in Multiple Myeloma Reveals DNA Hypermethylation of B Cell-Specific Enhancers’. Genome Research 25(4): 478–87. https: / / doi.org / 10.1101 / GR.180240.114. Alizadeh, Ash A., Michael B. Elsen, R. Eric Davis, Ch L. Ma, Izidore S. Lossos, Andreas Rosenwald, Jennifer C. Boldrick, et al. 2000. ‘Distinct Types of Diffuse Large B-Cell Lymphoma Identified by Gene Expression Profiling’. Nature 403 (6769): 503–11.https: / / doi.org / 10.1038 / 35000501. Andrieu, Christophe, and Gareth O. Roberts. 2009. ‘The Pseudo-Marginal Approach for Efficient Monte Carlo Computations’. Https: / / Doi.Org / 10.1214 / 07-AOS574 37 (2):697–725. https: / / doi.org / 10.1214 / 07-AOS574. Arber, Daniel A., Attilio Orazi, Robert P. Hasserjian, Michael J. Borowitz, Katherine R. Calvo, Hans Michael Kvasnicka, Sa A. Wang, et al. 2022. ‘International Consensus Classification of Myeloid Neoplasms and Acute Leukemias: Integrating Morphologic, Clinical, and Genomic Data’. Blood 140 (11): 1200–1228.https: / / doi.org / 10.1182 / BLOOD.2022015850. Aryee, Martin J, Andrew E Jaffe, Hector Corrada-Bravo, Christine Ladd-Acosta, Andrew P Feinberg, Kasper D Hansen, and Rafael A Irizarry. 2014. ‘Minfi: A Flexible and Comprehensive Bioconductor Package for the Analysis of Infinium DNA Methylation Microarrays’. Bioinformatics 30 (10): 1363–69.https: / / doi.org / 10.1093 / bioinformatics / btu049.P603125PC0075 Ashburner, Michael, Catherine A. Ball, Judith A. Blake, David Botstein, Heather Butler, J. Michael Cherry, Allan P. Davis, et al. 2000. ‘Gene Ontology: Tool for the Unification of Biology’. Nature Genetics 2000 25:1 25 (1): 25–29. https: / / doi.org / 10.1038 / 75556.Baele, Guy, Philippe Lemey, Trevor Bedford, Andrew Rambaut, Marc A. Suchard, and Alexander V. Alekseyenko. 2012. ‘Improving the Accuracy of Demographic and Molecular Clock Model Comparison While Accommodating Phylogenetic Uncertainty’. Molecular Biology and Evolution 29 (9): 2157–67.https: / / doi.org / 10.1093 / MOLBEV / MSS084. Beekman, Renée, Vicente Chapaprieta, Núria Russiñol, Roser Vilarrasa-Blasi, Núria Verdaguer-Dot, Joost H.A. Martens, Martí Duran-Ferrer, et al. 2018. ‘The Reference Epigenome and Regulatory Chromatin Landscape of Chronic Lymphocytic Leukemia’. Nature Medicine 2018 24:6 24 (6): 868–80. https: / / doi.org / 10.1038 / s41591-018-0028-4. Belsky, D. W., A. Caspi, D. L. Corcoran, K. Sugden, R. Poulton, L. Arseneault, A. Baccarelli, et al. 2022. ‘DunedinPACE, A DNA Methylation Biomarker of the Pace of Aging’. ELife 11 (January). https: / / doi.org / 10.7554 / ELIFE.73420. Bocklandt, Sven, Wen Lin, Mary E. Sehl, Francisco J. Sánchez, Janet S. Sinsheimer, Steve Horvath, and Eric Vilain. 2011. ‘Epigenetic Predictor of Age’. PLOS ONE 6 (6): e14821.https: / / doi.org / 10.1371 / JOURNAL.PONE.0014821. Boddicker, Nicholas J., Sameer A. Parikh, Aaron D. Norman, Kari G. Rabe, Rosalie Griffin, Timothy G. Call, Dennis P. Robinson, et al. 2023. ‘Relationship among Three Common Hematological Premalignant Conditions’. Leukemia 37 (8): 1719–22.https: / / doi.org / 10.1038 / S41375-023-01914-Z. Bohlin, J., S. E. Håberg, P. Magnus, S. E. Reese, H. K. Gjessing, M. C. Magnus, C. L. Parr, C. M. Page, S. J. London, and W. Nystad. 2016. ‘Prediction of Gestational Age Based on Genome-Wide Differentially Methylated Regions’. Genome Biology 17 (1): 1–9.https: / / doi.org / 10.1186 / S13059-016-1063-4 / TABLES / 5. Brocks, David, Yassen Assenov, Sarah Minner, Olga Bogatyrova, Ronald Simon, Christina Koop, Christopher Oakes, et al. 2014. ‘Intratumor DNA Methylation Heterogeneity Reflects Clonal Evolution in Aggressive Prostate Cancer’. Cell Reports 8 (3): 798–806.https: / / doi.org / 10.1016 / J.CELREP.2014.06.053 / ATTACHMENT / 0EB6C7A6-FCFD- 4977-A0EC-D8324921D744 / MMC1.PDF. Caravagna, Giulio, Timon Heide, Marc J. Williams, Luis Zapata, Daniel Nichol, Ketevan Chkhaidze, William Cross, et al. 2020. ‘Subclonal Reconstruction of Tumors by Using Machine Learning and Population Genetics’. Nature Genetics 2020 52:9 52 (9): 898–907. https: / / doi.org / 10.1038 / s41588-020-0675-5.P603125PC0076 Carpenter, Bob, Andrew Gelman, Matthew D Hoffman, Daniel Lee, Ben Goodrich, Michael Betancourt, Marcus Brubaker, Jiqiang Guo, Peter Li, and Allen Riddell. 2017. ‘Stan: A Probabilistic Programming Language’. Journal of Statistical Software; Vol 1, Issue 1 (2017), January. https: / / www.jstatsoft.org / v076 / i01. Chen, Yi An, Mathieu Lemire, Sanaa Choufani, Darci T. Butcher, Daria Grafodatskaya, Brent W. Zanke, Steven Gallinger, Thomas J. Hudson, and Rosanna Weksberg. 2013. ‘Discovery of Cross-Reactive Probes and Polymorphic CpGs in the Illumina Infinium HumanMethylation450 Microarray’. Epigenetics 8 (2): 203–9.https: / / doi.org / 10.4161 / epi.23470. Clarke, James, Hai Chen Wu, Lakmal Jayasinghe, Alpesh Patel, Stuart Reid, and Hagan Bayley. 2009. ‘Continuous Base Identification for Single-Molecule Nanopore DNA Sequencing’. Nature Nanotechnology 4 (4): 265–70.https: / / doi.org / 10.1038 / NNANO.2009.12. Damle, Rajendra N., Tarun Wasil, Franco Fais, Fabio Ghiotto, Angelo Valetto, Steven L. Allen, Aby Buchbinder, et al. 1999. ‘Ig V Gene Mutation Status and CD38 Expression As Novel Prognostic Indicators in Chronic Lymphocytic Leukemia’. Blood 94 (6): 1840–47. https: / / doi.org / 10.1182 / BLOOD.V94.6.1840. Díaz-Gay, Marcos, Raviteja Vangara, Mark Barnes, Xi Wang, S M Ashiqul Islam, Ian Vermes, Nithish Bharadhwaj Narasimman, et al. 2023. ‘Assigning Mutational Signatures to Individual Samples and Individual Somatic Mutations with SigProfilerAssignment’. BioRxiv, July, 2023.07.10.548264. https: / / doi.org / 10.1101 / 2023.07.10.548264. Dietrich, Sascha, Małgorzata Oleś, Junyan Lu, Leopold Sellner, Simon Anders, Britta Velten, Bian Wu, et al.2018. ‘Drug-Perturbation-Based Stratification of Blood Cancer’. The Journal of Clinical Investigation 128 (1): 427–45.https: / / doi.org / 10.1172 / JCI93801. Drummond, Alexei J., and Andrew Rambaut. 2007. ‘BEAST: Bayesian Evolutionary Analysis by Sampling Trees’. BMC Evolutionary Biology 7 (1).https: / / doi.org / 10.1186 / 1471-2148-7-214. Drummond, Alexei J., Marc A. Suchard, Dong Xie, and Andrew Rambaut. 2012. ‘Bayesian Phylogenetics with BEAUti and the BEAST 1.7’. Molecular Biology and Evolution 29(8): 1969–73. https: / / doi.org / 10.1093 / MOLBEV / MSS075. Dunham, Ian, Anshul Kundaje, Shelley F. Aldred, Patrick J. Collins, Carrie A. Davis, Francis Doyle, Charles B. Epstein, et al. 2012. ‘An Integrated Encyclopedia of DNA Elements in the Human Genome’. Nature 489 (7414): 57–74.https: / / doi.org / 10.1038 / NATURE11247.P603125PC0077 Duran-Ferrer, Martí, Guillem Clot, Ferran Nadeu, Renée Beekman, Tycho Baumann, Jessica Nordlund, Yanara Marincevic-Zuniga, et al. 2020. ‘The Proliferative History Shapes the DNA Methylome of B-Cell Tumors and Predicts Clinical Outcome’. Nature Cancer 1 (11): 1066–81. https: / / doi.org / 10.1038 / s43018-020-00131-2.Endicott, Jamie L., Paula A. Nolte, Hui Shen, and Peter W. Laird. 2022. ‘Cell Division Drives DNA Methylation Loss in Late-Replicating Domains in Primary Human Cells’. Nature Communications 2022 13:1 13 (1): 1–12. https: / / doi.org / 10.1038 / s41467-022-34268-8. Fortin, Jean-Philippe, Timothy J Triche Jr, and Kasper D Hansen. 2017. ‘Preprocessing, Normalization and Integration of the Illumina HumanMethylationEPIC Array with Minfi’. Bioinformatics 33 (4): 558–60.https: / / doi.org / 10.1093 / bioinformatics / btw691. Gabbutt, Calum, Ryan O. Schenck, Daniel J. Weisenberger, Christopher Kimberley, Alison Berner, Jacob Househam, Eszter Lakatos, et al. 2022. ‘Fluctuating Methylation Clocks for Cell Lineage Tracing at High Temporal Resolution in Human Tissues’. Nature Biotechnology 2022, January, 1–11. https: / / doi.org / 10.1038 / s41587-021-01109-w. Gaiti, Federico, Ronan Chaligne, Hongcang Gu, Ryan M Brand, Steven Kothen-Hill, Rafael C Schulman, Kirill Grigorev, et al. 2019. ‘Epigenetic Evolution and Lineage Histories of Chronic Lymphocytic Leukaemia’. Nature 569 (7757): 576–80.https: / / doi.org / 10.1038 / s41586-019-1198-z. Garagnani, Paolo, Maria G. Bacalini, Chiara Pirazzini, Davide Gori, Cristina Giuliani, Daniela Mari, Anna M. Di Blasio, et al. 2012. ‘Methylation of ELOVL2 Gene as a New Epigenetic Marker of Age’. Aging Cell 11 (6): 1132–34. https: / / doi.org / 10.1111 / ACEL.12005.Greaves, Mel, and Carlo C. Maley. 2012. ‘Clonal Evolution in Cancer’. Nature 2012 481:7381 481 (7381): 306–13. https: / / doi.org / 10.1038 / nature10762.Griffiths, R. C., and S. Tavaré. 1994. ‘Sampling Theory for Neutral Alleles in a Varying Environment’. Philosophical Transactions of the Royal Society of London. Series B, Biological Sciences 344 (1310): 403–10. https: / / doi.org / 10.1098 / RSTB.1994.0079.Gruber, Michaela, Ivana Bozic, Ignaty Leshchiner, Dimitri Livitz, Kristen Stevenson, Laura Rassenti, Daniel Rosebrock, et al. 2019. ‘Growth Dynamics in Naturally Progressing Chronic Lymphocytic Leukaemia’. Nature 570 (7762): 474–79.https: / / doi.org / 10.1038 / S41586-019-1252-X. Gutierrez, Catherine, Aziz M. Al’Khafaji, Eric Brenner, Kaitlyn E. Johnson, Satyen H. Gohil, Ziao Lin, Binyamin A. Knisbacher, et al. 2021. ‘Multifunctional Barcoding with ClonMapper Enables High-Resolution Study of Clonal Dynamics during Tumor Evolution and Treatment’. Nature Cancer 2 (7): 758–72.https: / / doi.org / 10.1038 / S43018-021-00222-8.P603125PC0078 Hamblin, Terry J., Zadie Davis, Anne Gardiner, David G. Oscier, and Freda K. Stevenson. 1999. ‘Unmutated Ig VH Genes Are Associated With a More Aggressive Form of Chronic Lymphocytic Leukemia’. Blood 94 (6): 1848–54.https: / / doi.org / 10.1182 / BLOOD.V94.6.1848. Hannum, Gregory, Justin Guinney, Ling Zhao, Li Zhang, Guy Hughes, Srini Vas Sadda, Brandy Klotzle, et al. 2013. ‘Genome-Wide Methylation Profiles Reveal Quantitative Views of Human Aging Rates’. Molecular Cell 49 (2): 359–67.https: / / doi.org / 10.1016 / J.MOLCEL.2012.10.016. Hao, Jia Jie, De Chen Lin, Huy Q. Dinh, Anand Mayakonda, Yan Yi Jiang, Chen Chang, Ye Jiang, et al. 2016. ‘Spatial Intratumoral Heterogeneity and Temporal Clonal Evolution in Esophageal Squamous Cell Carcinoma’. Nature Genetics 201648:1248 (12): 1500–1507. https: / / doi.org / 10.1038 / ng.3683. He, Xiaofei, Deng Cai, and Partha Niyogi. 2005. ‘Laplacian Score for Feature Selection’. In Advances in Neural Information Processing Systems, edited by Y Weiss, B Schölkopf, and J Platt. Vol. 18. MIT Press. https: / / proceedings.neurips.cc / paper / 2005 / file / b5b03f06271f8917685d14cea7c6c50 a-Paper.pdf. Heide, Timon, Luis Zapata, Marc J. Williams, Benjamin Werner, Giulio Caravagna, Chris P. Barnes, Trevor A. Graham, and Andrea Sottoriva. 2018. ‘Reply to “Neutral Tumor Evolution?”’ Nature Genetics 2018 50:12 50 (12): 1633–37.https: / / doi.org / 10.1038 / s41588-018-0256-z. Hong, You Jin, Paul Marjoram, Darryl Shibata, and Kimberly D. Siegmund. 2010. ‘Using DNA Methylation Patterns to Infer Tumor Ancestry’. PloS One 5 (8).https: / / doi.org / 10.1371 / JOURNAL.PONE.0012002. Horvath, Steve. 2013. ‘DNA Methylation Age of Human Tissues and Cell Types’. Genome Biology 14 (10): R115. https: / / doi.org / 10.1186 / gb-2013-14-10-r115.Horvath, Steve, Junko Oshima, George M. Martin, Ake T. Lu, Austin Quach, Howard Cohen, Sarah Felton, et al. 2018. ‘Epigenetic Clock for Skin and Blood Cells Applied to Hutchinson Gilford Progeria Syndrome and Ex Vivo Studies’. Aging 10 (7): 1758–75.https: / / doi.org / 10.18632 / AGING.101508. Jaiswal, Siddhartha, and Benjamin L. Ebert. 2019. ‘Clonal Hematopoiesis in Human Aging and Disease’. Science (New York, N.Y.) 366 (6465).https: / / doi.org / 10.1126 / SCIENCE.AAN4673.Kingman, J. F.C. 1982. ‘The Coalescent’. Stochastic Processes and Their Applications 13(3): 235–48. https: / / doi.org / 10.1016 / 0304-4149(82)90011-4. Knight, Anna K., Jeffrey M. Craig, Christiane Theda, Marie Bækvad-Hansen, Jonas Bybjerg- Grauholm, Christine S. Hansen, Mads V. Hollegaard, et al. 2016. ‘An Epigenetic ClockP603125PC0079 for Gestational Age at Birth Based on Blood Methylation Data’. Genome Biology 17(1): 1–11. https: / / doi.org / 10.1186 / S13059-016-1068-Z / FIGURES / 4. Knisbacher, Binyamin A., Ziao Lin, Cynthia K. Hahn, Ferran Nadeu, Martí Duran-Ferrer, Kristen E. Stevenson, Eugen Tausch, et al. 2022. ‘Molecular Map of Chronic Lymphocytic Leukemia and Its Impact on Outcome.’ Nature Genetics 54 (11): 1664–74. https: / / doi.org / 10.1038 / S41588-022-01140-W. Kulis, Marta, Angelika Merkel, Simon Heath, Ana C. Queirós, Ronald P. Schuyler, Giancarlo Castellano, Renée Beekman, et al. 2015. ‘Whole-Genome Fingerprint of the DNA Methylome during Human B Cell Differentiation’. Nature Genetics 2015 47:7 47 (7):746–56. https: / / doi.org / 10.1038 / ng.3291. Kumar, Ravin, Colin Carroll, Ari Hartikainen, and Osvaldo Martin. 2019. ‘ArviZ a Unified Library for Exploratory Analysis of Bayesian Models in Python’. Journal of Open Source Software 4 (33): 1143. https: / / doi.org / 10.21105 / JOSS.01143.Lee, Seung Tae, Marcus O. Muench, Marina E. Fomin, Jianqiao Xiao, Mi Zhou, Adam De Smith, José I. Martín-Subero, et al. 2015. ‘Epigenetic Remodeling in B-Cell Acute Lymphoblastic Leukemia Occurs in Two Tracks and Employs Embryonic Stem Cell-like Signatures’. Nucleic Acids Research 43 (5): 2590–2602.https: / / doi.org / 10.1093 / NAR / GKV103. Lee, Yunsung, Sanaa Choufani, Rosanna Weksberg, Samantha L. Wilson, Victor Yuan, Amber Burt, Carmen Marsit, et al. 2019. ‘Placental Epigenetic Clocks: Estimating Gestational Age Using Placental DNA Methylation Levels’. Aging 11 (12): 4238–53.https: / / doi.org / 10.18632 / AGING.102049. Lee-Six, Henry, Nina Friesgaard Øbro, Mairi S Shepherd, Sebastian Grossmann, Kevin Dawson, Miriam Belmonte, Robert J Osborne, et al. 2018. ‘Population Dynamics of Normal Human Blood Inferred from Somatic Mutations’. Nature 561 (7724): 473–78.https: / / doi.org / 10.1038 / s41586-018-0497-0. Leval, Laurence de, Ash A. Alizadeh, P. Leif Bergsagel, Elias Campo, Andrew Davies, Ahmet Dogan, Jude Fitzgibbon, et al. 2022. ‘Genomic Profiling for Clinical Decision Making in Lymphoid Neoplasms’. Blood 140 (21): 2193–2227.https: / / doi.org / 10.1182 / BLOOD.2022015854. Levine, Morgan E., Ake T. Lu, Austin Quach, Brian H. Chen, Themistocles L. Assimes, Stefania Bandinelli, Lifang Hou, et al. 2018. ‘An Epigenetic Biomarker of Aging for Lifespan and Healthspan’. Aging 10 (4): 573–91.https: / / doi.org / 10.18632 / AGING.101414. Liang, Xiaoyu, Amy C. Justice, Kaku So-Armah, John H. Krystal, Rajita Sinha, and Ke Xu. 2020. ‘DNA Methylation Signature on Phosphatidylethanol, Not on Self-ReportedAlcohol Consumption, Predicts Hazardous Alcohol Consumption in Two DistinctP603125PC0080 Populations’. Molecular Psychiatry 2020 26:6 26 (6): 2238–53.https: / / doi.org / 10.1038 / s41380-020-0668-x. Lin, Jyun Hong, Liang Chi Chen, Shu Chi Yu, and Yao Ting Huang. 2022. ‘LongPhase: An Ultra-Fast Chromosome-Scale Phasing Algorithm for Small and Large Variants’. Bioinformatics 38 (7): 1816–22.https: / / doi.org / 10.1093 / BIOINFORMATICS / BTAC058. Lin, Qiong, Carola I. Weidner, Ivan G. Costa, Riccardo E. Marioni, Marcelo R.P. Ferreira, Ian J. Deary, and Wolfgang Wagner.2016. ‘DNA Methylation Levels at Individual Age- Associated CpG Sites Can Be Indicative for Life Expectancy’. Aging 8 (2): 394–401.https: / / doi.org / 10.18632 / AGING.100908. Ling, Shaoping, Zheng Hu, Zuyu Yang, Fang Yang, Yawei Li, Pei Lin, Ke Chen, et al. 2015. ‘Extremely High Genetic Diversity in a Single Tumor Points to Prevalence of Non- Darwinian Cell Evolution’. Proceedings of the National Academy of Sciences of the United States of America 112 (47): E6496–6505.https: / / doi.org / 10.1073 / PNAS.1519556112 / SUPPL_FILE / PNAS.1519556112.SD01.X LSX. Loyfer, Netanel, Judith Magenheim, Ayelet Peretz, Gordon Cann, Joerg Bredno, Agnes Klochendler, Ilana Fox-Fisher, et al. 2023. ‘A DNA Methylation Atlas of Normal Human Cell Types’. Nature 2023 613:7943 613 (7943): 355–64.https: / / doi.org / 10.1038 / s41586-022-05580-6. Lu, Ake T., Austin Quach, James G. Wilson, Alex P. Reiner, Abraham Aviv, Kenneth Raj, Lifang Hou, et al. 2019. ‘DNA Methylation GrimAge Strongly Predicts Lifespan and Healthspan’. Aging 11 (2): 303–27. https: / / doi.org / 10.18632 / AGING.101684.Martinez, Pierre, Diego Mallo, Thomas G. Paulson, Xiaohong Li, Carissa A. Sanchez, Brian J. Reid, Trevor A. Graham, Mary K. Kuhner, and Carlo C. Maley. 2018. ‘Evolution of Barrett’s Esophagus through Space and Time at Single-Crypt and Whole-Biopsy Levels’. Nature Communications 2018 9:1 9 (1): 1–12.https: / / doi.org / 10.1038 / s41467-017-02621-x. Mayne, Benjamin T., Shalem Y. Leemaqz, Alicia K. Smith, James Breen, Claire T. Roberts, and Tina Bianco-Miotto. 2017. ‘Accelerated Placental Aging in Early Onset Preeclampsia Pregnancies Identified by DNA Methylation’. Epigenomics 9 (3): 279–89. https: / / doi.org / 10.2217 / EPI-2016- 0103 / SUPPL_FILE / SUPPLEMENTARY_FIGURE_1.PDF. McCartney, Daniel L., Robert F. Hillary, Anna J. Stevenson, Stuart J. Ritchie, Rosie M. Walker, Qian Zhang, Stewart W. Morris, et al.2018. ‘Epigenetic Prediction of Complex Traits and Death’. Genome Biology 19 (1): 136. https: / / doi.org / 10.1186 / S13059-018-1514-1 / TABLES / 3.P603125PC0081 McEwen, Lisa M., Kieran J. O’Donnell, Megan G. McGill, Rachel D. Edgar, Meaghan J. Jones, Julia L. MacIsaac, David Tse Shen Lin, et al. 2020. ‘The PedBE Clock Accurately Estimates DNA Methylation Age in Pediatric Buccal Cells’. Proceedings of the National Academy of Sciences of the United States of America 117 (38): 23329–35.https: / / doi.org / 10.1073 / PNAS.1820843116 / SUPPL_FILE / PNAS.1820843116.SAPP.PD F. Merlo, Lauren M.F., John W. Pepper, Brian J. Reid, and Carlo C. Maley. 2006. ‘Cancer as an Evolutionary and Ecological Process’. Nature Reviews Cancer. Nature Publishing Group. https: / / doi.org / 10.1038 / nrc2013. Nadeu, Ferran, David Martin-Garcia, Guillem Clot, Ander Díaz-Navarro, Martí Duran- Ferrer, Alba Navarro, Roser Vilarrasa-Blasi, et al. 2020. ‘Genomic and Epigenomic Insights into the Origin, Pathogenesis, and Clinical Behavior of Mantle Cell Lymphoma Subtypes’. Blood 136 (12): 1419–32. https: / / doi.org / 10.1182 / BLOOD.2020005289.Nadeu, Ferran, Romina Royo, Ramon Massoni-Badosa, Heribert Playa-Albinyana, Beatriz Garcia-Torre, Martí Duran-Ferrer, Kevin J. Dawson, et al. 2022. ‘Detection of Early Seeding of Richter Transformation in Chronic Lymphocytic Leukemia’. Nature Medicine 2022 28:8 28 (8): 1662–71. https: / / doi.org / 10.1038 / s41591-022-01927-8.Nik-Zainal, Serena, Ludmil B. Alexandrov, David C. Wedge, Peter Van Loo, Christopher D. Greenman, Keiran Raine, David Jones, et al. 2012. ‘Mutational Processes Molding the Genomes of 21 Breast Cancers’. Cell 149 (5): 979–93.https: / / doi.org / 10.1016 / j.cell.2012.04.024. Nordlund, Jessica, Christofer L. Bäcklin, Per Wahlberg, Stephan Busche, Eva C. Berglund, Maija Leena Eloranta, Trond Flaegstad, et al. 2013. ‘Genome-Wide Signatures of Differential DNA Methylation in Pediatric Acute Lymphoblastic Leukemia’. Genome Biology 14 (9): 1–15. https: / / doi.org / 10.1186 / GB-2013-14-9-R105 / TABLES / 3.Nowell, Peter C. 1976. ‘The Clonal Evolution of Tumor Cell Populations’. Science 194(4260): 23–28. https: / / doi.org / 10.1126 / science.959840. Oakes, Christopher C., and Jose I. Martin-Subero.2018. ‘Insight into Origins, Mechanisms, and Utility of DNA Methylation in B-Cell Malignancies’. Blood 132 (10): 999–1006.https: / / doi.org / 10.1182 / BLOOD-2018-02-692970. Oakes, Christopher C., Marc Seifert, Yassen Assenov, Lei Gu, Martina Przekopowitz, Amy S. Ruppert, Qi Wang, et al.2016. ‘DNA Methylation Dynamics during B Cell Maturation Underlie a Continuum of Disease Phenotypes in Chronic Lymphocytic Leukemia’. Nature Genetics 48 (3): 253–64. https: / / doi.org / 10.1038 / NG.3488.Plevova, Karla, Hana Skuhrova Francova, Katerina Burckova, Yvona Brychtova, Michael Doubek, Sarka Pavlova, Jitka Malcikova, Jiri Mayer, Boris Tichy, and Sarka Pospisilova. 2014. ‘Multiple Productive Immunoglobulin Heavy Chain Gene Rearrangements inP603125PC0082 Chronic Lymphocytic Leukemia Are Mostly Derived from Independent Clones’. Haematologica 99 (2): 329–38. https: / / doi.org / 10.3324 / HAEMATOL.2013.087593.Puente, Xose S., Silvia Beà, Rafael Valdés-Mas, Neus Villamor, Jesús Gutiérrez-Abril, José I. Martín-Subero, Marta Munar, et al. 2015. ‘Non-Coding Recurrent Mutations in Chronic Lymphocytic Leukaemia’. Nature 2015 526:7574 526 (7574): 519–24.https: / / doi.org / 10.1038 / nature14666. Queirós, Ana C., Renée Beekman, Roser Vilarrasa-Blasi, Martí Duran-Ferrer, Guillem Clot, Angelika Merkel, Emanuele Raineri, et al. 2016. ‘Decoding the DNA Methylome of Mantle Cell Lymphoma in the Light of the Entire B Cell Lineage’. Cancer Cell 30 (5):806–21. https: / / doi.org / 10.1016 / J.CCELL.2016.09.014. Rambaut, Andrew, Alexei J. Drummond, Dong Xie, Guy Baele, and Marc A. Suchard. 2018. ‘Posterior Summarization in Bayesian Phylogenetics Using Tracer 1.7’. Systematic Biology 67 (5): 901–4. https: / / doi.org / 10.1093 / SYSBIO / SYY032.Raudvere, Uku, Liis Kolberg, Ivan Kuzmin, Tambet Arak, Priit Adler, Hedi Peterson, and Jaak Vilo. 2019. ‘G:Profiler: A Web Server for Functional Enrichment Analysis and Conversions of Gene Lists (2019 Update)’. Nucleic Acids Research 47 (W1): W191–98. https: / / doi.org / 10.1093 / NAR / GKZ369. Reimand, Jüri, Meelis Kull, Hedi Peterson, Jaanus Hansen, and Jaak Vilo. 2007. ‘G:Profiler- -a Web-Based Toolset for Functional Profiling of Gene Lists from Large-Scale Experiments’. Nucleic Acids Research 35 (Web Server issue).https: / / doi.org / 10.1093 / NAR / GKM226. Reinius, Lovisa E., Nathalie Acevedo, Maaike Joerink, Göran Pershagen, Sven Erik Dahlén, Dario Greco, Cilla Söderhäll, Annika Scheynius, and Juha Kere.2012. ‘Differential DNA Methylation in Purified Human Blood Cells: Implications for Cell Lineage and Studies on Disease Susceptibility’. PLOS ONE 7 (7): e41361.https: / / doi.org / 10.1371 / JOURNAL.PONE.0041361. Rice, Siobhan, and Anindita Roy. 2020. ‘MLL-Rearranged Infant Leukaemia: A “thorn in the Side” of a Remarkable Success Story’. Biochimica et Biophysica Acta. Gene Regulatory Mechanisms 1863 (8). https: / / doi.org / 10.1016 / J.BBAGRM.2020.194564.Seale, Kirsten, Steve Horvath, Andrew Teschendorff, Nir Eynon, and Sarah Voisin. 2022. ‘Making Sense of the Ageing Methylome’. Nature Reviews. Genetics 23 (10): 585–605. https: / / doi.org / 10.1038 / S41576-022-00477-6. Shireby, Gemma L., Jonathan P. Davies, Paul T. Francis, Joe Burrage, Emma M. Walker, Grant W.A. Neilson, Aisha Dahir, et al. 2020. ‘Recalibrating the Epigenetic Clock: Implications for Assessing Biological Age in the Human Cortex’. Brain 143 (12): 3763–75. https: / / doi.org / 10.1093 / BRAIN / AWAA334.P603125PC0083 Siegmund, Kimberly D., Paul Marjoram, Yen Jung Woo, Simon Tavaré, and Darryl Shibata. 2009. ‘Inferring Clonal Expansion and Cancer Stem Cell Dynamics from DNA Methylation Patterns in Colorectal Cancers’. Proceedings of the National Academy of Sciences of the United States of America 106 (12): 4828.https: / / doi.org / 10.1073 / PNAS.0810276106. Skilling, John.2004. ‘Nested Sampling’. In AIP Conference Proceedings, 735:395–405. AIP Publishing. https: / / doi.org / 10.1063 / 1.1835238. Skilling, John. 2006. ‘Nested Sampling for General Bayesian Computation’. Bayesian Analysis 1 (4): 833–60. https: / / doi.org / 10.1214 / 06-BA127.Sottoriva, Andrea, Haeyoun Kang, Zhicheng Ma, Trevor A. Graham, Matthew P. Salomon, Junsong Zhao, Paul Marjoram, et al. 2015. ‘A Big Bang Model of Human Colorectal Tumor Growth’. Nature Genetics 2015 47:3 47 (3): 209–16.https: / / doi.org / 10.1038 / ng.3214. Speagle, Joshua S. 2019. ‘Dynesty: A Dynamic Nested Sampling Package for Estimating Bayesian Posteriors and Evidences’. Monthly Notices of the Royal Astronomical Society 493 (3): 3132–58. https: / / doi.org / 10.1093 / mnras / staa278. Suchard, Marc A, Philippe Lemey, Guy Baele, Daniel L Ayres, Alexei J Drummond, and Andrew Rambaut. 2018. ‘Bayesian Phylogenetic and Phylodynamic Data Integration Using BEAST 1.10’. Virus Evolution 4 (1). https: / / doi.org / 10.1093 / VE / VEY016.Tate, John G., Sally Bamford, Harry C. Jubb, Zbyslaw Sondka, David M. Beare, Nidhi Bindal, Harry Boutselakis, et al. 2019. ‘COSMIC: The Catalogue Of Somatic Mutations In Cancer’. Nucleic Acids Research 47 (D1): D941–47.https: / / doi.org / 10.1093 / NAR / GKY1015. Teschendorff, Andrew E. 2020. ‘A Comparison of Epigenetic Mitotic-like Clocks for Cancer Risk Prediction’. Genome Medicine 12 (1): 1–17. https: / / doi.org / 10.1186 / S13073-020-00752-3 / FIGURES / 4. Turajlic, Samra, Andrea Sottoriva, Trevor Graham, and Charles Swanton.2019. ‘Resolving Genetic Heterogeneity in Cancer’. Nature Reviews Genetics 2019 20:7 20 (7): 404–16. https: / / doi.org / 10.1038 / s41576-019-0114-6. Uhlén, Mathias, Linn Fagerberg, Bjö M. Hallström, Cecilia Lindskog, Per Oksvold, Adil Mardinoglu, Åsa Sivertsson, et al. 2015. ‘Tissue-Based Map of the Human Proteome’. Science 347 (6220).https: / / doi.org / 10.1126 / SCIENCE.1260419 / SUPPL_FILE / 1260419_UHLEN.SM.PDF. Vehtari, Aki, Andrew Gelman, and Jonah Gabry.2015. ‘Practical Bayesian Model Evaluation Using Leave-One-out Cross-Validation and WAIC’. Statistics and Computing 27 (5):1413–32. https: / / doi.org / 10.1007 / s11222-016-9696-4.P603125PC0084 Vidal-Bralo, Laura, Yolanda Lopez-Golan, and Antonio Gonzalez. 2016. ‘Simplified Assay for Epigenetic Age Estimation in Whole Blood of Adults’. Frontiers in Genetics 7 (JUL):209192. https: / / doi.org / 10.3389 / FGENE.2016.00126 / BIBTEX. Weidner, Carola I., Qiong Lin, Carmen M. Koch, Lewin Eisele, Fabian Beier, Patrick Ziegler, Dirk O. Bauerschlag, et al. 2014. ‘Aging of Blood Can Be Tracked by DNA Methylation Changes at Just Three CpG Sites’. Genome Biology 15 (2): 1–12.https: / / doi.org / 10.1186 / GB-2014-15-2-R24 / COMMENTS. Williams, Marc J., Benjamin Werner, Chris P. Barnes, Trevor A. Graham, and Andrea Sottoriva. 2016. ‘Identification of Neutral Tumor Evolution across Cancer Types’. Nature Genetics 48 (3): 238–44. https: / / doi.org / 10.1038 / ng.3489.Williams, Marc J., Benjamin Werner, Timon Heide, Christina Curtis, Chris P. Barnes, Andrea Sottoriva, and Trevor A. Graham. 2018. ‘Quantification of Subclonal Selection in Cancer from Bulk Sequencing Data’. Nature Genetics 50 (6): 895–903.https: / / doi.org / 10.1038 / s41588-018-0128-6. Williams, Nicholas, Joe Lee, Emily Mitchell, Luiza Moore, E. Joanna Baxter, James Hewinson, Kevin J. Dawson, et al.2022. ‘Life Histories of Myeloproliferative Neoplasms Inferred from Phylogenies’. Nature 2022 602:7895 602 (7895): 162–68.https: / / doi.org / 10.1038 / s41586-021-04312-6. Yang, Zhen, Andrew Wong, Diana Kuh, Dirk S. Paul, Vardhman K. Rakyan, R. David Leslie, Shijie C. Zheng, Martin Widschwendter, Stephan Beck, and Andrew E. Teschendorff. 2016. ‘Correlation of an Epigenetic Mitotic Clock with Cancer Risk’. Genome Biology 17 (1): 1–18. https: / / doi.org / 10.1186 / S13059-016-1064-3 / FIGURES / 6. Yao, Yuling, Aki Vehtari, Daniel Simpson, and Andrew Gelman. 2017. ‘Using Stacking to Average Bayesian Predictive Distributions’. Bayesian Analysis 13 (3).https: / / doi.org / 10.1214 / 17-BA1091. Yatabe, Yasushi, Simon Tavaré, and Darryl Shibata. 2001. ‘Investigating Stem Cells in Human Colon by Using Methylation Patterns’. Proceedings of the National Academy of Sciences 98 (19): 10839–44. https: / / doi.org / 10.1073 / pnas.191225998.Youn, Ahrim, and Shuang Wang. 2018. ‘The MiAge Calculator: A DNA Methylation-Based Mitotic Age Calculator of Human Tissue Types’. Epigenetics 13 (2): 192–206.https: / / doi.org / 10.1080 / 15592294.2017.1389361. Yu, Guangchuang, David K. Smith, Huachen Zhu, Yi Guan, and Tommy Tsan Yuk Lam. 2017. ‘Ggtree: An r Package for Visualization and Annotation of Phylogenetic Trees with Their Covariates and Other Associated Data’. Methods in Ecology and Evolution 8(1): 28–36. https: / / doi.org / 10.1111 / 2041-210X.12628. Zenz, Thorsten, Barbara Eichhorst, Raymonde Busch, Tina Denzel, Sonja Häbe, Dirk Winkler, Andreas Bühler, et al. 2010. ‘TP53 Mutation and Survival in ChronicP603125PC0085 Lymphocytic Leukemia’. Journal of Clinical Oncology 28 (29): 4473–79.https: / / doi.org / 10.1200 / JCO.2009.27.8762 / SUPPL_FILE / SUPPLEMENTAL_TABLE.XLS Zhang, Qian, Costanza L. Vallerga, Rosie M. Walker, Tian Lin, Anjali K. Henders, Grant W. Montgomery, Ji He, et al. 2019. ‘Improved Precision of Epigenetic Clock Estimates across Tissues and Its Implication for Biological Ageing’. Genome Medicine 11 (1): 1–11. https: / / doi.org / 10.1186 / S13073-019-0667-1 / FIGURES / 4. Zheng, Zhenxian, Shumin Li, Junhao Su, Amy Wing Sze Leung, Tak Wah Lam, and Ruibang Luo. 2022. ‘Symphonizing Pileup and Full-Alignment for Deep Learning-Based Long- Read Variant Calling’. Nature Computational Science 2022 2:12 2 (12): 797–803.https: / / doi.org / 10.1038 / s43588-022-00387-x. Zhou, Wanding, Huy Q. Dinh, Zachary Ramjan, Daniel J. Weisenberger, Charles M. Nicolet, Hui Shen, Peter W. Laird, and Benjamin P. Berman. 2018. ‘DNA Methylation Loss in Late-Replicating Domains Is Linked to Mitotic Cell Division’. Nature Genetics 201850:4 50 (4): 591–602. https: / / doi.org / 10.1038 / s41588-018-0073-4.
Claims
P603125PC0086 Claims1. A method of predicting the risk of progression of a cancer of a subject, wherein themethod comprises: obtaining DNA methylation data of fluctuating CpG loci of DNA derived from thecancer; using the DNA methylation data to generate a methylation distribution from whicha historical growth rate of the cancer can be determined, andpredicting the risk of progression of the cancer based on the historical growth rateof the cancer.
2. A method according to claim 1, wherein the cancer is a hematologic malignancy.
3. A method according to any preceding claim, wherein the cancer is a chroniclymphocytic leukemia.
4. A method according to any preceding claim, wherein the methylation distributionis generated by fitting a stochastic forward model to the methylation data using an inference method.
5. A method according to any preceding claim, wherein the risk of progression of thecancer based on the historical growth rate of the cancer is predicted using a survivalanalysis method, optionally wherein the survival analysis model is a Cox regression model.
6. A method according to any preceding claim, wherein the method further comprisestreating the subject with a therapeutic drug, a palliative drug, radiotherapy, surgical resection of the cancer, or a combination thereof.
7. A method according to any preceding claim, wherein the age of the subject whenthe most recent common ancestor (MRCA) of the sampling population emerged can bedetermined from the methylation distribution.
8. A method according to any preceding claim, wherein the inference method is aBayesian inference method.
9. A method according to claim 8, wherein the inference method is a pseudo-marginalBayesian inference method.P603125PC008710. A method according to any preceding claim, wherein the inference methodestimates a log-likelihood of the data.
11. A method according to any preceding claim, wherein the inference methodcomprises a penalisation function to penalise sizes of cancer that are non-physical.
12. A method according to any preceding claim, wherein the inference methodcomprises a correction function to account for finite sampling bias.
13. A method according to any preceding claim, wherein the inference methodcomprises a noise function to account for noise.
14. A method according to any preceding claim, wherein the step of obtaining DNAmethylation data comprises performing a DNA methylation microarray.
15. A method according to any preceding claim, wherein the method further comprisesa step of normalising the DNA methylation data.
16. A method according to claim 15, wherein the DNA methylation data is normalisedby single-sample Noob normalisation.
17. A method according to any preceding claim, wherein the method further comprisesa step of identifying fluctuating CpG loci of DNA derived from the cancer.
18. A method according to claim 17, wherein the CpG loci are identified as fluctuatingCpG loci by selecting CpG loci that: (i) are heterogeneous across different subjects withthe same disease, (ii) have an intermediate average methylation across different subjectswith the same disease, and (iii) are unlikely to be associated with specific cell-types, cancer type / subtype or epigenetic remodelling.
19. A method according to any preceding claim, wherein the fluctuating CpG locicomprise at least 70% of the group of consisting of: cg01768082, cg16526047, cg10126324, cg06036677, cg17132079, cg13354934, cg09368716, cg12208770, cg26863172, cg23288755, cg06770735, cg20312132, cg26220594, cg05129477, cg07981328, cg24792360, cg16585234, cg01811796, cg08134856, cg01815720, cg21151355, cg18931815, cg14076729, cg14523898, cg02851558, cg09847153,P603125PC0088 cg12051614, cg15986668, cg18174678, cg08706567, cg16342298, cg17428085, cg06320401, cg17742781, cg15546227, cg25189904, cg17083209, cg24724587, cg10593400, cg17108141, cg23627354, cg26523099, cg22374525, cg09815769, cg26050734, cg01963297, cg09037777, cg19321979, cg04843111, cg03390245, cg04415780, cg10963061, cg05843457, cg24509168, cg13229857, cg06775570, cg16029875, cg18438591, cg26354017, cg14159672, cg18086868, cg13986762, cg20704450, cg01730064, cg26536593, cg12726839, cg24598973, cg18224942, cg24557048, cg01894875, cg23656015, cg08808724, cg13885623, cg22698629, cg24876897, cg24730688, cg22077361, cg25593625, cg08133631, cg19633205, cg04673462, cg17434008, cg12410921, cg13393721, cg22345063, cg00405069, cg05821046, cg26013992, cg00597687, cg24049468, cg05694621, cg04896168, cg18143535, cg26537280, cg00729875, cg02605776, cg22697853, cg08477332, cg22941637, cg17935233, cg24058386, cg23640929, cg24329783, cg07580762, cg21691116, cg10092377, cg03948781, cg17178900, cg14893161, cg23615892, cg06939851, cg24860534, cg02012159, cg14099468, cg21397540, cg20433858, cg20510474, cg03429644, cg01621716, cg19584871, cg01931792, cg04082016, cg03350299, cg25071744, cg08965143, cg02716646, cg22488717, cg00956987, cg24209528, cg10588622, cg14223654, cg14758072, cg24686644, cg14339650, cg24220031, cg00058515, cg24222817, cg22111694, cg26874229, cg25924274, cg11942181, cg24980657, cg27664689, cg26817877, cg21361856, cg21627775, cg01217071, cg10634136, cg22308949, cg14157435, cg27369013, cg19075252, cg01898867, cg26302230, cg00680696, cg00899350, cg03727500, cg15371801, cg27661104, cg13716566, cg14853772, cg23531340, cg04131969, cg01583753, cg20097593, cg05414442, cg03813688, cg08201735, cg06383022, cg05723825, cg22022881, cg25079915, cg00280895, cg25223285, cg05353292, cg08719380, cg25487775, cg15742848, cg15161959, cg26853787, cg24715680, cg18887096, cg03817727, cg24330386, cg05868531, cg11559198, cg21054919, cg10883069, cg02555923, cg06234584, cg27364650, cg17658717, cg03855994, cg20540428, cg20232291, cg11874323, cg04865442, cg20988073, cg18026197, cg15709065, cg09373148, cg18005180, cg00472710, cg00010946, cg15180789, cg21030598, cg17036624, cg04207166, cg17546649, cg16930811, cg08072716, cg22158068, cg17057719, cg24576535, cg14218447, cg13511885, cg06726820, cg01324343, cg24960291, cg14463292, cg24793722, cg08222618, cg23855319, cg02754929, cg09573795, cg12948920, cg11266682, cg15175162, cg03238216, cg03313172, cg00236919, cg03463411, cg11388320, cg01883540, cg11991617, cg08125755, cg17069313, cg14016257, cg12494515, cg25932599, cg10432947, cg08644045, cg23139982, cg17814781, cg02136252, cg01096266, cg03315557, cg06634862,P603125PC0089 cg02170478, cg06483432, cg13210763, cg08572214, cg27242132, cg09511421, cg25654705, cg25886683, cg22025206, cg04156016, cg12518844, cg26690304, cg25102726, cg23522872, cg16079430, cg05530317, cg01784614, cg05465935, cg14789272, cg19590578, cg12980795, cg13003311, cg18665594, cg18586823, cg16988194, cg25218470, cg15574301, cg08470180, cg06536614, cg25340688, cg26896946, cg15548198, cg19727817, cg02147194, cg16391973, cg07751793, cg22675956, cg23521468, cg06060754, cg23248424, cg00619978, cg16035780, cg09075844, cg06836380, cg07401067, cg26875852, cg18394854, cg27586797, cg01809941, cg09560636, cg06837426, cg20451680, cg08336234, cg17503389, cg14996985, cg17253517, cg02250764, cg11620475, cg11809091, cg17763019, cg07017875, cg17303540, cg09546802, cg11609571, cg12352302, cg16006841, cg07763376, cg17974424, cg16935061, cg00101728, cg01474544, cg17204113, cg00942918, cg25063707, cg02631126, cg03350138, cg00219626, cg04522432, cg11406274, cg03228974, cg08710841, cg17931227, cg26668675, cg26818629, cg27547543, cg23906067, cg00872984, cg07885132, cg04749507, cg20980321, cg09459982, cg22297966, cg00116315, cg02377690, cg13602274, cg09140531, cg11611580, cg15234946, cg08313539, cg21855021, cg07571531, cg16699385, cg06864789, cg18136963, cg16073143, cg18322025, cg01238672, cg15279541, cg14115346, cg04767753, cg08660295, cg19579217, cg18278486, cg04612667, cg12623302, cg26968378, cg12584458, cg07734926, cg09012858, cg14132236, cg13143120, cg15070894, cg08497487, cg15946590, cg09876447, cg23252259, cg14036627, cg11811828, cg01869058, cg10223234, cg04142955, cg13069441, cg12672189, cg18816397, cg14590011, cg06386482, cg25716013, cg14989243, cg12602851, cg05237015, cg05567646, cg02270332, cg01883195, cg25399239, cg06388363, cg26850117, cg07257824, cg05779406, cg14662585, cg26012061, cg06955417, cg22750906, cg22330875, cg13710556, cg02207386, cg06166490, cg05929675, cg21004899, cg23068772, cg03552937, cg25783361, cg13578160, cg01830073, cg16453056, cg23192873, cg18209323, cg05667818, cg18577231, cg03027739, cg12560447, cg10157098, cg06761584, cg22047901, cg09916174, cg20104432, cg15935721, cg11945929, cg04816699, cg03454541, cg17264513, cg15575538, cg01499522, cg15881597, cg00506049, cg15285764, cg19435720, cg02984009, cg09669835, cg27203184, cg21406967, cg20665157, cg02193461, cg23117316, cg23305129, cg21724654, cg26967842, cg04671734, cg08242633, cg20463151, cg26041569, cg13346869, cg21505509, cg08247852, cg23245289, cg11491002, cg19632760, cg03003919, cg27053337, cg20366110, cg04724477, cg01385052, cg06070970, cg05807444, cg20845050, cg00643296, cg09945801, cg19806221, cg11583762, cg10327428, cg02980621, cg18570392, cg21244580,P603125PC0090 cg24636969, cg10039928, cg09981361, cg04966191, cg03894796, cg19682134, cg01319235, cg14182974, cg21891967, cg13891189, cg00220243, cg03191962, cg01126560, cg13573611, cg04824555, cg03576809, cg07881032, cg21177396, cg19519624, cg13518232, cg00219169, cg09955645, cg20503657, cg06188367, cg12845268, cg06850159, cg21012874, cg02971322, cg22704845, cg24143137, cg00492979, cg07262395, cg12765123, cg24696332, cg00153041, cg16977735, cg11444332, cg00142642, cg12428160, cg03444077, cg04004590, cg24160569, cg12450708, cg00169964, cg15233961, cg14251734, cg18103548, cg06620993, cg10475928, cg23331156, cg16351002, cg04309849, cg13221347, cg25258098, cg01794156, cg06974175, cg11976911, cg07238401, cg05061471, cg07934552, cg04862679, cg01421380, cg03101664, cg15887846, cg20051530, cg23188684, cg00864171, cg07716408, cg10283695, cg25132257, cg08743050, cg14201467, cg05566455, cg24953078, cg17840408, cg12110980, cg10026995, cg01694451, cg19415743, cg12904004, cg23065768, cg10570158, cg18052274, cg04106006, cg06705930, cg25948180, cg00233420, cg03606954, cg06962549, cg26418434, cg00089550, cg11132272, cg14933468, cg07539927, cg10130811, cg24851651, cg03349769, cg08355456, cg21570702, cg05353869, cg19624354, cg23149852, cg03222066, cg20118840, cg10961484, cg10262710, cg01094687, cg04468568, cg15235707, cg24829292, cg13147090, cg12675583, cg09328283, cg09601159, cg17570948, cg22566355, cg06868955, cg19893439, cg08409225, cg10306450, cg00700412, cg08784247, cg23390920, cg24460369, cg06840837, cg25752163, cg23469448, cg03495030, cg03554925, cg17566325, cg15878909, cg03353445, cg10809165, cg16267202, cg25134647, cg01844176, cg17292337, cg10317314, cg14687298, cg04348558, cg03320170, cg20402783, cg08455772, cg17295489, cg07910460, cg21961771, cg10651121, cg15996769, cg16624888, cg20092036, cg11367159, cg05134054, cg08800396, cg05361406, cg20050828, cg26750617, cg23591611, cg07610777, cg20334079, cg12206359, cg09084477, cg13868001, cg12436019, cg03898952, cg22731513, cg16032841, cg06897282, cg25658983, cg03593908, cg07974719, cg19759502, cg07937631, cg00750481, cg21805940, cg25444002, cg09460553, cg22721998, cg08183074, cg18624544, cg14184890, cg11071193, cg13097993, cg23511824, cg18771300, cg06892679, cg07922154, cg13099429, cg18313317, cg08072202, cg18333511, cg04329125, cg10775991, cg23099211, cg01214012, cg05225684, cg09989251, cg26421140, cg01115857, cg24036523, cg04716394, cg18220816, cg05150608, cg10095084, cg22454744, cg25440893, cg03945400, cg17221584, cg23634401, cg20481312, cg26710563, cg24382249, cg26218577, cg16146501, cg08637403, cg20246113, cg18105749, cg25985504, cg07615383, cg00098799, cg13812291, cg03380198, cg17074431,P603125PC0091 cg03822622, cg12756504, cg12163955, cg06891458, cg17082938, cg08133919, cg12036633, cg25481157, cg00474218, cg25924688, cg24795825, cg19708986, cg02054724, cg07695058, cg24667603, cg10135717, cg16548362, cg16367168, cg26538942, cg00639984, cg00671386, cg06334495, cg09242293, cg09798061, cg20035459, cg02873098, cg02859934, cg06138239, cg16014085, cg24642468, cg17484472, cg27268835, cg26580413, cg02574464, cg27014354, cg00618711, cg07307142, cg03072681, cg04810063, cg00511475, cg08521396, cg00591660, cg03264209, cg01680674, cg08230268, cg05787209, cg05725404, cg02451853, cg09758204, cg08638180, cg10433128, cg06655190, cg06521653, cg04245248, cg06471402, cg13768055, cg12770741, cg02866639, cg00668150, cg11684897, cg10553748, cg24768135, cg21358336, cg07391831, cg16893174, cg10092265, cg06418907, cg12156887, cg18048405, cg21197958, cg17033854, cg10747531, cg20291162, cg25067162, cg08471713, cg13947929, cg02219949, cg11732492, cg12165772, cg01992590, cg09889948, cg13655803, cg02237906, cg12407791, cg12700283, cg04498014, cg10759591, cg20067780, cg20932150, cg23246911, cg22635541, cg08103988, cg08750459, cg15660077, cg16696476, cg15041662, cg24156181, cg18091275, cg05451094, cg19865134, cg19789473, cg09760387, cg17757575, cg02633371, cg02481789, cg26782881, cg01405107, cg09983216, cg06602723, cg10237252, cg14192542, cg09601923, cg14952312, cg12443604, cg08084655, cg06697425, cg05865011, cg07010633, cg11043990, cg00423030, cg24207009, cg13694343, cg17563271, cg07016075, cg03889540, cg11520030, cg01908551, cg21266975, cg10092257, cg19453250, cg00057722, cg05458220, cg13287964, cg21237861, cg06511678, cg18086635, cg05872758, cg24864576, cg02978168, cg14887613, cg15002904, cg05308639, cg08350509, cg15255455, cg13617776, cg27555894, cg11056604, cg20587168, cg04351156, cg10750464, cg11738485, cg01632288, cg26914865, cg16701007, cg15428479, cg19492423, cg17969123, cg14265502, cg03657045, cg26643967, cg12216477, cg17069873, cg18458353, cg22376864, cg09438457, cg15173319, cg14061069, cg18861015, cg23448153, cg21198638, cg12068280, cg14780446, cg16194588, cg03630273, cg05696092, cg21723184, cg26651782, cg10635122, cg24794228, cg03153360, cg15636087, cg07786254, cg19907366, cg22603450, cg10503234, cg19870717, cg14111332, cg23671699, cg15043935, cg24084706, cg20467412, cg03489492, cg17490089, cg01272599, cg25506915, cg06189175, cg15050529, cg23899408, cg10384482, cg15010352, cg00706441, cg00003818, cg14021871, cg05516285, cg22548353, cg06942742, cg09007470, cg26920327, cg14881601, cg25786696, cg27331471, cg04382643, cg18741372, cg11237948, cg17716663, cg11773468, cg19851487, cg17071417, cg17729891, cg12289045, cg19014858, cg02592403,P603125PC0092 cg06878361, cg22331349, cg15721243, cg15061569, cg04358463, cg18282375, cg05386230, cg00216901, cg17083925, cg13928427, cg04307107, cg11598935, cg12751644, cg02544614, cg02519263, cg01735357, cg08389814, cg20964064, cg22463915, cg03173827, cg00237512, cg09990613, cg09123507, cg08424491, cg20102019, cg11133732, cg05389922, cg03539340, cg09607548, cg05139523, cg19348676, cg02174884, cg23303369, cg10797197, cg17029694, cg24818238, cg25421566, cg01481690, cg08431893, cg16501323, cg17356252, cg11113589, cg09417038, cg09374293, cg24435209, cg15498409, cg07331806, cg21401457, cg11141652, cg16210088, cg20548231, cg01704580, cg00276389, cg18982625, cg00610021, cg06961233, cg25163076, cg07459594, cg06942979, cg26730347, cg09696535, cg08178168, cg03688400, and cg06781532.
20. A computer system for implementing a prediction method for predicting the riskof progression of a cancer of a subject, the system comprising: one or more processor units; and a computer readable storage medium storing one or more computer programs such that when performed by the one or more processor units cause the computer system to operate in accordance with the method of any of claims 1-19.
21. A method of identifying fluctuating CpG (fCpG) loci, the method comprising:i. receiving a dataset of CpG methylation values from DNA derived from cancercells obtained from a large plurality of cancer patient subjects with different cancer presentations; ii. calculating a respective Laplacian Score feature selection metric for the CpGmethylation values in the dataset; andiii. identifying fCpG loci with a Laplacian Score greater than a predeterminedthreshold score.
22. A method according to claim 21, wherein the identified fCpG loci are then used asthe fCpG loci in a method of any of the preceding claims.
23. A method according to claim 21 or claim 22, wherein the calculating step furthercomprises performing a principal component analysis (PCA) on the CpG methylationvalues, wherein at least the first 2 principal components reveal disease specificclustering in methylation.P603125PC009324. A method according to claim 23, wherein the CpG methylation values with themost informative features for clustering and thereby have the lowest Laplacian Scoresshow disease-specific methylation, whereas the CpG methylation values that are leastinformative for clustering and thereby have the highest Laplacian Scores have either lowvariability or high variability that are well-dispersed across the PCA.
25. A method according to any of claims 21 to 24, wherein the predeterminedthreshold score is at least 10% and more preferably at least 5% of the CpG methylationvalues with the highest Laplacian Scores.
26. A method according to any of claims 21 to 25, and further comprising selectingthose fCpG loci that are heterogeneous across different cancer patient subjects with the same disease.
27. A method according to claim 26, wherein the selecting comprises accepting fCpGloci with average intra-disease standard deviation of methylation above 0.15 acrosscancer types.
28. A method according to any of claims 21 to 27, and further comprising selectingthose fCpG loci that are equally likely to be methylated or unmethylated.
29. A method according to claim 28, wherein the selecting comprises averagingacross all cases in a cancer type and accepting only fCpG sites with average methylation of approximately 0.5 in the cancer type.
30. A method according to any claims 1 to 19 or 21 to 29, wherein the method iscomputer implemented.
31. A computer system for identifying fluctuating CpG (fCpG) loci, the systemcomprising: one or more processor units; and a computer readable storage medium storing one or more computer programs such that when performed by the one or more processor units cause the computer system to operate in accordance with the method of any of claims 21- 29.
Citation Information
Patent Citations
Systems and methods for predicting hematological conditions using methylation data
WO2023172772A1