Integration of evolutionary, molecular and clinical data for prognostic modelling of clinical outcomes in neoplastic diseases
By recalibrating CCF to account for CNAs and integrating evolutionary and clinical covariates, the method enhances risk stratification in neoplastic diseases, improving predictive models and patient classification accuracy.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- UNIV DEGLI STUDI DI MILANO BICOCCA
- Filing Date
- 2025-10-29
- Publication Date
- 2026-06-04
AI Technical Summary
Existing risk stratification methods for neoplastic diseases, such as Myelodysplastic Syndromes (MDS), fail to integrate evolutionary and molecular co-occurrence features with clinical covariates effectively, leading to inadequate prognostic and predictive models due to high correlation and collinearity, and neglect the impact of copy number alterations (CNA) on Cancer Cell Fraction (CCF) estimation.
A computer-implemented method that recalibrates CCF computation to account for CNA, integrates evolutionary routes and clinical covariates using survival analysis, and applies an evolutionary model to identify statistically recurrent genomic and clinical covariates, enhancing predictive models with improved Harrell’s C-index and reclassifying patients into more accurate risk categories.
The method improves patient risk stratification by constructing a unified predictive model with higher discrimination performance, refining patient classification, and capturing the prognostic impact of evolutionary trajectories beyond static molecular or clinical features, as demonstrated by increased Harrell’s C-index and reclassification of over 40% of patients.
Smart Images

Figure IB2025061016_04062026_PF_FP_ABST
Abstract
Description
[0001] SIB - 1 - BW1367R
[0002] TITLE
[0003] Integration of Evolutionary, Molecular and Clinical Data for Prognostic Modelling of Clinical Outcomes in Neoplastic Diseases
[0004] DESCRIPTION
[0005] Field of the invention
[0006] The present invention provides a computer implemented method for the prediction of clinical outcomes in patients with cancer or pre-neoplastic conditions through the integration of genomic evolutionary, genomic and clinical. More specifically, the invention provides systems and algorithms that generate prognostic and predictive models based on the combined analysis of molecular features, inferred evolutionary routes, and clinical parameters, enabling patient risk stratification and individualized outcome estimation.
[0007] Background of the invention
[0008] Risk stratification involves categorizing patients based on their likelihood of disease progression and treatment response, allowing for tailored therapeutic approaches. This method utilizes advanced biomarkers and clinical data to identify high-risk individuals, enabling early intervention and optimized management. The advantages of this approach include improved patient outcomes, reduced healthcare costs, and enhanced resource allocation within oncology practices. By accurately determining risk profiles, healthcare providers can prioritize interventions, monitor disease progression more effectively, and ultimately improve survival rates.
[0009] Risk stratification and outcome prediction can be achieved through the development of predictive models. Several models have been designed to estimate clinical outcomes such as Overall Survival (OS) or Disease-Free Survival (DFS) based solely on clinical covariates (Greenberg et al., Blood, 2012; D’Agostino R. B. Sr. et al., Circulation, 2008; Goff D. C. Jr. et al., Circulation, 2014; Gage B. F. et al., JAMA, 2001; Lip G. Y. H. et al., Chest, 2010). More recent approaches have integrated clinical and molecular covariates to enhance clinical outcome prediction (Bernard et al., NEJM Evidence, 2023; D’Agostino et al., JCO, 2022; Sparano J. A. et al., Journal of Clinical Oncology, 2021; Spratt D. E. et al., Journal of Clinical Oncology, 2018; Wishart G. C. et al., Breast Cancer Research, 2010).
[0010] Molecular evolution, a discipline at the interface between biology and bioinformatics, investigates the temporal progression of tumorigenesis. It enables the identification of evolutionary routes and gene co-occurrences, where the presence of an early (“parent”) alteration significantly increases the probability of subsequent (“child”) mutations along SIB -2 - BW1367R
[0011] the evolutionary trajectory. These evolutionary routes therefore represent directed cooccurrences of genomic events characterized by temporal order and selective advantage. Some recent studies have examined the relationship between evolutionary routes and clinical outcomes such as overall survival (Civettini et al., AACR, 2024; Caravagna et al., Nature Methods, 2018; Fontana et al., Nature Communications, 2023; Bernard et al., Nature Communications, 2021). Nevertheless, to date, no methodology has been proposed that merges evolutionary and co-occurrence features together with molecular and clinical covariates into a single, statistically robust model for clinical outcome prediction. Simple additive or concatenative approaches are inadequate, as the high degree of correlation and collinearity between molecular variables and evolutionbased features prevents the construction of stable and generalizable prognostic and predictive models. Moreover, the dominant weight of clinical variables tends to obscure the contribution of evolutionary co-occurrence patterns in predicting patient outcomes. Moreover, although evolutionary models have been applied to intertemporal trajectories of mutation acquisition, their predictive integration within clinical frameworks remains limited. A major challenge arises from the intrinsic heterogeneity of genomic data, particularly in disorders such as MDS where cytogenetic abnormalities are frequent. Existing evolutionary models generally estimate the Cancer Cell Fraction (CCF) of each mutation directly from the Variant Allele Frequency (VAF) under the assumption of diploid ploidy. However, this assumption neglects the effect of copy number alterations (CNA), which can substantially distort the estimated CCF and consequently the inferred temporal order of mutations.
[0012] To achieve clinical-grade precision and allow evolutionary variables to compete in predictive strength with established clinical features, such as cytogenetic risk, it is necessary to adapt CCF computation to account for CNA-related allelic imbalance. By recalibrating the CCF according to locus-specific copy number states, the temporal inference of mutational events becomes more accurate and biologically consistent. Furthermore, while evolutionary algorithms can identify multiple potential routes, not all inferred relationships are equally stable across patient subsets. Therefore, quantitative measures of model stability are required.
[0013] The following represents a concrete example in which, within a pre-cancerous hematologic disorder, a predictive model for clinical outcomes, specifically Leukaemia-Free Survival (LFS), was first established using clinical covariates alone (IPSS-R), and subsequently refined through the integration of clinical and molecular covariates (IPSS- SIB - 3 - BW1367R
[0014] M). However, a unified framework combining clinical, molecular, and evolutionary or cooccurrence variables for outcome prediction had never been implemented.
[0015] The application of risk stratification to patient suffering from disease which can potentially evolve in cancerous one is of particular interest. Myelodysplastic Syndromes (MDS) are hematopoietic stem cell disorders characterized by ineffective haematopoiesis, dysplasia, and an increased risk of progression to acute myeloid leukaemia (AML). The evolution of MDS is complex and it follows a multi-step process, as described by M. Cazzola (NEJM, 2020, 383(14), 1358). Initially, a primary driver mutation occurs in a hematopoietic stem cell. This is followed by a phase of Clonal Haematopoiesis of Indeterminate Potential (CHIP), during which the mutant clone spreads. The duration of this phase can vary depending on the specific somatic mutation involved. The third phase is characterized by clonal dominance and the occurrence of additional mutations.
[0016] One of the tools available for risk stratification of MDS patients is The International Prognostic Scoring System-Revised (IPSS-R), developed by the International Working Group for Prognosis in MDS (IWG-PM), which relies on hematologic and cytogenetic features but does not consider gene mutations. Said tool has been refined into the IPSS-Molecular (IPSS-M), which incorporates mutational data from 31 MDS-related genes, thereby improving risk classification. The IPSS-M relies on clinical and cytogenetic data, as well as mutations identified through Next-Generation Sequencing (NGS). However, this method does not consider the potential temporal and evolutionary relationships between genomic alterations; furthermore, IPSS-M requires expert-driven categorizations and manual adjustments for re-clustering. As previously demonstrated, the application of the ASCETIC framework for the identification of cancer evolutionary signature (Nat. Commun. 2023 Sep 25;14(1):5982. doi: 10.1038 / s41467-023-41670-3) to MDS patient data allows one to delineate evolutionary trajectories and stratify patients into distinct clusters for Leukaemia- Free Survival (LFS) and Overall Survival (OS) (I. Civettini et al., 'Characterization of Molecular Evolution in Myelodysplastic Neoplasms,' AACR Annual Meeting 2024; Cancer Res 2024; 84(6_Suppl): Abstract nr 6208). However, when ASCETIC clusters are integrated with the IPSS-M score (analyzed as a continuous variable) using a dataset of 2,209 MDS patients from the 2022 IPSS-M publication (Bernard E et al., Molecular International Prognostic Scoring System for Myelodysplastic Syndromes, NEJM Evid. 2022; 1(7): EVIDoa2200008. doi: 10.1056 / EVIDoa2200008. PMID: 38319256; data available at https: / / github.com / papaemmelab / IPSSMstudy and https: / / www.cbioportal.org / ), no SIB -4 - BW1367R
[0017] additional stratification benefit is observed. The Harrell's C-index remains constant at 0.7501 for both the IPSS-M alone and the combined IPSS-M + ASCETIC model (Unpublished Data), indicating that the direct addition of ASCETIC clusters to IPSS-M does not enhance prognostic accuracy. Further modelling attempts, incorporating IPSS-M clinical and cytogenetic features (Age, Bone Marrow Blasts as a continuous variable, Cytogenetic Risk) with ASCETIC clusters, resulted in a C-index of 0.7261, again showing no improvement over IPSS-M alone. Consequently, such findings underscore that a straightforward additive approach, merging co-occurrence and / or evolutionary features with clinical or cytogenetic data, fails to refine patient outcome predictions.
[0018] As the accuracy of risk stratification tools impacts on the development of proper and effective personalized treatment protocols, the provision of an improved method for patient risk stratification is a felt needed.
[0019] Summary of the invention
[0020] In a first aspect, the present invention provides a computer implemented method for the calculation of hazard ratio of signature genomic covariates or signature genomic and clinical covariates or signature genomic, evolutionary routes and signature clinical covariates associated with a prognosis or outcome in a cancer patient comprising the following steps:
[0021] 1) providing a dataset of genomic covariates and prognosis data or a dataset of genomic and clinical covariates and prognosis data of a statistically significant number of cancer patients or individuals with somatic mutations identified through gene sequencing, irrespective of a formal cancer diagnosis.
[0022] 2) applying an evolutionary model to the dataset of step 1) for identifying statistically recurrent evolutionary genomic covariates according to said evolutionary method;
[0023] 3) processing the identified statistically recurrent evolutionary genomic covariates, for example single gene mutations, co-occurrence of mutations or evolutionary routes, of step 2) or recurrent evolutionary genomic covariates of step 2) and clinical covariates using a survival analysis model for selecting signature genomic covariates or signature genomic and clinical covariates related to a specific prognosis;
[0024] 4) processing the signature genomic covariates selected by step 3) or signature genomic and clinical covariates selected by step 3) by a survival analysis model for calculating the hazard ratio thereof; optionally further comprising
[0025] 5) calculating the risk-score for one or more cancer patients using said signature genomic covariates or signature genomic and clinical covariates associated with a SIB - 5 - BW1367R
[0026] prognosis of step 3) corrected for their relative hazard ratio obtained from step 4); optionally
[0027] 6) clustering patient with a clustering algorithm on the basis of calculated risk-score of step 5).
[0028] Specifically, the inclusion of biologically consistent evolutionary co-occurrence covariates, such as evolutionary trajectory, together with single-gene mutations and clinical covariates enables the construction of a unified predictive model that achieves higher discrimination performance and refined patient classification. When applied to MDS as an exemplary application, the model incorporating evolutionary routes (e.g., ASXL1 KRAS, SRSF2 NRAS) and validated co-occurrences (e.g., NRAS / RUNX1, ATRX, JAK2) improved Harrell’s C-index from 0.75 to 0.77 in the cBioPortal, with lower AIC values confirming better model fit. More than 40 % of patients were reclassified into more appropriate risk categories, demonstrating the method’s superior ability to capture the prognostic impact of evolutionary trajectories beyond static molecular or clinical features.
[0029] In a second aspect, the present invention provides the Hazard ratios of signature genomic covariates or signature genomic and clinical covariates calculated with the computer implemented method, object of the first aspect of the invention.
[0030] In a third aspect, the present invention relates to the use of the hazard ratio calculated according to the computer implemented method object of the first aspect of the invention in a risk stratification method of cancer patients or in individuals with somatic mutations identified through gene sequencing.
[0031] In a fourth aspect, the present invention relates to the use of signature genomic covariates, evolutionary routes or signature genomic and clinical covariates associated with a specific patient outcome, in risk stratification method of cancer or in individuals with somatic mutations identified through gene sequencing.
[0032] In a fifth aspect, the present invention provides a computer implemented method for riskstratification of a cancer patient or in individuals with somatic mutations identified through gene sequencing comprising a step of calculating the risk-stratification of said patient using the hazard ratios calculated by the computer implemented method according to the first aspect of the invention.
[0033] In other words; the invention provides computer implemented risk stratification method of cancer patients or individuals with somatic mutations identified through gene sequencing, irrespective of a formal cancer diagnosis comprising the following step: SIB -6 - BW1367R
[0034] 1) providing a dataset of genomic covariates and prognosis data or a dataset of genomic and clinical covariates and prognosis data of a statistically significant number of cancer patients or individuals with somatic mutations identified through gene sequencing, irrespective of a formal cancer diagnosis;
[0035] 2) applying an evolutionary model to the dataset of step 1) for identifying statistically recurrent evolutionary genomic covariates according to said evolutionary method;
[0036] 3) processing the identified statistically recurrent evolutionary genomic covariates of step 2) or statistically recurrent evolutionary genomic covariates of step 2) and clinical covariates using a survival analysis model for selecting signature genomic covariates or signature genomic and clinical covariates related to a specific prognosis;
[0037] 4) processing the signature genomic covariates or signature genomic and clinical covariates selected by step 3) by a survival analysis model for calculating the effect measures thereof;
[0038] 5) calculating the risk-score for one or more cancer patients using signature genomic and clinical covariates associated with a prognosis of step 3) corrected for their effect measures ratio obtained from step 4);
[0039] wherein said recurrent evolutionary genomic covariates include single gene mutations, co-occurrence of mutations and / or evolutionary routes.
[0040] In a preferred embodiment, in the method object of the fifth aspect of the invention, the evolutionary model comprises one or more of, preferably at least, the following phases:
[0041] computing the Cancer Cell Fraction (CCF) of each somatic mutation or gene mutation from its Variant Allele Frequency (VAF), wherein said computation is adapted to account for copy number alterations (CNA) affecting the corresponding genomic locus;
[0042] calculating a Cross-Validation Rank (CV-Rank) and a Cross-Validation Score (CV-Score) for each genomic mutation and evolutionary trajectory; selecting of evolutionary trajectory having an EVO-score greater than 0,5. In a sixth aspect, the invention provides a computer program comprising instructions which, when the program is executed by a computer, cause the computer to carry out the computer implemented methods object of the first and fifth aspects of the invention. In a seventh aspect, the invention provides a computer-readable data carrier having stored there on the computer program of the sixth aspect of the invention.
[0043] The term “hazard ratio” can be freely replaced by the term “effect measures” in any of the listed aspects. SIB -7 - BW1367R
[0044] Brief description of figures
[0045] Figure 1- Overview of Patients selection criteria derived from cBioPortal in order to perform the analysis on MDS patients.
[0046] Diagram summarizing the composition of the MDS patient cohort derived from cBioPortal.
[0047] Figure 2 - Comparison of Variant Allele Frequency (VAF) and Cancer Cell Fraction (CCF) Across Genes in the cBioPortal cohort.
[0048] Bar plots showing the mean distribution of VAF and CCF values calculated with Ploidy-adjusted CCF was as: CCF_estimate= VAF * CN (Copy Number). CN was set to 3 if can CNA (Copy number alteration) gain, 1 if CNAloss, and 2 if no CNA event was detected.
[0049] Figure 3 - Timing of driver mutations during the progression of MDS
[0050] This figure illustrates the identification of evolutionary routes in MDS using the ProgEvo framework. Panels A-C show the CV-Rank distribution of mutations within the training dataset (cBioPortal). CV-Rank is a continuous variable reflecting inferred mutation timing, with lower values corresponding to earlier events. Mutations were stratified into three temporal ranks: 1st (CV-Rank 0-0.5), 2nd (0.5-1.5), and 3rd (>1.5). Panel D displays 45 evolutionary routes identified by ProgEvo in the cBioPortal cohort (Only routes with an Evo-Score higher than 0.5 and at least 10 co-occurrence events are present). Each route represents a directional gene pair (parent — > child) with consistent temporal order and probabilistic dependence. The width of each link reflects route frequency across 1,765 co-occurring pairs. (View Supplementary Appendix Table 4S for details)
[0051] Figure 4: Impact of IPSS-M-Evo on Risk Reclassification.
[0052] The Bar plot shows the redistribution of patients across risk categories after applying IPSS-M-Evo in the cBioPortal cohort. Kaplan-Meier curves (panel B) illustrates LFS stratification using IPSS-M in patients who were reclassified, while panels C displays LFS stratification using IPSS-M-Evo.
[0053] Description of the invention
[0054] The Applicant has developed a computer implemented invention that, from a given dataset of patient suffering from a specific disease, furnishes the hazard ratio of signature covariates associated with a specific prognosis, calculates the Risk-factor associated with said hazard ratio and stratifies the patients belonging to said given dataset on the basis of said calculated Risk- factor. Moreover, the Applicant has found that is also possible to calculate the Risk-factor and to stratify patients that suffers from SIB - 8 - BW1367R
[0055] a specific disease, but do not belong to the given dataset, on the basis of the hazard ratio calculated with the completer implemented method object of the invention. In other words, the present invention provides a method for the identification of a set of clinical covariates, genomic covariates and evolutionary trajectories and for the calculation of the effect measures thereof that when used in risk stratification method allows to improve patients stratification compared to the method known in the art.
[0056] Unless otherwise defined herein, scientific and technical terms used in connection with the present invention shall have meanings that are commonly understood by those of ordinary skill in the art. Further, unless otherwise required by context, singular terms shall include pluralities and plural terms shall include the singular. Generally, nomenclature used in connection with, and techniques of, biology, molecular biology, genetics, medicine, statistic, biostatistics described herein are those well-known and commonly used in the art. The methods and techniques of the present invention are generally performed according to conventional methods well known in the art and as described in various general and more specific references that are cited and discussed throughout the present specification unless otherwise indicated.
[0057] The following terms, unless otherwise indicated, shall be understood to have the following meanings:
[0058] the term “covariate” as used herein refers to a continuous variable that may or may not be related to a specific prognosis;
[0059] the term “confounder” as used herein refers to a variable that is related to the risk factor and to a specific prognosis; for the sake of clarity, a confounder can increase or decrease the likelihood of a prognosis;
[0060] the term “genomic mutation” and “genomic alteration” are synonymous and can be used interchangeably;
[0061] the term “somatic mutation” as used herein refers to a genetic (DNA) alteration that occurs in non-germline cells and is therefore not inherited through the germline but acquired during the lifetime of an individual;
[0062] the term “evolutionary route” or “evolution trajectory” as used herein refers to a pair of co-occurring genomic alterations in which the alteration of a gene is predicted to arise earlier in the progression of the disease, and is predicted to increase the likelihood of a distinct mutation in the second gene at a later stage of the disease;
[0063] the term “ploidy” as used herein refers to the number of chromosomal couples in a specific genome region; SIB - 9 - BW1367R
[0064] the term “hazard ratio” as used herein refers to the statistical measure of the effect of a particular variable on the risk of an event occurring over time, typically expressed as a ratio comparing the hazard rates between different groups; said term and the term “effect measures” can be used alternatively and replaced one another throughout the specification;
[0065] the term “driver mutation” as uses herein refers to a mutation within a gene that confers a selective growth advantage, thus promoting cancer development;
[0066] the term “bootstrap resampling” as used herein refers to a statistical technique that involves generating synthetic datasets by randomly sampling with replacement from the original dataset; for example, if the original dataset contains n rows (samples) and m columns (features), multiple synthetic datasets of the same dimensions (n x m) can be created by randomly selecting rows from the original dataset and allowing each selected row to be re-sampled, or reintroduced, into the dataset pool for potential selection in subsequent draws;
[0067] the term “outcome” as used herein refers to a measurable result or endpoint related to a patient's health status or disease progression. This may include, but is not limited to, clinical outcomes such as survival metrics (e.g., overall survival (OS), leukemia-free survival (LFS)), response to therapy, disease remission, recurrence, progression, or other health indicators relevant to the effectiveness of a treatment or the natural course of the disease;
[0068] the term “overall survival” (OS) as used herein refers to the length of time from the date of diagnosis that patients diagnosed with the disease remain alive;
[0069] the term “leukaemia-free survival” (LFS) as used herein refers to the length of time from the date of diagnosis to either leukemic transformation or death from any cause, with censored observations extending until the last follow-up;
[0070] the term “disease-free survival” (DFS) as used herein refers to the length of time from the date of initial treatment or remission to the recurrence of the disease or death from any cause, with censored observations extending until the last follow-up;
[0071] the term “continuous variable” as used herein refers to a variable that can take on an infinite number of values within a given range, allowing for fractional or decimal values. This type of variable is often used to represent measurements such as age, weight, or temperature, where precision to any level of decimal place is possible; SIB - 10 - BW1367R
[0072] the term “categorical variable” as used herein refers to a variable that represents distinct categories or groups, with no inherent order or ranking. Examples include variables such as blood type, disease stage, or patient sex, where each category is mutually exclusive; the term “binary variable” as used herein refers to a type of categorical variable that has only two possible values, often represented as 0 and 1. Binary variables are used to indicate the presence or absence of a specific attribute, such as disease status (e.g., affected / unaffected) or treatment response (e.g., responder / non-responder);
[0073] the term “outcome” as used herein refers to prognosis, response or lack of response to therapy;
[0074] the term “ProgEvo” as used herein refers to the computer implemented method object of the first aspect of the present invention;
[0075] the term “IPSS-M-Evo” as used herein refers to the computer implemented method object of the fifth aspect of the present invention.
[0076] In the present invention with the term “first”, “second” and so on, it is referred to different aspect of the invention; with the term “embodiment” it is referred to different implementation of said aspects; moreover, the features of an embodiment of a specific aspect can be associated with a different aspect; for the sake of clarity, if in an embodiment of the first aspect of the invention are listed specific features, said same features can be applied to second or other aspects of the invention.
[0077] In a first aspect, the present invention provides a computer implemented method for the calculation of hazard ratio of signature genomic covariates or signature genomic and clinical covariates or signature genomic, evolutionary routes and clinical covariates associated with a prognosis in a cancer patient comprising the following steps:
[0078] 1) providing a dataset of genomic covariates and prognosis data or a dataset of genomic and clinical covariates and prognosis data of a statistically significant number of cancer patients or individuals with somatic mutations identified through gene sequencing, irrespective of a formal cancer diagnosis.
[0079] 2) applying an evolutionary model to the dataset of step 1) for identifying statistically recurrent evolutionary genomic covariates according to said evolutionary method;
[0080] 3) processing the identified statistically recurrent evolutionary genomic covariates of step 2), for example single gene mutations, co-occurrence of mutations or evolutionary routes, or statistically recurrent evolutionary genomic covariates of step 2) and clinical covariates using a survival analysis model for selecting signature genomic covariates or signature genomic and clinical covariates related to a specific prognosis; SIB - 11 - BW1367R
[0081] 4) processing the signature genomic covariates or signature genomic and clinical covariates selected by step 3) by a survival analysis model for calculating the hazard ratio thereof.
[0082] The method of the first aspect of the invention also refers to a computer implemented method for the calculation of hazard ratio of signature genomic covariates or signature genomic and clinical covariates or signature genomic, evolutionary routes and signature clinical covariates associated with a prognosis in individuals with somatic mutations identified through gene sequencing, irrespective of a formal cancer diagnosis, comprising steps 1)-4).
[0083] For the sake of clarity, the term “recurrent evolutionary genomic covariates” is herein used to describe the set genomic covariates that are observed, during the evolution of a specific type of cancer of the evolution of a specific disease precursor of cancer into a specific cancer, across different patients; the terms “signature genomic covariates” and “signature genomic and clinical covariates” is herein used to describe the set of covariates that are not only are observed during the evolution of a specific type of cancer or the evolution of a specific disease precursor of cancer into a specific cancer, across different patients or individuals, but also which are significantly associated with a specific prognosis in said patients or individuals; as mere example, to a set a recurrent evolutionary genomic covariates can be split in multiple sub-sets of signature genomic covariates on the basis of the prognosis considered.
[0084] In other words, in its first aspect, the present invention provides a computer implemented method for the calculation of effect measures of signature genomic covariates or signature genomic and clinical covariates or signature genomic, evolutionary routes and clinical covariates associated with a prognosis in a cancer patient comprising the following steps:
[0085] 1) providing a dataset of genomic covariates and prognosis data or a dataset of genomic and clinical covariates and prognosis data of a statistically significant number of cancer patients or individuals with somatic mutations identified through gene sequencing, irrespective of a formal cancer diagnosis.
[0086] 2) applying an evolutionary model to said dataset for identifying statistically recurrent evolutionary genomic covariates according to said evolutionary method; wherein recurrent evolutionary genomic covariates includes single gene mutations, cooccurring mutations and / or evolutionary routes, SIB - 12 - BW1367R
[0087] 3) processing said statistically recurrent evolutionary genomic covariates and or said recurrent evolutionary genomic covariates and clinical covariates using a survival analysis model for selecting signature genomic covariates or signature genomic and clinical covariates related to a specific prognosis;
[0088] 4) processing said signature genomic covariates or signature genomic and clinical covariates by a survival analysis model for calculating the effect measures thereof. Dataset
[0089] A cancer patient is a subject that suffers from cancer or a disease precursor of cancer; in a preferred embodiment, the cancer patient suffers from cancer; in an alternative preferred embodiment, the cancer patient suffers from a disease precursor of cancer. For example, the dataset can contain genomic covariates, clinical covariates for subjects diagnosed with cancer, with a precursor of cancer that has not already evolved in cancer or with a precursor of cancer that has evolved in cancer. Typically, a disease precursor of cancer can be any type of disease which can potentially evolve in cancer; in particular, said disease precursor of cancer presents genomic covariates; in a preferred embodiment, a disease precursor of cancer can be, for example, atypical ductal hyperplasia, lobular carcinoma in situ, pulmonary adenoma, adenomatous polyps, Lynch’s syndrome, prostatic intraepithelial neoplasia, hereditary prostate cancer, chronic lymphocytic leukaemia, myelodysplastic syndrome, chronic lymphocytic leukaemia, dysplastic nevi, pancreatic intraductal neoplasia, serous tubal intraepithelial carcinoma, oral premalignant lesions, cystitis and urothelial carcinoma in situ and Barrett’s oesophagus; in a more preferred embodiment, the disease precursor of cancer is myelodysplastic syndrome. Typically, cancer can be any type of cancer; in particular, cancer can be solid tumour or hematologic cancer; in a preferred embodiment, breast cancer, lung cancer, colorectal cancer, prostate cancer, leukaemia, lymphomas, melanoma, pancreatic cancer, ovarian cancer, head and neck cancer, bladder cancer and oesophageal cancer; a more preferred embodiment, cancer can be leukaemia. The dataset can be considered as comprising a statistically significant number of cancer patients if the number of patients is at least 30; in a preferred embodiment, the number of patients is at least 100; in an alternative preferred embodiment, the number of patients is at least 500; in a more preferred embodiment, the number of patients is at least 1000; in an alternative more preferred embodiment, the number of patients is at least 3000. The same range are also valid for the individuals above mentioned. SIB - 13 - BW1367R
[0090] In preferred embodiment, the dataset is The Molecular International Prognostic Scoring System for Myelodysplastic Syndromes Dataset. Still preferably, said dataset comprises profiling of 3323 patients (from the original IPSS-M publication) retrieved via cBioPortal (https: / / www.cbioportal.org / study / summary?id=mds_iwg_2022); for the sake of clarity, the expression “cBioPortal dataset” refers to said dataset.
[0091] In certain embodiments, the method herein disclosed can run more than one datasets; for example, the computer implemented method can use a first dataset as training dataset or training cohort, another dataset or two or more other dataset as validation datasets or validation cohorts. In certain embodiments, cBioPortal dataset (as defined above) is the training dataset.
[0092] Prognosis or prognosis data refers to the forecast or prediction of the likely course and outcome of a disease; in a preferred embodiment; prognosis can be overall survival or disease-free survival; in particular, depending on the cancer or disease precursor of cancer, disease-free survival can be breast cancer-free survival, lung cancer-free survival, colorectal cancer- free survival, colorectal cancer-free survival, prostate cancer-free survival, leukaemia-free survival, lymphomas-free survival, melanoma-free survival, pancreatic cancer-free survival, ovarian cancer-free survival, head and neck cancer-free survival, bladder cancer-free survival and oesophageal cancer-free survival; in a more preferred embodiment, prognosis can be overall survival or leukaemia-free survival; in an even more preferred embodiment; prognosis is leukaemia-free survival.
[0093] Clinical covariates are variables or factors that may influence the outcome of, for example, a clinical study or a clinical trial; in a preferred embodiment, clinical covariates can be continuous variables, categorical variables and binary variables; in a preferred embodiment, continuous covariates can be, for example, age, haemoglobin levels, blood cell count, white blood count, platelet count, bone marrow blasts, chromosomal abnormalities, body mass index, blood pressure, cholesterol level, blood glucose level, liver function tests, serum creatinine, tumour side, duration of illness, heart rate, ejection fraction, heart rate, physical activity level, nutritional intake, c-reactive protein levels and vascular measurement; in a more preferred embodiment, continuous variable can be age, haemoglobin levels, blood cell count, white blood count, platelet count, bone marrow blasts and chromosomal abnormalities; in a preferred embodiment, categorical variable can, for example, risk classification, gender, ethnicity, smoking status, alcohol consumption, comorbidities, treatment group, performance status, family history of disease, tumour stage, histological tumour classification, residence location, surgical SIB - 14 - BW1367R
[0094] history, medication use; in a more preferred embodiment, categorical variable can be risk classification, tumour stage, histological tumour classification; in an even more preferred embodiment, risk classification is risk classification according to The International Prognostic Scoring System-Molecular (also called “IPSS-M”) method, such as IPSS-M score; said clinical-molecular prognostic model is described in Bernard E et al., NEJM Evid 2022, 1(7), 1-14, which is incorporated in its entirety by reference in the present application; supplementary information of Bernard E et al., NEJM Evid 2022, 1(7), 1-14 are incorporated by reference as well; in a preferred embodiment, binary variable, for example, can be histopathological diagnosis, hypertension, obesity, radiation therapy history, chemotherapy history, vaccination status, presence of symptoms, infectious disease status; in a more preferred embodiment, binary variable is histopathological diagnosis. Clinical covariates can be one or more selected from the group consisting of age, haemoglobin levels, blood cell count, white blood count, platelet count, bone marrow blasts, chromosomal abnormalities, risk classification according IPSS-M method, such as IPSS-M score, and histopathological diagnosis. Clinical covariates can be selected from one or more from the group consisting of age, absolute neutrophil count, platelet count, haemoglobin, bone marrow blasts, presence of del(5q), presence of -7 / del(7q), presence of -17 / del(17p), presence of complex karyoptype, cytogenetics category according to IPSS-M method. In a preferred embodiment, clinical covariates are selected from the group consisting of age, haemoglobin, platelet count, absolute neutrophil count (ANC), bone marrow blasts (%), cytogenetics category, presence of del(5q), presence of -7 / del(7q), presence of -17 / del(17p), presence of complex karyoptype and IPSS-M score.
[0095] Genomic covariates are specific genetic or genomic factors that may influence the outcome of a study, particularly in the context of personalized medicine and genomic research; in a preferred embodiment, genomic covariates can be, for example, gene mutations, copy number variations (CNVs), single nucleotide polymorphisms (SNPs), methylation status, chromosomal rearrangements, fusion gene, structural variants such as insertions-deletions (indels) or others, somatic nonsynonymous single-nucleotide variants (SNVs), and any other type of somatic genomic alteration.
[0096] Genomic covariates can either be single mutations or couple of co-occurring single mutations; in a preferred embodiment, couple of co-occurring single mutations are evolutionary trajectories, such as couple of co-occurring mutations in which one mutation of the couple consistently occurs before the other mutation of the couple across different SIB - 15 - BW1367R
[0097] patients; in a more preferred embodiment; evolutionary trajectories are couples of cooccurring mutations in which one mutation of the couple consistently occurs before the other mutation of the couple across different patient and the occurrence of the former mutation increases the likelihood of the occurrence of the latter mutation; in an even more preferred embodiment; genomic mutation are single mutations and / or evolutionary trajectories.
[0098] Typically, genomic covariates are obtained using a method of genome sequencing; in a preferred embodiment, gene sequencing can be, for example, sanger sequencing, nextgeneration sequencing (NGS), RNA sequencing, single-cell sequencing, nanopore sequencing, methylation sequencing, hybrid capture sequencing; in a preferred embodiment, next-generation sequencing is bulk sequencing or multiregional-bulk sequencing; in an even more preferred embodiment; gene sequencing are single-cell sequencing, bulk sequencing and multiregional bulk sequencing.
[0099] In a preferred embodiment, before the application of the method for gene sequencing a variant calling standard pipeline with default parameters is applied and a reference genome is adopted.
[0100] For example, genomic covariates can be identified, after applying a variant calling standard pipeline (i.e. GATK, https: / / qatk.broadinstitute.org / hc / en-us) with default parameters and adopting a reference genome (i.e. hg19 o hg38), through a method selected form new generation bulk sequencing (also called “NGS”) on a single biological sample (i.e. biopsy), single-cell sequencing or, only for solid tumours, multi-regional bulk sequencing.
[0101] In a preferred embodiment, from dataset of step 1) one or more sub-dataset can be generated; in a more preferred embodiment, the sub-dataset are obtained by filtering the dataset of step 1) on the basis of selected clinical covariates that the a statistically significant number of cancer patients have in common; in particular, selected clinical covariates can be any of the clinical covariate above-mentioned; preferably, selected clinical covariates can be age, sex, risk classification and histological tumour classification. Said filtering advantageously renders the sub-datasets more homogeneous than dataset of step 1), allowing to identify eventual relation between genomic covariates, selected clinical covariates and prognosis data. For example, if the specific prognosis of interest is the disease-free survival and the aim is to identify signature genomic covariates or signature genomic and clinical covariates associated with said prognosis, the dataset is filtered for those samples for which said prognosis is SIB - 16 - BW1367R
[0102] present in association with a continuous variable (i.e. time of follow-up: from enrolment to censoring or event).
[0103] In a preferred embodiment, a sub-dataset is obtained by filtering the dataset of step A) on the basis of selected prognosis data, for example only for those patients who have a confirmed diagnosis of certain type of cancer or disease; in certain embodiment, a diagnosis of MDS according to WHO-2016 classification criteria (Arber D. A. et al. Blood, 2016, 127, 2391-2405). Said filtered dataset can used for an evolutionary analysis that can be later be included in a prognostic model. Therefore, the computer implemented method of the invention can comprise a step A.1) of filtration of the dataset of step A) for those patient diagnosed with a specific type of cancer or disease, preferably, MDS to give a sub-dataset A.1). Therefore, in a preferred embodiment, the dataset of step 1) is cBioPortal dataset that is filtered according to step A.1)
[0104] In certain embodiment, the sub-dataset A.1) can be further filtered on the bases of clinical covariates, preferably age (for example availability of age at diagnosis and age < 90 years) and IPSSM feature (for example availability of IPSS-M score and availability of all IPSS-M clinical input variables, i.e. age, absolute neutrophil count, platelet count, haemoglobin, bone marrow blasts, presence of del(5q), presence of-7 / del(7q), presence of -17 / del(17p), presence of complex karyoptype, cytogenetics category.
[0105] Moreover, in order to establish and identify a relation between signature genomic covariates or signature genomic and clinical covariates associated with a prognosis, it is essential that, for each patient of the dataset, clinical and genomic covariates and prognosis data are available.
[0106] Evolutionary model and recurrent evolutionary genomic covariates identification The mere co-occurrence of couple of single mutations does not necessarily imply an evolutionary relationship and said couple of single mutations cannot be considered an evolutionary trajectory regardless. With the application of an evolutionary model, it is possible to identify evolutionary trajectories (also called “evolutionary routes”) by detecting co-occurrence of single mutations in a significant temporal ordering, indeed defining a parental relationship between mutations in a "parent gene” which consistently occur earlier than in a "child" gene across most analysed patients. The identification of these evolutionary trajectories is advantageous, in particular if associated with a specific prognosis, and their use in risk stratification methods advantageously allows stratify, or re-stratify, patients in a more accurate way, essential elements for the development of personalized treatment strategies. SIB - 17 - BW1367R
[0107] With the application of an evolution model, recurrent evolutionary genomic covariates are identified.
[0108] The evolutionary model includes three categories of evolution-informed features: evolutionary routes or trajectories: these are pairs of genomic alterations (u — > v) for which a temporal directionality of accumulation, based on an extended version of the probabilistic causality, is inferred. Specifically, an evolutionary step is defined when both the Temporal Priority (TP) condition — formalized through an agony-minimized ranking — and the Probability Raising (PR) condition — are met. These criteria ensure that the earlier event u statistically precedes and increases the likelihood of observing the subsequent event v, within a probabilistically consistent structure (Suppes-Bayes Causal Network).
[0109] co-occurring mutations or same-rank co-occurrences: these are mutation pairs that are consistently assigned to the same temporal rank across agony-minimized hierarchical ordering, and that do not violate the global evolutionary model. This represents alterations that are not temporally distinguishable and are likely to co-occur as part of stable sub-clonal architectures, single mutation or root-derived single mutations: these covariates correspond to individual gene-level mutational events modeled in the evolutionary framework as edges originating from a dummy root node R^v, where R denotes the wild-type genotype (unmutated) and v represents the acquisition of a somatic mutation (e.g. presence of a mutation on the ASXL1 gene vs wild type; reported as Root_to_GENEp). In the clinical prognostic model these features are introduced as baseline covariates in the model and allow for direct comparison between the prognostic contribution of an alteration considered alone (root-derived) versus its behavior within an evolutionary structure (e.g., as part of a route or co-occurrence). If a route or cooccurrence significantly outperforms the two corresponding individual mutations in multivariable models, this supports its added predictive value.
[0110] An evolutionary model comprises at least the following phases: mutation profiling, temporal modelling, statistical inference and model selection, model refinement and validation.
[0111] In a preferred embodiment, in mutation profiling phase cancer cell fraction is estimated for each genomic mutation of patient of dataset 1) to generate a mutational profile; in a more preferred embodiment, a single mutational profile is generated for genomic mutations obtained using bulk sequencing and multiple mutational profiles are generated for genomic mutations obtained using single-cell sequencing or multiregional bulk SIB - 18 - BW1367R
[0112] sequencing. Cancer cell fraction is calculated as variant allele frequency corrected for ploidy. In particular, in bulk sequencings cancer cell fractions (also called “CCF” or “CCFs”), calculated based on variant allele frequency (also called “VAF”) corrected for ploidy, is considered. For example, if SNV”1000C> T” is found in a specific region of diploid genome (2 copies) and its variant allele frequency is 0,15, CFF is calculated as follow:
[0113] CFF = VAF x ploidy
[0114] CFF = 0,15 x 2 = 0,3
[0115] For single-cell sequencing the presence or absence of a genomic mutation is considered in each cell of the sample. Either for bulk and for single-cell sequencings, every mutation is recorded on a gene or on a specific genomic region. According to anyone of the embodiments herein disclosed, mutation profiling phase comprises a first selection of genes having a mutation frequency not less than a fixed threshold. Preferably, mutation profiling phase comprises selection of genes with a mutation frequency >1%. In a preferred embodiment, the genes selected from said first selection are the following: ASXL1; ATRX; BCOR; BCORL1; CBL; CEBPA; CREBBP; CSF3R; CUX1; DNMT3A; EP300; ETNK1; ETV6; EZH2; FLT3; GATA2; GNAS; GNB1; IDH1; IDH2; JAK2; KIT; KMT2D; KRAS; MPL; NF1; NOTCH1; NPM1; NRAS; PHF6; PPM1D; PRPF40B; PTPN11; RAD21; RUNX1; SETBP1; SF3B1; SMC1A; SMC3; SRSF2; STAG2; TET2; TP53; U2AF1; WT1; ZRSR2.
[0116] Mutation profiling phase comprises the calculation of cancer cell fractions (CCF) for each gene. Therefore, in a preferred embodiment, the evolutionary model comprises a phase of computation of the Cancer Cell Fraction (CCF) of each somatic mutation or gene mutation from its Variant Allele Frequency (VAF), wherein said computation is adapted to account for copy number alterations (CNA) affecting the corresponding genomic locus, thereby correcting for allelic imbalance and enabling accurate temporal inference.
[0117] To note, for each gene, if more than one mutation is present in the same patient, the mutation with the highest Variant Allele Frequency (VAF) is retained for the calculation of CFF. VAFs are then corrected for estimated ploidy to derive cancer cell fractions (CCF), enabling temporal ordering. VAFs is then corrected for estimated ploidy to derive cancer cell fractions (CCF), enabling temporal ordering. Specifically, given the possible impact of Copy Number Alteration (CNA), especially in certain diseases as in MDS, an adjustment is performed considering copy-neutral loss of heterozygosity (CN-LOH) and CNA estimated in Bernard et al (NEJM Evidence 2022, 1) through the output data of SIB - 19 - BW1367R
[0118] CNA and CN-LOH in the validation dataset, preferably, is manually corrected. In a preferred embodiment, CCF, preferably ploidy-adjusted CCF, is calculated according to the following equation:
[0119] CCF_estimate= VAF * CN
[0120] Where:
[0121] • CN (Copy Number) was set to 3 if Copy Number Alteration (CNA) gain, 1 if CNA loss, and 2 if no CNA event was detected; and preferably
[0122] •Tumour purity was assumed to be 100% for all samples, and / or
[0123] •Sex-based normalization was applied for mutations located on sex chromosomes. In cases of copy-neutral loss of heterozygosity (CN-LOH) events, CCF was assumed to be equal to the uncorrected VAF.
[0124] In a preferred embodiment, in temporal modelling phase the mutation profiles are combined to build an agony-ranking score (also called “ranking-score”) (Tatti, N. Hierarchies in directed networks. In 2015 IEEE International Conference on Data Mining, 991-996. IEEE, (2015)), of single genomic mutations; in particular, the agony-ranking score defines a temporal ordering among mutation.
[0125] According to anyone of the embodiments herein described, in temporal modelling phase for each patient mutational profile (such as the list of its genomic covariates) is represented by a binary matrix Dcp of size m * n, where m is the number of patients and n the number of driver alterations. The matrix element d(p(i,j) = 1 if the j-th mutation is detected in sample / , and 0 otherwise. CCFs, computed from VAFs corrected for purity and local copy number (as reported above), served as temporal surrogates, with more clonal (higher CCF) events assumed to occur earlier in tumour evolution.
[0126] A temporal model of clonal architecture is defined for each patient as a directed acyclic graph (DAG):
[0127] Gp= (Vp, Ap)
[0128] Where:
[0129] • Vp= {v ∈ n | dp(i,v) = 1} is the set of observed mutations.
[0130] • A(p £ V(p x V(p encodes directional edges (vi -> vj) consistent with CCF(vi) > CCF(vj).
[0131] This formulation yields patient-specific evolutionary DAGs describing the likely order of mutation acquisition. Then, to construct a cohort-wide view of temporal orderings, individual DAGs G<p are merged into a consensus graph GP= (VP, AP), where VP= ∪ SIB - 20 - BW1367R
[0132] Vpand AP= ∪Ap. Aggregation introduces cycles due to inter-patient variability. Then, an agony-based ranking strategy is applied to derive a consistent global hierarchy. Agony quantifies the inconsistency of edge direction relative to a proposed hierarchy given a ranking function r: VP→ ℕ, the agony of an arc (u,v) is defined as:
[0133] Agony of each arc (u,v) = max{0, 1 + r(u) - r(v)} total agony of the graph:
[0134] a(GP, r) = Σmax{0, 1 + r(u) - r(v)} overall (u,v) in APThe objective is to find the ranking r*that minimizes the agony of the graph:
[0135] a(GP) = minra(GP, r)
[0136] Each gene was assigned a discrete rank derived from the minimization of the Agony graph, reflecting its relative position within the inferred temporal hierarchy.
[0137] In a preferred embodiment, wherein the dataset is, cBioPortal MDS cohort, this procedure provides three discrete ranks or groups: 0 (first), 1 (second), and 2 (third) corresponding to early, intermediate, and late mutational events, respectively, along the evolutionary timeline. Thus, in a certain embodiment, single mutations are grouped in two or more groups basing on the agony-ranking score, preferably in three groups: the 1stgroup or rank includes genes having mutation emerging at the earliest stages of disease evolution - included KMT2D, NOTCH1, PRPF40B, EP300, CREBBP, KIT, and ATRX, the 2ndgroup or rank includes BCORL1, CSF3R, JAK2, GATA2, DNMT3A, SF3B1, SRSF2, TET2, TP53, U2AF1, ASXL1 and EZH2, the 3rdgroup or rank includes MPL, ETNK1, SMC3, RAD21, PTPN11, KRAS, SMC1A, STAG2, GNB1, PHF6, RUNX1, CBL, PPM1D, and NRAS.
[0138] In a preferred embodiment, to evaluate the robustness of the inferred hierarchy and to derive a continuous, gene-level temporal score suitable for clinical interpretation, a resampling-based refinement procedure is applied.
[0139] In a preferred embodiment, in the statistical inference and model selection phase the significance of identified evolutionary trajectories is determined using a likelihood-based approach based on the theory of probabilistic causation, displaying the most significant relationships among driver mutations within a Bayesian Network that depicts evolutionary trajectories. This involves evaluating probabilistic causality conditions, Temporal Priority (TP) and Probability Raising (PR), to assess whether mutations in the “parent” gene significantly increase the likelihood of subsequent mutations in the “child” gene. This step utilizes an extended version of probabilistic causation theory, incorporating cross-sectional data from the temporal graph (GP) that includes all SIB - 21 - BW1367R
[0140] observed mutations. More in detail, selective advantage relationships by applying Suppes1probabilistic theory of causation. For a directed edge from mutation u to mutation v to be retained in the final network, two statistical conditions must be satisfied:
[0141] • Temporal Priority (TP): mutation u must precede v in the inferred temporal ranking, i.e., r(u) < r(y)
[0142] • Probability Raising (PR): the presence of mutation u must significantly increase the probability of observing mutation v, evaluated as: P{u,v\tu<tv)> P{u,v\tu>tv) This implies that u is a prima facie cause of v if it both precedes and raises the conditional probability of v, consistent with causal patterns across patients.
[0143] Only mutation pairs satisfying both TP and PR criteria are retained as valid evolutionary routes and included in the set of edges. Once these pairwise constraints are satisfied, the ensemble of directed relations is structured into a Bayesian Network by maximum likelihood estimation. This probabilistic graphical model encodes the most significant and reproducible mutation trajectories, grounded in cohort-wide evidence of temporal directionality and conditional dependence.
[0144] According to anyone of the embodiments herein disclosed, the evolution model of step 2) comprises a model optimization phase, preferably regularized, to select the evolutionary model that best fits the data.
[0145] In preferred embodiments, the computer implemented method of the invention comprises structure learning phase comprising the use of Akaike Information Criterion (AIC) regularization to model fit and complexity. In detail, given a dataset D of N fully observed binary instances and a candidate network structure G, the score associated with G is defined as:
[0146] ScoreAIC(G) = sumjt1suiTij1=1logP ^x^ Pa (x^)J—Dim(G)
[0147]
[0148] Here^x^ denotes the value of variable Xj in the i-th data sample, and Pa^xf1-1) denotes
[0149]
[0150] the assignment of its parent variables. Since all variables are binary, the number of free parameters for each variable Xj is Dimj = (2 - 1) * 2lPa(xi)l = 2 lPa(xi)l, leading to a total model dimension of Dim(G) = sumj1=12lPa(xi)l. The network structure G that maximizes this AIC score and is consistent with our evolution model, is selected as the optimal model.
[0151] The final model is selected via maximum likelihood estimation with AIC regularization over a constrained space of candidate graphs, reducing the risk of including unsupported SIB - 22 - BW1367R
[0152] edges. In addition, a CV-Score is for each inferred evolutionary step, quantifying the stability of each relationship across resampling iterations.
[0153] In a preferred embodiment, the evolutionary model comprises a phase of calculation a Cross-Validation Rank (CV-Rank) and a Cross-Validation Score (CV-Score) for each genomic mutation and evolutionary trajectory, preferably wherein Cross-Validation Rank (CV-Rank) and a Cross-Validation Score (CV-Score) represent, respectively, the average inferred temporal order and the frequency of recovery of each evolutionary edge across resampling iterations
[0154] Preferably, to evaluate the robustness of the inferred hierarchy and to derive a continuous, gene-level temporal score suitable for clinical interpretation, a resamplingbased refinement procedure is applied to compute Cross-Validation Rank (CV-Rank). Specifically, a large number (preferably greater than 100) of subsampling iterations without replacement are performed, each time randomly selecting a fraction of the patient cohort (preferably greater than 80%) from the reference dataset (e.g., cBioPortal). For each subset, agony minimization is re-applied to calculate a discrete agony rank r. The final continuous CV-Rank for each gene is then calculated as the average of its r values across all subsampling iterations.
[0155] Preferably, to evaluate the stability of each evolutionary route and to ensure that only robust routes capable of competing with clinical variables in predicting clinical outcomes are retained, a Cross-Validation Score (CV-Score) was computed as follows.
[0156] A large number (preferably greater than 100) of subsampling iterations without replacement are performed, each time randomly selecting a fraction of the patient cohort (preferably greater than 80%) from the reference dataset (e.g., cBioPortal).
[0157] For each iteration, the full evolutionary model is recomputed. For every edge u— v, we computed the corresponding CV-Score as:
[0158] CV-Scoreu,t= (Number of iterations in which u— vis inferred) / Number of total iterations
[0159] For example, a CV-Score of 0.64 indicates that the route u — > v was recovered in 64 out of 100 iterations.
[0160] This procedure provides a quantitative measure of the uncertainty and reproducibility of each inferred evolutionary relationship, enabling the identification of stable routes with SIB - 23 - BW1367R
[0161] sufficient statistical robustness to be integrated alongside clinical covariates in the prediction of clinical outcomes.
[0162] This aspect allows one to quantitively estimate the uncertainty of the inferred model in a novel way with respect to existing approaches, such as ASCETIC (Fontana et Al. Nat. Com. 2023).
[0163] Significance of temporal relationships is quantified with an Evolutionary Score (also called “Evo-Score”); said score quantifies the likelihood that identified mutations follow a consistent evolutionary trajectory, based on their co-occurrence and CCF relationships. For each evolutionary route inferred by the model an Evo-Score is defined as the percentage of samples in which a given parent gene had a higher CCF in comparison with its child gene. An Evo-Score greater than 0.4 is considered to signify a strong directional relationship between mutations.
[0164] In a preferred embodiment, the evolutionary model comprises a phase of selection of evolutionary trajectory bases on their Evo-Score; preferably having an EVO-score greater than 0,5, more preferably occurring in at least 10 patients of the dataset and having an EVO-score greater than 0,5. In a preferred embodiment, evolutionary routes were considered validated if they satisfied both of the following criteria, preferably cBioPortal dataset:
[0165] • Evo-Score > 0.50
[0166] • Preferably, present in >10 patients
[0167] In other preferred embodiments, the evolutionary trajectory can be validated even when occurring in different number of patients; those skilled in the art would establish case by case the proper threshold according to their experience, the type of dataset provided to the computer implemented method, etc..
[0168] In a preferred embodiment, within the context of MDS, the evolutionary model provides the following evolutionary trajectories:
[0169] PARENT CHILD EVO_SCORE CV_SCORE COOCCURRENCE ASXL1 BCOR 0,78 0,77 51
[0170] ASXL1 CBL 0,67 0,91 49
[0171] ASXL1 CEBPA 0,78 0,87 41
[0172] ASXL1 KRAS 0,76 0,8 25
[0173] ASXL1 NF1 0,65 0,75 46
[0174] ASXL1 PHF6 0,71 0,9 49
[0175]
[0176] SIB -24 - BW1367R
[0177] ASXL1 PTPN11 0,65 0,69 20
[0178] ASXL1 RUNX1 0,67 0,91 180
[0179] ASXL1 SMC1A 0,73 0,53 15
[0180] ASXL1 STAG2 0,71 0,91 163
[0181] ASXL1 GNB1 0,50 0,3 6
[0182] ATRX TP53 0,58 0,49 12
[0183] ATRX DNMT3A 0,60 0,37 5
[0184] ATRX U2AF1 0,50 0,26 8
[0185] BCORL1 NRAS 0,86 0,3 7
[0186] BCORL1 BCOR 0,39 0,83 18
[0187] CREBBP PPM1D 1,00 0,69 3
[0188] CREBBP SF3B1 0,69 0,71 13
[0189] CREBBP ASXL1 0,50 0,91 36
[0190] CSF3R CEBPA 0,33 0,52 3
[0191] CUX1 PHF6 0,79 0,64 14
[0192] CUX1 MPL 0,63 0,47 8
[0193] DNMT3A BCOR 0,60 0,66 47
[0194] DNMT3A GNB1 0,60 0,41 10
[0195] EP300 CUX1 0,44 0,5 9
[0196] EZH2 ETNK1 0,62 0,84 13
[0197] EZH2 RUNX1 0,69 0,88 62
[0198] EZH2 STAG2 0,63 0,6 38
[0199] EZH2 ZRSR2 0,41 0,42 27
[0200] FLT3 SMC1A 0,67 0,57 3
[0201] FLT3 PTPN11 0,50 0,43 4
[0202] KIT TET2 0,65 0,3 17
[0203] KIT CUX1 0,57 0,59 7
[0204] KIT BCORL1 0,33 0,25 3
[0205] KMT2D FLT3 0,83 0,55 6
[0206] KMT2D NF1 0,62 0,19 13
[0207] KMT2D SF3B1 0,81 0,48 48
[0208] KMT2D U2AF1 0,67 0,8 9
[0209] KMT2D EZH2 0,43 0,29 7
[0210]
[0211] SIB - 25 - BW1367R
[0212] NOTCH 1 SRSF2 1,00 0,91 7
[0213] NOTCH 1 SMC3 1,00 0,49 4
[0214] NOTCH 1 SETBP1 0,89 0,43 9
[0215] NOTCH 1 DNMT3A 0,73 0,47 11
[0216] NPM1 SMC3 1,00 0,61 4
[0217] NPM1 KRAS 0,75 0,43 4
[0218] NPM1 NRAS 0,71 0,6 7
[0219] PRPF40B WT1 0,67 0,42 3
[0220] SETBP1 PTPN11 0,88 0,7 8
[0221] SETBP1 ETV6 0,79 0,52 19
[0222] SF3B1 GNB1 0,75 0,06 8
[0223] SF3B1 RUNX1 0,81 0,55 53
[0224] SF3B1 ZRSR2 0,41 0,85 17
[0225] SRSF2 CBL 0,90 0,75 21
[0226] SRSF2 CEBPA 0,83 0,88 24
[0227] SRSF2 ETNK1 0,76 0,84 17
[0228] SRSF2 ETV6 0,75 0,57 20
[0229] SRSF2 GNAS 0,64 0,99 14
[0230] SRSF2 IDH1 0,75 0,93 32
[0231] SRSF2 KRAS 1,00 0,5 16
[0232] SRSF2 NRAS 0,92 0,77 25
[0233] SRSF2 RAD21 0,60 0,55 10
[0234] SRSF2 RUNX1 0,75 1 121
[0235] SRSF2 STAG2 0,78 1 114
[0236] SRSF2 ZRSR2 0,44 0,87 9
[0237] TET2 MPL 0,71 0,56 28
[0238] TET2 STAG2 0,75 0,43 79
[0239] TET2 ZRSR2 0,72 0,88 90
[0240] TET2 IDH1 0,46 0,91 13
[0241] TP53 PHF6 0,75 0,26 4
[0242] TP53 NF1 0,59 0,03 17
[0243] TP53 PPM1D 0,67 1 21
[0244] U2AF1 BCOR 0,62 0,95 37
[0245]
[0246] SIB - 26 - BW1367R
[0247] U2AF1 ETNK1 0,60 0,8 15
[0248] U2AF1 ETV6 0,69 0,44 13
[0249] U2AF1 MPL 0,86 0,77 14
[0250] U2AF1 PHF6 0,75 0,99 28
[0251] WT1 CBL 0,86 0,69 7
[0252]
[0253] Tab e 1
[0254] Routes displaying an Evo-score greater than 0.5 and, preferably, supported by at least 10 co-occurrence events in the dataset are considered evolutionary trajectory included in the statistically recurrent evolutionary genomics covariates.
[0255] The evolution of a gene from not mutated (wild type) to mutated can be considered an evolutionary trajectory, as the couples of co-occurring single mutations above-mentioned. In other words, evolutionary trajectory are root to mutation, co-occurring mutations, couple of parent-child mutations. In one preferred embodiment, evolutionary trajectories are absence of mutations in a gene to presence of mutations in gene or mutations in a group with lower agony-rank score evolving to mutations in a group with higher agony- rank score. According to anyone of the embodiment of the invention, to validated the identified evolutionary routes or trajectory, the computer implemented method further comprises a step 2.1) in which evolutionary model of step 2) is applied a dataset different from the dataset of step 1) for identifying statistically recurrent evolutionary genomic covariates according to said evolutionary method; preferably, if step 2) was applied to a training dataset, step 2.1) is applied to a validation dataset. Evolutionary trajectories can originate from early mutation and progress to intermediate-early or intermediate-late or late mutation and from intermediated-early to intermediate-late or late mutation; in particular, evolutionary trajectories can originate from early mutation and progress to intermediate-early mutation, preferably from CREBBP to SF3B1, from KMT2D to U2AF1 or SF3B1, from NOTCH 1 to SRSF2 or SETBP1 or DNMT3A; in particular, evolutionary trajectories can originate from early mutation and progress to intermediate-late mutations, preferably from KMT2D to FLT3; in particular, evolutionary trajectories can originate from early mutation and progress to late mutations, preferably from CREBP to PPM1D and NOTCH1 to SMC3; in particular, evolutionary trajectories can originate from intermediate-early mutation and progress to intermediate-late mutations, preferably from ASXL1 to NF1 or CEBPA or BCOR, from SRSF2 to IDH1 or CEBPA or ETV6, from TET2 to ZRSR2; in particular, evolutionary trajectories can originate from intermediate-early mutation and progress to late mutations, preferably SIB - 27 - BW1367R
[0256] from ASXL1 to STAG2 or SMC1A or RUNX1 or PTPN11 or PHF6 or KRAS or CBL, from SRSF2 to RUNX1 or STAG2 or NRAS or CBL or ETNK1, from TET2 to STAG2 or MPL. In a preferred embodiment, genomic covariates are one more of the recurrent evolutionary genomic covariates; in a more preferred embodiment, recurrent evolutionary genomic covariates are single mutation and / or co-occurrence of mutation and / or evolutionary trajectories observed in oncologic patients or in individuals with somatic mutations identified through gene sequencing, irrespective of a formal cancer diagnosis. This embodiment encompasses applications, for example, in conditions such as clonal haematopoiesis of indeterminate potential (CHIP) and disorders with dysplastic changes, including colorectal and oesophageal dysplastic / pre-cancerous lesions. In a more preferred embodiment, these recurrent evolutionary genomic covariates include single mutations, co-occurring mutations, and / or evolutionary trajectories. Preferably, within the context of MDS, single mutations are those genes most frequently implicated in MDS or routinely assessed in clinical practice, thereby enhancing the model’s clinical applicability and translational utility. The choice of these particular trajectories is grounded in their high prevalence in MDS and their representation in established clinical and diagnostic panels, ensuring the model’s optimization for prognostic analysis with direct clinical relevance. (G. Hoog et Al. Cancer Genetics, 2023, Sauta et Al. JCO, 2023 Nazha et Al., JCO, 2021 ). The preferred mutations may include one or more of the following gene: ASXL1 to BCOR, ASXL1 to CBL, ASXL1 to IDH1, ASXL1 to IDH2, ASXL1 to KRAS, ASXL1 to NF1, ASXL1 to PHF6, ASXL1 to RUNX1, ASXL1 to STAG2, DNMT3A to BCOR, EZH2 to RUNX1, EZH2 to STAG2, EZH2 to ZRSR2, SF3B1 to RUNX1, SF3B1 to ZRSR2, SRSF2 to CBL, SRSF2 to IDH1, SRSF2 to IDH2, SRSF2 to NRAS, SRSF2 to RUNX1, SRSF2 to STAG2, TET2 to MPL, TET2 to STAG2, TET2 to ZRSR2; in a more preferred embodiment, evolutionary trajectories are absence of mutations in a gene to mutation in one or more of genes, also called root to mutation, selected from ASXL1, ATRX, BCOR, BCORL1, CBL, CEBPA, CUX1, DNMT3A, EP300, ETNK1, ETV6, EZH2, FLT3, GATA2, GNAS, IDH1, IDH2, JAK2, KIT, KRAS, MPL, NF1, NOTCH1, NPM1, NRAS, PHF6, PTPN11, RAD21, RUNX1, SETBP1, SF3B1, SMC1A, SMC3, SRSF2, STAG2, TET2, TP53, U2AF1, WT1, ZRSR2 or ASXL1 to BCOR, ASXL1 to CBL, ASXL1 to IDH1, ASXL1 to IDH2, ASXL1 to KRAS, ASXL1 to NF1, ASXL1 to PHF6, ASXL1 to RUNX1, ASXL1 to STAG2, DNMT3A to BCOR, EZH2 to RUNX1, EZH2 to STAG2, EZH2 to ZRSR2, SF3B1 to RUNX1, SF3B1 to ZRSR2, SRSF2 to CBL, SIB - 28 - BW1367R
[0257] SRSF2 to IDH1, SRSF2 to IDH2, SRSF2 to NRAS, SRSF2 to RUNX1, SRSF2 to STAG2, TET2 to MPL, TET2 to STAG2, TET2 to ZRSR2
[0258] In a preferred embodiment, in the model refinement and validation phase the final set of evolutionary trajectories is refined using resampling / bootstrap analysis techniques to ensure robustness. In detail, to assess the reproducibility of inferred evolutionary trajectory, 100 subsampling iterations without replacement has been performed, each time randomly selecting 80% of the cBioPortal cohort. For each iteration, the full evolutionary model was recomputed, in other words repeating all the phases here described of the evolutionary model. Strong probabilistic association between the two genes, such as where a mutation in the parent gene significantly increases the likelihood of a subsequent mutation in the child gene, is measured by a Resampling Score (also called “Cv-Score”); said score reflects the consistency of evolutionary steps across multiple iterations. A Cv-Score greater than 0.5 is considered to indicate high reliability. For every edge u v, we computed the corresponding CV-Score as:
[0259] CV-Scoreu^v = (Number of iterations in which u v is inferred) I Number of total iterations
[0260] For example, a CV-Score of 0.64 indicates that the route u v was recovered in 64 out of 100 iterations.
[0261] In a preferred embodiment, the evolutionary model can be, for example, ASCETIC (Agony-baSed Cancer EvoluTion Inference), OncoNEM, PINTA, AncesTree, Treeomics, PhyloWGS, ClonEvol, PyClone, CancerPathfinder, SCITE, TrAp, CITIIP, ClonePhy, CAPRI, CAPRESE, PMCE, LACE, Expands, SiFit, PhyloSub, Canopy, LICHeE, PhlSCS, Clomial, PhylogicNDT, MACHINA and ScGPS; preferably, the evolutionary model is the framework ASCETIC (Agony-baSed Cancer EvoluTion Inference), described in Fontana D. et al., Nature Communication 2023, 14, 5982, document which is in its entirety incorporated by reference and constitutes part of the present detailed description. Supplementary information of Fontana D. et al., Nature Communication 2023, 14, 5982 is incorporated by reference as well. In a more preferred embodiment, ASCETIC is applied to dataset of claim 1) to identify recurrent evolutionary genomic covariates; recurrent evolutionary genomic covariates are single mutation or evolutionary trajectories; preferably, single mutation can be mutation in one or more of the following genes ASXL1, ATRX, BCOR, BCORL1, CBL, CEBPA, CUX1, DNMT3A, EP300, ETNK1, ETV6, EZH2, FLT3, GATA2, GNAS, IDH1, IDH2, JAK2, KIT, KRAS, MPL, NF1, SIB - 29 - BW1367R
[0262] NOTCH1, NPM1, NRAS, PHF6, PTPN11, RAD21, RUNX1, SETBP1, SF3B1, SMC1A, SMC3, SRSF2, STAG2, TET2, TP53, LI2AF1, WT1, ZRSR2;, evolutionary trajectory can be selected from the group consisting of ASXL1 to BCOR, ASXL1 to CBL, ASXL1 to IDH1, ASXL1 to IDH2, ASXL1 to KRAS, ASXL1 to NF1, ASXL1 to PHF6, ASXL1 to RUNX1, ASXL1 to STAG2, DNMT3A to BCOR, EZH2 to RUNX1, EZH2 to STAG2, EZH2 to ZRSR2, SF3B1 to RUNX1, SF3B1 to ZRSR2, SRSF2 to CBL, SRSF2 to IDH1, SRSF2 to IDH2, SRSF2 to NRAS, SRSF2 to RUNX1, SRSF2 to STAG2, TET2 to MPL, TET2 to STAG2, TET2 to ZRSR2.
[0263] Survival analysis model and selecting signature covariates associated with a specific prognosis
[0264] In a preferred embodiment, to assess the prognostic relevance of the inferred evolutionary relationships, each stable route derived from the evolutionary model is evaluated through univariate Cox regression analysis for overall survival. This analysis included both individual parent genes and their downstream evolutionary trajectories, in order to determine whether the inclusion of temporal directionality affected the prognostic strength of specific genomic events.
[0265] Therefore, in preferred embodiment, the methods herein disclosed further comprise the application of a survival model, preferably Univariate Cox regression, to the evolutionary trajectories obtained from the evolutionary model and to genomic covariates, wherein said genomic covariates are the parent gene mutation of said evolutionary trajectory, to identify evolutionary trajectory associated with a specific prognosis, preferably overall survival.
[0266] Mutations in ASXL1, SRSF2, RUNX1, and STAG2 are associated with adverse prognosis, whereas SF3B1 retains a favorable effect on survival. However, when directional evolutionary information are incorporated, distinct variations in prognostic impact are observed. For instance, ASXL1 mutations shows a significant increase in hazard ratio when followed by downstream mutations such as KRAS or STAG2. Similarly, SRSF2 mutations evolving into NRAS orSTAG2 are linked to markedly poorer outcomes, while the route SF3B1 -> RUNX1 indicates that the late acquisition of transcriptional regulators could invert the favorable effect of early splicing mutations. These results support the hypothesis that the temporal sequence of mutation acquisition carries prognostic information beyond the mere presence of individual mutations.
[0267] The impact of stable evolutionary routes on overall survival, as estimated by univariate Cox regression analysis, is summarized below: SIB - 30 - BW1367R
[0268] Pare Evolutio Number in Number in Survival LFS 95% 95% nt nary Evolution Dataset Dataset (% of Hazard Cl Cl gen Routes (% of parent gene) parent gene) Ratio Lowe llppe e r r 2519 1987
[0269] ASX 651 517 1,67 1,46 1,91 L1
[0270] ASXL1 51 (7.8%) 40 (7.7%) 2,01 1,39 2,90 to
[0271] BCOR ASXL1 25 (3.8%) 17 (3.3%) 2,87 1,72 4,78 to
[0272] KRAS ASXL1 46 (7.1%) 35 (6.8%) 1,96 1,34 2,88 to NF1
[0273] ASXL1 49 (7.5%) 43 (8.3%) 1,78 1,20 2,63 to PHF6
[0274] ASXL1 163 (25.0%) 134 (25.9%) 2,51 2,04 3,08 to
[0275] STAG2
[0276] DN 436 352 1,09 0,93 1,28 MT3
[0277] A DNMT3 47 (10.8%) 39 (11.1%) 1,95 1,33 2,86 A to
[0278] BCOR EZH 168 129 1,69 1,35 2,12 2
[0279] EZH2 to 62 (36.9%) 46 (35.7%) 2,06 1,46 2,92 RUNX1
[0280] EZH2 to 38 (22.6%) 27 (20.9%) 2,44 1,57 3,81 STAG2
[0281]
[0282] SIB - 31 - BW1367R
[0283] SF3 615 492 0,61 0,52 0,71 B1
[0284] SF3B1 53 (8.6%) 40 (8.1%) 1,84 1,26 2,68 to
[0285] RUNX1
[0286] SRS 365 283 1,51 1,28 1,77 F2
[0287] SRSF2 21 (5.8%) 14 (4.9%) 1,68 0,87 3,24 to CBL
[0288] SRSF2 32 (8.8%) 23 (8.1%) 1,25 0,72 2,16 to IDH1
[0289] SRSF2 25 (6.8%) 17 (6.0%) 2,84 1,64 4,92 to
[0290] NRAS SRSF2 121 (33.2%) 93 (32.9%) 1,65 1,28 2,13 to
[0291] RUNX1
[0292] SRSF2 114 (31.2%) 97 (34.3%) 2,24 1,76 2,85 to
[0293] STAG2
[0294] TET 716 564 0,87 0,76 1,00 2
[0295] TET2 to 28 (3.9%) 19 (3.4%) 0,73 0,36 1,46 MPL TET2 to 79 (11.0%) 64 (11.3%) 2,72 2,03 3,66 STAG2
[0296] TET2 to 90 (12.6%) 75 (13.3%) 0,79 0,56 1,13 ZRSR2
[0297]
[0298] Table 2
[0299] The impact of stable evolutionary routes on overall survival, as estimated by univariate Cox regression analysis, is summarized below:
[0300] FEATURE HR LOWER UPPER NUM_MU FREQ_MU T T
[0301]
[0302] SIB - 32 - BW1367R
[0303] ASXL1 1,66 1,45 1,91 517 26,02% ASXL1_to_BCOR 2,00 1,36 2,93 40 2,01% ASXL1_to_KRAS 3,34 2,00 5,57 17 0,86% ASXL1_to_NF1 2,14 1,45 3,17 35 1,76% ASXL1_to_PHF6 1,72 1,16 2,56 43 2,16% ASXL1_to_STAG2 2,30 1,87 2,84 134 6,74% ATRX 1,46 1,00 2,13 44 2,21% BCOR 1,66 1,32 2,09 132 6,64% CBL 1,49 1,12 1,97 93 4,68% CEBPA 2,01 1,41 2,86 52 2,62% CUX1 0,97 0,74 1,29 111 5,59% DNMT3A 1,05 0,89 1,24 352 17,72% DNMT3A_to_BCO 1,65 1,08 2,52 39 1,96% R ETV6 1,76 1,21 2,56 44 2,21% EZH2 1,67 1,33 2,11 129 6,49% EZH2_to_RUNX1 2,18 1,52 3,13 46 2,32% EZH2_to_STAG2 2,31 1,45 3,69 27 1,36% GNAS 0,94 0,49 1,81 21 1,06% IDH1 0,95 0,66 1,36 63 3,17% KIT 1,11 0,67 1,84 32 1,61% KRAS 2,14 1,45 3,17 36 1,81% MPL 1,03 0,71 1,50 56 2,82% NF1 1,68 1,29 2,18 88 4,43% NOTCH 1 0,89 0,64 1,25 83 4,18% NRAS 1,92 1,36 2,71 49 2,47% PHF6 1,44 1,04 2,00 67 3,37% PTPN11 1,38 0,86 2,24 30 1,51% RAD21 1,40 0,82 2,37 27 1,36% RUNX1 1,91 1,62 2,27 249 12,53% SF3B1 0,64 0,55 0,75 492 24,76% SF3B1_to_RUNX1 1,58 1,05 2,37 40 2,01% SMC1A 1,67 1,01 2,79 25 1,26%
[0304]
[0305] SIB - 33 - BW1367R
[0306] SRSF2 1,45 1,22 1,71 283 14,24% SRSF2_to_CBL 1,76 0,91 3,39 14 0,70% SRSF2_to_IDH1 1,07 0,59 1,94 23 1,16% SRSF2_to_NRAS 3,34 1,93 5,79 17 0,86% SRSF2_to_RUNX1 1,51 1,16 1,97 93 4,68% SRSF2_to_STAG2 2,18 1,71 2,78 97 4,88% STAG2 2,13 1,77 2,56 197 9,91% TET2 0,90 0,78 1,04 564 28,38% TET2_to_MPL 0,82 0,41 1,64 19 0,96% TET2_to_STAG2 2,61 1,93 3,54 64 3,22% TET2_to_ZRSR2 0,85 0,60 1,21 75 3,77% TP53 2,63 2,23 3,09 247 12,43% U2AF1 1,58 1,30 1,93 188 9,46% ZRSR2 0,86 0,65 1,14 119 5,99%
[0307]
[0308] Table 3
[0309] Hazard ratios (HRs) for overall survival were estimated using univariate Cox proportional hazards models. Frequencies (FREQ_MUT) represent the number of patients exhibiting each evolutionary route or mutation over the total number of patients included in the survival analysis.
[0310] Beyond directional routes, the model also identified recurrent co-occurrence events that belong to the same evolutionary rank and do not violate the inferred temporal hierarchy. These represent parallel or convergent evolutionary processes that contribute independently to disease progression.
[0311] Gene pairs such as RUNX1-STAG2, BCOR-STAG2, and NRAS-STAG2 displayed strong prognostic associations, with hazard ratios exceeding 2.0, confirming their impact on overall survival.
[0312] The prognostic impact of gene co-occurrences on Leukemia Free Survival, as determined by univariate Cox regression, is reported in the following table:
[0313] FEATURE HR LOWER UPPER NUM_MUT FREQ_MU T ASXL1_and_DNMT3A 1,90 1,35 2,67 55 2,77% ASXL1_and_EZH2 1,79 1,35 2,36 83 4,18% ASXL1_and_IDH2 2,20 1,57 3,08 48 2,42%
[0314]
[0315] SIB - 34 - BW1367R
[0316] ASXL1_and_SF3B1 0,92 0,68 1,25 85 4,28% ASXL1_and_SRSF2 2,06 1,67 2,55 139 7,00% ASXL1_and_TET2 1,54 1,25 1,89 168 8,45% ASXL1_and_U2AF1 1,73 1,33 2,25 86 4,33% BCOR_and_RUNX1 2,22 1,61 3,06 54 2,72% BCOR_and_STAG2 2,54 1,79 3,60 42 2,11% CBL_and_RUNX1 2,93 1,69 5,07 16 0,81% DNMT3A_and_EZH2 1,03 0,57 1,87 23 1,16% DNMT3A_and_IDH2 1,41 0,76 2,63 18 0,91% DNMT3A_and_SF3B1 0,70 0,54 0,91 130 6,54% DNMT3A_and_SRSF2 2,09 1,33 3,30 22 1,11% DNMT3A_and_TET2 0,94 0,70 1,25 95 4,78% DNMT3A_and_U2AF1 1,26 0,78 2,03 34 1,71% EZH2_and_SF3B1 0,97 0,55 1,71 24 1,21% EZH2_and_TET2 1,96 1,42 2,71 53 2,67% IDH2_and_SRSF2 1,86 1,31 2,65 46 2,32% JAK2_and_SF3B1 1,05 0,50 2,20 14 0,70% JAK2_and_TET2 1,22 0,66 2,28 18 0,91% NRAS_and_RUNX1 1,99 1,07 3,72 14 0,70% NRAS_and_STAG2 2,44 1,38 4,32 17 0,86% PHF6_and_RUNX1 2,40 1,41 4,07 21 1,06% RUNX1_and_STAG2 2,42 1,85 3,15 77 3,88% SF3B1_and_SRSF2 1,14 0,57 2,30 17 0,86% SF3B1_and_TET2 0,62 0,48 0,79 174 8,76% SF3B1_and_TP53 1,22 0,81 1,85 34 1,71% SRSF2_and_TET2 1,21 0,96 1,53 132 6,64% TET2_and_TP53 1,57 1,03 2,37 38 1,91% TET2_and_U2AF1 1,68 1,09 2,58 36 1,81%
[0317]
[0318] Table 4
[0319] Hazard ratios (HRs) for overall survival were estimated using univariate Cox proportional hazards models. Frequencies (FREQ_MUT) represent the number of patients exhibiting each evolutionary route or mutation over the total number of patients included in the survival analysis. SIB - 35 - BW1367R
[0320] The prognostic impact of gene co-occurrences on Overall Survival, as determined by univariate Cox regression, is reported in the following table:
[0321] FEATURE HR LOWER UPPER NUM_MUT FREQ_MUT ASXL1_and_DNMT3A 1,72 1,20 2,46 55 2,77% ASXL1_and_EZH2 1,80 1,35 2,40 83 4,18% ASXL1_and_IDH2 1,74 1,22 2,49 48 2,42% ASXL1_and_SF3B1 0,99 0,73 1,35 85 4,28% ASXL1_and_SRSF2 1,94 1,56 2,40 139 7,00% ASXL1_and_TET2 1,57 1,27 1,94 168 8,45% ASXL1_and_U2AF1 1,59 1,21 2,10 86 4,33% BCOR_and_RUNX1 2,11 1,50 2,96 54 2,72% BCOR_and_STAG2 2,67 1,88 3,81 42 2,11% CBL_and_RUNX1 2,73 1,51 4,97 16 0,81% DNMT3A_and_EZH2 0,98 0,52 1,82 23 1,16% DNMT3A_and_IDH2 0,94 0,45 1,97 18 0,91% DNMT3A_and_SF3B1 0,71 0,54 0,94 130 6,54% DNMT3A_and_SRSF2 2,07 1,28 3,35 22 1,11% DNMT3A_and_TET2 0,98 0,73 1,31 95 4,78% DNMT3A_and_U2AF1 1,12 0,68 1,83 34 1,71% EZH2_and_SF3B1 0,89 0,48 1,67 24 1,21% EZH2_and_TET2 2,00 1,43 2,79 53 2,67% IDH2_and_SRSF2 1,51 1,04 2,20 46 2,32% JAK2_and_SF3B1 1,17 0,56 2,46 14 0,70% JAK2_and_TET2 1,38 0,74 2,57 18 0,91% NRAS_and_RUNX1 2,06 1,11 3,85 14 0,70% NRAS_and_STAG2 2,18 1,23 3,86 17 0,86% PHF6_and_RUNX1 1,85 1,04 3,27 21 1,06% RUNX1_and_STAG2 2,29 1,75 3,01 77 3,88% SF3B1_and_SRSF2 1,21 0,60 2,42 17 0,86% SF3B1_and_TET2 0,66 0,52 0,85 174 8,76% SF3B1_and_TP53 1,26 0,82 1,92 34 1,71% SRSF2_and_TET2 1,22 0,96 1,54 132 6,64% TET2_and_TP53 1,79 1,18 2,71 38 1,91%
[0322]
[0323] SIB - 36 - BW1367R
[0324] TET2_and_U2AF1 1,50 0,97 2,34 36 1,81%
[0325]
[0326] Table 5
[0327] Hazard ratios (HRs) for overall survival were estimated using univariate Cox proportional hazards models. Frequencies (FREQ_MUT) represent the number of patients exhibiting each evolutionary route or mutation over the total number of patients included in the survival analysis.
[0328] The application of a survival analysis model to the recurrent evolutionary genomic covariates, and clinical covariates of patients for which a specific prognosis is known allows to select, among said recurrent evolutionary genomic covariates and clinical covariates, those associated with the said specific prognosis.
[0329] This is particularly advantageous, in fact once identified signature genomic covariates and signature genomic covariates and clinical covariates associated with a specific prognosis, the observation of said signature genomic covariates or signature genomic covariates and clinical covariates in a patient for which the prognosis is not known allows physicians to associated said patient to a prognosis and to plan the proper treatment and follow-up plans accordingly.
[0330] In a preferred embodiment, clinical covariates can be one or more of the clinical covariates above-mentioned; preferably, clinical covariates are one or more selected from age, absolute neutrophil count, platelet count, haemoglobin, bone marrow blasts, presence of del(5q), presence of -7 / del(7q), presence of -17 / del(17p), presence of complex karyoptype, IPSS-M cytogenetics category. In another preferred embodiment, clinical covariates are age, haemoglobin, platelet count, absolute neutrophil count (ANC), bone marrow blasts (%), IPSS-M cytogenetics category, presence of del(5q), presence of -7 / del(7q), presence of -17 / del(17p), presence of complex karyoptype and IPSS-M score.
[0331] In a preferred embodiment, genomic covariates are selected from one or more recurrent evolutionary genomic covariates or alteration alterations observed in oncologic (cancer) patients or in individuals with somatic mutations identified through gene sequencing, irrespective of a formal cancer diagnosis. This embodiment encompasses applications in conditions such as clonal haematopoiesis of indeterminate potential (CHIP) and disorders with dysplastic changes, including certain lesions in the colorectal and oesophageal regions.
[0332] In a preferred embodiment, genomic covariates are one more of the recurrent evolutionary genomic covariates; in a more preferred embodiment, recurrent SIB - 37 - BW1367R
[0333] evolutionary genomic covariates are single mutation and / or co-occurrence of mutation and / or evolutionary trajectories observed in oncologic patients or in individuals with somatic mutations identified through gene sequencing, irrespective of a formal cancer diagnosis. This embodiment encompasses, for example, applications in conditions such as clonal haematopoiesis of indeterminate potential (CHIP) and disorders with dysplastic changes, including colorectal and oesophageal dysplastic / pre-cancerous lesions. In a more preferred embodiment, these recurrent evolutionary genomic covariates include single mutations, co-occurring mutations, and / or evolutionary trajectories. Preferably, within the context of MDS, single mutations are those genes most frequently implicated in MDS or routinely assessed in clinical practice, thereby enhancing the model’s clinical applicability and translational utility. The choice of these particular trajectories is grounded in their high prevalence in MDS and their representation in established clinical and diagnostic panels, ensuring the model’s optimization for prognostic analysis with direct clinical relevance. (G. Hoog et Al. Cancer Genetics, 2023, Sauta et Al. JCO, 2023 Nazha et Al., JCO, 2021 ). The preferred mutations may include one or more of the following genes: ASXL1, ATRX, BCOR, BCORL1, CBL, CEBPA, CUX1, DNMT3A, EP300, ETNK1, ETV6, EZH2, FLT3, GATA2, GNAS, IDH1, IDH2, JAK2, KIT, KRAS, MPL, NF1, NOTCH1, NPM1, NRAS, PHF6, PTPN11, RAD21, RUNX1, SETBP1, SF3B1, SMC1A, SMC3, SRSF2, STAG2, TET2, TP53, U2AF1, WT1, ZRSR2.
[0334] The evolutionary trajectories are preferably identified using the ASCETIC framework applied to selected genomic covariates. In an MDS context, these selected evolutionary trajectories may include, but are not limited to, pathways such as ASXL1 to BCOR, ASXL1 to CBL, ASXL1 to IDH1, ASXL1 to IDH2, ASXL1 to KRAS, ASXL1 to NF1, ASXL1 to PHF6, ASXL1 to RUNX1, ASXL1 to STAG2, DNMT3A to BCOR, EZH2 to RUNX1, EZH2 to STAG2, EZH2 to ZRSR2, SF3B1 to RUNX1, SF3B1 to ZRSR2, SRSF2 to CBL, SRSF2 to IDH1, SRSF2 to IDH2, SRSF2 to NRAS, SRSF2 to RUNX1, SRSF2 to STAG2, TET2 to MPL, TET2 to STAG2, TET2 to ZRSR2.
[0335] Survival analysis model can be, for example, selected from a range of approaches including, but not limited to, Kaplan-Meier estimation, parametric survival models, preferably Weibull, Gompertz, and Generalized Gamma, accelerated failure time (AFT) models, regularized Cox regression with LASSO or Ridge or Elastic Net penalties, Bayesian Cox models, multi-state models, machine learning methods, preferably Random Survival Forests, Survival Gradient Boosting, Survival Support Vector Machines, k-Nearest Neighbors Survival Analysis, Ensemble Survival Models, Bayesian SIB - 38 - BW1367R
[0336] Additive Regression Trees (BART) for Survival Analysis, Cox Proportional Hazards Model with Time-Dependent Covariates, Joint Models for Longitudinal and Survival Data, Recurrent Event Models, preferably, Andersen-Gill Model, Deep Learning-Based Survival Models, preferably, DeepHit, DeepSurv, Flexible Parametric Survival Models, preferably Royston- Parmar Model, Multi-Task Logistic Regression (MTLR), Aalen’s Additive Model; in a preferred embodiment, the survival analysis model is Cox regression; in a more preferred embodiment, the survival analysis model is Cox regression is regularized Cox regression with LASSO penalty; in an even more preferred embodiment, the survival analysis model is regularized Cox regression with LASSO penalty is carried out multiple times (preferably >= 100) with bootstrap resampling technique.
[0337] The survival analysis model furnishes a hazard ratio, or in other words a effect measure, for each recurrent evolutionary genomic covariates and clinical covariates analyzed; preferably, genomic covariates and clinical covariates analysed are considered signature genomic covariates or signature genomic covariates and clinical covariates if their hazard ratios is greater than zero in at least 10% of iteration on bootstrapped datasets; more preferably, if their hazard ratio is greater than zero in at least 50% of iteration on bootstrapped datasets; even more preferably, if their hazard ratios is greater than zero in at least 30% of iteration on bootstrapped datasets. In a still preferred embodiment, signature genomic covariates or signature genomic covariates and clinical covariates if their hazard ratios is greater than zero in at least 70% of iteration on bootstrapped datasets.
[0338] Regularized Cox regression with LASSO penalty combines Cox proportional hazard model with LASSO (Least Absolute Shrinkage and Selection Operator) penalty; said penalty, which impose sparsity, reduce the number of relevant variables in the model, advantageously preventing overfitting. In particular, said penalty forces less significant variable coefficients to zero, indeed removing them from the model. As a result, model returns a hazard ratio (also called “HR”) for each significant covariates, such as the signature ones, which determines the impact of that covariate on the prognosis; a hazard ratio greater than 1 indicates an increased relative risk of an event, while a hazard ratio less than 1 indicates a reduction in risk.
[0339] Regularized Cox regression with LASSO penalty is carried out multiple times with bootstrap resampling technique; preferably, in step 3) independent iterations of SIB - 39 - BW1367R
[0340] regularized Cox regression with LASSO penalty is carried out on 100 datasets generated with bootstrap resampling of dataset of step 1).
[0341] The procedure takes account for each patient clinical covariates and recurrent evolutionary genomic covariates; clinical covariates can be continuous values, categorical values or binary values; for bulk sequencing (single or multi-region) a gene is considered mutated in a patient if there is at least one mutation (SNV or indel) in a region covered by such gene with a CCF greater than threshold used in pipeline of variant calling (usually > 5%). For single cell sequencing, genes are considered mutate in a patient if there is at least a mutation in a region in at least a certain percentage of cells (default=10%).
[0342] Covariates are selected if their HR differs from 0 in a certain percentage of the iteration on bootstrapped datasets; preferably, in at least 30% of the iteration on bootstrapped datasets. Said threshold can be modified and lower threshold increases the probability of selection of a greater number of signature covariates, as a lower consistency between iteration is required, while a greater threshold ensures a stricter association between signature covariate and a specific prognosis, but it reduced the number of covariates. Signature genomic covariates are one or more from ASXL1 to KRAS, JAK2 mutations, NPM1 mutations, SF3B1 to RUNX1, TP53 CCF and TET2 to STAG2; preferably; signature genomic covariates associated with leukaemia- free survival are one or more from ASXL1 to KRAS, JAK2 mutations, NPM1 mutations, SF3B1 to RUNX1, TP53 CCF and TET2 to STAG2; more preferably, signature genomic covariates associated with leukaemia- free survival of a myelodysplastic syndrome patients are one or more from ASXL1 to KRAS, JAK2 mutations, NPM1 mutations, SF3B1 to RUNX1, TP53 CCF and TET2 to STAG2. Signature genomic covariates and clinical covariates are ASXL1 to KRAS, JAK2 mutations, NPM1 mutations, SF3B1 to RUNX1, TP53 CCF, TET2 to STAG2 and age; preferably; signature genomic covariates associated with leukaemia-free survival are one or more from ASXL1 to KRAS, JAK2 mutations, NPM1 mutations, SF3B1 to RUNX1, TP53 CCF. TET2 to STAG2 and age; more preferably, signature genomic covariates associated with leukaemia-free survival of a myelodysplastic syndrome patients are one or more from ASXL1 to KRAS, JAK2 mutations, NPM1 mutations, SF3B1 to RUNX1, TP53 CCF, TET2 to STAG2 and age.
[0343] In a preferred embodiment, to select the genomic and clinical covariates associated to LFS of MDS patients, the survival method applied to the recurrent genomic covariates obtained from step 2) and the clinical covariates is LASSO-penalized Cox regression SIB -40 - BW1367R
[0344] with 100 bootstrap iterations. With reference to said preferred embodiment, covariates were retained if they had a non-zero hazard ratio in >70% of iterations; This yielded 31 signature genomic and clinical covariates associate to a specific prognosis (i.e. LFS): 26 genomic features, including evolutionary routes, such as ASXL1_to_KRAS, Root_to_ATRX, NRAS_and_RUNX1, SRSF2_to_NRAS, TET2_and_U2AF1, Root_to_JAK2, ASXL1_and_DNMT3A, TET2_to_STAG2, Root_to_ZRSR2, Root_to_NPM1, Root_to_ASXL1, SRSF2_to_CBL, ASXL1_to_BCOR, SF3B1_to_RUNX1, BCOR_and_STAG2, EZH2_to_RUNX1, TET2_to_MPL Root_to_U2AF1, Root_to_RAD21, DNMT3A_and_SRSF2, Root_to_EP300, Root_to_ETV6, CBL_and_RUNX1, DNMT3A_and_U2AF1, Root_to_IDH2, SRSF2_to_RUNX1 and 5 clinical covariates IPSSM score, age, ANC, del5q, TP 53 CCF. By leveraging the inherent property of LASSO regression to select one predictor among correlated variables while shrinking others to zero, this approach aimed to achieve a parsimonious set of additive covariates for integration into the IPSS-M model. LASSO-penalized Cox regression was implemented using the glmnet package in R.
[0345] Calculation of hazard ratio for signature genomic covariates or signature genomic and clinical covariates
[0346] The hazard ratio of signature genomic covariates or signature genomic and clinical covariates is calculated applying survival analysis model to signature covariates obtained from step 3). Optionally, additional clinical and / or genomic covariates can be analysed with the survival analysis model in combination with signature genomic covariates or signature genomic and clinical covariates; preferably, the additional covariate is IPSS-M risk-score.
[0347] The survival analysis model of step 3) and step 4) can be the same or different survival analysis model; in a preferred embodiment, the survival analysis model are the same; in a more preferred embodiment, the survival analysis model is Cox regression; in an even more preferred embodiment, the survival analysis model is Cox regression is regularized Cox regression with LASSO penalty. In another embodiments, the survival analysis models of step 3) and 4) are different models; preferably, in step 3) the survival analysis model is regularized Cox regression with LASSO penalty and in step 4) is cox regression, such as a standard multivariate Cox regression.
[0348] According to certain embodiments, signature genomic and clinical covariates, provided by step 4) are one or more selected from IPSS-M score, ASXL1 to KRAS, JAK2 mutations, NPM1 mutations SF3B1 to RUNX1, TP53 CCF, Age, TET2 to STAG2, SIB - 41 - BW1367R
[0349] NRAS_and_RUNX1, Root_to_ATRX, Root_to_JAK2, SRSF2_to_NRAS; preferably PSS-M score, ASXL1 to KRAS, JAK2 mutations, Age.
[0350] According to a preferred embodiment, step 4) provides the covariates and the relative hazard ration reported in Table 1 and 2:
[0351] Covariates Hazard ratio
[0352] I PSS-M score 0.591
[0353] ASXL1 to KRAS 0.247
[0354] JAK2 mutations 0.143
[0355] NPM1 mutations -0.077
[0356] SF3B1 to RUNX1 0.025
[0357] TP53 CCF 0.01
[0358] Age 0.016
[0359] TET2 to STAG2 0.006
[0360]
[0361] Table 6
[0362] Covariates HAZARD RATIO (95% CI) COEFFICIENT IPSSM_SCORE 1.8571 (1.7474 -1.974) 0,61902
[0363] AGE 1.0208 (1.0147 - 1.027) 0,02 ASXL1_to_KRAS 1.7436 (1.0138 - 2.999) 0,55595 NRAS_and_RUNX1 0.4470 (0.2225 - 0.898) -0,80529 Root_to_ATRX 1.5037 (1.0230 - 2.210) 0,40791 Root_to_JAK2 1.5870 (1.1065 - 2.276) 0,46187 SRSF2_to_NRAS 1.9213 (1.0438 - 3.536) 0,65300
[0364]
[0365] Table 7
[0366] In table 7 “Coefficient” represent the regression weights of the Cox regression, while the Hazard ratio are the exponential function applied to such weights.
[0367] The Applicant noted several advantages associated with the present invention: first of all, the computer implemented method of the present invention allows to identify the evolutionary trajectories and to associate them to a specific prognosis. This technical effect is dramatically significant cause it allows to predict the outcome of a carcer or disease precursor of cancer by detecting evolutionary trajectories of mutation which can occur even in an early / intermediate early phase of the evolution of said disease, enabling physicians to define the proper therapeutic plant and to set up the follow-up. Moreover, surprisingly, even if the dataset employed for the development of I PSS-M SIB - 42 - BW1367R
[0368] method has been used, the computer implemented invention of the present invention has identified not only evolutionary trajectories, but also significant single mutation significantly associated with leukaemia- free survival, neglected by IPSS-M, as JAK2 mutations, NPM1 mutations, TP53 mutations. Furthermore, without bound the theory, in using a different dataset of patients suffering from the same disease, the computer implemented method could allow to identify additional clinical and / or genomic covariates associated with a specific prognosis.
[0369] The computer implemented method for the calculation of hazard ratio of signature genomic covariates or signature genomic and clinical covariates associated with a prognosis in a cancer patient, can further comprise at least the following steps:
[0370] 5) calculating the risk-score for one or more cancer patients using said signature genomic covariates or signature genomic and clinical covariates associated with a prognosis of step 3) corrected for their relative hazard ratio obtained from step 4); optionally
[0371] 6) clustering patient with a clustering algorithm on the basis of calculated risk-score of step.
[0372] In a preferred embodiment, the cancer patient can be a cancer patient that belongs to the dataset of step 1) or a cancer patient that does not belong to the dataset of step 1); preferably, the cancer patient is a patient suffering from a disease precursor of cancer; more preferably, the cancer patient suffers from myelodysplastic syndrome; more preferably, the cancer patient suffers from myelodysplastic syndrome and the prognosis associated with signature covariates is leukaemia-free survival. In certain embodiments, the cancer patient of step 5) can be a cancer patient that belongs to the dataset of step 1) or a cancer patient that does not belong to the dataset of step 1); preferably, the cancer patient is a patient suffering from a disease precursor of cancer; more preferably, the cancer patient suffers from myelodysplastic syndrome; more preferably, the cancer patient suffers from myelodysplastic syndrome and the prognosis associated with signature covariates is leukaemia- free survival.
[0373] In certain embodiments, the invention refers to a computer implemented method for the calculation of hazard ratio of signature genomic covariates or signature genomic and clinical covariates associated with a prognosis individuals with somatic mutations identified through gene sequencing, irrespective of a formal cancer diagnosis, further comprising following steps 5) and 6) SIB - 43 - BW1367R
[0374] The risk-score is calculated as the sum of the hazard ratios of signature genomic covariates or signature genomic, signature evolutionary routes, and / or clinical covariates associated with a patient outcome, which may include but are not limited to prognosis, response, or lack of response to therapy. Each covariate is assigned a numerical coefficient, herein referred to as a “value.” The term “value” as used herein is broadly defined to encompass a continuous numerical coefficient, such as effect measures, assigned to each selected covariate, which may include, but is not limited to, hazard ratios (HR), odds ratios (OR), coefficients, frequency percentages, clustering metrics or other statistical measures that capture the covariate’s impact on the outcome; therefore, in anyone of the embodiments herein described, effects measures are hazard ratios (HR), odds ratios (OR), coefficients, frequency percentages, clustering metrics or other statistical measures that capture the covariate’s impact on the outcome. These values are used to weight the contributions of each covariate in the Risk-score, thereby reflecting their relative influence on patient outcomes. In a preferred embodiment, the Risk score is computed as a weighted linear combination of Cox regression weights (beta) obtained from step 4) according to the following equation:
[0375] Risk-score = (Pi * xi) for I = 1 to n
[0376] where:
[0377] • Pi denotes the coefficient (log hazard ratio) associated with the i-th feature, as estimated from the multivariate Cox regression,
[0378] • Xi is the binary status (0 or 1) of the i-th feature for the patient,
[0379] • n = corresponds to the total number of signature covariates.
[0380] In a preferred embodiment, the Risk-score calculation incorporates the IPSS-M score as one of the clinical covariates. In a more preferred embodiment, the value of each covariate is represented as a normalized hazard ratio derived from a Cox proportional hazards regression analysis, with covariates pre-selected using LASSO regression to optimize relevance and predictive power. This approach enables the effective integration of genomic, evolutionary, and clinical information into a unified prognostic score, enhancing the model’s clinical utility and applicability across various patient outcomes. Therefore, in a preferred embodiment, the risk score of MDS patients can be calculated as, computed as a weighted linear combination of the regression weight of the seven covariates retained in the final multivariable Cox model:
[0381] IPSS-M-Evo Score = (Pi * x) for I = 1 to n where: SIB - 44 - BW1367R
[0382] • denotes the coefficient (log hazard ratio) associated with the i-th f
[0383]
[0384] eature, as estimated from the multivariate Cox regression, • Xi is the binary status (0 or 1) of the i-th feature for the patient,
[0385] • n = 7 corresponds to the total number of retained covariates.
[0386] Clustering algorithms can be, for example, k-means clustering, density-based spatial clustering, gaussian matrix models, self-organizing maps, fuzzy c-means clustering; preferably, clustering algorithms is k-means algorithm. In particular, such algorithm groups patient in two or more clusters, preferably in 6 clusters. In contrast to the methodology used in IPSS-M, which involved expert-driven categorizations and manual adjustments for re-clustering, k-means clustering was utilized. This approach offers a systematic and data-driven method for re-clustering, advantageously increasing the reproducibility of results across different datasets.
[0387] Advantageously, if the patient suffers from a disease for which hazard ratio of signature covariates have been already calculated with the computer implemented method object of the first aspect of the invention, even if said patient does not belongs to the dataset used for the calculation of said hazard ratio, the risk factor for said patient can be calculated and the patient stratified only carrying out steps 5) and 6).
[0388] In a preferred embodiment, the computer implemented method can further comprise a step of generation of graphical representation of results.
[0389] The Hazard ratios of signature genomic covariates or signature genomic and clinical covariates calculated with the computer implemented method, object of the first aspect of the invention, represent a second aspect of the invention. In a preferred embodiment, signature genomic covariates are one or more from ASXL1 to KRAS, JAK2 mutations, NPM1 mutations, SF3B1 to RUNX1, TP53 CCF and TET2 to STAG2; preferably; signature genomic covariates associated with leukaemia- free survival are one or more from ASXL1 to KRAS, JAK2 mutations, NPM1 mutations, SF3B1 to RUNX1, TP53 CCF and TET2 to STAG2; more preferably, signature genomic covariates associated with leukaemia- free survival of a myelodysplastic syndrome patients are one or more from ASXL1 to KRAS, JAK2 mutations, NPM1 mutations, SF3B1 to RUNX1, TP53 CCF and TET2 to STAG2. Signature genomic covariates and clinical covariates are ASXL1 to KRAS, JAK2 mutations, NPM1 mutations, SF3B1 to RUNX1, TP53 CCF, TET2 to STAG2 and age; preferably; signature genomic covariates associated with leukaemia-free survival are one or more from ASXL1 to KRAS, JAK2 mutations, NPM1 mutations, SF3B1 to RUNX1, TP53 CCF. TET2 to STAG2 and age; more preferably, signature SIB -45 - BW1367R
[0390] genomic covariates associated with leukaemia-free survival of a myelodysplastic syndrome patients are one or more from ASXL1 to KRAS, JAK2 mutations, NPM1 mutations, SF3B1 to RUNX1, TP53 CCF, TET2 to STAG2 and age. In other preferred embodiment, signature genomic covariates associated with leukaemia-free survival of a myelodysplastic syndrome patients are one or more IPSSM-score, age, ASXL1 to KRAS, NRAS and RUNX1, root to ATRX, root JAKW and SRSF2 to NRAS. In a third aspect, the present invention relates to the use of the hazard ratio calculated according to the computer implemented method object of the first aspect of the invention in a risk stratification method of cancer patients or in individuals with somatic mutations identified through gene sequencing. In a preferred embodiment, hazard ratio is the hazard ratios of signature genomic covariates or signature genomic and clinical covariates associated with a specific patient outcome, which may include, but is not limited to, prognosis, response, or lack of response to therapy. In a preferred embodiment, signature genomic covariates are one or more from ASXL1 to KRAS, JAK2 mutations, NPM1 mutations, SF3B1 to RUNX1, TP53 CCF and TET2 to STAG2; preferably; signature genomic covariates associated with leukaemia-free survival are one or more from ASXL1 to KRAS, JAK2 mutations, NPM1 mutations, SF3B1 to RUNX1, TP53 CCF and TET2 to STAG2; more preferably, signature genomic covariates associated with leukaemia-free survival of a myelodysplastic syndrome patients are one or more from ASXL1 to KRAS, JAK2 mutations, NPM1 mutations, SF3B1 to RUNX1, TP53 CCF and TET2 to STAG2. Signature genomic covariates and clinical covariates are ASXL1 to KRAS, JAK2 mutations, NPM1 mutations, SF3B1 to RUNX1, TP53 CCF, TET2 to STAG2 and age; preferably; signature genomic covariates associated with leukaemia-free survival are one or more from ASXL1 to KRAS, JAK2 mutations, NPM1 mutations, SF3B1 to RUNX1, TP53 CCF. TET2 to STAG2 and age; more preferably, signature genomic covariates associated with leukaemia-free survival of a myelodysplastic syndrome patients are one or more from ASXL1 to KRAS, JAK2 mutations, NPM1 mutations, SF3B1 to RUNX1, TP53 CCF, TET2 to STAG2 and age. In other preferred embodiment, signature genomic covariates associated with leukaemia-free survival of a myelodysplastic syndrome patients are one or more IPSSM-score, age, ASXL1 to KRAS, NRAS and RUNX1, root to ATRX, root JAKW and SRSF2 to NRAS.
[0391] In a fourth aspect, the present invention relates to the use of signature genomic covariates, evolutionary routes or signature genomic and clinical covariates associated with a specific patient outcome, which may include, but is not limited to, prognosis, SIB -46 - BW1367R
[0392] response, or lack of response to therapy, in risk stratification method of a cancer or patient with established somatic genetic mutations. In a preferred embodiment, the cancer patient suffers from myelodysplastic syndrome; in a preferred embodiment, signature genomic covariates can be at least one selected from the group of ASXL1 to KRAS, JAK2 mutations, NPM1 mutations, SF3B1 to RUNX1, TP53 CCF, and TET2 to STAG2; in a more preferred embodiment, signature genomic covariate and clinical covariates can be at least one selected from the group of ASXL1 to KRAS, JAK2 mutations, NPM1 mutations, SF3B1 to RUNX1, TP53 CCF, and TET2 to STAG2 and age. In a preferred embodiment, signature genomic covariate and clinical covariate are one or more of IPSSM-score, age, ASXL1 to KRAS, NRAS and RUNX1, root to ATRX, root JAKW and SRSF2 to NRAS.
[0393] In a preferred embodiment, signature genomic covariates or signature genomic and clinical covariates associated with a specific prognosis are used in a risk stratification method of a cancer patient in combination with one or more additional genomic and / or clinical covariates; preferably, the additional covariate is IPSS-M risk-score.
[0394] In a preferred embodiment, prognosis can be any one of the prognoses defined in the first aspect of the invention; preferably, prognosis is leukaemia-free survival or overall survival.
[0395] In a fifth aspect, the present invention provides a computer implemented method for riskstratification of a cancer patient or in individuals with somatic mutations identified through gene sequencing comprising a step of calculating the risk-stratification of said patient using the hazard ratios calculated by the computer implemented method according to the first aspect of the invention.
[0396] All definition given in the first aspect of the invention is included in the fifth aspect; in particular, the definition of prognosis, cancer patient, clinical covariates, genomic covariates, evolutionary trajectory, recurrent evolutionary genomic covariates, signature genomic covariates, signature genomic covariates and clinical covariates, gene sequencing.
[0397] Therefore, according to a preferred embodiment, refers to risk stratification method of cancer patients or individuals with somatic mutations identified through gene sequencing, irrespective of a formal cancer diagnosis comprising the following step: 1) providing a dataset of genomic covariates and prognosis data or a dataset of genomic and clinical covariates and prognosis data of a statistically significant number SIB - 47 - BW1367R
[0398] of cancer patients or individuals with somatic mutations identified through gene sequencing, irrespective of a formal cancer diagnosis;
[0399] 2) applying an evolutionary model to the dataset of step 1) for identifying statistically recurrent evolutionary genomic covariates according to said evolutionary method;
[0400] 3) processing the identified statistically recurrent evolutionary genomic covariates of step 2) or statistically recurrent evolutionary genomic covariates of step 2) and clinical covariates using a survival analysis model for selecting signature genomic covariates or signature genomic and clinical covariates related to a specific prognosis;
[0401] 4) processing the signature genomic covariates or signature genomic and clinical covariates selected by step 3) by a survival analysis model for calculating the effect measures, preferably the hazard ratio, thereof.
[0402] 5) calculating the risk-score for one or more cancer patients using signature genomic and clinical covariates associated with a prognosis of step 3) corrected for their relative effect measures, preferably hazard ratio, obtained from step 4); and optionally
[0403] 6) clustering patient with a clustering algorithm on the basis of calculated risk-score of step;
[0404] wherein said recurrent evolutionary genomic covariates include single gene mutations, co-occurrence of mutations and / or evolutionary routes.
[0405] Preferably, said risk stratification method is computer implemented.
[0406] Preferably, in risk stratification method the evolutionary model comprises one or more of, preferably at least, the following phases:
[0407] computing the Cancer Cell Fraction (CCF) of each somatic mutation or gene mutation from its Variant Allele Frequency (VAF), wherein said computation is adapted to account for copy number alterations (CNA) affecting the corresponding genomic locus, thereby correcting for allelic imbalance and enabling accurate temporal inference; calculating a Cross-Validation Rank (CV-Rank) and a Cross-Validation Score (CV-Score) for each genomic mutation and evolutionary trajectory
[0408] selecting of evolutionary trajectory having an EVO-score greater than 0,5, more preferably occurring in at least 10 patients of the dataset and having an EVO-score greater than 0,5.
[0409] In a still preferred embodiment, the invention refers to risk stratification method of MDS patients comprising the following step:
[0410] 1) providing a dataset of genomic and clinical covariates and prognosis data of a statistically significant number of MDS patients; SIB -48 - BW1367R
[0411] 2) applying an evolutionary model to the dataset of step 1) for identifying statistically recurrent evolutionary genomic covariates according to said evolutionary method;
[0412] 3) processing the identified statistically recurrent evolutionary genomic covariates of step 2) or statistically recurrent evolutionary genomic covariates of step 2) and clinical covariates using regularized Cox regression with LASSO penalty for selecting signature genomic covariates or signature genomic and clinical covariates related to a specific prognosis;
[0413] 4) processing the signature genomic covariates or signature genomic and clinical covariates selected by step 3) by Cox regression for calculating the hazard ratio thereof; 5) calculating the risk-score for one or more cancer patients using said signature genomic covariates or signature genomic and clinical covariates associated with a prognosis of step 3) corrected for their relative hazard ratio obtained from step 4), wherein said cancer patient, preferably MDS patient, belongs to the dataset of step A) or does not belong to the dataset of step A); and optionally
[0414] 6) clustering patient with a clustering algorithm on the basis of calculated risk-score of step;
[0415] wherein said recurrent evolutionary genomic covariates include single gene mutations, co-occurrence of mutations and / or evolutionary routes; and wherein clinical covariates comprise at least one or more selected from the group consisting of age, hemoglobin, platelet count, absolute neutrophil count, bone marrow blasts, IPSS-M cytogenetics category, presence of del(5q), presence of -7 / del(7q), presence of -17 / del(17p), presence of complex karyoptype.
[0416] In certain embodiments of the risk stratification method herein described, resamplingbased refinement procedure is applied.
[0417] In a preferred embodiment, the risk-stratification method further comprises a physical step of assessment of the presence or absence of signature genomic covariates or signature genomic covariates and clinical covariates for said cancer patient; preferably, the assessment of the presence or absence of genomic covariate is carried out by gene sequencing of a tumour sample of said patient; more preferably, sequencing is singlecell sequencing, bulk sequencing and multiregional bulk sequencing. In still preferred embodiment, the risk-stratification method object of the invention further comprises a physical step of assessment of the presence or absence of signature genomic covariates of cancer patient of step 5), when said cancer patient of step 5) does not belong to the dataset of step A); preferably, the assessment of the presence or absence of genomic SIB -49 - BW1367R
[0418] covariate is carried out by DNA sequencing of a tumour sample of said patient; more preferably, sequencing is single-cell sequencing, bulk sequencing and multiregional bulk sequencing
[0419] In a preferred embodiment, prognosis can be any one of the prognosis define in the first aspect of the invention; preferably, prognosis is leukaemia-free survival or overall survival.
[0420] In a preferred embodiment, cancer patient can be any one of the patients defined in the first aspect of the invention; preferably, cancer patient is a myelodysplastic syndrome patient.
[0421] In a preferred embodiment, hazard ratio is the hazard ratios of signature genomic covariates or signature genomic and clinical covariates associated with a specific prognosis. In a preferred embodiment, signature genomic covariates are one or more from ASXL1 to KRAS, JAK2 mutations, NPM1 mutations, SF3B1 to RUNX1, TP53 CCF and TET2 to STAG2; preferably; signature genomic covariates associated with leukemia-free survival are one or more from ASXL1 to KRAS, JAK2 mutations, NPM1 mutations, SF3B1 to RUNX1, TP53 CCF and TET2 to STAG2; more preferably, signature genomic covariates associated with leukaemia-free survival of a myelodysplastic syndrome patients are one or more from ASXL1 to KRAS, JAK2 mutations, NPM1 mutations, SF3B1 to RUNX1, TP53 CCF and TET2 to STAG2. Signature genomic covariates and clinical covariates are ASXL1 to KRAS, JAK2 mutations, NPM1 mutations, SF3B1 to RUNX1, TP53 CCF, TET2 to STAG2 and age; preferably; signature genomic covariates associated with leukaemia-free survival are one or more from ASXL1 to KRAS, JAK2 mutations, NPM1 mutations, SF3B1 to RUNX1, TP53 CCF. TET2 to STAG2 and age; more preferably, signature genomic covariates associated with leukaemia-free survival of a myelodysplastic syndrome patients are one or more from ASXL1 to KRAS, JAK2 mutations, NPM1 mutations, SF3B1 to RUNX1, TP53 CCF, TET2 to STAG2 and age. In a preferred embodiment, signature genomic covariate and clinical covariate are one or more of IPSSM-score, age, ASXL1 to KRAS, NRAS and RUNX1, root to ATRX, root JAKW and SRSF2 to NRAS.
[0422] In a preferred embodiment, hazard ratio of signature genomic covariates or signature genomic and clinical covariates associated with a specific prognosis are used in the computer implemented method for risk stratification method of a cancer patient in combination with one or more additional genomic and / or clinical covariates; preferably, the additional covariate is IPSS-M risk-score. SIB - 50 - BW1367R
[0423] In a preferred embodiment, the computer implemented method for risk stratification method of a cancer patient stratify said patient in a risk stratification group; preferably, preferably, the six groups are very low (also called “VL”), low (also called “L”), moderate-low (also called “ML”), moderate-high (also called “MH”), high (also called “high”) and very high (also called “VH”); more preferably; Very-Low (VL) score < 0.35, Low (L) score > 0.36 and < 0.90, Moderate-Low (ML) score > 0.91 and < 1.47, Moderate-High (MH) score > 1.48 and < 2.10, High (H) score > 2.11 and < 2.84 and Very High (VH) score > 2.85; even more preferably, the six categories are the same categories of IPSS-M. In certain embodiments, the risk stratification groups are the following: Very-Low (VL) score < 0.5, Low (L) score 0.50 to 1.00 (included), Moderate-Low (ML) score 1.00 to 1.50 (included), Moderate-High (MH) score 1.50 to 2.00 (included), High (H) score 2.00 to 3.00 (included) and Very High (VH) score >3.00.
[0424] Surprisingly, the Applicant noted that through the application of the computer implemented method for risk stratification method of a cancer to the same dataset of patient used for IPSS-M, from the comparison of the results of the two different methods results that 48% of patient were stratified in a different risk group by the computer implemented method of the present invention, dived in 43% reclassified into a lower risk group compared to the one of IPSS-M, while 5% in an higher one; as, the computer implemented invention consistently demonstrated improved C-index values compared to IPSS-M the re-clustered patients, that re-clustering was performed correctly. Moreover, although the feature selection was performed using LFS and the re-clustering, such as the re-stratification, of patients is based on these features, the patients are also correctly re-clustered, such as re-stratified, with respect to OS. This suggests that the features selected for LFS have a robust prognostic value, which is also reflected in their ability to discriminate overall survival. Specifically, the re-stratification assesses whether a patient has been assigned to a different risk class compared to the one defined by IPSS-M. This process is therefore patient-specific. The features used for re-stratification were selected based on their association with LFS. However, once a patient has been reassigned to a new risk class, it is possible to evaluate whether this new classification also improves the predictability of OS. Furthermore, the methods herein described, having an evolutionbased approach, may find application in early detection strategies, risk-adapted surveillance, and potentially even therapeutic anticipation. In this context, identifying the most likely downstream event and its associated risk in each patient could support SIB - 51 - BW1367R
[0425] individualized decisions regarding follow-up intensity or the timing of therapeutic intervention.
[0426] In a sixth aspect, the invention provides a computer program comprising instructions which, when the program is executed by a computer, cause the computer to carry out the computer implemented methods object of the first and fifth aspects of the invention. In a seventh aspect, the invention provides a computer-readable data carrier having stored there on the computer program of the sixth aspect of the invention.
[0427] Experimental section
[0428] Statistical Methods
[0429] To assess differences in variable distributions between two subgroups, the Chi-square test for categorical variables and the T-Test or Mann-Whitney test for continuous variables were employed, depending on the data distribution. For comparisons involving more than two subgroups, a mixed-model test was used for continuous variables. For the analysis of OS and LFS, time-to-event was measured from the time of sample collection to the occurrence of the event of interest. OS and LFS probabilities were estimated using the Kaplan-Meier method. Multivariable analysis was performed using Cox regularized regression. Model comparisons were conducted with Cox regression analysis, and differences between models were evaluated through Harrell’s Concordance Index (C-index) and the Akaike Information Criterion (AIC). A second round of c-index estimation was performed with a bootstrap analysis and differences between the median of inferred c-index were performed. Significant differences in the C-index were tested to determine statistical significance. A significance level of two-sided p-values <0.05 was used for all statistical tests. All analyses were conducted using the R statistical platform, version 4.3.1. Graphing of results was performed using the R statistical platform and Plotly version 5.23.0.
[0430] Example 1 – Application of the computer implemented method of the invention to MDS patients
[0431] Example 1.1 – Computer implemented method step 1)
[0432] The Molecular International Prognostic Scoring System for Myelodysplastic Syndromes Dataset from the IPSS-M 2022 publication (Bernard E et Al; Molecular International Prognostic Scoring System for Myelodysplastic Syndromes. NEJM Evid. 2022 Jul;1(7): EVIDoa2200008. doi: 10.1056 / EVIDoa2200008. Epub 2022 Jun 12. PMID: 38319256. / https: / / github.com / papaemmelab / IPSSMstudy and at https: / / www.cbioportal.org / .) has been employed as dataset of step 1). The data were SIB - 52 - BW1367R
[0433] downloaded from cBioPortal and consisted of 2,519 MDS patients, comprising sequencing information for a panel of 152 profiled mutations. All patients were older than 18 years, with a MDS diagnosis according to the 2016- WHO criteria. The training dataset included 234 patients with secondary / therapy-related MDS (s / t-MDS). All patients were considered for the evolutionary signature analysis of step 3.
[0434] Example 1.2 – Computer implemented method step 2)
[0435] For the association between evolutionary features and OS and Leukemia-Free Survival (LFS) a subset of the training dataset of 2,158 patients with available survival data and age <=90 years were employed.
[0436] Example 1.3 – Computer implemented method step 2) (Fontana, D. et al.; Evolutionary signatures of human cancers revealed via genomic analysis of over 35,000 patients. Nat Commun 14, 5982 (2023). https: / / doi. Org / 10.1038 / s41467-023-41670-3)
[0437] 1.3.1 - Temporal Modelling
[0438] To assess the timing of driver mutations, we selected 47 genes sequenced to validate potential evolutionary routes; 7,211 mutations were analyzed, representing 66.9% of all identified mutations. The Rank-Score, a continuous variable calculated within the dataset, indicates the predicted timing of a mutation: a lower Rank-Score corresponds to an earlier occurrence of the mutation during MDS progression (Figure 1A). Mutations were categorized by their timing into three groups: early (first agony rank), intermediate (second agony rank), and late (third agony rank)
[0439] Early and intermediate-early mutations predominantly involved genes associated with DNA methylation, RNA splicing, and transcription co-factors involved in histone acetyltransferase activity.
[0440] The early gene subset was further divided into two groups: a first subset with very low Rank-Scores (<0.29), including genes such as NOTCH1, PRPF40B, KMT2D, CREBBP, EP300, and ATRX and a second subset with slightly higher Rank-Scores (0.92-0.99), comprising CSF3R, BCORL1, JAK2, GATA2, and KIT. Although signaling genes are typically late mutations according to ProgEvo, JAK2, which constituted 1% of all analyzed mutations, had a Rank-Score of 0.96, indicating its involvement in the early phases of MDS.
[0441] Intermediate-early mutations are composed of genes frequently mutated in MDS, collectively accounting for 58% of all analyzed mutations. These mutations exhibited a narrow Rank-Score range (1.0 to 1.03). This group included key mutations in genes such as SRSF2, EZH2, TET2, U2AF1, TP53, DNMT3A, SF3B1, and ASXL1. The identification SIB - 53 - BW1367R
[0442] of these mutations in the intermediate-early phase, alongside their narrow Rank-Score range, aligns with their association with clonal haematopoiesis of indeterminate potential (CHIP) and suggests that they may form the backbone of the initial driver mutations in MDS, potentially leading to clonal dominance and subsequent secondary mutations. Gene Rank-Score Gene Rank-Score KMT2D 0,00 NPM1 1,19
[0443] NOTCH 1 0,00 IDH2 1,39
[0444] PRPF40B 0,00 ETV6 1,51
[0445] CREBBP 0,01 ZRSR2 1,76
[0446] EP300 0,02 CEBPA 1,86
[0447] ATRX 0,25 IDH1 1,87
[0448] KIT 0,29 GNAS 1,92
[0449] CSF3R 0,92 BCOR 1,97
[0450] BCORL1 0,94 NF1 1,97
[0451] JAK2 0,96 ETNK1 1,99
[0452] GATA2 0,99 MPL 1,99
[0453] CUX1 1,00 RAD21 1,99
[0454] DNMT3A 1,00 SMC3 1,99
[0455] EZH2 1,00 CBL 2,00
[0456] SF3B1 1,00 GNB1 2,00
[0457] SRSF2 1,00 KRAS 2,00
[0458] TET2 1,00 PHF6 2,00
[0459] TP53 1,00 PPM1D 2,00
[0460] U2AF1 1,00 PTPN11 2,00
[0461] WT1 1,00 RUNX1 2,00
[0462] SETBP1 1,01 SMC1A 2,00
[0463] ASXL1 1,03 STAG2 2,00
[0464] FLT3 1,16 NRAS 2,03
[0465]
[0466] Table 8
[0467] Intermediate-late mutations included transcription factors and signaling genes, show a broader Rank-Score range (1.16 to 1.97), reflecting variability in the timing of these mutations during the disease course. Genes in this group included FLT3, NPM1, IDH2, ETV6, ZRSR2, CEBPA, IDH1, GNAS, BCOR, and NF1. SIB - 54 - BW1367R
[0468] Late mutations were enriched for signalling and chromosome cohesion genes, with a narrow Rank-Score range (1.99 to 2.03). Genes present at this stage included MPL, ETNK1, SMC3, RAD21, PTPN11, KRAS, SMC1A, STAG2, GNB1, PHF6, RUNX1, CBL, PPM1D, and NRAS.
[0469] In conclusion, the similarity in Rank-Scores and the narrow Rs-Rank distribution within both early-intermediate and late gene categories suggest that these mutations cluster at critical timepoints in MDS evolution, with early mutations likely initiating the disease and later mutations contributing to its progression.
[0470] 1.3.2 - Statistical Inference and Model Selection
[0471] 1,625 co-occurring gene pairs, grouped into 46 distinct evolutionary routes, have been identified. These routes displayed convincing directional relationships (Evo-Score > 0.65) and statistical support in associating the parent with the child genes (Cv-Score > 0.4).
[0472] Parent Rank-Score Parent rank Child RankChild rank Evo- mutation parent group gene score child group Score CREBBP 0,01 early SF3B1 1,00 intermediat 0,6923 e early 1 KMT2D 0,00 early U2AF1 1,00 intermediat 0,6666 e early 7 KMT2D 0,00 early SF3B1 1,00 intermediat 0,8125 e early 0 NOTCH 1 0,00 early SRSF 1,00 intermediat 1,0000
[0473] 2 e early 0 NOTCH 1 0,00 early SETB 1,01 intermediat 0,8888
[0474] P1 e early 9 NOTCH 1 0,00 early DNMT 1,00 intermediat 0,7272
[0475] 3A e early 7 KMT2D 0,00 early FLT3 1,16 intermediat 0,8333 e late 3 CREBBP 0,01 early PPM1 2,00 late 1,0000
[0476] D 0 NOTCH 1 0,00 early SMC3 1,99 late 1,0000
[0477] 0
[0478]
[0479] SIB - 55 - BW1367R
[0480] ASXL1 1,03 intermediat RUNX 2,00 late 0,6813 e early 1 2 ASXL1 1,03 intermediat STAG 2,00 late 0,7116 e early 2 6 SRSF2 1,00 intermediat RUNX 2,00 late 0,7459 e early 1 0 SRSF2 1,00 intermediat STAG 2,00 late 0,7807 e early 2 0 TET2 1,00 intermediat ZRSR 1,76 intermediat 0,7333 e early 2 e late 3 TET2 1,00 intermediat STAG 2,00 late 0,7721 e early 2 5 EZH2 1,00 intermediat RUNX 2,00 late 0,7619 e early 1 0 SF3B1 1,00 intermediat RUNX 2,00 late 0,8113 e early 1 2 ASXL1 1,03 intermediat BCOR 1,97 intermediat 0,7843 e early e late 1 ASXL1 1,03 intermediat CBL 2,00 late 0,6600 e early 0 ASXL1 1,03 intermediat PHF6 2,00 late 0,7142 e early 9 ASXL1 1,03 intermediat NF1 1,97 intermediat 0,6521 e early e late 7 ASXL1 1,03 intermediat CEBP 1,86 intermediat 0,7804 e early A e late 9 EZH2 1,00 intermediat STAG 2,00 late 0,7631 e early 2 6 SRSF2 1,00 intermediat IDH1 1,87 intermediat 0,7500 e early e late 0 TET2 1,00 intermediat MPL 1,99 late 0,7500 e early 0
[0481]
[0482] SIB - 56 - BW1367R
[0483] U2AF1 1,00 intermediat PHF6 2,00 late 0,7500 e early 0 ASXL1 1,03 intermediat KRAS 2,00 late 0,7600 e early 0 SRSF2 1,00 intermediat NRAS 2,03 late 0,9200 e early 0 SRSF2 1,00 intermediat CEBP 1,86 intermediat 0,8333 e early A e late 3 SRSF2 1,00 intermediat CBL 2,00 late 0,8571 e early 4 TP53 1,00 intermediat PPM1 2,00 late 0,7142 e early D 9 SRSF2 1,00 intermediat ETV6 1,51 intermediat 0,7500 e early e late 0 ASXL1 1,03 intermediat PTPN 2,00 late 0,6500 e early 11 0 SETBP1 1,01 intermediat ETV6 1,51 intermediat 0,8421 e early e late 1 SRSF2 1,00 intermediat ETNK 1,99 late 0,7647 e early 1 1 ASXL1 1,03 intermediat SMC1 2,00 late 0,6875 e early A 0 CUX1 1,00 intermediat PHF6 2,00 late 0,7857 e early 1 U2AF1 1,00 intermediat MPL 1,99 late 0,8571 e early 4 EZH2 1,00 intermediat ETNK 1,99 late 0,7692 e early 1 3 SETBP1 1,01 intermediat PTPN 2,00 late 0,8888 e early 11 9 WT1 1,00 intermediat CBL 2,00 late 0,8571 e early 4
[0484]
[0485] SIB - 57 - BW1367R
[0486] SF3B1 1,00 intermediat KRAS 2,00 late 0,6666 e early 7 FLT3 1,16 intermediat SMC1 2,00 late 0,6666 e late A 7 NPM1 1,19 intermediat SMC3 1,99 late 1,0000 e late 0 NPM1 1,19 intermediat NRAS 2,03 late 0,7142 e late 9 NPM1 1,19 intermediat KRAS 2,00 late 0,7500 e late 0
[0487]
[0488] Table 9
[0489] Nine of these routes originated from early-occurring genes (Rank range 0.0-0.01), with the evolutionary signatures most frequently progressing towards intermediate-early genes (6 routes): CREBBP to SF3B1, KMT2D to U2AF1, KMT2D to SF3B1, NOTCH1 to SRSF2, NOTCH 1 to SETBP1, and NOTCH 1 to DNMT3A. One route progressed to an intermediate-late gene: KMT2D to FLT3. Two routes advanced to late genes: CREBP to PPM1D and NOTCH 1 to SMC3. The majority of evolutionary routes (33) originated from intermediate-early genes, leading in 8 instances to intermediate-late genes and in all other cases to late genes. The most frequently involved parent gene were ASXL1, SRSF2 and TET2 accounting respectively for 39.6%, 23.08% and 12.1% of all selected co-occurrent mutations. Most frequently child genes were RUNX1, STAG2 and ZRSR2 accounting respectively for 25.8, 24.2 and 5.5% of all selected evolutionary routes. In detail, patient carrying ASXL1 as a parent mutation have a high probability of having the following mutations as a child gene: STAG2, SMC1A, RUNX1, PTPN11, PHF6, NF1, KRAS, CEBPA, CBL, BCOR. Moreover, from SRSF2 the following child genes originate: STAG2 RUNX1 NRAS IDH1 ETV6 ETNK1 CEBPA CBL. Lastly from TET2 the following child genes were selected: ZRSR2 STAG2 MPL. TP53 mutations are frequently found in patients with therapy-related MDS. In this context, TP53 mutations present in rare subclones before therapy confer resistance to chemotherapy, providing a survival advantage. The common co-occurrent mutation found in therapy-related MDS patients TP53-PPM1D was selected by ProgEvo with TP53 acting as a parent gene and PPM1D as a child gene, with high Evo-score 0.71 and the highest CV-score. In summary, the training dataset analysis by ProgEvo identified 1,625 co-occurring gene pairs, forming 46 distinct evolutionary routes with significant directional relationships and statistical SIB - 58 - BW1367R
[0490] support. Despite the wide mutational landscape of MDS, certain genes frequently appeared as parent genes (ASXL1, SRSF2, and TET2) and others as child genes (RUNX1, STAG2, and ZRSR2). These genes are also high-frequency mutations in MDS, suggesting they could influence the occurrence of evolutionary signatures. However, it is noteworthy that other frequently mutated genes like SF3B1, DNMT3A, and TP53 were selected for inclusion in only a few evolutionary routes (e.g., TP53 to PPM1D; SF3B1 to KRAS or STAG2; NOTCH 1 to DNMT3A). This suggests that both the frequency of cooccurrence and the strength of evolutionary associations are crucial factors, not just the frequency of individual mutations.
[0491] Example 1.4 – Computer implemented method step 3 and 4)
[0492] 24 different evolutionary routes and 40 single genomic mutations have been selected to compare against clinical and cytogenetic variables present in IPSS-M. Using bootstrap sampling and regularized Cox regression with LASSO penalty, we selected only those features consistently associated with LFS, independent of the variables present in IPSS-M. Only covariates identified as significant (HR 0) in at least 70% of the bootstrap iterations were selected for final Cox-regression analysis. Final Cox regression with LASSO penalty has been exploited, adjusted for age, including all selected features plus the IPSS-M score as a continuous variable.
[0493] Example 2 – Re-stratification of IPSS-M patients
[0494] Example 2.1 – Computer implemented method step 5)
[0495] The IPSS-M-Evo was calculated as the sum of all Hazard Ratios, with the IPSS-M as a continuous variable accounting for an HR of 0.59 and the additional variables together accounting for an HR of 0.51 (86% of the IPSS-M score).
[0496] Example 2.2 – Computer implemented method step 6)
[0497] The clustering of IPSS-M-Evo was performed using k-means clustering with k=6 clusters to maintain the same number of clusters as the IPSS-M. The IPSS-M-Evo clustering was performed on the training dataset and is as follows: Very-Low (VL) score < 0.35 (314 patients, 15.6%), Low (L) score > 0.36 and < 0.90 (589 patients, 29.2%), Moderate-Low (ML) score > 0.91 and < 1.47 (427 patients, 21.1%), Moderate-High (MH) score > 1.48 and <2.10 (333 patients, 16.5%), High (H) score >2.11 and <2.84 (241 patients, 12.0%), and Very High (VH) score > 2.85 (111 patients, 5.5%).
[0498] IPSS-M-Evo obtained statistically significant higher C-indexes compared to IPSS-M (p<0.01) and with respect to both OS and LFS, with a 2.7-3.5% increase in the training dataset compared to IPSSM (OS c-index: Training IPSS-M 0.75 and IPSS-M-Evo 0.76 SIB - 59 - BW1367R
[0499] p<0.0001 with Delta-AIC 87 in favor of IPSS-M-Evo; LFS c-index: Training IPSS-M 0.75 and IPSSM-Evo 0.76 p=0.001 Delta-AIC in favor of 64). (Figure 2, Figure-1 S).
[0500] 52.1% (1,051) of patients remained in their original IPSS-M cluster, while 863 patients (43%) were reclassified into a lower risk category, and 101 patients (5%) were upgraded to a higher category. Most downgraded patients shifted by 1 to 2 risk classes, whereas upgraded patients were primarily within the VL and L IPSS-M risk groups, and all were upgraded by only 1 class.
[0501] Kaplan-Meier analysis for LFS and OS across re-clustered patients revealed significant overlap among the IPSS-M VL, L, and ML groups, indicating inadequate stratification. In contrast, applying the IPSS-M-Evo model resulted in a better distinct separation across all risk groups for both LFS and OS. These findings underscore the potential clinical impact of IPSS-M-Evo, as it identifies patients who may require significantly different therapeutic management compared to what would be suggested by IPSS-M. In conclusion, IPSS-M-Evo effectively identifies patients who are not well classified according to IPSS-M and re-stratifies them into more appropriate risk categories.
[0502] Example 3 - Risk stratification method according to one of the embodiment of the invention
[0503] Study Population
[0504] The training dataset included 2,519 patients from the original IPSS-M publication, with sequencing and clinical annotations retrieved via cBioPortal (https: / / www.cbioportal.org / study / summary?id=mds_iwg_2022).
[0505] Clinical validation of the prognostic model was carried out on an external cohort of 2,157 patients from the MCC study. (See Figure 1)
[0506] Statistical Methods
[0507] Chi-square has been used for categorical variables and T-Test or Mann- Whitney have been used for continuous variables, based on data distribution. For the analysis of OS and LFS, time-to-event was measured from the time of sample collection to the occurrence of the event of interest. For prognostic modeling, excluded patients without a computable IPSS-M (any missing IPSS-M component), with missing age, age >90 years, or survival time <1 month. Missingness at random was assumed (due to technical reasons rather than biologically related factors). Model comparisons were conducted with Cox regression analysis, and differences between models, specifically between IPSS-M and IPSS-M-Evo, were evaluated through Harrell’s Concordance Index (C-index) and the Akaike Information Criterion (AIC). A second round of C-index estimation SIB - 60 - BW1367R
[0508] was performed with a bootstrap analysis and differences between the median of inferred C-index were performed both in the training and testing dataset. Significant differences in the C-index were tested to determine statistical significance. A significance level of two-sided p-values <0.05 was used for all statistical tests. To address incomplete genetic features, deterministic best, worst, and average imputations were implemented and summarized uncertainty via the score range (worst-best), considering average assignments reliable when the score range <1.0. Robustness was evaluated with progressive-missingness bootstraps and single-variable removal. All analyses were conducted using the R statistical platform, version 4.3.1.
[0509] RESULTS
[0510] Patients
[0511] A total of 2,519 patients from the cBioPortal cohort were included in the evolutionary analysis. All patients had a confirmed diagnosis of MDS according to the 2016 WHO classification (Figure 1).
[0512] . Patients mean age was 69.6 (SD ±12.4), in the cBioPortal cohort. Mean bone marrow blast percentage was mean 4.7%; SD ± 4.7. Sex distribution was similar.
[0513] A total of 1,987 patients in the cBioPortal had complete survival data and were included in the prognostic modeling. In these subgroups, median LFS and 63.5 months. Median OS was 36.1 months
[0514] Driver Mutations Timing
[0515] For evolutionary analysis 46 genes sequenced in both the training and testing datasets were selcted. In the training dataset, 7,828 mutations, accounting for 70.3% of all identified mutations, were analyzed. To estimate the temporal sequence of mutation events, the CV-Rank score, a continuous variable indicating the relative timing of each mutation, was calculated: lower scores correspond to an earlier occurrence of the mutation during MDS progression. Mutations were further stratified into three discrete temporal ranks (0 to 2).
[0516] Early and intermediate-early mutations mainly involved DNA methylation, RNA splicing, and transcription factors genes.
[0517] In particular, the 1strank, which includes genes emerging at the earliest stages of disease evolution- included KMT2D, NOTCH 1, PRPF40B, EP300, CREBBP, KIT, and ATRX. The 2ndrank exhibited broader variability in CV-Rank scores. Within this rank, a first subset of early mutations could be identified, including BCORL1, CSF3R, JAK2, and GATA2. Although signaling genes are typically late mutations, JAK2, which constituted SIB - 61 - BW1367R
[0518] 3.5% of all analyzed mutations, was associated with the early-intermediate phase of MDS.
[0519] Following these, a larger subgroup of genes within the 2ndrank, accounting for 58% of all mutations, showed a narrow CV-Rank score range (1.00 to 1.12). This group included DNMT3A, SF3B1, SRSF2, TET2, TP53, U2AF1, ASXL1 and EZH2, genes previously linked to clonal hematopoiesis of indeterminate potential (CHIP).
[0520] A distinct backbone of mutations characterized by a narrow CV-Rank Score range (1.99 to 2.01) was identified in the 3rdrank. Genes present at this stage included MPL, ETNK1, SMC3, RAD21, PTPN11, KRAS, SMC1A, STAG2, GNB1, PHF6, RUNX1, CBL, PPM1D, and NRAS.
[0521] In conclusion, the narrow CV-Rank distributions observed among specific genes in the 2ndand 3rdranks suggest that these mutations cluster at critical timepoints in MDS evolution, with intermediate-early mutations likely initiating the disease and later mutations contributing to its progression.
[0522] Evolutionary Routes
[0523] Training Dataset, evolutionary routes analysis in the cBioPortal cohort:
[0524] In the training dataset, ProgEvo identified 1,765 co-occurring gene pairs, grouped into 45 distinct evolutionary routes.
[0525] Six of these routes originated from very early-occurring genes classified in the 1strank, representing 6.1% of all selected co-occurrent genes. The identified trajectories included: KMT2D to SF3B1, NOTCH1 to DNMT3A, CREBBP to SF3B1, KIT to TET2 and ATRX to TP53. Notably, the ATRX to TP53 route was the only one in which TP53 was identified as a downstream event. Except for the KMT2D to NF1 routes, all early-originating routes linked to the previously described central mutational backbone of 2ndrank genes.
[0526] Most evolutionary routes (39 out of 45; 87%) originated from genes in the 2ndrank. Among these, 21 routes were directed toward child genes within the previously described 3rdrank mutational backbone (from CV-rank 1.99 to 2.01).
[0527] The most frequently involved parent gene were ASXL1, SRSF2 and TET2 accounting respectively for 34.1% (n° 639), 22.1% (n° 414) and 10.1% (n° 190) of all selected co-occurrent mutations. In detail, patients carrying ASXL1 as a parent mutation have a high probability of having the following mutations as a child gene: STAG2, SMC1A, RUNX1, PTPN11, PHF6, NF1, KRAS, CEBPA, CBL, BCOR. Moreover, from SRSF2 the following child genes originate: STAG2, RAD21 RUNX1, NRAS, KRAS, IDH1, GNAS, ETV6, SIB - 62 - BW1367R
[0528] ETNK1, CEBPA, CBL. Lastly from TET2 the following child genes were selected: ZRSR2, STAG2, MPL. The common co-occurrent mutation found in therapy-related MDS patients TP53-PPM1D was selected by ProgEvo with TP53 acting as a parent gene and PPM1D as a child gene.
[0529] Prognostic Impact of Validated Evolutionary Routes and Co-occurrences Compatible with the Evolutionary Model
[0530] To preliminarily assess the prognostic impact of individual evolutionary trajectories, their association with LFS and OS were associated, analyzing both the validated routes and the parent genes independently.
[0531] As previously reported, ASXL1 mutations were associated with low LFS (HR = 1.67; 95% Cl, 1.46-1.91). However, the prognostic impact varied depending on the downstream evolutionary route. For example, when ASXL1 evolved into KRAS, the LFS hazard ratio increased to 2.87 (95% Cl, 1.72-4.78), whereas the ASXL1-STAG2 route was associated with an HR of 2.51 (95% Cl, 2.04-3.08).
[0532] This pattern was observed across other parent genes. SRSF2 mutations evolving into NRAS or STAG2 were associated with especially poor outcomes (HR: 2,84 Cl 1,64-4,92; HR: 2,24 Cl 1,76-2,85 respectively).
[0533] SF3B1 mutations, typically linked to favorable prognosis in MDS (HR = 0.61; 95% Cl, 0.52-0.71), demonstrated adverse prognostic associations when followed by RUNX1 (HR = 1.84; Cl 1,26-2,68), highlighting a potential shift in clinical significance depending on the downstream evolutionary context.
[0534] Beyond validated routes, "pure co-occurrence" events, defined as gene pairs within the same temporal rank that do not violate the inferred evolutionary order, were analysed. Features Selection for Implementing the IPSS-M Score with Evolutionary Routes After establishing a model of MDS evolution, variables associated with LFS were selected for integration into the IPSS-M score, while preserving its original structure. These included evolutionary routes, co-occurrences consistent with the inferred evolutionary hierarchy, and individual gene mutations. The primary aim was to preserve the original structure of the IPSS-M while minimizing the number of additional variables. To this end, we applied bootstrap resampling and LASSO-penalized Cox regression. Five variables were retained based on their recurrence and significance and added to IPSS-M in order to create IPSS-M-Evo: the evolutionary routes ASXL1 to KRAS and SRSF2 to NRAS, the co-occurrence of NRAS and RUNX1, ATRX and JAK2 mutation. SIB - 63 - BW1367R
[0535] The IPSS-M-Evo was calculated as the sum of all coefficients derived from a cox regression model, with the IPSS-M as a continuous variable accounting for an overall weight 0.62 and the additional variables together accounting for 2.9 (IPSS-M-Evo Variables weight: ASXL1 to KRAS = 0.55; SRSF2 to NRAS = 0.65; NRAS and RUNX1 -0.805; Root to ATRX = 0.41; Root to JAK2= 0.46). The model was subsequently adjusted for patient age (Variable Weight=0.02). (Figure 2).
[0536] the frequency of IPSS-M-Evo variables was relatively low: at least one variable was detected in 147 of 1,867 patients (7.9%) in the cBioPortal dataset. Co-occurrence of two or more IPSS-M-Evo variables was rare, observed in only 8 patients in cBioPortal. For categorical risk stratification, the IPSS-M-Evo score was partitioned into six discrete classes (k = 6), consistent with the original IPSS-M structure. Cutoffs were defined as follows: Very Low (VL) (score < 0.50; 253 patients; 12.7%), Low (L) (0.51-1.00; 467 patients; 23.5%), Moderate-Low (ML) (1.01-1.50; 421 patients; 21.2%), Moderate-High (MH) (1.51-2.00; 278 patients; 14.0%), High (H) (2.01-3.00; 409 patients; 20.1%), and Very High (VH) (>3.00; 159 patients; 8.0%).
[0537] The addition of evolutionary trajectories and gene mutations not originally included in the IPSS-M model significantly improved predictive performance. The IPSS-M-Evo consistently showed higher C-index and lower AIC values for LFS in the cBioPortal training cohort, Enhanced Risk Stratification with IPSS-M-Evo
[0538] In the cBioPortal cohort, 43% of patients (842 / 1,987) were reassigned to a different risk category by the IPSS-M-Evo model. This proportion increased to 58.5% (86 / 147) among those harboring at least one IPSS-M-Evo variable.
[0539] Reclassification predominantly affected patients in higher-risk IPSS-M categories and was generally associated with downgrading. In cBioPortal, 55.5% of patients initially classified as VH were reassigned to H, 31.2% of those in H to MH, and 31.3% of those in MH to ML.
[0540] These shifts reflected both methodological differences in how IPSS-M and IPSS-M-Evo define risk categories and the presence or absence of IPSS-M-Evo variables. For instance, among cBioPortal patients originally assigned to the VH category, the prevalence of at least one IPSS-M-Evo variable was 15% among those who remained in VH, but only 6% among those who were downgraded. A similar pattern was observed in the H risk group: 13.3% of patients who remained in H or were upstaged had at least one IPSS-M-Evo variable (24 / 180), compared to just 0.9% (1 / 108) among those who were downgraded. SIB -64 - BW1367R
[0541] IPSS-M-Evo consistently demonstrated improved C-index values compared to IPSS-M across both the training and testing datasets among the re-clustered patients, suggesting that re-clustering was performed correctly by the new model.
[0542] In conclusion, IPSS-M-Evo effectively identifies patients who are not well classified according to IPSS-M and restratifies them into more appropriate risk categories.
[0543] Web-Based Tool for IPSS-M-Evo Calculation and Patient-Specific Evolutionary Reconstruction
[0544] A dedicated web application (site link - available after publication), developed using R Shiny, enables clinicians to calculate the IPSS-M-Evo score for individual patients and to visualize estimated LFS and OS curves. Additionally, users can input patient-specific mutational profiles (presence / absence) to match them with the most probable evolutionary trajectory inferred by the ProgEvo framework. The proposed trajectories provide an informative approximation based on cohort-level evolutionary features and are not intended to reconstruct the evolution at the single-patient level.
[0545] DISCUSSION
[0546] The results obtained from the application of the method according to the invention confirm the roles of ASXL1, SRSF2, and TET2 as foundational elements in the early stages of evolutionary routes, serving as a backbone for MDS initiation and providing critical insights into disease’s progression.
[0547] SF3B1 is typically an early mutation in MDS. SF3B1 mutations can occur alone or alongside TET2, DNMT3A, RUNX1, or ASXL1. The analysis identified only one evolutionary route among those genes, from SF3B1 to RUNX1. This route was present in 6.9% to 8.1% of SF3B1 -mutated patients and was associated with significantly poorer outcomes.
[0548] A further refinement emerged in the differential evolutionary timing of BCOR and BCORL1. While these genes are often clustered together due to overlapping biological functions our model placed BCOR as a consistently late-acquired mutation, typically downstream of DNMT3A, U2AF1 and ASXL1. This was in line with previous studies; nevertheless, BCORL1 was inferred to occur earlier, suggesting a different evolutionary path compared to BCOR.
[0549] Previous studies underlined the co-occurrence of TET2 and ZRSR2; the present analysis clarified this relationship by identifying a directional evolutionary trajectory in which ZRSR2 consistently emerged as the downstream event. SIB - 65 - BW1367R
[0550] Each directional route in our model was evaluated through a two-step process: first by inferring the temporal sequence of mutations across the cohort, and subsequently by quantifying the clinical impact of each route. This dual-layered evaluation supports a temporally contextualized risk modeling framework, wherein the clinical significance of a mutation is informed not solely by its presence but also by its position within the clonal hierarchy.
[0551] IPSS-M-Evo features led to improved prognostic accuracy and enhanced risk stratification, particularly among high-risk subgroups. The ProgEvo framework enabled the integration of variables within an evolutionary context, including both validated directional routes and mutations not associated with validated trajectories, such as ATRX and JAK2. Although both genes were consistently detected in early-intermediate phases of disease evolution, neither gave rise to validated evolutionary routes.
[0552] In conclusion, the evolutionary modeling framework offers a biologically informed and clinically implementable refinement of existing prognostic tools in MDS. By incorporating directional relationships, temporal mutation acquisition, and clinical impact into a unified structure. By assigning gene-specific probabilities of downstream evolution, this model opens up the possibility for dynamic, personalized genomic monitoring.
[0553] Example 4 - Rank score and CV of evolutionary trajectory calculated according to a preferred embodiment of the invention
[0554] The values of the agony-derived rank and the CV-Rank, calculated by adjusting the VAF-derived Cancer Cell Fraction (CCF) to account for copy number alterations (CNA), are summarized below. Such adjustment enables accurate estimation of clonal hierarchies and improves the robustness of evolutionary inference compared with previously published evolutionary modelling frameworks (e.g., ASCETIC, Fontana et al., 2023), in which ploidy was assumed to be diploid regardless of the clinical variable of cytogenetic.
[0555] GENE RANK CV RANK KMT2D 0,00 0,00
[0556] NOTCH 1 0,00 0,00
[0557] PRPF40B 0,00 0,00
[0558] CREBBP 0,00 0,01
[0559] EP300 0,00 0,02
[0560] ATRX 0,00 0,21
[0561] KIT 0,00 0,31
[0562]
[0563] SIB - 66 - BW1367R
[0564] BCORL1 1,00 0,82
[0565] CSF3R 1,00 0,89
[0566] JAK2 1,00 0,94
[0567] GATA2 1,00 0,99
[0568] CUX1 1,00 1,00
[0569] DNMT3A 1,00 1,00
[0570] SF3B1 1,00 1,00
[0571] SRSF2 1,00 1,00
[0572] TET2 1,00 1,00
[0573] TP53 1,00 1,00
[0574] U2AF1 1,00 1,00
[0575] WT1 1,00 1,00
[0576] SETBP1 1,00 1,06
[0577] ASXL1 1,00 1,09
[0578] EZH2 1,00 1,12
[0579] FLT3 1,00 1,22
[0580] NPM1 1,00 1,32
[0581] IDH2 1,00 1,35
[0582] ETV6 2,00 1,58
[0583] ZRSR2 2,00 1,88
[0584] IDH1 2,00 1,93
[0585] BCOR 2,00 1,95
[0586] CEBPA 2,00 1,95
[0587] NF1 2,00 1,95
[0588] ETNK1 2,00 1,97
[0589] RAD21 2,00 1,98
[0590] PTPN11 2,00 1,99
[0591] SMC1A 2,00 1,99
[0592] CBL 2,00 2,00
[0593] GNAS 2,00 2,00
[0594] GNB1 2,00 2,00
[0595] MPL 2,00 2,00
[0596] PHF6 2,00 2,00
[0597]
[0598] SIB - 67 - BW1367R
[0599] PPM1D 2,00 2,00
[0600] RUNX1 2,00 2,00
[0601] SMC3 2,00 2,00
[0602] STAG2 2,00 2,00
[0603] KRAS 2,00 2,03
[0604] NRAS 2,00 2,09
[0605]
[0606] Table 10
[0607] All evolutionary routes derived using the evolutionary model, after adjusting the VAF-derived Cancer Cell Fraction (CCF) to account for copy number alterations (CNA), are summarized below.
[0608] PARENT CHILD EVO_SCORE CV_SCORE ASXL1 CBL 0,67347 0,91000
[0609] ASXL1 RUNX1 0,67222 0,91000
[0610] ASXL1 STAG2 0,71166 0,91000
[0611] ASXL1 PHF6 0,71429 0,90000
[0612] ASXL1 CEBPA 0,78049 0,87000
[0613] ASXL1 KRAS 0,76000 0,80000
[0614] ASXL1 BCOR 0,78431 0,77000
[0615] ASXL1 NF1 0,65217 0,75000
[0616] ASXL1 PTPN11 0,65000 0,69000
[0617] ASXL1 SMC1A 0,73333 0,53000
[0618] ASXL1 GNB1 0,50000 0,30000
[0619] ATRX TP53 0,58333 0,49000
[0620] ATRX DNMT3A 0,60000 0,37000
[0621] ATRX U2AF1 0,50000 0,26000
[0622] BCORL1 NRAS 0,85714 0,30000
[0623] BCORL1 BCOR 0,38889 0,83000
[0624] CREBBP PPM1D 1,00000 0,69000
[0625] CREBBP SF3B1 0,69231 0,71000
[0626] CREBBP ASXL1 0,50000 0,91000
[0627] CSF3R CEBPA 0,33333 0,52000
[0628] CUX1 PHF6 0,78571 0,64000
[0629] CUX1 MPL 0,62500 0,47000
[0630]
[0631] SIB - 68 - BW1367R
[0632] DNMT3A BCOR 0,59574 0,66000
[0633] DNMT3A GNB1 0,60000 0,41000
[0634] EP300 CUX1 0,44444 0,50000
[0635] EZH2 RUNX1 0,69355 0,88000
[0636] EZH2 ETNK1 0,61538 0,84000
[0637] EZH2 STAG2 0,63158 0,60000
[0638] EZH2 ZRSR2 0,40741 0,42000
[0639] FLT3 SMC1A 0,66667 0,57000
[0640] FLT3 PTPN11 0,50000 0,43000
[0641] KIT TET2 0,64706 0,30000
[0642] KIT CUX1 0,57143 0,59000
[0643] KIT BCORL1 0,33333 0,25000
[0644] KMT2D FLT3 0,83333 0,55000
[0645] KMT2D SF3B1 0,81250 0,48000
[0646] KMT2D NF1 0,61538 0,19000
[0647] KMT2D U2AF1 0,66667 0,80000
[0648] KMT2D EZH2 0,42857 0,29000
[0649] NOTCH 1 SRSF2 1,00000 0,91000
[0650] NOTCH 1 SMC3 1,00000 0,49000
[0651] NOTCH 1 SETBP1 0,88889 0,43000
[0652] NOTCH 1 DNMT3A 0,72727 0,47000
[0653] NPM1 SMC3 1,00000 0,61000
[0654] NPM1 KRAS 0,75000 0,43000
[0655] NPM1 NRAS 0,71429 0,60000
[0656] PRPF40B WT1 0,66667 0,42000
[0657] SETBP1 PTPN11 0,87500 0,70000
[0658] SETBP1 ETV6 0,78947 0,52000
[0659] SF3B1 GNB1 0,75000 0,06000
[0660] SF3B1 RUNX1 0,81132 0,55000
[0661] SF3B1 ZRSR2 0,41176 0,85000
[0662] SRSF2 RUNX1 0,75207 1,00000
[0663] SRSF2 STAG2 0,78070 1,00000
[0664] SRSF2 GNAS 0,64286 0,99000
[0665]
[0666] SIB - 69 - BW1367R
[0667] SRSF2 IDH1 0,75000 0,93000
[0668] SRSF2 CEBPA 0,83333 0,88000
[0669] SRSF2 ETNK1 0,76471 0,84000
[0670] SRSF2 NRAS 0,92000 0,77000
[0671] SRSF2 CBL 0,90476 0,75000
[0672] SRSF2 ETV6 0,75000 0,57000
[0673] SRSF2 RAD21 0,60000 0,55000
[0674] SRSF2 KRAS 1,00000 0,50000
[0675] SRSF2 ZRSR2 0,44444 0,87000
[0676] TET2 ZRSR2 0,72222 0,88000
[0677] TET2 MPL 0,71429 0,56000
[0678] TET2 STAG2 0,74684 0,43000
[0679] TET2 IDH1 0,46154 0,91000
[0680] TP53 PHF6 0,75000 0,26000
[0681] TP53 PPM1D 0,66667 1,00000
[0682] TP53 NF1 0,58824 0,03000
[0683] U2AF1 PHF6 0,75000 0,99000
[0684] U2AF1 BCOR 0,62162 0,95000
[0685] U2AF1 ETNK1 0,60000 0,80000
[0686] U2AF1 MPL 0,85714 0,77000
[0687] U2AF1 ETV6 0,69231 0,44000
[0688] WT1 CBL 0,85714 0,69000
[0689]
[0690] Table 11
[0691] The evolutionary routes with an Evo-Score greater than 0.5 and occurring in at least 10 co-occurrences within the cBioPortal dataset, derived using the evolutionary model after adjusting the VAF-derived Cancer Cell Fraction (CCF) to account for copy number alterations (CNA), are summarized below.
[0692] PARENT CHILD EVO_SCORE CV_SCORE ASXL1 CBL 0,67347 0,91000
[0693] ASXL1 RUNX1 0,67222 0,91000
[0694] ASXL1 STAG2 0,71166 0,91000
[0695] ASXL1 PHF6 0,71429 0,90000
[0696] ASXL1 CEBPA 0,78049 0,87000
[0697]
[0698] SIB - 70 - BW1367R
[0699] ASXL1 KRAS 0,76000 0,80000 ASXL1 BCOR 0,78431 0,77000 ASXL1 NF1 0,65217 0,75000 ASXL1 PTPN11 0,65000 0,69000 ASXL1 SMC1A 0,73333 0,53000
[0700] ATRX TP53 0,58333 0,49000 CREBBP SF3B1 0,69231 0,71000
[0701] CUX1 PHF6 0,78571 0,64000 DNMT3A BCOR 0,59574 0,66000 DNMT3A GNB1 0,60000 0,41000
[0702] EZH2 RUNX1 0,69355 0,88000
[0703] EZH2 ETNK1 0,61538 0,84000
[0704] EZH2 STAG2 0,63158 0,60000
[0705] KIT TET2 0,64706 0,30000 KMT2D SF3B1 0,81250 0,48000 KMT2D NF1 0,61538 0,19000 NOTCH 1 DNMT3A 0,72727 0,47000 SETBP1 ETV6 0,78947 0,52000 SF3B1 RUNX1 0,81132 0,55000 SRSF2 RUNX1 0,75207 1,00000 SRSF2 STAG2 0,78070 1,00000 SRSF2 GNAS 0,64286 0,99000 SRSF2 IDH1 0,75000 0,93000 SRSF2 CEBPA 0,83333 0,88000 SRSF2 ETNK1 0,76471 0,84000 SRSF2 NRAS 0,92000 0,77000 SRSF2 CBL 0,90476 0,75000 SRSF2 ETV6 0,75000 0,57000 SRSF2 RAD21 0,60000 0,55000 SRSF2 KRAS 1,00000 0,50000
[0706] TET2 ZRSR2 0,72222 0,88000
[0707] TET2 MPL 0,71429 0,56000
[0708] TET2 STAG2 0,74684 0,43000
[0709]
[0710] SIB - 71 - BW1367R
[0711] TP53 PPM1D 0,66667 1,00000
[0712] TP53 NF1 0,58824 0,03000
[0713] U2AF1 PHF6 0,75000 0,99000
[0714] U2AF1 BCOR 0,62162 0,95000
[0715] U2AF1 ETNK1 0,60000 0,80000
[0716] U2AF1 MPL 0,85714 0,77000
[0717] U2AF1 ETV6 0,69231 0,44000
[0718]
[0719] Table 12
[0720] Example 5 - Comparison of C-index and AIC for LFS and OS of IPSS-M vs IPSS-M-Evo In the following a comparison between the C-index and AIC between IPSS-M and IPSS-M-EVO is reported; the C-index measures how well a model’s predicted risk while the AIC measures the relative quality of a statistical model, balancing goodness of fit with model complexity.
[0721] OS ProgEvo c-index OS IPSS-M c-index OS ProgEvo AIC OS IPSS-M AIC 0,75993 0,74705 11572,65312 11682,04203 LFS ProgEvo c-index LFS IPSS-M c-index LFS ProgEvo AIC LFS IPSS-M AIC
[0722]
[0723] 0,75875 0,75201 12228,34432 12305,61087 Table 13 presents a comparison of the C-index and AIC values for IPSS-M and IPSS-M-Evo. The C-index measures a model's ability to stratify risk effectively within a population, with values ranging from 0 (no discrimination) to 1 (perfect discrimination), and 0.5 indicating random classification. Higher C-index values denote improved prognostic accuracy. The IPSS-M-Evo model shows statistically significant increases in C-index, indicating enhanced predictive precision.
Claims
SIB - 72 - BW1367RCLAIMS1. A computer implemented risk stratification method of cancer patients or individuals with somatic mutations identified through gene sequencing comprising the following step:1) providing a dataset of genomic covariates and prognosis data or a dataset of genomic and clinical covariates and prognosis data of a statistically significant number of cancer patients or individuals with somatic mutations identified through gene sequencing;2) applying an evolutionary model to the dataset of step 1) for identifying statistically recurrent evolutionary genomic covariates according to said evolutionary method;3) processing the identified statistically recurrent evolutionary genomic covariates of step 2) or statistically recurrent evolutionary genomic covariates of step 2) and clinical covariates using a survival analysis model for selecting signature genomic covariates or signature genomic and clinical covariates related to a specific prognosis;4) processing the signature genomic covariates or signature genomic and clinical covariates selected by step 3) by a survival analysis model for calculating the effect measures thereof;5) calculating the risk-score for one or more cancer patients using signature genomic and clinical covariates associated with a prognosis of step 3) corrected for their effect measures obtained from step 4);wherein said recurrent evolutionary genomic covariates include single gene mutations, co-occurrence of mutations and / or evolutionary routes.
2. The computer implemented method according to claim 1 further comprising the following step:6) clustering patient with a clustering algorithm on the basis of calculated risk-score of step 5).
3. The computer implemented method according to claim 1 or 2, wherein the evolutionary model comprises one or more of the following phases:- computing the Cancer Cell Fraction (CCF) of each somatic mutation or gene mutation from its Variant Allele Frequency (VAF), wherein said computation is adapted to account for copy number alterations (CNA) affecting the corresponding genomic locus;- calculating a Cross-Validation Rank (CV-Rank) and a Cross-Validation Score (CV-Score) for each genomic mutation and evolutionary trajectory;- selecting of evolutionary trajectory having an EVO-score greater than 0,5.SIB - 73 - BW1367R4. The computer implemented method according to claim 3, wherein CCF preferably ploidy-adjusted CCF, is calculated according to the following equation:CCF_estimate= VAF * CNWhere:CN (Copy Number) is set to 3 if Copy Number Alteration (CNA) gain, 1 if CNA loss, and 2 if no CNA event was detected;Tumour purity is assumed to be 100% for all samples; and optionallySex-based normalization is applied for mutations located on sex chromosomes.
5. The computer implemented method according to claim 3 wherein in cases of copyneutral loss of heterozygosity (CN-LOH) events, CCF is assumed to be equal to the uncorrected VAF.
6. The computer implemented method according any one of the previous claims, wherein the evolutionary model further comprises model optimization phase, preferably regularized, to select the model that best fits the data.
7. The computer implemented method according to any one of the previous claims further comprise a step of the application of a survival model, preferably Univariate Cox regression, to the evolutionary trajectories obtained from the evolutionary model and to genomic covariates, wherein said genomic covariates are the parent gene mutations of said evolutionary trajectories, to identify evolutionary trajectories associated with a specific prognosis, preferably overall survival.
8. The computer implemented method according to any one of the previous claims, wherein the cancer patient suffers from cancer or a precursor disease thereof; preferably, cancer is selected from breast cancer, lung cancer, colorectal cancer, prostate cancer, leukaemia, lymphomas, melanoma, pancreatic cancer, ovarian cancer, head and neck cancer, bladder cancer and oesophageal cancer and precursor of cancer are atypical ductal hyperplasia, lobular carcinoma in situ, pulmonary adenoma, adenomatous polyps, Lynch’s syndrome, prostatic intraepithelial neoplasia, hereditary prostate cancer, chronic lymphocytic leukaemia, myelodysplastic syndrome, chronic lymphocytic leukaemia, dysplastic nevi, pancreatic intraductal neoplasia, serous tubal intraepithelial carcinoma, oral premalignant lesions, cystitis and urothelial carcinoma in situ and Barrett’s oesophagus.
9. The computer implemented method according to any one of the previous claims, wherein said cancer patients suffers from myelodysplastic syndrome.SIB - 74 - BW1367R10. The computer implemented method according any one of the previous claims, wherein prognosis data can be overall survival breast cancer-free survival, lung cancer-free survival, colorectal cancer- free survival, colorectal cancer-free survival, prostate cancer-free survival, leukaemia-free survival, lymphomas-free survival, melanoma-free survival, pancreatic cancer-free survival, ovarian cancer-free survival, head and neck cancer-free survival, bladder cancer-free survival and oesophageal cancer-free survival; preferably, prognosis can be overall survival or leukaemia-free survival.
11. The computer implemented method according to any one of the previous claims, wherein prognosis data is overall survival or leukaemia-free survival.
12. The computer implemented method according to any one of the previous claims, wherein clinical covariates continuous variables, categorical variables and binary variables.
13. The computer implemented method according to claim 12, wherein continuous variable selected from the group consisting of age, haemoglobin levels, blood cell count, white blood count, platelet count, bone marrow blasts, chromosomal abnormalities, body mass index, blood pressure, cholesterol level, blood glucose level, liver function tests, serum creatinine, tumour side, duration of illness, heart rate, ejection fraction, heart rate, physical activity level, nutritional intake, c-reactive protein levels and vascular measurement.
14. The computer implemented method according to claim 12, wherein categorical variables are selected from the group consisting of risk classification, gender, ethnicity, smoking status, alcohol consumption, comorbidities, treatment group, performance status, family history of disease, tumour stage, histological tumour classification, residence location, surgical history and medication use.
15. The computer implemented method according to claim 12, wherein binary variable are selected from the group consisting of histopathological diagnosis, hypertension, obesity, radiation therapy history, chemotherapy history, vaccination status, presence of symptoms and infectious disease status.
16. The computer implemented method according anyone of the previous claims, wherein clinical covariates are one more selected from the group consisting of age, haemoglobin, platelet count, absolute neutrophil count, bone marrow blasts, cytogenetics category, presence of del(5q), presence of -7 / del(7q), presence of -17 / del(17p), presence of complex karyoptype and IPSS-M score.SIB - 75 - BW1367R17. The computer implemented method according to any one of the previous claims wherein the evolutionary model comprises at least the following phases: mutation profiling, temporal modelling, statistical inference and model selection, model refinement and validation; preferably the evolutionary model is selected from ASCETIC (Agony-baSed Cancer EvoluTion Inference), OncoNEM, PINTA, AncesTree, Treeomics, PhyloWGS, ClonEvol, PyClone, CancerPathfinder, SCITE, TrAp, CITIIP, ClonePhy, CAPRI, CAPRESE, PMCE, LACE, Expands, SiFit, PhyloSub, Canopy, LICHeE, PhlSCS, Clomial, PhylogicNDT, MACHINA and ScGPS; more preferably, the evolutionary model is ASCETIC.
18. The computer implemented method according to any one of the previous claims, wherein the survival analysis model selected from Kaplan-Meier estimation, parametric survival models, accelerated failure time (AFT) models, Cox regression, Bayesian Cox models, multi-state models, machine learning methods, Ensemble Survival Models, Bayesian Additive Regression Trees (BART) for Survival Analysis, Cox Proportional Hazards Model with Time-Dependent Covariates, Joint Models for Longitudinal and Survival Data, Recurrent Event Models, preferably, Andersen-Gill Model, Deep Learning-Based Survival Models, Flexible Parametric Survival Models, Multi-Task Logistic Regression (MTLR), Aalen’s Additive Model; preferably, the survival analysis model is regularized Cox regression with LASSO or Ridge or Elastic Net penalties; more preferably, Cox regression is regularized Cox regression with LASSO penalty; even more preferably, regularized Cox regression with LASSO penalty is carried out multiple times with bootstrap resampling technique.
19. The computer implemented method according to any one of the previous claims wherein signature genomic and clinical covariates are one or more selected from IPSS-M score, ASXL1 to KRAS, JAK2 mutations, NPM1 mutations SF3B1 to RUNX1, TP53 CCF, Age, TET2 to STAG2, NRAS_and_RUNX1, Root_to_ATRX, Root_to_JAK2, SRSF2_to_NRAS; preferably PSS-M score, ASXL1 to KRAS, JAK2 mutations, Age.
20. The computer implemented method according to any one of the previous claims, wherein signature genomic and clinical covariates areIPSS-M score, ASXL1 to KRAS, JAK2 mutations, NPM1 mutations, SF3B1 to RUNX1, TP53 CCF, Age, TET2 to STAG2; orIPSSM_SCORE, AGE, ASXL1_to_KRAS, NRAS_and_RUNX1, Root_to_ATRX, Root_to_JAK2, SRSF2_to_NRAS.SIB - 76 - BW1367R21. The computer implemented method according to any one of the previous claims, wherein the cancer patient of step 5) can be a cancer patient that belongs to the dataset of step 1) or a cancer patient that does not belong to the dataset of step 1).
22. The computer implemented method according to anyone of the previous claims further comprising a physical step of assessment of the presence or absence of signature genomic covariates or signature genomic covariates and clinical covariates for said cancer patient; preferably, the assessment of the presence or absence of genomic covariate is carried out by gene sequencing of a tumour sample of said patient; more preferably, sequencing is single-cell sequencing, bulk sequencing and multiregional bulk sequencing.
23. A computer program comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method according to any one of the previous claims.
24. A computer-readable data carrier having stored there on the computer program of claim 23.