Survival risk scoring model for acute myelogenous leukemia accompanied with myelodysplastic syndrome related genetic abnormality and application thereof
By constructing a survival risk scoring model based on Lasso regression and Cox regression, the problem of imperfect predictive scoring systems for AML patients with MRGA was solved, achieving more accurate prediction of treatment response and survival, providing individualized treatment plans, and improving the clinical outcomes of AML patients with MRGA.
Patent Information
- Application Number
- CN202511244116.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-02
- Publication Date
- 2026-02-03
AI Technical Summary
The existing predictive scoring system for acute myeloid leukemia with myelodysplastic syndrome-related genetic abnormalities (MRGA) is inadequate, making it difficult to make individualized treatment decisions and failing to effectively distinguish the impact of different genetic abnormalities on prognosis.
A survival risk scoring model based on Lasso regression and Cox regression analysis was constructed. By combining clinical data, molecular bioinformatics and cytogenetics data of 947 patients, key biomarkers were screened, and a treatment response model and a survival risk scoring model were established to predict the complete remission rate and overall survival of AML patients with MRGA.
It improves the accuracy of predicting treatment response and survival in patients with AML and MRGA. By using a combined scoring system to classify patients into low, intermediate, and high-risk groups, it provides a basis for individualized treatment decisions and improves clinical outcomes.
Smart Images

Figure CN121460140A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of molecular biology, and particularly relates to a survival risk scoring model for acute myeloid leukemia with myelodysplastic syndrome-related genetic abnormalities and application thereof. BACKGROUND
[0002] Acute myeloid leukemia (AML) with myelodysplastic syndrome-related genetic abnormalities (MRGA) including TP53 mutation, myeloid dysplasia-related gene mutation (MRG) and cytogenetic abnormality (MRC) accounts for 20-30% of all AML cases and is associated with poor prognosis. AML with MRGA is often secondary to cytotoxic or immunosuppressive therapy (treatment-related AML) or pre-existing hematological diseases such as myelodysplastic syndrome (MDS) or myelodysplasia / myeloproliferative neoplasm.
[0003] In 2022, the World Health Organization (WHO) classified these subtypes, including AML with myeloid dysplasia-related genetic features, myeloid neoplasms following cytotoxic therapy and myeloid neoplasms with predisposition to blast crisis. Similarly, the 2022 European Leukemia Net (ELN 2022) guidelines and the International Consensus Classification (ICC) of Myeloid Neoplasms and Acute Leukemia define a subgroup with poor prognosis mutations (ASXL1, BCOR, EZH2, RUNX1, SF3B1, SRSF2, STAG2, U2AF1 and ZRSR2) as AML with MRG. Cytogenetically, cases with MRC are identified as a separate subgroup.
[0004] Risk stratification is essential for AML with MRGA. Currently, individualized risk assessment and treatment decisions mainly rely on genetic risk classification at diagnosis, and these patients classified as poor risk groups have significantly different outcomes in subgroups with specific genetic abnormalities. Although several studies have attempted to develop predictive prognostic systems for AML, data specific to AML with MRGA are still limited. The relative impact of different genetic abnormalities on outcomes, or the hierarchy of these abnormalities in influencing prognosis, has not been fully elucidated. Over time, several studies have demonstrated the impact of gene mutations on the outcome of MRGA patients, including in the context of hematopoietic stem cell transplantation. However, these findings have not been integrated into a comprehensive prognostic system, limiting their consistent application in clinical practice. Therefore, it is necessary to develop a new scoring system to better guide treatment. SUMMARY
[0005] To solve this problem, the present application provides a survival risk score model for acute myeloid leukemia (AML) with myelodysplastic syndrome-related genetic abnormalities (MRGA) and its application, which includes clinical data, molecular bioinformatics and cytogenetic data of 947 patients, and establishes a treatment response model and a survival risk score model for AML patients with MRGA based on the key markers screened by Lasso regression and Cox regression analysis. Subsequently, the two models were verified using an external validation dataset to predict patient response to induction therapy and overall survival, aiming to improve risk stratification for AML with MRGA and guide individualized treatment decisions, ultimately improving the clinical outcomes of this high-risk population.
[0006] Although acute myeloid leukemia (AML) patients with TP53 mutations, myeloid dysplasia-related gene mutations (MRG) and cytogenetic abnormalities (MRC) have different outcomes, they generally have poor prognosis, and existing prediction scoring systems for these characteristics are not perfect. The present application analyzes data from 737 patients treated at Peking University People's Hospital (PKUPH) from January 2017 to December 2023, and develops a combined scoring system for AML patients with myelodysplastic syndrome-related genetic abnormalities (MRGA) treatment response and survival period by establishing a model using clinical data, cytogenetic and molecular biology data.
[0007] In the treatment response model provided by the present application, factors such as del(7q) / -7, complex karyotype, inv(3) or t(3;3), and U2AF1 mutation are significantly associated with lower CR / CRi rates; while white blood cell count <10x10 9 / L, bone marrow blast cells ≥45%, t(8;21), NPM1 and CEBPAbZIP mutations, and intensive induction therapy are associated with higher CR / CRi rates. The present application divides patients into low and high response groups through the treatment response model, with an area under the receiver operating characteristic curve (AUC) value of 0.63-0.79.
[0008] In addition, the present application also constructs a survival risk score model based on survival period, cytogenetic abnormalities, gene mutations and clinical factors, which divides patients into low, medium and high risk groups, with an AUC value of 0.789-0.831 in the 1-3 year period, and significant survival differences between different risk groups confirm the clinical applicability of the model in making treatment decisions for AML based on MRGA.
[0009] Compared with the existing AML prognosis model, the combined scoring system provided by the treatment response model and the survival risk scoring model is personalized (providing almost unique continuous scores for each patient), reproducible (using applicable mathematical formulas), and simplified (providing a simplified calculation formula for clinicians). Compared with the European Leukemia Network (ELN) and the International Consensus Classification (ICC), the combined scoring system provided by the present application exhibits better discrimination in terms of treatment response, survival, and re-stratification of MRGA patients. This re-stratification reveals significant outcome differences between survival evaluation categories within each ELN 2022 risk layer. The present application develops and verifies the treatment response and survival score using data from 947 patients, integrates blood cell counts, bone marrow blast cell proportions, cytogenetic abnormalities, and key effector gene information, and accurate risk scoring is crucial for advancing risk-adapted treatment strategies for AML and MRGA, helping to identify patient populations that are unlikely to benefit from intensive treatment, and guiding the development of targeted treatment strategies.
[0010] In order to achieve the above-mentioned purpose, the technical scheme adopted by the present application is as follows:
[0011] In one aspect, the present application provides a use of a marker for preparing a reagent for predicting a complete remission rate of acute myeloid leukemia with myelodysplastic syndrome related genetic abnormalities, the marker comprising a combination of any one or more of white blood cell count, bone marrow blast cell proportion, del(7q) / -7, complex karyotype, inv(3) or t(3;3), U2AF1 mutation, t(8;21), NPM1 mutation, and CEBPAbZIP mutation.
[0012] Further, the marker comprises a combination of white blood cell count, bone marrow blast cell proportion, del(7q) / -7, complex karyotype, inv(3) or t(3;3), U2AF1 mutation, t(8;21), NPM1 mutation, and CEBPAbZIP mutation.
[0013] Further, the complete remission rate comprises complete remission CR and complete remission with hematological incomplete recovery CRi, the CR being peripheral blood without blast cells appearing when bone marrow blast cells are <5%, and no leukemia invasion focus outside the bone marrow, neutrophil count ≥1×10 9 / L, and platelet count ≥100×10 9 / L; the CRi being meeting all CR criteria except neutrophil count <1×10 9 / L or platelet count <100×10 9 / L.
[0014] Further, the complete remission rate of acute myeloid leukemia with myelodysplastic syndrome-related genetic abnormalities refers to the probability of a patient with acute myeloid leukemia with myelodysplastic syndrome-related genetic abnormalities achieving complete remission after receiving induction treatment.
[0015] The induction treatment includes intensive treatment regimens (such as high harringtonine combined with cytarabine and aclacinomycin [HAA], idarubicin combined with cytarabine [IA]) and non-intensive treatment regimens (such as venetoclax combined with azacitidine [VEN+AZA]). Patients who achieve complete remission (CR) or complete remission with incomplete hematological recovery (CRi) receive consolidation treatment based on high-dose cytarabine for 3-4 cycles; for patients who do not achieve CR / CRi or relapse after receiving one cycle of induction treatment, a re-induction treatment regimen of medium-dose or high-dose cytarabine can be used, such as modified CLAG (cladribine, cytarabine and G-CSF) or FLAG (fludarabine, cytarabine and G-CSF).
[0016] In another aspect, the present application provides a treatment response model for predicting the complete remission rate of acute myeloid leukemia with myelodysplastic syndrome-related genetic abnormalities, which is constructed based on the markers as described above.
[0017] Further, the formula of the treatment response model is: CR Score = 0.843 x white blood cell count (if <10 x 10 9 / L, =1; otherwise, =0) + 0.757 x percentage of bone marrow blast cells (if ≥45%, =1; otherwise, =0) + 0.608 x intensive induction treatment (if intensive induction treatment, =1; otherwise, =0) + 2.024 x t(8;21) (if t(8;21), =1; otherwise, =0) + 0.701 x del(7q) / -7 (if del(7q) / -7, =0; otherwise, =1) + 1.011 x complex karyotype (if complex karyotype, =0; otherwise, =1) + 2.098 x inv(3) or t(3;3) (if inv(3) or t(3;3), =0; otherwise, =1) + 1.742 x NPM1 (if mutated, =1; otherwise, =0) + 1.385 x CEBPAbZIP (if mutated, =1; otherwise, =0) + 0.792 x U2AF1 (if mutated, =0; otherwise, =1) + 1.16.
[0018] In yet another aspect, the present application provides a combined scoring system for acute myeloid leukemia with myelodysplastic syndrome-related genetic abnormalities, comprising the therapeutic response model and the survival risk score model as described above.
[0019] Further, the survival risk score model is constructed based on a combination of markers of male, age, del(5q) / t(5q) / add(5q), del(7q) / -7, trisomy 8, t(9; 11), inv(3) or t(3;3), complex karyotype, t(v;l lq23.3), inv(16) or t(16; 16), t(8;21), ASXL1 mutation, DNMT3A mutation, TET2 mutation, KRAS mutation, ETV6 mutation, FLT3-ITD mutation, TP53 mutation, GATA2 mutation, BCOR mutation, EP300 mutation, and CEBPAbZIP mutation.
[0020] Further, the formula of the survival risk score model is: RiskScore = 0.529 x male (if male, =1; otherwise, =0) + age (if Age≤35, =0; 35<Age<60, =0.746; Age≥60, =1.347) + 0.551 x del(5q) / t(5q) / add(5q) (if del(5q) / t(5q) / add(5q), =1; otherwise, =0) + 0.424 x del(7q) / -7 (if del(7q) / -7, =1; otherwise, =0) + 0.373 x Trisomy 8 (if Trisomy 8, =1; otherwise, =0) + 0.809 x t(9;11) (if t(9;11), =1; otherwise, =0) + 0.759 x inv(3) or t(3;3) (if inv(3) or t(3;3), =1; otherwise, =0) + 0.691 x complex karyotype (if complex karyotype, =1; otherwise, =0) + 1.076 x t(v;11q23.3) (if t(v;11q23.3), =1; otherwise, =0) + 1.459 x inv(16) or t(16;16) (if inv(16) or t(16;16), =0; otherwise, =1) + 1.784 x t(8;21) (if t(8;21), =0; otherwise, =1) + 0.389 x ASXL1 (if mutated, =1; otherwise, =0) + 0.464 x DNMT3A (if mutated, =1; otherwise, =0) + 0.472 x TET2 (if mutated, =1; otherwise, =0) + 0.585 x KRAS (if mutated, =1; otherwise, =0) + 0.590 x ETV6 (if mutated, =1; otherwise, =0) + 0.591 x FLT3-ITD (if mutated, =1; otherwise, =0) + 0.615 x TP53 (if mutated, =1; otherwise, =0) + 0.729 x GATA2 (if mutated, =1; otherwise, =0) + 0.940 x BCOR (if mutated, =0; otherwise, =1) + 0.999 x EP300 (if mutated, =0; otherwise, =1) + 1.799 x CEBPAbZIP (if mutated, =0; otherwise, =1).
[0021] In another aspect, the present application provides a method for constructing a therapeutic response model for predicting complete remission rate of acute myeloid leukemia with myelodysplastic syndrome related genetic abnormalities, comprising the following steps:
[0022] (1) collecting clinical data, chromosome karyotype and gene mutation information of the patient;
[0023] (2) screening markers by Lasso and Cox regression analysis;
[0024] (3) constructing a therapeutic response model based on the screened markers.
[0025] In another aspect, the present application provides a use of a marker combination for constructing a therapeutic response model for predicting complete remission rate of acute myeloid leukemia with myelodysplastic syndrome related genetic abnormalities, wherein the marker combination comprises white blood cell count, bone marrow blast cell ratio, del(7q) / -7, complex karyotype, inv(3) or t(3;3), U2AF1 mutation, t(8;21), NPM1 mutation and CEBPAbZIP mutation.
[0026] Further, the formula of the therapeutic response model is: CR Score = 0.843 x white blood cell count (if <10 x 10 9 / L, =1; otherwise, =0) + 0.757 x bone marrow blast cell percentage (if ≥45%, =1; otherwise, =0) + 0.608 x intensive induction therapy (if intensive induction therapy, =1; otherwise, =0) + 2.024 x t(8;21) (if t(8;21), =1; otherwise, =0) + 0.701 x del(7q) / -7 (if del(7q) / -7, =0; otherwise, =1) + 1.011 x complex karyotype (if complex karyotype, =0; otherwise, =1) + 2.098 x inv(3) or t(3;3) (if inv(3) or t(3;3), =0; otherwise, =1) + 1.742 x NPM1 (if mutated, =1; otherwise, =0) + 1.385 x CEBPAbZIP (if mutated, =1; otherwise, =0) + 0.792 x U2AF1 (if mutated, =0; otherwise, =1) + 1.16.
[0027] In another aspect, the present invention provides a combination of biomarkers for predicting the complete remission rate of acute myeloid leukemia with genetic abnormalities associated with myelodysplastic syndromes, the combination of biomarkers including white blood cell count, bone marrow blast cell ratio, del(7q) / -7, complex karyotype, inv(3) or t(3;3), U2AF1 mutation, t(8;21), NPM1 mutation and CEBPAbZIP mutation.
[0028] Furthermore, the complete remission rate includes complete remission (CR) and complete remission (CRi) with incomplete hematological recovery. CR is defined as bone marrow blasts <5%, no blasts in peripheral blood, no extramedullary leukemia invasion lesions, and a neutrophil count ≥1×10⁻⁶. 9 / L, platelet count ≥100×10 9 / L; The CRi is defined as meeting all CR criteria except for a neutrophil count <1×10⁻⁶. 9 / L or platelet count <100×10 9 / L.
[0029] In another aspect, the present invention provides the use of biomarkers for preparing reagents to predict overall survival in acute myeloid leukemia with myelodysplastic syndrome-related genetic abnormalities, said biomarkers including any one or more combinations of male sex, age, del(5q) / t(5q) / add(5q), del(7q) / -7, trisomy 8, t(9;11), inv(3) or t(3;3), complex karyotype, t(v;11q23.3), inv(16) or t(16;16), t(8;21), ASXL1 mutation, DNMT3A mutation, TET2 mutation, KRAS mutation, ETV6 mutation, FLT3-ITD mutation, TP53 mutation, GATA2 mutation, BCOR mutation, EP300 mutation, and CEBPAbZIP mutation.
[0030] Furthermore, the markers include male sex, age, del(5q) / t(5q) / add(5q), del(7q) / -7, trisomy of chromosome 8, t(9;11), inv(3) or t(3;3), complex karyotype, t(v;11q23.3), inv(16) or t(16;16), t(8;21), ASXL1 mutation, DNMT3A mutation, TET2 mutation, KRAS mutation, ETV6 mutation, FLT3-ITD mutation, TP53 mutation, GATA2 mutation, BCOR mutation, EP300 mutation, and a combination of CEBPAbZIP mutation.
[0031] Furthermore, the predicted overall survival for acute myeloid leukemia with myelodysplastic syndrome-related genetic abnormalities refers to the predicted 30-month survival rate of patients with acute myeloid leukemia with myelodysplastic syndrome-related genetic abnormalities after receiving induction therapy.
[0032] In another aspect, the present invention provides a survival risk scoring model for predicting overall survival in acute myeloid leukemia with myelodysplastic syndrome-related genetic abnormalities, which is constructed based on the combination of biomarkers described above.
[0033] Furthermore, the formula of the survival risk scoring model is: RiskScore = 0.529 × male (if male, = 1; otherwise, = 0) + age (if Age ≤ 35, = 0; 35 < Age < 60, = 0.746; Age ≥ 60, = 1.347) + 0.551 × del(5q) / t(5q) / add(5q) (if del(5q) / t(5q) / add(5q), = 1; otherwise, = 0) + 0.424 × del(7q) / -7 (if del(7q) / -7, = 1; otherwise, = 0) + 0.373 × Trisomy 8 (if Trisomy 8, = 1; otherwise, = 0) + 0.809 × t(9;11) (if t(9;11), = 1; otherwise, = 0) + 0.759 × inv(3) or t(3;3) (if inv(3) or t(3;3), = 1; otherwise, = 0) + 0.691 × complex karyotype (if complex karyotype, = 1; otherwise, = 0) + 1.076 × t(v;11q23.3) (if t(v;11q23.3), = 1; otherwise, = 0) + 1.459 × inv(16) or t(16;16) (if inv(16) or t(16;16), = 0; otherwise, = 1) + 1.784 × t(8;21) (if t(8;21), = 0; otherwise, = 1) + 0.389 × ASXL1 (if mutated, = 1; otherwise, = 0) + 0.464 × DNMT3A (if mutated, = 1; otherwise, = 0) + 0.472 × TET2 (if mutated, = 1; otherwise, = 0) + 0.585 × KRAS (if mutated, = 1; otherwise, = 0) + 0.590 × ETV6 (if mutated, = 1; otherwise, = 0) + 0.591 × FLT3-ITD (if mutated, = 1; otherwise, = 0) + 0.615 × TP53 (if mutated, = 1; otherwise, = 0) + 0.729 × GATA2 (if mutated, = 1; otherwise, = 0) + 0.940 × BCOR (if mutated, = 0; otherwise, = 1) + 0.999 × EP300 (if mutated, = 0; otherwise, = 1) + 1.799 × CEBPAbZIP (if mutated, = 0; otherwise, = 1).
[0034] Furthermore, this invention provides a method for constructing a survival risk scoring model, as described above, for predicting overall survival in acute myeloid leukemia with myelodysplastic syndrome-related genetic abnormalities, comprising the following steps:
[0035] (1) Collect patients’ clinical data, chromosome karyotype and gene mutation information;
[0036] (2) Biomarkers were screened using Lasso and Cox regression analysis;
[0037] (3) Construct a survival risk scoring model based on the screened biomarkers.
[0038] In another aspect, the present invention provides the use of a combination of biomarkers for constructing a survival risk scoring model for predicting overall survival in acute myeloid leukemia with genetic abnormalities associated with myelodysplastic syndromes. The biomarker combination includes male sex, age, del(5q) / t(5q) / add(5q), del(7q) / -7, trisomy 8, t(9;11), inv(3) or t(3;3), complex karyotype, t(v;11q23.3), inv(16) or t(16;16), t(8;21), ASXL1 mutation, DNMT3A mutation, TET2 mutation, KRAS mutation, ETV6 mutation, FLT3-ITD mutation, TP53 mutation, GATA2 mutation, BCOR mutation, EP300 mutation, and CEBPAbZIP mutation.
[0039] Furthermore, the formula of the survival risk scoring model is: RiskScore = 0.529 × male (if male, = 1; otherwise, = 0) + age (if Age ≤ 35, = 0; 35 < Age < 60, = 0.746; Age ≥ 60, = 1.347) + 0.551 × del(5q) / t(5q) / add(5q) (if del(5q) / t(5q) / add(5q), = 1; otherwise, = 0) + 0.424 × del(7q) / -7 (if del(7q) / -7, = 1; otherwise, = 0) + 0.373 × Trisomy 8 (if Trisomy 8, = 1; otherwise, = 0) + 0.809 × t(9;11) (if t(9;11), = 1; otherwise, = 0) + 0.759 × inv(3) or t(3;3) (if inv(3) or t(3;3), = 1; otherwise, = 0) + 0.691 × complex karyotype (if complex karyotype, = 1; otherwise, = 0) + 1.076 × t(v;11q23.3) (if t(v;11q23.3), = 1; otherwise, = 0) + 1.459 × inv(16) or t(16;16) (if inv(16) or t(16;16), = 0; otherwise, = 1) + 1.784 × t(8;21) (if t(8;21), = 0; otherwise, = 1) + 0.389 × ASXL1 (if mutated, = 1; otherwise, = 0) + 0.464 × DNMT3A (if mutated, = 1; otherwise, = 0) + 0.472 × TET2 (if mutated, = 1; otherwise, = 0) + 0.585 × KRAS (if mutated, = 1; otherwise, = 0) + 0.590 × ETV6 (if mutated, = 1; otherwise, = 0) + 0.591 × FLT3-ITD (if mutated, = 1; otherwise, = 0) + 0.615 × TP53 (if mutated, = 1; otherwise, = 0) + 0.729 × GATA2 (if mutated, =
[0040] On another aspect, the present invention provides a survival risk scoring model for distinguishing acute myeloid leukemia with myelodysplastic syndrome-related genetic abnormalities as low, medium, or high-risk individuals. The survival risk scoring model is constructed based on a combination of biomarkers including male sex, age, del(5q) / t(5q) / add(5q), del(7q) / -7, trisomy 8, t(9;11), inv(3) or t(3;3), complex karyotype, t(v;11q23.3), inv(16) or t(16;16), t(8;21), ASXL1 mutation, DNMT3A mutation, TET2 mutation, KRAS mutation, ETV6 mutation, FLT3-ITD mutation, TP53 mutation, GATA2 mutation, BCOR mutation, EP300 mutation, and CEBPAbZIP mutation.
[0041] Furthermore, the formula of the survival risk scoring model is: RiskScore = 0.529 × male (if male, = 1; otherwise, = 0) + age (if Age ≤ 35, = 0; 35 < Age < 60, = 0.746; Age ≥ 60, = 1.347) + 0.551 × del(5q) / t(5q) / add(5q) (if del(5q) / t(5q) / add(5q), = 1; otherwise, = 0) + 0.424 × del(7q) / -7 (if del(7q) / -7, = 1; otherwise, = 0) + 0.373 × Trisomy 8 (if Trisomy 8, = 1; otherwise, = 0) + 0.809 × t(9;11) (if t(9;11), = 1; otherwise, = 0) + 0.759 × inv(3) or t(3;3) (if inv(3) or t(3;3), = 1; otherwise, = 0) + 0.691 × complex karyotype (if complex karyotype, = 1; otherwise, = 0) + 1.076 × t(v;11q23.3) (if t(v;11q23.3), = 1; otherwise, = 0) + 1.459 × inv(16) or t(16;16) (if inv(16) or t(16;16), = 0; otherwise, = 1) + 1.784 × t(8;21) (if t(8;21), = 0; otherwise, = 1) + 0.389 × ASXL1 (if mutated, = 1; otherwise, = 0) + 0.464 × DNMT3A (if mutated, = 1; otherwise, = 0) + 0.472 × TET2 (if mutated, = 1; otherwise, = 0) + 0.585 × KRAS (if mutated, = 1; otherwise, = 0) + 0.590 × ETV6 (if mutated, = 1; otherwise, = 0) + 0.591 × FLT3-ITD (if mutated, = 1; otherwise, = 0) + 0.615 × TP53 (if mutated, = 1; otherwise, = 0) + 0.729 × GATA2 (if mutated, = 1; otherwise, = 0) + 0.940 × BCOR (if mutated, = 0; otherwise, = 1) + 0.999 × EP300 (if mutated, = 0; otherwise, = 1) + 1.799 × CEBPAbZIP (if mutated, = 0; otherwise, = 1).
[0042] In another aspect, the present invention provides a system for predicting overall survival in acute myeloid leukemia with myelodysplastic syndrome-related genetic abnormalities, including a data analysis module for analyzing the detection values of the biomarkers described above.
[0043] Furthermore, the system also includes a RiskSore calculation module and a result output module. The RiskSore calculation module calculates a score based on the detection value of the marker, and the result output module outputs low-risk, medium-risk, and high-risk levels based on the score.
[0044] In another aspect, the present invention provides a combination of biomarkers for predicting overall survival in acute myeloid leukemia with genetic abnormalities associated with myelodysplastic syndromes. The combination of biomarkers includes male sex, age, del(5q) / t(5q) / add(5q), del(7q) / -7, trisomy 8, t(9;11), inv(3) or t(3;3), complex karyotype, t(v;11q23.3), inv(16) or t(16;16), t(8;21), ASXL1 mutation, DNMT3A mutation, TET2 mutation, KRAS mutation, ETV6 mutation, FLT3-ITD mutation, TP53 mutation, GATA2 mutation, BCOR mutation, EP300 mutation, and CEBPAbZIP mutation.
[0045] The present invention has the following beneficial effects:
[0046] 1. Using MRGA AML patient data from the Peking University People's Hospital cohort (PKUPH), Lasso and Cox regression analyses were used to identify covariates significantly associated with treatment response and overall survival. Treatment response models for predicting complete remission rates and survival risk scoring models for predicting overall survival were then constructed. The AUC values of the treatment response models reached 0.63-0.79, while the AUC values of the survival risk scoring models reached 0.789-0.831 over 1-3 years. The combined scoring system, integrating the treatment response and survival risk scoring models, demonstrated better discriminative ability in terms of treatment response, survival, and restratification of MRGA AML patients.
[0047] 2. Among the covariates selected in this invention for constructing the treatment response model, white blood cell count > 10 × 10⁻⁶. 9Factors such as / L, percentage of bone marrow blasts ≥45%, del(7q) / -7, complex karyotype, inv(3) or t(3;3), and U2AF1 mutation were significantly associated with lower CR / CRi rates; while factors such as t(8;21), NPM1 mutation, CEBPAbZIP mutation, and intensive induction therapy regimens were associated with higher CR / CRi rates. The ranking and weighting of these variables provides a detailed understanding of the complex relationships in MRGA AML and provides an individualized basis for clinical treatment decisions in AML.
[0048] 3. The survival risk scoring model constructed in this invention classifies AML patients into low-risk, intermediate-risk, and high-risk groups. Compared with the ELN 2022 risk classification, the survival risk scoring model provided by this invention has a higher Harrell C index, and its classification efficacy for overall survival (OS), relapse-free survival (RFS), and complete remission (CR) risk is improved in acute myeloid leukemia patients. Within the same risk stratum of ELN 2022, the survival risk scoring model provided by this invention can identify patient subgroups with different prognoses within the ELN strata. In addition, decision curve analysis shows that the survival risk scoring model has clinically practical net benefits in 1-, 2-, and 3-year survival prediction. Attached Figure Description
[0049] Figure 1 The flowchart illustrates the construction process of the combined scoring system of AML and MRGA provided by this invention.
[0050] Figure 2 Heatmap of gene and karyotype characteristics for AML patients with MRGA.
[0051] Figure 3A survival risk score model for AML MRGA was constructed. A density plot shows the survival risk scores calculated based on 737 patients, with the bottom x-axis displaying the survival risk score and the top x-axis displaying the hazard ratios corresponding to the assumed average patient. The vertical dashed line represents the critical value applied to the score, defining the three categories: low-risk, intermediate-risk, and high-risk. B represents the hazard ratios for overall survival in the low-risk, intermediate-risk, and high-risk categories, with the low-risk category corresponding to the reference value. CE represents the Kaplan-Meier probability estimates of overall survival for the three categories in the (C) training set, (D) validation set, and (E) BeatAML dataset, respectively, with p-values derived from the log-rank test. F shows the model's discriminative power as measured by the C-index, survival risk category, survival risk score, and CR score obtained using ELN 2022, covering the three endpoints (i.e., overall survival in blue, relapse-free survival in red, and complete remission in green). Risk is coded as a category (both ELN2022 and survival risk categories use three categories), while survival risk scores and CR scores are coded as scores; G is a stacked bar chart showing the distribution of AML patients in the three risk groups (good, intermediate, and poor) defined by the ELN 2022 classification. Each bar represents the percentage of patients in each group, with the corresponding number of patients (n) in parentheses. The groups are color-coded as follows: good (green), intermediate (blue), and poor (red); H is a Sankey diagram depicting the flow and reclassification of AML patients between the ELN 2022 and risk category systems. This diagram shows the transition of patients from one risk group to another, with the band width proportional to the number of patients.
[0052] Figure 4 To evaluate the survival probability and risk stratification model of AML patients under different induction therapies; where AC represents the Kaplan-Meier survival curves of different risk groups under intensified therapy in (A) all PKUPH groups, (B) intensive therapy training set, and (C) intensive therapy validation set, with p-values derived from the log-rank test; DF represents the Kaplan-Meier survival curves of different risk groups under non-intensive therapy in (D) all PKUPH groups, (E) non-intensive therapy training set, and (F) non-intensive therapy validation set. Similar to the intensive therapy group, the high-risk group has the lowest survival probability, and the difference is statistically significant, with p-values derived from the log-rank test; GJ represents the receiver operating characteristic (ROC) curves of the model's predictive ability for 1-year, 2-year, and 3-year survival rates in (G) intensive therapy training set, (H) intensive therapy validation set, and (I) non-intensive therapy training set and (J) non-intensive therapy validation set.
[0053] Figure 5The performance of the survival risk scoring model is shown in different cohorts and time points. AD represents the ROC curves for 1, 2, and 3 years of overall survival for (A) the entire cohort, (B) the training set, (C) the validation set, and (D) the BeatAML cohort. AUC values for each time point are provided, indicating the model's predictive accuracy. EH represents the 1-year overall survival calibration plots for (E) the entire cohort, (F) the training set, (G) the validation set, and (H) the BeatAML cohort. These plots show the consistency between predicted probabilities and observed outcomes, with the ideal model's points falling on the diagonal. IL represents the DCA of 1-year overall survival for (I) the entire cohort, (J) the training set, (K) the validation set, and (L) the BeatAML cohort. The DCA curves show the net benefit of using the prognostic model compared to treating all patients or not treating any patients. The threshold probability represents the probability of a change in treatment decision.
[0054] Detailed description
[0055] The technical terms involved in this invention will be further explained below. Unless otherwise specified, they shall be interpreted according to the general understanding of the terms in the art.
[0056] Overall survival (OS) : Defined as the time from diagnosis to death from any cause or from diagnosis to the last follow-up, calculated in months.
[0057] Kaplan-Meier curve Survival curves, also known as survival analysis, are a common method in survival analysis. They primarily analyze the impact of a single factor on survival time and are used to estimate patient survival rates and plot survival curves. The KM method, also known as the product-limit method, involves first calculating the probability (i.e., survival probability) of a patient who has survived a certain period and then survives to the next period. The product of these survival probabilities is then multiplied to obtain the survival rate for the corresponding time period. This is also the most commonly used method in survival analysis. A survival curve is a continuous, stepped curve plotted with the observation time on the x-axis and the survival rate on the y-axis. Each point on the curve corresponds to the patient's survival rate at that time point. Survival curves are generally smooth and horizontally extended. When a patient experiences a terminal event (such as death) at a certain time point, the curve will drop vertically. The faster the curve drops, the worse the prognosis.
[0058] Complete remission (CR) Bone marrow blasts <5%; no circulating blasts; no extramedullary disease; ANC ≥1.0×10 9 / L (1,000 / mL); platelet count ≥100×10 9 / L (100,000 / mL).
[0059] Complete remission with incomplete hematologic recovery (CRi) Except for residual neutropenia <1.0×109 / L (1,000 / mL) or thrombocytopenia <100×10 9 Except for / L (100,000 / mL), all CR standards.
[0060] Intensive treatment regimen Intensive chemotherapy for AML is a two-phase process (induction and consolidation) that uses high doses of chemotherapy drugs (usually cytarabine and anthracyclines) to kill leukemia cells. Due to the need for close monitoring and the risk of serious side effects, including a higher risk of infection, anemia, and bleeding, this chemotherapy is typically administered by healthcare professionals in a hospital. Its goal is to achieve and maintain remission by eliminating leukemia cells and preventing relapse.
[0061] Non-intensive treatment regimen Non-intensive chemotherapy for AML uses lower doses of drugs (such as azacitidine and veneclade) to control acute myeloid leukemia, resulting in fewer and milder side effects. For older or less fit patients This is a better option. Its goal is to induce remission and control leukemia for as long as possible, usually with outpatient treatment rather than complete cure. The main non-intensive drug combinations include azacitidine plus veneclade (VEN+AZA), or low-dose cytarabine plus veneclade. The criteria used in clinical trials to screen patients who are not suitable for intensive chemotherapy are as follows: (1) age ≥75 years (but this is not an absolute criterion; for example, patients with milder disease and no related comorbidities may benefit from intensive chemotherapy) or (2) ECOG performance status ≥2 and / or age-related comorbidities, such as severe cardiac disease (e.g., congestive heart failure requiring treatment, ejection fraction ≤50% or chronic stable angina), severe lung disease (e.g., DLCO ≤65% or FEV1 ≤65%), creatinine clearance ≤45 mL / min, liver disease with total bilirubin ≥1.5 times the upper limit of normal, or any other comorbidities assessed by a physician as unsuitable for intensive chemotherapy.
[0062] add(17p) / del(17p) / i(17q) / -17 : Chromosome 17 abnormality, add(17p) indicates short arm duplication of chromosome 17, del(17p) indicates short arm deletion, i(17q) indicates 17q allele, where the long arm is duplicated and the short arm is missing, -17 indicates complete deletion of chromosome 17 (monoclonus 17).
[0063] del(7q) / -7 :del(7q) / -7 indicates a deletion of the long arm of chromosome 7 (del(7q)) or a deletion of the entire chromosome 7 (monocytosis, -7).
[0064] del(5q) / t(5q) / add(5q): Chromosome 5 abnormality, del(5q) indicates deletion of the long arm of chromosome 5, t(5q) indicates translocation of the long arm of chromosome 5, add(5q) indicates duplication of the long arm of chromosome 12.
[0065] del(12p) / t(12p) / add(12p) : Chromosome 12 abnormality, del(12p) indicates deletion of the short arm of chromosome 12, t(12p) indicates translocation of the short arm of chromosome 12, add(12p) indicates duplication of the short arm of chromosome 12.
[0066] Complex karyotype In the absence of other categories of reproducible genetic abnormalities, there are ≥3 unrelated chromosomal abnormalities; excluding hyperdiploid karyotypes with three or more trisomic (or polysomic) abnormalities and no structural abnormalities.
[0067] inv(3) or t(3;3) : inv(3)(q21.3q26.2)or t(3;3)(q21.3;q26.2) / GATA2,MECOM(EVI1).
[0068] t(9;11) :t(9;11)(p21.3;q23.3) / MLLT3::KMT2A.
[0069] t(v;l lq23.3) :t(v;11q23.3) / KMT2A-rearranged.
[0070] inv(16) or t(16;16) : inv(16)(p13.1q22)or t(16;16)(p13.1;q22) / CBFB::MYH11.
[0071] t(8;21) :t(8;21)(q22;q22.1) / RUNX1::RUNX1T1. Detailed Implementation
[0072] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be noted that the embodiments described below are intended to facilitate understanding of the present invention and are not intended to limit it in any way. The reagents used in this embodiment are all known products and were obtained by purchasing commercially available products.
[0073] Example 1: Construction of a treatment response model for AML with MRGA provided by the present invention
[0074] I. Research Subjects
[0075] This invention selected data from 737 consecutive AML patients diagnosed and treated at Peking University People's Hospital (PKUPH) between January 2017 and December 2023, as well as 142 cases from the BeatAML dataset and 68 cases from the Cancer Genome Atlas (TCGA). The inclusion criteria for participants were as follows: (1) age > 18 years; (2) presence of myeloid dysplasia-associated gene (MRG) mutations, including ASXL1, BCOR, EZH2, RUNX1, SF3B1, SRSF2, STAG2, U2AF1, and ZRSR2; (3) presence of myeloid dysplasia-associated cytogenetic abnormalities (MRCs) as defined by the European Leukemia Net 2022 (ELN2022) criteria; and (4) TP53 mutation with a variant allele frequency (VAF) > 10%.
[0076] II. Research Methods
[0077] (1) Clinical data collection, cytogenetic and molecular biological analysis
[0078] Data collection and analysis were conducted on the subjects diagnosed with AML but before treatment. Diagnosis and monitoring were performed according to the ELN2022 criteria, using the following methods: Immunophenotypic analysis was performed using CD45 / lateral gating multiparameter flow cytometry. Cytogenetic analysis was performed using standard G-banding techniques. Molecular screening for leukemia-related fusion genes was performed on all patients. Demographic and clinical data, including complete blood counts and initial hematological, cytogenetic, and molecular analysis results, were extracted from medical records.
[0079] (2) Treatment plan
[0080] Treatment options for AML patients include induction therapy and consolidation therapy. Induction therapy includes intensive regimens (such as homoharringtonine combined with cytarabine and aclarubicin [HAA], idarubicin combined with cytarabine [IA]) and non-intensive regimens (such as veneclax combined with azacitidine [VEN+AZA]). For patients who achieve complete remission (CR) or complete remission with incomplete hematologic recovery (CRi), consolidation therapy based on high-dose cytarabine is administered for 3-4 cycles. For patients who do not achieve CR / CRi or relapse after one cycle of induction therapy, re-induction therapy with medium- or high-dose cytarabine can be used, such as modified CLAG (cladribine, cytarabine, and G-CSF) or FLAG (fludarabine, cytarabine, and G-CSF).
[0081] In FLT3-ITD-positive AML, FLT3 inhibitors (such as sorafenib or gipritinib) are used during induction therapy or subsequent treatment. Patients eligible for allogeneic hematopoietic stem cell transplantation (allo-HSCT) receive more than two cycles of consolidation chemotherapy. Donor selection includes HLA-matched sibling donors, HLA-matched unrelated donors, or HLA-haploidentical related donors.
[0082] After all AML patients received treatment, their physical condition was tracked and recorded, with a median follow-up time of 16 months.
[0083] (3) High-depth targeted region sequencing
[0084] DNA was extracted from cryopreserved bone marrow mononuclear cells using the QIAsymphony SP system (QIAGEN, Germany). Germline variants were identified using oral mucosal samples. Targeted DNA sequencing was performed using a validated, lab-designed hematologic malignancy gene pool (Twist Bioscience, USA) that covered the entire coding regions of genes associated with myeloid tumors. Sequencing was performed on an Illumina platform (ILLUMINA, USA) with an average coverage depth of 2000x. After demultiplexing, data quality was assessed by examining the FASTQ file, and low-quality or N-base-containing data were discarded. Qualified reads were aligned to the human reference genome (hg19) using the Burrows-Wheeler Aligner (BWA). Local realignment and base quality fraction recorrection around indels were performed using the Genomics Analysis Toolkit (GATK 3.4.0). Polymerase chain reaction (PCR) repetitive sequences were removed using Picard. VarScan2 was used to detect single nucleotide variants (SNVs) and insertions / deletions (indels), with a variant allele frequency (VAF) threshold of 1.0% set for SNVs and indels.
[0085] (4) Data Analysis
[0086] Categorical covariates were analyzed using the Pearson chi-square test, while continuous covariates were analyzed using the Student's t-test for normally distributed data or the Mann-Whitney U test for non-normally distributed data, depending on the data distribution. Subjects were randomly assigned to the training and validation sets in a 2:1 ratio. The survival ROC package on the R platform was used to evaluate time-dependent receiver operating characteristic (ROC) curves, and the cutoff value was determined based on the maximum Yangen index, following previous methods. Lasso regression was used to eliminate redundant prognostic variables based on coefficients and partial likelihood deviations. Cox regression models were used for multivariate analysis to identify covariates related to overall survival (OS), and multicollinearity among covariates in the Cox model was assessed using the variance inflation factor. To validate the established model, treatment response scores were divided into low-response and high-response groups based on the optimal cutoff value determined by the maximum Yangen index in the ROC analysis. The prognostic risk score for each subject was calculated using the above formula, and subjects were divided into high, intermediate, and low-risk groups based on the optimal cutoff value determined by X-tile software. Internal model validation was performed using the area under the receiver operating characteristic (AUROC) curve. Calibration plots were plotted to compare predicted probabilities with observed results. Decision curve analysis was used to calculate the model's net benefit, and the Harrell consistency index was used to assess the model's discriminative power. The Kaplan-Meier method was used to calculate relapse-free survival, overall survival, and response rate. Log-rank tests were used for intergroup comparisons.
[0087] III. Research Results
[0088] (1) Patient characteristics
[0089] This invention statistically analyzed 2,649 AML cases, including 1,776 from Peking University People's Hospital, 200 from The Cancer Genome Atlas (TCGA), and 673 from the BeatAML dataset. Among them, 947 participants with myeloid dysplasia-associated genetic abnormalities (MRGA) were identified: 737 from Peking University People's Hospital, 142 from the BeatAML dataset, and 68 from TCGA. The patient enrollment process is as follows... Figure 1 As shown.
[0090] Participants from Peking University People's Hospital were randomly assigned in a 2:1 ratio to the training set (n = 491, 67%) and the internal validation set (n = 246, 33%), with no significant differences in baseline covariates between the two cohorts. Patient characteristics are summarized in Table 1. The median age of this cohort was 51 years (interquartile range [IQR], 36–61 years), and 416 patients (56%) were male. According to the ELN2022 criteria, 122 patients (17%) were classified as having a good prognosis, 71 (10%) as having an intermediate risk, and 544 (74%) as having a poor risk.
[0091] Table 1. Clinically relevant covariates of participants
[0092]
[0093]
[0094] Note: AA: Aplastic anemia, AML: Acute myeloid leukemia, AMLRGA: Acute myeloid leukemia with recurrent genetic abnormalities, AMMRC: Acute myeloid leukemia with cytogenetic abnormalities, AMMRG: Acute myeloid leukemia with minimal residual disease, AMLTP53: Acute myeloid leukemia with TP53 mutation, AZA: Azacitidine, BM: Bone marrow, CAG: Cytarabine combined with aclarubicin and granulocyte colony-stimulating factor, CMML: Chronic myelomonocytic leukemia, CR: Complete myeloid leukemia. Remission, DA: Daunorubicin combined with cytarabine, ELN: European Leukemia Network, HAA: Homoharringtonine combined with cytarabine and aclarubicin, HB: Hemoglobin, HSCT: Hematopoietic stem cell transplantation, IA: Idarubicin combined with cytarabine, ICUS: Idiopathic cytopenic purpura of undetermined significance, IQR: Interquartile range, MDS: Myelodysplastic syndrome, MPN: Myeloproliferative neoplasm, MRD: Minimal residual disease, PB: Peripheral blood, PLT: Platelets, VEN: Veneclare, WBC: White blood cells
[0095] Regarding treatment, 389 patients (53%) received intensive induction therapy, while 348 (47%) received non-intensive induction therapy. The median follow-up was 16 months. 554 patients (75%) achieved complete remission or complete remission with incomplete hematologic recovery (CR / CRi), while the median survival for those who did not achieve CR / CRi was 0.50 months (IQR 0.27–0.83 months). Among the patients who achieved CR / CRi, 286 (47%) underwent allogeneic hematopoietic stem cell transplantation (allo-HSCT), and 144 (35%) experienced disease relapse.
[0096] A total of 255 deaths were recorded, with the following causes of death: 115 (45%) died from relapse, 78 (31%) died from non-remission, 22 (8.6%) died from transplant-related deaths, 18 (7%) died from chemotherapy-related deaths, 10 (4%) died from early death, 5 (2%) died from severe infection, and 3 (1%) died from other causes.
[0097] (2) Gene mapping analysis
[0098] In the internal dataset (Peking University People's Hospital cohort), a total of 4,329 hematologic malignancy-related pathogenic variants were detected in 727 patients (99%), with an average of 6 variants per patient. The most common mutation was RUNX1 mutation (n = 189, 26%), followed by ASXL1 mutation (n = 181, 25%) and NRAS mutation (n = 146, 20%) (see Table 2). The most common mutation subtype was missense mutation. Among MRGA patients in the hospital, 409 (56%) carried only myelodysplastic dysplasia-associated gene mutations (MRG), 214 (29%) carried only myelodysplastic dysplastic cytogenetic abnormalities (MRC), and 114 (15%) carried both MRG and MRC.
[0099] Table 2. Cytogenetic abnormalities and related covariates of participants.
[0100]
[0101] Note: add: Chromosomal fragment addition; ASXL1: Addition comb-like protein 1; ASXL3: Addition comb-like protein 3; BCOR: BCL6 co-repressor; BCORL1: BCL6 co-repressor-like protein 1; CBFB: Core binding factor β; CEBPA: CCAAT / enhancer binding protein α; CEBPAbZIP: CCAAT / enhancer binding protein α basic leucine zipper domain; CREBBP: CREB binding protein; CSF3R: Colony-stimulating factor 3 receptor; DDX41: DEAD box peptide 41; D EK: DEK proto-oncogene; DNMT3A: DNA methyltransferase 3α; EP300: E1A binding protein p300; ETV6: ETS variant transcription factor 6; EZH2: Zeste homolog enhancer 2; FLT3-ITD: FMS-like tyrosine kinase 3 internal tandem repeat; FLT3-TKD: FMS-like tyrosine kinase 3 tyrosine kinase domain; GATA2: GATA binding protein 2; IDH1: Isocitrate dehydrogenase 1; IDH2: Isocitrate dehydrogenase 2; inv: inversion; JAK2: Janus kinase 2; KAT6A: K (lysine) acetyltransferase 6A; KIT: KIT proto-oncogene receptor tyrosine kinase; KMT2A: lysine methyltransferase 2A; KRAS: KRAS proto-oncogene GTPase; MLLT3: myeloid / lymphoid or mixed lineage leukemia translocation to 3; MYH11: myosin heavy chain 11; NPM1: nucleophosphorin 1; NUP214: nucleoporin 214; NRAS: NRAS proto-oncogene GTPase; PHF6: plant homologous domain finger protein 6; PTPN11: non-receptor protein tyrosine phosphatase 11 RUNX1: Runt-related transcription factor 1; RUNX1T1: RUNX1 translocation chaperone 1; SF3B1: splicing factor 3B subunit 1; SRSF2: serine / arginine-enriched splicing factor 2; STAG2: matrix antigen 2; TET2: Tet methylcytosine dioxygenase 2; TP53: tumor protein 53; TPMT: thiopurine methyltransferase; U2AF1: U2 small nuclear RNA cofactor 1; WT1: nephroblastoma protein 1; ZRSR2: zinc finger CCCH type, RNA binding motif and serine / arginine-enriched protein 2.
[0102] Hematologic malignancy-related pathogenic variants were also analyzed in the TCGA and BeatAML datasets. In the TCGA dataset, a total of 1145 variants were detected in all 68 patients (100%), with an average of 17 variants per patient. Similarly, in the BeatAML dataset, a total of 2546 variants were detected in 133 patients (94%), with an average of 19 variants per patient. The most common mutation in the BeatAML dataset was RUNX1, followed by ASXL1 (n = 29, 20%), SRSF2 (n = 23, 16%), and NRAS (n = 23, 16%). The most common mutation in the TCGA dataset was TP53 (n = 16, 24%), followed by DNMT3A (n = 15, 22%), IDH2 (n = 11, 16%), and IDH1 (n = 9, 13%).
[0103] It is worth noting that significant differences (p<0.05) in mutation frequencies of multiple genes (including CEBPAbZIP, ASXL1, SRSF2, PHF6, CSF3R, GATA2, TET2, BCORL1, and EP300) were observed between the training / validation set (Peking University People's Hospital cohort), BeatAML, and TCGA datasets, as shown in Table 3.
[0104] Table 3. Genome maps of the training and validation sets.
[0105]
[0106]
[0107] (3) Cytogenetic abnormality analysis
[0108] G-banding analysis was performed on all 737 patients to assess cytogenetic abnormalities. Of these, 328 patients (45%) had detectable myeloid dysplasia-related cytogenetic abnormalities (MRCs). These were classified according to the ELN 2022 International Consensus Classification of AML (ICC) hierarchy.
[0109] 92 patients (12%) were classified as AML-MRC (i.e., AML that meets the morphological criteria related to MDS);
[0110] 219 cases (30%) were classified as AML with recurrent genetic abnormalities;
[0111] 64 cases (9%) were classified as AML with TP53 mutation;
[0112] 362 cases (50%) were classified as AML-MRG (i.e., AML that only meets the molecular genetic MRG criteria).
[0113] Among the observed cytogenetic abnormalities, complex karyotypes were the most common, occurring in 149 patients (20%); followed by trisomy 8, occurring in 135 patients (18%); and deletion of the long arm of chromosome 7 / complete deletion of chromosome 7 (del(7q) / -7), occurring in 101 patients (14%) (see Table 2).
[0114] In the BeatAML dataset, MRC was detected in 83 patients (58%). Complex karyotypes remained the most common abnormality, seen in 49 patients (35%); followed by deletions / translocations / duplications on the long arm of chromosome 5 (del(5q) / t(5q) / add(5q)) and Trisomy 8, each seen in 30 patients (21%); and del(7q), seen in 17 patients (12%). Similarly, in the TCGA dataset, MRC was detected in 45 patients (66%). Complex karyotypes again became the most common abnormality, seen in 24 patients (35%); followed by del(7q) / -7, seen in 20 patients (29%); Trisomy 8, seen in 19 patients (28%); and del(5q) / t(5q) / add(5q) / -5, seen in 15 patients (22%).
[0115] Significant differences (p<0.05) were observed in the incidence of the following anomalies among the training / validation sets (referring to the Peking University People's Hospital cohort), BeatAML, and TCGA datasets: complex karyotypes, del(7q) / -7, del(5q) / t(5q) / add(5q) / -5, and short arm duplication / deletion / complete deletion of chromosome 17 / 17q alleles, with long arm duplication and short arm loss (add(17p) / del(17p) / -17 / i(17q)). In contrast, no significant differences were found in the incidence of Trisomy 8, short arm deletion / translocation / duplication of chromosome 12 (del(12p) / t(12p) / add(12p)) or long arm deletion of chromosome 20 (del(20q)) among the datasets (p>0.05).
[0116] (3) Construction of treatment response model
[0117] No significant differences were observed between the training and validation sets in terms of baseline clinical covariates, genetic abnormalities, follow-up time, treatment response, and outcomes. Furthermore, no interactions were observed between covariates (including sex, white blood cell (WBC) count, percentage of bone marrow blasts, cytogenetic characteristics, and genetic variations) in the training dataset (variance inflation factor = 1.0–1.6). An optimal cutoff was determined using the maximum Youden index: a white blood cell count of 10 × 10⁻⁶. 9 / L, the percentage of bone marrow blasts is 45%.
[0118] The clinically relevant covariates and cytogenetic abnormality-related covariates mentioned above were included in the Lasso regression model for screening. The results showed that in the Lasso regression analysis, the following factors were significantly associated with a lower complete remission / complete remission with hematologic incomplete recovery (CR / CRi) rate: male sex, add(17p) / del(17p) / i(17q) / -17, del(7q) / -7, complex karyotype, inv(3) or t(3;3), single chromosome karyotype, KRAS mutation, ETV6 mutation, NF1 mutation, SETBP1 mutation, and U2AF1 mutation; conversely, the following factors were associated with a higher CR / CRi rate: white blood cell count <10×10 9 / L, percentage of bone marrow blasts ≥45%, t(8;21), BCORL1 mutation, NPM1 mutation, CEBPAbZIP gene mutation, or receiving intensive induction therapy.
[0119] The association between relevant covariates and the complete remission rate (CR / CRi) in AML patients can be distinguished by the p-value in Table 2 or by the weighting coefficients of the covariates. A smaller p-value and a higher weighting coefficient indicate a stronger association between the indicator and CR / CRi. Multivariate Cox regression analysis of the variables selected by Lasso regression showed that the following factors were significantly associated with a lower CR / CRi rate: white blood cell count ≥10 × 10⁻⁶. 9 / L, percentage of bone marrow blasts <45%, del(7q) / -7, complex karyotype, inv(3) or t(3;3), U2AF1 mutation; the following factors are associated with higher CR / CRi rates: t(8;21), NPM1 mutation, CEBPAbZIP mutation, and intensive induction therapy regimen (see Table 4).
[0120] The treatment response model uses the regression coefficient β to construct a score, with the formula: CR Score = 0.843 × white blood cell count (if < 10 × 10⁻⁶). 9 / L,=1;otherwise,=0)+0.757×% of bone marrow blasts (if≥45%,=1;otherwise,=0)+0.608×intensive induction therapy (if intensive induction therapy=1;otherwise,=0)+2.024×t(8;21)(if t(8;21),=1;otherwise,=0)+0.701×del(7q) / -7(if del(7q) / -7,=0;otherwise,=1)+1.011×complex karyotype (if complex karyotype=0;otherwise,=1)+2.098×inv(3) or t(3;3)(if inv(3) or t(3;3),=0;otherwise,=1)+1.742×NPM1(if mutated,=1;otherwise,=0)+1.385×CEBPAbZIP(if mutated,=1; otherwise,=0)+0.792×U2AF1(ifmutated,=0;otherwise,=1)+1.16.
[0121] Divide the regression coefficient β of each covariate by the smallest regression coefficient (i.e., 0.608), round to the nearest integer, and use this as the weight for constructing the treatment response model. The simplified formula is: CR Score = 1 × white blood cell count (if < 10 × 10⁻⁶). 9 / L,=1;otherwise,=0)+1×Bone marrow blast percentage(if≥45%,=1;otherwise,=0)+1×Intensive induction therapy(ifIntensive induction therapy,=1;otherwise,=0)+3×t(8;21)(if t(8;21),=1;otherwise,=0)+1×del(7q) / -7(if del(7q) / -7,=0;otherwise,=1)+2×Complex karyotype(if Complex karyotype,=0;otherwise,=1)+3×inv(3)ort(3;3)(if inv(3)ort(3;3),=0;otherwise,=1)+3×NPM1(if mutated,=1;otherwise,=0)+2×CEBPAbZIP(if mutated,=1;otherwise,=0)+1×U2AF1(if mutated,=0;otherwise,=1).
[0122] Table 4 shows the treatment response model for AML with MRGA constructed using Cox multivariate regression analysis.
[0123]
[0124]
[0125] (4) Validation of the treatment response model
[0126] Based on the weighted sum of predictors in the treatment response model CR Score, the training set was divided into two subgroups: a low response subgroup (score ≤ 8, n = 259): CR / CRi rate of 56%; and a high response subgroup (score ≥ 9, n = 232): CR / CRi rate of 91%, with a significant difference between the two groups (p < 0.01).
[0127] In the internal validation dataset, 53% of patients (n=130) were classified as the low-response subgroup (score ≤8) and 47% of patients (n=116) were classified as the high-response subgroup (score ≥9). There was a significant difference in CR / CRi rates between the two subgroups (p<0.01).
[0128] In the external validation dataset (BeatAML), 58% of patients (n=63) were classified as the low-response subgroup and 42% of patients (n=46) were classified as the high-response subgroup. There was a strong trend of difference in CR / CRi rates between the two subgroups (48% in the low-response group vs. 70% in the high-response group), but it did not reach statistical significance (p=0.106).
[0129] Because the TCGA dataset does not report treatment details or response outcomes, this validation could not be applied to that dataset. The area under the receiver operating characteristic (AUC) values for the training set, internal validation set, and BeatAML dataset were 0.79, 0.78, and 0.63, respectively.
[0130] The validation results of the internal training set, validation set, and external validation set were highly consistent with the follow-up results (median follow-up time was 16 months), indicating that the treatment response model constructed in this invention can be used to predict the probability of MRGAAML patients achieving complete remission after receiving induction therapy. Patients in the high response group had a higher probability of achieving complete remission, while those in the low response group had a lower probability of achieving complete remission, demonstrating the good predictive performance of the model.
[0131] (5) Comparison of treatment response models constructed with different combinations of covariates
[0132] To verify the superior performance of the treatment response model constructed based on covariates selected by Lasso regression and Cox regression analysis in this embodiment, this embodiment also used the internal training set and validation set (Peking University People's Hospital cohort) as samples to compare the performance differences of treatment response models constructed with the following 5 combinations of covariates.
[0133] Table 5. Performance comparison of treatment response models constructed with different combinations of covariates.
[0134]
[0135] The results in Table 5 show that different models have different AUC values for the same dataset. Compared with models 2, 3, and 4, model 1 has the highest AUC values on the training and validation sets, at 0.79 and 0.78, respectively. This indicates that the treatment response model constructed based on the 10 covariates selected in this invention has the best predictive performance on both datasets. The predictive performance of models constructed using fewer or more than 10 covariates or other covariates is reduced.
[0136] Example 2: Construction of the survival risk scoring model for AML with MRGA provided by the present invention
[0137] I. Construction of the Survival Risk Scoring Model
[0138] Effective risk stratification is crucial for AML with MRGA. Currently, when individual risk is assessed and categorized as adverse risk, there are significant differences in outcomes between subgroups with specific genetic abnormalities. According to the risk stratification of the European Leukemia Net 2022 (ELN 2022), AML with MRGA cannot distinguish between intermediate and adverse risk categories.
[0139] Based on the research objects and methods of Example 1, this example initially collected 458 survival-related covariates in the training set and then performed stepwise screening:
[0140] First, genes with an incidence rate of <5% were excluded. Then, Lasso regression was applied to remove redundant prognostic variables based on coefficients and partial likelihood deviance. After Lasso regression screening, 34 covariates were retained for Cox multivariate analysis. Following Cox multivariate analysis, 22 variables with p-values <0.1 were considered for inclusion in the survival risk scoring model: male sex, age, del(5q) / t(5q) / add(5q), -7 / del(7q), trisomy 8, t(9;11), inv(3) or t(3;3), complex karyotype, t(v;11q23.3), inv(16) or t(16;16), t(8;21), ASXL1 mutation, DNMT3A mutation, TET2 mutation, KRAS mutation, ETV6 mutation, FLT3-ITD mutation, TP53 mutation, GATA2 mutation, BCOR mutation, EP300 mutation, and CEBPAbZIP mutation (see Table 6).
[0141] Table 6. Construction of the Survival Risk Score Model for AML with MRGA
[0142]
[0143]
[0144] The strength of the association between relevant covariates and the overall survival of AML patients can be distinguished by the p-value or by the weight coefficient of the covariates. The smaller the p-value and the higher the weight coefficient, the stronger the association of the indicator with the overall survival. The treatment response model constructs a score value using the regression coefficient β, with the formula RiskScore = 0.529 × male (if male, × 1; otherwise, × 0) + age (if Age ≤ 35, = 0; 35 < Age < 60, = 0.746; Age ≥ 60, = 1.347) + 0.551 × del(5q) / t(5q) / add(5q) (if del(5q) / t(5q) / add(5q), = 1; otherwise, = 0) + 0.424 × del(7q) / -7 (if del(7q) / -7, = 1; otherwise, = 0) + 0.373 × Trisomy 8 (if Trisomy 8, = 1; otherwise, = 0) + 0.809 × t(9;11) (if t(9;11), = 1; otherwise, = 0) + 0.759 × inv(3) or t(3;3) (if inv(3) or t(3;3), = 1; otherwise, = 0) + 0.691 × complex karyotype (if complex karyotype, = 1; otherwise, = 0) + 1.076 × t(v;11q23.3) (if t(v;11q23.3), = 1; otherwise, = 0) + 1.459 × inv(16) or t(16;16) (if inv(16) or t(16;16), = 0; otherwise, = 1) + 1.784 × t(8;21) (if t(8;21), = 0; otherwise, = 1) + 0.389 × ASXL1 (if mutated, = 1; otherwise, = 0) + 0.464 × DNMT3A (if mutated, = 1; otherwise, = 0) + 0.472 × TET2 (if mutated, = 1; otherwise, = 0) + 0.585 × KRAS (if mutated, = 1; otherwise, = 0) + 0.590 × ETV6 (if mutated, = 1; otherwise, = 0) + 0.591 × FLT3-ITD (if mutated, = 1; otherwise, = 0) + 0.615 × TP53 (if mutated, = 1; otherwise, = 0) + 0.729 × GATA2 (if mutated, = 1; otherwise, = 0) + 0.940×BCOR(if mutated,=0;otherwise,=1)+0.999×EP300(ifmutated,=0;otherwise,=1)+1.799×CEBPAbZIP(if mutated,=0;otherwise,=1)。.
[0145] Round the regression coefficient β of each covariate to the nearest integer after dividing it by the smallest regression coefficient (i.e., 0.373), and use it as the weight for constructing the treatment response model. The simplified formula is RiskScore = 1 × male (if male, = 1; otherwise, = 0) + age (if Age ≤ 35, = 0; 35 < Age < 60, = 2; Age ≥ 60, = 4) + 1 × del(5q) / t(5q) / add(5q) (if del(5q) / t(5q) / add(5q), = 1; otherwise, = 0) + 1 × del(7q) / -7 (if del(7q) / -7, = 1; otherwise, = 0) + 1 × Trisomy 8 (if Trisomy 8, = 1; otherwise, = 0) + 2 × t(9;11) (if t(9;11), = 1; otherwise, = 0) + 2 × inv(3) or t(3;3) (if inv(3) or t(3;3), = 1; otherwise, = 0) + 2 × complex karyotype (if complex karyotype, = 1; otherwise, = 0) + 3 × t(v;11q23.3) (if t(v;11q23.3), = 1; otherwise, = 0) + 4 × inv(16) or t(16;16) (if inv(16) or t(16;16), = 0; otherwise, = 1) + 5 × t(8;21) (if t(8;21), = 0; otherwise, = 1) + 1 × ASXL1 (if mutated, = 1; otherwise, = 0) + 1 × DNMT3A (if mutated, = 1; otherwise, = 0) + 1 × TET2 (if mutated, = 1; otherwise, = 0) + 2 × KRAS (if mutated, = 1; otherwise, = 0) + 2 × ETV6 (if mutated, = 1; otherwise, = 0) + 2 × FLT3-ITD (if mutated, = 1; otherwise, = 0) + 2 × TP53 (if mutated, = 1; otherwise, = 0) + 2 × GATA2 (if mutated, = 1; otherwise, = 0) + 3 × BCOR (if mutated, = 0; otherwise, = 1) + 3 × EP300 (if mutated, = 0; otherwise, = 1) + 5 × CEBPAbZIP (if mutated, = 0; otherwise, = 1).
[0146] (2) Validation of the survival risk score model
[0147] The scores were calculated based on the survival risk scoring model formula, and survival curves were plotted using the Kaplan-Meier method. The actual overall survival (OS) differences between different risk groups were compared using the Log-rank test. Results showed that the training set was divided into three risk subgroups: low-risk group (score ≤ 24; n = 412, 56%), medium-risk group (28 ≥ score > 25; n = 253, 34%), and high-risk group (score > 29; n = 71, 10%). Figure 3 A). See the gene profiles, genetic abnormalities, and clinical information heatmaps for the three risk subgroups. Figure 2 Using the low-risk group as a reference, the overall survival (OS) hazard ratios (HRs) for the medium-risk and high-risk groups were 3.455 (95% confidence interval [CI]: 2.699, 4.422, p = 0.000) and 9.228 (95% CI: 6.674, 12.759, p < 0.001), respectively. Figure 3 B). Significant differences existed in 30-month survival rates among the three risk subgroups (low-risk group: 29% [n=81 survived] vs. intermediate-risk group: 9% [n=15 survived] vs. high-risk group: 0% [n=0 survived], p<0.001), as... Figure 3 As shown in C.
[0148] In the internal validation set, 135 patients (55%) were assigned to the low-risk group, 84 patients (34%) to the intermediate-risk group, and 27 patients (11%) to the high-risk group. Significant differences in 30-month survival rates were also observed among the three subgroups (low-risk group: 27% [n=37 survived] vs. intermediate-risk group: 8% [n=7 survived] vs. high-risk group: 4% [n=1 survived], p<0.001). Figure 3 D). The clinical outcomes of 737 patients classified by OS risk category are shown in Table 7.
[0149] Table 7 Summary of clinical outcomes of 737 patients classified by OS risk.
[0150]
[0151] In the external validation set BeatAML, 41 participants (28%) were classified as low-risk, 86 participants (59%) as intermediate-risk, and 20 participants (14%) as high-risk. Significant survival differences were also found among the three risk subgroups of BeatAML (p = 0.003). Figure 3 E).
[0152] In the external validation set TCGA, 10 participants (14%) were classified as low-risk, 29 as intermediate-risk, and 32 as high-risk. Due to the limited sample size of TCGA, the low-risk and intermediate-risk groups were combined and compared with the high-risk group. The results showed a significant survival difference between the two groups (p = 0.05).
[0153] The validation results from the internal training set, validation set, and external validation set show that the survival risk scoring model constructed in this invention predicts the lowest actual survival rate for the high-risk group and the highest actual survival rate for the low-risk group, and its risk grouping and predicted survival rate are consistent with the actual results.
[0154] To explore the ability of the survival scoring system to predict relapse-free survival (RFS), this invention also validated its predictive performance on RFS in a training cohort, an internal validation cohort, and an external validation cohort (BeatAML). The results show that the survival scoring system also possesses some predictive power for predicting disease relapse.
[0155] Example 3: Comparison of Treatment Response Model and Survival Risk Score Model with ELN 2022 Standards
[0156] like Figure 3 As shown in F, compared with the ELN 2022 risk classification, the survival risk scoring model and treatment response model (CR score) provided by this invention have higher Harrell C indices, indicating that the two sets of models provided by this invention have improved classification efficacy for overall survival (OS), relapse-free survival (RFS), and complete remission (CR) risk in patients with acute myeloid leukemia.
[0157] By mapping the three risk categories (good, intermediate, poor) of the ELN 2022 criteria to the three risk categories (low, intermediate, high) of the survival risk scoring system of this invention, the survival risk scoring model provided by this invention significantly enhances its ability to distinguish between various endpoint events (OS, RFS, CR). This restratification covered 71% of patients (620 out of 871), of whom 597 (69%) were risk-upgraded (i.e., assigned to a higher risk group) and 23 (4%) were risk-downgraded (i.e., assigned to a lower risk group) (see...). Figure 4 GH).
[0158] It is noteworthy that, within the same risk stratum of ELN 2022, significant differences in outcomes were still observed between different risk subgroups classified by the survival risk scoring system provided by this invention. This indicates that the scoring system is able to identify patient subgroups with different prognoses within the ELN strata.
[0159] Example 4: Performance Evaluation of Survival Risk Scoring Model
[0160] This invention further analyzes the response of acute myeloid leukemia (AML) patients to different induction chemotherapy regimens, validates the efficacy of the survival risk score model, and validates and refines the prognostic model by grouping the cohort according to chemotherapy intensity (intensive treatment vs. non-intensive treatment). Kaplan-Meier analysis was used to assess the prognostic correlation between risk groups (low risk L, intermediate risk M, high risk H) in the two treatment regimens.
[0161] The results showed that survival rates were significantly stratified in the intensive treatment group, with high-risk (H) patients having the worst prognosis, followed by intermediate-risk (M) and low-risk (L) patients. Figure 4 A similar significant stratification pattern was observed in the non-intensive treatment group (AC). Figure 4 This confirms the clinical applicability of risk stratification (DF).
[0162] Furthermore, the ROC curves of the survival risk scoring model's predictive ability for 1-year, 2-year, and 3-year survival rates on both the intensive treatment training set and the intensive treatment validation set, as well as on both the non-intensive treatment training set and the non-intensive treatment validation set, show (…). Figure 4 The AUC values of the intensive treatment group (G-4J) were 0.833, 0.850, and 0.876 on the training set and 0.705, 0.778, and 0.742 on the validation set. Figure 4 GH); the AUC values of the non-intensive treatment group were 0.821, 0.859, and 0.880 on the training set, and 0.767, 0.679, and 0.653 on the validation set. Figure 5 The ROC curves of 1-, 2-, and 3-year survival rates on the training and validation sets show that the survival risk scoring model constructed in this invention has good sensitivity and specificity. The 1-, 2-, and 3-year survival AUC values for all datasets are 0.789, 0.807, and 0.831, respectively, while the 1-, 2-, and 3-year survival AUC values for the training set are 0.827, 0.850, and 0.873, respectively. Figure 5 B), the corresponding AUC values for the internal validation set are 0.749, 0.761, and 0.783. Figure 5 C). The AUROC values for the 1-, 1.5-, and 2-year survival tests on the BeatAML validation set were 0.644, 0.681, and 0.658, respectively. Figure 5 D). Calibration curves for 1-year and 2-year survival probabilities show a high degree of agreement between predicted and observed values. Figure 5 The result (EH) indicates that the model can accurately predict the occurrence of outcome events. Decision curve analysis shows that the survival risk score model has a clinically practical net benefit in predicting 1-, 2-, and 3-year survival (EH). IL).
[0163] This embodiment further compares the predictive performance of the survival risk scoring model constructed based on the covariates selected in Embodiment 2 with the survival model constructed with other combinations of covariates. The covariates for constructing models 5-8 are shown in Table 8. The predictive performance of the four models is evaluated using the internal training set and validation set.
[0164] Table 8. Performance Comparison of Survival Risk Scoring Models Constructed with Different Combinations of Covariates
[0165]
[0166]
[0167] The results in Table 8 show that the survival risk scoring model (Model 5) constructed in this invention has the best sensitivity and specificity. The AUC values of the 1-year, 2-year, and 3-year survival rates in the training set are 0.827, 0.850, and 0.873, respectively, and the corresponding AUC values in the validation set are 0.749, 0.761, and 0.783, which are significantly higher than the other three models. This indicates that the predictive performance of survival models constructed using other combinations of covariates is not as good as the survival risk scoring model constructed using the 22 covariates selected in this invention.
[0168] While the present invention has been disclosed above, it is not limited thereto. Any person skilled in the art can make various modifications and alterations without departing from the spirit and scope of the invention; therefore, the scope of protection of the present invention should be determined by the scope defined in the claims.
Claims
1. The use of a biomarker in the preparation of a reagent for predicting overall survival in acute myeloid leukemia with associated genetic abnormalities of myelodysplastic syndromes, characterized in that, This includes any one or more combinations of male, age, del(5q) / t(5q) / add(5q), del(7q) / -7, trisomy of chromosome 8, t(9;11), inv(3) or t(3;3), complex karyotype, t(v;11q23.3), inv(16) or t(16;16), t(8;21), ASXL1 mutation, DNMT3A mutation, TET2 mutation, KRAS mutation, ETV6 mutation, FLT3-ITD mutation, TP53 mutation, GATA2 mutation, BCOR mutation, EP300 mutation, and CEBPAbZIP mutation.
2. The use as described in claim 1, characterized in that, The markers include male sex, age, del(5q) / t(5q) / add(5q), del(7q) / -7, trisomy of chromosome 8, t(9;11), inv(3) or t(3;3), complex karyotype, t(v;11q23.3), inv(16) or t(16;16), t(8;21), ASXL1 mutation, DNMT3A mutation, TET2 mutation, KRAS mutation, ETV6 mutation, FLT3-ITD mutation, TP53 mutation, GATA2 mutation, BCOR mutation, EP300 mutation, and a combination of CEBPAbZIP mutation.
3. A survival risk scoring model for predicting overall survival in acute myeloid leukemia with myelodysplastic syndrome-related genetic abnormalities, characterized in that, Constructed based on the combination of markers as described in any one of claims 1-2.
4. The survival risk scoring model as described in claim 3, characterized in that, The formula of the survival risk scoring model is: RiskScore = 0.529 × male (if male, = 1; otherwise, = 0) + age (if Age ≤ 35, = 0; 35 < Age < 60, = 0.746; Age ≥ 60, = 1.347) + 0.551 × del(5q) / t(5q) / add(5q) (if del(5q) / t(5q) / add(5q), = 1; otherwise, = 0) + 0.424 × del(7q) / -7 (if del(7q) / -7, = 1; otherwise, = 0) + 0.373 × Trisomy 8 (if Trisomy 8, = 1; otherwise, = 0) + 0.809 × t(9;11) (if t(9;11), = 1; otherwise, = 0) + 0.759 × inv(3) or t(3;3) (if inv(3) or t(3;3), = 1; otherwise, = 0) + 0.691 × complex karyotype (if complex karyotype, = 1; otherwise, = 0) + 1.076 × t(v;11q23.3) (if t(v;11q23.3), = 1; otherwise, = 0) + 1.459 × inv(16) or t(16;16) (if inv(16) or t(16;16), = 0; otherwise, = 1) + 1.784 × t(8;21) (if t(8;21), = 0; otherwise, = 1) + 0.389 × ASXL1 (if mutated, = 1; otherwise, = 0) + 0.464 × DNMT3A (if mutated, = 1; otherwise, = 0) + 0.472 × TET2 (if mutated, = 1; otherwise, = 0) + 0.585 × KRAS (if mutated, = 1; otherwise, = 0) + 0.590 × ETV6 (if mutated, = 1; otherwise, = 0) + 0.591 × FLT3-ITD (if mutated, = 1; otherwise, = 0) + 0.615 × TP53 (if mutated, = 1; otherwise, = 0) + 0.729 × GATA2 (if mutated, = 1; otherwise, = 0) + 0.940 × BCOR (if mutated, = 0; otherwise, = 1) + 0.999 × EP300 (if mutated, = 0; otherwise, = 1) + 5. The method for constructing a survival risk scoring model for predicting overall survival in acute myeloid leukemia with myelodysplastic syndrome-related genetic abnormalities as described in any one of claims 3-4, characterized in that, Includes the following steps: (1) Collect patients’ clinical data, chromosome karyotype and gene mutation information; (2) Biomarkers were screened using Lasso and Cox regression analysis; (3) Construct a survival risk scoring model based on the screened biomarkers.
6. The use of a combination of biomarkers to construct a survival risk scoring model for predicting overall survival in acute myeloid leukemia with myelodysplastic syndrome-related genetic abnormalities, characterized in that, The biomarker combination includes male sex, age, del(5q) / t(5q) / add(5q), del(7q) / -7, trisomy of chromosome 8, t(9;11), inv(3) or t(3;3), complex karyotype, t(v;11q23.3), inv(16) or t(16;16), t(8;21), ASXL1 mutation, DNMT3A mutation, TET2 mutation, KRAS mutation, ETV6 mutation, FLT3-ITD mutation, TP53 mutation, GATA2 mutation, BCOR mutation, EP300 mutation, and CEBPAbZIP mutation.
7. The use as described in claim 6, characterized in that, The formula of the survival risk scoring model is: RiskScore = 0.529 × male (if male, = 1; otherwise, = 0) + age (if Age ≤ 35, = 0; 35 < Age < 60, = 0.746; Age ≥ 60, = 1.347) + 0.551 × del(5q) / t(5q) / add(5q) (if del(5q) / t(5q) / add(5q), = 1; otherwise, = 0) + 0.424 × del(7q) / -7 (if del(7q) / -7, = 1; otherwise, = 0) + 0.373 × Trisomy 8 (if Trisomy 8, = 1; otherwise, = 0) + 0.809 × t(9;11) (if t(9;11), = 1; otherwise, = 0) + 0.759 × inv(3) or t(3;3) (if inv(3) or t(3;3), = 1; otherwise, = 0) + 0.691 × complex karyotype (if complex karyotype, = 1; otherwise, = 0) + 1.076 × t(v;11q23.3) (if t(v;11q23.3), = 1; otherwise, = 0) + 1.459 × inv(16) or t(16;16) (if inv(16) or t(16;16), = 0; otherwise, = 1) + 1.784 × t(8;21) (if t(8;21), = 0; otherwise, = 1) + 0.389 × ASXL1 (if mutated, = 1; otherwise, = 0) + 0.464 × DNMT3A (if mutated, = 1; otherwise, = 0) + 0.472 × TET2 (if mutated, = 1; otherwise, = 0) + 0.585 × KRAS (if mutated, = 1; otherwise, = 0) + 0.590 × ETV6 (if mutated, = 1; otherwise, = 0) + 0.591 × FLT3-ITD (if mutated, = 1; otherwise, = 0) + 0.615 × TP53 (if mutated, = 1; otherwise, = 0) + 0.729 × GATA2 (if mutated, = 1; otherwise, = 0) + 0.940 × BCOR (if mutated, = 0; otherwise, = 1) + 0.999 × EP300 (if mutated, = 0; otherwise, = 1) + 1.799 × CEBPAbZIP (if mutated, = 0; otherwise, = 1).
8. A system for predicting overall survival in acute myeloid leukemia with associated genetic abnormalities of myelodysplastic syndromes, characterized in that, It includes a data analysis module, which is used to analyze the detection value of the marker as described in any one of claims 1-2.
9. The system as described in claim 8, characterized in that, The system also includes a RiskSore calculation module and a result output module. The RiskSore calculation module calculates a score based on the detection value of the marker, and the result output module outputs low-risk, medium-risk, and high-risk levels based on the score.
10. A combination of biomarkers for predicting overall survival in acute myeloid leukemia with associated genetic abnormalities of myelodysplastic syndromes, characterized in that, The biomarker combination includes male sex, age, del(5q) / t(5q) / add(5q), del(7q) / -7, trisomy of chromosome 8, t(9;11), inv(3) or t(3;3), complex karyotype, t(v;11q23.3), inv(16) or t(16;16), t(8;21), ASXL1 mutation, DNMT3A mutation, TET2 mutation, KRAS mutation, ETV6 mutation, FLT3-ITD mutation, TP53 mutation, GATA2 mutation, BCOR mutation, EP300 mutation, and CEBPAbZIP mutation.