Chronic granulocytic leukemia risk level layering evaluation model and application
By constructing a risk stratification assessment model for CMML based on the Chinese population and using specific biomarkers to calculate risk scores, the problem of unclear applicability of existing models to the Chinese population is solved, enabling precise risk assessment and individualized management, and improving the accuracy and effectiveness of treatment decisions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANDONG UNIV QILU HOSPITAL
- Filing Date
- 2026-04-14
- Publication Date
- 2026-05-12
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The applicability of existing CMML prognostic models to the Chinese population is unclear, and the lack of a precise risk assessment system leads to difficulties in treatment decision-making, especially limiting their guiding role for intermediate-risk patients.
A risk stratification assessment model for chronic myelomonocytic leukemia was constructed. Peripheral blood platelet count, chromosome, U2AF1 gene, ASXL1_exon12_19721ul gene and SF3B1 gene were used as biomarkers. A risk score was calculated using a specific formula to classify patients into standard risk and high risk categories to guide individualized management.
It provides accurate risk assessment based on Chinese population data, simplifies treatment decisions, improves predictive efficacy, avoids missing treatment opportunities, and prolongs patient survival.
Smart Images

Figure CN122025191A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of bioinformatics and biomarker technology, and relates to a risk stratification assessment model for chronic myelomonocytic leukemia (CMML) and its application. Background Technology
[0002] Chronic myelomonocytic leukemia (CMML) is a malignant clonal disease characterized by both myelodysplastic syndromes and myeloproliferative neoplasms, exhibiting highly heterogeneous clinical presentations and prognoses. It is a rare hematologic malignancy with an incidence of approximately 0.4-1 cases per 100,000 people per year, primarily affecting the elderly. Survival outcomes for CMML patients vary greatly; some require only observation or oral medication after diagnosis, while others necessitate demethylation therapy or even hematopoietic stem cell transplantation.
[0003] Currently, there is a lack of a dedicated and objective scoring system for initiating treatment in CMML patients. Clinical decision-making mainly relies on prognostic stratification, with commonly used prognostic models including IPSS, CPSS, CPSS-Mol, the Mayo molecular model, and iCPSS. However, their area under the receiver operating characteristic (AUC) is only about 0.65, indicating unsatisfactory predictive efficacy. Existing scoring systems are numerous and inconsistent in their standards, and many are based on data from European and American populations, leading to difficulties in clinical selection. For example, CPSS-Mol only includes four gene mutations, while iCPSS covers ten, and there is a lack of unified standards in the weighting of parameters such as the proportion of primitive cells and chromosomal abnormalities. Furthermore, the applicability of these models to specific populations (such as those with myelofibrosis or oligomonocytic CMML) is unclear, limiting their guidance for treatment decisions—they are mostly used to predict survival and the risk of leukemia transformation, especially for "intermediate-risk" patients, where the decision to treat or not treat still requires clinical experience, making it difficult to provide clear decision guidance. Therefore, for patients diagnosed with CMML, there is an urgent clinical need for a risk stratification prognostic scoring system and model based on Chinese population data that is universal, efficient, and accurate, so as to provide new reference for the development of related products that help determine the timing of treatment initiation. Summary of the Invention
[0004] This invention addresses the problems existing in traditional CMML prognostic models by proposing a risk stratification assessment model for chronic myelomonocytic leukemia and its application.
[0005] To achieve the above objectives, the present invention is implemented using the following technical solution: A risk stratification assessment model for chronic myelomonocytic leukemia is provided, wherein the model is constructed using biomarker detection status as input variables; the biomarker detection status consists of peripheral blood platelet count, chromosome, U2AF1 gene, ASXL1_exon12_19721ul gene, and SF3B1 gene mutation.
[0006] Preferably, the model uses the following formula for risk scoring: Risk Score = (0.69 × peripheral blood platelet count) + (0.22 × score corresponding to chromosome alteration) + (0.64 × score corresponding to U2AF1 gene mutation) + (-1.37 × score corresponding to SF3B1 gene mutation) + (0.75 × score corresponding to ASXL1_exon12_19721ul gene mutation). Wherein, a peripheral blood platelet count ≥ 100 × 10⁻⁶ is considered a risk score. 9 / L is 0 points, and the peripheral blood platelet count is between 20 and 100 × 10⁹ / L. 9 A score of 1 is given for each interval between / L (excluding endpoints), and the peripheral blood platelet count is ≤20×10. 9 / L is scored as 2 points; normal chromosomes and -Y are scored as 0 points; -7, -5, +8, and complex karyotypes (≥3 abnormalities) are scored as 2 points; all other chromosomal changes are scored as 1 point; no mutation in the ASXL1_exon12_19721ul gene is scored as 0 points, and mutation is scored as 1 point; no mutation in the U2AF1 gene is scored as 0 points, and mutation is scored as 1 point; no mutation in the SF3B1 gene is scored as 0 points, and mutation is scored as 1 point. When the risk score is ≤1, the prognosis of CMML is standard risk; when the risk score is >1, the prognosis of CMML is high risk.
[0007] In another aspect, this invention proposes a prognostic risk marker for chronic myelomonocytic leukemia, which consists of peripheral blood platelets, chromosomes, the U2AF1 gene, the ASXL1_exon12_19721ul gene, and the SF3B1 gene.
[0008] The application process of the above model is as follows: For newly diagnosed CMML patients, peripheral blood is collected, bone marrow is aspirated, and bone marrow fluid is tested for genes. Simultaneously, core indicators of the patient are collected, including peripheral blood platelet count, chromosome analysis, and mutation detection of specific genes: U2AF1, ASXL1_exon12_19721ul, and SF3B1. Based on these indicators, an individualized risk score for the patient is calculated using the risk scoring model described in this invention, which is the sum of the products of the weights of each indicator and the corresponding scores of the detected results. Based on the score results, stratified management is implemented, with a total score of 1 as the dividing line: high-risk patients (greater than 1 point) receive timely targeted treatment, while standard-risk patients (less than or equal to 1 point) are managed with regular follow-up and close observation.
[0009] This invention, based on clinical data from the Chinese population, provides a novel prognostic stratification model that accurately integrates clinical characteristics and gene mutations, helping clinicians achieve precise risk assessment and individualized management. By constructing a CMML risk stratification model based on multi-dimensional data and validating its relationship with survival prognosis, it can effectively and accurately determine the patient's risk level. Treatment can be guided based on risk stratification, prompting clinicians to initiate demethylation therapy as early as possible. Compared to traditional risk assessment models, this model is simpler and more efficient at predicting disease progression, thereby assisting clinicians in providing appropriate treatment methods or approaches, preventing patients from missing treatment opportunities, and effectively prolonging patient survival.
[0010] Compared with the prior art, the advantages and positive effects of the present invention are as follows: The risk assessment model constructed in this invention is based on Chinese population data. Compared to traditional assessment methods, which are mostly based on statistical data from European and American populations, this model can more accurately determine the disease status of Chinese patients and predict prognosis and survival. It is also simple and efficient to operate, addressing the shortcomings of traditional models, such as unclear applicability to the Chinese population and weak treatment guidance, thus improving predictive efficacy. Innovatively, it classifies patient risk into only two categories: "standard risk" and "high risk." This simplified design eliminates the ambiguity in decision-making between "observation and waiting" and "overtreatment" for intermediate-risk patients found in traditional models. This provides new possibilities for the development of products that assess accuracy. Attached Figure Description
[0011] Figure 1 This is a survival curve predicted based on a risk model.
[0012] Figure 2 The ROC curve for the first-fold validation set of the model.
[0013] Figure 3 The ROC curve for the second-fold validation set of the model.
[0014] Figure 4 The ROC curve for the third-fold validation set of the model.
[0015] Figure 5 The ROC curve for the fourth-fold validation set of the model.
[0016] Figure 6 The ROC curve for the 5th fold validation set of the model.
[0017] Figure 7 The ROC curve for the sixth-fold validation set of the model.
[0018] Figure 8 The ROC curve for the 7th fold validation set of the model.
[0019] Figure 9 The ROC curve for the 8th-fold validation set of the model.
[0020] Figure 10 The ROC curve for the 9th-fold validation set of the model.
[0021] Figure 11 The ROC curve for the 10th fold validation set of the model. Detailed Implementation
[0022] To better understand the above-mentioned objectives, features, and advantages of the present invention, the present invention will be further described below with reference to specific embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.
[0023] Numerous specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways than those described herein, and therefore the invention is not limited to the specific embodiments disclosed in the following specification. Example 1
[0024] Data from 156 diagnosed CMML patients at Qilu Hospital of Shandong University over the past 10 years were collected, including clinical data, blood routine and bone marrow sample results, and next-generation sequencing data. Due to the limited sample size, a unified feature screening process was first completed based on the entire dataset. K-fold cross-validation was then used to evaluate the model's generalization performance, ensuring stable and reliable modeling results with a small sample size. Rv.4.4.0 software was used for analysis. Kaplan-Meier survival curves and receiver operating characteristic (ROC) curves were obtained using the "survival", "survminer", and "pROC" software packages, respectively. Lasso-Cox regression and feature screening were performed using the "glmnet" package. Survival time, outcome events, and clinical parameters were incorporated into Cox proportional hazards regression to identify independent factors influencing the survival time of CMML patients.
[0025] First, iterative outlier removal was performed on the entire dataset: based on all candidate clinical and biological indicators, the Martingale residuals were calculated using an iterative fitting of the basic Cox model. The sample with the largest absolute residual value (the most outlier sample) was identified, and the new dataset after removing outliers was temporarily stored (finally, a dataset of 143 cases remained). On the new dataset, Lasso-Cox was used to quickly filter features and construct scores.
[0026] For the data after processing outlier samples, potential parameters were selected as predictive factors from clinical and laboratory data. Cross-validation was used to calculate the regularization parameter of the Lasso-Cox model, and the value that minimized the cross-validation error was chosen to balance model fit and complexity. The optimally fitted Lasso-Cox model was then used to compress the regression coefficients of irrelevant variables to 0, initially retaining candidate indicators with predictive value.
[0027] The variables selected by Lasso were then incorporated into a multivariate Cox proportional hazards regression. Using P < 0.05 as the criterion, a final fixed set of features significantly associated with overall survival (OS) was further screened to ensure that subsequent modeling retained only core factors with independent predictive power for survival outcomes. These included: peripheral blood platelet count, chromosomes, U2AF1 gene, ASXL1_exon12_19721ul gene, and SF3B1 gene.
[0028] The gene sequences can be obtained from the following URL: ASXL1_exon12_19721ul gene: https: / / www.ncbi.nlm.nih.gov / nuccore / NC_000001.11?report=fasta.
[0029] U2AF1 gene: https: / / www.ncbi.nlm.nih.gov / nuccore / NM_006758.3?report=fasta.
[0030] SF3B1 gene: https: / / www.ncbi.nlm.nih.gov / nuccore / NM_012433.4?report=fasta.
[0031] Based on the feature data screened by Lasso-Cox, a multivariate Cox regression fitting, linear predictive value (lp) score calculation, and ROC curve analysis method were used to incorporate five core predictive factors into a multivariate Cox proportional hazards regression model. With overall survival (OS) as the outcome event, the model was fitted and the regression coefficient of each predictive factor was calculated to realize the construction of a comprehensive risk score and the evaluation of predictive efficacy. The predictive factors, scores, and weights are shown in Table 1 below.
[0032] Table 1. Correspondence between predictive factors, scores, and weights
[0033] Positive regression coefficients (weights) for peripheral blood platelet count, chromosome analysis, U2AF1 gene, and ASXL1_exon12_19721ul gene indicate an increased risk of death when these indicators are present. The predictor score is obtained by summing the products of the regression coefficients and the corresponding indicator variable values. Using the `coords` function in the `survminer` package of R, the optimal cutoff point of 1.097 was obtained, which is the critical value. A KM survival curve was then plotted. Figure 1 .
[0034] This divides the patients into two groups: Standard risk: Risk score ≤ 1.0, can be observed and followed up temporarily; High risk: Risk score > 1.0, short expected survival, requires initiation of demethylation therapy.
[0035] To fully utilize the small sample data and ensure robust results, a 10-fold stratified cross-validation was used for model evaluation: data was randomly divided into 10 subsets based on survival outcome. In each round, the Cox model was trained using 9 subsets and externally validated using 1 subset. The AUC and C-index of each validation set were calculated, and the model's predictive efficacy was evaluated using the average of the 10 results. The AUC was 0.744 ± 0.114, indicating that the prognostic model built based on these variables can accurately distinguish between standard-risk and high-risk individuals to a certain extent, demonstrating good predictive efficacy for patient prognosis. The C-index (concordance index) obtained using the concordance() function in the survival package was 0.674 ± 0.086, indicating that the model can distinguish between high-risk and low-risk samples to a certain extent. (See details...) Figure 2-11 This strategy maximizes data utilization and reduces assessment bias under small sample conditions. The constructed CMML prognostic model has good stability and risk discrimination ability, and can provide a reference for clinical prognostic judgment and treatment decision-making. Example 2
[0036] The following example, using a real clinical case from the dataset, further illustrates how to use the above model: The patient was diagnosed with CMML by the Department of Hematology at Qilu Hospital of Shandong University. Based on the traditional scoring systems IPSS and CPSS-Mol, the patient was determined to have high-risk CMML and required initiation of treatment. Peripheral blood was drawn from the patient for complete blood count and platelet count. Bone marrow aspiration results and gene mutation detection results were collected and calculated according to this risk stratification assessment model. The indicators are shown in Table 2 below: Table 2 Statistical analysis of predictive factor detection results Predictors Test results Peripheral blood platelet count (WBC) <![CDATA[8×10 9 / L]]> chromosome -7 U2AF1 gene mutation Mutation SF3B1 gene mutation No mutation ASXL1_exon12_19721ul gene mutation Mutation Based on the above statistical results, the values for each indicator are assigned as shown in Table 3 below: Table 3. Assignment of values to predictor factors
[0037] Substituting the values of the above predictive factors into the model, the calculation is as follows. Peripheral blood platelet count: 0.69 × 2 = 1.38; Chromosome: 0.22×2=0.44; U2AF1 gene mutation: 0.64 × 1 = 0.64; SF3B1 gene mutation: -1.37 × 0 = 0; ASXL1_exon12_19721ul gene mutation: 0.75×1=0.75.
[0038] The total risk score was calculated as follows: Total risk score = 1.38 + 0.44 + 0.64 + 0.00 + 0.75 = 3.21 points. Stratification results: The patient's risk score of 3.21 points > 1.0 points, classifying them as high-risk, consistent with traditional scoring systems. This patient underwent demethylation therapy after diagnosis, with good treatment outcomes, followed by hematopoietic stem cell transplantation. Currently, the disease is stable, and the patient is undergoing regular follow-up.
[0039] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments for application in other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
Claims
1. A risk stratification assessment model for chronic myelomonocytic leukemia, characterized in that, The model was constructed using biomarker detection status as input variables; the biomarker detection status consisted of peripheral blood platelet count, chromosome, U2AF1 gene, ASXL1_exon12_19721ul gene, and SF3B1 gene mutation.
2. The chronic myelomonocytic leukemia risk stratification assessment model according to claim 1, characterized in that, The model uses the following formula to calculate the risk score: Risk Score = (0.69 × score corresponding to peripheral blood platelet count) + (0.22 × score corresponding to chromosome alteration) + (0.64 × score corresponding to U2AF1 gene mutation) + (-1.37 × score corresponding to SF3B1 gene mutation) + (0.75 × score corresponding to ASXL1_exon12_19721ul gene mutation); where, peripheral blood platelet count ≥ 100 × 10⁻⁶ 9 / L is 0 points, 20×10 9 / L < Peripheral blood platelet count <100×10 9 / L is 1 point, peripheral blood platelet count ≤20×10 9 / L is 2 points; normal chromosomes and -Y are 0 points; -7, -5, +8, and complex karyotypes with ≥3 abnormalities are 2 points; excluding the aforementioned chromosome changes, all are 1 point; ASXL1_exon12_19721ul gene without mutation is 0 points, and mutation is 1 point; U2AF1 gene without mutation is 0 points, and mutation is 1 point; SF3B1 gene without mutation is 0 points, and mutation is 1 point.
3. The chronic myelomonocytic leukemia risk stratification assessment model according to claim 2, characterized in that, When the risk score is ≤1, the prognosis of CMML is standard risk; when the risk score is >1, the prognosis of CMML is high risk.
4. A group of prognostic risk markers for chronic myelomonocytic leukemia, characterized in that, It is composed of mutations in peripheral blood platelets, chromosomes, U2AF1 gene, ASXL1_exon12_19721ul gene, and SF3B1 gene.