An acute lymphoblastic leukemia scoring system
By employing a two-tiered data acquisition architecture and a multi-factor model scoring system, the problem of existing systems failing to integrate dynamic indicators has been solved, enabling real-time assessment and personalized treatment matching for patients with acute lymphoblastic leukemia, thereby improving the safety and suitability of treatment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- THE SECOND AFFILIATED HOSPITAL ARMY MEDICAL UNIV
- Filing Date
- 2026-03-26
- Publication Date
- 2026-07-03
AI Technical Summary
The existing acute lymphoblastic leukemia scoring system relies on static baseline indicators and fails to integrate dynamic evolution indicators during the treatment process. This results in dynamic efficacy assessments failing to reflect prognostic changes in a timely manner. Furthermore, the system does not adequately integrate indicators of comorbidities and treatment tolerance, leading to poor fit between the scoring results and the actual situation of patients and increasing the risk of adverse reactions.
A two-tiered data acquisition architecture of large-scale and sub-scale data acquisition is adopted. By combining baseline and dynamic data, digital PCR and NGS detection are used to construct a multivariate logistic regression model and a random forest algorithm to generate a stage-specific weight matrix. Risk scoring is performed by combining the SHAP value to match personalized treatment plans.
It enables real-time assessment of patients' conditions, reduces unexpected adverse reactions, improves the adaptability and safety of treatment plans, and optimizes system adaptability through closed-loop processes.
Smart Images

Figure CN122337586A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of acute lymphoblastic leukemia technology, specifically an acute lymphoblastic leukemia scoring system. Background Technology
[0002] Acute lymphoblastic leukemia (ALL) is a highly heterogeneous hematologic malignancy. Its pathogenesis involves abnormalities at multiple levels, including the genome, transcriptome, and epigenetics, and different subtypes show significant differences in treatment response and prognosis. Although various risk stratification scoring systems have been established in clinical practice, two unresolved technical limitations remain in practical applications.
[0003] Current clinical risk stratification scoring systems largely rely on static baseline indicators at diagnosis, such as fixed gene site mutations and initial clinical parameters, failing to integrate multi-dimensional dynamic evolution indicators during treatment—including temporal changes in gene expression profiles, immune function recovery curves, and dynamic data on drug exposure concentrations. This makes dynamic efficacy assessment unable to reflect prognostic changes in a timely manner during treatment and difficult to accurately capture real-time fluctuations in the patient's condition as treatment progresses.
[0004] Meanwhile, existing systems lack sufficient integration of patient comorbidities (such as cardiovascular disease, abnormal liver and kidney function, and metabolic syndrome) and treatment tolerance indicators (such as duration of myelosuppression, degree of organ dysfunction, and grade of adverse drug reactions), and the interpretability of the scoring models is weak. These systems typically only output risk levels without clearly defining the contribution weight and causal relationship of each indicator to the scoring results. This makes it difficult for clinicians to trace the scoring logic, ultimately leading to discrepancies between the risk stratification results and the patient's actual treatment tolerance and comorbidity suitability. This can easily cause unexpected adverse reactions or mismatches between treatment plans and organ function. Summary of the Invention
[0005] The purpose of this invention is to provide an acute lymphoblastic leukemia scoring system to solve the problems mentioned in the background art.
[0006] To achieve the above objectives, the present invention provides the following technical solution: an acute lymphoblastic leukemia scoring system, comprising a data acquisition module, an assessment module, a feature adaptation module, a weight calculation module, a risk score generation module, a matching module, and a feedback update module;
[0007] Preferably, the acquisition module adopts a two-level data acquisition architecture of large stages and sub-stages, dividing the entire treatment cycle of acute lymphoblastic leukemia into three major stages: induction remission period, consolidation treatment period, and maintenance treatment period. Each stage is further divided into sub-stages. The induction remission period is divided into early and late sub-stages, the consolidation treatment period is divided into dose escalation period and stable period sub-stages, and the maintenance treatment period is divided into early and late sub-stages. Each sub-stage corresponds to a dedicated dynamic indicator acquisition rule.
[0008] The data collection system includes baseline data and dynamic data throughout the entire process. Baseline data includes age, ECOGPS fitness score, fusion genes, chromosome karyotype, antigen expression profile and positive intensity; dynamic data throughout the entire process includes treatment response anchor points, toxicity warning anchor points, temporal changes in gene expression profiles, immune function indicators, and drug exposure concentrations.
[0009] It utilizes the connection between clinical laboratory information systems and gene testing platforms to achieve automatic data synchronization, has built-in automatic sub-stage identification function, and uses digital PCR and NGS combined detection to ensure the accuracy of rare mutation identification; the collected data is transmitted in real time to the evaluation module, feature adaptation module and weight calculation module.
[0010] Preferably, the evaluation module receives baseline comorbidity data and full-course tolerability data from the acquisition module, establishes a three-dimensional correlation model of comorbidity type, tolerability level, and efficacy attenuation coefficient, subclassifies comorbidities, and quantifies the correlation weights of the three through multivariate logistic regression.
[0011] Set up comorbidity grading correction rules, assess myelosuppression according to WHO grades 0-IV and recovery speed in two dimensions, define grade IV myelosuppression with recovery time >14 days as a high-risk tolerability event, and encode adverse reactions; output comorbidity and tolerability association results and grading correction coefficients are simultaneously transmitted to the weight calculation module and matching module.
[0012] Preferably, the feature adaptation module is based on a database, receives baseline data from the acquisition module including fusion genes and chromosome karyotypes, matches specific biomarkers, establishes an interaction weight matrix between rare subtypes and comorbidities, and clarifies the weight adjustment rules for subtypes and comorbidities.
[0013] The built-in subtype template update function automatically updates the interaction weight matrix after importing new subtype parameters; the output subtype-specific marker matching results and interaction weight data are transmitted to the weight calculation module and the matching module.
[0014] Preferably, the weight calculation module receives sub-stage dynamic data from the acquisition module, comorbidity and tolerability association results from the evaluation module, and interactive weight data from the feature adaptation module, and constructs a calculation model using a combination algorithm of random forest and SHAP value.
[0015] An anchor-triggered weight update mechanism was designed for each treatment sub-stage. When the MRD decline rate was <50% in the late induction sub-stage, the weight of gene expression profile indicators increased from 20% to 35%; when the drug AUC fluctuation was >20% in the consolidation stabilization phase, the weight of drug exposure concentration increased from 15% to 25%; and when the MRD was >0.01% in the late maintenance therapy phase, the weight of relapse-related biomarkers increased from 18% to 30%. SHAP values were graded as high, medium, and low to generate a stage-specific weight matrix, which was transmitted to the risk score generation module in real time.
[0016] Preferably, the risk score generation module receives the stage-specific weight matrix from the weight calculation module and, in conjunction with the real-time indicator data from the acquisition module, calculates a comprehensive risk score of 0-100 points, corresponding to five risk levels: extremely low risk, low risk, medium risk, high risk, and extremely high risk.
[0017] A correlation logic between SHAP values and clinical rules was constructed. When the SHAP value contribution was 0.4, it corresponded to Ph-like ALL with moderate renal insufficiency, and the correlation score increased by 25%. When the SHAP value contribution was -0.2, it corresponded to MRD <0.01% during the induction period, and the correlation score decreased by 15%. When the SHAP value contribution was 0.15, it corresponded to grade IV myelosuppression with a recovery time >14 days, and the correlation score increased by 12%. When the SHAP value contribution was -0.08, it corresponded to a sustained increase in CD4+ T cell count, and the correlation score decreased by 5%. A scoring report containing risk level, SHAP value heatmap, and intervention recommendations was generated and transmitted to the matching module.
[0018] Preferably, the matching module receives the risk level report from the risk score generation module, the comorbidity and tolerability association results from the assessment module, and the subtype matching results from the feature adaptation module. It then calls up the three-dimensional treatment plan library of risk level, subtype classification, and comorbidity status to match the corresponding treatment plan: standardized plans are used for common subtypes, targeted plans are matched for rare subtypes, and dose adjustment plans or drug substitution plans are used for patients with comorbidities.
[0019] The treatment protocol library incorporates a threshold-triggered optimization mechanism. When MRD > 0.1% for two consecutive cycles, the weight calculation module is triggered to increase the weight of relapse-related biomarkers. When the incidence of serious adverse reactions > 15%, low-toxicity alternative protocols are prioritized. The protocol library is updated quarterly based on NCCN guidelines and ELN recommendations, marking the level of evidence-based medicine, and incorporating RWD validation protocols through multicenter adverse reaction chi-square test. The output treatment protocols and patient treatment outcome data are synchronously transmitted to the feedback update module.
[0020] Preferably, the feedback update module receives patient treatment outcome data from the matching module, and sets the data collection frequency according to risk level and subtype: extremely high-risk patients are collected once every 2 weeks and the weight is adjusted once per cycle; high-risk patients are collected once every 4 weeks and the weight is adjusted once per 2 cycles; intermediate-low-risk patients are collected once every 6 weeks and the weight is adjusted once per 3 cycles; and extremely low-risk patients are collected once every 8 weeks and the weight is adjusted once per 4 cycles. Data integrity is ensured through outpatient and telephone follow-ups.
[0021] Data is protected by hierarchical encryption. Feedback data is used to correct the coefficients of indicators with large deviations in the weight calculation module, and to update the treatment suitability matching module for treatments with poor efficacy or excessively high adverse reaction rates. At the same time, data collection optimization suggestions are fed back to the collection module.
[0022] The beneficial effects of this invention are as follows:
[0023] 1. This invention divides the entire treatment cycle into three major stages: induction remission, consolidation therapy, and maintenance therapy. Each stage is further subdivided into sub-stages, forming a two-tiered data acquisition architecture. During data acquisition, not only baseline data such as age and fusion genes are included, but dynamic data such as the rate of MRD decline and drug exposure concentration are also continuously collected. At the same time, anchor triggering mechanisms are set in different sub-stages. For example, when the induction of late-stage MRD decline is slower than 50%, the weight of gene expression profile indicators will be increased, making the risk score more consistent with the patient's real-time condition and providing a more accurate reference for clinical selection of treatment strategies.
[0024] 2. The evaluation module of this invention analyzes the correlation between comorbidity type, tolerability level, and efficacy, and can also quantify the weight of the three factors. For example, it defines grade IV myelosuppression with recovery exceeding 14 days as a high-risk event. The feature matching module matches subtype-specific biomarkers and analyzes the interaction between rare subtypes and comorbidities. Based on this, the matching module calls a three-dimensional protocol library to select standardized protocols for common subtypes, targeted protocols for rare subtypes, and dose adjustment or drug substitution protocols for patients with comorbidities. This helps clinicians match more suitable protocols for different patients, reduce unexpected adverse reactions, and improve treatment safety.
[0025] 3. The risk score generation module of this invention uses a heatmap to show the influence of each indicator on the score by correlating SHAP values with clinical rules. Doctors can clearly understand the scoring logic. For example, the score will increase by 25% when Ph-like ALL is combined with moderate renal insufficiency. At the same time, the feedback update module sets the data collection frequency according to the risk level. Data is collected once every 2 weeks for very high-risk patients and once every 6 weeks for intermediate and low-risk patients. Then, the indicator coefficients are adjusted in reverse according to the treatment outcome and the treatment plan with poor efficacy is updated, forming a closed loop process of collection, evaluation and adjustment. This allows the system to be continuously optimized with the accumulation of clinical data and maintain good adaptability in the long term. Attached Figure Description
[0026] Figure 1 This is a flowchart of the acute lymphoblastic leukemia scoring system of the present invention. Detailed Implementation
[0027] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0028] like Figure 1 As shown in the figure, this embodiment of the invention provides an acute lymphoblastic leukemia scoring system, including a data acquisition module, an assessment module, a feature adaptation module, a weight calculation module, a risk score generation module, a matching module, and a feedback update module. The specific implementation of each module is as follows:
[0029] The data acquisition module employs a two-tiered data acquisition architecture, dividing the entire treatment cycle of acute lymphoblastic leukemia (ALL) into three major phases: induction remission, consolidation therapy, and maintenance therapy. Each phase is further subdivided into sub-phases: induction remission phase: 1-2 weeks in the early stage / 3-4 weeks in the late stage; consolidation therapy phase: dose escalation phase / stabilization phase; maintenance therapy phase: 1-6 months in the early stage / 7-12 months in the late stage. Each sub-phase corresponds to dynamic indicator acquisition rules. For example, during the dose escalation phase, data on "drug concentration - bone marrow suppression lag effect" is collected, and during the late maintenance therapy phase, dynamic changes in relapse-related biomarkers are collected.
[0030] Specific collection frequency and detection methods for dynamic indicators in each sub-stage:
[0031] ① Early stage of induction remission (1-2 weeks): MRD was measured every 7 days (digital PCR method), and drug exposure concentration was measured every 3 days (high performance liquid chromatography method).
[0032] ② Late stage of induction remission (3-4 weeks): Bone marrow blast cell count is tested every 5 days (bone marrow smear microscopy), and immune function indicators (CD4+ T cells) are tested weekly (flow cytometry).
[0033] ③ Consolidation treatment period: During the dose escalation period, the drug AUC should be tested every 2 days, and during the stabilization period, it should be tested every 7 days;
[0034] ④ Maintenance therapy period: MRD (NGS method) once a month in the early stage, and relapse marker (WT1 gene, real-time fluorescence quantitative PCR method) once every 2 weeks in the later stage.
[0035] The data collection system includes baseline data and dynamic data throughout the treatment process. Baseline data includes age, ECOGPS performance score, BCR-ABL1 / STIL-TAL1 / E2A-HLF fusion gene, chromosome karyotype, CD3 / CD10 / CD19 antigen expression profile and positive intensity. Dynamic data throughout the treatment process includes treatment response anchors (bone marrow blast cell decline rate on day 14 of induction, MRD clearance trend), toxicity warning anchors (neutrophil recovery time, platelet transfusion-dependent number of times), temporal changes in gene expression profiles (fluctuations in BCL-2 and MYC levels), immune function indicators (CD4+ T cell count, IL-6 / TNF-α levels), and drug exposure concentrations (methotrexate / cytarabine serum concentration AUC).
[0036] It utilizes the connection between clinical laboratory information systems and gene testing platforms to achieve automatic data synchronization, and has a built-in automatic sub-stage identification function (matching sub-stages based on treatment days and medication regimens). It employs digital PCR and NGS combined detection to ensure the accuracy of rare mutation identification; the collected data is transmitted in real time to the evaluation module, feature adaptation module, and weight calculation module.
[0037] The assessment module receives baseline comorbidity data (such as history of cardiovascular, liver and kidney diseases) and full-course tolerability data (degree of bone marrow suppression, time of adverse reaction occurrence) from the acquisition module. It establishes a three-dimensional correlation model of comorbidity type, tolerability level, and efficacy attenuation coefficient. Referring to ICD-11 and the assessment criteria for hematologic malignancies, it subcategories 32 comorbidities. For example, cardiovascular diseases are divided into stable coronary artery disease, acute coronary syndrome, and heart failure. The correlation weights among the three are quantified through multivariate logistic regression. For example, moderate renal insufficiency corresponds to a 20% prolongation of bone marrow suppression and a 15% attenuation of methotrexate efficacy.
[0038] The training parameters for the multivariate logistic regression model are as follows:
[0039] ① Data preprocessing: Z-score standardization was used for continuous variables, such as duration of myelosuppression and drug exposure concentration AUC, and one-hot encoding was used for categorical variables, such as comorbidity subclasses and tolerance grades.
[0040] ② Independent variable screening: Stepwise regression was used, with an inclusion criterion of α=0.05 and an exclusion criterion of α=0.10. Finally, 18 core independent variables were included, including 6 subcategories of comorbidities, 4 tolerability grades, and 8 efficacy-related indicators.
[0041] ③ Model regularization: L2 regularization (Ridge regression) is used with a regularization parameter λ=0.01 to avoid overfitting;
[0042] ④ Training and Validation: The training set and internal validation set were divided in a 7:3 ratio. The model parameters were optimized through 5-fold cross-validation. The validation set AUC was 0.89, accuracy was 0.83, and recall was 0.81. ⑤ Model Connection Relationship: Comorbidity type → Independent variable layer → Regression coefficient calculation layer → Tolerability level association layer → Efficacy attenuation coefficient output layer. Data transmission in each layer adopted tensor format, with dimensions of sample number × number of variables × time node to ensure the continuity and correlation of time series data.
[0043] The following rules are set for comorbidity grading correction: stable coronary artery disease is weighted with a correction coefficient of 1.0; acute coronary syndrome within the past 6 months is weighted with 1.8; mild liver function abnormalities are weighted with 1.2, and severe liver function abnormalities with 2.0. Myelosuppression is assessed using a dual dimension of WHO classification 0-IV and recovery speed. "Grade IV myelosuppression with a recovery time > 14 days" is defined as a high-risk tolerability event. Adverse reactions are coded using the CTCAE 5.0 standard. The output comorbidity-tolerance association results and grading correction coefficients are simultaneously transmitted to the weight calculation module and the matching module, providing comorbidity dimension parameters for weight calculation and tolerability basis for treatment plan adjustments.
[0044] The feature adaptation module is based on a database integrating nearly 5,000 rare subtype cases from 20 top-tier hospitals in China and abroad over the past 10 years. It receives baseline data from the acquisition module, including gene and chromosome karyotype data, and matches specific markers for subtypes such as complex karyotype ALL, Ph-like ALL variants, and biphenotype ALL, such as CRLF2 rearrangement in Ph-like ALL. It establishes an interaction weight matrix between rare subtypes and comorbidities, clarifying the weight adjustment rules for subtypes and comorbidities: the initial weight is based on the core indicator weight of the subtype without comorbidities. The baseline (initial weights were set based on multivariate regression analysis of multicenter data from the past 10 years) was adjusted by "initial weight × (1 + co-risk coefficient)" when there was a synergistic risk between the subtype and the comorbidity. For example, when Ph-like ALL was complicated with hypertension, the initial weight for cardiovascular risk was 1.2, the synergistic risk coefficient was 0.33, and the adjusted weight was 1.2 × 1.33 ≈ 1.6; when biphenotype ALL was complicated with liver injury, the initial weight for liver function injury was 1.0, the synergistic risk coefficient was 0.5, and the adjusted weight was 1.0 × 1.5 = 1.5.
[0045] The system has a built-in subtype template update function. The Excel template includes fields such as "subtype name - specific biomarker - comorbidity interaction rule - 5-year disease-free survival rate - treatment response rate". After the administrator imports the new subtype parameters, the interaction weight matrix is automatically updated. The output subtype specific biomarker matching results and interaction weight data are passed to the weight calculation module and the matching module to provide subtype dimension parameters for weight calculation and provide a basis for matching treatment plans for rare subtypes.
[0046] The weight calculation module receives sub-stage dynamic data (such as MRD decline rate and drug AUC fluctuation) from the acquisition module, the association results of comorbidities and tolerability from the assessment module, and the interactive weight data from the feature adaptation module, and constructs a calculation model using a combination algorithm of random forest and SHAP value.
[0047] An anchor-triggered weight update mechanism was designed for each treatment sub-stage. When the MRD decline rate was <50% in the late induction sub-stage, the weight of the gene expression profile index (MYC) was increased from 20% to 35%. When the drug AUC fluctuation was >20% in the consolidation and stabilization phase, the weight of the drug exposure concentration was increased from 15% to 25%. When the MRD was >0.01% in the late maintenance therapy phase, the weight of the relapse-related marker (WT1 gene) was increased from 18% to 30%.
[0048] Anchor point index calculation method:
[0049] ① MRD decrease rate = (Early induction baseline MRD value - Late induction MRD value) / Early induction baseline MRD value × 100%;
[0050] ② Drug AUC fluctuation = |Current period drug AUC measured value - Previous period drug AUC measured value| / Previous period drug AUC measured value × 100%;
[0051] ③ All anchor point index detections are based on the starting time point of "Day 1 of the sub-phase" and the calculation end time point of "Day 1 of the sub-phase".
[0052] Random Forest Model Construction and Training Parameters:
[0053] ① Model structure: Number of decision trees (n_estimators) = 200, maximum depth of each decision tree (max_depth) = 8, maximum number of features when splitting each tree (max_features) = "sqrt", minimum number of samples per leaf node (min_samples_leaf) = 5;
[0054] ② Splitting criteria: The Gini coefficient (for classification tasks) is combined with the mean squared error, with the classification task accounting for 60% and the regression task accounting for 40% in the weight calculation;
[0055] ③ Training process: Step 1, data splitting: The training set and validation set are split in an 8:2 ratio. The training set contains data from 3,000 multicenter ALL patients, and the validation set contains 750 independent samples.
[0056] The second step is model training: number of iterations = 500, learning rate = 0.005, and early stopping mechanism is used;
[0057] The third step is model evaluation: training set accuracy = 0.92, validation set accuracy = 0.87, and recall rate for identifying high-risk / very high-risk patients in the confusion matrix = 0.88.
[0058] ④ Module connection relationship: Dynamic data acquisition module → Data preprocessing layer → Feature input layer → Random forest model layer → Preliminary weight calculation layer → SHAP value correction layer → Stage-specific weight matrix output layer.
[0059] (2) The logic of combining SHAP values with random forests:
[0060] ① The TreeExplainer interpreter was used to calculate the SHAP value, with a sample size threshold of 100. The feature contribution ranking was based on the "absolute SHAP value mean".
[0061] ② Input data: intermediate outputs of the random forest model, clinical rule labels, such as "Grade IV myelosuppression" and "MRD > 0.01%";
[0062] ③ Output correlation: The mapping between SHAP values and clinical rules adopts a linear correlation function, SHAP_score=a×Clinical_label+b, where a is the correlation coefficient and b is the offset;
[0063] ④ Weight correction process: SHAP value classification → calculation of corresponding feature weight adjustment coefficients → correction of the initial weights of the random forest output → generation of the final stage specific weight matrix.
[0064] The SHAP value is graded according to the following rules: a SHAP value contribution > 0.3 is defined as high (positive contribution), corresponding to a significant increase in the risk score; a contribution between 0.1 and 0.3 is defined as medium (positive contribution), corresponding to a moderate increase in the risk score; a contribution between < 0.1 and ≥ 0 is defined as low (positive contribution), corresponding to a slight increase in the risk score. To cover scenarios where the SHAP value negatively influences the risk score, a further grading of negative contribution intervals is provided: a SHAP value contribution between -0.1 and 0 is defined as low (negative contribution), corresponding to a slight decrease in the risk score; a contribution between -0.3 and -0.1 is defined as medium (negative contribution), corresponding to a moderate decrease in the risk score; and a contribution between < -0.3 is defined as high (negative contribution), corresponding to a significant decrease in the risk score. Based on the above full-range grading rules, a stage-specific weight matrix is generated and transmitted in real time to the risk score generation module as the core weight basis for risk score calculation.
[0065] The risk score generation module receives the stage-specific weight matrix from the weight calculation module and combines it with the real-time indicator data from the acquisition module to calculate a comprehensive risk score of 0-100 points, corresponding to five risk levels: extremely low risk (0-20 points), low risk (21-40 points), medium risk (41-60 points), high risk (61-80 points), and extremely high risk (81-100 points).
[0066] The association logic between SHAP value and clinical rules is constructed as follows: When the SHAP value contribution is 0.4 (high), it corresponds to Ph-like ALL combined with moderate renal insufficiency, and the comprehensive baseline score at the current stage increases by 25% proportionally (i.e., baseline score × 1.25); when the SHAP value contribution is -0.2 (medium), it corresponds to MRD < 0.01% during the induction period, and the comprehensive baseline score at the current stage decreases by 15% proportionally (i.e., baseline score × 0.85); when the SHAP value contribution is 0.15 (medium), it corresponds to grade IV myelosuppression with a recovery time > 14 days, and the associated score increases by 12%; when the SHAP value contribution is -0.08 (low), it corresponds to a sustained increase in CD4+ T cell count, and the associated score decreases by 5%. A scoring report containing risk level, SHAP value heatmap, and intervention recommendations is generated and transmitted to the matching module to provide quantitative risk basis for treatment plan selection.
[0067] Mapping relationship between SHAP values and clinical rules:
[0068] Clinical scenario: Ph-like ALL combined with moderate renal insufficiency; SHAP value contribution: 0.4; correlation coefficient: 1.25; scoring adjustment rule: current stage comprehensive baseline score × 1.25; evidence-based basis: NCCNALL Guidelines 2024 Edition Chapter 4 + multicenter RWD data from 3 top-tier hospitals in China.
[0069] Clinical scenario: MRD < 0.01% on day 14 of induction remission; SHAP value contribution: -0.2; correlation coefficient: 0.85; scoring adjustment rule: current stage comprehensive baseline score × 0.85; evidence-based basis: ELN2023ALL prognostic stratification criteria, section 3.2 + European Society of Hematology clinical research data.
[0070] Clinical scenario: Grade IV myelosuppression with neutrophil recovery time > 14 days; SHAP value contribution: 0.15; correlation coefficient: 1.12; scoring adjustment rule: current stage comprehensive baseline score × 1.12; evidence-based basis: CTCAE 5.0 adverse reaction grading standard + data from the domestic hematologic oncology collaborative group.
[0071] Clinical scenario: CD4+ T cell count ≥500 / μL in two consecutive tests during maintenance therapy; SHAP value contribution: -0.08; correlation coefficient: 0.95; scoring adjustment rule: current stage comprehensive baseline score × 0.95; evidence-based basis: 2023 consensus on ALL immune function monitoring in Chinese Journal of Hematology (Peking University Journal, CSCD Journal) + multicenter follow-up data.
[0072] Clinical scenario: AUC fluctuation of drugs in the stable phase of consolidation therapy >20%; SHAP value contribution: 0.25; correlation coefficient: 1.20; scoring adjustment rule: current stage comprehensive baseline score × 1.20; evidence-based basis: International Therapeutic Drug Monitoring (TDM) Association ALL medication guidelines + North American multicenter TDM data.
[0073] Clinical scenario: WT1 gene expression <0.1% in the later stage of maintenance therapy; SHAP value contribution: -0.3; correlation coefficient: 0.80; scoring adjustment rule: current stage comprehensive baseline score × 0.80; evidence-based basis: 2024 ALL relapse biomarker monitoring guidelines in Leukemia and Lymphoma (Chinese Medical Association series journals, Chinese biomedical core journals) + joint study of 20 hospitals in China.
[0074] Clinical scenario: Biphenotype ALL combined with mild liver dysfunction; SHAP value contribution: 0.18; correlation coefficient: 1.15; scoring adjustment rule: current stage comprehensive baseline score × 1.15; evidence-based basis: international bone marrow transplant registry combined with comorbidity criteria + multicenter tolerability data.
[0075] Clinical scenario: Early bone marrow blast cell decline rate >70% during induction of remission; SHAP value contribution: -0.12; correlation coefficient: 0.90; scoring adjustment rule: current stage comprehensive baseline score × 0.90; evidence-based basis: EHA2023 ALL treatment response assessment guidelines + European multicenter clinical study.
[0076] Comprehensive risk score calculation formula:
[0077] Comprehensive risk score = Σ (weight of each feature in the stage-specific weight matrix × real-time detection value of the feature) × SHAP correlation coefficient × assessment module classification correction coefficient;
[0078] The real-time detection values of each feature have been standardized to ensure that the weights are consistent with the units of the detection values.
[0079] Output data format: The scoring report contains four columns of data: “Feature Name-Feature Weight-SHAP Contribution-Scoring Impact Value”. The SHAP value heatmap is drawn using Python matplotlib, which makes it easier for clinicians to trace the scoring logic.
[0080] The matching module receives risk level reports from the risk scoring generation module, comorbidity and tolerability association results from the assessment module, and subtype matching results from the feature adaptation module. It then calls upon a three-dimensional treatment plan library based on risk level, subtype classification, and comorbidity status to match corresponding treatment plans: standardized plans are used for common subtypes, and targeted plans are matched for rare subtypes. For STIL-TAL1 fusion genes, CDK4 / 6 inhibitors are used in combination with glucocorticoids, and for E2A-HLF fusion genes, BCL-2 inhibitors are used in combination with cytarabine. For patients with comorbidities, a dose adjustment plan is used. In cases of moderate renal insufficiency, the methotrexate dose is reduced by 30%, or a drug replacement plan is adopted. In cases of severe myelosuppression, liposomal doxorubicin is used to replace anthracyclines.
[0081] The 3D solution library has a simplified classification structure:
[0082] ① First dimension (risk level): divided into 5 categories: very low risk / low risk / medium risk / high risk / very high risk, with each category associated with 3-5 basic medication regimens;
[0083] ② Second dimension (subtype classification): Common subtypes correspond to 3 standardized treatment options, and rare subtypes correspond to 2 targeted treatment options;
[0084] ③ The third dimension (comorbidity status): divided into 4 categories: "no comorbidity / mild comorbidity / moderate comorbidity / severe comorbidity". Each category corresponds to a dosage adjustment ratio or a list of alternative drugs. For example, liposomal doxorubicin can be used to replace regular doxorubicin for liver damage.
[0085] The treatment protocol library incorporates a threshold-triggered optimization mechanism. When MRD > 0.1% for two consecutive cycles, the weight calculation module is triggered to increase the weight of relapse-related biomarkers. When the incidence of serious adverse reactions in a single treatment cycle is > 15%, low-toxicity alternative protocols are prioritized. The protocol library is updated quarterly based on NCCN guidelines and ELN recommendations, and the evidence-based medicine evidence level is marked as Class I / IIA / IIB. The protocol is also included in the RWD validation protocol through a multicenter chi-square test of adverse reactions. The output treatment protocols and patient treatment outcome data (response rate, relapse rate, adverse reaction rate) are synchronously transmitted to the feedback update module.
[0086] The feedback update module receives patient treatment outcome data from the matching module, including remission rate, relapse rate, OS, DFS, and adverse reaction incidence. Data collection frequency is set according to risk level and subtype (with 28 days as one standard treatment cycle): very high-risk patients are collected once every 2 weeks with weight adjustment once per cycle; high-risk patients are collected once every 4 weeks with adjustment once per 2 cycles; intermediate-low-risk patients are collected once every 6 weeks with adjustment once per 3 cycles; and very low-risk patients are collected once every 8 weeks with adjustment once per 4 cycles. OS / DFS data integrity is ensured through outpatient and telephone follow-ups.
[0087] Data is protected by hierarchical encryption. Routine data uses the hospital's existing encryption system, while sensitive data is subject to dynamic authorization. Feedback data is used to correct the coefficients of indicators with large deviations in the weight calculation module, and to update the treatment suitability matching module for treatments with poor efficacy or excessively high adverse reaction rates. At the same time, data collection optimization suggestions are fed back to the collection module, such as adding specific monitoring indicators for high-risk subtypes, forming a closed-loop process technology of data collection, assessment calculation, score matching, outcome feedback and parameter correction.
[0088] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0089] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. An acute lymphoblastic leukemia scoring system, characterized in that, It includes a data acquisition module, an evaluation module, a feature adaptation module, a weight calculation module, a risk score generation module, a matching module, and a feedback update module; Data Acquisition Module: Divides the entire treatment cycle of acute lymphoblastic leukemia into three major stages and corresponding sub-stages, collects baseline data and dynamic data and outputs them; Evaluation module: Establishes a three-dimensional correlation model, sets hierarchical correction rules, and outputs correlation results and correction coefficients; Feature adaptation module: Receives baseline data, matches subtype markers, establishes an interaction weight matrix, and outputs matching results and weight data; Weight calculation module: It uses a combination algorithm to build a calculation model, designs a sub-stage anchor point trigger weight update mechanism, generates a weight matrix and outputs it; Risk score generation module: Calculates the comprehensive risk score and corresponding level, constructs the logic linking SHAP value with clinical rules, and generates a report containing intervention recommendations; Matching module: calls the 3D treatment plan library to match plans, embeds a threshold-triggered optimization mechanism, updates the plan library, and outputs treatment plans and outcome data; Feedback and update module: Set the collection frequency according to risk level and subtype, use hierarchical encryption to protect data, reverse correct index coefficients and update schemes, and provide feedback optimization suggestions.
2. The acute lymphoblastic leukemia scoring system according to claim 1, characterized in that, The acquisition module adopts a two-level data acquisition architecture of large stages and sub-stages, which divides the entire treatment cycle of acute lymphoblastic leukemia into three major stages: induction remission period, consolidation treatment period, and maintenance treatment period. Each stage is further divided into sub-stages, and each sub-stage corresponds to a dedicated dynamic indicator acquisition rule. The data collection system includes baseline data and dynamic data throughout the entire process. Baseline data includes age, ECOGPS fitness score, fusion genes, chromosome karyotype, antigen expression profile and positive intensity; dynamic data throughout the entire process includes treatment response anchor points, toxicity warning anchor points, temporal changes in gene expression profiles, immune function indicators, and drug exposure concentrations. Automatic data synchronization is achieved by connecting with relevant clinical systems to ensure the accuracy of rare mutation identification; collected data is transmitted in real time to the evaluation module, feature adaptation module and weight calculation module.
3. The acute lymphoblastic leukemia scoring system according to claim 2, characterized in that, The evaluation module receives baseline comorbidity data and full-course tolerability data from the acquisition module, establishes a three-dimensional correlation model of comorbidity type, tolerability level and efficacy attenuation coefficient, subclassifies comorbidities and quantifies the correlation weight of the three. Set up a comorbidity grading correction rule, assess myelosuppression according to WHO grades 0-IV and recovery speed in two dimensions. Define grade IV myelosuppression with a recovery time >14 days as a high-risk tolerable event and encode adverse reactions. The output association results and grading correction coefficients are simultaneously transmitted to the weight calculation module and the matching module.
4. The acute lymphoblastic leukemia scoring system according to claim 3, characterized in that, The feature adaptation module receives baseline data from the acquisition module, including fusion genes and chromosome karyotypes, matches specific biomarkers, establishes an interaction weight matrix between rare subtypes and comorbidities, and clarifies the weight adjustment rules. It has a built-in subtype template update function, which automatically updates the interaction weight matrix after importing new subtype parameters; it outputs matching results and interaction weight data to the weight calculation module and the matching module.
5. The acute lymphoblastic leukemia scoring system according to claim 4, characterized in that, The weight calculation module receives sub-stage dynamic data from the acquisition module, the association results from the evaluation module, and the interactive weight data from the feature adaptation module, and constructs a calculation model using a combination algorithm of random forest and SHAP value. An anchor-triggered weight update mechanism was designed for each treatment sub-stage. When the MRD decrease rate in the late sub-stage was <50%, the weight of gene expression profile indicators was increased; when the drug AUC fluctuation in the consolidation and stabilization phase was >20%, the weight of drug exposure concentration was increased; when the MRD in the late maintenance treatment phase was >0.01%, the weight of relapse-related biomarkers was increased. A stage-specific weight matrix was generated and transmitted to the risk score generation module in real time.
6. The acute lymphoblastic leukemia scoring system according to claim 5, characterized in that, The risk score generation module receives the stage-specific weight matrix and real-time indicator data, and calculates a comprehensive risk score of 0-100 points, corresponding to five risk levels: extremely low risk, low risk, medium risk, high risk, and extremely high risk. Construct the logic linking SHAP values with clinical rules to clarify the direction of the impact of SHAP values on scores in different clinical scenarios; generate a scoring report containing risk level and intervention recommendations, and transmit the report to the matching module.
7. The acute lymphoblastic leukemia scoring system according to claim 6, characterized in that, The matching module receives risk level reports, comorbidity and tolerability association results, and subtype matching results. It then calls up a three-dimensional treatment plan library of risk level, subtype classification, and comorbidity status to match corresponding treatment plans: standardized plans are used for common subtypes, targeted plans are matched for rare subtypes, and dose adjustment plans or drug substitution plans are used for patients with comorbidities. The treatment plan library incorporates a threshold-triggered optimization mechanism, which triggers weight adjustments when MRD > 0.1% for two consecutive cycles; when the incidence of serious adverse reactions > 15%, low-toxicity alternatives are prioritized; the treatment plan library is updated regularly; and treatment plans and patient treatment outcome data are output to the feedback update module.
8. The acute lymphoblastic leukemia scoring system according to claim 7, characterized in that, The feedback update module receives patient treatment outcome data and sets the data collection frequency according to risk level and subtype. For very high-risk patients, data is collected once every 2 weeks and the weight is adjusted once per cycle; for high-risk patients, data is collected once every 4 weeks and the weight is adjusted once per 2 cycles; for intermediate-low-risk patients, data is collected once every 6 weeks and the weight is adjusted once per 3 cycles; and for very low-risk patients, data is collected once every 8 weeks and the weight is adjusted once per 4 cycles. Data is protected by hierarchical encryption. Feedback data is used to correct the coefficients of indicators with large deviations in the weight calculation module, and to update the treatment suitability matching module for treatments with poor efficacy or excessively high adverse reaction rates. At the same time, data collection optimization suggestions are fed back to the collection module.
9. The acute lymphoblastic leukemia scoring system according to claim 8, characterized in that, The acquisition module marks abnormal values in the synchronized data and triggers a supplementary recording reminder when an abnormality occurs.
10. The acute lymphoblastic leukemia scoring system according to claim 9, characterized in that, The risk scoring generation module's reports can be exported in PDF format.