Hematologic tumor cord blood transplantation treatment prognosis risk assessment method and system based on multi-modal data fusion, medium and electronic equipment

By employing a multimodal data fusion method, we screened and integrated multi-source data from patients undergoing umbilical cord blood transplantation for hematological malignancies, and constructed a gradient boosting decision tree model. This approach solved the problems of accuracy and interpretability in prognostic assessment in existing technologies, enabling precise risk stratification and personalized clinical decision support.

CN120954709APending Publication Date: 2025-11-14ANHUI PROVINCIAL HOSPITAL
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202511060223.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing technologies cannot effectively integrate multimodal data in the prognostic assessment of umbilical cord blood transplantation for hematological malignancies, resulting in poor predictive accuracy, lack of interpretability and dynamic update mechanisms, and difficulty in providing individualized risk assessment and treatment recommendations.

Method used

A multimodal data fusion approach is adopted, and a simplified feature set is selected through feature selection and dimensionality reduction algorithms. Combined with feature-level direct cascading, knowledge graph semantic association, and multi-model decision integration strategies, a gradient boosting decision tree model is constructed to generate a standardized risk assessment report to support clinical decision-making.

Benefits of technology

It enables precise stratification of prognostic risks for patients undergoing umbilical cord blood transplantation for hematological malignancies, providing highly accurate risk predictions and interpretable results to support personalized clinical decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120954709A_ABST
    Figure CN120954709A_ABST
Patent Text Reader

Abstract

The invention provides a hematologic tumor cord blood transplantation treatment prognosis risk assessment method and system based on multi-modal data fusion, a medium and electronic equipment, and relates to the technical field of medical artificial intelligence and clinical decision support systems. The method comprises the following steps: acquiring multi-modal data including unit characteristic data of cord blood of a patient, biomarker data and the like; screening out a simplified feature set from the preprocessed multi-modal data; fusing the simplified features based on three complementary strategies of feature level direct cascade, knowledge graph semantic association and multi-model decision integration; and a prediction model is constructed, and a standardized risk assessment report is generated based on the optimized hierarchical risk levels and is used for assisting clinical decision support. According to the method, accurate layering of the prognosis risk of the hematologic tumor umbilical cord blood transplantation patient is achieved, high-accuracy risk prediction is provided, good interpretability and personalized decision support capacity are achieved, and a scientific basis is provided for clinical practice.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to medical artificial intelligence and clinical decision support systems, specifically to a method, system, medium, and electronic device for prognostic risk assessment of umbilical cord blood transplantation for hematological malignancies based on multimodal data fusion. Background Technology

[0002] Umbilical cord blood hematopoietic stem cell transplantation, as an important treatment for hematologic malignancies, has made significant progress in clinical practice. Compared with traditional bone marrow transplantation, umbilical cord blood transplantation has advantages such as easier donor availability, relatively less stringent HLA matching requirements, and a lower risk of graft-versus-host disease (GVHD), providing treatment opportunities for many patients who have difficulty finding suitable bone marrow donors. However, the clinical prognosis of umbilical cord blood transplantation varies significantly, with some patients facing serious complications such as graft-versus-host disease, disease relapse, and infection after transplantation, even leading to treatment failure. Therefore, accurately assessing the prognostic risk of patients and rationally stratifying them is of great significance for optimizing treatment plans and improving clinical outcomes.

[0003] Currently, prognostic assessment of umbilical cord blood transplantation for hematological malignancies suffers from the following shortcomings: 1) Based on limited clinical parameters and relying on expert experience, it has significant limitations in predictive accuracy and individualization; 2) It cannot effectively integrate multimodal data from different dimensions and exhibiting heterogeneous characteristics, nor can it effectively explore complementary information in these multimodal data; 3) It cannot effectively identify features with real predictive value from numerous candidate biomarkers and integrate them with traditional clinical indicators; 4) Most prognostic assessment models lack interpretability and cannot clearly identify the specific factors leading to high risk; 5) Fifth, prognostic risk assessment models generally lack dynamic updating and maintenance mechanisms, which cannot ensure the continued effectiveness of the models.

[0004] Therefore, given the above issues, there is an urgent need to develop a prognostic risk stratification system that can integrate multimodal data, has high-precision predictive capabilities, provides interpretable results, and supports clinical decision-making, so as to provide individualized risk assessment and treatment recommendations for patients undergoing umbilical cord blood transplantation for hematological malignancies. Summary of the Invention

[0005] (a) Technical problems to be solved

[0006] To address the shortcomings of existing technologies, this invention provides a method, system, medium, and electronic device for prognostic risk assessment of umbilical cord blood transplantation for hematological malignancies based on multimodal data fusion, which solves at least one or more of the aforementioned technical problems.

[0007] (II) Technical Solution

[0008] To achieve the above objectives, the present invention provides the following technical solution:

[0009] Firstly, this application proposes a method for prognostic risk assessment of umbilical cord blood transplantation for hematological malignancies based on multimodal data fusion, the method comprising:

[0010] Acquire multimodal data including patient clinical characteristics data, laboratory test data, umbilical cord blood unit characteristic data, and biomarker data;

[0011] A simplified feature set is selected from the preprocessed multimodal data based on feature selection and dimensionality reduction algorithms;

[0012] The simplified features are fused based on three complementary strategies: direct feature-level concatenation, semantic association of knowledge graph, and multi-model decision integration.

[0013] A prediction model is constructed based on the gradient boosting decision tree algorithm, and a standardized risk assessment report is generated based on the optimized hierarchical risk level to assist clinical decision support.

[0014] A standardized risk assessment report is generated based on the stratified risk levels to support clinical decision-making.

[0015] In one embodiment, the preprocessing includes a missing value imputation algorithm, Z-score normalization, and outlier detection based on interquartile range.

[0016] In one embodiment, the feature selection and dimensionality reduction algorithm includes the LASSO regression algorithm and principal component analysis.

[0017] In one embodiment, the fusion of the simplified features based on three complementary strategies—feature-level direct cascading, knowledge graph semantic association, and multi-model decision integration—includes:

[0018] Based on the feature-level direct concatenation strategy, feature vectors of different modalities in the simplified feature set are merged into a single high-dimensional feature vector.

[0019] A semantic network containing disease-gene-drug relationships is constructed based on a knowledge graph semantic association strategy. In this semantic network, nodes represent specific entities, and edges represent the interaction relationships between entities. After mapping patient multimodal data to corresponding nodes, a multi-layer graph convolutional graph neural network is used to learn the semantic propagation path between nodes, generating a feature representation that integrates the semantic information of structured and unstructured data.

[0020] The multi-model decision ensemble strategy constructs independent expert models for each of the multimodal data: a random forest model for clinical characterization data, a temporal convolutional network for laboratory test data, a logistic regression model for umbilical cord blood unit characteristic data, and a support vector machine model for biomarker data.

[0021] The three complementary fusion strategies work together through a hierarchical integration architecture: feature-level direct cascading integration of structured features, knowledge graph mining of cross-modal semantic relationships, and multi-model decision integration to integrate prediction results from each layer.

[0022] In one embodiment, the optimization of the prediction model includes:

[0023] Construct an initial prediction model and calculate the residuals;

[0024] A new decision tree is trained based on residuals;

[0025] The prediction result of the new decision tree is multiplied by the learning rate and then added to the current model's prediction value;

[0026] Repeatedly calculate the residuals and train a new decision tree until the termination condition is met.

[0027] In a preferred embodiment, the optimization of the prediction model further includes:

[0028] The parameters of the prediction model are optimized by combining grid search with cross-validation, and early stopping is used to prevent the prediction model from overfitting the noise and random fluctuations in the training data.

[0029] In one embodiment, the method further includes:

[0030] The prediction model is updated using two methods: regular updates and triggered retraining.

[0031] The routine update is as follows: new case data is automatically included at regular intervals, and the model parameters are fine-tuned using incremental learning methods.

[0032] The triggered retraining is automatically initiated when the C-index on the validation set drops beyond a preset threshold or the deviation of the calibration curve exceeds a preset threshold, and the complete feature selection and model training process is re-executed.

[0033] Secondly, this application also proposes a prognostic risk assessment system for umbilical cord blood transplantation for hematological malignancies based on multimodal data fusion, the system comprising:

[0034] The data acquisition and processing module is configured to acquire multimodal data, including patient clinical characterization data, laboratory test data, umbilical cord blood unit characteristic data, and biomarker data.

[0035] The feature extraction and dimensionality reduction module is configured to filter a simplified feature set from the preprocessed multimodal data based on feature selection and dimensionality reduction algorithms;

[0036] The multimodal fusion module is configured to fuse the simplified features based on three complementary strategies: direct feature-level concatenation, knowledge graph semantic association, and multi-model decision integration.

[0037] The model training and risk assessment module is configured to analyze the fused simplified features based on a pre-built prediction model using a gradient boosting decision tree algorithm to obtain hierarchical risk levels.

[0038] The decision support module is configured to generate standardized risk assessment reports based on the stratified risk levels to assist in clinical decision support.

[0039] Thirdly, this application further proposes an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the prognostic risk assessment method for umbilical cord blood transplantation for hematological malignancies based on multimodal data fusion as described in any of the preceding claims.

[0040] Fourthly, a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the prognostic risk assessment method for umbilical cord blood transplantation for hematologic malignancies as described in any of the preceding claims.

[0041] (III) Beneficial Effects

[0042] This invention provides a method, system, medium, and electronic device for prognostic risk assessment of umbilical cord blood transplantation for hematological malignancies based on multimodal data fusion. Compared with existing technologies, it has at least the following advantages:

[0043] This application integrates multimodal data and combines advanced feature selection, data fusion, and machine learning algorithms to achieve precise stratification of prognostic risks for patients undergoing umbilical cord blood transplantation for hematological malignancies. It not only provides highly accurate risk prediction but also has good interpretability and personalized decision support capabilities, providing a scientific basis for clinical practice. Attached Figure Description

[0044] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0045] Figure 1 This is a flowchart of a method for assessing the prognostic risk of umbilical cord blood transplantation for hematological malignancies based on multimodal data fusion, as described in an embodiment of the present invention.

[0046] Figure 2 This is a schematic diagram of the prognostic risk assessment system for umbilical cord blood transplantation for hematological malignancies based on multimodal data fusion, as described in this embodiment of the invention.

[0047] Figure 3 This is a flowchart of data acquisition and preprocessing in an embodiment of the present invention;

[0048] Figure 4 This is a flowchart of feature extraction and dimensionality reduction in an embodiment of the present invention;

[0049] Figure 5 This is a flowchart of multimodal fusion in an embodiment of the present invention;

[0050] Figure 6 This is a flowchart of model training and risk assessment in an embodiment of the present invention;

[0051] Figure 7 This is a schematic diagram of a risk assessment report in an embodiment of the present invention. Detailed Implementation

[0052] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention are described clearly and completely. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0053] Umbilical cord blood hematopoietic stem cell transplantation, as an important treatment for hematologic malignancies, has made significant progress in clinical practice. Compared with traditional bone marrow transplantation, umbilical cord blood transplantation has advantages such as easier donor availability, relatively less stringent HLA matching requirements, and a lower risk of graft-versus-host disease (GVHD), providing treatment opportunities for many patients who have difficulty finding suitable bone marrow donors. However, the clinical prognosis of umbilical cord blood transplantation varies significantly, with some patients facing serious complications such as graft-versus-host disease, disease relapse, and infection after transplantation, even leading to treatment failure. Therefore, accurately assessing the prognostic risk of patients and rationally stratifying them is of great significance for optimizing treatment plans and improving clinical outcomes.

[0054] Currently, the prognostic assessment of umbilical cord blood transplantation for hematological malignancies mainly faces the following technical challenges:

[0055] First, existing clinical scoring systems primarily rely on limited clinical parameters, such as disease type, stage, age, and comorbidity scores, to construct risk assessment models. These traditional assessment methods depend on expert experience and have significant limitations in predictive accuracy and individualization. Particularly in borderline risk patient groups, the predictive results are often less than ideal, making it difficult to provide precise guidance for clinical decision-making.

[0056] Secondly, the prognosis of umbilical cord blood transplantation for hematological malignancies is influenced by a combination of factors, including the patient's underlying condition, the biological characteristics of the disease, the unit mass of umbilical cord blood, and molecular-level biomarkers. These data from different dimensions exhibit heterogeneity, differing in dimensions, distribution, and collection frequency. Effectively integrating these multimodal data and uncovering their complementary information is a key challenge for improving predictive accuracy.

[0057] Third, biomarkers, as prognostic indicators at the molecular level, have shown significant value in hematological malignancy research in recent years. However, such data are often high-dimensional, and identifying truly predictive features from numerous candidate biomarkers and effectively integrating them with traditional clinical indicators remains a technical challenge.

[0058] Fourth, most existing prognostic assessment models lack interpretability and fail to clearly identify the specific factors leading to high risk. This limits clinicians' understanding and acceptance of the prediction results and makes it difficult to develop targeted intervention measures based on them.

[0059] Fifth, there is a general lack of dynamic updating and maintenance mechanisms for prognostic risk assessment models. As new data accumulates and clinical practice develops, model performance may gradually decline. Ensuring the continued effectiveness of the model is a challenge for the long-term application of the system.

[0060] In view of the above problems, there is an urgent need to develop a prognostic risk stratification system that can integrate multimodal data, has high-precision predictive capabilities, provides interpretable results, and supports clinical decision-making, so as to provide individualized risk assessment and treatment recommendations for patients undergoing umbilical cord blood transplantation for hematological malignancies.

[0061] This application provides a method, system, medium, and electronic device for prognostic risk assessment of umbilical cord blood transplantation for hematological malignancies based on multimodal data fusion. It at least solves the problem that the existing technology cannot effectively integrate multimodal data, resulting in poor accuracy of prognostic prediction, and achieves the goal of precise stratification of prognostic risks for umbilical cord blood transplantation patients with hematological malignancies.

[0062] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.

[0063] Example 1:

[0064] Firstly, this invention proposes a method for prognostic risk assessment of umbilical cord blood transplantation for hematological malignancies based on multimodal data fusion, see [link to relevant documentation]. Figure 1 The method includes:

[0065] S1. Acquire multimodal data including patient clinical characteristics data, laboratory test data, umbilical cord blood unit characteristic data, and biomarker data;

[0066] S2. Based on feature selection and dimensionality reduction algorithms, a simplified feature set is selected from the preprocessed multimodal data;

[0067] S3. The simplified features are fused based on three complementary strategies: feature-level direct cascading, knowledge graph semantic association, and multi-model decision integration.

[0068] S4. Construct a prediction model based on the gradient boosting decision tree algorithm, and generate a standardized risk assessment report based on the optimized hierarchical risk level to assist clinical decision support;

[0069] S5. Generate a standardized risk assessment report based on the stratified risk levels to support clinical decision-making.

[0070] The method for prognostic risk assessment of umbilical cord blood transplantation for hematological malignancies proposed in this embodiment integrates multimodal data fusion and combines advanced feature selection, data fusion and machine learning algorithms to achieve precise stratification of prognostic risks for umbilical cord blood transplantation patients with hematological malignancies.

[0071] The following is in conjunction with the appendix Figure 1-7 The following details the implementation process of an embodiment of the present invention, including explanations of the specific steps S1-S5.

[0072] S1. Acquire multimodal data including patient clinical characterization data, laboratory test data, umbilical cord blood unit characteristic data, and biomarker data.

[0073] In practice, a system based on a multimodal data fusion approach for prognostic risk assessment of umbilical cord blood transplantation for hematological malignancies was deployed in the hematology department of a tertiary-level hospital. Figure 2 As shown, the system's data acquisition and processing module uses a data interface compatible with the hospital's electronic medical record system. It collects four types of multimodal raw data—patient clinical characteristics, laboratory test data, umbilical cord blood unit characteristic data, and biomarker data—according to Good Clinical Practice (GCP) standards. The collected raw data is then preprocessed using data standardization methods such as missing value imputation algorithms, Z-score normalization, and outlier detection based on interquartile range (IQR). The specific process of data acquisition and preprocessing is as follows: Figure 3 As shown.

[0074] For newly enrolled acute myeloid leukemia patients, clinical data were stored in structured JSON format, including 18 basic information items such as age, sex, body mass index, ECOG performance status, demographic characteristics, disease type, disease stage, previous treatment history, and comorbidities. This data reflects the patient's baseline condition and disease progression, providing basic clinical information for predicting transplant prognosis. Laboratory test data were organized in time-series matrix form, including indicators such as complete blood count, biochemistry, and coagulation function at different time points before, during, and after transplantation, reflecting the patient's organ function and hematopoietic reconstitution process. Umbilical cord blood unit characteristic data were recorded in standardized CSV format, including key indicators such as HLA typing degree, nucleated cell count, CD34+ cell content, and cryopreservation time. These parameters are directly related to graft quality and implantation success rate. Biomarker data were stored in high-dimensional matrix form, including data on cytokine levels, gene expression profiles, and microbiome, used to predict the risk of immune rejection and graft-versus-host disease at the molecular level.

[0075] After collecting the raw data for the four types of multimodal data mentioned above, data quality and consistency are ensured through missing value imputation algorithms, Z-score normalization, and outlier detection methods based on interquartile ranges. Specifically:

[0076] Missing value imputation algorithms are applied to the unavoidable data loss during clinical data collection to ensure data integrity and model training quality. Different processing strategies are adopted according to the data loss mechanism and proportion.

[0077] Z-score standardization transforms the original data value x into a standardized value z = (x-μ) / σ by calculating the mean (μ) and standard deviation (σ) of each feature. This transforms clinical indicators, laboratory parameters, and biomarker data of different dimensions into the same distribution interval with a mean of 0 and a standard deviation of 1, eliminating the influence of dimensional differences on model training while preserving the relative distribution characteristics of the data.

[0078] Outlier detection methods based on the interquartile range (IQR) calculate the first quartile (Q1) and third quartile (Q3) of each feature. Data points smaller than Q1 - 1.5 × IQR or greater than Q3 + 1.5 × IQR are identified as potential outliers. For continuous medical indicators, cross-validation with clinical normal reference ranges is used for further confirmation. Data points identified as outliers are either truncated and replaced with critical values ​​or corrected through expert review to prevent extreme values ​​from adversely affecting subsequent model training and improve the robustness of system predictions. Outlier detection methods can identify and handle outliers, ensuring data quality.

[0079] In a preferred embodiment, the missing value imputation algorithm first classifies and evaluates the missing data, categorizing the missing mechanisms into three types: completely random missing (MCAR), random missing (MAR), and non-random missing (MNAR). For data items with a missing proportion of less than 5%, the mean / median / mode imputation method is used, selecting the appropriate statistic for simple replacement based on the distribution characteristics of the feature. For continuous variables with a missing proportion between 5% and 15%, multiple imputation is used, generating multiple possible complete datasets through Markov chain Monte Carlo (MCMC) technology, and taking the average as the final imputation result. For missing values ​​in time series test data, a forward or backward imputation strategy based on time correlation is used, inferring reasonable values ​​for missing points using historical or subsequent data of the same indicator from patients.

[0080] Missing value imputation algorithms are designed with differentiated imputation strategies for different types of clinical data: For categorical variables such as disease classification and HLA typing, a random forest-based missing value imputation method is used, which uses other relevant features to construct a decision tree model to predict the possible values ​​of missing items; for continuous variables such as laboratory test indicators, the K nearest neighbor (KNN) algorithm is used, which finds the K most similar samples based on the similarity of patients in other features, and calculates the missing value by weighted average; for some highly correlated clinical indicators, a multiple linear regression model is used, which uses other existing indicators to predict the reasonable values ​​of missing indicators. The missing value imputation algorithm sets strict missing value handling rules: when the missing rate of non-critical features in a single patient sample exceeds 30%, the sample is removed from the training set; when the missing rate of a single feature in the overall dataset exceeds 40%, the predictive importance of the feature is assessed, non-critical features are removed, and critical features are retained or removed through expert review; the missing value of the following core features will prevent the system from making reliable predictions: HLA typing, umbilical cord blood nucleated cell count, patient age, disease type and stage, pretreatment protocol type, etc., these features are marked as "required". The system employs a mandatory data collection mechanism to ensure the integrity of the data. The missing value imputation algorithm uses an iterative imputation framework to handle multivariate missing values: first, all variables are initialized and filled; then, each variable is finely imputed, with a single variable containing a missing value used as the dependent variable and the others as independent variables to build the prediction model. This process is repeated through multiple iterations until the imputation results converge or the maximum number of iterations is reached. The system performs anomaly detection on the imputed data to ensure the imputation results are medically sound. Normal reference ranges are set for clinical laboratory indicators, and manual review is triggered when the imputation results exceed the reference range. The imputation quality of the missing value imputation algorithm is assessed in two ways: internal assessment uses cross-validation, randomly marking some known data as "artificially missing," comparing the imputed values ​​with the true values ​​after applying the imputation algorithm, and calculating the root mean square error (RMSE) and mean absolute error (MAE) to evaluate imputation accuracy; external assessment uses sensitivity analysis to compare the stability of the model's prediction results after using different imputation methods. When different imputation methods lead to significant differences in prediction results, the system triggers a warning and suggests collecting more complete patient data.

[0081] S2. Based on feature selection and dimensionality reduction algorithms, a simplified feature set is selected from the preprocessed multimodal data.

[0082] The standardized multimodal data from step S1 is input into the preliminary screening index library for prognostic factors. Then, starting from this library, a subset of features with predictive value is selected using the LASSO regression algorithm and principal component analysis to form a simplified feature set. See details in [link to documentation]. Figure 4 .

[0083] In practice, key features are selected from a preliminary screening index database of prognostic factors. This database comprises four main categories: patient baseline characteristics, disease characteristics, transplantation characteristics, and biomarkers, totaling 100 potential predictive factors. Patient baseline characteristics include demographic indicators such as age and performance status; disease characteristics include indicators such as typing, staging, and molecular markers; transplantation characteristics include transplantation-related indicators such as matching accuracy and cell count; and biomarkers include molecular-level indicators such as cytokines and gene expression. All these features have been reviewed and confirmed by clinical experts to be relevant to transplant prognosis.

[0084] The patient baseline characteristics category includes 18 demographic and clinical baseline indicators: age (in years), sex (male / female), body mass index (kg / m²). 2 ECOG performance status score (0-4 points), Karnofsky performance status score (0-100 points), comorbidity score (HCT-CI score, 0-29 points), history of hypertension (yes / no), history of diabetes (yes / no), history of malignancy (yes / no), cardiac function status (left ventricular ejection fraction %), pulmonary function status (FEV1% predictive value), liver function status (Child-Pugh score), and renal function status (eGFR ml / min / 1.73m³). 2 Nutritional status score (NRS2002 score), smoking history (pack-years), family genetic background (East Asia / Europe / Africa / Other), time from initial diagnosis to transplant (months), number of previous lines of treatment (numerical value);

[0085] The disease characteristics category includes 21 biological indicators: disease subtype (acute myeloid leukemia / acute lymphoblastic leukemia / myelodysplastic syndrome / multiple myeloma / other), disease stage (initial diagnosis / first complete remission / relapsed / refractory), WHO classification subtype, chromosome karyotype (normal / complex / deletion / translocation), cytogenetic risk stratification (low / intermediate / high risk), bone marrow blast cell percentage (%), peripheral blood blast cell count (×10⁻¹⁰). 9 / L), minimal residual disease status (negative / positive), FLT3-ITD mutation (present / absent), NPM1 mutation (present / absent), CEBPA biallelic mutation (present / absent), IDH1 / 2 mutation (present / absent), RUNX1 mutation (present / absent), ASXL1 mutation (present / absent), TP53 mutation (present / absent), BCR-ABL fusion gene (present / absent), PML-RARα fusion gene (present / absent), MLL rearrangement (present / absent), degree of myelofibrosis (0-3), degree of myelohematopoietic suppression (mild / moderate / severe), previous chemotherapy efficacy (complete remission / partial remission / stable disease / progressive disease);

[0086] The transplantation characteristics category includes 24 indicators related to the transplantation process: transplant type (umbilical cord blood / bone marrow / peripheral blood stem cells), single / double umbilical cord blood (single / double), HLA typing (matched / incompatible, 5 / 10, 6 / 10, 7 / 10, 8 / 10, 9 / 10), ABO blood type compatibility (matched / major incompatible / minor incompatible / bidirectional incompatible), conditioning regimen (myeloablative / reduced intensity / non-myeloablative), total dose in conditioning regimen (mg / kg or Gy), anti-thymocyte globulin use (present / absent and dose), total body irradiation use (present / absent and dose), GVHD prophylaxis regimen (cyclosporine + short-course methotrexate / tacrolimus + short-course methotrexate / tacrolimus + mycophenolate mofetil / cyclosporine + mycophenolate mofetil + others), and total nucleated cell count in umbilical cord blood (×10). 7 / kg), umbilical cord blood CD34+ cell count (×10) 5 / kg), number of umbilical cord blood mononuclear cells (×10) 7 / kg), umbilical cord blood cryopreservation time (months), donor / recipient gender matching (male to male / male to female / female to male / female to female), donor / recipient CMV status (+ / +, + / -, - / +, - / -), stem cell reinfusion date and reinfusion time (day x during transplantation), neutrophil engraftment time (days), platelet engraftment time (days), post-transplant infection (CMV / EBV / bacteria / fungus), occurrence of acute GVHD within 100 days post-transplantation (grades 0-IV), red blood cell transfusion dependence time (days), platelet transfusion dependence time (days), time interval from the last chemotherapy before transplantation to transplantation (days), umbilical cord blood unit source (public bank / private bank);

[0087] The biomarker category includes 37 molecular-level indicators: sCD26 / DDPIV (ng / ml), interleukin-1α (IL-1α, pg / ml), interleukin-1β (IL-1β, pg / ml), interleukin-6 (IL-6, pg / ml), interleukin-10 (IL-10, pg / ml), tumor necrosis factor-α (TNF-α, pg / ml), interferon-γ (IFN-γ, pg / ml), C-reactive protein (CRP, mg / L), serum ferritin (ng / ml), D-dimer (μg / ml), serum albumin (g / L), lactate dehydrogenase (LDH, U / L), total bilirubin (μmol / L), serum creatinine (μmol / L), blood urea nitrogen (mmol / L), aspartate aminotransferase (AST, U / L), and alanine. Transaminase (ALT, U / L), alkaline phosphatase (ALP, U / L), gamma-glutamyl transferase (GGT, U / L), serum sodium (mmol / L), serum potassium (mmol / L), serum calcium (mmol / L), serum magnesium (mmol / L), serum phosphorus (mmol / L), percentage of CD3+ cells (%), percentage of CD4+ cells (%), percentage of CD8+ cells (%), percentage of CD19+ cells (%), percentage of CD16+56+ cells (%), percentage of regulatory T cells (%), WT1 gene expression level, BAALC gene expression level, plasma EVs concentration, miR-29a expression level, miR-146a expression level, miR-155 expression level, telomerase activity, telomere length detection, and whole blood gene expression profile (previously validated 78-gene combination score).

[0088] The LASSO regression algorithm with L1 regularization automatically filters for a subset of features with predictive value. Specifically, for historical data from 300 AML patients in a cohort, the system applies the LASSO regression algorithm with L1 regularization. By adding a penalty term to the loss function and determining the optimal regularization parameter λ through cross-validation, the regression coefficients of non-important features are reduced to zero, automatically filtering for a subset of features with significant prognostic predictive value. For example, after filtering, 5 features from the patient's basic features class, 7 features from the disease features class, 8 features from the transplantation features class, and 12 features from the biomarkers class are retained, for a total of 32 features. The LASSO regression algorithm effectively solves the multicollinearity problem and overfitting risk in high-dimensional feature selection.

[0089] For the retained biomarker features, principal component analysis (PCA) is used to reduce the dimensionality of the high-dimensional biomarker data, and cross-validation is employed to ensure the stability of the selected features. Specifically, PCA transforms the original high-dimensional biomarker features into linearly independent principal components, where the cumulative contribution rate represents the proportion of original information retained by the selected principal components. When a preset threshold is reached, the selected principal components are considered to fully represent the original feature information, achieving data dimensionality reduction while preserving key information. For example, the top 5 principal components with a cumulative contribution rate of 85% are selected to represent the original 12 biomarker features. It should be noted that cross-validation ensures that the selected features have consistent predictive ability across different data subsets, and the reliability of the features is assessed by calculating a feature selection stability index.

[0090] The streamlined feature set ultimately contains high-predictive-value features from four major categories. These features have independent predictive contributions and complementarity, which significantly improves model training efficiency and reduces the risk of overfitting.

[0091] In one embodiment, the final simplified feature set contains 28 features (23 original features and 5 principal component features).

[0092] S3. The simplified features are fused based on three complementary strategies: feature-level direct cascading, knowledge graph semantic association, and multi-model decision integration.

[0093] Three complementary strategies—feature-level direct cascading, knowledge graph semantic association, and multi-model decision ensemble—are employed to integrate a concise feature set selected from multi-source heterogeneous data. See also Figure 5 Specifically:

[0094] First, a feature-level direct cascading strategy is used to merge feature vectors from different modalities into a single high-dimensional feature vector. This strategy directly combines feature vectors from different modalities into a single high-dimensional feature vector, and an adaptive weight allocation mechanism dynamically adjusts the weight ratios based on the predictive performance metrics (such as AUC value or C-index) of each modality on the validation set. This ensures that modalities with high predictive performance receive a larger weight, enhancing the overall predictive ability of the fused features. For example, in one embodiment, clinical characterization data has a weight of 0.25, laboratory test data has a weight of 0.20, cord blood unit characteristic data has a weight of 0.30, and biomarker data has a weight of 0.25.

[0095] Secondly, the knowledge graph semantic association strategy constructs a semantic network containing disease-gene-drug relationships, where nodes represent specific entities and edges represent interactions between entities. After mapping patient multimodal data to corresponding nodes, a multi-layer graph convolutional neural network learns the semantic propagation paths between nodes, generating feature representations that integrate semantic information from both structured and unstructured data, effectively discovering cross-modal interactions that are difficult to capture using traditional methods. For example, in one embodiment, 196 disease entities, 324 gene entities, and 128 drug entities related to hematological malignancies, along with their interrelationships, are encoded into a semantic network. A three-layer graph convolutional neural network learns the semantic propagation paths between nodes, generating feature representations that integrate semantic information from both structured and unstructured data.

[0096] Furthermore, the multi-model decision ensemble strategy constructs independent expert models for each of the four data sources: a random forest model for clinical characterization data, a temporal convolutional network for laboratory test data, a logistic regression model for umbilical cord blood unit characteristic data, and a support vector machine model for biomarker data. A weighted voting method is used to integrate the prediction results of each model, and the optimal weight combination is determined through Bayesian optimization to achieve multi-model decision ensemble. Simultaneously, a model discrepancy analysis mechanism is introduced to identify key features that lead to prediction differences and give them extra attention in the final decision, preserving the unique predictive advantages of each modality.

[0097] It should be noted that in the above embodiments, the three fusion strategies work collaboratively through a hierarchical integration architecture: direct cascading integration of structured features at the feature level, cross-modal semantic relationships mined from knowledge graphs, and integration of prediction results from each layer through multi-model decision fusion. The system evaluates the fusion effect of each strategy in real time and dynamically adjusts its weight ratio. The final fusion effect significantly improves the prediction accuracy compared to the single-modal prediction model, especially showing a significant improvement in the classification effect for patients at borderline risk.

[0098] In a preferred embodiment, the system evaluates the predictive performance of each strategy and dynamically adjusts the weights: feature-level fusion weight 0.4, knowledge graph fusion weight 0.3, and multi-model integration weight 0.3.

[0099] S4. Construct a prediction model based on the gradient boosting decision tree algorithm, and generate a standardized risk assessment report based on the optimized hierarchical risk level to assist clinical decision support.

[0100] like Figure 6As shown, the Gradient Boosting Decision Tree (GBDT) algorithm achieves high-precision prediction by integrating multiple decision trees. Its process comprises four key steps: First, an initial prediction model is constructed and residuals are calculated; then, a new decision tree is trained based on the residuals; next, the prediction result of the new tree is multiplied by the learning rate and added to the current model's prediction value; finally, the residuals are repeatedly calculated and new trees are trained until the termination condition is met. In a preferred embodiment, for the collected 400 historical cases, the system is divided into a training set and a validation set in a 7:3 ratio. This algorithm effectively handles mixed-type features, captures complex nonlinear relationships between features, and is highly robust to data containing missing values, making it particularly suitable for clinical prediction problems involving multiple factors, such as hematological malignancies and umbilical cord blood transplantation.

[0101] Model parameters are optimized using grid search combined with cross-validation, and early stopping is employed to prevent overfitting. Model parameter optimization utilizes grid search combined with k-fold cross-validation to determine the optimal parameter combination, primarily focusing on key parameters such as the maximum tree depth, minimum number of leaf node samples, learning rate, regularization coefficient, and feature sampling ratio. After defining the parameter search space, k-fold cross-validation is performed on the training set, selecting the parameter combination with the best performance on the validation set as the final model parameters. Simultaneously, Bayesian optimization is introduced to accelerate the parameter search process, effectively reducing computational resource consumption. Early stopping is a technique to prevent overfitting. By monitoring the model's performance changes on the validation set, training is terminated early when performance begins to decline. The system divides the training data into training and validation sets, evaluating the model's performance metrics on the validation set after each iteration. Training is stopped and the best model is saved when performance no longer improves after n consecutive iterations or when the performance metrics begin to decline. Avoiding overfitting improves the model's generalization ability on unseen data, prevents the model from overfitting noise and random fluctuations in the training data, and ensures stable predictive performance on new patient data.

[0102] In a preferred embodiment, the system determines the optimal parameter combination through grid search combined with 5-fold cross-validation: a maximum tree depth of 5, a minimum number of leaf node samples of 30, a learning rate of 0.05, and a feature sampling ratio of 0.8. To prevent overfitting, the system implements early stopping, stopping training when performance no longer improves after 10 consecutive iterations on the validation set.

[0103] The model outputs a continuous risk score (0-100 points) and categorizes patients into three risk levels—low risk (0-30 points), intermediate risk (31-70 points), and high risk (71-100 points)—based on clinically validated thresholds. Overall survival and progression-free survival were used as the primary clinical endpoints to evaluate model performance.

[0104] The risk scoring system converts the predicted values ​​output by the model into risk probabilities in the 0-1 range using a logistic function, linearly mapping them to a scoring range of 0-100 points for easy understanding and use by clinicians. The scoring formula is: Risk Score = 100 × Predicted Probability, where the predicted probability is converted from the logarithmic probability value output by GBDT into a probability value in the 0-1 range using a logistic function. A higher score indicates a greater risk of poor prognosis after transplantation.

[0105] The threshold setting was determined based on clinical validation data. By analyzing the relationship between risk scores and actual prognostic outcomes, the optimal cutoff point was determined using the Youden Index maximization method. The system classifies patients into three risk levels: low risk (0-30 points), intermediate risk (31-70 points), and high risk (71-100 points). This three-tiered stratification method aligns with clinical decision-making practices and provides differentiated treatment recommendations for patients at different risk levels.

[0106] Overall survival is defined as the time from the date of transplantation to death from any cause, while progression-free survival is defined as the time from the date of transplantation to disease progression or death from any cause. The difference between the two is that progression-free survival takes into account both disease relapse and death events.

[0107] The model performance evaluation methods included: using the time-dependent area under the receiver operating characteristic (AUC) curve to assess the discriminative ability at different time points; applying Harrell's C-index to assess the accuracy of the model in ranking survival time; evaluating the consistency between predicted and actual risks using calibration curves; and using decision curve analysis to assess the clinical net benefit of the model at different decision thresholds. In an external validation cohort (150 cases), the model demonstrated excellent predictive performance: the AUC for 2-year overall survival prediction was 0.83 (95% CI: 0.78–0.88), the C-index was 0.81 (95% CI: 0.76–0.86), and the calibration curves showed good consistency between predicted and actual risks.

[0108] S5. Generate a standardized risk assessment report based on the stratified risk levels to support clinical decision-making.

[0109] In practice, standardized risk assessment reports include patient risk scores, percentile rankings, key risk factor analysis, and personalized treatment recommendations. Differentiated treatment guidance is provided based on risk levels, and the SHAP value calculation framework enables interpretable analysis of risk factors. The system maintains predictive accuracy through both routine updates and triggered retraining, ensuring continued effectiveness in clinical practice. Risk assessment reports are as follows: Figure 7 As shown. Wherein:

[0110] The standardized risk assessment report adopts a hierarchical structure and includes four main components: the patient basic information area displays the patient's demographic characteristics and basic disease information; the risk quantification area displays the patient's risk score and percentile ranking in the same patient group, using red, yellow, and green to identify different risk levels, intuitively presenting the patient's risk level; the key influencing factors area lists the five factors that contribute most to the prediction of patient risk and their influence weights, visually demonstrating the positive and negative contributions of each factor to the final prediction result through SHAP values; and the personalized recommendations area provides differentiated treatment recommendations based on the patient's risk level and specific characteristics.

[0111] The risk assessment report provides guidance to patients in three aspects: For low-risk patients (0-30 points), it recommends a standard-intensity conditioning regimen and routine GVHD prevention measures, and reduces the use of immunosuppressants in the early post-transplant period. The report details monitoring indicators and follow-up frequencies. For intermediate-risk patients (31-70 points), it recommends adjusting the conditioning intensity or GVHD prevention regimen based on specific risk factors. The report provides personalized risk mitigation strategies for specific high-risk factors, such as recommending specific immunomodulatory therapy for patients with high levels of inflammatory factors. For high-risk patients (71-100 points), it recommends considering an enhanced conditioning regimen or adopting novel GVHD prevention strategies, increasing the frequency of post-transplant complication monitoring, and the report details early warning signs and emergency treatment procedures.

[0112] The interpretability analysis of key risk factors is based on the SHAP (SHapley Additive exPlanations) value calculation framework, which quantifies the contribution of each feature to the final prediction as a SHAP value. The system generates local interpretability analysis in the form of horizontal bar charts, visually displaying positive influencing factors (increasing risk) and negative influencing factors (reducing risk) and their degree of contribution. The system also provides clinical explanatory texts for key risk factors, transforming technical indicators into clinically understandable meanings for physicians. The level of detail in the explanation varies depending on the type of factor: brief explanations are provided for common clinical indicators such as age and disease stage; moderately detailed mechanistic explanations are provided for complex biomarkers such as cytokine levels and gene expression characteristics; and the most detailed literature support and mechanism of action explanations are provided for unconventional high-impact factors identified by the system.

[0113] Taking a 45-year-old male AML patient as an example, the risk score generated by the system was 65 points, which belongs to the medium risk level, and the percentile ranking in the same patient group was 72%.

[0114] Key influencing factor analysis revealed that the main contributing factors to the patient's risk score included: incomplete HLA matching (6 / 10), high IL-6 level (42.3 pg / ml), first-time relapse, low CD34+ cell count (1.8 × 10^5 / kg), and age. The system visualized the contribution of each factor using SHAP values ​​and provided clinical interpretation.

[0115] Based on the patient's risk level and specific characteristics, the system generates personalized treatment recommendations: "A reduced-intensity conditioning regimen containing ATG is recommended, along with enhanced GVHD prevention measures. Consideration should be given to double umbilical cord blood transplantation to increase the stem cell dose. Post-transplant IL-6 level monitoring should be strengthened, and IL-6 receptor antagonists should be considered as a potential intervention. Minimal residual disease monitoring should be performed weekly for the first three months post-transplantation to detect early signs of disease relapse."

[0116] In a preferred embodiment, a model update mechanism is implemented to update the model. This mechanism includes two methods: regular updates and triggered retraining. Regular updates automatically incorporate new case data at intervals (e.g., quarterly), using incremental learning to fine-tune model parameters and maintain the predictive model's adaptability to new data distributions. Triggered retraining automatically initiates when the C-index on the validation set drops by more than 3% or the calibration curve deviation exceeds a preset threshold, re-executing the complete feature selection and model training process. The system maintains an external clinical validation dataset, regularly evaluating the model's performance across different centers. When performance differences are identified, the causes are analyzed and the model is optimized to ensure stable predictive performance across various clinical environments. For example, after one year of operation, the system has incorporated 150 new cases, completed four regular updates and one triggered retraining. Through continuous evaluation by the external validation cohort, the system has maintained stable predictive performance, with the C-index remaining above 0.80. The application of this system significantly improves the scientific rigor and individualization of clinical decision-making, providing strong support for optimizing umbilical cord blood transplantation treatment plans.

[0117] This completes the entire process of the prognostic risk assessment method for umbilical cord blood transplantation for hematological malignancies based on multimodal data fusion proposed in this embodiment and the preferred embodiment.

[0118] Example 2:

[0119] Secondly, this invention also provides a prognostic risk assessment system for umbilical cord blood transplantation for hematological malignancies based on multimodal data fusion, see [link to relevant documentation]. Figure 2 The system includes:

[0120] The data acquisition and processing module is configured to acquire multimodal data, including patient clinical characterization data, laboratory test data, umbilical cord blood unit characteristic data, and biomarker data.

[0121] The feature extraction and dimensionality reduction module is configured to filter a simplified feature set from the preprocessed multimodal data based on feature selection and dimensionality reduction algorithms;

[0122] The multimodal fusion module is configured to fuse the simplified features based on three complementary strategies: direct feature-level concatenation, knowledge graph semantic association, and multi-model decision integration.

[0123] The model training and risk assessment module is configured to analyze the fused simplified features based on a pre-built prediction model using a gradient boosting decision tree algorithm to obtain hierarchical risk levels.

[0124] The decision support module is configured to generate standardized risk assessment reports based on the stratified risk levels to assist in clinical decision support.

[0125] It is understood that the prognostic risk assessment system for hematologic malignancies based on umbilical cord blood transplantation based on multimodal data fusion provided in this embodiment of the invention corresponds to the above-mentioned prognostic risk assessment method for hematologic malignancies based on umbilical cord blood transplantation based on multimodal data fusion. The explanation, examples, and beneficial effects of the relevant content can be referred to the corresponding content in the prognostic risk assessment method for hematologic malignancies based on umbilical cord blood transplantation based on multimodal data fusion, and will not be repeated here.

[0126] Example 3:

[0127] Thirdly, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps of the method for prognostic risk assessment of umbilical cord blood transplantation for hematological malignancies based on multimodal data fusion as described in any of the above embodiments and their preferred embodiments. The method mainly includes:

[0128] S1. Acquire multimodal data including patient clinical characteristics data, laboratory test data, umbilical cord blood unit characteristic data, and biomarker data;

[0129] S2. Based on feature selection and dimensionality reduction algorithms, a simplified feature set is selected from the preprocessed multimodal data;

[0130] S3. The simplified features are fused based on three complementary strategies: feature-level direct cascading, knowledge graph semantic association, and multi-model decision integration.

[0131] S4. Construct a prediction model based on the gradient boosting decision tree algorithm, and generate a standardized risk assessment report based on the optimized hierarchical risk level to assist clinical decision support;

[0132] S5. Generate a standardized risk assessment report based on the stratified risk levels to support clinical decision-making.

[0133] It is understood that the electronic device for prognostic risk assessment of umbilical cord blood transplantation for hematological malignancies based on multimodal data fusion provided in this embodiment of the invention corresponds to the above-mentioned method and system for prognostic risk assessment of umbilical cord blood transplantation for hematological malignancies based on multimodal data fusion. The explanation, examples, and beneficial effects of the relevant content can be referred to the corresponding content in the method and system for prognostic risk assessment of umbilical cord blood transplantation for hematological malignancies based on multimodal data fusion, and will not be repeated here.

[0134] Example 4:

[0135] Fourthly, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method for prognostic risk assessment of hematologic malignancy umbilical cord blood transplantation based on multimodal data fusion as described in any of the foregoing embodiments and their preferred embodiments, the method comprising:

[0136] S1. Acquire multimodal data including patient clinical characteristics data, laboratory test data, umbilical cord blood unit characteristic data, and biomarker data;

[0137] S2. Based on feature selection and dimensionality reduction algorithms, a simplified feature set is selected from the preprocessed multimodal data;

[0138] S3. The simplified features are fused based on three complementary strategies: feature-level direct cascading, knowledge graph semantic association, and multi-model decision integration.

[0139] S4. Construct a prediction model based on the gradient boosting decision tree algorithm, and generate a standardized risk assessment report based on the optimized hierarchical risk level to assist clinical decision support;

[0140] S5. Generate a standardized risk assessment report based on the stratified risk levels to support clinical decision-making.

[0141] It is understood that the storage medium for prognostic risk assessment of hematologic malignancies based on umbilical cord blood transplantation based on multimodal data fusion provided in this embodiment of the invention corresponds to the above-mentioned method and system for prognostic risk assessment of hematologic malignancies based on umbilical cord blood transplantation based on multimodal data fusion. The explanation, examples, and beneficial effects of the relevant content can be referred to the corresponding content in the method and system for prognostic risk assessment of hematologic malignancies based on umbilical cord blood transplantation based on multimodal data fusion, and will not be repeated here.

[0142] In summary, compared with existing technologies, it has the following beneficial effects:

[0143] 1. This application integrates multimodal data and combines advanced feature selection, data fusion and machine learning algorithms to achieve precise stratification of prognostic risks for patients undergoing umbilical cord blood transplantation for hematological malignancies. It not only provides highly accurate risk prediction, but also has good interpretability and personalized decision support capabilities, providing a scientific basis for clinical practice.

[0144] 2. This application achieves effective fusion of multimodal data, comprehensively considering multi-dimensional data such as patient clinical characteristics, laboratory tests, umbilical cord blood unit characteristics and biomarkers. Through a three-level fusion strategy of feature level, knowledge level and decision level, it maximizes the mining of complementary information from different data sources, significantly improving the accuracy and robustness of the prediction model. Compared with traditional single-modal assessment methods, this application has achieved a significant improvement in prediction accuracy.

[0145] 3. This application adopts a refined feature engineering strategy. Through the construction of a preliminary screening index library under the guidance of expert knowledge, combined with feature selection methods such as LASSO regression and principal component analysis, the problem of feature redundancy and overfitting in high-dimensional data processing is effectively solved. Finally, a simplified feature set with high predictive value is selected, which not only ensures model performance but also reduces the burden of clinical data collection.

[0146] 4. This application constructs a prediction model based on gradient boosting decision trees. This algorithm can effectively handle mixed-type features, capture complex nonlinear relationships between features, and has strong robustness to data containing missing values. Through parameter optimization and early stopping techniques, it avoids overfitting while ensuring high accuracy and has good generalization ability.

[0147] 5. This application emphasizes the interpretability of the prediction results. It identifies and quantifies the contribution of key risk factors through the SHAP value analysis framework, providing clinicians with a transparent explanation of the prediction results. The risk assessment report not only displays the overall risk score but also details the specific factors affecting patient prognosis and their clinical interpretations, providing a clear direction for targeted interventions.

[0148] 6. This application provides personalized clinical decision support, offering differentiated treatment recommendations based on the patient's risk level and specific combination of risk factors. These recommendations cover multiple aspects, including pretreatment options, GVHD prevention strategies, and post-transplant monitoring protocols, which help optimize treatment strategies and improve patient outcomes.

[0149] 7. This application has a sound model update mechanism. Through regular updates and triggered retraining, it ensures that the system can adapt to changes in data distribution and medical advancements, and maintain long-term predictive accuracy. In addition, an external validation mechanism has been established to regularly evaluate the model’s performance in different centers, further improving the applicability and reliability of the system.

[0150] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0151] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for prognostic risk assessment of umbilical cord blood transplantation for hematological malignancies based on multimodal data fusion, characterized in that, The method includes: Acquire multimodal data including patient clinical characteristics data, laboratory test data, umbilical cord blood unit characteristic data, and biomarker data; A simplified feature set is selected from the preprocessed multimodal data based on feature selection and dimensionality reduction algorithms; The simplified features are fused based on three complementary strategies: direct feature-level concatenation, semantic association of knowledge graph, and multi-model decision integration. A prediction model is constructed based on the gradient boosting decision tree algorithm, and a standardized risk assessment report is generated based on the optimized hierarchical risk level to assist clinical decision support. A standardized risk assessment report is generated based on the stratified risk levels to support clinical decision-making.

2. The method as described in claim 1, characterized in that, The preprocessing includes missing value imputation algorithm, Z-score normalization, and outlier detection based on interquartile range.

3. The method as described in claim 1, characterized in that, The feature selection and dimensionality reduction algorithms include LASSO regression algorithm and principal component analysis.

4. The method as described in claim 1, characterized in that, The fusion of the simplified features based on three complementary strategies—feature-level direct cascading, knowledge graph semantic association, and multi-model decision integration—includes the following: Based on the feature-level direct concatenation strategy, feature vectors of different modalities in the simplified feature set are merged into a single high-dimensional feature vector. A semantic network containing disease-gene-drug relationships is constructed based on a knowledge graph semantic association strategy. In this semantic network, nodes represent specific entities, and edges represent the interaction relationships between entities. After mapping patient multimodal data to corresponding nodes, a multi-layer graph convolutional graph neural network is used to learn the semantic propagation path between nodes, generating a feature representation that integrates the semantic information of structured and unstructured data. The multi-model decision ensemble strategy constructs independent expert models for each of the multimodal data: a random forest model for clinical characterization data, a temporal convolutional network for laboratory test data, a logistic regression model for umbilical cord blood unit characteristic data, and a support vector machine model for biomarker data. The three complementary fusion strategies work together through a hierarchical integration architecture: feature-level direct cascading integration of structured features, knowledge graph mining of cross-modal semantic relationships, and multi-model decision integration to integrate prediction results from each layer.

5. The method as described in claim 1, characterized in that, The optimization of the prediction model includes: Construct an initial prediction model and calculate the residuals; A new decision tree is trained based on residuals; The prediction result of the new decision tree is multiplied by the learning rate and then added to the current model's prediction value; Repeatedly calculate the residuals and train a new decision tree until the termination condition is met.

6. The method as described in claim 5, characterized in that, The optimization of the prediction model also includes: The parameters of the prediction model are optimized by combining grid search with cross-validation, and early stopping is used to prevent the prediction model from overfitting the noise and random fluctuations in the training data.

7. The method as described in claim 1, characterized in that, The method further includes: The prediction model is updated using two methods: regular updates and triggered retraining. The routine update is as follows: new case data is automatically included at regular intervals, and the model parameters are fine-tuned using incremental learning methods. The triggered retraining is automatically initiated when the C-index on the validation set drops beyond a preset threshold or the deviation of the calibration curve exceeds a preset threshold, and the complete feature selection and model training process is re-executed.

8. A prognostic risk assessment system for umbilical cord blood transplantation for hematological malignancies based on multimodal data fusion, characterized in that, The system includes: The data acquisition and processing module is configured to acquire multimodal data, including patient clinical characterization data, laboratory test data, umbilical cord blood unit characteristic data, and biomarker data. The feature extraction and dimensionality reduction module is configured to filter a simplified feature set from the preprocessed multimodal data based on feature selection and dimensionality reduction algorithms; The multimodal fusion module is configured to fuse the simplified features based on three complementary strategies: direct feature-level concatenation, knowledge graph semantic association, and multi-model decision integration. The model training and risk assessment module is configured to analyze the fused simplified features based on a pre-built prediction model using a gradient boosting decision tree algorithm to obtain hierarchical risk levels. The decision support module is configured to generate standardized risk assessment reports based on the stratified risk levels to assist in clinical decision support.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method for prognostic risk assessment of umbilical cord blood transplantation for hematological malignancies based on multimodal data fusion as described in any one of claims 1-7.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the method for prognostic risk assessment of umbilical cord blood transplantation for hematological malignancies based on multimodal data fusion as described in any one of claims 1-7.

Citation Information

Cited By

  • Experimental test method and system for evaluating TFF3 liver protection based on NF-kappa B pathway inhibition

    CN121459956A

  • Complication risk prediction system and device based on agent workflow series scheduling

    CN121460192A

  • Urinary tract disease prediction system based on multimodal chromosome abnormality and clinical data

    CN121483568A