Method and device for training and applying medulloblastoma survival rate prediction model and medium

By training multiple initial models and selecting the optimal model, combined with detailed patient information of domestic MB patients, the problem of low prediction accuracy of medulloblastoma survival rate in the prior art is solved, and a higher-precision survival rate prediction is achieved.

CN120199469AActive Publication Date: 2025-06-24BEIJING TIANTAN HOSPITAL AFFILIATED TO CAPITAL MEDICAL UNIV +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510667855.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-06-24
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

The prior art predicts the survival rate of medulloblastoma patients with poor prediction accuracy, and the limitations of multivariable Cox regression models cannot meet more complex clinical needs.

Method used

By obtaining detailed patient information of domestic MB patients, including molecular subtype, gender, age, metastasis status, surgical resection status, histological subtype, treatment plan, radiotherapy dosage, etc., multiple initial models (such as Cox proportional hazard regression model, random survival forest model, extreme gradient enhancement model, etc.), and selecting the optimal model as the prediction model for survival of medulloblastoma through model performance comparison.

Benefits of technology

It significantly improves the accuracy of predicting survival rates of medulloblastoma, can consider a variety of clinical and molecular characteristics more accurately, and provides more reliable prediction results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199469A_ABST
    Figure CN120199469A_ABST
Patent Text Reader

Abstract

The invention discloses a medulloblastoma survival rate prediction model training and application method and device and a medium, and relates to the technical field of survival rate prediction.The method comprises the steps that a first training data set is obtained, the first training data set serves as input, multiple initial models are trained, and the first training data set is obtained; a plurality of first trained models and the model performance of each first trained model are obtained, the first trained model with the best model performance is selected as a medulloblastoma survival rate prediction model, and the medulloblastoma survival rate prediction model is used for predicting the survival rate of the MB patient at the prediction time point. The prediction precision can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of survival rate prediction, and particularly to a method, device and medium for training and applying a medulloblastoma survival rate prediction model. Background Art

[0002] Medulloblastoma (MB) is a common tumor in the central nervous system of children, most common in children aged 5 - 10 years old, with high malignancy, poor prognosis and high fatality rate. Although there have been advances in the treatment management of MB, the increased chances of endocrine - related dysfunction, neurocognitive deficits and the development of secondary tumor lesions will significantly affect the quality of life of MB patients. Therefore, prognosis prediction plays a crucial role in optimizing the treatment strategy for MB patients. Most scholars use Cox regression (i.e., Cox proportional hazards regression model) to establish a nomogram model based on clinical information, radiomic data or genomic data to predict the specific survival rate of MB patients. However, multivariate Cox regression is a semi - parametric model, which assumes that the death risk of MB patients is a linear combination of its covariates, and has certain limitations. Due to its limitations, some scholars consider introducing machine learning methods to predict the specific survival rate of MB patients based on clinical information, radiomic data or genomic data, but there are still problems with poor prediction accuracy. Summary of the Invention

[0003] The purpose of this application is to provide a method, device and medium for training and applying a medulloblastoma survival rate prediction model, which can improve the prediction accuracy.

[0004] To achieve the above purpose, the following solutions are provided in this application.

[0005] In the first aspect, this application provides a method for training a medulloblastoma survival rate prediction model. The method for training a medulloblastoma survival rate prediction model includes: Obtain a first training data set; the first training data set is a data set obtained by collecting patient information of domestic MB patients. The first training data set includes multiple first samples, and each first sample includes first feature data and label data. The first feature data includes the molecular subtype, gender, age, metastasis status, surgical resection situation, histological subtype, treatment plan, whole - brain and whole - spinal cord radiotherapy dose, and posterior fossa whole or tumor bed local boost radiotherapy dose of MB patients. The label data includes the survival status of MB patients; all MB patients have undergone surgical resection and received radiotherapy and / or chemotherapy after surgery; Use the first training data set as input to train multiple initial models, and obtain multiple first - trained models and the model performance of each first - trained model; the initial model is a model that can complete the prediction function; Select one of the first trained models with the best model performance as the medulloblastoma survival rate prediction model; the medulloblastoma survival rate prediction model is used to predict the survival rate of MB patients at the prediction time point.

[0006] In a second aspect, the present application provides a method for training a medulloblastoma survival rate prediction model, and the method for training a medulloblastoma survival rate prediction model includes: Obtain a second training data set; the second training data set is a data set obtained by collecting patient information of domestic MB patients, the second training data set includes a plurality of second samples, the second samples include second feature data and label data, and the second feature data includes the molecular subtype, gender, age, metastasis status, histological subtype, whole brain and whole spinal cord radiotherapy dose, posterior fossa whole or tumor bed local boost radiotherapy dose, high or low MYC expression level, high or low MYCN expression level, high or low OTX2 expression level, and high or low GFI1 expression level of MB patients, and the label data includes the survival status of MB patients; all MB patients have undergone surgical resection and received radiotherapy and / or chemotherapy after surgery; Using the second training data set as input, train a plurality of initial models to obtain a plurality of second trained models and the model performance of each second trained model; the initial model is a model capable of completing a prediction function; Select one of the second trained models with the best model performance as the medulloblastoma survival rate prediction model; the medulloblastoma survival rate prediction model is used to predict the survival rate of MB patients at the prediction time point.

[0007] In a third aspect, the present application provides a method for applying a medulloblastoma survival rate prediction model, and the method for applying a medulloblastoma survival rate prediction model includes: Obtain the patient information of the MB patient to be predicted; the patient information includes the molecular subtype, gender, age, metastasis status, surgical resection situation, histological subtype, treatment plan, whole brain and whole spinal cord radiotherapy dose, and posterior fossa whole or tumor bed local boost radiotherapy dose of the MB patient to be predicted, or the patient information includes the molecular subtype, gender, age, metastasis status, histological subtype, whole brain and whole spinal cord radiotherapy dose, posterior fossa whole or tumor bed local boost radiotherapy dose, high or low MYC expression level, high or low MYCN expression level, high or low OTX2 expression level, and high or low GFI1 expression level of the MB patient to be predicted; Using the patient information as input, use the medulloblastoma survival rate prediction model to predict the survival rate of the MB patient to be predicted at the prediction time point to obtain the predicted survival rate of the MB patient to be predicted; the medulloblastoma survival rate prediction model is a medulloblastoma survival rate prediction model trained by using the above-mentioned method for training a medulloblastoma survival rate prediction model.

[0008] In a fourth aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the above-mentioned medulloblastoma survival rate prediction model training method or the above-mentioned medulloblastoma survival rate prediction model application method.

[0009] In a fifth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the above-mentioned medulloblastoma survival rate prediction model training method or the above-mentioned medulloblastoma survival rate prediction model application method.

[0010] According to the specific embodiments provided by the present application, the present application has the following technical effects.

[0011] The present application provides a method, device, and medium for training and applying a medulloblastoma survival rate prediction model. When establishing a medulloblastoma survival rate prediction model through training, the first feature data used includes the molecular subtype, gender, age, metastasis status, surgical resection situation, histological subtype, treatment plan, whole brain and whole spinal cord radiotherapy dose, and posterior fossa whole or tumor bed local boost radiotherapy dose of MB patients. By selecting more complete patient information as the first feature data, the resulting medulloblastoma survival rate prediction model can consider more information, and the prediction accuracy is significantly improved. In addition, multiple initial models are trained to obtain multiple first trained models, and the model performances of the multiple first trained models are compared, and the first trained model with the best model performance is selected as the medulloblastoma survival rate prediction model. At this time, the resulting medulloblastoma survival rate prediction model has better model performance and significantly improved prediction accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0013] Figure 1 It is an application environment diagram of a method for training and applying a medulloblastoma survival rate prediction model provided by the present application.

[0014] Figure 2 It is a flowchart of a method for training a medulloblastoma survival rate prediction model provided by the present application.

[0015] Figure 3Schematic diagram of performance comparison of ROC curves (Receiver Operating Characteristic curve) of six models for 5-year survival rate prediction provided by this application.

[0016] Figure 4 Schematic diagram of performance comparison of ROC curves of six models for 10-year survival rate prediction provided by this application.

[0017] Figure 5 Schematic diagram of performance comparison of clinical decision curves (Decision Curve Analysis, DCA) of six models for 5-year survival rate prediction provided by this application.

[0018] Figure 6 Schematic diagram of performance comparison of clinical decision curves of six models for 10-year survival rate prediction provided by this application.

[0019] Figure 7 Schematic diagram of performance comparison of calibration curves between the XGBoost (Extreme Gradient Boosting) model and the Coxph model (Cox Proportional Hazards Regression Model, that is, the Cox proportional hazards regression model) for 5-year survival rate prediction provided by this application.

[0020] Figure 8 Schematic diagram of performance comparison of calibration curves between the XGBoost model and the Coxph model for 10-year survival rate prediction provided by this application.

[0021] Figure 9 Bar chart of variable importance of SHAP (SHapley Additive exPlanation) of the XGBoost model provided by this application.

[0022] Figure 10 Summary diagram of SHAP of the XGBoost model provided by this application.

[0023] Figure 11 Schematic diagram of the external validation ROC curve of the XGBoost model provided by this application.

[0024] Figure 12 Schematic diagram of screening of molecular characteristics provided by this application.

[0025] Figure 13 Schematic diagram of the process of a method for training a survival rate prediction model for medulloblastoma provided by this application.

[0026] Figure 14 Schematic diagram for comparing the ROC curve performance of six models for predicting the 5-year survival rate of patients with medulloblastoma of Group_3 subtype (Group 3 subtype) or Group_4 subtype (Group 4 subtype) provided by this application.

[0027] Figure 15 Schematic diagram for comparing the ROC curve performance of six models for predicting the 10-year survival rate of patients with medulloblastoma of Group_3 subtype or Group_4 subtype provided by this application.

[0028] Figure 16 Schematic flow diagram of a method for applying a medulloblastoma survival rate prediction model provided by this application.

[0029] Figure 17 Schematic diagram of the structure of a computer device provided by this application. Detailed implementation manners

[0030] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.

[0031] Embodiment 1.

[0032] The medulloblastoma survival rate prediction model training method provided in the embodiments of this application can be applied to, for example Figure 1In the application environment shown. Among them, the terminal communicates with the server through the network. The data storage system can store the data that the server needs to process. The data storage system can be set up separately, integrated on the server, or placed on the cloud or other servers. The terminal can send a to-be-processed establishment request to the server. After receiving the to-be-processed establishment request, for the to-be-processed establishment request, the server obtains a first training data set; using the first training data set as input, trains multiple initial models to obtain multiple first-trained models and the model performance of each first-trained model; selects one first-trained model with the best model performance as the medulloblastoma survival rate prediction model, and the medulloblastoma survival rate prediction model is used to predict the survival rate of MB patients at the prediction time point. At this time, a prediction model is established based on clinical data, or the server obtains a second training data set; using the second training data set as input, trains multiple initial models to obtain multiple second-trained models and the model performance of each second-trained model; selects one second-trained model with the best model performance as the medulloblastoma survival rate prediction model, and the medulloblastoma survival rate prediction model is used to predict the survival rate of MB patients at the prediction time point. At this time, a prediction model is established based on clinical data and molecular data. The server can feedback the established result, which is the medulloblastoma survival rate prediction model for the establishment request, to the terminal.

[0033] In addition, in some embodiments, the method for training the medulloblastoma survival rate prediction model can also be implemented separately by the server or the terminal. For example, the terminal can directly process the to-be-processed establishment request, or the server can obtain the to-be-processed establishment request from the data storage system and process the to-be-processed establishment request.

[0034] Among them, the terminal can be, but is not limited to, various desktop computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server can be implemented by an independent server or a server cluster composed of multiple servers, and can also be a cloud server.

[0035] In an exemplary embodiment, as Figure 2 shown, a method for training a medulloblastoma survival rate prediction model is provided. This method is executed by a computer device, and can be specifically executed separately by a computer device such as a terminal or a server, or jointly executed by the terminal and the server. In the embodiments of the present application, taking this method applied to the Figure 1 server in as an example for description, the method includes the following steps.

[0036] Step S101: Obtain the first training dataset. The first training dataset is a dataset obtained by collecting patient information of domestic MB patients. The first training dataset includes multiple first samples. Each first sample includes first feature data and label data. The first feature data includes the molecular subtype, gender, age, metastasis status, surgical resection situation, histological subtype, treatment plan, whole brain and whole spinal cord radiotherapy dose, and whole posterior fossa or local tumor bed boost radiotherapy dose of MB patients. The label data includes the survival status of MB patients. All MB patients have undergone surgical resection and received radiotherapy and / or chemotherapy after surgery.

[0037] Step S102: Use the first training dataset as input to train multiple initial models to obtain multiple first-trained models and the model performance of each first-trained model. The initial model is a model capable of completing the prediction function.

[0038] Step S103: Select one first-trained model with the best model performance as the medulloblastoma survival rate prediction model. The medulloblastoma survival rate prediction model is used to predict the survival rate of MB patients at the prediction time point.

[0039] By implementing the above steps S101 to S103, when establishing the medulloblastoma survival rate prediction model through training in this embodiment, the first feature data used includes the molecular subtype, gender, age, metastasis status, surgical resection situation, histological subtype, treatment plan, whole brain and whole spinal cord radiotherapy dose, and whole posterior fossa or local tumor bed boost radiotherapy dose of MB patients. By selecting more complete patient information as the first feature data, the obtained medulloblastoma survival rate prediction model can consider more information, and the prediction accuracy is significantly improved. In addition, multiple initial models are trained to obtain multiple first-trained models, and through model performance comparison, one first-trained model with the best model performance is selected as the medulloblastoma survival rate prediction model. At this time, the model performance of the obtained medulloblastoma survival rate prediction model is better, and the prediction accuracy is significantly improved.

[0040] After obtaining the medulloblastoma survival rate prediction model, the medulloblastoma survival rate prediction model training method of this embodiment further includes: obtaining a first external data set, which is a data set obtained by collecting patient information of global MB patients. The first external data set includes multiple first samples, and each first sample includes first feature data and label data. The first feature data includes the molecular subtype, gender, age, metastasis status, surgical resection situation, histological subtype, treatment plan, whole brain and spinal cord radiotherapy dose, and posterior fossa whole or tumor bed local boost radiotherapy dose of the MB patient, and the label data includes the survival status of the MB patient; using the first training data set as input, performing internal validation on the medulloblastoma survival rate prediction model to obtain an internal validation result and test the model stability; using the first external data set as input, performing external validation on the medulloblastoma survival rate prediction model to obtain an external validation result and test the model extrapolation ability.

[0041] The medulloblastoma survival rate prediction model training method of this embodiment specifically includes the following steps.

[0042] (1) Obtain the first training data set and the first external data set.

[0043] This embodiment collects domestic clinical data (which can also be called clinicopathological data), specifically collecting a multi-institutional cohort of 1043 MB patients diagnosed between September 2001 and April 2023 in China. Among them, 729 MB patients have complete clinical data, and the clinical data includes molecular subtype, gender, age, metastasis status, surgical resection situation, histological subtype, treatment plan, whole brain and spinal cord radiotherapy dose, posterior fossa whole or tumor bed local boost radiotherapy dose, surgery date, follow-up cut-off date, and survival status. The molecular subtypes include WNT, SHH, Gr.3 (i.e., Group_3), Gr.4 (i.e., Group_4), and unknown molecular subtype. The gender includes male and female. The age includes infants (generally 0 - 3 years old), children (generally 3 - 10 years old), adolescents (generally 10 - 17 years old), and adults (generally 18 years old and above). The metastasis status includes non-metastasis and metastasis. The surgical resection situation includes partial resection and subtotal / total resection. The histological subtypes include classical type, desmoplastic / nodular type, extensive nodular type, large cell / anaplastic type, and unknown histological subtype. The treatment plan includes radiotherapy only after surgery and concurrent chemoradiotherapy after surgery (i.e., radiotherapy and chemotherapy are performed simultaneously after surgery). The unit of the whole brain and spinal cord radiotherapy dose is Gy, and the unit of the posterior fossa whole or tumor bed local boost radiotherapy dose is Gy. The survival duration of the MB patient is defined as the time from the surgery time to the death time / follow-up cut-off time (August 31, 2023). Subsequently, the first training data set is constructed based on the above domestic clinical data.

[0044] This embodiment collects clinical data from all over the world. Specifically, a multi-center cohort of 116 MB patients from 23 institutions worldwide is collected. The clinical data includes molecular subtype, gender, age, metastasis status, surgical resection status, histological subtype, treatment regimen, whole brain and spinal cord radiotherapy dose, posterior fossa whole or tumor bed local boost radiotherapy dose, survival duration, and survival status. Subsequently, the above-mentioned global clinical data is used as external validation data, and a first external dataset is constructed based on the above-mentioned global clinical data.

[0045] It should be noted that all the above MB patients underwent surgical resection and received radiotherapy and / or chemotherapy after surgery. MB patients who only underwent biopsy or did not receive radiotherapy are not included. The median age at diagnosis was 8 years, and the IQR (Interquartile Range) was 6 to 11 years.

[0046] During training, this embodiment first selects data for model training. Specifically, the first feature data is designed to include molecular subtype, gender, age, metastasis status, surgical resection status, histological subtype, treatment regimen, whole brain and spinal cord radiotherapy dose, and posterior fossa whole or tumor bed local boost radiotherapy dose. Therefore, this embodiment can design a first training dataset and a first external dataset. The first training dataset is a dataset obtained by collecting patient information of domestic MB patients, and the first external dataset is a dataset obtained by collecting patient information of global MB patients. Both the first training dataset and the first external dataset include multiple first samples. The first sample includes first feature data and label data. The first feature data includes the molecular subtype, gender, age, metastasis status, surgical resection status, histological subtype, treatment regimen, whole brain and spinal cord radiotherapy dose, and posterior fossa whole or tumor bed local boost radiotherapy dose of MB patients. The label data includes the survival status of MB patients. Based on the survival status, it is determined whether an MB patient survives at the prediction time point. All MB patients underwent surgical resection and received radiotherapy and / or chemotherapy after surgery.

[0047] (2) Using the first training dataset as input, multiple initial models are trained to obtain multiple first trained models and the model performance of each first trained model.

[0048] This embodiment performs data cleaning on the first training dataset and the first external dataset. Data cleaning includes: deleting duplicate data, reviewing abnormal data (i.e., outliers) to remove logical errors, and processing the units of numerical data to maintain data consistency. Finally, the data in the first cleaned training dataset and the first cleaned external dataset are high-quality, consistent, non-missing, and non-abnormal data after cleaning, processing, and formatting, and can be directly used for model training and validation.

[0049] At this time, in this embodiment, using the first training dataset as the input, training multiple initial models specifically includes: cleaning the data of the first training dataset to obtain the first cleaned training dataset; using the first cleaned training dataset as the input to train multiple initial models. Data cleaning includes: deleting duplicate data, rechecking abnormal data, and unifying the units of numerical data of the same type.

[0050] In this embodiment, the initial model is designed as a model capable of completing the prediction function. As an example, the initial models include: Cox proportional hazards regression model (abbreviated as Coxph), random survival forests model (Random Survival Forests, abbreviated as RSF), extreme gradient boosting model (abbreviated as XGBoost), elastic net model (Elastic Net Regularization, abbreviated as ENET), DeepSurv model (Deep Survival Model), and gradient boosting machine model (Gradient Boosting Machine, abbreviated as GBM). Of course, other models capable of completing the prediction function can be used as the initial model, and this embodiment does not impose any restrictions on this.

[0051] During training, in this embodiment, the first training dataset is used to train each initial model respectively. During the training process, the 5-fold cross-validation method is combined with the grid search algorithm for hyperparameter tuning, and the finally obtained optimal hyperparameters are used to retrain the initial model on the first training dataset.

[0052] At this time, in this embodiment, using the first training dataset as the input, training multiple initial models to obtain multiple first-trained models and the model performance of each first-trained model specifically includes: for each initial model, using the first training dataset as the input, using the 5-fold cross-validation method and the grid search algorithm to perform hyperparameter tuning on the initial model to obtain the optimal hyperparameters of the initial model; using the first training dataset to train the initial model with the optimal hyperparameters to obtain the first-trained model corresponding to the initial model and the model performance of each first-trained model.

[0053] Multiple performance evaluation indicators are used to evaluate the model performance, including the area under the receiver operating characteristic curve (Area Under the Curve, AUC), decision curve analysis, and calibration curve analysis. The model performance is compared through the performance evaluation indicators, and the best prediction model (i.e., the model with the best model performance) is selected as the medulloblastoma survival rate prediction model. At this time, in this embodiment, the model performance is characterized by the performance evaluation indicators, and the performance evaluation indicators include the area under the receiver operating characteristic curve, decision curve analysis, and calibration curve analysis.

[0054] (3) Select the first trained model with the best model performance as the medulloblastoma survival rate prediction model, which is used to predict the survival rate of MB patients at the prediction time point.

[0055] After experiments, the medulloblastoma survival rate prediction model is the XGBoost model. When tuning the hyperparameters of the XGBoost model, the parameter range of grid search is as follows: the range of the number of iterations (nrounds) is (200, 300), the range of the maximum depth (max_depth) is (2, 6), and the range of the learning rate (eta) is (0.01, 0.3). The obtained optimal hyperparameters are: the number of iterations = 207, the maximum depth = 2, and the learning rate = 0.08004. After obtaining the medulloblastoma survival rate prediction model, an online survival rate prediction calculator based solely on clinical information (i.e., clinical data) is constructed based on the medulloblastoma survival rate prediction model and deployed on the software interface, thereby providing a method for predicting the survival rate of medulloblastoma based on multimodal data and the XGBoost model. When predicting, the inputs of the medulloblastoma survival rate prediction model include molecular subtype, gender, age, metastasis status, surgical resection situation, histological subtype, treatment plan, whole brain and whole spinal cord radiotherapy dose, and posterior fossa whole or tumor bed local boost radiotherapy dose, and the outputs include the predicted survival rates at prediction time points such as 5 years, 10 years, and 15 years. At the same time, the online survival rate prediction calculator can also perform risk stratification judgment. The standard risk is defined as: age > 3 years, no metastasis, and total resection / near total resection. This situation is the low-risk population, and the rest are high-risk populations, as Figure 3 、 Figure 4 、 Figure 5 、 Figure 6 、 Figure 7 、 Figure 8 shown. Figure 3 In, CI refers to Confidence Interval, representing the confidence interval. Figure 5 In, the All line refers to the clinical decision curve obtained by assuming that all patients are given active treatment interventions without considering their individual predicted risks, and the None line refers to the clinical decision curve obtained by not giving any additional active treatment interventions to any patients.

[0056] It should be noted that the medulloblastoma survival rate prediction models corresponding to different prediction time points may be different. For each prediction time point, the medulloblastoma survival rate prediction model is determined according to the training method.

[0057] (4) Use the first training dataset as the input to perform internal validation on the medulloblastoma survival rate prediction model to obtain the internal validation result.

[0058] Since the first training dataset and the first external dataset have been pre - cleaned, at this time, in this embodiment, using the first training dataset as the input, internal validation of the medulloblastoma survival prediction model is carried out, specifically including: cleaning the data of the first training dataset to obtain the first cleaned training dataset; using the first cleaned training dataset as the input to carry out internal validation of the medulloblastoma survival prediction model. Data cleaning includes: deleting duplicate data, rechecking abnormal data, and unifying the units of numerical data of the same type.

[0059] In this embodiment, using the first training dataset as the input, the Bootstrap method (resampling 1000 times) is used for internal validation. By sampling with replacement in the model training queue, Bootstrap resampling samples of the same sample size are constructed to evaluate the model performance. This process is repeated 1000 times to obtain the stability of the model in internal validation, test the repeatability of the model development process, and prevent overfitting of the model from overestimating the model performance.

[0060] (5) Using the first external dataset as the input, external validation of the medulloblastoma survival prediction model is carried out to obtain the external validation result.

[0061] Using the first external dataset as the input, external validation of the medulloblastoma survival prediction model is carried out, specifically including: cleaning the data of the first external dataset to obtain the first cleaned external dataset; using the first cleaned external dataset as the input to carry out external validation of the medulloblastoma survival prediction model. Data cleaning includes: deleting duplicate data, rechecking abnormal data, and unifying the units of numerical data of the same type. By using the first external dataset as the input for external validation, the prediction performance of the best prediction model on external data is tested, and specifically, external validation is carried out through this first external dataset to verify the extrapolability of the model.

[0062] This embodiment can also calculate the SHAP values of all features of all first samples, analyze the importance ranking of each feature according to the SHAP values, and how the features affect the prediction results of the medulloblastoma survival prediction model, so as to interpret the medulloblastoma survival prediction model.

[0063] At this time, in this embodiment, after obtaining the medulloblastoma survival rate prediction model, the medulloblastoma survival rate prediction model training method of this embodiment further includes: using the first training data set as the input, calculating the SHAP value of each feature in each first sample by using the SHapley Additive exPlanation method, and determining the importance of each feature and the influence on the prediction result of the medulloblastoma survival rate prediction model based on the SHAP value of each feature in each first sample, so as to interpret the medulloblastoma survival rate prediction model, where the feature is a kind of data in the first feature data. The SHAP variable importance bar chart and SHAP summary chart of the XGBoost model are as Figure 9 and Figure 10 shown, and the prediction performance of the XGBoost model applied to external validation is as Figure 11 shown.

[0064] Through the above process, a prediction model can be established through clinical data. In this embodiment, genomic data (i.e., molecular data) is further introduced. Since it is difficult to distinguish between Group_3 and Group_4 tumors, Group_3 / Group_4 patients are combined to establish a prediction model that combines clinical data and molecular data. The gene expression profile of the medulloblastoma tissue sample is measured through the BGISEQ-50 platform, and the readings of the gene expression profile are aligned with the hg38 human reference genome, and then the gene expression profile is standardized to obtain a tpm (Transcripts Per Million) matrix. Then, the tpm matrix is logarithmically transformed (log2(tpm + 1)) to obtain a transformed matrix. Based on the transformed matrix, the expression level of each gene is determined to obtain molecular data. At this time, when the molecular subtype is the Group_3 / Group_4 subtype, the variable screening process is as follows: To improve the accuracy and prediction efficiency of the prediction model, fully consider the molecular risk stratification conditions of the Group_3 / Group_4 subtype. Based on previous studies, key molecular events are included: GLI2 activation, MYC activation, MYCN activation, OTX2 activation, SNCAIP activation, PRDM6 activation, GFI1B activation, GFI1 activation, CDK6 activation. Artificial feature selection is performed according to clinical expertise and previous studies, and univariate Cox analysis is included to select candidate molecular features. Specifically, molecular features with P (used to measure whether the association between a certain molecular feature and the survival outcome is statistically significant) < 0.1 are selected, such as Figure 12As shown, since MYC, MYCN, and OTX2 are common driver genes in medulloblastoma, the second feature data finally included are: molecular subtype, gender, age, metastasis status, histological subtype, whole brain and spinal cord radiotherapy dose, posterior fossa whole or tumor bed local boost radiotherapy dose, and the high or low expression levels of MYC, MYCN, OTX2, and GFI1 in MB patients (specifically, the high or low expression levels are divided using the median. Greater than or equal to the median is high expression, and less than the median is low expression).

[0065] At this time, when constructing a survival prediction model for medulloblastoma of the Group_3 / Group_4 subtype (194 cases), the second feature data included are: molecular subtype, gender, age, metastasis status, histological subtype, whole brain and spinal cord radiotherapy dose, posterior fossa whole or tumor bed local boost radiotherapy dose, high or low expression level of MYC (median = 1.66216), high or low expression level of MYCN (median = 2.885574), high or low expression level of OTX2 (median = 7.725381), and high or low expression level of GFI1 (median = 0.09083753). Except for the differences in the second feature data, other training processes, internal validation processes, and external validation processes are the same as the above training method based on clinical data, and a prediction model based on clinical data and molecular data can be obtained.

[0066] As Figure 13 shown, the training method for the medulloblastoma survival prediction model in this embodiment includes the following steps.

[0067] Step S201, obtain a second training dataset; the second training dataset is a dataset obtained by collecting the patient information of domestic MB patients. The second training dataset includes multiple second samples, and each second sample includes second feature data and label data. The second feature data includes the molecular subtype, gender, age, metastasis status, histological subtype, whole brain and spinal cord radiotherapy dose, posterior fossa whole or tumor bed local boost radiotherapy dose, high or low expression level of MYC, high or low expression level of MYCN, high or low expression level of OTX2, and high or low expression level of GFI1 in MB patients, and the label data includes the survival status of MB patients; all MB patients have undergone surgical resection and received radiotherapy and / or chemotherapy after surgery.

[0068] Step S202, use the second training dataset as input to train multiple initial models to obtain multiple second trained models and the model performance of each second trained model; the initial model is a model capable of completing the prediction function.

[0069] Step S203: Select one of the second trained models with the best model performance as the medulloblastoma survival rate prediction model; the medulloblastoma survival rate prediction model is used to predict the survival rate of MB patients at the prediction time point.

[0070] By implementing the above steps S201 to S203, when establishing the medulloblastoma survival rate prediction model through training in this embodiment, the second feature data used includes the molecular subtype, gender, age, metastasis status, histological subtype, whole brain and whole spinal cord radiotherapy dose, posterior fossa whole or tumor bed local boost radiotherapy dose, high or low MYC expression level, high or low MYCN expression level, high or low OTX2 expression level, and high or low GFI1 expression level of MB patients. At this time, the obtained medulloblastoma survival rate prediction model can consider more information, fully consider clinical data and molecular data, the prediction accuracy is significantly improved, and multiple initial models are trained to obtain multiple second trained models. By comparing the model performances, select one of the second trained models with the best model performance as the medulloblastoma survival rate prediction model. At this time, the obtained medulloblastoma survival rate prediction model has better model performance and the prediction accuracy is significantly improved.

[0071] The training method of the medulloblastoma survival rate prediction model in this embodiment specifically includes the following steps.

[0072] (1) Obtain the second training dataset and the second external dataset.

[0073] On the basis of the first training dataset, further add molecular data (that is, high or low MYC expression level, high or low MYCN expression level, high or low OTX2 expression level, and high or low GFI1 expression level), then the second training dataset can be obtained. On the basis of the first external dataset, further add molecular data, then the second external dataset can be obtained.

[0074] (2) Use the second training dataset as the input to train multiple initial models to obtain multiple second trained models and the model performance of each second trained model.

[0075] Using the second training dataset as the input to train multiple initial models specifically includes: cleaning the data of the second training dataset to obtain the second cleaned training dataset; using the second cleaned training dataset as the input to train multiple initial models. Data cleaning includes: deleting duplicate data, rechecking abnormal data, and unifying the units of numerical data of the same type.

[0076] Using the second training dataset as input, train multiple initial models to obtain multiple post-training models and the model performance of each post-training model. Specifically, for each initial model, use the second training dataset as input, and use the 5-fold cross-validation method and grid search algorithm to optimize the hyperparameters of the initial model to obtain the optimal hyperparameters of the initial model; use the second training dataset to train the initial model with the optimal hyperparameters to obtain the post-training model corresponding to the initial model and the model performance of each post-training model.

[0077] (3)Select one post-training model with the best model performance as the medulloblastoma survival rate prediction model, and the medulloblastoma survival rate prediction model is used to predict the survival rate of MB patients at the prediction time point.

[0078] (4)Using the second training dataset as input, perform internal validation on the medulloblastoma survival rate prediction model to obtain the internal validation result.

[0079] Using the second training dataset as input, perform internal validation on the medulloblastoma survival rate prediction model. Specifically, clean the data of the second training dataset to obtain the second cleaned training dataset; use the second cleaned training dataset as input to perform internal validation on the medulloblastoma survival rate prediction model. Data cleaning includes: deleting duplicate data, reviewing abnormal data, and unifying the units of numerical data of the same type.

[0080] (5)Using the second external dataset as input, perform external validation on the medulloblastoma survival rate prediction model to obtain the external validation result.

[0081] Using the second external dataset as input, perform external validation on the medulloblastoma survival rate prediction model. Specifically, clean the data of the second external dataset to obtain the second cleaned external dataset; use the second cleaned external dataset as input to perform external validation on the medulloblastoma survival rate prediction model. Data cleaning includes: deleting duplicate data, reviewing abnormal data, and unifying the units of numerical data of the same type.

[0082] Use the second training dataset for training and internal validation respectively, use the second external dataset for external validation, and use the receiver operating characteristic curve to compare the model performance, as Figure 14 and Figure 15 shown.

[0083] After obtaining the medulloblastoma survival rate prediction model, calculate the SHAP values of all features of all second samples, analyze the importance ranking of each feature according to the SHAP values, and how the features affect the prediction results to interpret the medulloblastoma survival rate prediction model.

[0084] After experiments, the survival prediction model for medulloblastoma of Group_3 / Group_4 subtype is the XGBoost model. When performing hyperparameter tuning, the parameter range of grid search is as follows: the range of the number of iterations (nrounds) is (20, 28), the range of the maximum depth (max_depth) is (2, 6), and the range of the learning rate (eta) is (0.01, 0.3). The obtained optimal hyperparameters are: the number of iterations = 21, the maximum depth = 4, and the learning rate = 0.189. An online survival prediction calculator integrating clinical information (i.e., clinical data) and molecular event information (i.e., molecular data) is constructed and deployed on the software interface. When making predictions, the inputs of the medulloblastoma survival prediction model include molecular subtype, gender, age, metastasis status, histological subtype, whole brain and whole spinal cord radiotherapy dose, posterior fossa whole or tumor bed local boost radiotherapy dose, high or low expression levels of MYC, high or low expression levels of MYCN, high or low expression levels of OTX2, and high or low expression levels of GFI1. The outputs include the predicted survival rates at prediction time points such as 5 years, 10 years, and 15 years, and further risk stratification judgment is carried out.

[0085] This embodiment proposes a method for predicting the survival rate of medulloblastoma patients based on multimodal data. The steps are as follows: clean the clinical data and molecular data, construct and train six initial models (including Cox proportional hazards regression model, random survival forest model, extreme gradient boosting model, elastic net model, DeepSurv model, and gradient boosting machine model) by using clinical data alone or integrating clinical data and molecular data, perform model training in the training dataset, use the 5-fold cross-validation method combined with the grid search algorithm for hyperparameter tuning, retrain the initial models on the training dataset with the final optimal hyperparameters, compare the model performances by using the area under the receiver operating characteristic curve, decision curve analysis, and calibration curve analysis, determine the final medulloblastoma survival prediction model, perform internal validation by using the Bootstrap method, perform external validation by using an external dataset of a multicenter cohort from 23 units around the world, rank the feature importance by using the SHapley Additive exPlanation method, and explain the final medulloblastoma survival prediction model. Use the final medulloblastoma survival prediction model to construct an online survival prediction calculator and deploy it on the software interface. Input the observed data of the individual patient to be analyzed into the online survival prediction calculator for survival prediction and obtain the prediction results. This embodiment can perform survival analysis and prediction by using complete patient information, with good performance and accurate prediction results.

[0086] The online survival rate prediction calculator is deployed in an open and free software interface. By selecting individual information and clicking the "predict" button, the 5-year and 10-year survival rates can be predicted. On the right interface, the risk grouping is first displayed, followed by the predicted 5-year and 10-year survival rates. At the bottom, a line chart showing the changes in the predicted 5-year and 10-year survival rates is displayed. Currently, there is no clinical data from research covering radiotherapy doses or integrating clinical data with molecular data to predict survival rates. Compared with existing survival prediction models, this method has higher prediction performance, and the specific advantages come from the following technical points.

[0087] (1) Multi-modal data fusion and high-precision prediction: This method introduces radiotherapy dose data not included in other models and also introduces molecular data, achieving deep fusion of multi-source data. By using machine learning algorithms, the accuracy and precision of survival rate prediction have been significantly improved. Through internal validation, the stability of the prediction model has been verified, and through external validation, the generalization of the model has been improved.

[0088] (2) Optimization of personalized treatment plans: By using the online version of the online survival rate prediction calculator based on the XGBoost model, clinicians can accurately predict patient survival and provide personalized treatment plans, which can help clinicians make precise treatment decisions and reduce patient risks.

[0089] Example 2.

[0090] The method for applying the medulloblastoma survival rate prediction model provided by the embodiment of this application can be applied to the application environment as Figure 1 shown. Among them, the terminal communicates with the server through the network. The data storage system can store the data that the server needs to process. The data storage system can be set up separately, integrated on the server, or placed on the cloud or other servers. The terminal can send the prediction request to be processed to the server. After receiving the prediction request to be processed, for the prediction request to be processed, the server obtains the patient information of the MB patient to be predicted; using the patient information as input, the medulloblastoma survival rate prediction model is used to predict the survival rate of the MB patient to be predicted at the prediction time point, and the predicted survival rate of the MB patient to be predicted is obtained. The server can feedback the prediction result, which is the predicted survival rate for the prediction request, to the terminal.

[0091] In addition, in some embodiments, the method for applying the medulloblastoma survival rate prediction model can also be implemented by the server or the terminal alone. For example, the terminal can directly process the prediction request to be processed, or the server can obtain the prediction request to be processed from the data storage system and process the prediction request to be processed.

[0092] In an exemplary embodiment, asFigure 16 As shown, a method for applying a medulloblastoma survival rate prediction model is provided. This method is executed by a computer device, specifically, it can be executed independently by a computer device such as a terminal or a server, or jointly executed by a terminal and a server. In the embodiments of the present application, taking the case where this method is applied to the Figure 1 server in as an example for illustration, it includes the following steps.

[0093] Step T1, obtain the patient information of the MB patient to be predicted; the patient information includes the molecular subtype, gender, age, metastasis status, surgical resection situation, histological subtype, treatment plan, whole brain and whole spinal cord radiotherapy dose, and whole posterior fossa or tumor bed local boost radiotherapy dose of the MB patient to be predicted. Alternatively, the patient information includes the molecular subtype, gender, age, metastasis status, histological subtype, whole brain and whole spinal cord radiotherapy dose, whole posterior fossa or tumor bed local boost radiotherapy dose, high or low MYC expression level, high or low MYCN expression level, high or low OTX2 expression level, and high or low GFI1 expression level of the MB patient to be predicted.

[0094] Step T2, using the patient information as input, utilize the medulloblastoma survival rate prediction model to predict the survival rate of the MB patient to be predicted at the prediction time point, and obtain the predicted survival rate of the MB patient to be predicted; the medulloblastoma survival rate prediction model is a medulloblastoma survival rate prediction model trained by using the medulloblastoma survival rate prediction model training method described in Embodiment 1.

[0095] When the patient information includes the molecular subtype, gender, age, metastasis status, surgical resection situation, histological subtype, treatment plan, whole brain and whole spinal cord radiotherapy dose, and whole posterior fossa or tumor bed local boost radiotherapy dose of the MB patient to be predicted, the medulloblastoma survival rate prediction model is a medulloblastoma survival rate prediction model trained by using clinical data. When the patient information includes the molecular subtype, gender, age, metastasis status, histological subtype, whole brain and whole spinal cord radiotherapy dose, whole posterior fossa or tumor bed local boost radiotherapy dose, high or low MYC expression level, high or low MYCN expression level, high or low OTX2 expression level, and high or low GFI1 expression level of the MB patient to be predicted, the medulloblastoma survival rate prediction model is a medulloblastoma survival rate prediction model trained by using clinical data and molecular data.

[0096] The present application also provides an application scenario, which applies the above method for applying a medulloblastoma survival rate prediction model. Specifically, the method for applying a medulloblastoma survival rate prediction model provided in this embodiment can be applied in a survival prediction scenario. The survival prediction scenario includes a prediction link and a display link. The prediction link is used to predict the survival rate of a medulloblastoma patient at a prediction time point to obtain a predicted survival rate, and the display link is used to display the predicted survival rate to the user. The method for applying a medulloblastoma survival rate prediction model provided in this embodiment belongs to the prediction link.

[0097] Embodiment 3.

[0098] In an exemplary embodiment, a computer device is provided. The computer device can be a server or a terminal, and its internal structure diagram can be as Figure 17 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a method for training a medulloblastoma survival rate prediction model or a method for applying a medulloblastoma survival rate prediction model.

[0099] Those skilled in the art can understand that Figure 17 the structure shown in

[0100] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0101] Embodiment 4.

[0102] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program which, when executed by a processor, implements the medulloblastoma survival rate prediction model training method in Embodiment 1 or the medulloblastoma survival rate prediction model application method in Embodiment 2.

[0103] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0104] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0105] Specific examples are used in this article to elaborate on the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.

Claims

1. A method for training a survival rate prediction model for medulloblastoma, characterized in that, The method for training the medulloblastoma survival rate prediction model includes: Obtain a first training data set; the first training data set is a data set obtained by collecting the patient information of domestic MB patients. The first training data set includes multiple first samples. The first sample includes first feature data and label data. The first feature data includes the molecular subtype, gender, age, metastasis status, surgical resection situation, histological subtype, treatment plan, whole brain and whole spinal cord radiotherapy dose, and posterior fossa whole or tumor bed local boost radiotherapy dose of MB patients. The label data includes the survival status of MB patients; all the MB patients have undergone surgical resection and received radiotherapy and / or chemotherapy after surgery; Using the first training data set as input, train multiple initial models to obtain multiple first trained models and the model performance of each first trained model; the initial model is a model capable of completing the prediction function; Select one first trained model with the best model performance as the medulloblastoma survival rate prediction model; the medulloblastoma survival rate prediction model is used to predict the survival rate of MB patients at the prediction time point.

2. The method for training a medulloblastoma survival rate prediction model according to claim 1, characterized in that After obtaining the medulloblastoma survival rate prediction model, the method for training the medulloblastoma survival rate prediction model further includes: Obtain a first external data set; the first external data set is a data set obtained by collecting the patient information of global MB patients. The first external data set includes multiple of the first samples; Using the first training data set as input, perform internal verification on the medulloblastoma survival rate prediction model to obtain an internal verification result and test the model stability; Using the first external data set as input, perform external verification on the medulloblastoma survival rate prediction model to obtain an external verification result and test the model extrapolation.

3. The method for training a medulloblastoma survival rate prediction model according to claim 2, wherein Using the first training data set as input to train multiple initial models specifically includes: cleaning the data of the first training data set to obtain a first cleaned training data set; using the first cleaned training data set as input to train multiple initial models; Using the first training data set as input to perform internal verification on the medulloblastoma survival rate prediction model specifically includes: cleaning the data of the first training data set to obtain a first cleaned training data set; using the first cleaned training data set as input to perform internal verification on the medulloblastoma survival rate prediction model; Using the first external data set as input to perform external verification on the medulloblastoma survival rate prediction model specifically includes: cleaning the data of the first external data set to obtain a first cleaned external data set; using the first cleaned external data set as input to perform external verification on the medulloblastoma survival rate prediction model; Among them, the data cleaning includes: deleting duplicate data, rechecking abnormal data, and unifying the units of numerical data of the same type.

4. The method for training a medulloblastoma survival rate prediction model according to claim 1, wherein, The initial models include: Cox proportional hazards regression model, random survival forest model, extreme gradient boosting model, elastic net model, DeepSurv model, and gradient boosting machine model; The performance of the models is characterized by performance evaluation metrics, and the performance evaluation metrics include area under the receiver operating characteristic curve, decision curve analysis, and calibration curve analysis.

5. The method for training a medulloblastoma survival rate prediction model according to claim 1, characterized in that, Using the first training dataset as input, training multiple initial models to obtain multiple first-trained models and the model performance of each first-trained model, specifically including: For each initial model, using the first training dataset as input, tuning the hyperparameters of the initial model using the 5-fold cross-validation method and grid search algorithm to obtain the optimal hyperparameters of the initial model; training the initial model with the optimal hyperparameters using the first training dataset to obtain the corresponding first-trained model of the initial model and the model performance of the first-trained model.

6. The method for training a medulloblastoma survival rate prediction model according to claim 1, wherein After obtaining the medulloblastoma survival rate prediction model, the medulloblastoma survival rate prediction model training method further includes: Using the SHapley Additive exPlanation method to calculate the SHAP value of each feature in each of the first samples with the first training dataset as input, and determining the importance of each feature and its impact on the prediction result of the medulloblastoma survival rate prediction model based on the SHAP value of each feature in each of the first samples, so as to interpret the medulloblastoma survival rate prediction model; wherein, the feature is a type of data in the first feature data.

7. A method for training a survival rate prediction model for medulloblastoma, characterized in that, The medulloblastoma survival rate prediction model training method includes: Obtaining a second training dataset; the second training dataset is a dataset obtained by collecting the patient information of domestic MB patients, the second training dataset includes multiple second samples, the second samples include second feature data and label data, the second feature data includes the molecular subtype, gender, age, metastasis status, histological subtype, whole brain and whole spinal cord radiotherapy dose, posterior fossa whole or tumor bed local boost radiotherapy dose, high or low MYC expression level, high or low MYCN expression level, high or low OTX2 expression level, and high or low GFI1 expression level of MB patients, and the label data includes the survival status of MB patients; all MB patients have undergone surgical resection and received radiotherapy and / or chemotherapy after surgery; Using the second training dataset as input, training multiple initial models to obtain multiple second-trained models and the model performance of each second-trained model; the initial models are models capable of completing prediction functions; Selecting one second-trained model with the best model performance as the medulloblastoma survival rate prediction model; the medulloblastoma survival rate prediction model is used to predict the survival rate of MB patients at the prediction time point.

8. A method for applying a medulloblastoma survival rate prediction model, characterized in that, The application method of the medulloblastoma survival rate prediction model includes: Obtain the patient information of the MB patient to be predicted; the patient information includes the molecular subtype, gender, age, metastasis status, surgical resection situation, histological subtype, treatment plan, whole brain and whole spinal cord radiotherapy dose, and whole posterior fossa or tumor bed local boost radiotherapy dose of the MB patient to be predicted. Alternatively, the patient information includes the molecular subtype, gender, age, metastasis status, histological subtype, whole brain and whole spinal cord radiotherapy dose, whole posterior fossa or tumor bed local boost radiotherapy dose, high or low MYC expression level, high or low MYCN expression level, high or low OTX2 expression level, and high or low GFI1 expression level of the MB patient to be predicted. Using the patient information as input, use the medulloblastoma survival rate prediction model to predict the survival rate of the MB patient to be predicted at the prediction time point, and obtain the predicted survival rate of the MB patient to be predicted; the medulloblastoma survival rate prediction model is a medulloblastoma survival rate prediction model trained by using the medulloblastoma survival rate prediction model training method described in any one of claims 1-6, or the medulloblastoma survival rate prediction model is a medulloblastoma survival rate prediction model trained by using the medulloblastoma survival rate prediction model training method described in claim 7.

9. A computer device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that the processor executes the computer program to implement the medulloblastoma survival rate prediction model training method described in any one of claims 1-6, the medulloblastoma survival rate prediction model training method described in claim 7, or the medulloblastoma survival rate prediction model application method described in claim 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the medulloblastoma survival rate prediction model training method described in any one of claims 1-6, the medulloblastoma survival rate prediction model training method described in claim 7, or the medulloblastoma survival rate prediction model application method described in claim 8.

Citation Information

Patent Citations

  • Genome for molecular typing of medulloblastoma and application thereof

    CN109182517A

  • Osteosarcoma survival prediction method based on machine learning model

    CN117672522A

  • Prediction system for bone marrow suppression risk after radiotherapy of nephroblastoma patient

    CN119694568A

  • Prognostic model of hepatocellular carcinoma based on DDR and ICD gene expression and construction method and application thereof

    US20230383364A1

  • Composition for predicting chance of brain tumor recurrence and survival prognosis, and kit containing same

    WO2011122859A2