Deep learning-based aGVHD risk prediction model training method and device
Through the deep learning-based GBGNN model, aGVHD risk prediction data is trained, combined with the advantages of GBDT and GNN, the problem of low accuracy of aGVHD risk prediction in the prior art is solved, and more efficient and accurate risk prediction is achieved.
Patent Information
- Application Number
- CN202411889163.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-05-16
AI Technical Summary
The traditional statistical methods used in the prior art for aGVHD risk prediction have low prediction accuracy and lack more effective auxiliary prediction methods.
A deep learning-based method is used to train the user sample data carrying aGVHD hierarchical tags through the GBGNN model (combined with GBDT and GNN) to build an aGVHD risk prediction model. This model uses GBDT to process table data and GNN to process graph structure data, improving the understanding and prediction accuracy of complex data patterns.
It improves the accuracy and efficiency of aGVHD risk prediction, and can more accurately predict users' aGVHD risk level, assisting clinical decision-making.
Smart Images

Figure CN120011807A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a deep learning-based aGVHD risk prediction model training method and device. Background Art
[0002] Allogeneic hematopoietic stem cell transplantation is a way to treat malignant or non-malignant blood problems, but the mortality rate of transplantation is high, and acute graft-versus-host disease (aGVHD) is one of the most common serious complications after HSCT.
[0003] Although various types of predictive scoring systems have been created using traditional statistical methods, their predictive accuracy is still unsatisfactory.
[0004] Therefore, how to better assist aGVHD risk prediction has become an urgent problem to be solved in the industry. Summary of the invention
[0005] The present invention provides a deep learning-based aGVHD risk prediction model training method and device, which are used to solve the problem of how to assist aGVHD risk prediction in a better way in the prior art.
[0006] The present invention provides a deep learning-based aGVHD risk prediction model training method, comprising: Obtain multiple original multi-dimensional user sample data with aGVHD classification labels from a real user database; After preprocessing the original multi-dimensional user sample data, target user sample data is screened from the original multi-dimensional user sample data through feature engineering analysis; The GBGNN model is trained based on sample data of each target user carrying an aGVHD classification label, and when a first preset training condition is met, the training is stopped to obtain an aGVHD risk prediction model; wherein the aGVHD risk prediction model includes: a GBDT model and a GNN model; The aGVHD risk prediction model is used to predict the aGVHD risk level based on user data.
[0007] According to a deep learning-based aGVHD risk prediction model training method provided by the present invention, the GBGNN model is trained based on sample data of each target user carrying an aGVHD classification label, comprising: Inputting the target user sample data carrying the aGVHD classification label into the GBDT model for training, and stopping the training when the GBDT model meets the second preset training condition to obtain a trained GBDT model; The output of the GBDT model is used as the input of the GNN model to output the final prediction information of the aGVHD risk level of the target user sample data; wherein the output of the GBDT model carries a first grading label or a second grading label, the first grading label is the aGVHD grading label, and the second grading label is the residual of the parameters in the GBDT model training; When the GNN model meets the first preset training condition, the training is stopped to obtain an aGVHD risk prediction model including the GBDT model and the GNN model.
[0008] According to a deep learning-based aGVHD risk prediction model training method provided by the present invention, the target user sample data carrying the aGVHD classification label is input into the GBDT model for training, and when the GBDT model meets the second preset training condition, the training is stopped to obtain a trained GBDT model; comprising: Inputting the target user sample data into the GBDT model for training, and inputting preliminary prediction information of aGVHD risk level of the target user sample data; Training a new decision tree according to the residual between the preliminary prediction information of the aGVHD risk level and the aGVHD grading label; The new decision tree is added to the GBDT model, and the target user sample data is continued to be trained according to the updated GBDT model until the second preset training condition is met to obtain a trained GBDT model.
[0009] According to a deep learning-based aGVHD risk prediction model training method provided by the present invention, after data preprocessing, the original multi-dimensional user sample data is screened from the original multi-dimensional user sample data through feature engineering analysis, including: After checking and processing missing values, outliers and noise in the original multi-dimensional user sample data, the original multi-dimensional user sample data is format-converted by one-hot encoding or sequence encoding to obtain the original multi-dimensional user sample data after data preprocessing; The importance scores of the user sample data of each dimension in the original multi-dimensional user sample data are calculated by the XGBoost algorithm, and the target user sample data are determined according to the ranking of the importance scores.
[0010] The present invention also provides a method for predicting aGVHD risk based on deep learning, comprising: Acquiring user data to be analyzed, wherein the user data includes at least one of the following: matching type data, transplant type data, user treatment plan information, and user demographic information; Inputting the user data to be analyzed into an aGVHD risk prediction model, and outputting aGVHD risk level prediction corresponding to the user data; Among them, the aGVHD risk prediction model includes: a GBDT model and a GNN model; the aGVHD risk prediction model is trained based on original multi-dimensional user sample data carrying aGVHD grading labels.
[0011] The present invention also provides an aGVHD risk prediction method based on deep learning, which includes, before the step of inputting the user data to be analyzed into the aGVHD risk prediction model and outputting the aGVHD risk level prediction corresponding to the user data: Inputting a plurality of target user sample data carrying aGVHD classification labels into the GBDT model for training, and stopping the training when the GBDT model meets a second preset training condition to obtain a trained GBDT model; The output of the GBDT model is used as the input of the GNN model to output the final prediction information of the aGVHD risk level of the target user sample data; wherein the output of the GBDT model carries a first grading label or a second grading label, the first grading label is the aGVHD grading label, and the second grading label is the residual of the parameters in the GBDT model training; When the GNN model meets the first preset training condition, the training is stopped to obtain an aGVHD risk prediction model including the GBDT model and the GNN model.
[0012] The present invention also provides a deep learning-based aGVHD risk prediction method, which obtains user data to be analyzed, including: Obtain multi-dimensional user data to be analyzed; Convert the format of the multi-dimensional user data to be analyzed by one-hot encoding or sequence encoding to obtain the multi-dimensional user data to be analyzed after data preprocessing; The importance score of each dimension of the multi-dimensional user data to be analyzed is calculated by the XGBoost algorithm, and the user data to be analyzed is obtained according to the ranking of the importance scores.
[0013] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the aGVHD risk prediction model training method based on deep learning or the aGVHD risk prediction method based on deep learning as described above is implemented.
[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the aGVHD risk prediction model training methods based on deep learning or the aGVHD risk prediction methods based on deep learning as described above.
[0015] The present invention also provides a computer program product, comprising a computer program, which, when executed by a processor, implements any of the aGVHD risk prediction model training methods based on deep learning or the aGVHD risk prediction methods based on deep learning as described above.
[0016] The aGVHD risk prediction model training method and device based on deep learning provided by the present invention provide rich training data for the model by obtaining multiple original multi-dimensional user sample data carrying aGVHD grading labels from a real user database. The diversity and dimensionality of these data help the deep learning model capture complex patterns and relationships, thereby improving the accuracy of prediction. The GBGNN model (combining GBDT and GNN) is used for training. The model can effectively learn the decision space in the tabular data and simultaneously consider the neighborhood information and node features of the node for prediction. This combined method takes advantage of the advantages of GBDT in processing tabular data and GNN in processing graph structure data, improves the model's understanding of complex data patterns and the accuracy of prediction, and ultimately achieves accurate prediction of the aGVHD risk level. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0018] Figure 1 It is a flowchart of the aGVHD risk prediction model training method based on deep learning provided by the present invention; Figure 2 The GBGNN architecture diagram provided by the present invention; Figure 3 A schematic diagram of the process of aGVHD risk prediction method based on deep learning provided by the present invention; Figure 4 The overall solution flow chart provided by the present invention; Figure 5 The aGVHD risk prediction model training device based on deep learning provided by the present invention; Figure 6A schematic diagram of the structure of an aGVHD risk prediction device based on deep learning provided by the present invention; Figure 7 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0019] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0020] Figure 1 is a flow chart of the aGVHD risk prediction model training method based on deep learning provided by the present invention, such as Figure 1 As shown, the method includes the following: Step 110, obtaining a plurality of original multi-dimensional user sample data carrying aGVHD classification labels from a real user database; First, the original multi-dimensional user sample data in the real user database contains information of multiple dimensions, which may involve various clinical characteristics of patients, such as demographic information (age, gender, etc.), medical history, treatment process, laboratory test results (including indicators such as liver function and kidney function), and imaging data, etc. The multi-dimensional characteristics of these data provide rich information for the model, enabling it to capture potential patterns and associations related to the occurrence of aGVHD from different angles.
[0021] Secondly, each sample data carries an aGVHD grading label, which is assessed by doctors based on the actual clinical manifestations of patients after receiving allogeneic hematopoietic stem cell transplantation (HSCT). aGVHD grading usually ranges from 0 (asymptomatic) to IV (the most severe symptoms). This grading label provides a clear training target for the model, which is to predict the risk of patients developing different degrees of aGVHD after receiving HSCT.
[0022] During the data extraction process, ensure the quality and integrity of the data, exclude samples with missing values or outliers, to ensure the effectiveness of subsequent model training. In addition, ensure the privacy and security of the data in this application, comply with relevant data protection regulations, and desensitize sensitive information.
[0023] In the present invention, a dataset containing rich clinical information and clear labels is constructed to provide a solid data foundation for subsequent feature engineering, model training and risk prediction. These data will be used for the training of deep learning models in order to achieve high-precision aGVHD risk prediction, thereby assisting clinical decision-making.
[0024] Step 120, after preprocessing the original multi-dimensional user sample data, target user sample data is screened from the original multi-dimensional user sample data through feature engineering analysis; The present invention involves preprocessing the original multi-dimensional user sample data and screening the target user sample data through feature engineering analysis. This process first includes cleaning the data and processing missing values and outliers in the data to ensure the quality and consistency of the data set. Then, the data conversion step encodes the categorical features and standardizes or normalizes the numerical features so that all features are at the same level for easy model processing.
[0025] The present invention further extracts the most informative features from the original features through feature selection and construction steps, and may construct new features to capture more complex patterns in the data. The application of dimensionality reduction technology reduces the number of features while retaining the most important information as much as possible, which helps to improve the efficiency and effectiveness of model training.
[0026] Finally, the present invention screens out target user sample data from the data after preprocessing and feature engineering analysis. These data not only contain the most relevant features for aGVHD risk prediction, but are also divided into training sets, validation sets, and test sets so that they can be used for model training and evaluation in subsequent steps. Such data processing flow lays a solid foundation for building an accurate and efficient aGVHD risk prediction model. Through these steps, the present invention can ensure that the data used for model training is clean, orderly, and informative, thereby improving the prediction accuracy and generalization ability of the model.
[0027] Step 130, training the GBGNN model based on the sample data of each target user carrying the aGVHD classification label, and stopping the training when the first preset training condition is met to obtain the aGVHD risk prediction model; wherein the aGVHD risk prediction model includes: a GBDT model and a GNN model; The aGVHD risk prediction model is used to predict the aGVHD risk level based on user data.
[0028] The present invention trains the GBGNN model by utilizing the target user sample data carrying aGVHD grading labels. The model integrates the advantages of the GBDT and GNN models to build a prediction model that can accurately predict the risk of aGVHD. During the training process, the model learns to identify key features related to the occurrence of aGVHD from multi-dimensional data, and gradually adjusts the model parameters to minimize the prediction error. The present invention also sets preset training conditions, such as reaching a specific accuracy rate or after a certain number of iterations, the training will stop automatically. This is to avoid overfitting of the model and ensure that it can also have good prediction performance on unknown data. When these conditions are met, the training process will terminate, and the obtained aGVHD risk prediction model can combine the ability of GBDT to process structured data with the ability of GNN to process graph structured data to provide accurate risk assessment. The model can predict its aGVHD risk level based on new user data.
[0029] The present invention provides rich training data for the model by obtaining multiple original multi-dimensional user sample data carrying aGVHD grading labels from a real user database. The diversity and dimensionality of these data help the deep learning model capture complex patterns and relationships, thereby improving the accuracy of predictions. The GBGNN model (combining GBDT and GNN) is used for training, which can effectively learn the decision space in tabular data and simultaneously consider the neighborhood information and node features of the node for prediction. This combined method takes advantage of the advantages of GBDT in processing tabular data and GNN in processing graph structure data, improves the model's understanding of complex data patterns and the accuracy of prediction, and ultimately achieves accurate prediction of aGVHD risk levels.
[0030] Optionally, the training of the GBGNN model based on sample data of each target user carrying an aGVHD classification label includes: Inputting the target user sample data carrying the aGVHD classification label into the GBDT model for training, and stopping the training when the GBDT model meets the second preset training condition to obtain a trained GBDT model; The output of the GBDT model is used as the input of the GNN model to output the final prediction information of the aGVHD risk level of the target user sample data; wherein the output of the GBDT model carries a first grading label or a second grading label, the first grading label is the aGVHD grading label, and the second grading label is the residual of the parameters in the GBDT model training; When the GNN model meets the first preset training condition, the training is stopped to obtain an aGVHD risk prediction model including the GBDT model and the GNN model.
[0031] Figure 2 The GBGNN architecture diagram provided by the present invention is as follows: Figure 2 As shown in the figure, the GBGNN architecture consists of two modules, the GBDT module and the GNN module.
[0032] The GBGNN model training process involved in the present invention uses target user sample data carrying aGVHD classification labels for training, wherein the GBDT model is first trained until it meets preset training conditions, such as accuracy or number of iterations, and then stops training.
[0033] At this point, the output of the GBDT model, including the original aGVHD grading labels or parameter residuals generated during training, serves as the input of the GNN model.
[0034] The GNN model continues to train based on these input and graph structure data until another set of preset training conditions is met. Finally, when the training of the GNN model also meets the conditions, the entire GBGNN model, including the trained GBDT and GNN parts, is integrated into an aGVHD risk prediction model. This model can use the information obtained from GBDT and GNN to predict the aGVHD risk level of new user data and provide assistance for clinical decision-making. Through this training method combining GBDT and GNN, the present invention aims to improve the accuracy and efficiency of aGVHD risk prediction.
[0035] Optionally, the target user sample data carrying the aGVHD classification label is input into the GBDT model for training, and when the GBDT model meets the second preset training condition, the training is stopped to obtain a trained GBDT model; including: Inputting the target user sample data into the GBDT model for training, and inputting preliminary prediction information of aGVHD risk level of the target user sample data; Training a new decision tree according to the residual between the preliminary prediction information of the aGVHD risk level and the aGVHD grading label; The new decision tree is added to the GBDT model, and the target user sample data is continued to be trained according to the updated GBDT model until the second preset training condition is met to obtain a trained GBDT model.
[0036] The GBDT model training process described in the present invention is an iterative enhancement process, in which each step is intended to improve the model's prediction accuracy for aGVHD risk level. At the beginning of the training process, the target user sample data with aGVHD grading labels is input into the GBDT model, which includes preliminary prediction information and actual aGVHD grading labels. The goal of the model is to minimize the difference between the predicted information and the actual label.
[0037] In each iteration, the GBDT model calculates the residuals between the current prediction and the aGVHD grade label, which represent a measure of the model's prediction error. Then, a new decision tree is trained using these residuals as new targets. This tree is designed to correct the prediction errors in the previous iteration, i.e., it tries to predict the errors in the previous iteration.
[0038] After the new decision tree is completed, it is added to the GBDT model to enhance the model's predictive ability. Subsequently, the updated GBDT model is used to continue training the target user sample data, and this process is repeated until the model meets the second preset training conditions. These conditions may include that the improvement in model performance is no longer significant, reaching a predetermined accuracy rate or an upper limit on the number of iterations.
[0039] Through this iterative training and step-by-step enhancement method, the GBDT model can gradually approach the optimal solution, and finally obtain a trained GBDT model that can effectively predict the risk level of aGVHD. This model can not only identify the key features that affect the risk of aGVHD, but also provide accurate predictions in practical applications.
[0040] Optionally, GNN consists of multiple layers, each layer represents a nonlinear function, and this nonlinear function needs to be aggregated to obtain : in, is the expression of node A on level n, and UNION and MASS are aggregation functions from the local domain.
[0041] The graph structure with enhanced features is sent to GNN for further processing. At this time, information is transmitted through the connection relationship in the multi-layer graph of GNN, and the context information provided by the GBDT module is integrated, so as to further establish the complex dependencies in the graph data and understand the complex data patterns.
[0042] Optionally, after the original multi-dimensional user sample data is preprocessed, target user sample data is screened from the original multi-dimensional user sample data through feature engineering analysis, including: After checking and processing missing values, outliers and noise in the original multi-dimensional user sample data, the original multi-dimensional user sample data is format-converted by one-hot encoding or sequence encoding to obtain the original multi-dimensional user sample data after data preprocessing; The importance scores of the user sample data of each dimension in the original multi-dimensional user sample data are calculated by the XGBoost algorithm, and the target user sample data are determined according to the ranking of the importance scores.
[0043] In the present invention, first, data cleaning is performed on the original multi-dimensional user sample data, which includes checking the missing values, outliers and noise in the data. Missing values will be filled or deleted, outliers will be identified and processed, and noise data will be filtered or cleared accordingly to ensure the quality and accuracy of the data set.
[0044] Then, the data format is converted. For categorical variables, the original multi-dimensional user sample data is converted into a format suitable for machine learning model processing by using one-hot encoding or sequence encoding. This step helps the model better understand and utilize these variables, because many algorithms are more efficient when processing numerical data.
[0045] Next, the XGBoost algorithm is used to perform feature importance analysis on the preprocessed data. The XGBoost algorithm can calculate the importance score of each feature in the model, which is based on the contribution of the feature in the process of building the tree model. By analyzing these scores, the features that have the greatest impact on the model's predictive ability can be identified.
[0046] Finally, the features are ranked according to their importance scores, for example, the top 20 variables are selected to determine the target user sample data. This step helps reduce the data dimension and remove redundant or irrelevant features, thereby improving the efficiency of model training and the accuracy of prediction.
[0047] In an optional embodiment, the model evaluation index used in the present invention is Accuracy, and the Accuracy formula is as follows: Among them, the classification targets are counted as positive examples and negative examples respectively: True positives (TP): The number of correctly classified as positive examples, that is, the number of instances (number of samples) that are actually positive examples and are classified as positive examples by the classifier; False positives (FP): The number of incorrectly classified as positive examples, that is, the number of instances that are actually negative examples but are classified as positive examples by the classifier; False negatives (FN): The number of incorrectly classified as negative examples, that is, the number of instances that are actually positive examples but are classified as negative examples by the classifier; True negatives (TN): The number of correctly classified as negative examples, that is, the number of instances that are actually negative examples and are classified as negative examples by the classifier. The above model can effectively evaluate the predictive effectiveness of the aGVHD risk prediction model in this application.
[0048] Figure 3 A schematic diagram of the aGVHD risk prediction method based on deep learning provided by the present invention is shown in FIG. Figure 3 As shown, including: Step 310, obtaining user data to be analyzed, wherein the user data includes at least one of the following: matching type data, transplant type data, user treatment plan information, and user demographic information; The present invention involves obtaining comprehensive data of the user to be analyzed, which is crucial to the accuracy of the aGVHD risk prediction model. The collected user data covers key clinical and demographic information, including matching type data, such as details of HLA matching; transplant type data, indicating the source and type of transplant; user's treatment plan information, including pre-treatment and post-transplant treatment plans; and user's demographic information, such as age, gender, etc.
[0049] The collection of these data provides a comprehensive view for the model, enabling it to deeply understand the unique situation of each patient and predict the risk level of aGVHD accordingly. Through the analysis of this comprehensive data, the present invention can provide a more accurate individualized risk assessment.
[0050] Optionally, obtaining the user data to be analyzed includes: Obtain multi-dimensional user data to be analyzed; Convert the format of the multi-dimensional user data to be analyzed by one-hot encoding or sequence encoding to obtain the multi-dimensional user data to be analyzed after data preprocessing; The importance score of each dimension of the multi-dimensional user data to be analyzed is calculated by the XGBoost algorithm, and the user data to be analyzed is obtained according to the ranking of the importance scores.
[0051] The present invention will collect multi-dimensional user data to be analyzed, which includes patient-related information collected from different sources and different dimensions, such as clinical data, treatment history, laboratory test results, etc.
[0052] In order to make these multi-dimensional data suitable for model analysis, the present invention converts the data format through one-hot encoding or sequence encoding. This conversion converts the categorical data into numerical data that can be better processed by the model, thereby preparing for subsequent analysis.
[0053] The present invention uses the XGBoost algorithm to calculate the importance score of each dimension in the multi-dimensional user data to be analyzed. The XGBoost algorithm can evaluate the contribution of each feature to the model's predictive ability, thereby determining which features are most important to the model.
[0054] Finally, according to the importance score ranking obtained by the XGBoost algorithm, the present invention selects the features with the highest scores, which are considered to be the most valuable user data for aGVHD risk prediction. Through this process, the present invention can filter out the most critical information from a large amount of multidimensional data for subsequent risk prediction model training and analysis.
[0055] In the present invention, not only the efficiency of data processing is improved, but also the prediction ability of the model is enhanced through feature selection, ensuring that the model can focus on the most relevant data for learning and prediction.
[0056] Step 320, inputting the user data to be analyzed into an aGVHD risk prediction model, and outputting an aGVHD risk level prediction corresponding to the user data; Among them, the aGVHD risk prediction model includes: a GBDT model and a GNN model; the aGVHD risk prediction model is trained based on original multi-dimensional user sample data carrying aGVHD grading labels.
[0057] In the present invention, the user data to be analyzed is input into the aGVHD risk prediction model, which combines the two technologies of GBDT and GNN, which are trained based on multi-dimensional user sample data with aGVHD grade labels. The model analyzes the input data and outputs the aGVHD risk level prediction corresponding to the user data, helping medical professionals assess the risk of acute graft-versus-host disease in patients.
[0058] Figure 4 The overall solution flow chart provided by the present invention is as follows: Figure 4 As shown, including: First, we need to obtain the real data of hematopoietic stem cell transplant patients, which can include demographic information, treatment plans, disease information, and surgical information. Then, we need to establish a real-world database of hematopoietic stem cell transplant patients, and ensure that the data comes from a legitimate source.
[0059] Then, after data preprocessing and feature engineering, a risk prediction model for aGVHD after HSCT patients was constructed and optimized. Finally, personalized push of aGVHD risk prediction after HSCT was achieved, and the risk influencing factors of aGVHD after HSCT were analyzed.
[0060] In the present invention, by applying GBGNN technology to the prediction of the risk of aGVHD after hematopoietic stem cell transplantation, the technology not only reduces the model prediction time and improves the accuracy of the prediction, but also supports online learning to update the model. In the face of the problem that the patient data is high in dimension and seriously missing, this method can efficiently screen feature variables and process missing data, so that the model can maximize the use of feature data and obtain the best prediction effect with fewer features. In addition, this method improves the accuracy of the prediction model and assists doctors in intervening in patient treatment in advance, thereby reducing the risk of aGVHD in patients.
[0061] The device provided by the present invention is described below. The device described below and the method described above can be referenced to each other.
[0062] Figure 5 The aGVHD risk prediction model training device based on deep learning provided by the present invention is as follows: Figure 5 As shown, including: The first acquisition module 510 is used to acquire a plurality of original multi-dimensional user sample data carrying aGVHD classification labels from a real user database; The screening module 520 is used to pre-process the original multi-dimensional user sample data and then screen the target user sample data from the original multi-dimensional user sample data through feature engineering analysis; The training module 530 is used to train the GBGNN model based on the sample data of each target user carrying the aGVHD classification label, and stop the training when the first preset training condition is met to obtain the aGVHD risk prediction model; wherein the aGVHD risk prediction model includes: a GBDT model and a GNN model; The aGVHD risk prediction model is used to predict the aGVHD risk level based on user data.
[0063] Optionally, the device is also used for: Inputting the target user sample data carrying the aGVHD classification label into the GBDT model for training, and stopping the training when the GBDT model meets the second preset training condition to obtain a trained GBDT model; The output of the GBDT model is used as the input of the GNN model to output the final prediction information of the aGVHD risk level of the target user sample data; wherein the output of the GBDT model carries a first grading label or a second grading label, the first grading label is the aGVHD grading label, and the second grading label is the residual of the parameters in the GBDT model training; When the GNN model meets the first preset training condition, the training is stopped to obtain an aGVHD risk prediction model including the GBDT model and the GNN model.
[0064] Optionally, the device is also used for: Inputting the target user sample data into the GBDT model for training, and inputting preliminary prediction information of aGVHD risk level of the target user sample data; Training a new decision tree according to the residual between the preliminary prediction information of the aGVHD risk level and the aGVHD grading label; The new decision tree is added to the GBDT model, and the target user sample data is continued to be trained according to the updated GBDT model until the second preset training condition is met to obtain a trained GBDT model.
[0065] Optionally, the device is also used for: After checking and processing missing values, outliers and noise in the original multi-dimensional user sample data, the original multi-dimensional user sample data is format-converted by one-hot encoding or sequence encoding to obtain the original multi-dimensional user sample data after data preprocessing; The importance scores of the user sample data of each dimension in the original multi-dimensional user sample data are calculated by the XGBoost algorithm, and the target user sample data are determined according to the ranking of the importance scores.
[0066] Figure 6 A schematic diagram of the structure of the aGVHD risk prediction device based on deep learning provided by the present invention is shown in FIG. Figure 6 As shown, including: The second acquisition module 610 is used to acquire user data to be analyzed, wherein the user data includes at least one of the following: matching type data, transplant type data, user treatment plan information, and user demographic information; The output module 620 is used to input the user data to be analyzed into the aGVHD risk prediction model, and output the aGVHD risk level prediction corresponding to the user data; Among them, the aGVHD risk prediction model includes: a GBDT model and a GNN model; the aGVHD risk prediction model is trained based on original multi-dimensional user sample data carrying aGVHD grading labels.
[0067] Optionally, the device further comprises: a pre-training module; The pre-training module is specifically used for: Inputting a plurality of target user sample data carrying aGVHD classification labels into the GBDT model for training, and stopping the training when the GBDT model meets a second preset training condition to obtain a trained GBDT model; The output of the GBDT model is used as the input of the GNN model to output the final prediction information of the aGVHD risk level of the target user sample data; wherein the output of the GBDT model carries a first grading label or a second grading label, the first grading label is the aGVHD grading label, and the second grading label is the residual of the parameters in the GBDT model training; When the GNN model meets the first preset training condition, the training is stopped to obtain an aGVHD risk prediction model including the GBDT model and the GNN model.
[0068] Figure 7 is a schematic diagram of the structure of the electronic device provided by the present invention, such as Figure 7 As shown, the electronic device may include: a processor 710, a communication interface 720, a memory 730 and a communication bus 740, wherein the processor 710, the communication interface 720, and the memory 730 communicate with each other through the communication bus 740. The processor 710 may call the logic instructions in the memory 730 to execute the aGVHD risk prediction model training method based on deep learning or the aGVHD risk prediction method based on deep learning, the method comprising: obtaining a plurality of original multi-dimensional user sample data carrying aGVHD classification labels from a real user database; After preprocessing the original multi-dimensional user sample data, target user sample data is screened from the original multi-dimensional user sample data through feature engineering analysis; The GBGNN model is trained based on sample data of each target user carrying an aGVHD classification label, and when a first preset training condition is met, the training is stopped to obtain an aGVHD risk prediction model; wherein the aGVHD risk prediction model includes: a GBDT model and a GNN model; The aGVHD risk prediction model is used to predict the aGVHD risk level based on user data.
[0069] or, obtaining user data to be analyzed, wherein the user data includes at least one of the following: matching type data, transplant type data, user treatment plan information, and user demographic information; Inputting the user data to be analyzed into an aGVHD risk prediction model, and outputting aGVHD risk level prediction corresponding to the user data; Among them, the aGVHD risk prediction model includes: a GBDT model and a GNN model; the aGVHD risk prediction model is trained based on original multi-dimensional user sample data carrying aGVHD grading labels.
[0070] In addition, the logic instructions in the above-mentioned memory 730 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present invention can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0071] On the other hand, the present invention further provides a computer program product, the computer program product includes a computer program, the computer program can be stored on a non-transitory computer-readable storage medium, when the computer program is executed by a processor, the computer can execute the aGVHD risk prediction model training method based on deep learning or the aGVHD risk prediction method based on deep learning provided by the above methods, the method comprising: obtaining a plurality of original multi-dimensional user sample data carrying aGVHD grading labels from a real user database; After preprocessing the original multi-dimensional user sample data, target user sample data is screened from the original multi-dimensional user sample data through feature engineering analysis; The GBGNN model is trained based on sample data of each target user carrying an aGVHD classification label, and when a first preset training condition is met, the training is stopped to obtain an aGVHD risk prediction model; wherein the aGVHD risk prediction model includes: a GBDT model and a GNN model; The aGVHD risk prediction model is used to predict the aGVHD risk level based on user data.
[0072] or, obtaining user data to be analyzed, wherein the user data includes at least one of the following: matching type data, transplant type data, user treatment plan information, and user demographic information; Inputting the user data to be analyzed into an aGVHD risk prediction model, and outputting aGVHD risk level prediction corresponding to the user data; Among them, the aGVHD risk prediction model includes: a GBDT model and a GNN model; the aGVHD risk prediction model is trained based on original multi-dimensional user sample data carrying aGVHD grading labels.
[0073] In another aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which is implemented when the computer program is executed by a processor to execute the aGVHD risk prediction model training method based on deep learning or the aGVHD risk prediction method based on deep learning provided by the above methods, the method comprising: obtaining a plurality of original multi-dimensional user sample data carrying aGVHD classification labels from a real user database; After preprocessing the original multi-dimensional user sample data, target user sample data is screened from the original multi-dimensional user sample data through feature engineering analysis; The GBGNN model is trained based on sample data of each target user carrying an aGVHD classification label, and when a first preset training condition is met, the training is stopped to obtain an aGVHD risk prediction model; wherein the aGVHD risk prediction model includes: a GBDT model and a GNN model; The aGVHD risk prediction model is used to predict the aGVHD risk level based on user data.
[0074] or, obtaining user data to be analyzed, wherein the user data includes at least one of the following: matching type data, transplant type data, user treatment plan information, and user demographic information; Inputting the user data to be analyzed into an aGVHD risk prediction model, and outputting aGVHD risk level prediction corresponding to the user data; Among them, the aGVHD risk prediction model includes: a GBDT model and a GNN model; the aGVHD risk prediction model is trained based on original multi-dimensional user sample data carrying aGVHD grading labels.
[0075] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0076] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0077] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A deep learning-based aGVHD risk prediction model training method, characterized in that: include: Obtain multiple original multi-dimensional user sample data with aGVHD classification labels from a real user database; After preprocessing the original multi-dimensional user sample data, target user sample data is screened from the original multi-dimensional user sample data through feature engineering analysis; The GBGNN model is trained based on sample data of each target user carrying an aGVHD classification label, and when a first preset training condition is met, the training is stopped to obtain an aGVHD risk prediction model; wherein the aGVHD risk prediction model includes: a GBDT model and a GNN model; The aGVHD risk prediction model is used to predict the aGVHD risk level based on user data.
2. The aGVHD risk prediction model training method based on deep learning according to claim 1, characterized in that: The GBGNN model is trained based on sample data of each target user carrying aGVHD classification label, including: Inputting the target user sample data carrying the aGVHD classification label into the GBDT model for training, and stopping the training when the GBDT model meets the second preset training condition to obtain a trained GBDT model; The output of the GBDT model is used as the input of the GNN model to output the final prediction information of the aGVHD risk level of the target user sample data; wherein the output of the GBDT model carries a first grading label or a second grading label, the first grading label is the aGVHD grading label, and the second grading label is the residual of the parameters in the GBDT model training; When the GNN model meets the first preset training condition, the training is stopped to obtain an aGVHD risk prediction model including the GBDT model and the GNN model.
3. The aGVHD risk prediction model training method based on deep learning according to claim 2, characterized in that: Inputting the target user sample data carrying the aGVHD classification label into the GBDT model for training, and stopping the training when the GBDT model meets the second preset training condition to obtain a trained GBDT model; comprising: Inputting the target user sample data into the GBDT model for training, and inputting preliminary prediction information of aGVHD risk level of the target user sample data; Training a new decision tree according to the residual between the preliminary prediction information of the aGVHD risk level and the aGVHD grading label; The new decision tree is added to the GBDT model, and the target user sample data is continued to be trained according to the updated GBDT model until the second preset training condition is met to obtain a trained GBDT model.
4. The aGVHD risk prediction model training method based on deep learning according to claim 1, characterized in that: After the original multi-dimensional user sample data is preprocessed, target user sample data is screened from the original multi-dimensional user sample data through feature engineering analysis, including: After checking and processing missing values, outliers and noise in the original multi-dimensional user sample data, the original multi-dimensional user sample data is format-converted by one-hot encoding or sequence encoding to obtain the original multi-dimensional user sample data after data preprocessing; The importance scores of the user sample data of each dimension in the original multi-dimensional user sample data are calculated by the XGBoost algorithm, and the target user sample data are determined according to the ranking of the importance scores.
5. A deep learning-based aGVHD risk prediction method, characterized in that: include: Acquiring user data to be analyzed, wherein the user data includes at least one of the following: matching type data, transplant type data, user treatment plan information, and user demographic information; Inputting the user data to be analyzed into an aGVHD risk prediction model, and outputting aGVHD risk level prediction corresponding to the user data; Among them, the aGVHD risk prediction model includes: a GBDT model and a GNN model; the aGVHD risk prediction model is trained based on original multi-dimensional user sample data carrying aGVHD grading labels.
6. The aGVHD risk prediction method based on deep learning according to claim 5, characterized in that: Before the step of inputting the user data to be analyzed into the aGVHD risk prediction model and outputting the aGVHD risk level prediction corresponding to the user data, the method further includes: Inputting a plurality of target user sample data carrying aGVHD classification labels into the GBDT model for training, and stopping the training when the GBDT model meets a second preset training condition to obtain a trained GBDT model; The output of the GBDT model is used as the input of the GNN model to output the final prediction information of the aGVHD risk level of the target user sample data; wherein the output of the GBDT model carries a first grading label or a second grading label, the first grading label is the aGVHD grading label, and the second grading label is the residual of the parameters in the GBDT model training; When the GNN model meets the first preset training condition, the training is stopped to obtain an aGVHD risk prediction model including the GBDT model and the GNN model.
7. The aGVHD risk prediction method based on deep learning according to claim 5, characterized in that: The obtaining of the user data to be analyzed includes: Obtain multi-dimensional user data to be analyzed; Convert the format of the multi-dimensional user data to be analyzed by one-hot encoding or sequence encoding to obtain the multi-dimensional user data to be analyzed after data preprocessing; The importance score of each dimension of the multi-dimensional user data to be analyzed is calculated by the XGBoost algorithm, and the user data to be analyzed is obtained according to the ranking of the importance scores.
8. A deep learning-based aGVHD risk prediction model training device, characterized in that: include: The first acquisition module is used to acquire a plurality of original multi-dimensional user sample data carrying aGVHD classification labels from a real user database; A screening module, configured to pre-process the original multi-dimensional user sample data and then screen target user sample data from the original multi-dimensional user sample data through feature engineering analysis; A training module, used for training the GBGNN model based on sample data of each target user carrying an aGVHD classification label, and stopping the training when a first preset training condition is met to obtain an aGVHD risk prediction model; wherein the aGVHD risk prediction model includes: a GBDT model and a GNN model; The aGVHD risk prediction model is used to predict the aGVHD risk level based on user data.
9. A deep learning-based aGVHD risk prediction device, characterized in that: include: A second acquisition module is used to acquire user data to be analyzed, wherein the user data includes at least one of the following: matching type data, transplant type data, user treatment plan information, and user demographic information; An output module, used to input the user data to be analyzed into the aGVHD risk prediction model, and output the aGVHD risk level prediction corresponding to the user data; Among them, the aGVHD risk prediction model includes: a GBDT model and a GNN model; the aGVHD risk prediction model is trained based on original multi-dimensional user sample data carrying aGVHD grading labels.
10. The aGVHD risk prediction device based on deep learning according to claim 9, characterized in that: The device also includes: a pre-training module; The pre-training module is specifically used for: Inputting a plurality of target user sample data carrying aGVHD classification labels into the GBDT model for training, and stopping the training when the GBDT model meets a second preset training condition to obtain a trained GBDT model; The output of the GBDT model is used as the input of the GNN model to output the final prediction information of the aGVHD risk level of the target user sample data; wherein the output of the GBDT model carries a first grading label or a second grading label, the first grading label is the aGVHD grading label, and the second grading label is the residual of the parameters in the GBDT model training; When the GNN model meets the first preset training condition, the training is stopped to obtain an aGVHD risk prediction model including the GBDT model and the GNN model.
11. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the aGVHD risk prediction model training method based on deep learning as described in any one of claims 1 to 4 or the aGVHD risk prediction method based on deep learning as described in any one of claims 5 to 7 is implemented.
12. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the aGVHD risk prediction model training method based on deep learning as described in any one of claims 1 to 4 or the aGVHD risk prediction method based on deep learning as described in any one of claims 5 to 7.
Citation Information
Cited By
Prostate cancer life cycle management system based on early screening database
CN121191775A