Cancer recurrence probability deep learning prediction system based on multi-modal time sequence characteristics

Through multimodal timing data fusion and LSTM model optimization, the problems of insufficient data utilization and lack of feature fusion in cancer recurrence prediction are solved, high-precision cancer recurrence probability prediction and personalized treatment decisions are achieved, and initiative and accuracy of cancer diagnosis and treatment are improved.

CN120473183AActive Publication Date: 2025-08-12BEIJING YAOYUN DATA TECH CO LTD

Patent Information

Application Number
CN202510940844.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2025-08-12
Estimated Expiration
2045-07-09

AI Technical Summary

Technical Problem

In the prior art, cancer recurrence prediction relies on single type of data and traditional statistical analysis, and fails to make full use of multimodal timing data, resulting in limited prediction accuracy and reliability, and there are problems such as insufficient utilization of dynamic treatment data, lack of fusion of multimodal timing features, and interference from outliers and redundant features.

Method used

By collecting multimodal timing data of patients, missing processing and outlier elimination, key features are screened using a random forest model, radiotherapy accumulation toxicity model is constructed, four-dimensional timing tensors are generated, and LSTM model is used for training, combined with a random search algorithm to optimize hyperparameters, and high-precision prediction of cancer recurrence probability is achieved.

Benefits of technology

It has significantly improved the initiative and accuracy of cancer recurrence diagnosis and treatment, promoted the use of passive response to active prevention and control, and achieved high-precision cancer recurrence probability prediction and personalized treatment decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120473183A_ABST
    Figure CN120473183A_ABST
Patent Text Reader

Abstract

The invention discloses a cancer recurrence probability deep learning prediction system based on multi-modal time sequence characteristics, and relates to the technical field of medical data processing. The method comprises the following steps: constructing a multi-modal time sequence data set by collecting clinical static data, dynamic treatment data and time sequence monitoring data of a patient; missing value filling and abnormal value elimination based on sliding Z-score are carried out on the data to form an effective data set; screening key clinical static features by using a random forest model; constructing a radiotherapy cumulative toxicity model to quantify residual toxicity, and generating dynamic characteristics by combining the dose change rate three-state characteristics; converting the features into a four-dimensional time sequence tensor through a dynamic sliding window; optimizing LSTM hyper-parameters by adopting random search, and training a high-precision prediction model to output a recurrence probability; finally, differentiated clinical intervention is triggered based on risk grading, a'prediction-decision 'closed-loop system is formed, and the initiative and accuracy of cancer recurrence diagnosis and treatment are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical data processing, and more specifically, relates to a deep learning prediction system for cancer recurrence probability based on multimodal time series features. Background Art

[0002] This invention lies at the intersection of medical information technology and artificial intelligence in the pharmaceutical industry. Specifically, it relates to a system that uses multimodal time-series data to accurately predict the probability of cancer recurrence through a deep learning algorithm. This system is directly applicable to prognostic assessment during clinical cancer treatment, providing a scientific basis for developing personalized treatment plans to improve the survival rate and quality of life of cancer patients.

[0003] Currently, cancer recurrence prediction primarily relies on a single type of data and traditional statistical analysis methods. In clinical practice, recurrence risk assessment is often based on limited information such as a patient's clinical stage and pathological features. For example, the likelihood of cancer recurrence is determined solely based on indicators such as tumor size, grade, and the presence of lymph node metastasis. However, this approach ignores the dynamic changes and complex interactions between multiple factors during cancer progression. Regarding data sources, most existing methods only consider static patient clinical information, such as basic demographics and initial tumor diagnosis, while failing to fully utilize multimodal time-series data such as medication, radiotherapy, follow-up, and tumor event information during treatment. This data captures the dynamic changes during cancer treatment and is crucial for accurately predicting the probability of cancer recurrence. Regarding analytical methods, traditional statistical models, such as logistic regression and decision trees, struggle to handle high-dimensional, complex, multimodal data and nonlinear relationships. These models often assume linear relationships between data and fail to capture complex patterns and underlying information within the data, resulting in limited accuracy and reliability of predictions.

[0004] The existing technologies have problems such as insufficient utilization of dynamic treatment data, lack of multimodal time series feature fusion, interference from outliers and redundant features, and lack of deep learning applications. Summary of the Invention

[0005] (1) Technical problems solved In response to the problems in the related art, the present invention provides a deep learning prediction method for cancer recurrence probability based on multimodal time series features to overcome the above-mentioned technical problems existing in the existing related art.

[0006] (2) Technical solution To solve the above technical problems, the present invention is achieved through the following technical solutions: S1. Collect the patient's clinical static data, dynamic treatment data, and time series monitoring data to obtain the patient's multimodal time series data; S2. Perform missing and outlier processing on the patient's multimodal time series data to obtain valid patient multimodal time series data; S3. Selected clinical static features are screened from the effective patient multimodal time series data based on the random forest model to obtain patient selected static feature data; S4. Calculate the cumulative toxicity index and dose change rate characteristics of drug treatment in the multimodal time series data of valid patients to obtain patient dynamic characteristic data; Perform time series formatting operations on the patient's selected static feature data and the patient's dynamic feature data to obtain a four-dimensional time series feature tensor; S5. Build an LSTM model. Train the LSTM model based on the labeled historical four-dimensional time series feature tensor data combined with a random search algorithm to obtain an LSTM cancer recurrence prediction model. The four-dimensional time series feature tensor is input into the LSTM cancer recurrence prediction model to obtain the predicted probability of cancer recurrence in patients; S6. Setting a recurrence probability treatment rule; based on the recurrence probability treatment rule, predicting the patient's recurrence probability, and treating the cancer patient; This invention uses multimodal time series data fusion, combined with intelligent screening of key features and an innovative radiotherapy cumulative toxicity model, to construct an LSTM deep learning model with optimized four-dimensional time series tensor input, thereby achieving high-precision prediction of cancer recurrence probability. It also triggers differentiated clinical intervention based on risk grading, forming a "prediction-decision-making" closed loop, significantly improving the initiative and accuracy of cancer recurrence diagnosis and treatment, and promoting the transformation of cancer recurrence diagnosis and treatment from passive response to active prevention and control.

[0007] Preferably, the S1 comprises the following steps: S11. Collect the patient's age, gender, tumor stage, and pathological classification from the hospital information system to obtain the patient's static clinical data; S12. Extracting the patient's medication sequence and radiotherapy parameters from the electronic medical record to obtain the patient's dynamic treatment data; the medication sequence includes the name of the drug, the single dose, and the administration time point; the radiotherapy parameters include each radiotherapy dose, the irradiation site, and the cumulative dose change curve; S13. Obtain the patient's tumor marker test values and test timestamps, as well as key indicator changes in imaging examination results, from the follow-up system to obtain the patient's time-series monitoring data; The patient's clinical static data, patient's dynamic treatment data and patient's time series monitoring data together constitute the patient's multimodal time series data; This invention integrates three sources: hospital information systems, electronic medical records, and follow-up systems to construct a multimodal time series dataset containing dynamic and static features and precise timestamps, providing a comprehensive and coherent picture of the patient's treatment process for subsequent analysis, effectively avoiding data fragmentation and supporting accurate recurrence prediction.

[0008] Preferably, said S2 comprises the following steps: S21. Filling continuous variables in the patient's multimodal time series data with time series linear interpolation to obtain filled patient multimodal time series data; S22. For the categorical variables in the filled multimodal time series data of the patient, mode filling is performed to obtain the filled multimodal time series data of the patient; S23, setting an invalid data threshold; calculating the sliding Zscore of the tumor marker detection value in the padded patient multimodal time series data to obtain a sliding Zscore set; taking the value in the sliding Zscore set where the sliding Zscore is greater than the invalid data threshold as an invalid sliding Zscore to obtain an invalid sliding Zscore set; Eliminate the tumor marker detection values in the padded patient multimodal time series data corresponding to the invalid sliding Zscore set to obtain valid patient multimodal time series data; The present invention solves the problem of missing data through linear interpolation of continuous variables in time series and mode filling of categorical variables, and then accurately identifies and eliminates abnormal detection values of tumor markers based on the sliding Z-score dynamic threshold, significantly improving the integrity and reliability of multimodal data, and laying a high-quality data foundation for subsequent modeling.

[0009] Preferably, the step S3 includes the following steps: S31. Setting clinical characteristics of the tumor to obtain a tumor clinical characteristic set; the tumor clinical characteristic set includes all clinical characteristics of this type of tumor; Based on the tumor clinical feature set, historical tumor clinical feature data were collected; S32. Initialize the random forest model and set the feature importance evaluation criterion of the random forest model to the Gini impurity reduction; The random forest model was trained using historical tumor clinical feature data combined with a cross-validation algorithm to obtain a trained random forest classifier. S33. Calculate the sum of purity improvements of each tumor clinical feature in the effective patient multimodal time series data at the time of decision tree node splitting using the trained random forest classifier to obtain a set of improvement sums; S34. Sort all features in the historical tumor clinical feature data in descending order according to the size of the total lift in the total lift set, perform normalization, record the importance scores of the historical tumor clinical features, and obtain a tumor clinical feature importance ranking; Set the correlation threshold and the number of key features; select the key features with the highest importance scores based on the importance ranking of tumor clinical features and the number of key features to obtain the key tumor clinical feature set; Verify the clinical significance of key features in the key tumor clinical feature set, exclude redundant features with correlation greater than the correlation threshold, and obtain a selected static feature set; Based on the selected static feature set, the selected static features of the patient are extracted from the effective patient multimodal time series data to obtain the patient selected static feature data; The present invention uses a random forest model combined with cross-validation training to quantify the importance of each clinical feature and rank it; selects key features, verifies clinical significance and eliminates redundant variables, and ultimately extracts high-value static features; significantly improves feature representativeness and reduces model complexity, providing a core indicator set for accurate recurrence prediction.

[0010] Preferably, the S4 comprises the following steps: S41. Collecting the patient's medication time data from the valid patient multimodal time series data to obtain a standardized medication time series; Set drug attenuation parameters and construct a radiotherapy cumulative toxicity model; calculate the residual toxicity of each drug administration at the current time point through the radiotherapy cumulative toxicity model to obtain the real-time toxicity weight of each drug administration; S42. Based on the real-time toxicity weight of each drug administration, the toxicity weights of all effective drug administrations at the same time point are accumulated, and toxicity superposition correction is performed on the combined drug administration scenario to generate a time-toxicity curve to obtain a dynamic toxicity accumulation characteristic curve; Based on the dynamic toxicity accumulation characteristic curve, the actual liver and kidney function test values of the patients are compared, and the drug attenuation parameters are adjusted to make the curve consistent with the clinical toxicity manifestations. The toxicity levels are then divided to obtain a calibrated clinically applicable toxicity curve. S43. Extracting patient radiotherapy dose time series records from valid patient multimodal time series data, arranging the dose data in chronological order of treatment to form a time series; verifying data continuity to obtain a standardized dose time series; The difference between each dose in the standardized dose time series was calculated using the changes in adjacent time periods; dose increases were set as positive values and dose decreases as negative values to obtain a dose change series; S44, setting a change threshold; based on the dose change sequence and the change threshold, converting the change into a treatment intensity adjustment flag, generating a three-state characteristic value, and obtaining a dose change rate characteristic vector; S45, aligning all the patient static feature data, the calibrated clinically available toxicity curve, and the dose change rate feature vector according to the time point to obtain an original time series feature matrix; The dynamic sliding window method is used to perform window slicing and filling operations on the original feature matrix to obtain a four-dimensional time series feature tensor; The present invention innovatively constructs a radiotherapy cumulative toxicity model, generates a clinical calibration curve through toxicity superposition correction, extracts the three-state characteristics of dose variation to characterize the adjustment of treatment intensity, and finally uses a dynamic sliding window to convert multi-source features into a four-dimensional time series tensor, which not only accurately quantifies the dynamic effects of treatment, but also solves the problem of variable-length time series data input, providing high-value structured features for the LSTM model.

[0011] Preferably, the S5 comprises the following steps: S51, constructing an LSTM model and setting hyperparameters for training the LSTM model; S52. Collecting four-dimensional feature tensor data of historical cancer patients and label data indicating whether the historical cancer patients have relapsed, to obtain labeled historical four-dimensional time series feature tensor data; S53. Use labeled historical four-dimensional time series feature tensor data to train the LSTM model. During the training process, use a random search algorithm to find the optimal hyperparameters of the LSTM model and obtain the optimal solution. Use the optimal solution as the hyperparameters of the LSTM model to obtain the LSTM cancer recurrence prediction model. S54, inputting the four-dimensional time series feature tensor into the LSTM cancer recurrence prediction model to obtain the predicted patient recurrence probability; The present invention constructs an LSTM model, utilizes labeled four-dimensional time series feature tensor data, and combines it with a random search algorithm to efficiently optimize hyperparameters to train a high-precision cancer recurrence prediction model. This model can deeply explore the complex associations between multimodal time series features and output individualized recurrence probabilities, providing a reliable quantitative basis for clinical decision-making and significantly improving prediction accuracy and practicality.

[0012] Preferably, in the training process in S53, a random search algorithm is used to find the optimal hyperparameters of the LSTM model, and obtaining the optimal solution includes the following steps: S531, set the parameter search space to L , set the maximum number of random searches to M , the number of cross validation folds is E , initialize the optimal hyperparameters; S532. Under the optimal hyperparameters, use the labeled historical four-dimensional time series feature tensor data to train and evaluate the LSTM model to obtain the best prediction accuracy of the LSTM model. d ; S533, define a candidate hyperparameter set; randomly select a hyperparameter from the candidate hyperparameter set f , the training data in the labeled historical four-dimensional time series feature tensor data is divided into K Fold, for each foldk , ; Use except k The training data in all the labeled historical four-dimensional time series feature tensor data outside the fold are used to train the LSTM model from scratch. k The training data evaluation model in the folded labeled historical four-dimensional time series feature tensor data is used to obtain the prediction accuracy. The mean of the prediction accuracy of all folds is calculated to obtain the average prediction accuracy. g ; If the average prediction accuracy g >Best prediction accuracy d , then the updated best prediction accuracy is g , update the optimal hyperparameters f Otherwise, the original best prediction accuracy and optimal hyperparameters are maintained; S534, repeat S533, and when the maximum number of random searches M is reached, stop the iteration and obtain the optimal solution; The present invention efficiently explores LSTM hyperparameter combinations by combining this random search algorithm with K-fold cross-validation: hyperparameters are randomly selected each time, trained with K-1 fold data, one fold is left for validation, and the average accuracy g is calculated; when g exceeds the historical best d, the optimal hyperparameters are updated; this mechanism locks in the optimal configuration with limited computational cost, significantly improving the model's generalization ability and prediction stability, and avoiding the risk of overfitting.

[0013] Preferably, the S6 comprises the following steps: S61. Set treatment rules for recurrence probability; S62. Treating cancer patients based on a set recurrence probability treatment rule and a predicted recurrence probability of the patient; The present invention converts the predicted probability into differentiated clinical intervention measures by setting recurrence probability treatment rules, achieving seamless conversion of predicted results to treatment decisions, improving the accuracy of diagnosis and treatment, and avoiding over-medicalization or delayed treatment.

[0014] A cancer recurrence probability deep learning prediction system based on multimodal time series features is used to implement the above-mentioned cancer recurrence probability deep learning prediction method based on multimodal time series features, including a data acquisition module, a data processing module, a static feature screening module, a dynamic feature calculation and formatting module, a model building and prediction module, and a treatment decision module; The data acquisition module collects the patient's original data from multiple medical systems, including clinical static data, dynamic treatment data, and time series monitoring data; these data together constitute the patient's multimodal time series data, providing basic input for subsequent analysis; The data processing module is used to preprocess the collected multimodal time series data to address data quality issues; use time series linear interpolation to fill missing values of continuous variables, use the mode to fill missing values of categorical variables, and then calculate the sliding Z-score of tumor marker detection values to identify and eliminate outliers; and finally output valid patient multimodal time series data to ensure data integrity and reliability; The feature screening module uses a random forest model to select the most predictive static features from valid data. The random forest model is initialized, trained using historical tumor clinical feature data, and feature importance scores are calculated based on Gini impurity reduction. The features are then sorted and normalized, and the key features with the highest importance scores are selected. Their clinical significance is verified and redundant features are excluded. Finally, the patient's selected static feature data is output, and the feature set is optimized to improve prediction accuracy. The feature calculation and formatting module calculates dynamic treatment features and converts all features into a format that can be input by the model; constructs a radiotherapy cumulative toxicity model to calculate the cumulative toxicity index, generates a time-toxicity curve, and calibrates the toxicity level based on clinical data; extracts dose change rate features from radiotherapy dose time series records, calculates the variation sequence by adjacent dose differences, and converts it into a three-state feature value; aligns static features, toxicity curves, and dose change features by time points, and uses a dynamic sliding window method to generate a four-dimensional time series feature tensor for easy processing by deep learning models; The model building and prediction module is used to build and optimize an LSTM deep learning model to predict the probability of cancer recurrence. By collecting labeled historical four-dimensional time series feature tensor data, the LSTM model is built and hyperparameters are optimized using a random search algorithm. After training, the LSTM cancer recurrence prediction model is obtained. The input is the current patient's four-dimensional time series feature tensor, and the output is the predicted patient's cancer recurrence probability, achieving high-precision time series prediction. The treatment decision module formulates clinical treatment strategies based on the predicted recurrence probability; sets recurrence probability treatment rules; and performs personalized treatment interventions on cancer patients based on the predicted probability to ensure that the predicted results are converted into actual clinical actions.

[0015] (3) Beneficial effects The present invention has the following beneficial effects: This invention uses multimodal time series data fusion, combined with intelligent screening of key features and an innovative radiotherapy cumulative toxicity model, to construct an LSTM deep learning model with optimized four-dimensional time series tensor input, thereby achieving high-precision prediction of cancer recurrence probability. It also triggers differentiated clinical intervention based on risk grading, forming a "prediction-decision-making" closed loop, significantly improving the initiative and accuracy of cancer recurrence diagnosis and treatment, and promoting the transformation of cancer recurrence diagnosis and treatment from passive response to active prevention and control.

[0016] The present invention constructs a multimodal dataset with a unified timeline through structured collection of clinical static data, dynamic treatment data, and time-series monitoring data; uses a sliding Z-score to eliminate tumor marker outliers, and combines random forest feature screening to compress redundant static features, significantly improving data quality and feature representativeness.

[0017] The present invention introduces drug attenuation parameters to create a radiotherapy cumulative toxicity index model, quantifies the residual toxicity of drug administration at different time points, generates a time-toxicity curve through toxicity superposition correction, and calibrates the model based on liver and kidney function test values; extracts the three-state characteristics of the dose change rate, converts the treatment intensity adjustment into a calculable timing mark, and provides the model with key treatment decision signals.

[0018] The present invention converts the original feature matrix into a four-dimensional tensor through a dynamic sliding window to solve the problem of variable-length time series data input; the window overlapping design retains the continuity of the treatment stage, and the front-end zero-padding strategy is compatible with short-cycle patient data; combined with random search hyperparameter optimization, the cross-validation average accuracy is used as an indicator to efficiently lock the optimal LSTM configuration and improve the reliability of recurrence probability prediction.

[0019] The present invention sets hierarchical treatment rules based on predicted probability, avoids over-medicalization, can screen out signs of recurrence earlier, and promptly treat patients with a high probability of recurrence.

[0020] Of course, any product implementing the present invention does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions of the embodiments of the invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the invention. For ordinary technicians in this field, they can also obtain drawings based on these drawings without paying any creative work.

[0022] Figure 1 Schematic diagram of the process of the deep learning prediction method for cancer recurrence probability based on multimodal time series features of the present invention; Figure 2 Schematic diagram of the process of processing patient data in the deep learning prediction method for cancer recurrence probability based on multimodal time series features of the present invention; Figure 3 This is a schematic diagram of the process of collecting historical cancer recurrence patient time series feature data in the cancer recurrence probability deep learning prediction method based on multimodal time series features of the present invention; Figure 4 This is a module diagram of the deep learning prediction system for cancer recurrence probability based on multimodal time series features of the present invention. DETAILED DESCRIPTION

[0023] The following will clearly and completely describe the technical solutions in the embodiments of the invention in conjunction with the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of the invention, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.

[0024] In the description of the present invention, it should be understood that the terms "opening", "upper", "lower", "top", "middle", "inside" and the like indicating orientation or positional relationship are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the components or elements referred to must have a specific orientation, be constructed and operate in a specific orientation, and therefore cannot be understood as limiting the invention.

[0025] Example 1:

[0026] See also Figure 1 、 Figure 2 、 Figure 3 The present invention discloses a deep learning prediction method for cancer recurrence probability based on multimodal time series features, comprising the following steps: S1. Collect the patient's clinical static data, dynamic treatment data, and time series monitoring data to obtain the patient's multimodal time series data; Said S1 comprises the following steps: S11. Collect the patient's age, gender, tumor stage, and pathological classification from the hospital information system (HIS) to obtain the patient's static clinical data; S12. Extract the patient's medication sequence and radiotherapy parameters from the electronic medical record (EMR) to obtain the patient's dynamic treatment data; the medication sequence includes the drug name, single dose, and administration time (accurate to the day); the radiotherapy parameters include each radiotherapy dose, irradiation site, and cumulative dose change curve; S13. Obtain multiple test values and test timestamps of the patient's tumor markers (such as CEA and CA125) and changes in key indicators of imaging examination results (CT / MRI) from the follow-up system to obtain the patient's time-series monitoring data; The patient's clinical static data, patient's dynamic treatment data and patient's time series monitoring data together constitute the patient's multimodal time series data; S2. Perform missing and outlier processing on the patient's multimodal time series data to obtain valid patient multimodal time series data; The S2 comprises the following steps: S21. For continuous variables (such as drug concentration) in the patient's multimodal time series data, use time series linear interpolation to fill in the data to obtain the filled patient multimodal time series data; S22. For the categorical variables (such as radiotherapy type) in the filled multimodal time series data of the patient, mode filling is performed to obtain the filled multimodal time series data of the patient; S23, setting an invalid data threshold; calculating the sliding Zscore of the tumor marker detection value in the padded patient multimodal time series data to obtain a sliding Zscore set; taking the value in the sliding Zscore set where the sliding Zscore is greater than the invalid data threshold as an invalid sliding Zscore to obtain an invalid sliding Zscore set; Eliminate the tumor marker detection values in the padded patient multimodal time series data corresponding to the invalid sliding Zscore set to obtain valid patient multimodal time series data; S3. Selected clinical static features are screened from the effective patient multimodal time series data based on the random forest model to obtain patient selected static feature data; The S3 includes the following steps: S31. Setting clinical characteristics of the tumor to obtain a tumor clinical characteristic set; the tumor clinical characteristic set includes all clinical characteristics of this type of tumor, such as TNM stage and Ki67 index; Based on the tumor clinical feature set, historical tumor clinical feature data were collected; S32. Initialize a random forest model containing 100 decision trees and set the feature importance evaluation criterion of the random forest model to Gini impurity reduction; The random forest model was trained using historical tumor clinical feature data combined with a cross-validation algorithm to obtain a trained random forest classifier. S33. Calculate the sum of purity improvements of each tumor clinical feature in the effective patient multimodal time series data at the time of decision tree node splitting using the trained random forest classifier to obtain a set of improvement sums; S34. Sort all features in the historical tumor clinical feature data in descending order according to the size of the total lift in the total lift set, perform normalization, record the importance scores of the historical tumor clinical features, and obtain a tumor clinical feature importance ranking; The correlation threshold and the number of key features were set to 10. Based on the importance ranking of tumor clinical features and the number of key features, the top 10 features with the highest importance scores were selected to obtain the key tumor clinical feature set. Verify the clinical significance of key features in the key tumor clinical feature set (such as the Ki-67 index reflecting cell proliferation activity), exclude redundant features with correlations greater than the correlation threshold, and obtain a selected static feature set (including core indicators such as EGFR mutation status); Based on the selected static feature set, the selected static features of the patient are extracted from the effective patient multimodal time series data to obtain the patient selected static feature data; S4. Calculate the cumulative toxicity index and dose change rate characteristics of drug treatment in the multimodal time series data of valid patients to obtain patient dynamic characteristic data; Perform time series formatting operations on the patient's selected static feature data and the patient's dynamic feature data to obtain a four-dimensional time series feature tensor; The S4 comprises the following steps: S41. Collecting the patient's medication time data from the valid patient multimodal time series data to obtain a standardized medication time series; Setting drug decay parameters λ , construct a radiotherapy cumulative toxicity model; the radiotherapy cumulative toxicity model formula is as follows, ; in, TOX cum ( t ) indicates the current time point t The residual toxicity n The total number of radiotherapy treatments, D i Indicates the i The dose of radiotherapy, t i Indicates the i The timing of the dose of radiotherapy e represents the base of natural logarithms; The radiotherapy cumulative toxicity model is used to calculate the residual toxicity of each administration at the current time point, and the real-time toxicity weight of each administration is obtained. If the drug is recently administered, the toxicity remains above 90%; if it was administered 30 days ago, the toxicity decreases to about 35%; if it was administered 60 days ago, the toxicity decreases to about 12%; S42. Based on the real-time toxicity weight of each drug administration, the toxicity weights of all effective drug administrations at the same time point are accumulated, and toxicity superposition correction is performed on the combined drug administration scenario to generate a time-toxicity curve to obtain a dynamic toxicity accumulation characteristic curve; Based on the dynamic toxicity accumulation characteristic curve, compared with the patient's actual liver and kidney function test values, the drug attenuation parameters are adjusted to make the curve consistent with clinical toxicity manifestations, and the toxicity levels are divided to obtain a calibrated clinically applicable toxicity curve; for example, mild is 0% to 20%, moderate is 20% to 50%, and severe is >50%; For example, a breast cancer patient received 8 cycles of chemotherapy with a cumulative paclitaxel dose of 960 mg / m². The calculated cumulative toxicity value was 58.3% (severe), and the actual manifestation was grade 3 neutropenia (verification of consistency). S43. Extract patient radiotherapy dose time series records (e.g., single dose values recorded by day or week) from valid patient multimodal time series data, arrange the dose data in chronological order of treatment to form a time series; verify data continuity (ensure there are no date gaps) to obtain a standardized dose time series; The difference between each dose in the standardized dose time series was calculated using the changes in adjacent time periods; dose increases were set as positive values and dose decreases as negative values to obtain a dose change series; S44. Set a dose change threshold; based on the dose change sequence and the dose change threshold, convert the dose change into a treatment intensity adjustment flag (e.g., an increase threshold > 1 Gy indicates intensive treatment, i.e., a dose reduction adjustment), and generate three-state characteristic values (an increase threshold > 1 Gy is marked as +1, indicating a significant dose increase; a dose change between +1 and +1 is marked as 0, indicating a stable dose; and a decrease threshold < -1 Gy is marked as -1, indicating a significant dose reduction), thereby obtaining a dose change rate characteristic vector; S45, aligning all the patient static feature data, the calibrated clinically available toxicity curve, and the dose change rate feature vector according to the time point to obtain an original time series feature matrix; The dynamic sliding window method is used to perform window slicing and filling operations on the original feature matrix to obtain a four-dimensional time series feature tensor. The specific steps are as follows: Construct the original time series feature matrix (rows represent time points, columns represent features); set the window length to cover 10 consecutive time points (approximately 3-6 months of treatment), set the sliding step size to move 2 time points each time (the new and old windows overlap by 50%), and set the edge processing to fill zeros when the front end is insufficient and retain new data at the end; Based on the time series feature matrix, complete window slicing is performed. For example, if a patient has 24 months of data, 7 windows can be sliced out. A front-end zero padding operation is performed. For example, if a patient has only 5 months of data, 5 zero values need to be padded to the window. The end retention rule is implemented, and actual data (such as the end window) is retained without forced padding, to obtain a window sequence set. For all the window sequences of patients in the window sequence set, stack them by dimension to obtain the stacked feature time series tensor; for example, dimension 1 is the number of patient samples (450 cases), dimension 2 is the number of time windows (e.g., 7 per patient), dimension 3 is the time step within the window (fixed at 10 steps), and dimension 4 is the feature dimension (18 dimensions); The missing values in the stacked feature time series tensor are padded with 0 to obtain a four-dimensional time series feature tensor; S5. Build an LSTM model. Train the LSTM model based on the labeled historical four-dimensional time series feature tensor data combined with a random search algorithm to obtain an LSTM cancer recurrence prediction model. The four-dimensional time series feature tensor is input into the LSTM cancer recurrence prediction model to obtain the predicted probability of cancer recurrence in patients; The S5 comprises the following steps: S51, constructing an LSTM model and setting hyperparameters for training the LSTM model; S52. Collecting four-dimensional feature tensor data of historical cancer patients and label data indicating whether the historical cancer patients have relapsed, to obtain labeled historical four-dimensional time series feature tensor data; S53. Use labeled historical four-dimensional time series feature tensor data to train the LSTM model. During the training process, use a random search algorithm to find the optimal hyperparameters of the LSTM model and obtain the optimal solution. Use the optimal solution as the hyperparameters of the LSTM model to obtain the LSTM cancer recurrence prediction model. In the training process in S53, a random search algorithm is used to find the optimal hyperparameters of the LSTM model. Obtaining the optimal solution includes the following steps: S531, set the parameter search space to L ( L is a candidate set of learning rates, i.e. a set of predefined hyperparameters), and the maximum number of random searches is set to M , the number of cross validation folds is E , initialize the optimal hyperparameters; S532. Under the optimal hyperparameters, use the labeled historical four-dimensional time series feature tensor data to train and evaluate the LSTM model to obtain the best prediction accuracy of the LSTM model; S533, define the candidate hyperparameter set (the candidate learning rate set is all candidates in L); randomly select hyperparameters from the candidate hyperparameter set as f , the training data in the labeled historical four-dimensional time series feature tensor data is divided into K Fold, for each fold k , ; Use the training data in all the labeled historical four-dimensional time series feature tensor data except the w-th fold to train the LSTM model from scratch, and use the k The training data evaluation model in the folded labeled historical four-dimensional time series feature tensor data is used to obtain the prediction accuracy. The mean of the prediction accuracy of all folds is calculated to obtain the average prediction accuracy. g , the calculation formula is as follows, ; in, g represents the average prediction accuracy, k Indicates the k A discount, K represents the number of cross validation folds, T k Indicates thek The prediction accuracy of each fold; If the average prediction accuracy g >Best prediction accuracy d , then the updated best prediction accuracy is g , update the optimal hyperparameters f Otherwise, the original best prediction accuracy and optimal hyperparameters are maintained; S534, repeat S533, and when the maximum number of random searches M is reached, stop the iteration and obtain the optimal solution; S54, inputting the four-dimensional time series feature tensor into the LSTM cancer recurrence prediction model to obtain the predicted patient recurrence probability; S6. setting a recurrence probability treatment rule; treating the cancer patient based on the set recurrence probability treatment rule and the predicted recurrence probability of the patient; The S6 comprises the following steps: S61. Set treatment rules for the probability of recurrence: if the risk is low (<0.3), routine 3-month follow-up is recommended; if the risk is moderate (between 0.3 and 0.7), additional circulating tumor DNA testing is recommended; if the risk is high (>0.7), a second-line treatment plan is initiated. S62. Treat cancer patients based on the recurrence probability level and the predicted recurrence probability of the patients.

[0027] Example 2:

[0028] See also Figure 4 , a deep learning prediction system for cancer recurrence probability based on multimodal time series features, used to implement the above-mentioned deep learning prediction method for cancer recurrence probability based on multimodal time series features, including a data acquisition module, a data processing module, a static feature screening module, a dynamic feature calculation and formatting module, a model building and prediction module, and a treatment decision module; The data acquisition module collects the patient's original data from multiple medical systems, including clinical static data, dynamic treatment data, and time series monitoring data; these data together constitute the patient's multimodal time series data, providing basic input for subsequent analysis; The data processing module is used to preprocess the collected multimodal time series data to address data quality issues; use time series linear interpolation to fill missing values of continuous variables, use the mode to fill missing values of categorical variables, and then calculate the sliding Z-score of tumor marker detection values to identify and eliminate outliers; and finally output valid patient multimodal time series data to ensure data integrity and reliability; The feature screening module uses a random forest model to select the most predictive static features from valid data. The random forest model is initialized, trained using historical tumor clinical feature data, and feature importance scores are calculated based on Gini impurity reduction. The features are then sorted and normalized, and the key features with the highest importance scores are selected. Their clinical significance is verified and redundant features are excluded. Finally, the patient's selected static feature data is output, and the feature set is optimized to improve prediction accuracy. The feature calculation and formatting module calculates dynamic treatment features and converts all features into a format that can be input by the model; constructs a radiotherapy cumulative toxicity model to calculate the cumulative toxicity index, generates a time-toxicity curve, and calibrates the toxicity level based on clinical data; extracts dose change rate features from radiotherapy dose time series records, calculates the variation sequence by adjacent dose differences, and converts it into a three-state feature value; aligns static features, toxicity curves, and dose change features by time points, and uses a dynamic sliding window method to generate a four-dimensional time series feature tensor for easy processing by deep learning models; The model building and prediction module is used to build and optimize an LSTM deep learning model to predict the probability of cancer recurrence. By collecting labeled historical four-dimensional time series feature tensor data, the LSTM model is built and hyperparameters are optimized using a random search algorithm. After training, the LSTM cancer recurrence prediction model is obtained. The input is the current patient's four-dimensional time series feature tensor, and the output is the predicted patient's cancer recurrence probability, achieving high-precision time series prediction. The treatment decision module formulates clinical treatment strategies based on the predicted recurrence probability; sets recurrence probability treatment rules; and performs personalized treatment interventions on cancer patients based on the predicted probability to ensure that the predicted results are converted into actual clinical actions.

[0029] Throughout this specification, references to terms such as "one embodiment," "example," or "specific example" indicate that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in at least one embodiment or example of the invention. In this specification, schematic representations of these terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0030] The preferred embodiments of the invention disclosed above are intended only to help illustrate the invention. These preferred embodiments do not exhaust all details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations are possible based on the content of this specification. These embodiments are selected and described in detail in this specification to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention.

Claims

1. A deep learning prediction method for cancer recurrence probability based on multimodal time series features, characterized by: The following steps are involved: S1. Collect the patient's clinical static data, dynamic treatment data, and time series monitoring data to obtain the patient's multimodal time series data; S2. Perform missing and outlier processing on the patient's multimodal time series data to obtain valid patient multimodal time series data; S3. Selected clinical static features are screened from the effective patient multimodal time series data based on the random forest model to obtain patient selected static feature data; S4. Calculate the cumulative toxicity index and dose change rate characteristics of drug treatment in the multimodal time series data of valid patients to obtain patient dynamic characteristic data; Perform time series formatting operations on the patient's selected static feature data and the patient's dynamic feature data to obtain a four-dimensional time series feature tensor; S5. Build an LSTM model. Train the LSTM model based on the labeled historical four-dimensional time series feature tensor data combined with a random search algorithm to obtain an LSTM cancer recurrence prediction model. The four-dimensional time series feature tensor is input into the LSTM cancer recurrence prediction model to obtain the predicted probability of cancer recurrence in patients; S6. Set a recurrence probability treatment rule; based on the recurrence probability treatment rule, predict the patient's recurrence probability and treat the cancer patient.

2. The method for predicting cancer recurrence probability by deep learning based on multimodal time series features according to claim 1, characterized in that: Said S1 comprises the following steps: S11. Collect the patient's age, gender, tumor stage, and pathological classification from the hospital information system to obtain the patient's static clinical data; S12, extracting the patient's drug treatment sequence and radiotherapy parameters from the electronic medical record to obtain the patient's dynamic treatment data; S13. Obtain the patient's tumor marker test values and test timestamps, as well as key indicator changes in imaging examination results, from the follow-up system to obtain the patient's time-series monitoring data; The patient's clinical static data, patient's dynamic treatment data and patient's time series monitoring data together constitute the patient's multimodal time series data.

3. The method for predicting cancer recurrence probability by deep learning based on multimodal time series features according to claim 1, characterized in that: The S2 comprises the following steps: S21. Filling continuous variables in the patient's multimodal time series data with time series linear interpolation to obtain filled patient multimodal time series data; S22. For the categorical variables in the filled multimodal time series data of the patient, mode filling is performed to obtain the filled multimodal time series data of the patient; S23, setting an invalid data threshold; calculating the sliding Zscore of the tumor marker detection value in the padded patient multimodal time series data to obtain a sliding Zscore set; taking the value in the sliding Zscore set where the sliding Zscore is greater than the invalid data threshold as an invalid sliding Zscore to obtain an invalid sliding Zscore set; The tumor marker detection values in the padded patient multimodal time series data corresponding to the invalid sliding Zscore set are eliminated to obtain the valid patient multimodal time series data.

4. The method for predicting cancer recurrence probability by deep learning based on multimodal time series features according to claim 1, wherein: The S3 includes the following steps: S31. Setting tumor clinical characteristics to obtain a tumor clinical characteristic set; and collecting historical tumor clinical characteristic data based on the tumor clinical characteristic set; S32. Initialize the random forest model and set the feature importance evaluation criterion of the random forest model to the Gini impurity reduction; The random forest model was trained using historical tumor clinical feature data combined with a cross-validation algorithm to obtain a trained random forest classifier. S33. Calculate the sum of purity improvements of each tumor clinical feature in the effective patient multimodal time series data at the time of decision tree node splitting using the trained random forest classifier to obtain a set of improvement sums; S34. Sort all features in the historical tumor clinical feature data in descending order according to the size of the total lift in the total lift set, perform normalization, record the importance scores of the historical tumor clinical features, and obtain a tumor clinical feature importance ranking; Set the correlation threshold and the number of key features; select the key features with the highest importance scores based on the importance ranking of tumor clinical features and the number of key features to obtain the key tumor clinical feature set; Verify the clinical significance of key features in the key tumor clinical feature set, exclude redundant features with correlation greater than the correlation threshold, and obtain a selected static feature set; Based on the selected static feature set, the patient's selected static features are extracted from the effective patient multimodal time series data to obtain the patient's selected static feature data.

5. The method for predicting cancer recurrence probability by deep learning based on multimodal time series features according to claim 1, characterized in that: The S4 comprises the following steps: S41. Collecting the patient's medication time data from the valid patient multimodal time series data to obtain a standardized medication time series; Set drug attenuation parameters and construct a radiotherapy cumulative toxicity model; calculate the residual toxicity of each drug administration at the current time point through the radiotherapy cumulative toxicity model to obtain the real-time toxicity weight of each drug administration; S42. Based on the real-time toxicity weight of each drug administration, the toxicity weights of all effective drug administrations at the same time point are accumulated, and toxicity superposition correction is performed on the combined drug administration scenario to generate a time-toxicity curve to obtain a dynamic toxicity accumulation characteristic curve; Based on the dynamic toxicity accumulation characteristic curve, the actual liver and kidney function test values of the patients are compared, and the drug attenuation parameters are adjusted to make the curve consistent with the clinical toxicity manifestations. The toxicity levels are then divided to obtain a calibrated clinically applicable toxicity curve. S43. Extracting patient radiotherapy dose time series records from valid patient multimodal time series data, arranging the dose data in chronological order of treatment to form a time series; verifying data continuity to obtain a standardized dose time series; The difference between each dose in the standardized dose time series was calculated using the changes in adjacent time periods; dose increases were set as positive values and dose decreases as negative values to obtain a dose change series; S44, setting a change threshold; based on the dose change sequence and the change threshold, converting the change into a treatment intensity adjustment flag, generating a three-state characteristic value, and obtaining a dose change rate characteristic vector; S45, aligning all the patient static feature data, the calibrated clinically available toxicity curve, and the dose change rate feature vector according to the time point to obtain an original time series feature matrix; The dynamic sliding window method is used to perform window slicing and filling operations on the original feature matrix to obtain a four-dimensional time series feature tensor.

6. The method for predicting cancer recurrence probability by deep learning based on multimodal time series features according to claim 1, characterized in that: The S5 comprises the following steps: S51, constructing an LSTM model and setting hyperparameters for training the LSTM model; S52. Collecting four-dimensional feature tensor data of historical cancer patients and label data indicating whether the historical cancer patients have relapsed, to obtain labeled historical four-dimensional time series feature tensor data; S53. Use labeled historical four-dimensional time series feature tensor data to train the LSTM model. During the training process, use a random search algorithm to find the optimal hyperparameters of the LSTM model and obtain the optimal solution. Use the optimal solution as the hyperparameters of the LSTM model to obtain the LSTM cancer recurrence prediction model. S54. Input the four-dimensional time series feature tensor into the LSTM cancer recurrence prediction model to obtain the predicted patient recurrence probability.

7. The method for predicting cancer recurrence probability by deep learning based on multimodal time series features according to claim 6, characterized in that: In the training process in S53, a random search algorithm is used to find the optimal hyperparameters of the LSTM model. Obtaining the optimal solution includes the following steps: S531, set the parameter search space to L , set the maximum number of random searches to M , the number of cross validation folds is E , initialize the optimal hyperparameters; S532. Under the optimal hyperparameters, use the labeled historical four-dimensional time series feature tensor data to train and evaluate the LSTM model to obtain the best prediction accuracy of the LSTM model. d ; S533, define a candidate hyperparameter set; randomly select a hyperparameter from the candidate hyperparameter set f , the training data in the labeled historical four-dimensional time series feature tensor data is divided into K Fold, for each fold k , ; Use except k The training data in all the labeled historical four-dimensional time series feature tensor data outside the fold are used to train the LSTM model from scratch. k The training data evaluation model in the folded labeled historical four-dimensional time series feature tensor data is used to obtain the prediction accuracy. The mean of the prediction accuracy of all folds is calculated to obtain the average prediction accuracy. g ; If the average prediction accuracy g >Best prediction accuracy d , then the updated best prediction accuracy is g , update the optimal hyperparameters to f Otherwise, the original best prediction accuracy and optimal hyperparameters are maintained; S534, repeat S533, and when the maximum number of random searches M is reached, stop the iteration and obtain the optimal solution; S54. Input the four-dimensional time series feature tensor into the LSTM cancer recurrence prediction model to obtain the predicted patient recurrence probability.

8. The method for predicting cancer recurrence probability by deep learning based on multimodal time series features according to claim 1, characterized in that: The S6 comprises the following steps: S61. Set treatment rules for recurrence probability; S62. Treat cancer patients based on the set recurrence probability treatment rules and the predicted recurrence probability of the patients.

9. A deep learning prediction system for cancer recurrence probability based on multimodal time series features, characterized by: A method for predicting cancer recurrence probability through deep learning based on multimodal time series features as described in any one of claims 1 to 8 is implemented, wherein the system includes a data acquisition module, a data processing module, a static feature screening module, a dynamic feature calculation and formatting module, a model building and prediction module, and a treatment decision module.

10. A storage medium, characterized in that: A program is stored thereon, and when the program is executed by the processor, the method for predicting cancer recurrence probability by deep learning based on multimodal time series features as described in any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Radiation-induced toxicity and machine learning

    CN114974609A

  • Medical data sales prediction method and system of hybrid model based on time series

    CN116703455A

  • Prediction model-based breast cancer drug dosage ratio prediction system and method

    CN118571404A

  • Learner cognitive state recognition method and device, electronic equipment and storage medium

    CN119202565A

  • Bladder cancer postoperative recurrence risk prediction system

    CN119541870A

Cited By

  • Medical health care patient portrait typing improvement method and related device

    CN121964159A

  • A medical health care patient image typing improvement method and related device

    CN121964159B