Method for predicting side effects of gastrointestinal stromal tumor

By standardizing and feature extraction of sample data of patients with gastrointestinal stromal tumors, and using logistic regression and ARIMA models for training, the problems of insufficient side effects management and lack of individual differentiated treatment plans in GIST treatment were solved, and accurate prediction of side effects and dynamic monitoring of patient conditions were achieved.

CN119943437APending Publication Date: 2025-05-06HANGZHOU FIRST PEOPLES HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411787894.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-06
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

There are problems in the existing gastrointestinal stromal tumor (GIST) treatment with insufficient side effects management, lack of dynamic monitoring and prediction capabilities, serious data silos, lack of individual differentiated treatment plans, and insufficient application of traditional Chinese medicine therapy.

Method used

By obtaining sample data from multiple patients, standardized processing and feature extraction, comprehensive feature vectors are calculated, and training is achieved using logistic regression model and ARIMA model to accurately predict GIST side effects and dynamic monitoring of patient condition.

Benefits of technology

It improves the ability to predict and prevent GIST side effects, realizes dynamic monitoring and predicts patients' future conditions, enhances individual differences and flexibility of treatment plans, and integrates traditional Chinese medicine therapy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943437A_ABST
    Figure CN119943437A_ABST
Patent Text Reader

Abstract

The invention relates to a side effect prediction method for gastrointestinal stromal tumor, and the method comprises the steps: obtaining the sample data of a plurality of patients, carrying out the data integration of the obtained sample data of a plurality of patients, carrying out the standardization of the format of each sample data, and carrying out the prediction of the gastrointestinal stromal tumor. Then, feature vector calculation, time sequence feature vector calculation and individual difference sign vector calculation are carried out on each piece of standardized sample data; fusing the feature vector, the time sequence feature vector and the individual difference sign vector in each piece of standardized sample data to obtain a comprehensive feature vector of each piece of standardized sample data; and inputting the plurality of comprehensive feature vectors into a logistic regression model to train the logistic regression model, so that the trained logistic regression model has the function of accurately predicting the side effect probability, and training the plurality of comprehensive feature vectors for an ARIMA model to obtain the trained ARIMA model with the function of predicting the future illness state of the patient.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of gastrointestinal stromal tumors, and in particular to a method for predicting side effects of gastrointestinal stromal tumors. Background Art

[0002] Gastrointestinal stromal tumor (GIST) refers to a type of tumor that originates from the mesenchymal tissue of the gastrointestinal tract, accounting for the majority of digestive tract mesenchymal tumors. Mazur et al. first proposed the concept of gastrointestinal stromal tumor in 1983. GIST is similar to the interstitial cells of Cajal (ICC) around the gastrointestinal myenteric plexus, and both have positive expressions of c-kit gene, CD117 (tyrosine kinase receptor), and CD34 (bone marrow stem cell antigen).

[0003] The existing treatment of gastrointestinal stromal tumors (GIST) has defects, which are mainly reflected in the following aspects:

[0004] 1. Inadequate side effect management: Although molecular targeted drugs such as imatinib and sunitinib have made significant progress in the treatment of GIST, the side effects of these drugs may affect patients' quality of life and treatment compliance. Each person is different, and there is a lack of effective personalized prediction and prevention methods.

[0005] 2. The existing system lacks dynamic monitoring and prediction capabilities: it is unable to track patients’ treatment progress and drug responses in real time, resulting in the inability to adjust treatment plans in a timely manner and prevent the occurrence of side effects.

[0006] 3. The data island phenomenon is serious: The current system relies on a single data source, such as electronic medical records or genetic test results, and is unable to fully integrate and analyze multi-source data, resulting in incomplete and inaccurate diagnosis and treatment information.

[0007] 4. Lack of individual treatment options: Existing treatment options are usually based on standardized treatment guidelines, lack consideration of individual differences, and are difficult to meet the specific needs of different patients, especially in terms of drug side effect management, which lacks flexibility and real-time performance.

[0008] 5. Insufficient application of traditional Chinese medicine (TCM) therapies: Although TCM has unique advantages in alleviating drug side effects and improving patients’ overall health, it has not been fully integrated and utilized, and it is impossible to provide customized TCM treatment plans based on the specific conditions of patients. Summary of the invention

[0009] Based on this, it is necessary to provide a side effect prediction method for gastrointestinal stromal tumors to address the problem of lack of monitoring and prevention of side effects during the treatment of traditional gastrointestinal stromal tumors.

[0010] The present application provides a method for predicting side effects of gastrointestinal stromal tumors, comprising:

[0011] Acquire sample data of a plurality of samples, wherein the sample data includes drug use information of sample patients, gene information of sample patients, imaging information of sample patients, clinical information of sample patients, treatment plan information of sample patients, traditional Chinese medicine information of sample patients, and side effect information of sample patients;

[0012] Performing standardization processing on each sample data to obtain standardized sample data of each sample data;

[0013] Perform feature extraction and feature fusion on the drug use information of the sample patient, the gene information of the sample patient, the imaging information of the sample patient, the clinical information of the sample patient, the treatment plan information of the sample patient, the traditional Chinese medicine information of the sample patient, and the side effect information of the sample patient in each standardized sample data to obtain a feature vector of each standardized sample data;

[0014] Extracting time series features from the drug use information, gene information, imaging information, clinical information, treatment plan information, traditional Chinese medicine information, and side effect information of the sample patient in each standardized sample data to obtain a time series feature vector for each standardized sample data;

[0015] Extract individual difference features of drug use information, gene information, imaging information, clinical information, treatment plan information, TCM information and side effect information of sample patients in each standardized sample data, and obtain individual difference feature vectors of each standardized sample data;

[0016] Based on the feature vector of each standardized sample data, the time series feature vector of each standardized sample data and the individual difference feature vector of each standardized sample data, a comprehensive feature vector of each standardized sample data is calculated;

[0017] Create a logistic regression model;

[0018] The comprehensive feature vector of each standardized sample data is used as training data to train the logistic regression model to obtain a trained logistic regression model;

[0019] Create an ARIMA model;

[0020] The time series feature vector of each standardized sample data is used as training data to train the ARIMA model to obtain the trained ARIMA model.

[0021] The present application relates to a method for predicting side effects of gastrointestinal stromal tumors, which comprises obtaining sample data of multiple patients, integrating and processing the obtained sample data of the multiple patients, standardizing the format of each sample data, and then performing feature vector calculation, time series feature vector calculation and individual difference sign vector calculation on each standardized sample data, and then fusing the feature vector, time series feature vector and individual difference sign vector in each standardized sample data to obtain a comprehensive feature vector of each standardized sample data; inputting multiple comprehensive feature vectors into a logistic regression model to train the logistic regression model, so that the trained logistic regression model has the function of accurately predicting the probability of side effects, and using multiple comprehensive feature vectors to train an ARIMA model, so as to obtain an ARIMA model with the function of predicting the patient's future condition after training. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 A flowchart of a method for predicting side effects of gastrointestinal stromal tumors provided in one embodiment of the present application. DETAILED DESCRIPTION

[0023] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0024] like Figure 1 As shown, in one embodiment of the present application, the method for predicting side effects of gastrointestinal stromal tumors includes the following S001 to S010:

[0025] S001, obtaining sample data of a plurality of samples, wherein the sample data includes drug use information of sample patients, gene information of sample patients, imaging information of sample patients, clinical information of sample patients, treatment plan information of sample patients, traditional Chinese medicine treatment information of sample patients, and side effect information of sample patients.

[0026] Specifically, the drug usage information of the sample patients includes: information on the drugs used by the patients, such as the type, dosage and usage time of the drugs; these records reflect the specific drug usage of the patients during the treatment process and may contain multiple records, indicating the usage history of different time points, different dosages or different drugs.

[0027] The genetic information of the sample patients includes: the patient's gene sequence, mutation information, and gene expression level, etc.

[0028] The imaging information of sample patients includes: CT, MRI and other imaging examination results, which are usually stored in the form of matrix or tensor, and can reflect the development of the patient's lesions and disease.

[0029] The clinical information of the sample patients includes: basic information of the patients (such as age, gender), medical history, physical examination results and laboratory test data.

[0030] The treatment plan information of the sample patients includes: the treatment plan formulated by the patient, such as the specific treatment method used, the course of treatment, and the follow-up plan.

[0031] The TCM treatment information of the sample patients includes: the TCM treatment plans received by the patients, such as TCM prescriptions, TCM diagnosis results and efficacy records.

[0032] The side effect information of the sample patients includes: treatment effect evaluation, side effect records and patient feedback.

[0033] S002, performing standardization processing on each sample data to obtain standardized sample data for each sample data.

[0034] Specifically, each sample data is standardized to obtain the standardized sample data including:

[0035] Select a sample data;

[0036] Obtain drug use information, gene information, imaging information, clinical information, treatment plan information, traditional Chinese medicine information, and side effect information of the sample patient of the sample data;

[0037] The obtained drug use information of the sample patients, the genetic information of the sample patients, the imaging information of the sample patients, the clinical information of the sample patients, the treatment plan information of the sample patients, the traditional Chinese medicine information of the sample patients, and the side effect information of the sample patients are all converted based on the same data specification to obtain standardized drug use information of the sample patients, standardized genetic information of the sample patients, standardized imaging information of the sample patients, standardized clinical information of the sample patients, standardized treatment plan information of the sample patients, standardized traditional Chinese medicine information of the sample patients, and standardized side effect information of the sample patients;

[0038] Return to the step of selecting a sample data until each sample data has been selected once.

[0039] S003, extracting and fusing features of the sample patient's drug use information, the sample patient's genetic information, the sample patient's imaging information, the sample patient's clinical information, the sample patient's treatment plan information, the sample patient's traditional Chinese medicine treatment information, and the sample patient's side effect information in each standardized sample data, and obtaining a feature vector for each standardized sample data.

[0040] S004, extracting time series features from the sample patient's drug use information, the sample patient's genetic information, the sample patient's imaging information, the sample patient's clinical information, the sample patient's treatment plan information, the sample patient's traditional Chinese medicine information, and the sample patient's side effect information in each standardized sample data, and obtaining a time series feature vector for each standardized sample data.

[0041] S005, extract individual difference features of the sample patient's drug use information, the sample patient's genetic information, the sample patient's imaging information, the sample patient's clinical information, the sample patient's treatment plan information, the sample patient's traditional Chinese medicine treatment information and the sample patient's side effect information in each standardized sample data, and obtain the individual difference feature vector of each standardized sample data.

[0042] S006, based on the feature vector of each standardized sample data, the time series feature vector of each standardized sample data and the individual difference feature vector of each standardized sample data, a comprehensive feature vector of each standardized sample data is calculated.

[0043] S007, create a logistic regression model.

[0044] S008, using the comprehensive feature vector of each standardized sample data as training data to train the logistic regression model, to obtain a trained logistic regression model.

[0045] S009, create an ARIMA model.

[0046] S010, using the time series feature vector of each standardized sample data as training data to train the ARIMA model, and obtaining a trained ARIMA model.

[0047] In this embodiment, sample data of multiple patients are first obtained, and the obtained sample data of the multiple patients are integrated and processed together, the format of each sample data is standardized, and then feature vector calculation, time series feature vector calculation and individual difference sign vector calculation are performed on each standardized sample data respectively, and then the feature vector, time series feature vector and individual difference sign vector in each standardized sample data are fused to obtain a comprehensive feature vector of each standardized sample data; multiple comprehensive feature vectors are input into a logistic regression model to train the logistic regression model, so that the trained logistic regression model has the function of accurately predicting the probability of side effects, and multiple comprehensive feature vectors are used to train the ARIMA model to obtain the trained ARIMA model with the function of predicting the patient's future condition.

[0048] In one embodiment of the present application, the drug use information of the sample patient, the gene information of the sample patient, the imaging information of the sample patient, the clinical information of the sample patient, the treatment plan information of the sample patient, the traditional Chinese medicine information of the sample patient, and the side effect information of the sample patient in each standardized sample data are subjected to feature extraction and feature fusion to obtain a feature vector of each standardized sample data, including the following S003a to S003e:

[0049] S003a, select a standardized sample data.

[0050] S003b, obtaining the sample patient's drug use information, the sample patient's genetic information, the sample patient's imaging information, the sample patient's clinical information, the sample patient's treatment plan information, the sample patient's traditional Chinese medicine treatment information, and the sample patient's side effect information in the standardized sample data.

[0051] S003c, feature extraction is performed on the sample patient's drug usage information, the sample patient's genetic information, the sample patient's imaging information, the sample patient's clinical information, the sample patient's treatment plan information, the sample patient's traditional Chinese medicine information, and the sample patient's side effect information to obtain the sample patient's drug usage information features, the sample patient's genetic information features, the sample patient's imaging information features, the sample patient's clinical information features, the sample patient's treatment plan information features, the sample patient's traditional Chinese medicine information features, and the sample patient's side effect information features.

[0052] S003d, the obtained drug use information characteristics of the sample patients, the gene information characteristics of the sample patients, the imaging information characteristics of the sample patients, the clinical information characteristics of the sample patients, the treatment plan information characteristics of the sample patients, the traditional Chinese medicine information characteristics of the sample patients and the side effect information characteristics of the sample patients are integrated to obtain the feature vector of the standardized sample data.

[0053] S003e, returning to the step of selecting a standardized sample data, until each standardized sample data has been selected once.

[0054] The obtained sample patient's drug use information features, sample patient's gene information features, sample patient's imaging information features, sample patient's clinical information features, sample patient's treatment plan information features, sample patient's traditional Chinese medicine information features and sample patient's side effect information features are fused to obtain the feature vector of the standardized sample data, including the following S013d:

[0055] S013d, creating a fusion formula for sample patient's drug use information features, sample patient's gene information features, sample patient's imaging information features, sample patient's clinical information features, sample patient's treatment plan information features, sample patient's traditional Chinese medicine information features and sample patient's side effect information features as shown in Formula 1;

[0056] Z=W1X+W2G+W3I+W4C+W5T+W6M+W7R+b Formula 1;

[0057] Among them, Z is the characteristic vector of the standardized sample data obtained by fusion; X is the drug use information feature of the sample patient; G is the gene information feature of the sample patient; I is the image information feature of the sample patient; C is the clinical information feature of the sample patient; T is the treatment plan information feature of the sample patient; M is the traditional Chinese medicine information feature of the sample patient; R is the side effect information feature of the sample patient; W1 is the weight matrix of the drug use information feature of the sample patient; W2 is the weight matrix of the gene information feature of the sample patient; W3 is the weight matrix of the image information feature of the sample patient; W4 is the weight matrix of the clinical information feature of the sample patient; W5 is the weight matrix of the treatment plan information feature of the sample patient; W6 is the weight matrix of the traditional Chinese medicine information feature of the sample patient; W7 is the weight matrix of the side effect information feature of the sample patient; b is the bias constant.

[0058] Specifically, the weight matrix W is usually learned through model training. For example, the value of W is continuously adjusted through an optimization algorithm (such as gradient descent) to minimize the loss function of the model. This optimization process enables W to accurately capture the most valuable features in multidimensional data, thereby extracting more accurate feature vectors.

[0059] Bias b is a constant term used to adjust the fused feature vector Z so that the data distribution is more consistent with the model requirements. Usually, bias b is automatically adjusted through the model training process. Specifically, during the training process, bias b is learned together with the weight matrix W through gradient descent or other optimization methods to minimize the loss function and ensure that Z can best represent the data features.

[0060] W is to extract the eigenvector with the largest amount of information by performing eigenvalue decomposition on the covariance matrix C of the comprehensive data matrix D. First, construct the data matrix D = [X, G, I, C, T, M]. Then calculate the covariance matrix:

[0061]

[0062] Next, perform eigenvalue decomposition on the covariance matrix C to obtain the eigenvalues ​​and corresponding eigenvectors:

[0063]

[0064] Among them, λ is the eigenvalue, υ is the eigenvector. Select the eigenvectors υ1, υ2, ..., υ corresponding to the first k largest eigenvalues k , and use them as the weight matrix W of different data sources, namely:

[0065] W=[υ1,υ2,....,υ k ].

[0066] In this embodiment, feature extraction is first performed on the sample patients' drug usage information, the sample patients' genetic information, the sample patients' imaging information, the sample patients' clinical information, the sample patients' treatment plan information, the sample patients' traditional Chinese medicine information and the sample patients' side effect information in the standardized sample data, and then the obtained sample patients' drug usage information features, sample patients' genetic information features, sample patients' imaging information features, sample patients' clinical information features, sample patients' treatment plan information features, sample patients' traditional Chinese medicine information features and sample patients' side effect information features are fused to obtain a feature vector of the standardized sample data.

[0067] In one embodiment of the present application, the drug use information of the sample patient, the gene information of the sample patient, the imaging information of the sample patient, the clinical information of the sample patient, the treatment plan information of the sample patient, the traditional Chinese medicine information of the sample patient, and the side effect information of the sample patient in each standardized sample data are extracted to obtain a time series feature vector of each standardized sample data, including the following S004a to S004g:

[0068] S004a, select a standardized sample data.

[0069] S004b, obtaining the sample patient's drug use information, the sample patient's genetic information, the sample patient's imaging information, the sample patient's clinical information, the sample patient's treatment plan information, the sample patient's traditional Chinese medicine treatment information, and the sample patient's side effect information in the standardized sample data.

[0070] S004c, select a moment.

[0071] S004d, analyze the moment and obtain at least one time series feature of the sample patient's drug use information features, the sample patient's gene information features, the sample patient's imaging information features, the sample patient's clinical information features, the sample patient's treatment plan information features, the sample patient's traditional Chinese medicine information features, and the sample patient's side effect information features at the moment.

[0072] S004e, all the obtained time series features are integrated to obtain the time series feature vector of the standardized sample data at that moment.

[0073] S004f, returning to the step of selecting a moment until each moment has been selected once.

[0074] S004g, returning to the step of selecting a standardized sample data until each standardized sample data has been selected once.

[0075] The step of fusing all the obtained time series features to obtain the time series feature vector of the standardized sample data at the moment includes the following S014e:

[0076] S014e, create a fusion formula for all time series features as shown in Formula 2;

[0077] Z t =W t X t +W t G t +W t I t +W t C t +W t T t +W t M t +W t R t +b t

[0078] Formula 2;

[0079] Among them, Z tis the time series feature vector of the standardized sample data obtained by fusion at time t; X t is the drug use information feature of the sample patients at time t; G t is the genetic information feature of the sample patient at time t; I t is the image information feature of the sample patient at time t; C t is the clinical information characteristics of the sample patients at time t; T t is the treatment plan information feature of the sample patient at time t; M t is the information feature of TCM therapy for sample patients at time t; R t is the side effect information feature of the sample patients at time t; W t1 is the weight matrix of the drug use information features of the sample patients at time t; W t2 is the weight matrix of the gene information features of the sample patients at time t; W t3 is the weight matrix of the image information features of the sample patient at time t; W t4 is the weight matrix of the clinical information features of the sample patients at time t; W t5 is the weight matrix of the treatment plan information features of the sample patients at time t; W t6 is the weight matrix of the TCM information features of the sample patients at time t; W t7 is the weight matrix of the side effect information features of the sample patients at time t; b t is the bias constant at time t.

[0080] Specifically, Z t is the feature vector extracted at time t, representing the characteristics of the multi-source data of the patient at that time point. Assuming that there are records from different data sources (such as X, G, I) at time t, then Z t is the feature vector obtained by weighting the information of these data sources at time t. Furthermore, only multi-source data containing timestamps will participate in Z t If some data sources (such as gene data) do not have timestamps, they will not be used in time series feature extraction and will only be used in the static feature extraction stage. The final Z t It is a vector containing the features of multi-source data at that moment.

[0081] W t By constructing a temporal autoencoder model, the encoder weight matrix (W{enc}) and the decoder weight matrix (W{dec}) are the parameters to be learned, and the input time series data (X t ) passes through the encoder and decoder, using the reconstruction error as the loss function:

[0082]

[0083] Where n is the number of time nodes; is the reconstruction error, used to calculate the input X t and the reconstructed X t The squared error measures the accuracy of the model in reconstructing the original data. The smaller the error, the better the reconstruction effect.

[0084] The weight matrices (W{enc}), (W{dec}) and biases (b{enc}), (b{dec}) of the encoder and decoder are adjusted through the back-propagation algorithm to minimize the loss function. Finally, the encoder weight matrix (W{enc}) is used to extract time series features:

[0085] Z t =W enc ·X t +b enc

[0086] The bias b{enc} is a parameter optimized during autoencoder training and is used to adjust the distribution of feature vectors. It is continuously adjusted through the back-propagation algorithm to minimize the loss function.

[0087] The bias b{dec} is used in the decoder stage to adjust the reconstructed output; therefore, it does not appear in the final formula during the feature extraction stage.

[0088] The back propagation algorithm process is:

[0089] 1. Forward propagation: Transform the time series data X t Input encoder to get encoding feature X t .

[0090] 2. Decoding and reconstruction: encoding feature X t Input decoder, reconstruct X t .

[0091] 3. Calculate the loss: By reconstructing the error (i.e. ) Calculate the loss function L.

[0092] 4. Back propagation: According to the gradient of the loss function, update the weights W{enc} and W{dec} of the encoder and decoder and the biases b{enc} and b{dec}.

[0093] Iteration and termination conditions: The back-propagation process minimizes the loss function through iterations. Training is terminated when the loss function value reaches a preset threshold or the number of iterations reaches an upper limit.

[0094] In one embodiment of the present application, the individual difference feature extraction is performed on the drug use information of the sample patient, the gene information of the sample patient, the imaging information of the sample patient, the clinical information of the sample patient, the treatment plan information of the sample patient, the traditional Chinese medicine information of the sample patient, and the side effect information of the sample patient in each standardized sample data to obtain the individual difference feature vector of each standardized sample data, including the following S005a to S005d:

[0095] S005a, select a standardized sample data.

[0096] S005b, obtaining the drug use information characteristics of the sample patients, the gene information characteristics of the sample patients, the imaging information characteristics of the sample patients, the clinical information characteristics of the sample patients, the treatment plan information characteristics of the sample patients, the traditional Chinese medicine treatment information characteristics of the sample patients and the side effect information characteristics of the sample patients in the standardized sample data.

[0097] S005c, based on the attention mechanism, extract and fuse the obtained drug use information features of the sample patients, the genetic information features of the sample patients, the imaging information features of the sample patients, the clinical information features of the sample patients, the treatment plan information features of the sample patients, the traditional Chinese medicine information features of the sample patients and the side effect information features of the sample patients to obtain the individual difference feature vector of the standardized sample data.

[0098] S005d, returning to the step of selecting a standardized sample data, until each standardized sample data has been selected once.

[0099] The attention mechanism is used to extract and fuse the obtained drug use information features of the sample patients, the gene information features of the sample patients, the imaging information features of the sample patients, the clinical information features of the sample patients, the treatment plan information features of the sample patients, the traditional Chinese medicine information features of the sample patients, and the side effect information features of the sample patients to obtain the individual difference feature vector of the standardized sample data, including the following S015c:

[0100] S015c, creating a fusion formula based on the attention mechanism for the obtained sample patient's drug use information features, sample patient's gene information features, sample patient's imaging information features, sample patient's clinical information features, sample patient's treatment plan information features, sample patient's traditional Chinese medicine information features and sample patient's side effect information features, as shown in Formula 3;

[0101] Z i =α i (W i X i +b i ) Formula 3;

[0102] Among them, Z i is the individual difference feature vector fused in the standardized sample data of patient i; α i X i is the input feature of patient i; W i is the weight matrix of the input features of patient i; b i is the bias constant.

[0103] Specifically, individual difference features are mainly used to reflect the different characteristics of each patient and are therefore patient-specific. The weights of feature extraction are dynamically adjusted through the attention mechanism and personalized neural network (PNN), so as to give customized feature representation according to the specific situation of the patient.

[0104] X i Individualized elements for patients can be included, such as personalized dosage or specific treatment regimens in medication use records.

[0105] W i Refers to the weight matrix in LSTM, which is used to process time series data; LSTM uses different weight matrices (such as Wi, Wf, Wo, Wc, etc.) to capture long-term dependencies in time series and generate time series features. These weight matrices are continuously updated through the back-propagation algorithm and are ultimately used to extract the patient's time series features.

[0106] In the LSTM model, i, f, o, and c represent different gating units, and their functions are as follows:

[0107] i (input gate): controls the proportion of the current input information being written.

[0108] f(Forget gate): controls the proportion of information retained in the previous state.

[0109] o (output gate): controls the output of the current time step.

[0110] c(candidate state): calculate new state information.

[0111] These gating mechanisms help LSTM maintain and adjust the information transfer and forgetting ratio between different time steps in the time series, so as to better capture the long-term dependencies in the time series.

[0112] In this embodiment, individual difference feature extraction focuses on individual differences of patients, dynamically adjusts feature weights through attention mechanism and personalized neural network (PNN), and generates personalized features.

[0113] The weight matrix of the LSTM layer is used to process time series features, focusing on the time dependency of data, and personalized feature extraction are two complementary processes.

[0114] In one embodiment of the present application, the comprehensive feature vector of each standardized sample data is calculated based on the feature vector of each standardized sample data, the time series feature vector of each standardized sample data, and the individual difference feature vector of each standardized sample data, including the following S006a:

[0115] S006a, creating a fusion formula based on the feature vector of each standardized sample data, the time series feature vector of each standardized sample data, and the individual difference feature vector of each standardized sample data, such as Formula 4;

[0116] Z final =W final Z+W time Z t +W personal Z i +b final Formula 4;

[0117] Among them, Z final is the comprehensive feature vector of the standardized sample data; Z is the feature vector of the standardized sample data; Z t is the time series feature vector of the standardized sample data; Z i is the individual difference feature vector of the standardized sample data; W final is the weight matrix of the eigenvector of the standardized sample data; W time is the weight matrix of the time series feature vector of the standardized sample data; W personal is the weight matrix of the individual difference feature vector of the standardized sample data; b final is the bias constant of the comprehensive eigenvector.

[0118] In this embodiment, W final It is used to balance the importance of different static data sources (such as genes, demographic information, basic clinical data, etc.) in the final features. It can also give appropriate weights to static features in the comprehensive feature vector to ensure that the patient's static data has an impact on the final decision of the model; W time It is used to capture the dynamic trend of patients, such as the time series changes of treatment effects, and ensure that the time series data occupies an appropriate proportion in the comprehensive feature vector, so that the model can reflect the dynamic trend of the patient's condition; W personal It can be used to improve adaptability to different patients and allow the model to capture the unique characteristics of each patient, ensuring that personalized features can contribute appropriate weight to the final prediction.

[0119] In one embodiment of the present application, the comprehensive feature vector of each standardized sample data is used as training data to train the logistic regression model to obtain the trained logistic regression model, including the following S008a:

[0120] S008a, create a logistic regression model formula as shown in Formula 5;

[0121]

[0122] Where P(Y=1|Z final ) is the probability of side effects; Z1, Z2...Z n are different feature combinations; β is the regression coefficient.

[0123] Specifically, in order to evaluate the side effect risk of each treatment regimen, Z needs to be further subdivided into feature vectors related to specific treatment regimens. In other words, Z represents the comprehensive features of all diagnostic and treatment elements, but when predicting the side effect risk, sub-features related to each specific treatment regimen can be extracted.

[0124] For example, suppose there are multiple treatment options Z1, Z2, ... Z n ,in:

[0125] Z1 represents the first treatment option (such as the dosage of a specific drug, biomarker levels, and other related characteristics).

[0126] Z2 represents the characteristics of the second treatment regimen (e.g., using a different dose of the drug or combining it with other therapies).

[0127] In this embodiment, in order to more accurately explore the side effects of each treatment regimen, the comprehensive feature Z can be split into feature vectors Z of different treatment regimens. i Then for each Z i The logistic regression model was applied to calculate the corresponding risk of side effects, thus helping medical staff to evaluate and select appropriate treatment options.

[0128] In one embodiment of the present application, the time series feature vector of each standardized sample data is used as training data to train the ARIMA model to obtain the trained ARIMA model, including the following S010a:

[0129] S010a, create the ARIMA model formula as shown in Formula 6;

[0130] Y t =φ1Y t-1 +φ2Y t-2 +...+φ p Y t-p +θ1∈t-1 +θ2∈ t-2 +...+θ q ∈ t-q +∈ t Formula 6;

[0131] Among them, Y t is the disease status at time t; φ i is the autoregression coefficient; θ i is the moving average coefficient; ∈ t is the error of the current time node.

[0132] Specifically, Y t Indicates the condition at time t, which can be expressed as a specific indicator (such as biomarker value, blood pressure, body temperature, etc.). If the model needs to integrate multiple features, the feature vector Z can be used t As a comprehensive expression of the disease state, or from Z t Extract a specific dimension from

[0133] φ i Indicates the past condition Y t-1 Current medical condition Y t The impact of

[0134] θ i represents the past error term ∈ t-1 Current medical condition Y t The impact of

[0135] ∈ t Represents the difference between the model prediction value and the true value. During the model training process, the error term ∈ t is the actual condition Y at the current moment. t and the model predicted value Y t It is calculated by the difference between , usually expressed as:

[0136] ∈ t = Actual condition status Y t -Model predicted value Y t

[0137] In this embodiment, the error term is the part of the time series model used to adjust future predictions. By including the moving average part of the past error terms, the prediction is adjusted to improve the accuracy of the prediction.

[0138] The technical features of the above-described embodiments may be arbitrarily combined, and the execution order of the method steps is not limited. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0139] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A method for predicting side effects of gastrointestinal stromal tumors, characterized in that: The method for predicting side effects of gastrointestinal stromal tumors comprises: Acquire sample data of a plurality of samples, wherein the sample data includes drug use information of sample patients, gene information of sample patients, imaging information of sample patients, clinical information of sample patients, treatment plan information of sample patients, traditional Chinese medicine information of sample patients, and side effect information of sample patients; Performing standardization processing on each sample data to obtain standardized sample data of each sample data; Perform feature extraction and feature fusion on the drug use information of the sample patient, the gene information of the sample patient, the imaging information of the sample patient, the clinical information of the sample patient, the treatment plan information of the sample patient, the traditional Chinese medicine information of the sample patient, and the side effect information of the sample patient in each standardized sample data to obtain a feature vector of each standardized sample data; Extracting time series features from the drug use information, gene information, imaging information, clinical information, treatment plan information, traditional Chinese medicine information, and side effect information of the sample patient in each standardized sample data to obtain a time series feature vector for each standardized sample data; Extract individual difference features of drug use information, gene information, imaging information, clinical information, treatment plan information, traditional Chinese medicine information and side effect information of sample patients in each standardized sample data, and obtain individual difference feature vectors of each standardized sample data; Based on the feature vector of each standardized sample data, the time series feature vector of each standardized sample data and the individual difference feature vector of each standardized sample data, a comprehensive feature vector of each standardized sample data is calculated; Create a logistic regression model; The comprehensive feature vector of each standardized sample data is used as training data to train the logistic regression model to obtain a trained logistic regression model; Create an ARIMA model; The time series feature vector of each standardized sample data is used as training data to train the ARIMA model to obtain the trained ARIMA model.

2. The method for predicting side effects of gastrointestinal stromal tumors according to claim 1, characterized in that: The feature vector of each standardized sample data is obtained by extracting and fusing the drug use information of the sample patient, the gene information of the sample patient, the imaging information of the sample patient, the clinical information of the sample patient, the treatment plan information of the sample patient, the traditional Chinese medicine information of the sample patient, and the side effect information of the sample patient, including: Select a standardized sample data; Obtaining drug use information, gene information, imaging information, clinical information, treatment plan information, traditional Chinese medicine information, and side effect information of the sample patient from the standardized sample data; Feature extraction is performed on the sample patient's drug use information, the sample patient's genetic information, the sample patient's imaging information, the sample patient's clinical information, the sample patient's treatment plan information, the sample patient's traditional Chinese medicine information, and the sample patient's side effect information to obtain the sample patient's drug use information features, the sample patient's genetic information features, the sample patient's imaging information features, the sample patient's clinical information features, the sample patient's treatment plan information features, the sample patient's traditional Chinese medicine information features, and the sample patient's side effect information features; The obtained drug use information features of the sample patients, the gene information features of the sample patients, the imaging information features of the sample patients, the clinical information features of the sample patients, the treatment plan information features of the sample patients, the traditional Chinese medicine information features of the sample patients, and the side effect information features of the sample patients are integrated to obtain the feature vector of the standardized sample data; The step of selecting a standardized sample data is returned until each standardized sample data is selected once.

3. The method for predicting side effects of gastrointestinal stromal tumors according to claim 2, characterized in that: The obtained sample patient's drug use information features, sample patient's gene information features, sample patient's imaging information features, sample patient's clinical information features, sample patient's treatment plan information features, sample patient's traditional Chinese medicine information features and sample patient's side effect information features are integrated to obtain the feature vector of the standardized sample data, including: Create a fusion formula for the sample patient's drug use information features, the sample patient's gene information features, the sample patient's imaging information features, the sample patient's clinical information features, the sample patient's treatment plan information features, the sample patient's traditional Chinese medicine information features, and the sample patient's side effect information features, as shown in Formula 1; Z=W1X+W2G+W3I+W4C+W5T+W6M+W7R+b Formula 1; Among them, Z is the characteristic vector of the standardized sample data obtained by fusion; X is the drug use information feature of the sample patient; G is the gene information feature of the sample patient; I is the image information feature of the sample patient; C is the clinical information feature of the sample patient; T is the treatment plan information feature of the sample patient; M is the traditional Chinese medicine information feature of the sample patient; R is the side effect information feature of the sample patient; W1 is the weight matrix of the drug use information feature of the sample patient; W2 is the weight matrix of the gene information feature of the sample patient; W3 is the weight matrix of the image information feature of the sample patient; W4 is the weight matrix of the clinical information feature of the sample patient; W5 is the weight matrix of the treatment plan information feature of the sample patient; W6 is the weight matrix of the traditional Chinese medicine information feature of the sample patient; W7 is the weight matrix of the side effect information feature of the sample patient; b is the bias constant.

4. The method for predicting side effects of gastrointestinal stromal tumors according to claim 3, characterized in that: The time series feature extraction is performed on the drug use information of the sample patient, the gene information of the sample patient, the imaging information of the sample patient, the clinical information of the sample patient, the treatment plan information of the sample patient, the traditional Chinese medicine information of the sample patient and the side effect information of the sample patient in each standardized sample data to obtain a time series feature vector of each standardized sample data, including: Select a standardized sample data; Obtaining drug use information, gene information, imaging information, clinical information, treatment plan information, traditional Chinese medicine information, and side effect information of the sample patient from the standardized sample data; Pick a moment; Analyze the moment to obtain at least one time series feature of the sample patient's drug use information feature, the sample patient's gene information feature, the sample patient's imaging information feature, the sample patient's clinical information feature, the sample patient's treatment plan information feature, the sample patient's traditional Chinese medicine information feature, and the sample patient's side effect information feature at the moment; All the obtained time series features are integrated to obtain the time series feature vector of the standardized sample data at that moment; Returning to the step of selecting a moment until each moment has been selected once; The step of selecting a standardized sample data is returned until each standardized sample data is selected once.

5. The method for predicting side effects of gastrointestinal stromal tumors according to claim 4, characterized in that: The step of fusing all the obtained time series features to obtain the time series feature vector of the standardized sample data at that moment includes: Create all time series feature fusion formulas as shown in Formula 2; Z t = W t X t + W t G t + W t I t + W t C t + W t T t + W t M t + W t R t + b t Formula 2; Among them, Z t is the time series feature vector of the standardized sample data obtained by fusion at time t; X t is the drug use information feature of the sample patients at time t; G t is the genetic information feature of the sample patient at time t; I t is the image information feature of the sample patient at time t; C t is the clinical information characteristics of the sample patients at time t; T t is the treatment plan information feature of the sample patient at time t; M t is the information feature of TCM therapy for sample patients at time t; R t is the side effect information feature of the sample patients at time t; W t1 is the weight matrix of the drug use information features of the sample patients at time t; W t2 is the weight matrix of the gene information features of the sample patients at time t; W t3 is the weight matrix of the image information features of the sample patient at time t; W t4 is the weight matrix of the clinical information features of the sample patients at time t; W t5 is the weight matrix of the treatment plan information features of the sample patients at time t; W t6 is the weight matrix of the TCM information features of the sample patients at time t; W t7 is the weight matrix of the side effect information features of the sample patients at time t; b t is the bias constant at time t.

6. The method for predicting side effects of gastrointestinal stromal tumors according to claim 5, characterized in that: The individual difference feature extraction is performed on the drug use information of the sample patient, the gene information of the sample patient, the imaging information of the sample patient, the clinical information of the sample patient, the treatment plan information of the sample patient, the traditional Chinese medicine information of the sample patient and the side effect information of the sample patient in each standardized sample data to obtain the individual difference feature vector of each standardized sample data, including: Select a standardized sample data; Obtaining drug use information characteristics of sample patients, gene information characteristics of sample patients, imaging information characteristics of sample patients, clinical information characteristics of sample patients, treatment plan information characteristics of sample patients, traditional Chinese medicine information characteristics of sample patients, and side effect information characteristics of sample patients in the standardized sample data; Based on the attention mechanism, the obtained drug use information features of the sample patients, the gene information features of the sample patients, the imaging information features of the sample patients, the clinical information features of the sample patients, the treatment plan information features of the sample patients, the traditional Chinese medicine information features of the sample patients, and the side effect information features of the sample patients are extracted and integrated to obtain the individual difference feature vector of the standardized sample data; The step of selecting a standardized sample data is returned until each standardized sample data is selected once.

7. The method for predicting side effects of gastrointestinal stromal tumors according to claim 6, characterized in that: The attention mechanism is used to extract and fuse the obtained drug use information features of the sample patients, the gene information features of the sample patients, the imaging information features of the sample patients, the clinical information features of the sample patients, the treatment plan information features of the sample patients, the traditional Chinese medicine information features of the sample patients, and the side effect information features of the sample patients to obtain the individual difference feature vector of the standardized sample data, including: Create a fusion formula based on the attention mechanism for the obtained sample patient's drug use information features, sample patient's gene information features, sample patient's imaging information features, sample patient's clinical information features, sample patient's treatment plan information features, sample patient's traditional Chinese medicine information features, and sample patient's side effect information features, as shown in Formula 3; Z i =α i (W i X i +b i ) Formula 3; Among them, Z i is the individual difference feature vector fused in the standardized sample data of patient i; α i X i is the input feature of patient i; W i is the weight matrix of the input features of patient i; b i is the bias constant.

8. The method for predicting side effects of gastrointestinal stromal tumors according to claim 7, characterized in that: The method of calculating a comprehensive feature vector of each standardized sample data based on the feature vector of each standardized sample data, the time series feature vector of each standardized sample data and the individual difference feature vector of each standardized sample data includes: Create a fusion formula based on the feature vector of each standardized sample data, the time series feature vector of each standardized sample data, and the individual difference feature vector of each standardized sample data, such as Formula 4; Z final = W final Z + W time Z t + W personal Z i + b final Formula 4; Among them, Z final is the comprehensive feature vector of the standardized sample data; Z is the feature vector of the standardized sample data; Z t is the time series feature vector of the standardized sample data; Z i is the individual difference feature vector of the standardized sample data; W final is the weight matrix of the eigenvector of the standardized sample data; W time is the weight matrix of the time series feature vector of the standardized sample data; W personal is the weight matrix of the individual difference feature vector of the standardized sample data; b final is the bias constant of the comprehensive eigenvector.

9. The method for predicting side effects of gastrointestinal stromal tumors according to claim 8, characterized in that: The method of using the comprehensive feature vector of each standardized sample data as training data to train the logistic regression model to obtain the trained logistic regression model includes: Create a logistic regression model formula as shown in Formula 5; Where P(Y=1|Z final ) is the probability of side effects; Z1, Z2...Z n are different feature combinations; β is the regression coefficient.

10. The method for predicting side effects of gastrointestinal stromal tumors according to claim 9, characterized in that: The method of using the time series feature vector of each standardized sample data as training data to train the ARIMA model to obtain the trained ARIMA model includes: Create the ARIMA model formula as shown in Formula 6; Y t = φ1Y t-1 + φ2Y t-2 +... + φ p Y t-p + θ1∈ t-1 + θ2∈ t-2 +... + θ q ∈ t-q + ∈ t Formula 6; Among them, Y t is the disease status at time t; φ i is the autoregression coefficient; θ i is the moving average coefficient; ∈ t is the error of the current time node.