Immune checkpoint blockade response prediction method and system based on diffusion model
By combining the diffusion model with the ODE solver and feedforward network to learn the multi-dimensional characteristic data of patients, the accuracy and individualization problems of ICB efficacy prediction are solved, and a more accurate ICB treatment response prediction is achieved.
Patent Information
- Application Number
- CN202510902791.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-07-01
AI Technical Summary
Existing ICB efficacy prediction methods have limitations in accuracy and individualized prediction, and it is difficult to effectively capture the latent mapping rules between patient characteristics and treatment response.
A diffusion model-based approach is used to learn the potential representation and intrinsic distribution of patients' multi-dimensional feature data through a combination of ODE solver and feedforward network, and to calculate the probability of patients' response to immune checkpoint blockade therapy.
It achieves more accurate prediction of ICB treatment response, improves the interpretability and accuracy of prediction results, and fills the application gap in low-dimensional data space.
Smart Images

Figure CN120413084B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computational biology and bioinformatics, and in particular to a method and system for predicting immune checkpoint blockade responses based on a diffusion model. Background Art
[0002] Cancer is a malignant disease caused by abnormal cell proliferation with the ability to invade and metastasize, posing a serious threat to human health. Due to its high heterogeneity and dynamic evolution, traditional cancer treatments generally have certain limitations in clinical applications. For example, surgical treatment is mainly suitable for patients with early-stage tumors that have not yet metastasized, and it is difficult to completely eliminate potential micrometastatic lesions. Although radiotherapy and chemotherapy can control the progression of the disease to a certain extent, they are often accompanied by significant toxic side effects, affecting the patient's quality of life. Targeted therapy, while having a certain degree of specificity, is only suitable for patient groups with clear molecular targets, and long-term use can easily lead to drug resistance.
[0003] In recent years, immunotherapy, particularly immune checkpoint blockade (ICB) therapy, has brought new breakthroughs in cancer treatment. This approach inhibits immune checkpoint signaling pathways between tumor cells and immune cells (such as the immune checkpoints PD-1 / PD-L1 and the immunomodulatory protein CTLA-4), thereby activating the host immune system to recognize and kill tumor cells. However, the clinical response rate of ICB therapy is only approximately 20%-30%, and efficacy varies significantly between individuals. Therefore, accurately predicting patient response to ICB therapy has become a key issue to improve treatment efficacy, avoid ineffective treatment and delayed disease progression, and promote the development of personalized precision medicine.
[0004] Traditional methods for predicting ICB efficacy, such as those based on tumor tissue PD-L1 expression levels or tumor mutational burden (TMB), have limitations. These methods often only predict treatment response in a subset of patients and are susceptible to factors such as tissue sampling and testing methods. Therefore, the discovery of more relevant features is needed to aid prediction. Currently, clinical, pathological, and genomic features considered to be associated with ICB efficacy include: programmed cell death 1 (PD-1) and programmed cell death 1 ligand 1 (PD-L1) expression in tumor cells, microsatellite instability (MSI), human leukocyte antigen class I evolutionary divergence (HED), fraction of copy number alteration (FCNA), blood neutrophil–lymphocyte ratio (NLR), blood albumin level, body mass index (BMI), sex, and age.
[0005] Previous studies have attempted to improve predictive accuracy by introducing statistical learning methods (such as logistic regression and support vector machines) and combining multiple clinical features for comprehensive modeling. However, these methods are limited by the expressive power of the model structure and often struggle to fully explore the complex relationships between variables when processing data, nor to capture the underlying mapping patterns between patient characteristics and treatment responses. Summary of the Invention
[0006] Based on the technical problems existing in the background technology, the present invention proposes an immune checkpoint blockade response prediction method and system based on a diffusion model, which can effectively learn the potential representation and intrinsic distribution in the patient's multi-dimensional feature data, thereby achieving a more accurate prediction of the ICB treatment response.
[0007] The present invention proposes a method for predicting immune checkpoint blockade response based on a diffusion model, comprising:
[0008] Obtain the multidimensional indicator features of the patient and splice them to obtain 7-dimensional spliced data, which is then input into the ODE solver;
[0009] The ODE solver inputs the current time step and the current intermediate data generated by the solution into the trained diffusion model, and feeds back the current prediction score obtained into the ODE solver. The ODE solver iteratively inputs the next time step and the intermediate data corresponding to the next time step generated by the solution into the diffusion model until the ordinary differential equation in the ODE solver is solved. The ODE solver calculates the data sample and the integral term result that conforms to the Gaussian distribution. The intermediate data is the noise disturbance feature corresponding to the current time step.
[0010] Calculating the joint probability using the sum of the log-likelihood of the data sample and the integral term result;
[0011] Based on the fact that the patient's efficacy to immune checkpoint blockade therapy is divided into two states: response and non-response, two joint probabilities are obtained, and the ratio of the joint probability in the response state to the sum of the two joint probabilities is used as the patient's response probability to immune checkpoint blockade therapy.
[0012] Furthermore, the multi-dimensional indicator features of the patient are obtained and spliced to obtain 7-dimensional spliced data, specifically:
[0013] Obtain the patient's 6-dimensional indicator features and normalize them as original 6-dimensional data;
[0014] The patient's efficacy on immune checkpoint blockade therapy is preset as the 7th dimension data and spliced to the end of the original 6-dimensional data after normalization to obtain 7-dimensional spliced data.
[0015] Furthermore, in the process of the ODE solver solving the ordinary differential equation, one ODE solver gradually accumulates and calculates the integral term results based on the prediction score feedback from the diffusion model corresponding to each time step; the other ODE solver is used to add noise to the 7-dimensional spliced data at each time step to obtain intermediate data. After all time steps are completed, data samples that conform to the Gaussian distribution are obtained, and the logarithmic likelihood of the data samples is calculated.
[0016] Furthermore, the diffusion model uses a feedforward network as a noise prediction network, and the feedforward network includes two hidden layers connected in series and one output layer;
[0017] The training process of the diffusion model is as follows:
[0018] Obtain the patient's 6-dimensional indicator characteristics and the preset patient's efficacy on immune checkpoint blockade therapy to obtain 7-dimensional spliced data; randomly sample the time step , after adding noise to the 7-dimensional spliced data, it is used as a noisy sample to construct a training dataset, where For uniform distribution, The maximum end point of time;
[0019] Input the training dataset into the first hidden layer of the feedforward network;
[0020] An overlay layer is set between two adjacent hidden layers, and an overlay layer is set between the second hidden layer and the output layer. In the two overlay layers, the corresponding time steps are embedded into the noisy samples through Gaussian Fourier projection;
[0021] Construct a loss function for the diffusion model to adjust the trainable parameters in the diffusion model.
[0022] Furthermore, the time step Gaussian Fourier projection operation for:
[0023] ;
[0024] in,
[0025] ;
[0026] in, are the linear layer network parameters, represents the phase information of the concatenated sine and cosine functions, is a randomly generated Gaussian matrix, For splicing operation.
[0027] Furthermore, after adding noise to the 7-dimensional spliced data, the training data set is constructed as a noisy sample and the time step is calculated. The operations corresponding to the noisy samples are as follows:
[0028] ;
[0029] in,
[0030] ;
[0031] in, Represents the time step The corresponding noise intensity, σ is a hyperparameter, indicating the maximum noise intensity, represents randomly sampled standard Gaussian noise, represents the original sample without noise, Represents the noisy sample after noise addition.
[0032] Furthermore, the loss function of the diffusion model is as follows:
[0033] ;
[0034] in, represents the prediction score of the diffusion model, Represents the noisy sample after noise addition.
[0035] Furthermore, the log-likelihood of the data sample The calculation formula is as follows:
[0036] ;
[0037] in, is the likelihood of a Gaussian distributed data sample, is a hyperparameter, For data samples that conform to the Gaussian distribution, for The second norm of .
[0038] Furthermore, in the joint probability calculated by the sum of the log-likelihood and the integral term of the data sample, when the patient is responsive to the immune checkpoint blockade therapy, the logarithm of the joint probability is The specific calculation formula is:
[0039] ;
[0040] in, for In the response group The joint probability in is the original sample without noise, Time during the diffusion process The noise disturbance characteristics when For patients who respond to immune checkpoint blockade therapy, is the prediction score, is the divergence of the prediction scores, represents the model parameters of the diffusion model after training, and The diffusion start time and diffusion end time are respectively, For the integral arrive Cumulative divergence effect, is a hyperparameter, indicating the maximum noise intensity, is the integration variable.
[0041] An immune checkpoint blockade response prediction system based on a diffusion model, including a data acquisition module, a prediction module, a joint probability module, and a response probability module;
[0042] The data acquisition module is used to obtain the multi-dimensional indicator characteristics of the patient and splice them to obtain 7-dimensional spliced data, which is then input into the ODE solver;
[0043] In the prediction module, the ODE solver inputs the current time step and the current intermediate data generated by the solution into the trained diffusion model, and feeds the obtained current prediction score back to the ODE solver. The ODE solver iteratively inputs the next time step generated by the solution and the intermediate data corresponding to the next time step as the current time step into the diffusion model until the ordinary differential equation in the ODE solver is solved. The ODE solver calculates the data samples and integral term results that conform to the Gaussian distribution. The intermediate data is the noise disturbance feature corresponding to the current time step.
[0044] In the joint probability module, the joint probability is calculated using the sum of the log-likelihood of the data sample and the integral term result;
[0045] In the response probability module, two joint probabilities are obtained based on the patient's response and non-response to immune checkpoint blockade therapy. The ratio of the joint probability in the response state to the sum of the two joint probabilities is used as the patient's response probability to immune checkpoint blockade therapy.
[0046] The advantages of the immune checkpoint blockade response prediction method and system based on the diffusion model provided by the present invention are: it can effectively learn the potential representation and intrinsic distribution in the patient's multi-dimensional feature data, thereby achieving more accurate prediction of ICB treatment response; it uses the diffusion model to generate 7-dimensional medical data and calculate sample conditional probabilities to predict ICB efficacy, filling a gap in the application of diffusion models in low-dimensional data spaces; and it theoretically provides an accurate calculation method for the response probability of predicted efficacy. Compared with other methods that only provide confidence levels for prediction problems, this embodiment has greater advantages in the interpretability of prediction results. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 It is a schematic diagram of the process of the present invention;
[0048] Figure 2 Schematic diagram of the ROC curve and AUC performance corresponding to the prediction results of this embodiment on the test set;
[0049] Figure 3 Schematic diagram of the precision-recall curve and AUPRC performance corresponding to the prediction results of this embodiment on the test set. DETAILED DESCRIPTION
[0050] The technical solutions of the present invention are described in detail below through specific embodiments. Numerous specific details are set forth in the following description to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways than those described herein, and those skilled in the art may make similar modifications without departing from the scope of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0051] like Figures 1 to 3 As shown, the immune checkpoint blockade response prediction method proposed by the present invention based on the diffusion model includes:
[0052] Step 1: Obtain the multidimensional indicator features of the patient and splice them to obtain 7-dimensional spliced data, which is then input into the ODE solver;
[0053] Step 2: The ODE solver inputs the current time step and the current intermediate data generated by the solution into the trained diffusion model, and feeds the obtained current prediction score back to the ODE solver. The ODE solver iteratively inputs the next time step generated by the solution and the intermediate data corresponding to the next time step as the current time step into the diffusion model until the ordinary differential equation in the ODE solver is solved. The ODE solver calculates the data sample and the integral term result that conforms to the Gaussian distribution. The intermediate data is the noise disturbance feature corresponding to the current time step.
[0054] Step 3: Calculate the joint probability using the sum of the log-likelihood of the data sample and the integral term result;
[0055] Step 4: Based on the two states of response and non-response of the patient's efficacy to immune checkpoint blockade therapy, two joint probabilities are obtained, and the ratio of the joint probability in the response state to the sum of the two joint probabilities is used as the patient's response probability to immune checkpoint blockade therapy.
[0056] Given the first six dimensions of vector data (patient indicators), it is desirable to determine the seventh dimension of the data (the patient's response to ICB therapy). This embodiment utilizes a combination of an ODE solver (i.e., a differential equation solver) and a diffusion model, combined with a joint probability generation process, to effectively learn the latent representations and intrinsic distributions of the patient's multidimensional feature data, thereby achieving more accurate predictions of the response to ICB therapy.
[0057] In this embodiment, step 1 is to obtain the multi-dimensional indicator features of the patient and splice them to obtain 7-dimensional spliced data, which is then input into the ODE solver. Specifically,
[0058] Obtain the patient's 6-dimensional indicator features and normalize them as original 6-dimensional data;
[0059] The patient's efficacy on immune checkpoint blockade therapy is preset as the 7th dimension data and spliced to the end of the original 6-dimensional data after normalization to obtain 7-dimensional spliced data.
[0060] Among them, the original six-dimensional data include tumor mutation burden (TMB), patient systemic treatment history (PSTH), blood albumin, blood neutrophil-lymphocyte ratio (NLR), age (age), and cancer type (cancertype).
[0061] This embodiment divides the patient's response to immune checkpoint blockade therapy into two states: "responsive" and "non-responsive". The joint probability of these two states is calculated using the ODE solver and diffusion model process, thereby calculating the patient's response probability to immune checkpoint blockade therapy.
[0062] It should be noted that, in this embodiment, after the ordinary differential equation is solved in the ODE solver, two ODE solvers are set. The solution range of the two ODE solvers is [,0.59]. One ODE solver gradually accumulates the integral term results based on the prediction score of the diffusion model feedback corresponding to each time step. The integral term results are used for the subsequent calculation of the joint probability, that is:
[0063] ;
[0064] in, is the prediction score, is the divergence of the predicted scores, σ is a hyperparameter representing the maximum noise intensity, is the original sample without noise.
[0065] Another ODE solver adds noise to the 7-dimensional spliced data at each time step to obtain intermediate data. After all time steps are completed, data samples that conform to the Gaussian distribution are obtained, and the logarithmic likelihood of the data samples is calculated, that is:
[0066] ;
[0067] in, Time during the diffusion process The noise disturbance characteristics when Represents the parameters of the diffusion model after training.
[0068] Solve the above two-way ODE solver in the time range [0,0.59] to obtain a data sample that conforms to the Gaussian distribution. and the integral term The result of the data sample is calculated according to the following formula The log-likelihood of :
[0069] ;
[0070] in, is the likelihood of a Gaussian distributed data sample, For data samples that conform to the Gaussian distribution, for The second norm of .
[0071] The logarithm of the joint probability is obtained according to the following formula :
[0072] ;
[0073] in, for In the response group The joint probability in is the input data sample, Time during the diffusion process The noise disturbance characteristics when For patients who respond to immune checkpoint blockade therapy, is the prediction score, is the divergence of the prediction scores, represents the model parameters of the diffusion model after training, and are the diffusion start time and diffusion end time, respectively.
[0074] Part 1 The log-likelihood of the data sample corresponding to the Gaussian distribution at the end of the diffusion process. The second part is the integral term , For the integral arrive The cumulative divergence effect, in this embodiment, Preferably 0, Preferably it is 0.59.
[0075] In the above calculation formula of joint probability, As the base likelihood, it is used to evaluate the matching degree between the patient's baseline characteristics and the response group. As a dynamic correction, it is used to quantify the likelihood adjustment of biological characteristics evolution during treatment. For positive adjustment, the feature evolution conforms to the response mode. For negative adjustment, the feature deviates from the response mode; the final joint probability The higher the level, the more likely the patient is to respond to ICB (immune checkpoint blockade) therapy.
[0076] Based on the logarithm of the joint probability , and then get the joint probability :
[0077] .
[0078] because In the non-responding group The joint probability in The calculation formula and The calculation process is consistent with that of , so this embodiment will not be described in detail. and The response probability can be calculated :
[0079] ;
[0080] The response probability is a prediction of the patient's therapeutic effect. Specifically, it means the probability that the patient will "respond" to ICB therapy under the premise that the patient's relevant indicators are known.
[0081] This embodiment transforms the diffusion model from an implicit generative model to an explicit probability density calculation model. The calculation process has a solid theoretical foundation and guaranteed results. Furthermore, the use of an ODE solver allows for faster convergence and controllable accuracy.
[0082] In one embodiment, the training process of the diffusion model is as follows:
[0083] S100, obtain the patient's 6-dimensional indicator characteristics and the preset patient's efficacy on immune checkpoint blockade therapy to obtain 7-dimensional spliced data; random sampling time step , after adding noise to the 7-dimensional spliced data, it is used as a noisy sample to construct a training dataset, where For uniform distribution, The maximum end point of time;
[0084] First, the 6-dimensional indicator features are normalized, and the patient's efficacy on immune checkpoint blockade therapy is used as the 7th dimension data. The 7th dimension data is normalized, and the two normalized features are spliced to obtain 7-dimensional spliced data.
[0085] The normalization process of the 6-dimensional indicator features is consistent with the 7th-dimensional data, and the formula is:
[0086] ;
[0087] in, is the original data, which corresponds to the 6-dimensional indicator features or the 7th dimension data, is the normalized feature, and are the mean and variance of the training set used to train the diffusion model.
[0088] S200, input the training data set into the first hidden layer of the feedforward network;
[0089] An overlay layer is set between two adjacent hidden layers, and an overlay layer is set between the second hidden layer and the output layer. In the two overlay layers, the corresponding time steps are embedded into the noisy samples through Gaussian Fourier projection;
[0090] Time step Gaussian Fourier projection operation for:
[0091] ;
[0092] in,
[0093] ;
[0094] in, are the linear layer network parameters, represents the phase information of the concatenated sine and cosine functions, is a randomly generated Gaussian matrix, For splicing operation.
[0095] After adding noise to the 7-dimensional spliced data, the time step is calculated as the noisy sample to construct the training data set. The operations corresponding to the noisy samples are as follows:
[0096] ;
[0097] in,
[0098] ;
[0099] in, Represents the time step The corresponding noise intensity, σ is a hyperparameter, indicating the maximum noise intensity, represents randomly sampled standard Gaussian noise, , is a normal distribution, is the identity matrix, represents the original sample without noise, Represents the noisy sample after noise addition.
[0100] During diffusion model training, the model essentially learns how to predict noise for noisy samples at different time steps and noise intensities, thereby performing denoising. By embedding time steps in the overlay layer, the diffusion model can more accurately predict the corresponding scores for noisy input samples.
[0101] S300: Construct a loss function of the diffusion model to adjust the trainable parameters in the diffusion model.
[0102] Loss function of the diffusion model as follows:
[0103] ;
[0104] in, represents the prediction score of the diffusion model, Represents the noisy sample after noise addition.
[0105] As an embodiment, the training process of the diffusion model in this embodiment is as follows:
[0106] (a1) Obtain patient data;
[0107] In this embodiment, the diffusion model was trained using the publicly available Chowell_train dataset, which includes 11 indicators of 964 patients. In the testing phase, the diffusion model was used to predict efficacy and evaluate the prediction results on the Chowell_test dataset, which also includes 11 indicators of 515 patients.
[0108] (a2) training the diffusion model;
[0109] In this example, the diffusion model is trained in Python 3.10.14 and torch 2.5.0, and the Adam optimizer is used for optimization. The relevant hyperparameter settings are: epoch is set to 1200, learning rate is set to 0.0001, and maximum noise intensity is set to 0. Set to 25, Epoch refers to the process of inputting the entire training data set into the diffusion model for one forward propagation and back propagation.
[0110] (a3) predicting patient outcomes;
[0111] according to Figure 1 The process extracts six indicators of patients in the test set (tumor mutation burden (TMB), patient systemic treatment history (PSTH), blood albumin, blood neutrophil-lymphocyte ratio (NLR), age, and cancer type) as raw 6-dimensional data; after normalization, the data is spliced with the 7th dimension data to obtain 7-dimensional spliced data; the 7-dimensional spliced data is input into the ODE solver, which solves the ODE solver with the help of the trained diffusion model; the joint probability is calculated using the solution result, and the response probability is calculated as a prediction of the patient's "response" to the ICB efficacy.
[0112] (a4) prediction results;
[0113] To evaluate the prediction results of the diffusion model, this embodiment uses two indicators for evaluation: Area Under the ROC Curve (AUC) and Area Under the Precision-Recall Curve (AUPRC).
[0114] The ROC curve is plotted with the false positive rate (FPR) on the horizontal axis and the true positive rate (TPR) on the vertical axis. The AUC value assesses the overall ability of the diffusion model to distinguish whether a patient responds to treatment. Its value range is [0, 1], with 0.5 representing random classification and 1.0 representing perfect classification.
[0115] A precision-recall curve is plotted with recall on the horizontal axis and precision on the vertical axis. The AUPRC range is [0, 1], with the baseline being the proportion of positive samples. AUPRC reflects the model's ability to identify positive samples in highly imbalanced data.
[0116] Random guessing, a decision tree, and a support vector machine (SVM) are used as comparisons. Random guessing means the classifier has no bias towards all samples, and positive and negative samples are predicted as positive or negative with equal probability. Therefore, in the ROC curve, the true positive rate (TPR) and false positive rate (FPR) are always equal. In the precision-recall curve, the precision remains constant, consistent with the proportion of positive samples in the dataset.
[0117] The datasets used for the decision tree and support vector machine are identical to those used in the model of this embodiment. The implementation relies on the DecisionTreeClassifier function and the support vector machine classifier (SVC function) provided by the sklearn tool library. DecisionTreeClassifier is a function in the scikit tool library used to build decision tree classification models, using the CART algorithm (Classification and Regression Trees).
[0118] according to Figure 2 , it can be seen that the performance indicator AUC of the method of this embodiment (corresponding to Ours in the figure) is significantly better than random guessing, and also outperforms traditional classification methods such as decision trees and support vector machines, verifying the effectiveness of the method of this embodiment. Furthermore, the ROC curve of the method of this embodiment rises significantly in the low false positive rate range, indicating that the method of this embodiment is more sensitive to positive examples.
[0119] according to Figure 3, it can be seen that the performance metric AUPRC of the method of this embodiment (corresponding to Ours in the figure) is significantly better than random guessing, and also outperforms decision trees and support vector machines, verifying the ability of this embodiment to identify positive samples in imbalanced datasets. Furthermore, the precision-recall curve shows high precision in the low recall range, indicating that this embodiment's method is highly reliable in identifying samples that are confirmed to be positive examples.
[0120] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.
Claims
1. A method for predicting immune checkpoint blockade response based on a diffusion model, characterized in that: include: Obtain the multidimensional indicator features of the patient and splice them to obtain 7-dimensional spliced data, which is then input into the ODE solver; The ODE solver inputs the current time step and the current intermediate data generated by the solution into the trained diffusion model, and feeds back the current prediction score obtained into the ODE solver. The ODE solver iteratively inputs the next time step and the intermediate data corresponding to the next time step generated by the solution into the diffusion model until the ordinary differential equation in the ODE solver is solved. The ODE solver calculates the data sample and the integral term result that conforms to the Gaussian distribution. The intermediate data is the noise disturbance feature corresponding to the current time step. Calculating the joint probability using the sum of the log-likelihood of the data sample and the integral term result; Based on the fact that the patient's efficacy to immune checkpoint blockade therapy is divided into two states: response and non-response, two joint probabilities are obtained, and the ratio of the joint probability in the response state to the sum of the two joint probabilities is used as the patient's response probability to immune checkpoint blockade therapy.
2. The response prediction method according to claim 1, characterized in that The multi-dimensional indicator features of the patient are obtained and spliced to obtain 7-dimensional spliced data, specifically: Obtain the patient's 6-dimensional indicator features and normalize them as original 6-dimensional data; The patient's efficacy on immune checkpoint blockade therapy is preset as the 7th dimension data and spliced to the end of the original 6-dimensional data after normalization to obtain 7-dimensional spliced data.
3. The response prediction method according to claim 1, wherein: In the process of solving ordinary differential equations by the ODE solver, one ODE solver gradually accumulates and calculates the integral term results based on the predicted score feedback from the diffusion model corresponding to each time step; the other ODE solver is used to add noise to the 7-dimensional spliced data at each time step to obtain intermediate data. After all time steps are completed, data samples that conform to the Gaussian distribution are obtained, and the logarithmic likelihood of the data samples is calculated.
4. The response prediction method according to claim 1, wherein: The diffusion model uses a feedforward network as a noise prediction network, and the feedforward network includes two hidden layers arranged in series and one output layer; The training process of the diffusion model is as follows: Obtain the patient's 6-dimensional indicator characteristics and the preset patient's efficacy on immune checkpoint blockade therapy to obtain 7-dimensional spliced data; randomly sample the time step , after adding noise to the 7-dimensional spliced data, it is used as a noisy sample to construct a training dataset, where For uniform distribution, The maximum end point of time; Input the training dataset into the first hidden layer of the feedforward network; An overlay layer is set between two adjacent hidden layers, and an overlay layer is set between the second hidden layer and the output layer. In the two overlay layers, the corresponding time steps are embedded into the noisy samples through Gaussian Fourier projection; Construct a loss function for the diffusion model to adjust the trainable parameters in the diffusion model.
5. The response prediction method according to claim 4, characterized in that: Time step Gaussian Fourier projection operation for: in, in, are the linear layer network parameters, represents the phase information of the concatenated sine and cosine functions, is a randomly generated Gaussian matrix, For splicing operation.
6. The response prediction method according to claim 4, characterized in that: After adding noise to the 7-dimensional spliced data, the time step is calculated as the noisy sample to construct the training data set. The operations corresponding to the noisy samples are as follows: in, in, Represents the time step The corresponding noise intensity, σ is a hyperparameter, indicating the maximum noise intensity, represents randomly sampled standard Gaussian noise, represents the original sample without noise, Represents the noisy sample after noise addition.
7. The response prediction method according to claim 6, characterized in that: Loss function of the diffusion model as follows: in, represents the prediction score of the diffusion model, Represents the noisy sample after noise addition.
8. The response prediction method according to claim 1, wherein: The log-likelihood of the data sample The calculation formula is as follows: in, is the likelihood of a Gaussian distributed data sample, is a hyperparameter, For data samples that conform to the Gaussian distribution, for The second norm of .
9. The response prediction method according to claim 8, characterized in that: In the joint probability calculated by the sum of the log-likelihood and the integral term results of the data sample, when the patient responds to the efficacy of immune checkpoint blockade therapy, the logarithm of the joint probability is The specific calculation formula is: in, for In the response group The joint probability in is the original sample without noise, Time during the diffusion process The noise disturbance characteristics when For patients who respond to immune checkpoint blockade therapy, is the prediction score, is the divergence of the prediction scores, represents the model parameters of the diffusion model after training, and The diffusion start time and diffusion end time are respectively, For the integral arrive Cumulative divergence effect, is a hyperparameter, indicating the maximum noise intensity, is the integration variable.
10. An immune checkpoint blockade response prediction system based on a diffusion model, characterized in that: It includes data acquisition module, prediction module, joint probability module and response probability module; The data acquisition module is used to obtain the multi-dimensional indicator characteristics of the patient and splice them to obtain 7-dimensional spliced data, which is then input into the ODE solver; In the prediction module, the ODE solver inputs the current time step and the current intermediate data generated by the solution into the trained diffusion model, and feeds the obtained current prediction score back to the ODE solver. The ODE solver iteratively inputs the next time step generated by the solution and the intermediate data corresponding to the next time step as the current time step into the diffusion model until the ordinary differential equation in the ODE solver is solved. The ODE solver calculates the data samples and integral term results that conform to the Gaussian distribution. The intermediate data is the noise disturbance feature corresponding to the current time step. In the joint probability module, the joint probability is calculated using the sum of the log-likelihood of the data sample and the integral term result; In the response probability module, two joint probabilities are obtained based on the patient's response and non-response to immune checkpoint blockade therapy. The ratio of the joint probability in the response state to the sum of the two joint probabilities is used as the patient's response probability to immune checkpoint blockade therapy.