An artificial intelligence-based pre-evaluation method for gastric cancer immunotherapy

By combining the PPI-RWR and RGP algorithms to screen immunotherapy-related gene features, and using generative diffusion models and active learning for data enhancement, the problems of low biomarker detection efficiency and insufficient data in traditional methods were solved, and efficient and accurate prediction of the efficacy of gastric cancer immunotherapy was achieved, supporting personalized treatment.

CN119920455BActive Publication Date: 2025-09-19SICHUAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411705216.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-26
Publication Date
2025-09-19
Estimated Expiration
2044-11-26

AI Technical Summary

Technical Problem

Existing technologies make it difficult to effectively predict the efficacy of gastric cancer immunotherapy. Traditional biomarker testing is inefficient and costly, gene interactions are complex, and data sample size is insufficient, resulting in poor prediction accuracy.

Method used

The protein interaction restart random walk (PPI-RWR) and random gene perturbation (RGP) algorithms were combined with a generative diffusion model and active learning to perform feature screening and data enhancement to construct an immunotherapy efficacy prediction model.

Benefits of technology

It improves the accuracy of immunotherapy efficacy prediction and the generalization performance of the model, identifies biomarkers closely related to immunotherapy, and supports the formulation of individualized treatment plans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119920455B_ABST
    Figure CN119920455B_ABST
Patent Text Reader

Abstract

The present invention discloses a pre-evaluation method for gastric cancer immunotherapy based on an artificial intelligence method, which relates to the field of computational prediction of the efficacy of gastric cancer immunotherapy, and comprises the following steps: S100: analyzing data of gastric cancer patients receiving immunotherapy and constructing an immune feature data set; S200: writing a Python program to screen features based on a PPI-RWR algorithm and a random gene perturbation algorithm; S300: using a data enhancement algorithm based on a diffusion model to amplify samples; S400: using six machine learning algorithms to model and predict the standardized data obtained in step S300, and finding the optimal parameters of each model through ten-fold cross-validation and grid parameter search; S500: performing ten-fold cross-validation on each machine learning algorithm under the optimal hyperparameters, comparing and selecting the optimal machine learning model, and constructing the final model; S600: constructing an external test set to verify the model obtained in step S400.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computational prediction of the efficacy of gastric cancer immunotherapy, and in particular to a gastric cancer immunotherapy pre-evaluation method based on artificial intelligence methods. Background Art

[0002] In recent years, with the progress of immunotherapy research and the exploration of new immunotherapy targets in gastric cancer, major breakthroughs have been made in the treatment of gastric cancer. Compared with current traditional treatments, immunotherapy has significantly improved the survival benefits of gastric cancer patients, breaking the previous monopoly of chemotherapy and targeted therapy. However, the biggest challenge of this therapy is the huge disparity in patient efficacy, and only a small number of patients can benefit from it. Therefore, accurately predicting the patient groups that can benefit from immunotherapy, formulating corresponding individualized treatment plans, and avoiding the harm and burden caused by inappropriate treatment to patients are issues that urgently need to be addressed in clinical practice.

[0003] A major challenge in predicting immunotherapy efficacy in patients is identifying reliable predictive biomarkers in patients who have previously received immunotherapy. The current mainstream approach is to perform auxiliary diagnostic tests using experimental testing of known biomarkers (tumor mutation burden, PD-1 / PD-L1, peripheral blood biomarkers, etc.). However, traditional experiments are time-consuming and labor-intensive, and their detection efficiency is no longer sufficient to meet the needs of relevant research. Furthermore, as research in this field deepens, existing biomarkers are no longer able to meet the growing demand. Studies have reported no clear correlation between PD-L1 expression and response to ICI therapy, and some studies have even found that patients who respond to immunotherapy have lower PD-L1 expression levels. Litchfield et al. recently found that traditional biomarkers can only predict immunotherapy response in approximately 60%, indicating that new biomarkers remain to be discovered. Therefore, it is imperative to develop effective methods to identify biomarkers in patients who have previously received immunotherapy, ultimately maximizing the effectiveness of immunotherapy. With the rapid development of computer technology, machine learning, with its flexibility and powerful learning capabilities, has begun to be applied to biomarker screening. Chen et al. used a support vector machine model to develop a lncRNAs model composed of 16 lncRNA signatures for automated classification of microsatellite instability. These signatures may serve as targets for PD-1 / PD-L1 immunotherapy. Wang et al. used a random forest algorithm to identify a microbial signature among 32 bacterial genera associated with early gastric cancer. Saito et al. used LC / ESI-MS to measure the levels of 236 lipid molecules, including phosphatidylcholine and sphingomyelin, in 1592 plasma samples and applied a logistic regression algorithm to identify these signatures. The results showed a diagnostic accuracy of gastric cancer of 94.4%. However, while these previous approaches have been successful, they often overlook the interconnections and interactions between genes. Genes do not function independently; complex network interactions among them significantly influence disease development, progression, and response to treatment. Ignoring these interactions may result in the selection of signatures limited to the single gene level, failing to fully reflect potential synergistic effects. Therefore, it is crucial to utilize modern complex network theory, combined with gene interactions, to more comprehensively analyze the efficacy of immunotherapy.

[0004] Another major challenge in predicting the efficacy of immunotherapy in patients is the high cost and long experimental cycles of immunotherapy experiments, resulting in a small amount of valid data samples, which significantly affects the accuracy of prediction models. Directly applying machine learning will inevitably lead to problems such as poor model prediction performance and severe overfitting. Therefore, it is necessary to develop corresponding technologies to overcome the limitations of data volume. Currently, data augmentation technology has achieved great success in the field of computer vision. It reduces the risk of model overfitting and improves model generalization performance, and has become a common method in computer vision. However, biomedical data is significantly different from image data and cannot be used for data augmentation operations such as image flipping. Therefore, in the analysis of high-throughput sequencing data in the biomedical field, especially for gene expression data, there are few reports on data generation and data augmentation.

[0005] In summary, given the scarcity of immunotherapy data samples, the complex interactions between genes, and the difficulty of existing biomarkers in comprehensively predicting treatment effects, it is particularly necessary to develop a method that can quickly and accurately predict the efficacy of immunotherapy. This not only overcomes the shortcomings of traditional methods in marker screening and data volume limitations, but also provides a theoretical basis for personalized immunotherapy plans, thereby further guiding clinical decision-making and efficacy evaluation, and maximizing the benefits of immunotherapy. Summary of the Invention

[0006] The object of the present invention is to provide a method for pre-evaluation of gastric cancer immunotherapy based on an artificial intelligence method, and propose a feature screening strategy of a protein interaction restart random walk (Protein-ProteinInteraction Random Walk With Restart, PPI-RWR) algorithm combined with a random network walk and a random gene perturbation (Random Gene Perturbation, RGP) algorithm, and introduce a data enhancement method combining a generative diffusion model and active learning, which is applied to the pre-evaluation of the efficacy of gastric cancer immunotherapy. Through this method, while effectively solving the problems of insufficient data and limited sample size in traditional experimental determination processes, biomarkers closely related to the efficacy of immunotherapy can be screened out, and the excellent extrapolation performance of the model can be ensured. This method not only provides a new idea for the prediction of the efficacy of immunotherapy, but also provides reliable support for clinical treatment decision-making, helping the development of individualized and precise immunotherapy.

[0007] In order to achieve the above object, the present invention adopts the following technical solutions:

[0008] A new method for pre-evaluation of gastric cancer immunotherapy based on artificial intelligence methods, including:

[0009] Step S100: Analyze data from gastric cancer patients receiving immunotherapy and construct an immune signature dataset. RNA-seq data from 78 gastric cancer patients who received immunotherapy were downloaded from ICBaltas. The treatment efficacy of these patients was determined using the Response Evaluation in Solid Tumors (RECIST) criteria. To ensure the effectiveness of subsequent modeling, the data was preprocessed using the following procedures: data filtering, missing value imputation, data normalization, and data conversion. Subsequently, differentially expressed genes between patients who responded to treatment and those who did not were extracted. The data were then aggregated into behavioral genes and listed as a sample matrix.

[0010] Step S200: Write a Python program to screen features using the PPI-RWR algorithm and the random gene perturbation algorithm. Select known target genes and pathway genes related to immunotherapy as seed genes for network propagation, and select genes with the top 10% propagation scores. Finally, continue screening for biomarkers using the random gene perturbation algorithm to further narrow the feature space and obtain the optimal gene set for subsequent model construction.

[0011] Step S300: Amplify the sample using a data augmentation algorithm based on a diffusion model. This technology applies the diffusion model to immunotherapy sample generation for the first time and innovatively combines active learning with data augmentation, taking into account the quality of new data during the data augmentation process, effectively addressing the limitation of the scarce immunotherapy data.

[0012] Step S400, use six machine learning algorithms to model and predict the standardized data obtained in step S300, including: K-nearest Neighbors (KNN), support vector machine (SVM), decision tree (DT), random forest (RF), extreme gradient boosting (XGB), and logistic regression (LR). The optimal parameters of each model are found through ten-fold cross validation and grid parameter search.

[0013] Step S500: Perform ten-fold cross validation on each machine learning algorithm under the optimal hyperparameters, compare and select the optimal machine learning model, and build the final model.

[0014] S600: Construct an external test set to verify the model obtained in step S400 to prove the good generalization ability of the model.

[0015] The step S100 specifically includes:

[0016] Step S110 : Download RNA-seq data of gastric cancer patients who have received immunotherapy from ICBaltas as an initial data set.

[0017] Step S120: Preprocess the data: use the Python Pandas package to implement data filtering, missing value filling, and min-max data standardization.

[0018] Step S130: Using the R package limma, differentially expressed genes were screened and retained according to the conditions of |log2FC|≥1 and FDR<0.05.

[0019] Step S140: In order to compress the data value range, smooth the distribution and improve the stability of the model, the data is logarithmized using the Python Numpy package.

[0020] The step S200 specifically includes:

[0021] Step S210: Build a target network. Download the human PPI network from the STRING database v.12.0.

[0022] Step S220: Select existing immunotherapy target genes (PD-1, PD-L1, CTLA-4) and immune-related pathway (T cell receptor, B cell receptor, antigen presentation) genes as seed genes, use the page-rank method of PythonNetworkX, and determine immunotherapy-related genes through the PPI-RWR algorithm.

[0023] Step S230: iteratively run the random gene perturbation algorithm until the evaluation index begins to decrease, and continue screening key biomarkers.

[0024] The step S300 specifically includes:

[0025] Step S310: Use the data set D0 obtained in step S200 to construct an agent model C0 for predicting the efficacy of immunotherapy.

[0026] Step S320: Run the diffusion model to perform data enhancement to obtain 500 enhanced samples, which constitute the sample pool U0 required for active learning.

[0027] Step S330: Use the proxy model obtained in step S310 to predict the sample pool obtained in step S320 and calculate the expected improvement value EI of each sample to be selected. Then sort the expected improvement values ​​of each sample to be selected in descending order, select the top 50 samples and add them to the original dataset D0 to form the enhanced dataset D1. At this time, there are 450 samples left in the sample pool U1 to be selected.

[0028] Step S340: Use the enhanced dataset D1 in step S330 to retrain and obtain an updated proxy model C1. Use the proxy models C1 and C1 to predict the sample pool U1 and calculate the expected improvement EI. After sorting in descending order, select the top 50 samples and add them to the dataset D1 to form the enhanced dataset D2. At this time, there are 400 samples left in the sample pool U2 to be selected.

[0029] Step S350: Use the enhanced dataset D2 from step S340 to retrain the updated proxy model C2, perform predictions on the sample pool U2, and calculate the expected improvement EI. After sorting in descending order, select the top 50 samples and add them to the dataset D2 to form the enhanced dataset D3. At this point, there are 350 samples left in the sample pool U3 to be selected.

[0030] Step S360: Use the enhanced dataset D3 in step S330 to retrain and obtain the updated proxy model C3, make predictions on the sample pool U3, and calculate the expected improvement EI. After sorting in descending order, select the top 50 samples and add them to the dataset D3 to form the enhanced dataset D4. At this time, there are 300 samples left in the sample pool U4 to be selected.

[0031] The step S400 specifically includes:

[0032] Step S410: determining the parameter optimization space of each machine learning algorithm;

[0033] Step S420: Divide the data set according to ten-fold cross validation to ensure that the distribution of each data is as even as possible, with 9 data sets used as training sets and the remaining 1 data set used as a test set;

[0034] In step S430, when executing the parameter optimization process, taking XGB as an example, for each set of hyperparameters, the model is first trained on the training set of step S420, and the evaluation indicators of the classification model are calculated on the test set. The entire process is repeated 10 times to ensure that each piece of data is used as a training set and a test set, and finally 10 sets of evaluation indicator results are obtained. The average of these 10 sets of results is taken as the performance of the model under this set of hyperparameters. Next, all hyperparameter combinations defined in step S410 are traversed, the evaluation indicators of each combination are calculated, and finally the set of hyperparameters with the best performance is selected as the optimal hyperparameters for XGB. Other machine learning algorithms follow the same steps to determine their respective optimal hyperparameters.

[0035] Step S440: Execute steps S420 and S430 on the enhanced datasets D2, D3, and D4 in sequence to obtain the optimal hyperparameters of each machine learning algorithm under different datasets.

[0036] The step S500 specifically includes:

[0037] Step S510: Based on the enhanced dataset D1, ten-fold cross validation is performed on each machine learning algorithm according to the optimal hyperparameters obtained in step S400, and the ten average evaluation indicators R2, RMSE, and MAE are obtained. The optimal model is selected based on the criteria of maximum R2, minimum RMSE, and MAE.

[0038] Step S520: Similarly, step S510 is performed on the enhanced data sets D2, D3, and D4 in sequence to select the optimal model corresponding to each data set;

[0039] Step S530: Compare the model performance under different enhanced data sets obtained in step S510 and step S520, and select the optimal enhanced data set and the optimal machine learning algorithm;

[0040] Step S540: Based on the comparison result of step S530, it is determined that the optimal enhanced data set is D4 and the optimal friction sensitivity prediction algorithm is XGB. Therefore, the final immunotherapy efficacy prediction model is constructed based on all the data of D4 using the optimal hyperparameters.

[0041] The step S600 specifically includes:

[0042] S610. Collect gastric cancer patient samples from the GEO database. First, use the combat package of R language to eliminate the batch effect. Then use the TIDER package of R language to simulate the results of patients after receiving immunotherapy, and construct an external test set for subsequent verification.

[0043] S620 , using the optimal XGB model constructed in step S540 to predict the external test set generated in step S610 , and calculating the evaluation index of the model to evaluate the generalization ability of the model.

[0044] The beneficial effects of the pre-evaluation method for gastric cancer immunotherapy based on artificial intelligence disclosed in this application include but are not limited to:

[0045] 1. This invention innovatively incorporates the PPI-RWR algorithm and the RGP algorithm into feature screening for predicting immunotherapy efficacy. Starting from target genes, PPI-RWR explores inter-gene interactions through random walks, revealing potential synergistic effects and identifying genes closely related to immunotherapy from a global network. This approach not only considers the importance of individual genes but also captures the complex interactions between genes, significantly enhancing the biological significance of screening results.

[0046] 2. To address the limitations of insufficient immunotherapy data samples, this paper proposes the first AL-Diffusion data augmentation technique, combining a generative deep diffusion model with active learning. The diffusion model generates new samples that approximate real data, while active learning monitors their quality, ensuring the effectiveness of data augmentation while avoiding the over-reliance on experimental annotations that is common in conventional active learning.

[0047] 3. The present invention has a wide range of applications, is easy to operate, and provides fast and accurate predictions. The optimal model constructed based on 79 gene signatures is not only applicable to immunotherapy research, but can also be widely applied to other related fields, enabling rapid and accurate predictions based on gene expression values. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 is a flow chart of the present invention,

[0049] Figure 2 This is an example diagram of the results of the data processing part of the present invention.

[0050] Figure 3 This is a schematic diagram of the core algorithm PPI-RWR algorithm of the present invention.

[0051] Figure 4 This is an example diagram of the PPI-RWR algorithm results.

[0052] Figure 5 This is an example diagram of the RGP algorithm results.

[0053] Figure 6 This is a schematic diagram of the AL-Diffusion data enhancement technology, the core technology of this invention.

[0054] Figure 7 Schematic diagram of the diffusion model used in the present invention. DETAILED DESCRIPTION

[0055] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0056] On the contrary, this application covers any alternatives, modifications, equivalents, and solutions made within the spirit and scope of this application as defined by the claims. Furthermore, to facilitate a better understanding of this application, certain specific details are described in detail below in the detailed description of this application. Those skilled in the art will be able to fully understand this application without these details.

[0057] The following is a detailed description of a method for monitoring the growth process and quality of peach blossoms according to an embodiment of the present invention. It should be noted that the following embodiments are merely used to explain the present invention and do not constitute a limitation of the present invention.

[0058] like Figure 1 As shown, a pre-evaluation method for gastric cancer immunotherapy based on artificial intelligence method comprises:

[0059] Step 1: Download gene expression data of gastric cancer patients who received immunotherapy from ICBaltas. The data was pre-processed by pre-data filtering, missing value filling, data normalization, and data conversion. Subsequently, differentially expressed genes between patients with and without treatment response were extracted. Data with a large numerical scale range will not only greatly slow down the training process, but also seriously affect the model performance and cause biased results. Therefore, the data was min-max normalized according to eq (1) to obtain a dataset, and the data was logarithmized:

[0060]

[0061] The final dataset is as follows Figure 2 As shown, the rows of the dataset represent genes and the columns represent patient samples.

[0062] Step 2: Use PPI-RWR technology and RGP algorithm to screen important immune-related genes according to eq (2), such as Figure 3 Network propagation is performed as shown:

[0063] P t+1 =(1-r)WP t +rP0(2)

[0064] Where: P t represents the probability vector of the node at step t, P0 represents the initial probability vector, r is the restart probability, and W is the transition matrix, i.e., the column-normalized adjacency matrix of the network. The walk process starts randomly from a set of nodes in the network, repeatedly spreading from the current node to randomly selected adjacent nodes, or returning to the starting node. Let the network propagation reach a stable state (P t and P t+1 The change between the two is less than 10 -30 ), the probability vector can represent the association score of all genes in the network with the starting point. The results of the PPI-RWR algorithm are as follows Figure 4 .

[0065] Next, we first selected the top 10% of genes in the association score as our candidate genes. We then used the following pseudocode to execute the random gene perturbation algorithm to screen biomarkers and obtain the optimal gene set for subsequent modeling. Finally, after 10 rounds of iteration, 79 genes were selected as the optimal feature set. The results of the screening process are shown below. Figure 5 shown.

[0066] Algorithm: Random Gene Perturbation (RGP) algorithm

[0067] Input: gene expression dataset R, estimator E

[0068] Output: Screened biomarker gene set D

[0069] 1: Begin

[0070] 2: Input R into E and calculate the AUC value A of A;

[0071] 3: Randomly shuffle the expression value of gene g and calculate the new AUC value A';

[0072] 4: Calculate the importance score of g I = A'-A;

[0073] 5: Select genes G whose importance scores rank in the top 20%;

[0074] 6: Repeat

[0075] 7: Use feature recursive elimination to screen genes in G;

[0076] 8: Remove the last 10% of genes to obtain a new gene set G';

[0077] 9: UntilAUC value starts to decrease;

[0078] 10: Return the optimal biomarker gene set D;

[0079] 11: End

[0080] Step 3: Use AL-Diffusion data enhancement technology to amplify the data set and generate high-quality samples. Specifically, first perform the diffusion process according to eq (3) and add noise to disturb the original sample: The core technology of the present invention, AL-Diffusion data enhancement technology, is shown in the schematic diagram. Figure 6 .

[0081]

[0082] Among them, x0 represents the original sample, x t represents the sample with noise added after step t, α tis a predefined noise scheduling parameter, ∈ is the Gaussian noise of the standard normal distribution, and through forward diffusion, the sample is gradually perturbed by the noise.

[0083] Then, the reverse diffusion process is performed according to eq (4) to finally generate new noise-free samples.

[0084]

[0085] in, The noise predicted by the neural network at step t helps to gradually remove the noise, returning the sample to a state close to the true distribution and generating new high-quality new sample points. Through the reverse process, we can effectively restore the original characteristics of the sample and generate enhanced samples that are highly similar to the original data.

[0086] Then, valuable samples are selected from the augmented samples through active learning technology, and the initial proxy model Cf is constructed using the original dataset D0. The sample labels of the sample pool U0 are predicted and the expected improvement value EI of each sample to be selected is calculated according to eq(5)(6). The EI value of each sample to be selected is sorted in descending order, and the top 50 samples are selected and added to the original dataset D0 to form the enhanced dataset D1. At this time, there are 450 samples left in the sample pool U1 to be selected.

[0087] Then, valuable samples are selected from the augmented samples through active learning technology, and the initial proxy model Cf is constructed using the original dataset D0. The sample labels of the sample pool U0 are predicted and the expected improvement value EI of each sample to be selected is calculated according to eq(5)(6). The EI value of each sample to be selected is sorted in descending order, and the top 50 samples are selected and added to the original dataset D0 to form the enhanced dataset D1. At this time, there are 450 samples left in the sample pool U1 to be selected.

[0088]

[0089]

[0090] Where σ is the standard deviation, is the probability density function, Φ(z) is the cumulative distribution function, μ * It refers to the maximum value expressed in the data set.

[0091] Step 4: Perform parameter optimization of different machine learning algorithms on the enhanced dataset. Specifically, determine the parameter optimization space for each machine learning algorithm, as shown in Table 1:

[0092] Table 1. Hyperparameter optimization space for different machine learning algorithms

[0093]

[0094]

[0095] Combined with ten-fold cross-validation, grid search was performed, and Acc, Pre, Recall, F1, and AUC were selected as model evaluation indicators. The indicator values ​​under each set of hyperparameters were calculated according to the following formula, and the set of hyperparameters corresponding to the optimal value was selected as the optimal hyperparameters:

[0096]

[0097]

[0098]

[0099]

[0100] Step 5: Compare the performance of various machine learning algorithms and select the best algorithm to build the final model. Specifically:

[0101] Comparing the impact of different augmentation times on model building, the overall performance of the model shows a gradual improvement trend as the number of data augmentation increases. When 50 data augmentations are performed, the AUC value of most models is approximately 0.7. As the number of augmentations increases to 100 and 150, the performance of each model improves to varying degrees, with the highest AUC value exceeding 0.9. However, when the number of augmentations reaches 200, the performance improvement begins to level off. This may be because when the original sample size is small, too many augmentation operations will lead to data redundancy in the sample space, thereby deviating from the actual distribution of the original data. Therefore, 200 data augmentations may be the optimal number of augmentations.

[0102] After determining the optimal number of augmentations, the performance of each machine learning algorithm was further compared. The Acc, Pre, Recall, F1, and AUC values ​​of each machine learning algorithm under 10-fold cross-validation with 200 data augmentations showed that XGB was the optimal immunotherapy efficacy prediction model. The hyperparameters and optimal results of the optimal model are shown in Table 2.

[0103] Table 2. Optimal model parameters and performance

[0104]

[0105] Step 6: Construct an external test set according to step S600 to verify the optimal XGBoost model. The analysis results are listed in Table 3. It can be seen that the constructed prediction model still shows good prediction performance in the external test set, which proves the effectiveness of our model.

[0106] Table 3. External test results of the optimal model

[0107] Performance indicators Acc Pre Recall F1 AUC Numerical 0.922 0.948 0.948 0.948 0.901

[0108] Step 7: Use the pickle.dump method to save the normalizer as a .dat file and the final model as a .pkl file. Use the pickle.load method to load it when predicting new samples.

[0109] In summary, the present invention achieves:

[0110] (1) Introduce PPI-RWR algorithm and RGP algorithm for feature screening.

[0111] A major challenge in predicting the efficacy of immunotherapy in patients lies in identifying reliable predictive biomarkers in patients who have received immunotherapy. The current mainstream approach is to use immunohistochemistry to detect the expression of known biomarkers as auxiliary diagnostic tests. However, current detection efficiency is no longer sufficient to meet the needs of related research. As research in this field deepens, existing biomarkers are no longer able to meet the growing demand. There is an urgent need to discover new biomarkers to maximize the effectiveness of immunotherapy. Previous studies have often overlooked the interactions between genes. Biological networks are powerful resources for discovering genes or modules that drive disease development. Modern complex network theory provides comprehensive tools to help us understand the occurrence and development of complex diseases such as tumors. Currently, many bioinformatics methods based on random walks, information diffusion, and resistance have been proposed. These methods use networks for genetic analysis and have been successfully applied to identify pathogenic genes, therapeutic modules, and drug targets. Currently, these technologies have not been widely used in immunotherapy research.

[0112] To effectively identify gene associations and identify effective immune gene signatures, this study incorporates the PPI-RWR algorithm into immunotherapy research. Starting from a given PPI network, the algorithm moves along the network with equal probability, ranking genes based on their stabilized scores. Compared to previous methods, our algorithm considers interactions between genes, revealing potential synergistic effects through these interactions. This approach offers exceptional flexibility and scalability, enabling rapid adaptation to the growing demands of immunotherapy research.

[0113] Furthermore, to further extract key gene features, we proposed the RGP algorithm, based on the PPI-RWR algorithm. This algorithm shuffles gene expression values ​​and compares them with the original expression data to calculate gene importance scores, thereby determining the gene's impact on immunotherapy. This method effectively identifies genes with potential impact on immunotherapy efficacy, resulting in a smaller and more precise feature set, which greatly improves subsequent computational efficiency and accuracy.

[0114] By combining the PPI-RWR and RGP algorithms, this paper can identify genetic signatures closely related to immunotherapy efficacy from multiple perspectives. This not only provides new insights into existing immunotherapy research but also offers strong support for future personalized treatment strategies. The biomarkers identified by this method can be further applied in clinical diagnosis, helping doctors to more accurately predict a patient's immune response before treatment and develop more effective treatment plans.

[0115] (2) AL-Diffusion data enhancement technology is proposed to effectively amplify small sample data sets.

[0116] Diverse and abundant data are fundamental to training high-performance machine learning models. However, immunotherapy clinical trials are difficult and expensive, and lack standardized standards across different experiments, resulting in limited available data. Because machine learning relies heavily on data, the quality and quantity of data directly impact the predictive model's ability to predict unknown samples. Therefore, the limitations of existing experimental data have become a bottleneck hindering the prediction of immunotherapy efficacy.

[0117] Data augmentation is an effective method for expanding the size of data samples. Its core idea is to reduce the risk of overfitting and improve the generalization ability of the model by appropriately modifying existing data or generating new samples based on existing data. Data augmentation technology has achieved remarkable success in the field of computer vision and has become a common method for improving model performance. Unlike data augmentation, which focuses on expanding the size of data, active learning focuses on selecting the most valuable samples to achieve the best model performance with the least data. Active learning evaluates the value of unlabeled data based on the current model and selects high-value samples for expert annotation, reducing the need for extensive experimental annotation and improving the model's predictive performance. Although data augmentation and active learning have developed independently in their respective fields, research combining the two is relatively rare, especially in the field of immunotherapy efficacy prediction, where this combination has not yet been fully utilized.

[0118] See also Figure 7To address this problem, this paper introduces generative diffusion models from the field of computer vision to immunotherapy research for the first time, and combines them with active learning to propose an adaptive AL-Diffusion method to achieve higher-quality data enhancement. The advantage of the diffusion model is that it can generate new samples with high biological significance by gradually adding and removing noise, effectively simulating the potential distribution of immunotherapy data, and thus generating samples close to real patient data. In image data, the diffusion model generates realistic new images through denoising; in immunotherapy data, it generates new samples that conform to biological characteristics by learning the potential distribution of existing patient samples, thereby effectively addressing the challenge of insufficient data.

[0119] The advantage of the diffusion model is that the new samples it generates are more continuous and smoothly distributed in the latent space, thus addressing the deficiencies of insufficient sample size and uneven distribution in clinical trials and improving the model's robustness to out-of-distribution samples. Furthermore, the diffusion model can flexibly capture subtle differences between samples, generating samples that are not only highly diverse but also retain important biometric information, which plays a significant role in improving the accuracy of immunotherapy efficacy predictions.

[0120] However, generative models may also produce some samples that deviate from biological significance. In this case, relying solely on data augmentation may not improve model performance, and may even burden the training process and reduce prediction accuracy. To this end, combining active learning can effectively solve this problem. Active learning selects the most valuable samples for training to ensure that the model's generalization ability is further improved, while avoiding the negative impact of generated data on model performance. Active learning uses adaptive design strategies combined with optimization algorithms to intelligently select the most valuable samples in each round of training based on the model's prediction results and uncertainty. The AL-Diffusion method cleverly combines active learning and generative diffusion models. Active learning screens and monitors the quality of generated samples to ensure the high quality of the generated data. At the same time, the diffusion model reduces active learning's dependence on experimental annotations. The two complement each other's strengths to achieve efficient data augmentation.

[0121] This method not only effectively addresses the limited sample size issue in the immunotherapy field, but also further enhances the model's generalization and extrapolation capabilities by generating high-quality samples. This innovative data augmentation strategy provides a new solution for predicting immunotherapy efficacy, significantly improving the ability to process small sample data and is expected to promote the personalized and precise development of immunotherapy.

[0122] (3) A prediction model with superior performance was constructed.

[0123] In the context of predicting the efficacy of immunotherapy, the differences in raw gene expression data are large. Directly using these data for model training may cause highly expressed genes to have an excessive impact on the model, while the effects of low-expressed genes may be weakened or even ignored. In order to reduce this impact, the present invention adopts a min-max normalization method. By scaling the expression values ​​of each gene to the same interval, it ensures that each gene has the same weight for the objective function of the model and balances the impact of each gene on the prediction results. This processing method not only improves the comparability between features, but also effectively reduces the deviation caused by different feature scales, thereby enhancing the stability of model training and the accuracy of prediction.

[0124] On the basis of completing min-max normalization and AL-Diffusion data enhancement, the present invention also used grid search and ten-fold cross-validation to compare the modeling effects of six different machine learning algorithms. Ultimately, the XGBoost (XGB) model performed best, with an AUC value of 0.978 on the training set, and was selected as the optimal model. To further evaluate the generalization ability of the model, we verified it on a constructed external test set, and the AUC value reached 0.901, demonstrating the good robustness and predictive ability of the model in predicting the efficacy of immunotherapy.

Claims

1. A method for constructing a gastric cancer immunotherapy efficacy prediction model based on artificial intelligence, characterized in that: The following steps are involved: S100: Analyze data from gastric cancer patients receiving immunotherapy and construct an immune signature dataset; S200: Write a Python program to screen features based on the protein interaction restart random walk PPI-RWR algorithm and the random gene perturbation algorithm; The random gene perturbation algorithm calculates the importance score of the gene by disrupting the expression value of the gene and comparing it with the original expression data to determine the impact of the gene in immunotherapy; S300: Amplify samples using a data augmentation algorithm based on a generative deep diffusion model combined with active learning; S400: Using six machine learning algorithms to perform modeling and prediction on the standardized data obtained in step S300, and finding the optimal parameters of each model through ten-fold cross validation and grid parameter search; S500: Each machine learning algorithm performs ten-fold cross-validation under the optimal hyperparameters, compares and selects the optimal machine learning model, and builds the final model; S600: Construct an external test set to verify the model obtained in step S500; The step S300 specifically includes: Step S310: constructing an agent model C0 for predicting the efficacy of immunotherapy using the data set D0 obtained in step S200; Step S320: Run the diffusion model to perform data enhancement to obtain X enhanced samples, which constitute the sample pool U0 required for active learning; Step S330: Use the proxy model obtained in step S310 to predict the sample pool obtained in step S320 and calculate the expected improvement value EI of each sample to be selected. Then sort the expected improvement values ​​of each sample to be selected in descending order, select the top 50 samples and add them to the original dataset D0 to form the enhanced dataset D1. At this time, there are (X-50) samples left in the sample pool U1 to be selected. Step S340: Use the enhanced dataset D1 from step S330 to retrain and obtain an updated proxy model C1. Use the proxy model C1 to predict the sample pool U1 and calculate the expected improvement EI. After sorting in descending order, select the top 50 samples and add them to the dataset D1 to form the enhanced dataset D2. At this point, there are (X-100) samples left in the sample pool U2 to be selected. Step S350: Use the enhanced dataset D2 from step S340 to retrain the updated proxy model C2, perform predictions on the sample pool U2, and calculate the expected improvement EI. After sorting in descending order, select the top 50 samples and add them to the dataset D2 to form the enhanced dataset D3. At this point, there are (X-150) samples left in the sample pool U3 to be selected. Step S360: Use the enhanced dataset D3 in step S330 to retrain and obtain the updated proxy model C3, make predictions on the sample pool U3, and calculate the expected improvement EI. After sorting in descending order, select the top 50 samples and add them to the dataset D3 to form the enhanced dataset D4. At this time, there are (X-200) samples left in the sample pool U4 to be selected.

2. The method for constructing an artificial intelligence-based gastric cancer immunotherapy efficacy prediction model according to claim 1, characterized in that: The step S100 specifically includes: Step S110, downloading RNA-seq data of gastric cancer patients who have received immunotherapy from ICBaltas as an initial data set; Step S120: pre-process the data: use the Python Pandas package to implement data filtering, missing value filling and min-max data standardization; Step S130, using the R package limma, according to the conditions of |log2FC| ≥ 1 and FDR < 0.05, screen and retain differentially expressed genes; Step S140: Use Python's Numpy package to perform logarithmic processing on the data.

3. The method for constructing an artificial intelligence-based gastric cancer immunotherapy efficacy prediction model according to claim 1, characterized in that: The step S200 specifically includes: Step S210: construct a target network and download the human PPI network from the STRING database v.12.0; Step S220: Select existing immunotherapy target genes PD-1, PD-L1, and CTLA-4, as well as immune-related pathway T cell receptor, B cell receptor, and antigen presentation genes as seed genes, and use the page-rank method of PythonNetworkX to determine immunotherapy-related genes using the PPI-RWR algorithm; Step S230: iteratively run the random gene perturbation algorithm until the evaluation index begins to decrease, and continue screening key biomarkers.

4. The method for constructing an artificial intelligence-based gastric cancer immunotherapy efficacy prediction model according to claim 1, characterized in that: The step S400 specifically includes: Step S410: determining the parameter optimization space of each machine learning algorithm; Step S420: Divide the data set according to ten-fold cross validation to ensure that the distribution of each data is as even as possible, with 9 data sets used as training sets and the remaining 1 data set used as a test set; Step S430: When performing the parameter optimization process, for each set of hyperparameters, first train the model on the training set of step S420, and calculate the evaluation index of the classification model on the test set; the entire process is repeated 10 times, ensuring that each data set is used as a training set and a test set, and finally obtain the results of 10 sets of evaluation indicators; take the average of these 10 sets of results as the performance of the model under this set of hyperparameters; traverse all hyperparameter combinations defined in step S410, calculate the evaluation index of each combination, and finally select the set of hyperparameters with the best performance as the optimal hyperparameters; other machine learning algorithms follow the same steps to determine their respective optimal hyperparameters; Step S440: Execute steps S420 and S430 on the enhanced datasets D2, D3, and D4 in sequence to obtain the optimal hyperparameters of each machine learning algorithm under different datasets.

5. The method for constructing an artificial intelligence-based gastric cancer immunotherapy efficacy prediction model according to claim 4, characterized in that: The step S500 specifically includes: Step S510: Based on the enhanced data set D1, ten-fold cross validation is performed on each machine learning algorithm according to the optimal hyperparameters obtained in step S400 to obtain the ten-fold average evaluation index R 2 , RMSE and MAE, according to the maximum R 2 The optimal model is selected based on the criteria of minimum RMSE and MAE; Step S520: Execute step S510 on the enhanced datasets D2, D3, and D4 in sequence to select the optimal model corresponding to each dataset; Step S530: Compare the model performance under different enhanced data sets obtained in step S510 and step S520, and select the optimal enhanced data set and the optimal machine learning algorithm; Step S540: Based on the comparison result of step S530, it is determined that the optimal enhanced data set is D4 and the optimal prediction algorithm is XGB. Therefore, the final immunotherapy efficacy prediction model is constructed based on all the data of D4 using the optimal hyperparameters.

6. The method for constructing an artificial intelligence-based gastric cancer immunotherapy efficacy prediction model according to claim 5, characterized in that: The step S600 specifically includes: S610: Gastric cancer patient samples were collected from the GEO database. The R language combat package was first used to eliminate batch effects. The R language TIDER package was then used to simulate the outcomes of patients receiving immunotherapy. An external test set was constructed for subsequent validation. S620 , using the optimal XGB model constructed in step S540 to predict the external test set generated in step S610 , and calculating the evaluation index of the model to evaluate the generalization ability of the model.

Citation Information

Patent Citations

  • Safety performance prediction method of RDX modified double-base propellant

    CN117766058A

  • Gastric cancer immunotherapy response prediction method based on TRP scoring system

    CN118072954A