A method, device, medium and program product for constructing a model reflecting the degree of immunotherapy response
By constructing a HAPIR model based on Hallmark gene set, the problem of inaccurate prediction of existing ICI responses is solved, and more efficient immunotherapy response prediction and drug resistance prediction are achieved, enhancing the effect of ICI treatment.
Patent Information
- Application Number
- CN202510371615.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-03-27
AI Technical Summary
Existing immunocheckpoint inhibitors (ICI) treatment response predicts that biomarkers such as PD-L1 expression are not robust enough, resulting in 60% to 80% of patients not responding to immunotherapy, and more robust and specific biomarkers are urgently needed to predict responses and overcome drug resistance.
HAPIR, a prediction method based on Hallmark gene set refining, was adopted to obtain transcriptome data, perform differential expression analysis and gene set enrichment, identify significantly enriched Hallmark gene sets, calculate activity values and input machine learning models to construct a model reflecting the degree of immunotherapy response.
It improves the accuracy and robustness of ICI response prediction, can more accurately predict the drug resistance probability of patients, and has a wide range of application prospects.
Smart Images

Figure CN119905138B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent medicine, and more specifically, to a method, device, medium and program product for constructing a model reflecting the degree of immunotherapy response. Background Art
[0002] In recent years, immune checkpoint inhibitors (ICIs) have significantly improved survival rates for cancer patients. ICIs have been used to treat a variety of cancers, such as melanoma, breast cancer, and esophageal squamous cell carcinoma. These inhibitors boost the patient's immune system, enabling it to effectively recognize and attack cancer cells. However, 60% to 80% of treated patients do not respond well to immunotherapy. This creates an urgent need to identify highly sensitive and specific biomarkers that can predict response before treatment or facilitate the development of combination therapies to overcome ICI resistance.
[0003] Current predictive biomarkers for ICI response, such as programmed cell death ligand 1 (PD-L1) expression, have been clinically validated but are not robust enough. Due to this limitation, researchers are increasingly focusing on gene sets that have shown greater robustness and accuracy in predicting ICI response. For example, pan-cancer stemness signatures have been successfully used to predict the response of multiple cancer types to immunotherapy; there are also studies that use the expression levels of multiple biological pathways as biomarkers, combined with logistic regression models to predict ICI response. Similarly, researchers have demonstrated the strong predictive ability of 14 signaling pathway activities for ICI. However, the predictive performance is still limited, and further exploration of other gene sets is needed to enhance the prediction of ICI response. Summary of the Invention
[0004] The present invention aims to address at least one of the technical problems existing in the prior art. To this end, the present invention provides a method, apparatus, medium, and program product for constructing a model reflecting the degree of immunotherapy response. The present method proposes a method for predicting immunotherapy response based on Hallmark gene set refinement, namely HAPIR, and validates it using multiple cohorts, demonstrating that HAPIR is an effective tool for predicting ICI response and guiding the development of new immunotherapy strategies.
[0005] The first aspect of the present application discloses a method for constructing a model reflecting the degree of immunotherapy response, the method comprising:
[0006] S101, obtain the transcriptome dataset and clinical classification information of the training set samples; the clinical classification information includes sensitivity and resistance;
[0007] S102, performing differential expression analysis on the transcriptome dataset to obtain N differentially expressed genes;
[0008] S103, performing gene set enrichment analysis based on the differentially expressed genes to identify M significantly enriched Hallmark gene sets;
[0009] S104, taking the intersection of a single Hallmark gene set and the differentially expressed genes to obtain M refined Hallmark gene sets; and calculating the activity value of each refined Hallmark gene set;
[0010] S105, inputting the activity value into a machine learning model to obtain a predicted classification result, comparing it with the clinical classification information, optimizing the model based on the comparison result, and constructing a model reflecting the degree of immunotherapy response.
[0011] In some embodiments, the activity value calculation method includes: calculating the expression level of each gene in the refined Hallmark gene set and weighting it according to its importance in the gene set; summing the weighted expression levels of all genes in the gene set to obtain the activity value of the corresponding refined Hallmark gene set.
[0012] In some embodiments, the training set samples include immunotherapy-sensitive samples and drug-resistant samples; and the differentially expressed genes include significantly upregulated and downregulated differentially expressed genes.
[0013] In some embodiments, the method used by the machine learning model includes any one or more of the following: glm function, lm function, nls function, glmnet function, coxph function, mgcv function, randomForest function, gbm function, xgboost function, neuralnet function.
[0014] The second aspect of the present application discloses a method for predicting drug resistance probability based on HAPIR, the method comprising:
[0015] S201, obtaining transcriptome data of the subject;
[0016] S202, inputting the transcriptome data into the model reflecting the degree of immunotherapy response described in the first aspect of the present application, and outputting a result indicating whether the probability of drug resistance of the subject is high or low.
[0017] In some embodiments, the method further includes: if a result with a high probability of drug resistance is output, an auxiliary prediction result with a shorter subject survival period is obtained; if a result with a low probability of drug resistance is output, an auxiliary prediction result with a longer subject survival period is obtained.
[0018] The third aspect of the present application discloses a method for predicting drug resistance probability based on HAPIR, the method comprising:
[0019] S301, obtaining transcriptome data of the subject; the transcriptome data includes any one or more of the following genes: S100A1, TYR, SERPINE2, SGCD, PHLDA1;
[0020] S302, inputting the transcriptome data into the model reflecting the degree of immunotherapy response described in the first aspect of the present application, and outputting a result indicating whether the probability of drug resistance of the subject is high or low.
[0021] In a fourth aspect, the present application discloses a computer device, comprising: a memory and a processor; the memory is used to store a computer program; and the processor executes the computer program to implement the steps of the above method.
[0022] In a fifth aspect, the present application discloses a computer-readable storage medium having a computer program stored thereon, which implements the steps of the above-mentioned method when the computer program is executed by a processor.
[0023] In a sixth aspect, the present application discloses a computer program product, comprising a computer program, which implements the steps of the above method when executed by a processor.
[0024] This application has the following beneficial effects:
[0025] 1. This application innovatively proposes an improved method for predicting immunotherapy response (HAPIR) based on the Hallmark gene set. We first refined seven Hallmark gene sets, which are enriched in genes that are differentially expressed between responders and non-responders. Then, a machine learning model was trained based on the activity of the seven gene sets. As a result, HAPIR outperformed 13 existing ICI response prediction biomarkers, including biomarkers based on PD-1 and PD-L1. In addition, HAPIR also showed higher accuracy compared with gene-based models and other gene set-based models. Validation in multiple cancer type cohorts confirmed its robustness and significant correlation with patient survival.
[0026] 2. This application innovatively discloses a method for constructing a model that reflects the degree of immunotherapy response, and a method for predicting the probability of drug resistance using the model. During the model construction process, the transcriptome data of samples sensitive to immunotherapy and samples resistant to immunotherapy are obtained, and the differentially expressed genes are first analyzed. Then, by introducing the Hallmark gene set, 7 refined Hallmark gene sets containing differentially expressed genes are extracted, and the activity values of the 7 refined Hallmark gene sets are calculated. The algorithm is used to mechanically process the activity values and the clinical classification information of sensitivity or resistance corresponding to the samples to construct the HAPIR model.
[0027] 3. This application has discovered a gene set that can accurately predict the probability of drug resistance. Based on this gene set, the probability of drug resistance in patients can be predicted more accurately and efficiently, and it has good application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0029] Figure 1 This is a schematic diagram of the method flow provided by the first aspect of the embodiment of the present invention;
[0030] Figure 2 is a schematic flow chart of the method provided by the second aspect of an embodiment of the present invention;
[0031] Figure 3 is a schematic flow chart of a method provided in the third aspect of an embodiment of the present invention;
[0032] Figure 4 is a schematic diagram of a computer device provided by an embodiment of the present invention;
[0033] Figure 5 is a schematic diagram of the architecture of an exemplary computing device provided by an embodiment of the present invention;
[0034] Figure 6 is a schematic diagram of a storage medium provided by an embodiment of the present invention;
[0035] Figure 7 : is a schematic diagram of the identification process of the refined Hallmark gene set provided by an embodiment of the present invention; wherein, Figure 7 A is the identification of genes related to immunotherapy response. Orange indicates significantly upregulated differentially expressed genes; blue indicates significantly downregulated differentially expressed genes; and gray indicates non-significant genes. Figure 7 B is the functional enrichment analysis results of genes related to immunotherapy response. Figure 7 C is the refined Hallmark gene set and the genes it contains;
[0036] Figure 8 1 is a comparison result diagram of HAPIR and different parameters provided by an embodiment of the present invention;
[0037] Figure 9 is a comparison result diagram of HAPIR provided by an embodiment of the present invention and a model based on genes and gene sets; wherein, Figure 9A is a comparison between HAPIR and gene-based models. The heat map below shows the AUROC value of each dataset, and the bar chart above shows the average AUROC value of the six datasets. Figure 9 B is a comparison of HAPIR and other gene set-based models. The radar chart shows the AUROC value of each data set;
[0038] Figure 10 The results are compared with the characteristics of the effects of immunotherapy reported in the embodiments of the present invention. Figure 10 A is the comparison result of AUROC between HAPIR and other methods in the training set. Figure 10 B is the comparison result of the accuracy of HAPIR and other methods in the training set. Figure 10 C is the AUROC comparison results of HAPIR and other methods in the five validation sets. * indicates p < 0.05; ** indicates p < 0.01; ns indicates no significant difference;
[0039] Figure 11 This is the survival analysis of immunotherapy patients provided by the embodiment of the present invention. Figure 11 A is the Kaplan-Meier survival curve of HAPIR in the Riaz et al. data set, Figure 11 B is the Kaplan-Meier survival curve of HAPIR in the Gide et al. dataset. Figure 11 C is the Kaplan-Meier survival curve for the HAPIR dataset from Lauss et al. Patients were divided into high- and low-resistance groups based on their predicted probability of immunotherapy resistance. DETAILED DESCRIPTION
[0040] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0041] In some of the processes described in the specification and claims of the present invention and the above-mentioned figures, multiple operations that appear in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this article or may be executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish between different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this article are used to distinguish different messages, devices, modules, etc., and do not represent the order of precedence, nor do they limit "first" and "second" to be different types.
[0042] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.
[0043] Figure 1 1 is a flow chart of a method for constructing a model reflecting the degree of immunotherapy response provided by an embodiment of the present invention. Specifically, the method includes the following steps:
[0044] S101, obtain the transcriptome dataset and clinical classification information of the training set samples; the clinical classification information includes sensitivity and resistance;
[0045] In some embodiments, the training set samples include immunotherapy-sensitive samples and drug-resistant samples;
[0046] In one embodiment, the transcriptome dataset is from TIGER;
[0047] S102, performing differential expression analysis on the transcriptome dataset to obtain N differentially expressed genes;
[0048] In some embodiments, the differentially expressed genes include significantly up-regulated and down-regulated differentially expressed genes. N is a natural number greater than 1.
[0049] S103, performing gene set enrichment analysis based on the differentially expressed genes to identify M significantly enriched Hallmark gene sets;
[0050] In some embodiments, M is a natural number greater than 1. All Hallmark gene sets (50) are directly obtained from the database, and then the differentially expressed genes are used for enrichment analysis.
[0051] S104, taking the intersection of a single Hallmark gene set and the differentially expressed genes to obtain M refined Hallmark gene sets; and calculating the activity value of each refined Hallmark gene set;
[0052] In some embodiments, the activity value calculation method includes: for each sample, calculating the expression level of each gene in the refined Hallmark gene set, and weighting it according to its importance in the gene set; and for each sample, summing the weighted expression levels of all genes in the gene set to obtain the activity value of the refined Hallmark gene set corresponding to the sample. The weight can be the frequency of occurrence of the gene in the gene set or other indicators, which are not limited here.
[0053] S105, inputting the activity value into a machine learning model to obtain a predicted classification result, comparing it with the clinical classification information, optimizing the model based on the comparison result, and constructing a model reflecting the degree of immunotherapy response.
[0054] In some embodiments, the machine learning model uses any one or more of the following methods: glm function, lm function, nls function, glmnet function, coxph function, mgcv function, randomForest function, gbm function, xgboost function, neuralnet function, preferably glm function.
[0055] The second aspect of the present application discloses a method for predicting drug resistance probability based on HAPIR, the method comprising:
[0056] S201, obtaining transcriptome data of the subject;
[0057] S202, inputting the transcriptome data into the model reflecting the degree of immunotherapy response described in the first aspect of the present application, and outputting a result indicating whether the probability of drug resistance of the subject is high or low.
[0058] In some embodiments, the method further includes: if a result with a high probability of drug resistance is output, an auxiliary prediction result with a shorter subject survival period is obtained; if a result with a low probability of drug resistance is output, an auxiliary prediction result with a longer subject survival period is obtained.
[0059] The third aspect of the present application discloses a method for predicting drug resistance probability based on HAPIR, the method comprising:
[0060] S301, obtaining transcriptome data of the test subject; the transcriptome data includes any one or more of the following genes: S100A1, TYR, SERPINE2, SGCD, PHLDA1;
[0061] In some embodiments, the transcriptome data further includes any one or more of the following genes: CLU, CD79A, C3, CXCL13, SERPINA1;
[0062] In some embodiments, the transcriptome data further includes any one or more other genes in the refined Hallmark gene set, and the specific gene names are shown in Table 1.
[0063] S302, inputting the transcriptome data into the model reflecting the degree of immunotherapy response described in the first aspect of the present application, and outputting a result indicating whether the probability of drug resistance of the subject is high or low.
[0064] In some embodiments, the method further includes: if a result with a high probability of drug resistance is output, an auxiliary prediction result with a shorter subject survival period is obtained; if a result with a low probability of drug resistance is output, an auxiliary prediction result with a longer subject survival period is obtained.
[0065] In some embodiments, the auxiliary prediction results include but are not limited to paper or electronic report forms. The results are only obtained by the intelligent machine based on the analysis of relevant data of the subject, and are only used as a reference for medical staff, and are not the final diagnosis results of the subject.
[0066] In some embodiments, the terms "subject," "subject," or "test sample" as used herein refer to any animal (e.g., mammal), including but not limited to humans, non-human primates, rodents, etc., that will be the recipient of a specific treatment. Generally, the terms "subject" and "patient" are used interchangeably herein when referring to a human subject. Preferably, the subject is a human. In some embodiments, the subject is a patient clinically undergoing a prognostic assessment.
[0067] Figure 4 is a schematic diagram of a computer device provided by an embodiment of the present invention, such as Figure 4 As shown, the device 2000 may include: one or more processors 2010, and one or more memories 2020; wherein the memories store computer-readable codes, and when the computer-readable codes are run by the one or more processors, they may execute the method described above.
[0068] The processor in this embodiment can be an integrated circuit chip with signal processing capabilities. The above-mentioned processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. It can implement or execute the various methods, operations, and logic block diagrams disclosed in the embodiments of the present disclosure. The general-purpose processor can be a microprocessor or any conventional processor, etc., and can be an X86 architecture or an ARM architecture.
[0069] In general, various example embodiments of the present disclosure may be implemented in hardware or dedicated circuitry, software, firmware, logic, or any combination thereof. Certain aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device. When various aspects of the disclosed embodiments are illustrated or described as block diagrams, flow charts, or using some other graphical representation, it will be understood that the blocks, devices, systems, techniques, or methods described herein may be implemented, as non-limiting examples, in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or a controller or other computing device, or some combination thereof.
[0070] For example, the method or apparatus according to the embodiment of the present disclosure may also be implemented by Figure 5 The architecture of the computing device 3000 shown in FIG. Figure 5 As shown, the computing device 3000 may include a bus 3010, one or more CPUs 3020, a read-only memory (ROM) 3030, a random access memory (RAM) 3040, a communication port 3050 connected to a network, an input / output component 3060, a hard disk 3070, etc. The storage device in the computing device 3000, such as the ROM 3030 or the hard disk 3070, may store various data or files used for processing and / or communication of the method provided in the present disclosure, as well as program instructions executed by the CPU. The computing device 3000 may also include a user interface 3080. Of course, Figure 5 The architecture shown is only exemplary and can be omitted according to actual needs when implementing different devices. Figure 5 One or more components of a computing device are shown.
[0071] The embodiment of the present invention further provides a computer-readable storage medium, such as Figure 6As shown, it is a schematic diagram of a storage medium 4000 provided in an embodiment of the present invention, and computer-readable instructions 4010 are stored on the computer storage medium 4020. When the computer-readable instructions 4010 are executed by the processor, the method according to the embodiment of the present disclosure described with reference to the above figures can be executed. The computer-readable storage medium in the embodiment of the present disclosure can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. The non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM) or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus random access memory (DR RAM). It should be noted that the memory of the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory. It should be noted that the memory of the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0072] The embodiments of the present disclosure further provide a computer program product or system, including a computer program, which implements the steps of the above method when executed by a processor.
[0073] In some embodiments, this embodiment further discloses a system for constructing a model reflecting the degree of immunotherapy response, the system comprising:
[0074] A first acquisition module is used or configured to acquire a transcriptome dataset and clinical classification information of a training set sample; the clinical classification information includes sensitivity and resistance;
[0075] a differentially expressed gene screening module, used for or configured to perform differential expression analysis on the transcriptome dataset to obtain N differentially expressed genes;
[0076] a gene set identification module, configured to perform gene set enrichment analysis based on the differentially expressed genes and identify M significantly enriched Hallmark gene sets;
[0077] an activity value calculation module, configured to or configured to take the intersection of a single Hallmark gene set and the differentially expressed genes to obtain M refined Hallmark gene sets; and calculate the activity value of each refined Hallmark gene set;
[0078] The model training module is used or configured to input the activity value into the machine learning model to obtain the predicted classification result, compare it with the clinical classification information, optimize the model according to the comparison result, and construct a model reflecting the degree of immunotherapy response.
[0079] In some embodiments, this embodiment further discloses a system for predicting drug resistance probability based on HAPIR, the system comprising:
[0080] A second acquisition module is used or configured to acquire transcriptome data of the subject;
[0081] The first drug resistance probability output module is used or configured to input the transcriptome data into the model reflecting the degree of immunotherapy response described in the first aspect of the present application, and output a result of whether the test subject has a high or low probability of drug resistance.
[0082] In some embodiments, the system further includes: a first survival prediction module, which is used or configured to obtain an auxiliary prediction result of a shorter subject survival if a result with a high probability of drug resistance is output; and to obtain an auxiliary prediction result of a longer subject survival if a result with a low probability of drug resistance is output.
[0083] In some embodiments, this embodiment further discloses a system for predicting drug resistance probability based on HAPIR, the system comprising:
[0084] A third acquisition module is used or configured to acquire transcriptome data of the subject; the transcriptome data includes any one or more of the following genes: S100A1, TYR, SERPINE2, SGCD, PHLDA1;
[0085] The second drug resistance probability output module is used or configured to input the transcriptome data into the model reflecting the degree of immunotherapy response described in the first aspect of the present application, and output a result of whether the test subject has a high or low probability of drug resistance.
[0086] In some embodiments, the system further includes: a second survival prediction module, which is used or configured to obtain an auxiliary prediction result of a shorter subject survival if a result with a high probability of drug resistance is output; and to obtain an auxiliary prediction result of a longer subject survival if a result with a low probability of drug resistance is output.
[0087] Example 1 Construction of a model for predicting the degree of immunotherapy response
[0088] 1.1 Experimental Materials
[0089] Next-generation sequencing data (RNA-seq) of immunotherapy-sensitive and -resistant patients.
[0090] 1.2 Experimental methods
[0091] (1) The Limma method was used to perform differential expression (DE) analysis on immunotherapy-sensitive and -resistant samples, and 200 significant (p < 0.05) most upregulated and most downregulated differentially expressed genes (DEGs) were obtained, respectively.
[0092] (2) The clusterProfiler R package was used to perform gene set enrichment analysis on these 400 differentially expressed genes, and seven significantly enriched Hallmark gene sets were identified (FDR < 0.05).
[0093] (3) For each significantly enriched Hallmark gene set, only 400 genes in the differentially expressed gene set were retained. These enriched Hallmark gene sets were refined to obtain 7 "refined Hallmark gene sets".
[0094] (4) Use the AUCell algorithm to score each refined Hallmark gene set to obtain the refined Hallmark gene set activity.
[0095] (5) A logistic regression model was constructed using the glm function to predict the response to immunotherapy, and the model was named HAPIR.
[0096] 1.3 Experimental Results
[0097] Genes associated with immunotherapy response such as Figure 7 A. Functional enrichment analysis results of genes related to immunotherapy response are shown in Figure 7 As shown in B, there are 7 significantly enriched Hallmark gene sets. The 7 refined Hallmark gene sets include genes such as Figure 7 C and Table 1.
[0098] Table 1 Name, number of genes included, and gene names of the refined Hallmark gene set
[0099]
[0100] Example 2 Validation of the HAPIR model
[0101] 1.1 Experimental Materials
[0102] Six data sets were downloaded from the TIGER database (http: / / tiger.canceromics.org / # / download) to verify the accuracy of HAPIR. As shown in Table 2, each data set contains immunotherapy drug response signatures and gene expression information, specifically including the following data:
[0103] Table 2. Summary of data for HAPIR model validation
[0104]
[0105] 1.2 Experimental methods
[0106] (1) First, we compared HAPIR with models with other parameters, using the Riaz et al. dataset as the training set and the remaining datasets as the validation sets. For the training set, we used the ten-fold cross-validation method for validation. We compared the accuracy of different models and used the area under the receiver operating characteristic curve (AUROC) as the evaluation metric. The specific parameters are as follows: number of differentially expressed genes (100, 200, 300, 400, 500); Hallmark gene set scoring method (AUCell, ssGSEA, GSVA, Addmodulescore, gene expression average); machine learning model (logistic regression, elastic net regression, ridge regression, XGboost, support vector machine, random forest, K-nearest neighbor algorithm).
[0107] (2) Secondly, we also compared HAPIR with gene-based models. For genes, HAPIR involved a refined set of 77 genes and 400 differentially expressed genes for predicting immunotherapy response, and combined them with seven machine learning models (logistic regression, elastic net regression, ridge regression, XGboost, support vector machine, random forest, and K-nearest neighbor algorithm) for prediction. We used AUROC values to compare the prediction performance of different methods on six data sets.
[0108] (3) Again, we compared HAPIR with models based on other gene sets. For gene sets, the complete 50 Hallmark gene set (50_Hallmark) and the original 7 Hallmark gene set (7_Hallmark) were involved. In addition, we collected 386 tumor microenvironment-related gene sets (TME_Bio) and identified 32 refined gene sets using the same workflow as HAPIR. We also combined Hallmark and tumor microenvironment-related gene sets (TME_Bio+50_Hallmark) and identified 40 refined gene sets using the same method. The immunotherapy response of these gene sets was predicted using the AUCell scoring method and the logistic regression model. We used the AUROC value to compare the prediction performance of different methods on the six data sets.
[0109] (4) Finally, to further evaluate the predictive performance, we compared HAPIR with reported immunotherapy response signatures, including four pan-cancer signatures (PD-1, PD-L1, C_ECM, and ADO) and nine melanoma-specific signatures (KDM5A, CD8_SF, B_cell_Helmink, ImmuneCells, CRMA, T_cell_exclusion, IMPRES, MPS, and Myeloid_DC). Detailed information on these signatures and their corresponding algorithms is shown in Table 3. For HAPIR, 0.5 was used as the cutoff value for distinguishing between sensitive and resistant patients. For each compared method, the average value was used to distinguish between sensitive and resistant patients. We used accuracy and AUROC values to compare the predictive performance of different methods on the six datasets.
[0110] (4) Survival analysis of immunotherapy patients was performed using the survival R package, and Kaplan-Meier curves were drawn using the survminer R package. Receiver operating characteristic (ROC) curves were drawn using the pROC R package, and AUROC values were calculated.
[0111] Table 3. Summary of reported imaging features of immunotherapy
[0112]
[0113] 1.3 Experimental Results
[0114] To demonstrate the superiority of the HAPIR parameter, we conducted a comparative analysis with other parameters. The average AUROC value of the prediction results of HAPIR in 6 data sets was about 0.8, which was better than the results predicted by other parameters ( Figure 8 ).
[0115] Then, to demonstrate the superiority of the refined Hallmark gene set, we compared HAPIR with models based on genes and other functional gene sets. For the gene-based models, HAPIR showed the highest average AUROC values in the six datasets ( Figure 9 A), indicating that the use of refined Hallmark gene set activity scores provides superior predictive power compared to single gene expression. For gene set comparison, we first compared the performance of the refined Hallmark gene set with the 7 original Hallmark gene sets (7_Hallmark) and the 50 Hallmark gene sets (50_Hallmark). To reduce the complexity of the comparison, we used the same parameters "AUCell" and "Logistic Regression" as HAPIR to predict the degree of response to immunotherapy. The results showed that HAPIR showed the highest AUROC value in all six datasets ( Figure 9 B). This highlights the effectiveness of gene set refinement in improving model performance. In addition, we also compared with a model based on the TME-related gene set (TME_Bio) and a model combining TME_Bio with the 50 Hallmark gene set (TME_Bio+50_Hallmark), and HAPIR achieved the highest AUROC ( Figure 9 B). Taken together, these results demonstrate that HAPIR provides superior predictive performance compared to models based on genes, Hallmark gene sets, and TME-associated gene sets, enhancing its utility in predicting immunotherapy response.
[0116] Next, we compared HAPIR with other previously reported signature biomarkers of immunotherapy response, including melanoma-specific markers such as KDM5A and CD8_SF, and pan-cancer biomarkers such as C_ECM and ADO. In the training set, HAPIR achieved the highest AUROC and accuracy ( Figure 10 A and Figure 10 B). In the external validation set, HAPIR still achieved the best prediction performance ( Figure 10 C). These results suggest that HAPIR outperforms previously reported immunotherapy response-associated biomarkers in predicting immunotherapy response and patient survival.
[0117] Finally, we performed a survival analysis, and the results showed that among the three data sets with survival information, the survival time of patients with a high probability of drug resistance was significantly (p < 0.05) lower than that of patients with a low probability ( Figure 11 ). These results further illustrate the effectiveness of the HAPIR model.
[0118] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architectures, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the boxes can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified functions or operations, or can be implemented using a combination of dedicated hardware and computer instructions.
[0119] In general, various example embodiments of the present disclosure may be implemented in hardware or dedicated circuitry, software, firmware, logic, or any combination thereof. Certain aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device. When various aspects of the disclosed embodiments are illustrated or described as block diagrams, flow charts, or using some other graphical representation, it will be understood that the blocks, devices, systems, techniques, or methods described herein may be implemented, as non-limiting examples, in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or a controller or other computing device, or some combination thereof.
[0120] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0121] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0122] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0123] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0124] The exemplary embodiments of the present disclosure described in detail above are merely illustrative and not restrictive. Those skilled in the art will appreciate that various modifications and combinations may be made to these embodiments or their features without departing from the principles and spirit of the present disclosure, and such modifications should fall within the scope of the present disclosure.
Claims
1. A method for constructing a model reflecting the degree of immunotherapy response, characterized in that: The method comprises: S101, obtain the transcriptome dataset and clinical classification information of the training set samples; the clinical classification information includes sensitivity and resistance; S102, performing differential expression analysis on the transcriptome datasets of sensitive and resistant samples to obtain N differentially expressed genes; S103, performing gene set enrichment analysis based on the differentially expressed genes to identify M significantly enriched Hallmark gene sets; S104, taking the intersection of a single Hallmark gene set and the differentially expressed genes to obtain M refined Hallmark gene sets; calculating the activity value of each refined Hallmark gene set using an AUCell algorithm; S105, inputting the activity value into a machine learning model constructed by logistic regression to obtain a predicted classification result, comparing it with the clinical classification information, optimizing the model based on the comparison result, and constructing a model reflecting the degree of immunotherapy response.
2. The method for constructing a model reflecting the degree of immunotherapy response according to claim 1, characterized in that: The differentially expressed genes include significantly up-regulated and down-regulated differentially expressed genes.
3. A method for predicting drug resistance probability based on HAPIR, characterized in that: The method comprises: S201, obtaining transcriptome data of the subject; S202, inputting the transcriptome data into the HAPIR, and outputting a result indicating whether the subject has a high or low probability of drug resistance, wherein the HAPIR is a model reflecting the degree of immunotherapy response obtained by the construction method according to any one of claims 1-2.
4. The method for predicting drug resistance probability based on HAPIR according to claim 3, characterized in that: The method further includes: if a result with a high probability of drug resistance is output, obtaining an auxiliary prediction result with a shorter survival period of the subject; if a result with a low probability of drug resistance is output, obtaining an auxiliary prediction result with a longer survival period of the subject.
5. A method for predicting drug resistance probability based on HAPIR, characterized in that: The method comprises: S301, obtaining transcriptome data of a subject; the transcriptome data includes any one or more of the following genes: S100A1, TYR, SERPINE2, SGCD, PHLDA1; S302, inputting the transcriptome data into the HAPIR, and outputting a result indicating whether the subject has a high or low probability of drug resistance, wherein the HAPIR is a model reflecting the degree of immunotherapy response obtained by the construction method according to any one of claims 1-2.
6. A computer device, characterized in that: The device comprises: a memory and a processor; the memory is used to store a computer program; the processor executes the computer program to implement the steps of the method according to any one of claims 1 to 5.
7. A computer-readable storage medium, characterized in that A computer program is stored thereon, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
8. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Method, system and equipment for predicting drug resistance state of drug
CN118380163A
Biomarkers for predicting response to immunotherapy and uses thereof
CN119339805A