Establishment and validation of a gene pair predicting response to anti-PD-1 immunotherapy in melanoma

By screening gene pair combinations that affect melanoma anti-PD-1 immunotherapy response based on the rank relationship of relative gene expression, the problem of difficulty in cross-platform verification was solved, and personalized prediction of patient drug response was achieved, which has clinical application value.

CN119560003BActive Publication Date: 2025-09-26GANNAN MEDICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411681846.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-22
Publication Date
2025-09-26
Estimated Expiration
2044-11-22

AI Technical Summary

Technical Problem

In existing technologies, the efficacy of anti-PD-1 immunotherapy is related to multiple molecular markers, but its application is easily affected by factors such as expression levels and batch effects, resulting in greater difficulty in cross-dataset and cross-platform verification, which affects the actual clinical application of biomarkers.

Method used

A method based on the rank relationship of relative gene expression within the sample was used to screen out gene pair combinations with the maximum sample coverage through a greedy algorithm as predictive markers for anti-PD-1 immunotherapy response in melanoma, and personalized predictions were performed using candidate predictive gene pairs.

Benefits of technology

It can be robustly applied to independent clinical samples evaluated by different laboratories at an individualized level, accurately predicting the occurrence and development of cancer, achieving truly personalized prediction of patient drug responses, and has clinical application value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119560003B_ABST
    Figure CN119560003B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for establishing and validating gene pairs for predicting the response of melanoma anti-PD-1 immunotherapy. A melanoma anti-PD-1 immunotherapy dataset is selected as a training set, and a gene pair matrix is ​​formed based on the relative expression rank relationship of genes within the sample to screen candidate prediction gene pairs. Based on the candidate prediction gene pairs, a greedy algorithm is used to find the gene pair combination with the maximum sample coverage, and the gene pair combination with the best prediction efficiency is selected as the prediction gene pair. For a sample to be predicted, the gene pair in the prediction gene pair combination is calculated. a >gene b According to the voting rules, if gene a >gene b If the number of gene pairs in the prediction is more than half, the sample to be predicted can be classified as a responsive sample, otherwise it is a non-responsive sample. This has been verified in sample validation sets from other platforms.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of prediction of drug response to anti-PD-1 immunotherapy for melanoma, and specifically relates to a method for establishing and verifying a gene pair for predicting response to anti-PD-1 immunotherapy for melanoma. Background Art

[0002] Melanoma is a highly aggressive malignant tumor that often develops in areas such as the skin, mouth, intestines, and eyes. It is prone to metastasis and difficult to treat. In recent years, anti-PD-1 immunotherapy has been used to treat patients with advanced melanoma with some success. However, due to the heterogeneity of tumors and individual patients, anti-PD-1 immunotherapy is not suitable for all patients. Some patients do not benefit from this treatment and may even experience accelerated cancer progression. Furthermore, anti-PD-1 immunotherapy is expensive, resulting in significant medical costs.

[0003] Studies have shown that the efficacy of anti-PD-1 immunotherapy is correlated with multiple molecular markers (such as PD-L1, LAG-3, and CXCR3), tumor mutational burden, tumor lymphocyte infiltration, and neoantigens. However, the practical application of these biomarkers is susceptible to factors such as expression levels and batch effects, making cross-dataset and cross-platform validation difficult. These factors can easily hinder the actual clinical application of biomarkers. Summary of the Invention

[0004] In view of the fact that the efficacy of anti-PD-1 immunotherapy in the existing technology is related to multiple molecular markers (such as PD-L1, LAG-3, CXCR3, etc.), tumor mutation load, tumor lymphocyte infiltration level, new antigens and other factors, but the actual application of these biomarkers is easily affected by factors such as expression level and batch effect, resulting in greater difficulty in cross-dataset and cross-platform verification, the present invention provides a method for establishing and verifying gene pairs that predict the response to melanoma anti-PD-1 immunotherapy, and provides a method for finding stable predictive drug response markers with clinical translation value in melanoma anti-PD-1 immunotherapy samples.

[0005] In order to solve the above technical problems, the present invention adopts the following technical solutions:

[0006] A method for establishing a gene pair predicting response to anti-PD-1 immunotherapy for melanoma comprises the following steps:

[0007] Step A: obtaining expression profile data and clinical information of melanoma patients and melanoma anti-PD-1 immunotherapy patients from the Gene Expression Comprehensive Database and The Cancer Genome Atlas, and preprocessing the expression profile data and clinical information;

[0008] The melanoma expression profile datasets from different detection platforms were divided into two data subsets, one melanoma anti-PD-1 immunotherapy dataset was used as the training set, and the other data subset was used as the validation set;

[0009] Step B: Pairwise combinations of genes in the training set were performed to compare the relative expression rank relationships of gene pairs in progressive disease (PD) samples and partial or complete response (PRCR) samples. Gene pairs with significantly reversed relative expression rank relationships were screened to identify candidate predictive gene pairs with the potential to predict response to anti-PD-1 immunotherapy in melanoma.

[0010] Step C: Use a greedy algorithm to find the gene pair combination with the maximum sample coverage;

[0011] Starting from each gene pair in the candidate prediction gene pairs, use it as a seed and add the remaining gene pairs to the current combination in sequence; after each addition, calculate the sample coverage of the new combination; this process continues until the sample coverage can no longer be increased by adding candidate gene pairs; the final combination is the characteristic gene pair combination with the maximum sample coverage identified based on the candidate prediction gene pairs;

[0012] Calculate the frequency of each characteristic gene pair in all combinations. The frequency reflects the importance of the characteristic pair. Select the best prediction gene pair as the final prediction gene pair.

[0013] For a sample to be predicted, calculate the gene in the predicted gene pair combination a >gene b The number of gene pairs; According to the voting rules, if gene a >gene b If the number of gene pairs is more than half, the sample to be predicted can be classified as a responsive sample (responder, R), otherwise it can be classified as a non-responder (non-responder, NR); the prediction gene pair is then verified using the validation set, and the obtained prediction gene pair is used as a marker for predicting the response to anti-PD-1 immunotherapy for melanoma.

[0014] Furthermore, step A involves obtaining expression profile data and clinical information of melanoma patients and melanoma anti-PD-1 immunotherapy patients from a comprehensive gene expression database and the Cancer Genome Atlas, and preprocessing the expression profile data and clinical information, specifically comprising the following steps:

[0015] Step A1: Download expression profile data from the GEO database and map probe IDs to gene IDs using the platform annotation file;

[0016] Step A2: Studies have shown that during anti-PD-1 immunotherapy for melanoma, the subsequent efficacy of samples in the stable disease (SD) state is unstable and may develop into PD samples or PRCR samples during subsequent treatment. Therefore, SD samples are removed from the training set;

[0017] Step A3: Remove low-expression genes, which are genes whose expression is absent or zero in more than 70% of the sample genes.

[0018] Furthermore, in step A1, the expression profile data are downloaded from the GEO database, and the probe ID is mapped to the gene ID using the platform annotation file. The melanoma anti-PD-1 immunotherapy expression profile dataset GSE91061 (including samples before and during treatment) from the GPL9052 detection platform is used as the training set (GSE91061_Pre) and validation set (GSE91061_On), respectively, and the other three melanoma expression profile datasets are used as independent validation sets; the three validation sets are the melanoma expression profile data from TCGA, GSE115821, and GSE22153 of the detection platforms GPL11154, GPL18573, and GPL6102.

[0019] Furthermore, the genes in the training set are paired together as described in step B, and the differences in the relative expression rank relationships of the gene pairs in the PD samples and PRCR samples are compared. Gene pairs with significantly reversed relative expression rank relationships between genes are screened to determine candidate predictive gene pairs with potential ability to predict the response to anti-PD-1 immunotherapy for melanoma, specifically including the following steps:

[0020] Step B1: Based on the rank relationship of gene expression levels between gene pairs, the gene pairs with stable expression between genes in PD samples were screened as reference stable gene pairs. The genes in PD samples were combined in pairs to form n(n-1) / 2 gene pairs. According to the calculation formula (1), P(gene a >gene b ) ≥ 0.8 as the reference stable gene pairs;

[0021] (1)

[0022] in m represents the total number of PD samples, k Represents genes in PD samples a >gene b The number of samples, gene a>gene b Indicates gene a and genes b Rank relationship of expression levels;

[0023] Step B2: Based on the reference stable gene pairs obtained in PD samples (gene a , gene b ), calculate the gene of each reference stable gene pair in PD samples and PRCR samples a >gene b Number of samples m 1. n 1 and gene a <gene b Number of samples m 2. n 2. Then, Fisher's exact test was used to determine whether there was a significant difference in the REO distribution of the reference stable gene pairs between PD and PRCR samples. Gene pairs with a false positive discovery rate (FDR) of less than 5% were defined as reversal gene pairs. For each reversal gene pair, the formula △P=P PD (gene a >gene b )- P PRCR (gene a >gene b ) to calculate the reversal rate △P; the larger the △P value, the greater the difference in REO of the gene pair; △P=1 means that the REO of the gene pair in PD samples is the same as that of the gene pair. a >gene b , in PRCR samples, both genes a ≤gene b ; Given a threshold △Pt such as 0.7, the gene pairs that satisfy △P>△Pt are taken as candidate prediction gene pairs.

[0024] Furthermore, the step C specifically includes the following steps:

[0025] A greedy algorithm is used to find the gene pair combination with the maximum sample coverage. Starting from each gene pair in the candidate prediction gene pair, it is used as a seed and the remaining gene pairs are added to the current combination in sequence. After each addition, the sample coverage of the new combination is calculated. This process continues until the sample coverage can no longer be increased by adding candidate gene pairs. The final combination is the characteristic gene pair combination with the maximum sample coverage identified based on the candidate prediction gene pairs. The frequency of each characteristic gene pair in all combinations is calculated. The frequency reflects the importance of the characteristic pair. The optimal prediction gene pair is selected as the final prediction gene pair. For a sample to be predicted, the gene pair in the prediction gene pair combination is calculated.a >gene b The number of gene pairs; According to the voting rules, if gene a >gene b If the number of gene pairs is more than half, the sample to be predicted can be classified as a responder (R), otherwise it can be classified as a non-responder (NR). The prediction gene pair is then verified using the validation set, and the obtained prediction gene pair is used as a marker for predicting the response to anti-PD-1 immunotherapy for melanoma; including:

[0026] The REO reference pattern of a gene pair is defined as the gene expressed in PD samples. a >gene b Or in PRCR samples, it appears as a gene a ≤gene b When at least one gene pair in a sample shows the gene pair REO reference pattern, the combination is considered to cover the sample, and the percentage of covered samples to the total number of samples is calculated;

[0027] To screen a set of characteristic gene pairs that can predict therapeutic efficacy, a greedy algorithm is used to find the gene pair combination with the maximum sample coverage. Starting from each gene pair in the candidate prediction gene pair, it is used as a seed and the remaining gene pairs are added to the current combination in sequence. After each addition, the sample coverage of the new combination is calculated. This process continues until the sample coverage can no longer be increased by adding candidate gene pairs. The final combination is the characteristic gene pair combination with the maximum sample coverage identified based on the candidate prediction gene pairs.

[0028] Calculate the frequency of each characteristic gene pair in all combinations. The frequency reflects the importance of the characteristic pair. Sort the frequencies from large to small and select the ones with the highest frequency first. k feature pairs; k It is a predefined parameter with a value of 1~ n ( n is the total number of characteristic gene pairs); for the selected k feature pairs, calculate the geometric mean of their negative predictive value and positive predictive value; select the one that first reaches the maximum geometric mean and tends to be stable. k gene pairs as the final prediction gene pair combination; for a sample to be predicted, calculate the gene a >gene b The number of gene pairs; According to the voting rules, if gene a >gene bIf the number of gene pairs is more than half, the sample to be predicted can be classified as a responsive sample (responder, R), otherwise it can be classified as a non-responder (non-responder, NR).

[0029] The present invention also relates to a method for validating a gene pair predicting response to anti-PD-1 immunotherapy for melanoma, comprising the following steps:

[0030] Step A: Obtain expression profile data and clinical information of melanoma patients and melanoma anti-PD-1 immunotherapy patients from the Gene Expression Comprehensive Database and The Cancer Genome Atlas, preprocess the expression profile data and clinical information, and divide the melanoma expression profile datasets from different detection platforms into two data subsets, one melanoma anti-PD-1 immunotherapy data subset as a training set and the other data subset as a validation set;

[0031] Step B: Pairwise combinations of genes in the training set were performed to compare the relative expression rank relationships of gene pairs in PD samples and PRCR samples. Gene pairs with significantly reversed relative expression rank relationships were screened to identify candidate predictive gene pairs with the potential to predict response to anti-PD-1 immunotherapy in melanoma.

[0032] Step C: Use the greedy algorithm to find the gene pair combination with the maximum sample coverage; select the optimal prediction gene pair combination as the final prediction gene pair; for a sample to be predicted, calculate the gene a >gene b The number of gene pairs; According to the voting rules, if gene a >gene b If the number of gene pairs is more than half, the sample to be predicted can be classified as a responder (R), otherwise it can be classified as a non-responder (NR); the prediction gene pair is then verified using the validation set, and the obtained prediction gene pair is used as a marker for predicting the response to anti-PD-1 immunotherapy for melanoma;

[0033] Step D: Validating the melanoma anti-PD-1 immunotherapy response prediction gene pair in a validation set, specifically comprising: calculating evaluation indicators (accuracy, positive predictive value, negative predictive value, specificity, sensitivity) for a data set containing anti-PD-1 immunotherapy drug responses using the melanoma anti-PD-1 immunotherapy response prediction gene pair; performing survival analysis and multivariate Cox regression analysis on the melanoma anti-PD-1 immunotherapy response prediction gene pair; performing immune infiltration analysis on the melanoma anti-PD-1 immunotherapy response prediction gene pair; and performing immunotherapy response analysis on the melanoma anti-PD-1 immunotherapy response prediction gene pair.

[0034] Furthermore, the calculation of evaluation indicators (accuracy, positive predictive value, negative predictive value, specificity, and sensitivity) for a dataset containing anti-PD-1 immunotherapy drug responses using the melanoma anti-PD-1 immunotherapy response prediction gene pairs described in step D specifically includes the following steps:

[0035] Use a 2x2 confusion matrix to represent the prediction results of gene pairs, where the rows represent the actual labels and the columns represent the predicted labels;

[0036] According to Table 1, the confusion matrix is ​​established and the prediction index calculation formulas are:

[0037] 、 、 、 、

[0038]

[0039] Note: TP (True positive); TN (True negative); FN (False negative); FP (Falsepositive); P (positive); N (negative)

[0040] Furthermore, the survival analysis and multivariate Cox regression analysis of the melanoma anti-PD-1 immunotherapy response prediction gene pairs described in step D specifically include the following steps:

[0041] Kaplan-Meier (KM) survival analysis was performed using the R packages “survival” and “survminer” to verify whether there was a survival difference between the two subgroups predicted by the predicted gene pairs in the validation set; the expression rank relationship of the gene pairs in the predicted gene pair combination in the sample was gene a Greater than gene b A multivariate Cox regression analysis was performed on the validation set, incorporating factors such as age, gender, and predictive gene pair scores.

[0042] Furthermore, the immune infiltration analysis of the melanoma anti-PD-1 immunotherapy response prediction gene pair described in step D specifically includes the following steps:

[0043] The ssGSEA (Single Sample Gene Set Enrichment Analysis) algorithm was used to evaluate the immune cell content between the response group and the non-response group predicted by the prediction gene pairs in the validation set. The immune cells between the two subgroups predicted by the prediction gene pairs in the validation set included 28 types: activated CD4+ / CD8+ T cells, central memory CD4+ / CD8+ T cells, effector memory CD4+ / CD8+ T cells, γδT cells, helper T cells (Th1), Th2 cells, Th17 cells, regulatory T cells (Treg), follicular helper T cells (Tfh), activated B cells, naive B cells, memory B cells, activated dendritic cells (DC), naive DC, plasmacytoid DC, natural killer cells (NK), natural killer T cells (NKT), myeloid-derived suppressor cells (MDSC), macrophages, monocytes, mast cells, eosinophils, neutrophils, CD56bright NK cells, CD56dim NK cells. cells; the Wilcoxon test was used for comparison.

[0044] Furthermore, the immunotherapy response analysis of the melanoma anti-PD-1 immunotherapy response prediction gene pair described in step D specifically includes the following steps:

[0045] Methods: Ten published immune checkpoint-related genes (CD274 (PD-1), CTLA-4, CXCR3, LAG3, PDCD1 (PD-L1), CD80, ICOS, IFNG, IL10, and TNFRSF9) were collected and their expression levels in the response and non-response groups predicted by the validation set were evaluated. The differences in the levels of immunotherapy response-related genes between the two subgroups predicted by the prediction gene pairs in the validation set were investigated by comparing their expression levels. The Wilcoxon test was used for comparison.

[0046] Compared with the prior art, the present invention has the following beneficial effects:

[0047] The method for establishing and verifying a gene pair for predicting the response of melanoma to anti-PD-1 immunotherapy described in the present invention is also a method for establishing and verifying a gene pair for predicting the response of melanoma to anti-PD-1 immunotherapy. A melanoma anti-PD-1 immunotherapy dataset is selected as a training set, and a gene pair matrix is ​​formed based on the rank relationship of relative expression of genes within the sample to screen candidate prediction gene pairs. Based on the candidate prediction gene pairs, a greedy algorithm is used to find the gene pair combination with the maximum sample coverage, and the gene pair combination with the best prediction efficiency is screened as the prediction gene pair. For a sample to be predicted, the gene pair matrix in the prediction gene pair combination is calculated. a >gene b According to the voting rules, if gene a >gene b If more than half of the gene pairs are positive, the sample to be predicted is classified as a responder (R), while if not, it is classified as a non-responder (NR). This approach has been validated in validation sets of samples from other platforms. This approach, based on the relative rank relationships between genes, can be robustly applied to independent clinical samples evaluated in different laboratories at an individualized level, accurately predicting the development and progression of cancer. Predictive gene pairs enable truly personalized prediction of patient drug response and have clinical application value in predicting response to anti-PD-1 immunotherapy in melanoma. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 This is a schematic diagram of a process for establishing a gene pair predicting response to anti-PD-1 immunotherapy for melanoma according to Example 1 of the present invention;

[0049] Figure 2 Schematic diagram of the process of establishing a training set and a validation set for anti-PD-1 immunotherapy for melanoma according to Example 2 of the present invention;

[0050] Figure 3 This is a flow chart of a method for validating a gene pair predicting response to anti-PD-1 immunotherapy for melanoma as described in Example 2 of the present invention;

[0051] Figure 4 This is a schematic diagram of the survival analysis and multivariate Cox regression analysis of genes predicting response to anti-PD-1 immunotherapy for melanoma described in Example 2 of the present invention;

[0052] Figure 5 This is a schematic diagram of immune infiltration analysis of genes predicting response to anti-PD-1 immunotherapy for melanoma according to Example 2 of the present invention;

[0053] Figure 6This is a schematic diagram of the immunotherapy response analysis of melanoma anti-PD-1 immunotherapy response prediction genes described in Example 2 of the present invention. DETAILED DESCRIPTION

[0054] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0055] The present application provides a method for establishing and validating gene pairs for predicting the response to anti-PD-1 immunotherapy of melanoma. Biomarkers related to the response to anti-PD-1 immunotherapy are identified based on the relative expression rank relationship between genes within a sample. The resulting set of gene pairs is used as the gene pairs for predicting the response to anti-PD-1 immunotherapy of melanoma.

[0056] Example 1:

[0057] like Figure 1 As shown, a method for establishing a gene pair for predicting the response to anti-PD-1 immunotherapy of melanoma is provided, specifically comprising:

[0058] Step 102: obtaining expression profile data and clinical information of melanoma patients and melanoma anti-PD-1 immunotherapy patients from a comprehensive gene expression database and the Cancer Genome Atlas, preprocessing the expression profile data and clinical information, and dividing the melanoma expression profile datasets from different detection platforms into two data subsets, one melanoma anti-PD-1 immunotherapy data subset as a training set, and the other data subset as a validation set;

[0059] Gene expression profiles from four independent datasets, comprising 639 samples, were retrospectively collected. The data analyzed were from the Gene Expression Omnibus (GEO, http: / / www.ncbi.nlm.nih.gov / geo / ) and The Cancer Genome Atlas (TCGA, https: / / portal.gdc.cancer.gov / ). The GSE91061 melanoma anti-PD-1 immunotherapy expression profile dataset, derived from the GPL9052 platform, includes both pre-treatment and on-treatment samples and served as the training set (GSE91061_Pre) and validation set (GSE91061_On), respectively. The remaining three melanoma expression profile datasets, derived from the TCGA, GSE115821, and GSE22153 platforms (GPL11154, GPL18573, and GPL6102 platforms), were used as independent validation sets. Among them, GSE115821 also contains samples before and during treatment, while GSE22153 and TCGA-SKCM do not contain anti-PD-1 treatment information.

[0060] Before constructing the predicted gene pairs, the data needs to be preprocessed as follows:

[0061] 1) Download expression profile data from the GEO database and use the platform annotation file to map the probe ID to the gene ID;

[0062] 2) Remove SD samples from the training set;

[0063] 3) Remove low-expression genes, which are genes whose expression is absent or zero in more than 70% of the sample genes.

[0064] Step 104, performing pairwise combinations of the training set genes, comparing the relative expression rank relationships of the gene pairs in the PD samples and the PRCR samples, and determining candidate predictive gene pairs with potential for predicting anti-PD-1 immunotherapy for melanoma;

[0065] The genes in the training set samples are combined into pairs to form gene pairs. a and gene b Representative genes a and b The expression level of these two genes is a >gene b or gene a ≤gene b Screening for genes that are present in more than 80% of samples a >gene b The gene pairs were used as reference stable gene pairs for PD samples in anti-PD-1 immunotherapy for melanoma, and a total of 83,634,592 reference stable gene pairs were screened. The gene expression of each reference stable gene pair in PD samples and PRCR samples was calculated. a >gene b Number of samples m 1. n 1 and gene a <gene b Number of samples m 2. n 2. Fisher's exact test was then used to determine whether the REO distribution of the reference stable gene pairs was significantly different between PD and PRCR samples. Gene pairs with a false positive discovery rate (FDR) of less than 5% were defined as reverse gene pairs, and a total of 453,714 reverse gene pairs were screened. For each reverse gene pair, the formula △P=P PD (gene a >gene b )- P PRCR (gene a >geneb ) to calculate the reversal rate △P. The larger the △P value, the greater the difference in REO of the gene pair. △P=1 means that the REO of the gene pair in PD samples is gene a >gene b , in PRCR samples, both genes a ≤gene b Given a threshold △Pt such as 0.7, gene pairs that satisfy △P>Pt are taken as candidate prediction gene pairs, and 320 reversal gene pairs are obtained as candidate prediction gene pairs.

[0066] Step 106: Based on the candidate prediction gene pairs, a greedy algorithm is used to find the gene pair combination with the largest sample coverage, and the gene pair combination with the best prediction efficiency is selected as the prediction gene pair. a >gene b According to the voting rules, if gene a >gene b If the number of gene pairs is more than half, the sample to be predicted can be classified as a responsive sample (responder, R), otherwise it can be classified as a non-responder (non-responder, NR).

[0067] To construct predictive gene pairs for melanoma anti-PD-1 immunotherapy response, a greedy algorithm was used to identify the gene pair combination with the highest sample coverage based on candidate predictive gene pairs. The gene pair combination with the best predictive performance was selected as the predictive gene pair. The specific approach was as follows: Starting with each gene pair in the candidate predictive gene pair, this gene pair was used as a seed, and the remaining gene pairs were sequentially added to the current combination. After each addition, the sample coverage of the new combination was calculated. This process continued until no further candidate gene pairs could be added to increase sample coverage. The resulting combination was the signature gene pair combination with the highest sample coverage identified based on the candidate predictive gene pairs.

[0068] Calculate the frequency of each characteristic gene pair in all combinations. The frequency reflects the importance of the characteristic pair. Sort the frequencies from large to small and select the ones with the highest frequency first. k feature pairs. Here, k It is a predefined parameter with a value of 1~ n ( n is the total number of characteristic gene pairs). k The geometric mean of the negative and positive predictive values ​​of the feature pairs is calculated. The first 15 gene pairs that reach the maximum geometric mean and tend to be stable are selected as the final prediction gene pairs (15-pairs). For a sample to be predicted, the genea >gene b According to the voting rules, if gene a >gene b If the number of gene pairs in the prediction is more than half, the sample to be predicted can be classified as a responder (R), otherwise it can be classified as a non-responder (NR). The 15-pairs information is shown in Table 2.

[0069] Table 2.15-pairs

[0070]

[0071] In the above-mentioned method for establishing a gene pair for predicting the response of melanoma anti-PD-1 immunotherapy, the melanoma anti-PD-1 immunotherapy data set is used as the training set. Based on the difference in the relative expression rank relationship of gene pairs in PD samples and PRCR samples, candidate prediction gene pairs are screened. Based on the candidate prediction gene pairs, a greedy algorithm is used to find the gene pair combination with the maximum sample coverage, and the gene pair combination with the best prediction efficiency is selected as the prediction gene pair. For a sample to be predicted, the gene expression in the prediction gene pair combination is calculated. a >gene b According to the voting rules, if gene a >gene b If more than half of the gene pairs are positive, the sample is classified as a responder (R), whereas if not, it is classified as a non-responder (NR). This approach has been validated in validation sets from other platforms. This approach, based on the relative rank relationships between genes, can be robustly applied to independent clinical samples evaluated in different laboratories at an individualized level, accurately predicting the development and progression of cancer. Predictive gene pairs enable truly personalized prediction of patient drug response and have clinical application value in predicting response to anti-PD-1 immunotherapy in melanoma.

[0072] Example 2:

[0073] like Figure 2 As shown, a method for validating gene pairs for predicting response to anti-PD-1 immunotherapy for melanoma is provided, specifically comprising:

[0074] Step 302: obtaining expression profile data and clinical information of melanoma patients and melanoma anti-PD-1 immunotherapy patients from a comprehensive gene expression database and the Cancer Genome Atlas, preprocessing the expression profile data and clinical information, and dividing the melanoma expression profile datasets from different detection platforms into two data subsets, one melanoma anti-PD-1 immunotherapy data subset as a training set, and the other data subset as a validation set;

[0075] Step 304: Pairwise combinations of genes in the training set are performed to compare the relative expression rank relationships of gene pairs in PD samples and PRCR samples, screening gene pairs with significantly reversed relative expression rank relationships between genes, and identifying candidate predictive gene pairs with potential for predicting response to anti-PD-1 immunotherapy in melanoma.

[0076] Step 306: Based on the candidate prediction gene pairs, a greedy algorithm is used to find the gene pair combination with the largest sample coverage, and the gene pair combination with the best prediction efficiency is selected as the prediction gene pair. a >gene b According to the voting rules, if gene a >gene b If the number of gene pairs is more than half, the sample to be predicted can be classified as a responsive sample (responder, R), otherwise it can be classified as a non-responder (non-responder, NR).

[0077] Step 308 is to verify the melanoma anti-PD-1 immunotherapy response prediction gene pair in a validation set, specifically including: using the melanoma anti-PD-1 immunotherapy response prediction gene pair to calculate evaluation indicators (accuracy, positive predictive value, negative predictive value, specificity, sensitivity) for a data set containing anti-PD-1 immunotherapy drug responses; performing a survival analysis on the melanoma anti-PD-1 immunotherapy response prediction gene pair; performing an immune infiltration analysis on the melanoma anti-PD-1 immunotherapy response prediction gene pair; and performing an immunotherapy response analysis on the melanoma anti-PD-1 immunotherapy response prediction gene pair.

[0078] Evaluation metrics (accuracy, positive predictive value, negative predictive value, specificity, and sensitivity) were calculated for a dataset containing responses to anti-PD-1 immunotherapy drugs using the 15-pair gene pair. A 2x2 confusion matrix was used for GSE91061_Pre, GSE91061_On, GSE115821_Pre, and GSE115821_On. The validation results are shown in Table 3. The results demonstrate that the 15-pair gene pair has good predictive efficacy for anti-PD-1 therapy in melanoma.

[0079] Table 3. Verification results of training set and validation set

[0080]

[0081] Survival analysis and multivariate Cox regression analysis were performed on the validation set using the gene pairs predicting response to anti-PD-1 immunotherapy in melanoma. After removing samples with a survival time of less than 1 month, 451 samples remained in the TCGA-SKCM data set and 47 samples remained in the GSE22153 data set. In the TCGA-SKCM data set, 193 cases were predicted to be NR samples and 258 cases were predicted to be R samples. In the GSE22153 data set, 31 cases were predicted to be NR samples and 16 cases were predicted to be R samples. Figure 4 As shown in the KM survival curve, there are significant survival differences between the two types of samples in the TCGA-SKCM and GSE22153 datasets ( Figure 4 A, B), and non-responders had worse survival (p < 0.05, Log-rank test). This result indicates that the 15-pairs-predicted survival time between the two subgroups was significantly different. Multivariate Cox regression analysis was performed on GSE22153 and TCGA-SKCM, incorporating age, sex, and 15-pairs.

[0082] The results showed that 15-pairs was significantly associated with the overall survival of patients (HR=9.5, 95%CI: 3.44-26.5, P value <0.001; HR=51.1, 95%CI: 2.73-954.9, P value = 0.008, Wald test) ( Figure 4 C, D) The results showed that 15-pairs is an independent influencing factor for anti-PD-1 immunotherapy in melanoma. This analysis further proves the effectiveness of 15-pairs.

[0083] Based on the immune cell signature gene set containing the signature genes of immune cells, the ssGSEA (Single Sample Gene Set Enrichment Analysis) algorithm was used to perform immune cell infiltration analysis on the response group and non-response group of the validation set TCGA-SKCM and GSE22153. The relative abundance of specific immune cell types in each sample was evaluated by the ssGSEA algorithm to compare the differences in the degree of immune cell infiltration between the two subgroups divided by 15-pairs. The analysis results showed that there were significant differences in T cells, B cells, dendritic cells, macrophages, eosinophils, mast cells, neutrophils, natural killer cells, etc. between the two subgroups of TCGA-SKCM divided by 15-pairs (p < 0.05, Wilcoxon test), and the response group had a higher infiltration level ( Figure 5A). At the same time, there were significant differences in B cells, dendritic cells, macrophages, eosinophils, mast cells, monocytes, myeloid-derived suppressor cells, regulatory T cells, and natural killer cells between the two subgroups of GSE22153 divided by 15-pairs (p<0.05, Wilcoxon test), and the response group had a higher infiltration level ( Figure 5 B). These results indicate that the responder group based on 15-pairs division showed higher levels of immune cell infiltration, suggesting that immunotherapy drug response may be associated with the level of immune cell infiltration.

[0084] To evaluate whether there are differences in immunotherapy between the response group and the non-response group divided by 15-pairs: ten published immune checkpoint-related genes (CD274 (PD-1), CTLA-4, CXCR3, LAG3, PDCD1 (PD-L1), CD80, ICOS, IFNG, IL10, TNFRSF9) were collected and the expression levels of these genes in the two subgroups were evaluated. The results showed that the expression values ​​of CD274 (PD-L1), CTLA-4, CXCR3, LAG3, PDCD1 (PD-1), CD80, ICOS, IFNG, IL10, and TNFRSF9 in the response group of the TCGA-SKCM dataset were significantly higher than those in the non-response group (p < 0.05, Wilcoxon test, Figure 6 A). However, since the expression of CD274 and IFNG was not detected in the GSE22153 dataset, this study analyzed the expression of the remaining eight immune checkpoints. The results showed that the expression values ​​of CTLA-4, CXCR3, PDCD1, CD80, ICOS, IL10, and TNFRSF9 in the response group of GSE22153 were significantly higher than those in the non-response group (p<0.05, Wilcoxon test, Figure 6 B). Taking into account the prediction results of 15-pairs and the expression of immunotherapy-related genes, we can more comprehensively evaluate the differences in immunotherapy response among different subgroups, indicating that 15-pairs can indeed predict patients' response to anti-PD-1 immunotherapy.

[0085] It should be understood that although Figure 1 and Figure 3 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. In addition, Figure 1 and Figure 3At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.

[0086] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0087] The above embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.

Claims

1. A method for establishing gene pairs that predict response to anti-PD-1 immunotherapy for melanoma, characterized by: The following steps are involved: Step A: obtaining expression profile data and clinical information of melanoma patients and melanoma anti-PD-1 immunotherapy patients from the Gene Expression Comprehensive Database and The Cancer Genome Atlas, and preprocessing the expression profile data and clinical information; The melanoma expression profile datasets from different detection platforms were divided into two data subsets, one melanoma anti-PD-1 immunotherapy dataset was used as the training set, and the other data subset was used as the validation set; Acquiring expression profile data and clinical information of melanoma patients and melanoma anti-PD-1 immunotherapy patients from the Gene Expression Comprehensive Database and The Cancer Genome Atlas as described in step A, and preprocessing the expression profile data and clinical information, specifically comprises the following steps: Step A1: Download expression profile data from the GEO database and map probe IDs to gene IDs using the platform annotation file; Step A2: Studies have shown that during anti-PD-1 immunotherapy for melanoma, the subsequent efficacy of SD samples in a stable disease state is unstable and may develop into PD samples or PRCR samples during subsequent treatment. Therefore, SD samples are removed from the training set; Step A3: removing low-expression genes, wherein the low-expression genes are genes whose expression is absent or zero in more than 70% of the sample genes; Step B: Pairwise combinations of genes in the training set were performed to compare the relative expression rank relationships between gene pairs in progressive PD samples and those in partial or complete response PRCR samples. Gene pairs with significantly reversed relative expression rank relationships were screened to identify candidate predictive gene pairs with the potential to predict response to anti-PD-1 immunotherapy in melanoma. Step C: Use a greedy algorithm to find the gene pair combination with the maximum sample coverage; Starting from each gene pair in the candidate prediction gene pairs, use it as a seed and add the remaining gene pairs to the current combination in sequence; after each addition, calculate the sample coverage of the new combination; this process continues until the sample coverage can no longer be increased by adding candidate gene pairs; the final combination is the characteristic gene pair combination with the maximum sample coverage identified based on the candidate prediction gene pairs; Calculate the frequency of each characteristic gene pair in all combinations. The frequency reflects the importance of the characteristic gene pair. Select the best prediction gene pair as the final prediction gene pair. For a sample to be predicted, calculate the gene in the predicted gene pair combination a >gene b The number of gene pairs; According to the voting rules, if gene a >gene b If the number of gene pairs is more than half, the sample to be predicted is classified as a responsive sample R, otherwise it is a non-responsive sample NR; the prediction gene pair is then verified using the validation set, and the obtained prediction gene pair is used as a marker for predicting the response to anti-PD-1 immunotherapy for melanoma.

2. The method for establishing a gene pair predicting response to anti-PD-1 immunotherapy for melanoma according to claim 1, characterized in that: In step A1, the expression profile data is downloaded from the GEO database, and the probe ID is mapped to the gene ID using the platform annotation file. The melanoma anti-PD-1 immunotherapy expression profile dataset GSE91061 from the GPL9052 detection platform, including samples before and during treatment, is used as the training set GSE91061_Pre and the validation set GSE91061_On, respectively. The other three sets of melanoma expression profile datasets are used as independent validation sets; the three sets of validation sets are the melanoma expression profile data of TCGA, GSE115821, and GSE22153 of the detection platforms GPL11154, GPL18573, and GPL6102.

3. The method for establishing a gene pair predicting response to anti-PD-1 immunotherapy for melanoma according to claim 1, characterized in that: The steps described in step B are to combine the genes in the training set in pairs, compare the differences in the relative expression rank relationships of the gene pairs in the PD samples and the PRCR samples, screen for gene pairs with significantly reversed relative expression rank relationships between the genes, and identify candidate predictive gene pairs with potential for predicting the response to anti-PD-1 immunotherapy in melanoma, specifically comprising the following steps: Step B1: Based on the rank relationship of gene expression levels between gene pairs, the gene pairs with stable expression between genes in PD samples were screened as reference stable gene pairs. The genes in PD samples were combined in pairs to form n(n-1) / 2 gene pairs. According to the calculation formula (1), P(gene a >gene b ) ≥ 0.8 as the reference stable gene pairs; (1) in m represents the total number of PD samples, k Represents genes in PD samples a >gene b The number of samples, gene a >gene b Indicates gene a and genes b Rank relationship of expression levels; Step B2: Gene pairing based on the reference stable genes obtained in PD samples a , gene b Calculate the gene expression of each reference stable gene pair in PD samples and PRCR samples a >gene b Number of samples m 1. n 1 and gene a <gene b Number of samples m 2. n 2. Then, Fisher's exact test was used to determine whether there was a significant difference in the REO distribution of the reference stable gene pairs between PD and PRCR samples. Gene pairs with a false positive discovery rate (FDR) less than 5% were defined as reversal gene pairs. For each reversal gene pair, the formula △P=P PD (gene a >gene b )- P PRCR (gene a >gene b ) to calculate the reversal rate △P; the larger the △P value, the greater the difference in REO of the gene pair; △P=1 means that the REO of the gene pair in PD samples is the same as that of the gene pair. a >gene b , in PRCR samples, both genes a ≤gene b ; Given a threshold △Pt of 0.7, gene pairs that satisfy △P>△Pt are taken as candidate prediction gene pairs.

4. The method for establishing a gene pair predicting response to anti-PD-1 immunotherapy for melanoma according to claim 1, characterized in that: The step C specifically comprises the following steps: A greedy algorithm is used to find the gene pair combination with the maximum sample coverage. Starting from each gene pair in the candidate prediction gene pair, it is used as a seed and the remaining gene pairs are added to the current combination in sequence. After each addition, the sample coverage of the new combination is calculated. This process continues until the sample coverage can no longer be increased by adding candidate gene pairs. The final combination is the characteristic gene pair combination with the maximum sample coverage identified based on the candidate prediction gene pairs. The frequency of each characteristic gene pair in all combinations is calculated. The frequency reflects the importance of the characteristic gene pair. The optimal prediction gene pair is selected as the final prediction gene pair. For a sample to be predicted, the gene pair in the prediction gene pair combination is calculated. a >gene b The number of gene pairs; According to the voting rules, if gene a >gene b If the number of gene pairs is more than half, the sample to be predicted is classified as a responsive sample R, otherwise it is classified as a non-responsive sample NR; the prediction gene pair is then verified using the validation set, and the obtained prediction gene pair is used as a marker for predicting the response to anti-PD-1 immunotherapy for melanoma; including: The REO reference pattern of a gene pair is defined as the gene expressed in PD samples. a >gene b Or in PRCR samples, it appears as a gene a ≤gene b When at least one gene pair in a sample shows the gene pair REO reference pattern, the combination is considered to cover the sample, and the percentage of covered samples to the total number of samples is calculated; To screen a set of characteristic gene pairs that can predict therapeutic efficacy, a greedy algorithm is used to find the gene pair combination with the maximum sample coverage. Starting from each gene pair in the candidate prediction gene pair, it is used as a seed and the remaining gene pairs are added to the current combination in sequence. After each addition, the sample coverage of the new combination is calculated. This process continues until the sample coverage can no longer be increased by adding candidate gene pairs. The final combination is the characteristic gene pair combination with the maximum sample coverage identified based on the candidate prediction gene pairs. Calculate the frequency of each characteristic gene pair in all combinations. The frequency reflects the importance of the characteristic gene pair. Sort the frequencies from large to small and select the ones with the highest frequency first. k characteristic gene pairs; k It is a predefined parameter with a value of 1~ n , n is the total number of characteristic gene pairs; for the selected k The geometric mean of the negative predictive value and the positive predictive value of the characteristic gene pairs were calculated; the top pair that first reached the maximum geometric mean and tended to be stable was selected. k gene pairs as the final prediction gene pair combination; for a sample to be predicted, calculate the gene a >gene b The number of gene pairs; According to the voting rules, if gene a >gene b If the number of gene pairs is more than half, the sample to be predicted is classified as a responsive sample R, otherwise it is a non-responsive sample NR.

5. A method for validating gene pairs for predicting response to anti-PD-1 immunotherapy in melanoma, characterized by: The following steps are involved: Step A: obtaining expression profile data and clinical information of melanoma patients and melanoma anti-PD-1 immunotherapy patients from the Gene Expression Comprehensive Database and The Cancer Genome Atlas, and preprocessing the expression profile data and clinical information; The melanoma expression profile datasets from different detection platforms were divided into two data subsets, one melanoma anti-PD-1 immunotherapy dataset was used as the training set, and the other data subset was used as the validation set; Acquiring expression profile data and clinical information of melanoma patients and melanoma anti-PD-1 immunotherapy patients from the Gene Expression Comprehensive Database and The Cancer Genome Atlas as described in step A, and preprocessing the expression profile data and clinical information, specifically comprises the following steps: Step A1: Download expression profile data from the GEO database and map probe IDs to gene IDs using the platform annotation file; Step A2: Studies have shown that during anti-PD-1 immunotherapy for melanoma, the subsequent efficacy of SD samples in a stable disease state is unstable and may develop into PD samples or PRCR samples during subsequent treatment. Therefore, SD samples are removed from the training set; Step A3: removing low-expression genes, wherein the low-expression genes are genes whose expression is absent or zero in more than 70% of the sample genes; Step B: Pairwise combinations of genes in the training set were performed to compare the relative expression rank relationships between gene pairs in progressive PD samples and those in partial or complete response PRCR samples. Gene pairs with significantly reversed relative expression rank relationships were screened to identify candidate predictive gene pairs with the potential to predict response to anti-PD-1 immunotherapy in melanoma. Step C: Use a greedy algorithm to find the gene pair combination with the maximum sample coverage; Starting from each gene pair in the candidate prediction gene pairs, use it as a seed and add the remaining gene pairs to the current combination in sequence; after each addition, calculate the sample coverage of the new combination; this process continues until the sample coverage can no longer be increased by adding candidate gene pairs; the final combination is the characteristic gene pair combination with the maximum sample coverage identified based on the candidate prediction gene pairs; Calculate the frequency of each characteristic gene pair in all combinations. The frequency reflects the importance of the characteristic gene pair. Select the best prediction gene pair as the final prediction gene pair. For a sample to be predicted, calculate the gene in the predicted gene pair combination a >gene b The number of gene pairs; According to the voting rules, if gene a >gene b If the number of gene pairs is more than half, the sample to be predicted is classified as a responsive sample R, otherwise it is classified as a non-responsive sample NR; the prediction gene pair is then verified using the validation set, and the obtained prediction gene pair is used as a marker for predicting the response to anti-PD-1 immunotherapy for melanoma; Step D: Validating the melanoma anti-PD-1 immunotherapy response prediction gene pair in a validation set, specifically comprising: using the melanoma anti-PD-1 immunotherapy response prediction gene pair to evaluate the data set containing anti-PD-1 immunotherapy drug responses: calculation of accuracy, positive predictive value, negative predictive value, specificity, and sensitivity; performing survival analysis and multivariate Cox regression analysis on the melanoma anti-PD-1 immunotherapy response prediction gene pair; performing immune infiltration analysis on the melanoma anti-PD-1 immunotherapy response prediction gene pair; and performing immunotherapy response analysis on the melanoma anti-PD-1 immunotherapy response prediction gene pair.

6. The method for validating a gene pair for predicting response to anti-PD-1 immunotherapy for melanoma according to claim 5, characterized in that: Step D, wherein the evaluation indicators of the dataset containing the response to anti-PD-1 immunotherapy drugs are calculated using the melanoma anti-PD-1 immunotherapy response prediction gene pairs: accuracy, positive predictive value, negative predictive value, specificity, and sensitivity, specifically comprises the following steps: Use a 2x2 confusion matrix to represent the prediction results of gene pairs, where the rows represent the actual labels and the columns represent the predicted labels; The confusion matrix is ​​established based on the predicted genes and the prediction index calculation formulas are: 、 、 、 、 ; Confusion matrix of predicted genes and predicted results: The actual PRCR, predicted PRCR, is TP; the actual PRCR, predicted PD, is FN; the actual PRCR, predicted total, is P; The actual patient is PD, and the predicted patient is PRCR, which is FP; the actual patient is PD, and the predicted patient is PD, which is TN; the actual patient is PD, and the predicted total patient is N; The actual total, predicted as PRCR, is TP+FN; the actual total, predicted as PD, is FP+TN; the actual total, predicted as P+N; Note: TP, True positive; TN, True negative; FN, False negative; FP, False positive; P, positive; N, negative.

7. The method for validating gene pairs for predicting response to anti-PD-1 immunotherapy for melanoma according to claim 5, characterized in that: The step D of performing survival analysis and multivariate Cox regression analysis on the melanoma anti-PD-1 immunotherapy response prediction gene pair specifically comprises the following steps: Kaplan-Meier KM survival analysis was performed using the "survival" and "survminer" packages in the R package to verify whether there was a survival difference between the two subgroups predicted by the predicted gene pairs in the validation set; the expression rank relationship of the gene pairs in the predicted gene pair combination in the sample was gene a Greater than gene b A multivariate Cox regression analysis was performed on the validation set, incorporating factors such as age, gender, and predictive gene pair scores.

8. The method for validating gene pairs for predicting response to anti-PD-1 immunotherapy for melanoma according to claim 5, characterized in that: The immune infiltration analysis of the melanoma anti-PD-1 immunotherapy response prediction gene pair described in step D specifically includes the following steps: The ssGSEA algorithm was used to evaluate the immune cell content between the response group and the non-response group predicted by the predictive gene pairs in the validation set. The immune cell content between the two subgroups predicted by the predictive gene pairs in the validation set included 28 types: activated CD4+ / CD8+ T cells, central memory CD4+ / CD8+ T cells, effector memory CD4+ / CD8+ T cells, γδT cells, helper T cells Th1, Th2 cells, Th17 cells, regulatory T cells Treg, follicular helper T cells Tfh, activated B cells, naive B cells, memory B cells, activated dendritic cells DC, naive DC, plasmacytoid DC, natural killer cells NK, natural killer T cells NKT, myeloid-derived suppressor cells MDSC, macrophages, monocytes, mast cells, eosinophils, neutrophils, CD56bright NK cells, and CD56dim NK cells. The Wilcoxon test was used for comparison.

9. The method for validating gene pairs for predicting response to anti-PD-1 immunotherapy for melanoma according to claim 5, characterized in that: The immunotherapy response analysis of the melanoma anti-PD-1 immunotherapy response prediction gene pair described in step D specifically includes the following steps: Methods: Ten published immune checkpoint-related genes (CD274 PD-1, CTLA-4, CXCR3, LAG3, PDCD1PD-L1, CD80, ICOS, IFNG, IL10, and TNFRSF9) were collected, and their expression levels in the response and non-response groups predicted by the validation set were evaluated. The differences in the levels of immunotherapy response-related genes between the two subgroups predicted by the prediction gene pairs in the validation set were investigated by comparing their expression levels. The Wilcoxon test was used for comparison.

Citation Information

Patent Citations

  • Mononuclear cells expressing p21 for cancer cell therapy

    CN114450012A

  • Prognostic and treatment response predictive method

    US20210139999A1