A method for establishing and validating a prognostic biomarker model for serous ovarian cancer.

By constructing a prognostic biomarker model for serous ovarian cancer based on the relative expression order among genes, the instability problem of existing models was solved, and robust prediction and personalized prognostic assessment were achieved across different laboratories.

CN115331812BActive Publication Date: 2025-10-31GANNAN MEDICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211153210.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-21
Publication Date
2025-10-31
Estimated Expiration
2042-09-21

AI Technical Summary

Technical Problem

Existing prognostic prediction models for serous ovarian cancer cannot accurately predict patient outcomes and are affected by biases in gene expression level measurement systems and batch effects in laboratories, resulting in instability when applied across different laboratories.

Method used

By acquiring data from a comprehensive gene expression database, Cox regression model was used to screen out significantly related gene pairs. Combined with a greedy algorithm and the consistency index C-index, a prognostic biomarker model for serous ovarian cancer was constructed. Based on the relative expression order relationship between genes, a gene pair matrix was formed to screen out stable prognostic biomarkers.

Benefits of technology

It enables accurate prediction of prognosis for patients with serous ovarian cancer at the individualized level, avoiding systematic bias and laboratory batch effects during risk stratification, and has clinical application value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115331812B_ABST
    Figure CN115331812B_ABST
Patent Text Reader

Abstract

This application relates to a method for establishing and validating a prognostic biomarker model for serous ovarian cancer. Samples from different platforms are merged into a training set. A gene pair matrix is ​​formed based on the relative expression order of genes within the samples. A Cox regression model is used to obtain gene pairs with significant prognostic relevance. Then, through positive selection, a greedy algorithm, and testing, a set of gene pairs is selected as the prognostic biomarker model for serous ovarian cancer. This model is validated on test sets and validation sets from other platforms. This method, based on the relative order of genes, can be applied at the individualized level to independent clinical samples from different laboratories, accurately predicting the occurrence and development of cancer. It avoids the systematic bias and batch effect of laboratory factors in threshold selection during risk stratification of prognostic biomarkers. It comprehensively considers the impact of various genes on prognosis during the development of serous ovarian cancer, and has clinical application value for the prognosis of patients with serous ovarian cancer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of ovarian cancer prognostic technology, and in particular to a method for establishing and validating a prognostic biomarker model for serous ovarian cancer. Background Technology

[0002] Serous ovarian cancer is often discovered at an advanced stage and has a poor prognosis, making it the second leading cause of gynecological cancer death. The variable biology and complex molecular characteristics of serous ovarian cancer present the greatest challenge to its prognosis by enabling personalized precision medicine.

[0003] With the development of big data and gene technology, techniques have emerged to identify prognostic biomarkers for ovarian cancer based on genome-wide expression changes. Currently, these techniques are mainly divided into two categories: the first assesses the impact of a single gene marker on ovarian cancer prognosis. This type of study, based on a single gene, does not consider the influence of gene-gene interactions and cannot accurately predict the prognosis of patients with serous ovarian cancer. The second type evaluates the prognostic value of a group of genes based on a specific characteristic level. This method ignores the heterogeneity of individual patients and the complexity of influencing factors, easily leading to overfitting of prognostic markers and making them unsuitable for clinical application.

[0004] Currently, most ovarian cancer prognostic prediction models, whether based on a single gene or a group of functionally related genes, rely on risk scores and preset risk thresholds to determine patient risk. However, due to batch effects and platform differences, gene expression levels are highly sensitive to systematic biases in microarray measurements, and risk thresholds generated from training datasets cannot be directly applied to independent datasets. Furthermore, risk scoring methods based on gene expression levels are uncertain in classifying samples as high or low risk, making them unsuitable for clinical application. Finally, most current studies share a common problem: the sample size is too small, resulting in very poor robustness of the identified gene markers. Summary of the Invention

[0005] Therefore, it is necessary to provide a method for identifying stable prognostic risk markers with clinical translational value in serous ovarian cancer samples to address the aforementioned technical problems.

[0006] A method for establishing a prognostic biomarker model for serous ovarian cancer, the method comprising:

[0007] Step A: Obtain expression profile data and clinical information of patients with serous ovarian cancer from the gene expression comprehensive database, preprocess the expression profile data and clinical information, and divide the serous ovarian cancer expression profile datasets from different detection platforms into two data subsets, one data subset as the training set and the other data subset as the test set;

[0008] Step B: Genes significantly associated with overall survival of patients with serous ovarian cancer are screened using a Cox regression model. The significantly associated genes are paired to obtain candidate prognostic gene pairs. A gene pair matrix is ​​obtained based on the relative expression order of the candidate prognostic gene pairs. Gene pairs significantly associated with the prognosis of serous ovarian cancer are screened using a Cox regression model based on the gene pair matrix.

[0009] Step C: Select the top N gene pairs with the highest consistency index C-index value from the gene pairs that are significantly associated with the prognosis of serous ovarian cancer in a forward selection order. The consistency index C-index is an index that describes the predictive ability of the prognostic model. Use a greedy algorithm to screen each gene pair in the top N gene pairs, and then use the test set to test the gene pairs. The resulting set of gene pairs is used as the prognostic biomarker model for serous ovarian cancer.

[0010] In one embodiment, the preprocessing of the expression profile data and clinical information includes:

[0011] Step A1: Remove tumor samples with no clinical information and an overall survival of 0 days;

[0012] Step A2: Remove normal samples;

[0013] Step A3: Remove low-expression genes, which are genes whose expression is absent or zero in more than half of the samples.

[0014] In one embodiment, the step of using serous ovarian cancer expression profile datasets from different detection platforms as training sets and using serous ovarian cancer expression profile datasets from another dataset as test sets includes: the training set being a collection of serous ovarian cancer expression profile datasets from detection platforms GPL570, GPL8300, GPL96 (GSE18520, GSE19829, and TCGA), and the test set being a collection of serous ovarian cancer expression profile datasets from GPL7759 platform (GSE13876).

[0015] In one embodiment, step B involves: screening for genes significantly associated with overall survival in patients with serous ovarian cancer using a Cox regression model; pairing these significantly associated genes to obtain candidate prognostic gene pairs; obtaining a gene pair matrix based on the relative expression order of the candidate prognostic gene pairs; and screening for gene pairs significantly associated with the prognosis of serous ovarian cancer using a Cox regression model based on the gene pair matrix, including:

[0016] The Cox regression model was used to identify individual genes in the training set. When the p-value was less than 0.05, the gene was identified as a gene that was significantly associated with the overall survival of ovarian cancer patients.

[0017] The significantly related genes are paired to obtain candidate prognostic gene pairs, and a gene pair matrix is ​​obtained based on the relative expression order of the candidate prognostic gene pairs.

[0018] The candidate prognostic-related gene pairs were identified using a Cox regression model. When the p-value was less than 0.05, the gene pair was considered to be significantly associated with the prognosis of ovarian cancer.

[0019] In one embodiment, step C involves selecting the top N gene pairs with the highest concordance index (C-index) values ​​from the gene pairs significantly associated with the prognosis of serous ovarian cancer, following a forward selection order. The C-index is an index describing the predictive ability of the prognostic model. A greedy algorithm is used to screen each gene pair in the top N groups, and the gene pairs are then tested using the test set. The resulting set of gene pairs serves as a prognostic biomarker model for serous ovarian cancer, comprising:

[0020] Using each gene pair significantly associated with ovarian cancer prognosis as a seed, the remaining gene pairs are added to the combination one by one. If the consistency index C-index value increases after adding a gene pair, gene pairs are added to the combination. If the consistency index C-index value decreases after adding a gene pair, no gene pairs are added. This process continues until the consistency index C-index value no longer increases, resulting in N sets of gene pairs. The consistency index C-index is an index that describes the predictive ability of the prognostic model.

[0021] Based on the N gene pairs, they are sorted from largest to smallest according to their consistency index (C-index). The N gene pairs are then tested using the test set to screen out the gene pair with the highest consistency index (C-index) that is associated with the prognosis of serous ovarian cancer. This gene pair is then used as the final prognostic biomarker model for serous ovarian cancer.

[0022] Validation method of a prognostic biomarker model for serous ovarian cancer

[0023] Step A: Obtain expression profile data and clinical information of patients with serous ovarian cancer from the gene expression comprehensive database, preprocess the expression profile data and clinical information, and divide the serous ovarian cancer expression profile datasets from different detection platforms into three data subsets: one data subset as the training set, one data subset as the test set, and another data subset as the validation set.

[0024] Step B: Genes significantly associated with overall survival of patients with serous ovarian cancer are screened using a Cox regression model. The significantly associated genes are paired to obtain candidate prognostic gene pairs. A gene pair matrix is ​​obtained based on the relative expression order of the candidate prognostic gene pairs. Gene pairs significantly associated with the prognosis of serous ovarian cancer are screened using a Cox regression model based on the gene pair matrix.

[0025] Step C: Select the top N gene pairs with the highest consistency index C-index value from the gene pairs that are significantly associated with the prognosis of serous ovarian cancer in a forward selection order. Use a greedy algorithm to screen each gene pair in the top N gene pairs. Then use the test set to test the combination of gene pairs. Use the resulting gene pair as a prognostic biomarker model for serous ovarian cancer.

[0026] Step D: Validate the prognostic biomarker model for serous ovarian cancer on the validation set, specifically including: performing survival analysis, functional enrichment analysis, and immune infiltration analysis on the prognostic biomarker model for serous ovarian cancer.

[0027] In one embodiment, the survival analysis of the prognostic biomarker model for serous ovarian cancer includes:

[0028] When the gene affects G a G b The order of expression in the sample is G. a Greater than G b If the gene affects G, the sample is considered high-risk. a G b The order of expression in the sample is G. a Less than or equal to G b If the sample is deemed low-risk, then a multivariate Cox regression analysis is performed, taking into account age and stage information.

[0029] In one embodiment, the functional enrichment analysis of the prognostic biomarker model for serous ovarian cancer includes:

[0030] Functional enrichment analysis was performed on the prognostic biomarker model of serous ovarian cancer based on the Kyoto Gene, Genome Encyclopedia, and Gene Ontology databases in Metascape. Gene annotation and pathway analysis were performed on the prognostic gene markers in the Human Disease-Related Genes and Mutation Sites Database to search for pathways related to cancer occurrence and development in gene marker pairs.

[0031] In one embodiment, the immune infiltration analysis of the serous ovarian cancer prognostic marker model includes:

[0032] The content of immune cells between high- and low-risk groups in the training set was estimated using TIMER2.0, and the Wilcoxon test was used for comparison.

[0033] In one embodiment, the immune cells between the high- and low-risk groups in the training set include six types: B cells, CD4+ T cells, CD8+ T cells, neutrophils, macrophages, and dendritic cells.

[0034] The aforementioned method for establishing and validating a prognostic biomarker model for serous ovarian cancer involves merging samples from different platforms into a training set. A gene pair matrix is ​​formed based on the relative gene expression order within the samples. A Cox regression model is used to obtain gene pairs with significant prognostic relevance. Then, through positive selection, a greedy algorithm, and testing, a set of gene pairs is selected as the prognostic biomarker model for serous ovarian cancer. This model is validated on test and validation sets from other platforms. This method, based on the relative order between genes, can be robustly applied at the individual level to independent clinical samples evaluated in different laboratories, accurately predicting the occurrence and development of cancer. Furthermore, this method selects prognostic gene markers based on the gene expression order within the whole-genome expression profiles of samples from different platforms, avoiding the systematic bias and batch effect of laboratory factors in threshold selection during risk stratification. It comprehensively considers the impact of various genes on prognosis during the development of serous ovarian cancer, enabling truly personalized prediction of patient prognosis and demonstrating clinical application value for the prognosis of serous ovarian cancer patients. Attached Figure Description

[0035] Figure 1 This is a flowchart illustrating a method for establishing a prognostic biomarker model for serous ovarian cancer in one embodiment.

[0036] Figure 2 This is a schematic diagram illustrating the process of establishing the training set, test set, and validation set for serous ovarian cancer in one embodiment.

[0037] Figure 3 This is a schematic diagram illustrating the establishment of a gene pair matrix based on the relative order of genes in one embodiment;

[0038] Figure 4 This is a flowchart illustrating a method for validating a prognostic biomarker model for serous ovarian cancer in one embodiment. Detailed Implementation

[0039] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0040] This application provides a method for establishing and validating a prognostic biomarker model for serous ovarian cancer. Based on the relative expression order relationship between genes within a sample, the method identifies disease molecular markers and obtains a set of gene pairs as a prognostic biomarker model for serous ovarian cancer.

[0041] In one embodiment, such as Figure 1 As shown, a method for establishing a prognostic biomarker model for serous ovarian cancer is provided, specifically including:

[0042] Step 102: Obtain expression profile data and clinical information of patients with serous ovarian cancer from the gene expression comprehensive database, preprocess the expression profile data and clinical information, and divide the serous ovarian cancer expression profile datasets from different detection platforms into two data subsets, one data subset as the training set and the other data subset as the test set.

[0043] Gene expression profiles from eight independent datasets, comprising 1493 samples, were retrospectively collected. The analyzed data were from the Gene Expression Integrated Database (GEO, http: / / www.ncbi.nlm.nih.gov / geo / ) and the University of California, Santa Cruz Xena database (UCSC Xena, https: / / xenabrowser.net / datapages / ), and were all microarray data.

[0044] Raw data downloaded from GEO was preprocessed using the RMA algorithm. Probe IDs were mapped to gene IDs using annotation files from various platforms. Probes not mapped to genes were removed. For different probes mapped to the same gene, the average value of the different probes was used as the final expression value of that gene. Microarray expression profiles and clinical data from 630 ovarian cancer patients from the Cancer Genome Atlas (TCGA) project were downloaded from the UCSC-Xena database. Before constructing prognostic gene marker pairs, the following preprocessing was performed: 1) tumor samples with no clinical information and overall survival of 0 days were removed; 2) normal samples were removed; 3) low-expression genes were removed (more than half of the samples showed missing or zero gene expression).

[0045] The serous ovarian cancer expression profile datasets from different detection platforms GPL570, GPL8300, GPL96 (GSE18520, GSE19829, and TCGA) were merged as the training set; the GSE13876 dataset from the GPL7759 platform was used as the test set; the GSE14764 and GSE26712 datasets from the GPL96 platform were merged as validation set 1; and the GSE26193 and GSE53963 datasets from the GPL570 and GPL6480 platforms were used as validation sets 2 and 3, respectively. For detailed research procedures, please refer to [link to research documentation]. Figure 2 .

[0046] In this embodiment, the three training datasets jointly detected 6934 genes. For the commonly detected genes, candidate genes significantly associated with the prognosis of ovarian cancer were identified in each of the three training datasets.

[0047] Step 104: Use the Cox regression model to screen for genes that are significantly associated with the overall survival of patients with serous ovarian cancer. Combine the significantly associated genes in pairs to obtain candidate prognostic gene pairs. Obtain a gene pair matrix based on the relative expression order of the candidate prognostic gene pairs. Use the Cox regression model to screen for gene pairs that are significantly associated with the prognosis of serous ovarian cancer based on the gene pair matrix.

[0048] Univariate Cox regression analysis was used to screen for prognostic-related genes and gene pairs. When Cox regression analysis was applied to analyze single genes, a p-value less than 0.05 was considered significantly associated with overall survival in ovarian cancer patients. Pairwise pairing of significantly prognostic-related genes yielded candidate prognostic-related gene pairs.

[0049] Based on the relative expression order of each gene pair in each sample, the actual expression level in each sample is replaced with the relative size ranking of genes, thus obtaining the relative expression order matrix X of candidate prognostic related gene pairs, such as... Figure 3 As shown. This matrix is ​​a 0-1 matrix, where x i,ab =1 indicates that the expression order of gene pair (Ga, Gb) in sample i is that gene Ga is greater than gene Gb, x i,ab =0 indicates that the expression order of gene pair (Ga, Gb) in sample i is that gene Ga is less than or equal to gene Gb. When Cox regression analysis is applied to analyze gene pairs, gene pairs with a p-value less than 0.05 after Benjamini-Hochberg correction are identified as gene pairs significantly associated with the prognosis of ovarian cancer.

[0050] Step 106: Select the top N gene pairs with the highest consistency index (C-index) values ​​from the gene pairs that are significantly associated with the prognosis of serous ovarian cancer in a forward selection order. The consistency index (C-index) is an index that describes the predictive ability of the prognostic model. Use a greedy algorithm to screen each gene pair in the top N gene pairs, and then use a test set to test the gene pairs. The resulting set of gene pairs is used as the prognostic biomarker model for serous ovarian cancer.

[0051] To construct a prognostic model for gene pair markers, all significantly prognostically related gene pairs were selected, and the top N pairs with the highest consistency index (C-index) were chosen in a forward selection order. A greedy algorithm was used to select an optimal prognostic gene pair combination for each gene pair. Specifically, each gene pair was used as a seed, and all remaining gene pairs were added to the combination one by one. If the C-index value increased after adding a gene pair, more gene pairs were added; otherwise, no more were added, until the C-index value no longer increased. This yielded N locally optimal combinations of prognostic biomarkers. Then, for each combination, it was sorted by C-index value from largest to smallest, and the test set GSE13876 was used to select the combination with the highest C-index value and the highest prognostic correlation (p < 0.05) as the final prognostic-related biomarker gene pair.

[0052] In the aforementioned method for establishing a prognostic biomarker model for serous ovarian cancer, samples from different platforms are merged into a training set. A gene pair matrix is ​​formed based on the relative gene expression order within the samples. A Cox regression model is used to obtain gene pairs with significant prognostic relevance. Then, through positive selection, a greedy algorithm, and testing, a set of gene pairs is selected as the prognostic biomarker model for serous ovarian cancer. This method, based on the relative order relationship between genes, can be robustly applied at the individual level to independent clinical samples evaluated in different laboratories, accurately predicting the occurrence and development of cancer. Furthermore, this method selects prognostic gene markers based on the gene expression order relationship within the whole-genome expression profile samples from different platforms, avoiding the systematic bias and batch effect of laboratory factors in threshold selection during risk stratification of prognostic markers. It comprehensively considers the impact of various genes on prognosis during the occurrence and development of serous ovarian cancer, enabling truly personalized prediction of patient prognosis and possessing clinical application value for the prognosis of serous ovarian cancer patients.

[0053] In one embodiment, such as Figure 4 As shown, a method for validating a prognostic biomarker model for serous ovarian cancer is provided, specifically including:

[0054] Step 402: Obtain expression profile data and clinical information of patients with serous ovarian cancer from the gene expression comprehensive database, preprocess the expression profile data and clinical information, and divide the serous ovarian cancer expression profile datasets from different detection platforms into three data subsets: one data subset as the training set, one data subset as the test set, and another data subset as the validation set.

[0055] Step 404: Genes significantly associated with overall survival of patients with serous ovarian cancer are screened using a Cox regression model. Significantly associated genes are paired to obtain candidate prognostic gene pairs. Gene pair matrix is ​​obtained based on the relative expression order of the candidate prognostic gene pairs. Gene pairs significantly associated with prognosis of serous ovarian cancer are screened based on the gene pair matrix using a Cox regression model.

[0056] Step 406: Select the top N gene pairs with the highest consistency index (C-index) values ​​that are significantly associated with the prognosis of serous ovarian cancer in a forward selection order. Use a greedy algorithm to screen each gene pair in the top N gene pairs. Then, use a test set to test the gene pair combinations. The resulting gene pair is used as a prognostic biomarker model for serous ovarian cancer.

[0057] Step 408: Validate the prognostic biomarker model for serous ovarian cancer on the validation set, specifically including: survival analysis, functional enrichment analysis, and immune infiltration analysis of the prognostic biomarker model for serous ovarian cancer.

[0058] Survival analysis of a prognostic biomarker model for serous ovarian cancer: when the gene affects G... a G b The order of expression in the sample is G. a Greater than G b If the gene affects G, the sample is considered high-risk. a G b The order of expression in the sample is G. a Less than or equal to G bIf the risk level is low, the sample is then classified as low-risk. Multivariate Cox regression analysis is then performed, incorporating age and stage information. A prognostic biomarker model for serous ovarian cancer is used to score the risk of samples in each dataset, determining high or low risk based on the voting principle of the half-gene pair. In the training set, based on the voting order of the prognostic biomarker model, 298 and 452 samples were assigned to high and low-risk groups, respectively. Survival analysis between the two groups showed a significant difference (p<0.0001, HR=0.23, 95%CI: 0.18–0.29). In the three validation sets, 292 and 136, 44 and 145, and 129 and 150 samples were assigned to high-risk and low-risk groups, respectively, and survival analyses showed significant differences (p = 0.0077, HR = 0.60, 95% CI: 0.41–0.88), (p = 0.028, HR = 0.54, 95% CI: 0.31–0.94), and (p = 0.0062, HR = 0.59, 95% CI: 0.41–0.87). Kaplan-Meier survival curves showed that in the training set, high-risk ovarian cancer patients had lower overall survival rates, while low-risk patients generally exhibited longer survival times, with significant differences between the two groups. Since the overall survival time distribution of patients exceeded 5 years, the AUC of the model was evaluated at 3, 5, and 7 years. The average AUC of the training set was 0.756; the average AUC of validation set 1 was 0.59; the average AUC of validation set 2 was 0.630; and the average AUC of validation set 3 was 0.680. This indicates that the gene markers have significant prognostic value.

[0059] Functional enrichment analysis was performed on a prognostic biomarker model for serous ovarian cancer. Based on the online databases Metascape (Kyoto Gene, Genome Encyclopedia, and Gene Ontology), functional enrichment analysis was conducted on gene pairs in the serous ovarian cancer prognostic biomarker model. Gene annotation and pathway analysis were performed on prognostic gene markers in the Human Disease-Related Genes and Mutation Sites Database to search for pathways related to cancer development and progression among gene marker pairs. Functional analysis was performed on genes in the serous ovarian cancer prognostic biomarker model using Metascape. Enrichment results showed that eight genes were significantly enriched in biological pathways regulating cytokine production. Cytokines are small polypeptide molecules mainly secreted by immune cells that regulate cellular function. Cytokine and cytokine receptor processes are mainly involved in regulating the body's immune response, hematopoietic function, and inflammatory response during the immune response. They can inhibit tumor occurrence and progression and have also been shown to be effective in cancer treatment. Three genes are involved in the biological pathway of viral entry into host cells. Pathway analysis in the DisGeNet database revealed five genes associated with tumor recurrence, confirming the role of the prognostic biomarker model for serous ovarian cancer in cancer development and progression. Further enrichment analysis was then performed on differentially expressed genes in high-risk and low-risk groups. Results showed significant enrichment in the PI3K-Akt signaling pathway, the proteoglycan pathway in cancer, and the AGE-RAGE signaling pathway.

[0060] Immune infiltration analysis was performed on a prognostic biomarker model for serous ovarian cancer. The content of immune cells between high- and low-risk groups in the training set was estimated using TIMER 2.0, and the Wilcoxon test was used for comparison. The high- and low-risk samples in the training set were analyzed using the Estimation function of the Timer database. The immune cells in the high- and low-risk groups in the training set included six types: B cells, CD4+ T cells, CD8+ T cells, neutrophils, macrophages, and dendritic cells. The results showed significant new differences in the infiltration levels of CD8+ T cells, neutrophils, macrophages, and dendritic cells in the prognostic gene marker risk groups (P < 0.001, Wilcoxon test), with the low-risk group exhibiting a higher infiltration level.

[0061] It should be understood that, although Figure 1 and Figure 4 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figure 1 and Figure 4At least some of the steps in the process may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.

[0062] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0063] The above embodiments merely illustrate several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A method for establishing a prognostic biomarker model for serous ovarian cancer, characterized in that, The method includes: Step A: Obtain expression profile data and clinical information of patients with serous ovarian cancer from the gene expression comprehensive database, preprocess the expression profile data and clinical information, and divide the serous ovarian cancer expression profile datasets from different detection platforms into two data subsets, one data subset as the training set and the other data subset as the test set; Step B: Genes significantly associated with overall survival of patients with serous ovarian cancer are screened using a Cox regression model. The significantly associated genes are paired to obtain candidate prognostic gene pairs. A gene pair matrix is ​​obtained based on the relative expression order of the candidate prognostic gene pairs. Gene pairs significantly associated with the prognosis of serous ovarian cancer are screened using a Cox regression model based on the gene pair matrix. Step C: Select the top N gene pairs with the highest consistency index C-index value from the gene pairs that are significantly associated with the prognosis of serous ovarian cancer in a forward selection order. The consistency index C-index is an index that describes the predictive ability of the prognostic model. Use a greedy algorithm to screen each gene pair in the top N gene pairs, and then use the test set to test the gene pairs. The resulting set of gene pairs is used as the prognostic biomarker model for serous ovarian cancer.

2. The method according to claim 1, characterized in that, The preprocessing of expression profile data and clinical information includes: Step A1: Remove tumor samples with no clinical information and an overall survival of 0 days; Step A2: Remove normal samples; Step A3: Remove low-expression genes, which are genes whose expression is absent or zero in more than half of the samples.

3. The method according to claim 1, characterized in that, The process of dividing the serous ovarian cancer expression profile dataset from different detection platforms into two subsets, one subset serving as the training set and the other subset as the test set, includes: The training set is a collection of serous ovarian cancer expression profiles from GSE18520, GSE19829, and TCGA on the GPL570, GPL8300, and GPL96 platforms, while the test set is a collection of serous ovarian cancer expression profiles from GSE13876 on the GPL7759 platform.

4. The method according to claim 1, characterized in that, Step B: Genes significantly associated with overall survival in patients with serous ovarian cancer are screened using a Cox regression model. These significantly associated genes are then paired to obtain candidate prognostic gene pairs. A gene pair matrix is ​​obtained based on the relative expression order of these candidate prognostic gene pairs. Based on this gene pair matrix, gene pairs significantly associated with the prognosis of serous ovarian cancer are screened using a Cox regression model, including: The Cox regression model was used to identify individual genes in the training set. When the p-value was less than 0.05, the gene was identified as a gene that was significantly associated with the overall survival of ovarian cancer patients. The significantly related genes are paired to obtain candidate prognostic gene pairs, and a gene pair matrix is ​​obtained based on the relative expression order of the candidate prognostic gene pairs. The candidate prognostic-related gene pairs were identified using a Cox regression model. When the p-value was less than 0.05, the gene pair was considered to be significantly associated with the prognosis of ovarian cancer.

5. The method according to any one of claims 1-4, characterized in that, Step C: The gene pairs significantly associated with the prognosis of serous ovarian cancer are selected in a forward selection order, choosing the top N gene pairs with the highest concordance index (C-index). The C-index is an index describing the predictive ability of the prognostic model. A greedy algorithm is used to screen each gene pair in the top N groups, and then the gene pairs are tested using the test set. The resulting set of gene pairs serves as the prognostic biomarker model for serous ovarian cancer, including: Using each gene pair significantly associated with ovarian cancer prognosis as a seed, the remaining gene pairs are added to the combination one by one. If the consistency index C-index value increases after adding a gene pair, gene pairs are added to the combination. If the consistency index C-index value decreases after adding a gene pair, no gene pairs are added. This process continues until the consistency index C-index value no longer increases, resulting in N sets of gene pairs. The consistency index C-index is an index that describes the predictive ability of the prognostic model. Based on the N gene pairs, they are sorted from largest to smallest according to their consistency index (C-index). The N gene pairs are then tested using the test set to screen out the gene pair with the highest consistency index (C-index) that is associated with the prognosis of serous ovarian cancer. This gene pair is then used as the final prognostic biomarker model for serous ovarian cancer.

6. A method for validating a prognostic biomarker model for serous ovarian cancer, characterized in that, The method includes: Step A: Obtain expression profile data and clinical information of patients with serous ovarian cancer from the gene expression comprehensive database, preprocess the expression profile data and clinical information, and divide the serous ovarian cancer expression profile datasets from different detection platforms into three data subsets: one data subset as the training set, one data subset as the test set, and another data subset as the validation set. Step B: Genes significantly associated with overall survival of patients with serous ovarian cancer are screened using a Cox regression model. The significantly associated genes are paired to obtain candidate prognostic gene pairs. A gene pair matrix is ​​obtained based on the relative expression order of the candidate prognostic gene pairs. Gene pairs significantly associated with the prognosis of serous ovarian cancer are screened using a Cox regression model based on the gene pair matrix. Step C: Select the top N gene pairs with the highest consistency index C-index value from the gene pairs that are significantly associated with the prognosis of serous ovarian cancer in a forward selection order. Use a greedy algorithm to screen each gene pair in the top N gene pairs. Then use the test set to test the combination of gene pairs. Use the resulting gene pair as a prognostic biomarker model for serous ovarian cancer. Step D: Validate the prognostic biomarker model for serous ovarian cancer on the validation set, specifically including: performing survival analysis, functional enrichment analysis, and immune infiltration analysis on the prognostic biomarker model for serous ovarian cancer.

7. The method according to claim 6, characterized in that, The survival analysis of the prognostic biomarker model for serous ovarian cancer includes: When the gene affects G a G b The order of expression in the sample is G. a Greater than G b If the gene affects G, the sample is considered high-risk. a G b The order of expression in the sample is G. a Less than or equal to G b If the sample is deemed low-risk, then a multivariate Cox regression analysis is performed, taking into account age and stage information.

8. The method according to claim 7, characterized in that, The functional enrichment analysis of the prognostic biomarker model for serous ovarian cancer includes: Functional enrichment analysis was performed on the prognostic biomarker model of serous ovarian cancer based on the Kyoto Gene, Genome Encyclopedia, and Gene Ontology databases in Metascape. Gene annotation and pathway analysis were performed on the prognostic gene markers in the Human Disease-Related Genes and Mutation Sites Database to search for pathways related to cancer occurrence and development in gene marker pairs.

9. The method according to claim 8, characterized in that, The immune infiltration analysis of the prognostic biomarker model for serous ovarian cancer includes: The content of immune cells between high- and low-risk groups in the training set was estimated using TIMER2.0, and the Wilcoxon test was used for comparison.

10. The method according to claim 9, characterized in that, The training set includes six types of immune cells between high- and low-risk groups: B cells, CD4+ T cells, CD8+ T cells, neutrophils, macrophages, and dendritic cells.

Citation Information

Patent Citations

  • Stomach cancer prognostic marker screening and classifying method based on gene expression profile

    CN106407689A

  • Molecular model for judging prognosis of ovarian cancer patient and application

    CN112680523A