Gene marker combinations for predicting ovarian cancer prognosis, immunotherapy response, or immunophenotype and uses thereof

CN122445798APending Publication Date: 2026-07-24NANJING MEDICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANJING MEDICAL UNIV
Filing Date
2026-06-23
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

现有技术中,针对卵巢癌特定癌相关成纤维细胞亚群的稳定基因标志物组合仍较少,难以将单细胞水平发现的功能性细胞亚群转化为可用于样本层面检测和评估的分子工具

Benefits of technology

本申请的基因标志物组合为MMP11+ CAF细胞亚群的特征基因集,基于所述基因标志物组合的表达水平,能够获得MMP11+ CAF细胞亚群富集评分。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122445798A_ABST
    Figure CN122445798A_ABST
Patent Text Reader

Abstract

This application discloses a combination of gene markers for predicting ovarian cancer prognosis, immunotherapy response, or immunophenotype, and their applications, belonging to the field of molecular diagnostic technology for ovarian cancer. The combination of gene markers includes... MMP1 , GJB2 , PLPP4 , MME , GREM1 , COL10A1 , CTHRC1 , CEMIP , MMP11 and NTM The combination of gene markers presented in this application can not only predict the prognostic risk of ovarian cancer, but also predict the immunotherapy response and immunophenotype of ovarian cancer samples, with excellent performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of molecular diagnostic technology for ovarian cancer, specifically relating to a combination of gene markers for predicting the prognosis, immunotherapy response, or immune phenotype of ovarian cancer and their applications. Background Technology

[0002] Ovarian cancer is one of the most common gynecological malignancies, characterized by its insidious onset, rapid progression, high recurrence rate, and significant prognostic variability. Most patients are already in advanced stages at the time of diagnosis. Although surgery combined with platinum-based chemotherapy remains the primary treatment, some patients still experience recurrence, drug resistance, or limited treatment benefit. Therefore, establishing molecular markers that can assist in the diagnosis, recurrence monitoring, and prognostic assessment of ovarian cancer is of great significance for improving patient stratification and individualized treatment.

[0003] Currently, clinical assessment of ovarian cancer mainly relies on imaging examinations, serum tumor marker detection, histopathological examination, and clinical staging. While these methods are valuable in disease diagnosis and treatment decisions, their ability to indicate differences in tumor microenvironment, recurrence risk, and treatment response remains limited. In particular, single serum markers or traditional clinicopathological indicators are insufficient to fully reflect the complex cellular composition and molecular heterogeneity within tumor tissue.

[0004] In recent years, advancements in single-cell sequencing and spatial transcriptomics technologies have demonstrated that cancer-associated fibroblasts, immune cells, and extracellular matrix components within the tumor microenvironment can collectively influence the progression, recurrence, and treatment response of ovarian cancer. Among these, cancer-associated fibroblasts exhibit significant heterogeneity, with different subpopulations displaying varying distribution locations, activation states, and biological functions within tumor tissues. Current technologies lack a sufficient set of stable gene markers targeting specific cancer-associated fibroblast subpopulations in ovarian cancer, hindering the transformation of functional cell subpopulations identified at the single-cell level into molecular tools suitable for sample-level detection and evaluation.

[0005] Therefore, there is still a need in the field for a combination of gene markers that can reflect the pro-tumor stroma activation status in the ovarian cancer tumor microenvironment, and further be used for ovarian cancer diagnosis, recurrence monitoring, prognostic risk stratification, treatment response assessment, or immunophenotypic indication, in order to make up for the shortcomings of existing assessment methods in identifying tumor microenvironment characteristics. Summary of the Invention

[0006] To solve at least one of the above-mentioned technical problems, the technical solution adopted in this application is as follows: The first aspect of this application provides a combination of gene markers for predicting the prognosis, immunotherapy response, or immune phenotype of ovarian cancer, the combination of gene markers comprising: MMP1 , GJB2 , PLPP4 , MME , GREM1 , COL10A1 , CTHRC1 , CEMIP , MMP11 and NTM .

[0007] Furthermore, the combination of genetic markers also includes TMEM158 , CA12 , RSAD2 , AL139393 2. EVA1A , COL11A1 , SFRP2 , ACKR4 , ABHD17C and SERINC2 At least one gene in the sample, for example, including one, two, three, four, five, six, seven, eight, nine, or ten of them.

[0008] In this application, including MMP1 , GJB2 , PLPP4 , MME , GREM1 , COL10A1 , CTHRC1 , CEMIP , MMP11 and NTMI A combination of gene markers can serve as MMP11 +Cancer-associated fibroblast subsets ( MMP11 The characteristic gene set of the + CAF cell subset, the MMP11 + CAF cell subsets are used as a single indicator to predict ovarian cancer prognosis, immunotherapy response, or immune phenotype.

[0009] In this application, predicting the prognosis of ovarian cancer refers to predicting a good or bad prognosis for ovarian cancer patients. A good prognosis means that the disease does not progress (including disease recurrence or patient death) within 180 days after ovarian cancer treatment; a bad prognosis means that the disease progresses within 180 days after ovarian cancer treatment; disease progression includes disease recurrence or patient death.

[0010] In this application, predicting the response to immunotherapy for ovarian cancer refers to predicting whether an ovarian cancer patient will respond to immunotherapy.

[0011] In this application, predicting the immune phenotype of ovarian cancer refers to predicting whether an ovarian cancer patient is immune-rejecting or non-immune-rejecting.

[0012] The second aspect of this application provides a cancer-associated fibroblast subpopulation whose characteristic genes include any combination of gene markers described in the first aspect of this application.

[0013] The third aspect of this application provides the use of a reagent for detecting the expression level of the combination of gene markers described in the first aspect of this application or a reagent for detecting cancer-associated fibroblast subsets described in the second aspect of this application in the preparation of a kit for predicting the prognosis, immunotherapy response, or immunophenotype of ovarian cancer.

[0014] In some embodiments of this application, the expression levels of the gene biomarker combination are obtained based on RNA samples by at least one method selected from the group consisting of reverse transcription PCR, quantitative real-time PCR, in situ hybridization, RNA sequencing, and microarray chips.

[0015] in: Reverse transcription polymerase chain reaction (RT-PCR): This is a commonly used method that transcribes RNA into cDNA through reverse transcription, then amplifies the target gene or genomic fragment using PCR, and finally quantifies the RNA level using gel electrophoresis or quantitative PCR.

[0016] Real-time quantitative PCR (qPCR) is a highly sensitive method for quantitatively detecting RNA levels. It can monitor the increase in DNA during the PCR process in real time, thereby quantifying the RNA content.

[0017] In situ hybridization (ISH) is a method used to detect the location and distribution of specific RNA sequences in tissue sections or cells. It typically uses a labeled probe with affinity to bind to the RNA sequence, which is then observed under a microscope.

[0018] RNA sequencing (RNA-Seq) is a high-throughput method that transcribes RNA into cDNA and then sequences the cDNA to obtain comprehensive information about the RNA, including the expression levels of different genes and splicing variations.

[0019] Microarray: A commonly used high-throughput technology that uses probes fixed on the surface of a chip to detect RNA in a sample, thereby analyzing RNA expression levels.

[0020] In some embodiments of this application, the detection reagent includes primers and / or probes.

[0021] In some embodiments of this application, the expression level is the transcriptional level obtained using RNA sequencing methods, specifically including: RNA samples of the test samples were obtained, and after library construction, high-throughput sequencing was performed to obtain sequencing results. The sequencing results are compared to a reference genome to obtain the transcriptional level.

[0022] In some embodiments of this application, the transcription level includes, but is not limited to, Count, FPKM, TPM, CPM, etc.

[0023] The fourth aspect of this application provides a kit for predicting the prognosis, immunotherapy response, or immunophenotype of ovarian cancer, including a kit for detecting the expression level of a combination of gene markers as described in the first aspect of this application or a kit for detecting cancer-associated fibroblast subsets as described in the second aspect of this application.

[0024] The fifth aspect of this application provides a system for predicting the prognosis, immunotherapy response, or immune phenotype of ovarian cancer, comprising: The data input module is used to obtain the expression level of each gene in the gene marker combination described in the first aspect of this application or the level of the cancer-associated fibroblast subset described in the second aspect of this application; The prediction module, connected to the data input module, is used to predict the prognosis, immunotherapy response, or immune phenotype of ovarian cancer based on the expression levels of the genes or the levels of the cancer-associated fibroblast subsets using statistical or artificial intelligence methods.

[0025] Artificial Intelligence (AI) is an interdisciplinary field based on computer science, integrating multiple subjects such as mathematics, informatics, and control.

[0026] Furthermore, the artificial intelligence method is implemented through machine learning. Machine learning encompasses many different algorithms, with deep learning being one of them. Other methods include logistic regression, decision trees, random forests, support vector machines, Naive Bayes, K-nearest neighbors, and neural networks.

[0027] In some embodiments of this application, a gene set scoring method is used to calculate the enrichment score of the gene biomarker combination in the test sample. The gene set scoring method refers to: obtaining the expression information of each gene in the gene biomarker combination in the test sample, and after standardization, sorting, weighting, or enrichment calculation, converting the expression information of multiple genes into a score value that reflects the overall level of the gene biomarker combination.

[0028] In some embodiments of this application, the gene expression matrix of the sample to be tested, such as transcriptome data, is first obtained, and the expression information of each gene in the combination of gene markers is extracted.

[0029] In this application, the score obtained based on the combination of the gene markers can be used to represent the strength of the transcriptional features associated with the MMP11+ CAF cell subset in the sample to be tested.

[0030] In some preferred embodiments of this application, the gene set scoring method employs Gene Set Variation Analysis (GSVA) ​​to calculate the enrichment score of the MMP11+ CAF cell subset in each sample based on the combination of gene markers. GSVA can estimate the relative level of the combination of gene markers in each sample based on the sample expression matrix without the need for pre-setting sample grouping, thereby obtaining the enrichment score of the MMP11+ CAF cell subset.

[0031] In other embodiments of this application, the gene set scoring method may also employ single-sample gene set enrichment analysis, gene expression rank-based scoring, summation or averaging based on standardized expression values, weighted expression value-based methods, principal component analysis-based methods, or other methods capable of converting the expression information of multiple genes in the feature gene set into a single-sample score. Any method that can obtain a score characterizing the strength of transcriptional features associated with the MMP11+ CAF cell subset based on the gene marker combination can be used as the calculation method for the MMP11+ CAF cell subset enrichment score of this application.

[0032] It should be noted that this invention does not limit the MMP11+ CAF cell subset enrichment score to be obtained through a specific software package, a specific algorithm name, or specific parameters. For the same test sample, as long as the scoring calculation process is based on the expression level of the gene marker combination and outputs a value that reflects the overall expression activity, relative enrichment degree, or strength of transcriptional features related to the MMP11+ CAF cell subset, it is considered an equivalent implementation of the scoring calculation method described in this application.

[0033] In some embodiments of this application, if the enrichment score calculated using the combination of the gene markers is higher than a preset threshold, it is predicted that the ovarian cancer patient will have a poor prognosis, or that the ovarian cancer patient will not respond to immunotherapy (is insensitive), or that the ovarian cancer patient will be immune-rejected.

[0034] In some specific embodiments of this application, the preset threshold is obtained based on a population sample. Specifically, the enrichment score is calculated in the population sample using the combination of the gene markers, and the preset threshold is further determined based on all the enrichment scores obtained in the population sample.

[0035] A sixth aspect of this application provides a computer device, comprising: a memory for storing a computer program; and a processor for executing the computer program to perform the following steps: The expression levels of each gene in the combination of gene markers described in the first aspect of this application or the expression levels of the cancer-associated fibroblast subsets described in the second aspect of this application are obtained in the sample to be tested. Based on the expression levels of the aforementioned genes or the levels of the aforementioned cancer-associated fibroblast subsets, statistical or artificial intelligence methods are used to predict the prognosis, immunotherapy response, or immunophenotype of ovarian cancer.

[0036] A seventh aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, performs the following steps: The expression levels of each gene in the combination of gene markers described in the first aspect of this application or the expression levels of the cancer-associated fibroblast subsets described in the second aspect of this application are obtained in the sample to be tested. Based on the expression levels of the aforementioned genes or the levels of the aforementioned cancer-associated fibroblast subsets, statistical or artificial intelligence methods are used to predict the prognosis, immunotherapy response, or immunophenotype of ovarian cancer.

[0037] This application also provides a method for predicting the prognosis, immunotherapy response, or immunophenotype of ovarian cancer, the method comprising the following steps: A predictive model is constructed in a computer using expression level data of any combination of gene markers described in the first aspect of this application or level data of cancer-associated fibroblast subpopulations described in the second aspect of this application in biological samples of the population, wherein the population includes ovarian cancer patients and non-ovarian cancer subjects. The expression level data of the combination of the aforementioned gene markers or the level data of the aforementioned cancer-associated fibroblast subsets in the subject's biological samples are input into the prediction model to obtain prediction results for ovarian cancer prognosis, immunotherapy response, or immune phenotype. The prediction model is constructed using statistical or artificial intelligence methods.

[0038] This application also provides the use of the cancer-associated fibroblast subsets described in the second aspect of this application as targets in screening immunotherapeutic drugs for the treatment of ovarian cancer.

[0039] Compared with the prior art, this application has the following technical effects: The gene marker combination in this application is a characteristic gene set of the MMP11+ CAF cell subpopulation. Based on the expression level of the gene marker combination, the enrichment score of the MMP11+ CAF cell subpopulation can be obtained.

[0040] The enrichment score of MMP11+ CAF cell subsets can be used to predict the prognostic risk of ovarian cancer; the higher the score, the worse the prognosis.

[0041] The MMP11+ CAF cell subset enrichment score can also be used to predict the immunotherapy response of ovarian cancer samples. An increased MMP11+ CAF cell subset enrichment score tends to indicate a decreased immunotherapy response in ovarian cancer samples, while a low MMP11+ CAF cell subset enrichment score suggests that the sample has a higher tendency to respond to immunotherapy.

[0042] The MMP11+ CAF cell subset enrichment score can also be used to predict immune phenotypes, specifically, to predict whether an immune rejection type is present. When the MMP11+ CAF cell subset enrichment score exceeds a preset threshold, an immune rejection type is predicted. Attached Figure Description

[0043] Figure 1 The ten cell types and their proportions identified in Example 1 of this application are shown; Figure 2 The cell types and proportions in normal and tumor tissues in Example 1 of this application are shown; Figure 3 This application illustrates the various fibroblast subpopulations and the cell origins of the F1 and F2 subpopulations in Example 1. Figure 4 The differentiation potential and pseudo-temporal trajectory analysis results of the MMP11+ CAF cell subset in Example 1 of this application are shown. Figure 5 The image shows fibroblast UMP images of clinical data from Example 1 of this application; Figure 6 This paper illustrates the expression distribution of 10 core characteristic genes of the MMP11+ CAF cell subpopulation in the fibroblast subpopulation of a clinically derived ovarian cancer single-cell sample in Example 2 of this application. Figure 7 The characteristic gene set of the MMP11+ CAF cell subset in Example 2 of this application is enriched in the extracellular matrix-receptor interaction pathway; Figure 8 The MMP11+ CAF cell subset in Example 2 of this application expresses typical myCAF-related marker molecules and exhibits obvious matrix remodeling characteristics; Figure 9 The results of prognostic stratification Kaplan-Meier survival analysis based on the characteristic gene set of the MMP11+ CAF cell subset in Example 3 of this application are shown. Figure 10 The distribution of MMP11+ CAF cell subset enrichment scores in normal samples, borderline tumor samples, malignant tumor samples, and ovarian cancer samples at different stages is shown in Example 3 of this application. Figure 11The results of TIDE analysis of the characteristic gene set of the MMP11+ CAF cell subset in Example 4 of this application for low immunotherapy response propensity in an ovarian cancer cohort are shown. Figure 12 The results of a low-response association analysis of the characteristic gene set of the MMP11+ CAF cell subset in an external immunotherapy cohort are shown in Example 4 of this application. Figure 13 This illustrates the relationship between high and low enrichment scores of MMP11+ CAF cell subsets in clinically derived ovarian cancer single-cell samples and chemotherapy sensitivity status in Example 4 of this application. Figure 14 The results of the prediction of ovarian cancer immune phenotype by the characteristic gene set of the MMP11+ CAF cell subset in Example 5 of this application are shown. Figure 15 The results of spatial distribution analysis of the characteristic gene set of the MMP11+ CAF cell subset in the spatial transcriptome of ovarian cancer are shown in Example 5 of this application. Figure 16 The results of spatial expression module analysis co-localization with the MMP11+ CAF cell enrichment region in Example 5 of this application are shown. Figure 17 The results of correlation analysis of MMP11+ CAF cell subset enrichment score, extracellular matrix-receptor interaction pathway score and TGF-β signaling pathway score in clinically derived ovarian cancer single-cell samples in Example 5 of this application are shown. Detailed Implementation

[0044] Unless otherwise stated, implied from the context, or as is customary in the art, all parts and percentages in this application are based on weight, and all testing and characterization methods used are concurrent with the filing date of this application. Where applicable, any patent, patent application, or disclosure relating to this application is incorporated herein by reference in its entirety, and its equivalent patent families are also incorporated herein by reference, particularly the definitions of relevant terms in the art disclosed in such documents. If any definition of a specific term disclosed in the prior art is inconsistent with any definition provided in this application, the definition provided in this application shall prevail.

[0045] The following examples are used to illustrate preferred embodiments of this application. Those skilled in the art will understand that the techniques disclosed in the examples represent technologies discovered by the inventors that can be used to implement this application, and therefore can be considered preferred embodiments of this application. However, those skilled in the art should understand from this specification that many modifications can be made to the specific embodiments disclosed herein, still yielding the same or similar results, without departing from the spirit or scope of this application.

[0046] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains, and all materials cited herein and referenced by them are incorporated herein by reference.

[0047] Those skilled in the art will recognize, or can learn through routine experimentation, many equivalents of the specific embodiments of the invention described herein. These equivalents will be included in the claims.

[0048] Unless otherwise specified, the experimental methods used in the following examples are conventional methods. Unless otherwise specified, the instruments and equipment used in the following examples are all conventional laboratory instruments and equipment; unless otherwise specified, the experimental materials used in the following examples were all purchased from conventional biochemical reagent stores.

[0049] Example 1: Identification of cancer-associated cell subsets in the ovarian cancer microenvironment This embodiment illustrates how to identify cancer-associated fibroblast subpopulations enriched in tumors from single-cell transcriptome data of ovarian cancer, and provides a foundation for subsequent feature gene set extraction and application analysis.

[0050] Step 1. Acquisition, quality control, and standardization of single-cell sequencing data Single-cell transcriptome data of ovarian tissues publicly available in the GEO database were obtained: GSE211956, GSE184880, and GSE173682, which contain a total of 19 ovarian-related samples, including normal ovarian tissue samples and ovarian cancer tissue samples.

[0051] The single-cell transcriptome data were imported into the Seurat software package (version 5.0.3) for preprocessing. Low-quality cells were removed, specifically: (1) Remove cells with a mitochondrial gene expression rate greater than 20% and a number of genes less than 200 or greater than 5000; (2) Remove genes that are expressed in fewer than 3 cells.

[0052] Subsequently, NormalizeData was used for standardization, FindVariableFeatures was used to screen for the top 2000 hypervariable genes, ScaleData was used for data scaling, and RunPCA was used for principal component analysis, retaining the top 50 principal components for subsequent analysis. After quality control and integration, a final 66,347 high-quality cells were retained.

[0053] Step 2. Cell clustering, cell type annotation, and fibroblast population extraction Based on the single-cell objects obtained in step 1, each sample was processed using NormalizeData and FindVariableFeatures, and 2000 hypervariable genes were screened using the VST method. Then, SelectIntegrationFeatures was used to select integration features, and FindIntegrationAnchors was used to identify anchor points. Canonical Correlation Analysis (CCA) was used to integrate the whole-cell objects. After anchor point identification, IntegrateData was used to complete data integration, with the default assay set to integrated. ScaleData, RunPCA, FindNeighbors, RunTSNE, and RunUMAP analyses were then performed, with the principal component count set to 50, and the nearest neighbor graph construction and dimensionality reduction analysis using a dims=1:30 ratio. Finally, FindClusters was used for cluster analysis, with a cluster resolution set to 1.5.

[0054] Based on the differentially expressed genes of each cell population and the classic cell type marker genes, the clustering results were annotated with cell types, as shown in Table 1.

[0055] Table 1: Cell types and their molecular markers .

[0056] Based on the above analysis, 10 major cell subpopulations can be identified, such as Figure 1 As shown, compared with normal tissue, the number of fibroblasts in tumor tissue is significantly increased, such as... Figure 2 As shown, it is one of the main stromal components that constitute the ovarian cancer microenvironment.

[0057] Step 3. Reclustering of fibroblast subsets and identification of cancer-associated fibroblast subsets The fibroblast population extracted in step 2 was further standardized, dimensionality reduced, and clustered to obtain a refined fibroblast subpopulation structure. In this embodiment, the Seurat analysis workflow was followed, and hypervariable gene screening, principal component analysis, proximity graph construction, and unsupervised clustering were re-executed. Further RunUMAP, RunTSNE, FindNeighbors, and FindClusters analyses were performed with a dims=1:30 and a clustering resolution of 0.5.

[0058] Reclustering results showed that fibroblasts could be further divided into 8 different subpopulations, each named after its highly expressed gene. Specifically: F0-TCF7-CAF represents the F0-TCF7 positive cancer-associated fibroblast subpopulation, F1-OGN-Fibroblast represents the F1-OGN positive fibroblast subpopulation, F2-MMP11-CAF represents the F2-MMP11 positive cancer-associated fibroblast subpopulation, F3-MYH11-CAF represents the F3-MYH11 positive cancer-associated fibroblast subpopulation, F4-RGS5-CAF represents the F4-RGS5 positive cancer-associated fibroblast subpopulation, F5-CCL5-CAF represents the F5-CCL5 positive cancer-associated fibroblast subpopulation, F6-WFDC2-CAF represents the F6-WFDC2 positive cancer-associated fibroblast subpopulation, and F7-RAMP1-CAF represents the F7-RAMP1 positive cancer-associated fibroblast subpopulation. Among them, the F2 subgroup is mainly characterized by high expression of MMP11, while the F1 subgroup is mainly characterized by high expression of OGN. Further comparison of the distribution of each subgroup in different tissue origins shows that the F2 subgroup is mainly distributed in tumor tissues, while the F1 subgroup is mainly distributed in normal tissues, such as... Figure 3 As shown in the figure, HGSOC represents high-grade serous ovarian cancer, and NC represents a normal control.

[0059] Based on the above results, the F2 subset can be defined as a tumor-enriched cancer-associated fibroblast (CAF) cell subset (MMP11+ CAF cell subset) in ovarian cancer tissue.

[0060] Step 4. Cell state of the MMP11+ CAF cell subset To further confirm the cellular state of the MMP11+ CAF cell subset, pseudo-temporal analysis and differentiation potential analysis were performed on the fibroblast subset. In this embodiment, CytoTRACE2 (version 0.3.3) was used to assess cell differentiation potential, and Monocle2 (version 2.22.0) was used to reconstruct the pseudo-temporal trajectory of transformation from resident fibroblasts to cancer-associated fibroblasts. In the pseudo-temporal analysis, differentially expressed genes were used as input features, and the DDRTree algorithm was used for dimensionality reduction.

[0061] Analysis results as follows Figure 4 As shown, the MMP11+ CAF cell subset is located at the activation end of the trajectory from OGN+ resident fibroblasts to the FAP-high state, indicating that this subset belongs to the activated fibroblast state. Combined with its enrichment characteristics in tumor tissue, this cell subset can be further identified as a tumor-enriched activated cancer-associated fibroblast subset in the ovarian cancer microenvironment.

[0062] Step 5. Validation of the MMP11+ CAF cell subset in clinically derived ovarian cancer single-cell samples To verify the presence of the MMP11+ CAF cell subset in clinically derived ovarian cancer samples, eight clinically derived ovarian cancer single-cell samples were further selected for verification.

[0063] Table 2: Information on single-cell samples of ovarian cancer from clinical sources .

[0064] Using the previously annotated fibroblast subpopulations as a reference dataset (where the F2 subpopulation is defined as the MMP11+CAF cell subpopulation), the transcriptome data of the above-mentioned clinically derived ovarian cancer single-cell samples were used as the test dataset. The fibroblast subpopulation labels in the reference dataset were mapped to the test dataset through anchor matching and label transfer methods.

[0065] Mapping results showed that a type of fibroblast predicted to be F2-MMP11-CAF-like existed in the transcriptome data of clinically derived single-cell samples. Figure 5 As shown in the figure. This result suggests that the MMP11+ CAF subset identified in publicly available data can be detected in clinically derived ovarian cancer single-cell samples.

[0066] The above results indicate that the MMP11+ CAF cell subset identified in publicly available data can be mapped to clinically derived ovarian cancer single-cell samples and is detectable in real clinical samples.

[0067] Example 2: Extraction of characteristic gene set of MMP11+ CAF cell subset and analysis of its biological function This embodiment illustrates how to extract the characteristic gene set of the MMP11+ CAF cell subpopulation identified in Example 1, and how to perform biological function analysis on the characteristic gene set.

[0068] The MMP11+ CAF cell subpopulation identified in Example 1 was used as the target cell subpopulation, and differential expression analysis was performed on it and other fibroblast subpopulations. The differential expression analysis was performed using the FindAllMarkers function in Seurat software, and the screening criteria were set as follows: the target gene was expressed in at least 25% of the target cell subpopulation, and |log2FC|>0.25.

[0069] To determine the optimal number of genes in the characteristic gene set of the MMP11+ CAF cell subset, differentially expressed genes of the F2-MMP11-CAF subset relative to other fibroblast subsets were sorted from largest to smallest based on |log2FC|>0.25. The top 30 genes are shown in Table 3.

[0070] Table 3: Differentially expressed genes in the F2-MMP11-CAF subset compared to other fibroblast subsets .

[0071] First, the top 10 ranked genes were selected to construct a Top10 characteristic gene set, and the enrichment score (signature score) of the MMP11+ CAF cell subpopulation was calculated using Gene Set Variation Analysis (GSVA). Then, based on the Top10 characteristic gene set, candidate genes ranked lower were added sequentially to construct Top11, Top12, Top13, and up to Top20 characteristic gene sets, and their enrichment scores in the MMP11+ CAF cell subpopulation were calculated. The results are shown in Table 4.

[0072] Table 4: Gene sets and enrichment scores of different characteristic gene subsets of MMP11+ CAF cells .

[0073] As shown in Table 4, the enrichment score of the MMP11+ CAF cell subpopulation calculated using the Top10 gene set was higher than that of other fibroblast subpopulations in different samples, indicating that the Top10 gene set can stably identify the MMP11+ CAF cell subpopulation.

[0074] Based on the Top10 gene set, one gene ranked 11 was added to form the Top11 gene set. Based on the Top11 gene set, one gene ranked 12 was added to form the Top12 gene set, and so on, forming the Top11~Top20 gene sets. The enrichment scores were calculated for each set. It was found that although the enrichment scores of the Top11~Top20 gene sets were always higher than those of other fibroblast subpopulations (as shown in Table 5), the enrichment scores did not increase with the increase of the number of genes. For some gene sets, the enrichment scores even decreased, indicating that increasing the number of genes in the gene sets did not produce a significant additional subpopulation distinguishing advantage.

[0075] Table 5: Enrichment scores calculated using different characteristic gene sets in different cell subpopulations .

[0076] After obtaining the Top 10 characteristic gene set of the MMP11+ CAF cell subset, which includes 10 genes, the Top 10 characteristic gene set was further validated using transcriptome data from the aforementioned clinically derived ovarian cancer single-cell samples. Specifically: Expression distribution maps of 10 genes in different fibroblast subpopulations of single-cell ovarian cancer samples from clinical sources were plotted. The results are as follows: Figure 6 As shown, the 10 genes are mainly highly expressed in the F2-MMP11-CAF cell subset, while their expression is relatively low in other fibroblast cell subsets. This indicates that the above Top 10 characteristic gene set can well characterize the MMP11+ CAF cell subset in clinically derived ovarian cancer single-cell samples and can be used as a combination of gene markers to characterize the MMP11+ CAF cell subset.

[0077] After obtaining the aforementioned Top 10 characteristic gene set, further biological function analysis was performed. The results showed that the MMP11+ CAF cell subset was significantly enriched in extracellular matrix-related processes, such as... Figure 7 The extracellular matrix-receptor interaction pathway is shown.

[0078] Meanwhile, the MMP11+ CAF cell subset also expresses typical myofibroblastic CAF (myCAF)-related molecules, including the periosteum protein gene POSTN and the collagen-related gene COL11A1, and exhibits significant matrix remodeling characteristics, such as... Figure 8 As shown.

[0079] This indicates that the Top 10 characteristic gene set is not a typical collection of differentially expressed genes, but rather a myCAF-like transcriptional program that characterizes a significant extracellular matrix feature and matrix remodeling capability.

[0080] Based on the above analysis, the characteristic gene set of the MMP11+ CAF cell subset obtained in this embodiment is defined as: a set of gene markers derived from a tumor-enriched activated cancer-associated fibroblast subset in the ovarian cancer tumor microenvironment, capable of characterizing myCAF-like status, extracellular matrix remodeling capacity, and pro-tumor stroma activation features. This gene set can be used as a whole as a set of gene markers for subsequent prognostic risk stratification, chemotherapy sensitivity assessment, and immunophenotypic indication.

[0081] Example 3: Application of Gene Marker Combinations in Prognostic Risk Stratification of Ovarian Cancer This example illustrates that the characteristic gene set of the MMP11+ CAF cell subset obtained in Example 2 can be used for prognostic risk stratification of ovarian cancer samples.

[0082] Multiple independent bulk transcriptome cohorts were selected to validate the characteristic gene set. These cohorts included GSE211669, GSE135820, GSE51088, and TCGA-OV, with the TCGA-OV cohort containing 317 serous ovarian cancer samples. Gene set variation analysis (GSVA) ​​was used to calculate the enrichment score (signature score) of each CAF subpopulation. For the MMP11+ CAF cell subpopulation, the enrichment score of the MMP11+ CAF cell subpopulation for each sample was calculated using the Top 10 gene set determined in Example 2. After obtaining the enrichment scores of the MMP11+ CAF cell subpopulations, the samples were divided into high-score and low-score groups based on the median enrichment score. Kaplan-Meier curves were used for survival analysis, and inter-group differences were assessed using the log-rank test. The results are as follows: Figure 9 As shown, in both the GSE211669 and GSE135820 datasets, the high-scoring group had a worse survival rate.

[0083] Further comparisons were made of the distribution of MMP11+ CAF cell subset enrichment scores in the GSE51088 and TCGA-OV datasets across different disease categories and clinical stages. Results are as follows: Figure 10 As shown, the enrichment score of MMP11+ CAF cell subsets showed an increasing trend in different disease categories, including normal samples, borderline tumor samples, and malignant tumor samples, and also showed an increasing trend in stage II–IV samples.

[0084] This indicates that the enrichment score of the MMP11+ CAF cell subset corresponding to the gene marker combination can distinguish ovarian cancer samples with different prognostic risks, with a high score indicating a poor prognosis. Therefore, the MMP11+ CAF cell subset enrichment score can be used as a prognostic risk stratification indicator for ovarian cancer.

[0085] Example 4: Application of Gene Marker Combinations in the Adjunctive Assessment of Low Response Propensity to Immunotherapy in Ovarian Cancer This example illustrates that the characteristic gene set of the MMP11+ CAF cell subset obtained in Example 2 can be used as an auxiliary assessment of the tendency of ovarian cancer samples to have a low response to immunotherapy.

[0086] The characteristic gene set was validated using an ovarian cancer bulk transcriptome cohort, which included GSE32062, GSE49997, and GSE211669. Enrichment scores for each subset of cancer-associated fibroblasts were calculated using gene set variation analysis. For the MMP11+ CAF cell subset, the enrichment score for each sample was calculated using the Top 10 gene set determined in Example 2. Furthermore, the median enrichment score of the MMP11+ CAF cell subset was used as the grouping criterion to divide the samples into a high-score group and a low-score group.

[0087] In the aforementioned ovarian cancer cohort, the TIDE method was further used to predict immunotherapy response propensity, and the distribution of predicted responders and predicted non-responders was statistically analyzed in both the high-score and low-score groups. The results are as follows: Figure 11 As shown, in the three cohorts GSE32062, GSE49997, and GSE211669, the proportion of predicted non-responders was higher in the high-score group, while the proportion of predicted responders was relatively higher in the low-score group. This result indicates that elevated MMP11+ CAF cell subset enrichment scores are associated with a lower propensity for immunotherapy response in ovarian cancer samples.

[0088] Further investigation revealed an association between higher MMP11+ CAF cell subset enrichment scores and lower predicted immunotherapy responses, supported not only in the ovarian cancer cohort but also in external tumor cohorts receiving immune checkpoint blockade therapy. Specifically, in the melanoma cohort GSE78220 and the colorectal cancer cohort GSE203071, after grouping by median MMP11+ CAF cell subset enrichment scores, the high-score group showed a higher proportion of predicted non-responders, while the low-score group showed a relatively higher proportion of predicted responders. Figure 12 As shown in the figure. This result further supports the view that a higher MMP11+ CAF cell subset enrichment score is associated with a low response to immunotherapy.

[0089] To further evaluate the applicability of the MMP11+ CAF cell subset enrichment score in stratifying treatment response in real clinical samples, supplementary analysis was performed using transcriptomic data from the aforementioned clinically sourced ovarian cancer single-cell samples and their corresponding clinical treatment response information. Specifically, following the methods described in Examples 1 and 2, MMP11+ CAF-like cells were first identified in the clinically sourced ovarian cancer single-cell data, and the MMP11+ CAF cell subset enrichment score for each clinical sample was calculated. Since the number of chemotherapy-resistant samples in this clinical cohort is relatively small, while the number of chemotherapy-sensitive samples is relatively large, directly comparing all cells might result in a significant imbalance in cell numbers between groups. Therefore, in this example, all fibroblasts in the chemotherapy-resistant samples were first retained, and random downsampling was performed from the fibroblasts corresponding to the chemotherapy-sensitive samples to extract cells with the same or similar numbers as the chemotherapy-resistant source cells, thus forming a relatively balanced clinical source comparison set. Subsequently, in the aforementioned balanced clinical source comparison set, cells were divided into high-score and low-score groups based on the median enrichment score of the MMP11+ CAF cell subset, and the distribution of chemotherapy-sensitive and chemotherapy-resistant cells in the two groups was compared. The results are as follows: Figure 13 As shown, in the balanced clinically derived cell set, the MMP11+ CAF cell subset enrichment score (high / low) and chemotherapy sensitivity status can be combined for stratified analysis. The composition of chemotherapy-sensitive and chemotherapy-resistant samples differs between the high-score and low-score groups, indicating that the MMP11+ CAF cell subset enrichment score can be further combined with clinical treatment response information for auxiliary stratification of treatment response-related status in ovarian cancer samples.

[0090] It should be noted that the sample size of the aforementioned clinical sample cohort is limited, and the chemotherapy sensitivity analysis is mainly used to illustrate that the MMP11+ CAF cell subset enrichment score and actual clinical treatment response information can be jointly evaluated. This embodiment does not limit the MMP11+ CAF cell subset enrichment score to a single indicator for judging chemotherapy sensitivity or chemotherapy resistance, nor does it limit it to a direct indicator for judging the efficacy of a specific treatment regimen.

[0091] Therefore, in this embodiment, the MMP11+ CAF cell subset enrichment score can be used as an auxiliary assessment indicator of the tendency for low immunotherapy response in ovarian cancer samples. When the MMP11+ CAF cell subset enrichment score of the sample is elevated, it suggests that the sample has a higher tendency for non-response to immunotherapy or a lower probability of benefit from immunotherapy; when the MMP11+ CAF cell subset enrichment score is low, it suggests that the sample has a higher tendency for immunotherapy response. Simultaneously, combined with chemotherapy sensitivity information from clinically sourced samples, the MMP11+ CAF cell subset enrichment score can also be used for auxiliary stratified analysis of ovarian cancer treatment response-related status.

[0092] Example 5: Application of Gene Marker Combinations in Predicting Ovarian Cancer Immunophenotype This example illustrates that the characteristic gene set (Top 10 characteristic gene set) of the MMP11+ CAF cell subset obtained in Example 2 can be used to predict the immune phenotype in the ovarian cancer tumor microenvironment.

[0093] In the ovarian cancer transcriptome data (TCGA-OV), the corresponding MMP11+ CAF cell subset enrichment scores were calculated based on the Top 10 characteristic gene set. Further comparison of the MMP11+ CAF cell subset enrichment scores in samples with different immunophenotypes revealed that the MMP11+ CAF cell subset enrichment scores were significantly higher in immune-excluded samples than in immune-inflamed and immune-desert samples. Figure 14 As shown in the figure. The above results indicate that a combination of gene markers can reflect the tumor stroma state associated with immune rejection and can serve as a predictive indicator for immune rejection in ovarian cancer.

[0094] To further verify the spatial distribution characteristics of the above gene marker combinations, spatial transcriptome samples from ovarian cancer (GSE213699, PMID: 36249907) were analyzed, and the enriched regions of the MMP11+ CAF cell subset were mapped to the spatial transcriptome data. The analysis results are as follows: Figure 15 As shown, in the immune rejection sample (sample S2), the enriched areas of the MMP11+ CAF cell subset were mainly located around the tumor and at the tumor-stromal junction, forming a dense stromal boundary around the tumor epithelial region. However, no similar obvious distribution pattern was observed in the immune inflammation sample (sample S1).

[0095] This embodiment also analyzes the spatial representation module co-located with the MMP11+ CAF niche. The results are as follows: Figure 16As shown, the spatial expression module co-localizes with the MMP11+ CAF enrichment region, and genes related to this module are enriched in extracellular matrix-related processes. Furthermore, the spatial expression module score is strongly correlated with the MMP11+ CAF enrichment region derived from single cells. These results further demonstrate that the immune rejection pattern predicted by the above gene marker combination is closely related to extracellular matrix enrichment and interstitial remodeling.

[0096] Furthermore, to verify whether the MMP11+ CAF enrichment region in clinically derived ovarian cancer samples could reflect immune rejection-related matrix remodeling characteristics, supplementary analysis was performed using transcriptomic data from the aforementioned clinically derived ovarian cancer single-cell samples. Specifically, firstly, following the methods described in Examples 1 and 2, fibroblast populations were identified in the transcriptomic data of clinically derived ovarian cancer single-cell samples, and the MMP11+ CAF cell subset enrichment score for each clinical sample was calculated. Subsequently, pathway scores for each clinical sample were calculated based on the extracellular matrix-receptor interaction pathway (hsa04512) related gene set from the KEGG PATHWAY database, and pathway scores for each clinical sample were calculated based on the TGF-β signaling pathway (hsa04350) related gene set. After obtaining the MMP11+ CAF cell subset enrichment score, extracellular matrix-receptor interaction pathway score, and TGF-β signaling pathway score for each clinical sample, the correlation between these scores was further analyzed. The results are as follows: Figure 17 As shown, the enrichment score of the MMP11+ CAF cell subset in clinical samples showed a positive correlation with the extracellular matrix-receptor interaction pathway score, and also with the TGF-β signaling pathway score. These results suggest that in clinically derived ovarian cancer cell samples, samples with higher MMP11+ CAF cell subset enrichment scores exhibit both stronger extracellular matrix remodeling and TGF-β signaling characteristics.

[0097] The transcriptomic data from the aforementioned clinically derived ovarian cancer single-cell samples further support the view that the MMP11+ CAF cell subset enrichment score can not only characterize the enrichment status of the MMP11+ CAF cell subset, but also reflect the transcriptional status of the tumor microenvironment related to the stromal barrier, extracellular matrix remodeling, TGF-β signaling, and immune rejection. It should be noted that the transcriptomic data from clinically derived ovarian cancer single-cell samples are primarily used to validate the relationship between MMP11+ CAF cell subset enrichment and immune rejection-related transcriptional features, and are not limited to a direct determination of spatial immune rejection structures.

[0098] Therefore, the MMP11+ CAF cell subset enrichment score can be used as a predictive indicator of ovarian cancer immunophenotype. When the MMP11+ CAF cell subset enrichment score of a sample increases, and it simultaneously exhibits characteristics such as immune rejection, spatial matrix boundary formation, increased scores in extracellular matrix-receptor interaction pathways, or increased scores in TGF-β signaling pathways, the sample is predicted to have an immune rejection-related microenvironment tendency. It should be noted that the prediction of immune rejection in this embodiment is mainly based on transcriptome feature scores, spatial distribution analysis, and correlation analysis, used to characterize the stromal state related to immune rejection in the tumor microenvironment; this embodiment does not limit the direct determination of the efficacy of specific immunotherapies.

[0099] Furthermore, it should be understood that after reading the foregoing contents of this application, those skilled in the art can make various alterations or modifications to this application, and these equivalent forms also fall within the scope defined by the appended claims.

Claims

1. A combination of gene markers for predicting ovarian cancer prognosis, immunotherapy response, or immune phenotype, characterized in that, The combination of genetic markers includes MMP1 , GJB2 , PLPP4 , MME , GREM1 , COL10A1 , CTHRC1 , CEMIP , MMP11 and NTM .

2. The combination of gene markers according to claim 1, characterized in that, Also includes TMEM158 , CA12 , RSAD2 , AL139393 2. EVA1A , COL11A1 , SFRP2 , ACKR4 , ABHD17C and SERINC2 At least one gene in it.

3. The use of the gene marker combination expression level detection reagent according to claim 1 or 2 in the preparation of a kit for predicting ovarian cancer prognosis, immunotherapy response or immunophenotype.

4. The application according to claim 3, characterized in that, The expression levels of the gene biomarker combination were obtained based on RNA samples using at least one method selected from reverse transcription PCR, quantitative real-time PCR, in situ hybridization, RNA sequencing, and microarray chips.

5. The application according to claim 4, characterized in that, The detection reagents include primers and / or probes.

6. A kit for predicting ovarian cancer prognosis, immunotherapy response, or immunophenotype, characterized in that, The kit includes an expression level detection reagent for the combination of gene markers as described in claim 1 or 2.

7. A system for predicting the prognosis, immunotherapy response, or immune phenotype of ovarian cancer, characterized in that, include: A data input module is used to obtain the expression level of each gene in the gene marker combination as described in claim 1 or 2; The prediction module, connected to the data input module, is used to predict the prognosis, immunotherapy response, or immune phenotype of ovarian cancer based on the expression levels of each gene using statistical or artificial intelligence methods.

8. The system according to claim 7, characterized in that, The artificial intelligence method is a machine learning method, which is selected from one of the following: logistic regression, decision tree, random forest, support vector machine, Naive Bayes, K-nearest neighbor, and neural network.

9. A computer device, characterized in that, include: Memory, used to store computer programs; A processor, used to perform the following steps when executing the computer program: Obtain the expression level of each gene in the combination of gene markers described in claim 1 or 2 in the sample to be tested; Based on the expression levels of each gene, statistical or artificial intelligence methods are used to predict the prognosis, immunotherapy response, or immune phenotype of ovarian cancer.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the following steps: Obtain the expression level of each gene in the combination of gene markers described in claim 1 or 2 in the sample to be tested; Based on the expression levels of each gene, statistical or artificial intelligence methods are used to predict the prognosis, immunotherapy response, or immune phenotype of ovarian cancer.