A method for screening tumor prognostic molecular markers specifically expressed in cell subpopulations of individual samples

Through gene expression data decomposition and co-expression network analysis, subpopulations of cells with significant survival differences were screened out, and hub genes were screened out as molecular markers of tumor prognosis, which solved the problem of inaccurate breast cancer typing and achieved accurate breast cancer molecular typing and prognosis evaluation.

CN115295071BActive Publication Date: 2025-09-02HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210928358.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-03
Publication Date
2025-09-02
Estimated Expiration
2042-08-03

AI Technical Summary

Technical Problem

Existing breast cancer typing methods cannot accurately reflect the differences in cell subpopulations in individual samples, resulting in inconsistent treatment effects. Traditional typing methods cannot provide accurate individualized treatment plans.

Method used

By decomposing the gene expression data, subpopulations of cells with significant survival differences were screened, co-expression networks were constructed, and hub genes were screened as molecular markers of tumor prognosis, and used for molecular typing and prognosis evaluation of breast cancer.

Benefits of technology

Accurate typing and prognostic evaluation of breast cancer is achieved, providing a better foundation for the design of treatment plans, and revealing the mechanism of the role of cell subpopulations in breast cancer heterogeneity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115295071B_ABST
    Figure CN115295071B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for screening tumor prognosis molecular markers specifically expressed by individual sample cell subpopulations. First, the gene expression data is analyzed and decomposed to obtain estimates of population cell subpopulation-specific expression; then, by combining it with clinical information, a prognostic evaluation of population cell subpopulation-specific expression is obtained, and cell subpopulations with significant survival differences are further decomposed to obtain estimates of individual sample cell subpopulation-specific expression; finally, by constructing a co-expression module of cell subpopulations, genes with a high degree of connectivity are obtained as tumor prognosis molecular markers for tumor molecular typing and prognosis evaluation. The invention is applied to breast cancer tumors. The tumor prognosis molecular markers obtained by this method are used to type breast cancer. This is a molecular typing of breast cancer at the level of individual cell subpopulation heterogeneity. This typing method is more capable of explaining the changes and development of the disease in a complex biological environment, laying the foundation for the design of targeted drugs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of biotechnology and medicine, and relates to a method for screening tumor prognosis molecular markers specifically expressed by cell subpopulations of individual samples. Background Art

[0002] Breast cancer is a highly heterogeneous disease, with variations in biological behavior, treatment responses, and prognosis. Even with the same clinical stage and treatment regimen, patients may not respond identically to treatment, making both diagnosis and treatment extremely challenging. Traditional pathological morphology and classification are no longer sufficient to fully understand the true nature of breast cancer. To better predict patient prognosis and provide precise, personalized treatment plans, researchers are using various high-throughput molecular techniques to investigate the rational classification of breast cancer.

[0003] Currently, there are five generally accepted molecular classifications of breast cancer: luminal A, luminal B, HER2-positive, basal-like, and normal-like. Studies have shown that these five molecular classifications accurately predict prognosis. For example, luminal A and luminal B have a lower risk of recurrence and metastasis and a better prognosis, while HER2-positive and basal-like have a higher risk of malignancy and a worse prognosis. However, the prognostic power of these five classifications is limited. For example, the normal-like type is found in all four other breast cancer types, with similar expression patterns to normal breast tissue and adenofibromas, leading to debate about whether a normal-like type exists. Furthermore, the five molecular classifications derived from this classification method may share common molecular features, making it impossible to determine which type warrants specific drug treatment. Therefore, more precise classification of breast cancer is essential.

[0004] Cancer tissue is a heterogeneous mixture composed of different cell subtypes, so it is crucial to understand the heterogeneity of breast cancer based on cell subpopulations. Currently, only researchers have reclassified breast cancer by estimating the average cell subpopulation-specific expression in the population. Because the composition and expression of cell subpopulations in individual samples will vary under different biological states or conditions, and changes in specific cell subpopulations can lead to changes and development of the disease. Therefore, estimating the unique cell subtype expression of individual samples and identifying molecular markers with clinical and biological significance from the differential expression and co-expression networks of cell subpopulations are of great significance for studying the molecular typing of breast cancer and providing precise clinical treatment. Summary of the Invention

[0005] The purpose of the present invention is to address the shortcomings of the prior art and provide a method for screening tumor prognostic molecular markers specifically expressed by cell subpopulations in individual samples. The tumor prognostic molecular markers are obtained by screening cell subpopulations with significant survival differences. The cell subpopulation ratio matrix reflects the proportion of gene expression information in each subpopulation and is the main reason for the differences in patient expression data. The present invention is applied to breast cancer tumors and can be used to classify breast cancer from the perspective of genomic heterogeneity. It can be used to assist in genotyping and prognostic assessment of breast cancer and has good application prospects in precision medicine.

[0006] The present invention provides a method for screening tumor prognostic molecular markers specifically expressed by cell subpopulations in individual samples, and applies the method to breast cancer typing and prognosis assessment, comprising the following steps:

[0007] Step (1), preprocessing of gene expression data:

[0008] Obtain gene expression data, remove non-randomly missing data, and select training and test sets.

[0009] Step (2), estimation of cell subpopulation-specific expression:

[0010] First, the optimal number of subpopulations K was determined based on the minimum description length. Then, the mixture convex analysis model (debCAM model) was used to decompose the preprocessed gene expression data to obtain the optimal decomposition model, thereby obtaining the estimated value of the average cell subpopulation-specific expression in the population and the subpopulation ratio matrix.

[0011] Step (3), survival analysis of cell subpopulations:

[0012] First, the cell subpopulation ratio matrix obtained in step (2) is combined with the sample clinical information into survival information and preprocessed; then, the Kaplan-Meier survival curve is estimated based on the optimal classification threshold of each subpopulation, and combined with the COX univariate analysis, the cell subpopulations with significant survival differences are screened as characteristic subpopulations for prognosis prediction.

[0013] Step (4), estimation of cell subset-specific expression in individual samples:

[0014] According to the estimated values ​​of population cell subpopulation-specific expression and the subpopulation ratio matrix in step (2), the preprocessed gene expression data are decomposed using the extended mixture convex analysis model (swCAM model) to obtain estimates of the specific expression of individual sample cell subpopulations.

[0015] Step (5): Construction of subpopulation co-expression network and screening of modules:

[0016] 5-1 Screening out cell subpopulations with significant survival differences according to step (3), and selecting individual cell subpopulation-specific expression values ​​of cell subpopulations with significant survival differences from the estimation of cell subpopulation-specific expression of individual samples in step (4);

[0017] 5-2 The weighted gene co-expression network analysis method (WGCNA) was used to construct a co-expression network based on the individual cell subpopulation-specific expression values ​​of cell subpopulations with significant survival differences;

[0018] 5-3 The Kruskal-Wallis test (KW test) was used to analyze the first principal component of the gene module in the co-expression network and screen out the modules with significant differences.

[0019] Step (6), screening of molecular markers:

[0020] The connectivity between genes in the significant difference module is calculated, and genes with higher connectivity are selected as hub genes. The hub genes are the molecular markers for tumor molecular typing.

[0021] Preferably, in step (1), the method for processing the non-randomly missing data is: deleting genes that are not expressed in more than 10% of the samples; and the ratio of the training set to the test set is 2:1.

[0022] Preferably, in step (2), the optimal number of subgroups obtained according to the minimum description length is K=7.

[0023] Preferably, in step (3), the clinical information of the sample is ten-year follow-up time data; and the preprocessing is to treat the data with a survival time of more than ten years as censored data.

[0024] Preferably, in step (4), the parameter lambda in the swCAM model is 1400, and the parameter Iteradmm is 1000.

[0025] As an advantage, in step (5), the module fusion threshold selected when constructing the network is 0.25

[0026] Preferably, in step (6), the top 20 genes in terms of connectivity are selected as hub genes.

[0027] Preferably, in step (6), the tumor is breast cancer.

[0028] Compared with the prior art, the present invention has the following beneficial effects:

[0029] This invention uses gene expression data to decompose the specific expression of cell subpopulations in individual samples, providing molecular markers for these cell subpopulations to classify breast cancer into two subtypes and perform prognostic assessments. By studying the individual differential expression of cell subpopulations, this invention obtains molecular markers for tumor prognosis in these cell subpopulations. Compared with molecular markers derived from the differential expression of cell subpopulations as a whole, these markers can accurately molecularly subtype breast cancer and provide better prognostic predictions. This group of molecular markers is conducive to exploring the mechanism of action of cell subpopulations on breast cancer, laying the foundation for the classification of breast cancer at the level of cell subpopulation heterogeneity and the design of targeted drugs. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 It is a box plot of the cell subset ratio in Example 2.

[0031] Figure 2 These are the survival curves of the cell subpopulations with significant survival differences in Example 2, (a) is the survival curve of subpopulation 1, and (b) is the survival curve of subpopulation 4.

[0032] Figure 3 This is the survival curve of breast cancer molecular typing in Example 5. DETAILED DESCRIPTION

[0033] The present invention will be further described below with reference to the accompanying drawings and implementation examples.

[0034] As mentioned above, in view of the shortcomings of the existing technology, the inventors of this case have proposed the technical solution of the present invention after long-term research. Its main basis includes at least: first, analyzing the gene expression data and decomposing it to obtain estimates of the specific expression of population cell subpopulations; then, by combining it with clinical information, a prognostic evaluation of the specific expression of population cell subpopulations is obtained, and the cell subpopulations with significant survival differences are further decomposed to obtain estimates of the specific expression of individual sample cell subpopulations; finally, by constructing a co-expression module of cell subpopulations, genes with high connectivity are obtained as molecular markers for molecular typing and prognostic evaluation of breast cancer. Compared with the existing breast cancer typing, the molecular markers obtained by this method for breast cancer typing are molecular typing of breast cancer at the level of cell subpopulation heterogeneity. This typing method is more capable of explaining the changes and development of diseases in complex biological environments, laying the foundation for the design of targeted drugs.

[0035] At the same time, the present invention proposes a Kruskal-Wallis test (KW test) to further extract significant differences in the co-expression modules of differentially expressed subgroups, which facilitates the subsequent screening of molecular markers.

[0036] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and Examples. It should be understood that the specific embodiments described herein are merely for the purpose of explaining the present invention and are not intended to limit the present invention. In addition, the analytical methods in the following examples are conventional methods unless otherwise specified.

[0037] Example 1 Preprocessing of gene expression data

[0038] The gene expression matrix of 298 breast cancer patients (GSE17705) was preprocessed to remove non-randomly missing genes, that is, genes that were not expressed in more than 10% of the samples were deleted; and 200 patients were randomly sampled at a ratio of 2:1 as the training set.

[0039] Example 2 Estimation of Breast Cancer Cell Subpopulations and Ratio Matrix

[0040] 1) Estimation of population cell subpopulation-specific expression: The training set randomly sampled in Example 1 was decomposed using a mixture convex analysis model (debCAM model), and the minimum value of the minimum description length was obtained when K = 7; K = 7 was used as the optimal number of subpopulations and an optimal decomposition model was constructed to obtain the average cell subpopulation-specific expression and subpopulation ratio matrix in the population. Figure 1 The proportions of seven cell subpopulations are shown, among which the average proportion of subpopulation 1 and subpopulation 7 is larger, and the average proportion of subpopulation 5 and subpopulation 6 is smaller, indicating that subpopulation 1 and subpopulation 7 are the most abundant.

[0041] 2) Survival Analysis of Cell Subpopulations: First, the clinical information corresponding to the training set obtained in Example 1 was preprocessed, including setting the follow-up period to ten years and censoring the survival time. This information was then combined with the subpopulation ratio matrix obtained in step 1), and the seven subpopulations were analyzed one by one using the Cox univariate method to identify two cell subpopulations that had an impact on survival. Finally, the Kaplan-Meier method was used to perform survival analysis on these two cell subpopulations, screening for cell subpopulations with differential survival and plotting KM curves.

[0042] Table 1 shows the results of the COX univariate analysis of the subgroups, which shows that subgroups 1 and 4 have a significant impact on sample survival. Figure 2 The survival curves of subgroup 1 and subgroup 4 are shown. Figure 2 (a) The survival curve of subgroup 1 shows that the optimal threshold for subgroup 1 is 0.19. Subgroup 1 is divided into two groups based on the threshold, and the survival difference between the two groups is statistically significant (P = 0.0016); Figure 2 (b) The survival curve of subgroup 4 shows that the optimal threshold for subgroup 4 is 0.15. Subgroup 4 is divided into two groups based on the threshold, and the survival difference between the two groups is statistically significant (P = 0.0014).

[0043] Table 1: COX univariate analysis results of cell subsets

[0044]

[0045] Example 3 Estimation of cell subpopulation-specific expression in individual samples

[0046] The training set obtained in Example 1, the cell subpopulation ratio matrix obtained in Example 2, and the average cell subpopulation-specific expression in the population were used as input data, and the mixture convex analysis extended model (swCAM model) was used to decompose them to obtain the subpopulation differential expression of the seven subpopulations in individual samples.

[0047] Example 4 Construction of subpopulation co-expression network and module screening

[0048] a) Construction of a co-expression network: The subpopulations 1 and 4 obtained in Example 2 were mapped one-to-one to the subpopulation differential expressions in the individual samples obtained in Example 3. The weighted gene co-expression network analysis method (WGCNA) was used, and the module merging threshold was set to 0.25. The co-expression modules of subpopulations 1 and 4 were obtained. Subpopulations 1 and 4 had 38 and 23 co-expression modules, respectively.

[0049] b) Module screening: The first principal component matrix (MEs) of each module in subpopulation 1 and subpopulation 4 was derived from the co-expression network constructed in step a), and the Kruskal-Wallis test (KW test) was used to screen the modules with significant differences.

[0050] Table 2 shows the P values ​​of subgroup 1 and subgroup 4 in the Kruskal-Wallis test. It can be seen that the P value of the dark magenta module in subgroup 1 is the most significant (P = 8.42E-04), and the P value of the dark turquoise module in subgroup 4 is the most significant (P = 4.59E-04).

[0051] Table 2: KW test results of subgroup 1 and subgroup 4

[0052]

[0053] Example 5 Screening of molecular markers

[0054] The connectivity between the genes in the darkmagenta module and the darkturquoise module obtained in Example 4 was calculated, and the top 20 genes with the highest connectivity were selected as hub genes. These hub genes are molecular markers for molecular typing of breast cancer. Table 3 shows the molecular markers for subgroups 1 and 4.

[0055] Table 3: Molecular markers of subpopulations

[0056]

[0057] Example 6 KEGG analysis of molecular markers

[0058] The molecular markers obtained in Example 5 were uploaded to the KOBAS3.0 online tool, and the top four KEGG pathways with the highest enrichment were selected. By analyzing the biological pathways of the modules, it is helpful to find the biological explanation for the prognostic differences. Table 4 below shows the molecular markers in the darkmagenta module and the darkturquoise module and the top four KEGG pathways with the highest enrichment. It can be seen that the molecular markers of subgroup 1 are mainly enriched in the glutathione metabolism-related pathway (P = 5.69E-09), and the molecular markers of subgroup 4 are mainly enriched in the vascular smooth muscle contraction-related pathway (P = 4.60E-05).

[0059] Table 4: KEGG analysis of molecular markers

[0060]

[0061] Example 7 Breast cancer classification and prognosis assessment

[0062] i) Breast cancer classification: Consensus clustering was used, the clustering algorithm was set to Hierarchical Clustering, and the distance was calculated as Spearman distance. The first principal component matrix (ME) with significant differences in subpopulations 1 and 4 obtained in Example 3 was clustered to obtain two breast cancer molecular classifications.

[0063] ii) Prognostic evaluation of breast cancer typing: The breast cancer typing results obtained in step i) were combined with the clinical information obtained in Example 2, where 1 and 2 represent subtype 1 and subtype 2, respectively. Prognostic evaluation was performed using the Kaplan-Meier method.

[0064] Figure 3 The survival curves of sample typing were displayed. It can be seen that the samples were typed using molecular markers screened out from the specific expression of individual sample cell subpopulations, and the two categories (subtype 1 and subtype 2) obtained had significant survival differences (P=4.59E-06), among which subtype 2 had a better prognosis than subtype 1. At the same time, compared with the use of molecular markers screened out from the specific expression of population cell subpopulations for prognosis prediction, combining subpopulation 1 (glutathione metabolism-related pathway) and subpopulation 4 (vascular smooth muscle contraction-related pathway) using molecular markers screened out from the specific expression of individual sample cell subpopulations for prognosis prediction had a better effect. This further shows that the molecular markers of these two subpopulations can be used as potential prognostic molecular markers, which is beneficial for our understanding of the pathogenesis of the disease.

[0065] The above embodiments are not limitations of the present invention, and the present invention is not limited to the above embodiments. As long as the requirements of the present invention are met, they belong to the protection scope of the present invention.

Claims

1. A method for screening tumor prognostic molecular markers specifically expressed by cell subpopulations in individual samples, characterized in that The following steps are involved: Step (1), preprocessing of gene expression data: Obtain gene expression data, remove non-randomly missing data, and select training and test sets; Step (2), estimation of cell subpopulation-specific expression: First, the optimal number of subpopulations, K, was determined based on the minimum description length. Then, a mixture convex analysis model was used to decompose the preprocessed gene expression data to obtain the optimal decomposition model, thereby obtaining the estimated value of the average cell subpopulation-specific expression in the population and the subpopulation ratio matrix. Step (3), survival analysis of cell subpopulations: First, the cell subpopulation ratio matrix obtained in step (2) is combined with the sample clinical information into survival information and preprocessed; then, the Kaplan-Meier survival curve is estimated based on the optimal classification threshold of each subpopulation, and combined with the COX univariate analysis, the cell subpopulations with significant survival differences are screened as characteristic subpopulations for prognosis prediction; Step (4), estimation of cell subset-specific expression in individual samples: Based on the estimated values ​​of population cell subpopulation-specific expression and the subpopulation ratio matrix in step (2), the preprocessed gene expression data are decomposed using the mixture convex analysis extension model to obtain estimates of the specific expression of individual sample cell subpopulations; Step (5): Construction of subpopulation co-expression network and screening of modules: 5-1 Screening out cell subpopulations with significant survival differences according to step (3), and selecting individual cell subpopulation-specific expression values ​​of cell subpopulations with significant survival differences from the estimation of cell subpopulation-specific expression of individual samples in step (4); 5-2 Use weighted gene co-expression network analysis method to construct a co-expression network based on the individual cell subpopulation-specific expression values ​​of cell subpopulations with significant survival differences; 5-3 The Kruskal-Wallis test was used to analyze the first principal component of the gene module in the co-expression network and screen out the modules with significant differences; Step (6), screening of molecular markers: The connectivity between genes in the significant difference module is calculated, and genes with higher connectivity are selected as hub genes. The hub genes are the molecular markers for tumor molecular typing.

2. The method according to claim 1, characterized in that In step (1), the method for processing the non-random missing data is to delete genes that are not expressed in more than 10% of the samples; the ratio of the training set to the test set is 2:

1.

3. The method according to claim 1, characterized in that In step (2), the optimal number of subgroups obtained according to the minimum description length is K=7.

4. The method according to claim 1, characterized in that In step (3), the clinical information of the sample is the ten-year follow-up time data; the preprocessing is to treat the data with a survival time of more than ten years as missing data.

5. The method according to claim 1, characterized in that In step (4), the parameter lambda in the mixture convex analysis extension model is 1400 and the parameter Iteradmm is 1000.

6. The method according to claim 1, characterized in that In step (5), the module fusion threshold selected when constructing the network is 0.

25.

7. The method according to claim 1, characterized in that In step (6), the top 20 genes with the highest connectivity were selected as hub genes.

8. The method according to claim 1, characterized in that In step (6), the tumor is breast cancer.

Citation Information

Patent Citations

  • Cell imaging and analysis to differentiate clinically relevant sub-populations of cells

    CA2977262A1

  • T-cell subset for cancers and characteristic genes

    CN109081866A