A prognosis model based on breast tumor stem cell related genes and application thereof
By constructing a prognostic model based on genes related to breast cancer stem cells, screening out characteristic genes and combining them with clinical variables, the problem of low accuracy in existing breast cancer prognostic models has been solved, achieving efficient and accurate prognostic assessment and treatment guidance for breast cancer.
Patent Information
- Application Number
- CN202310557532.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-17
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2043-05-17
AI Technical Summary
Existing breast cancer prognostic models suffer from low accuracy and high cost due to heterogeneity, making them unsuitable for effective clinical application. There is a lack of prognostic models based on genes related to breast cancer stem cells.
A prognostic model based on breast cancer stem cell-related genes was constructed. By screening genes such as BRD4, RPS24, SERPINA3, SKP1, NTRK3, CD79A, JAK1, NT5E, NDRG1, and CD24, a risk scoring model was constructed in combination with clinical variables. LASSO regression and multivariate Cox regression analysis were used to construct the BCSCRS score, and the survival probability of patients was predicted by nomogram.
It improved the predictive accuracy and precision of breast cancer prognostic models, reduced treatment costs, provided better treatment guidance, enhanced sensitivity to immunotherapy and chemotherapy, and improved predictive performance.
Smart Images

Figure CN116656820B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of medical artificial intelligence technology, specifically relating to a prognostic model based on breast tumor stem cell-related genes and its application. Background Technology
[0002] Breast cancer is a highly heterogeneous malignant tumor that occurs within breast tissue. Although traditional treatments such as surgery, chemotherapy, and radiotherapy, as well as emerging immunotherapies, have significantly improved prognosis, the heterogeneity of cancer leads to recurrence, metastasis, drug resistance, and immune escape, posing a significant challenge to clinical treatment of breast cancer. Therefore, a thorough understanding of cancer heterogeneity and its characteristics applied to clinical diagnosis and treatment will contribute to further improving clinical outcomes for cancer patients.
[0003] Existing reports include predictive models constructed using genes related to immune function (CN 113862363 A), ferroptosis, autophagy (CN 113593648A), apoptosis, lactate metabolism, macrophage characteristic genes, copper-dependent genes, and single-cell transcriptome sequencing analysis of breast cancer (CN 109072481 B, the constructed model involves 95 genes, making the model too complex). Some prognostic models involve the detection of nearly 100 genes, resulting in high industrialization costs and unsuitability for clinical application. Most prognostic models are constructed targeting specific signaling pathways or targets, often leading to low accuracy and precision in practical applications. This may be due to the heterogeneity of breast cancer. Therefore, there is currently no accurate and reliable prognostic model for breast cancer applied clinically.
[0004] Previous studies have speculated that breast cancer stem cells may be a source of heterogeneity; however, there are currently no reports or clinical applications of constructing breast cancer prognostic models based on genes related to breast cancer stem cells. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention provides a prognostic model based on breast tumor stem cell-related genes and its application.
[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0007] This invention provides a prognostic characteristic gene for constructing a breast cancer prognostic model, wherein the prognostic characteristic gene is a breast cancer stem cell-related gene. The prognostic characteristic gene includes BRD4, RPS24, SERPINA3, SKP1, NTRK3, CD79A, JAK1, NT5E, NDRG1, and CD24.
[0008] Based on the application of the prognostic feature genes in constructing a breast cancer prognostic model, the specific formula of the model is as follows: Breast Cancer Stem Cell Related Gene Risk Score (BCSCRS) = (-0.53045×BRD4) + (-0.26259×RPS24) + (-0.31334×SERPINA3) + (0.434039×SKP1) + (-0.53742×NTRK3) + (-0.23344×CD79A) + (-0.40628×JAK1) + (0.192005×NT5E) + (0.152866×NDRG1) + (0.194872×CD24); when BCSCRS < 0.985569, it is judged as a low-risk state, and when BCSCRS > 0.985569, it is judged as a high-risk state.
[0009] Furthermore, a nomogram was constructed based on risk scores of breast cancer stem cell-related genes and clinical variables to predict patients' overall survival and prognosis.
[0010] The clinical variables included gender, TNM stage, and age.
[0011] By adding the risk scores of breast cancer stem cell-related genes to clinical variables, an overall score was calculated for each patient, and a transformation function was used to estimate the survival probability of each patient at 1, 3, and 5 years. Patients with lower overall scores had a higher survival probability.
[0012] The method for constructing the prognostic model includes the following steps:
[0013] (1) Data acquisition: Transcriptome and clinical data of breast cancer in TCGA and GEO were acquired, normal tissue samples and samples from the same patient were excluded, and the createDataPartition function in the R package caret was used to randomly divide TCGA patients into training cohort and internal validation cohort in a 7:3 ratio. Patients in GSE20685 were used as external validation cohort.
[0014] (2) Breast cancer stem cell gene screening: Breast cancer stem cell related genes were collected from the GeneCards database, and genes with a correlation score greater than 30 were retained. Univariate Cox analysis was performed to screen breast cancer stem cell genes related to breast cancer prognosis.
[0015] (3) Construction of the prognostic model: In the TCGA training cohort, the Least Absolute Shrinkage Selection Operator (LASSO) Cox regression method was used to further screen prognostic characteristic genes for breast cancer. A breast cancer prognostic model was constructed through multivariate Cox regression analysis. The risk score of the prognostic model was BCSCRS = (-0.53045×BRD4) + (-0.26259×RPS24) + (-0.31334×SERPINA3) + (0.434039×SKP1) + (-0.53742×NTRK3) + (-0.23344×CD79A) + (-0.40) 628×JAK1)+(0.192005×NT5E)+(0.152866×NDRG1)+(0.194872×CD24); Based on the gene expression levels and corresponding regression coefficients in the Cox regression model, the risk score for each sample is calculated using the predict function in the survival package; Based on the median BCSCRS (0.985569), all patients are divided into high-risk and low-risk groups, i.e., when BCSCRS < 0.985569, they are assigned to the low-risk group, and when BCSCRS > 0.985569, they are assigned to the high-risk group;
[0016] (4) Validation of the prognostic model: The TCGA test cohort and the GSE20685 cohort were used as internal and external validation cohorts, respectively, to validate the accuracy of the prognostic model. ROC, C-index and DCA indicators were used to evaluate the accuracy of the prognostic model. (5) Construction of nomogram: The rms package was used to construct nomograms and calibration curves including patient age, gender, TNM stage and breast cancer stem cell related gene risk score (BCSCRS). The timeROC and ggDCA packages were used to perform receiver operating characteristic (ROC) and decision curve analysis (DCA) to compare the predictive accuracy of nomograms with other prognostic factors. The nomograms constructed with BCSCRS and clinical variables were used as a quantitative method to predict the survival rate of breast cancer patients.
[0017] This invention also provides the application of a detection reagent in the preparation of a kit for assessing the prognostic survival of breast cancer. The detection reagent is composed of reagents for detecting the expression levels of the following 10 genes: BRD4, RPS24, SERPINA3, SKP1, NTRK3, CD79A, JAK1, NT5E, NDRG1, and CD24. This detection reagent is the sole key component for the kit to achieve its function of assessing the prognostic survival of breast cancer. The kit also includes an instruction manual, which states the following: BCSCRS = (-0.53045 × BRD4) + (-0.26259 × RPS24) + (-0.31334×SERPINA3)+(0.434039×SKP1)+(-0.53742×NTRK3)+(-0.23344×CD79A)+(-0.40628×JAK1)+(0.192005×NT5E)+(0.152866×NDRG1)+(0.194872×CD24); When BCSCRS < 0.985569, it is judged as a low-risk state; when BCSCRS > 0.985569, it is judged as a high-risk state. Patients with low-risk state have a better prognosis than patients with high-risk state.
[0018] Preferably, the specimens tested using the kit are fresh tissue tumor specimens.
[0019] The advantages of this invention are:
[0020] 1. This invention obtains transcriptomic and clinical data on breast cancer from TCGA and GEO, screens for survival-related breast cancer stem cell genes, and constructs a prognostic model related to breast cancer stem cells using LASSO regression and multivariate Cox regression. ROC, C-index, and DCA indicators are used to compare the accuracy of this model with other existing prognostic models. Based on this research, to better apply the prognostic model clinically, we constructed a nomogram based on risk scores and clinical variables to predict overall patient survival. The prognostic model based on breast cancer stem cell-related genes constructed in this study showed good predictive accuracy in both the training and validation cohorts (AUC values of 0.7 at 1, 3, and 5 years). This model has better predictive accuracy compared to other models, with the nomogram incorporating clinical variables showing better predictive precision and accuracy (AUC = 0.758, C-index = 0.744). Furthermore, based on the BCSCRS risk score, the clinical tumor microenvironment and immune landscape analysis of immune infiltration and the assessment of the responsiveness to immune checkpoint inhibitor therapy were applied. The results showed that the low-risk group had a better response to immunotherapy and chemotherapy, which also proved that the prognostic model constructed based on breast cancer stem cell-related genes can be effectively applied to the clinical practice of breast cancer treatment and prognosis.
[0021] 2. The prognostic model based on breast tumor stem cell-related genes obtained by this method has the advantages of high sensitivity, good specificity and high accuracy. It can provide effective guidance for clinicians to make treatment decisions for breast cancer patients, reduce the occurrence of ineffective treatment, thereby reducing the treatment cost for patients and promoting the development and application of precision medicine in clinical practice.
[0022] 3. The nomogram constructed based on risk scores and clinical variables in this invention improves the predictive performance of the prognostic model for breast cancer prognosis. Attached Figure Description
[0023] Figure 1 This is a flowchart illustrating the construction of a prognostic model based on genes related to breast tumor stem cells.
[0024] Figure 2 For the development and validation of BCSCs-related prognostic models: (AB) Minimum absolute contraction and selection operators were used to further screen prognostic characteristic genes of breast cancer; (C) Forest plot of 10 prognostic genes of BCSCs; (D) Regression coefficients of each gene in the prognostic model; Scatter plots of 1-year, 3-year and 5-year survival in the TCGA training cohort (E), TCGA testing cohort (F), TCGA-total cohort (G), and GSE20685 testing cohort (H), Kaplan-Meier analysis, and time-dependent ROC curve analysis.
[0025] Figure 3 For the development and validation of nomograms: (A) Forest plot of univariate Cox regression analysis; (B) Forest plot of multivariate Cox regression analysis; (C) Nomograms predicting 1-year, 3-year, and 5-year survival probabilities of breast cancer patients based on risk scores and clinical factors; (D) Calibration curves of nomograms; (EG) 1-year, 3-year, and 5-year ROC curves of nomograms, BCSCRS, and clinical factors; (HJ) Decision curve analysis (DCA) of nomograms, BCSCRS, and clinical factors.
[0026] Figure 4 Figures showing the comparison of the prognostic value of different prognostic models for breast cancer: (AD) Kaplan-Meier survival curves for the prognostic models constructed by BCSCRS, Li et al., Wang et al., and Zhang et al., respectively; ROC curves (EH) for 1-year, 3-year, and 5-year overall survival; (I) Comparison of the overall ROC curves of the prognostic models involved in this study; (JL) C-index, RMS, and DCA analysis.
[0027] Figure 5Figures showing the results of the immune landscape analysis of the tumor microenvironment and immune infiltration: (A) GSEA enrichment analysis of high-risk and low-risk groups; (B) Heatmap showing the overall immune landscape of different risk groups; (C) Difference analysis of tumor microenvironment between different risk groups; (D) Difference analysis of immune infiltrating cells between two risk groups; (E) Correlation analysis of BCSCRS with immune infiltrating cells.
[0028] Figure 6 Figures showing the results of responsiveness to immune checkpoint inhibitor therapy: (A) Expression of 27 immune checkpoint molecules; (B) IPS analysis between the two risk groups.
[0029] Figure 7 The relationship between BCSCRS and tumor progression: (A) The relationship between BCSCRS and stage, age, sex and TNM stage in the TCGA cohort; (B) The relationship between BCSCRS and TNM stage in the GSE20685 cohort. Detailed Implementation
[0030] The abbreviations involved in this invention are:
[0031] BCSCs: Breast Cancer Stem Cells; IARC: International Agency for Research on Cancer; WHO: World Health Organization; PD-1: Programmed Death 1; CTLA-4: Cytotoxic T-Lymphocyte-Associated Antigen-4; GEO: Gene Expression Comprehensive Database; TCGA: Cancer Genome Atlas; NMF: Nonnegative Matrix Factorization; PCA: Principal Component Analysis; TMB: Tumor Mutation Burden; OS: Overall Survival; NES: Normalized Enrichment Score; WGCNA: Weighted Gene Co-expression Network Analysis; GO: Gene Ontology; KEGG: Kyoto Genome and Genome Hundreds Encyclopedia of Science; LASSO: Least Absolute Contraction Selection Operator; ROC: Receiver Operating Characteristic Curve; AUC: Area Under the Curve; DCA: Decision Curve Analysis; C-index: Concordance Index; GSEA: Gene Set Enrichment Analysis; TME: Tumor Microenvironment; ICIs: Immune Checkpoint Inhibitors; IPS: Immunophenotypic Score; IC50: 50% Maximum Inhibitory Concentration; BCSCRS: Breast Cancer Stem Cell-Related Risk Score; NK: Natural Killer Cells; TAMs: Tumor-Associated Macrophages; APC: Antigen Presenting Cells; TCR: T Cell Receptor;
[0032] Example 1: Construction of a prognostic model based on genes related to breast tumor stem cells.
[0033] The prognostic model construction process described in this invention is as follows: Figure 1 As shown.
[0034] 1. Data Acquisition:
[0035] Transcriptomic and clinical phenotypic data of breast cancer patients were obtained from the UCSC Genome Browser database (https: / / xenabrowser.net / datapages / ) and the GDC-TCGA-BRCA project in the Gene expression Omnibus (GEO) database (https: / / www.ncbi.nlm.nih.gov / geo / ). After excluding normal tissue samples and samples from the same patient, transcriptomic data from a total of 1069 breast cancer patients with complete clinical information were obtained from the TCGA database. Using the createDataPartition function in the R package caret, TCGA patients were randomly assigned to a training cohort (N=749) and an internal validation cohort (N=320) in a 7:3 ratio, and 327 patients from GSE20685 were used as an external validation cohort. Table 1 shows the clinical characteristics of breast cancer patients in all cohorts. Finally, breast cancer stem cell-related genes (BCSCGs) were collected from the GeneCards database (https: / / www.genecards.org / ), and genes with a correlation score greater than 30 were retained for subsequent analysis. Subsequently, univariate Cox analysis was performed to screen stem cell genes related to breast cancer prognosis for use in the construction of subsequent models.
[0036] Table 1. Clinical characteristics of breast cancer patients in all cohorts.
[0037]
[0038] 2. Construction and validation of prognostic models:
[0039] In this study, we calculated the risk score for each sample using the predict function in the survival package based on the gene expression levels and their corresponding regression coefficients in the model formula.
[0040] Least Absolute Shrunk Selection Operator (LASSO) Cox regression was used to further screen prognostic genes for breast cancer. Subsequently, LASSO Cox regression was used to further screen prognostic genes for breast cancer in the TCGA training cohort (N=749). A breast cancer prognostic model was constructed through multivariate Cox regression analysis. The TCGA test cohort (N=320) and the GSE20685 cohort (N=327) were used as internal and external validation cohorts, respectively, to verify the accuracy of the prognostic model. Then, based on the gene expression levels and their corresponding regression coefficients in the Cox regression model, the predict function in the survival package was used to calculate the risk score for each sample.
[0041] From 749 patients in the training cohort, we identified 45 breast cancer stem cell-related genes associated with survival using univariate Cox analysis. Next, we performed LASSO regression analysis to select 17 breast cancer stem cell-related genes for constructing a multivariate Cox regression model. Figure 2 (A and B). Then, we used multivariate Cox regression analysis to construct a prognostic model and identified 10 genes as prognostic characteristic genes (A and B). Figure 2 -C): BRD4, RPS24, SERPINA3, SKP1, NTRK3, CD79A, JAK1, NT5E, NDRG1, CD24. like Figure 2 - As shown in D and Table 2, the following formula was used to calculate the risk score (BCSCRS) for breast cancer stem cell-related genes:
[0042] BCSCRS=(-0.53045×BRD4)+(-0.26259×RPS24)+(-0.31334×SERPINA3)+(0.434039×SKP1)+(-0.53742×N TRK3)+(-0.23344×CD79A)+(-0.40628×JAK1)+(0.192005×NT5E)+(0.152866×NDRG1)+(0.194872×CD24).
[0043] Table 2. Regression coefficients of 10 characteristic genes in the prognostic model.
[0044]
[0045] The accuracy of the Cox regression model was assessed by calculating the area under the receiver operating characteristic (ROC) curve (AUC). Risk curves for all cohorts and survival maps for all patients were plotted using the pheatmap package. Kaplan-Meier (KM) analysis was used to compare overall survival (OS) between the high- and low-risk groups.
[0046] All patients were divided into high-risk and low-risk groups based on the median BCSCRS score. Specifically, patients with a BCSCRS score < 0.985569 (median risk score) were classified into the low-risk group, and patients with a BCSCRS score > 0.985569 were classified into the high-risk group.
[0047] In the TCGA training cohort (N=749), the overall survival rate of the low-risk group (N=374) was significantly higher than that of the high-risk group (N=375). Time-dependent ROC curve analysis showed that BCSCRS had good predictive accuracy in the TCGA training cohort, with AUCs of 0.733 (1 year), 0.742 (3 years), and 0.741 (5 years), respectively. Figure 2 -E); the 1-year, 3-year, and 5-year AUCs of the TCGA test queue were 0.808, 0.689, and 0.646, respectively. Figure 2 -F); the AUCs for the 1-year, 3-year, and 5-year total cohort of TCGA were 0.751, 0.728, and 0.707, respectively; the AUCs for the 1-year, 3-year, and 5-year external validation cohort of GSE20685 were 0.765, 0.718, and 0.692, respectively, indicating that BCSCRS has good accuracy in predicting survival. Figure 2 (G and H). This indicates that the prognostic model constructed based on 10 genes related to breast cancer stem cells has good accuracy and can be used in clinical practice for the prognosis and treatment of breast cancer.
[0048] The key feature of this invention in screening target genes is that it employs a dual dimensionality reduction screening method. Specifically, after using univariate Cox analysis to identify genes associated with breast cancer prognosis, it further uses LASSO regression to screen for prognostic characteristic genes of breast cancer. This helps us further reduce the data dimensionality and obtain more accurate prognostic genes.
[0049] Example 2: Construction of a Column Chart:
[0050] To improve the accuracy of the prognostic model and better apply it to clinical practice, we further developed a nomogram model that includes a risk score (i.e., the risk score of breast cancer stem cell-related genes, BCSCRS) and clinical variables (such as age and tumor stage).
[0051] First, univariate and multivariate Cox regression analyses were performed to assess whether risk scores and clinical variables could serve as independent prognostic factors. Then, nomograms and calibration curves incorporating patient age, sex, TNM stage, and risk scores were constructed using the rms package. To compare the predictive accuracy of the nomogram with other prognostic factors, receiver operating characteristic (ROC) and decision curve (DCA) analyses were performed using the timeROC and ggDCA packages, respectively.
[0052] The potential independence of BCSCRS as a prognostic factor was investigated using univariate and multivariate Cox regression analysis (Table 3).
[0053] Table 3. Univariate and multivariate Cox regression analysis of risk scores and clinical characteristics.
[0054]
[0055] *Independent prognostic factor.
[0056] The results showed that risk score, age, stage, and T, N, M stages were significantly associated with cancer prognosis. Figure 3 -A, p<0.001), multivariate Cox regression analysis showed that risk score and age could both serve as independent prognostic factors for breast cancer patients (-A, p<0.001). Figure 3 -B, p<0.001).
[0057] To explore the potential associations between BCSCRS and multiple clinical variables, Wilcoxon and Kruskal-Wallis tests were performed. Results showed that in the TCGA cohort, BCSCRS increased with increasing tumor stage, with significant differences between stages. Figure 7 -A). Risk scores for stages T and N showed an increasing trend with significant differences between groups, but the opposite was true for stage N3. Furthermore, patients with advanced stage M and those over 65 years of age had significantly higher BCSCRS scores. There were no statistically significant differences in risk scores between genders, and similar results were obtained in the GSE20685 cohort, where risk scores for advanced TNM stages were significantly higher. Figure 7 -B).
[0058] These findings indicate that the BCSCRS differs significantly among different clinical variable groups, with higher risk scores indicating a poorer pathological status in patients with breast cancer.
[0059] A nomogram was constructed by incorporating risk scores and clinical variables as a quantitative method for predicting the survival rate of breast cancer patients. Figure 3 -C). An overall score was calculated for each patient by summing the risk score and clinical variables (including sex, TNM stage, and age). Patients with lower overall scores had a higher probability of survival. This was determined using a calibration curve ( Figure 3 The accuracy of nomograms was assessed using the area under the ROC curve (AUC) and the nomogram itself. Nominal plots showed better predictive accuracy compared to other clinical features and raw risk scores. In the TCGA cohort, the 1-year, 3-year, and 5-year AUCs for nomograms were 0.805, 0.746, and 0.758, respectively. Figure 3 (E, F, G). Furthermore, decision curve analysis (DCA) confirms that the nomogram has better predictive accuracy compared to other predictive indicators. Figure 3 (H, I, J). This improves the accuracy of the prognostic model and allows for better application in clinical practice.
[0060] Example 3: Comparison and Analysis with Existing Breast Cancer Prediction Models
[0061] While we have demonstrated the accuracy of BCSCRS from multiple perspectives, the most important aspect of a clinical prognostic model is its superior predictive performance in clinical practice. To verify that the breast cancer prognostic model constructed in this study has better predictive performance, we compared it with three different prognostic models.
[0062] The first model is the ferroptosis-related prognostic model established by Wang et al. [1], which involves 9 genes including ALOX15, CISD1, CS, GCLC, GPX4, SLC7A11, EMC2, G6PD and ACSF2; the second model is the macrophage characteristic gene model constructed by Li et al. [2], which involves 7 genes including SERPINA1, CD74, STX11, ADAM9, CD24, NFKBIA and PGK1; the third model is the lactate metabolism-related prognostic model proposed by Zhang et al. [3], which involves 3 genes including LDHD, LYRM7 and PNKD. The model construction method is consistent with the literature, and in order to reduce the error caused by different data dimensions, we analyze the same transcriptome data, extract the expression level of genes in each model, and perform multivariate Cox regression to obtain the regression coefficient of each gene. Subsequently, the risk score of each sample is calculated, and the predictive ability and clinical utility of each model are evaluated by the consistency index (C-index) and decision curve analysis (DCA), as well as the receiver operating characteristic (ROC) curve and survival analysis. All analyses were performed using timeROC and the survival package in R software.
[0063] The references for the construction of the three different prognostic models compared in this invention are as follows:
[0064] [1].Wang D, Wei G, Ma J et al. Identification of the prognostic value offerroptosis-related gene signature in breast cancer patients, BMC Cancer 2021; 21:645.
[0065] [2]. Li Y, Zhao X, Liu Q et al. Bioinformatics reveal macrophages markergenes signature in breast cancer to predict prognosis, Ann Med 2021;53:1019-1031.
[0066] [3]. Zhang Z, Fang T, Lv YA novel lactate metabolism-related signatures predicts prognosis and tumor immune microenvironment of breast cancer, FrontGenet 2022;13:934830.
[0067] like Figure 4 As shown, the survival curves indicate that the low-risk group has a higher survival rate. Figure 4 (AD). Apart from the features described by Zhang et al. (AUC = 0.502, 0.522, 0.568), other features, based on the area under the receiver operating characteristic curve, showed good potential in predicting 1-year, 3-year, and 5-year survival rates for breast cancer. Figure 4 (EH). The BCSCRS (AUC = 0.694) and nomogram (AUC = 0.758) established in this study have higher accuracy than other features (EH). Figure 4 The clinically optimized plots are not included in the feature comparisons and are only used for auxiliary validation. C-index, RMS, and DCA analyses further confirm the superior accuracy of BCSCRS in predicting breast cancer survival. Figure 4 :JL).
[0068] Example 4: Application of prognostic models based on breast tumor stem cell-related genes - BCSCRS-based cancer immune landscape analysis.
[0069] Given the significant correlation between BCSC core genes and immune activity observed in our research and analysis, we performed GSVA and GSEA analyses to further explore this correlation. We found that in the high-risk group, signaling pathways were significantly enriched in biological processes such as steroid biosynthesis, fructose and mannose metabolism, protein export, proteasome and citrate cycle, and TCA cycle. Conversely, the low-risk group was characterized by primary immunodeficiency and T cell receptor signaling pathways (…). Figure 5 A) This suggests a potential link between the low-risk group and immunity.
[0070] To further explore this relationship, we investigated features associated with the immune landscape, including TME and immune infiltration. Figure 5 The results showed that the estimated scores, immune scores, and matrix scores (these scores are expressed directly in English) of the low-risk group were significantly higher than those of the high-risk group, while the tumor purity results were the opposite. Figure 5These findings indicate that the levels of stromal cells and immune cells in the TME were higher in the low-risk group than in the high-risk group. Subsequent analysis using ssGSEA showed that, except for macrophages, the expression levels of immune cells in the TME were higher in the low-risk group. Figure 5 (B) The CIBERSORT algorithm was used to further analyze the infiltration abundance of 22 specific immune cells in the high-risk and low-risk groups. It was observed that B cells, plasma cells, memory-activated CD4 T cells, CD8 T cells, and γ-δ T cells were more abundant in the low-risk group, while M0 and M2 macrophages were more abundant in the high-risk group. Figure 5 :D). Furthermore, the infiltration levels of B cells, plasma cells, memory-activated CD4 T cells, CD8 T cells, and γ-δ T cells were negatively correlated with the risk score. Figure 5 :E).
[0071] These results indicate a close relationship between BCSCRS and immune cells, with lower risk scores suggesting higher expression of stromal cells and immune cells in the TME.
[0072] Example 5: Application of prognostic models based on breast tumor stem cell-related genes - BCSCRS-based assessment of immunotherapy response.
[0073] To further investigate the relationship between BCSCRS and immunotherapy response, we evaluated the following indicators:
[0074] First, we analyzed the expression of immune checkpoint molecules, and the results showed that the expression of 27 immune checkpoints was significantly higher in the low-risk group, suggesting that these patients may have a stronger response to immune checkpoint inhibitors. Figure 6 :A).
[0075] We also used IPS scores for PD1 and CTLA4 as quantitative indicators to further evaluate the efficacy of immune checkpoint inhibitors. The results showed that the low-risk group had significantly higher IPS-CTLA4, IPS-PD1, and IPS-PD1-CTLA4 scores, indicating that these patients would have better efficacy when treated with PD-1 and CTLA4 inhibitors. Figure 6 :B).
[0076] Furthermore, since existing literature reports that breast cancer stem cells (BCSCs) are involved in the drug resistance process of cancer, we also analyzed the drug sensitivity of commonly used chemotherapy drugs for breast cancer in different risk groups. The results showed that the low-risk group was more sensitive to chemotherapy drugs such as cisplatin, doxorubicin, gemcitabine, methotrexate, paclitaxel, and vinorelbine. This suggests that patients in the low-risk group have better efficacy when receiving chemotherapy and are less likely to develop drug resistance.
[0077] Therefore, our findings suggest that the low-risk group may respond better to both immunotherapy and chemotherapy, which has significant clinical implications and demonstrates that the prognostic model based on breast cancer stem cell-related genes constructed in this invention can be effectively applied to the clinical practice of breast cancer treatment and prognosis.
[0078] Breast cancer stem cells (BCSCs) may be the origin of breast cancer heterogeneity and are thought to be involved in regulating responses to breast cancer immunotherapy. Therefore, understanding the prognostic value and immune reactivity of BCSCs is crucial for identifying patients who may benefit from immunotherapy. We first obtained transcriptomic and clinical data on breast cancer from TCGA and GEO, then screened for survival-related breast cancer stem cell genes. We then constructed a breast cancer stem cell prognostic model using LASSO regression and multivariate Cox regression, and used ROC, C-index, and DCA indicators to compare the accuracy of this model with other existing prognostic models. Building on this study, to better apply the prognostic model clinically, we constructed a nomogram based on risk scores and clinical variables to predict overall patient survival, improving the accuracy of the prognostic model and enabling better clinical application.
[0079] Results: Our study constructed a 10-gene prognostic model related to breast cancer stem cells. The prognostic model constructed in this study showed good predictive accuracy in both the training and validation cohorts (AUC values of 0.7 at 1, 3, and 5 years). In addition, this model had better predictive accuracy compared with other models. The nomogram that included clinical variables showed better predictive precision and accuracy (AUC = 0.758, C-index = 0.744).
[0080] This invention is based on the prognostic characteristic genes of breast cancer stem cells and the corresponding prediction model (i.e., the sum of the products of the expression levels and coefficients of 10 prognostic characteristic genes of breast cancer stem cells, combined with the patient's clinical characteristics) to determine the prognosis of patients. It has the advantages of efficient and accurate prediction of the prognosis of breast cancer patients, providing effective guidance for clinicians to make treatment decisions for breast cancer patients, reducing the occurrence of ineffective treatment, and thus reducing the treatment costs and discomfort of patients.
[0081] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Obviously, other related modifications can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.
Claims
1. The application of a breast cancer prognostic gene in constructing a breast cancer prognostic model, characterized in that: The prognostic characteristic genes are breast cancer stem cell-related genes, namely BRD4, RPS24, SERPINA3, SKP1, NTRK3, CD79A, JAK1, NT5E, NDRG1, and CD24. The prognostic model includes the following: Breast Cancer Stem Cell-Related Gene Risk Score BCSCRS = (−0.53045×BRD4) + (−0.26259×RPS24) + (−0.31334×SERPINA3) + (0.434039×SKP1) + (−0.53742×NTRK3) + (−0.23344×CD79A) + (−0.40628×JAK1) + (0.192005×NT5E) + (0.152866×NDRG1) + (0.194872×CD24); When BCSCRS < 0.985569, it is judged as a low-risk state, and when BCSCRS > 0.985569, it is judged as a high-risk state.
2. The application according to claim 1, characterized in that: A nomogram was constructed based on risk scores and clinical variables related to breast cancer stem cell-related genes to predict overall patient survival and prognosis.
3. The application according to claim 2, characterized in that: The clinical variables included gender, TNM stage, and age.
4. The application according to claim 3, characterized in that: By adding the risk scores of breast cancer stem cell-related genes and the scores of various clinical variables, an overall score is calculated for each patient, thereby predicting the survival probability of breast cancer patients at 1 year, 3 years, and 5 years. Patients with lower total scores have a higher survival probability.
5. The application according to claim 4, characterized in that, The method for constructing a prognostic model includes the following steps: (1) Data acquisition: Transcriptome and clinical data of breast cancer from TCGA and GEO were acquired. Normal tissue samples and samples from the same patient were excluded. The createDataPartition function in the R package caret was used to randomly divide TCGA patients into training cohort and internal validation cohort in a 7:3 ratio. Patients in GSE20685 were used as external validation cohort. (2) Breast cancer stem cell gene screening: Breast cancer stem cell related genes were collected from the GeneCards database, and genes with a correlation score greater than 30 were retained. Univariate Cox analysis was performed to screen breast cancer stem cell genes related to breast cancer prognosis. (3) Construction of the prognostic model: In the TCGA training cohort, the Least Absolute Shrinkage Selection Operator (LASSO) Cox regression method was used to further screen prognostic characteristic genes of breast cancer, and a breast cancer prognostic model was constructed by multivariate Cox regression analysis; the risk score of the prognostic model was BCSCRS=(−0.53045×BRD4) + (−0.26259×RPS24) + (−0.31334×SERPINA3) +(0.434039×SKP1) + (−0.53742×NTRK3) + (−0.23344×CD79A)+ (−0.40628×JAK1) + (0.192005×NT5E) + (0.152866×NDRG1) + (0.194872×CD24); Based on the gene expression levels and corresponding regression coefficients in the Cox regression model, the risk score for each sample is calculated using the predict function in the survival package; All patients are divided into high-risk and low-risk groups based on the median BCSCRS, where the median is 0.985569. That is, when BCSCRS < 0.985569, the patient is assigned to the low-risk group, and when BCSCRS > 0.985569, the patient is assigned to the high-risk group. (4) Validation of the prognostic model: The TCGA test queue and the GSE20685 queue were used as the internal validation queue and the external validation queue, respectively, to validate the accuracy of the prognostic model. ROC, C-index and DCA indicators were used to evaluate the accuracy of the prognostic model. (5) Construction of nomograms: The rms package was used to construct nomograms and calibration curves including patient age, sex, TNM stage and breast cancer stem cell-related gene risk score BCSCRS; the timeROC and ggDCA packages were used to perform receiver operating characteristic (ROC) and decision curve analysis (DCA) to compare the predictive accuracy of nomograms with other prognostic factors; the nomograms constructed with BCSCRS and clinical variables were used as a quantitative method to predict the survival rate of breast cancer patients.
6. The application of the detection reagent in the preparation of a kit for assessing the prognostic survival of breast cancer, characterized in that, The detection reagent consists of reagents for detecting the expression levels of the following 10 genes: BRD4, RPS24, SERPINA3, SKP1, NTRK3, CD79A, JAK1, NT5E, NDRG1, and CD24. This detection reagent is the sole key component of the kit for assessing the prognostic survival function of breast cancer. The kit also includes an instruction manual, which states the following: BCSCRS = (−0.53045×BRD4) + (−0.26259×RPS24) + (−0.31334×SERPINA3) + (0.434039×SKP1) + (−0.53742×NTRK3) + (−0.23344×CD79A) + (−0.40628×JAK1) + (0.192005×NT5E) + (0.152866×NDRG1) + (0.194872×CD24); When BCSCRS < 0.985569, it is judged as a low-risk state, and when BCSCRS > 0.985569, it is judged as a high-risk state. Patients with low-risk state have a better prognosis than patients with high-risk state.
7. The application according to claim 6, characterized in that: The specimens tested using the kit are fresh tissue tumor specimens.
Citation Information
Patent Citations
Genetic characteristics of residual risk after endocrine therapy in early-stage breast cancer
CN109072481B
Breast cancer prognosis evaluation method and system based on autophagy related lncRNA model
CN113593648A
Application of immune related genes in breast cancer prognosis kit and system
CN113862363A
Kit for detecting in situ hybridization of romodomain-4 (BRD-4) genes, and detection method and application thereof
CN101993943A
Methods and kit for the prognosis of breast cancer
CN102443627A