Gene marker related to breast cancer prognosis and detection kit
By constructing a breast cancer prognosis model related to ferroptosis genes and using gene expression profile data to identify and detect cell subtypes, the problems of low accuracy and high cost of existing models were solved, and effective prediction and diagnosis of breast cancer prognosis were achieved.
Patent Information
- Application Number
- CN202510706971.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-06-26
- Filing Date
- 2025-05-29
- Publication Date
- 2025-09-23
AI Technical Summary
The existing breast cancer prognosis model has low accuracy and precision, and has high industrialization costs. It is not suitable for clinical promotion and cannot effectively predict the prognosis of triple-negative breast cancer.
A breast cancer prognosis model based on ferroptosis genes was constructed. Macrophage and T cell subtypes were identified through gene expression profile data. Cell timing analysis and communication analysis were performed. 23 genes closely related to breast cancer prognosis were screened as markers. A breast cancer prognosis model was established and detected using fluorescent probes or hybridization probes.
It improves the accuracy and reliability of breast cancer prognosis prediction, provides a new method for the diagnosis and survival prognosis of breast cancer patients, and can effectively judge the high-risk or low-risk prognosis of patients.
Smart Images

Figure BDA0005426107770000121 
Figure BDA0005426107770000131 
Figure BDA0005426107770000171
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of cancer treatment and prognosis evaluation, and particularly to gene markers associated with breast cancer prognosis. Background Art
[0002] Breast cancer is a common malignancy in women and a major threat to their health. Triple-negative breast cancer (TNBC) is a subtype of breast cancer that is negative for both estrogen receptor (ER) and progesterone receptor (PR) and lacks human epidermal growth factor receptor 2 (HER2) amplification. It is characterized by high invasiveness, high rates of recurrence and metastasis, and poor prognosis.
[0003] Therefore, searching for biomarkers for breast cancer metastasis and prognosis prediction has important theoretical guidance and clinical significance for discovering new therapeutic targets, improving efficacy, and improving patient prognosis.
[0004] Existing reports have used immune-related genes, autophagy-related genes, apoptosis-related genes, lactate metabolism genes, macrophage characteristic genes, copper-dependence-related genes, and breast cancer single-cell transcriptome sequencing analysis to construct prediction models. Some prognostic models involve the detection of nearly 100 genes, and the industrialization cost is too high, making them unsuitable for clinical promotion. Existing prognostic models are basically constructed for specific signaling pathways and targets. In actual application, problems such as low prediction accuracy and low precision often occur, which may be caused by the heterogeneity of breast cancer. Therefore, there is an urgent need for an accurate and reliable breast cancer prognostic model for clinical application. Summary of the Invention
[0005] The purpose of this disclosure is to provide a breast cancer prognosis model that uses a set of ferroptosis genes as genetic markers to predict breast cancer prognosis. Using gene expression profile data from breast cancer patients, this disclosure identifies macrophage and T cell subtypes based on ferroptosis genes, performs cell timing analysis, communication analysis, establishes a gene regulatory network, and performs differential detection of transcription factor activity and functional enrichment analysis. The model identifies 23 genes closely associated with breast cancer prognosis as markers, and constructs a breast cancer prognosis model.
[0006] According to a first aspect of the present disclosure, gene markers associated with breast cancer prognosis are provided, wherein the gene markers include one or more of the following genes: TMEM160, EWSR1, BCAT2, PNKP, MLEC, SAP30BP, TB L1XR1, STAG1, NR4A1, TNFRSF9, PSD3, BAD, MIIP, HIST3H2A, CDC25B, TCEB 1, HMGCS1, SPC25, TKT, PTTG1, ADA, AK1 and AIMP2.
[0007] According to a second aspect of the present disclosure, a detection kit for evaluating the prognosis of breast cancer is provided, comprising: a reagent for detecting the gene marker described in the first aspect.
[0008] In some embodiments, the reagents for detecting the gene markers include primer pairs or probes targeting the gene markers;
[0009] In some embodiments, the probe is a fluorescent probe or a hybridization probe.
[0010] In some embodiments, the fluorescent probe is selected from one or more of the following: an MGB probe, an LNA probe, a molecular beacon probe, and a TaqMAN probe.
[0011] As a preferred embodiment of the present disclosure, the kit is used to detect gene markers, and the reagents include a reverse transcription primer pair for the gene marker, and preferably also include a specific fluorescent probe for the gene marker. By real-time monitoring of the fluorescent signal in the PCR system, the expression level of the gene marker in the patient sample is quantitatively detected.
[0012] According to a third aspect of the present disclosure, a method for evaluating the prognosis of breast cancer is provided, the method comprising: using a computational model to calculate a risk score (RiskScore) of a breast cancer patient to be tested, wherein the risk score (RiskScore) = TME M160×(-0.4689)+EWSR1×(-0.4237)+BCAT2×(-0.2346)+PNKP×(-0.1592)+MLE C×(-0.1567)+SAP30BP×(-0.1469)+TBL1XR1×(-0.0894)+STAG1×(-0.0695)+NR 4A1×(-0.0630)+TNFRSF9+(-0.0503)+PSD3×(-0.0153)+BAD×(-0.0094)+MIIP×0.0104+HIST3H2A×0.0294+CDC25B×0.0528+TCEB1×0.0670+HMGCS1×0.0988+SPC25×0.1317+TKT×0.2234+PTTG1×0.2960+ADA×0.3097+AK1×0.4003+AIMP2×0.4081; wherein, each gene in the calculation model represents the expression level of the gene in the sample; when the risk score of the breast cancer patient to be tested is higher than the threshold, the prognosis result is judged to be high risk; if it is lower than the threshold, the prognosis result is judged to be low risk.
[0013] In some embodiments, the threshold value can be obtained by collecting breast cancer patient data, calculating the risk score, and aggregating and summarizing the risk scores of multiple patients to obtain a median value as the threshold value. In some embodiments, the breast cancer patient data can be obtained from a database known in the art or from clinical samples.
[0014] In some embodiments, the threshold value = 3.0-3.5.
[0015] In some embodiments, the threshold value = 3.0, 3.1, 3.2, 3.3, 3.4, or 3.5.
[0016] According to a fourth aspect of the present disclosure, a system for evaluating the prognosis of breast cancer is provided, the system comprising:
[0017] (1) a data input module, used to input the expression level of the gene marker;
[0018] (2) Model calculation module, based on the data input in step (1), the calculation model is used, risk score (RiskScore) = TMEM160×(-0.4689)+EWSR1×(-0.4237)+BCAT2×(-0.2346)+PNKP×(-0.1592)+MLEC×(-0.1567)+SAP30BP×(-0.1469)+TBL1XR1×(-0.0894)+STAG1×(-0.0695)+NR4A1×(-0.0630)+TNFRSF9+(-0.0 503)+PSD3×(-0.0153)+BAD×(-0.0094)+MIIP×0.0104+HIST3H2A×0.0294+CDC25B×0.0528+TCEB1×0.0670+HMGCS1×0.0988+SPC25×0.1317+TKT×0.2234+PTTG1×0.2960+ADA×0.3097+AK1×0.4003+AIMP2×0.4081, and calculation was performed; wherein, each gene in the calculation model represents the expression level of the gene in the sample;
[0019] (3) Result output module: when the patient's risk score is higher than the threshold, the prognosis is judged to be high risk; if it is lower than the threshold, the prognosis is judged to be low risk.
[0020] In some embodiments, the threshold value can be obtained by the following steps: collecting breast cancer patient data, calculating the risk score, and summarizing and counting the risk scores of the multiple patients to obtain the median as the threshold value.
[0021] In some embodiments, the breast cancer patient data can be from a database known in the art, or from clinical samples.
[0022] In some embodiments, the threshold value = 3.0-3.5.
[0023] In some embodiments, the threshold value = 3.0, 3.1, 3.2, 3.3, 3.4, or 3.5.
[0024] According to a fifth aspect of the present disclosure, provided is the use of the gene marker, the kit, the evaluation method or the evaluation system in the diagnosis of breast cancer.
[0025] According to a sixth aspect of the present disclosure, there is provided a use of the gene markers of the first aspect in preparing a kit for evaluating the likelihood of occurrence of breast cancer, wherein the breast cancer includes triple-negative breast cancer or non-triple-negative breast cancer.
[0026] According to a seventh aspect of the present disclosure, there is provided use of the gene markers of the first aspect in preparing a kit for prognosis evaluation of breast cancer, wherein the breast cancer includes triple-negative breast cancer or non-triple-negative breast cancer.
[0027] This disclosure is based on triple-negative breast cancer single-cell data, exploring the effects of ferroptosis-related genes on immune cells in triple-negative breast cancer, and further exploring the differences between different cell subtypes through ferroptosis-mediated macrophage and T cell subtype identification, and finally explaining that ferroptosis-related genes do have some effect on the immune cell activity of triple-negative breast cancer. Afterwards, a prognostic model was established based on ferroptosis-mediated subtype-specific genes through univariate COX regression analysis, Lasso regression analysis, ROC curve analysis, and independent prognostic factor analysis, and ferroptosis genes significantly correlated with breast cancer prognosis were screened out as gene markers. Combined with bulk data, it was shown that the characteristic genes of different immune cell subtypes have a significant effect on prognosis, further proving that cancer disrupting the immune microenvironment will affect the patient's prognosis. The gene markers provided by this disclosure can well reflect the prognosis of breast cancer patients, and provide a new method for improving the diagnosis and survival prognosis of clinical breast cancer patients. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 The flowchart of the construction of the triple-heterogeneous breast cancer prognostic model is shown.
[0029] Figure 2 Cell types and ferroptosis-related gene maps are shown. Figure 2 A shows the tSNE of major cells; Figure 2 B shows the tSNE of minor cells; Figure 2 C shows tSNE of 12 integrated cells; Figure 2 D shows the average expression of ferroptosis-related genes in different clinical groups; Figure 2 E shows the average expression of ferroptosis-related genes in 12 cell types; Figure 2 F shows cell communication analysis of all cells.
[0030] Figure 3 Shown are the distributions of ferroptosis-mediated macrophage and T cell subtypes in different samples.
[0031] Figure 4 Shown is the identification of a macrophage subset that mediates ferroptosis in triple-negative breast cancer. Figure 4 A shows pseudo-time series analysis of macrophages; Figure 4 B shows a scatter plot of temporal analysis of ferroptosis-mediated macrophage subtype grouping; Figure 4 C shows tSNE of ferroptosis-mediated macrophage subtypes; Figure 4D shows the average expression levels of ferroptosis-related genes in macrophage subtypes (Top 50); Figure 4 EF shows tSNE of macrophage subtype enrichment scores; Figure 4 GH shows tSNE of macrophage subtype marker expression; Figure 4 I shows the correlation analysis of macrophage subtype enrichment scores; Figure 4 J shows analysis of communication between macrophage subtypes and tumor epithelium.
[0032] Figure 5 The figure shows the differences in activity of the top 20 transcription factors between macrophage subtypes. The horizontal axis represents the transcription factor, and the _extended suffix indicates the gene regulatory network composed of the transcription factor and all target genes. The number of target genes is in parentheses.
[0033] Figure 6 Identification of ferroptosis-mediated T cell subsets in triple-negative breast cancer is shown. Figure 6 A shows pseudo-temporal analysis of T cells; Figure 6 B shows a scatter plot of temporal analysis of the original annotated T cell subtypes; Figure 6 C shows a scatter plot of temporal analysis of ferroptosis-mediated T cell subtype grouping; Figure 6 D shows the original annotation T cell subtype identification tSNE; Figure 6 E shows the proportion of the originally annotated T cell subtypes in the ferroptosis-mediated T cell subtype grouping; Figure 6 F shows the average expression levels of ferroptosis-related genes in ferroptosis-mediated T cell subtypes (Top 50); Figure 6 G shows the average expression levels of characteristic genes in the T cell functional signature model in 15 T cell subtypes; Figure 6 H shows analysis of T cell subtypes communicating with tumor epithelium.
[0034] Figure 7 Shown are differences in Top20 transcription factor activity between T cell subtypes.
[0035] Figure 8 Hallmark pathway enrichment scores for T cell subtypes are shown.
[0036] Figure 9 The expression of ferroptosis-mediated cell subtype-specific markers in bulk data is shown.
[0037] Figure 10 Prognostic differences based on ferroptosis-mediated TME subset scores are shown. Figure 10 A shows the results based on OS; Figure 10 B shows the results based on RFS.
[0038] Figure 11 Figure 3: Ferroptosis-mediated TME subset scores differentially respond to immunotherapy efficacy across cohorts.
[0039] Figure 12 Shows the batch univariate cox regression Top20 forest plot.
[0040] Figure 13 The lasso regression results are shown. Figure 13 A shows the corresponding changes of lambda and variable coefficients; Figure 13 B shows the cross-validation to obtain lambda.min; Figure 13 C shows the regression coefficients corresponding to the filtered variables.
[0041] Figure 14 Validation of the risk model on GSE25066 data is shown. Figure 14 A shows the KM curve verification of the lasso regression model; Figure 14 B shows the ROC curve validation; Figure 14 C shows the risk score curve of all samples; Figure 14 D shows the scatter plot of survival time of all samples; Figure 14 E shows the expression heatmap of genes within the model.
[0042] Figure 15 Validation of the risk model on the GSE86166 data is shown. Figure 15 A shows the KM curve verification of the lasso regression model; Figure 15 B shows the ROC curve validation; Figure 15 C shows the risk score curve of all samples; Figure 15 D shows the scatter plot of survival time of all samples; Figure 15 E shows the expression heatmap of genes within the model.
[0043] Figure 16 Shows the TCGA-BRCA model validation results, Figure 16 A shows the KM curve display, Figure 16 B shows the ROC curve display based on time.
[0044] Figure 17 The GSE25066 data were used to verify whether high-risk and low-risk groups were independent prognostic factors.
[0045] Figure 18 The GSE86166 data were used to verify whether high-risk and low-risk groups were independent prognostic factors.
[0046] Figure 19 Chemotherapy resistance among different signature model groups is shown. Figure 19A shows the presentation of drugs that are negatively correlated with risk scores; Figure 19 B shows drug presentation that is positively correlated with risk score.
[0047] Figure 20 The ROC curves for judging the breast cancer risk of 22 clinical patients with triple-negative breast cancer and 8 non-triple-negative breast cancer patients using the 23 gene markers shown in Table 6 and the risk scoring model obtained in Example 6 are shown. DETAILED DESCRIPTION
[0048] To make the objectives, technical solutions, and advantages of the present disclosure more clearly understood, the present disclosure is further described in detail below with reference to the following embodiments. The specific embodiments described herein are intended only to illustrate the present disclosure and are not intended to limit the present disclosure in any way. In addition, descriptions of well-known structures and technologies are omitted in the following description to avoid unnecessary confusion about the concepts of the present disclosure. Such structures and technologies are also described in many publications.
[0049] definition
[0050] Unless otherwise defined, all technical and scientific terms used in this disclosure have the same meaning as commonly used in the field to which this disclosure belongs. For the purpose of interpreting this specification, the following definitions will apply, and where appropriate, terms used in the singular will also include the plural form, and vice versa.
[0051] As used herein, the articles "a," "an," and "an" include plural referents unless the context clearly dictates otherwise.
[0052] The expression "about" as used in this disclosure is understood by one of ordinary skill in the art and varies within certain ranges depending on the context in which it is used. If the use of the term is not understood by one of ordinary skill in the art based on the context in which it is used, "about" will mean up to plus or minus 10% of the specified value.
[0053] As used herein, the term "triple-negative breast cancer" refers to a type of breast cancer whose cancer cells do not have estrogen or progesterone receptors and do not make much of a protein called human epidermal growth factor receptor 2 (HER2). Test results for estrogen (ER) receptors, progesterone (PR) receptors, and HER2 / neu (HER2) receptors are negative.
[0054] The term "ferroptosis" used in this disclosure is a new form of programmed cell death, discovered by Stockwell in 2012, which is characterized by iron-dependent cell membrane lipid peroxidation. The main mechanism of ferroptosis is that under the action of divalent iron or esteroxygenase, unsaturated fatty acids highly expressed on the cell membrane are catalyzed to cause lipid peroxidation, thereby inducing cell death; in addition, it is also manifested as a decrease in the regulatory core enzyme GPX4 of the antioxidant system (glutathione system). Cysteine is an important negative regulatory factor in the ferroptosis pathway. Studies have shown that lymphocytes cannot synthesize cysteine. This important discovery may become a breakthrough point in the treatment of lymphoma. Studies have shown that in rat and mouse animal models, the ferroptosis inducer imidazole-keton e-erastin (IKE) can significantly reduce the growth of diffuse large B-cell lymphoma (DLBCL). Ferroptosis has been found to play a positive role in radiotherapy, chemotherapy and immunotherapy. Therefore, ferroptosis activation may be a potential strategy to overcome the mechanism of cancer treatment resistance.
[0055] The term "prognosis" as used in this disclosure refers to a prediction of the likely consequences of cancer, including the expectation of cancer recovery. As used herein, the terms prognostic information and predictive information interchangeably represent information that can be used to predict any aspect of the course of a disease or condition, whether or not treatment is being performed. Such information may include, but is not limited to, the average life expectancy of a patient, the likelihood that a patient will survive a given time (e.g., 6 months, 1 year, 5 years, etc.), the likelihood that the patient's disease will be cured, the likelihood that the patient's disease will respond to a specific treatment (wherein response can be defined in any different ways). Both prognostic and predictive information are included in the broad category of diagnostic information.
[0056] In some embodiments, "prognosis" refers to a prediction of the likely course and outcome of a clinical condition or disease. A patient's prognosis is typically determined by evaluating factors or symptoms that indicate a favorable or unfavorable disease course or outcome. The term "prognosis" does not imply a 100% accurate prediction of the course or outcome of a condition. Rather, as will be understood by those skilled in the art, the term "prognosis" refers to an increased probability that a particular course or outcome will occur; that is, the course or outcome is more likely to occur in a patient exhibiting a given condition than in an individual who does not exhibit the condition. Prognosis can be expressed as the amount of time a patient is expected to survive. Alternatively, prognosis can refer to the likelihood that a disease will remit or the amount of time a disease is expected to remain in remission. Prognosis can be expressed in various ways; for example, prognosis can be expressed as the percentage probability that a patient will be alive after one year, five years, ten years, etc. Alternatively, prognosis can be expressed as the average number of months a patient can be expected to survive due to the condition or disease. A patient's prognosis can be considered a relative expression, as many factors influence the ultimate outcome. For example, for a patient with a particular condition, the prognosis may be appropriately expressed as the likelihood that the condition is treatable or curable, or the likelihood that the disease will go into remission, whereas the prognosis for a patient with a more severe condition may be more appropriately expressed as the likelihood of surviving for a specified period of time.
[0057] As used herein, the terms "expression amount" or "expression level" are generally used interchangeably and generally refer to the amount of a polynucleotide, mRNA, or amino acid product or protein in a biological sample. "Expression" generally refers to the process by which the information encoded by a gene is converted into structures that exist and function in a cell.
[0058] The term "probe" as used in this disclosure refers to a target sequence labeled with a capture label or a detection label. The probe sequence is used to hybridize with the target sequence. The appropriate length of the probe depends on the desired use of the probe and is generally in the range of 80 to 200 nucleotides (nt). Preferably, the appropriate length of the primer includes 100 to 150 nucleotides. In this disclosure, the probe can be one of a TaqMan probe, an MGB probe, or a hybridization probe.
[0059] In this disclosure, the term "hybridize" or "specifically hybridize" means that a molecule can bind, duplex or hybridize only to a specific polynucleotide sequence under stringent conditions when that sequence is present in a complex mixture (e.g., total cellular) DNA or RNA.
[0060] In the present disclosure, the term "complementary" refers to the concept of sequence complementarity between regions of two polynucleotide chains or between two regions of the same polynucleotide chain. It is known that the adenine base of the first region of a polynucleotide can form specific hydrogen bonds ("base pairs") with the base of the second region of the polynucleotide that is antiparallel to the first region (if the base is thymine or uracil). Similarly, it is known that the cytosine base of the first polynucleotide chain can base pair with the base of the second polynucleotide chain that is antiparallel to the first region (if the base is guanine). If when the first region of a polynucleotide is arranged in an antiparallel manner with the second region of the same or another different polynucleotide, at least one nucleotide in the first region can base pair with the base of the second region, then the two regions are complementary. Therefore, two complementary polynucleotides do not need to base pair at every nucleotide position. "Complementary" means that the first polynucleotide is 100% or "completely" complementary to the second polynucleotide, and therefore forms base pairing at every nucleotide site. "Complementary" also refers to a first polynucleotide that is not 100% complementary (e.g., 90%, or 80% or 70%, or 60%, or 50% complementary) and comprises mismatched nucleotides at one or more nucleotide positions. In one embodiment, two complementary polynucleotides are capable of hybridizing to each other under high stringency hybridization conditions.
[0061] In the present disclosure, the term "stringent hybridization conditions" refers to conditions under which a probe hybridizes to its target subsequence, typically in a complex mixture of nucleic acids, but does not hybridize to other sequences. Stringent conditions are sequence-dependent and will be different in different situations. Longer sequences specifically hybridize at higher temperatures. Typically, the stringent conditions are selected to be about 5-10°C lower than the thermal melting point (Tm) of the specific sequence at a defined ionic strength pH. Tm is the temperature (under specified ionic strength, pH, and nucleic acid concentration) at which 50% of the probes complementary to the target hybridize to the target sequence at equilibrium (because the target sequence is present in excess, at Tm, 50% of the probes are occupied at equilibrium). Stringent conditions can also be achieved by adding destabilizing agents (e.g., formamide). For selective or specific hybridization, the positive signal is at least twice the background, preferably 10 times the background hybridization. Exemplary stringent hybridization conditions may be as follows: 50% formamide, 5×SSC, and 1% SDS, incubation at 42° C., or 5×SSC, 1% SDS, incubation at 65° C., with a wash in 0.2×SSC and 0.1% SDS at 65° C.
[0062] The term "clustering" as used in this disclosure refers to grouping data into corresponding categories based on the similarities between the data. The same category has a high degree of similarity, and the difference between different categories is the greatest.
[0063] The term "NMF clustering" as used in this disclosure is a type of cluster analysis, which refers to the process of grouping a collection of transcriptome RNA sequencing data from a sample into multiple categories consisting of similar gene expression profiles. Conventional cluster analysis methods include hierarchical clustering, K-means clustering, and second-order clustering. The preferred cluster analysis method in this disclosure is NMF cluster analysis. The term "NMF" refers to non-negative matrix factorization.
[0064] As used herein, the term "cell subpopulation" refers to any group of cells in a biological sample whose presence is characterized by one or more characteristics, such as gene expression at the RNA level, protein expression, genomic mutations, biomarkers, etc. A cell subpopulation can be, for example, a cell type or a cell subtype.
[0065] The term "risk scores" used in this disclosure is also called Clinical Predicti on Models (CPMs), Clinical prediction rules, Risk prediction models or Predictive models, which refers to the use of mathematical formulas to estimate the probability that a specific individual currently suffers from a certain disease or will have a certain outcome in the future. Clinical prediction models include diagnostic models and prognostic models. Diagnostic models focus on the probability of diagnosing a certain disease based on the clinical symptoms and characteristics of the research subjects, and are more common in cross-sectional studies; prognostic models focus on the probability of a certain outcome (onset, recurrence, death, disability and complications) occurring in the future under the current health status (healthy or sick), and are more common in cohort studies. In a narrow sense, prognostic models specifically refer to the probability of disease recurrence, death, disability or complications in the future after being diagnosed with a disease. In clinical prediction models, common outcomes are illness, onset, recurrence of disease, death, disability and complications. Common predictors include sociodemographic characteristics (such as age and gender), medical history, medication history, physical examination results, imaging, electrophysiological tests, blood and urine samples, pathological examinations, disease stage and characteristics, and omics (genomics, proteomics, metabolomics, transcriptomics, and pharmacogenomics) indicators. Predictors are also referred to in other literature as prognostic factors, determinants, or, more statistically, covariates and independent variables. The distinction between diagnostic and prognostic models is based solely on the application scenario; from a statistical modeling perspective, there is no essential difference between the two. For example, model outcomes are often binary variables, indicating whether or not a disease occurs. Study effect indicators are the absolute risk of the outcome, that is, the probability of occurrence, rather than relative effect indicators such as relative risk (RR), odds ratio (OR), or hazard ratio (HR). Technically, both models require the selection of predictors, the development of modeling strategies, and the evaluation of model performance.
[0066] The term "KM curve (Kaplan-Meier survival curve)" used in this disclosure is a survival analysis graph used to show the probability of a certain event (such as survival, recurrence-free or disease progression) occurring in a specific population over time. The population is divided into different groups or subgroups, and a comparison of the survival of multiple groups can be displayed simultaneously in the same graph. The horizontal axis of the KM curve represents time, and the vertical axis represents the survival rate or recurrence-free rate. Each point or step on the curve represents the survival rate of those who did not experience an event during that time period. When a patient experiences an event, a decrease in the point or step indicates a decrease in survival rate. The KM curve can intuitively display the time and extent of the effect of drug treatment and is commonly used in drug clinical trials and later evaluations.
[0067] The term "HR (Hazard Ratio)" used in this disclosure is an indicator used to compare the difference in risk of certain events between two groups of people. An HR value greater than 1 indicates that one group has a higher risk of occurrence, and a value less than 1 indicates a lower risk. An HR value equal to 1 indicates that the risk of an event occurring in both groups is equal. In the KM curve, the HR value can be calculated by taking the logarithm of the survival rate ratio at two time points. The HR value is the survival rate ratio, and the larger the HR value, the better the drug treatment effect. The reliability of the HR value increases with the increase of sample size, but decreases with the extension of time.
[0068] The term "ROC curve" or "receiver operating characteristic curve" as used in the present disclosure refers to a graphical curve that shows the performance of a binary classifier system as its discrimination threshold changes. The curve is created by plotting the true positive rate against the false positive rate at various threshold settings. The true positive rate is also called sensitivity. The false positive rate is calculated as 1-specificity. Therefore, the ROC curve is a graphical display of the true positive rate against the false positive rate (sensitivity vs (1-specificity)) within a range of cutoff values and a way to select the best cutoff value for clinical use. Accuracy is expressed as the area under the ROC curve (AUC), which provides a useful parameter for comparing test performance. An AUC close to 1 indicates that the test is highly sensitive and has a high specificity, while an AUC close to 0.5 indicates that the test is neither sensitive nor specific.
[0069] As used herein, the term "IC50" refers to a concentration that produces a measurable phenotype or response, such as a concentration at which the growth of a cell, such as a tumor cell, is inhibited by 50%. IC50 values can be estimated from a suitable dose-response curve, for example, by visual fitting or using suitable curve fitting or statistical software. More precisely, IC50 values can be determined using nonlinear regression analysis.
[0070] The term "gene" used in the present disclosure may be a natural (such as genomic) gene including transcriptional and / or translational regulatory sequences, and / or coding regions, and / or non-translated sequences (such as introns, 5'- and 3'-untranslated sequences). The coding region of a gene may be a nucleotide sequence encoding an amino acid sequence or a functional RNA (such as tRNA, rRNA, catalytic RNA, siRNA, miRNA or antisense RNA). The term "gene" has the meaning understood in the art. However, one of ordinary skill in the art will appreciate that the term "gene" has many meanings in the art, some of which include gene regulatory sequences (such as promoters, enhancers, etc.) and / or intron sequences, while others are limited to coding sequences. It should also be understood that the definition of "gene" includes nucleic acids that do not encode proteins but encode functional RNA molecules (such as tRNA). For the sake of clarity, we note that the term "gene" used in this application generally refers to a part of a nucleic acid that encodes a protein; the term also optionally includes a regulatory sequence. This definition is not intended to exclude that the term "gene" may be used for non-protein-coding expression units, but rather to clarify that, in most cases, the term is used herein to refer to nucleic acids that encode proteins.
[0071] The term "TMEM160 gene," as used herein, encodes transmembrane protein 160. Overexpression of TMEM160 promotes gastric cancer cell proliferation and enhances chemotherapy drug resistance; knockdown of TMEM160 inhibits gastric cancer cell proliferation and increases chemotherapy drug sensitivity. TMEM160 has also been found to promote chemotherapy resistance in gastric cancer cells by inhibiting ferroptosis through the Keap1 / Nrf2 pathway. There are also reports of using this gene as a marker for lung cancer risk assessment, and another report using its expression level to determine colon cancer prognosis. Currently, there are no reports of this gene sequence being associated with breast cancer prognosis.
[0072] The term "EWSR1 gene" used in the present disclosure encodes Ewing sarcoma breakpoint region 1, which plays an important role in mitotic cell separation, spindle formation, microtubule stability, DNA repair, and cell senescence. It is a member of the TET family of protein cytokines that control cell growth.
[0073] The term "BCAT2 gene," as used herein, encodes branched-chain amino acid transaminase 2, a key enzyme in BCAA catabolism and associated with estrogen receptor signaling. It is highly expressed in ductal epithelial cells during the early PanIN stage of pancreatic cancer development and may play an important role in the development and progression of PanIN. Furthermore, in malic enzyme-deficient pancreatic ductal adenocarcinoma, overexpression of BCAT2 significantly promotes tumor cell proliferation. There are also reports that this gene is used as a molecular marker for diffuse gastric cancer classification. Currently, there are no reports that this gene sequence is associated with breast cancer prognosis.
[0074] The term "PNKP gene" as used herein refers to polynucleotide 5' kinase-3' phosphatase. Expression of PNKP in pancreatic cancer tissue and adjacent adjacent tissues has been shown to be abnormally low in pancreatic cancer patient tumor tissue and cells, with lower PNKP expression associated with higher pancreatic cancer malignancy. While there are reports of polynucleotide 5' kinase-3' phosphatase being used as a pancreatic cancer marker, there are currently no reports linking this gene sequence with breast cancer prognosis.
[0075] As used herein, the term "MLEC gene" refers to a galectin gene encoding malectin, a type I membrane-anchored ER protein. MLEC has an affinity for Glc2Man9GlcNAc2 (G2M9) N-glycans and is involved in regulating ER glycosylation. MLEC has also been shown to interact with ribosome-associated glycoprotein I and may be involved in the degradation of misfolded proteins. MLEC is dysregulated in colorectal cancer and enhanced in glioblastoma. MLEC may be a biomarker for papillary thyroid carcinoma, and there are also reports of this gene being used as a diagnostic biomarker for pancreatic cancer. Currently, there are no reports of this gene sequence being associated with breast cancer prognosis.
[0076] The term "SAP30BP gene" used in the present disclosure encodes SAP30 binding protein. There are reports that this gene is associated with the prognosis of non-small cell lung cancer. Currently, there are no reports that this gene sequence is associated with the prognosis of breast cancer.
[0077] The term "TBL1XR1 gene" as used herein encodes transducin β1X-linked receptor protein 1. The gene product is a nuclear protein ubiquitous in most tissue markers, affecting tumor cell proliferation, migration, invasion, and progression. TBL1XR1 expression is associated with the clinical progression of various human cancers and patient survival. TBL1XR1 is overexpressed in a variety of malignancies, and detecting TBL1XR1 can help predict cancer prognosis, such as ovarian cancer. This gene has been reported to be a lung cancer-related driver gene or targetable drug resistance gene, and a gene specific for single-nodule liver cancer. Other reports have also used this gene for esophageal cancer classification.
[0078] The term "STAG1 gene" used in the present disclosure encodes a subunit of the cohesin complex, which plays a vital role in controlling chromosome segregation during cell division. There has been no report that this gene sequence is associated with breast cancer prognosis.
[0079] The term "NR4A1 gene" used in the present disclosure encodes nuclear receptor subfamily 4A member 1. NR4A1 is often overexpressed in various cancer cells, including lung cancer, prostate cancer, breast cancer and colon cancer, and is believed to represent a survival signal that mediates cancer cell growth.
[0080] The term "TNFRSF9 gene" used in this disclosure refers to tumor necrosis factor receptor superfamily member 9. It has been reported that this gene is a marker gene for predicting the efficacy of immunotherapy for esophageal squamous cell carcinoma. Currently, there are no reports that this gene sequence is associated with the prognosis of breast cancer.
[0081] The term "PSD3 gene" used in the present disclosure encodes Sec7 domain 3, which has been reported to be a genetic marker for predicting the efficacy of adjuvant radiotherapy in breast cancer patients and a major biomarker for identifying benign and malignant thyroid cancers.
[0082] The term "BAD gene" as used in this disclosure encodes the BCL2-associated death promoter, and the BAD pathway is associated with the development, progression, response to treatment, and overall survival of cancer, and these effects may be due to changes in the phosphorylation state of the BAD protein. The effects of the BAD pathway and phosphorylated proteins appear to apply to a wide range of human cancer types, as well as to the sensitivity of these cancers to a wide range of cytotoxic drugs. Whole-genome expression analysis determined that the BCL2-associated death promoter (BAD) apoptosis pathway and the phosphorylation state of the BAD protein are associated with the development of platinum resistance in ovarian cancer. It was also found that BAD phosphorylation is closely related to the chemotherapy sensitivity and patient survival rate of ovarian cancer. In addition, the expression of the BAD pathway is associated with the overall survival of ovarian cancer patients. To date, there have been no reports of this gene sequence being associated with the prognosis of breast cancer.
[0083] The term "MIIP gene," as used herein, refers to a tumor suppressor gene. Located on chromosome 1p36.2, the MIIP gene is 12,595 base pairs long and contains 10 exons. It encodes a 2,024-base pair of mRNA and a 388-amino acid protein, which is widely expressed in various tissues in the human body. MIIP expression is reduced in prostate cancer and is associated with CRPC progression, suggesting its potential role as a regulatory factor in prostate cancer. There are also reports that this gene can be used in a predictive model for liver cancer prognosis. Currently, there are no reports linking this gene sequence with breast cancer prognosis.
[0084] The term "HIST3H2A gene," as used herein, encodes histone cluster 3 H2a. This gene has been reported to be useful as a protein marker for gastric cancer typing or as a biomarker for characterizing colorectal cancer by comparison with a reference. HIST3H2A has also been reported to be a potential biomarker for pancreatic cancer. Currently, there are no reports linking this gene sequence to breast cancer prognosis.
[0085] The term "CDC25B gene" used in the present disclosure encodes cell division cycle 25B. It has been reported that this gene can be used as a marker for predicting the prognosis of gastric cancer. Another report uses this gene for early screening of esophageal cancer. There are also reports that this gene can predict esophageal squamous cell carcinoma and its immunotherapy efficacy.
[0086] The term "TCEB1 gene" used in the present disclosure encodes transcription elongation factor B1. It has been reported that TCEB1 is significantly overexpressed in liver cancer tissues, and the higher the expression level, the worse the patient's tumor grade and prognosis.
[0087] The term "HMGCS1 gene" as used herein encodes 3-hydroxy-3-methylglutaryl-CoA synthetase, a key metabolic enzyme in the cholesterol synthesis pathway (the mevalonate pathway). Reports indicate that this gene, as a terpenoid backbone biosynthesis gene and protein, is associated with prostate cancer prognosis. GEPIA database analysis shows that HMGCS1 is highly correlated with Rad17 in cervical cancer tissue. Rad17 influences the development and progression of cervical cancer by regulating 3-hydroxy-3-methylglutaryl-CoA synthase 1 (HMGCS1). HMGCS1 is closely associated with the migration and invasion of cervical cancer cells. Currently, there are no reports of this gene sequence being associated with breast cancer prognosis.
[0088] The term "SPC25 gene" used in the present disclosure encodes a component of the NDC80 kinetochore complex. Compared with normal mucosa, SPC25, together with CDCA1, KNTC2 and APC24, is overexpressed in colorectal cancer and gastric cancer. This gene has been reported as a prognostic marker for lung cancer.
[0089] The term "TKT gene" used in this disclosure encodes pituitary tumor transforming gene 1. It has been reported that this gene can be used to predict the prognosis of liver cancer. It has also been reported that this gene is associated with smoking status or serves as a metabolite-related feature of gastric cancer. Currently, there are no reports that this gene sequence is associated with the prognosis of breast cancer.
[0090] The term "PTTG1 gene" used in the present disclosure is pituitary tumor transforming gene 1, which encodes separase inhibitory protein. Studies have shown that PTTG1 is significantly highly expressed in various tumors, and is also expressed in large quantities and highly active in testicular and thymic tissues, but is lowly expressed or not expressed in other normal human tissues.
[0091] The term "ADA gene" used in the present disclosure encodes adenosine deaminase. There are reports that this gene is used as a marker for molecular typing of diffuse gastric cancer. Currently, there are no reports that this gene sequence is associated with the prognosis of breast cancer.
[0092] The term "AK1 gene" used in the present disclosure encodes adenylate kinase 1. Currently, there is no report that this gene sequence is associated with the prognosis of breast cancer.
[0093] The term "AIMP2 gene" used in the present disclosure encodes aminoacyl-tRNA synthetase complex interacting multifunctional protein 2. It has been reported that AIMP2, as a new type of tumor suppressor, has the function of enhancing TGF-β signaling through direct interaction with Smad2 / 3. At present, there are no reports that this gene sequence is associated with breast cancer prognosis.
[0094] This paper uses databases such as GEO and TCGA to obtain single-cell sequencing data of triple-negative breast cancer with phenotypic information; then displays the expression of ferroptosis-related genes based on sample clinical data; divides subpopulations based on NMF clustering; cell timing analysis, subtype identification, communication analysis, TF activity difference detection, functional enrichment analysis of macrophage subpopulations mediated by ferroptosis genes in negative breast cancer, correlation analysis of TME patterns mediated by ferroptosis genes in triple-negative breast cancer with prognosis and immunotherapy response, univariate Cox analysis, Lasso regression, external data verification, and independent prognostic factor verification; ROC curve analysis and risk score screen out genes closely related to the prognosis of triple-negative breast cancer.
[0095] Example 1. Process for screening markers
[0096] The process of screening markers is as follows Figure 1 As shown, the following steps are included:
[0097] 1. Data download and preprocessing
[0098] 1.1. Data download:
[0099] (1) GSE176078 was downloaded from the GEO database as single-cell data: https: / / www.ncbi.nlm.nih.gov / geo / q uery / acc.cgi, from which triple-negative breast cancer samples were extracted for analysis. The information of GSE176078 triple-negative breast cancer samples is shown in Table 1.
[0100] GSE176078 was downloaded from the GEO data set, and a total of 9 triple-negative breast cancer adult samples were used for subsequent analysis. After processing the raw data, a 38,985 gene matrix representing 29,733 cells was obtained for subsequent analysis.
[0101] (2) TCGA-BRCA data download:
[0102] Expression data: https: / / gdc-hub.s3.us-east-1.amazonaws.com / download / TCGA-BRCA.htseq_fpk m.tsv.gz;
[0103] Survival data: https: / / tcga-xena-hub.s3.us-east-1.amazonaws.com / download / survival%2FBRC A_survival.txt .
[0104] (3) Other bulk data GSE10893, GSE25066, GSE45725, GSE86166, GSE96058, GSE42664, GSE100797, GSE126044, GSE135222, and GSE165252 were downloaded from the GEO database: https: / / www.ncbi.nlm.nih.gov / geo / query / acc.cgi. Bulk data sample information is shown in Table 2.
[0105] (4) IMvigor210 data were downloaded from the R package IMvigor210CoreBiologies (v1.0.0).
[0106] Table 1. GSE176078 triple-negative breast cancer sample information display
[0107]
[0108]
[0109] Table 2. Bulk data sample information display.
[0110] Dataset Number of samples Grouping TCGA-BRCA 116 Overall survival / disease-free interval GSE10893 31 Overall survival / recurrence-free survival GSE25066 178 Recurrence-free survival GSE45725 36 Recurrence-free survival GSE86166 52 Overall survival / recurrence-free survival GSE96058 151 Overall survival GSE42664 37 Immunotherapy GSE100797 25 Immunotherapy GSE126044 16 Immunotherapy GSE135222 27 Immunotherapy GSE165252 70 Immunotherapy IMvigor210 298 Immunotherapy
[0111] Note 1: The disease-free interval (DFI) refers to the period from the completion of tumor resection or elimination to the diagnosis of recurrence. The disease-free interval is an independent factor affecting prognosis and is applicable to the survival evaluation of tumors that can be radically cured.
[0112] Note 2: Relapse-free survival (RFS) refers to the time from when a patient achieves complete remission after anti-tumor treatment to when relapse occurs or the end of follow-up.
[0113] 1.2. Single-cell data processing:
[0114] The breast cancer single-cell dataset GSE176078 contains triple-negative breast cancer data. We extracted relevant data for subsequent analysis. Because the original literature contains cell annotation metadata, we used the original cell annotation results directly.
[0115] The R package Seurat (v4.1.1) was used to perform quality control and analysis of the data:
[0116] (1) First, based on the relevant feature data (nFeature_RNA, nCount_RNA, percent.mt), we removed extreme cells (cells with the number of expressed genes (nFeature_RNA) less than 200 and the proportion of mitochondrial RNA (percent.mt) higher than 20%), and drew scatter plots and violin plots for feature display;
[0117] (2) Normalize the expression data of each sample cell using the NormalizeData() function, and use the FindVariabl eFeatures() function to identify the highly variable genes in each sample (genes with the largest expression differences between cells, 2000 genes are obtained by default);
[0118] (3) Due to the batch effect between samples, we first use FindIntegrationAnchors() to find anchor points and use Integrate Data() to integrate the standardized data;
[0119] (4) Normalize the expression values using the ScaleData() function; perform linear dimensionality reduction based on highly variable genes and the first 100 PCs (principal components) using RunPCA();
[0120] (5) Use the RunTSNE() function to reduce the dimensionality of the integrated samples and draw a tSNE diagram based on the samples for display.
[0121] The processed data is saved as an Rds file for subsequent analysis.
[0122] 2. Ferroptosis gene map in triple-negative breast cancer TME cells
[0123] 2.1. Ferroptosis gene map:
[0124] We then used the Seurat R package to identify cell types. We first performed tSNE visualization of cell types and subtypes based on the cell annotation file. We then used the VlnPlot() and FeaturePlot() functions to plot violin plots and tSNE plots of cell type-related marker expression levels.
[0125] Communication analysis of annotated cells was performed using the R package iTALK (v0.1.0). The rawParse() function was first used to identify the top 50% of highly expressed genes in each cell type based on mean expression. The FindLR() function was then used to identify ligand-receptor interactions among these highly expressed genes. Finally, an interaction diagram was constructed based on all interactions.
[0126] 471 ferroptosis-related genes were obtained from the FerrDb database (http: / / www.zhounan.org / ferrdb / current / ). The expression of ferroptosis-related genes was then displayed based on the clinical data of the samples. The gene expression values of samples of the same type were averaged and then displayed in a heat map. The average expression values of ferroptosis genes were then averaged across different cell types and displayed in a heat map.
[0127] 2.2. Subgroup division based on NMF clustering:
[0128] Macrophage and T cell expression data were extracted from the total cell count data. Subtypes of macrophages and T cells were then identified based on ferroptosis genes. Subtype identification was primarily performed using the R package NMF (v0.26). The proportions of different subtypes were then displayed within each sample.
[0129] 3. Ferroptosis-mediated macrophage subsets in triple-negative breast cancer
[0130] 3.1. Cellular timing analysis, subtype identification, and communication analysis:
[0131] Then, R package monocle (v2.28.0) was used to perform pseudo-time series analysis on macrophages:
[0132] (1) First, use the newCellDataSet() function to construct a CDS object;
[0133] (2) Use the estimateSizeFactors() and estimateDispersions() functions to calculate the size factor (SF) and dispersion;
[0134] (3) Then, the FindVariableFeatures() function was used to identify macrophage highly variable genes, and setOrderingFilter() was used to perform sorting based on characteristic gene markers;
[0135] (4) Use the reduceDimension() function to reduce the dimensionality of the data based on the reverse graph embedding (DDRTree) algorithm;
[0136] (5) Based on the PCBC database, the OCLR algorithm was used to train and derive the stemness index. The stemness score of each cell was calculated based on the Spearman correlation. The stemness scores of different stages were counted. The stage with the highest score was selected as the starting point for trajectory construction. The cells were sorted using the orderCells() function to complete the trajectory construction.
[0137] (6) Finally, the cell trajectory map was drawn based on the ferroptosis subtypes identified by time series analysis and non-negative matrix factorization (NMF).
[0138] A heat map was further drawn to display the average expression levels of ferroptosis-related genes in ferroptosis-mediated macrophage subtypes. Due to the large number of ferroptosis genes, the 50 genes with the highest expression levels were selected for display.
[0139] Ferroptosis-mediated macrophage subtypes were then identified based on PMIDs 35325594 and 32783918. The identified macrophage subtypes and subtype-specific markers were then displayed using tSNE. The AddModuleScore() function in the R package Seurat was used to obtain the ferroptosis-mediated macrophage subtype score for each cell. A tSNE display was then performed based on the subtype score, and the correlation between subtype cells was calculated based on the subtype score. Finally, the R package iTALK was used to analyze the ferroptosis-mediated communication between macrophages and tumor epithelium, and an interaction diagram was drawn based on all interactions.
[0140] 3.2. Detection of TF activity differences:
[0141] Establishing a gene regulatory network requires the following steps: identifying potential targets of each transcription factor based on co-expression (selecting a gene expression matrix and running GENIE3 / GRNBOOST, formatting targets into co-expression modules within GENIE3 / GRNBOOST); selecting potentially directly related targets based on DNA-motif analysis (determining cell states and their regulators); and analyzing the network activity of each cell (scoring regulators in each cell and calculating the AUC, converting the network activity into a two-dimensional matrix). The R package SCENIC (v1.3.1) was used to determine whether subtype marker genes were specific transcription factors and construct the gene regulatory network:
[0142] (1) First, the motif scoring data of 500 bp upstream of human TSS (hg19-500bp-upstream-7species.mc9nr.feather) and 20 kb around human TSS (hg19-tss-centered-10kb-7species.mc9nr.feather) were downloaded from the cisTarget database (https: / / resources.aertslab.org / cistarget / databases / old / homo_sapie ns / hg19 / refseq_r45 / mc9nr / gene_based / );
[0143] (2) Then use the initializeScenic() function to create a SCENIC object;
[0144] (3) Gene filtering was performed using the geneFiltering() function based on minCountsPerGene and minSamples = 5% of the number of samples;
[0145] (4) Perform correlation analysis using the runCorrelation() function;
[0146] (5) Use the runGenie3() function to construct the co-expression network;
[0147] (6) Use runSCENIC_1_coexNetwork2modules() to obtain co-expression modules;
[0148] (7) Use runSCENIC_2_createRegulons() to obtain regulators;
[0149] (8) Use the runSCENIC_3_scoreCells() function to perform scoring based on the AUCell algorithm;
[0150] (9) Use the runSCENIC_4_aucell_binarize() function to convert the data into two dimensions.
[0151] After the calculation process is completed, the transcription factor activity information of the subtype cells is stored in the file "3.1_regulons_forAUCell.Rds", from which the differences in transcription factor activity between macrophage subtype cells mediated by ferroptosis are detected and a heat map is drawn for display.
[0152] 4. Ferroptosis-mediated T cell subsets in triple-negative breast cancer
[0153] 4.1 Then, the R package monocle (v2.28.0) was used to perform pseudo-time analysis on T cells. Based on the same analysis method as for macrophages, a cell trajectory diagram was finally drawn based on pseudotime, the original cell annotated T cell subtypes, and the ferroptosis subtypes identified by NMF. The original annotated T cell subtypes were then displayed by tSNE. The proportion of the original annotated T cell subtypes in the ferroptosis-mediated T cell subtypes was subsequently counted and a bar chart was drawn for display. The ferroptosis subtypes identified by NMF were further subdivided based on the original annotated T cell subtypes to obtain the final ferroptosis-mediated T cell subtypes.
[0154] A heat map was drawn to display the average expression levels of ferroptosis-related genes in ferroptosis-mediated T cell subtypes. Due to the large number of ferroptosis genes, the 50 genes with the highest expression levels were selected for display.
[0155] The genes related to 8 types of T cell functional characteristic models (Co-inhibitors, Co-stimulations, Cytotoxic score, Exhaustion score, Immune checkpoints, T-fuciotnGenes, T effect score, Tevasion score) were downloaded from PMID: 35509079, and then a heat map was drawn to show the average expression levels of the related genes in subtype cells.
[0156] Finally, the communication between ferroptosis-mediated T cell subtypes and tumor epithelium was analyzed based on the R package iTALK, and an interaction diagram was drawn based on all the interactions.
[0157] 4.2. Detection of TF activity differences:
[0158] Based on the same process as above, the differences in transcription factor activity between T cell subtypes mediated by ferroptosis were detected. However, due to the large number of T cells (11,784), 50% of the cells (evenly including various subtypes) were randomly selected for calculation and a heat map was drawn for display.
[0159] 4.3. Functional enrichment analysis:
[0160] Single-cell functional enrichment analysis was performed using the R package irGSEA (v2.1.5). Based on the ssGSEA algorithm, the HALLMARK pathway scores in ferroptosis-mediated T cell subtypes were calculated, and a score heat map was drawn for display.
[0161] 5. Combining bulk RNA seq data to analyze the correlation between ferroptosis gene-mediated TME patterns and prognosis and immunotherapy response in triple-negative breast cancer
[0162] 5.1.Bulk data processing:
[0163] We downloaded prognostic data from GSE10893, GSE25066, GSE45725, GSE86166, GSE96058, and TCGA_BRCA datasets and extracted triple-negative breast cancer data. TCGA provides prognostic data for OS and DFI, while GSE 10893 and GSE86166 provide prognostic data for OS and RFS. GSE25066 and GSE45725 only have RFS data, and GSE96058 only has OS data. TCGA DFI data can be considered RFS data. Therefore, the overall prognostic analysis was divided into two types: OS-based and RFS / DFI-based. We then calculated cell subtype signature scores using the GSVA algorithm based on the signature genes of ferroptosis-mediated cell subtypes for subsequent analysis.
[0164] Download immunotherapy response data from GSE42664, GSE100797, GSE126044, GSE135222, GSE165252, and IMvigor210. Some samples were categorized as progressive, stable, partial, or complete responses. Therefore, progressive and stable samples were considered non-responders, while partial and complete responses were considered responders. Based on the signature genes of ferroptosis-mediated cell subtypes, the GSVA algorithm was used to calculate cell subtype signature scores for subsequent analysis.
[0165] 5.2. Prognostic differences based on ferroptosis-mediated TME subpopulation scores:
[0166] Based on the above prognostic data, we used the R packages survival (v3.2-7) and survminer (v0.4.8) to perform prognostic analysis of cell subtype signature scores (univariate Cox regression, grouping nodes by median). We then plotted a bubble chart based on the HR values and p-values for visualization.
[0167] 5.3. Evaluating the Differences in Response of Ferroptosis-Mediated TME Subpopulation Scores to Immunotherapy Efficacy in Cohorts:
[0168] Based on the immunotherapy response data above, we used the glm() function to perform a logistic regression analysis of the cell subtype signature scores. We then plotted a bubble chart based on the ORs and p-values for visualization. For optimal visualization, we only displayed results with an OR greater than 10.
[0169] 6. Constructing a model of TME cell subpopulation-related characteristics mediated by ferroptosis genes
[0170] 6.1. Single factor cox:
[0171] Combined with the RFS data from GSE25066, we performed batch Cox regression analysis on the subtype signature markers using the R packages survival (v3.2-7) and survminer (v0.4.8). After regression analysis, genes significantly associated with RFS were screened using a threshold of p < 0.01 for subsequent analysis.
[0172] 6.2.Lasso regression:
[0173] We then performed Lasso regression dimensionality reduction on the prognostic subtype signature markers and constructed a risk score model. This process primarily relied on the R package glmnet (v4.0-2). To build a more accurate regression model, we first used cross-validation to perform lambda screening, then selected the model corresponding to lamdba.min. We then extracted the expression matrix of the relevant genes in the model and calculated the risk score for each sample based on the following formula:
[0174]
[0175] Where exp represents the expression of the corresponding gene, β represents the regression coefficient (coef) of the corresponding gene in the lasso regression result, RScore represents the sum of the expression of the significantly correlated gene in each sample multiplied by the coef of the corresponding gene, i represents the sample, and j represents the gene.
[0176] To verify the effectiveness of the model, the risk score of the GSE25066 sample was used to divide the high-risk and low-risk groups with the median as the node. Combined with the RFS data, the KM curve was drawn and the p-value was calculated. A p-value < 0.05 was considered to indicate a significant difference between the high-risk and low-risk groups. The sample risk score was then used as the model prediction result, and the AUC value of the model was calculated in combination with the survival data. The ROC curve was then drawn. The AUC values for 3, 4, and 5 years were greater than 0.8, indicating that the model was effective.
[0177] 6.3. External Data Verification:
[0178] The results were further validated in the triple-negative breast cancer patient data from GSE86166. The risk score was calculated based on the model genes, and the high-risk and low-risk groups were divided using the median as the node. The KM curve was drawn in combination with the RFS. A p-value < 0.05 indicated that the difference between the high-risk and low-risk groups was significant. The sample risk score was then used as the model prediction result, and the AUC value of the model was calculated in combination with the survival data. The ROC curve was then drawn. The 3-year, 4-year, and 5-year AUC values were greater than 0.65, indicating that the model was effective.
[0179] The patient information statistics are shown in Table 3.
[0180] Table 3. Patient information
[0181]
[0182] In addition, we used DFI to validate and present relevant results using patient data from TCGA-BRCA patients diagnosed with triple-negative invasive breast cancer (ER- / PR- / Her2-) with detailed genomic expression profiles and complete clinical follow-up information. The patient data for triple-negative breast cancer samples is shown in Table 4.
[0183] Table 4. Patient information of triple-negative breast cancer samples
[0184]
[0185]
[0186] 6.4. Validation of Independent Prognostic Factors:
[0187] To verify whether the GSE25066 data risk score grouping was an independent prognostic factor, we first performed a Cox univariate regression analysis in combination with other prognostic factors, including age, grade, and stage. We then used a Cox multivariate regression analysis to analyze the overall prognosis of these four factors (including the risk score grouping), further demonstrating that the risk score factor was an independent prognostic factor. Patient information is shown in Table 5.
[0188] Table 5. Patient information
[0189]
[0190] The grading and staging of GSE86166 data were then combined to further verify that the risk score grouping was an independent prognostic factor.
[0191] Example 2: Ferroptosis gene map in triple-negative breast cancer TME cells
[0192] The analysis was performed according to the method described in step 2 of Example 1, and the results were as follows:
[0193] 1. Ferroptosis gene map:
[0194] Based on the original cell annotation file, 9 major cells and 29 minor cells were obtained. Among the 9 major cells, myeloid cells include circulating myeloid cells, dendritic cells, macrophages, and monocytes. Subsequently, myeloid cells were split into 4 types of cells to complete the analysis, that is, 12 types of cells were finally obtained: B cells, cancer-associated fibroblasts (CAFs), tumor epithelial cells, circulating myeloid cells, dendritic cells, endothelial cells, macrophages, monocytes, normal epithelial cells, plasmablasts, perivascular-like cells (PVL), T cells ( Figure 2 A-2C). For the other eight cell types besides myeloid cells, their corresponding cell markers are MS4A1 (B-cells), CD79A (Plasmablasts), PDGFRA (CAFs), EPCAM (Cancer Epithelial), PECAM1 (Endothelial), JCHAIN (Plasmablasts), ACTA2 (PVL), and CD3D (T-cells). The cell marker corresponding to myeloid cells is CD68.
[0195] Afterwards, 471 ferroptosis-related genes were obtained based on the FerrDb database, of which 391 genes had expression information in the single-cell matrix. According to the steps described in Example 1, a ferroptosis gene map was drawn. Based on the clinical information of different samples, the expression differences of ferroptosis-related genes in different clinical information were displayed. Based on the heat map display, it was found that the average expression levels of ferroptosis-related genes in various groups of age, cell type, and different T stages were significantly different, suggesting that the expression of ferroptosis-related genes is related to the above clinical characteristics ( Figure 2 D). The average expression of ferroptosis-related genes in 12 cell types is then displayed. The heat map shows that there are significant differences in the average expression levels of related genes in the 12 cell types ( Figure 2 E).
[0196] According to the steps described in Example 1, the communication between the 12 types of cells was analyzed. The results showed that the communication between the 12 types of cells was relatively strong ( Figure 2 F).
[0197] 2. Subgroup division based on NMF clustering:
[0198] According to the steps recorded in Example 1, the expression matrix of macrophage and T cell ferroptosis-related genes was extracted from all cells, and then ferroptosis-mediated macrophage and T cell subtype identification was performed based on the NMF algorithm. The rank at which cophenetic begins to decline is taken as the best rank. First, in the identification of macrophage subtypes, the best rank = 4, but the characteristic markers of the subsequent four subtypes indicate that the four subtypes are similar in pairs, so the best classification number is set to 2 here, and two ferroptosis-mediated macrophage subtypes (M_C1 and M_C2) are finally obtained. Afterwards, in the identification of T cell subtypes, the best rank = 3, so three ferroptosis-mediated T cell subtypes (T_C1, T_C2, T_C3) are obtained. A bar graph is drawn to show the proportion of related subtype cells in the sample, see Figure 3 .
[0199] Example 3: Macrophage Subpopulations Mediated by Ferroptosis Genes in Triple-Negative Breast Cancer
[0200] The analysis was performed according to the method described in step 3 of Example 1, and the results were as follows:
[0201] 1. Cell timing analysis, subtype identification, and communication analysis:
[0202] There are 3671 macrophages in the single-cell sequencing data. First, we use highly variable genes to perform pseudo-time series analysis on macrophages, and draw a pseudo-time series diagram of cells based on cell stemness ( Figure 4 A). Then, a pseudo-time sequence diagram was drawn based on the ferroptosis-mediated macrophage subtypes identified by NMF. In the figure, the distinction between M_C1 subtype cells and M_C2 is more obvious ( Figure 4 B). Further drawing the tSNE diagram of the two subtypes, the tSNE dimensionality reduction division of the two subtypes is not very obvious. M_C2 is mainly concentrated in the center, and M_C1 is mainly distributed around ( Figure 4 C). The expression differences of the top 50 ferroptosis-related genes in the two cell subtypes are then displayed. The centralized heat map shows that the expression of related genes is significantly different between different cell subtypes ( Figure 4 D). The FindAll markers() function was then used to identify markers of ferroptosis-mediated macrophage subtypes. Among them, the highly expressed markers in M_C1 include FOLR2, SEPP1, MRC1, LYVE1, SLC40A1, and CD163, so it can be defined as FOLR2+MAC. The highly expressed markers in M_C2 include TREM2, FN1, CXCR4, C3, S100A8, IFI6, and SPP1, so it can be defined as TREM2+MAC. The enrichment scores of the two subtypes of cells were then calculated for each cell, and tSNE display was performed based on the enrichment scores of the two subtypes of cells. At the same time, tSNE display of the expression of the two subtype markers was performed ( Figure 4EH). Pearson correlation analysis was further used to calculate the correlation of enrichment scores. The bubble chart showed that there was a significant negative correlation (-0.779) between TREM2+MAC and FOLR2+MAC ( Figure 4 I). Based on the communication analysis between macrophages and tumor epithelium, it was found that both macrophage subtypes communicated actively with tumor epithelium ( Figure 4 J).
[0203] 2. Detection of TF activity differences:
[0204] Based on macrophages, the activity of transcription factors in each cell was detected. Then, the activity mean was integrated based on the macrophage subtype, and the top 20 transcription factors were selected to display the average activity. The centralized heat map results are as follows: Figure 5 As shown, most of the Top 20 transcription factors have higher transcription factor activities in FOLR2+MAC cells.
[0205] Example 4: T cell subsets mediated by ferroptosis genes in triple-negative breast cancer
[0206] The analysis was performed according to the method described in step 4 of Example 1, and the results were as follows:
[0207] 1. Cell timing analysis, subtype identification, and communication analysis:
[0208] There are 11,784 T cells in the single-cell sequencing data, including cycling T-cells, NK cells, NKT cells, T cells CD4+, and T cells CD8+ subtypes. First, a pseudo-sequential analysis of T cells was performed using highly variable genes. Then, a pseudo-sequential analysis scatter plot was performed based on T cell subtypes. The results showed that cycling T-cells had the strongest stemness and were clearly distinguished from other subtypes. Then, a pseudo-sequential analysis scatter plot was performed based on ferroptosis-mediated subtypes. The results showed that the distinction between ferroptosis-mediated subtypes was not obvious ( Figure 6 AC). Then, based on the cell subtypes annotated in the original text, tSNE was performed to show that the five cell types were clearly distinguished ( Figure 6 D) Further demonstrating the differences in the proportions of the five types of originally annotated T cells in ferroptosis-mediated T cell subtypes, the bar graph shows that the proportions of the five types of T cells in each subtype are not much different, and the proportion of NK cells in the T_C2 subtype is significantly higher than that in the other two subtypes ( Figure 6 E). The expression differences of the top 50 ferroptosis-related genes in the three cell subtypes are then displayed. The heat map shows that the expression differences of related genes between different cell subtypes are quite obvious ( Figure 6F). The five cell types were then integrated with ferroptosis-mediated T cell subtypes for analysis, resulting in a total of 15 T cell subtypes. Based on the characteristic genes in the eight T cell functional characteristic models, the average expression of related genes in the 15 T cell subtypes was displayed. The heat map showed that the related genes differed among different T cell subtypes, with the expression of related genes in the cycling T-cell subtype being significantly higher than that in other cell types ( Figure 6 G). Based on the communication analysis between T cells and tumor epithelium, it was found that all subtypes of T cells communicated more actively with tumor epithelium ( Figure 6 H).
[0209] 2. Detection of TF activity differences:
[0210] Based on T cells, the activity of transcription factors in each cell was detected. Then, the activity mean was integrated based on the T cell subtype, and the top 20 transcription factors were selected to display the average activity. The heat map is as follows Figure 7 As shown in the figure, among the top 20 transcription factors, some have higher transcription factor activities in NKT cells and NK cell subtype cells (TBX21_extended, CEBPB_extended), and some have higher transcription factor activities in T cell CD8+ subtype cells (CEBPB_extended, STAT2_extended).
[0211] 3. Functional enrichment analysis:
[0212] Based on the HALLMARK dataset, the enrichment scores of T cell subtypes were calculated, and the heat map is shown as follows: Figure 8As shown in the figure, most of the average enrichment scores of Cycling T-cells subtype cells are relatively high, such as HALLMARK-DNA-REPAIR, HALLMARK-ADIPOGENESIS, HALLMARK-ADIPOGENESIS, etc. The average enrichment scores of HALLMARK-APICAL-SURFACE, HALLMARK-INFLAMMATORY-RESPONSE, HALLMARK-TNFA-SIGNALING-VIA-NFKBHALLMARK-TNFA-SIGNALING-VIA-NFKB and other pathways in NK cells subtype cells are relatively high, while most of the enrichment scores of the remaining subtypes are relatively low, but the enrichment scores of HALLMARK-EPITHE LIAL-MESENCHYMAL-TRANSITION, HALLMARK-KRAS-SIGNALING-DN, HALLMARK-KRAS-SIGNALING-UP and other pathways are higher than those of Cycling T-cells subtypes. Overall, there was no significant difference in the enrichment scores between ferroptosis-mediated T cell subtypes.
[0213] Example 5: Analysis of the correlation between ferroptosis gene-mediated TME pattern and prognosis and immunotherapy response in triple-negative breast cancer by combining bulk RNA seq data
[0214] The analysis was performed according to the method described in step 5 of Example 1, and the results were as follows:
[0215] 1. Prognostic differences based on ferroptosis-mediated TME subpopulation scores:
[0216] Download GSE10893, GSE25066, GSE45725, GSE86166, GSE96058, TCGA_BRCA data, based on specific markers of macrophage subtypes and T cell subtypes, see Figure 9 . Calculate the functional enrichment score of each subtype of cells in each sample, and then perform single factor Cox regression. Figure 10 As shown in A, the results based on OS showed that all cells in the TCGA_BRCA data had no significant prognostic effect; M_C1 cells in the GSE10893 data had a significant prognostic effect, and the HR of M_C1 was 9.372; T cells CD4+T_C3 and T cells CD8+T_C2 cells in the GSE86166 data had a significant prognostic effect, and the HRs were both greater than 1; T cells CD4+T_C3 and T cells CD8+T_C2 in the GSE96058 data had a significant prognostic effect, but the HRs were both less than 1, which was contrary to the conclusion of GSE86166. Figure 10 As shown in B, the results based on RFS showed that all cells in TCGA_BRCA, GSE10893, GSE45725 and GSE86166 data had no significant prognostic effect; Cycling T-cells T_C2 cells in GSE25066 data had a significant prognostic effect, with HR = 0.556.
[0217] 2. Evaluate the differences in ferroptosis-mediated TME subset scores in response to immunotherapy efficacy in the cohorts:
[0218] Download GSE42664, GSE100797, GSE126044, GSE135222, GSE165252, and IMvigor210 data. Based on the specific markers of macrophage subtypes and T cell subtypes, the functional enrichment scores of each subtype of cells in each sample were calculated, and then logistic regression was performed. The results are as follows: Figure 11 As shown, all enrichment scores had no significant effect on immunotherapy response.
[0219] Example 6: Construction of a model of TME cell subpopulation-related characteristics mediated by ferroptosis genes
[0220] The analysis was performed according to the method described in step 6 of Example 1, and the results were as follows:
[0221] 1. Single factor cox:
[0222] Due to the poor consistency of OS data, RFS was used for model construction. Given the large number of samples in GSE25066, a batch Cox univariate regression analysis of 8371 ferroptosis-mediated subtype-specific markers was performed on cancer samples based on GSE25066 combined with RFS data. Figure 12 The forest plot of the top 20 genes in the bulk Cox univariate regression analysis is shown. 41 genes significantly associated with survival (p < 0.01) were screened and used for subsequent analysis.
[0223] 2. LASSO:
[0224] Further Lasso regression dimensionality reduction was performed on the single factor cox regression results ( Figure 13A to C), 23 genes were obtained to construct the risk score model, namely: RiskScore = TMEM160×(-0.4689)+EWSR1×(-0.4237)+BCAT2×(-0.2346)+PNKP×(-0.1592)+MLEC×(-0.1567)+SAP30BP×(-0.1469)+TBL1XR1×(-0.0894)+STAG1×(-0.0695)+NR4A1×(-0.0630)+TN FRSF9+(-0.0503)+PSD3×(-0.0153)+BAD×(-0.0094)+MIIP×0.0104+HIST3H2A×0.0294+CDC25B×0.0528+TCEB1× 0.0670+HMGCS1×0.0988+SPC25×0.1317+TKT×0.2234+PTTG1×0.2960+ADA×0.3097+AK1×0.4003+AIMP2×0.4081.
[0225] Detailed information of the 23 genes is shown in Table 6 .
[0226] Table 6. Information of 23 genes
[0227] Serial number Gene GeneID chromosome NC Number Location 1 TMEM160 54958 19 NC_000019.10 47045909..47048624 2 EWSR1 2130 22 NC_000022.11 29268268..29300521 3 BCAT2 587 19 NC_000019.10 48795064..48811029 4 PNKP 11284 19 NC_000019.10 49861204..49867576 5 MLEC 9761 12 NC_000012.12 120687149..120701859 6 SAP30BP 29115 17 NC_000017.11 75667338..75708059 7 TBL1XR1 79718 3 NC_000003.12 177019344..177201800 8 STAG1 10274 3 NC_000003.12 136336236..136752378 9 NR4A1 3164 12 NC_000012.12 52022832..52059503 10 TNFRSF9 21942 4 NC_000070.7 151004612..151030561 11 PSD3 23362 8 NC_000008.11 18527303..19084805 12 BAD 572 11 NC_000011.10 64269828..64284704 13 MIIP 60672 1 NC_000001.11 12019498..12032045 14 HIST3H2A 92815 1 NC_000001.11 228457364..228457873 15 CDC25B 994 20 NC_000020.11 3786951..3806115 16 TCEB1 420187 2 NC_052533.1 117730456..117744263 17 HMGCS1 3157 5 NC_000005.10 43287470..43313412 18 SPC25 57405 2 NC_000002.12 168861521..168890430 19 TKT 7086 3 NC_000003.12 53224712..53256022 20 PTTG1 9232 5 NC_000005.10 160421855..160428744 21 ADA 100 20 NC_000020.11 44619522..44651699 22 AK1 203 9 NC_000009.12 127866480..127879621 23 AIMP2 7965 7 NC_000007.14 6009272..6023834
[0228] Based on the GSE25066 dataset, the sample risk score was calculated according to the formula. Based on the sample risk score, the high-risk group and the low-risk group were divided with the median 3.345 as the node. When the sample risk score was higher than the median, it was divided into the high-risk group. When the sample risk score was lower than the median, it was divided into the low-risk group. Combined with the RFS data, the KM curve of the high-risk group and the low-risk group was drawn ( Figure 14 A), the survival curves of the two groups were significantly different (p<0.05); then the risk score of the sample was further used as the model prediction result ( Figure 14 B) Draw the ROC curve and divide all samples of the GSE25066 dataset into 3-year survival patient group, 4-year survival patient group and 5-year survival patient group. The AUC of the three groups is greater than 0.8 ( Figure 14 C), indicating that the model is effective. The survival time scatter plot of all samples is shown in Figure 14 D, expression heat map of genes in the model is shown in Figure 14 E.
[0229] 3. External data verification:
[0230] The GSE86166 data were used to verify the model effectiveness. Based on the risk score of the sample, the median of 3.0427 was used as the node for the high-risk group and the low-risk group. The KM curve ( Figure 15A) shows that the difference between the two groups of data reached a significant level (p < 0.05); then further draw the ROC curve ( Figure 15 B), 3-year survival patient group, 4-year survival patient group and 5-year survival patient group, the AUC of the three groups is greater than 0.65, indicating that the model has good predictive efficiency. The risk score curve of all samples is shown in Figure 15 C; Scatter plot of survival time of all samples is shown in Figure 15 D; The expression heat map of genes in the model is shown in Figure 15 E.
[0231] The TCGA-BRCA triple-negative breast cancer samples were used for validation, and the prognostic analysis of the risk score was performed based on the disease-free interval (DFI). The KM curve p = 0.082 was obtained, and the ROC curve showed that the AUC of the 4-year survival group and the 5-year survival group of patients were greater than 0.65, proving that the risk score has a certain prognostic efficacy ( Figure 16 A and Figure 16 B).
[0232] 4. Validation of independent prognostic factors:
[0233] GSE25066 data were further used to test the independent prognostic effects of high-risk and low-risk groups. Figure 17 As shown in the results, the risk score groups have good prognostic power and are independent of each other. In addition, the stage also has good prognostic power.
[0234] We also combined other factors, such as grading and staging, to verify whether the risk score grouping in the GSE86166 data is an independent prognostic factor. Figure 18 As shown in the results, risk score grouping and stage grouping have good prognostic efficacy and are independent of each other.
[0235] 5. Differences in chemotherapy resistance between different feature model groups
[0236] The R package pRRophetic (v0.5) was used to predict the responses of the samples to 138 drugs. The predicted IC50 values were obtained. The correlation between the IC50 and the risk score was calculated. Significant correlations (p < 0.05) were selected for subsequent analysis. Positively and negatively correlated drugs were then distinguished and presented in a lollipop plot. The differences in IC50 values between high- and low-risk groups were further analyzed and presented in a boxplot. The significance of the differences was tested using the wilcox.test() test.
[0237] The analysis results selected 50 drugs that were significantly correlated with the risk score, of which 16 were positively correlated with the risk score and the remaining 34 were negatively correlated with the risk score. The differences in IC50 of related drugs in high-risk and low-risk groups were further statistically analyzed. Among the drugs positively correlated with the risk score, 13 drugs had significantly higher IC50 in the high-risk group, suggesting that the sensitivity of related drugs to samples in the high-risk group was significantly lower than that in the low-risk group ( Figure 19 A), while 27 drugs that were negatively correlated with risk scores had significantly higher IC50 in the low-risk group, suggesting that the sensitivity of the related drugs to the low-risk group was significantly lower than that to the high-risk group ( Figure 19 B).
[0238] The above results show that the above risk scoring model is closely related to the prognosis of triple-negative breast cancer and has the potential to serve as a characteristic factor for the prognosis prediction and analysis of triple-negative breast cancer patients. The evaluation model of the prognostic effect of triple-negative breast cancer based on the above 23 genes can effectively evaluate the prognostic effect of triple-negative breast cancer.
[0239] Example 7: Validation Experiment of Risk Score Model for Detecting Triple-Negative Breast Cancer
[0240] Based on the clinical data of 22 triple-negative breast cancer clinical patients (positive cases) and 8 non-triple-negative breast cancer patients (negative cases) from the Cancer Hospital of the Chinese Academy of Medical Sciences, the risk score model obtained using Example 6 was verified, and the triple-negative breast cancer risk of 22 patients was judged according to the obtained RiskScore value. The threshold value 3.0427 used in the external data verification was used as the threshold value of this verification experiment. Patients with RiskScore values higher than the threshold value were judged to be high risk, and patients with RiskSco re values lower than the threshold value were judged to be low risk. The results were compared with the clinical diagnosis results, and statistical sensitivity and specificity demonstrated that the gene markers disclosed herein and the prediction model showed good diagnostic performance. Patient information is shown in Table 7.
[0241] Table 7. Patient information of 22 clinical patients with triple-negative breast cancer
[0242]
[0243] Verification results are as follows Figure 20 As shown, the sensitivity reached about 0.857, which means that 85.7% of patients in actual positive cases can be correctly identified; the specificity was about 0.750, which means that 75.0% of patients in actual negative cases can be accurately distinguished.
[0244] The technical solution of the present disclosure is not limited to the above-mentioned specific embodiments. Any technical variations made according to the technical solution of the present disclosure fall within the protection scope of the present disclosure.
Claims
1. A gene marker associated with breast cancer prognosis, characterized in that: The gene markers include one or more of the following genes: TMEM160, EWSR1, BCAT2, PNKP, MLEC, SAP30BP, TBL1XR1, STAG1, NR4A1, TNFRSF9, PSD3, BAD, MIIP, HIST3H2A, CDC25B, TCEB1, HMGCS1, SP C25, TKT, PTTG1, ADA, AK1 and AIMP2.
2. A detection kit for evaluating the prognosis of breast cancer, characterized in that: The kit comprises: a reagent for detecting the gene marker according to claim 1.
3. The kit according to claim 2, wherein The reagents for detecting the gene markers include primer pairs or probes targeting the gene markers.
4. The kit according to claim 3, wherein The probe is a fluorescent probe or a hybridization probe.
5. A method for evaluating the prognosis of breast cancer, characterized in that: The method comprises: The calculation model was used to calculate the risk score of the breast cancer patient to be tested, and the risk score = TMEM160×(-0.4689)+EWSR1×(-0.4237)+BCAT2×(-0.2346)+PNKP×(-0.1592)+MLEC×(-0.1567)+SAP30BP×(-0.1469)+TBL1XR1×(-0.0894)+STAG1×(-0.0695)+NR4A1×(-0.0630)+TNFRSF9+(-0.0503)+P SD3×(-0.0153)+BAD×(-0.0094)+MIIP×0.0104+HIST3H2A×0.0294+CDC25B×0.0528+TCEB1×0.0670+HMGCS1×0.0988+SPC25×0.1317+TKT×0.2234+PTTG1×0.2960+ADA×0.3097+AK1×0.4003+AIMP2×0.4081; wherein, each gene in the calculation model represents the expression level of the gene in the sample; When the risk score of the breast cancer patient to be tested is higher than the threshold, the prognosis is judged to be high risk; if it is lower than the threshold, the prognosis is judged to be low risk.
6. The evaluation method according to claim 5, characterized in that The threshold is obtained by collecting breast cancer patient data, calculating the risk score, and summarizing and counting the risk scores of the multiple patients to obtain the median as the threshold.
7. A breast cancer prognosis evaluation system, characterized in that: The system comprises: (1) a data input module for inputting the expression level of the gene marker according to claim 1; (2) Model calculation module, based on the data input in step (1), the risk score of breast cancer patients is calculated using the calculation model = TMEM160×(-0.4689)+EWSR1×(-0.4237)+BCAT2×(-0.2346)+PNKP×(-0.1592)+MLEC×(-0.1567)+SAP30BP×(-0.1469)+TBL1XR1×(-0.0894)+STAG1×(-0.0695)+NR4A1×(-0.0630)+T NFRSF9+(-0.0503)+PSD3×(-0.0153)+BAD×(-0.0094)+MIIP×0.0104+HIST3H2A×0.0294+CDC25B×0.0528+TCEB1×0.0670+HMGCS1×0.0988+SPC25×0.1317+TKT×0.2234+PTTG1×0.2960+ADA×0.3097+AK1×0.4003+AIMP2×0.4081 were calculated; (3) Result output module: when the patient's risk score is higher than the threshold, the prognosis is judged to be high risk; if it is lower than the threshold, the prognosis is judged to be low risk.
8. The evaluation system according to claim 7, characterized in that The threshold is obtained by collecting breast cancer patient data, calculating the risk score, and summarizing and counting the risk scores of the multiple patients to obtain the median as the threshold.
9. Use of the gene marker according to claim 1 in the preparation of a kit for evaluating the likelihood of occurrence or prognosis of breast cancer.
10. The application according to claim 9, characterized in that: The breast cancer includes triple-negative breast cancer or non-triple-negative breast cancer.
Citation Information
Patent Citations
SE10893C1