Marker combination based on ubiquitination-related genes and application thereof in breast cancer prognosis and chemotherapy sensitivity evaluation
Patent Information
- Application Number
- CN202511320065.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-16
- Publication Date
- 2026-08-18
- Estimated Expiration
- 2045-09-16
AI Technical Summary
预测效能较低,难以满足临床对高精度个体化预后评估的需求
[0024] 1. Screening predictive biomarkers (E4F1, CBLL1, RNF13, TRIM59) based on treatment response of taxane drugs to improve the clinical applicability of the model.
Smart Images

Figure CN121294655B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of biomedical engineering, specifically relating to a combination of ubiquitination-related gene markers constructed based on taxane resistance and their application in the prognosis and chemotherapy sensitivity assessment of breast cancer. Background Technology
[0002] Breast cancer is the most common malignant tumor among women worldwide. In 2020, there were over 416,000 new cases of breast cancer in Chinese women, with over 117,000 related deaths. Although various treatments for breast cancer have been developed, including targeted drug therapy and immunotherapy, taxanes (including paclitaxel and docetaxel) are the first-line chemotherapy regimen for breast cancer. However, the development of resistance often leads to treatment failure and rapid tumor recurrence. Therefore, screening for differentially expressed genes associated with taxane resistance is essential. Furthermore, analyzing the impact of these differentially expressed genes on prognosis and drug sensitivity is crucial for identifying new targets to reverse resistance and improve prognosis.
[0003] Ubiquitination-related genes play a central role in biological processes such as cell cycle progression, cell development, DNA transcription, intracellular transport, and apoptosis by regulating protein turnover. Disruption of the ubiquitination-deubiquitination balance is closely related to the development and progression of various malignant tumors. Notably, in docetaxel-resistant prostate cancer cells, protein ubiquitination is the most differentially expressed regulatory pathway. Studies have reported that URGs such as FBXW2, USP7, and USP18 are involved in regulating the sensitivity of prostate cancer cells (BCs) to paclitaxel. This indicates that ubiquitination modification plays an important role in the development of taxane resistance in BCs, and ubiquitination-related molecules can serve as important biomarkers and therapeutic targets. Although two studies have reported the role of ubiquitin-related genes in breast cancer prognostic models, research on the role of differentially expressed ubiquitin genes related to taxane resistance in predicting the prognosis and clinicopathological features of breast cancer patients is lacking.
[0004] The closest prior art to this invention is the prognostic model for triple-negative breast cancer ubiquitination genes developed by Zhao et al. (2023). This approach obtains gene expression data (FPKM) and clinical information of TNBC patients from the METABRIC and GEO (GSE58812) databases, downloads a list of ubiquitination-related genes (URGs) from the UUCD database, and merges them with the expression data to obtain an expression matrix of 525 URGs. Univariate Cox regression (P<0.01) was used to screen for 17 URGs significantly associated with overall survival (OS). LASSO-Cox regression was then used to screen 11 key genes (HECTD3, PCG F1, RNF123, STC1, GRWD1, USP30, OTUB2, ATXN3L, BIRC3, STAMBP1, PARP11) from the 17 URGs to construct a risk scoring model. Patients were divided into high-risk and low-risk groups based on median risk values. The model's prognostic predictive ability was validated on the training set (META BRIC) and test set (GSE58812) using methods including Kaplan-Meier, ROC curves, and PCA / t-SNE analysis. Independent prognostic factors were identified using multivariate Cox regression, combining risk scores and clinicopathological factors (age, lymph node status, etc.), and nomograms were constructed to predict 1-, 3-, and 5-year survival rates. The accuracy and consistency of the nomograms were evaluated using C-index, calibration curves, and ROC curves. The sensitivity (IC50) of chemotherapy drugs to high- and low-risk groups was predicted using the GDSC database and the pRRophetic package. 50 (The title of this technical publication is: Identification of ubiquitination-related gene classification and a novel ubiquitination-related gene signature for patients with triple-negative breast cancer, published in Front Genet. 2023 Jan 6:13:932027. Link:) https: / / pubmed.ncbi.nlm.nih.gov / 36685836 / )
[0005] Current technologies for predicting prognosis and assessing chemotherapy sensitivity in breast cancer patients still have the following shortcomings:
[0006] 1. Existing prognostic models have limited predictive accuracy:
[0007] Current research (such as the prognostic model based on 11 ubiquitination-related genes constructed by Zhao et al. in 2023) has AUC values of 0.708, 0.702, and 0.744 for predicting 3-year, 5-year, and 8-year survival rates in the training set. A nomogram model constructed using age, menopausal status, lymph node status, tumor size, surgical history, Nottingham Prognostic Index (NPI), and risk score has AUC values of 0.695, 0.733, and 0.760 for predicting 1-year, 3-year, and 5-year overall survival, respectively. These predictive powers are low and fail to meet the clinical need for high-precision, personalized prognostic assessment.
[0008] 2. Selection of genes based on a single prognostic criterion failed to integrate drug resistance phenotypes:
[0009] Existing studies have mostly focused on prognostic genes and have failed to screen biomarkers based on the treatment response of taxanes, a first-line chemotherapy regimen. This has resulted in insufficient practicality of the models in predicting chemotherapy resistance and guiding medication use, which may be one of the reasons for the limited accuracy of predictions.
[0010] 3. Weak mechanistic research and lack of potential for therapeutic translation:
[0011] While existing technologies have identified multiple prognosis-related ubiquitination genes, they have not thoroughly verified their functional mechanisms in taxane resistance, nor have they provided specific targeted therapy strategies or experimental evidence, thus limiting their clinical translational value. Summary of the Invention
[0012] To address the aforementioned issues, this invention provides a combination of biomarkers for ubiquitination-related genes constructed based on taxane resistance and their application in breast cancer prognosis and chemotherapy sensitivity assessment, which relates to the biomedical engineering industry.
[0013] The first objective of this invention is to provide the application of taxane resistance-related ubiquitination-associated genes in constructing a prognostic model for breast cancer, said ubiquitination-associated genes including CBLL1, RNF13, TRIM59 and E4F1.
[0014] The second objective of this invention is to provide a method for constructing a prognostic model of breast cancer based on taxane resistance ubiquitination-related genes, which includes the following steps: constructing a nomogram using the rms package in R language based on the patient's RiskScore and TNM stage, wherein the RiskScore = expression level of E4F1 × 0.040 - expression level of CBLL1 × 1.194 - expression level of RNF13 × 0.777 - expression level of TRIM59 × 0.826.
[0015] Preferably, the breast cancer prognostic model is a breast cancer prognostic model that predicts the 1-year, 3-year and / or 5-year survival rate of patients.
[0016] A third objective of this invention is to provide the use of formulations that inhibit E4F1 gene expression in the preparation of anti-breast cancer drugs.
[0017] Preferably, it is the use of preparations that inhibit E4F1 gene expression and paclitaxel in the preparation of anti-breast cancer drugs.
[0018] Preferably, it involves the use of siRNA that inhibits E4F1 gene expression and paclitaxel in the preparation of anti-breast cancer drugs.
[0019] Preferably, the siRNA sequence is 5′-TGAAGCTACTGGTGAACAA-3′, 5′-GCAAGCGCTACAAGACTAA-3′ and / or 5′-GGACAAGTGGCACTGAACA-3′.
[0020] Preferably, the anti-breast cancer drug is a drug that inhibits the invasion or migration of breast cancer cells.
[0021] Preferably, the breast cancer cells are MDA-MB-231 cells.
[0022] Preferably, the siRNA that inhibits E4F1 gene expression enhances the cytotoxicity of paclitaxel or increases the sensitivity of breast cancer cells to paclitaxel.
[0023] Advantages of this invention:
[0024] 1. Screening predictive biomarkers (E4F1, CBLL1, RNF13, TRIM59) based on treatment response of taxane drugs to improve the clinical applicability of the model.
[0025] 2. This invention uses taxane-based drug sensitivity screening genes to construct a prognostic model, improving the accuracy of the model and providing a high-precision prognostic prediction model for breast cancer.
[0026] The existing Zhao model is constructed based on 11 prognostic-related ubiquitin genes. The AUC values for predicting 3-year, 5-year, and 8-year survival rates in the training set were 0.708, 0.702, and 0.744, respectively. A nomogram model constructed using age, menopausal status, lymph node status, tumor size, surgical history, Nottingham Prognostic Index (NPI), and risk score showed AUCs of 0.695, 0.733, and 0.760 for predicting 1-year, 3-year, and 5-year overall survival, respectively.
[0027] This invention, for the first time, screened four differentially expressed ubiquitin genes (E4F1, CBLL1, RNF13, and TRIM59) associated with prognosis in paclitaxel-sensitive and resistant breast cancer samples for constructing a prognostic model. This model achieved AUC values of 0.896, 0.866, and 0.893 for 1-year, 3-year, and 5-year survival predictions on the TCGA training set, significantly improving the accuracy of these predictions. A nomogram model constructed using risk scores and TNM staging achieved AUC values of 0.992, 0.890, and 0.897 for 1-, 3-, and 5-year overall survival (OS), demonstrating extremely high predictive accuracy, far exceeding existing technologies.
[0028] 3. This invention reveals the core role of E4F1 in taxane resistance and provides a targeted treatment strategy.
[0029] This invention reveals for the first time the core mechanism of E4F1 in paclitaxel resistance and provides a targeted treatment strategy. Through multi-dimensional experimental verification (bioinformatics, tissue microarray, in vitro functional experiments, animal models), it clarifies that E4F1 is a key driver gene in taxane resistance and tumor progression, and provides experimental evidence for its use as a therapeutic target (such as siRNA sequences and in vitro and in vivo efficacy data), demonstrating clear clinical translational potential.
[0030] While existing Zhao models mention multiple ubiquitin genes, they lack in-depth verification of their functional mechanisms and do not provide targeted therapy strategies. This invention, through tissue microarrays, in vitro cell experiments, and animal models, comprehensively demonstrates that E4F1 is a key gene promoting paclitaxel resistance and metastasis; it provides specific siRNA sequences and in vitro / in vivo experimental data, clarifying the feasibility of E4F1 as a therapeutic target and demonstrating direct translational potential. Attached Figure Description
[0031] Figure 1 This study identifies differentially expressed ubiquitination-related genes used to construct prognostic biomarkers. (a) Heatmap of ubiquitination-related gene expression in patients in the taxane-sensitive group (taxane-S) and the taxane-resistant group (taxane-R); (b) Volcano plot of differentially expressed genes in the two groups (|log2FC|>0.263, *p*<0.05); (c) Univariate Cox regression screening identified 17 ubiquitination-related genes associated with breast cancer prognosis; (d) λ parameter of the LASSO model was determined by 10-fold cross-validation; (e) LASSO regression coefficients of the 17 ubiquitination-related genes.
[0032] Figure 2This represents the prognostic efficacy of ubiquitination-related biomarkers in the training and validation cohorts. (ab) Risk scores and survival status distribution for each patient in the TCGA and GSE25055 cohorts; (cd) KM survival curves for low-risk and high-risk patients in the TCGA and GSE25055 cohorts; (ef) ROC curves predicting 1-year, 3-year, and 5-year overall survival (OS).
[0033] Figure 3 This section describes the construction and validation of a nomogram model for breast cancer. (a) Univariate Cox regression analysis of risk score and clinical characteristics; (b) Multivariate Cox regression analysis of risk score and clinical characteristics; (c) Nomogram: integrating risk score and clinicopathological variables; (d) Kaplan-Meier survival curves for nomogram grouping; (e) ROC curves predicting 1 / 3 / 5-year overall survival (OS); (fh) Calibration curves of nomograms predicting 1 / 3 / 5-year OS.
[0034] Figure 4 This is an analysis of drug sensitivity and clinical relevance in breast cancer patients. (ac) Half-maximal inhibitory concentration (IC50) in high- and low-risk groups. 50 (d) Risk score distribution between sensitive and resistant patients; (e) Risk score distribution between patients with different TNM stages; (f) Risk score distribution between patients with different T stages.
[0035] Figure 5 This study examines the differences in E4F1 expression and its prognostic value between normal and breast cancer tissues in public databases. (a) TCGA database: E4F1 mRNA expression level in breast cancer tissues is higher than in normal tissues; (b) HPA database: E4F1 protein expression level in breast cancer tissues is higher than in normal tissues; (c) TCGA database: Breast cancer patients with high E4F1 expression have a worse prognosis.
[0036] Figure 6 The expression of E4F1 and its prognostic value were validated through tissue microarray analysis. (ab) Immunohistochemical (IHC) images: expression of E4F1 and relative histochemical scores in breast cancer tumor tissue and adjacent normal tissue; (cd) Survival analysis: overall survival (OS) and disease-free survival (DFS) of patients with different E4F1 expression levels.
[0037] Figure 7This study validated the in vitro biological function of E4F1. (a) Western blot analysis of E4F1 protein expression levels in breast cancer cells (SK-BR-3, MCF-7, BT-474, BT-549, MDA-MB-231) and human breast epithelial cells (MCF-10A); (b) E4F1 knockout efficiency in MDA-MB-231 cells 48 hours after siRNA transfection (Western blot); (cd) Effect of E4F1 knockout combined with paclitaxel (PTX) treatment on cell migration / invasion ability.
[0038] Figure 8 This is to verify the biological function of E4F1 in vivo. (a) Bioluminescence imaging of tumor-bearing mice on day 16 and day 40; (b) Quantitative analysis of tumor bioluminescence in mice in different treatment groups on day 40; (c) HE-stained images of tumor sections in mice in different treatment groups on day 40; (d) Immunohistochemical (IHC) images of E4F1 expression in tumor sections in mice in different treatment groups on day 40. Detailed Implementation
[0039] This invention constructs a paclitaxel resistance prediction and treatment decision-making system for breast cancer, comprising three major technical modules:
[0040] (1) Construction and validation of a prognostic model for breast cancer based on ubiquitination-related genes of taxane resistance: Extracting samples from public databases → screening differentially expressed genes → constructing a prognostic model;
[0041] (2) Construction and validation of nomogram model based on risk score and TNM staging: Screening independent prognostic factors → Integrating risk score and clinical staging to construct nomogram model → Validating the performance of nomogram model;
[0042] (3) Drug sensitivity analysis and clinical relevance analysis based on prognostic models;
[0043] (4) Functional validation of E4F1 as a therapeutic target: clinical evidence (public database / tissue microarray) → in vitro and in vivo functional validation.
[0044] Key points of the technical solution of this invention:
[0045] 1. A pioneering combination of molecular markers
[0046] This invention, for the first time, screened differentially expressed ubiquitinated genes associated with prognosis based on the treatment response (sensitivity vs. resistance) of taxane drugs. A prognostic model consisting of four ubiquitination-related genes—CBLL1 (Gene ID: 79872), RNF13 (Gene ID: 11342), TRIM59 (Gene ID: 286827), and E4F1 (Gene ID: 1877)—was constructed and validated. The risk score formula is: Risk score = (E4F1 expression level × 0.040) + (CBLL1 expression level × -1.194) + (RNF13 expression level × -0.777) + (TRIM59 expression level × -0.826), where expression level refers to the gene expression level. This model demonstrated extremely high predictive accuracy, achieving AUC values of 0.896, 0.866, and 0.893 for 1-year, 3-year, and 5-year survival rates, respectively, on the training set (TCGA).
[0047] 2. Clinically translatable integrated predictive and decision-making systems
[0048] A nomogram model integrating genetic risk scores and TNM clinical staging improved the 5-year survival prediction AUC to 0.897, significantly outperforming traditional clinical models. Furthermore, drug sensitivity analysis based on the model showed significant differences in sensitivity to different antitumor drugs between high- and low-risk groups, indicating that this model can not only achieve individualized survival prediction and accurate identification of high-risk patients, but may also provide important evidence for the precise selection of clinical chemotherapy regimens.
[0049] 3. Drug sensitivity analysis and treatment recommendations based on risk stratification
[0050] This invention utilizes a risk scoring model to predict a patient's sensitivity to specific drugs. Analysis using the GDSC database and pRRophetic package revealed differences in sensitivity to 73 drugs across high and low risk groups. This allows for targeted treatment recommendations for high-risk patients.
[0051] 4. E4F1 as a key driver of paclitaxel resistance and a novel therapeutic target
[0052] Through bioinformatics, tissue microarray (IHC), in vitro cell function experiments (migration / invasion), and in vivo animal models, this study comprehensively demonstrated for the first time that E4F1 is a key gene promoting paclitaxel resistance and migration / invasion in breast cancer. It also verified that targeted inhibition of E4F1 (such as using siRNA) can significantly enhance the efficacy of paclitaxel, demonstrating clear therapeutic translational potential.
[0053] Any alternative solutions that achieve the following core functions are within the scope of protection of this invention:
[0054] 1. Substitution of gene combinations
[0055] The present invention preferably uses four genes, E4F1, CBLL1, RNF13, and TRIM59, to construct a prognostic model, but it can also be extended to include combinations of other ubiquitinated genes (such as UBA7, USP1, RAD18, etc.) that are related to taxane drug sensitivity or breast cancer prognosis, as long as they can achieve similar prognostic prediction or chemotherapy sensitivity assessment functions.
[0056] 2. Algorithm Model Replacement
[0057] This invention uses LASSO-Cox regression to screen key genes and construct a risk scoring model, but other machine learning or statistical methods (such as random forest, support vector machine, Cox proportional hazards model, neural network, etc.) can also be used for gene screening and model construction, as long as they can achieve similar predictive accuracy.
[0058] 3. Replacement of testing platforms
[0059] This invention is based on mRNA expression level for gene detection, but similar functions can also be achieved through protein level detection (such as immunohistochemistry (IHC), Western blot, ELISA, etc.) or epigenetic level (such as methylation status). The detection platform is not limited to microarrays or RNA-seq, but may also include qPCR, nanostring technology, liquid biopsy (such as ctDNA detection), etc.
[0060] 4. Sample type substitution
[0061] This invention uses tumor tissue samples for detection, but it can also be extended to non-invasive sample types such as blood, saliva, and pleural effusion. By detecting the expression or variation of relevant genes in these samples, similar prognostic or drug sensitivity prediction functions can be achieved.
[0062] 5. Alternatives to targeted therapy strategies
[0063] This invention validates the potential of E4F1 as a therapeutic target by knocking down E4F1 with siRNA. However, other gene silencing or inhibition methods (such as shRNA, CRISPR / Cas9, small molecule inhibitors, antibody drugs, etc.) can also be used to target E4F1 or other core genes (such as CBLL1, RNF13, TRIM59) to achieve the effect of reversing taxane resistance or inhibiting tumor progression.
[0064] 6. Alternatives to nomogram building factors
[0065] This invention integrates risk scores and TNM staging to construct a nomogram, and can also incorporate other clinicopathological factors (such as age, tumor grade, hormone receptor status, Ki-67 index, etc.) to further improve prediction accuracy or adapt to different clinical scenarios.
[0066] 7. Alternatives to drug sensitivity analysis methods
[0067] This invention uses the GDSC database and the pRRophetic package for drug sensitivity prediction, but other databases (such as CTD) can also be used. 2 Validate or expand upon using methods such as CCLE (commonly known as in vitro drug sensitivity testing, organoid models, etc.) or experimental methods.
[0068] 8. Alternative System Implementation Methods
[0069] The prediction system of this invention can be implemented through local software, cloud platform, mobile terminal or embedded device. The data processing flow can be completed based on programming languages such as R, Python, Java or commercial analysis tools (such as SPSS, SAS), as long as they can achieve the same gene expression analysis, risk score calculation and prognosis output functions.
[0070] The following embodiments are further illustrations of the present invention, but not limitations thereof.
[0071] Example 1: Construction and validation of a prognostic model for breast cancer based on taxane resistance ubiquitination-related genes
[0072] Step 1: Data Acquisition and Preprocessing
[0073] Gene expression level data and corresponding clinical information of breast cancer samples were downloaded from the TCGA database (https: / / xenabrowser.net / ). Samples treated with taxanes (including Docetaxel, Taxol, TAXOTERE, Paclitaxel, Abraxane, Tesetaxel, Doxetaxel, Paclitaxelum, Paclitaxel (Protein-Bound), and Taxane) with prognostic information were selected, resulting in 133 breast cancer samples as the training set. These 133 samples were grouped according to "measure_of_response," with Sensitive, Complete Response, and Partial Response being the sensitive group (Sensitive), and Clinical Progressive Disease and Stable Disease being the resistant group (Resistant). This resulted in 119 samples in the sensitive group and 14 samples in the resistant group (e.g., ...). Figure 1 (as shown in a).
[0074] The dataset GSE25055 was downloaded from the NCBI GEO (Gene Expression Omnibus, GEO, http: / / www.ncbi.nlm.nih.gov / geo / ) database. The sequencing platforms were GPL96[HG-U133A]Affymetrix HumanGe nome U133A Array, the species was Homo sapiens, and it contained 310 tumor samples as a validation set.
[0075] Step 2: Screening for differentially expressed ubiquitination-related genes
[0076] Clustering heatmap of differentially expressed genes between the sensitive and resistant groups in TCGA is shown below. Figure 1 As shown in Figure a, 1357 ubiquitination-related genes were obtained from the iUUCD database (integrated annotations for Ubiquitin and Ubiquitin-like Conjugation Database, ver 2.0, http: / / iuucd.biocuckoo.org). The intersection of these genes with the 19938 genes from TCGA in step 1 yielded 1318 ubiquitination-related genes. Based on the sample information, the samples were divided into sensitive and resistant groups. Differential analysis was performed using the limma package (Version 3.52.4, https: / / bioconductor.org / packages / release / bioc / html / limma.html). A p-value < 0.05 and |log2FC| > 0.263 were selected as the threshold for screening differentially expressed genes. A total of 59 ubiquitination-related genes were identified, of which 48 genes were upregulated, including DTL, USP1, IFIH1, CHAF1B, RNF213, NUP153, BARD1, NCF2, MAP3K21, RFWD3, FBXO28, UCHL5, RAD23B, USP13, FANCD2, FBXO45, TAF5L, RNF20, AHCTF1, UBE2O, RBBP5, POLH, USP37, and WDR. 76. DHX36, UBQLN1, RANBP2, KDM2A, MARK2, PRPF4, CBLL1, USP36, DCAF6, WDR3, PIK3R4, NPLOC4, MSL2, RNF168, CRKL, TRIM59, SEH1L, SPRTN, RNF13, SF3A1, RAD18, ILRUN, GGA2, MALT1.
[0077] Eleven genes showed downregulated expression, including MINDY3, ZBTB47, BTBD19, CORO2B, E4F1, HE CTD2, LDB2, BCL6B, HIC1, WDR86, and RNF39 (e.g., Figure 1 (as shown in b).
[0078] Step 3: Screening of genes related to breast cancer prognosis and model construction
[0079] Based on the 59 differentially expressed genes identified in step 2, univariate Cox regression analysis was performed using the R language's `survival` package. Using p < 0.05 as the threshold, 17 genes with significant prognostic correlations were identified (CBLL1, RNF13, MSL2, RAD23B, UBQLN1, PIK3R4, KDM2A, DHX36, RNF20, RNF168, TRIM59, FBXO28, RAD18, RANBP2, BARD1, E4F1, FBXO45, etc.). Figure 1 (as shown in c).
[0080] Based on the expression data of the 17 prognostic genes mentioned above, survival regression analysis was performed using the LASSO-Cox algorithm in the R language glmnet package (Version 1.2, https: / / cran.r-project.org / web / packages / glmnet / index.html). Through 10-fold cross-validation, lambda.min was selected to identify four key genes (CBLL1, RNF13, TRIM59, and E4F1) and their corresponding regression coefficients (-1.194, -0.777, -0.826, and 0.040). A risk score formula was established based on the regression coefficients of each gene, resulting in: Risk Score = E4F1 expression level × 0.040 - CBLL1 expression level × 1.194 - RNF13 expression level × 0.777 - TRIM59 expression level × 0.826 (e.g., E4F1 expression level × 0.040 - CBLL1 expression level × 1.194 - RNF13 expression level × 0.777 - TRIM59 expression level × 0.826). Figure 1 As shown in d-1e), the expression level is the gene expression level.
[0081] Step 4: Model Validation and Performance Evaluation
[0082] (1) Risk scores were calculated for samples in both the TCGA training set and the GEO validation set. Samples with scores higher than the median were included in the high-risk group, and samples with scores lower than the median were included in the low-risk group. The median risk score in the training set was -7.733856, containing 67 high-risk samples and 66 low-risk samples. The median risk score in the validation set was -21.85604, containing 155 high-risk samples and 155 low-risk samples.
[0083] (2) Using R language, the risk score distribution and survival status of each sample were plotted. The results showed that the overall survival time of the high-risk group in the TCGA training set and GEO validation set was shorter than that of the low-risk group (e.g., Figure 2 (as shown in a-2b).
[0084] (3) Kaplan-Meier curves were plotted using the R language's `survival` package on the TCGA training set and the GEO validation set to assess the correlation between risk scores and actual patient prognostic information. Results from both the training and validation sets showed that the high-risk group had shorter overall survival and worse prognosis (e.g., ...). Figure 2 (as shown in c-2d).
[0085] (4) ROC curves were plotted using R language, and the area under the curve (AUC) was used to quantify the predictive efficacy of the prognostic model for 1-year, 3-year, and 5-year survival rates. The results showed that the AUC values for predicting 1-year, 3-year, and 5-year overall survival rates in the training set were 0.896, 0.866, and 0.893, respectively, while those in the validation set were 0.62, 0.624, and 0.618, respectively. These data confirm that the prognostic model has high predictive accuracy (e.g., ...). Figure 2 (as shown in e-2f).
[0086] Example 2: Construction and Validation of a Nomogram Model Based on Risk Scoring and TNM Staging
[0087] Step 1: Screening for independent prognostic factors
[0088] Clinical information from all TCGA samples was compiled. Using the R language's `survival` package, univariate and multivariate Cox regression analyses were performed on clinical factors (including age, race, and TNM stage) and risk scores to identify independent prognostic factors, and forest plots were generated. Both univariate and multivariate Cox regression analyses showed that risk score and TNM stage were independent prognostic factors (e.g., age, race, TNM stage). Figure 3 (as shown in a-3b).
[0089] Step 2: Construct the Column Chart Model
[0090] To achieve more accurate prediction of patient prognosis and provide a more precise basis for clinical decision-making, a nomogram was constructed using the rms package in R (Version 5.1-2, https: / / cran.r-project.org / web / packages / rms / index.html) based on the selected independent prognostic factors (risk score and TNM stage). Figure 3c. Nomograms constructed based on RiskScore and TNM staging can achieve individualized survival prediction. For example, when a patient's RiskScore = -5 (score 70) and TNM stage = 3 (score approximately 35), the total score is approximately 105, and the predicted 1-year, 3-year, and 5-year survival rates are approximately 95%, 70%, and 60%, respectively. It is worth noting that a total score >140 indicates high risk (1-year survival rate <50%).
[0091] Step 3: Validate the prognostic performance of the nomogram model from multiple dimensions.
[0092] (1) Using the survival package in R, Kaplan-Meier curves were plotted. The results showed that the overall survival of patients in the low-risk group was significantly prolonged, consistent with the Kaplan-Meier curve results in the training and validation sets (e.g., ...). Figure 3 (as shown in d).
[0093] (2) ROC curves were plotted using R language, and the area under the curve (AUC) was used to quantify the predictive power of the prognostic model for 1-year, 3-year, and 5-year survival rates. The AUC values for predicting 1-year, 3-year, and 5-year overall survival rates reached 0.992, 0.89, and 0.897, respectively, indicating that the model has good prognostic predictive performance (e.g., Figure 3 (as shown in e).
[0094] (3) The nomogram survival calibration curve was plotted using R language. The results showed that the survival rate of breast cancer patients predicted by the nomogram was highly consistent with the actual survival rate. The above results verified the predictive accuracy of the nomogram model (e.g., Figure 3 (as shown in f-3h).
[0095] Example 3: Model-based analysis of drug sensitivity and clinical relevance
[0096] To investigate differences in drug sensitivity between low-risk and high-risk groups, we estimated the sensitivity of each patient to 138 drugs based on the Genomics of Cancer Drug Sensitivity (GDSC) database, and analyzed the IC50 using the pRRophetic package in R. 50 (Half maximal inhibitory concentration) was quantified. Results showed that the IC50 of 73 drugs differed between the high-risk and low-risk groups. 50 There were significant differences (P < 0.05) (as shown in Table 1), with the three drugs showing the most statistically significant differences being Bryostatin.1, PHA.665752, and Salubrinal (e.g. Figure 4(As shown in a-4c). The three drugs showed significantly higher sensitivity in the high-risk group than in the low-risk group, suggesting that these drugs may be effective in treating patients with high-risk scores. These results reveal a significant correlation between breast cancer risk typing and drug sensitivity, providing important evidence for personalized clinical medication.
[0097] Furthermore, by integrating clinical information data, we analyzed the distribution of risk scores across different clinical characteristics. The results showed that the risk scores of patients in the drug-resistant group were significantly higher than those of sensitive patients (e.g., ...). Figure 4 (As shown in d). Patients with TNM stage III had significantly higher risk scores than those with stage I-II (e.g., ...). Figure 4 (as shown in e). Similarly, patients with pathological stage T3 had significantly higher risk scores than patients with stage T1-T2 (e.g., ...). Figure 4 (as shown in f). These results indicate that the risk score is significantly correlated with drug sensitivity and tumor progression.
[0098] Table 1. IC in high-risk and low-risk groups 50 Different drugs
[0099]
[0100]
[0101] Example 4: Functional Mechanism and Clinical Prognostic Value of E4F1 in Paclitaxel Resistance in Breast Cancer
[0102] Step 1: Validate the expression and prognostic value of E4F1 in public databases
[0103] (1) Download BRCA gene expression level data from the TCGA database (https: / / xenabrowser.net / ). Based on the clinical information of samples downloaded at the same time, samples with a survival time of less than 30 days were excluded, resulting in a primary tumor group of 1054 cases and a solid tissue normal group of 111 cases. The expression levels of E4F1 mRNA in the two groups were compared, and the results are as follows: Figure 5 The results showed that the mRNA level of E4F1 in the breast cancer group was significantly higher than that in the normal group.
[0104] (2) The differences in E4F1 protein levels between the breast cancer group and the normal control group were compared using the Human Protein Atlas database (HPA) (https: / / www.proteinatlas.org / ). The results are as follows: Figure 5 b shows that the protein expression level of E4F1 in the breast cancer group was higher than that in the normal group.
[0105] (3) Based on the clinical information of the samples obtained in step 1, Kaplan-Meier curves were plotted using the survival package in R language to evaluate the survival prognosis of E4F1 in the tumor sample group. The optimal cutpoint for gene expression (calculated using the survminer package by entering the surv_cutpoint command) was used as the basis for high and low grouping. The results are as follows: Figure 5 c shows that breast cancer patients with high E4F1 expression have a worse prognosis.
[0106] Step 2: Histological experiments to verify the clinical significance of E4F1
[0107] This step was completed by Shanghai Xinchao Biotechnology Co., Ltd., and the steps are as follows:
[0108] Breast cancer tissue microarrays (product numbers TFBrec-01 and TFBrec-02) provided by Shanghai Xinchao Biotechnology Co., Ltd. were used. TFBrec-01 contained 143 breast cancer tissue samples and 18 adjacent normal tissue samples, while TFBrec-02 contained 67 adjacent normal tissue samples. Immunohistochemical staining was performed using Sigma Aldrich's E4F1 primary antibody (catalog number HPA072309, volume dilution 1:500). The specific immunohistochemical staining procedure was as follows: the breast cancer tissue microarrays were placed in an oven (manufacturer: Shanghai Yiheng Scientific Instruments Co., Ltd.; model: PH-070A) at 63℃ for one hour. Then, dewaxing was completed in a fully automated staining machine (manufacturer: LEICA; model: LEICAST5020). The dewaxing times were as follows: (1) 2 xylene tanks, 15 minutes each; (2) 2 100% alcohol tanks, 7 minutes each; (3) 1 90% alcohol tank, 5 minutes; (4) 1 80% alcohol tank, 5 minutes; (5) 1 70% alcohol tank, 5 minutes. According to the "PTLink Concise Operating Procedures", the slides were placed in the antigen retrieval solution. After retrieval, the slides were placed in distilled water at room temperature and allowed to cool naturally for more than 10 minutes. The chip was removed and rinsed with PBST buffer. The diluted primary antibody working solution (1:500 volume ratio, Sigma Aldrich, catalog number #HPA072309) was added and incubated overnight at 4°C. The slides were removed from the refrigerator, warmed to room temperature for 45 minutes, and washed with PBST buffer. Place the slides into the DAKO Autostainer Link 48 fully automated immunohistochemical staining system. Use the immunochromogenic reagent (Dako; catalog number: K8002), and follow the "Autostainer Link 48 User Guide" to select the appropriate program for blocking, incubating at room temperature for 20 minutes to allow the E4F1 primary antibody to bind to the ready-to-use peroxidase dextran polymer and goat anti-rabbit secondary antibody molecules, followed by the 3,3'-diaminobenzidine (DAB) staining program. Stain with hematoxylin (Gene, catalog number GT100540) for 1 minute, then immerse in 0.25% hydrochloric acid-alcohol solution (400ml 70% ethanol + 1ml concentrated hydrochloric acid) for 5 seconds, and rinse with tap water for 3 minutes. Air dry the slides at room temperature and mount them with neutral resin. Finally, scan the prepared slides using an Aperio scanner (LEICA; catalog number: Aperio XT) to obtain immunohistochemical images. Data collection was conducted under a strict double-blind principle to ensure objectivity. The scoring criteria for the percentage of positive cells were: 0 (no positive cells), 1 (≤25% positive cells), 2 (26-50%), 3 (51-75%), and 4 (>75%). The scoring criteria for staining intensity were: 0 (colorless), 1 (light brown), 2 (brown), and 3 (dark brown).Three independent fields of view were randomly selected for each sample. The single-field score was calculated using the formula "staining index = positive cell proportion score × staining intensity score". The final score of the three fields of view was taken as the sample staining index. The results of 140 tumor tissues and 61 adjacent normal tissues were included in the statistics.
[0109] Results: The expression level of E4F1 protein in breast cancer tumor tissues was significantly higher than that in adjacent normal tissues (e.g., breast cancer tumor tissues). Figure 6 (as shown in a-6b). Kaplan-Meier survival analysis showed that high E4F1 expression was significantly associated with shorter overall survival (OS) and disease-free survival (DFS) (e.g., ...). Figure 6 (as shown in c-6d). Univariate analysis revealed an association between OS and E4F1 expression, tumor size, N stage, M stage, and TNM stage; multivariate analysis further confirmed that E4F1 expression, M stage, and TNM stage are independent prognostic factors affecting patient survival (Table 2).
[0110] Table 2. Univariate and multivariate analyses of factors related to overall survival in breast cancer patients.
[0111]
[0112] Univariate analysis showed that the difference was statistically significant.
[0113] b. Multivariate analysis showed that the difference was statistically significant.
[0114] Step 3: In vitro functional experiments: E4F1 promotes breast cancer invasion and paclitaxel resistance
[0115] (1) The expression level of E4F1 protein in breast cancer cells SK-BR-3 (Catalog No. HTB-30), MCF-7 (Catalog No. HTB-22), BT-474 (Catalog No. HTB-20), BT-549 (Catalog No. HTB-122), MDA-MB-231 (Catalog No. HTB-26) and human breast epithelial cells MCF-10A (Catalog No. CRL-10317) derived from ATCC (https: / / www.atcc.org / ) was detected by Western blot. The cell culture conditions were as follows: non-tumorigenic breast cells MCF-10A were cultured in DME M-F12 medium (Gibco, 11320033) with the addition of 0.5 μg / mL hydrocortisone (MCE, HY-N0583), 10 μg / mL insulin (Sigma-Aldrich, Catalog No. I2643), and 20 ng / mL human epidermal growth factor (hEGF) (Sigma-Aldrich, Catalog No. I2643). ma-Aldrich (catalog number E5036) and 10% fetal bovine serum (Gibco, catalog number 10091148) were used to culture SK-BR-3, MCF-7, BT-549, and MDA-MB-231 cells in Duchenne Modified Eagle Medium (DMEM, Gibco, catalog number 11965092) supplemented with 10% fetal bovine serum and 1% penicillin-streptomycin antibiotics (Gibco, catalog number 15070063). BT-474 cells were cultured in RPMI 1640 medium (Gibco, catalog number 11875093) supplemented with an equal volume of serum and antibiotics. All cells were routinely cultured in a 37°C, 5% CO2 incubator. The specific procedure for Western blot is as follows: After washing cells with pre-cooled PBS, lysis was performed using RIPA lysis buffer (Beyotime, catalog number P0013B). Cells were collected in 1.5 mL EP tubes using a cell scraper. After sonication (parameters set to 30% amplitude, 3 seconds sonication, 10-second interval, repeated for 5 cycles), the cells were incubated on ice for 10 min, centrifuged at 12000g for 30 min at 4°C, and the supernatant was transferred to a new EP tube to obtain the protein sample. An equal volume of the protein sample was separated by SDS-PAGE electrophoresis and then transferred to a PVDF membrane (Merck Millipore, 0.45 μm). After being blocked with 5% skim milk powder, the transfer membranes were incubated overnight at 4°C with E4F1 primary antibody (1:1,000 v / v, Sigma Aldrich, catalog number HPA072309) and β-tubulin internal control antibody (1:10,000 v / v, CST, catalog number 15115). Subsequently, the membranes were incubated at room temperature for 2 hours with horseradish peroxidase-conjugated goat anti-rabbit secondary antibody (1:10,000 v / v, Jackson, catalog number 111-035-003).Finally, colorimetric analysis was performed using a chemiluminescent HRP substrate kit (Merck Millipore, catalog number WBULP-100ML) with ImageQuant LAS 4000. TM The imaging system (GE Healthcare) acquires signals. For example... Figure 7 As shown in figure a, the expression level of E4F1 protein in breast cancer cells SK-BR-3, BT-549, and MDA-MB-231 was significantly higher than that in human breast epithelial cells MCF-10A.
[0116] (2) MDA-MB-231 cells were transfected with siRNA to knock down the expression of the E4F1 gene, and the knockdown effect was verified by Western blot. The siRNA transfection steps were as follows: After MDA-MB-231 cells were seeded to an appropriate density (30-50% confluence), Lipofectamine was used. TM Thermo Fisher Scientific (Catalog No. L3000015) transfected negative control siRNA (si-NC, Ribobio, Catalog No. siN0000001-1-5) and E4F1-targeting siRNA (si-E4F1, Ribobio, Catalog No. stB0006201A-C) with 3000 transfection reagent. Transfection was performed strictly according to the manufacturer's instructions, and the cells were incubated at 37°C with 5% CO2 for 48 hours after transfection. The E4F1 siRNA sequence design is as follows:
[0117] si-E4F1 1#:5′-TGAAGCTACTGGTGAACAA-3′
[0118] si-E4F1 2#:5′-GCAAGCGCTACAAGACTAA-3′
[0119] si-E4F1 3#: 5′-GGACAAGTGGCACTGAACA-3′
[0120] The Western blot method is as described above.
[0121] like Figure 7 As shown in b, si-E4F11#, si-E4F12# and si-E4F13# can significantly reduce the expression level of E4F1 protein in MDA-MB-231 cells.
[0122] (3) In MDA-MB-231 cells, the effect of knocked-down E4F1 gene on the invasion and migration functions of breast cancer was studied. The invasion and migration abilities of breast cancer cells were determined by transwell assay. The specific operation was as follows: the migration experiment used bare membrane chambers without Matrigel coating (Merck Millipore, catalog number PTMP24H48), and the invasion experiment used Matrigel coated chambers (Corning, catalog number CB-40234C). After transfecting MDA-MB-231 cells with si-E4F11# or si-NC for 48 hours, the cells were collected by trypsin digestion and resuspended in serum-free medium (DMEM, Gibco, catalog number 11965092) containing 1 μM paclitaxel (MedChemExpress, catalog number HY-B0015) at a concentration of 2×10⁻⁶. 4 Cells were seeded at a density per well into the upper chamber of the microcell. Serum-containing complete culture medium, prepared by mixing serum-free medium with fetal bovine serum (DMEM, Gibco, catalog number 11965092) at a volume ratio of 9:1, was added to the lower chamber. After 24 hours of incubation (37°C, 5% CO2), cells that migrated / invaded to the lower membrane surface were fixed with 4% paraformaldehyde (Beyotime, P0099), stained with 0.1% crystal violet, and counted and analyzed under an Olympus IX71 inverted microscope.
[0123] The results showed that knocking down the E4F1 gene significantly inhibited the migration and invasion abilities of MDA-MB-231 cells. After paclitaxel treatment, the migration and invasion abilities of cells in the E4F1-silenced group further deteriorated, indicating that inhibiting E4F1 enhances the cytotoxic effects of paclitaxel (e.g., ...). Figure 7 (as shown in c-7d).
[0124] Step 4: Validate the tumor-promoting effect and therapeutic potential of E4F1 using animal models.
[0125] A xenograft tumor model was established using BALB / c-nu female mice (4 weeks old, purchased from Beijing Vital River Laboratory Animal Technology Co., Ltd.). Mice were housed in CP-4 cages (290×178×160mm, 5 mice / cage) under controlled environmental conditions of 24±2℃, 40-60% humidity, and a 12-hour light-dark cycle, with free access to food and water. After one week of acclimatization, the model was established by tail vein injection of 4×10⁶ MDA-MB-231 cells. Sixteen days later, 20 mice were randomly divided into four groups (n=5): (1) in vivo si-control group (Raybot Biotech, catalog number siN0000005-4-10); (2) paclitaxel group (MedChemExpress, catalog number HY-B0015); (3) si-E4F1 group (Raybot Biotech, catalog number siBDMV002, sequence 5'-TGAAGCTACTGGT GAACAA-3'); (4) paclitaxel + si-E4F1 combination group. Paclitaxel (10 mg / kg per injection) and in vivo siRNA (5 nmol / time) were both administered via tail vein injection every two days. Mice were treated for 10 consecutive times. On days 16 and 40, mice were intraperitoneally injected with D-fluorescein (Yisheng Biotechnology, catalog number 40902ES02, 150 mg / kg). Whole-mouse fluorescence images were acquired using the IVIS Lumina in vivo imaging system (Xenogen, USA). After imaging on day 40, mice were euthanized, and mammary tissue was harvested, fixed in 10% formalin for 48 hours, embedded in paraffin, and then subjected to HE staining and immunohistochemical analysis. This experiment was approved by the Animal Experiment Ethics Committee of Sun Yat-sen University (SYSU-IACUC-2024-001996), and all procedures complied with relevant animal research regulations.
[0126] The immunohistochemistry procedure is shown in step 2.
[0127] The specific procedure for HE staining is as follows: Mouse tumor tissue was embedded in paraffin and then continuously sectioned, dewaxed, and hydrated. The sections were immersed in hematoxylin staining solution (Zhongshan Jinqiao, catalog number ZLI-9610) for 5 minutes, rinsed with running water, and then transferred to eosin staining solution (Zhongshan Jinqiao, catalog number ZLI-9613) for 2 minutes. After dehydration with graded alcohols and clearing with xylene, the sections were mounted with neutral resin and finally observed under an optical microscope (Olympus IX71, Tokyo, Japan).
[0128] The results showed that the tumor volume in the si-E4F1 group was significantly smaller than that in the si-control group, and the tumor volume in the si-E4F1 group combined with paclitaxel treatment was further reduced compared with the paclitaxel-only group (e.g., Figure 8 a-8c). Immunohistochemical analysis of tumor tissue (e.g.) Figure 8As shown in d), the combined treatment with si-E4F1 and paclitaxel confirmed that the expression level of E4F1 decreased significantly.
[0129] SEQ ID NO.1
[0130] TGAAGCTACTGGTGAACAA
[0131] SEQ ID NO.2
[0132] GCAAGCGCTACAAGACTAA
[0133] SEQ ID NO.3
[0134] GGACAAGTGGCACTGAACA
Claims
1. The application of an agent that inhibits E4F1 gene expression and paclitaxel in the preparation of an anti-breast cancer drug, wherein the agent that inhibits E4F1 gene expression is siRNA, and the siRNA sequence is 5′-TGAAGCTACTGGTGAACAA-3′, 5′-GCAAGCGCTACAAGACTAA-3′ and / or 5′-GGACAAGTGGCACTGAACA-3′.
2. The application according to claim 1, characterized in that, The aforementioned anti-breast cancer drug is a drug that inhibits the invasion or migration of breast cancer cells.
3. The application according to claim 2, characterized in that, The breast cancer cells were MDA-MB-231 cells.
4. The application according to claim 1, characterized in that, The siRNA enhances the cytotoxicity of paclitaxel or increases the sensitivity of breast cancer cells to paclitaxel.