Method for analyzing action mechanism of plasticizer acetyl tri-n-butyl citrate for inducing breast cancer based on network toxicology

By combining network toxicology and machine learning methods, this study systematically elucidates the molecular association mechanism of acetyl tributyl citrate (ATBC) inducing breast cancer, identifies MAOA and ADRA2A as core targets, solves the problem of unclear target identification in existing studies, and realizes the establishment of a breast cancer prognostic prediction model and the functional validation of the targets.

CN121963949APending Publication Date: 2026-05-01HUAIYIN INSTITUTE OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUAIYIN INSTITUTE OF TECHNOLOGY
Filing Date
2026-01-14
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

The mechanism by which acetyl tributyl citrate (ATBC) induces breast cancer is unclear in existing studies. There is a lack of a target screening system that integrates network toxicology and machine learning, resulting in incomplete target identification, insufficient screening accuracy, lack of molecular-level validation, and an inability to systematically analyze its association with breast cancer.

Method used

Using a combination of network toxicology and machine learning, we obtained ATBC-related targets, performed differentially expressed gene analysis and weighted gene co-expression network analysis for breast cancer, constructed a protein-protein interaction network, conducted functional and pathway enrichment analysis, and used various machine learning algorithms to screen core targets. Finally, we verified their binding activity through molecular docking.

Benefits of technology

The system revealed the molecular association mechanism of ATBC-induced breast cancer, identified MAOA and ADRA2A as core targets, improved the comprehensiveness and specificity of target screening, provided a prognostic prediction model for breast cancer, and verified the functional association of targets through molecular docking, thus enhancing the scientific validity and persuasiveness of the research results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The invention discloses a method for analyzing an action mechanism of a plasticizer acetyl tri-n-butyl citrate (ATBC) for inducing breast cancer based on network toxicology. The method comprises the following steps: firstly, integrating multiple databases to obtain and standardize ATBC targets, and identifying disease-related genes in combination with differential expression of TCGA data and weighted gene co-expression network analysis (WGCNA); overlapping target spots are obtained through intersection of the three, a protein interaction network is constructed, and function enrichment analysis is carried out. TCGA is used as a training set, GEO is used as a verification set, random forest and Lasso regression are combined to screen out core targets MAOA and ADRA2A, and a prognosis model is constructed. And finally verifying that the ATBC can be stably combined with the two target spots through molecular docking (the combination energy is 1t;-5.0 kcal / mol). According to the invention, a full-process scheme from target prediction, function analysis, machine learning screening to molecular docking verification is established, and a standardized normal form is provided for the study of the carcinogenic mechanism of environmental chemicals.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of bioinformatics and disease risk assessment technology, specifically relating to a method for analyzing the mechanism of action of the plasticizer tributyl acetyl citrate (ATBC) in inducing breast cancer based on network toxicology. Background Technology

[0002] Tributyl acetylglucose citrate (ATBC) is a plasticizer widely used in polyvinyl chloride (PVC) food packaging containers, children's toys, medical products, films, and sheets. Its hydrophobic properties and resistance to environmental degradation make it easy to accumulate in biological and ecological systems, which has led to extensive discussions about its endocrine disruption effects and chronic health hazards. Globally, breast cancer is one of the most common malignant tumors among women. Its pathogenesis is jointly regulated by genetic factors, behavioral patterns, and external environmental factors. Long-term exposure to environmental chemicals, especially plasticizers, has been shown to be potentially associated with an increased risk of breast cancer.

[0003] While research on the association between environmental chemicals and breast cancer has made some progress, several technical shortcomings and research gaps remain: First, existing studies do not clearly elucidate the molecular link between exposure to tributyl acetyl citrate and the development of breast cancer, and the core regulatory targets in the association process have not yet been systematically identified. Second, traditional disease-related target screening methods often rely on single databases or single analysis techniques, resulting in incomplete target coverage and insufficient screening accuracy, failing to achieve a complete analysis of environmental chemical targets, disease targets, and interaction mechanisms. Third, there is a lack of standardized target screening systems that integrate network toxicology, weighted gene co-expression network analysis, and multi-algorithm machine learning models, making it difficult to balance the comprehensiveness of target screening with the specificity of core targets. Fourth, for the selected candidate targets, there is a lack of standardized procedures to verify their direct interaction with tributyl acetyl citrate at the molecular level, leading to insufficient evidence of functional association between the targets.

[0004] Network toxicology, as an interdisciplinary technology integrating computational biology, systems analysis, and cheminformatics, can achieve a panoramic analysis of the interaction network between environmental chemicals and disease targets. Meanwhile, machine learning technology can accurately screen core regulatory factors from a massive number of targets. The combination of the two can effectively make up for the shortcomings of existing technologies. However, there is currently no technical solution to apply the two in combination to the screening of targets for acetyl citrate-induced breast cancer.

[0005] Therefore, there is an urgent need to establish a method based on network toxicology and machine learning to analyze the mechanism of action of tributyl citrate in inducing breast cancer. This is of great significance for solving the problem of unclear mechanism of action of tributyl citrate in inducing breast cancer and lack of core target screening system. Summary of the Invention

[0006] The present invention aims to provide a method for analyzing the mechanism of action of the plasticizer tributyl citrate (ATBC) in inducing breast cancer based on network toxicology, so as to solve the technical problems of unclear mechanism of action of tributyl citrate in inducing breast cancer and lack of core target screening system in related technologies.

[0007] To achieve the above objectives, the technical solution adopted by the present invention is as follows: The first aspect of this invention provides a method for analyzing the mechanism of action of the plasticizer tributyl acetylcitrate in inducing breast cancer based on network toxicology, comprising the following steps: (1) Acquisition and standardization of targets related to tributyl acetyl citrate; (2) Differential gene expression analysis and weighted gene co-expression network analysis in breast cancer; (3) Screening overlapping targets of acetyl tributyl citrate and breast cancer and constructing a protein-protein interaction network; (4) Perform functional and pathway enrichment analysis on overlapping targets; (5) Screening core targets from overlapping targets based on a combination of multiple machine learning algorithms; (6) Verify the binding activity of the core target with tributyl acetyl citrate.

[0008] Further, step (1) above specifically involves: using acetyl tributyl citrate as the search term, obtaining the chemical structure and SMILES information of acetyl tributyl citrate from the PubChem database; using acetyl tributyl citrate as the search term, screening for human-specific potential targets in the ChEMBL database, and simultaneously submitting the SMILES information of acetyl tributyl citrate to the Swiss Target Prediction database and the STITCH database for complementary potential target identification; summarizing all obtained target identifiers, performing standardized mapping using the UniProt database, obtaining a unified UniProt accession number and standard gene symbol, and forming a set of potential human targets for acetyl tributyl citrate.

[0009] Furthermore, step (2) above specifically involves: obtaining transcriptome data of breast cancer tissue and its paired normal breast tissue from the TCGA database; performing differential expression analysis using the limma package in R language, with |log2FC|≥0.585 and adjusted... pA value <0.05 was used as the screening criterion for differentially expressed genes, and a set of differentially expressed genes in breast cancer was obtained. Weighted gene co-expression network analysis was performed on the same transcriptome data, with an optimal soft threshold of 18. A topological overlap matrix was constructed, and a dynamic tree slicing algorithm was used to identify co-expressed gene modules. The correlation between characteristic genes of each module and breast cancer phenotype was calculated, and correlations were screened out. p Value < 0.05 and |correlation coefficient r |>0.36 module, extract all non-repetitive genes within the module to form a set of candidate targets related to breast cancer.

[0010] Further, step (3) above specifically involves: performing an intersection analysis on the set of target points related to tributyl acetyl citrate obtained in step (1) and the set of differentially expressed genes and candidate target points related to breast cancer obtained in step (2) to screen out the set of overlapping target points between tributyl acetyl citrate and breast cancer; inputting the standard gene symbols corresponding to the overlapping target points into the STRING database, setting the species as Homo sapiens, and generating a protein-protein interaction network with a minimum interaction confidence threshold of 0.4; using Cytoscape software to visualize and analyze the topology of the protein-protein interaction network, and visualizing it based on the degree centrality of the nodes.

[0011] Further, step (4) above specifically involves: using the DAVID and FUMA platforms to perform GO functional annotation and KEGG pathway enrichment analysis on the overlapping target set, wherein the significance screening criteria are set as follows: the corrected false detection rate FDR < 0.05, and each enrichment entry contains at least 3 target genes.

[0012] Furthermore, step (5) above specifically involves: using the gene expression profiles of breast cancer samples in the TCGA database and the corresponding clinical survival data as the training set, and using an independent breast cancer gene expression dataset from the GEO database as the external validation set; constructing and evaluating models using combinations of different machine learning algorithms; using the consistency index C-index as the evaluation metric to select combination strategies with a C-index greater than 0.7; using the combination strategies to select core targets from the overlapping target set, and constructing a prognostic prediction model based on the core targets.

[0013] Furthermore, the process of using the combined strategy of random forest initial screening and Lasso regression fine screening to screen core targets from the overlapping target set includes: first, using the random forest model on the training set to perform preliminary dimensionality reduction on the overlapping targets and screen out the candidate key genes most related to survival; second, using Lasso regression on the training set to finely screen these candidate key genes and determine the final core targets.

[0014] Further, step (6) above specifically involves: obtaining the three-dimensional crystal structure of the core target protein from the RCSB PDB database, preprocessing it using PyMOL and AutoDockTools software, and converting it into PDBQT format as a receptor file; obtaining the molecular structure of tributyl acetyl citrate from the PubChem database and converting it into PDBQT format as a ligand file; performing molecular docking using AutoDock Vina software, using a binding energy <-5.0 kcal / mol as the criterion for stable binding, and predicting the binding energy and binding mode; and using PyMOL software to perform visualization analysis on the conformation with the lowest binding energy.

[0015] A second aspect of the present invention provides the application of the core targets screened by the above method in constructing a prognostic prediction model for breast cancer.

[0016] Furthermore, the core targets mentioned above are MAOA and ADRA2A.

[0017] The beneficial effects of this invention are as follows: Compared with traditional methods, this invention innovatively integrates six core links: multi-source target acquisition, weighted gene co-expression network analysis, protein-protein interaction network construction, functional pathway enrichment, multi-algorithm machine learning selection, and molecular docking verification. It forms a complete technology chain of "target screening - functional analysis - core identification - binding verification". For the first time, it systematically reveals the molecular association mechanism of plasticizer ATBC inducing breast cancer, solves the technical bottleneck of unclear identification of core regulatory targets and ambiguous mechanism explanation in existing studies, and fills the technical gap in the study of the association between ATBC and breast cancer.

[0018] Unlike traditional methods that rely on a single or limited set of algorithms, this invention selects algorithms such as random forest, elastic network, and Lasso regression to construct different combined models. Using TCGA as the training set and datasets from the GEO database as the validation set, a two-stage screening strategy—initial screening with random forest and precise optimization with Lasso regression—is employed to ultimately identify the two core targets, MAOA and ADRA2A. This approach avoids the bias of a single algorithm and ensures the reliability and specificity of the screening results through validation on multiple datasets.

[0019] In the breast cancer target screening stage, this invention not only uses differentially expressed gene analysis, but also combines weighted gene co-expression network analysis. A gene co-expression network is constructed with the optimal soft threshold of 18 to screen out core modules that are significantly related to the breast cancer phenotype. Then, overlapping targets are obtained by taking the intersection of Venn diagrams, which significantly improves the correlation between targets and ATBC and breast cancer, providing a more solid data foundation for subsequent screening.

[0020] Compared to traditional methods that rely solely on bioinformatics analysis and lack direct molecular-level verification, this invention, after screening core targets, conducts molecular docking experiments using AutoDock Vina. With a binding energy of <-5.0 kcal / mol as the standard for stable binding, it confirms that ATBC can spontaneously form stable bindings with ADRA2A and MAOA. This provides direct structural biology evidence for the functional association of core targets at the atomic scale, significantly enhancing the scientific rigor and persuasiveness of the research results.

[0021] In summary, this invention systematically elucidates for the first time the potential mechanism of ATBC inducing breast cancer. By combining network toxicology with machine learning, it provides researchers with a new method to explore the potential mechanism of ATBC inducing breast cancer based on the interaction between specific compounds and target proteins, and also provides new research ideas for the etiological prevention and targeted intervention of breast cancer. Attached Figure Description

[0022] Figure 1 : Technical flowchart of the present invention.

[0023] Figure 2 This invention presents a heatmap showing the correlation between gene co-expression modules and breast cancer / normal breast tissue phenotypes. The heatmap illustrates the varying degrees of correlation between nine co-expression gene modules and breast cancer and normal breast tissue phenotypes. Six core modules significantly associated with breast cancer are clearly identified: brown, red, pink, purple, black, and blue-green. "Normal" represents the normal breast tissue group, and "Tumor" represents the breast cancer tissue group.

[0024] Figure 3 This invention presents a heatmap of the top 50 differentially expressed genes in breast cancer. In this figure, "Group" represents a subgroup, with "Normal" representing normal breast tissue and "Tumor" representing breast cancer tissue. The vertical axis represents the gene name, visually demonstrating the difference in expression levels of ATBC-related breast cancer genes between the two groups, showing the upregulation or downregulation trend of gene expression.

[0025] Figure 4 The Venn diagram of overlapping targets in the ATBC potential human target set, the breast cancer differentially expressed gene set, and the breast cancer-related candidate target set in this invention.

[0026] Figure 5 This invention presents a PPI network and target annotation diagram. In this diagram, nodes represent target proteins, and lines represent interactions between proteins. Node size is positively correlated with degree centrality.

[0027] Figure 6: GO enrichment analysis results of overlapping target sets in this invention. This figure encompasses enrichment results across three dimensions: biological processes, cellular components, and molecular functions. Specifically, Process (BP) represents biological processes; negative regulation of transcription by DNA-templated; positive regulation of transcription by DNA polymerase II; methylation; DNA damage response; heterochromatin formation; chromatin organization; protein deacetylation; Cellular Component (CC); nucleoplasm; nucleus; chromatin; histone deacetylase complex; histone-containing complex; transcription repressor complex; pericentric heterochromatin; cytosol; Molecular Function (MF); DNA-binding transcription factor binding; and enzyme. binding: enzyme binding; chromatin binding: chromatin binding; histone methyltransferase activity: histone methyltransferase activity; histone binding: histone binding; histone deacetylase activity: histone deacetylase activity; histone lysine deacetylase activity: histone lysine deacetylase activity; transcription coactivator activity: transcription coactivator activity; transcription corepressor activity: transcription corepressor activity.

[0028] Figure 7This invention presents a KEGG pathway analysis of overlapping target sets. The figure illustrates the distribution and statistical significance of gene products within the cell localization of the studied gene set. The Y-axis represents cellular component entries, including cytosol (cytoplasm), transcription repressor complex, pericentric heterochromatin, protein-containing complex, histone deacetylase complex, chromosome, heterochromatin, chromatin, and nucleus. The X-axis represents -log[…]. 10 (pvalue), representing the negative logarithm of the enrichment significance; the larger the value, the more significant the enrichment result. p The smaller the value, the less likely it is to be random; the size of the point (Count) represents the number of genes enriched in that cell component; the color of the point represents the significance of enrichment in a gradient, from green (low significance) to red (extremely high significance).

[0029] Figure 8 The predictive performance evaluation of machine learning models generated based on different algorithms in this invention (using C-index as the indicator).

[0030] Figure 9 a: A graph showing the relationship between the number of trees and model performance in the RF model, indicating the optimal number of trees; b: A variable importance graph based on the RF algorithm, showing the ranking of genes that contribute the most to the model; c: A LASSO coefficient spectrum, showing the path of the predictor variable coefficients as the regularization parameter (λ) changes; d: A validation graph of the LASSO model tuning parameter (λ) selection, determining the optimal λ value through cross-validation.

[0031] Figure 10 Venn diagram of the intersection of the overlapping target set with key genes in RF and Lasso models.

[0032] Figure 11 a: Comparison of survival curves for high- and low-risk groups based on prognostic models in the training set; b: Time-dependent ROC curves of the prediction accuracy of the models in the training set for 1 year, 3 years, and 5 years.

[0033] Figure 12 a: Comparison of survival curves for high- and low-risk groups based on prognostic models in the validation set; b: Time-dependent ROC curves of the model's prediction accuracy over 1, 3, and 5 years in the validation set. Detailed Implementation

[0034] The technical subject matter of this invention will now be described in conjunction with exemplary embodiments. It should be understood that the discussion of the foregoing embodiments is merely to help those skilled in the art fully understand and implement the technical solutions of this invention. Without departing from the scope of protection defined in this specification, the functional configuration and connection relationships of the technical elements involved can be adjusted or changed. Each exemplary embodiment can omit, replace, or add corresponding process steps or structural components according to actual needs; furthermore, the technical features described in different exemplary embodiments can be combined to form new technical solutions. Example 1

[0035] A method for analyzing the mechanism of action of the plasticizer tributyl acetylcitrate (ATBC) in inducing breast cancer based on network toxicology includes the following steps: (1) Acquisition and standardization of ATBC-related targets: First, using the PubChem database (https: / / pubchem.ncbi.nlm.nih.gov / ) with acetyltributyl citrate as the search term, we retrieved the standard chemical structure and simplified molecular linear input specification (SMILES) information for acetyltributyl citrate (ATBC).

[0036] Subsequently, a systematic identification of potential human targets for ATBC was initiated. The ChEMBL database (https: / / www.ebi.ac.uk / chembl / ) was accessed, and a search was performed using acetyl tributyl citrate as the keyword, limiting the species to *Homo sapiens*. Information on reported or predicted human targets of ATBC was collected to obtain basic data on its known targets. Considering the potential limitations of a single database in terms of coverage, a complementary strategy was adopted to further expand data sources and improve the comprehensiveness and robustness of target predictions. The previously acquired ATBC SMILES information was simultaneously submitted to the Swiss Target Prediction database (https: / / www.swisstargetprediction.ch / ) and the STITCH database (https: / / stitch-db.org / ) for target prediction analysis.

[0037] Next, target identifiers were collected, integrated, and standardized. Due to significant differences in source and inconsistent formats among the original target identifiers from the CheEMBL, SwissTarget Prediction, and STITCH databases, direct merging could easily lead to data redundancy and ambiguity. To address this issue, all original identifiers were submitted to the UniProt database (https: / / www.uniprot.org / ) for standardized mapping. This mapping process automatically identifies different format identifiers pointing to the same protein entity and merges them into a unique UniProt accession number. This efficiently removes duplicate targets while simultaneously verifying the consistency of target entities, ensuring that each target corresponds to unique protein information. Based on the standardized UniProt accession numbers, the standard gene symbol corresponding to each accession number was further extracted, achieving comprehensive standardization of target names and avoiding subsequent analysis errors caused by name differences.

[0038] Ultimately, a set of potential human targets for ATBC was obtained, using the UniProt accession number-standard gene symbol as a unified format and after integration, deduplication, and strict quality control.

[0039] (2) Differential expression analysis and WGCNA analysis: First, gene expression profiling data were screened from the TCGA database (https: / / www.cancer.gov / ccg / research / genome-sequencing / tcga) to obtain gene expression profiles and corresponding clinical data from the Breast Cancer Research and Development (BRCA) project. To identify stable and abnormal gene expression patterns in breast cancer, breast cancer samples with paired adjacent normal breast tissue were selected. The selection criteria were as follows: Program selected as TCGA, Project as BRCA, and Data Category limited to Transcriptome Profiling. The gene expression data of the final included breast cancer tissues (n=104) and their paired normal breast tissues (n=104) served as the foundation dataset for identifying disease-related genes in this step.

[0040] To identify genes with significantly altered expression in breast cancer, the gene expression profile data were standardized and differentially expressed using the limma package in R (version 4.1.2). During the results screening phase, relatively strict statistical criteria were set, requiring an absolute value of the logarithmic fold change in gene expression |log2FC| ≥ 0.585, and a p-value < 0.05 after multiple validation. Based on these criteria, statistically significant differentially expressed genes were finally identified, resulting in a set of differentially expressed genes in breast cancer.

[0041] Meanwhile, to delve deeper into functional modules related to disease phenotypes from the perspective of gene co-expression, weighted gene co-expression network analysis (WGCNA) was performed on the same set of gene expression profile data. At the initial stage of network construction, different soft threshold parameters were systematically evaluated. By examining the relationship between scale independence indices and average connectivity, a soft threshold of 18 was ultimately selected. At this threshold, the scale-free topology fit index R0 of the network was [value missing]. 2 =0.85, indicating that the network has good scale-free characteristics. Based on this parameter, the weighted adjacency matrix between genes is calculated and further converted into a topological overlap matrix (TOM) to more accurately represent the strength of association between genes. Subsequently, a dynamic tree cutting algorithm (minClusterSize=100, deepSplit=2, pamRespectsDendro=TRUE) is used to perform hierarchical clustering and module partitioning of genes, initially identifying several co-expression modules; to simplify the module structure and enhance its biological significance, adjacent modules with a height below 0.25 in the clustering dendrogram are merged. After this step, nine relatively stable gene co-expression modules are finally obtained. To screen out the core modules directly related to the pathological state of breast cancer, the correlation between the characteristic genes of each module and the breast cancer phenotype is calculated, and the correlation is... p Modules with a correlation value < 0.05 and a correlation coefficient r ≥ 0.36 were identified as disease-related modules. Based on this criterion, six modules were selected: brown, red, pink, purple, black, and blue-green. All non-repetitive genes contained in these six modules were extracted and merged to obtain the final set of candidate targets related to breast cancer.

[0042] (3) Screening overlapping targets of ATBC and breast cancer and constructing a PPI network: First, using R (version 4.1.2), an intersection analysis was performed on the potential human target set for ATBC obtained in step (1), the set of differentially expressed genes for breast cancer obtained in step (2), and the set of candidate targets related to breast cancer. The results were visualized using a Venn diagram, identifying 112 overlapping targets that coexist in all three sets, forming an overlapping target set. These co-occurring targets are considered to have a high potential association with breast cancer in the context of ATBC.

[0043] To explore the potential functional connections between these overlapping targets, their corresponding standard gene symbols were submitted to the STRING database (https: / / cn.string-db.org / ). During the search, the species was limited to *Homosapiens*, and the minimum interaction confidence threshold was set to 0.4 according to the database's recommended parameters. This setting helps to effectively reduce background noise interference from low-confidence interactions or isolated nodes while ensuring network reliability. Furthermore, isolated nodes with no connections were explicitly removed from the network, thus constructing a PPI network focused on core protein-protein interactions.

[0044] Subsequently, Cytoscape software (version 3.9.0) was used to visualize and analyze the topology of the aforementioned PPI network. When generating the network graph, the visual representation was adjusted based on the degree centrality of the nodes: the higher the degree centrality, the larger the graph size. This visualization method makes key nodes with extensive connections stand out more prominently in the graph, facilitating intuitive identification of core members in the network. The degree centrality index objectively reflects the number of connections a node has in the network. The higher the connectivity of a target, the more extensive its interactions with other proteins in the network, suggesting that it may be located at a key position in functional regulatory pathways. Therefore, we speculate that these highly centralized targets are likely to play a more crucial regulatory role in the biological processes by which ATBC potentially influences the development and progression of breast cancer.

[0045] (4) Perform functional and pathway enrichment analysis on overlapping targets: To systematically analyze the biological functions and pathway background of the overlapping target set obtained in step (3), GO functional annotation and KEGG pathway enrichment analysis were performed using the FUMA platform (https: / / fuma.ctglab.nl / ), and cross-validation and deep functional clustering were conducted on key results using the DAVID platform (https: / / david.ncifcrf.gov / ). The significance screening criteria were set as follows: corrected false detection rate (FDR) < 0.05, and each enriched entry was required to contain at least three target genes. Finally, the significantly enriched entries that met the criteria were sorted by fold enrichment for visualization and result interpretation.

[0046] GO analysis revealed that these targets were significantly enriched at multiple functional levels: at the biological process level, they are mainly involved in RNA polymerase II-mediated transcriptional regulation (including inhibition and activation) and methylation-related activities; at the cellular component level, they are mainly located in the nucleus, nucleoplasm, and nuclear compartments; and at the molecular function level, they are mainly involved in chromatin binding, DNA-binding transcription factor binding, and histone methyltransferase activity. KEGG pathway analysis further revealed that overlapping targets were significantly enriched in pathways such as the transcriptional repression complex and pericentromere heterochromatin organization.

[0047] The combined results of the above analysis suggest that ATBC may exert its cancer-promoting effect by interfering with two core biological processes closely related to the occurrence and development of breast cancer: transcriptional regulation and chromatin remodeling.

[0048] (5) Screening potential key targets related to breast cancer prognosis from overlapping targets using multiple machine learning algorithms: First, the data foundation for the model was determined. The gene expression profiles of the breast cancer samples used in step (2) (excluding their paired normal samples) and their corresponding clinical survival follow-up information were used to construct a training set for identifying prognostic-related genes. Simultaneously, the GSE42568, GSE45827, and GSE20685 datasets were obtained from the GEO database (http: / / www.ncbi.nlm.nih.gov / geo / ), and gene expression profiles and corresponding clinical survival follow-up information of a total of 327 breast cancer samples were extracted to form an external validation set. To ensure data quality and comparability, all gene expression data input to the model were standardized, and batch effect correction was applied to eliminate technical variations caused by different platforms or experimental batches, thereby ensuring the reliability and consistency of subsequent analyses.

[0049] In terms of model construction methodology, a multi-algorithm combination strategy is employed to optimize prediction performance. Specifically, the system selects various machine learning algorithms, including Random Forest (RF), Elastic Network (enet), Lasso Regression (Lasso), Stepwise Cox Regression (Stepglm), Cox Boost (gimBoost), Cox Partial Least Squares Regression (plsRglm), Generalized Boosting Regression Model (gimBoost), Survival Support Vector Machine (SVM), Ridge Regression (Ridge), XGBoost, Naive Bayes, and Linear Discriminant Analysis (LDA). Different algorithm combination models are constructed and evaluated using the mime package in R (version 4.1.2). Model performance is evaluated using the consistency index C-index as the core metric, and quantitative ranking is performed on the training set through 10-fold cross-validation. Evaluation results show that the combined algorithm model exhibits excellent performance with a C-index exceeding 0.9 on the training set, and the C-index also remains above 0.9 on the external validation datasets (GSE42568 dataset with 104 breast cancer samples and GSE45827 dataset with 130 breast cancer samples). In this embodiment, the combined strategy of random forest for initial screening and Lasso regression for fine screening is selected as the final scheme for subsequent target selection.

[0050] To ensure the model's rigor and reliability, the selection of all core targets and the determination of model coefficients were completed within the training set, completely avoiding data leakage. The above combined screening strategy was implemented: First, a random forest model was used on the training set to perform preliminary dimensionality reduction on the overlapping target set obtained in step (3). This model constructs and integrates multiple survival decision trees, assesses the importance of each gene to patient survival outcomes based on the degree of reduction in node impurity, and sorts them according to variable importance scores, selecting 14 candidate key genes most relevant to survival. Subsequently, on the same training set, Lasso regression was used for fine-tuning these candidate key genes. This algorithm automatically selects the optimal penalty coefficient through L1 regularization and 10-fold cross-validation, compressing the regression coefficients of unimportant variables to zero, thereby achieving automatic feature selection. This process screened MAOA and ADRA2A from the overlapping target set, identifying them as two potential key targets significantly related to the prognosis of breast cancer patients, and fixed their coefficients in the prognostic model.

[0051] Based on the two potential key targets and their coefficients determined on the training set, a final prognostic prediction model was constructed. The risk score calculation formula is: Risk Score = -0.4266 × ADRA2A expression level + 0.7683 × MAOA expression level. Model validation was conducted in two steps. First, within the training set, patients were divided into high-risk and low-risk groups based on the median calculated risk score, and internal performance validation was performed using survival analysis and the area under the time-dependent ROC curve (AUC). Subsequently, to rigorously test the model's generalization ability, the same method was used to validate the model's effectiveness on the validation set. Results showed that in the training set, the low-risk group (fen=0) had a better prognosis than the high-risk group (fen=1); compared to the low-risk group (fen=0), the high-risk group (fen=1) had a higher risk of death and shorter survival time, with statistically significant differences between the two groups. p =0.0012); the AUC of this dual-target model for predicting 1-year, 3-year, and 5-year risks reached 0.768, 0.725, and 0.728, respectively. In the validation set, there was a significant difference in survival outcomes between the low-risk group (fen=0) and the high-risk group (fen=1). p =0.0063); compared to the low-risk group (fen=0), the high-risk group (fen=1) had a higher risk of death and shorter survival time; the predictive ability of this dual-target model improved over time, with AUCs for 1-year, 3-year, and 5-year risks reaching 0.567, 0.638, and 0.704, respectively. Results from the external validation set confirmed that the model possesses good predictive performance and generalization ability. Therefore, MAOA and ADRA2A, as genes closely related to breast cancer prognosis screened from potential ATBC targets, warrant further investigation regarding their association, as they may constitute a potential bridge between ATBC exposure and adverse breast cancer outcomes.

[0052] (6) Verify the binding activity of the core target to ATBC: Search for the two potential key targets MAOA (PDB ID: 1QAK) and ADRA2A (PDB ID: 9CBI) selected in step (5) in the RCSB PDB database (https: / / www.rcsb.org / ), download the corresponding three-dimensional crystal structures and save them in PDBQT format. Search for ATBC in the PubChem database (https: / / pubchem.ncbi.nlm.nih.gov / ), view and save its two-dimensional molecular structure. Use PyMOL software to dehydrate and remove ligands from the macromolecules, and use AutoDock software for preprocessing such as hydrogenation, charge calculation and atomic rigidity setting. After processing, predict active sites and construct docking boxes, set docking programs and parameters and start docking to obtain the corresponding binding energies, and use PyMOL software for docking visualization analysis. When the binding energy is below -5kJ / mol, it indicates that the binding activity of the ligand to the target protein is stronger, and the lower the binding energy, the tighter the binding between the ligand and the target protein. The docking results showed that the binding energy of ATBC with ADRA2A was -5.6 kcal / mol, and with MAOA was -5.1 kcal / mol. Both values ​​are lower than the set reference threshold of -5 kJ / mol. In summary, these calculations theoretically suggest that ATBC molecules may achieve good binding with the active pockets of ADRA2A and MAOA, respectively. This result provides preliminary molecular-level structural clues and computational simulation support for ATBC to exert its biological effects by acting on ADRA2A and MAOA, two targets related to breast cancer prognosis. It enhances the possibility that these two targets are potential action nodes for ATBC and lays a preliminary foundation for subsequent experimental verification to explore the specific mechanisms of action of ATBC in the development and progression of breast cancer.

Claims

1. A method for analyzing the mechanism of action of the plasticizer tributyl citrate in inducing breast cancer based on network toxicology, characterized in that: Includes the following steps: (1) Acquisition and standardization of targets related to tributyl acetyl citrate; (2) Differential gene expression analysis and weighted gene co-expression network analysis in breast cancer; (3) Screening overlapping targets of acetyl tributyl citrate and breast cancer and constructing a protein-protein interaction network; (4) Perform functional and pathway enrichment analysis on overlapping targets; (5) Screening core targets from overlapping targets based on a combination of multiple machine learning algorithms; (6) Verify the binding activity of the core target with tributyl acetyl citrate.

2. The method according to claim 1, characterized in that: Step (1) specifically involves: using acetyl tributylcitrate as the search term, obtaining the chemical structure and SMILES information of acetyl tributyl citrate from the PubChem database; using acetyl tributyl citrate as the search term, screening for human-specific potential targets in the ChEMBL database, and simultaneously submitting the SMILES information of acetyl tributyl citrate to the Swiss Target Prediction database and the STITCH database for complementary potential target identification; summarizing all obtained target identifiers, performing standardized mapping using the UniProt database, obtaining a unified UniProt accession number and standard gene symbol, and forming a set of potential human targets for acetyl tributyl citrate.

3. The method according to claim 1, characterized in that: Step (2) specifically involves: obtaining transcriptome data of breast cancer tissue and its paired normal breast tissue from the TCGA database; performing differential expression analysis using the limma package in R language, with |log2FC|≥0.585 and adjusted for 0.

585. p A value <0.05 was used as the screening criterion for differentially expressed genes, and a set of differentially expressed genes in breast cancer was obtained. Weighted gene co-expression network analysis was performed on the same transcriptome data, with the optimal soft threshold being 18. A topological overlap matrix was constructed, and a dynamic tree cutting algorithm was used to identify co-expressed gene modules. Calculate the correlation between characteristic genes of each module and breast cancer phenotype, and screen out the correlations. p Value < 0.05 and |correlation coefficient r |>0.36 module, extract all non-repetitive genes within the module to form a set of candidate targets related to breast cancer.

4. The method according to claim 1, characterized in that: The specific steps (3) are as follows: the intersection analysis of the set of target points related to tributyl citrate obtained in step (1) and the set of differentially expressed genes and candidate target points related to breast cancer obtained in step (2) is performed to screen out the set of overlapping target points of tributyl citrate and breast cancer; the standard gene symbols corresponding to the overlapping target points are input into the STRING database, the species is set as Homo sapiens, and a protein-protein interaction network is generated with a minimum interaction confidence threshold of 0.4; The protein-protein interaction network was visualized and its topology analyzed using Cytoscape software, and the visualization was presented based on the degree centrality of the nodes.

5. The method according to claim 1, characterized in that: Step (4) specifically involves using the DAVID and FUMA platforms to perform GO functional annotation and KEGG pathway enrichment analysis on the overlapping target set. The significance screening criteria are set as follows: the corrected false detection rate (FDR) is <0.05, and each enrichment entry contains at least 3 target genes.

6. The method according to claim 1, characterized in that: Step (5) specifically involves: using the gene expression profiles and corresponding clinical survival data of breast cancer samples in the TCGA database as the training set, and using an independent breast cancer gene expression dataset from the GEO database as the external validation set; constructing and evaluating models using different machine learning algorithm combinations; using the consistency index C-index as the evaluation metric, selecting combination strategies with a C-index greater than 0.7; using the combination strategies to select core targets from the overlapping target set, and constructing a prognostic prediction model based on the core targets.

7. The method according to claim 6, characterized in that: The process of using a combination strategy of random forest initial screening and Lasso regression fine screening to screen core targets from overlapping target sets includes: first, using a random forest model on the training set to perform preliminary dimensionality reduction on overlapping targets and screen out candidate key genes most relevant to survival; second, using Lasso regression on the training set to finely screen these candidate key genes and determine the final core targets.

8. The method according to claim 1, characterized in that: The specific steps (6) are as follows: obtain the three-dimensional crystal structure of the core target protein from the RCSB PDB database, preprocess it using PyMOL and AutoDockTools software and convert it into PDBQT format as a receptor file; The molecular structure of tri-n-butyl acetyl citrate was obtained from the PubChem database and converted into PDBQT format as a ligand file. Molecular docking was performed using AutoDock Vina software, with a binding energy <-5.0 kcal / mol as the criterion for stable binding, and the binding energy and binding mode were predicted. The conformation with the lowest binding energy was visualized and analyzed using PyMOL software.

9. The application of the core targets screened by the method of any one of claims 1 to 8 in constructing a prognostic prediction model for breast cancer.

10. The application according to claim 9, characterized in that: The core targets are MAOA and ADRA2A.