A method and device for screening potential substitutes of bisphenol a and a storage medium

By combining a hybrid deep learning architecture with molecular descriptors and molecular graphs, the problem of insufficient accuracy and generalization ability in screening bisphenol A (BPA) alternatives in traditional methods is solved. This approach identifies safer lignin derivatives as BPA alternatives, improving the effectiveness of screening and designing environmentally friendly alternatives.

CN117292766BActive Publication Date: 2026-04-17JIANGHAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
JIANGHAN UNIVERSITY
Filing Date
2023-05-19
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing technologies are insufficient for efficiently screening out bisphenol A alternatives with lower risks to human health, and traditional methods suffer from insufficient accuracy and generalization ability when predicting AR antagonistic activity.

Method used

A hybrid deep learning architecture is adopted, combining molecular descriptors and molecular graphs. By acquiring a global AR antagonistic dataset and a local BPA analog dataset, the machine learning algorithm model is optimized to construct a hybrid model to predict the AR antagonistic activity of potential BPA alternatives.

Benefits of technology

It improves the accuracy and generalization of BPA alternative predictions, identifies lignin derivatives that are safer than bisphenol analogs as potential alternatives, and enhances the screening and design capabilities for environmentally friendly BPA alternatives.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117292766B_ABST
    Figure CN117292766B_ABST
Patent Text Reader

Abstract

The application relates to a bisphenol A potential substitute screening method, device and storage medium, the method comprising the following steps: acquiring a global AR antagonistic data set and a local BPA analogue data set; acquiring molecular characterization of the global and local BPA analogue data sets; acquiring a machine learning algorithm model; optimizing parameters of the machine learning algorithm model; processing the global AR antagonistic data set and the local BPA analogue data set by using the machine learning algorithm model; and acquiring a potential BPA substitute predicted by an optimal mixed model. The application provides a mixed deep learning architecture, which combines molecular descriptors and molecular graphs to predict the antagonistic activity of a compound on AR; compared with previous models, the mixed model can extract a large amount of chemical information from different molecular features, so that the generalization ability of the model for predicting BPA substitutes is improved; and the prediction result also shows that lignin derivatives are safer than bisphenol analogues as BPA substitutes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of bisphenol A, and more particularly to a method, apparatus and storage medium for screening potential alternatives to bisphenol A. Background Technology

[0002] Bisphenol A (BPA) is a ubiquitous chemical substance commonly used in the synthesis of epoxy resins, plastics, and polycarbonates. These BPA-synthesized materials are widely used in various products, including medical devices, building materials, thermal receipt paper, food containers, and baby bottles. However, free BPA monomers can leach from these materials and enter the air, water, and soil, being absorbed by the human body through inhalation of dust, drinking water, and skin contact. Over the past few decades, research has shown that BPA can cause various health hazards, such as increasing the risk of prostate cancer and reducing sperm motility. Therefore, environmental scientists have also studied its potential toxic mechanisms. For example, low concentrations of BPA can inhibit the action of endogenous hormones, thereby interfering with the normal functioning of the endocrine system and affecting cell proliferation and differentiation.

[0003] Given the toxicity and health risks of bisphenol A (BPA), the European Union issued a regulation in 2011 banning the addition of BPA to baby bottles. Simultaneously, EU member states have strictly limited the use of BPA, successively banning its use in various food packaging materials. In response to these regulations, new consumer products have been advertised by manufacturers as "BPA-free." However, while these products may not contain BPA, they often contain BPA substitutes, such as bisphenol S (BPS), bisphenol F (BPF), and bisphenol AF (BPAF). These substitutes, with molecular structures similar to BPA, are called BPA analogs or bisphenol A analogs. Their levels in environmental media are even higher than BPA, and they have already been detected in human samples (e.g., blood, urine, and tissues of children and adults). For these, the relevant authorities have only set a limit of 0.05 mg / kg for BPS. Studies have shown that BPA analogs also cause environmental pollution and have endocrine disruptor characteristics similar to BPA. Unlike BPA, there are other BPA alternatives with different structural features, such as Pergafast201, which has a larger molecular volume than BPA, and SYR-EPO, which is derived from eugenol. Therefore, the challenge in designing BPA alternatives lies not only in reducing their potential health risks to humans, but also in avoiding simply using them as "replacements" to comply with regulations.

[0004] Regarding the toxicity assessment of androgen receptor (AR), many studies have reported that bisphenol A (BPA) analogues can exhibit AR antagonistic activity, leading to prostate-related diseases by modulating AR-mediated signaling pathways. Although current in vitro assay techniques have reduced the cost of AR toxicity assessment, high-throughput screening of potential BPA alternatives remains impractical. Therefore, current research has proposed computational methods, such as molecular docking, molecular dynamics simulations, and quantitative structure-activity relationship (QSAR) models, to aid in compound determination. These efforts expand the scope of traditional toxicology and accelerate the development of computational toxicology as an interdisciplinary field. Applying computer-aided prediction in chemical safety assessments can better assist relevant departments in assessing the AR toxicity of BPA alternatives.

[0005] Among these computational methods, quantitative structure-activity relationship (QSAR) models based on machine learning (ML) are an emerging artificial intelligence technique that can be used to assist in drug design and chemical safety assessment. It can learn features and rules of activity labels from known datasets and transfer these features and rules to new, unlabeled datasets for prediction. ML-based QSAR models can handle "big data" and predict their activity or toxicity, which is very difficult and costly for traditional in vivo and in vitro assays. For predicting the chemical properties of environmental chemicals, traditional ML algorithms, such as random forests (RF), support vector machines (SVM), and graph neural networks (GNN), have proven highly effective in previous studies.

[0006] For predicting androgen receptor (AR) activity, the Tox21 and ToxCast datasets collected by the US Environmental Protection Agency (USEPA) have been widely used to build classification models to identify AR binders, agonists, and antagonists from thousands of chemicals with different molecular structures. Subsequently, the USEPA's National Center for Computational Toxicology launched the Cooperative Modeling Project for Androgen Receptor Activity (CoMPARA) to develop consensus models based on traditional machine learning algorithms for predicting the impact of man-made chemicals on AR activity. Previous studies have shown that RF and SVM models using molecular descriptors or molecular fingerprints exhibit good predictive capabilities for AR activity. Furthermore, on the Tox21 NR-AR and NR-AR-LBD datasets, gradient boosting decision tree models with multi-scale weighted color maps achieved the highest balanced accuracy compared to models with different molecular fingerprints. Evaluations by Walter et al. showed that the XGB model has higher performance when using Morgan fingerprints for multi-task deep neural networks on the same dataset. For 11 datasets of the AR signaling pathway in the ToxCast / Tox21 assay, a Bayesian machine learning model employing an extended connection fingerprint (ECFP) with a diameter of 6 showed better statistical performance than other ML models such as RF and SVM. However, most previous studies have only used traditional ML algorithms to build single models; therefore, more advanced strategies are needed to further improve the predictive accuracy and generalization ability of BPA alternatives. Summary of the Invention

[0007] In order to solve the above-mentioned technical problems, or at least partially solve the above-mentioned technical problems, this application provides a method, apparatus and storage medium for screening potential alternatives to bisphenol A.

[0008] In a first aspect, this application provides a method for screening potential alternatives to bisphenol A, the method comprising the steps of:

[0009] Obtain the global AR antagonist dataset and the local BPA analogue dataset;

[0010] Obtain molecular characterizations of the global and local BPA analog datasets;

[0011] Obtain machine learning algorithm models;

[0012] Optimize the parameters of the machine learning algorithm model;

[0013] The global AR antagonist dataset and the local BPA analog dataset are processed using the machine learning algorithm model described above.

[0014] Obtain potential BPA alternatives predicted by the optimal hybrid model.

[0015] Preferably, obtaining the global AR antagonist dataset and the local BPA analogue dataset includes the following steps:

[0016] Obtain the NuRA dataset;

[0017] The NuRA dataset is divided according to AR antagonistic activity;

[0018] Obtain the binary classification dataset after partitioning.

[0019] Obtain a dataset of BPA analogues and alternatives.

[0020] Preferably, the molecular characterization of the global and local BPA analog datasets includes the following steps:

[0021] Obtain the representation model;

[0022] Obtain compounds from the global and local datasets;

[0023] The compound was characterized using the characterization model described above.

[0024] Preferably, obtaining the machine learning algorithm model includes the following steps:

[0025] Obtain the GNN algorithm;

[0026] Obtain the ML algorithm;

[0027] A binary classification model is constructed using the GNN algorithm and the ML algorithm.

[0028] Preferably, optimizing the parameters of the machine learning algorithm model includes the following steps:

[0029] Obtain the NuRA-AR dataset and local BPA analogue dataset;

[0030] The NuRA-AR dataset and the local BPA analog dataset are divided into training set, validation set and test set;

[0031] Adjust the hyperparameters of the machine learning algorithm model;

[0032] Calculate the correlation coefficient of the machine learning algorithm model.

[0033] Preferably, the step of processing the global AR antagonist dataset and the local BPA analogue dataset using the machine learning algorithm model includes the following steps:

[0034] Obtain the global AR antagonistic dataset and the machine learning algorithm model;

[0035] Obtain the local BPA analogue dataset and the machine learning algorithm model;

[0036] Obtain the compounds from the global AR antagonistic dataset;

[0037] Obtain compounds from the local BPA analogue dataset;

[0038] The compound is input into the machine learning algorithm model;

[0039] Obtain the optimal model from the machine learning algorithm models;

[0040] The optimal hybrid model is obtained by fusing features from the optimal model.

[0041] Preferably, obtaining the potential BPA alternatives predicted by the optimal hybrid model includes the following steps:

[0042] Find potential BPA alternatives;

[0043] Construct an application prediction set using the BPA alternative;

[0044] Obtain the optimal hybrid machine learning model;

[0045] The optimal hybrid machine learning algorithm model is used to predict the application prediction set.

[0046] Secondly, a device for screening potential alternatives to bisphenol A is provided, comprising:

[0047] The dataset acquisition module is used to acquire the global AR antagonist dataset and the local BPA analogue dataset;

[0048] The molecular characterization acquisition module is used to acquire the molecular characterization of the global and local BPA analogue datasets;

[0049] The algorithm model acquisition module is used to acquire machine learning algorithm models;

[0050] The parameter optimization module is used to optimize the parameters of the machine learning algorithm model.

[0051] A dataset processing module is used to process the global AR antagonist dataset and the local BPA analog dataset using the machine learning algorithm model;

[0052] The alternative prediction module is used to obtain potential BPA alternatives predicted by the optimal hybrid model.

[0053] Thirdly, an electronic device is provided, the electronic device comprising:

[0054] At least one processor; and,

[0055] A memory communicatively connected to the at least one processor; wherein,

[0056] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform any of the aforementioned methods for screening potential alternatives to bisphenol A.

[0057] Fourthly, a non-transitory computer-readable storage medium is provided that stores computer instructions for causing the computer to perform any of the aforementioned methods for screening potential alternatives to bisphenol A.

[0058] The technical solutions provided in this application have the following advantages compared with the prior art:

[0059] This application provides a method, apparatus, and storage medium for screening potential BPA alternatives. It offers a hybrid deep learning architecture that combines molecular descriptors and molecular graphs to predict the antagonistic activity of compounds against bisphenol A (BPA). Compared to previous models, this hybrid model can extract a large amount of chemical information from different molecular features, thereby improving the model's generalization ability to predict BPA alternatives. Furthermore, the prediction results indicate that using lignin derivatives as BPA alternatives is safer than bisphenol analogs. Overall, this research contributes to screening BPA alternatives and helps design environmentally friendly BPA alternatives. Attached Figure Description

[0060] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0061] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0062] Figure 1 This is a schematic flowchart of a method for screening potential alternatives to bisphenol A provided in an embodiment of the present invention;

[0063] Figure 2 This is a schematic diagram of a potential alternative screening device for bisphenol A provided in an embodiment of the present invention;

[0064] Figure 3 This is a schematic diagram of the structure of an electronic device provided by the present invention;

[0065] Figure 4 This is a schematic diagram of the structure of a non-transitory computer-readable storage medium provided by the present invention;

[0066] Figure 5 This is a schematic diagram of the chemical spatial distribution of the dataset for a screening method for potential substitutes for phenol A provided in an embodiment of the present invention;

[0067] Figure 6 This is a schematic diagram illustrating the prediction of potential BPA substitutes using a bisphenol A (BPA) potential substitute screening method provided in an embodiment of the present invention. Detailed Implementation

[0068] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0069] Figure 1 This is a flowchart illustrating a method for screening potential alternatives to bisphenol A, provided in an embodiment of this application.

[0070] This application provides a method for screening potential alternatives to bisphenol A, the method comprising the steps of:

[0071] S1: Obtain the global AR antagonist dataset and the local BPA analogue dataset;

[0072] In this embodiment of the application, obtaining the global AR antagonist dataset and the local BPA analogue dataset includes the following steps:

[0073] Obtain the NuRA dataset;

[0074] The NuRA dataset is divided according to AR antagonistic activity;

[0075] Obtain the binary classification dataset after partitioning.

[0076] Obtain a dataset of BPA analogues and alternatives.

[0077] Specifically, in this study, the global AR antagonist dataset was considered for constructing a predictive model for AR antagonists. This dataset, carefully processed as part of the NuRA (nuclear receptor activity) dataset (referred to as NuRA-AR), contains 15,247 compound entries and is categorized into five classes based on AR antagonist activity: (1) active, (2) weakly active, (3) inactive, (4) uncertain, and (5) missing data. In this study, missing and uncertain entries were removed, and weakly active entries were merged into the active category, resulting in a binary classification dataset of 6,108 chemicals (1,166 active and 4,942 inactive). In addition, we collected 66 BPA analogs and alternatives from previous studies. Of these 66 compounds, 26 were reported to have no AR antagonist activity, while the remaining 40 showed AR antagonist activity. The aforementioned BPA dataset was used as an external validation set (EVS) to evaluate the generalization ability of the developed AR antagonist predictive model.

[0078] S2: Obtain the molecular characterization of the global and local BPA analog datasets;

[0079] In this embodiment of the application, obtaining the molecular characterization of the global and local BPA analog datasets includes the following steps:

[0080] Obtain the representation model;

[0081] Obtain compounds from the global and local datasets;

[0082] The compound was characterized using the characterization model described above.

[0083] Specifically, for GNN models (such as Graph Convolutional Networks (GCN), Graph Attention Networks (GAT), Attention Fingerprint (AFP), and Message Passing Neural Networks (MPNN), as well as derivative models of MPNN, such as Directed MPNN (D-MPNN), they are able to learn feature representations of molecular structures, where atoms are nodes and bonds are edges as inputs. Nodes are described by atom type, atom element, number of additional hydrogen atoms, number of valence electrons, aromaticity, and other properties. In this work, a set of atomic and bond descriptors based on the SMILES of molecules was generated using RDKit software, which served as input nodes and edge features for the GNN. For the chemical representation of the molecular descriptors, we used ChemDraw software to generate 2D representations of the compounds in this study. The molecular structures were obtained in D form and converted into 3D structures using Chem3D software with an energy minimization strategy. Subsequently, 5666 alvaDesc descriptors for each compound were calculated using the online chemical database OCHEM (https: / / ochem.eu / home / show.do) based on the 3D structures. Simultaneously, 1446 molecular descriptors, MACSS, and PubChem fingerprints for all compounds were calculated using PaDEL software. The RDKit package was used to calculate ECFP4 fingerprints (1024 bits). During feature selection, irrelevant variables and constants were excluded using correlation coefficients (r>0.95) because the original descriptors contained useless information, and a traditional ML-based model was developed using the remaining descriptors.

[0084] S3: Obtain the machine learning algorithm model;

[0085] In this embodiment of the application, obtaining the machine learning algorithm model includes the following steps:

[0086] Obtain the GNN algorithm;

[0087] Obtain the ML algorithm;

[0088] A binary classification model is constructed using the GNN algorithm and the ML algorithm.

[0089] Specifically, this study uses a local AR antagonistic dataset and develops local classification models based on different traditional ML algorithms. These traditional ML algorithms include RF, SVM, Gradient Boosting Classifier (GBC), Extra Tree Classifier (EXT), XGB, Lightweight Gradient Boosting (LGB), Classification Boosting (CAT), and Multilayer Perceptron (MLP). The global classification model simultaneously employs four GNN algorithms and eight traditional ML algorithms to construct a binary classification model, with its source code referenced from PyTorch, Deep Graph Library (DGL), DGL-LifeSci, and chemprop.

[0090] S4: Optimize the parameters of the machine learning algorithm model;

[0091] In this embodiment of the application, optimizing the parameters of the machine learning algorithm model includes the following steps:

[0092] Obtain the NuRA-AR dataset and local BPA analogue dataset;

[0093] The NuRA-AR dataset and the local BPA analog dataset are divided into training set, validation set and test set;

[0094] Adjust the hyperparameters of the machine learning algorithm model;

[0095] Calculate the correlation coefficient of the machine learning algorithm model.

[0096] Specifically, the NuRA-AR dataset and the local BPA analog dataset were split into training, validation, and test sets in an 8:1:1 ratio, based on different random seeds. After hyperparameter tuning, the optimal model was defined based on the highest receiver operating characteristic (AUC) value of each model in the validation set. Statistical metrics, accuracy (ACC), F1 score (F1), and Matthew correlation coefficient (MCC) were calculated to evaluate model performance.

[0097] S5: Process the global AR antagonist dataset and the local BPA analogue dataset using the machine learning algorithm model;

[0098] In this embodiment of the application, the step of processing the global AR antagonistic dataset and the local BPA analog dataset using the machine learning algorithm model includes the following steps:

[0099] Obtain the global AR antagonistic dataset and the machine learning algorithm model;

[0100] Obtain the local BPA analogue dataset and the machine learning algorithm model;

[0101] Obtain the compounds from the global AR antagonistic dataset;

[0102] Obtain compounds from the local BPA analogue dataset;

[0103] The compound is input into the machine learning algorithm model;

[0104] Obtain the optimal model from the machine learning algorithm models;

[0105] The optimal hybrid model is obtained by feature fusion of the optimal model.

[0106] Specifically, a comparison of classification models based on different methods is presented. Based on ten independent runs (random seeds: 0–9), the statistical criteria for comparing the overall performance of different models are summarized below. It can be seen that all GNN models exhibit acceptable predictive ability on the test set. The average prediction accuracy of GNN models ranges from 85.44% to 90.64%, with an average AUC exceeding 90%. In particular, among GNN models, D-MPNN shows better predictive performance, with higher F1 and MCC scores, which are widely used as evaluation metrics for classification accuracy on imbalanced datasets. For traditional ML-based models, XGB and LGB achieve comparable predictive performance relative to D-MPNN, and even SVM has the best average accuracy, AUC, F1, and MCC values ​​among all models. Of course, this does not mean that traditional ML models have surpassed GNN models. Typically, GNN models tend to have high data quality requirements, but in the field of molecular attribute prediction, data is often relatively limited in terms of chemical diversity and training size. Therefore, the current results indicate that traditional ML-based models can still provide reliable predictions of AR antagonistic activity even with small datasets (<10K).

[0107] Previous research has shown that the stepwise approach of multiple linear regression performs well in predicting the anti-androgenic activity of bisphenol A chemicals using three descriptors on small datasets. Therefore, this application also uses EVS to construct a traditional machine learning-based classification model. During feature selection, EVS is randomly split into 80% training set and 20% test set. Subsequently, recursive feature elimination based on RF is used to select the best descriptor from the remaining descriptors in the training set. The classification model is then constructed using either three or ten best descriptors. Notably, in the current data processing, this application uses different random seeds to repeat the data split for error analysis so that different best descriptors emerge in the ten independent builds of the model. The results show that the model based on the small dataset performs poorly, and the best model is the EXT model based on three best descriptors, with an average accuracy and AUC of only 77.14% and 80.08%, respectively. Based on the above results, the global AR antagonism dataset is considered more suitable for developing predictive models than the local BPA dataset.

[0108] Hybrid classification models. Hybrid models based on multiple types of features have proven effective in improving predictive power. Recently, different model architectures integrating molecular descriptors, molecular fingerprints, and molecular graphs have shown good predictive power on some drug discovery-related datasets, such as FP-GNN and CheMixNet. Therefore, in this study, four models that performed well in individual model comparisons—SVM, XGB, LGB, and D-MPNN—were selected as representative algorithmic architectures to develop hybrid models. Molecular descriptors and molecular graphs were concatenated as features and then imported into the architectures of the aforementioned models. Hyperparameter tuning was performed for each hybrid model, and the best model was retained for performance evaluation. Table 1 shows that most hybrid models (e.g., XGB, LGB, and D-MPNN) showed a small performance improvement on the test set compared to their respective individual models. This result indicates that hybrid architectures that learn molecular representations simultaneously from molecular descriptors and molecular graphs can improve the predictive power of the models. Furthermore, it is noteworthy that although D-MPNN showed a higher average evaluation metric than other hybrid models (Table 1), there were no significant differences among these hybrid models.

[0109] Recently, Ramaprasad et al. proposed a machine learning-based web service model (NR-ToxPred) that predicted the chemical bindings of nine different NRs using the NuRA dataset and achieved good performance. However, NR-ToxPred is trained based on traditional ML algorithms and molecular fingerprints, and does not use molecular descriptors as input, nor is its performance compared with GNN models. Therefore, this application compares the performance of the hybrid model developed in this study with NR-ToxPred on the AR antagonism dataset based on the statistical metric MCC. The average MCC value of the four hybrid models achieved better performance on the test set (69.99%–70.92%), which is better than the best model in previous work (67.14%) (Table 1). In addition, this application also compares the performance of FP-GNN and CheMixNet models on the current AR antagonism dataset based on the same data splitting method. By comparing accuracy, AUC, F1, and MCC values, this application found that the hybrid model developed in this application outperforms the FP-GNN and CheMixNet models in prediction performance on the test set (Table 1). This means that the hybrid model developed in this application significantly improves predictive ability through the combination of molecular descriptors and molecular diagrams.

[0110] Furthermore, previous studies have shown that consensus models generally outperform individual models. Therefore, this application calculated the average predicted probability of each compound in the current hybrid model as the consensus model. The results show that, based on statistical evaluation of ten independent runs, the consensus model performs comparably to the hybrid model on the test set (Table 1). Among the results of different runs, this application found that the corresponding consensus model has relatively high accuracy, F1, and MCC values ​​when a random seed of 5 is selected. Therefore, the specific model (seed = 5) was selected as the final prediction model for subsequent consistency and reproducibility comparisons.

[0111] Comparison of Generalization Ability on External Validation Set. BPA analogues and alternatives with experimental antagonistic activity were used as external validation sets (EVS) to evaluate the generalization ability of the developed hybrid models. The corresponding results are shown in Table 2. In addition to the four hybrid models and consensus models in this study, this application also compared the generalization ability of previous models to predict AR antagonistic activity on EVS. The results show that the accuracy values ​​of NR-ToxPred, FP-GNN, CheMixNet, CoMPARA, and NRMEA are 56.25%, 62.5%, 46.88%, 61.29%, and 68.75%, respectively, which are lower than the current hybrid models and consensus models (78.13% and 81.25%) (Table 2). Compared with traditional ML models and GNN models, the hybrid models and consensus models developed in this application also show superior performance in the evaluation of accuracy, AUC, F1, and MCC. Furthermore, the consensus model has the highest AUC value among the above models. Therefore, based on the overall statistical data of the test set and the external validation set, this application believes that the consensus model developed in this study exhibits excellent performance in predicting AR antagonistic activity and demonstrates a strong ability to distinguish structurally similar molecules from different data resources, with a wider range of applicability.

[0112] Application Domain Analysis. To evaluate the application domain (AD) of the predictive model, this application first implements an Euclidean distance method based on the alvaDesc descriptor. Results show that in 10 independent builds of the model, no compounds in the test set exceeded the AD, consistent with previous research suggesting that the Euclidean distance-based method may be too lenient to define the model's AD. Therefore, this application further executes a fingerprint similarity method to evaluate the AD. Statistical results show that as the maximum similarity threshold increases from 0.3 to 0.45, very few compounds in the test set (average: 7.8–106) exceed the AD. Compared to the Euclidean distance-based method, this application argues that fingerprint similarity evaluation may be more conservative and reliable, and can be used to define the AD of the trained model.

[0113] Therefore, this application subsequently used the fingerprint similarity method to assess the reliability of the predictions for BPA analogs and alternatives. For EVS, when the threshold was set relatively high at 0.45, approximately two molecules exceeded AD. In most cases, the threshold was set to 0.4, at which point the number of compounds exceeding AD decreased to a mean of 0.1. This indicates that only one compound exceeded AD in 10 independent assessments. For APS, when the threshold was set to 0.4, more compounds were found to exceed AD compared to EVS (mean: 2.7). Overall, approximately 0.15% and 3.0% of compounds were considered to exceed AD in EVS and APS, respectively. This is a relatively low percentage, lower than the percentage in the test set (10%), indicating that the model developed in this study can reliably predict the antagonistic activity of BPA analogs and alternatives against AR. However, when the prediction model was built based on EVS, the AD assessment showed that over 54% of compounds in APS exceeded AD as defined by the fingerprint similarity method. This result suggests that the major molecular structures of compounds in APS differ from those in EVS, consistent with the analysis of chemical spatial distribution. Therefore, it can be concluded that by considering only known BPA analogues when constructing models, small adversarial values ​​(AD) may limit the model's predictive ability for BPA substitutes that have significant structural differences from bisphenol analogues.

[0114] S6: Obtain potential BPA alternatives predicted by the optimal mixture model.

[0115] In this embodiment of the application, obtaining the potential BPA alternatives predicted by the optimal hybrid model includes the following steps:

[0116] Find potential BPA alternatives;

[0117] Construct an application prediction set using the BPA alternative;

[0118] Obtain the optimal hybrid machine learning model;

[0119] The optimal hybrid machine learning algorithm model is used to predict the application prediction set.

[0120] Specifically, this application collected 89 potential BPA alternatives as an Application Prediction Set (APS), the antagonistic activities of which have not yet been experimentally determined. The toxicity and health risks of these compounds are unknown, and they may become new pollutants. Therefore, we used the best classification model developed in this study to predict the AR antagonistic activities of these compounds.

[0121] Recently, guaiacol, extracted from lignin, has been reported as a potentially safer and more renewable alternative. In the current evaluation, this application found that the predicted results for the BPA and BPF scaffold structures showed significantly different results for the o-methoxy group substituted in the phenolic hydroxyl group. Figure 6For BPA-based guaiacol, the results showed that p,p′-GPA, p,p′-BGA, p,p′-GSA, and p,p′-BSA were predicted to be the active compounds. Figure 6 However, as BPF-based compounds, p,p′-GPF, p,p′-BGF, p,p′-GSF, and p,p′-BSF were predicted to be inactive compounds. Figure 6 These findings suggest that adding an o-methoxy group to the aromatic ring of BPF may sterically hinder the binding of key phenolic hydroxyl groups to AR, and that the dimethyl-substituted bridging carbon of bisphenol a may weaken the binding affinity and antagonistic activity of the o-methoxy group to the compound by increasing the hydrophobic interaction of the compound to AR. As previously reported, the guaiacol F obtained from lignin showed no detected estrogenic activity at environmental concentrations. Therefore, this application hypothesizes that guaiacol may be suitable as a renewable resource and scaffold structure for the design of more environmentally friendly BPA alternative molecules. In addition, furans derived from lignocellulose biomass (e.g., 5-hydroxymethylfurfural (HMF) and its derivatives) are also considered as renewable starting materials. These compounds exhibited inactivity in current predictions, suggesting that low molecular weight alternative bisphenol parent structures may be potential more sustainable BPA alternatives. Figure 6 ).

[0122] Environmental Implications. In this study, this application proposes a novel hybrid deep learning architecture that utilizes molecular descriptors and molecular graphs to predict AR antagonistic activity. Through comprehensive comparison, this application demonstrates that multi-type feature fusion based on various molecular representations can improve the model's generalization ability to BPA alternatives. Typically, model prediction studies in the environmental field deal with data from a single modality, such as images, audio, text, and signals. For example, toxicity prediction of environmental pollutants can extract structural features from 2D or 3D molecular images as input to graph convolutional neural networks, and various signals from infrared spectroscopy, Raman spectroscopy, and mass spectrometry are also identified as graphical or digital features. These feature extractions differ from traditional ML algorithms based on molecular descriptors and molecular fingerprints. However, real-world problems require capturing large amounts of chemical information from multiple data modalities, rather than a single modality.

[0123] Recently, ChatGPT, as an emerging AI technology, is revolutionizing all aspects of life, with the potential to significantly advance research in environmental chemistry and toxicity screening. Therefore, for data science and deep learning, attention-based language models (such as bidirectional encoder Transformers (BERT) and generative pre-trained Transformers (GPT)) hold promise for representing molecules and improving model predictive performance. Furthermore, using high-quality large datasets as pre-training sets for transfer learning also has the potential to enhance performance in downstream tasks such as chemical persistence, bioaccumulation, and toxicity prediction. Combining other computational models, such as molecular docking and MD simulations, can reveal the fundamental mechanisms of chemically induced AR toxicity at the atomic level. Overall, this research's workflow based on various computational methods contributes to the toxicological screening of BPA alternatives and plays a role in the design of future environmentally friendly BPA alternatives.

[0124] Table 1 Comparison of AR suppression prediction performance of the hybrid model and previous research models on the test set.

[0125]

[0126] Table 2 Comparison of generalization performance of the hybrid model and previous research models on EVS (seed = 5)

[0127]

[0128] Figure 5 t-SNE distributions for NuRA-AR (6108 compounds), EVS (66 compounds), and APS (89 compounds). Tanimoto coefficient distributions for active and inactive compounds in EVS using BPA as the standard (B). Tanimoto coefficient distributions for EVS and APS using BPA as the standard (C). Figure 6 BPA-based derivatives with o-methoxy groups substituted at the phenolic hydroxyl groups (A). BPF-based derivatives with o-methoxy groups substituted at the phenolic hydroxyl groups (B). 5-hydroxymethylfurfural (HMF) and its derivatives derived from lignocellulose biomass (C).

[0129] like Figure 2 This application provides a screening device for potential alternatives to bisphenol A, comprising:

[0130] Dataset acquisition module 10 is used to acquire the global AR antagonist dataset and the local BPA analogue dataset;

[0131] Molecular characterization acquisition module 20 is used to acquire the molecular characterization of the global and local BPA analog datasets;

[0132] Algorithm model acquisition module 30 is used to acquire machine learning algorithm models;

[0133] The parameter optimization module 40 is used to optimize the parameters of the machine learning algorithm model;

[0134] Dataset processing module 50 is used to process the global AR antagonistic dataset and the local BPA analog dataset using the machine learning algorithm model;

[0135] The alternative prediction module 60 is used to obtain potential BPA alternatives predicted by the optimal hybrid model.

[0136] The bisphenol A potential substitute screening device provided in this application can perform the bisphenol A potential substitute screening method provided in the above steps.

[0137] It should be understood that the specific embodiments described above are merely illustrative or explanatory of the principles of the invention and do not constitute a limitation thereof. Therefore, any modifications, equivalent substitutions, improvements, etc., made without departing from the spirit and scope of the invention should be included within the protection scope of the invention. Furthermore, the appended claims are intended to cover all variations and modifications falling within the scope and boundaries of the appended claims, or equivalent forms of such scope and boundaries.

[0138] The following is for reference. Figure 3 The diagram illustrates a structural schematic of an electronic device 100 suitable for implementing embodiments of the present disclosure. The electronic devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 3 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0139] like Figure 3 As shown, the electronic device 100 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 101, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 102 or a program loaded from a storage device 108 into a random access memory (RAM) 103. The RAM 103 also stores various programs and data required for the operation of the electronic device 100. The processing unit 101, ROM 102, and RAM 103 are interconnected via a bus 104. An input / output (I / O) interface 105 is also connected to the bus 104.

[0140] Typically, the following devices can be connected to I / O interface 105: input devices 106 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 107 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 108 including, for example, magnetic tapes, hard disks, etc.; and communication devices 109. Communication device 109 allows electronic device 100 to communicate wirelessly or wiredly with other devices to exchange data. Although electronic device 100 with various devices is shown in the figure, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0141] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 109, or installed from storage device 108, or installed from ROM 102. When the computer program is executed by processing device 101, it performs the functions defined in the methods of embodiments of this disclosure.

[0142] The following is for reference. Figure 4 The diagram illustrates a computer-readable storage medium suitable for implementing embodiments of the present disclosure, the computer-readable storage medium storing a computer program that, when executed by a processor, enables the implementation of the bisphenol A potential alternative screening method as described above.

[0143] This application provides a method, apparatus, and storage medium for screening potential BPA alternatives. It offers a hybrid deep learning architecture that combines molecular descriptors and molecular graphs to predict the antagonistic activity of compounds against bisphenol A (BPA). Compared to previous models, this hybrid model can extract a large amount of chemical information from different molecular features, thereby improving the model's generalization ability to predict BPA alternatives. Furthermore, the prediction results indicate that using lignin derivatives as BPA alternatives is safer than bisphenol analogs. Overall, this research contributes to screening BPA alternatives and helps design environmentally friendly BPA alternatives.

[0144] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0145] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A method for screening potential alternatives to bisphenol A, characterized in that, The method includes the following steps: S1. Construct a global AR antagonist dataset and a local BPA analogue dataset: 1a: Obtain various chemicals from the NuRA database, remove missing and uncertain data, and merge weakly active data to form a first binary classification dataset. The first binary classification dataset includes chemicals with AR antagonistic activity and chemicals without AR antagonistic activity. 1b: Collect a variety of reported BPA analogs and alternatives as an external validation set (EVS), which includes chemicals with AR antagonistic activity and chemicals without AR antagonistic activity; S2. Perform two molecular characterizations on the chemicals in the dataset: 2a: Graph Neural Network Characterization: Using RDKit, the SMILES of each chemical are converted into an atom-bond graph. Node features include atom type, number of valence electrons, and aromaticity, while edge features include bond type. 2b: Descriptor representation: First, based on Chem3D software, the descriptors are converted into 3D structures using an energy minimization strategy. 5666 alvaDesc descriptors are calculated using OCHEM, 1446 molecular descriptors and MACSS and PubChem fingerprints are calculated using PaDEL, and 1024 ECFP4 fingerprints are calculated using RDKit. Irrelevant variables are excluded with a correlation coefficient r>0.95, and a traditional ML model is constructed based on the remaining descriptors. S3. Constructing Machine Learning Algorithm Models: Eight traditional ML algorithms and four GNN algorithms are used to establish binary classification models. S4. Optimize model parameters: Divide the first binary classification dataset and EVS data into training set, validation set and test set according to a preset ratio. After adjusting the hyperparameters, define the optimal model based on the highest AUC value of the receiver operating characteristic curve of each model in the validation set. Evaluate the optimal single model corresponding to each algorithm by accuracy ACC, F1 score and Matthews correlation coefficient MCC. S5. Establish a hybrid model: Connect the remaining descriptors obtained in S2 with the molecular graph features, and input them into the selected multiple models to form multiple hybrid models; take the average of the prediction probabilities of the multiple hybrid models as the consensus model, and select the consensus model corresponding to the preset random seed as the final prediction model; S6. Using the final prediction model obtained in S5, potential BPA alternatives are predicted to obtain AR antagonistic activity classification results. The application domain is defined by the Tanimoto fingerprint similarity threshold to screen out candidate compounds that can safely replace bisphenol A.

2. The method of screening for potential bisphenol A replacements according to claim 1, wherein, The first binary classification dataset contains 6108 chemicals, of which 1166 have AR antagonistic activity and 4942 have no activity; the EVS contains 66 chemicals, of which 40 have AR antagonistic activity and 26 have no activity.

3. A bisphenol A potential replacement screening device suitable for use in the method of claim 1 or 2, characterized in that, include: The dataset acquisition module is used to acquire the global AR antagonist dataset and the local BPA analogue dataset; The molecular characterization acquisition module is used to acquire the molecular characterization of the global and local BPA analogue datasets; The algorithm model acquisition module is used to acquire machine learning algorithm models; The parameter optimization module is used to optimize the parameters of the machine learning algorithm model. A dataset processing module is used to process the global AR antagonist dataset and the local BPA analog dataset using the machine learning algorithm model; The alternative prediction module is used to obtain potential BPA alternatives predicted by the optimal hybrid model.

4. An electronic device, comprising: The electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the bisphenol A potential alternative screening method as described in claim 1 or 2.

5. A non-transitory computer-readable storage medium storing computer instructions for causing the computer to perform the bisphenol A potential alternative screening method as described in claim 1 or 2.

Citation Information

Patent Citations

  • Novel and efficient Graph neural network (GNN) for accurate chemical property prediction

    US20220406416A1

  • Prediction method and device for drug molecular feature attribute

    WO2022222492A1