Model for predicting endocrine disruptors by integrating pharmacological and toxicological profiles, method for constructing the same, and use thereof

By integrating network pharmacology and machine learning methods, a combined target profile was constructed, which solved the challenges of screening endocrine disruptors, achieved high-accuracy prediction of multi-target proteins, simplified the evaluation process of large-scale compounds, and improved the efficiency and accuracy of endocrine disruptor discovery.

CN116246718BActive Publication Date: 2026-02-17EAST CHINA UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211414808.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-11
Publication Date
2026-02-17
Estimated Expiration
2042-11-11

AI Technical Summary

Technical Problem

Existing technologies are insufficient for effectively screening and predicting potential endocrine disruptors, especially considering the interactions of multiple target proteins. This makes the discovery of endocrine disruptors a serious challenge, and traditional methods are often limited to common nuclear receptors, lacking comprehensiveness and accuracy.

Method used

By integrating network pharmacology and machine learning methods, we construct network-based target profiles and machine learning-based target profiles, generate combined target profiles, build endocrine disruptor prediction models using machine learning methods, and combine molecular fingerprint information to evaluate whether a compound is an endocrine disruptor.

Benefits of technology

It improves the accuracy and breadth of endocrine disruptor prediction, covering a variety of target proteins. The prediction model performs well on both the training and test sets, achieving a prediction accuracy of 80% in practical applications, and simplifies the prediction process for large-scale compounds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_1
    Figure SMS_1
  • Figure SMS_2
    Figure SMS_2
  • Figure SMS_3
    Figure SMS_3
Patent Text Reader

Abstract

The present application provides a model for predicting endocrine disruptors by integrating pharmacology and toxicology profiles, and a construction method and application thereof. Specifically, the present application constructs an endocrine disruptor prediction model based on pharmacology and toxicology data and in combination with molecular fingerprints of substructures of compounds. The endocrine disruptor prediction model of the present application can more accurately and efficiently evaluate whether a to-be-tested compound is an endocrine disruptor. The present application also develops an endocrine disruptor prediction system based on the endocrine disruptor prediction model.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of cheminformatics and drug safety assessment, in particular to a model for predicting endocrine disruptors by integrating pharmacology and toxicology profiles and a construction method and application thereof. BACKGROUND

[0002] Endocrine disruptor chemical (EDC) exposure can lead to adverse events, including birth defects, reproductive dysfunction, obesity, cancer, diabetes, and neurodevelopmental changes. Because of its huge harm to the environment, a large amount of funds and costs are used for endocrine disruptor related treatment. In the United States, the cost of diseases caused by endocrine disruptors is as high as 340 million US dollars per year, accounting for 2.33% of the domestic GDP. In Europe, the cost of diseases related to endocrine disruptor exposure is 163 million euros per year, accounting for 1.28% of the gross domestic product of the European Union. Therefore, it is of great significance to find potential endocrine disruptors to reduce the social and economic burden and reduce the risk of human health. However, most compounds lack endocrine disruptor related experimental tests, which leads to the discovery of potential endocrine disruptors is still a serious challenge.

[0003] Endocrine disruption has been a focus of toxicology research in recent decades, and several screening programs have been proposed for screening and predicting potential endocrine disruptors. As early as 1999, the U.S. Environmental Protection Agency (EPA) established the Endocrine Disruptor Screening Program to assess the risk of endocrine disorders caused by exposure to potential contaminants in pesticides and drinking water. Since then, the EPA has published 11 low-throughput in vitro and in vivo detection methods for endocrine disruptor screening programs in 2009. In addition, SPEED'98 in Japan and European Cluster in the European Union have been proposed to identify potential endocrine disruptors. Based on these screening programs and other experimental tests on endocrine disruption-related endpoints, a variety of databases have been established to collect information on structurally diverse endocrine disruptors, such as the Endocrine Disruptor Knowledge Base (EDKB) and DEDuCT. Therefore, various computational methods have been proposed and widely used for large-scale prediction of endocrine disruptors, with the advantages of low cost and high accuracy. For example, methods such as quantitative structure-activity relationship and compound structural feature fragments have been used to predict potential endocrine disruptors. Although these methods have been shown to have some reliability, most studies are usually only directed at common nuclear receptors, especially androgen receptor (AR) and estrogen receptor (ER). With the deepening of research on endocrine disruption, endocrine disruption is not only related to these nuclear receptors, but also to other target proteins, such as enzymes related to hormone synthesis and transporters related to hormone transport. Therefore, in addition to using traditional toxicological knowledge, integrating some pharmacological knowledge may help to discover endocrine disruptors and explain their mechanisms of action.

[0004] With the rapid development of systems biology and network pharmacology, the drug discovery model has shifted to a network model of'multiple drugs → multiple targets → multiple diseases'. For a compound, its interaction with multiple targets can lead to therapeutic effects and adverse events. Therefore, experimentally determined target profiles have been introduced into computational studies to explore potential safety issues of compounds. However, due to the incompleteness of high-quality experimental data, experimental target profiles are sometimes incomplete. To address this deficiency, some recent studies on computational target profiling have provided some new ideas. Compared to experimental target profiles, this computational target profiling can characterize a wider biological space of compounds. In recent years, a new computational method called network-based method has achieved great success in predicting drug-target interactions (DTIs), with two advantages, i.e., independence on the three-dimensional structure of the target and negative samples.

[0005] Therefore, integrating the computational target profiles obtained by network-based methods and machine learning-based methods to predict potential endocrine disruptors can be a novel and effective strategy. SUMMARY

[0006] The present application provides a model for predicting endocrine disruptors by integrating pharmacological and toxicological profiles, and a method for constructing the model and applications thereof.

[0007] In a first aspect of the present application, a method for constructing an endocrine disruptor prediction model is provided, comprising the following steps:

[0008] (S1) providing a first data set comprising data information of determined endocrine disruptors and determined non-endocrine disruptors, and the data information comprises (i) structure information and (ii) molecular fingerprint information corresponding to the compounds of the endocrine disruptors and non-endocrine disruptors; wherein the molecular fingerprint is based on the molecular fingerprint of the substructure of the compound;

[0009] (S2) constructing a substructure-drug-target network, and then constructing a network prediction model based on the substructure-drug-target network according to a network inference algorithm, and inputting the endocrine disruptors and non-endocrine disruptors in the first data set into the prediction model to obtain a network-based target profile (NBTP) of the compounds;

[0010] (S3) providing endocrine-related toxicology-related target activity data; based on the target activity data, using a machine learning method to construct a machine learning-based prediction model; and inputting the endocrine disruptors and non-endocrine disruptors in the first data set into the machine learning-based prediction model to obtain a machine learning-based target profile (MBTP) of the compounds;

[0011] (S4) combining the network-based target profile and the machine learning-based target profile to generate a combined target profile (CTP);

[0012] (S5) using a machine learning method to construct an endocrine disruptor prediction model for the combined target profile, the endocrine disruptor prediction model being used to evaluate whether a to-be-tested compound is an endocrine disruptor.

[0013] In another preferred embodiment, the order of steps (S2) and (S3) can be interchanged or performed simultaneously.

[0014] In another preferred embodiment, in step (S5), the endocrine disruptors and non-endocrine disruptors in the first data set are characterized by using molecular features of each compound in the combined target profile, and based on the molecular features, a machine learning method is used to build the endocrine disruptor prediction model; wherein the molecular features include molecular fingerprints and NBTPs, or molecular fingerprints and CTPs.

[0015] In another preferred embodiment, in step (S1), the following sub-steps are included:

[0016] (S1a) Collect information of endocrine disruptors and non-endocrine disruptors from public databases or screening projects;

[0017] (S1b) According to the endocrine disruptors and non-endocrine disruptors, match the structural information of these compounds; data processing is performed on the structure of the compounds, and the specific steps are as follows: desalination, standardization of coordination bond, de-mixture, and retention of compounds with one carbon or more;

[0018] (S1c) Collect marketed oral drugs from drug databases (such as DrugBank database) for supplementing non-endocrine disruptors;

[0019] (S1d) Calculate the MACCS fingerprint, PubChem fingerprint, KR (Klekota-Roth) fingerprint, FP4 (Substructure) fingerprint, CDK fingerprint, and AP2D (Atom Pairs 2D) fingerprint of the compounds on the PaDEL-Descriptor software; calculate two connectivity fingerprints, including ECFP4 fingerprint and FCFP4 fingerprint, on the RDKit software.

[0020] In another preferred embodiment, in step (S2), the following sub-steps are included:

[0021] (S2a) Construct a DTI network using known DTI data; based on the chemical structure information of drugs in the DTI network, calculate the MACCS fingerprint, PubChem fingerprint, KR fingerprint, FP4 fingerprint, and FCFP4 fingerprint of the compounds, and then obtain drug-substructure correlation associations according to the molecular fingerprints of the compounds, thereby constructing a substructure-drug interrelation network; integrate the DTI network and the substructure-drug interrelation network to construct a substructure-drug-target network;

[0022] (S2b) constructing a network prediction model based on network inference algorithm: according to the substructure-drug-target network, for any drug, the target nodes and substructure nodes connected to it are each assigned an initial resource with a certain weight, and an initial resource matrix based on network inference algorithm is constructed; then in each resource diffusion process, the substructure nodes and target nodes in the network that have initial resources, the existing resources of the nodes are evenly distributed to the neighbor nodes connected to them, and a transition matrix based on network inference algorithm is constructed according to the number of resource diffusion; based on the transition matrix and the substructure-drug-target network, a network prediction model is constructed;

[0023] (S2c) constructing a substructure-compound network based on endocrine disruptors and non-endocrine disruptors; inputting the network into the network prediction model, thereby obtaining the predicted scores of endocrine disruptors and non-endocrine disruptors with all protein targets in the DTI network as the network-based target profile (NBTP).

[0024] In another preferred example, in step (S2), the following sub-steps are included:

[0025] (S2a) providing drug-target interaction (DTI) data and chemical structure information of the corresponding drug, the chemical structure information including substructure information corresponding to the drug;

[0026] (S2b) constructing a substructure-drug-target network based on the drug-target interaction data and the chemical structure information of the drug;

[0027] (S2c) constructing a network prediction model based on the substructure-drug-target network according to network inference algorithm;

[0028] (S2d) inputting endocrine disruptors and non-endocrine disruptors in the first data set into the prediction model to obtain the network-based target profile (NBTP) of the compounds.

[0029] In another preferred example, in step (S3), the following sub-steps are included:

[0030] (S3a) collecting different target activity data sets related to endocrine disruption from a toxicology data set (such as Tox21, Toxicology in the 21st Century); wherein the original data set has been pre-processed including the following steps: desalination, standardization of coordination bond, reservation of compound data with carbon atom number greater than 1, removal of duplicate data and deletion of ambiguous label data;

[0031] (S3b) using random sampling, dividing each of the different target activity data sets into a training set and a test set in a ratio of 4:1;

[0032] wherein, on the training set, 5-fold cross-validation and grid search are used to find the optimal parameters on each target, and then the optimal prediction model is built for each target according to the optimal parameters, and verified on the test set;

[0033] (S3c) Based on the divided different target activity data sets, a prediction model is built using machine learning methods including Support vector machines (SVM), Decision tree (DT), Random forest (RF), K nearest neighbors (kNN), Linear regression (LR) and Extreme Gradient Boosting (XGB).

[0034] In another preferred example, the specific steps of generating the combined target profile are as follows:

[0035] The network-based target profile and the machine learning-based target profile are combined to obtain a combined target profile (CTP), which is represented by the following mathematical formula:

[0036] Z = concat(X, Y),

[0037] wherein Z is a matrix of the CTP containing all compounds, X is a matrix of the NBTP containing all compounds, and Y is a matrix of the MBTP containing all compounds.

[0038] In another preferred example, in step (S5), the following sub-steps are included:

[0039] (1) The first data set is divided into a training set and a test set in a ratio of 4:1; the compounds in the data set are represented using the following molecular features: molecular fingerprint and NBTP, or molecular fingerprint and CTP, respectively;

[0040] wherein, in the training set, grid search and five-fold cross-validation are used to find the optimal parameters of these models; then the optimal model is built according to the optimal parameters, and verified on the test set;

[0041] (2) Based on the three different molecular features of the compounds, an endocrine disruptor prediction model is built using machine learning methods including SVM, DT, RF, kNN, LR and XGB.

[0042] In another preferred example, the compounds in the first data set are represented using molecular fingerprint and CTP.

[0043] In another preferred embodiment, the following common evaluation indexes are used to evaluate the performance of each of the models: accuracy, precision, recall, F1 score, Matthews correlation coefficient (MCC), and area under the curve (AUC) of the receiver operating characteristic curve; wherein the AUC is used to select the optimal model.

[0044] In a second aspect of the present application, an evaluation method for evaluating whether a test compound is an endocrine disruptor is provided, comprising the steps of:

[0045] (a) providing a test compound;

[0046] (b) inputting the test compound into an endocrine disruptor prediction model constructed by the construction method of claim 1;

[0047] (c) evaluating the test compound by using the endocrine disruptor prediction model, thereby obtaining an evaluation result of whether the test compound is an endocrine disruptor; and

[0048] (d) outputting the evaluation result.

[0049] In another preferred embodiment, the test compound is selected from the group consisting of a determined endocrine disruptor, a determined non-endocrine disruptor, and an unknown compound.

[0050] In a third aspect of the present application, an endocrine disruptor prediction system for evaluating whether a test compound is an endocrine disruptor is provided, comprising:

[0051] (M1) an input module: the input module is configured to input information of a test compound;

[0052] (M2) an evaluation module: the evaluation module is configured to evaluate the test compound based on an endocrine disruptor prediction model, thereby obtaining an evaluation result of whether the test compound is an endocrine disruptor; wherein the endocrine disruptor prediction model is constructed by a machine learning method based on molecular fingerprints of compound substructures and combined target profiles; the combined target profiles are generated by combining network-based target profiles and machine learning-based target profiles;

[0053] Alternatively, the endocrine disruptor prediction model is an endocrine disruptor prediction model constructed by the construction method of claim 1; and

[0054] (M3) output module: the output module is configured to output the evaluation result.

[0055] In another preferred example, the evaluation module comprises the following sub-modules:

[0056] (M2a) calculation feature sub-module: when the search module fails to search for a relationship between the input compound and the known endocrine disruptor, i.e., the input compound is completely new, the calculation feature sub-module calculates the network-based target spectrum for the new input compound based on the constructed prediction model and the substructure-compound correlation; and then constructs a relevant prediction model according to the optimal parameters of the prediction model of the endocrine disruptor-related target, and calculates the machine learning-based target spectrum for the new input compound, respectively;

[0057] (M2b) prediction sub-module: when the input compound is completely new, the prediction sub-module combines the NBTP and MBTP of the known drug and the new input compound in the storage module to obtain a combined target spectrum; and then uses these molecular features (NBTP and CTP) and molecular fingerprints to characterize the drug and the new compound, and constructs a machine learning model based on the optimal parameters of the endocrine disruptor prediction model to calculate the endocrine disruptor label of the new input compound and determine the endocrine disruptor potential thereof.

[0058] In another preferred example, the system further comprises the following modules:

[0059] (M4) search module: the search module is used to search for a relationship between the input compound and the collected endocrine disruptor or non-endocrine disruptor; if the search is successful, the relevant information is called from the storage module and directly enters the output display module; if the search fails, the molecular fingerprint is calculated;

[0060] (M5) construction module: the construction module constructs a DTI network depending on the collected DTI data, and obtains a substructure-compound correlation network according to the molecular fingerprint, and then constructs a network prediction model based on the substructure-drug-target network according to the network reasoning-based algorithm;

[0061] (M6) storage module: the storage module is used to store the substructure-drug-target network, the network reasoning-based algorithm, the optimal parameters of the prediction model of the endocrine disruptor-related target, the optimal parameters of the endocrine disruptor prediction model, the three molecular features (including the molecular fingerprint, the NBTP and the MBTP) of the known compound, the statistical test method, and the information of the known endocrine disruptor;

[0062] (M7) key protein identification module: the key protein identification module uses the statistical test method to identify the endocrine disruptor-related key protein according to the combined target spectrum of the optimal prediction model in the prediction module;

[0063] (M8) Optional display module: the display module is configured to display the predicted key proteins related to endocrine disruptors, the combined target spectrum and the prediction results.

[0064] In another preferred embodiment, the system further comprises (M9) a control module: the control module is configured to control the operation of each module.

[0065] In another preferred embodiment, the control module is configured to control the operation of the input module, the evaluation module, the output module, the search module, the construction module, the storage module, the calculation feature module, the construction module, the prediction module, the key protein identification module, and the display module.

[0066] In another preferred embodiment, the input module is further configured to input information of known endocrine disruptors or non-endocrine disruptors or structural information of compounds to be predicted.

[0067] In another preferred embodiment, the storage module is further configured to record and store the output results of each module in real time, while providing the calling of relevant information.

[0068] It should be understood that, within the scope of the present application, each of the above technical features of the present application and each of the technical features specifically described below (e.g., in the examples) can be combined with each other to form new or preferred technical solutions. Due to the limited space, they will not be listed one by one here. BRIEF DESCRIPTION OF DRAWINGS

[0069] Figure 1 A flow chart showing the strategy of integrating pharmacology and toxicology spectrum to predict endocrine disruptors.

[0070] Figure 2 a. Percentage of 1844 target proteins in NBTP in different target protein types; b. Distribution of EDC and non-EDC prediction scores of NBTP on nine target protein types; c. Distribution of EDC and non-EDC prediction scores of NBTP in 39 nuclear receptors; d. AUC values of 11 endocrine disruptor-related target prediction models in 5-fold cross-validation and test set validation; e. Distribution of EDC and non-EDC prediction scores of MBTP in 11 endocrine disruptor-related targets.

[0071] Figure 3 AUC values of the optimal prediction model constructed by three molecular features (molecular fingerprint, NBTP, MBTP) in 5-fold cross-validation.

[0072] Figure 4 a. Pathway enrichment analysis results of 267 target proteins discovered on KEGG Pathway; b. Pathway enrichment analysis results of 267 target proteins discovered on WikiPathway.

[0073] Figure 5 The experimental labels of nine potential EDCs on five nuclear receptors are shown. Among them, 1: the prediction result of the optimal prediction model is consistent with the experimental result in the literature; 0: the prediction result of the optimal prediction model is inconsistent with the experimental result in the literature; N / A: these compounds have no clear interaction with the target in the literature.

[0074] Figure 6 The prediction ranking results of nine compounds in the steroid hormone biosynthesis pathway are shown; among them, the higher the ranking, the greater the possibility of the compound and the target, and vice versa; the white part means that the compound has no clear interaction with the target in the literature.

[0075] Figure 7 The relationship between the average AUC value of the EDC prediction model in 5-fold cross-validation and different fingerprints is shown, including (a) NBTP, (b) NBTP_norm, (c) molecular fingerprint (Fingerprint, FP).

[0076] Figure 8 The generation schematic diagram of drug-substructure correlation is shown. DETAILED DESCRIPTION

[0077] After extensive and in-depth research, a high-accuracy and precision endocrine disruptor prediction model is developed for the first time. Specifically, by combining the knowledge of network pharmacology and computational toxicology, a combination target spectrum is constructed to represent the compounds by applying network-based methods and machine learning methods, and a prediction model is constructed by applying machine learning methods to predict potential endocrine disruptors. The present inventors have completed the present invention on this basis.

[0078] Endocrine disruptor

[0079] Endocrine disruptors (EDCs) are defined as exogenous compounds that interfere with the normal physiological processes of blood-borne hormones in the body during growth, development, and reproduction, leading to dysfunction of the endocrine system. The mechanism of action of endocrine disruptors is complex, including but not limited to targeting hormone receptors such as nuclear receptors and affecting hormone-responsive cells to interfere with the synthesis, transport, action, metabolism, and clearance of hormones. In addition, exposure to endocrine disruptors can lead to adverse events, including birth defects, reproductive dysfunction, obesity, cancer, diabetes, and neurodevelopmental changes.

[0080] Endocrine disruptor prediction model construction method

[0081] The present application provides a method for constructing an endocrine disruptor prediction model, comprising the following steps:

[0082] (S1) providing a first dataset comprising data information of determined endocrine disruptors and determined non-endocrine disruptors, and the data information includes (i) structure information and (ii) molecular fingerprint information corresponding to the compounds of the endocrine disruptors and non-endocrine disruptors; wherein the molecular fingerprint is based on the molecular fingerprint of substructures of the compounds;

[0083] (S2) constructing a substructure-drug-target network, and based on the substructure-drug-target network, constructing a network prediction model according to a network inference algorithm, and inputting the endocrine disruptors and non-endocrine disruptors in the first dataset into the prediction model to obtain a network-based target profile (NBTP) of the compounds;

[0084] (S3) providing endocrine-related toxicology-related target activity data; based on the target activity data, constructing a machine learning-based prediction model using a machine learning method; and inputting the endocrine disruptors and non-endocrine disruptors in the first dataset into the machine learning-based prediction model to obtain a machine learning-based target profile (MBTP) of the compounds;

[0085] (S4) combining the network-based target profile and the machine learning-based target profile to generate a combined target profile (CTP);

[0086] (S5) using a machine learning method to construct an endocrine disruptor prediction model for the combined target profile, the endocrine disruptor prediction model being used to assess whether a to-be-tested compound is an endocrine disruptor.

[0087] In another preferred example, the order of steps (S2) and (S3) can be interchanged or performed simultaneously.

[0088] In another preferred example, in step (S2), the following sub-steps are included:

[0089] (S2a) providing drug-target interaction (DTI) data and chemical structure information of the corresponding drugs, the chemical structure information including substructure information corresponding to the drugs;

[0090] (S2b) constructing a substructure-drug-target network based on the drug-target interaction data and the chemical structure information of the drugs;

[0091] (S2c) based on the substructure-drug-target network, constructing a network prediction model according to a network inference algorithm;

[0092] (S2d) inputting the endocrine disruptors and non-endocrine disruptors in the first data set into the prediction model to obtain a network-based target profile (NBTP) of the compound.

[0093] In another preferred embodiment, in step (S5), the endocrine disruptors and non-endocrine disruptors in the first data set are characterized by molecular features of each compound in the combined target profile, and based on the molecular features, a machine learning method is used to build the endocrine disruptor prediction model; wherein the molecular features include molecular fingerprints and NBTP, or molecular fingerprints and CTP.

[0094] Endocrine disruptor prediction system

[0095] The present application also provides an endocrine disruptor prediction system (or device) for evaluating whether a test compound is an endocrine disruptor, the system comprising:

[0096] (M1) an input module: the input module is configured to input information of a test compound;

[0097] (M2) an evaluation module: the evaluation module is configured to evaluate the test compound based on the endocrine disruptor prediction model, thereby obtaining an evaluation result of whether the test compound is an endocrine disruptor; wherein the endocrine disruptor prediction model is built by a machine learning method based on molecular fingerprints of compound substructures and combined target profiles; the combined target profile is generated by combining a network-based target profile and a machine learning-based target profile;

[0098] or the endocrine disruptor prediction model is an endocrine disruptor prediction model built by the building method of claim 1; and

[0099] (M3) an output module: the output module is configured to output the evaluation result.

[0100] In another preferred embodiment, the evaluation module comprises the following sub-modules:

[0101] (M2a) a feature calculation sub-module: when the search module does not search for the relationship between the newly input compound and the known endocrine disruptors, i.e. the input is a brand new compound, the network-based target profile is calculated for the newly input compound based on the built prediction model and the substructure-compound correlation; then the relevant prediction model is built according to the optimal parameters of the endocrine disruptor-related target prediction model, and the machine learning-based target profile is calculated for the newly input compound;

[0102] (M2b) prediction submodule: the prediction submodule combines the NBTP and MBTP of known drugs and new input compounds in the storage module to obtain a combined target spectrum; then uses these molecular characteristics (NBTP and CTP) and molecular fingerprints to characterize drugs and new compounds, and constructs a machine learning model based on the optimal parameters of the endocrine disruptor prediction model to calculate the endocrine disruption label of the new input compound and determine its endocrine disruption potential.

[0103] In another preferred example, the system further comprises the following modules:

[0104] (M4) search module: the search module is used to search for the relationship between the input compound and the collected endocrine disruptors or non-endocrine disruptors; if found, the relevant information is called from the storage module and directly enters the output display module; if failed, the molecular fingerprint is calculated;

[0105] (M5) construction module: the construction module constructs a DTI network depending on the collected DTI data, and obtains a substructure-compound correlation network according to the molecular fingerprint, and then constructs a network prediction model based on the substructure-drug-target network according to the network inference-based algorithm;

[0106] (M6) storage module: the storage module is used to store the substructure-drug-target network, the network inference-based algorithm, the optimal parameters of the prediction model of endocrine disruption related targets, the optimal parameters of the endocrine disruptor prediction model, the three molecular characteristics (including molecular fingerprints, NBTP and MBTP) of known compounds, statistical test methods, and information of known endocrine disruptors;

[0107] (M7) key protein identification module: the key protein identification module identifies the key proteins related to endocrine disruption using the statistical test method according to the combined target spectrum of the optimal prediction model in the prediction module;

[0108] (M8) optional display module: the display module displays the predicted key proteins related to endocrine disruption, the combined target spectrum and the prediction results.

[0109] In another preferred example, the system further comprises (M9) a control module: the control module is configured to control the operation of each module.

[0110] In another preferred example, the control module is configured to control the operation of the following modules: input module, evaluation module, output module, search module, construction module, storage module, feature calculation module, construction module, prediction module, key protein identification module, and display module.

[0111] In another preferred embodiment, the input module is further configured to input information of known endocrine disruptors or non-endocrine disruptors or structure information of compounds to be predicted. In another preferred embodiment, the storage module further records and stores the output results of each module in real time, while providing the calling of relevant information.

[0112] The main advantages of the present application include:

[0113] 1. The present application applies the knowledge of network pharmacology to the field of endocrine disruptor prediction, and creatively constructs a network-based target spectrum for characterizing compounds using a network-based method. Compared with common molecular fingerprints, the network-based target spectrum proposed in the present application can cover various targets, thereby expanding the application range of computational research on the mechanism of action of endocrine disruptors, rather than being limited to a few common nuclear receptors.

[0114] 2. The present application combines the knowledge of network pharmacology and computational toxicology, and respectively applies a network-based method and a machine learning method to construct a combined target spectrum for characterizing compounds, and applies a machine learning method to construct a prediction model to predict potential endocrine disruptors. Compared with the network-based target spectrum, this combined target spectrum can improve the prediction performance of the prediction model on endocrine disruptors by introducing the knowledge of computational toxicology.

[0115] 3. Compared with traditional molecular fingerprints, the molecular features, i.e., the combined target spectrum, constructed by the present application, and the prediction model constructed by the present application, both show better performance in the results of the training set and the test set, indicating the high accuracy of the prediction model of endocrine disruptors. At the same time, the prediction accuracy of the endocrine disruptor prediction model reaches 80% in the case study of 15 new endocrine disruptors, indicating that this molecular feature exhibits good generalization ability in practical application.

[0116] 4. Compared with the method based on molecular docking, the prediction model developed by the present application is simple to calculate and can efficiently output comprehensive prediction results. It only takes 40 seconds, 1 minute and 8 minutes to predict the labels and combined target spectrum of 10, 100 and 999 compounds, respectively, wherein the combined target spectrum contains 1855 target information, which means that toxicology researchers can make decisions based on this target information to determine whether further in vitro and in vivo tests are needed. Therefore, the prediction model developed by the present application shows the potential for large-scale prediction of endocrine disruptors and helps to explore the potential mechanism of action of endocrine disruptors.

[0117] The present application will be further described in conjunction with specific examples. It should be understood that these examples are only used to illustrate the present application and not used to limit the scope of the present application. The experimental methods in the following examples without specific conditions are generally according to the conventional conditions, or according to the conditions suggested by the manufacturers.

[0118] Example 1. Construction of endocrine disruptor prediction model

[0119] The flow chart of constructing the endocrine disruptor prediction model of the present application is shown in Figure 1 The specific process is as follows.

[0120] 1.1 Collection and processing of endocrine disruptors

[0121] The data of endocrine disruptors used in the present application is derived from four public databases of endocrine disruption part of Eurpean commission, TEDX, EDKB, DEDuCT (version 1.0). After data processing, finally 1334 endocrine disruptors and 102 non-endocrine disruptors were obtained.

[0122] Then, the present application collected the marketed oral drugs from DrugBank database as non-endocrine disruptors, and finally 1334 endocrine disruptors and 1474 non-endocrine disruptors were obtained.

[0123] Based on the structure information of the above compounds, the MACCS fingerprint, PubChem fingerprint, KR fingerprint, FP4 fingerprint, CDK fingerprint, AP2D fingerprint of the compounds were calculated using PaDEL-Descriptor (version 2.21), and the ECFP4 fingerprint and FCFP4 fingerprint of the compounds were calculated using RDKit (version 2018.09). The data of endocrine disruptors were stored in the storage module.

[0124] 1.2 Generation and analysis of network-based target profile

[0125] 1.2.1 Generation of network-based target profile

[0126] Firstly, in terms of network construction: the drug-target network of the present application is derived from Chemical Science 2022, 13: 1060-1079, containing 45018 drug-target interactions connecting 12751 drugs and 1844 target proteins. In constructing the substructure-drug correlation network, the MACCS fingerprint of all drugs in the DTI network is calculated to represent the substructure of the drug, and the drug and substructure are connected to construct the drug-substructure correlation network. The basic attributes of the constructed drug-substructure network are shown in Table 1, and the drug-substructure correlation is shown in Figure 8

[0127] Secondly, constructing the prediction model: in order to obtain a target spectrum that can cover more target types, the present application uses a network-based DTI prediction method to construct a prediction model, and the specific method and model hyperparameters are derived from the literature British Journal of Pharmacology 2016, 173: 3372-3385.

[0128] Finally, generating network-based target spectrum: after constructing the model, the molecular fingerprints of endocrine disruptors and non-endocrine disruptors are calculated, and the substructure- compound network is constructed. Input this compound-substructure network into the network-based prediction model, and the predicted scores of each compound with the 1844 target proteins are obtained as the network-based target spectrum. These network-based target spectra are used for subsequent analysis and stored in the storage module.

[0129] 1.2.2 Analysis of network-based target spectrum

[0130] In order to show the diversity of the network-based target spectrum, the 1844 target proteins are classified according to the categories of the IUPHAR / BPS Guide to PHARMACOLOGY.

[0131] As Figure 2 ​As shown in FIG. a, these target proteins cover a variety of target types, including 857 enzymes (Enzymes), 207 G protein-coupled receptors (GPCRs), 39 nuclear receptors (NRs), 82 catalytic receptors, 78 transporters, 47 ligand-gated ion channels (LGICs), 70 voltage-gated ion channels (VGICs), 5 other ICs, and 460 other proteins. While the 39 nuclear receptors only account for a small fraction of the total 1844 target proteins (2.1%), they have covered more than 80% of the nuclear receptors.

[0132] Based on the above 9 different protein target types, the predicted scores in the NBTPs of endocrine disruptors and non-endocrine disruptors were compared, resulting in Figure 2 b.

[0133] On all target types, endocrine disruptors have higher predicted scores than non-endocrine disruptors. This phenomenon is most obvious on nuclear receptors compared to the other eight target types, which means that endocrine disruptors tend to target nuclear receptors, consistent with the known mechanism of action of endocrine disruptors.

[0134] In addition, the 39 nuclear receptors were ranked according to the NBTP predicted scores of endocrine disruptors, resulting in Figure 2 c. Obviously, ER, AR, peroxisome proliferator-activated receptors (PPAR), glucocorticoid receptors (GR), and progesterone receptors (PR) rank in the top 10, respectively, which means that endocrine disruptors tend to target these nuclear receptors again. The difference between the predicted scores of endocrine disruptors and non-endocrine disruptors on ER and AR is greater than 0.2, further proving the importance of these two nuclear receptors in endocrine disruption.

[0135] In summary, these results show that this network-based target profile not only covers a variety of protein targets, but also helps to discover the potential mechanism of action of endocrine disruption.

[0136] 1.3 Generation and analysis of target profiles based on machine learning

[0137] 1.3.1 Generation of target profiles based on machine learning

[0138] The present application is based on the data set of 11 endocrine disruptor related target activities from Tox21 AR, AR-LBD, AhR, Aromatase, ER, ER-LBD, PPAR-gamma, ARE, ATAD5, HSE, p53 in the literature Journal of Medicinal Chemistry, 2021, 64, 10:6924-6936. After data processing and division, based on the 11 target activity data sets, six machine learning methods are used to build models and perform five-fold cross-validation and test set validation.

[0139] The AUC results of the optimal model of 11 respective targets in five-fold cross-validation and test set are shown in Figure 2 d. As shown in the figure, the 11 optimal prediction models have good performance in both validation methods, with AUC greater than 0.8. Among them, the model results for AR and AhR are the best, with AUC values reaching 0.89 in five-fold cross-validation and test set validation.

[0140] Based on the 11 optimal prediction models built, the machine learning-based target profiles are predicted for endocrine disruptors and non-endocrine disruptors. These machine learning-based target profiles are used for subsequent analysis and stored in the storage module.

[0141] 1.3.2 Analysis of machine learning-based target profiles

[0142] According to the 11 optimal prediction models built, the prediction scores of MBTPs for endocrine disruptors and non-endocrine disruptors are generated.

[0143] The results are shown in Figure 2 e. The results show that in almost all targets, the prediction scores of EDCs are higher than those of non-EDCs, especially in PPAR-gamma, AR, and AhR receptors, which shows the potential in predicting endocrine disruptors.

[0144] 1.4 Construction of endocrine disruptor prediction model using five molecular features

[0145] 1.4.1 Based on the NBTP and MBTP calculated previously, a combined target profile (CTP) is generated by combining (or splicing).

[0146] The two molecular features (NBTP and CTP) can be used for subsequent construction of endocrine disruptors and stored in the storage module.

[0147] 1.4.2 Construction of endocrine disruptor prediction model

[0148] AsFigure 3 To compare the influence of different molecular features on the performance of EDC prediction models, the inventors evaluated the performance of models built with three different molecular features, as shown in Table 5.

[0149] As shown in Figure 7 a, for models built with network-based target profiles, among five molecular fingerprint (FP)-based network-based target profiles, the model built with MACCS fingerprint-based target profiles, i.e., MACCS-NBTP, performed the best.

[0150] As shown in Figure 7 b, among five prediction models built with normalized network-based target profiles, the model built with MACCS fingerprint-based target profiles, i.e., MACCS-NBTP_norm, also performed the best.

[0151] Therefore, in the subsequent model comparison and combination target profile construction, the inventors used both MACCS-NBTP and MACCS-NBTP_norm to build CTP and CTP_norm. As shown in Figure 7 c, the MACCS fingerprint-based model outperformed the other seven fingerprints.

[0152] Based on the three feature types (molecular fingerprints, NBTPs, and CTPs), the inventors used six machine learning methods to build 18 best models. Then, the performance of these optimized models was further explored.

[0153] As shown in Figure 3 , the models built with CTPs outperformed the models built with NBTPs. This indicates that the generation of CTPs by combining NBTPs with MBTPs indeed helps EDC prediction.

[0154] In addition, the models built with CTPs also outperformed the models built with fingerprints, further demonstrating the advantage of CTPs in EDC prediction.

[0155] Finally, the inventors selected the model built with CTP and XGB method as the best model. The AUC value of this model reached 0.92( Figure 3 ), indicating its high performance.

[0156] To further demonstrate its high performance, the inventors evaluated the best model and other best models by the previously divided test set. As shown in Table 1, the best model (CTP-based model) showed a high AUC = 0.907, indicating its high generalization ability.

[0157] Table 1. Results of the best endocrine disruptor prediction models built with five molecular features on the test set

[0158] Molecular features Method Accuracy Precision Recall F1 score MCC AUC FP RF 0.824 0.833 0.779 0.806 0.646 0.900 NBTP XGB 0.851 0.865 0.806 0.835 0.700 0.909 NBTP_norm XGB 0.836 0.838 0.806 0.822 0.671 0.889 CTP XGB 0.840 0.865 0.779 0.820 0.679 0.907 CTP_norm XGB 0.843 0.857 0.798 0.827 0.686 0.907

[0159] Example 2. Identification of key proteins related to endocrine disruptors

[0160] To evaluate the significance of the proposed combined target profile in helping to explain the mechanism of action of endocrine disruptors, key targets were identified on the combined target profile consisting of 1855 targets. The specific steps are as follows: Wilcoxon rank-sum test was performed on the combined target profile of the optimal model and corrected, and P values < 0.001 were used as the screening condition, and finally 277 key target proteins were obtained, and ranked according to P values.

[0161] According to the ranking results as shown in Table 2, it was found that 7 and 9 of the top 10 and 20 key targets, respectively, were included in the MBTP, which indicated the importance of the MBTP in the endocrine disruptor prediction model.

[0162] Similar to the results obtained by the two target profile analyses, AR, AhR and PPAR-gamma targets were found to be important key targets, and these targets have also been widely proven to play an important role in endocrine disruption.

[0163] Table 2. Ranking results of the top 20 key proteins according to P values

[0164] Target name Odds ratio P value Adjusted P value AR 1019670 1.37E-143 2.54E-140 HSE 1003940 1.88E-132 1.75E-129 Aromatase 1002830 1.11E-131 6.85E-129 PPAR-gamma 992562 1.11E-124 5.16E-122 AhR 982526 5.06E-118 1.88E-115 Q02318 966064 1.68E-107 5.19E-105 p53 960246 6.66E-104 1.76E-101 Q15466 941259 1.36E-92 3.16E-90 ATAD5 933698 2.85E-88 5.87E-86 O00748 924982 2.00E-83 3.72E-81 P15374 921216 2.26E-81 3.81E-79 ARE 910908 6.86E-76 1.06E-73 P55055 908834 8.23E-75 1.17E-72 O15357 901306 5.85E-71 7.75E-69 P23141 900980 8.55E-71 1.06E-68 O15392 890160 1.90E-65 2.20E-63 Q15119 885774 2.42E-63 2.64E-61 P23470 885715 2.59E-63 2.66E-61 Q07817 882968 5.17E-62 5.04E-60 Q8WWX8 882258 1.11E-61 1.03E-59

[0165] The 267 important target proteins found previously (excluding the 10 targets included in the MBTP) were input into the DAVID website for pathway enrichment, and KEGG Pathway and WikiPathway pathway analysis results were obtained, respectively, and P values < 0.05 were used as the screening condition to obtain the results as shown in Table 2. Figure 4 a.

[0166] Figure 4 a. WikiPathway pathway enrichment results, after screening, these important proteins were significantly enriched in 8 important pathways, among which the first ranked nuclear receptor pathway and the fifth ranked nuclear receptor meta-pathway and the seventh ranked estrogen signaling pathway are all important pathways known to be related to endocrine disruption, which directly proves that the proposed combined target profile can indeed help to elucidate the mechanism of action related to endocrine disruption.

[0167] In addition, some proteins in the third ranked prostaglandin synthesis and regulation pathway were also found to be related to endocrine disruption, which will be studied as a potential mechanism of action in the future.

[0168] AndFigure 4 b is the pathway enrichment result on KEGG Pathway, from the figure we can find that 267 important target proteins are significantly enriched into 14 important pathways, among which the third ranked steroid hormone biosynthesis and the fourth ranked ovarian steroidogenesis are both related to endocrine disruptors. The steroid hormone biosynthesis pathway plays an important role in endocrine disruption.

[0169] Example 3. Application of endocrine disruptor prediction model

[0170] According to the important pathways and target proteins obtained in the identification step of endocrine disruptor related key proteins, two case studies were carried out to prove the practical application value of the prediction model.

[0171] 3.1 Case study 1

[0172] The first case study is to predict the endocrine disruptors related to nuclear receptors, based on the review Environment International 2021, 153, 106550, five nuclear receptors are selected for case study, including AR, ER, AhR, GR, PR.

[0173] Ten endocrine disruptors were collected from this review and other literature and input into the optimal model for prediction. Then analyze the CTP generated by these endocrine disruptors, among which three targets can be found in the MBTP part of CTP, and the last two targets can be found in the NBTP part of CTP.

[0174] The final results of these endocrine disruptors in the prediction model are shown in Figure 5 The prediction model achieved an accuracy of 75%, 71.4% and 50% on AR, ER and AhR respectively.

[0175] 3.2 Case study 2

[0176] The second case study is to predict endocrine disruptors of other mechanisms. Considering that the target proteins in this case study are all from the NBTP part of CTP, the top 50 targets are used as a threshold to analyze the relationship between 14 endocrine disruptors and these target proteins.

[0177] First of all, the study of steroid hormone biosynthesis pathway related mechanism, using the prediction model to predict the related targets of 9 potential endocrine disruptors, including CYP19, 17β-HSD, CYP11B1, CYP11B2.

[0178] The results are shown in Figure 6As shown in Table 3, the prediction results of nine potential endocrine disruptors in the prostaglandin synthesis and regulation pathway were shown. It was found that the optimal prediction model had good prediction ability, and the accuracy rates for CYP19, 17β-HSD, CYP11B1 and CYP11B2 were 100%, 66.7%, 66.7% and 83.3%, respectively. Then, the present application also studied the synthesis and regulation pathway of prostaglandin. As shown in Table 3, mainly for the PTGS2 target protein, it was found that the model could successfully predict this target for the five compounds.

[0179] In summary, the case analysis results show that the target spectrum proposed in the present application can not only find nuclear receptor related endocrine disruptors, but also help to find endocrine disruptors of other mechanisms.

[0180] In addition, all 15 endocrine disruptors used in this part (in which the compounds repeated in the data set have been deleted) were input into the optimal prediction model to determine whether they are endocrine disruptors, and 0.5 was selected as the threshold.

[0181] As shown in Table 4, the prediction accuracy of the 15 endocrine disruptors reached 80%, further indicating the high performance of the prediction model and the practicability of the present application.

[0182] Table 3. Prediction ranking results of nine potential endocrine disruptors in the prostaglandin synthesis and regulation pathway

[0183] Compound name PTGS2 rank Bifenthrin 22 Dichlorodiphenyltrichloroethane 28 Phthalic acid (2-ethylhexyl) ester 34 Diethyl phthalate 21 Mono-2-ethylhexyl ester 28

[0184] Table 4. Prediction results of EDC prediction model of all potential EDCs

[0185]

[0186]

[0187] Note: 1: The compound is considered to be an endocrine disruptor.

[0188] Table 5. Results of the optimal EDC prediction model constructed using five features respectively in 5-fold cross-validation

[0189]

[0190]

[0191] All the documents mentioned in the present application are cited as references in the present application, as if each document is cited as a reference individually. In addition, it should be understood that those skilled in the art can make various modifications or changes to the present application after reading the above teachings of the present application, and these equivalent forms also fall within the scope defined by the claims attached to the present application.

Claims

1. A method of constructing an endocrine disruptor prediction model, characterized by, The method comprises the following steps: (S1) providing a first data set containing data information of determined endocrine disruptors and determined non-endocrine disruptors, and the data information comprises (i) structure information and (ii) molecular fingerprint information of compounds corresponding to the endocrine disruptors and non-endocrine disruptors; wherein the molecular fingerprint is a molecular fingerprint based on a substructure of the compound; (S2) constructing a substructure-drug-target network, and then constructing a network prediction model based on the network inference algorithm, and inputting the endocrine disruptors and non-endocrine disruptors in the first data set into the prediction model, so as to obtain a network-based target spectrum, i.e. NBTP, of the compounds; wherein the following sub-steps are included: (S2a) constructing a DTI network by using known drug-target interaction data, i.e. DTI data; calculating MACCS fingerprint, PubChem fingerprint, KR fingerprint, FP4 fingerprint and FCFP4 fingerprint of the compounds based on the chemical structure information of the drugs in the DTI network, and then obtaining drug-substructure correlation association according to the molecular fingerprint of the compounds, so as to construct a substructure-drug correlation association network; and integrating the DTI network and the substructure-drug correlation association network to construct a substructure-drug-target network; (S2b) constructing a network prediction model based on the network inference algorithm: according to the substructure-drug-target network, for any drug, the target nodes and substructure nodes connected thereto are each assigned an initial resource with a certain weight to construct an initial resource matrix based on the network inference algorithm; then in each resource diffusion process, the substructure nodes and target nodes in the network that have initial resources evenly distribute the existing resources to the neighbor nodes connected thereto, and then a transfer matrix based on the network inference algorithm is constructed according to the number of resource diffusion; and based on the transfer matrix and the substructure-drug-target network, a network prediction model is constructed; (S2c) constructing a substructure-compound network based on the molecular fingerprint of the endocrine disruptors and non-endocrine disruptors; inputting the network into the network prediction model, so as to obtain the predicted scores of the endocrine disruptors and non-endocrine disruptors with all protein targets in the DTI network as a network-based target spectrum, i.e. NBTP; (S3) providing endocrine-related toxicology-related target activity data; based on the target activity data, a machine learning method is used to construct a machine learning-based prediction model; and the endocrine disruptors and non-endocrine disruptors in the first data set are input into the machine learning-based prediction model to obtain a machine learning-based target spectrum, i.e. MBTP, of the compounds; (S4) combining the network-based target spectrum and the machine learning-based target spectrum to generate a combined target spectrum, i.e. CTP; and (S5) using a machine learning method to construct an endocrine disruptor prediction model for the combined target spectrum, and the endocrine disruptor prediction model is used to evaluate whether a to-be-tested compound is an endocrine disruptor.

2. The method of claim 1, wherein, The order of steps (S2) and (S3) can be interchanged or performed simultaneously.

3. The method of claim 1, wherein, In step (S1), the following sub-steps are included: (S1a) Collect information of endocrine disruptors and non-endocrine disruptors from public databases or screening projects; (S1b) Match the structure information of these compounds according to the endocrine disruptors and non-endocrine disruptors; data processing is performed on the structure of the compounds, and the specific steps are as follows: desalination, standardization of coordination bonds, removal of mixtures, and retention of compounds with one carbon atom or more; (S1c) Collect marketed oral drugs from drug databases for supplementing non-endocrine disruptors; (S1d) Calculate the MACCS fingerprint, PubChem fingerprint, KR fingerprint, FP4 fingerprint, CDK fingerprint, and AP2D fingerprint of the compounds on the PaDEL-Descriptor software; and calculate two connectivity fingerprints, including ECFP4 fingerprint and FCFP4 fingerprint, on the RDKit software.

4. The method of claim 1, wherein, In step (S3), the following sub-steps are included: (S3a) Collect different target activity data sets related to endocrine disruption from a toxicology data set; wherein the original data set is pre-processed by including the following steps: desalination, standardization of coordination bonds, retention of compound data with more than one carbon atom, removal of duplicate data, and deletion of ambiguous label data; (S3b) Use random sampling to divide each of the different target activity data sets into a training set and a test set in a ratio of 4:1; wherein, on the training set, 5-fold cross-validation and grid search are used to find the optimal parameters on each target, and then the optimal prediction model is constructed for each target according to the optimal parameters, and verified on the test set; (S3c) Based on the divided different target activity data sets, use machine learning methods including support vector machine (SVM), decision tree (DT), random forest (RF), k-nearest neighbor (kNN), linear regression (LR), and extreme gradient boosting (XGB) to construct prediction models.

5. The method of claim 1, wherein, The specific steps for generating the combined target spectrum are as follows: The combined target spectrum, i.e. CTP, is obtained by combining the network-based target spectrum and the machine learning-based target spectrum, and is represented by the following mathematical formula: , wherein Z is a matrix of CTP containing all compounds, X is a matrix of NBTP containing all compounds, and Y is a matrix of MBTP containing all compounds.

6. The method of claim 1, wherein, In step (S5), the following sub-steps are included: (1) Divide the first data set into a training set and a test set in a ratio of 4:1; and use the following molecular features to represent the compounds in the data set: molecular fingerprint and NBTP, or molecular fingerprint and CTP; wherein, in the training set, grid search and five-fold cross-validation are used to find the optimal parameters of these models; then the optimal model is constructed according to the optimal parameters, and verified on the test set; (2) Based on the three different molecular features of the compounds, use machine learning methods including SVM, DT, RF, kNN, LR, and XGB to construct endocrine disruptor prediction models.

7. The method of claim 1, wherein, The first data set is characterized using molecular fingerprints and CTP.

8. The method of claim 6, wherein, The performance of each of the models is evaluated using the following common evaluation metrics: accuracy, precision, recall, F1 score, Matthews correlation coefficient, and the area under the receiver operating characteristic curve, i.e., AUC; wherein the AUC is used to select the optimal model.

9. An evaluation method for evaluating whether or not a test compound is an endocrine disrupter, characterized by, The method comprises the steps of: (a) providing a test compound; (b) inputting the test compound into the endocrine disruptor prediction model constructed using the construction method of claim 1; (c) evaluating the test compound using the endocrine disruptor prediction model to obtain an evaluation result of whether the test compound is an endocrine disruptor; and (d) outputting the evaluation result. The test compound is selected from the group consisting of a determined endocrine disruptor, a determined non-endocrine disruptor, or an unknown compound.

10. The method of claim 9, wherein, The system comprises:

11. An endocrine disruptor prediction system for assessing whether a test compound is an endocrine disruptor, characterized by, (M1) an input module configured to input information of a test compound; (M2) an evaluation module configured to evaluate the test compound based on an endocrine disruptor prediction model to obtain an evaluation result of whether the test compound is an endocrine disruptor; wherein the endocrine disruptor prediction model is an endocrine disruptor prediction model constructed using the construction method of claim 1; and (M3) an output module configured to output the evaluation result. The evaluation module comprises the following sub-modules:

12. The system of claim 11, wherein, (M2a) a feature calculation sub-module: when the search module fails to search for a relationship between a newly input compound and a known endocrine disruptor, i.e., a brand-new compound is input, the feature calculation sub-module calculates a network-based target profile for the newly input compound based on the constructed prediction model and the substructure-compound correlation; then, a prediction model related to endocrine disruption is constructed according to the optimal parameters of the prediction model of the endocrine disruption-related target, and a machine learning-based target profile is calculated for the newly input compound; (M2b) a prediction sub-module: when a brand-new compound is input, the prediction sub-module combines the NBTP and MBTP of the known drug and the newly input compound in the storage module to obtain a combined target profile; then, the molecules are characterized using these molecular features, i.e., NBTP and CTP, and molecular fingerprints, and a machine learning model is constructed based on the optimal parameters of the endocrine disruptor prediction model, and the endocrine disruption label of the newly input compound is calculated, and the potential of endocrine disruption is determined. The system further comprises the following modules:

13. The system of claim 11, wherein, (M4) a search module: the search module is used to search for a relationship between an input compound and a collected endocrine disruptor or non-endocrine disruptor; if a relationship is found, the related information is called from the storage module and directly enters the output display module; if a relationship is not found, the molecular fingerprint is calculated; (M5) a construction module: the construction module constructs a DTI network based on the collected DTI data, and obtains a substructure-compound correlation network based on the molecular fingerprint, and then constructs a network prediction model based on the substructure-drug-target network according to a network reasoning-based algorithm; ​ (M6) storage module: the storage module is used for storing the substructure-drug-target network, the algorithm based on network inference, the optimal parameters of the prediction model of endocrine disruptor related targets, the optimal parameters of the endocrine disruptor prediction model, three molecular characteristics of known compounds, i.e. molecular fingerprints, NBTP and MBTP, statistical test methods and information of known endocrine disruptors; (M7) key protein identification module: the key protein identification module identifies endocrine disruptor related key proteins using statistical test methods according to the optimal prediction model in the prediction module. (M8) optional display module: the display module displays the predicted endocrine disruptor related key proteins, the combined target spectrum and the prediction results.

14. The system of claim 11, wherein, The system further comprises (M9) a control module: the control module is configured to control the operation of each module.

15. The system of claim 14, wherein, The control module is configured to control the operation of the input module, the evaluation module, the output module, the search module, the construction module, the storage module, the calculation feature module, the construction module, the prediction module, the key protein identification module and the display module.

16. The system of claim 11, wherein, The input module is further configured to input information of known endocrine disruptors or non-endocrine disruptors or structural information of compounds to be predicted.

17. The system of claim 15, wherein, The storage module further records and stores the output results of each module in real time, and provides calling of related information.