An endocrine disruptor prediction method based on harmful outcome path network
By constructing an endocrine disruptor prediction method based on harmful outcome pathway networks, and utilizing machine learning algorithms and experimental data, the problem of the lack of biological connections in traditional models is solved, achieving efficient qualitative and quantitative prediction of EDCs and meeting EU standards.
Patent Information
- Application Number
- CN202411725946.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-28
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2044-11-28
AI Technical Summary
Existing technologies lack methods that can simultaneously qualitatively identify and quantitatively classify endocrine disruptors (EDCs), and traditional data-driven models lack biological connections, making it difficult to meet the three EU standards.
By employing machine learning algorithms combined with adverse outcome pathway networks, qualitative and quantitative prediction models for targets were constructed. By acquiring experimental activity data and endocrine disruptor inventory data, the biological link between the toxicological mechanisms of EDCs and adverse outcomes was established, and qualitative and quantitative prediction models for targets were constructed to predict the endocrine disrupting effects of compounds and classify their hazards.
It achieves efficient qualitative and quantitative prediction of EDCs, improves prediction performance, and can transparently show the complete process from compound-protein interaction to individual phenotypic toxicity. The qualitative prediction ROC-AUC value reaches 0.90, and the quantitative prediction mean square error is controlled within 0.08, meeting the standards of government agencies.
Smart Images

Figure CN119811528B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of endocrinology, more specifically, to a method for predicting endocrine disruptors based on harmful outcome pathway networks. BACKGROUND
[0002] With the acceleration of industrialization, humans are exposed to exogenous chemicals, i.e. endocrine disruptors (EDCs). These chemicals can disrupt the complex endocrine system and have adverse effects on health. On March 31, 2023, the European Union (EU) passed a revision of the Classification, Labeling and Packaging Regulation, requiring substances and mixtures to be labeled as endocrine disruptors by 2028. Therefore, assessing the endocrine disrupting properties of more than 100,000 commercial chemicals is a major challenge that has yet to be addressed.
[0003] With the advancement of experimental techniques, the United States Environmental Protection Agency (US EPA), the Organization for Economic Cooperation and Development (OECD) and the EU have established a wide range of in vitro and in vivo testing techniques (such as the EDSP program and the ToxCastTM / Tox21 program) to monitor the endocrine disrupting mechanisms and adverse effects of chemicals. The field of artificial intelligence has made efforts to develop new machine learning algorithms and deep learning architectures, providing new opportunities for high-throughput virtual screening of EDCs. Many platforms (such as the VEGA platform and the Endocrine Disruptome) or models (such as CoMPARA, CERAPP and NRMEA) have been developed and are freely available. However, all experimental techniques and virtual screening methods are only targeted at a specific effect, either the endocrine disrupting mechanism or the physiologically measurable adverse effect. So far, the reasonable biological connection between the evidence still needs to be judged manually according to expert experience, which makes it difficult for tools to meet all three criteria at the same time.
[0004] Adverse outcome pathway networks (AOP networks) help to establish the above-mentioned causal links. They can condense multi-level toxicology information of molecular initiating events (MIEs), key events (KEs) and adverse outcomes (AOs) after chemical exposure observed in various laboratory experiments and human epidemiology. Although recent progress has been made in experimental and computational techniques, it is still challenging to develop and apply new technologies to identify EDCs, because there is a lack of biologically reasonable links between chemical effects detected at the macromolecular level in biological tissues and adverse effects and top outcomes traditionally used in EDC hazard identification and classification. In the face of this difficulty, a new technical solution is urgently needed, which can combine advanced artificial intelligence algorithms and hierarchical biological knowledge in AOP networks to build a predictive model with biological information, realize high-throughput screening, qualitative hazard identification and quantitative hazard classification of EDCs by encoding hierarchical biological knowledge (including toxicological mechanisms, adverse outcomes and biological causal links), meet the three standards of government agencies, improve the prediction performance, and at the same time reveal the endocrine disrupting patterns of EDCs.
[0005] Chinese patent application, application number CN201811597767.1, published on March 29, 2019, discloses a method for high-throughput screening of endocrine disruptors based on hierarchical warning structure, which extracts the first warning structure of active data compounds by using substructure frequency analysis and substructure proportion analysis; extracts the second warning structure of the compounds meeting the first warning structure by using SARpy software; extracts the third warning structure by using SARpy software; combines the first warning structure and the second warning structure to form an active prediction module, and screens out compounds with characteristic structures in the prediction module, and then screens out warning compounds with potential endocrine disrupting effects based on the second warning structure; takes the third warning structure as an interference activity prediction module, and then screens the interference activity based on the interference activity prediction module. However, the extraction of warning structure in this scheme is only based on molecular fingerprint and substructure analysis, and lacks consideration of the mechanism of action of endocrine disruptors and adverse outcomes. SUMMARY
[0006] 1. Technical problems to be solved
[0007] In view of the problem that the prior art lacks simultaneous qualitative identification and quantitative classification of EDCs in the European Union, the present application provides an endocrine disruptor prediction method based on a harmful outcome pathway network, which obtains experimental activity data and endocrine disruptor list data, and pre-processes the experimental activity data, constructs a target qualitative prediction model and an endocrine disruptor effect qualitative prediction model, and predicts whether a test compound has an endocrine disruptor effect; meanwhile, a target quantitative prediction model is constructed to predict the target quantitative activity of the compound with the endocrine disruptor effect, obtain a quantitative harmful outcome pathway, and identify sensitive harmful outcomes and sensitive quantitative harmful outcome pathways according to the quantitative harmful outcome pathway; and finally, the endocrine disruptor is classified according to the sensitive harmful outcomes and the sensitive quantitative harmful outcome pathways.
[0008] 2. Technical solution
[0009] The purpose of the present application is achieved by the following technical solutions.
[0010] One aspect of the present application provides an endocrine disruptor prediction method based on a harmful outcome pathway network, which comprises: obtaining experimental activity data as a first data set according to a pre-constructed harmful outcome pathway network, and pre-processing the first data set; wherein the first data set contains qualitative data and quantitative data; obtaining endocrine disruptor list data as a second data set; constructing a target qualitative prediction model using a machine learning algorithm according to the qualitative data in the pre-processed first data set; predicting target activity data using the target qualitative prediction model for the second data set to obtain a target activity data set; constructing an endocrine disruptor effect qualitative prediction model using a machine learning algorithm using the target activity data set; constructing a target quantitative prediction model using a machine learning algorithm according to the quantitative data in the pre-processed first data set; predicting whether a test compound has an endocrine disruptor effect using the endocrine disruptor effect qualitative prediction model; predicting the corresponding target quantitative activity of the compound with the endocrine disruptor effect using the target quantitative prediction model to obtain a quantitative harmful outcome pathway; obtaining sensitive harmful outcomes and sensitive quantitative harmful outcome pathways according to the quantitative dose-effect relationship of each event in the quantitative harmful outcome pathway; and classifying the endocrine disruptor according to the sensitive harmful outcomes and the sensitive quantitative harmful outcome pathways to obtain the prediction result of the endocrine disruptor.
[0011] Further, the experimental activity data includes molecule initiation event, key event and harmful outcome event data; the molecule initiation event represents the starting event in the harmful outcome pathway network; the key event represents the event node in the harmful outcome pathway network that reflects the change of the chemical substance and ultimately leads to the harmful outcome; and the harmful outcome event represents the terminal event in the harmful outcome pathway network.
[0012] Further, the first data set is preprocessed, including: quality assessment is performed on the experimental activity data; the experimental activity data is filtered according to the quality assessment result; and the filtered experimental activity data is taken as the preprocessed first data set.
[0013] Further, a target point quantitative prediction model or a target point qualitative prediction model is constructed by using a machine learning algorithm, including: a target point quantitative prediction model or a target point qualitative prediction model is constructed by using a modeling method based on molecular descriptors or a modeling method based on molecular graphs.
[0014] Further, the target point quantitative prediction model or the target point qualitative prediction model is constructed by using a modeling method based on molecular descriptors, including: the chemical structure SMILES of the compound is converted into numerical molecular descriptors or molecular fingerprints; and the target point quantitative prediction model or the target point qualitative prediction model is constructed by using a machine learning algorithm according to the molecular descriptors or the molecular fingerprints.
[0015] Further, the target point quantitative prediction model or the target point qualitative prediction model is constructed by using a modeling method based on molecular graphs, including: the chemical structure SMILES of the compound is converted into a molecular graph structure, in which atoms are represented as nodes and bonds are represented as edges; and the target point quantitative prediction model or the target point qualitative prediction model is constructed by using a graph neural network GNN algorithm according to the molecular graph structure.
[0016] Further, an endocrine disrupting effect qualitative prediction model is constructed by using a graph neural network-based machine learning algorithm GCNConvEdge and by using a target point activity data set, including: a harmful outcome path network is taken as a directed acyclic graph, target point activity data is taken as node input information, a causal relationship between key events is taken as edge input information, and whether an endocrine disrupting effect exists is taken as output information.
[0017] Further, a sensitive harmful outcome and a sensitive quantitative harmful outcome path are obtained according to a quantitative dose-effect relationship of each event in a quantitative harmful outcome path, including: a harmful outcome with the smallest NOAEL is obtained as a sensitive harmful outcome according to a quantitative dose-effect relationship of each event in a quantitative harmful outcome path; a corresponding quantitative harmful outcome path is determined as a sensitive quantitative harmful outcome path based on the sensitive harmful outcome, the sensitive quantitative harmful outcome path including: a sensitive molecular starting event, a sensitive key event and a sensitive harmful outcome; and whether the NOAEL values of the sensitive molecular starting event, the sensitive key event and the sensitive harmful outcome in the sensitive quantitative harmful outcome path satisfy: the NOAEL value of the sensitive molecular starting event is less than the NOAEL value of the sensitive key event, which is less than the NOAEL value of the sensitive harmful outcome; if yes, the corresponding path is taken as the sensitive quantitative harmful outcome path.
[0018] Further, according to the sensitive harmful outcomes and the sensitive quantitative harmful outcome path, the endocrine disruptor is classified in harm, and a prediction result of the endocrine disruptor is obtained, including: obtaining a highest no observable effect level (NOAEL) of each key event in the sensitive quantitative harmful outcome path; using an uncertainty factor method, the highest no observable effect level (NOAEL) is extrapolated to the sensitive harmful outcomes to obtain a highest no observable effect level (NOAEL) of the sensitive harmful outcomes based on each key event; obtaining a minimum value in the highest no observable effect level (NOAEL) of the sensitive harmful outcomes of each key event as a harm classification threshold value of the corresponding compound; and classifying the to-be-tested compound according to the harm classification threshold value.
[0019] Further, using the target point activity data set, a machine learning algorithm is used to construct an endocrine disruptor effect qualitative prediction model: the machine learning algorithm uses a machine learning algorithm based on a graph neural network, GCNConvEdge.
[0020] Another aspect of the present application also provides an endocrine disruptor prediction system, including: a data acquisition module, which acquires experimental activity data as a first data set, acquires endocrine disruptor list data as a second data set, and pre-processes the first data set; wherein the first data set contains qualitative data and quantitative data; a first data set processing module, which uses a machine learning algorithm to construct a target point qualitative prediction model according to the qualitative data in the pre-processed first data set; a second data set processing module, which uses the target point qualitative prediction model to predict target point activity data of the second data set to obtain a target point activity data set; a third data set processing module, which uses the target point activity data set to construct an endocrine disruptor effect qualitative prediction model using a machine learning algorithm; a fourth data set processing module, which uses a machine learning algorithm to construct a target point quantitative prediction model according to the quantitative data in the pre-processed first data set; a prediction module, which uses the endocrine disruptor effect qualitative prediction model to predict whether a to-be-tested compound has an endocrine disruptor effect; for the compound with the endocrine disruptor effect, the target point quantitative prediction model is used to predict the corresponding target point quantitative activity to obtain a quantitative harmful outcome path; according to the quantitative dose-effect relationship of each event in the quantitative harmful outcome path, a sensitive harmful outcome and a sensitive quantitative harmful outcome path are obtained; according to the sensitive harmful outcome and the sensitive quantitative harmful outcome path, the endocrine disruptor is classified in harm, and a prediction result of the endocrine disruptor is obtained.
[0021] 3. Beneficial effects
[0022] Compared with the prior art, the present application has the following advantages:
[0023] (1) The application adopts machine learning and machine learning algorithms, combines experimental activity data and endocrine disruptor list data, and constructs a target qualitative prediction model, an endocrine disruptor effect qualitative prediction model, and a target quantitative prediction model, which can effectively establish the biological link between the toxicological mechanism and the harmful outcome of EDCs, thereby overcoming the "black box" defect of traditional data-driven models;
[0024] (2) The application provides a visual interactive form, which can directly encode hierarchical prior knowledge into directed acyclic graph language, and convert the hierarchical knowledge into graph neural network (GNN). By adopting the machine learning algorithm GCNConvEdge based on graph neural network, the endocrine disruptor effect qualitative prediction model constructed by the application can transparently show the complete process from compound-protein interaction to individual phenotype toxicity, and provide comprehensive knowledge information for prediction and mechanism elucidation in the context of environmental toxicology;
[0025] (3) Compared with traditional data-driven models, the application organizes and learns the complete hierarchical biological knowledge of key event (KE) information (including toxicological mechanism and harmful outcome) and key event relationship (KER) information (biologically credible link) related to the harmful outcome path network, and achieves excellent performance in qualitative and quantitative hazard identification and hazard classification tasks. Specifically, in the qualitative prediction task, the method of the application can achieve an ROC-AUC value of 0.90; in the quantitative prediction task, the method of the application can control the mean square error (MSE) to a low level of 0.08. BRIEF DESCRIPTION OF DRAWINGS
[0026] Figure 1 An exemplary flowchart of an endocrine disruptor prediction method based on a harmful outcome path network of the application;
[0027] Figure 2 An EDC screening detection method collected by the application;
[0028] Figure 3 Qualitative data collected according to the standardized screening assay (top) and the ratio of the amount of data used to the amount of data collected (%) (bottom) of the application;
[0029] Figure 4 Quantitative data collected according to the standardized screening assay (top) and the ratio of the amount of data used to the amount of data collected (%) (bottom) of the application;
[0030] Figure 5 74 optimal target qualitative prediction models constructed by the application;
[0031] Figure 6This invention relates to a method for constructing a qualitative prediction model of endocrine disruption effects based on harmful outcome path networks;
[0032] Figure 7 To achieve optimal prediction performance for the novel GCNConvEdge architecture of this application;
[0033] Figure 8 A quantitative prediction model for 52 optimal targets was constructed for this application. Detailed Implementation
[0034] The present application will now be described in detail with reference to the accompanying drawings and specific embodiments.
[0035] Example 1
[0036] like Figure 1 As shown, an endocrine disruptor prediction method based on a harmful outcome pathway network includes: acquiring experimental activity data as a first dataset according to a pre-constructed harmful outcome pathway network, and preprocessing the first dataset; wherein the first dataset contains qualitative and quantitative data; acquiring an endocrine disruptor inventory as a second dataset; constructing a target qualitative prediction model using a machine learning algorithm based on the qualitative data in the preprocessed first dataset; predicting target activity data in the second dataset using the target qualitative prediction model to obtain a target activity dataset; and constructing a qualitative endocrine disruption effect model using a machine learning algorithm based on the target activity dataset. Predictive Model: Based on the quantitative data in the preprocessed first dataset, a target quantification prediction model is constructed using machine learning algorithms. A qualitative prediction model for endocrine disruption effects is used to predict whether the test compound has endocrine disruption effects. For compounds with endocrine disruption effects, the target quantification prediction model is used to predict the corresponding target quantitative activity, obtaining the quantitative adverse outcome pathway. Based on the quantitative dose-response relationship of each event in the quantitative adverse outcome pathway, sensitive adverse outcomes and sensitive quantitative adverse outcome pathways are obtained. Based on the sensitive adverse outcomes and sensitive quantitative adverse outcome pathways, the endocrine disruptors are classified into hazard levels, obtaining the prediction results for the endocrine disruptors.
[0037] Specifically, in this embodiment, estrogen (E), androgen (A), thyroid (T) and steroid production (S) mediated EDCs were selected as the research subjects. After collecting the harmful outcome pathway network of EATS-mediated EDCs, the model was constructed.
[0038] According to the pre-constructed adverse outcome pathway (AOP) network, obtain experimental activity data as the first data set: use the established AOP network to extract experimental activity data related to key events in the AOP network from relevant databases, literature and other sources. Experimental activity data includes qualitative data (such as positive / negative results) and quantitative data (such as half-effective concentration EC50, half-inhibition concentration IC50, etc.). Organize the collected experimental activity data into a structured first data set. Preprocess the first data set: clean the experimental activity data in the first data set to remove missing values, outliers and other noise data. Standardize quantitative data, such as unify concentration units, normalize numerical ranges, etc. Encode qualitative data, such as convert positive / negative results to 1 / 0, etc.
[0039] Obtain endocrine disruptor (EDC) list data as the second data set: obtain information of known or suspected EDCs from EDC lists published by authoritative agencies, relevant databases and other channels. Organize the EDC list data and extract key information such as compound name, molecular structure, physicochemical properties, etc. Organize the EDC list data as the second data set.
[0040] Collect EDC standard test methods: based on the OECD document "150th revised guidance document on standardized test guidelines for assessing endocrine disruptors", the ECHA document "Guidelines for identifying endocrine disruptors under Regulation (EU) No. 528 / 2012 and Regulation (EC) No. 1107 / 2009" and the EPA document "Endocrine Disruptor Screening Program Test Guidelines", collect EDC standard test methods. According to the above guidance documents, delete standardized tests with non-mammalian test species (such as birds), non-EATS related tests and repeated test methods. Collect test endpoints of the screened standardized test methods and compile a list of test endpoints. Correspond the test endpoints to the events in the AOP network: analyze the correspondence between the test endpoints and the molecular initiation events, key events and adverse outcomes in the AOP network. Establish a one-to-one mapping relationship between each test endpoint and a specific event in the AOP network.
[0041] As shown in Figure 2 , 26 detection methods are finally selected: according to the above three documents and the correspondence between test endpoints and AOP network events, 26 detection methods are finally selected, including 7 in vitro detection methods and 19 in vivo detection methods.
[0042] As shown in Figure 3Collection of EDCs qualitative data: Based on the selected detection methods, the collection of EDCs qualitative data was performed. The data of selected standardized tests came from five public databases, including eChemPortal, DrugBank, EADB, Endocrine Disruptor Knowledge Base, and CompTox Chemicals Dashboard. We collected a total of 21,700 experimental results from the five databases. Considering the importance of data quality for model construction, the collected experimental results data were screened: the reliability of each experimental result was evaluated with / without limitation conditions. Finally, less than half of the experimental results data were selected to ensure the high quality and reliability of the data.
[0043] Mapping of 26 detection methods' toxicity endpoints to key events in AOP networks: The key events in AOP networks include molecular initiating events (MIE), key events (KE), and adverse outcomes (AO). The correspondence between the toxicity endpoints of each detection method and the specific key events in AOP networks was analyzed. A mapping table between the toxicity endpoints of detection methods and the key events in AOP networks was established. According to the mapping relationship between the toxicity endpoints of detection methods and the key events in AOP networks, the screened experimental results data were classified into the corresponding key events: using the mapping table, the AOP network key events corresponding to each experimental result were determined. The experimental results were classified and organized according to the corresponding key events. The mapping and classification results were counted: a total of 74 key events in AOP networks were matched to the toxicity endpoints of detection methods. 132,947 toxicology records were successfully classified into different key events.
[0044] Collection of EDCs quantitative data, as shown in Figure 4 Selection of the highest no observed adverse effect level (NOAEL) as the indicator of quantitative data: NOAEL is an important toxicology parameter for evaluating the safety of compounds, representing the highest dose or exposure level that does not cause significant adverse effects under experimental conditions. It is generally believed that the exposure dose of compounds below NOAEL is safe for human body. Based on the same five databases as the collection of qualitative data, the NOAEL quantitative data of EDCs were collected: eChemPortal, DrugBank, EADB, Endocrine Disruptor Knowledge Base, CompTox Chemicals Dashboard; from the above databases, a total of 10,200 experimental results data of 18 detection methods were collected. It is noted that there is no available NOAEL data for other 8 detection methods. The collected NOAEL experimental results data were evaluated and screened for quality: the quality of toxicology records of each experimental result was evaluated. Finally, 6,713 high-quality toxicology records were selected, accounting for 65.18% of the total data.
[0045] Mapping toxicity endpoints of 18 detection methods to key events in AOP networks: Analyzing the correspondence between toxicity endpoints of each detection method and specific key events in AOP networks. Establishing a mapping table between toxicity endpoints of detection methods and key events in AOP networks. According to the mapping relationship between toxicity endpoints of detection methods and key events in AOP networks, classifying the screened NOAEL experimental result data into the corresponding key events: Using the mapping table, determining the corresponding AOP network key events for each NOAEL experimental result. Classifying and organizing NOAEL experimental results according to the corresponding key events. Statistics of mapping and classification results: The toxicity endpoints of 18 detection methods are mapped to 52 key events in AOP networks. 54200 toxicology records are successfully classified into different key events. Extensive collection of NOAEL quantitative experimental data of EDCs from five authoritative databases, combined with AOP network for data screening, mapping and classification, finally obtained high-quality, key event-related structured quantitative data set. These NOAEL data can be used for subsequent quantitative prediction model construction
[0046] Building key event-based target qualitative prediction model, data set division: Randomly dividing the qualitative data in the preprocessed first data set into training set and test set. Select 80% of the compounds as the training set for QSAR modeling, and the remaining 20% of the compounds as the test set for evaluating the model performance. Traditional QSAR modeling method: filtering and standardizing the chemical structure, generating QSAR-ready structure through the standardization workflow in KNIME. Using PaDEL-descriptor software to convert the standardized chemical structure into molecular descriptors and molecular fingerprints (2756 bits). Strict data cleaning is performed on the converted data to remove blank values, outliers and error values. Use Z-score to normalize the features, so that the features of the training set follow a normal distribution with mean 0 and variance 1. Remove features with small variance (var=0), and use univariate feature selection methods (such as mutual information method) to reduce variables and reduce information noise.
[0047] Selecting machine learning algorithms: Selecting five representative machine learning algorithms, including random forest (RF), support vector machine (SVM), decision tree (DT), naive Bayes (NB) and k-nearest neighbor (kNN) algorithm. Hyperparameter optimization: Using learning curve and grid search method to optimize the hyperparameters of each machine learning algorithm.
[0048] Optimized hyperparameters include: RF: n_estimators (number of trees) and max_depth (maximum depth of trees); SVM: C (regularization parameter) and kernel (type of kernel used in the algorithm); DT: max_depth (maximum depth of trees); NB: feature_prior (feature likelihood, such as Gaussian, multinomial, complementary, Bernoulli); kNN: n_neighbors (number of neighbors). Handling of imbalanced class problems: Synthetic Minority Over-sampling Technique (SMOTE) is used to address the potential imbalanced class problem in the dataset. SMOTE balances the dataset by synthesizing new minority class samples, improving the model's prediction performance on minority classes. Model training and evaluation: Using the optimized hyperparameters and balanced training set, models of the five machine learning algorithms are trained respectively. The performance of each model is evaluated on the test set, and the model with the best performance is selected as the final target qualitative prediction model. Using pre-processed qualitative data, traditional QSAR modeling methods and various machine learning algorithms are used to build target qualitative prediction models based on key events. These models can be used to predict whether a compound has a specific key event-related biological activity.
[0049] Building target qualitative prediction models based on molecular graphs: Convert the SMILES string of each compound into a molecular graph. Use the open-source package RDKit to determine the set of atoms and chemical bonds in the molecular graph. Feature vector calculation and initialization: Calculate and initialize the feature vectors in the molecular graph. Atomic features include 8 attributes, such as atomic number and formal charge. Chemical bond features include 4 attributes, such as bond type and conjugation.
[0050] Selecting graph neural network (GNN) algorithm: Use the Directed Message Passing Neural Network (DMPNN) as the graph neural network algorithm. DMPNN can directly learn from the graph structure of chemical substances to predict chemical toxicity. Information transmission and feature learning: DMPNN aggregates information from adjacent atoms and chemical bonds through a series of message passing steps. In each message passing step, DMPNN updates the feature vector of the current atom based on the feature vectors of adjacent atoms and chemical bonds. Through multiple message passing, DMPNN gradually establishes an understanding of local chemical properties.
[0051] Hyperparameter optimization: Hyperopt python package was used to optimize the hyperparameters through Bayesian optimization. The optimized hyperparameters include: depth: the number of information passing steps; dropout probability: a regularization technique to prevent overfitting; hidden size: the size of the key information vector; number of feedforward network layers: the number of fully connected layers used for feature transformation and prediction. Handling imbalanced class problem: SMOTE up-sampling method was used to balance the proportion of active and inactive compounds. SMOTE balances the dataset by synthesizing new minority class (active) samples to improve the model's prediction performance on active compounds. Model training and evaluation: The DMPNN model was trained using the optimized hyperparameters and the balanced training set. The performance of the model was evaluated on the test set and compared with the models built by traditional QSAR methods. A target qualitative prediction model based on molecular graph was constructed by using directed message passing neural network (DMPNN) algorithm. Compared with traditional QSAR methods, this method can learn molecular features directly from the graph representation of chemical structure without the need for manual design of molecular descriptors, and has stronger expression ability and flexibility.
[0052] As Figure 5As shown, six binary QSAR classification models were established for each event in this embodiment and the best model was selected. Specifically, for each key event, models were established using the two QSAR modeling methods mentioned earlier (traditional method and molecular graph-based method) respectively. For each modeling method, three different data partitioning methods (such as random partitioning, time series partitioning, and chemical structure similarity partitioning) were used to establish models. Therefore, for each key event, a total of 6 binary QSAR classification models were established (2 modeling methods x 3 data partitioning methods). The performance of each QSAR classification model was evaluated using three indicators: F1-score, ROC-AUC value, and accuracy. F1-score is the harmonic mean of precision and recall, suitable for evaluating imbalanced data sets. ROC-AUC value represents the area under the receiver operating characteristic curve, reflecting the overall performance of the classifier at different thresholds. Accuracy represents the proportion of correctly predicted samples to the total number of samples. For each key event, the QSAR classification model with the best performance according to the F1-score indicator was selected as the best target qualitative prediction model. A total of 74 best target qualitative prediction models corresponding to 74 key events were selected. Statistical analysis was performed on the performance indicators of the 74 best target qualitative prediction models. The average and standard deviation of F1-score, ROC-AUC value, and accuracy were calculated. The results showed that the average F1-score was 0.42 with a standard deviation of 0.19. The average ROC-AUC value was 0.71 with a standard deviation of 0.12. The average accuracy was 0.88 with a standard deviation of 0.11. The statistical and visual analysis results of the performance indicators of these best models showed that the QSAR models constructed had good overall prediction ability and could be used to predict the effects of compounds on specific key events.
[0053] The target qualitative prediction model was used for prediction and the endocrine disrupting effect qualitative prediction model was constructed, including: for the compounds in the second data set, the best target qualitative prediction model constructed in the S2 step was used for prediction. For each compound, the corresponding best model was used to predict its effect on each key event (i.e., target activity). The prediction results were integrated to obtain a complete target activity data set, including original data and predicted data. Known EDCs were collected as positive samples, and since there was a lack of clear non-EDC data, it was necessary to construct hypothetical non-EDCs as negative samples. 81 chemicals identified by ECHA as EDCs were selected as positive samples. 810 hypothetical non-EDCs were generated as negative samples using the DUD-E method, which were similar to EDCs in physical characteristics but different from EDCs in topological structure to minimize the possibility of being actual EDCs. The target activity data of the positive samples (EDCs) and negative samples (hypothetical non-EDCs) were preprocessed.
[0054] For missing target activity qualitative data, the best performing model is used to fill in. A total of 891 chemicals (81 EDCs + 810 hypothesized non-EDCs) are used to fill in 64336 missing target activity qualitative data. Feature engineering is performed on the preprocessed target activity data to extract meaningful feature representations. Features can be generated using methods such as molecular descriptors, molecular fingerprints, or graph representations of compounds. Necessary preprocessing operations such as normalization, standardization, etc. are performed on the features. Machine learning algorithms suitable for processing target activity data are selected, such as feedforward neural networks, convolutional neural networks, or graph neural networks, etc. According to the characteristics of the data and the requirements of the task, a suitable network architecture and hyperparameters are designed. The preprocessed target activity dataset is divided into training set, validation set and test set. The endocrine disrupting effect qualitative prediction model is trained using the training set, and the hyperparameters are optimized on the validation set. The performance of the model is evaluated on the test set using appropriate evaluation metrics such as accuracy, precision, recall, F1 value, etc.
[0055] As shown in Figure 6 , different architectures are used to build endocrine disrupting effect qualitative prediction models and compare performance. A novel GCNConvEdge architecture developed by the present application is used to build an endocrine disrupting effect qualitative prediction model based on the adverse outcome pathway (AOP) network. GCNConvEdge architecture is a variant of graph convolutional neural network (GCN) that can consider both node features and edge features, better capturing information in the AOP network. The AOP network is represented as a graph, where nodes represent key events and edges represent causal relationships between key events. GCNConvEdge layers are used to extract features and pass information in the AOP network, learning node and edge representations. After the GCNConvEdge layer, a fully connected layer and a softmax activation function are used for classification, predicting whether a compound has an endocrine disrupting effect. A method based on traditional GNN architecture is used to build an endocrine disrupting effect qualitative prediction model. Traditional GNN architectures such as graph convolutional network (GCN), graph attention network (GAT), etc. only consider node features and do not consider edge features. The AOP network is represented as a graph, where nodes represent key events and edges represent connections between key events. Traditional GNN layers are used to extract features and pass information in the AOP network, learning node representations. After the GNN layer, a fully connected layer and a softmax activation function are used for classification, predicting whether a compound has an endocrine disrupting effect.
[0056] A qualitative prediction model of endocrine disrupting effect is constructed using an architecture based on machine learning (ML) ensemble methods. Ensemble methods such as random forest, gradient boosting decision tree, etc. improve performance by combining the prediction results of multiple base models. For each key event in the AOP network, the best target point qualitative prediction model is used to predict the impact of the compound. The prediction results of all key events are used as features to train the ML ensemble model for classification to predict whether the compound has an endocrine disrupting effect. The same training set, validation set and test set are used to train and evaluate three different architecture models respectively. Using appropriate evaluation metrics such as accuracy, precision, recall, F1 value, etc. to compare the performance of different models. Endocrine disrupting effect qualitative prediction models are constructed using new GCNConvEdge architecture, traditional GNN architecture and ML ensemble architecture respectively, and their performance is compared. This comparison can reveal the key role of AOP network in EDC identification and demonstrate the advantage of GCNConvEdge architecture-based model in capturing AOP network information.
[0057] In short, in the GCNConvEdge architecture, the present example uses the adverse outcome pathway network as a typical directed acyclic graph and embeds the target activity as a node and the causal relationship between events (KER, i.e. biologically reasonable connection) as an edge into a three-layer GCNConvEdge architecture, and finally performs EDC qualitative prediction based on a fully connected neural network layer. Since GCNConvEdge can integrate the hierarchical biological knowledge of the adverse outcome pathway network into a unified representation, the present example believes that it can provide the most comprehensive biological knowledge for EDC prediction. Unlike GCNConvEdge, the traditional GNN-based architecture (the second architecture) ignores the causal relationship between events and uses the adverse outcome pathway network as an undirected acyclic graph. This architecture learns node embeddings from the adverse outcome pathway network, ignoring the causal relationship between events, and uses four different GNN algorithms for EDC prediction. In the ML-based ensemble architecture (the third architecture), the KER information is completely removed, only the event information is retained, and it is converted into fixed table data. The third architecture only learns node information from the adverse outcome pathway network for EDC prediction.
[0058] Three different architectures are adopted to process Adverse Outcome Pathway (AOP) networks and predict Endocrine Disrupting Chemicals (EDCs): GCNConvEdge-based architecture: The AOP network is represented as a directed acyclic graph (DAG), where nodes represent key events and edges represent causal relationships between events (KER, i.e., biologically plausible connections). The target activity is embedded as a node, and the KER is embedded as an edge. These information is input into the GCNConvEdge architecture. The GCNConvEdge architecture consists of three layers of GCNConvEdge layers, which can handle both node and edge embeddings simultaneously and capture the hierarchical biological knowledge in the AOP network. After the GCNConvEdge layers, a fully connected neural network layer is used for EDC qualitative prediction. By integrating the hierarchical biological knowledge of the AOP network into a unified representation, the GCNConvEdge architecture can provide the most comprehensive biological knowledge for EDC prediction.
[0059] Traditional GNN-based architecture: The AOP network is represented as an undirected acyclic graph, where nodes represent key events and edges represent connections between events (ignoring causal relationships). The target activity is embedded as a node, and this information is input into the traditional GNN architecture. Traditional GNN architectures such as Graph Convolutional Networks (GCN), Graph Attention Networks (GAT), etc., only consider node embeddings and ignore edge embeddings (i.e., causal relationships between events). Four different GNN algorithms (such as GCN, GAT, GraphSAGE, GIN) are used to extract features and pass information in the AOP network, learning the representation of nodes. After the GNN layers, a fully connected neural network layer is used for EDC qualitative prediction. Due to the neglect of causal relationships between events, the traditional GNN architecture may not be able to fully utilize the biological knowledge in the AOP network.
[0060] Based on ML ensemble architecture: Extract the key event information in the AOP network and convert it into fixed table data, completely delete the KER information. Use the best target point qualitative prediction model to predict the target point activity corresponding to each key event. Use the predicted target point activity as a feature to build a model based on machine learning (ML) ensemble methods, such as random forests, gradient boosting decision trees, etc. Use the ML ensemble model to perform EDC qualitative prediction, and improve performance by combining the prediction results of multiple base models. Since the KER information is completely deleted, the ML ensemble architecture may not be able to capture the causal relationships and biological knowledge in the AOP network. By comparing the above three architectures, it can be seen that the GCNConvEdge architecture has a clear advantage in utilizing the hierarchical biological knowledge of the AOP network. It can consider both node information (key events) and edge information (KER), capture the causal relationships between events, and provide more comprehensive biological knowledge for EDC prediction. In contrast, the traditional GNN architecture ignores the causal relationships between events, while the ML ensemble architecture completely deletes the KER information and may not be able to fully utilize the biological knowledge in the AOP network. This comparison highlights the advantages and potential of the GCNConvEdge architecture in endocrine disruptor identification.
[0061] Performance comparison and results of different architectures on the endocrine disruptor effect qualitative prediction task as Figure 7The superior performance of the GCNConvEdge architecture: The models based on the GCNConvEdge architecture outperformed other architectures in all evaluation metrics (ROC-AUC, accuracy, and F1-score). The superior performance of the GCNConvEdge architecture can be attributed to its ability to leverage the complete hierarchical biological knowledge in the AOP network, including key event information and causal relationships between events (KER). Specifically, the models based on the GCNConvEdge architecture achieved 0.90, 0.96, and 0.78 in ROC-AUC, accuracy, and F1-score, respectively. In comparison, other architectures (traditional GNN architectures and ML ensemble architectures) performed worse in each metric. The other architectures achieved a range of 0.86-0.89 in ROC-AUC, 0.92-0.95 in accuracy, and 0.65-0.75 in F1-score. The performance differences among these architectures can be attributed to their limitations in handling AOP network information, such as ignoring causal relationships between events or completely removing KER information. The indispensability of each component in the AOP network: The superior performance of the GCNConvEdge architecture in all metrics highlights the indispensability of each component (key events and KER) in the AOP network. By considering both key event information and causal relationships between events, the GCNConvEdge architecture can more comprehensively leverage the hierarchical biological knowledge in the AOP network, resulting in more accurate predictions. The importance of event information: When the ensemble model only learned half of the event information, its prediction performance decreased significantly compared to the GCNConvEdge architecture. Compared to the GCNConvEdge architecture, the ensemble model decreased by 0.14, 0.10, and 0.16 in ROC-AUC, accuracy, and F1-score, respectively. This result indicates that event information is crucial for accurate prediction of endocrine disruptor effects, and removing part of the event information can significantly affect the performance of the model. The models based on the GCNConvEdge architecture performed best in the qualitative prediction task of endocrine disruptor effects, thanks to their ability to fully leverage the complete hierarchical biological knowledge in the AOP network.
[0062] Constructing target quantitative prediction models, such as Figure 8The quantitative data in the pre-processed first dataset were randomly divided into training and test sets. 80% of the compounds were selected as the training set for QSAR modeling and model training. The remaining 20% of the compounds were used as the test set to evaluate the performance of the developed model. Five classic machine learning (ML) algorithms and one deep learning (DL) algorithm were selected to build the target quantitative prediction model. The classic ML algorithms included random forest (RF), decision tree (DT), Xgboost, naive Bayes (NB), and k-nearest neighbors (kNN). The DL algorithm was fully connected neural network (FCNN). Six quantitative QSAR models were built for each key event, including five models based on ML algorithms and one model based on DL algorithm. A total of 312 regression models were built (52 key events x 6 algorithms). For each key event, the same training and test sets were used to train and evaluate the six models. The performance of the six models for each key event was evaluated using the R2index (coefficient of determination). For each key event, the model with the highest R2index was selected as the best prediction model. The performance indicators of each best model were recorded, including mean absolute error (MAE), mean squared error (MSE), and R2index. The performance of the best prediction models for the 52 key events was analyzed. The average and standard deviation of the MAE, MSE, and R2indices of the 52 models were calculated. The results showed that the average MAE of the 52 models was 0.18 ± 0.08, the average MSE was 0.08 ± 0.05, and the average R2was 0.58 ± 0.20, indicating that these models had high prediction performance. Further analysis found that among the 52 best models, 50 (96.15%) were built using tree-based algorithms such as RF, DT, and Xgboost. Six quantitative QSAR models were built for each key event using machine learning algorithms and machine learning algorithms, and the best-performing model was selected as the final target quantitative prediction model. The results showed that these models had high prediction performance, with tree-based algorithms accounting for the highest proportion of models.
[0063] The endocrine disrupting effect qualitative prediction model is used to predict whether the test compound has an endocrine disrupting effect. The test compound is input into the constructed endocrine disrupting effect qualitative prediction model. The model outputs the prediction result of whether the test compound has an endocrine disrupting effect. Compounds predicted to have an endocrine disrupting effect are input into the constructed target quantitative prediction model. The model outputs the quantitative activity value of each key event, forming a quantitative adverse outcome pathway (qAOP). The quantitative dose-effect relationship of each event in the qAOP is analyzed. The most likely to be activated by EDCs at the lowest environmental concentration is determined as the sensitive AO and qAOP.
[0064] Hazard classification of EDCs based on sensitive AO and sensitive qAOP: EDCs are classified based on the type and severity of sensitive AO. The lowest observed adverse effect level (NOAEL) of EDCs is determined based on the quantitative dose-effect relationship in sensitive qAOP. EDCs are classified as high hazard (NOAEL≤1 mg / kg bw / day), medium hazard (NOAEL>1 mg / kg bw / day) and so on based on NOAEL. The qualitative prediction results of EDCs, sensitive AO, sensitive qAOP and hazard classification results are comprehensively considered. The prediction results of endocrine disrupting effect of each test compound are output, including whether it has endocrine disrupting effect, main sensitive AO, sensitive qAOP and hazard level.
[0065] The present application predicts potential EDCs from the global existing chemical substance list: Three chemical substance lists of China, the United States and the European Union are collected and sorted, a total of 225814 structurally unique chemical substances. The qualitative prediction model of endocrine disrupting effect based on harmful outcome path network is used to predict whether each chemical substance has endocrine disrupting effect. The results show that a total of 6461 (2.86%) chemical substances are qualitatively predicted as potential EDCs.
[0066] Prioritization and hazard classification of potential EDCs: 6461 EDCs are prioritized using three quantitative rules. The results show that a total of 293 qAOPs are activated by at least one EDC, of which 73 qAOPs can be activated by more than half of EDCs. EDCs tend to activate 40 sensitive qAOPs, which in turn produce 40 sensitive AO, with NOAEL from 430 mg / kg bw / day (medium hazard) to 1 mg / kg bw / day (high hazard). Reproductive effects (3986 EDCs, 61.69%) and developmental effects (765 EDCs, 11.84%) are the most sensitive AOs.
[0067] Analysis of gender-specific toxicity of EDCs on the reproductive system: 2535 EDCs mainly have harmful effects on the male reproductive system (such as testis and seminal vesicle). 1435 EDCs mainly have harmful effects on the female reproductive system (such as ovary and uterus). Only 99 EDCs have harmful effects on both male and female reproductive systems (such as impaired fertility). The remaining 1710 EDCs (26.47%) mainly affect the adrenal gland and pituitary gland, which play a key role in HPG and HPA axis and are essential for endocrine balance.
[0068] The qualitative prediction model of endocrine disrupting effect and the quantitative prediction model of target point were used to predict and classify the endocrine disrupting effect of the tested compounds. The case analysis results show that this method can effectively screen out potential EDCs from the global chemical substance list, and determine the sensitive AO, sensitive qAOP and hazard grade. At the same time, the analysis also reveals the gender difference in the toxicity of EDCs to the reproductive system, which provides an important basis for the risk assessment and management of endocrine disruptors.
[0069] The above describes the application creation and its implementation in a schematic manner, which is not restrictive, and the application can be realized in other specific forms without departing from the spirit or essential characteristics of the application. The embodiments shown in the drawings are only one of the embodiments of the application creation, and the actual structure is not limited thereto, and any reference signs in the claims should not limit the claims involved. Therefore, if a person skilled in the art is inspired by it, without departing from the spirit of the creation, similar structural forms and embodiments can be designed without creative design, which should belong to the protection scope of the patent. In addition, the word "comprising" does not exclude other elements or steps, and the word "one" before the element does not exclude the inclusion of "multiple" elements. The multiple elements stated in the product claim can also be realized by one element through software or hardware. The words "first", "second" and the like are used to represent names, and do not represent any specific order.
Claims
1. A method for predicting endocrine disruptors based on harmful outcome path networks, characterized in that, include: Experimental activity data were obtained based on a pre-constructed harmful outcome pathway network and used as the first dataset. The first dataset was then preprocessed, and it contained both qualitative and quantitative data. Obtain a list of endocrine disruptors as a second dataset; Based on the qualitative data in the preprocessed first dataset, a target qualitative prediction model is constructed using machine learning algorithms. Using a target qualitative prediction model, target activity data are predicted in the second dataset to obtain the target activity dataset. Using a target activity dataset and machine learning algorithms, a qualitative prediction model for endocrine disruption effects was constructed. Based on the quantitative data in the preprocessed first dataset, a target quantitative prediction model is constructed using machine learning algorithms. A qualitative prediction model for endocrine disruption effects is used to predict whether a test compound has endocrine disruption effects. For compounds with endocrine-disrupting effects, a target quantification prediction model is used to predict the corresponding target quantitative activity and obtain the quantitative harmful outcome pathway. Based on the quantitative dose-response relationship of each event in the quantitative adverse outcome pathway, sensitive adverse outcomes and sensitive quantitative adverse outcome pathways are obtained; Based on sensitive adverse outcomes and sensitive quantitative adverse outcome pathways, endocrine disruptors are classified into hazard levels to obtain prediction results for endocrine disruptors. Using a target activity dataset and machine learning algorithms, a qualitative prediction model for endocrine disruption effects is constructed, including: The harmful outcome path network is treated as a directed acyclic graph, with target activity data as node input information, causal relationships between key events as edge input information, and whether or not it has an endocrine disruption effect as output information. Specifically, based on the GCNConvEdge architecture: the harmful outcome path (AOP) network is represented as a directed acyclic graph (DAG), where nodes represent key events and edges represent causal relationships (KERs) between events; target activity is embedded as nodes and KERs are embedded as edges; the GCNConvEdge architecture consists of three GCNConvEdge layers, which simultaneously handle node embedding and edge embedding to capture hierarchical biological knowledge in the AOP network; after the GCNConvEdge layers, a fully connected neural network layer is used for qualitative prediction of endocrine disruptors (EDCs).
2. The method for predicting endocrine disruptors based on harmful outcome path networks according to claim 1, characterized in that: Experimental activity data include data related to molecular initiation events, critical events, and harmful outcome events; Molecular initiation events represent the starting events in a harmful outcome path network; Critical events represent event nodes in the harmful outcome path network that reflect changes in chemical substances that ultimately lead to harmful outcomes; Harmful ending events represent endpoint events in a harmful ending path network.
3. The method for predicting endocrine disruptors based on harmful outcome path networks according to claim 2, characterized in that: The first dataset is preprocessed, including: Quality assessment of experimental activity data; Based on the quality assessment results, the experimental activity data were screened. The selected experimental activity data were used as the first dataset after preprocessing.
4. The method for predicting endocrine disruptors based on harmful outcome path networks according to claim 1, characterized in that: Machine learning algorithms are used to construct quantitative or qualitative target prediction models, including: Target quantitative prediction models or target qualitative prediction models are constructed using molecular descriptor-based or molecular graph-based modeling methods.
5. The method for predicting endocrine disruptors based on harmful outcome path networks according to claim 4, characterized in that: Using molecular descriptor-based modeling methods, quantitative or qualitative target prediction models are constructed, including: Convert the chemical structure of a compound into a numerical molecular descriptor or molecular fingerprint; Based on molecular descriptors or molecular fingerprints, machine learning algorithms are used to construct quantitative or qualitative target prediction models.
6. The method for predicting endocrine disruptors based on harmful outcome path networks according to claim 4, characterized in that: Sampling-based molecular graph modeling methods are used to construct quantitative or qualitative target prediction models, including: The chemical structure of a compound is transformed into a molecular diagram structure, in which atoms are represented as nodes and bonds as edges. Based on the molecular graph structure, a graph neural network (GNN) algorithm is used to construct a quantitative or qualitative prediction model for the target.
7. The method for predicting endocrine disruptors based on harmful outcome path networks according to claim 1, characterized in that: Pathways for obtaining sensitive adverse outcomes and sensitive quantitative adverse outcomes include: Based on the quantitative dose-response relationship of each event in the quantitative adverse outcome pathway, the adverse outcome with the lowest highest ineffective concentration is obtained as the sensitive adverse outcome; Based on sensitive harmful outcomes, the corresponding quantitative harmful outcome pathways are identified as sensitive quantitative harmful outcome pathways. Sensitive quantitative harmful outcome pathways include: sensitive molecular initiation events, sensitive key events, and sensitive harmful outcomes. Compare whether the highest no-effect concentration values for the sensitive molecule initiation event, sensitive key event, and sensitive adverse outcome in the sensitive quantitative adverse outcome pathway meet the following requirements: The highest no-effect concentration of the sensitive molecule initiation event is less than the highest no-effect concentration of the sensitive critical event, which is less than the highest no-effect concentration of the sensitive adverse outcome. If this condition is met, the corresponding pathway is considered the sensitive quantitative adverse outcome pathway.
8. The method for predicting endocrine disruptors based on harmful outcome path networks according to claim 1, characterized in that: The predicted results for endocrine disruptors include: Obtain the highest no-response concentration for each key event in the sensitive quantitative adverse outcome pathway; By using the uncertainty factor method, the highest no-effect concentration was extrapolated to the sensitive adverse outcome, thus obtaining the highest no-effect concentration based on the sensitive adverse outcome for each key event; The minimum of the highest ineffective concentrations for sensitive adverse outcomes of each key event is obtained and used as the hazard classification threshold for the corresponding compound. The test compounds are classified according to the hazard classification threshold.
9. A system for predicting endocrine disruptors, used to implement the method according to any one of claims 1 to 8, characterized in that, include: The data acquisition module acquires experimental activity data as the first dataset and endocrine disruptor list data as the second dataset, and preprocesses the first dataset; the first dataset contains qualitative and quantitative data. The first dataset processing module uses machine learning algorithms to construct a target qualitative prediction model based on the qualitative data in the preprocessed first dataset. The second dataset processing module uses a target qualitative prediction model to predict target activity data in the second dataset, thus obtaining the target activity dataset. The third dataset processing module uses the target activity dataset and employs machine learning algorithms to construct a qualitative prediction model for endocrine disruption effects. The fourth dataset processing module uses machine learning algorithms to construct a target quantitative prediction model based on the quantitative data in the preprocessed first dataset. The prediction module uses a qualitative prediction model of endocrine disruption effects to predict whether the test compound has endocrine disruption effects; for compounds with endocrine disruption effects, it uses a target quantification prediction model to predict the corresponding target quantitative activity to obtain a quantitative harmful outcome pathway; based on the quantitative dose-response relationship of each event in the quantitative harmful outcome pathway, it obtains sensitive harmful outcomes and sensitive quantitative harmful outcome pathways; based on sensitive harmful outcomes and sensitive quantitative harmful outcome pathways, it classifies the hazard of endocrine disruptors to obtain the prediction results of endocrine disruptors.
Citation Information
Patent Citations
A method for high-throughput screening of endocrine disruptors based on a hierarchical alert structure
CN109545289B
Method for performing high-flux screening on endocrine disruptors on basis of hierarchical warning structure
CN109545289A
Hybrid mimetic and resistant glucocorticoid interferent identifying method based on enhanced sampling molecular dynamics simulation
CN110501510A