A prediction method and device for obesity-causing chemicals based on machine learning
By establishing a database of key molecular event and obesity-causing chemicals database, and using machine learning algorithms to build a prediction model, the problem that obesity-causing chemicals prediction in the existing technology depends on in vitro experimental data, achieving rapid, batch and accurate prediction, with an accuracy of 74.5%.
Patent Information
- Application Number
- CN202310907765.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-24
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2043-07-24
AI Technical Summary
The prior art is difficult to predict obesity-causing chemicals quickly, batch and accurately, and it relies on in vitro test data, which is expensive and has a long cycle.
By establishing a database of key molecular event and obesity-causing chemicals database, using machine learning algorithms to build a predictive model, and integrating predictive sub-models of multiple key molecular events are carried out to achieve the prediction of obesity-causing chemicals.
It has achieved batch, fast and accurate prediction of obesity-causing chemicals without relying on in vitro test data, with an accuracy of 74.5%, providing an effective tool for compound risk assessment and control.
Smart Images

Figure CN117133375B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of compound prediction, and particularly relates to a method and device for predicting obesity-causing chemicals based on machine learning. Background Art
[0002] Statistical data shows that the number of overweight and obese people in the world has increased from more than 800 million in 1980 to 2.1 billion in 2013 and has continued to increase in recent years. Obesity increases the risk of type 2 diabetes, cardiovascular diseases and cancer and has now become a major global health challenge. It is worth noting that the obesity epidemic has occurred simultaneously with the widespread use of chemicals. More and more studies have shown that environmental pollutants can cause obesity. These chemical substances that can interfere with the lipid metabolism homeostasis of the human body to promote fat formation and lipid accumulation are called obesity-causing chemicals, and there is still a lack of corresponding regulatory policies.
[0003] Currently, the generally recognized criterion for determining obesity-causing chemicals in toxicology is: the ability to promote the adipogenic differentiation process of 3T3-L1 cells. Existing mechanistic studies have found that the interaction with peroxisome proliferator-activated receptor gamma (PPARγ), glucocorticoid receptor (GR), liver X receptor (LXR), retinoic acid receptor (RXR), CCAAT / enhancer binding protein beta (C_EBP), and sterol regulatory element binding transcription factor 1 (SREBP) is the key molecular mechanism for promoting the adipogenic differentiation process of 3T3-L1 cells.
[0004] Based on the above key molecular mechanisms, two methods for scoring the obesity-inducing potential of chemicals were proposed: the 5-slice and 8-slice methods. These two methods are based on the interaction data of compounds with the above six receptors and proteins obtained from batch in vitro tests in the US EPA's ToxCast program. The 5-slice method divides the data into 5 parts (PPARγ, GR, LXR, RXR, Other), calculates the average value of the data within each part, and sums up each part to obtain the total score. The 8-slice method divides the data into 8 parts (PPRE, PPARγ, GR, LXR, LXRE, RXR, C_EBP, SREBP), calculates the average value of the data within each part, and sums up each part to obtain the total score. Chemicals with higher scores have a higher obesity-inducing potential. The main limitation of the above methods is that they rely on data from in vitro tests, which are costly and time-consuming, and cannot quickly and batch predict obesity-inducing chemicals. Every year, a large number of chemical substances enter the environmental medium. In the fields of new compound declaration and environmental risk assessment and control, it is necessary to identify the chronic toxic effects of chemical substances on human health. Therefore, there is an urgent need to develop a technical method that can batch, quickly, and accurately predict obesity-inducing chemicals, and screen out obesity-inducing chemicals through this method to serve the environmental risk assessment and control of new compounds.
[0005] Patent document CN114171137A discloses a method for predicting the environmental hazard of compounds based on machine learning, which includes the following steps: (1) establishing screening criteria for the environmental hazard of compounds; (2) collecting sample labels and sample data, and preprocessing the sample data; (3) constructing a prediction model based on a machine learning algorithm, and optimizing the parameters of the prediction model using the preprocessed sample data; (4) predicting whether the compound to be tested has environmental hazards. This method uses chemical properties and acute toxicity to predict the input substances, and only labels substances that meet all requirements as hazardous substances. However, this method has the problem of missed detection, that is, actual hazardous substances are labeled as non-hazardous substances when they only meet one or two conditions. At the same time, this method cannot identify the toxic action mechanism of chemicals.
[0006] Patent document CN109360610A discloses a prediction model algorithm for the biological toxicity of chemical molecules based on a fuzzy neural network, with the hydrophobicity of different chemical structures as the control quantity and the biological toxicity quantity as the controlled quantity, including the following steps: (1) establishing a QSAR model for biological toxicity and the octanol / water partition coefficient, that is, establishing a model for biological toxicity and the octanol / water partition coefficient during the chemical molecule synthesis process; (2) establishing a prediction model for the biological toxicity of chemical molecules based on a fuzzy neural network; (3) establishing a genetic algorithm model for optimizing the parameters of the NFN to correct the parameters of the fuzzy neural network; (4) finally, using the optimized fuzzy neural network model to calculate and predict the biological toxicity value of new molecules. Summary of the Invention
[0007] The object of the present invention is to provide a method and device for predicting obesity-causing chemicals based on machine learning, which can predict obesity-causing chemicals without relying on in vitro test data.
[0008] To achieve the first object of the present invention, a method for predicting obesity-causing chemicals based on machine learning is provided. The obesity-causing chemicals refer to compounds that can promote the adipogenic differentiation process of 3T3-L1 cells, and the method includes the following steps:
[0009] Judge whether a compound has an agonist effect on key molecular events according to multiple in vitro test data related to the key molecular events, and construct a key molecular event database.
[0010] Based on the definition of obesity-causing chemicals, label the compounds, and construct a corresponding obesity-causing chemical database.
[0011] Extract part of the compounds from the key molecular event database as samples, and extract the corresponding specific molecular features based on the molecular structural formulas of the samples.
[0012] Based on the specific molecular features of the samples and the labels of whether there is an agonist effect, form a sample data set, and train a pre-constructed model based on the sample data set to obtain a prediction model for predicting whether a compound has an agonist effect on key molecular events. The prediction model includes prediction sub-models corresponding to multiple key molecular events.
[0013] Perform compound similarity matching on the key molecular event database and the obesity-causing chemical database, and use the overlapping compounds as the standard data set.
[0014] Based on the standard data set, calculate scores using the 5-slice method and the 8-slice method respectively, draw an ROC curve with the scores and the actual obesity-causing results, compare the AUC areas, select the 5-slice method with higher prediction accuracy for sub-model integration, and train all the prediction sub-models and integrate the training results to obtain the cut-off point of the prediction model.
[0015] Input the molecular structural formula of the compound to be tested into the prediction model, and judge the prediction result based on the cut-off point to obtain the judgment result of whether the compound to be tested is an obesity-causing compound, and identify the key molecular mechanism of the obesity-causing effect of the compound to be tested.
[0016] The present invention constructs a quantitative structure-activity model of compounds through a machine learning algorithm to predict whether a chemical can trigger key molecular events (interacting with PPARγ, GR, LXR, RXR, C_EBP, SREBP), and then integrates the results of multiple models to comprehensively predict whether a chemical can promote the adipogenic differentiation process of 3T3-L1 cells, realizing the prediction of obesity-causing chemicals without relying on in vitro test data.
[0017] Specifically, for the compounds in the obesity-causing chemical database, if they can increase the lipid droplet content in 3T3-L1 cells, the label is 1; otherwise, it is 0.
[0018] Specifically, the key molecular events include peroxisome proliferator-activated receptor, glucocorticoid receptor, liver X receptor, retinoic acid receptor, CCAAT / enhancer-binding protein, or sterol regulatory element-binding transcription factor.
[0019] Specifically, the discrimination criterion for whether a compound has an agonistic effect on a key molecular event:
[0020] In the in vitro test data related to the key molecular event, more than half of the test results show an agonistic effect.
[0021] Specifically, the specific molecular features include molecular descriptors and / or molecular fingerprints.
[0022] Specifically, the sample dataset needs to be preprocessed before being input into the prediction model, including standardization and normalization. Standardization processes the data according to the columns of the feature matrix to convert the feature values of the samples to the same dimension, and normalization processes the data according to the rows of the feature matrix to map the data to a specified range.
[0023] Specifically, the prediction model is constructed using one of the random forest algorithm, support vector machine algorithm, or extreme gradient boosting tree algorithm.
[0024] Specifically, according to the prediction result scores of each substance in the integrated training results, they are sorted from high to low and compared with the actual results, and the score that maximizes the accuracy is used as the cut-off point.
[0025] To achieve the second object of the present invention, an obesity-causing chemical prediction device is provided, including a computer memory, a computer processor, and a computer program stored in the computer memory and executable on the computer processor. The computer memory adopts the above-mentioned machine learning-based obesity-causing chemical prediction method.
[0026] When the computer processor executes the computer program, the following steps are implemented:
[0027] Input the molecular structural formula of the compound to be tested into the computer and analyze it through the method for predicting obesogenic chemicals to output the judgment result on whether the compound to be tested is an obesogenic chemical.
[0028] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0029] (1) By establishing a key molecular event database and an obesogenic chemical database, the present invention constructs an optimized prediction model based on sample data and sample labels using machine learning algorithms.
[0030] (2) According to the molecular expression of the substance to be tested, it is possible to batch, quickly, and accurately predict whether the compound belongs to an obesogenic chemical, providing guidance for the risk assessment and control of compounds. Description of the Drawings
[0031] Figure 1 It is a schematic flowchart of the method for predicting obesogenic chemicals provided in this embodiment. Detailed Embodiments
[0032] The present invention will be further clarified below in conjunction with the drawings and embodiments. It should be understood that these embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention.
[0033] As Figure 1 shown, a method for predicting obesogenic chemicals provided in this embodiment includes the following steps:
[0034] Establish a key molecular event database. That is, download the ToxCast in vitro test data from the EPA website. The data in the ToxCast in vitro test database are the specific concentrations at which compounds can produce agonist effects in specific receptor transcriptional activity tests (one key molecular event corresponds to multiple test results), and further judgment is required to convert them into binary classification values for model construction. The judgment criterion is: if more than half of the test results corresponding to a key molecular event show agonist effects, the label is 1; otherwise, the label is 0. In this embodiment, a total of 6 key molecular event databases are established. Among them, the PPARγ (peroxisome proliferator-activated receptor) database contains 1823 compounds, the GR (glucocorticoid receptor) database contains 2328 compounds, the LXR (liver X receptor) database contains 3563 compounds, the RXR (retinoic acid receptor) database contains 3563 compounds, the C_EBP (CCAAT / enhancer-binding protein) database contains 3563 compounds, and the SREBP (sterol regulatory element-binding transcription factor) database contains 3563 compounds.
[0035] Build an obesity-causing chemical database. Query compounds related to 3T3-L1 cell experiments in Web of Science, and determine whether the compounds are obesity-causing chemicals through the abstract content. The obesity-causing chemical database in this embodiment contains 384 substances.
[0036] Extract some compounds from the key molecular event database as samples, and then extract the corresponding specific molecular features based on the molecular structural formulas of the samples.
[0037] Combine the specific molecular features of the samples and the labels of whether there is an agonist effect to form a sample data set. Based on this sample data set, train the pre-constructed model. Before inputting the sample data set, preprocessing is required. Taking the PPARγ (peroxisome proliferator-activated receptor) receptor as an example, a total of 1697 substances are extracted as samples. Export the molecular structural formulas and linear molecular structures of the samples from the PubChem database, and use the RDKit software to extract the specific structural features of the samples. Use the linear molecular structure of the samples as the input and the RDKit molecular descriptors as the output. Each column of data corresponds to a molecular descriptor. Finally, obtain 208 columns of molecular descriptors, that is, a feature matrix of 1697 rows and 208 columns. The number 1 represents a substance that has an agonist effect on the PPARγ receptor, and the number 0 represents a substance that has no agonist effect on the PPARγ receptor. Establish a sample data set, and standardize and normalize the sample data set.
[0038] Use the preprocessed sample data to train the prediction model under the supervision of the sample labels to optimize the prediction model parameters.
[0039] Randomly divide the sample data into a test set and a training set, where the ratio of the training set to the test set is 8:2. The machine learning algorithm uses a binary classifier in the random forest algorithm, support vector machine algorithm, and extreme gradient boosting tree algorithm. Use the test set to evaluate the goodness of the prediction model, optimize the prediction model parameters, and select the best-performing model from the three models as the prediction model for this key molecular event. This prediction model includes prediction sub-models corresponding to each key molecular event.
[0040] The goodness evaluation indicators of the model include: accuracy, recall rate, and the area under the curve AUC value. The accuracy rate refers to the proportion of all correctly predicted sample data in the total sample data; the recall rate refers to the proportion of correctly predicted positives in all actual positive samples; the AUC value refers to the area under the ROC curve. The ROC curve is a curve with the false positive rate as the abscissa and the true positive rate as the ordinate.
[0041] Find the overlapping compounds in the key molecular event database and the obesity-causing chemical database, and set them as the standard data set. The standard data set in this embodiment contains 126 compounds.
[0042] The scores were calculated using the 5-slice method and the 8-slice method respectively. The ROC curve was plotted using the scores and the actual obesity-causing results, and the AUC areas were compared. It was found that the prediction effect of integrating with the 5-slice method on the final result was better (5-slice: AUC = 0.735; 8-slice: AUC = 0.565). The 5-slice method was selected for sub-model integration.
[0043] All substances were scored according to the 5-slice method. Different cut-off points were selected to divide the scores. If the score was higher than the cut-off point, the prediction result was an obesity-causing chemical; if it was lower than the cut-off point, the prediction was a non-obesity-causing chemical. The balanced accuracy was calculated by comparing the prediction results with the true results, and 1.75 with the highest balanced accuracy was selected as the cut-off point.
[0044] Among them, each sub-model would get results of 0 and 1. The results of the 6 models for each substance were summed to obtain a score from 0 to 5 (according to the 5-slice calculation method, the results of SREBP and CEBP needed to take the average value). Then the substances were sorted from high to low according to the scores, compared with the actual results, and the score that could maximize the accuracy was found as the cut-off point. Only those higher than this were obesity-causing chemicals.
[0045] After exporting the molecular structure formula of the compound to be tested, the specific molecular features of the compound to be tested were extracted and input into the parameter-optimized prediction sub-model as the data to be tested. Then the prediction results of the sub-models were integrated, and it was determined whether the compound to be tested was an obesity-causing chemical through the cut-off point.
[0046] This example also provides an obesity-causing chemical prediction device, including a computer memory, a computer processor, and a computer program stored in the computer memory and executable on the computer processor.
[0047] The computer memory adopts the obesity-causing chemical prediction method based on machine learning proposed in the above embodiment, and the following steps are implemented during its execution:
[0048] The molecular structure formula of the compound to be tested is input into the computer and analyzed through the obesity-causing chemical prediction method to output a judgment result on whether the compound to be tested is an obesity-causing chemical.
[0049] It should be understood that the processor in this embodiment may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0050] In summary, the present invention fills the gap in the field of predicting obesity-causing chemicals in the current technology by establishing a key molecular event database and an obesity-causing chemical database, and constructing an optimized prediction model based on machine learning algorithms using sample data and sample labels, providing a batch, fast, and accurate method for predicting obesity-causing chemicals, with an accuracy of 74.5%.
[0051] The above-described embodiments have elaborated on the technical solutions of the present invention. It should be understood that the above are only specific embodiments of the present invention and do not limit the present invention. Any modifications, supplements, or substitutions in a similar manner made within the scope of the principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for predicting obesity-causing chemicals based on machine learning, where the obesity-causing chemicals refer to compounds that can promote the adipogenic differentiation process of 3T3-L1 cells. Characterized in that, It includes the following steps: Construct a key molecular event database according to whether a compound has an agonistic effect on key molecular events; Based on the definition of obesity-causing chemicals, label the compounds to construct a corresponding obesity-causing chemical database; Extract some compounds from the key molecular event database as samples, and extract corresponding specific molecular features based on the molecular structural formulas of the samples; Based on the specific molecular features of the samples and the labels of whether there is an agonistic effect, form a sample data set, and train a pre-constructed model based on the sample data set to obtain a prediction model for predicting whether a compound has an agonistic effect on key molecular events. The prediction model includes prediction sub-models corresponding to multiple key molecular events; Perform compound similarity matching on the key molecular event database and the obesity-causing chemical database, and use the overlapping compounds as the standard data set; Based on the standard data set, train all prediction sub-models and integrate the training results to obtain the cut-off point of the prediction model; Input the molecular structural formula of the compound to be tested into the prediction model, and judge the prediction result based on the cut-off point to obtain the judgment result of whether the compound to be tested is an obesity-causing compound.
2. The method for predicting obesity-causing chemicals based on machine learning according to claim 1, Characterized in that, For the compounds in the obesity-causing chemical database, if they can increase the lipid droplet content of 3T3-L1 cells, the label is 1, otherwise it is 0.
3. The method for predicting obesity-causing chemicals based on machine learning according to claim 1, Characterized in that, The key molecular events include peroxisome proliferator-activated receptor, glucocorticoid receptor, liver X receptor, retinoic acid receptor, CCAAT / enhancer-binding protein, or sterol regulatory element-binding transcription factor.
4. The method for predicting obesity-causing chemicals based on machine learning according to claim 1, Characterized in that, The specific molecular features include molecular descriptors and / or molecular fingerprints.
5. The method for predicting obesity-causing chemicals based on machine learning according to claim 1, Characterized in that, The sample data set needs to be pre-processed before being input into the prediction model, including standardization and normalization.
6. The method for predicting obesity-causing chemicals based on machine learning according to claim 1, Characterized in that, The prediction model is constructed using one of the random forest algorithm, support vector machine algorithm, or extreme gradient boosting tree algorithm.
7. The method for predicting obesity-causing chemicals based on machine learning according to claim 1, Characterized in that, Sort the prediction result scores of each substance in the integrated training results from high to low and compare with the actual results, and use the score that maximizes the accuracy as the cut-off point.
8. An obesity-causing chemical prediction device, including a computer memory, a computer processor, and a computer program stored in the computer memory and executable on the computer processor, Characterized in that, The computer memory adopts the machine learning-based method for predicting obesogenic chemicals as described in any one of claims 1 to 7; When the computer processor executes the computer program, the following steps are implemented: input the molecular structural formula of the compound to be tested into the computer, and analyze it through the method for predicting obesogenic chemicals, so as to output a judgment result on whether the compound to be tested is an obesogenic chemical.
Citation Information
Patent Citations
Prediction model algorithm for biological toxicity of chemical molecules based on fuzzy neural network
CN109360610A
Method for predicting environmental harmfulness of compounds based on machine learning
CN114171137A
Interrogatory cell-based assays and uses thereof
US20130259847A1
Systems and Methods for Creating an Optimal Prediction Model and Obtaining Optimal Prediction Results Based on Machine Learning
US20200074325A1