An ensemble learning method to predict lethal effects of short-term exposure to neurotoxicants
By constructing an ensemble learning model that combines the unique heat encoding features of different test animals and exposure routes, the limitations of existing technologies in predicting the lethal effects of short-term exposure to neurotoxins are overcome. This achieves efficient and accurate prediction results, expands the application scope of the model, and reduces experimental costs.
Patent Information
- Application Number
- CN202310122078.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-16
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2043-02-16
AI Technical Summary
Existing technologies struggle to accurately predict the short-term lethal effects of neurotoxic substances on different test animals and through multiple exposure routes under conditions of high efficiency, low cost, and low animal numbers, and the models lack sufficient applicability and robustness.
An ensemble learning model is constructed by combining one-hot encoded features from different test animals and multiple exposure pathways. An ensemble learning method with a hard voting strategy is adopted to comprehensively consider molecular structure and biological exposure features. The model is constructed using KNN, SVM, RF and GBDT algorithms, which breaks through the limitations of traditional single machine learning.
It achieves efficient and accurate prediction of the lethal effects of short-term exposure to neurotoxins, has a wide range of applications and good robustness, saves experimental time and costs, and reduces the number of animals used.
Smart Images

Figure CN116030905B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of computational toxicology for chemicals health risk assessment and management, and relates to a method for high-throughput prediction of short-term exposure lethal effects of neurotoxicants based on quantitative structure-activity relationship (QSAR) and ensemble learning. BACKGROUND
[0002] A neurotoxicant is any chemical, biological or physical substance that can cause neurotoxicity. More than hundreds of neurotoxicants have been reported to cause poisoning events in test animals worldwide. The toxic mechanisms of these neurotoxicants involve acetylcholinesterase inhibition and adrenergic receptor blockade, etc., leading to severe complications related to central and peripheral nervous systems of test animals, such as anorexia, convulsions, skeletal paralysis, ataxia, tremors, loss of righting reflex and death. Short-term exposure lethal effect is one of the core contents of chemicals health risk assessment and management, which can be evaluated by the median lethal dose (LD 50 ) of test animals after chemical exposure. The number of neurotoxicants is large, and the short-term exposure lethal effects of most of them are still unknown, and the experimental test results are uneven. In view of the harmful effects of neurotoxicants on health, it is necessary to predict the short-term exposure lethal effects before they enter the environment under the same method.
[0003] The relevant guideline (OECD guideline 424) issued by the Organization for Economic Cooperation and Development (OECD) is a test guideline for rodent neurotoxicity studies, which can be used for acute neurotoxicity studies. The guideline requires to evaluate a series of behaviors of animals that may be affected by neurotoxicants during each experimental observation period, and death is one of the key observation indicators. However, this experimental test method is expensive and time-consuming, and is contrary to animal ethics, making it difficult to achieve one-by-one evaluation of the short-term exposure lethal effects of a large number of neurotoxicants, and it is necessary to develop high-throughput quantitative prediction technology. Based on QSAR, the correlation between the molecular structure characteristics of neurotoxicants and the induced short-term exposure lethal effects can be established, which can realize the efficient prediction of the short-term exposure lethal effects of neurotoxicants, save time, reduce cost and reduce the number of animals required for experimental tests.
[0004] Traditional linear QSAR cannot meet the prediction requirements because it cannot handle the nonlinear relationship between molecular structure characteristics and short-term exposure lethal effects. In recent years, the rapidly developing machine learning algorithm has data adaptive characteristics, and can provide superior performance for analyzing high-order, high-dimensional and nonlinear relationship data, and has been applied to mining the internal correlation between molecular structure characteristics and short-term exposure lethal effects. Ensemble learning can obtain significantly superior generalization performance than single machine learning algorithm by combining multiple machine learning algorithms, and is expected to play an active role in quantitative prediction of short-term exposure lethal effects of neurotoxicants, and realize efficient filling of data gaps of short-term exposure lethal effects of neurotoxicants.
[0005] Currently, there are some studies that have constructed QSAR prediction models for lethal effects of short-term exposure to neurotoxicants (toxicity endpoint value is LD 50 ). Literature 1 “Chemical Research in Toxicology, 2006, 19(2): 209-216.” considered the lethal effects of short-term exposure of 38 organophosphorus compounds in male rats via oral route, and the toxicity mechanisms of these compounds mainly include acetylcholinesterase inhibition. Literature 1 constructed a QSAR model to predict the lethal effects of short-term exposure of organophosphorus compounds based on descriptors describing absorption, distribution, metabolism, and excretion processes, and the leave-one-out cross-validation coefficient was 0.82. Literature 2 “Toxicology Research, 2020, 9(3): 164-172.” constructed an extreme tree regression model for lethal effects of short-term exposure of 422 organic compounds in the autonomic nervous system of mice via intraperitoneal injection based on data collected from the ChemIDplus database using PyBioMed descriptors, and the external validation coefficient was 0.784. In addition, Literature 3 “Environmental Science & Technology, 2022, 56(1): 335-348.” collected data on the toxicity of 128 compounds to multiple bird species from the Ecotoxicity Database of Pesticides of the Office of Pesticide Programs of the U.S. Environmental Protection Agency and the literature, and the toxicity mechanisms of these compounds involve acetylcholinesterase inhibition. Literature 3 constructed a QSAR model for predicting the acute oral toxicity of pesticides to birds using 11 two-dimensional descriptors, and the external validation coefficient of the toxicity prediction model for American quail was 0.648. In summary, although the above models are all prediction models for lethal effects of short-term exposure to neurotoxicants, they all use traditional statistical methods or a single machine learning algorithm for modeling, and the robustness and prediction ability of the models need to be improved. Moreover, the biological exposure routes considered are single, which limits the application scope of the models to toxicity prediction under specific exposure conditions. Furthermore, the data sets of these models contain less structural diversity of neurotoxicants, which leads to a small application domain of the models. Specifically, Literature 1 has a small data set and contains non-diverse compounds, only involving 38 organophosphorus compounds in male rats via oral short-term exposure, the prediction model has a narrow application scope, and the application domain is not explicitly characterized. Literature 2 has a large data set, but only covers lethal effects of intraperitoneal injection in mice; Literature 3 considers experimental data from multiple test animals, but only involves oral exposure toxicity in birds. Literature 2 and 3 cannot be used for prediction of neurotoxicity effects in multiple organisms and exposure routes. Therefore, it is necessary to break through the research paradigm of single algorithm modeling and develop an integrated learning model for lethal effects of short-term exposure to neurotoxicants based on heterogeneous data sets from different test animals and multiple exposure routes.
[0006] Based on the above reasons, 574 high-dimensional and diverse short-term exposure lethal effect data of neurotoxicants covering different test animals and various exposure routes were collected and sorted from the PubChem database. The experimental information of different test animals and various exposure routes was one-hot encoded, and the chemical information of the calculated molecular structure features was coupled to obtain the modeling features. Three different machine learning algorithms were used as the base regressor, and the hard voting strategy was used to construct an integrated learning model for predicting the short-term exposure lethal effect of neurotoxicants, and the application domain of the model was characterized. SUMMARY
[0007] The present application constructs an efficient and low-cost integrated learning method for predicting the short-term exposure lethal effect of neurotoxicants. This method can directly predict the short-term exposure lethal effect according to the molecular features calculated from the molecular structure of neurotoxicants and the specified biological exposure route. This method is based on a variety of neurotoxicants, breaking through the research paradigm of traditional machine learning modeling, which only considers molecular structure features and ignores biological exposure features. Different test animals and various exposure routes are considered and one-hot encoded, coupled with molecular structure features, and an integrated learning prediction model based on the hard voting strategy is developed. The model established in the present application has high internal robustness and external prediction ability, is easy to operate, and can save time, cost and animal number for experimental testing; it can be used as an efficient tool for predicting the short-term exposure lethal effect of neurotoxicants, and provides basic toxicity data for human health risk assessment and management of chemicals.
[0008] Technical scheme of the present application:
[0009] An integrated learning method for predicting the short-term exposure lethal effect of neurotoxicants, comprising the following steps:
[0010] (1) Construction of short-term exposure lethal effect data set of neurotoxicants
[0011] Short-term exposure lethal effect data of neurotoxicants were collected and sorted from the PubChem database, including two kinds of rodents, rats and mice, and four exposure routes, intraperitoneal injection, intravenous injection, oral injection and subcutaneous injection, a total of 574 toxicity test data; the short-term exposure lethal effect endpoint value (LD 50 , mg / kg) was converted to molar concentration (mol / kg), and then to negative logarithm (pLD 50); the toxic mechanism involves alpha adrenergic receptor blockade, beta adrenergic receptor blockade, ganglionic blockade, acetylcholinesterase inhibition, etc.; the collected neurotoxicants belong to pesticides, antipsychotic drugs, dye and rubber auxiliary intermediates, etc., including organic acids, ethers, esters, ketones, alcohols, amides, anilines, polycyclic aromatic hydrocarbons and their substitutes, halogenated alkanes, halogenated alkenes, heterocyclic compounds and their derivatives, etc., excluding inorganic compounds, organometallic compounds, mixtures (mainly macromolecular salt compounds);
[0012] (2) Molecular structure characteristics and biological exposure characteristics of neurotoxicants
[0013] Based on PubChemPy, the PubChem CID and the corresponding 2D structure of 574 neurotoxicants were batch obtained and saved in SDF format file; the SDF file was input into the PaDEL-Descriptor 2.21 version software to calculate the 1D and 2D molecular structure characteristics; the one-hot encoding was performed for two test animals and four exposure routes; the features with variance of 0 were removed, and the features with Pearson correlation coefficient greater than 0.9 were removed, and the recursive feature elimination method was used for feature selection;
[0014] (3) Model training
[0015] The 1D and 2D molecular structure characteristics of neurotoxicants and the experimental characteristics of different test animals and multiple exposure routes were coupled as the feature input of the model, and the pLD 50 value was taken as the prediction endpoint of the model to construct a machine learning regression model; the data set was randomly divided into training set and validation set in the ratio of 4:1, in order to reduce random error, ten-fold cross-validation was adopted for internal validation, and the validation set was used for external validation of the model; four machine learning algorithms, namely K nearest neighbors (KNN), support vector machine (SVM), random forest (RF) and gradient boosting decision tree (GBDT), were used to construct models, the three algorithms with the best model performance were taken as the base regressors, and the hard voting strategy was used to construct an ensemble learning model, and the weights of each base regressor in the hard voting strategy were the same; the best hyperparameters of the machine learning algorithm were determined by grid search, and the best hyperparameters of the ensemble learning model were derived from the best hyperparameters of each base regressor;
[0016] The following are the model optimization hyperparameters, the best hyperparameters of KNN: the number of neighboring points (n_neighbors) is 4, the weight function (weights) is distance; the best hyperparameters of SVM: radial basis as kernel function, regularization parameter (C) is 200, kernel coefficient (gamma) is 0.005, epsilon is 0.1; the best hyperparameters of RF: the number of decision trees (n_estimators) is 290, the maximum number of features (max_features) of each decision tree is 11, the random seed (random_state) is set to 86; the best hyperparameters of GBDT: n_estimators is 200, max_features is 20, random_state is set to 0; the remaining parameters of the above algorithms are default values;
[0017] (4) Model evaluation and verification
[0018] The fitting ability of the model is evaluated using the square of the correlation coefficient adjusted for degrees of freedom (R 2 adj ) and the root mean square error (RMSE). The robustness of the model is evaluated using the ten-fold cross-validation coefficient (Q 2 CV ) of the training set. The predictive ability of the model is evaluated using the external validation coefficient (Q 2 ext ) of the validation set.
[0019] The predictive performance of the model is:
[0020] The performance of model 1 constructed by KNN algorithm: training set R 2 adj = 0.772, RMSE CV = 0.394, Q 2 CV = 0.783; validation set RMSE ext = 0.414, Q 2 ext = 0.799;
[0021] The performance of model 2 constructed by SVM algorithm: training set R 2 adj = 0.821, RMSE CV = 0.346, Q 2 CV = 0.830; validation set RMSE ext = 0.386, Q 2 ext = 0.825;
[0022] The performance of model 3 constructed by RF algorithm: training set R 2adj = 0.838, RMSE CV = 0.335, Q 2 CV = 0.846; validation set RMSE ext = 0.371, Q 2 ext = 0.838;
[0023] Model 4 performance built by GBDT algorithm: training set R 2 adj = 0.844, RMSE CV = 0.326, Q 2 CV = 0.852; validation set RMSE ext = 0.383, Q 2 ext = 0.827;
[0024] The three algorithms with the best model effect are SVM, RF, and GBDT, which are used as base regressors, and the ensemble learning model 5 performance built by adopting hard voting strategy: training set R 2 adj = 0.863, RMSE CV = 0.307, Q 2 CV = 0.870; validation set RMSE ext = 0.344, Q 2 ext = 0.861; the ensemble learning model has the best performance, and is used as the final model for predicting the lethal effect of short-term exposure to neurotoxicants;
[0025] (5) Application domain representation
[0026] The Williams plot is drawn based on the standardized residual and the leverage distance (h i ) of the ensemble learning model 5, wherein the training set features include biological exposure features in addition to molecular structure features; the calculation formula of h i is as follows:
[0027] H = X (X T X) -1 X T (1)
[0028] h i = [H] ii (2)
[0029]
[0030] Wherein, X is a matrix of multiple neurotoxicants and multiple features in the training set, H is a hat matrix, and hi Leverage distance of the ith neurotoxicant, [H] ii h* is the defined warning value of the leverage distance, p is the number of features of the training set, n is the number of neurotoxicants of the training set;
[0031] If the neurotoxicant h i is greater than h*, the neurotoxicant is considered to be outside the application domain. Therefore, the ensemble learning model 5 is suitable for predicting the pLD i value of the neurotoxicant with h 50 less than 0.137.
[0032] The beneficial effects of the present application are:
[0033] (1) The ensemble learning model 5 combines the advantages of three different machine learning algorithms (SVM, RF, GBDT), and has better robustness and prediction ability compared with the models of documents 1, 2 and 3 and the models 1, 2, 3 and 4 constructed by a single machine learning algorithm, and has a clear application domain;
[0034] (2) The modeling data set covers toxicity data of two test animals (rats and mice) and four exposure routes (intraperitoneal, intravenous, oral and subcutaneous injection), breaking through the research paradigm of traditional machine learning modeling which only considers molecular structure features and ignores biological exposure features, and the application domain of the model is wider;
[0035] (3) The number of neurotoxicants used for modeling is large and the structure is diverse, including organic acids, ethers, esters, ketones, alcohols, amides, anilines, polycyclic aromatic hydrocarbons and their substitutes, halogenated alkanes, halogenated alkenes, heterocyclic compounds and their derivatives, etc., and the model has good generalization ability;
[0036] (4) The method is simple and efficient, which can save the time, cost and number of animals for experimental testing, and is expected to play an important role in the quantitative prediction of neurotoxicity caused by short-term exposure, fill the data gap of neurotoxicity caused by short-term exposure, provide a basic tool for chemical health risk assessment and management, and serve the national major needs of chemical risk control and new pollutant management. BRIEF DESCRIPTION OF DRAWINGS
[0037] Figure 1 is the fitting plot of the experimental value and the predicted value of the pLD 50 of the training set, and the training set is 459 toxicity data.
[0038] Figure 2 is the fitting plot of the experimental value and the predicted value of the pLD 50 of the validation set, and the validation set is 115 toxicity data.
[0039] Figure 3This is the Williams diagram of the model. Circles represent neurotoxins in the training set, triangles represent neurotoxins in the validation set, and the warning value h* is 0.137. DETAILED DESCRIPTION
[0040] The specific implementation of the present invention is further described below in conjunction with the accompanying drawings and technical solutions.
[0041] Example 1
[0042] Given a neurotoxin, diisopropanolamine (CAS: 110-97-4 / PubChem CID: 8086), we need to predict its short-term lethal effect (endpoint value pLD 50 ). First, the SDF structure file was obtained based on the PubChem CID of diisopropanolamine, and its 1D and 2D molecular structures were calculated using PaDEL-Descriptor software. The h value calculated according to formulas (1) and (2) is 0.032, which is less than Figure 3 The lever alert value h* of the neurotoxicant is (0.137), so diisopropanolamine is within the application domain of the model. The test animal is specified as mouse, the exposure route is intraperitoneal injection, and the values of the molecular structure characteristics calculated above are input into the model to obtain pLD 50 The predicted value is 3.003, and the experimentally determined pLD 50 The value is 3.142, which is in good agreement with the experimental data.
[0043] Example 2
[0044] Given a neurotoxin, guaifenesin (CAS: 93-14-1 / PubChem CID: 3516), we need to predict its short-term lethal effect (endpoint value pLD 50 ). First, the SDF structure file was obtained based on the PubChem CID of guaifenesin, and its 1D and 2D molecular structures were calculated using PaDEL-Descriptor software. The h value calculated according to formulas (1) and (2) is 0.042, which is less than Figure 3 The lever alert value h* for a medium neurotoxicant is (0.137), so guaifenesin is within the application domain of the model. The test animals are designated as mice, the exposure route is subcutaneous injection, and the values of the molecular structure features calculated above are input into the model to obtain pLD 50 The predicted value is 2.545, and the experimentally determined pLD 50 The value is 2.394, which is in good agreement with the experimental data.
[0045] Example 3
[0046] Given a neurotoxicant 3-(m-tolyl)octanone (CAS: 3483-17-8 / PubChem CID: 198874), the lethal effect of its short-term exposure (the endpoint value is pLD 50 ) is to be predicted. First, the SDF structure file is obtained according to the PubChem CID of 3-(m-tolyl)octanone, and its 1D and 2D molecular structure is calculated using PaDEL-Descriptor software. The h value calculated according to formulas (1) and (2) is 0.033, which is less than Figure 3 the lever alert value h* (0.137) of the neurotoxicant in the middle, so 3-(m-tolyl)octanone is within the application domain of the model. The test animal is specified as mouse, the exposure route is oral injection, and the values of the above calculated molecular structure features are input into the model, and the predicted value of pLD 50 is 2.551, and the experimentally determined pLD 50 value is 2.550, and the data of the predicted value and the experimental value are very consistent.
Claims
1. An ensemble learning method for predicting the lethal effects of short-term exposure to neurotoxicants, characterized in that: Here are the steps: (1) Constructing a dataset of lethal effects of short-term exposure to neurotoxins The short-term lethal effect data of neurotoxicants were collected and sorted from the PubChem database. The toxicity endpoint value was the median lethal dose (LD). 50 , convert it to molar concentration and then to negative log pLD 50 The dataset covers two test animals, rats and mice, and four exposure routes: intraperitoneal injection, intravenous injection, oral injection, and subcutaneous injection, totaling 574 toxicity test data. The dataset includes multiple toxicity mechanisms, specifically α-adrenergic receptor blockade, β-adrenergic receptor blockade, ganglion blockade, and acetylcholinesterase inhibition. Neurotoxicants include organic acids, ethers, esters, ketones, alcohols, amides, anilines, polycyclic aromatic hydrocarbons and their substituents, halogenated alkanes, halogenated alkenes, heterocyclic compounds and their derivatives. (2) Expression of molecular structural characteristics and biological exposure characteristics of neurotoxicants Based on the PubChem database, 574 neurotoxicants' PubChem CIDs and their corresponding 2D structures were obtained in batches. The PaDEL-Descriptor software was used to calculate the molecular 1D and 2D structural features. One-hot encoding was performed for two test animal species and four exposure routes. The preprocessing process includes first removing features with a variance of 0 and then removing features with a Pearson correlation coefficient greater than 0.9, and using recursive feature elimination to perform feature screening; (3) Model training The 1D and 2D molecular structure characteristics of neurotoxicants are coupled with experimental characteristics of different test animals and multiple exposure pathways as feature inputs of the model. 50 The value was used as the prediction endpoint of the model to construct a machine learning regression model; the dataset was randomly split into a training set and a validation set in a 4:1 ratio, and ten-fold cross-validation was used for internal validation, while the validation set was used for external validation of the model; four machine learning algorithms, namely K-nearest neighbor, support vector machine, random forest, and gradient boosting decision tree, were used to construct models respectively. The three algorithms with the best model performance were used as base regressors, and a hard voting strategy was adopted to construct an ensemble learning model. In the hard voting strategy, the weights of each base regressor were the same; the optimal hyperparameters of the machine learning algorithm were determined through grid search, and the ensemble learning model was constructed based on the optimal hyperparameters; The following are the hyperparameters for model optimization: the optimal hyperparameters for K-nearest neighbor are: 4 neighbors, weight function: distance; the optimal hyperparameters for support vector machine are: radial basis function as kernel function, regularization parameter: 200, kernel coefficient: 0.005, epsilon: 0.1; the optimal hyperparameters for random forest are: 290 decision trees, 11 maximum features per tree, and 86 random seeds; the optimal hyperparameters for gradient boosted decision tree are: 200 decision trees, 20 maximum features per tree, and 0 random seeds; The three algorithms with the best model performance, namely support vector machine, random forest, and gradient boosting decision tree, are used as base regressors, and a hard voting strategy is adopted to build an ensemble learning model; (4) Model evaluation, validation, and application domain characterization Use the square of the correlation coefficient R after adjusting for degrees of freedom 2 adj The fitting ability of the model is evaluated by the root mean square error RMSE, and the internal ten-fold cross validation coefficient Q 2 CV Evaluation model robustness and external validation coefficient Q 2 ext Evaluate the predictive ability of the model; use Williams graphs to characterize the application domain of the model, where the training set features include biological exposure features in addition to molecular structure features; Ensemble learning model training set R 2 adj =0.863,RMSE CV =0.307,Q 2 CV =0.870; validation set RMSE ext =0.344,Q 2 ext =0.861; this model was used as the final model for predicting the lethal effects of short-term exposure to neurotoxins.
2. The method according to claim 1, characterized in that In addition to molecular structure features, the training set features in the model application domain characterization also include biological exposure features, and the defined application domain threshold is: leverage warning value h* = 0.
137.
3. The method according to claim 1, characterized in that The data on the lethal effects of short-term exposure to neurotoxicants cover two test animals, rats and mice, and four exposure routes: intraperitoneal injection, intravenous injection, oral injection and subcutaneous injection.
4. The method according to claim 1, wherein Neurotoxicants include pesticides, antipsychotic drugs, dyes and rubber chemical intermediates.
5. The method according to claim 1, wherein Nerve poisons include organic acids, ethers, esters, ketones, alcohols, amides, anilines, polycyclic aromatic hydrocarbons and their substituents, halogenated alkanes, halogenated alkenes, heterocyclic compounds and their derivatives.
Citation Information
Patent Citations
Intelligent drug toxicity judgment method and device and computer readable storage medium
CN110322972A
Integrated learning method for screening carcinogenic chemicals
CN114743614A