In silico toxicity risk evaluation method using machine learning technology
By using a machine learning model to optimize similarity indices based on molecular descriptors and in vitro test data, the method achieves high accuracy and transparency in toxicity risk assessment, addressing the limitations of conventional methods.
Patent Information
- Application Number
- JP2024098796
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-19
- Publication Date
- 2026-01-07
AI Technical Summary
Conventional in silico toxicity assessment methods, such as RASAR and RAID, face challenges in achieving high prediction accuracy due to unweighted similarity indices and lack of clarity in the prediction basis, particularly in read-across methods, which rely on chemical structure similarity and experimental data agreement without computational optimization for toxicity.
A machine learning model is constructed using molecular descriptors and in chemico or in vitro test data to predict toxicity risk intensity, with the predicted values serving as a toxicity-specific similarity index, and the difference in these values used as a distance metric for optimized similarity evaluation.
This approach enhances prediction accuracy and clarity in the basis for toxicity risk assessment, enabling reliable and transparent predictions.
Smart Images

Figure 2026001453000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an in silico toxicity risk assessment method using machine learning technology. [Background technology]
[0002] In silico toxicity assessment methods, which use computational science to evaluate the toxicity of chemical substances, have been attracting considerable attention since the EU Directive on Alternatives to Animal Testing for the Evaluation of Cosmetic Ingredients was enacted in 2013. Representative in silico toxicity assessment methods include quantitative structure-activity relationship (QSAR) methods and read-across methods.
[0003] As shown in Figure 5, the QSAR method estimates the toxicity of untested substances by computationally linking the toxicity of a chemical substance with its physicochemical and biological properties. Recent advances in machine learning technology have significantly improved the accuracy of optimization for predicting toxicity from chemical properties. Meanwhile, the read-across method estimates the toxicity of untested substances based on the idea that similar chemicals should also have similar toxicity. In recent years, the toxicity of untested substances has been estimated using the Tanimoto coefficient, Euclidean distance based on physical property values, and the degree of agreement with in vitro test results as indicators of similarity. Toxicity predicted by these toxicity assessment methods can be used by some regulatory authorities as one of the bases for safety assessments of chemicals.
[0004] Several attempts have been made to combine QSAR and read-across methods. For example, the RASAR method has been devised as an example of applying the concepts of the read-across method to QSAR. In this method, the Jaccard distance of fingerprints based on chemical structure is used as a similarity index, and the similarity between chemicals in a database, combined with their toxic hazard, is used as an explanatory variable. This method allows for the construction of a QSAR model using the results of the read-across method as explanatory variables. Another application of read-across prediction, called regression analysis-based inductive DNA microarray (RAID), has also been developed. In this method, multiple gene expression is predicted using a QSAR model based on chemical structure and in vitro data, and DNA microarray analysis is performed in silico. Based on the prediction results, toxicity prediction is performed using the read-across method using principal component analysis (PCA) and enrichment analysis.
[0005] Furthermore, to improve the accuracy and robustness of evaluations, an integrated approach to testing and assessment (IATA) based on the adverse outcome pathway (AOP) of the target toxicity is recommended. As an example of an IATA, the Defined Approach (DA) for skin sensitization assessment has been published, and the 2 out of 3 DA and ITS DA are included in OECD Guideline No. 497, which establish criteria that integrate in chemico, in vitro, and in silico assessments. Skin sensitization is one of the toxicities for which the development of alternative methods to animal testing has progressed the most, and many test methods for evaluating key events (KE) based on AOPs are included in the OECD Guideline.
[0006] For example, the evaluation criteria for a 2 out of 3 DA include the Direct Peptide Reactivity Assay (DPRA), which evaluates KE 1; the KeratinoSens test, which evaluates KE 2; and the human cell line activation test (h-CLAT), which evaluates KE 3. The mouse local lymph node assay (LLNA), which is commonly used to evaluate KE 4, is an animal test. In recent years, efforts have been reported to predict LLNA EC3 values with high accuracy by using in chemico- and in vitro test data in addition to chemical structure as explanatory variables in machine learning models. The EC3 value is the concentration at which a chemical substance is exposed to the ear of a mouse and induces three times the immunity of the control, and is a quantitative indicator of skin sensitization potency.
[0007] The LLNA EC3 value is used in human risk assessment of skin sensitization as the No Expected Sensitization Induction Level (NESIL), an index corresponding to the level at which no observable adverse effects are observed. The key steps in skin sensitization risk assessment are determining the NESIL, applying the sensitization assessment factor (SAF), calculating the consumer exposure (CEL) from product use, calculating the acceptable exposure level (AEL), and comparing the CEL with the AEL. The Research Institute for Fragrance Materials, Inc. (RIFM) and the International Fragrance Association (IFRA) adopted this QRA (quantitative risk assessment) approach to assess the safe use level of fragrance ingredients that have skin sensitization potential. Several cases have been reported in which a QRA approach was used based on NESILs predicted by QSAR models, as a highly practical Next Generation Risk Assessment (NGRA) study (see Non-Patent Documents 1 and 2). [Prior art documents] [Patent documents]
[0008] [Non-Patent Document 1] Ashikaga, Takao, et al. "Establishment of a Threshold of Toxicological Concern Concept for Skin Sensitization by in Vitro / in Silico Approaches." JOURNAL OF JAPANESE COSMETIC SCIENCE SOCIETY 45.4 (2021): 331-335. [Non-patent document 2] Otsubo, Yuki, et al. "Adjustment of a no expected sensitization induction level derived from Bayesian network integrated testing strategy for skin sensitization risk assessment." The Journal of Toxicological Sciences 45.1 (2020): 57-67. Summary of the Invention [Problem to be solved by the invention]
[0009] Conventional read-across methods, including the RASAR and RAID methods mentioned above, use differences in Tanimoto coefficients, Euclidean distances based on physical property values, and equivalence of in vitro test results as indices of similarity. This method clearly identifies similar substances and makes predictions, making the prediction process highly readable. However, the similarity index used here utilizes the similarity of chemical structures and the degree of agreement of experimental data. In other words, the similarity index is not weighted for each variable, and is not computationally optimized for the toxicity of the target substance. This makes it difficult to demonstrate high prediction accuracy.
[0010] On the other hand, in conventional QSAR methods, prediction values are optimized by weighting each variable using machine learning technology, but it is difficult to show the process and basis of the prediction.
[0011] In view of the above situation, the present invention aims to provide an in silico toxicity risk assessment method that achieves both high prediction accuracy and clear indication of the basis for prediction. [Means for solving the problem]
[0012] In light of this current situation, the inventors conducted extensive research and discovered that by defining toxicity-specific similarities using machine learning technology and applying them to the read-across method, it is possible to both optimize similarity indices and clearly state the basis for prediction, thereby completing the present invention.
[0013] That is, the present invention includes the following inventions. (1) An in silico toxicity risk assessment method in which a machine learning model is constructed using an information processing system to predict the intensity of the toxicity risk of a chemical substance using molecular descriptors based on chemical structure and in chemico or in vitro test data, the predicted results of the toxicity risk intensity output by the information processing system using the machine learning model are used as an index of similarity specific to toxicity, and the toxicity risk is evaluated using a read-across method based on the similarity determined by the index.
[0014] (2) The in silico toxicity risk assessment method according to claim 1, wherein the machine learning model is constructed using molecular descriptors based on chemical structure and in chemico or in vitro test data as explanatory variables and the intensity of toxicity risk as a response variable.
[0015] (3) The in silico toxicity risk assessment method described in (1), in which the difference between the predicted values of toxicity risk intensity output by the machine learning model is used as an index of the similarity, as a distance between substances in a space optimized for the toxicity risk.
[0016] (4) An in silico toxicity risk assessment method according to (3), in which, for each evaluation data, a predetermined number of substances are collected as neighboring substances in the learning data in order of the smallest absolute value of the difference in predicted values of toxicity risk intensity between chemical substances, scores are assigned to the toxicity risk intensity of the collected neighboring substances, and the toxicity risk intensity of the evaluation data is determined based on the scores of the predetermined number of substances. [Effects of the Invention]
[0017] According to the present invention as described above, it is possible to provide an in silico toxicity risk assessment method that achieves both high prediction accuracy and clarity of the prediction basis. [Brief explanation of the drawings]
[0018] [Figure 1] FIG. 1 is a flow diagram illustrating an in silico toxicity risk assessment method. [Figure 2] FIG. 1 is an explanatory diagram of the optimization of similarity indices in an in silico toxicity risk assessment method. [Figure 3] FIG. 1 is an explanatory diagram of the distance between substances using a machine learning model. [Figure 4] This is an explanatory diagram of multi-class classification of skin sensitization using the read-across method. [Figure 5] FIG. 1 is an explanatory diagram of the conventional read-across method and QSAR method. DETAILED DESCRIPTION OF THE INVENTION
[0019] Next, an embodiment of the present invention will be described in detail with reference to the accompanying drawings. In the following embodiment, skin sensitization is evaluated as the toxicity risk of a chemical substance, and an example of predicting LLNA Category 4 class will be described. However, the in silico toxicity risk assessment method of the present invention is not limited to such prediction of skin sensitization, and can be applied to any toxicity assessment having a quantitative index.
[0020] As shown in Figures 1 and 2, the in silico toxicity risk assessment method of the present invention collects molecular descriptors based on chemical structure and in chemico or in vitro test data (S1). Using the molecular descriptors based on chemical structure and the in chemico or in vitro test data as explanatory variables and the toxicity risk intensity as a response variable, a machine learning model is constructed to predict the toxicity risk intensity of a chemical substance (S2). This machine learning model is then used to obtain a predicted value for the toxicity risk intensity (S3). In other words, the machine learning model that predicts toxicity indicators is used as a function that transforms a space of data representing the properties of chemical substances into a toxicity-specific space.
[0021] Furthermore, the toxicity risk is evaluated by a read-across method using the predicted values of the toxicity risk intensity. More specifically, the difference between the predicted values of the toxicity risk intensity output by the machine learning model is used as the distance between substances in a space optimized for the toxicity risk, substances with small absolute values of the difference between the predicted values of the toxicity risk intensity are collected as neighboring substances (S4), scores are assigned to the toxicity risk intensity of the collected neighboring substances (S5), and the toxicity risk intensity is determined based on the scores (S6).
[0022] In this example, skin sensitization is predicted as a toxic risk of chemicals, and the LLNA Category 4 class is predicted. Using chemical-based molecular descriptors and test data from various alternative methods as explanatory variables, a LightGBM model is constructed to predict the LLNA EC3 value, which represents skin sensitization strength. In other words, because this machine learning model is a function that converts chemical property data into skin sensitization strength, the difference in predicted values can be treated as the distance between substances in a space optimized for skin sensitization.
[0023] The machine learning model that predicts the LLNA EC3 value, which expresses skin sensitization potency, is a function that converts the properties of a chemical substance into skin sensitization potency. Therefore, the predicted value represents the skin sensitization-specific similarity of the chemical substance in a one-dimensional manner. In other words, as shown in Figure 3, data A about chemical substances can be used to obtain a one-dimensional quantitative predicted value P by machine learning model F. Then, by using the difference in these predicted values P as a similarity index, toxicity prediction can be performed using the read-across method based on the distance between substances optimized for skin sensitization using machine learning technology.
[0024] The predicted values from the machine learning model are used only as an indicator of similarity, and the sensitizing potency of similar substances is determined using the results of LLNA (in vivo).
[0025] This method can be applied to predict any toxicity that has a quantitative indicator. Utilizing the results of reliable in chemico / in vitro tests in QSAR is an approach that strengthens the predictive power of these tests. However, with conventional EC3 value prediction methods using machine learning models, it was difficult to demonstrate the basis for the prediction. By applying the predictions from machine learning models to the read-across method, it is possible to clearly identify substances that serve as the basis for the prediction.
[0026] (Data Source) In addition to the LLNA results, data on 195 chemical substances that have results from DPRA, KeratinoSens, and h-CLAT, which are in chemico / in vitro test methods corresponding to KE1, KE2, and KE3, will be obtained from existing literature. Specific literature includes the following (1) to (3). Note that Log means common logarithm.
[0027] (1) M. Hirota, et al., Development of an artificial neural network model for risk assessment of skin sensitization using human cell line activation test, direct peptide reactivity assay, KeratinoSens and in silico structure alert parameter. J. Appl. Toxicol., 38 (2018), pp. 514-526 (2) K. Ambe et al., Development of quantitative model of a local lymph node assay for evaluating skin sensitization potency applying machine learning CatBoost. Regulatory Toxicology and Pharmacology, 125 (2021) 105109 (3) Asai, T., Umeshita, K., Sakurai, M., & Sakane, S. (2024). Development of an in silico evaluation system that quantitatively predicts skin sensitization using OECD Guideline No. 497 ITSv2 defined approach for skin sensitization classification. Food and Chemical Toxicology, 185, 114444.
[0028] (Data arrangement) All 195 substances are sorted in descending order of EC3 (%) and assigned consecutive numbers R001 to R195. Furthermore, starting from R001, numbers 1 to 10 are assigned in order, with numbers 1 to 4 and 6 to 10 designated as training data, and number 5 designated as test data. The training data is used to train a machine learning model using three-fold cross-validation. The evaluation data is then used to objectively evaluate the predictive accuracy of the constructed model.
[0029] Using Mordred (ver. 1.2.0), molecular descriptors that can be calculated for all 195 substances and do not show the same values for all substances are obtained. Using the obtained descriptors, LLNA Positive / Negative predictions are made using Random Forest for only the training data. In this case, the variables to be used are narrowed down using Boruta (ver. 0.3), a method for variable selection using variable importance in Random Forest. The descriptors selected here are defined as "Molecular Descriptors" and used as part of the explanatory variables.
[0030] The data obtained from the literature, Log(MIT(μM)), Log(CV75(μM)), Log(Adjusted Cys depletion), Log(Adjusted Lys depletion), and Log(KEC1.5), are defined as "In chemico / in vitro tests descriptors."
[0031] (similarity index) LLNA Log(EC3(μmol / cm 2 )) is used as a similarity index specific to the skin sensitization potential of chemicals. Log(EC3(μmol / cm)) is calculated using machine learning technology. 2 )) and calculate the absolute value of the difference between the predicted values between the chemical substances as the substance distance D ij Use as.
[0032]
number
[0033] [Machine learning model] LightGBM, a commonly used GBDT algorithm with improved computational efficiency due to leaf-wise trees and exclusive feature bundling, is used as the machine learning algorithm. All models were built using Python for Windows v3.11.5 with lightgbm(4.1.0). Parameter tuning was performed using optuna(3.4.0), a parameter tuning tool based on the Bayesian optimization algorithm, with tuning performed on 'lamda_l1', 'lamda_l2', 'num_leaves', 'feature_fraction', 'bagging_fraction', 'bagging_freq', and 'min_child_samples'.
[0034] Log(EC3(μmol / cm 2 The performance of the model to predict the value of 2 The value is evaluated by RMSE and is defined as:
[0035]
number
[0036] R 2 The higher the value, the better, with 1 as the upper limit, and the lower the RMSE, with 0 as the lower limit, the better.
[0037] (Prediction by read-across method) For the evaluation data, the inter-substance distance D ij The 10 substances with the smallest valence are listed as similar substances.
[0038] Similar substances are assigned a score of 0 for NS, 1 for weak, 2 for moderate, and 3 for extreme or strong, based on the four LLNA categories (NS (non-sensitizer): EC3 >100%, weak: 100% EC3 ≥10%, moderate: 10% EC3 ≥1%, extreme or strong: 1% EC3).
[0039] The LLNA category of the target substance is then predicted based on the median score of the 10 similar substances (NS: 1> score ≧0, weak: 2> score ≧1, moderate: 3> score ≧2, extreme or strong: 4> score ≧3).
[0040] Defining the scope of application is an important factor in evaluating the reliability of a prediction method. While this invention does not explicitly define the scope of application, a previous application proposed a method for defining the scope of application using machine learning technology. Defining the scope of application for a machine learning model that predicts LLNA EC3 values can improve prediction accuracy. Another possible approach is to define the scope of application for the nearby substances collected and select only appropriate substances as nearby substances. With the latter method, prediction becomes impossible only if no suitable similar substances are obtained. Therefore, it is possible to build a prediction system with a wide range of targets that rarely produces unpredictable substances. Future challenges include defining the scope of application and expanding the training data.
[0041] Although the embodiments of the present invention have been described above, the present invention is not limited to these examples, and it goes without saying that any toxicity evaluation can be carried out within the scope of the present invention. [Example]
[0042] [Variable Selection] Using the aforementioned Mordred, 701 molecular descriptors that could be calculated for all 195 substances and had non-zero variance were obtained. LLNA Positive / Negative was used as the objective variable, and all molecular descriptors were used as explanatory variables, and variable selection was performed using Boruta on only the training data. Of the selected variables, those with correlation coefficients of 0.9 or higher between variables were excluded, leaving 9 descriptors selected and defined as "Molecular Descriptors." Table 1 shows the variables selected as "Molecular Descriptors."
[0043] [Table 1]
[0044] [Machine learning model to predict quantitative indicators of skin sensitization intensity] Using data from 175 study substances, the LLNA Log(EC3(μmol / cm)) was used as a quantitative index to express the skin sensitization strength of chemical substances. 2 A machine learning model was constructed using the objective variable, Molecular Descriptors and In chemico / in vitro test Descriptors, as explanatory variables. Table 2 shows the machine learning model for quantitatively predicting skin sensitization intensity. To avoid overfitting, three-fold cross-validation was performed, and optuna was used to optimize the hyperparameters. The prediction accuracy of the constructed LightGBM model was evaluated using R. 2 The values were 0.810 for the training data and 0.692 for the evaluation data.
[0045] [Table 2]
[0046] The prediction accuracy of the machine learning model for predicting LLNA EC3 values was measured using R 2 The value exceeded 0.8, and even in the evaluation data it was close to 0.7. It is thought that overfitting was suppressed by cross-validation.
[0047] We also calculated the variable importance of the constructed model (Table 3). The variable with the greatest contribution was "ATSC3pe (centered moreau-broto autocorrelation of lag 3 weighted by Pauling EN)." This variable represents the distribution of electronegativity of atoms in a chemical structure based on graph theory. Therefore, the arrangement of characteristic substituents, such as those containing fluorine or oxygen atoms, may contribute to the quantitative nature of skin sensitization intensity. However, it is difficult to identify the causative chemical structure based on this single variable. Table 3 shows the variable importance of the constructed LightGBM model.
[0048] [Table 3]
[0049] [Multi-class classification of skin sensitization using the read-across method] As shown in Figure 4, for each evaluation data set, 10 substances from the training data with the smallest absolute difference in predicted values between chemical substances were collected as neighboring substances (Table 4). A score was assigned to the LLNA category of the collected neighboring substances, and the LLNA category of each evaluation data set was predicted using the median score of the 10 substances. In other words, the LLNA category of the evaluation data set was predicted from the LLNA categories of the 10 neighboring substances collected based on the predicted values from the machine learning model. Table 4 ([Table 4-1] to [Table 4-4]) shows the prediction results for the evaluation data.
[0050] [Table 4-1]
[0051] [Table 4-2]
[0052] [Table 4-3]
[0053] [Table 4-4]
[0054] For the 20 substances in the evaluation data, the agreement rate for prediction of the LLNA Category 4 class was 0.60 (Table 5). Table 5 shows the evaluation of the prediction accuracy of the evaluation data.
[0055] [Table 5]
[0056] Among the substances with incorrect predictions, the overestimation and underestimation rates were both 0.20. When an error of one class of LLNA category was allowed, the concordance rate was 0.95. Only one substance had an error of two or more classes, and it was predicted as Moderate when the correct category was NS (CAS No. 108-95-2, Phenol).
[0057] This represents an overestimation of toxicity, and is a relatively tolerable error in terms of reducing the risk caused by uncertainty in the prediction. On the other hand, the errors for the four underestimated substances were all one class or less. Since underestimation of toxicity carries the risk of harming human health, particular attention should be paid to this as a prediction error.
[0058] In this method, the median of the LLNA category of the 10 substances is used as the predicted value, but the balance between overestimation and underestimation can be adjusted by adjusting the criteria, such as using the mean or the third quartile. In particular, by adopting strict criteria such as the third quartile, it is possible to design the method to more easily avoid false negatives and to avoid underestimating toxicity that is of concern.
[0059] These results demonstrate the usefulness of applying predictions from machine learning models to the read-across method. Furthermore, our previous research has shown that using skin sensitization data for similar substances as explanatory variables in a machine learning model improves the accuracy of EC3 value predictions. Therefore, we believe that predictive accuracy can be improved by reconstructing a machine learning model using the results of a read-across method using a machine learning model. By reusing the reconstructed machine learning model in the read-across method, it is possible to build a multi-layered prediction system.
[0060] A QRA approach is being discussed to determine the safe concentrations of fragrance ingredients and preservatives based on skin sensitization predicted by QSAR and read-across methods. With this read-across method, the agreement rate is 0.95 when an error of one class is allowed, and a sufficiently safe NESIL can be set by using the smallest EC3 (%) of the class that is overestimated by one class relative to the predicted value. We hope that this discussion will progress toward realizing a more appropriate assessment system, such as a method for setting SAFs that take into account the effects of volatility and retention factors.
[0061] Although this study focused on skin sensitization, this approach can be applied to any toxicity assessment. We hope that this new approach, which combines QSAR and read-across methods, will contribute to the advancement of in silico toxicity prediction and evaluation research.
Claims
1. Using molecular descriptors based on chemical structure and in chemico or in vitro test data, we construct a machine learning model that predicts the level of toxicity risk of chemical substances using an information processing system. The prediction result of the toxicity risk intensity output by the information processing system using the machine learning model is used as an index of similarity specific to toxicity, The in silico toxicity risk assessment method assesses the toxicity risk using a read-across technique based on the similarity of the index.
2. The machine learning model: The explanatory variables are molecular descriptors based on chemical structure and in chemico or in vitro test data. The in silico toxicity risk assessment method according to claim 1, wherein the method is constructed using the intensity of toxicity risk as a response variable.
3. The in silico toxicity risk assessment method according to claim 1, wherein the difference between the predicted values of toxicity risk intensity output by the machine learning model is taken as the distance between substances in a space optimized for the toxicity risk, and this is used as an index of the similarity.
4. For each evaluation data, a predetermined number of substances are collected as neighboring substances from the learning data in order of smallest absolute value of difference in predicted value of toxicity risk intensity between chemical substances; A score is assigned to the toxicity risk intensity of the collected nearby substances, The in silico toxicity risk assessment method according to claim 3 , wherein the toxicity risk intensity of the assessment data is determined based on the scores of the predetermined number of substances.