Information processing system for predicting correctness of evaluation by evaluation method for predicting positivity / negativity to show application range of the evaluation method, control program, and information processing method

The information processing system addresses the challenge of defining the applicable range for toxicity prediction models by using machine learning to predict evaluation correctness, thereby improving reliability and regulatory acceptance.

JP2025091722APending Publication Date: 2025-06-19SUNSTAR INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023207144
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-07
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Current toxicity prediction methods, such as QSAR models, face challenges in defining the applicable range of their predictions, which is crucial for reliability and regulatory acceptance.

Method used

An information processing system and method that utilize machine learning to predict the correctness of evaluations and define the applicable range by sorting and processing objective and explanatory variables, including oversampling to address imbalances in data.

Benefits of technology

The system effectively predicts the correctness of toxicity evaluations and defines the applicable range, enhancing the reliability and transparency of toxicity prediction models, thus supporting regulatory compliance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025091722000001_ABST
    Figure 2025091722000001_ABST
Patent Text Reader

Abstract

To provide an information processing system for predicting the correctness of an evaluation by an evaluation method for predicting positivity / negativity to show an application range of the evaluation method, a control program, and an information processing method.SOLUTION: An information processing system stores information the correctness of an evaluation by an evaluation method as an objective variable, stores at least one piece of information among information based on an experiment, information about a chemical structure, and information on substances similar to a test substance as an explanatory variable, distributes each information of the objective variable and the explanatory variable into learning data and data for verification to store them, constructs a prediction model for predicting the correctness of the evaluation by using the objective variable and the explanatory variable of the learning data by a machine learning method, evaluating the prediction model by using the objective variable and the explanatory variable for the data for verification, performs necessary restructure of a prediction model, and outputs the correctness of the evaluation by the evaluation method by using the prediction model.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing system, a control program, and an information processing method that predict the correctness of an evaluation by an evaluation method for positive / negative prediction and indicate the applicable range of the evaluation method.

Background Art

[0002] Triggered by the full implementation of the 2013 EU directive on the replacement of animal experiments in the safety testing of cosmetic ingredients, animal experiments related to cosmetic development need to be replaced by alternative testing methods. As one of such alternative testing methods, a toxicity prediction method using machine learning technology has attracted attention and has been rapidly developed in recent years.

[0003] (Quantitative) Structure-Activity Relationship ((Q)SAR) is a method that mathematically expresses the relationship between the toxicity of a chemical substance and the characteristic quantities obtained from its chemical structure, and predicts the toxicity of substances with unknown toxicity. By incorporating an algorithm using machine learning technology, an improvement in its predictability is expected.

[0004] On the other hand, with the complication of the (Q)SAR model, the difficulty of appropriately evaluating the model has increased. For the acceptance of (Q)SAR models using highly novel algorithms, the "OECD Principles for the Validation of (Q)SAR Models for Regulatory Purposes" (Non-Patent Document 1) for evaluating the reliability and transparency of toxicity prediction models has been agreed upon.

[0005] As the "OECD Principles for the Validation of (Q)SAR Models for Regulatory Purposes", 1: Definition of endpoint (a defined endpoint) 2: Unambiguous algorithm (An unambiguous algorithm) 3: Definition of applicable range (a defined domain applicability) 4: Appropriate measures of goodness-of-fit, robustness and predictivity 5: a mechanistic interpretation, if possible are mentioned.

[0006] Among these, as the "definition of the scope of application", although thresholds such as "LogP < 3.5" and "molecular weight < 500" may be set, the prediction by the (Q)SAR model may strongly depend on the range of toxicity values, chemical structures, and mechanisms of action of the training data. There is still no defined method for the "definition of the scope of application", and it has not been possible to show the basis for determining that the prediction result is calculated within the scope of application.

Prior Art Documents

Non-Patent Documents

[0007]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0008] In view of such problems, an object of the present invention is to provide an information processing system, a control program, and an information processing method that predict the correctness of an evaluation by an evaluation method for positive / negative prediction and show the scope of application of the evaluation method.

Means for Solving the Problems

[0009] The present invention includes the following inventions. (1) An information processing system that predicts the correctness of an evaluation by an evaluation method for positive / negative prediction and shows the applicable range of the evaluation method, comprising: an objective variable storage unit that stores information on the correctness / incorrectness of the evaluation by the evaluation method as an objective variable; an explanatory variable storage unit that stores at least one of information based on an experiment used in the evaluation method, information on a chemical structure, and information on a substance similar to a test substance as an explanatory variable; a sorting processing unit that sorts and stores each piece of information on the objective variable and the explanatory variable into learning data and verification data; a prediction model construction processing unit that constructs a prediction model for predicting the correctness / incorrectness of the evaluation using the objective variable and the explanatory variable of the learning data by a machine learning method, evaluates the prediction model using the objective variable and the explanatory variable of the verification data, and performs necessary reconstruction of the prediction model; and an output processing unit that outputs the correctness / incorrectness of the evaluation by the evaluation method using the prediction model.

[0010] (2) The information processing system according to (1), further comprising an objective variable generation processing unit that, when there is an imbalance exceeding a predetermined range in the sample amounts of the information on the correctness / incorrectness, which is the objective variable, or true positives correctly evaluated as positive, true negatives correctly evaluated as negative, false negatives incorrectly evaluated as negative, and false positives incorrectly evaluated as positive, uses an oversampling method for the minority samples to detect neighboring data, internally generate data, and add this as a minority sample to the objective variable.

[0011] (3) The information processing system according to (2), wherein the correctness / incorrectness of the evaluation, which is the objective variable, is information on the true positives, true negatives, false positives, and false negatives, and the output processing unit outputs which of the true positives, true negatives, false positives, and false negatives of the evaluation by the evaluation method using the prediction model.

[0012] An information processing system that predicts the risk of an error in an evaluation by an evaluation method for performing positive / negative prediction and shows the scope of application of the evaluation method, comprising: an objective variable storage unit that stores information on the risk of an error in the evaluation by the evaluation method as an objective variable; an explanatory variable storage unit that stores at least one of information based on an experiment used in the evaluation method, information on a chemical structure, and information on a substance similar to a test substance as an explanatory variable; a first sorting processing unit that sorts and stores each piece of information on the objective variable and the explanatory variable into learning data and verification data; a prediction model construction processing unit that constructs a prediction model for predicting the risk of the evaluation error using the objective variable and the explanatory variable of the learning data by a machine learning method, evaluates the prediction model using the objective variable and the explanatory variable of the verification data, and performs necessary reconstruction of the prediction model; and an output processing unit that outputs a value of the risk that the evaluation by the evaluation method is in error using the prediction model.

[0013] (5) The information processing system according to (4), further comprising a second sorting processing unit that sorts the correctness of the evaluation by the evaluation method into any one of true positive correctly evaluated as positive being positive, true negative correctly evaluated as negative being negative, false negative wrongly evaluated as positive being negative, and false positive wrongly evaluated as negative being positive, wherein the risk information as the objective variable includes distribution information of the true positive, the true negative, the false positive, and the false negative assumed in a classification of a boundary-type error existing near a threshold within a range where its error can be correctly classified, and distribution information of the true positive, the true negative, the false positive, and the false negative assumed in a classification of an outlier-type error deviating from the overall trend due to excessive properties, and the prediction model construction processing unit constructs a prediction model for predicting the risk of the boundary-type error and the risk of the outlier-type error using the objective variable and the explanatory variable of the learning data by a machine learning method, and the output processing unit outputs a value of the risk that the evaluation by the evaluation method is the boundary-type error and a value of the risk that the evaluation by the evaluation method is the outlier-type error using the prediction model.

[0014] (6) The prediction model predicts the distribution (Y B ) of the boundary type and the distribution (Y O ) of the outlier type, and the output processing unit calculates the value (R B ) of the risk that becomes an error of the boundary type and the value (R O ) of the risk that becomes an error of the outlier type according to the following formula (1), and outputs them. The information processing system according to (5).

[0015]

Number

[0016] (7) The output processing unit integrates the value (R B ) of the risk that becomes an error of the boundary type and the value (R O ) of the risk that becomes an error of the outlier type according to the following formula (2), and calculates the value (R M ) of the risk that becomes an error in the evaluation according to the evaluation method. The information processing system according to (6).

[0017]

Number

[0018] (8) The output processing unit outputs the predicted correct / incorrect of the evaluation according to the evaluation method based on whether the value of the risk that becomes an error in the evaluation according to the evaluation method exceeds a predetermined threshold. The information processing system according to (4).

[0019] (9) A control program for causing a computer to function as the information processing system according to (1), the control program for causing a computer to function as the above sorting processing unit, the above prediction model construction processing unit, and the above output processing unit.

[0020] (10) A control program for causing a computer to function as the information processing system according to (4), the control program for causing a computer to function as the above first sorting processing unit, the above prediction model construction processing unit, and the above output processing unit.

[0021] (11) An information processing method executed by a system that has a target variable storage unit that stores information on the correctness / incorrectness of evaluation by an evaluation method for predicting positive / negative, as a target variable, and an explanatory variable storage unit that stores at least one piece of information among information based on experiments used in the evaluation method, information on chemical structures, and information on substances similar to the test substance, as an explanatory variable, predicts the correctness / incorrectness of evaluation by an evaluation method for predicting positive / negative, and indicates the applicable range of the evaluation method. The information processing method includes a distribution procedure for distributing and storing each piece of information on the target variable and the explanatory variable into learning data and verification data, a prediction model construction procedure for constructing a prediction model that predicts the correctness / incorrectness of the evaluation using the target variable and the explanatory variable of the learning data by a machine learning method, evaluating the prediction model using the target variable and the explanatory variable of the verification data, and performing necessary reconstruction of the prediction model, and an output procedure for outputting the correctness / incorrectness of evaluation by the evaluation method using the prediction model.

[0022] (12) An information processing method executed by a system that has a target variable storage unit that stores information on the correctness / incorrectness of evaluation by an evaluation method for predicting positive / negative, as a target variable, and an explanatory variable storage unit that stores at least one piece of information among information based on experiments used in the evaluation method, information on chemical structures, and information on substances similar to the test substance, as an explanatory variable, predicts the risk of error in evaluation by an evaluation method for predicting positive / negative, and indicates the applicable range of the evaluation method. The information processing method includes a distribution procedure for distributing and storing each piece of information on the target variable and the explanatory variable into learning data and verification data, a prediction model construction procedure for constructing a prediction model that predicts the risk of error in the evaluation using the target variable and the explanatory variable of the learning data by a machine learning method, evaluating the prediction model using the target variable and the explanatory variables of the verification data, and performing necessary reconstruction of the prediction model, and an output procedure for outputting the value of the risk that the evaluation by the evaluation method is incorrect using the prediction model.

Advantages of the Invention

[0023] According to the present invention configured as described above, it is possible to predict the correctness of an evaluation by an evaluation method for performing positive / negative prediction and to show the applicable range of the evaluation method.

[0024] "Definition of applicable range" in the toxicity evaluation of chemical substances is an important item indicating the reliability of the hazard evaluation method, and it is possible to provide a machine learning model that predicts its correctness / incorrectness as a threshold value for the applicable range of the hazard evaluation method.

Brief Description of the Drawings

[0025]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Modes for Carrying Out the Invention

[0026] (Embodiment 1) The information processing system according to the present invention is an information processing system that predicts the correctness of an evaluation by an evaluation method for predicting positive / negative by a computer and shows the applicable range of the evaluation method. As shown in FIG. 1, it is composed of one or a plurality of information processing devices including a processing device 2, a storage means 3, an input means 4, and an information display unit 5. Specifically, it is a computer device centered on the processing device 2 and including a storage means 3, an input means 4 such as a pointing device, a keyboard, or a touch panel, an information display unit 5 such as a display, and other communication control units (not shown). Note that, in this embodiment, a system that defines the applicable range of the skin sensitization hazard evaluation method will be described as an example, but the present invention is not limited to this, and among evaluation methods for events caused by chemical substances such as other toxicities such as repeated-dose toxicity and metabolic reactivity, it can show the applicable range of an evaluation method for predicting positive / negative.

[0027] The processing device 2 is mainly composed of a CPU such as a microprocessor, has a storage unit composed of a RAM and a ROM (not shown), and stores a program and processing data that define the procedures of various processing operations.

[0028] The storage means 3 is composed of a memory, a hard disk, etc. inside and outside the information processing system 1. It may store the contents of part or all of the storage unit in the memory, hard disk, etc. of other computers communicatively connected to the information processing system 1. The storage means 3 has an objective variable storage unit 31 and an explanatory variable storage unit 32.

[0029] The objective variable storage unit 31 stores the information on the correctness / incorrectness of the evaluation by the evaluation method as an objective variable, and the explanatory variable storage unit 32 stores at least one of the information based on the experiments used in the evaluation method, the information on the chemical structure, and the information on substances similar to the test substance as an explanatory variable.

[0030] The data sources of the target variable and the explanatory variables obtain the data of 195 chemical substances having the information of the results of ITSv2 hazard assessment, the results of LLNA (Murine Local Lymph Node Assay), DPRA (Direct Peptide Reactivity Assay), KeratinoSens, and h-CLAT from existing literature. Specific literatures include the following (1) to (2), etc. Note that P / N means Positive / Negative, and Log means common logarithm.

[0031] (1) M. Hirota, et al., Development of an artificial neural network model for risk assessment of skin sensitization using human cell line activation test, direct peptide reactivity assay, KeratinoSens and in silico structure alert parameter. J. Appl. Toxicol., 38 (2018), pp. 514-526 (2) K. Ambe et al., Development of quantitative model of a local lymph node assay for evaluating skin sensitization potency applying machine learning CatBoost. Regulatory Toxicology and Pharmacology, 125 (2021) 105109

[0032] In a machine learning model that predicts the results of ITSv2 hazard assessment, information regarding whether it is correct (True) or incorrect (False), or information regarding true positives (TP) correctly evaluated as positive, true negatives (TN) correctly evaluated as negative, false negatives (FN) incorrectly evaluated as negative, and false positives (FP) incorrectly evaluated as positive is used as the target variable. However, the finally obtained predicted value is either correct (True) or incorrect (False).

[0033] Using Mordred (ver. 1.2.0), it is possible to calculate for all 195 substances and perform correct / incorrect prediction. At this time, Boruta (ver. 0.3), which is a method of variable selection using the variable importance of Random Forest, is used to narrow down the variables to be used. The descriptors selected here are defined as "molecular descriptor" and used as part of the explanatory variables.

[0034] Furthermore, in addition to Log(EC150(CD86)(μM)), Log(EC200(CD54)(μM)), Log(MIT(μM)), Log(CV75(μM)), Log(Adjusted Cys depletion), Log(Adjusted Lys depletion), Log(KEC1.5), h-CLAT P / N, CD86 P / N, CD54 P / N, DPRA P / N, KeratinoSens P / N which are data obtained from the literature, QSAR TB DASS AW P / N which is the result of implementing the QSAR Toolbox automated workflow ”Skin sensitisation for defined approaches” (QSAR TB DASS AW) using QSAR Toolbox ver.4.5 is defined as "in chemico / vitro / silico descriptor". Also, in the Euclidean space formed by the data normalized based on the learning data for "in chemico / vitro / silico descriptor" and "molecular descriptor", the shortest distances ("min_dist_TP", "min_dist_TN", "min_dist_FP", "min_dist_FN") between each data and the substances with true positive (TP), true negative (TN), false positive (FP), false negative (FN) in the ITSv2 hazard assessment in the learning data, and the ITSv2 hazard assessment results of the nearest neighbor substances ("Nearest T_0 / F_1", "Nearest TP_0 / TN_1 / FP_2 / FN_3") of six variables are defined as "distance-based descriptor". The above "in chemico / vitro / silico descriptor", "molecular descriptor", "distance-based descriptor" are used as explanatory variables for constructing the prediction model.

[0035] The processing device 2 includes a sorting processing unit 21, an objective variable generation processing unit 22, a prediction model construction processing unit 23, and an output processing unit 24.

[0036] The distribution processing unit 21 distributes and stores each piece of information on the target variable and the explanatory variables into learning data and verification data. When distributing the learning data and the verification data, it is preferable to perform a process of adjusting the scale, such as compressing to 0 to 1 by normalization / standardization. This enables data with different digits to be handled by the same calculation formula. Further, the distribution processing unit 21 further distributes the verification data into two types: internal verification data and external verification data. Specifically, for each piece of information on the target variable and the explanatory variables, after adjusting the scale (number of digits) by compressing through normalization / standardization, it is distributed and stored into the learning data and the verification data.

[0037] More specifically, based on the above-described data source, after arranging 195 substances in descending order with respect to EC3 (%), the ITSv2 hazard evaluation is arranged in the order of true positive (TP), true negative (TN), false positive (FP), and false negative (FN) of the ITSv2 hazard evaluation, and sequential numbers from SS001 to SS195 are assigned. Further, numbers from 1 to 4 are sequentially assigned in order from SS0001, and 1 and 4 are set as learning data, 2 as internal verification data, and 3 as external verification data. Thus, with respect to the sensitization intensity and the ITSv2 hazard evaluation result, they can be evenly allocated to each data set. Model construction is performed based on the learning data, and overfitting is avoided by adopting a model that fits well to the internal verification data, and the prediction accuracy is objectively evaluated using the external verification data.

[0038] When there is an imbalance exceeding a predetermined range in the information regarding whether the target variable is True or False, or in the sample sizes of the information on true positives (TP), true negatives (TN), false positives (FP), and false negatives (FN), the target variable generation processing unit 22 uses an oversampling method for the minority samples to detect neighboring data, internally generate data, and add this as minority samples. Examples of oversampling methods include SMOTE (Synthetic Minority Over-sampling Technique), Borderline SMOTE, ADASYN (Adaptive Synthetic sampling), and Safe-level SMOTE. In this embodiment, incorrect samples of the learning data are virtually increased by the oversampling method. Specifically, Borderline SMOTE (v0.10.1), which is a derivative of SMOTE and performs sampling considering the ratio of the majority class to the minority class in the neighborhood of the sample, is used. Also, for regression, SMOGN (v0.1.2), which is applicable to regression problems, is used to make the density of the data uniform.

[0039] The prediction model construction processing unit 23 constructs a prediction model that predicts the correctness of the evaluation using the target variable and the explanatory variables of the learning data by a machine learning method, evaluates the prediction model using the internal validation data, reconstructs the prediction model, and adopts a model that fits well to the internal validation data to avoid overfitting. The CatBoost model, which is a machine learning model, is a GBDT algorithm published in 2017 and is a method that can handle categorical variables well by a permutation-driven approach and is less prone to overfitting. In the present invention, since several categorical variables are handled as explanatory variables, it is preferable to use the CatBoost model.

[0040] All model construction uses Python for Windows v3.9.12 and CatBoostClassifier / Regressor for CatBoost (1.0.6). Parameter tuning is performed using the parameter tuning tool optuna (2.10.1) based on the Bayesian optimization algorithm, and tuning is performed for 'depth', 'learning_rate', 'bagging_temperature' and 'od_type'. Also, the predicted misclassification risk R M For 146 substances combining the training data and the internal validation data, the threshold for classifying positive (True) / negative (False) is determined by ROC analysis or any numerical value for the predicted misclassification risk R.

[0041] The classification is evaluated by sensitivity, specificity, accuracy, and balanced accuracy, and the evaluation of positive (True) / negative (False) classification is defined by the following formula (3).

[0042]

Number

[0043] Similarly, the evaluation of the classification of true positive (TP), true negative (TN), false positive (FP), and false negative (FN) is defined by the following formula (4).

[0044]

Number

[0045] The regression is evaluated by the following formula (5) R 2 value and RMSE, and is defined as in the following formula (5).

[0046]

Number

[0047] R 2 The value is preferably higher with 1 as the upper limit, and the RMSE is preferably lower with 0 as the lower limit.

[0048] The output processing unit 24 performs processing such as displaying and outputting the evaluation of the classification of positive (True) / negative (False) according to the evaluation method, and the evaluation of the classification of true positive (TP), true negative (TN), false positive (FP), and false negative (FN) on a display or the like, or transmitting the result to a communication-connected client computer.

[0049] As described above, the embodiments of the present invention have been described. However, the present invention is not limited to such examples. For example, instead of configuring the processing device by software processing by a computer, it is also preferable to configure part or all of it by a hardware processing circuit. In this case, an artificial intelligence processing circuit can also be used as the machine learning mechanism, and it goes without saying that the present invention can be implemented in various forms without departing from the gist of the present invention.

Example

[0050] (Variable selection) Using the aforementioned Mordred (ver. 1.2.0), 701 molecular descriptors that can be calculated for all 195 substances and whose variance is not zero were obtained. Using the correct / incorrect data of the ITSv2 hazard assessment as the target variable and all molecular descriptors as the explanatory variables, variable selection was performed using the aforementioned Boruta (ver. 0.3) only for the learning data. Twelve descriptors were selected and defined as "molecular descriptor". Table 1 is a table for the variables selected as "molecular descriptor".

[0051]

Table 1

[0052] Table 1 calculates the "distance-based descriptor" based on the "in chemico / vitro / silico descriptor" and the "molecular descriptor".

[0053] In addition, among the "in chemico / vitro / silico descriptor", "molecular descriptor", and "distance-based descriptor", variables with continuous values were normalized and principal component analysis (PCA) was performed to confirm the chemical space of each dataset. The results are shown in Figure 2. As shown in Figure 2, the distributions of 49 substances each for internal and external validation did not differ significantly from the distribution of 97 substances for training.

[0054] (Qualitative ITSv2 Hazard Assessment Correct / Incorrect Prediction Model) A classification model was constructed with the correct (True) / incorrect (False) of the ITSv2 hazard assessment as the target variable and the "in chemico / vitro / silico descriptor", "molecular descriptor", and "distance-based descriptor" as explanatory variables. In addition, the False samples of the training data were oversampled by Borderline SMOTE, changing the ratio of the number of True substances / number of False substances from 77 / 20 to 77 / 77. The CatBoost model without data imbalance correction was named Model C1, and the corrected model was named Model C2. However, the data for evaluation of Model C2 does not include the data added by oversampling and is the same as Model C1. Furthermore, considering that TP and TN, FN and FP are separate events, a multi-class classification model with the target variable as TP / TN / FP / FN was constructed. Those without oversampling of the training data were constructed as Model C3, and those with oversampling were constructed as Model C4. However, the prediction accuracy was evaluated as a binary classification with true positive (TP) and true negative (TN) as correct (True), and false positive (FP) and false negative (FN) as incorrect (False). The evaluation results are shown in Table 2.

[0055]

Table 2

[0056] When comparing Model C1 with C2, and C3 with C4 respectively, it was found that Model C1 and C3 had higher sensitivity, while Model C2 and C4 had higher specificity. This difference is considered to be caused by the virtual increase in false data by SMOTE, which made it easier to detect false data. Also, overall, the specificity and balanced accuracy (hereinafter referred to as "BA") of Model C4 were the highest in all datasets. By using oversampling by SMOTE and dealing with more detailed target variables, a model with relatively good prediction accuracy could be obtained. However, the specificity of the internal and external validation data was below 0.5 for all prediction models, and it was difficult to detect false data even after performing SMOTE.

[0057] Also, Table 3 shows the results of calculating the variable importance in Model C1 - C4.

[0058]

Table 3

[0059] In all models, the variable importance of ATSC2i (centered moreau - broto autocorrelation of lag 2 weighted by ionization potential), which is one of the "molecular descriptors", was within the top 10 variables. This is a variable that represents the distribution of ionization energies of atoms in a chemical structure based on graph theory. Therefore, the arrangement of characteristic substituents such as fluorine atoms and nitrogen atoms may contribute to the classification of correct and incorrect results by ITSv2. However, it is difficult to identify the causative chemical structure based on just this one variable. Also, when comparing Model C1 with C2, and C3 with C4, in Models C1 and C3 where SMOTE was not performed, there was a tendency for the variable importance of "distance - based descriptor" to be relatively high. Since SMOTE performs oversampling by virtually creating analogous substances that do not actually exist, it was shown that there may be a poor compatibility with the "distance - based descriptor", which is an explanatory variable based on the concept of lead - across that utilizes the properties of similar substances.

[0060] (Embodiment 2) Next, an information processing system that predicts the risk of evaluation errors by an evaluation method for performing positive / negative prediction and shows the applicable range of the evaluation method, a control program for causing a computer to function as the information processing system, and an information processing method executed by a system that shows the applicable range of the evaluation method will be described.

[0061] (Regarding boundary - type risk and outlier - type risk) As risks of classification (evaluation) errors in the evaluation method for positive / negative prediction, there are cases where a negative substance is included within the positive range or vice versa. As situations where misclassified substances are placed, there are cases where they exist near the threshold within the correctly classifiable range and cases where they deviate from the overall trend due to excessive properties. The former is, for example, the case where, regarding the toxicity that is likely to occur in substances with a molecular weight of 300 or more, the molecular weight of the substance to be evaluated is exactly around 300. The latter is, for example, the case where, regarding the toxicity that is likely to occur in substances with a molecular weight of 300 or more, the molecular weight of the substance to be evaluated is quite high, such as 1200, and no toxicity occurs. The former is defined as "borderline-type misclassification", and the latter is defined as "outlier-type misclassification".

[0062] The so-called borderline-type misclassification means that, as shown in Fig. 3(a), when data is plotted one-dimensionally, it is assumed to be arranged in the order of true negative (TN), false negative (FN), false positive (FP), true positive (TP), or vice versa. This is because it is considering the case where misjudgments exist near the judgment threshold. On the other hand, the outlier-type misclassification means that, as shown in Fig. 3(b), it is assumed to be arranged in the order of false negative (FN), true negative (TN), true positive (TP), false positive (FP), or vice versa. In addition, "○" in Fig. 3(a) and Fig. 3(b) indicates positive (TP + FN), and "●" indicates negative (TF + FP), respectively.

[0063] Regarding the borderline-type distribution as Y B give TN: 0, FN: 0.333, FP: 0.667, TP: 1 respectively. Similarly, regarding the outlier-type distribution as Y O give TN: 0, FN: 0.333, FP: 0.667, TP: 1 respectively. Predict these as target variables, and calculate the value of the risk of becoming a borderline-type error (hereinafter referred to as "borderline risk") R B and the value of the risk of becoming an outlier-type error (hereinafter referred to as "outlier risk") R O The borderline risk R B and the outlier risk RO Let it be the following formula (1).

[0064]

Number

[0065] Also, R B and R O are combined to obtain a value that becomes the risk of evaluation error according to the evaluation method (hereinafter referred to as "misclassification risk") R M Let it be the following formula (2).

[0066]

Number

[0067] Assuming the relationship between the reliability C that can be correctly determined and the misjudgment risk R is C = 1 - R, the harmonic mean of the boundary type reliability C B and the outlier type reliability C O is taken, so that if either one of R B and R O is high, the misclassification risk R M is also calculated to be high.

[0068] The information processing system according to the present invention is an information processing system that predicts the risk of evaluation error according to an evaluation method for performing positive / negative prediction by a computer and shows the application range of the evaluation method. As shown in FIG. 4, it is composed of one or a plurality of information processing devices including a processing device 2, a storage means 3, an information display unit 5, and an input means 4.

[0069] Since the processing device 2, the storage means 3, the input means 4, and the information display unit 5 have already been described based on FIG. 1 of "Embodiment 1", the description thereof will be omitted.

[0070] The objective variable storage unit 31 stores, as an objective variable, information on risks that would result in errors in evaluation by the evaluation method, and the explanatory variable storage unit 32 stores, as an explanatory variable, at least one piece of information among information based on experiments used in the evaluation method, information on chemical structures, and information on substances similar to the test substance.

[0071] The data sources of the objective variable and the explanatory variable are, as described above, information on the results of the ITSv2 hazard assessment, data on 195 chemical substances having the results of LLNA, DPRA, KeratinoSens, and h-CLAT, etc., and are obtained from the existing literature described above. Also, in the machine learning model for predicting the results of the ITSv2 hazard assessment, as described above, information on whether it is correct (True) or incorrect (False), or information on true positive (TP), true negative (TN), false negative (FN), and false positive (FP) is used as the objective variable. However, the finally obtained predicted value is either correct (True) or incorrect (False).

[0072] Similar to Embodiment 1, in addition to Log(EC150(CD86)(μM)), Log(EC200(CD54)(μM)), Log(MIT(μM)), Log(CV75(μM)), Log(Adjusted Cys depletion), Log(Adjusted Lys depletion), Log(KEC1.5), h-CLAT P / N, CD86 P / N, CD54 P / N, DPRA P / N, KeratinoSens P / N, which are data obtained from the literature, QSAR TB DASS AW P / N, which is the result of implementing the QSAR Toolbox automated workflow "Skin sensitisation for defined approaches" (QSAR TB DASS AW) using QSAR Toolbox ver.4.5, is defined as "in chemico / vitro / silico descriptor". Also, in the Euclidean space formed by the data obtained by normalizing "in chemico / vitro / silico descriptor" and "molecular descriptor" based on the learning data, the shortest distances ("min_dist_TP", "min_dist_TN", "min_dist_FP", "min_dist_FN") between each data and the substances with true positive (TP), true negative (TN), false positive (FP), and false negative (FN) in the ITSv2 hazard assessment in the learning data, and the ITSv2 hazard assessment results of the nearest substances ("Nearest T_0 / F_1", "Nearest TP_0 / TN_1 / FP_2 / FN_3"), these six variables are defined as "distance-based descriptor". The above "in chemico / vitro / silico descriptor", "molecular descriptor", and "distance-based descriptor" are used as explanatory variables for constructing the prediction model.

[0073] The processing device 2 includes a first sorting processing unit 25, an objective variable generation processing unit 26, a prediction model construction processing unit 27, a second sorting processing unit 28, and an output processing unit 29.

[0074] The first distribution processing unit 25 is the same as the aforementioned distribution processing unit 21, and thus the description thereof will be omitted.

[0075] The objective variable generation processing unit 26 is the same as the aforementioned objective variable generation processing unit 22, and thus the description thereof will be omitted.

[0076] The prediction model construction processing unit 27 constructs a prediction model that predicts the risk of the error in the evaluation using the objective variable and the explanatory variable of the learning data by a machine learning method, evaluates the prediction model using the internal verification data, reconstructs the prediction model, and adopts a model that fits well with the internal verification data to avoid overfitting.

[0077] Furthermore, when the risk information as the objective variable includes the distribution information of the true positive, the true negative, the false positive, and the false negative assumed in the classification of the boundary type error where the error exists near the threshold within the range where the error can be correctly classified, and the distribution information of the true positive, the true negative, the false positive, and the false negative assumed in the classification of the outlier type error that deviates from the overall trend due to excessive properties, the prediction model construction processing unit 27 constructs a prediction model that predicts the risk of the boundary type error and the risk of the outlier type error, evaluates the prediction model using the internal verification data, reconstructs the prediction model, and adopts a model that fits well with the internal verification data.

[0078] Similar to the aforementioned output processing unit 24, the output processing unit 29 performs processes such as displaying and outputting the evaluation of the positive (True) / negative (False) classification according to the evaluation method, and the evaluation of the classification of the true positive (TP), true negative (TN), false positive (FP), and false negative on a display or the like, or transmitting the result to a communication-connected client computer.

[0079] Also, when the prediction model construction processing unit 27 constructs a prediction model that predicts the risk of the boundary type error and the risk of the outlier type error, the output processing unit 29 determines that the value (R) of the risk of the boundary type error is such that the evaluation according to the evaluation method B) and the value of the risk (R) of becoming an outlier-type error O ) is output.

Example

[0080] (Variable selection) Variable selection was performed in the same manner as the "variable selection" in the above-mentioned "Example 1".

[0081] (Quantitative misclassification risk prediction model) As situations where substances misclassified by binary classification are placed, there are cases where they exist near the decision threshold as shown in Fig. 3(a) and cases where they deviate from the overall trend due to excessive properties as shown in Fig. 3(b). Based on this concept, a boundary-type distribution (Y B ) and an outlier-type distribution (Y O ) were set, and Table 4 shows the results of quantitatively predicting them. The target variable was the boundary-type distribution (Y B ) or the outlier-type distribution (Y O ), and the explanatory variables were set as "in chemico / vitro / silico descriptor", "molecular descriptor", and "distance-based descriptor" in the same way as the qualitative prediction of the correctness of the ITSv2 hazard assessment. Model parameter tuning was performed using optuna. Also, the density of the distribution of the target variable in the learning data was corrected by SMOGN. For the CatBoost model that predicts Y B , the one without applying SMOGN was designated as Model R1, and the one with applying SMOGN was designated as Model R2. Similarly, for the CatBoost model that predicts Y O , the one without applying SMOGN was designated as Model R3, and the one with applying SMOGN was designated as Model R4.

[0082]

Table 4

[0083] As shown in Table 4, when comparing Model R1 with R2 and R3 with R4 respectively, overall, the prediction accuracies of Model R1 and R3 exceeded those of Model R2 and R4. Different from the classification model, in this regression model under this assumption, oversampling by SMOTE may not be effective. Therefore, Model R1 without implementing SMOGN was used as the model for predicting the boundary type distribution Y B , and Model R3 was used as the model for predicting the outlier type distribution Y O .

[0084] Table 5 shows the results of calculating the variable importance from Model R1 to R4.

[0085]

Table 5

[0086] When comparing Model R1 and R2 in Table 5, in Model R2 with SMOTE performed, the variable importance of "distance-based descriptor" tended to be low. It was confirmed that there was a tendency for poor compatibility between the virtual oversampling by SMOTE and the explanatory variable "distance-based descriptor" based on the concept of lead across using the properties of similar substances. Also, in all models, the variable importances of DPRA P / N, Log(MIT(μM)), and Log(EC150 (CD86)(μM)) were within the top 10 variables. In particular, the variable importances of DPRA P / N and Log(MIT(μM)) were stably high, suggesting their contribution to the misclassification risk. Also, the prediction accuracies of Model R3 and R4, which are the prediction models for the outlier type distribution Y O , were clearly low. This is considered to be due to the fact that outlier data has already been removed according to the applicable ranges of each test method, which is the existing applicable range of ITSv2.

[0087] (Classification based on misclassification risk R M ) The boundary type distribution Y predicted by Model R1 and R3 B, Outlier-type distribution Y O to boundary-type risk R B , outlier-type risk R O was calculated, and misclassification risk R M was calculated. For R M which is a continuous value, a threshold was set, and the results of classifying and evaluating ITSv2 hazard assessment as positive (True) / negative (False) are shown in Table 6. Figure 5 is a graph explaining the optimal threshold based on the Youden Index for 146 substances combining learning data and internal validation data. Also, any threshold can be set for misclassification risk R M . A threshold of 0.25 was set as a strict threshold for misclassification detection, and a threshold of 0.5 was set as a loose threshold and evaluated.

[0088]

Table 6

[0089] In the external validation data, in Models C1 - C4, the specificity did not exceed 0.5. However, when classified by the threshold (0.385977) calculated based on the Youden Index for misclassification risk R M , the specificity was 0.545. The specificity exceeded 0.5 in all datasets, and the problem of the classification model that False substances are extremely difficult to detect in the classification model was overcome. Also, when a threshold of 0.25 was set as an arbitrary threshold, the specificity was high, and when a threshold of 0.5 was set, the sensitivity was high. Thus, the threshold can be adjusted according to the operation purpose and used as an evaluation method with high sensitivity or specificity. From the above, it was shown that the classification using misclassification risk R M calculated based on the assumed distribution of misclassified substances is effective as a method for classifying the correctness of ITSv2 hazard assessment which is non-uniform data.

[0090] (Exclusion of substances with high misclassification risk) Table 7 shows the results of evaluating how much the prediction accuracy of the ITSv2 hazard assessment is improved using the above qualitative and quantitative prediction models for misclassified substances.

[0091] [Table 7]

[0092] Substances predicted to be positive (True) for the ITSv2 hazard assessment by previous prediction methods can be correctly evaluated by the ITSv2 hazard assessment and are within the applicable range. Conversely, substances predicted to be false for the ITSv2 hazard assessment are outside the applicable range of the ITSv2 hazard assessment. By setting a new applicable range using Model C4 or misclassification risk R M , the prediction accuracy of the ITSv2 hazard assessment was clearly improved compared to "Original". However, it was difficult to clearly distinguish the advantages and disadvantages of the applicable ranges set by Model C4 and misclassification risk R M . In the prediction by Model C4, relatively few substances were excluded, and more substances could be accurately evaluated by the ITSv2 hazard assessment. On the other hand, since the classification based on misclassification risk R M is a quantitative evaluation, it has the advantage that an arbitrary threshold can be set to determine the substance to be evaluated.

[0093] As an example of making use of these advantages, a method of combining a qualitative prediction model and a quantitative prediction model can be considered. Substances predicted to be within the applicable range by Model C4, which predicts a relatively wide applicable range, and with a relatively loose misclassification risk R M < 0.5 were evaluated as being within the applicable range. In this case, for the 36 substances of the internal validation data, the coincidence rate was 0.917 and the BA was 0.857, and for the 33 substances of the external validation data, the coincidence rate was 0.879 and the BA was 0.819. Thus, misclassification risk R MBy adjusting the threshold value, a more effective applicable range can be set without overly narrowing the applicable range.

[0094] (Discussion) The present invention considered that a model for predicting the correctness of a certain evaluation method should function as a threshold value for the applicable range of that evaluation method. When the evaluation result is predicted to be correct, the sample is considered to be within the applicable range of that evaluation method. That is, if the correctness of the evaluation result can be defined, this method can determine the applicable range for any evaluation method regardless of whether it is qualitative or quantitative. This time, an example of the ITSv2 hazard evaluation, which is a qualitative evaluation method, was shown. Also in the case of a quantitative evaluation method, it is possible to define the correctness by setting a threshold value for the prediction error. It is also possible to predict the prediction error itself as a quantitative target variable and set a threshold value for the predicted value to determine the applicable range.

[0095] Also, by setting the boundary type distribution Y B and the outlier type distribution Y O as target variables, it was realized to handle the result of qualitative evaluation as a quantitative variable. By obtaining the quantitative misclassification risk R M , it became possible to set an arbitrary threshold value according to the purpose, and the degree of freedom of true / false classification was greatly improved.

[0096] As the situation where the misclassified substance is placed, the distribution when it exists near the threshold value of the range where correct classification is possible (boundary type distribution Y B ) and the distribution when it deviates from the overall trend due to excessive properties (outlier type distribution Y O) was assumed. For example, when considering a boundary type distribution, when plotting data in one dimension, it is assumed that they are arranged in the order of TN, FN, FP, TP. This distribution does not matter even in the reverse order when treating it as the target variable of prediction. By substituting 0, 0.333, 0.667, and 1 into TN, FN, FP, and TP respectively, the result of qualitative evaluation can be predicted quantitatively. However, since these numerical values were set experimentally so as to be evenly distributed between 0 and 1, there is room for consideration especially for 0.333 and 0.667. Y B and Y O from the boundary type risk R B and the outlier type risk R O were calculated respectively, and integrated into the misclassification risk R M . This R M is defined as the harmonic mean of 1 - R B and 1 - R O so as to show a high value if either one of R B and 1 - R O is high. For R M too, since the correct value can be defined such as 1 if True and 0 if False, there is room for optimization such as constructing a multiple regression model to calculate R B from the predicted R O and R M . However, since using a complex model may impair readability and robustness, careful consideration is required.

[0097] In the classification model, the prediction accuracy of Model C4 for which the target variable was subdivided and oversampling was performed was the highest. The coincidence rate (Accuracy) of the prediction for the external verification data was 0.755. However, the specificity was 0.455, and it was difficult to detect misclassified data even when oversampling was performed. In the prediction using the quantitative misclassification risk R M obtained by the regression model, by optimizing the threshold by ROC analysis, the coincidence rate (Accuracy) for the external verification data was 0.842 and the BA was 0.796. R MBy setting an arbitrary threshold for [X], it is also possible to obtain a model that exhibits high specificity. Also, Y B Compared with Y O the prediction accuracy was clearly low. This is presumably due to the fact that outlier data has already been excluded by the existing application range of ITSv2, such as substances with LogP > 3.5 being outside the applicable range of h-CLAT.

[0098] The classification model and R M By newly defining the application range based on [X], the accuracy and BA of the ITSv2 hazard assessment were clearly improved. There is no clear superiority or inferiority between the method using the classification model and R M Each has its own advantages, as the classification model can accurately evaluate more substances as being within the applicable range, and the method based on R M allows for the determination of an arbitrary threshold.

Industrial Applicability

[0099] An in silico model was constructed to define the application range of an evaluation method by predicting the correctness of the evaluation method using machine learning techniques. Although this study focused on the ITSv2 hazard assessment, the approach of defining the application range by predicting correctness and quantitatively predicting the results of qualitative evaluations is applicable to any qualitative / quantitative evaluation method. We are confident that the results of this study will contribute to the advancement of in silico toxicity prediction evaluation research.

Explanation of Symbols

[0100] 1 Information processing system 2 Processing device 3 Storage means 4 Input means 5 Information display unit 21 Sorting processing unit 22 Objective variable generation processing unit 23 Prediction model construction processing unit 24 Output processing unit 25 First sorting processing unit 26 Objective variable generation processing unit 27 Prediction model construction processing unit 28 Second sorting processing unit 29 Output processing unit 31 Objective variable storage unit 32 Explanatory variable storage unit

Claims

1. An information processing system that predicts the correctness of an evaluation by an evaluation method for performing positive / negative prediction and shows the application range of the evaluation method, comprising: An objective variable storage unit that stores information on the correctness / incorrectness of the evaluation by the evaluation method as an objective variable; An explanatory variable storage unit that stores at least one of information based on an experiment used in the evaluation method, information on a chemical structure, and information on a substance similar to the test substance as an explanatory variable; A distribution processing unit that distributes and stores each piece of information on the objective variable and the explanatory variable into learning data and verification data; A prediction model construction processing unit that constructs a prediction model for predicting the correctness of the evaluation using the objective variable and the explanatory variable of the learning data by a machine learning method, evaluates the prediction model using the objective variable and the explanatory variable of the verification data, and performs necessary reconstruction of the prediction model; An information processing system comprising an output processing unit that outputs the correctness / incorrectness of the evaluation by the evaluation method using the prediction model.

2. When there is an imbalance exceeding a predetermined range in the sample size of the information on the correctness / incorrectness as the objective variable, or true positives correctly evaluated as positive, true negatives correctly evaluated as negative, false negatives incorrectly evaluated as negative, and false positives incorrectly evaluated as positive, an objective variable generation processing unit that uses an oversampling method for minority samples to detect neighboring data, internally generate data, and add this as minority samples; The information processing system according to Claim 1.

3. The correctness / incorrectness of the evaluation as the objective variable is information on the true positive, the true negative, the false positive, and the false negative; The output processing unit outputs which of the true positive, true negative, false positive, and false negative of the evaluation by the evaluation method using the prediction model; The information processing system according to Claim 2.

4. An information processing system that predicts the risk of evaluation errors by an evaluation method for performing positive / negative prediction and shows the applicable range of the evaluation method, a target variable storage unit that stores information on the risk of evaluation errors by the evaluation method as a target variable, an explanatory variable storage unit that stores at least one of information based on experiments used in the evaluation method, information on chemical structures, and information on substances similar to the test substance as explanatory variables, a first sorting processing unit that sorts and stores each piece of information on the target variable and the explanatory variable into learning data and verification data, a prediction model construction processing unit that constructs a prediction model for predicting the risk of evaluation errors using the target variable and the explanatory variable of the learning data by a machine learning method, evaluates the prediction model using the target variable and the explanatory variable of the verification data, and performs necessary reconstruction of the prediction model, an information processing system comprising an output processing unit that outputs a value of the risk of evaluation errors by the evaluation method using the prediction model.

5. further comprising a second sorting processing unit that sorts the correctness of the evaluation by the evaluation method into any one of true positives correctly evaluated as positive being positive, true negatives correctly evaluated as negative being negative, false negatives incorrectly evaluated as positive being negative, and false positives incorrectly evaluated as negative being positive, the risk information as the target variable includes distribution information of the true positives, true negatives, false positives, and false negatives assumed in the classification of boundary-type errors existing near the threshold within the range where the errors can be correctly classified, and distribution information of the true positives, true negatives, false positives, and false negatives assumed in the classification of outlier-type errors deviating from the overall trend due to excessive properties, the prediction model construction processing unit constructs a prediction model for predicting the risk of boundary-type errors and the risk of outlier-type errors using the target variable and the explanatory variable of the learning data by a machine learning method, The output processing unit uses the prediction model to output the value of the risk that the evaluation by the evaluation method becomes the boundary type error and the value of the risk that becomes the outlier type error. The information processing system according to claim 4.

6. The prediction model predicts the distribution (Y B ) of the boundary type and the distribution (Y O ) of the outlier type. The output processing unit uses the following formula (1) to output the value of the risk (R B ) that becomes the boundary type error and the value of the risk (R O ) that becomes the outlier type error. The information processing system according to claim 5. 【Equation 1】

7. The output processing unit integrates the value of the risk (R B ) that becomes the boundary type error and the value of the risk (R O ) that becomes the outlier type error according to the following formula (2) to calculate the value of the risk (R M ) that becomes the error in the evaluation by the evaluation method. The information processing system according to claim 6. 【Equation 2】

8. The output processing unit Outputs the predicted correct / error of the evaluation by the evaluation method based on whether or not the value of the risk that becomes the error in the evaluation by the evaluation method exceeds a predetermined threshold. The information processing system according to claim 4.

9. A control program for causing a computer to function as the information processing system according to claim 1, A control program for causing a computer to function as the above sorting processing unit, the above prediction model construction processing unit, and the above output processing unit.

10. A control program for causing a computer to function as the information processing system according to claim 4, A control program for causing a computer to function as the above-described first distribution processing unit, the above-described prediction model construction processing unit, and the above-described output processing unit.

11. An information processing method executed by a system that has an objective variable storage unit that stores information on the correctness / incorrectness of evaluation by an evaluation method for predicting positive / negative as an objective variable, and an explanatory variable storage unit that stores at least one of information based on an experiment used in the evaluation method, information on a chemical structure, and information on a substance similar to a test substance as an explanatory variable, and predicts the correctness / incorrectness of evaluation by an evaluation method for predicting positive / negative and indicates the applicable range of the evaluation method, the method comprising: A distribution procedure for distributing and storing each piece of information on the objective variable and the explanatory variable into learning data and verification data; A prediction model construction procedure for constructing a prediction model that predicts the correctness / incorrectness of the evaluation using the objective variable and the explanatory variable of the learning data by a machine learning technique, evaluating the prediction model using the objective variable and the explanatory variable of the verification data, and performing necessary reconstruction of the prediction model; An output procedure for outputting the correctness / incorrectness of the evaluation by the evaluation method using the prediction model.

12. An information processing method executed by a system that has an objective variable storage unit that stores information on the correctness / incorrectness of evaluation by an evaluation method for predicting positive / negative as an objective variable, and an explanatory variable storage unit that stores at least one of information based on an experiment used in the evaluation method, information on a chemical structure, and information on a substance similar to a test substance as an explanatory variable, and predicts the risk of error in evaluation by an evaluation method for predicting positive / negative and indicates the applicable range of the evaluation method, the method comprising: A distribution procedure for distributing and storing each piece of information on the objective variable and the explanatory variable into learning data and verification data; Using a machine learning method, a prediction model is constructed to predict the risk of evaluation error using the target variable and the explanatory variables of the learning data. The prediction model is evaluated using the target variable and the explanatory variables of the verification data, and necessary reconstruction of the prediction model is performed. This is a prediction model construction procedure, An information processing method including an output procedure for outputting a risk value at which the evaluation by the evaluation method becomes an error using the prediction model.