Method for predicting human health toxicity of small molecule compounds based on machine learning

By building a QSAR model based on machine learning, combining multiple algorithms and molecular description methods, the problems of insufficient data sources and insufficient model generalization capabilities in the existing technology are solved, and accurate prediction and rapid evaluation of the human health toxicity of small molecule compounds are achieved.

CN120299554APending Publication Date: 2025-07-11NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510372993.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

When the existing QSAR model predicts the human health toxicity of small molecule compounds, the data source is small and the data collation principles are unclear, and the modeling algorithm is relatively single, resulting in poor prediction performance of the model for applied extraterritorial compounds.

Method used

Build a QSAR model based on machine learning, and ensure the broadness and reliability of the data by obtaining data from international toxicity databases and literature, combining 15 machine learning algorithms and three molecular description methods.

Benefits of technology

It achieves accurate prediction of the toxicity of a variety of specific human health of small molecule compounds, has the ability to quickly evaluate new chemicals, and is in line with the OECD's QSAR modeling principle, and can be widely used worldwide.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299554A_ABST
    Figure CN120299554A_ABST
Patent Text Reader

Abstract

The invention discloses a method for predicting human health toxicity of small molecule compounds based on machine learning. The method comprises the following steps: acquiring a data set from a wide source; sorting the data set; calculating molecular descriptors and molecular fingerprints of the small molecular compounds; the molecular descriptors, the molecular fingerprints and the combination of the molecular descriptors and the molecular fingerprints are used as input, N machine learning algorithms are used for constructing QSAR models, and 3N QSAR models are constructed for each modeling end point; the 3N QSAR models are trained; an optimal QSAR model set is formed; and inputting a to-be-detected small molecule compound into the optimal QSAR model set for prediction so as to obtain the human health toxicity of the to-be-detected small molecule compound. According to the method, the universality of data sources is ensured, the reliability of the data is ensured through a directional arrangement rule, the QSAR model is constructed by three molecular description modes and 15 machine learning algorithms, and various specific human health toxicities of small molecular compounds can be accurately predicted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of toxicity prediction, and more specifically, to a method for predicting the human health toxicity of small molecule compounds based on machine learning. Background Art

[0002] Small molecule compounds generally refer to chemical substances with relatively small molecular weights, and their molecular weights are generally less than 900 Daltons (Da). In the fields of medicinal chemistry, materials science, biochemistry, etc., the term "small molecule" is often used to describe organic or inorganic compounds that can interact with biological macromolecules (such as proteins, nucleic acids, etc.) and may affect the functions of these biological macromolecules.

[0003] With the rapid development of technology and industry, the types and production of chemicals in the world have increased rapidly, becoming an important risk source threatening human health. Chemicals may be released into the environment through production, storage, sales, application, waste disposal, etc., thus directly or indirectly exposing to the human body and causing a series of adverse effects on the human body and animal body, such as reproductive disorders, endocrine disorders, skin necrosis, and tissue canceration, which are highly concerned by management departments in various countries. To prevent the health threats of chemicals to human exposure, systematic toxicity tests are urgently needed to obtain sufficient toxicity data for hazard identification and to provide necessary information for the risk assessment of chemicals.

[0004] However, in vivo toxicity tests have problems such as long screening time, violation of animal ethics, and high experimental costs, while in vitro toxicity tests also require high costs and cannot provide experimental tests one by one for the hazard identification of tens of thousands of existing chemicals and new pollutants. Therefore, at present, there is a gradual shift from the toxicity test paradigm based on in vitro and in vivo to computer-based virtual screening.

[0005] Currently, some international legislations, institutions, and organizations support, recognize, and recommend the use of the Quantitative Structure-Activity Relationship (QSAR) method for toxicity prediction under the regulatory framework.

[0006] QSAR, as a mature method, has been widely used in various toxicity predictions. For example, in the patent application for invention titled "A Method for Screening Hepatotoxic Chemicals Based on Random Forest Algorithm" with publication number CN116312858A, the described method collects hepatotoxicity endpoint data of chemicals from literature and databases (DILIrank, ChemIDplus), then based on the SMILES codes of the chemicals, uses PaDEL-Descriptor 2.21 software to calculate Pubchem fingerprints and perform feature screening, adopts the random forest algorithm to establish a QSAR model for screening hepatotoxic chemicals, and uses molecular similarity to characterize the application domain. However, the data used in this method is relatively single, making it difficult to include more chemical data, which may lead to insufficient generalization ability of the model, unable to cover a wide chemical space, and the data is not screened. In addition, Pubchem fingerprint is a fixed fingerprint representation method, and the features it contains may not be fully applicable to the QSAR model for hepatotoxicity screening. Moreover, when the random forest algorithm processes complex chemical data, overfitting may occur, especially when the number of samples in the training dataset is limited and the dimension of chemical structure features (such as Pubchem fingerprint) is high, the model may overlearn the noise and local features in the training data, resulting in poor performance on new and unseen data. Finally, using molecular similarity to define the application domain depends on the quality and diversity of the selected reference compound set. If the reference set is not extensive or appropriate, it may misevaluate whether a new compound is within the application domain of the model, thus affecting the prediction reliability.

[0007] For another example, in the patent application for invention titled "A Method and Prediction Model for Predicting the Carcinogenicity of PAHs Based on 2D Molecular Descriptors" with publication number CN114530213A, the described method collects carcinogenicity data from the Carcinogenic Potency Database (CPDB), and based on 2D molecular descriptors, uses ordinary least squares (OLS) combined with multiple linear regression (MLR) to establish a QSAR model to predict the carcinogenicity of polycyclic aromatic hydrocarbons (PAHs) to rodents. However, the data coverage of the described method is also limited, and OLS assumes that the error terms are independently and identically distributed normal distributions, and the variance of the error terms is constant, but in actual carcinogenicity data, these assumptions may not hold. In addition, MLR assumes a linear relationship between the independent variables (2D molecular descriptors) and the dependent variable (carcinogenicity), however, the relationship between the carcinogenicity of PAHs and their molecular structure may be non-linear. Summary of the Invention

[0008] The present invention overcomes the deficiencies in the prior art, including fewer data sources, unclear data sorting principles, relatively single algorithms selected for modeling, relatively single dimensions of model inputs, and poor prediction performance for compounds outside the application domain, and constructs an optimal QSAR model to accurately predict the human health toxicity of small molecule compounds.

[0009] The present invention discloses a method for predicting the human health toxicity of small molecule compounds based on machine learning, including: Step S1, obtaining a data set containing the structures of small molecule compounds and their corresponding toxicity data from a toxicity database and public literature; Step S2, sorting the obtained data set to obtain a sorted data set; Step S3, calculating the molecular descriptors and molecular fingerprints of the small molecule compounds in the sorted data set; Step S4, using N machine learning algorithms to construct QSAR models with molecular descriptors, molecular fingerprints, and combinations of molecular descriptors and molecular fingerprints as inputs respectively, so as to construct 3N initial QSAR models for each modeling endpoint; Step S5, dividing the sorted data set into a training data set and a test data set at a preset ratio, using the molecular descriptors, molecular fingerprints, and combinations of molecular descriptors and molecular fingerprints of the small molecule compounds in the training data set as inputs, and the corresponding toxicity data of the small molecule compounds in the training data set as outputs, training the 3N initial QSAR models, and obtaining 3N target QSAR models after training; Step S6, using the test data set to evaluate the performance of the 3N target QSAR models based on multiple evaluation indicators, and selecting the optimal QSAR model for each modeling endpoint from them to form an optimal QSAR model set; and Step S7, inputting the molecular descriptors and molecular fingerprints corresponding to the small molecule compound to be tested into the optimal QSAR model set for prediction, so as to obtain the human health toxicity of the small molecule compound to be tested.

[0010] In one embodiment of the present invention, the sorting of the obtained data set in Step S2 includes one or more of the following operations: removing one or more of mixtures, inorganic substances, organometallic compounds, polymer types, and counterions; standardizing specific chemical types; neutralizing salt compounds; removing data with conflicting activity results; processing chemical structures using a standardized workflow; integrating the data in the database; removing missing values, undefined / non-standard endpoints, and data with unclear units; classifying chemical substances; merging the screened data sets; deleting duplicate data; and deleting duplicate chemical substances; performing special processing on chemical substances with multiple duplicate records: (1) when the duplicate results are the same, only retaining one; (2) when the duplicate results are different, following the principle of the minority obeying the majority; (3) when the duplicate results are different and the numbers on both sides are the same, deleting all records.

[0011] In one embodiment of the present invention, the machine learning algorithms include: Random Forest (RF), Support Vector Machine (SVM), Decision Tree (DT), Naive Bayes (NB), and K-Nearest Neighbor (KNN); among them, the Support Vector Machine (SVM) includes five SVM algorithms constructed based on Linear Kernel, Polynomial Kernel, Radial Basis Function (RBF), non-linear activation function kernel function of neurons (Sigmoid tanh), and custom kernel function (precomputed), namely: SVM-linear, SVM-poly, SVM-RBF, SVM-sigmoid, SVM-precomputed; among them, Naive Bayes (NB) includes four types, namely: Gaussian Naive Bayes (GaussianNB), Bernoulli Naive Bayes (BernoulliNB), Multinomial Naive Bayes (MultinomialNB), and Complement Naive Bayes (ComplementNB); among them, K-Nearest Neighbor (KNN) includes four KNN algorithms constructed based on "ball_tree", "kd_tree", "brute", and "auto", namely: kNN-algorithm='auto', kNN-algorithm=ball_tree, kNN-algorithm=kd_tree, and kNN-algorithm='brute'.

[0012] In one embodiment of the present invention, calculating the molecular descriptors and molecular fingerprints of the small molecule compound in step S3 includes preprocessing the molecular descriptors and molecular fingerprints of the small molecule compound, including one or more of the following operations: deleting features containing missing values and over-large data; filtering out features with a variance of 0; performing correlation filtering using the mutual information method to avoid overfitting; performing normalization using the following z-score normalization method: where x is the feature value, μ is the mean of each feature, σ is the standard deviation of each feature, and wherein the molecular descriptors include 1D molecular descriptors and 2D molecular descriptors; the molecular fingerprints include the following 12 types of molecular fingerprints: FP, ExtFP, EStateFP, GraphFP, MACCSFP, PubchemFP, SubFP, SubFPC, KRFP, KRFPC, AP2D, and APC2D.

[0013] In one embodiment of the present invention, step S4 includes selecting the hyperparameters of the machine learning algorithm using a grid search method.

[0014] In one embodiment of the present invention, the multiple evaluation metrics in the step S6 include true positive (TP), false positive (FP), true negative (TN), false negative (FN), sensitivity (Sn), specificity (Sp), and accuracy.

[0015] In a preferred embodiment of the present invention, the application domain of the optimal QSAR model is further constructed, wherein the Euclidean distance algorithm is used to calculate the model application domain, and the first and second principal components of the principal component analysis are used to visualize and compare the application domain of the optimal QSAR model.

[0016] In one embodiment of the present invention, for mutagenicity, data of the following 5 types of strains in the NIEHS Division of Toxicology Transformation (DTT) and in vitro mutagenicity of Salmonella typhimurium (ISSSTY) are extracted: ① Salmonella typhimurium TA1535; ② Salmonella typhimurium TA1537 or TA97 or TA97a; ③ Salmonella typhimurium TA98; ④ Salmonella typhimurium TA100; ⑤ Salmonella typhimurium TA102 or Escherichia coli WP2 uvrA or Escherichia coli WP2 uvrA (pKM101). 3N*5 primary models are respectively constructed with the test conclusions of each type of strain as the end points, and then a secondary model is constructed based on the 5 primary models with the best prediction results to predict Ames test active substances.

[0017] In a further embodiment of the present invention, for the secondary model of mutagenicity - Ames test, an independent external validation set is used to evaluate its prediction ability, and the independent external validation set includes the combined results of the 5 types of strains in the Hanson database.

[0018] In various embodiments of the present invention, the human health toxicity includes carcinogenicity, mutagenicity, acute oral toxicity in rats, reproductive toxicity, endocrine disruptor - androgen disruptor, endocrine disruptor - estrogen disruptor, skin corrosion / irritation, skin sensitization, and eye corrosion / irritation; the modeling endpoints include the following twenty modeling endpoints of the human health toxicity: carcinogenicity, mutagenicity - Ames test - TA1535, mutagenicity - Ames test - TA1537 / TA97 / T97a, mutagenicity - Ames test - TA98, mutagenicity - Ames test - TA100, mutagenicity - Ames test - TA102 / E.coli WP2 uvrA / E.coli WP2 uvrA(pKM101), mutagenicity - Ames test - secondary model, mutagenicity - micronucleus test, acute oral toxicity in rats - extremely toxic, acute oral toxicity in rats - non - toxic, reproductive toxicity - male, reproductive toxicity - female, endocrine disruptor - androgen disruptor - agonist, endocrine disruptor - androgen disruptor - antagonist, endocrine disruptor - androgen disruptor - binding, endocrine disruptor - estrogen disruptor - agonist, endocrine disruptor - estrogen disruptor - antagonist, endocrine disruptor - estrogen disruptor - binding, skin corrosion / irritation, skin sensitization, eye corrosion / irritation.

[0019] The beneficial effects of the present invention are as follows:

[0020] The present invention combines international toxicity databases and literature data to ensure the wide range, diversity, and comprehensiveness of data sources, ensures the reliability of data through targeted sorting rules, and selects 15 machine learning algorithms to construct a QSAR model in three molecular description ways, enabling the constructed QSAR model to accurately predict various specific human health toxicities of small - molecule compounds and having the ability to rapidly evaluate and identify newly emerging chemical substances. Moreover, following the OECD's QSAR modeling principles enables the present invention to be widely accepted and applied globally.

[0021] In addition, the present invention separately models the 5 strains of the Ames test specified by OECD. It can not only improve the recognition rate of active substances in the Ames test, but also the prediction results for each strain can give the mutagenicity mechanism of the chemical, providing effective mechanism information for the development of green chemicals. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The drawings constituting a part of the present invention are used to provide a further understanding of the present invention, making other features, objectives, and advantages of the present invention more obvious. The schematic embodiments and descriptions of the present invention are used to explain the present invention and do not constitute an improper limitation of the present invention.

[0023] In the accompanying drawings:

[0024] Figure 1 is a flowchart of the method for predicting the human health toxicity of small molecule compounds based on machine learning of the present invention.

[0025] Figure 2 is a schematic diagram of the construction of the QSAR model of the present invention.

[0026] Figure 3 is a flowchart of the construction of the mutagenicity - Ames test model.

[0027] Figure 4 is a view of the accuracy of the internal validation of 15 models for carcinogenicity under three inputs.

[0028] Figure 5 is the average value, the highest value, and the optimal model (b) of the accuracy (a) of the external validation of 15 models for carcinogenicity under three inputs.

[0029] Figure 6 is the accuracy of the internal validation for 5 types of strains in the mutagenicity Ames test under three inputs.

[0030] Figure 7 is the accuracy (a) of the external validation for 5 types of strains in the mutagenicity Ames test under three inputs, as well as the average value, the highest value, and the optimal model (b) of the accuracy.

[0031] Figure 8 is the accuracy of the internal validation of 15 models for mutagenicity - in vivo micronucleus test (rat) under three inputs.

[0032] Figure 9 is the accuracy (a) of the external validation of 15 models for mutagenicity - in vivo micronucleus test (rat) under three inputs, as well as the average value, the highest value, and the optimal model (b) of the accuracy.

[0033] Figure 10 is the accuracy of the internal validation for rat acute oral toxicity under three inputs.

[0034] Figure 11 is the accuracy (a) of the external validation for rat acute oral toxicity under three inputs, as well as the average value, the highest value, and the optimal model (b) of the accuracy.

[0035] Figure 12 is the accuracy of the internal validation for reproductive toxicity under three inputs.

[0036] Figure 13Are the accuracy rates (a) of the external validation for reproductive toxicity under three inputs, as well as the average value, the highest value, and the optimal model of the accuracy rates (b).

[0037] Figure 14 Is the accuracy rate of the internal validation for endocrine disruptors - androgen disruptors under three inputs.

[0038] Figure 15 Are the accuracy rates (a) of the external validation for endocrine disruptors - androgen disruptors under three inputs, as well as the average value, the highest value, and the optimal model of the accuracy rates (b).

[0039] Figure 16 Is the accuracy rate of the internal validation for endocrine disruptors - estrogen disruptors under three inputs.

[0040] Figure 17 Are the accuracy rates (a) of the external validation for endocrine disruptors - estrogen disruptors under three inputs, as well as the average value, the highest value, and the optimal model of the accuracy rates (b).

[0041] Figure 18 Is the accuracy rate of the internal validation for skin corrosion / irritation of 15 models under three inputs.

[0042] Figure 19 Are the accuracy rates (a) of the external validation for skin corrosion / irritation of 15 models under three inputs, as well as the average value, the highest value, and the optimal model of the accuracy rates (b).

[0043] Figure 20 Is the accuracy rate of the internal validation for skin sensitization of 15 models under three inputs.

[0044] Figure 21 Are the accuracy rates (a) of the external validation for skin sensitization of 15 models under three inputs, as well as the average value, the highest value, and the optimal model of the accuracy rates (b).

[0045] Figure 22 Is the accuracy rate of the internal validation for eye corrosion / irritation of 15 models under three inputs.

[0046] Figure 23 Are the accuracy rates (a) of the external validation for eye corrosion / irritation of 15 models under three inputs, as well as the average value, the highest value, and the optimal model of the accuracy rates (b).

[0047] Figure 24 Is the confusion matrix of the nine human health toxicity prediction models (HHTox).

[0048] Figure 25 Is the application domain of each model.

[0049] Figure 26 It is a relationship graph of the 10 key feature structures of the optimal model based on 15 toxicity endpoints. Specific implementation manner

[0050] The present invention will be described in detail below in conjunction with embodiments and drawings.

[0051] As Figure 1 shown, in an embodiment of the present invention, a method for predicting the human health toxicity of small molecule compounds based on machine learning is provided. The method includes: Step S1, obtaining a data set containing the structures of small molecule compounds and their corresponding toxicity data from a toxicity database and public literature; and Step S2, sorting out the obtained data set to obtain a sorted data set.

[0052] 1. Data source and compound information

[0053] 1.1 Data source

[0054] According to OECD Principle 1, "QSAR should have a defined endpoint", the endpoints studied in the present invention are 20 toxicity endpoints of the following nine human health toxicities.

[0055] 1.1.1 Carcinogenicity

[0056] The present invention selects four large databases, namely the Carcinogenic Potency Database (CPDB), the Long-term carcinogenicity bioassay on rodents of biocides and plant protection products (ISSBIOC), the Chemical Carcinogens: "Structures and Experimental Data" database of the Istituto Superiore di Sanita (ISSCAN), and the Integrated Chemical Environment (ICE), to extract experimental data.

[0057] According to OECD 451 and OECD 453, the test species was selected as rats. If there was at least one positive result in either female or male animals, the substance was considered a carcinogenic active substance. In addition, two tests, "NTP Carcinogenicity" and "Two year cancer bioassay", were selected from the ICE database. It was stipulated that "clear evidence" and "some evidence" in "Level of evidence" were active data, while "no evidence" was inactive data.

[0058] After that, the extracted data was sorted out, mainly including: removing mixtures and inorganic substances; removing high-molecular types and counter ions from each database; standardizing specific chemical types; neutralizing salt compounds; and extracting SMILES (Simplified Molecular Input Line Entry System) formulas from PubChem and CompTox chemical dashboards according to the CAS number (Chemical Abstracts Service Registry Number) and deleting duplicate data; removing data with conflicting activity results. Finally, a carcinogenicity data set with unified activity classification was obtained.

[0059] 1.1.2 Mutagenicity

[0060] For gene mutation and chromosome mutation, the Ames test and in vitro micronucleus test (human cells) were respectively selected to collect data. The specific collection and sorting methods are as follows:

[0061] I. Ames test

[0062] The Ames test, namely the Salmonella typhimurium reverse mutation test, is a simple method to detect whether there are chemical carcinogens in the environment or food by using the reverse mutant strains of bacterial auxotrophs. This test usually selects a variety of different strains for testing to ensure the comprehensiveness and accuracy of the results. The following are the 5 common types of strains in the Ames test:

[0063] 1. Salmonella typhimurium

[0064] TA100: Belongs to the base pair substitution type strain and is sensitive to a variety of mutagens. Its characteristic is that there is a specific mutation in the operon of histidine biosynthesis, resulting in the inability to synthesize histidine by itself.

[0065] TA98: It belongs to the frameshift mutant strain and is particularly sensitive to frameshift mutagens. There are repetitive sequences at specific positions in its genome, making it highly sensitive to mutagens that cause base deletions in these sequences.

[0066] TA1535: It also belongs to the base pair substitution type strain, similar to TA100, but may show different sensitivities in the detection of certain mutagens.

[0067] TA1537 or TA97 or TA97a: Frameshift mutant strains, similar to TA98, but may be more sensitive to certain specific types of mutagens.

[0068] 2. Escherichia coli

[0069] WP2uvrA (or WP2uvrA(pKM101)): This is a specially modified Escherichia coli strain, usually used as a supplementary strain in the Ames test. It contains the plasmid pKM101, and the muc+ gene on this plasmid participates in the DNA repair pathway induced by DNA damage, increasing the sensitivity of the strain to mutagens. The WP2uvrA strain is defective in tryptophan biosynthesis and thus also depends on exogenous tryptophan.

[0070] These strains have different mutations in the operons encoding histidine / tryptophan biosynthesis, so they have different sensitivities to different types of mutagens. The selection of strains may vary in domestic and foreign laboratories, but usually includes the above several strains to ensure the comprehensiveness of the test. For example, more foreign laboratories choose WP2uvrA, while in China, TA102 is more commonly used (although not directly mentioned in the question, TA102 is similar to WP2uvrA in function and use).

[0071] This invention collects data based on three major databases of the Ames test, namely the Ames test database (hereinafter referred to as DTT for short) sorted out by the NIEHS Division of Translational Toxicology (DTT), the in vitro mutagenicity in Salmonella typhimurium (ISSSTY), and the Ames test constructed by Hanson et al. Among them, DTT and ISSSTY have specific strain information, so they are used as the training set and the test set, while the Hanson database only has the combined results of 5 strains, so it is used as an additional validation set.

[0072] According to the provisions of OECD Test No. 471, extract the experimental data of the following 5 types of strains in DTT and ISSSTY: ① Salmonella typhimurium TA1535; ② Salmonella typhimurium TA1537 or TA97 or TA97a; ③ Salmonella typhimurium TA98; ④ Salmonella typhimurium TA100; ⑤ Salmonella typhimurium TA102 or Escherichia coli WP2 uvrA or Escherichia coli WP2 uvrA(pKM101). For each strain, if the result of ≥1 test is positive, the substance is considered positive under this strain. For cases ② and ⑤ with multiple strains, as long as the result of ≥1 strain is positive, the substance is considered positive under this type of strain.

[0073] After that, the Ames test data is sorted out, mainly including: removing mixtures and inorganic substances; extracting SMILES formulas from PubChem and CompTox Chemical Dashboard according to CAS numbers and deleting duplicate data; removing data with conflicting activity results. Finally, an Ames test data set with unified activity classification is obtained.

[0074] II. In Vivo Micronucleus Test (Rats)

[0075] This invention selects 6 major databases, namely ECVAM(P), ECVAM(N), the database of biological test genotoxicity conclusions (hereinafter referred to as DTT-Genetox) and in vivo micronucleus database (hereinafter referred to as DTT-MN) sorted out by the NIEHS Division of Translational Toxicology (DTT), ISSMIC, and the in vivo micronucleus test database published by Yoo et al. in the literature (hereinafter referred to as Yoo), to collect in vivo micronucleus test data of rats. Among them, ECVAM(P) and ECVAM(N) have collected and sorted out mammalian micronucleus test data from NTP, EFSA, SCCS, CosE, BASF, GSK, ECHA, ISSTox, and the literature, including in vivo micronucleus test data of 7 and 25 compounds for rats respectively. DTT-Genetox is the in vivo micronucleus test information collected for different genders of rats and mice. During the collection process, the test species selected is rats, and the results of different genders are combined (if at least one result of female / male is active, the compound is considered an active substance).

[0076] The test information of DTT-MN is similar to that of DTT-Genetox, so similar data collation rules are adopted: the test results of different genders of rats are combined, and the data of "uncertain" and "no evidence" are removed from the "confidence level" column. The same collation rules also apply to the ISSMIC database. Yoo contains the in vivo micronucleus test results of 1,131 substances, and the substances with rats as the test species are selected.

[0077] After collating the data of the six major databases according to the operations described in 1.1.1, the in vivo micronucleus test results with high confidence of rats are obtained, and finally an in vivo micronucleus test data set with unified activity classification is obtained.

[0078] 1.1.3 Acute oral toxicity of rats

[0079] The present invention selects the information of the rat acute oral systemic toxicity test in the Collaborative Acute Toxicity Modeling Suite (CATMoS). This data set was initially compiled by NICEATM and CCTE of the US Environmental Protection Agency, and the data sources include OECD's eChemPortal, the National Library of Medicine's Hazardous Substances Data Bank (NLM HSDB), ChemIDplus database, and the AcutoxBase of the European Commission Joint Research Center (JRC). Mansouri et al. reviewed the data to ensure its accuracy. Then, the chemical structures were processed using a standardized workflow to obtain structures that can be directly used for QSAR modeling, and finally a rat acute oral toxicity data set with unified activity classification was obtained.

[0080] The present invention divides the above databases into two modeling endpoints according to the median lethal dose (LD50) value to meet different screening requirements:

[0081] (1) Very toxic (VT)

[0082] Activity: LD50 ≤ 50 mg / kg

[0083] Inactivity: LD50 > 50 mg / kg

[0084] (2) Nontoxic (NT)

[0085] Activity: LD50 > 2,000 mg / kg

[0086] Inactive: LD50 ≤ 2,000 mg / kg

[0087] 1.1.4 Reproductive toxicity

[0088] The present invention obtains animal toxicity data based on five major databases (ToxRefDB, NTP-ICE, GHS-J, C&L inventory, REACH).

[0089] For the ToxRefDB database, the second edition updated in 2019 (toxrefdb_v2) is selected, and two types of experiments related to reproductive toxicity, multigeneration reproductive (MGR) and reproductive (REP), are selected for data extraction. The test species is selected as rats.

[0090] For the NTP-ICE database, the rat test data of "DART, Developmental malformation", "DART, Offspring survival early", "DART, Reproductive performance", "DART, Offspring survival late", "DART, Organ weight", "DART, Developmental landmark", "Reproductive" and "Developmental" are selected. In particular, for the "DART, Organ weight" test, only reproductive organs are selected as the endpoints, such as the testis, epididymis, seminal vesicle, prostate in males, and the uterus and ovaries in females.

[0091] For GHS-J and C&L inventory, according to the GHS system, category 1A (known human reproductive toxicants) is selected, and the toxicity data is retrieved using the screening function of the eChemPortal website, compared with each other, and the compounds with inconsistent classifications are deleted. And the origin analysis of the final compounds is carried out to distinguish genders.

[0092] For the REACH database, retrieve the rat reproductive toxicity test data according to the rule of "experimental reliability is 1", and the test should be carried out in accordance with at least one of the following test guidelines: EPA OPPTS 870.3550, EPA OPPTS 870.3650, EPA OPPTS 870.3700, EPA OPPTS 870.3800, EPA OPPTS 885.3650, EU Method B.34, EU Method B.35, EU Method B.56, OECD Guideline 414, OECD Guideline 415, OECD Guideline 416, OECD Guideline 421, OECD Guideline 422 and OECD Guideline 443, and conduct traceability analysis on the retrieved compounds to distinguish genders.

[0093] Activity classification rule: If there are any effects related to reproductive toxicity in the animal test results of a certain compound, it is considered an active substance, and substances in GHS categories 1 and 2 are all active substances; other substances are non-active substances.

[0094] Integrate the above five major databases into two major databases for males / females according to gender, and organize the data in the two major databases according to the operations described in 1.1.1. Finally, obtain a reproductive toxicity data set with unified activity classification.

[0095] 1.1.5 Endocrine disruptors - Androgen disruptors

[0096] The U.S. Environmental Protection Agency (U.S. EPA) organized and carried out the Collaborative Modeling Project for Androgen Receptor Activity (CoMPARA). The present invention uses the test information related to androgen disruptors in this project as the androgen disruptor data set. This data set consists of the following two parts of data:

[0097] (1) ToxCast / Tox21 in vitro test data based on the AR pathway. Eleven groups of ToxCast TM / Tox21 in vitro test data were selected. TM / Tox21 in vitro assays: 3 receptor binding assays, 2 cofactor recruitment assays, 1 RNA transcription assay, 3 agonist-mode protein syntheses, and 2 antagonist-mode protein syntheses. An AUC score representing the overall AR activity was calculated based on the above in vitro assay results to simulate in vivo results. One of the antagonist-mode assays was conducted with two different concentrations of the stimulating ligand to provide confirmation data on receptor-mediated activity and to further distinguish true antagonist activity from loss of function mediated by cellular stress or cytotoxicity. Based on the above approach, CoMPARA conducted assays on 1855 ToxCast TM chemicals and obtained AUC scores for AR agonist and antagonist activities, thus constructing a high-quality database based on the AR pathway.

[0098] The activity classification criteria are as follows: active: AUC ≥ 0.1 (corresponding to an active concentration up to 100 μM); inactive: AUC < 0.001.

[0099] In addition, due to the limited activity data, active chemicals collected from DrugBank were also added to the dataset and their AUC scores were considered to be 1.

[0100] Since the AUC score is only applicable to agonist and antagonist activities, it is stipulated that if a chemical is an agonist or antagonist activity, it is considered a binder activity.

[0101] (2) Other in vitro assay data: The present invention selected PubChem data (64 sources) collected and collated by the National Center for Computational Toxicology (NCCT) of the US Environmental Protection Agency and grouped them according to agonist / antagonist and active / inactive. The data were divided into three categories: agonists, antagonists, and binders (each category contains active and inactive compounds) according to the following rules:

[0102] If there are multiple assay records for a chemical in this data, at least 2 of the 3 assay results must be consistent to be retained; agonist or antagonist active substances are considered binder active substances; agonist inactive and antagonist inactive substances are defined as binder inactive substances. Chemicals that are duplicates of the ToxCast TM / Tox21 in vitro assay data based on the AR pathway were deleted.

[0103] For the above data, the KNIME software standardization workflow was used to standardize their chemical structures, and finally an androgen disruptor dataset with a unified activity classification was obtained.

[0104] 1.1.6 Endocrine disruptors - Estrogen disruptors

[0105] The present invention selects the test information related to estrogen disruptors in the Collaborative Estrogen Receptor Activity Prediction Project (CERAPP) organized by the United States Environmental Protection Agency

[48] . Similar to CoMPARA, this dataset is also composed of two parts of data.

[0106] (1) ToxCast / Tox21 in vitro test data based on the ER pathway TM The present invention selects 18 in vitro HTS tests based on the mammalian ER pathway. These tests include biochemical and cell-based in vitro assays that can detect perturbations in the ER pathway response at intracellular sites: receptor binding, receptor dimerization, chromatin binding of mature transcription factors, gene transcription, and changes in ER-induced cell growth kinetics. According to the information of each test result, a mathematical model is used to integrate them to form an area under the curve (AUC) score to quantify the ER agonist, antagonist, and binding activities of chemicals. The AUC ranges from 0 to 1. If the agonist or antagonist score of a chemical is higher than 0.1, the substance is considered an active substance.

[0107] (2) Other in vitro test data:

[0108] HTS data from Tox21: including approximately 8,000 chemicals evaluated in four assays;

[0109] Estrogenic Activity Database (EADB) developed by the U.S. Food and Drug Administration (U.S. FDA): containing literature-derived ER data of approximately 8,000 chemicals;

[0110] Database of the Ministry of Economy, Trade and Industry, Japan (METI): estrogen activity data of approximately 2,000 chemicals;

[0111] ChEMBL database maintained by the European Bioinformatics Institute (EMBL-EBI): estrogen activity data of approximately 2,000 chemicals.

[0112] Integrate the data from the above 4 databases, and remove in vivo data, cytotoxicity information, and all ambiguous entries (missing values, undefined / non-standard endpoints, and data with unclear units). Delete the chemicals that overlap with the first part. Then classify the chemicals into three test categories: binding, reporter gene / transcriptional activation, or cell proliferation. For compounds with a detected concentration, convert the concentration unit to μM. If the compound is reported as active, consider it as the Half Maximal Activity Concentration (AC50) value. For cell proliferation assays, if a chemical shows a proliferation exceeding the threshold of 125%, it is considered active. The AC50 value for all inactive compounds is 1M.

[0113] Finally, for the data in the two parts, use the KNIME software's standardization workflow to standardize the chemical structures, and finally obtain an estrogenic disruptor dataset with a unified activity classification.

[0114] 1.1.7 Skin corrosion / irritation

[0115] This invention selects the dataset collected and sorted by the Systemic and Topical chemical Toxicity (STopTox) portal website (https: / / stoptox.mml.unc.edu / ). The skin corrosion / irritation data in STopTox mainly comes from the REACH database. According to the classification information of the Globally Harmonized System of Classification and Labelling of Chemicals (GHS) (compliant with OECD Test No. 404), it is stipulated that substances classified as GHS 1 - 3 are positive substances for skin corrosion / irritation.

[0116] After that, organize the chemical structures of the data in the dataset, mainly including: removing mixtures, inorganic substances, and organometallic compounds; cleaning and neutralizing salts; normalizing specific chemical types; and performing special processing on chemicals with multiple duplicate records: (1) when the duplicate results are the same, only retain one; (2) when the duplicate results are different, follow the principle of the minority obeying the majority; (3) when the duplicate results are different and the quantities on both sides are the same, delete all records. Finally, obtain a skin corrosion / irritation dataset with a unified activity classification.

[0117] 1.1.8 Skin sensitization

[0118] The present invention uses the database of human predictive patch tests (HPPTs) collected by NICEATM and the German Federal Institute for Risk Assessment (BfR). This database includes two types of human skin sensitization tests: human repeated insult patch tests and human maximization tests.

[0119] Chemicals containing "QSAR.Ready.SMILES" are selected because their chemical structures have been desalted, stereochemistry removed, and tautomers standardized, so they can be directly used for QSAR model prediction. Other chemicals that cannot be predicted by the QSAR model, such as metals, inorganic substances, and mixtures, are deleted.

[0120] Since the database contains test results of test substances at different concentrations, we stipulate that if a positive result appears for the test substance at any concentration, then the substance is considered a skin sensitizer. Through the above operations, a skin sensitization dataset with a unified activity classification is finally obtained.

[0121] 1.1.9 Eye corrosion / irritation

[0122] The present invention constructs a dataset based on the in vivo eye irritation tests in the ICE database and the eye corrosion / irritation data in STopTox. The test data of the former comply with the OECD test guideline 405 specifically for evaluating the eye irritation of chemicals, and the test results are presented according to the classification criteria of GHS (as stipulated in OECD test guideline 405); the latter comes from the integration of data from the REACH database and literature data, including: first screening data from each source separately, then merging the screened datasets, and deleting duplicate substances, which also follows the OECD test guideline 405.

[0123] Activity classification criteria: Substances classified as GHS category 1, 2 (including 2A, 2B) are active substances, and other substances are inactive substances.

[0124] Integrate the data of the above two databases, and sort out the data of the databases according to the operations described in 1.1.1. Finally, an eye corrosion / irritation dataset with a unified activity classification is obtained.

[0125] 1.2 Compound information

[0126] Based on the various datasets with unified activity classification collected and sorted in 1.1, in an embodiment of the present invention, the compound information of the top 20 modeling endpoints of nine types of human health toxicity is finally obtained, as shown in Table 1 below. It includes a total of 43 tests, covering 90,367 toxicity information of 21,439 compounds.

[0127] Table 1

[0128]

[0129]

[0130] 2. Modeling method

[0131] The method for predicting the human health toxicity of small molecule compounds based on machine learning according to the present invention includes: Step S3, calculating the molecular descriptors and molecular fingerprints of the small molecule compounds in the sorted dataset; additionally or optionally, preprocessing the calculated molecular descriptors and the molecular fingerprints of the small molecule compounds.

[0132] 2.1 Calculation and preprocessing of molecular descriptors and molecular fingerprints

[0133] In an embodiment of the present invention, PaDEL-Descriptor software is used to calculate the molecular descriptors and molecular fingerprints of compounds. Select all 1D and 2D molecular descriptors, including 1444 molecular descriptors of types such as acidic group count, ALOGP, aromatic atom count, and carbon type; calculate all molecular fingerprints, a total of 12 types of molecular fingerprints: FP, ExtFP, EStateFP, GraphFP, MACCSFP, PubchemFP, SubFP, SubFPC, KRFP, KRFPC, AP2D, and APC2D, with a total fingerprint bit number of 16092 bits.

[0134] In an embodiment of the present invention, the preprocessing of molecular descriptors and molecular fingerprints includes: deleting features with missing values and overly large data (≥10000); filtering out features with a variance of 0; using the mutual information method for correlation filtering to avoid overfitting; in order to prevent certain features from contributing too much to the model, using the z-score normalization method to normalize them (Formula 1):

[0135]

[0136] Where x is the feature value, μ is the mean of each feature, and σ is the standard deviation of each feature.

[0137] 2.2 Modeling method for nine human health toxicity prediction models

[0138] According to OECD Principle 2, "The QSAR model should have a clear algorithm." The method for predicting the human health toxicity of small molecule compounds based on machine learning according to the present invention includes: Step S4, using N machine learning algorithms to construct a QSAR model with molecular descriptors, molecular fingerprints, and a combination of molecular descriptors and molecular fingerprints as inputs respectively, so that 3N initial QSAR models are constructed for each modeling endpoint.

[0139] In an embodiment of the present invention, fifteen (i.e., N = 15) representative machine learning algorithms are used to construct a binary classification model. The fifteen machine learning algorithms are based on the following five basic algorithms: Random Forest (RF), Support Vector Machine (SVM), Decision Tree (SVM), Naive Bayes (NB), and K Nearest Neighbor (KNN).

[0140] For SVM, five SVMs constructed based on five kernel functions (Linear Kernel, Polynomial Kernel, Radial Basis Function (RBF), non - linear activation function kernel function of neurons (Sigmoid tanh), and custom kernel function (precomputed)) are adopted;

[0141] For NB, four NB algorithms are adopted: GaussianNB, BernoulliNB, MultinomialNB, and ComplementNB;

[0142] Meanwhile, KNN constructed based on the following four algorithms is adopted: "ball_tree", "kd_tree", "brute", and "auto".

[0143] Finally, the following fifteen algorithms were selected for the present invention: RF, SVM-linear, SVM-poly, SVM-RBF, SVM-sigmoid, SVM-precomputed, DT, NB-GaussianNB, NB-BernoulliNB, NB-MultinomialNB, NB-ComplementNB, kNN-algorithm='auto', kNN-algorithm=ball_tree, kNN-algorithm=kd_tree, and kNN-algorithm='brute'. These fifteen machine learning algorithms were combined pairwise with the following three inputs for model construction: molecular descriptors, molecular fingerprints, and molecular descriptors_molecular fingerprints. Therefore, 3*15 initial QSAR models can be constructed for each modeling endpoint ( Figure 2 ).

[0144] Additionally or optionally, the hyperparameters of the machine learning algorithms were selected to achieve the best model performance. The present invention uses the grid search method to select the hyperparameters.

[0145] In one embodiment of the present invention, due to the imbalance in the number of active compounds and inactive compounds such as acute oral toxicity, androgen / estrogen disruptors, and skin sensitizers, oversampling techniques can be used for treatment.

[0146] As an innovative aspect of the present invention, for mutagenicity, the OECD test guidelines stipulate that the results of 5 types of strains must be comprehensively considered in the Ames test. If the result of one strain is positive, the substance is a mutagenic active substance. Existing studies have all constructed models with the integrated results of 5 types of strains as the endpoint.

[0147] See Figure 3 , in one embodiment of the present invention, 45*5 primary models were constructed respectively with the test conclusions of each type of strain as the endpoint, and then a secondary model was constructed based on the 5 best primary models with the best prediction results to predict the final Ames test active substances. Specifically, if the prediction result of one of the 5 best primary models for a certain substance is active, then the substance is a mutagenic active substance.

[0148] The method for predicting the human health toxicity of small molecule compounds based on machine learning according to the present invention includes: Step S5, dividing the sorted data set into a training data set and a test data set at a preset ratio, using the molecular descriptors, molecular fingerprints of the small molecule compounds in the training data set, and the combination of molecular descriptors and molecular fingerprints as inputs, and the corresponding toxicity data of the small molecule compounds in the training data set as outputs, training 3N initial QSAR models, and obtaining 3N target QSAR models after the training is completed. In an embodiment of the present invention, the obtained activity data set is arbitrarily divided into a training set (80%) and a test set (20%) at a ratio of 4:1. The training set is used for model construction and internal verification, and the test set is used for external verification. The training data set is used to train 3 * 15 initial QSAR models.

[0149] 2.3 Evaluation of Nine Human Health Toxicity Prediction Models

[0150] According to OECD Principle 4, "the QSAR model should have appropriate fitness, robustness, and predictive ability". The method for predicting the human health toxicity of small molecule compounds based on machine learning according to the present invention includes: Step S6, using the test data set to evaluate the performance of 3N target QSAR models based on multiple evaluation metrics, and selecting the optimal QSAR model for each modeling endpoint therefrom, so as to form an optimal QSAR model set.

[0151] In an embodiment of the present invention, the test data set is used to externally verify the prediction effect of 3 * 15 QSAR models using the following 7 metrics: true positive (TP), false positive (FP), true negative (TN), false negative (FN), sensitivity (Sn), specificity (Sp), and accuracy. Among them, the application of sensitivity (Sn) and specificity (Sp) ensures that the model can accurately identify compounds with and without toxicity; the analysis of true positive (TP) and false positive (FP) helps to optimize the model and reduce the possibility of misjudgment; the measurement of true negative (TN) and false negative (FN) provides important information about the model's ability to identify non-toxic compounds; accuracy, as a comprehensive metric, reflects the overall prediction efficiency of the model.

[0152] Accuracy can be calculated by the following formula (2):

[0153]

[0154] The sensitivity (Sn) or true positive rate (TPR) can be calculated by the following formula (3):

[0155]

[0156] The specificity (Sp) can be calculated by the following formula (4):

[0157]

[0158] It should be noted that for the secondary model of the mutagenicity - Ames test, an independent external validation set was used to evaluate its predictive ability.

[0159] Finally, the predictive performance of the model was verified using the above seven indicators, and the model with the best predictive ability was selected as the optimal QSAR model. A total of 20 optimal QSAR models were selected for the 20 modeling endpoints of the nine toxicity endpoints to form an optimal QSAR model set. The model performance details are shown in Figures 4 to 23 and Table 2, and the accuracy of each optimal model is between 0.724 - 0.996.

[0160] 3. Model Results and Model Characterization Methods

[0161] In Figure 4 , Figure 5 , Figure 8 , Figure 9 , Figures 18 to 23 , models 1 - 15 represent RF, SVM - liner, SVM - poly, SVM - RBF, SVM - sigmoid, SVM - precomputed, DT, NB - GaussianNB, NB - BernoulliNB, MultinomialNB, ComplementNB, kNN - algorithm = 'auto', kNN - algorithm = 'ball_tree', kNN - algorithm = 'kd_tree', and kNN - algorithm = 'brute', respectively.

[0162] The results show that the prediction effects of the following fourteen models are the best using the Random Forest (RF) algorithm: carcinogenicity (Model 1), mutagenicity - Ames test - TA1535 (Model 2), mutagenicity - Ames test - TA1537 / TA97 / T97a (Model 3), mutagenicity - Ames test - TA98 (Model 4), mutagenicity - Ames test - TA100 (Model 5), acute oral toxicity in rats - extremely toxic (Model 8), acute oral toxicity in rats - non-toxic (Model 9), reproductive toxicity - male (Model 10), reproductive toxicity - female (Model 11), endocrine disruptor - androgen disruptor - antagonist (Model 13), endocrine disruptor - estrogen disruptor - agonist (Model 15), endocrine disruptor - estrogen disruptor - binding (Model 17), skin corrosion / irritation (Model 18), and eye corrosion / irritation (Model 20);

[0163] For skin corrosion / irritation (Model 18), the SVM - linear and SVM - precomputed algorithms can achieve the same prediction effect as RF;

[0164] Endocrine disruptor - androgen disruptor - agonist (Model 12) achieves the best prediction ability under the SVM - RBF algorithm;

[0165] The optimal model algorithms for mutagenicity - Ames test - TA102 / E.coli WP2 uvrA / E.coli WP2 uvrA(pKM101) (Model 6), endocrine disruptor - androgen disruptor - binding (Model 14), and skin sensitization (Model 19) are kNN - algorithm = 'auto', kNN - algorithm = 'ball_tree', and kNN - algorithm = 'auto', respectively;

[0166] The two modeling endpoints of mutagenicity - micronucleus test (Model 7) and endocrine disruptor - estrogen disruptor - antagonist (Model 16) achieve the optimal performance under the Decision Tree (DT) algorithm.

[0167] Looking at the optimal models under the three molecular description methods, it is found that molecular descriptors - molecular fingerprints are more likely to obtain high-performance models. When using molecular descriptors and molecular fingerprints in combination, the optimal models for the following nine modeling endpoints are obtained: Mutagenicity - Ames test - TA1537 / TA97 / T97a (Model 3), Mutagenicity - Ames test - TA100 (Model 5), Mutagenicity - Ames test - TA102 / E. coli WP2 uvrA / E. coli WP2 uvrA (pKM101) (Model 6), Rat acute oral toxicity - non-toxic (Model 9), Reproductive toxicity - male (Model 10), Reproductive toxicity - female (Model 11), Endocrine disruptor - estrogen disruptor - agonist (Model 15), Endocrine disruptor - estrogen disruptor - antagonist (Model 16), and Endocrine disruptor - estrogen disruptor - binding (Model 17).

[0168] Using molecular descriptors, the optimal models for the following seven modeling endpoints are obtained: Mutagenicity - Ames test - TA1535 (Model 2), Mutagenicity - Ames test - TA98 (Model 4), Mutagenicity - Micronucleus test (Model 7), Endocrine disruptor - androgen disruptor - agonist (Model 12), Endocrine disruptor - androgen disruptor - binding (Model 14), Skin corrosion / irritation (Model 18), and Eye corrosion / irritation (Model 20).

[0169] Using molecular fingerprints, the optimal models for the following four modeling endpoints are obtained: Carcinogenicity (Model 1), Rat acute oral toxicity - extremely toxic (Model 8), Endocrine disruptor - androgen disruptor - antagonist (Model 13), and Skin sensitization (Model 19).

[0170] The model evaluation results of each optimal model are shown in Table 2 below and Figure 24 .

[0171] Table 2

[0172]

[0173]

[0174]

[0175]

[0176]

[0177] a: Abbreviation, k Nearest Neighbor (kNN); Naive Bayes( Bayes, NB); Random Forest (RF); Support Vector Machine (SVM); Decision Tree (SVM); True Negative (TN), False Positive (FP), False Negative (FN), True Positive (TP), Accuracy (Acc, acc), Sensitivity (Sn), and Specificity (Sp).

[0178] 4. Application Domain Characterization Method and Results

[0179] According to OECD Principle 3, "The QSAR model should have a defined application domain." The QSAR model has a high dependence on the training set, and its reliability is always limited to a specific range. Therefore, it is necessary to define the model application domain. In one embodiment of the present invention, the Euclidean distance algorithm is used to calculate the model application domain. Compounds are represented as a point formed by high-dimensional molecular descriptors in Euclidean space, and the average value of each scaled descriptor is the centroid of the training compounds.

[0180] Then calculate the Euclidean distance d k,centroid between compound i and the centroid x Euc,i :

[0181]

[0182] where x k,i is the k-th scaled descriptor of compound i, and x k,centroid is the k-th descriptor of the centroid. First, calculate the distances between all training set compounds and the centroid, and define the maximum distance as the threshold d t . When d Euc,i > d t , the compound is considered to be within the application domain.

[0183] The first principal component and the second principal component of principal component analysis are used to visualize and compare the application domain of the model. For the consensus model of the mutagenicity - Ames test, its application domain is the union of the 5 optimal models.

[0184] The application domains of the models of the present invention are as Figure 25 shown, and the first principal component and the second principal component of PCA can explain at least 60% of the data.

[0185] 5. Model Mechanism Revelation

[0186] According to OECD Principle 5, "if possible, the mechanism should be explained". Table 3 below lists the 10 key feature structures of the optimal models for 20 toxicity endpoints, and the relationship diagrams of each key feature are shown in Figure 26 Among them. The molecular descriptors that intersect with more than 3 modeling endpoints are ALogP, ALogP2, AMR, apol, naAromAtom, nAromBond, nAtom, nHeavyAtom, nH, nC, nN, and ATS0e, which represent the partition coefficient between octanol and water, the square of the octanol-water partition coefficient, molar refractivity, molecular polarizability, number of aromatic atoms, number of free bonds of aromatic bonds, number of atoms, number of heavy atoms, number of hydrogen atoms, number of carbon atoms, number of nitrogen atoms, and normalized Moreau-Broto autocorrelation, respectively; the molecular fingerprint that intersects with more than 3 modeling endpoints is PubchemFP72, which means the presence of the chemical element niobium.

[0187] Table 3

[0188]

[0189]

[0190]

[0191] The method for predicting the human health toxicity of small molecule compounds based on machine learning according to the present invention includes: Step S7, inputting the molecular descriptors and molecular fingerprints corresponding to the small molecule compound to be tested into the set of optimal QSAR models for prediction, so as to obtain the human health toxicity of the small molecule compound to be tested.

[0192] Examples

[0193] Given a compound: cyclopentylamine (CAS No.: 1003-03-8), to predict its nine human health toxicities, first calculate the molecular descriptors and molecular fingerprints using PaDEL-Descriptor software according to its SMILES formula, and then determine that it is within the application domain of each optimal model. According to the corresponding molecular descriptors and molecular fingerprints, use each optimal model for prediction. The prediction results are shown in Table 3 below. The results show that this compound has potential micronucleus test mutagenicity, acute oral toxicity (but not extremely toxic), female reproductive toxicity, skin corrosion and irritation, skin sensitization, and eye corrosion and irritation, and no Ames test mutagenic toxicity. Existing literature shows that cyclopentylamine has acute oral toxicity (but not extremely toxic), skin corrosion and irritation, skin sensitization, and eye corrosion and irritation, and no Ames test mutagenic toxicity. There is a lack of experimental data for other toxicity endpoints, which indicates that the existing experimental results are the same as the prediction results.

[0194] Table 3

[0195]

[0196] In summary, the present invention selects 15 machine learning algorithms, combines the international toxicity database and literature data to ensure the wide range, diversity and comprehensiveness of the data sources, ensures the reliability of the data through the targeted sorting rules, and constructs a QSAR model in three molecular description ways, so that the constructed QSAR model can accurately predict various specific human health toxicities of small molecule compounds and has the ability to rapidly evaluate and identify newly emerging chemical substances. Moreover, following the QSAR modeling principles of OECD enables the present invention to be widely accepted and applied globally.

[0197] The above description is only some preferred embodiments of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) disclosed in the embodiments of the present disclosure that have similar functions.

Claims

1. A method for predicting the human health toxicity of small molecule compounds based on machine learning, characterized in that Including: Step S1: Obtain a dataset containing the structures of small molecule compounds and their corresponding toxicity data from a toxicity database and public literature; Step S2: Organize the obtained dataset to get an organized dataset; Step S3: Calculate the molecular descriptors and molecular fingerprints of the small molecule compounds in the organized dataset; Step S4: Using the molecular descriptors, molecular fingerprints, and combinations of molecular descriptors and molecular fingerprints as inputs respectively, use N machine learning algorithms to construct QSAR models, so as to construct 3N initial QSAR models for each modeling endpoint; Step S5: Divide the organized dataset into a training dataset and a test dataset at a preset ratio. Using the molecular descriptors, molecular fingerprints, and combinations of molecular descriptors and molecular fingerprints of the small molecule compounds in the training dataset as inputs, and the corresponding toxicity data of the small molecule compounds in the training dataset as outputs, train the 3N initial QSAR models. After training, obtain 3N target QSAR models; Step S6: Use the test dataset to evaluate the performance of the 3N target QSAR models based on multiple evaluation metrics, and select the optimal QSAR model for each modeling endpoint from them, thus forming an optimal QSAR model set; And Step S7: Input the molecular descriptors and molecular fingerprints corresponding to the small molecule compound to be tested into the optimal QSAR model set for prediction, so as to obtain the human health toxicity of the small molecule compound to be tested.

2. The method for predicting the human health toxicity of small molecule compounds based on machine learning according to claim 1, wherein The organizing the obtained dataset in Step S2 includes one or more of the following operations: Removing one or more of mixtures, inorganic substances, organometallic compounds, macromolecular types, and counterions; Standardizing specific chemical types; Neutralizing salt compounds; Removing data with conflicting activity results; Standardizing chemical structures using a standardized workflow; Normalizing specific chemical types; Integrating the data in the database; Removing missing values, undefined / non-standard endpoints, and data with unclear units; Classifying chemical substances; Merging the screened datasets; Deleting duplicate data; And Special processing for chemical substances with multiple duplicate records: (1) When the duplicate results are the same, only keep one; (2) When the duplicate results are different, follow the principle of the minority obeying the majority; (3) When the duplicate results are different and the numbers on both sides are the same, delete all records.

3. The method for predicting the human health toxicity of small molecule compounds based on machine learning according to claim 1, wherein The machine learning algorithms include: Random forest, support vector machine, decision tree, naive Bayes, and K-nearest neighbor; Among them, the support vector machine includes 5 SVM algorithms constructed based on linear kernel, polynomial kernel, radial basis kernel function, non-linear activation function kernel function of neurons, and custom kernel function, namely: SVM-linear, SVM-poly, SVM-RBF, SVM-sigmoid, SVM-precomputed; Among them, naive Bayes includes 4 types, namely: Gaussian naive Bayes, Bernoulli naive Bayes, multinomial naive Bayes, and complementary naive Bayes; Among them, K-nearest neighbors includes 4 KNN algorithms constructed based on "ball_tree", "kd_tree", "brute" and "auto", namely: kNN-algorithm = 'auto', kNN-algorithm = ball_tree, kNN-algorithm = kd_tree, and kNN-algorithm = 'brute'.

4. The method for predicting the human health toxicity of small molecule compounds based on machine learning according to claim 1, wherein Calculating the molecular descriptors and molecular fingerprints of the small molecule compounds in step S3 further includes preprocessing the calculated molecular descriptors and the molecular fingerprints, including one or more of the following operations: Deleting features containing missing values and over-large data; Filtering out features with a variance of 0; Using the mutual information method for correlation filtering to avoid overfitting; Normalizing using the following z-score normalization method: where x is the feature value, μ is the mean of each feature, and σ is the standard deviation of each feature, and among them, the molecular descriptors include 1D molecular descriptors and 2D molecular descriptors; The molecular fingerprints include the following 12 types of molecular fingerprints: FP, ExtFP, EStateFP, GraphFP, MACCSFP, PubchemFP, SubFP, SubFPC, KRFP, KRFPC, AP2D, and APC2D.

5. The method for predicting the human health toxicity of small molecule compounds based on machine learning according to claim 4, wherein, Step S4 further includes using the grid search method to select the hyperparameters of the machine learning algorithm.

6. The method for predicting the human health toxicity of small molecule compounds based on machine learning according to claim 1, wherein The multiple evaluation metrics in step S6 include true positive, false positive, true negative, false negative, sensitivity, specificity, and accuracy.

7. The method for predicting the human health toxicity of small molecule compounds based on machine learning according to claim 1, wherein Further, the Euclidean distance algorithm is used to calculate the application domain of the optimal QSAR model.

8. The method for predicting the human health toxicity of small molecule compounds based on machine learning according to claim 7, wherein It also includes visualizing and comparing the application domain of the optimal QSAR model using the first principal component and the second principal component of the principal component analysis.

9. The method for predicting the human health toxicity of small molecule compounds based on machine learning according to any one of claims 1 to 8, characterized in that, For mutagenicity, data of the following 5 types of strains in the in vitro mutagenicity of the NIEHS Division of Translational Toxicology and Salmonella typhimurium are extracted: ① Salmonella typhimurium TA1535; ② Salmonella typhimurium TA1537 or TA97 or TA97a; ③ Salmonella typhimurium TA98; ④ Salmonella typhimurium TA100; ⑤ Salmonella typhimurium TA102 or Escherichia coli E. coli WP2 uvrA or Escherichia coli E. coli WP2 uvrA (pKM101). 3N*5 primary models are constructed respectively with the test conclusions of each type of strain as the endpoints, and then a secondary model is constructed based on the 5 primary models with the best prediction results to predict Ames test active substances.

10. The method for predicting the human health toxicity of small molecule compounds based on machine learning according to claim 1, wherein The human health toxicities include carcinogenicity, mutagenicity, acute oral toxicity in rats, reproductive toxicity, endocrine disruptor - androgen disruptor, endocrine disruptor - estrogen disruptor, skin corrosion / irritation, skin sensitization, and eye corrosion / irritation; the modeling endpoints include the following twenty modeling endpoints of the human health toxicities: carcinogenicity, mutagenicity - Ames test - TA1535, mutagenicity - Ames test - TA1537 / TA97 / T97a, mutagenicity - Ames test - TA98, mutagenicity - Ames test - TA100, mutagenicity - Ames test - TA102 / E. coli WP2 uvrA / E. coli WP2 uvrA(pKM101), mutagenicity - Ames test - secondary model, mutagenicity - micronucleus test, acute oral toxicity in rats - extremely toxic, acute oral toxicity in rats - non - toxic, reproductive toxicity - male, reproductive toxicity - female, endocrine disruptor - androgen disruptor - agonist, endocrine disruptor - androgen disruptor - antagonist, endocrine disruptor - androgen disruptor - binding, endocrine disruptor - estrogen disruptor - agonist, endocrine disruptor - estrogen disruptor - antagonist, endocrine disruptor - estrogen disruptor - binding, skin corrosion / irritation, skin sensitization, eye corrosion / irritation.

Citation Information

Patent Citations

  • PAHs carcinogenicity prediction method and prediction model based on 2D molecular descriptors

    CN114530213A

  • Method for screening hepatotoxic chemicals based on random forest algorithm

    CN116312858A