Harmful outcome path-based chemical reproductive toxicity transfer learning model prediction method
By constructing a deep neural network source model based on harmful outcome paths and performing transfer learning, the accuracy and robustness of predicting live reproductive toxicity in the prior art are solved, and efficient screening and regulatory support for chemical reproductive toxicity is achieved.
Patent Information
- Application Number
- CN202510190620.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-07-04
AI Technical Summary
The prior art is difficult to efficiently and at low cost to integrate massive ex vivo test data for live reproductive toxicity prediction, and the existing machine learning models are prone to overfitting on small data sets, and their accuracy and robustness are insufficient.
A deep neural network source model is constructed based on harmful outcome paths, and the reproductive toxicity of chemicals is predicted by stepping freezing the network layer for transfer learning. SMILES is used to generate MACCS molecular fingerprints, and the model performance is evaluated based on the area under the curve, F1 score and equilibrium accuracy, and the application domain is characterized.
It has achieved high accuracy prediction of live reproductive toxicity, improved the robustness of the model, and has clear application domains. It is suitable for high-throughput screening of reproductive toxicity of chemicals, and supports chemical health risk assessment and supervision.
Smart Images

Figure CN120260716A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular, to a prediction method for a chemical reproductive toxicity transfer learning model based on an adverse outcome pathway. Background Art
[0002] The evaluation of the potential reproductive toxicity of chemicals is a prerequisite for their health risk control and plays a key role in formulating scientific regulatory strategies. Reproductive toxicity refers to the adverse effects of chemical substances or mixtures on the sexual function, fertility of adults, and the development of offspring after exposure. For example, chemicals such as endocrine disruptors (EDCs) enter the human body and cause reproductive toxicity by interfering with the normal function of the endocrine system. However, there are numerous chemical substances, and only less than 1% of them have reproductive toxicity information, which hinders their management and health risk prevention and control. Therefore, it is particularly important to adopt an efficient method to screen the reproductive toxicity of chemicals.
[0003] The Organization for Economic Cooperation and Development advocates the evaluation of the reproductive toxicity of chemicals based on animal experiments on rats and mice, and has issued test guidelines such as one-generation reproductive toxicity, two-generation reproductive toxicity, and extended one-generation reproductive toxicity. The half inhibitory concentration obtained from the test is compared with the designated standard threshold to determine whether a chemical has reproductive toxicity.
[0004] Animal experiments can well evaluate the reproductive toxicity of chemicals, but it is difficult to clarify the complex toxic action mechanism, and the test costs are high, the cycle is long, and it involves animal ethics issues. In recent years, the high-throughput screening method HTS based on in vitro cells or nuclear receptors has been generally regarded as an alternative evaluation method for the reproductive toxicity of chemicals. HTS can achieve rapid screening of reproductive toxicity, but it depends on chemical standards and equipment, and cannot reflect the in vivo pharmacokinetic process, resulting in deviations between the results and the in vivo effects. At the same time, relying solely on experimental tests is difficult to meet the needs of evaluating the reproductive toxicity of a large number of chemicals, and it is necessary to develop an efficient and low-cost screening technology.
[0005] Computational toxicology constructs mathematical or computer models to achieve efficient prediction and evaluation of the hazards and risks of chemicals. With the development of artificial intelligence technology, some machine learning methods such as random forest, support vector machine, and k-nearest neighbor have been widely used to construct computational toxicology models. These models have the advantages of low cost, fast calculation rate, and high throughput, and have gradually become alternative methods for experimental tests. Existing studies have developed reproductive toxicity machine learning prediction models. For example, the literature "Front. Toxicol., 2022, 4, 981928" and "Front. Pharmacol., 2022, 13, 1018226" respectively modeled in vitro and in vivo reproductive toxicity, but the prediction accuracy of the models is not ideal. On the one hand, the HTS data is difficult to reflect the in vivo environment, and on the other hand, the amount of in vivo toxicity data is too small, resulting in poor model robustness.
[0006] How to integrate massive in vitro test data for in vivo toxicity prediction with small data samples has become a trend in the development of computational toxicology models. The Adverse Outcome Pathway (AOP) framework provides multi-level toxic effect information from initiating events (MIEs) at the molecular level, to key events (KEs) in the middle, and to adverse outcomes (AOs) at the individual level, providing a large amount of data for reproductive toxicity modeling. The literature "Environ. Sci. Technol., 2022, 56, 12391" constructs AOP multi-level endpoint models and stacks them to predict reproductive toxicity; the literature "Environ. Sci. Technol., 2021, 55, 10875" proposes a model framework that integrates reproductive toxicity AOP data into deep neural networks. Although these models have improved accuracy and interpretability, the model architecture is complex and requires modeling for multiple biological tests. It is also prone to overfitting due to the small number of in vivo data samples (dozens of chemicals). Therefore, it is necessary to provide a new transfer learning strategy to solve the small data set modeling problem by modeling on a large data set (source domain) and migrating it to a small data set (target domain), thereby providing an effective strategy for reproductive toxicity prediction. Summary of the invention
[0007] According to the technical problems raised above, a method for predicting the reproductive toxicity of chemicals based on harmful outcome pathways through transfer learning models is provided. The present invention is based on in vitro test data of androgen / estrogen α receptor (AR / ERα). By constructing a deep neural network source model, and gradually freezing the network layers, fine-tuning on different events downstream of AOP, it finally migrates and predicts reproductive toxicity in vivo. The model can directly predict whether a chemical substance has reproductive toxicity based on its SMILES. The transfer learning prediction model created by the present invention has a wider application domain and high prediction accuracy, and can be used for high-throughput prediction of reproductive toxicity of chemicals. This method is expected to be widely expanded to the modeling and prediction of other in vivo toxicity endpoints, and to assist in the evaluation and supervision of health risks of chemical substances.
[0008] The technical means adopted by the present invention are as follows:
[0009] A method for predicting chemical reproductive toxicity transfer learning model based on harmful outcome pathways, including:
[0010] Collect test data on the reproductive toxicity of chemicals in men and women at different levels in the harmful outcome pathway and pre-process the test data;
[0011] The preprocessed test data is randomly split into training set, validation set and test set, and the source model is trained using a deep neural network. After migration, the target model is obtained.
[0012] The prediction performance of the target model was evaluated using the area under the curve, F1 score, and balanced accuracy as indicators;
[0013] Characterize the application domain of the target model according to the index similarity density and weighted ruggedness;
[0014] Use the target model to predict whether there is reproductive toxicity of chemicals in the application domain to living organisms.
[0015] Further, the test data includes androgen-mediated male reproductive toxicity adverse outcome pathways and estrogen-mediated female reproductive toxicity adverse outcome pathways;
[0016] The androgen-mediated male reproductive toxicity adverse outcome pathway includes: androgen binding, co-regulator recruitment, chromatin binding, transcription factor activation, up-regulation of gene expression, cell proliferation, and increase in sex organ weight; the estrogen-mediated female reproductive toxicity adverse outcome pathway includes: estrogen binding, estrogen dimerization, chromatin binding, transcription factor activation, up-regulation of gene expression, cell proliferation, and increase in sex organ weight.
[0017] Further, the preprocessing of the test data includes:
[0018] Generate the SMILES of each chemical through the PubChem database, and standardize it to remove metal salts, inorganic substances, and mixtures; for in vitro data, mark the chemicals that are negative in all experiments of the same event as negative, and mark the chemicals that are positive in at least one experiment as positive; for in vivo data, only select the chemicals with consistent results in all experiments.
[0019] Further, the training of the source model using a deep neural network and obtaining the target model after migration includes:
[0020] Process the SMILES of the chemical to generate MACCS molecular fingerprints, use a deep neural network to train the source model, use a single event as the source domain, and train the deep neural network source model through the TensorFlow architecture;
[0021] The deep neural network consists of 6 neural network layers and an output layer, uses sigmoid, tanh, elu, and relu as the activation functions of neurons in different network layers, optimizes the hyperparameters through cross-validation and grid search, determines the optimal hyperparameters according to the minimum loss function value, and obtains the target model after migration.
[0022] Further, evaluate the prediction performance of the target model with the area under the curve, F1 score, and balanced accuracy rate, including:
[0023] The area under the curve is the area under the ROC curve, which is used to reflect the discrimination ability of the target model for positive and negative samples. The ROC curve visualizes the trade-off between the true positive rate and the false positive rate. The area under the curve AUC is calculated based on the true positive rate and the false positive rate:
[0024] TPR = TP / (TP + FN)
[0025] FPR = FP / (FP + TN)
[0026] where TPR is the true positive rate, FPR is the false positive rate; TP is the true positive chemical sample, TN is the true negative chemical sample; FP is the false positive chemical sample, and FN is the false negative chemical sample;
[0027] The F1 score comprehensively considers precision and recall. Precision is used to measure the probability of true positive samples among the samples predicted as positive by the target model; Recall is used to measure the probability of correctly predicting true positive samples by the target model:
[0028] Precision = TP / (TP + FP)
[0029] Recall = TP / (TP + FN)
[0030] F1Score = 2 * (Precision * Recall) / (Precision + Recall)
[0031] where F1Score represents the F1 score;
[0032] The balanced accuracy rate is used to reflect the overall classification performance of the target model in the case of unbalanced dataset where the number of samples in each category is uneven:
[0033] BA = (TPR + TNR) / 2
[0034] where BA represents the balanced accuracy rate.
[0035] Furthermore, the application domain of the target model is characterized according to the index similarity density and weighted ruggedness, including:
[0036] Based on the chemicals in the training set of the step-by-step transfer learning model, the similarity density ρ s and the weighted ruggedness I A of the target model are determined:
[0037]
[0038]
[0039] where ρs,q is the weighted similarity density between the chemical q to be tested and the chemicals in the training set, I A,q is the weighted activity discontinuity between the chemical q to be tested and the chemicals in the training set; w q,t is the weight function of the chemical q to be tested and the chemical pair t in the training set; T represents the training set; S WD,t represents the weighted local discontinuity score of the chemical t in the training set; S M,q,t is the Tanimoto similarity coefficient between the chemical q to be tested and the chemical t in the training set calculated based on the MACCS molecular fingerprint; S cutoff is the threshold of the Tanimoto similarity coefficient;
[0040] Set the similarity density ρ s and the weighted ruggedness I A thresholds. When the chemical to be predicted meets the thresholds of the similarity density ρ s and the weighted ruggedness I A in the target model, it is considered that the chemical to be predicted is within the application domain of the target model.
[0041] Furthermore, predicting whether the chemicals within the application domain have reproductive toxicity to living organisms by using the target model includes:
[0042] Calculate the similarity density ρ s and the weighted ruggedness I A of the chemical to be predicted and the chemicals in the training set. When the chemical to be predicted is within the application domain, use the target model to predict the chemical to be predicted. When the output prediction value is greater than 0.5, the chemical to be predicted is identified as positive and has reproductive toxicity; when the output prediction value is less than 0.5, the chemical to be predicted is identified as negative and does not have reproductive toxicity.
[0043] Compared with the prior art, the present invention has the following advantages:
[0044] A prediction method for a chemical reproductive toxicity transfer learning model based on an adverse outcome pathway provided by the present invention collects test data of male and female reproductive toxicity of chemicals at different levels in the adverse outcome pathway and preprocesses the test data; randomly splits the preprocessed test data into a training set, a validation set, and a test set, trains a source model using a deep neural network, and obtains a target model after migration; evaluates the prediction performance of the target model using the area under the curve, F1 score, and balanced accuracy as indicators; characterizes the application domain of the target model according to the index similarity density and weighted ruggedness; uses the target model to predict whether there is reproductive toxicity of chemicals in the living body. The transfer learning model in the present invention fully integrates multi-level in vitro test toxicity data of chemicals from molecular initiating events to the generation of adverse outcomes, and can accurately predict in vivo reproductive toxicity. Compared with existing machine learning models, it not only has a significant improvement in accuracy, but also has better model robustness. In addition, the model of the present invention has a clear characterization of the application domain, ensuring the reliability of model application, and is expected to play an important role in high-throughput screening of chemical toxicity, providing technical support for national chemical regulation and new pollutant treatment.
[0045] For the above reasons, the present invention can be widely promoted in the fields of data processing and the like. Brief Description of the Drawings
[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0047] Figure 1 It is a flowchart of a prediction method for a chemical reproductive toxicity transfer learning model based on an adverse outcome pathway in the present invention.
[0048] Figure 2 It is a schematic diagram of the relationship between the area under the curve AUC and the true positive rate and false positive rate of the ROC curve in the present invention.
[0049] Figure 3 It is the performance of the transfer learning model created in the embodiment of the present invention on the training set and the external test set. Detailed Embodiments
[0050] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other. The following will refer to the drawings and combine the embodiments to detail the present invention.
[0051] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only a part rather than all of the embodiments of the present invention. The following description of at least one exemplary embodiment is actually illustrative only and in no way limits the present invention and its application or use. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0052] It should be noted that the terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the exemplary embodiments of the present invention. As used herein, unless the context clearly dictates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they specify the presence of the stated features, steps, operations, devices, components, and / or combinations thereof.
[0053] Unless otherwise specifically stated, the relative arrangements, numerical expressions and values of the components and steps set forth in these embodiments do not limit the scope of the present invention. At the same time, it should be clear that, for the sake of convenience of description, the dimensions of the various parts shown in the drawings are not drawn in actual proportional relationship. Technologies, methods and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the said technologies, methods and devices should be regarded as part of the authorized specification. In all the examples shown and discussed here, any specific values should be construed as merely exemplary and not as a limitation. Therefore, other examples of the exemplary embodiments may have different values. It should be noted that like reference numerals and letters denote like items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.
[0054] In the description of the present invention, it should be understood that the orientation terms such as "front, rear, upper, lower, left, right", "lateral, vertical, perpendicular, horizontal" and "top, bottom", etc. generally refer to the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description. Without contrary description, these orientation terms do not indicate and imply that the device or element referred to must have a specific orientation or be constructed and operated in a specific orientation, and thus should not be construed as limiting the protection scope of the present invention. The orientation terms "inside, outside" refer to the inside and outside relative to the contour of each component itself.
[0055] For ease of description, spatial relative terms, such as "above", "over", "on the upper surface", "upper", etc., may be used herein to describe the spatial positional relationship of one device or feature to other devices or features as shown in the figures. It should be understood that spatial relative terms are intended to encompass different orientations in use or operation in addition to the orientation depicted in the figures. For example, if the device in the figures is inverted, a device described as "above" or "over" other devices or structures will then be positioned "below" or "under" the other devices or structures. Thus, the exemplary term "above" can include both orientations of "above" and "below". The device may also be positioned in other different ways (rotated 90 degrees or in other orientations), and the corresponding explanations for the spatial relative descriptions used herein will be made accordingly.
[0056] In addition, it should be noted that the use of terms such as "first", "second", etc. to limit components is only for the convenience of differentiating the corresponding components. Without additional statements, the above terms have no special meanings, and thus should not be construed as limiting the protection scope of the present invention.
[0057] As Figure 1 shown, the present invention provides a prediction method for a chemical reproductive toxicity transfer learning model based on an adverse outcome pathway, including:
[0058] Collecting test data of male and female reproductive toxicities of chemicals at different levels in the adverse outcome pathway, and preprocessing the test data;
[0059] In specific implementation, as a preferred implementation manner of the present invention, the test data includes an adverse outcome pathway of androgen-mediated male reproductive toxicity and an adverse outcome pathway of estrogen-mediated female reproductive toxicity;
[0060] The adverse outcome pathway of androgen-mediated male reproductive toxicity (AR-AOP) includes: androgen binding (MIE), co-regulator recruitment (KE1), chromatin binding (KE2), transcription factor activation (KE3), up-regulation of gene expression (KE4), cell proliferation (KE5), and increase in sex organ weight (AO); the adverse outcome pathway of estrogen-mediated female reproductive toxicity (ERα-AOP) includes: estrogen binding (MIE), estrogen dimerization (KE1), chromatin binding (KE2), transcription factor activation (KE3), up-regulation of gene expression (KE4), cell proliferation (KE5), and increase in sex organ weight (AO).
[0061] In implementation, the in vitro test data (MIE-KE5) is sourced from the CompTox Chemicals Dashboard database, and the AO data comes from the rodent Hershberger database, the Rodent Uterotrophic Bioactivity database, and the ToxRefDB database.
[0062] Specifically in implementation, as a preferred implementation manner of the present invention, the preprocessing of the test data includes:
[0063] Generate the SMILES of each chemical through the PubChem database and standardize it, removing metal salts, inorganic substances, and mixtures. This process is completed in the Python 3.8 environment; for in vitro data, chemicals that are negative in different experiments of the same event are marked as negative, and chemicals that are positive in at least one experiment are marked as positive; for in vivo data, only chemicals with consistent results in all experiments are selected. In implementation, the positive label is 1 and the negative label is 0.
[0064] Randomly split the preprocessed test data into a training set, a validation set, and a test set, and use a deep neural network to train the source model, and obtain the target model after migration;
[0065] Specifically in implementation, as a preferred implementation manner of the present invention, the use of a deep neural network to train the source model and obtain the target model after migration includes:
[0066] Use the RDKit toolkit in Python 3.8 to process the SMILES of chemicals to generate MACCS molecular fingerprints, which altogether include 166-bit binary molecular feature representations. Use the deep neural network DNN to train the source model, use a single event as the source domain, and train the deep neural network source model through the TensorFlow architecture;
[0067] In implementation, the DNN is a multi-layer artificial neural network. By learning data features layer by layer, it can effectively capture complex non-linear relationships. Randomly split the collected reproductive toxicity AOP hierarchical test data sets into a training set, a validation set, and a test set according to the ratio of 7:1:2, which are respectively used for model training, hyperparameter selection, and external prediction ability evaluation.
[0068] The deep neural network consists of 6 neural network layers and an output layer. Use sigmoid, tanh, elu, and relu as the activation functions of neurons in different network layers, and optimize the hyperparameters through 10-fold cross-validation and grid search. The training step size batch_size is 32, and the learning rate of the neural network is 0.0001. Determine the optimal hyperparameters according to the minimum loss function value, and obtain the target model after migration.
[0069] In implementation, transfer learning is based on the MIE source model. By freezing network layers layer by layer, that is, setting the parameters of the network layer to "Frozen", and fine-tuning on the target domain dataset (such as KE1, KE2, etc.). For the first transfer, the first network layer is frozen, and the model is optimized with KE1 as the target domain; subsequently, the optimized model is used as the new source model, the first two network layers are frozen, and further fine-tuning is performed with KE2 as the target domain. Transfer in this way successively, gradually passing the toxicity knowledge of each level of AOP along the neural network, and finally transferring it to the target domain AO.
[0070] The prediction performance of the target model is evaluated using the area under the curve, F1 score, and balanced accuracy as metrics;
[0071] Specifically in implementation, as a preferred implementation manner of the present invention, evaluating the prediction performance of the target model using the area under the curve, F1 score, and balanced accuracy as metrics includes:
[0072] The area under the curve is the area under the ROC curve, which is a comprehensive metric for evaluating the performance of a binary classification model and is usually used to measure the accuracy of model classification. In the present invention, it is used to reflect the discrimination ability of the target model for positive and negative samples, and its value is between 0.5 and 1. The closer the AUC value of the area under the curve is to 1, the better the performance of the model. The ROC curve visualizes the trade-off between the true positive rate and the false positive rate; the area under the curve AUC is calculated based on the true positive rate and the false positive rate:
[0073] TPR = TP / (TP + FN)
[0074] FPR = FP / (FP + TN)
[0075] Wherein, TPR is the true positive rate, and FPR is the false positive rate; TP is the true positive chemical sample, TN is the true negative chemical sample; FP is the false positive chemical sample, and FN is the false negative chemical sample;
[0076] The F1 score is a commonly used metric for evaluating the performance of a classification model, which comprehensively considers precision and recall. The precision Precision is used to measure the probability of true positive cases among the samples predicted as positive by the target model; the recall Recall is used to measure the probability of correctly predicting true positive cases by the target model:
[0077] Precision = TP / (TP + FP)
[0078] Recall = TP / (TP + FN)
[0079] F1Score = 2 * (Precision * Recall) / (Precision + Recall)
[0080] Among them, F1Score represents the F1 score; the closer the value of F1Score is to 1, the better the performance of the model. F1Score combines the accuracy and comprehensiveness of the model and is often applicable to the evaluation of models with data imbalance.
[0081] The balanced accuracy is a comprehensive index for evaluating the performance of a classification model in an imbalanced class dataset. By averaging the accuracy of each class, it can provide a more comprehensive measure of the model's performance. The value range of the balanced accuracy is also from 0 to 1, the same as other accuracy metrics. In the present invention, the balanced accuracy is used to reflect the overall classification performance of the target model in the case of uneven sample numbers of each class in the imbalanced dataset:
[0082] BA = (TPR + TNR) / 2
[0083] Among them, BA represents the balanced accuracy.
[0084] Characterize the application domain of the target model according to the index similarity density and weighted ruggedness;
[0085] Specifically in implementation, as a preferred implementation manner of the present invention, the characterizing the application domain of the target model according to the index similarity density and weighted ruggedness includes:
[0086] Based on the chemicals in the step-by-step transfer learning model training set, determine the similarity density ρ s and weighted ruggedness I A :
[0087]
[0088]
[0089] Among them, ρ s,q is the weighted similarity density between the chemical q to be measured and the chemicals in the training set, and the larger its value, the higher the similarity; I A,q is the weighted activity discontinuity between the chemical q to be measured and the chemicals in the training set; w q,t is the weight function of the chemical q to be measured and the chemical pair t in the training set; T represents the training set; S WD,t represents the weighted local discontinuity score of the chemical t in the training set; S M,q,t is the Tanimoto similarity coefficient between the chemical q to be measured and the chemical t in the training set calculated based on the MACCS molecular fingerprint; S cutoff is the threshold of the Tanimoto similarity coefficient, taking 0.65;
[0090] Set the similarity density ρ s and the weighted ruggedness I A thresholds. When the chemical to be predicted meets the similarity density ρ s and the weighted ruggedness I A threshold ranges set in the target model, then it is considered that the chemical to be predicted is within the application domain of the target model.
[0091] During implementation, set the threshold ρ s ≥15 and I A ≤0.14, then it is considered to be within the model application domain, otherwise it is outside the application domain.
[0092] Use the target model to predict whether the chemicals within the application domain have reproductive toxicity to living organisms.
[0093] Specifically in implementation, as a preferred implementation manner of the present invention, the use of the target model to predict whether the chemicals within the application domain have reproductive toxicity to living organisms includes:
[0094] Calculate the similarity density ρ s and the weighted ruggedness I A of the chemical to be predicted and the chemicals in the training set. When the chemical to be predicted is within the application domain, use the target model to predict the chemical to be predicted. When the output prediction value is greater than 0.5, the chemical to be predicted is identified as positive and has reproductive toxicity; when the output prediction value is less than 0.5, the chemical to be predicted is identified as negative and does not have reproductive toxicity.
[0095] Example 1
[0096] As Figure 1 shown, the present invention provides a method for predicting the reproductive toxicity of chemicals based on the adverse outcome pathway migration learning model. In this example, the given chemical nandrolone (CAS No.: 10161 - 33 - 8) is used as the research object to screen whether it has male reproductive toxicity.
[0097] First, according to the SMILES code of nandrolone, in the Python 3.8 environment, use the RDKit toolkit to generate MACCS fingerprints, and then calculate ρ s and I A between this fingerprint and the chemicals in the training set. The result is ρ s = 17, I A = 0.12, and it is determined that this substance is within the application domain. Use the migration learning model created by the present invention for prediction, and the result is: the prediction value is 0.842 (greater than 0.5 is identified as positive, otherwise negative), and the corresponding experimental value is positive, and the prediction value is consistent with the experimental value.
[0098] Example 2
[0099] In this embodiment, the given chemical 4-dodecylphenol (CAS No.: 1478-61-1) is taken as the research object to screen whether it has female reproductive toxicity. First, according to the SMILES code of 4-dodecylphenol, in the Python 3.8 environment, the MACCS fingerprint is generated using the RDKit toolkit, and then the fingerprint is used to calculate ρ with the training set chemicals s and I A . The result is ρ s = 18, I A = 0.10, and it is determined that this substance is within the application domain. The migration learning model created by the present invention is used for prediction, and the result is obtained: the predicted value is 0.986, the corresponding experimental value is positive, and the predicted value is consistent with the experimental value.
[0100] Example 3
[0101] In this embodiment, the given chemical benzophenone (CAS No.: 119-61-9) is taken as the research object to screen whether it has male reproductive toxicity. First, according to the SMILES code of benzophenone, in the Python 3.8 environment, the MACCS fingerprint is generated using the RDKit toolkit, and then the fingerprint is used to calculate ρ s and I A . The result is ρ s = 6, I A = 0.24, and it is determined that this substance is not within the application domain. The migration learning model created by the present invention is used for prediction, and the result is obtained: the predicted value is 0.287, the corresponding experimental value is positive, and the predicted value is not consistent with the experimental value. Since the substance is not within the application domain, the credibility of the prediction result is relatively low.
[0102] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A prediction method for a chemical reproductive toxicity transfer learning model based on an adverse outcome pathway, characterized in that, Including: Collecting the test data of the male and female reproductive toxicity of chemicals at different levels in the adverse outcome pathway, and preprocessing the test data; Randomly splitting the preprocessed test data into a training set, a validation set and a test set, training the source model using a deep neural network, and obtaining the target model after migration; Evaluating the prediction performance of the target model with the area under the curve, F1 score and balanced accuracy as indicators; Characterizing the application domain of the target model according to the index similarity density and weighted ruggedness; Using the target model to predict whether the chemicals in the application domain have reproductive toxicity to living organisms.
2. The method for predicting a chemical reproductive toxicity transfer learning model based on an adverse outcome pathway according to claim 1, wherein The test data includes the adverse outcome pathway of androgen-mediated male reproductive toxicity and the adverse outcome pathway of estrogen-mediated female reproductive toxicity; The adverse outcome pathway of androgen-mediated male reproductive toxicity includes: androgen binding, co-regulator recruitment, chromatin binding, transcription factor activation, up-regulation of gene expression, cell proliferation and increase in sex organ weight; The adverse outcome pathway of estrogen-mediated female reproductive toxicity includes: estrogen binding, estrogen dimerization, chromatin binding, transcription factor activation, up-regulation of gene expression, cell proliferation and increase in sex organ weight.
3. The method for predicting a chemical reproductive toxicity transfer learning model based on an adverse outcome pathway according to claim 1, wherein The preprocessing of the test data includes: Generating the SMILES of each chemical through the PubChem database, and standardizing it to remove metal salts, inorganic substances and mixtures; for in vitro data, chemicals that are negative in all experiments of the same event are marked as negative, and chemicals that are positive in at least one experiment are marked as positive; for in vivo data, only chemicals with consistent results in all experiments are selected.
4. The prediction method of the chemical reproductive toxicity transfer learning model based on the adverse outcome pathway according to claim 1, wherein The training of the source model using a deep neural network and obtaining the target model after migration includes: Processing the SMILES of the chemical to generate MACCS molecular fingerprints, training the source model using a deep neural network, using a single event as the source domain, and training the deep neural network source model through the TensorFlow architecture; The deep neural network consists of 6 neural network layers and an output layer, using sigmoid, tanh, elu and relu as the activation functions of neurons in different network layers, optimizing the hyperparameters through cross-validation and grid search, determining the optimal hyperparameters according to the minimum loss function value, and obtaining the target model after migration.
5. The method for predicting a chemical reproductive toxicity transfer learning model based on an adverse outcome pathway according to claim 1, wherein The evaluation of the prediction performance of the target model with the area under the curve, F1 score and balanced accuracy as indicators includes: The area under the curve is the area under the ROC curve, which is used to reflect the discrimination of the target model for positive and negative samples. The ROC curve visualizes the trade-off between the true positive rate and the false positive rate; calculating the area under the curve AUC according to the true positive rate and the false positive rate: TPR = TP / (TP + FN) FPR = FP / (FP + TN) Where, TPR is the true positive rate, FPR is the false positive rate; TP is the true positive chemical sample, TN is the true negative chemical sample; FP is the false positive chemical sample, FN is the false negative chemical sample; The F1 score comprehensively considers precision and recall. The precision is used to measure the probability of true positive examples among the samples predicted as positive examples by the target model; the recall is used to measure the probability of correctly predicting true positive examples by the target model: Precision = TP / (TP + FP) Recall = TP / (TP + FN) F1Score = 2 * (Precision * Recall) / (Precision + Recall) where F1Score represents the F1 score; The balanced accuracy is used to reflect the overall classification performance of the target model in the case of unbalanced sample numbers of various categories in an imbalanced dataset: BA = (TPR + TNR) / 2 where BA represents the balanced accuracy.
6. The prediction method of the chemical reproductive toxicity transfer learning model based on the adverse outcome pathway according to claim 1, wherein Characterizing the application domain of the target model according to the index similarity density and weighted ruggedness includes: Based on the chemicals in the stepwise transfer learning model training set, determine the similarity density ρ of the target model s and the weighted ruggedness I A : Among them, ρ s,q is the weighted similarity density between the chemical q to be measured and the chemicals in the training set, and I A,q is the weighted activity discontinuity between the chemical q to be measured and the chemicals in the training set; w q,t is the weight function of the pair of the chemical q to be measured and the training set chemical t; T represents the training set; S WD,t represents the weighted local discontinuity score of the training set chemical t; S M,q,t is the Tanimoto similarity coefficient between the chemical q to be measured and the training set chemical t calculated based on the MACCS molecular fingerprint; S cutoff is the threshold of the Tanimoto similarity coefficient; Set the similarity density ρ s and the weighted ruggedness I A thresholds. When the chemical to be predicted meets the similarity density ρ s and the weighted ruggedness I A threshold ranges set in the target model, it is considered that the chemical to be predicted is within the application domain of the target model.
7. The prediction method of the chemical reproductive toxicity transfer learning model based on the adverse outcome pathway according to claim 1, wherein Using the target model to predict whether there is reproductive toxicity of chemicals in the application domain to living organisms includes: Calculate the similarity density ρ between the chemical to be predicted and the chemicals in the training set s and the weighted ruggedness I A , when the chemical to be predicted is within the application domain, use the target model to predict the chemical to be predicted. When the output prediction value is greater than 0.5, the chemical to be predicted is identified as positive and has reproductive toxicity; when the output prediction value is less than 0.5, the chemical to be predicted is identified as negative and does not have reproductive toxicity.