A method and device for generative prediction of blood biochemical and blood routine indexes in rat chemical poisoning experiments based on multimodal fusion and GANs model
Through multimodal fusion and GANs model, the SMILES-form of chemicals is used to generate blood biochemical and blood conventional indicators of rat poisoning experiments, which solves the shortcomings of toxicity prediction of new synthetic chemicals, achieves efficient and accurate toxicity identification and management, and improves prediction accuracy and animal welfare.
Patent Information
- Application Number
- CN202411372807.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-29
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-09-29
AI Technical Summary
In the prior art, the number of new synthetic chemicals is much higher than the number of compounds that have been included in the toxicological effect evaluation, resulting in insufficient understanding of the impact of environmental pollutants and lack of effective toxicity experimental research strategies, affecting human health and ecological security.
Using a method based on multimodal fusion and GANs model, the SMILES formula of chemicals is obtained, molecular mechanism, biological activity and batch molecular docking activity data are calculated, and a pre-trained alternative rat experimental prediction model is input to generate predictor values of blood biochemistry and blood conventional indexes, and discretized through self-programming functions to predict the affected pathways.
It improves the prediction accuracy and efficiency of chemical toxicity identification, saves experimental time and animal resources, improves animal welfare, and provides rich chemical toxicity characteristics to support the management of new chemicals.
Smart Images

Figure CN119274662B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for generative prediction of blood biochemical and blood routine indexes in rat toxicology experiments on chemicals based on multi-modal fusion and GANs models, and also relates to a corresponding device for generative prediction of blood biochemical and blood routine indexes in rat toxicology experiments on chemicals based on multi-modal fusion and GANs models, belonging to the field of artificial intelligence-assisted prediction of chemical toxic effects, and belonging to the interdisciplinary research of computer science, toxicology, chemoinformatics, and bioinformatics. Background Art
[0002] There is no doubt that synthetic chemicals have significantly improved food production and living standards, but the current situation is that the number of newly synthesized chemicals is much higher than the number of chemicals that have been included in the evaluation of toxicological effects. According to the research of the European Commission on public data, even after a more careful inspection of 2,465 high-production chemicals, 21% of the compounds have no toxicological data, and the data of 65 compounds is insufficient. At the same time, with the progress of the measurement ability of low environmental concentration chemicals, people's understanding of a series of effects of environmental pollutants on organisms has been greatly improved, which means that the toxicological test research strategy should be adjusted to a certain extent. Chemical substances that may have adverse effects on human health and ecological safety without sufficient testing will inevitably enter the environment after production and use, and some of these substances exhibit the behavior of environmental organic pollutants.
[0003] In the past few decades, the synthesis and diversification of man-made chemicals have increased sharply, exceeding other well-known global challenge factors such as climate change. It is estimated that about 200-300 new chemicals are produced and sold in large quantities every year. Limited information based on laboratory toxicity data indicates that some emerging pollutants can have harmful effects on exposed humans and mammals, such as cytotoxicity, developmental toxicity, hepatotoxicity, and endocrine disruption. Relevant chemicals that deserve special attention include but are not limited to developmental neurotoxins, endocrine disruptors, chemical herbicides, new pesticides, pharmaceutical waste, and nanomaterials. These statistics emphasize the necessity of studying and clarifying the ecological impact and health hazards of chemical pollution. Exposure to environmental organic pollutants has caused significant ecological impacts and adverse health outcomes.
[0004] Based on artificial intelligence-based toxicity prediction models, the use of chemical databases, molecular descriptors, fingerprint maps, and model algorithms are all important factors in model development. With the development of information technology, multi-modal methods have entered the field of scientific research. Multi-modal deep learning has been proposed, and due to its high non-linearity, it has proven its advantages in representing multi-modal data. If multi-modal fusion can provide a new breakthrough for computational toxicology prediction models, it can promote the development of intelligent systems towards a more intelligent direction. Summary of the Invention
[0005] The primary technical problem to be solved by the present invention is to provide a method for generative prediction of blood biochemical and blood routine indexes in rat toxicology experiments on chemicals based on multimodal fusion and GANs models.
[0006] Another technical problem to be solved by the present invention is to provide a device for generative prediction of blood biochemical and blood routine indexes in rat toxicology experiments on chemicals based on multimodal fusion and GANs models.
[0007] To achieve the above object, the present invention adopts the following technical solutions:
[0008] According to the first aspect of the embodiments of the present invention, there is provided a generative prediction of blood biochemical and blood routine indexes in rat toxicology experiments on chemicals based on multimodal fusion and GANs models, including the following steps:
[0009] Step S1, obtain the SMILES formula of the chemical to be predicted;
[0010] Step S2, the device calculates the data of three modalities of the chemical, including molecular structure, biological activity, and batch key molecular docking activities;
[0011] Step S3, merge the calculated data of the three modalities with the dosing dose and dosing time in a linearly additive manner and input them into a pre-trained alternative rat experiment prediction model;
[0012] Step S4, the output result of the alternative rat experiment prediction model is the generative prediction values of 25 indexes of blood biochemistry and blood routine under the set dosing dose and set dosing time of the chemical from the data source;
[0013] Furthermore, the present invention can achieve a sub-classified prediction of the developmental toxicity of chemicals through the following steps:
[0014] Step S5, further call a self-written function to discretize the batch of key molecular docking data to obtain the genes corresponding to the variables with a value of 1;
[0015] Step S6, screen out the genes corresponding to the proteins with higher docking scores in the key molecular docking for annotation and function analysis to achieve the purpose of predicting the affected pathways, and output the effect diagram.
[0016] Preferably, when calculating the required data, the following processing is performed:
[0017] Step S21: Retrieve key genes related to human blood disease phenotypes from the HPO phenotype database, use Alpha Fold to predict a series of corresponding protein structure pdbqt files, and save them in the corresponding device after removing ligands, adding hydrogen, etc.;
[0018] Step S22: According to the SMILES formula of the chemical to be tested, use relevant libraries such as Python to calculate the 3D molecular structure of the chemical (the chemical structure can be characterized by a set of numerical values, which are called molecular fingerprints or descriptors. They may characterize the properties of the molecule, such as log P, molecular weight, hydrogen bond donors, acceptors, rotatable bonds, etc., and these properties can be related to the experimental evidence of the molecule; for each level of molecular characterization, hundreds or thousands of structural features can be calculated. There are various molecular descriptors and fingerprints, encoding structural, topological, geometric, electrostatic, quantum chemical, thermodynamic, fragment features, etc.), and abstractly extract thousands of features, which is the first modal data;
[0019] Step S23: According to the SMILES formula of the chemical to be tested, use relevant libraries such as Python to calculate the bioactivity data of the chemical (the bioactivity data used in the present invention is calculated by the Chemical Checker tool. Chemical Checker extends the principle of small molecule similarity to all levels of biology. CC classifies the data into five levels, with increasing complexity from the chemical properties of the compound to the clinical results. Express the bioactivity data in a general vector format to obtain the detailed information of the vectors in 25 CC spaces), and extract thousands of features, which is the second modal data;
[0020] Step S24: According to the SMILES formula of the chemical to be tested, call relevant libraries such as Python to calculate the 3D molecular structure of the chemical, and perform ligand removal and hydrogen addition treatment to generate a pbdqt file. Call the autodock vina program package for batch molecular docking to generate a docking activity sequence, with a total of about 1k features, which is the third modal data;
[0021] Step S25: Combine the data of the exposure dose and exposure time by linearly adding the three modalities according to the requirements of the established algorithm model.
[0022] Preferably, when calculating the first modal data:
[0023] Step 221: Save the modal features as a 50×50 matrix, fill all numerical values towards the center, and fill the edges with random noise.
[0024] Preferably, when calculating the second modal data:
[0025] Step 231, the modal feature is saved as a 25×128 matrix.
[0026] Preferably, when calculating the third modal data:
[0027] Step 241, the present invention pre-additionally trains an LDA topic model based on text data mining of the gene cards database for gene classification;
[0028] Step 242, according to the gene classification model in step 241, the data of the third modality is split into several sequences for storage.
[0029] Preferably, for the three modalities of data calculated, we perform the following processing:
[0030] The original feature data obtained by calculation in this study is sequentially subjected to outlier processing, recoding, standardization, and resampling.
[0031] Preferably, in S3, the combination of the three modalities of data with the dosing dose and dosing time specifically includes: the data is further standardized between 0 and 1 to prevent errors caused by dimensional differences between different variables; the three modalities of data are trained by separate networks and further sent to the subsequent network; the data of the dosing dose and dosing time is directly linearly added to the subsequent network through a fully connected layer.
[0032] Preferably, the alternative rat experiment prediction model is obtained through the following steps:
[0033] Step S31, obtain multiple groups of chemical SMILES numbers and 25 blood biochemical and blood routine indexes, and calculate the three modalities of data for each group of data according to the above method. And integrate these data into the corresponding format for modeling.
[0034] Step S32, obtain multiple groups of rat experiment data to train a pre-designed adversarial generative neural network (GANs) model structure, and further perform parameter tuning and optimization to obtain the optimal alternative rat experiment prediction model.
[0035] Preferably, the alternative rat experiment prediction model adopts a multi-modal fusion structure and uses a Dropout layer for regularization. The model complexity is reduced by adjusting the number of nodes in the intermediate layer. The model adopts an adversarial generative neural network, including a generator and a discriminator, which are jointly trained in an adversarial manner to achieve the effect of generating realistic data:
[0036] After the first modality data is input, the structure of a convolutional neural network (CNN) is used;
[0037] After the second modality data is input, the structure of a convolutional neural network (CNN) is used;
[0038] After the third modal data is input, a structure using the attention mechanism (Transformer) is adopted.
[0039] The above three modalities combine the dosing dose and dosing time and input them into the neural network. The model adopts an adversarial generative neural network, including a generator and a discriminator, which are jointly trained in an adversarial manner to achieve the effect of generating realistic data.
[0040] Preferably, after the alternative rat experiment prediction model makes a prediction, it automatically calculates the distribution composition of each generated prediction index. Considering laboratory errors, when the sum of the proportions of white blood cells (i.e., neutrophils, lymphocytes, monocytes, eosinophils, basophils) exceeds 105%, the data generated by the generator is identified as low-quality data.
[0041] Preferably, after the alternative rat experiment prediction model makes a prediction, there are subsequent processing and functions for predicting affected pathways:
[0042] Step S51: After the index prediction is completed, the data of the third modality is automatically discretized according to a set threshold, and the corresponding gene sequence is returned.
[0043] Step S61: According to the gene sequence returned in S51, enrichment analysis is performed to predict the pathways that may be affected.
[0044] According to the second aspect of the embodiments of the present invention, there is provided a device for generative prediction of blood biochemical and blood routine indexes in a rat chemical poisoning experiment based on multimodal fusion and GANs model, including a processor and a memory. The processor reads the computer program in the memory, and the result is displayed on a display for performing the following operations:
[0045] Obtain the data to be collected, which only includes the SMILES formula, dosing dose, and dosing time of the chemical to be predicted.
[0046] Input the SMILES into the receiving interface of the entire device;
[0047] The output results of the alternative rat experiment prediction model are 25 blood biochemical and blood routine indexes, and the results of enrichment analysis predicting affected pathways are displayed on the display.
[0048] Advantages of the present invention
[0049] The method and device for generative prediction of blood biochemical and blood routine indexes in rat chemical poisoning experiments based on multimodal fusion and GANs models provided by the present invention only use the SMILES formula of the chemical, the preset poisoning dose and poisoning time to provide rich relevant features for predicting the blood biochemical and blood routine indexes after rat chemical poisoning experiments, greatly improving the prediction accuracy and efficiency. On the one hand, it can effectively assist in the toxicity identification and management of new chemicals, on the other hand, it effectively saves the time and energy of experimenters, and at the same time saves animal resources and improves animal welfare. Brief Description of the Drawings
[0050] Figure 1 It is a schematic flowchart of the method for generative prediction of blood biochemical and blood routine indexes in rat chemical poisoning experiments based on multimodal fusion and GANs models provided by an embodiment of the present invention.
[0051] Figure 2 It is the technical route and data flow of in vitro bioactivity calculation provided by an embodiment of the present invention
[0052] Figure 3 It is the technical route and data flow of molecular descriptor calculation provided by an embodiment of the present invention
[0053] Figure 4 It is the technical route and data flow of batch molecular docking calculation provided by an embodiment of the present invention
[0054] Figure 5 It is the module design and block function distribution provided by an embodiment of the present invention
[0055] Figure 6 It is the t-SNE dimensionality reduction clustering visualization diagram of the built-in training data provided by an embodiment of the present invention
[0056] Figure 7 It is the training curve and loss function curve of the built-in model provided by an embodiment of the present invention
[0057] Figure 8 It is a schematic structural diagram of the device for generative prediction of blood biochemical and blood routine indexes in rat chemical poisoning experiments based on multimodal fusion and GANs models provided by an embodiment of the present invention Detailed Embodiments
[0058] The present invention will be further described below in conjunction with embodiments, but the protection scope of the present invention is not limited thereto:
[0059] For emerging chemicals, in order to provide toxicity prediction values for identification and management as early as possible, such as Figure 1As shown in the figure, the generative prediction method for blood biochemical and blood routine indexes of rat chemical poisoning experiments based on multi-modal fusion and GANs model provided by the embodiments of the present invention includes the following steps:
[0060] Step S1, obtain the SMILES formula of the chemical to be predicted;
[0061] Step S2, the device calls internal files to calculate the data of three modalities of the chemical, including molecular structure, biological activity, and batch key molecular docking activities (such as Figure 5 shown in - module 2a);
[0062] When calculating the data of three modalities required by the model, after the chemical SMILES is passed in, the following steps are included:
[0063] Step S21, retrieve the key genes related to human blood disease phenotypes through the HPO phenotype database, use Alpha Fold to predict a series of corresponding protein structure pdbqt files, and save them in the corresponding device after removing ligands, adding hydrogen, etc.;
[0064] Step S22, according to the SMILES formula of the chemical to be tested, use relevant libraries such as Python to calculate the 3D molecular structure of the chemical, and abstract and extract thousands of features (preferably refer to https: / / www.rdkit.org / docs / index.html for implementation), which is the first modality data (the method is as Figure 2 shown);
[0065] Step 221, save the modality features as a 50×50 matrix, fill all values towards the center, and fill the edges with random noise.
[0066] Step S23, according to the SMILES formula of the chemical to be tested, use relevant libraries such as Python to calculate the biological activity data of the chemical, and extract thousands of features, which is the second modality data (the method is as Figure 3 shown, and the method technology comes from the literature Nat Biotechnol. 2020; 38(9): 1087-1096);
[0067] Step 231, save the modality features as a 25×128 matrix.
[0068] Step 231, the present invention has pre-trained an additional LDA topic model based on text data mining of the gene cards database for gene classification;
[0069] Step S24: According to the SMILES formula of the chemical substance to be measured, call relevant libraries such as Python to calculate the 3D molecular structure of the chemical substance, and perform ligand removal and hydrogen addition treatment to generate a pbdqt file. Call the autodock vina program package for batch molecular docking to generate a docking activity sequence, with a total of about 1k features, which is the third modal data (the method is as Figure 4 shown);
[0070] Step 241: The present invention has pre-trained an LDA topic model based on text data mining of the gene cards database for gene classification (where the LDA topic model is based on Bayesian statistical theory and works by constructing a generative model with a three-layer structure (document - topic - vocabulary); its principle includes two parts: the document generation process (first, randomly select a series of topics for each document; then, for each vocabulary in the document, randomly select a topic according to the topic distribution at the current position; finally, randomly generate a vocabulary according to the vocabulary distribution under the selected topic) and model training (the LDA model adjusts the parameters of the topic - vocabulary distribution and the document - topic distribution to maximize the probability of generating the document. This is usually achieved using algorithms such as Gibbs Sampling or Variational Inference));
[0071] Step 242: According to the gene classification model in Step 241, split the data of the third modality into several sequences for storage.
[0072] Using the method for obtaining the three modal data of the chemical substance calculated by Step S2, obtain multiple groups of chemical substance SMILES numbers and 25 indicators of blood biochemistry and blood routine. For each group of data, calculate the three modal data according to the above method. And integrate these data into the corresponding format for modeling.
[0073] Step S3: Combine the calculated three modal data, the exposure dose, and the exposure time in a linearly additive manner and input them into a pre-trained alternative rat experiment prediction model;
[0074] Such as Figure 5As shown in Module 2b, the pre-designed alternative rat experiment prediction model adopts a multi-modal fusion method and uses the Dropout layer for regularization. The model complexity is reduced by adjusting the number of nodes in the intermediate layer. The last Dense layer outputs a probability using the sigmoid activation function for binary classification problems, where: after the first modal data is input, the structure of a convolutional neural network (CNN) is used; after the second modal data is input, the structure of a convolutional neural network (CNN) is used; after the third modal data is input, the structure of an attention mechanism (Transformer) is used; the training data of the entire model is visualized after t-SNE dimensionality reduction as Figure 6 shown, where the OPEN TG GATES database is the training set and the DrugMatrix database is the external validation set; Figure 7 It shows the changes in the accuracy and loss function during the training of the entire project with the training batches, revealing good convergence of this study.
[0075] Step S4, as Figure 5 As shown in Module 2c, the result of the alternative rat experiment prediction model is the generative prediction values of 25 indicators of blood biochemistry and blood routine at the set dosing dose and set dosing time of the chemical substance from this data source.
[0076] Step S5, further call the self-written function to discretize a batch of key molecular docking data and obtain the genes corresponding to the variables with a value of 1;
[0077] Step S6, screen out the genes corresponding to the proteins with higher docking scores in the key molecular docking for annotation and function analysis to achieve the purpose of predicting the affected pathways and output the effect diagram.
[0078] Merge the collected chemical SMILES formulas obtained in Step S1 with the set dosing time and dosing dose and input them into the receiving interface of the entire device, and the model will output the results of 25 generative prediction indicators of blood biochemistry and blood routine. After the alternative rat experiment prediction model makes a prediction, it automatically calculates the distribution composition of each generated prediction indicator. Considering laboratory errors, when the sum of the proportions of white blood cells (i.e., neutrophils, lymphocytes, monocytes, eosinophils, basophils) exceeds 105%, the data generated by the generator is identified as low-quality data.
[0079] To implement the method for generative prediction of blood biochemistry and blood routine indicators in rat poisoning experiments on chemical substances based on multi-modal fusion and GANs models provided by the present invention, the present invention also provides a device for generative prediction of blood biochemistry and blood routine indicators in rat poisoning experiments on chemical substances based on multi-modal fusion and GANs models. As Figure 8As shown in the figure, the device for generative prediction of blood biochemical and blood routine indexes of rats in chemical poisoning experiments based on multimodal fusion and GANs models includes a memory 31, a processor 32, and a display 33. The processor 32 reads the computer program in the memory 31 to enable the prediction device to execute the prediction model, and the results are displayed on the display 33.
[0080] The specific embodiments described herein are merely illustrative of the spirit of the present invention. Those skilled in the art to which the present invention pertains can make various modifications or supplements to the described specific embodiments or use similar methods for substitution, but will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.
Claims
1. A method for generative prediction of blood biochemical and blood routine indexes in rat toxicology experiments of chemicals based on multimodal fusion and GANs model, characterized in that It includes the following steps: Step S1: Obtain the SMILES formula of the chemical substance to be predicted; Step S2: Obtain three-modal data of the chemical substance, including molecular structure, biological activity, and batch key molecular docking activities; obtaining three-modal data of the chemical substance to be predicted includes the following steps: Step S21: Retrieve genes related to human blood disease phenotypes through the HPO phenotype database, use AlphaFold to predict a series of corresponding pdbqt files of protein structures, and save them for subsequent calls after removing ligands and hydrogenation treatment; Step S22: Calculate the 3D molecular structure of the chemical substance according to the SMILES formula of the chemical substance to be predicted, and abstract and extract features, which is the first-modal data; Step S23: Calculate the biological activity data of the chemical substance according to the SMILES formula of the chemical substance to be predicted, and extract features, which is the second-modal data; Step S24: Calculate the 3D molecular structure of the chemical substance according to the SMILES formula of the chemical substance to be predicted, and perform ligand removal and hydrogenation treatment to generate a pbdqt file; Call the pdbqt file of the protein macromolecule in Step S21 to perform batch molecular docking together to generate a docking activity sequence, which is the third-modal data; Step S3: Combine the three-modal data with the dosing dose and dosing time, and input them into a pre-trained prediction model to replace the rat experiment; Step S4: The output result of the prediction model to replace the rat experiment is the generative prediction value of 25 indicators of blood biochemistry and blood routine under the set dosing dose and set dosing time of the chemical substance from this data source.
2. The method according to claim 1, wherein: When obtaining the first-modal data: Step 221: Save the modal features as a 50×50 matrix, fill all values towards the center, and fill the edges with random noise; When obtaining the second-modal data: Step 231: Save the modal features as a 25×128 matrix; When obtaining the third-modal data: Step 241: Perform gene classification based on the LDA topic model for text data mining of the gene cards database; the LDA topic model is obtained by the following steps: Ⅰ Randomly select a series of topics for the description documents of each gene; Ⅱ For each vocabulary in the document, randomly select a topic according to the topic distribution at the current position; Ⅲ Randomly generate a vocabulary according to the vocabulary distribution under the selected topic; Ⅳ Use the Gibbs sampling or variational inference algorithm to adjust the parameters of the topic-vocabulary distribution and the document-topic distribution to maximize the probability of generating the document, thereby obtaining the LDA model; Step 242: According to the LDA topic model in Step 241, split the data of the third modal into several sequences for storage.
3. The method according to claim 1, wherein: For the original feature data of Steps S22, S23, and S24, outlier processing, re-encoding, standardization, and resampling are performed in sequence.
4. The method according to claim 1, characterized in that In step S3, the combined exposure dose and exposure time of the three modal data specifically include: the data is further normalized to between 0 and 1; after the three modal data are respectively subjected to feature extraction by separate feature extraction networks, they are further transmitted to the downstream network; the data of the exposure dose and exposure time are directly added to the subsequent network; The alternative rat experiment prediction model is obtained through the following steps: Step S31: Obtain multiple groups of chemical SMILES formulas and 25 blood biochemical and blood routine indexes, and for each group of data, three modal data are obtained according to step S2; Step S32: Divide the multiple groups of data obtained in S31 into a training set and a test set. Using the three modal data as the input and the corresponding 25 blood biochemical and blood routine indexes as the output, train the pre-designed adversarial generative neural network GANs model structure, and further perform parameter tuning and optimization to obtain the optimal alternative rat experiment prediction model.
5. The method according to claim 1 or 4, wherein: The pre-designed adversarial generative neural network GANs model structure of the alternative rat experiment prediction model adopts a multi-modal fusion structure and uses a Dropout layer for regularization; the model complexity is reduced by adjusting the number of nodes in the intermediate layer; the model adopts an adversarial generative neural network, including a generator and a discriminator, which are jointly trained in an adversarial manner to achieve the effect of generating realistic data: After the first modal data is input, the structure of a convolutional neural network CNN is used; After the second modal data is input, the structure of a convolutional neural network CNN is used; After the third modal data is input, the structure of an attention mechanism Transformer is used; The combined exposure dose and exposure time of the three modal data are input into the generator.
6. The method according to claim 1 or 4, wherein: After the alternative rat experiment prediction model makes a prediction, it automatically calculates the distribution composition of each generated prediction index. When the sum of the white blood cell proportions exceeds 105%, the data generated by the generator is identified as low-quality data and discarded.
7. The method according to claim 1, wherein It further includes: Step S5: Discretize the key molecular docking data of the batch of the chemical substance to obtain the gene sequence corresponding to the variable with a value of 1; Step S6: According to the gene sequence returned by S5, perform GO or KEGG annotation on the obtained related genes, conduct enrichment analysis. The more genes enriched in a certain pathway, the more likely this pathway is to be affected, and a detailed prediction is made for the possibly affected pathways.
8. A generative prediction device for blood biochemical and blood routine indexes in rat chemical poisoning experiments based on multimodal fusion and GANs model, characterized in that It includes a processor, a memory, and a display. The processor reads the computer program in the memory to enable the prediction device to execute the prediction method according to any one of claims 1-7, and the result is displayed on the display.