Drug toxicity prediction method, device, equipment and readable storage medium

By converting drug molecular formulas into drug molecular graphs and using generative-adversarial neural networks to learn toxic functional group features, a hypergraph model is constructed to predict drug toxicity. This solves the problems of lack of interpretability and accuracy in existing technologies and achieves more comprehensive toxicity prediction.

CN119314588BActive Publication Date: 2025-10-28UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411335449.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-24
Publication Date
2025-10-28
Estimated Expiration
2044-09-24

AI Technical Summary

Technical Problem

Existing technologies lack interpretability in drug toxicity prediction, making it difficult to clearly understand how the model uses input features for prediction, and they ignore the interactions between molecular functional groups, affecting the accuracy and completeness of drug toxicity prediction.

Method used

The drug molecular formula is converted into a drug molecular graph. Generative-adversarial neural networks are used to learn the features of toxic functional groups. By calculating the influence of functional groups on drug toxicity, a hypergraph model is constructed for prediction, providing interpretability and accuracy.

Benefits of technology

This improves the accuracy and completeness of drug toxicity prediction by clarifying the importance and influence of functional groups in drug toxicity, and provides interpretability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119314588B_ABST
    Figure CN119314588B_ABST
Patent Text Reader

Abstract

The present invention relates to a field, and particularly to a drug toxicity prediction method, apparatus, device, and readable storage medium. The method comprises obtaining a drug data set; converting a drug molecular formula included in the drug data set into a corresponding drug molecular graph to obtain first drug molecular graph information; sending the first drug molecular graph information to a trained neural network to obtain second drug molecular graph information; performing calculations based on the first drug molecular graph information and the second drug molecular graph information to obtain first score information, the first score information being used to indicate the degree of influence of functional groups on drug toxicity; and predicting drug toxicity based on the first score information. The present invention characterizes drug molecules through hypergraph technology, more comprehensively captures the high-level connection features between functional groups in drug molecules, and more accurately captures the structural features of the molecules, thereby improving the accuracy and interpretability of toxicity prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of drug toxicity prediction, and more specifically, to a method, apparatus, device, and readable storage medium for predicting drug toxicity. Background Technology

[0002] In drug development, predicting drug toxicity is a crucial step because toxicity assessment is directly related to drug safety. Existing technologies typically consider the correlation between multiple chemical toxicities and share important features between various toxicities during model training to discover potential associations between different toxicity outcomes, thereby enabling the prediction of drug toxicity. However, this approach not only lacks interpretability but also easily overlooks the interactions between molecular functional groups, affecting the accuracy and completeness of drug toxicity prediction. Summary of the Invention

[0003] The purpose of this invention is to provide a method, apparatus, device, and readable storage medium for predicting drug toxicity, in order to improve the above-mentioned problems.

[0004] To achieve the above objectives, the embodiments of this application provide the following technical solutions:

[0005] On the one hand, embodiments of this application provide a method for predicting drug toxicity, the method comprising:

[0006] Obtain drug datasets;

[0007] The drug molecular formulas included in the drug dataset are converted into corresponding drug molecular diagrams to obtain the first drug molecular diagram information;

[0008] The first drug molecule map information is sent to the trained neural network to obtain the second drug molecule map information;

[0009] A first score is calculated based on the first drug molecule map information and the second drug molecule map information. The first score information is used to represent the degree of influence of functional groups on drug toxicity.

[0010] Drug toxicity is predicted based on the first score information.

[0011] Secondly, embodiments of this application provide a drug toxicity prediction device, the device comprising:

[0012] The acquisition module is used to acquire drug datasets;

[0013] The first processing module is used to convert the drug molecular formulas included in the drug dataset into corresponding drug molecular diagrams to obtain first drug molecular diagram information.

[0014] The second processing module is used to send the first drug molecule map information to the trained neural network to obtain the second drug molecule map information.

[0015] The calculation module is used to calculate based on the first drug molecule map information and the second drug molecule map information to obtain first score information, which is used to represent the degree of influence of functional groups on drug toxicity.

[0016] The prediction module is used to predict drug toxicity based on the first score information.

[0017] Thirdly, embodiments of this application provide a drug toxicity prediction device, the device including a memory and a processor. The memory is used to store a computer program; the processor is used to execute the computer program to implement the steps of the above-described drug toxicity prediction method.

[0018] Fourthly, embodiments of this application provide a readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the above-described drug toxicity prediction method.

[0019] The beneficial effects of this invention are as follows:

[0020] This invention converts the drug molecular formulas included in the drug dataset into corresponding drug molecular graphs to obtain first drug molecular graph information. The first drug molecular graph information is then sent to a trained neural network to obtain second drug molecular graph information. The influence of functional groups on drug toxicity is calculated using the first and second drug molecular graph information, thereby predicting drug toxicity based on the characteristics of different functional group combinations, which improves the accuracy and completeness of toxicity prediction.

[0021] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing embodiments of the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings. Attached Figure Description

[0022] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a schematic diagram of the drug toxicity prediction method described in the embodiments of the present invention.

[0024] Figure 2 This is a schematic diagram of the drug toxicity prediction device described in an embodiment of the present invention.

[0025] Figure 3 This is a schematic diagram of the structure of the drug toxicity prediction device described in an embodiment of the present invention.

[0026] The diagram is labeled as follows: 901, Acquisition Module; 902, First Processing Module; 903, Second Processing Module; 904, Calculation Module; 905, Prediction Module; 800, Drug Toxicity Prediction Device; 801, Processor; 802, Memory; 803, Multimedia Component; 804, I / O Interface; 805, Communication Component. Detailed Implementation

[0027] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0028] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this invention, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0029] Example 1:

[0030] This embodiment provides a method for predicting drug toxicity. It can be understood that a scenario can be set up in this embodiment, such as a scenario where drug toxicity testing is required.

[0031] See Figure 1 The figure shows that the method includes steps S1, S2, S3, S4 and S5.

[0032] Step S1: Obtain the drug dataset;

[0033] Step S2: Convert the drug molecular formulas included in the drug dataset into corresponding drug molecular diagrams to obtain the first drug molecular diagram information;

[0034] In this step, the drug molecular formulas included in the drug dataset are converted into corresponding drug molecular diagrams using Python's RDKit library.

[0035] Step S3: Send the first drug molecule map information to the trained neural network to obtain the second drug molecule map information;

[0036] In this step, the trained neural network is a generative adversarial neural network. The second drug molecule map is similar to the first drug molecule map, but the second drug molecule map includes toxic functional group features learned from the generative adversarial neural network.

[0037] Step S3 further includes steps S31, S32, S33, and S34, which specifically include:

[0038] Step S31: Obtain the preset first loss function;

[0039] In this step, the preset first loss function is specifically as follows:

[0040]

[0041] In the above formula, L1 represents the preset first loss function, and n represents the total number of functional group categories in all sample sets. d K is the number of samples, x is a drug sample, G(·) is the generator, D(·) is the discriminator, E(·) is the encoder, and ||·|| 2 The squared Euclidean distance is used to bring samples from the same class closer together in the latent space, while samples from different classes are further apart. The encoder ultimately obtains features specific to drug molecules, enabling the identification of functional group information in drug data.

[0042] Step S32: Train the encoder on the drug dataset according to the preset first loss function to obtain the trained encoder;

[0043] In this step, an encoder is trained on a drug dataset. The main function of the encoder is to map real drug molecule data into the latent space.

[0044] Step S33: Generate the latent vector of the real drug molecule using the trained encoder;

[0045] In this step, the latent vector includes random noise generated from a Gaussian distribution.

[0046] Step S34: Train the neural network based on the latent vectors to obtain the trained neural network.

[0047] Step S34 further includes steps S341, S342, S343, and S344, which specifically include:

[0048] Step S341: Using the latent vector, a first poison molecule diagram is obtained, wherein the first poison molecule diagram includes poison molecules generated by the generator;

[0049] Step S342: Obtain poison molecule information from a preset database;

[0050] In this step, one specific implementation involves obtaining the targeted real poison molecule formula (e.g., hepatotoxicity) from the PubChem database.

[0051] Step S343: Convert the poison molecule information into a second poison molecule diagram;

[0052] Step S344: Send the first poison molecule map and the second poison molecule map to the discriminator and train it using a preset second loss function to obtain the trained neural network.

[0053] In this step, the preset second loss function is specifically as follows:

[0054]

[0055] In the above formula, L2 represents the preset second loss function, and λ is the gradient penalty coefficient. Represents the gradient. ε is a random number sampled from a uniform distribution [0, 1), and ||·||² is the sum of squared gradients. x and represent the features included in the first and second poison molecule diagrams, respectively. During training, the discriminator will distinguish the differences between the poison molecule features generated by the generator and the real poison molecule features.

[0056] Step S4: Calculate the first score information based on the first drug molecule map information and the second drug molecule map information. The first score information is used to represent the degree of influence of functional groups on drug toxicity.

[0057] In this step, the first score information can intuitively demonstrate the importance or influence of functional groups on drug toxicity and provide a certain degree of interpretability.

[0058] Step S4 further includes steps S41 and S42, which specifically include:

[0059] Step S41: Preprocess the first drug molecule map information and the second drug molecule map information respectively to obtain the first functional group map information and the second functional group map information;

[0060] In this step, the preprocessing involves segmenting the molecule using BR I CS and Harper graph reduction methods to obtain functional groups. It should be noted that the functional groups obtained by segmenting the same molecule using multiple segmentation methods in this step are all considered as the functional group set of that molecule, providing different functional group combination perspectives to improve the model's generalization ability to different types of drug molecules.

[0061] Step S42: Calculate the number of times each functional group appears in each sample based on the first functional group map information and the second functional group map information to obtain the first score information.

[0062] In this step, the specific calculation process for the first score information is as follows:

[0063]

[0064] Among them, S f(i) This represents the toxicity score of each functional group, i.e., the first score information. n represents the total number of functional group categories in the sample, α and β are learnable parameters, and N(Sample) represents the number of sample pairs. N represents the number of the i-th functional groups in the second drug molecule diagram of the j-th sample pair. j (x(i)) represents the number of the i-th functional groups in the first drug molecule graph of the j-th sample pair.

[0065] Step S5: Predict drug toxicity based on the first score information.

[0066] In this embodiment, since existing technologies typically consider the correlation between multiple chemical toxicities and share important features between multiple toxicities during model training to discover potential correlations between different toxicity outcomes, thereby achieving drug toxicity prediction, this approach cannot clearly understand how the model uses input features for prediction, and which features play an important role in the prediction, thus lacking interpretability. Therefore, this invention provides interpretability for the model by calculating first score information to characterize the importance or influence of functional groups in drug toxicity.

[0067] Step S5 further includes steps S51, S52, S53, S54, and S55, which specifically include:

[0068] Step S51: Determine the atoms that make up the functional groups in the drug molecule diagram based on the first drug molecule diagram information to obtain atomic information;

[0069] Step S52: Determine the nodes in the hypergraph based on the atomic information to obtain the node position information;

[0070] Step S53: Determine the hyperedges between nodes based on the first fraction information to obtain hyperedge information;

[0071] Step S54: Construct a hypergraph based on the node position information and the hyperedge information to obtain at least one sub-hypergraph information, where each sub-hypergraph information corresponds to a functional group;

[0072] Step S55: Predict drug toxicity based on the sub-hypergraph information.

[0073] In this embodiment, each atom constituting a functional group in the first drug molecule graph information is connected by a hyperedge to construct a sub-hypergraph information, with each sub-hypergraph information corresponding to a functional group. The weights of the hyperedges are first score information, which reflect the importance or influence of the functional group on the toxicity of the drug molecule. By processing the first drug molecule graph information using hypergraph technology, the functional groups in the drug molecule are displayed in a network form, and the toxic effects of these functional groups are expressed through weights. This allows for better acquisition of the characteristics of different functional group combinations, more comprehensive capture of the high-level connections between functional groups in the drug molecule, and more accurate capture of the structural features of the molecule, thereby improving the accuracy and completeness of toxicity prediction.

[0074] Step S55 further includes steps S551, S552, S553, S554, S555, and S556, which specifically include:

[0075] Step S551: Send the information of each sub-hypergraph to the graph attention network to obtain at least two feature vectors;

[0076] In this step, graph neural networks are used to obtain the feature vector information corresponding to each sub-hypergraph.

[0077] Step S552: Calculate the attention score between the two feature vectors to obtain the second score information;

[0078] In this step, the calculation process for the second score information is as follows:

[0079] e ij =LeakyRuLU(a T [W*h i ,W*h j ])

[0080] Among them, e ij Represents the second fractional information, h i and h jThese are the feature vector information corresponding to the i-th and j-th sub-hypergraphs, respectively, a T The weights represent the weight parameters of the neural network, W is the weight matrix, and LeakyReLU is an activation function.

[0081] Step S553: ​​Calculate the attention coefficient based on the second score information to obtain the attention coefficient information;

[0082] In this step, the calculation of the attention coefficient is a technical solution well known to those skilled in the art, and therefore will not be described in detail here.

[0083] Step S554: Update the feature vector information according to the attention coefficient information to obtain the updated feature vector;

[0084] In this step, the attention coefficients of the two sub-hypergraphs can reflect the toxicity score of the combination of these two functional groups. Then, the feature vectors of the sub-hypergraphs are further updated using the attention coefficient information to obtain the updated feature vectors.

[0085] Step S555: Pool the updated feature vector to obtain the pooled feature vector;

[0086] Step S556: Send the pooled feature vector to the prediction layer to obtain the prediction result.

[0087] Example 2:

[0088] like Figure 2 As shown, this embodiment provides a drug toxicity prediction device, which includes an acquisition module 901, a first processing module 902, a second processing module 903, a calculation module 904, and a prediction module 905, specifically including:

[0089] Module 901 is used to acquire drug datasets;

[0090] The first processing module 902 is used to convert the drug molecular formulas included in the drug dataset into corresponding drug molecular diagrams to obtain first drug molecular diagram information.

[0091] The second processing module 903 is used to send the first drug molecule graph information to the trained neural network to obtain the second drug molecule graph information.

[0092] The calculation module 904 is used to calculate based on the first drug molecule map information and the second drug molecule map information to obtain first score information, which is used to represent the degree of influence of functional groups on drug toxicity.

[0093] The prediction module 905 is used to predict drug toxicity based on the first score information.

[0094] In one specific embodiment of this disclosure, the second processing module further includes a first acquisition unit, a first training unit, a first processing unit, and a second processing unit, specifically including:

[0095] The first acquisition unit is used to acquire a preset first loss function;

[0096] The first training unit is used to train the encoder on the drug dataset according to the preset first loss function to obtain the trained encoder.

[0097] The first processing unit is used to generate potential vectors of real drug molecules using the trained encoder.

[0098] The second processing unit is used to train the neural network based on the latent vectors to obtain the trained neural network.

[0099] In one specific embodiment of this disclosure, the second processing unit further includes a third processing unit, a second acquisition unit, a fourth processing unit, and a second training unit, specifically including:

[0100] The third processing unit is used to obtain a first poison molecule diagram using the latent vector, the first poison molecule diagram including poison molecules generated by the generator;

[0101] The second acquisition unit is used to acquire poison molecule information from a preset database;

[0102] The fourth processing unit is used to convert the poison molecule information into a second poison molecule diagram;

[0103] The second training unit is used to send the first poison molecule map and the second poison molecule map to the discriminator and train them using a preset second loss function to obtain the trained neural network.

[0104] In one specific embodiment of this disclosure, the computing module further includes a fifth processing unit and a first computing unit, specifically including:

[0105] The fifth processing unit is used to preprocess the first drug molecule map information and the second drug molecule map information respectively to obtain the first functional group map information and the second functional group map information.

[0106] The first calculation unit is used to calculate the number of times each functional group appears in each sample based on the first functional group map information and the second functional group map information, and obtain the first score information.

[0107] In one specific embodiment of this disclosure, the prediction module further includes a sixth processing unit, a seventh processing unit, an eighth processing unit, a ninth processing unit, and a tenth processing unit, specifically comprising:

[0108] The sixth processing unit is used to determine the atoms that make up the functional groups in the drug molecule diagram based on the first drug molecule diagram information, and obtain the atomic information;

[0109] The seventh processing unit is used to determine the nodes in the hypergraph based on the atomic information and obtain the node position information;

[0110] The eighth processing unit is used to determine the hyperedges between nodes based on the first fraction information, and obtain the hyperedge information;

[0111] The ninth processing unit is used to construct a hypergraph based on the node position information and the hyperedge information, and obtain at least one sub-hypergraph information, wherein one sub-hypergraph information corresponds to one functional group;

[0112] The tenth processing unit is used to predict drug toxicity based on the sub-hypergraph information.

[0113] In one specific embodiment of this disclosure, the tenth processing unit further includes an eleventh processing unit, a second calculation unit, a third calculation unit, a twelfth processing unit, a pooling unit, and a prediction unit, specifically including:

[0114] The eleventh processing unit is used to send the information of each sub-hypergraph to the graph attention network to obtain at least two feature vector information;

[0115] The second computational unit is used to calculate the attention score between two feature vectors to obtain the second score information.

[0116] The third calculation unit is used to calculate the attention coefficient based on the second score information to obtain the attention coefficient information.

[0117] The twelfth processing unit is used to update the feature vector information according to the attention coefficient information to obtain the updated feature vector;

[0118] A pooling unit is used to pool the updated feature vector to obtain a pooled feature vector.

[0119] The prediction unit is used to send the pooled feature vector to the prediction layer to obtain the prediction result.

[0120] It should be noted that the specific manner in which each module performs its operation in the apparatus described in the above embodiments has been described in detail in the embodiments of the method, and will not be elaborated here.

[0121] Example 3:

[0122] Corresponding to the above method embodiments, this embodiment also provides a drug toxicity prediction device. The drug toxicity prediction device described below and the drug toxicity prediction method described above can be referred to each other.

[0123] Figure 3 This is a block diagram illustrating a drug toxicity prediction device 800 according to an exemplary embodiment. Figure 3 As shown, the drug toxicity prediction device 800 may include: a processor 801 and a memory 802. The drug toxicity prediction device 800 may also include one or more of a multimedia component 803, an I / O interface 804, and a communication component 805.

[0124] The processor 801 controls the overall operation of the drug toxicity prediction device 800 to complete all or part of the steps in the aforementioned drug toxicity prediction method. The memory 802 stores various types of data to support the operation of the drug toxicity prediction device 800. This data may include, for example, instructions for any application or method operating on the drug toxicity prediction device 800, as well as application-related data such as contact data, sent and received messages, images, audio, video, etc. The memory 802 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The multimedia component 803 may include a screen and an audio component. The screen may be, for example, a touchscreen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signals may be further stored in the memory 802 or transmitted via the communication component 805. The audio component also includes at least one speaker for outputting audio signals. I / O interface 804 provides an interface between processor 801 and other interface modules, such as a keyboard, mouse, buttons, etc. These buttons can be virtual or physical buttons. Communication component 805 is used for wired or wireless communication between the drug toxicity prediction device 800 and other devices. Wireless communication includes, for example, Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, or 4G, or a combination thereof; therefore, the corresponding communication component 805 may include a Wi-Fi module, a Bluetooth module, or an NFC module.

[0125] In an exemplary embodiment, the drug toxicity prediction device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the drug toxicity prediction method described above.

[0126] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided, which, when executed by a processor, implement the steps of the drug toxicity prediction method described above. For example, the computer-readable storage medium may be the memory 802 including the program instructions described above, which may be executed by the processor 801 of the drug toxicity prediction device 800 to complete the drug toxicity prediction method described above.

[0127] Example 4:

[0128] Corresponding to the above method embodiments, this embodiment also provides a readable storage medium. The readable storage medium described below can be referred to in conjunction with the drug toxicity prediction method described above.

[0129] A readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the drug toxicity prediction method described in the above method embodiments.

[0130] Specifically, the readable storage medium can be a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, or any other readable storage medium capable of storing program code.

[0131] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

[0132] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for predicting drug toxicity, characterized in that, include: Obtain drug datasets; The drug molecular formulas included in the drug dataset are converted into corresponding drug molecular diagrams to obtain the first drug molecular diagram information; The first drug molecule map information is sent to the trained neural network to obtain the second drug molecule map information. The trained neural network is a generative-adversarial neural network. A first score is calculated based on the first drug molecule map information and the second drug molecule map information. The first score information is used to represent the degree of influence of functional groups on drug toxicity. Predicting drug toxicity based on the first score information includes: Based on the information from the first drug molecule diagram, the atoms that make up the functional groups in the drug molecule diagram are determined, and the atomic information is obtained; Based on the atomic information, the nodes in the hypergraph are determined, and the node position information is obtained; Based on the first score information, determine the hyperedges between nodes to obtain the hyperedge information; A hypergraph is constructed based on the node position information and the hyperedge information to obtain at least one subhypergraph information, and one subhypergraph information corresponds to one functional group. Predicting drug toxicity based on the sub-hypergraph information includes: Information from each sub-hypergraph is sent to a graph attention network to predict drug toxicity.

2. The drug toxicity prediction method according to claim 1, characterized in that, Sending the first drug molecule map information to the trained neural network includes: Obtain the preset first loss function; The encoder is trained on the drug dataset according to the preset first loss function to obtain the trained encoder; The trained encoder is used to generate latent vectors for real drug molecules. The neural network is trained based on the latent vectors to obtain the trained neural network.

3. The drug toxicity prediction method according to claim 2, characterized in that, Training the neural network based on the latent vectors includes: Using the latent vector, a first poison molecule diagram is obtained, which includes poison molecules generated by the generator; Retrieve poison molecule information from a pre-set database; The poison molecule information is converted into a second poison molecule diagram; The first poison molecule map and the second poison molecule map are sent to the discriminator and trained using a preset second loss function to obtain the trained neural network.

4. The drug toxicity prediction method according to claim 1, characterized in that, Calculations are performed based on the first drug molecule map information and the second drug molecule map information, including: The first drug molecule map information and the second drug molecule map information are preprocessed respectively to obtain the first functional group map information and the second functional group map information; The first score information is obtained by calculating the number of times each functional group appears in each sample based on the first functional group map information and the second functional group map information.

5. A drug toxicity prediction device, characterized in that, include: The acquisition module is used to acquire drug datasets; The first processing module is used to convert the drug molecular formulas included in the drug dataset into corresponding drug molecular diagrams to obtain first drug molecular diagram information. The second processing module is used to send the first drug molecule map information to the trained neural network to obtain the second drug molecule map information. The trained neural network is a generative-adversarial neural network. The calculation module is used to calculate based on the first drug molecule map information and the second drug molecule map information to obtain first score information, which is used to represent the degree of influence of functional groups on drug toxicity. The prediction module is used to predict drug toxicity based on the first score information; The prediction module includes: The sixth processing unit is used to determine the atoms that make up the functional groups in the drug molecule diagram based on the first drug molecule diagram information, and obtain the atomic information; The seventh processing unit is used to determine the nodes in the hypergraph based on the atomic information and obtain the node position information; The eighth processing unit is used to determine the hyperedges between nodes based on the first fraction information, and obtain the hyperedge information; The ninth processing unit is used to construct a hypergraph based on the node position information and the hyperedge information, and obtain at least one sub-hypergraph information, wherein one sub-hypergraph information corresponds to one functional group; The tenth processing unit is used to predict drug toxicity based on the sub-hypergraph information; The prediction of drug toxicity based on the sub-hypergraph information includes: Information from each sub-hypergraph is sent to a graph attention network to predict drug toxicity.

6. The drug toxicity prediction device according to claim 5, characterized in that, The second processing module includes: The first acquisition unit is used to acquire a preset first loss function; The first training unit is used to train the encoder on the drug dataset according to the preset first loss function to obtain the trained encoder. The first processing unit is used to generate potential vectors of real drug molecules using the trained encoder. The second processing unit is used to train the neural network based on the latent vectors to obtain the trained neural network.

7. The drug toxicity prediction device according to claim 6, characterized in that, The second processing unit includes: The third processing unit is used to obtain a first poison molecule diagram using the latent vector, the first poison molecule diagram including poison molecules generated by the generator; The second acquisition unit is used to acquire poison molecule information from a preset database; The fourth processing unit is used to convert the poison molecule information into a second poison molecule diagram; The second training unit is used to send the first poison molecule map and the second poison molecule map to the discriminator and train them using a preset second loss function to obtain the trained neural network.

8. The drug toxicity prediction device according to claim 5, characterized in that, The computing module includes: The fifth processing unit is used to preprocess the first drug molecule map information and the second drug molecule map information respectively to obtain the first functional group map information and the second functional group map information. The first calculation unit is used to calculate the number of times each functional group appears in each sample based on the first functional group map information and the second functional group map information, and obtain the first score information.

9. A drug toxicity prediction device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the drug toxicity prediction method as described in any one of claims 1 to 4.

10. A readable storage medium, characterized in that: The readable storage medium stores a computer program that, when executed by a processor, implements the steps of the drug toxicity prediction method as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Drug small molecule property prediction method, device and equipment based on graph neural network

    CN113707236A

  • Drug multi-toxicity prediction method based on hierarchical attention fusion of multiple features

    CN116759106A