Drug molecule screening method and system based on the fusion of graph neural network block structure and multi-head attention mechanism
By using the method of integrating graph neural network block structure and multi-head attention mechanism in the prediction of the molecular properties of drugs, the problems of high cost, long time and low efficiency in the existing technology are solved, and more efficient and accurate molecular properties prediction are achieved.
Patent Information
- Application Number
- CN202211374345.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-04
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-11-04
AI Technical Summary
The prior art has problems of high cost and long time in predicting the molecular properties of drugs, and traditional methods require manual feature processing and low efficiency.
The molecular graph is processed using a method based on the fusion of the block structure of the graph neural network and the multi-head attention mechanism. Through the message delivery block and node environment mixing module, the characteristic information of the molecular map is extracted, and the hidden information of the chemical environment is polymerized using transformer structure to generate a molecular fingerprint for molecular properties prediction.
It improves the efficiency and accuracy of drug molecular properties prediction, reduces the need for manual feature processing, reduces costs, and alleviates network degradation problems.
Smart Images

Figure CN115938505B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a drug molecule screening method and system based on the fusion of a graph neural network block structure and a multi-head attention mechanism, belonging to the field of biochemical information technology. Background Art
[0002] The prediction of molecular properties is important for drug discovery and design. In the past, the properties of molecules were usually calculated through specialized prediction equations, which was an expensive and time-consuming process. Therefore, using a computer to simulate drug theory is very helpful for controlling costs and accelerating the progress of overall drug research and development. Algorithms for molecular property prediction include traditional machine learning methods and graph neural networks. Traditional methods describe the chemical environment or overall conformation of each atom through artificially crafted molecular descriptors obtained from chemical domain expertise. These descriptors are usually processed using classical methods to obtain targets on atoms or structures. Graph neural networks are deep learning-based methods that are processed based on graphs. When applying graph neural networks to molecular property prediction, the graph neural network only takes the molecular graph as input and hardly requires artificial feature processing. Although the overall process is simplified, it shows significantly better performance than previous machine learning models, which demonstrates the good prospects of using molecular graphs for molecular property prediction. Summary of the Invention
[0003] The object of the present invention is to provide a drug molecule screening method and system based on the fusion of a graph neural network block structure and a multi-head attention mechanism. The present invention combines the block structure commonly adopted by neural networks with the multi-head attention mechanism to process the molecular graph. The multi-head attention mechanism can extract more feature information in subspaces, and the block structure can alleviate the network degradation problem that commonly exists as the network accuracy increases.
[0004] To achieve the above object, the present invention is realized through the following technical solutions:
[0005] A drug molecule screening method based on the fusion of a graph neural network block structure and a multi-head attention mechanism, comprising the following steps:
[0006] S1. The molecular graph is processed through a message passing block, and the message passing block is composed of I residual blocks. Each residual block is a block-based graph neural network, and the overall architecture of the block-based graph neural network is respectively composed of a directional bond message passing part, a feed-forward neural network, and an edge-level batch normal layer;
[0007] S2. Information processing is performed through a node environment mixing module to obtain the hidden information of the chemical environment around atom b, and a method similar to the combination of multi-head attention and residual blocks in the message passing block is adopted to obtain the target information;
[0008] S3. Obtain the chemical fingerprint through the chemical fingerprint reading part to get the molecular fingerprint for predicting molecular properties:
[0009] S31. The coordinate information of the atoms is fused to participate in the feature expression of the atoms;
[0010] S32. Use the transformer structure to aggregate the hidden information of the chemical environment around each atom, that is, the molecular fingerprint.
[0011] Based on the above drug molecule screening method that combines the graph neural network block structure and the multi-head attention mechanism, in the overall architecture of the graph neural network, the chemical bond ab is the initial hidden state pointing from atom a to atom b. The calculation formula is as follows:
[0012] ,
[0013] ,
[0014] ,
[0015] ,
[0016] Among them, represents the edge feature information of the chemical bond ab in the initial state;
[0017] represents the feature information of atom a;
[0018] represents the feature information of the chemical bond ab. When transmitting information for the first time, the input chemical bond information is which is the initial information of the chemical bond ab Get In the subsequent information transmission process, is obtained from the previous message passing block;
[0019] , , is the dimension of the output vector of a single attention head in the multi-head attention mechanism, is the product of represents the position feature information of the chemical bond ab, represents the position vector of the chemical bond ab, , , , represents layer normalization, is the activation function.
[0020] Based on the above drug molecule screening method that combines the graph neural network block structure and the multi-head attention mechanism, during the directed bond message passing process, and are used as the input of the multi-head attention mechanism to obtain . Finally, after L iterations, the hidden message represented by the chemical bond ab is obtained.
[0021] Based on the above drug molecule screening method that combines the graph neural network block structure and the multi-head attention mechanism, the relevant formulas are as follows:
[0022] a. The query matrix Q, the key matrix K, and the value matrix V are obtained through learning matrix processing. The query vector and the key vector of the scaled dot product are used as the input values for processing the attention mechanism pointing from the chemical bond ca to the chemical bond ab;
[0023] b. The softmax function is used to process the attention scores represented by the chemical bonds in this part;
[0024] c. All the value vectors are weighted and summed, and the sum result is used as the message passed by the chemical bond ab to the next step to obtain the output value of the nth attention head.
[0025] Based on the above drug molecule screening method that combines the graph neural network block structure and the multi-head attention mechanism, the calculation method of the output value of the nth attention head is as follows:
[0026] ,
[0027] where , , , , ;
[0028] The outputs of all the attention heads are concatenated, and the concatenated result is processed to obtain : , where , , ; After obtaining , a transition process is carried out. This part consists of a feed-forward neural network, an edge-level batch normalization layer, and an activation function. Through this part, the residual information of the hidden information of each chemical bond is obtained , , represents the edge-level batch normalization operation;
[0029] After obtaining After that, skip connections are adopted to obtain , and the formula is as follows: , which enters the next residual block as the input. After being processed by I residual blocks, the final output is the hidden feature vector of chemical bond ab ; Based on the above drug molecule screening method that combines the graph neural network block structure and the multi-head attention mechanism, the specific process of S2 is as follows:
[0030] First, using the initial information of atom b, calculate the initial hidden information of the chemical environment around atom b , and the formula is as follows:
[0031] ,
[0032] where , ;
[0033] The multi-head attention mechanism used in this part is similar to the message passing block. The query vector is obtained from the atom hidden , and the corresponding key vector and value vector are the hidden vectors represented by chemical bond ab and the hidden vector of the atom surrounding environment calculated;
[0034] The output formula of each attention head is as follows:
[0035] ,
[0036] where , , ;
[0037] Next, first concatenate the outputs of each attention head, and then obtain the residual vector through the transition operation ,
[0038] ,
[0039] where , , , represents the node-level batch normalization operation;
[0040] Finally, perform a residual connection on and to obtain , ;
[0041] After being updated by F residual blocks, the hidden information of the chemical characteristics around atom b is obtained 。
[0042] Based on the above drug molecule screening method that combines the graph neural network block structure and the multi-head attention mechanism, the specific process of S32 is as follows:
[0043] First, obtain the initial input of the transformer, the initial state of the hidden information of the chemical environment around atom b , and the formula is as follows:
[0044] ,
[0045] Next, assume that a molecule contains n atoms, and embed a random vector into of a molecule, and the formula is as follows:
[0046] ,
[0047] ,
[0048] Among them, , , , , is a learnable random vector used as a chemical fingerprint, obtained through continuous calculation of T transformer layers , select as the molecular chemical fingerprint, and use ANN to for processing, and finally obtain the property prediction result of the atom.
[0049] A system for implementing the drug molecule screening method that combines the graph neural network block structure and the multi-head attention mechanism includes a message passing block, a node environment mixing module, and a molecular fingerprint acquisition module. The message passing block is composed of I residual blocks, and each residual block is a block-based graph neural network. The overall architecture of the block-based graph neural network is composed of a directional bond message passing part, a feed-forward neural network, and an edge-level batch normal layer; the node environment mixing module adopts a method similar to the combination of multi-head attention and residual blocks in the message passing block to obtain target information; the molecular fingerprint acquisition module uses a transformer structure to aggregate the hidden information of the chemical environment around each atom.
[0050] The advantages of the present invention are as follows:
[0051] (1) In the message passing block, the overall residual block design is adopted, and the problem of network degradation is reduced through skip connections and normalization processing. According to the characteristics of the transfer network in this article, edge-level batch normalization is proposed.
[0052] (2) In the message passing block, a message passing mode centered on directional chemical bonds is adopted, which can avoid the repeated transmission of messages that occurs during node message passing.
[0053] (3) When the framework utilizes directional chemical bond message passing and obtains molecular chemical fingerprints, a multi-head attention mechanism is adopted. Compared with the commonly used methods of direct summation and set2set, this method can extract richer chemical feature information and fuse this information more reasonably, ultimately obtaining chemical information with stronger expressive power. Description of the Drawings
[0054] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention.
[0055] Figure 1 It is a schematic diagram of the framework of the embodiment of the present invention;
[0056] Figure 2 It is the overall architecture of the graph neural network based on the block structure of the embodiment of the present invention;
[0057] Figure 3 It is a schematic diagram of directional bond message passing of the embodiment of the present invention;
[0058] Figure 4 It is a schematic diagram of different normalization methods of the embodiment of the present invention;
[0059] Figure 5 It is a schematic diagram of the node environment mixing module of the embodiment of the present invention;
[0060] Figure 6 It is a schematic diagram of the multi-head attention mechanism of the node mixing module of the embodiment of the present invention;
[0061] Figure 7 It is a schematic diagram of initial information aggregation of the embodiment of the present invention;
[0062] Figure 8 It is a schematic diagram of the transformer layer structure of the embodiment of the present invention;
[0063] Figure 9 It is a comparison diagram of the evaluation effect provided for the embodiment of the present invention and the current advanced classification model;
[0064] Figure 10 It is a comparison of the ROC curves of the experimental results of the BBBP dataset of the embodiment of the present invention. Specific Embodiments
[0065] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0066] In this embodiment, the BBBP (Blood-brain barrier penetration) dataset is selected. The dataset consists of molecular formulas and molecular formula labels. There are two types of molecular formula classification labels, 0 and 1, and the final prediction results are also represented by outputting 0 and 1.
[0067] The overall architecture of the drug molecule screening method based on the fusion of the graph neural network block structure and the multi-head attention mechanism of the present invention is as Figure 1 shown. The generation of the molecular orientation information descriptor and the operation of the convolutional neural network included in the present invention include the following steps:
[0068] 1. In the first step, the molecular graph is processed by the message passing block. The message passing block is composed of I residual blocks, and each residual block is a block-based graph neural network. Figure 2 The overall architecture of the shown block-based graph neural network is respectively composed of the directional bond message passing part, the feed-forward neural network, and the edge-level batch normal layer.
[0069] In the embodiment, the directional bond message passing part in the overall architecture of the graph neural network has a multi-head attention mechanism, and its architecture is as Figure 3 shown. The chemical bond ab is the initial hidden state from atom a to atom b. The calculation formula is as follows:
[0070] ,
[0071] ,
[0072] ,
[0073] ,
[0074] Among them, represents the edge feature information of the chemical bond ab in the initial state; represents the feature information of atom a; represents the feature information of the chemical bond ab. When transmitting information for the first time, the input chemical bond information is which is the initial information of the chemical bond ab to obtain , during the subsequent information transmission process, is obtained from the previous message transmission block;
[0075] , , is the dimension of the output vector of a single attention head in the multi-head attention mechanism, is the product of and the number of attention heads, representing the position feature information of chemical bond ab, represents the position vector of chemical bond ab, , , , represents layer normalization, is the activation function.
[0076] In this embodiment, during the directed bond message transmission process, and are used as the input of the multi-head attention mechanism to obtain , and finally, after L iterations, the hidden message represented by chemical bond ab is obtained ; the relevant formula is as follows:
[0077] a. The query matrix Q, key matrix K, and value matrix V are obtained by processing the learning matrix. The query vector and the key vector of the scaled dot product are used as the input values for processing the attention mechanism pointing from chemical bond ca to chemical bond ab;
[0078] b. The softmax function is used to process the attention scores represented by the chemical bonds in this part;
[0079] c. All the value vectors are weighted and summed, and the sum result is used as the message passed by chemical bond ab to the next step to obtain the output value of the nth attention head;
[0080] In this embodiment, the calculation method of the output value of the nth attention head is as follows:
[0081] ,
[0082] where , , , , ;
[0083] The outputs of all attention heads are concatenated, and the concatenated result is processed to obtain : , where, , , ;
[0084] As shown in Figure 2 , after obtaining , a transition process is carried out. This part consists of a feed-forward neural network, an edge-level batch normalization layer, and an activation function. Through this part, the residual information of the hidden information of each chemical bond is obtained. , , represents the edge-level batch normalization operation;
[0085] After obtaining , a skip connection is adopted to obtain , and the formula is as follows: , enters the next residual block as the input. After being processed by I residual blocks, the final output, the hidden feature vector of chemical bond ab, is obtained. .
[0086] 2. In the second step, information processing is carried out through the node environment mixing module. As shown in Figure 5 , in this module, the purpose is to obtain the hidden information of the chemical environment around atom b, and a method combining multi-head attention and residual blocks similar to that in the message passing block is adopted to obtain the target information.
[0087] First, using the initial information of atom b, the initial hidden information of the chemical environment around atom b is calculated , and the formula is as follows:
[0088] ,
[0089] Among them, , ;
[0090] The multi-head attention mechanism used in this part is similar to that in the message passing block. The query vector is obtained from the atom hidden , and the corresponding key vector and value vector are calculated from the hidden vector represented by chemical bond ab and the hidden vector of the atom's surrounding environment .
[0091] As shown in Figure 6 , the output formula of each attention head is as follows:
[0092] ,
[0093] Among them, , , ;
[0094] Next, the output of each attention head is concatenated first, and then a residual vector is obtained through a transition operation. ,
[0095] ,
[0096] Among them, , , , represents the node-level batch normalization operation;
[0097] Finally, for and a residual connection is performed to obtain , ;
[0098] After being updated through F residual blocks, the hidden information of the chemical features around atom b is obtained .
[0099] S3. In the last step, the chemical fingerprint is obtained through the chemical fingerprint reading part to get the molecular fingerprint for molecular property prediction, specifically as follows:
[0100] S31. The coordinate information of the atom is fused to participate in the feature expression of the atom;
[0101] S32. The transformer structure is used to aggregate the hidden information of the chemical environment around each atom, that is, the molecular fingerprint, so as to obtain a molecular fingerprint with stronger expression ability.
[0102] The specific process of S32 is as follows:
[0103] First, obtain the initial input of the transformer, the initial state of the hidden information of the chemical environment around atom b ;
[0104] The overall calculation process of Figure 7 is as shown in
[0105] ,
[0106] Next, assume that a molecule contains n atoms, and a random vector is embedded into in a molecule, and the formula is as follows:
[0107] ,
[0108] ,
[0109] Among them, , , , , is a learnable random vector used as a chemical fingerprint, obtained through consecutive calculations of T transformer layers , select as the molecular chemical fingerprint, and use ANN to process it to finally obtain the predicted results of atomic properties.
[0110] To verify the advantages of the present invention for molecular property screening, the present invention conducted molecular property prediction experiments on the BBBP (Blood-brain barrier penetration) dataset, ClinTox dataset, SIDER (Side Effect Resource) dataset, and Tox21 (Toxicology in the 21st Century) dataset. The experimental results are as Figure 9 and Figure 10 shown. From Figure 9 and Figure 10 , it can be seen that the drug molecule screening method based on the fusion of the graph neural network block structure and the multi-head attention mechanism established by the present invention has achieved good results in molecular property prediction. The AUC value of the present invention is significantly higher than that of other methods. The higher the AUC value, the stronger the classification ability. This indicates that the classification ability of the present invention is significantly stronger than that of other methods. The prediction of the present invention for molecular properties is effective, providing a better method for molecular property screening and having certain practical value.
[0111] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A drug molecule screening method based on the fusion of graph neural network block structure and multi-head attention mechanism, characterized in that, it includes the following steps: S1. The molecular graph is processed by a message passing block, and the message passing block is composed of I residual blocks. Each residual block is a block-based graph neural network, and the overall architecture of the block-based graph neural network is composed of a directed bond message passing part, a feed-forward neural network, and an edge-level batch normal layer; During the process of directional key message passing, the initial hidden state where atom a points to atom b and the hidden state where atom a points to atom b after the l-th iteration are used as the input of the multi-head attention mechanism to obtain , and finally, after L iterations, the hidden message represented by the chemical bond ab is obtained ; The relevant formula is as follows: a. The query matrix Q, key matrix K, and value matrix V are obtained after being processed by the learning matrix, and the query vector and the key vector are used as the input values for the attention mechanism that processes the chemical bond ca pointing to the chemical bond ab; b. The softmax function is used to process the attention scores represented by the chemical bonds contained in this part; c. Sum all the value vectors and sum the weights. The sum result is used as the message passed by the chemical bond ab to the next step, and the output value of the nth attention head is obtained; The calculation method of the output value of the nth attention head is as follows: , Among them, , , , , ; Concatenate the outputs of all attention heads and process the concatenated result to obtain : , Among them, , , ; After obtaining , perform a transition process, which consists of a feed-forward neural network, an edge-level batch normalization layer, and an activation function. Through this part, the residual information of the hidden information of each chemical bond is obtained , , represents an edge-level batch normalization operation; Obtained After that, a skip connection is adopted to obtain , and the formula is as follows: , As the input enters the next residual block. After being processed by I residual blocks, the final output is the hidden feature vector of chemical bond ab ; S2. Information processing is carried out through the node environment mixing module to obtain the hidden information of the chemical environment around atom b, and a method similar to the combination of multi-head attention and residual blocks in the message passing block is adopted to obtain the target information; S3. Chemical fingerprints are obtained through the chemical fingerprint reading part to obtain molecular fingerprints for molecular property prediction: S31. The coordinate information of the atom is fused to participate in the feature expression of the atom; S32. The transformer structure is used to aggregate the hidden information of the chemical environment around each atom, that is, the molecular fingerprint.
2. The drug molecule screening method based on the fusion of graph neural network block structure and multi-head attention mechanism according to claim 1, characterized in that: In the overall architecture of the graph neural network, the chemical bond ab is the initial hidden state pointing from atom a to atom b , and the calculation formula is as follows: , , , , Among them, represents the edge feature information of the chemical bond ab in the initial state; Characteristic information representing atom a; Representing the characteristic information of chemical bond ab, when the information is transmitted for the first time, the input chemical bond information is , which is the initial information of chemical bond ab Obtain , during the subsequent information transmission process, is obtained from the previous message passing block; , , is the dimension of the output vector of a single attention head in the multi-head attention mechanism, is the product of and the number of attention heads, represents the position feature information of chemical bond ab, , , , represents layer normalization, is the activation function.
3. The drug molecule screening method based on the fusion of graph neural network block structure and multi-head attention mechanism according to claim 2, characterized in that: The specific process of S2 is as follows: First, using the initial information of atom b, calculate the initial hidden information of the chemical environment around atom b , and the formula is as follows: , Among them, , ; The multi-head attention mechanism used in this part is similar to the message passing block, and the query vector is obtained from the atomic hidden The corresponding key vector and value vector are the hidden vectors represented by the chemical bond ab and the hidden vector of the atomic surrounding environment calculated; The output formula of each attention head is as follows: , Among them, , , ; Next, each attention head is concatenated and output, and then a residual vector is obtained through a transition operation , , Among them, , , , represent node-level batch normalization operations; Finally, for and perform a residual connection to obtain , ; After being updated by F residual blocks, the hidden information of the surrounding chemical features of atom b is obtained .
4. The drug molecule screening method based on the fusion of graph neural network block structure and multi-head attention mechanism according to claim 3, characterized in that: The specific process of S32 is as follows: First, obtain the initial input of the transformer, which is the initial state of the hidden information about the chemical environment around atom b , and the formula is as follows: , Next, assume that a molecule contains n atoms, and embed a random vector into a molecule, as follows: , , Among them, , , , , is a learnable random vector used as a chemical fingerprint, obtained through consecutive calculations of T transformer layers , select as the molecular chemical fingerprint, and use ANN to process it to finally obtain the prediction results of atomic properties.
5. A system for implementing the drug molecule screening method based on the fusion of graph neural network block structure and multi-head attention mechanism according to any one of claims 1 to 4, characterized in that: It includes a message passing block, a node environment mixing module, and a molecular fingerprint acquisition module. The message passing block is composed of I residual blocks. Each residual block is a block-based graph neural network, and the overall architecture of the block-based graph neural network is composed of a directed bond message passing part, a feed-forward neural network, and an edge-level batch normal layer; the node environment mixing module adopts a method similar to the combination of multi-head attention and residual blocks in the message passing block to obtain the target information; the molecular fingerprint acquisition module uses the transformer structure to aggregate the hidden information of the chemical environment around each atom.