Drug molecule screening method and system based on a neural network with dual self-attention mechanism

By introducing a dual self-attention mechanism neural network into drug molecular screening technology, the hidden information inside and between drug molecules is extracted, and the problem of strong subjectivity of molecular feature selection in the prior art is solved, achieving more accurate prediction of drug molecular properties and more efficient molecular screening.

CN115424680BActive Publication Date: 2025-06-27SHANDONG AOWANGDE INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211202242.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-29
Publication Date
2025-06-27
Estimated Expiration
2042-09-29

AI Technical Summary

Technical Problem

In the existing drug molecular screening technology, molecular feature selection is highly subjective, resulting in large errors, and deep learning algorithms are less used in this field, resulting in low molecular properties prediction efficiency.

Method used

The drug molecule screening method based on the dual self-attention mechanism neural network is adopted. Through the juxtaposed first self-attention mechanism branch and second self-attention mechanism branch, the molecular formula of the drug molecule is processed, hidden information inside and between the molecules is extracted, and fusion is carried out to obtain more objective molecular characteristics.

Benefits of technology

This method can more accurately predict the molecular properties of drugs, reduce errors, shorten experimental time, and improve molecular screening efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115424680B_ABST
    Figure CN115424680B_ABST
Patent Text Reader

Abstract

The present invention discloses a drug molecule screening method and system based on a double self-attention mechanism neural network; obtaining drug molecules to be screened; using the trained double self-attention mechanism neural network to process the molecular formula of the drug molecules to obtain a prediction result of the drug molecule properties; the double self-attention mechanism neural network includes: a first and a second self-attention mechanism branches arranged in parallel; inputting the molecular formula of the drug molecules into the first self-attention mechanism branch to output a first molecular descriptor; the first molecular descriptor includes the molecular hidden information inside each atomic subunit in the molecular formula; inputting the molecular formula of the drug molecules into the second self-attention mechanism branch to output a second molecular descriptor; the second molecular descriptor includes the molecular hidden information between all atomic subunits in the molecular formula; fusing the first and second molecular descriptors to obtain a third molecular descriptor; classifying the third molecular descriptor to obtain a prediction result of the drug molecule properties.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of drug molecule screening, and particularly to a drug molecule screening method and system based on a dual self-attention mechanism neural network. Background Art

[0002] The statements in this section merely mention the background art related to the present invention and do not necessarily constitute prior art.

[0003] Currently, the cost of developing a new drug averages between $2 billion and $3 billion. Molecular screening is a crucial step in the successful development of a new drug. To reduce time and save costs, machine learning techniques have been introduced into molecular screening. Researchers extract fingerprints or manually engineered features, predict molecular properties, and screen molecules through machine learning methods. It can help predict molecular properties and select the most likely molecules for molecular screening. However, most molecular features are artificially set according to the corresponding tasks, resulting in excessive subjectivity in the selection of corresponding features and prone to errors. Summary of the Invention

[0004] To solve the deficiencies of the prior art, the present invention provides a drug molecule screening method and system based on a dual self-attention mechanism neural network; applying a deep learning algorithm to the molecular property prediction experiment, the deep learning algorithm can obtain molecular features more objectively and shorten the experimental time.

[0005] In the first aspect, the present invention provides a drug molecule screening method based on a dual self-attention mechanism neural network;

[0006] The drug molecule screening method based on a dual self-attention mechanism neural network includes:

[0007] Obtain the drug molecules to be screened;

[0008] Use the trained dual self-attention mechanism neural network to process the molecular formula of the drug molecule to obtain the prediction result of the drug molecule property;

[0009] Among them, the trained dual self-attention mechanism neural network includes: a first self-attention mechanism branch and a second self-attention mechanism branch arranged in parallel; input the molecular formula of the drug molecule into the first self-attention mechanism branch to output a first molecular descriptor; the first molecular descriptor includes: the molecular hidden information inside each atomic subunit in the molecular formula; input the molecular formula of the drug molecule into the second self-attention mechanism branch to output a second molecular descriptor; the second molecular descriptor includes: the molecular hidden information between all atomic subunits in the molecular formula; fuse the first molecular descriptor and the second molecular descriptor to obtain a third molecular descriptor; classify the third molecular descriptor to obtain the prediction result of the drug molecule property.

[0010] In a second aspect, the present invention provides a drug molecule screening system based on a dual self-attention mechanism neural network;

[0011] The drug molecule screening system based on a dual self-attention mechanism neural network includes:

[0012] An acquisition module configured to acquire drug molecules to be screened;

[0013] A prediction module configured to process the molecular formula of the drug molecule using a trained dual self-attention mechanism neural network to obtain a prediction result of the drug molecule property;

[0014] Wherein, the trained dual self-attention mechanism neural network includes: a first self-attention mechanism branch and a second self-attention mechanism branch arranged in parallel; inputting the molecular formula of the drug molecule into the first self-attention mechanism branch to output a first molecular descriptor; the first molecular descriptor includes: molecular hidden information inside each atomic subunit in the molecular formula; inputting the molecular formula of the drug molecule into the second self-attention mechanism branch to output a second molecular descriptor; the second molecular descriptor includes: molecular hidden information between all atomic subunits in the molecular formula; fusing the first molecular descriptor and the second molecular descriptor to obtain a third molecular descriptor; classifying the third molecular descriptor to obtain a prediction result of the drug molecule property.

[0015] In a third aspect, the present invention further provides an electronic device, including:

[0016] A memory for non-temporarily storing computer-readable instructions; and

[0017] A processor for running the computer-readable instructions,

[0018] Wherein, when the computer-readable instructions are run by the processor, the method described in the first aspect above is executed.

[0019] In a fourth aspect, the present invention further provides a storage medium that non-temporarily stores computer-readable instructions, wherein when the non-temporary computer-readable instructions are executed by a computer, the instructions for executing the method described in the first aspect are executed.

[0020] In a fifth aspect, the present invention further provides a computer program product including a computer program, and the computer program is used to implement the method described in the first aspect above when running on one or more processors.

[0021] Compared with the prior art, the beneficial effects of the present invention are:

[0022] (1) The dual self - attention mechanism not only focuses on the relationships between various subunits in the molecule but also on the relationships between the atoms and chemical bonds contained in each subunit. This method extracts more detailed information.

[0023] (2) On the directed molecular graph, an information transfer method centered on directed molecular bonds is adopted. This method reduces redundant information and avoids repeated extraction of information. Brief Description of the Drawings

[0024] The attached drawings forming a part of this invention are used to provide a further understanding of the invention. The schematic embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0025] Figure 1 is the overall architecture of the drug molecule screening algorithm based on the dual self - attention fusion information neural network in the first embodiment of the invention;

[0026] Figure 2 is a schematic diagram of the process of transferring molecular information by directed chemical bonds in the first embodiment of the invention;

[0027] Figure 3 is a schematic diagram of the molecular information aggregation process in the first embodiment of the invention;

[0028] Figure 4 is a schematic diagram of the molecular information reading - out process in the first embodiment of the invention;

[0029] Figure 5 is a comparison graph of the AUC curves of the experimental results of this patent and the SAMPN algorithm on the BBBP dataset in the first embodiment of the invention;

[0030] Figures 6(a) and 6(b) are comparison graphs of the running results of this patent and other algorithms on eight publicly available datasets in the first embodiment of the invention. Detailed Description of the Invention

[0031] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the invention. Unless otherwise specified, all technical and scientific terms used in this invention have the same meaning as commonly understood by those of ordinary skill in the technical field to which this invention belongs.

[0032] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units need not be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0033] In the case of no conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.

[0034] All data acquisition in this embodiment is based on compliance with laws, regulations and user consent, and is a legal application of the data.

[0035] Embodiment 1

[0036] This embodiment provides a drug molecule screening method based on a dual self-attention mechanism neural network;

[0037] As Figure 1 shown, the drug molecule screening method based on the dual self-attention mechanism neural network includes:

[0038] S101: Obtain the drug molecules to be screened;

[0039] S102: Use the trained dual self-attention mechanism neural network to process the molecular formula of the drug molecule to obtain the prediction result of the drug molecule properties;

[0040] Among them, the trained dual self-attention mechanism neural network includes:

[0041] Parallel first self-attention mechanism branch and second self-attention mechanism branch;

[0042] Input the molecular formula of the drug molecule into the first self-attention mechanism branch, and output the first molecular descriptor; the first molecular descriptor includes: the molecular hidden information inside each atomic subunit in the molecular formula;

[0043] Input the molecular formula of the drug molecule into the second self-attention mechanism branch, and output the second molecular descriptor; the second molecular descriptor includes: the molecular hidden information between all atomic subunits in the molecular formula;

[0044] Fuse the first molecular descriptor and the second molecular descriptor to obtain the third molecular descriptor;

[0045] Classify the third molecular descriptor to obtain the prediction result of the drug molecular property.

[0046] Among them, a subunit refers to an atom and all chemical bonds connected to it in a molecule, which together form a subunit.

[0047] Furthermore, input the molecular formula of the drug molecule into the first self-attention mechanism branch to output the first molecular descriptor; the working process of the first self-attention mechanism branch includes:

[0048] S102-a1: Obtain the hidden information of the chemical bonds between atom b and all adjacent atoms through chemical information aggregation; where chemical information aggregation means aggregating multiple collected chemical vector information into one chemical vector information by addition or splicing.

[0049] S102-a2: Process the chemical information of atom b and the hidden information of the chemical bonds between atom b and all adjacent atoms through the self-attention mechanism layer respectively, sum the output results after the self-attention mechanism processing, and process the sum result through an activation function to obtain the molecular hidden information of the subunit where atom b is located.

[0050] Similarly, obtain the molecular hidden information of the subunits where other atoms of the drug molecule are located.

[0051] S102-a3: Sum the molecular hidden information of the subunits where all atoms are located to obtain the molecular descriptor K.

[0052] Furthermore, as Figure 3 shown, S102-a2: Process the chemical information of atom b and the hidden information of the chemical bonds between atom b and all adjacent atoms through the self-attention mechanism layer respectively, sum the obtained results, and process the sum result through an activation function to obtain the molecular hidden information of the subunit where atom b is located; specifically including:

[0053] Apply the self-attention mechanism in the process of aggregating the molecular hidden information inside the subunits of the polymer molecule. The information involved in this process comes from the information of atoms and the information of directional chemical bonds:

[0054]

[0055] Among them, SeAt(·) represents the operation process of the self-attention mechanism, is the chemical information of atom b, is a learning matrix, and its size S a is determined by the size of x b of. is the hidden information of the ab bond, It is the hidden information of the db key, It is the hidden information of the fb key, K b It is the molecular hidden information of the subunit where atom b is located. ε represents the Relu activation function.

[0056] The input of the self-attention mechanism is W a x b , These four vectors are input simultaneously in a concatenated form. After being processed by the self-attention mechanism, the output is still four vectors of the same scale. Summation is to add these four vectors pairwise.

[0057] Furthermore, in step S102-a3: Summing up the molecular hidden information of all the subunits where atoms are located to obtain the molecular descriptor K, specifically including:

[0058] Adding up the molecular hidden information of each subunit, and finally obtaining the molecular descriptor K. The formula is as follows:

[0059]

[0060] M represents the set of all atoms contained in the molecular unit.

[0061] Furthermore, as Figure 2 shown, in step S102-a1: Obtaining the hidden information of the chemical bonds between atom b and all adjacent atoms through chemical information aggregation, specifically including:

[0062] In the first iteration process, assuming that the adjacent atom of atom b is a, then in the initial state, calculate the hidden molecular information of chemical bond ab:

[0063]

[0064] Among them, represents the molecular hidden information of chemical bond ab in its initial state. represents the chemical information x of atom a a and the chemical information l of chemical bond ab ab concatenated fusion. W f is a learning matrix, and the default value of s is 300. s i is determined by the size of f(x a ,l ab ). ε represents the Relu activation function.

[0065] In step t + 1, the molecular hidden information passed to chemical bond ab comes from the bonds around atom a ( Figure 1 the atoms c and e in it), pointing to atom a, but excluding atom b; the information transfer equation is as follows:

[0066]

[0067] Represents the molecular hidden information that has been updated in step t. Among them Represents the hidden information passed from the molecule to the molecular bond ab. Q(a) represents all the atoms around atom a, and x ∈ {Q(a)\b} represents the set of atoms around atom a except atom b. G t (·) is the information transfer function, and the calculation method is as follows:

[0068]

[0069] Among them, Represents x a 、x c and The fusion of is selected as vector concatenation fusion. s j is obtained by adding the sizes of x a 、x c and . ε represents the Relu activation function. is a learning matrix. s j is determined by the size of .

[0070] In step t + 1, the molecular hidden information is updated to The calculation method is as follows:

[0071]

[0072] Represents the molecular hidden information of the chemical bond ab in its initial state. Represents the hidden information passed from the molecule to the molecular bond ab. W g ∈R s×s is the learning matrix. ε represents the Relu activation function. After T iterations, the hidden information

[0073] Apply the self-attention mechanism to the molecular aggregation process. After T iterations, the chemical information It comes from the chemical bond. The aggregation of and x y is the first step of chemical information aggregation, which is the aggregation of information within the atomic subunit.

[0074] Further, input the molecular formula of the drug molecule into the second self-attention mechanism branch to output a second molecular descriptor; the working process of the second self-attention mechanism branch includes:

[0075] S102-b1: Calculate the molecular hidden information of each atomic subunit;

[0076] S102-b2: Use the self-attention mechanism to process the molecular hidden information of each atomic subunit to obtain the molecular descriptor J.

[0077] Further, the S102-b1: Calculate the molecular hidden information of each atomic subunit, including:

[0078] Calculate the molecular hidden information of each subunit and explain the hidden information contained in each subunit:

[0079]

[0080] is the chemical information of atom b, represents the hidden information contained in all chemical bonds pointing to atom b from atom b. is a learning matrix, and its size S a is determined by the size of. represents x b and the concatenated fusion of. ε represents the Relu activation function.

[0081] Further, the S102-b2: Use the self-attention mechanism to process the molecular hidden information of each atomic subunit to obtain the molecular descriptor J, including:

[0082] Use the self-attention mechanism to process the molecular hidden information of each subunit and obtain the molecular descriptor J. The formula is as follows:

[0083] J = ε(∑SeAt(J a , J b , ···, J j ))

[0084] where SeAt(·) represents the running process of the self-attention mechanism, and J x represents the hidden information of the subunit where atom x is located.

[0085] Further, as Figure 4 shown, the fusion of the first molecular descriptor and the second molecular descriptor to obtain the third molecular descriptor specifically includes:

[0086] Input the molecular descriptor K and the molecular descriptor J into the fusion layer for parallel fusion;

[0087] By learning the output of the matrix processing fusion layer and then processing it through the activation function Relu, the molecular descriptor P is obtained.

[0088] P = ε(W f f(K, J))

[0089] ε represents the Relu activation function, is a learning matrix, and its size S a is determined by the size of f(K, J), and P ∈ R S , S = 300. f(K, J) represents the concatenated fusion of the molecular descriptor K and the molecular descriptor J.

[0090] Furthermore, the classification of the third molecular descriptor to obtain the prediction result of the drug molecular property specifically includes:

[0091] P is processed through a feedforward neural network to realize the prediction of molecular properties.

[0092] Furthermore, for the trained double self-attention mechanism neural network, the training process includes:

[0093] Construct a training set and a test set; both the training set and the test set are drug molecular structural formulas with known drug molecular properties;

[0094] Input the training set into the double self-attention mechanism neural network to train the network. When the value of the cross-entropy loss function no longer decreases, stop training; obtain the preliminarily trained network;

[0095] Input the test set into the double self-attention mechanism neural network to test the network. The evaluation index selects the area under the curve AUC (Area under the Curve of ROC) value. When the AUC value reaches the set threshold, the training ends, and the trained double self-attention mechanism neural network is obtained.

[0096] In the experimental model training of the present invention, the loss function selects the categorical cross-entropy loss function, the optimizer selects the adam optimizer, and the evaluation index selects the AUC value.

[0097] The present invention conducts experiments on eight publicly available data sets, and the experimental results are as Figure 5 shown. Train on the labeled data set, select the best-performing model parameters, and then save the best-performing model to predict molecules with unknown labels, thereby obtaining the prediction results.

[0098] In this implementation case, the BBBP (Blood-brain barrier penetration) dataset is selected. The dataset consists of molecular formulas and molecular formula labels. There are two types of molecular formula classification labels, 0 and 1, and the final prediction result is also represented by outputting 0 and 1.

[0099] The experimental results of the present invention are compared with multiple methods. As can be seen from Figures 6(a) and 6(b), the drug molecule screening method based on molecular orientation information descriptors and self-attention mechanism established by the present invention has achieved good results in molecular property prediction. The AUC value of the present invention is significantly higher than that of other methods. The higher the AUC value, the stronger the classification ability. This indicates that the classification ability of the present invention is significantly stronger than that of other methods. The present invention is effective in predicting molecular properties, provides a better method for molecular property screening, and has certain practical value.

[0100] The present invention relates to a drug molecule screening algorithm based on a dual self-attention fusion information neural network (DSFMNN), which is a method for classifying molecular properties and belongs to the field of bioinformatics processing. DSFMNN adopts the combination of a dual self-attention mechanism and a graph convolutional neural network. Advantages of DSFMNN:

[0101] (1) The dual self-attention mechanism not only focuses on the relationships between various subunits in the molecule, but also focuses on the relationships between atoms and chemical bonds contained in each subunit.

[0102] (2) On the directed molecular graph, an information transfer method centered on directed molecular bonds is adopted. The present invention tests the performance of the model on eight publicly available datasets and compares it with the performance of several models. The drug molecule screening algorithm based on the dual self-attention fusion information neural network (DSFMNN) has obtained better molecular property prediction results, providing a method with obvious advantages for the field of drug screening.

[0103] The features of the present invention are: (1) The dual self-attention mechanism not only focuses on the relationships between various subunits in the molecule, but also focuses on the relationships between atoms and chemical bonds contained in each subunit. (2) On the directed molecular graph, an information transfer method centered on directed molecular bonds is adopted.

[0104] The present invention proposes a dual self-attention fusion information neural network. The present invention uses a directed information centered on chemical bonds to transfer molecular information to reduce the redundant information of the molecule. Specifically, the present invention improves the message passing part and the readout part of the D-MPNN framework to perform the molecular information collection function to obtain molecular graph descriptors. The present invention applies a self-attention mechanism in the network to improve the importance of the key parts of molecular information.

[0105] Example Two

[0106] This embodiment provides a drug molecule screening system based on a dual self-attention mechanism neural network;

[0107] A drug molecule screening system based on a dual self-attention mechanism neural network includes:

[0108] An acquisition module, which is configured to: acquire drug molecules to be screened;

[0109] A prediction module, which is configured to: use the trained dual self-attention mechanism neural network to process the molecular formula of the drug molecule to obtain a prediction result of the drug molecule properties;

[0110] Among them, the trained dual self-attention mechanism neural network includes: a first self-attention mechanism branch and a second self-attention mechanism branch arranged in parallel; input the molecular formula of the drug molecule into the first self-attention mechanism branch to output a first molecular descriptor; the first molecular descriptor includes: molecular hidden information inside each atomic subunit in the molecular formula; input the molecular formula of the drug molecule into the second self-attention mechanism branch to output a second molecular descriptor; the second molecular descriptor includes: molecular hidden information between all atomic subunits in the molecular formula; fuse the first molecular descriptor and the second molecular descriptor to obtain a third molecular descriptor; classify the third molecular descriptor to obtain a prediction result of the drug molecule properties.

[0111] It should be noted here that the above acquisition module and prediction module correspond to steps S101 to S102 in Embodiment One. The examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the content disclosed in the above Embodiment One. It should be noted that the above modules, as part of the system, can be executed in a computer system such as a set of computer-executable instructions.

[0112] In the above embodiments, the descriptions of each embodiment have their own emphases. For parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0113] The proposed system can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the above division of modules is only a logical function division. In actual implementation, there can be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.

[0114] Example Three

[0115] This embodiment also provides an electronic device, including: one or more processors, one or more memories, and one or more computer programs; wherein, the processor is connected to the memory, the above one or more computer programs are stored in the memory, and when the electronic device runs, the processor executes the one or more computer programs stored in the memory, so that the electronic device executes the method described in the first embodiment above.

[0116] It should be understood that in this embodiment, the processor may be a central processing unit CPU, and the processor may also be other general-purpose processors, digital signal processors DSP, application-specific integrated circuits ASIC, off-the-shelf programmable gate arrays FPGA or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0117] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random memory. For example, the memory may also store information about the device type.

[0118] In the implementation process, each step of the above method may be completed by the integrated logic circuit in the hardware of the processor or the instructions in the form of software.

[0119] The method in the first embodiment may be directly embodied as being executed and completed by the hardware processor, or executed and completed by the combination of the hardware and software modules in the processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.

[0120] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with this embodiment can be implemented by electronic hardware or the combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.

[0121] Embodiment 4

[0122] This embodiment also provides a computer-readable storage medium for storing computer instructions, and when the computer instructions are executed by a processor, the method described in the first embodiment is completed.

[0123] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A drug molecule screening method based on a neural network with dual self-attention mechanisms, characterized in that Including: Obtain drug molecules to be screened; Use the trained dual self-attention mechanism neural network to process the molecular formula of the drug molecule to obtain the prediction result of the drug molecule properties; Among them, the trained dual self-attention mechanism neural network includes: a first self-attention mechanism branch and a second self-attention mechanism branch in parallel; input the molecular formula of the drug molecule into the first self-attention mechanism branch, and output the first molecular descriptor; the working process of the first self-attention mechanism branch includes: Obtain the hidden information of the chemical bond between atom b and all adjacent atoms through chemical information aggregation; where chemical information aggregation refers to aggregating multiple collected chemical vector information into one chemical vector information by addition or splicing; Process the chemical information of atom b and the hidden information of the chemical bond between atom b and all adjacent atoms through the self-attention mechanism layer respectively, sum the output results processed by the self-attention mechanism, and process the sum result through the activation function to obtain the molecular hidden information of the subunit where atom b is located; Similarly, obtain the molecular hidden information of the subunits where other atoms of the drug molecule are located; Sum the molecular hidden information of the subunits where all atoms are located to obtain the molecular descriptor K; The first molecular descriptor includes: the molecular hidden information inside each atom subunit in the molecular formula; input the molecular formula of the drug molecule into the second self-attention mechanism branch, and output the second molecular descriptor; the working process of the second self-attention mechanism branch includes: Calculate the molecular hidden information of each atom subunit; Use the self-attention mechanism to process the molecular hidden information of each atom subunit to obtain the molecular descriptor J; The second molecular descriptor includes: the molecular hidden information between all atom subunits in the molecular formula; fuse the first molecular descriptor and the second molecular descriptor to obtain the third molecular descriptor; classify the third molecular descriptor to obtain the prediction result of the drug molecule properties.

2. The drug molecule screening method based on the neural network with double self-attention mechanism according to claim 1, characterized in that, The process of processing the chemical information of atom b and the hidden information of the chemical bond between atom b and all adjacent atoms through the self-attention mechanism layer respectively, summing the output results processed by the self-attention mechanism, and processing the sum result through the activation function to obtain the molecular hidden information of the subunit where atom b is located; Specifically includes: The input to the self-attention mechanism is four such vectors. The four vectors are input simultaneously in a concatenated form. After being processed by the self-attention mechanism, the output is still four vectors of the same scale. Then, these four vectors are added pairwise, and then the activation function is applied to the addition result to obtain the molecular hidden information of the subunit where atom b is located.

3. The drug molecule screening method based on the neural network with double self-attention mechanism according to claim 1, characterized in that, The process of obtaining the hidden information of the chemical bond between atom b and all adjacent atoms through chemical information aggregation specifically includes: In the first iteration process, assume that the adjacent atom of atom b is a, then in the initial state, calculate the hidden molecular information of the chemical bond ab: Among them, represents the molecular hidden information of chemical bond ab in its initial state; represents the chemical information of atom a and the chemical information of chemical bond ab in a concatenated fusion; , is a learning matrix, and the default value of 𝑠 is 300; is determined by the size of; 𝜀 represents the Relu activation function; In step t + 1, the molecular hidden information passed to the chemical bond ab comes from the bonds around atom a, pointing to atom a, but excluding atom b; the information transfer equation is as follows: represents the molecular hidden information updated in step t; where represents the hidden information passed by the molecule to the molecular bond ab, and Q(a) represents all the atoms around atom a, represents the set of atoms around atom a except atom b; is the information transfer function, and the calculation method is as follows: Among them, ( ) , ( ) represents , and 's fusion, and the fusion method is selected as vector connection fusion. is obtained by adding the sizes of , and . 𝜀 represents the Relu activation function. is a learning matrix. is determined by the size of ( ). In step t+1, the molecular hidden information is updated to ; is calculated as follows: Represents the molecular hidden information of chemical bond ab in its initial state, represents the hidden information transferred by the molecule to molecular bond ab; is the learning matrix; 𝜀 represents the Relu activation function; after T iterations, the hidden information is obtained .

4. The drug molecule screening method based on the neural network with dual self-attention mechanism according to claim 1, characterized in that, The calculation of the molecular hidden information of each atom subunit includes: Calculate the molecular hidden information of each subunit and explain the hidden information contained in each subunit: is the chemical information of atom b, indicating the hidden information contained in all chemical bonds pointing to atom b by atom b; is a learning matrix, whose size is determined by the dimensions of; represents and the concatenated fusion of; 𝜀 represents the Relu activation function.

5. The drug molecule screening method based on the neural network with dual self-attention mechanism according to claim 1, characterized in that, The process of using the self-attention mechanism to process the molecular hidden information of each atom subunit to obtain the molecular descriptor J includes: The self-attention mechanism is used to process the molecular hidden information of each subunit and obtain the molecular descriptor J, and its formula is as follows: Among them, represents the running process of the self-attention mechanism, represents the hidden information of the subunit where atom x is located.

6. A drug molecule screening system based on a neural network with a dual self-attention mechanism, characterized in that, Including: An acquisition module, which is configured to: acquire drug molecules to be screened; A prediction module, which is configured to: process the molecular formula of the drug molecule by using a trained dual self-attention mechanism neural network to obtain a prediction result of the drug molecule property; Among them, the trained dual self-attention mechanism neural network includes: a first self-attention mechanism branch and a second self-attention mechanism branch arranged in parallel; the molecular formula of the drug molecule is input into the first self-attention mechanism branch, and a first molecular descriptor is output; the working process of the first self-attention mechanism branch includes: Through chemical information aggregation, the hidden information of the chemical bond between atom b and all adjacent atoms is obtained; among them, chemical information aggregation means that multiple collected chemical vector information is aggregated into one chemical vector information by addition or splicing; The chemical information of atom b and the hidden information of the chemical bond between atom b and all adjacent atoms are respectively processed through the self-attention mechanism layer, the output results after the self-attention mechanism processing are summed, and the sum result is processed through the activation function to obtain the molecular hidden information of the subunit where atom b is located; Similarly, the molecular hidden information of the subunits where other atoms of the drug molecule are located is obtained; The molecular hidden information of the subunits where all atoms are located is summed to obtain the molecular descriptor K; The first molecular descriptor includes: the molecular hidden information inside each atomic subunit in the molecular formula; the molecular formula of the drug molecule is input into the second self-attention mechanism branch, and a second molecular descriptor is output; the working process of the second self-attention mechanism branch includes: Calculate the molecular hidden information of each atomic subunit; Use the self-attention mechanism to process the molecular hidden information of each atomic subunit to obtain the molecular descriptor J; The second molecular descriptor includes: the molecular hidden information between all atomic subunits in the molecular formula; the first molecular descriptor and the second molecular descriptor are fused to obtain a third molecular descriptor; the third molecular descriptor is classified to obtain a prediction result of the drug molecule property.

7. An electronic device, characterized in that it includes: A memory for non-temporarily storing computer-readable instructions; And A processor for running the computer-readable instructions, Among them, when the computer-readable instructions are run by the processor, the method described in any one of the above claims 1-5 is executed.

8. A storage medium, characterized in that, Non-temporarily store computer-readable instructions, where when the non-temporary computer-readable instructions are executed by a computer, the instructions for executing the method described in any one of claims 1-5.

Citation Information

Patent Citations

  • Drug molecule screening method and system

    CN114530210A

  • Prediction method and device for ligand-protein interaction

    WO2021218791A1