Method and device for predicting affinity between protein and ligand molecule

Through the application of feature extraction and prediction model based on the three-dimensional structural diagram of protein binding to ligand molecules, the problems of inefficient computing efficiency and lack of interpretability in the prior art are solved, and efficient and accurate affinity prediction and interpretability interaction map generation are achieved.

CN115148279BActive Publication Date: 2025-05-09TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210734651.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-24
Publication Date
2025-05-09
Estimated Expiration
2042-06-24

AI Technical Summary

Technical Problem

The existing affinity prediction methods for proteins and small molecules have problems of inefficient computational efficiency and lack of interpretability, making it difficult to accurately grasp the key interactions between proteins and small molecules.

Method used

Based on the three-dimensional structural diagram of protein binding to ligand molecules, node features, edge features and geometric features are extracted, affinity prediction is performed using a pre-trained prediction model, and interaction maps are generated to achieve efficient and accurate affinity prediction and impart the model interpretability.

Benefits of technology

More efficient and accurate affinity predictions of protein and small molecules are achieved, while the generated interaction map makes the predictions interpretable and reflect the atomic-level interaction between protein and small molecules.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115148279B_ABST
    Figure CN115148279B_ABST
Patent Text Reader

Abstract

The embodiments of the present disclosure provide a method, device, equipment and computer-readable storage medium for predicting the affinity between a protein and a ligand molecule. The method provided by the embodiments of the present disclosure is based on the node features, edge features and geometric features extracted from the three-dimensional structural diagram of the binding of the protein and the ligand molecule, and obtains the affinity of the protein and the ligand molecule through a pre-trained prediction model, and obtains an interaction graph for indicating the interaction between the atoms of the protein and the ligand molecule. On the basis of improving the affinity prediction performance, it is possible to judge whether the predicted atomic interaction between the protein and the ligand molecule is correct, so that the prediction result is interpretable. Among them, the prediction model is obtained by error correction of affinity prediction and interaction graph prediction. Therefore, through the method of the embodiments of the present disclosure, more accurate affinity prediction and interaction graph prediction can be learned on the basis of error correction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence, and more specifically, to a method, device, equipment and storage medium for predicting the affinity between a protein and a ligand molecule. Background Art

[0002] The interaction between proteins and small molecules is the basis for drug design and development. In-depth research on the binding mechanism between proteins and drug molecules at the molecular level can help to quickly screen out effective drug candidate molecules, greatly shorten the new drug development process, and reduce the risk of new drug failure. Therefore, it is very necessary to study the interaction between proteins and small molecules. By exploring the relationship between protein molecular structure and small molecule affinity and predicting the affinity between proteins and small molecules, it is possible to quickly screen effective drug candidate molecules in batches, thereby accelerating the process of drug development and reducing the cost of drug development.

[0003] Existing technologies for predicting protein-small molecule affinity include using a three-dimensional (3D) convolutional neural network (CNN) model to divide the 3D structure of proteins and small molecules into three-dimensional rectangular grids, and using various chemical information blocks encoded in each grid as input to the 3D-CNN. In addition, in order to further improve the accuracy and generalization ability of deep learning-based methods to predict the interaction between proteins and small molecules, algorithms based on molecular graphs that add three-dimensional structural information (i.e., graph neural network (GNN) algorithms) have also emerged to achieve protein-small molecule affinity prediction. However, these technologies have some significant disadvantages. For example, for 3D CNNs trained on three-dimensional structural grids, the three-dimensional rectangular grid points are some high-dimensional and sparse three-dimensional matrices, resulting in low computational efficiency and difficulty in capturing key interactions. The existing GNN model is not interpretable and cannot reflect the key interactions between proteins and small molecules.

[0004] Therefore, an efficient and accurate method for predicting the affinity between proteins and small ligand molecules is needed. Summary of the invention

[0005] In order to solve the above problems, the present invention determines the affinity of protein and ligand molecule and generates an interaction map based on the three-dimensional structural diagram of the binding of protein and ligand molecule, thereby achieving efficient and accurate affinity prediction, and the generated interaction map makes the model interpretable.

[0006] Embodiments of the present disclosure provide a method, apparatus, device and computer-readable storage medium for predicting affinity between a protein and a ligand molecule.

[0007] An embodiment of the present disclosure provides a method for predicting the affinity of a protein and a ligand molecule, comprising: obtaining a three-dimensional structural diagram of the binding of a protein and a ligand molecule, wherein the three-dimensional structural diagram has atoms of the protein and the ligand molecule as nodes; determining node features of atoms of each of the protein and the ligand molecule from the three-dimensional structural diagram, and determining edge features and geometric features of the three-dimensional structural diagram based on each node in the three-dimensional structural diagram; and determining the affinity of the protein and the ligand molecule through a pre-trained prediction model based on the node features, edge features and geometric features of the three-dimensional structural diagram, and obtaining an interaction diagram of the protein and the ligand molecule, wherein the interaction diagram is used to indicate the interaction between atoms of the protein and the ligand molecule; wherein the prediction model is trained by error correction of affinity prediction and interaction diagram prediction.

[0008] An embodiment of the present disclosure provides an affinity prediction device for a protein and a ligand molecule, comprising: a data acquisition module, configured to acquire a three-dimensional structure diagram of the binding of a protein and a ligand molecule, wherein the three-dimensional structure diagram has atoms of the protein and the ligand molecule as nodes; a feature extraction module, configured to determine node features of atoms of each of the protein and the ligand molecule from the three-dimensional structure diagram, and determine edge features and geometric features of the three-dimensional structure diagram based on each node in the three-dimensional structure diagram; and a prediction module, configured to determine the affinity of the protein and the ligand molecule based on the node features, edge features and geometric features of the three-dimensional structure diagram through a pre-trained prediction model, and obtain an interaction diagram of the protein and the ligand molecule, wherein the interaction diagram is used to indicate the interaction between atoms of the protein and the ligand molecule; wherein the prediction model is trained by error correction of affinity prediction and interaction diagram prediction.

[0009] An embodiment of the present disclosure provides an affinity prediction device for a protein and a ligand molecule, comprising: one or more processors; and one or more memories, wherein a computer executable program is stored in the one or more memories, and when the computer executable program is executed by the processor, the affinity prediction method for a protein and a ligand molecule as described above is executed.

[0010] An embodiment of the present disclosure provides a computer-readable storage medium having computer-executable instructions stored thereon, which, when executed by a processor, are used to implement the above-mentioned method for predicting the affinity between a protein and a ligand molecule.

[0011] The embodiments of the present disclosure provide a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method for predicting the affinity between a protein and a ligand molecule according to the embodiments of the present disclosure.

[0012] Compared with the existing protein and small molecule affinity prediction methods, the method provided by the embodiments of the present disclosure can achieve protein and small molecule affinity prediction more efficiently and accurately, while generating an interaction map that can reflect the atomic interaction between proteins and small molecules, making the prediction results of the method provided by the embodiments of the present disclosure explainable.

[0013] The method provided by the embodiments of the present disclosure is based on node features, edge features and geometric features extracted from the three-dimensional structural diagram of the binding of proteins and ligand molecules. The affinity of proteins and ligand molecules is obtained through a pre-trained prediction model, and an interaction diagram for indicating the interaction between atoms of proteins and ligand molecules is obtained. On the basis of improving the affinity prediction performance, it is possible to determine whether the predicted atomic interaction between proteins and ligand molecules is correct, so that the prediction results are interpretable. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some exemplary embodiments of the present disclosure, and a person of ordinary skill in the art can obtain other drawings based on these drawings without creative work.

[0015] Figure 1 is a schematic diagram of a scenario showing the processing of an affinity prediction request initiated from a user terminal according to an embodiment of the present disclosure;

[0016] Figure 2 is a flow chart showing a method 200 for predicting affinity between a protein and a ligand molecule according to an embodiment of the present disclosure;

[0017] Figure 3 is a schematic flowchart showing a method for predicting affinity between a protein and a ligand molecule according to an embodiment of the present disclosure;

[0018] Figure 4 is a schematic diagram showing the determination of an interaction graph based on the attention vectors of the protein and the ligand molecule respectively and the error between the interaction graph and the true interaction graph according to an embodiment of the present disclosure;

[0019] Figure 5 is a flow chart illustrating a training prediction model according to an embodiment of the present disclosure;

[0020] Fig. 6A is a schematic diagram showing the role of cofactor molecules in the binding of proteins and ligand molecules according to an embodiment of the present disclosure;

[0021] Figure 6B is a schematic diagram showing the predicted results of the interaction map according to an embodiment of the present disclosure and the actual binding structure of the protein ligand molecule;

[0022] Figure 7 is a schematic diagram showing a device for predicting affinity between a protein and a ligand molecule according to an embodiment of the present disclosure;

[0023] Figure 8 A schematic diagram showing an affinity prediction device for a protein and a ligand molecule according to an embodiment of the present disclosure is shown;

[0024] Fig. 9 A schematic diagram illustrating the architecture of an exemplary computing device according to an embodiment of the present disclosure; and

[0025] Fig.10 A schematic diagram of a storage medium according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0026] In order to make the purpose, technical solution and advantages of the present disclosure more obvious, the exemplary embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all the embodiments of the present disclosure, and it should be understood that the present disclosure is not limited to the exemplary embodiments described here.

[0027] In this specification and the accompanying drawings, substantially the same or similar steps and elements are represented by the same or similar reference numerals, and repeated descriptions of these steps and elements will be omitted. At the same time, in the description of the present disclosure, the terms "first", "second", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance or ranking.

[0028] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which the present disclosure pertains. The terms used herein are only for the purpose of describing the embodiments of the present invention and are not intended to limit the present invention.

[0029] To facilitate description of the present disclosure, concepts related to the present disclosure are introduced below.

[0030] The affinity prediction method of the protein and ligand molecule disclosed in the present invention can be based on artificial intelligence (AI). Artificial intelligence is a theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. For example, for the affinity prediction method of protein and ligand molecule based on artificial intelligence, it can find the atomic pairs with interaction and determine the affinity of protein and configuration molecule in a way similar to that of human beings identifying the atomic interaction between protein and ligand molecule by naked eyes. Artificial intelligence enables the affinity prediction method of protein and ligand molecule disclosed in the present invention to quickly and accurately determine the contribution of each atom in protein and ligand molecule to its binding affinity and determine the interaction between atoms therefrom.

[0031] The affinity prediction method of the protein and ligand molecule disclosed in the present invention can be based on deep learning. Deep learning is an algorithm in machine learning based on characterization and learning of data. Observations (for example, an image) can be represented in a variety of ways, such as a vector of intensity values ​​of each pixel, or more abstractly represented as a series of edges, regions of specific shapes, etc. Using certain specific representation methods makes it easier to learn tasks from examples (for example, image recognition). The benefit of deep learning is to replace manual feature acquisition with efficient algorithms for unsupervised or semi-supervised feature learning and hierarchical feature extraction. Among them, optionally, the affinity prediction method of the protein and ligand molecule disclosed in the present invention can be based on a graph neural network. Graph neural network is a framework that uses deep learning to directly learn graph structure data in recent years, and its excellent performance has attracted high attention and in-depth exploration. By formulating certain strategies on the nodes and edges in the graph, GNN converts graph structure data into a standardized and standard representation, and inputs it into a variety of different neural networks for training, achieving excellent results in tasks such as node classification, edge information propagation, and graph clustering. In the method disclosed in the present invention, a graph is a data structure that is very suitable for characterizing molecules. The two component structures of nodes and edges correspond to atoms and chemical bonds in molecules, respectively. Unlike the rasterization method that defines a regular cubic range, the number of nodes and edges in the graph is not limited, and molecules of different sizes can be flexibly and completely represented. Therefore, a graph neural network can be used to process the irregular topological relationship structure between proteins and ligand molecules, which requires that the molecular data be represented as a graph before entering the network. Therefore, before entering the graph neural network processing, the three-dimensional structure graph of proteins and ligand molecules can be represented as consisting of node feature vectors, edge feature vectors, and connection relationships between nodes.

[0032] Optionally, the protein-ligand affinity prediction method disclosed in the present invention can apply the attention mechanism to the graph neural network to determine the attention weight of each atom based on the self-attention mechanism, that is, its contribution to the affinity of the protein and ligand molecule binding. The essence of the attention mechanism originates from the human visual mechanism and belongs to the brain signal processing mechanism unique to human vision, that is, human attention.

[0033] In addition, the terms that may be involved in the method for predicting affinity between a protein and a ligand molecule disclosed in the present invention are described below.

[0034] PLIP (Protein-Ligand Interaction Profiler): It is an analytical tool for the non-covalent interaction between proteins and ligand molecules. It can analyze the non-covalent interactions between proteins and ligand molecule complexes at the atomic level, including hydrogen bonds, water bridges, salt bridges, halogen bonds, hydrophobic interactions, π-stacking, π-ion interactions and metal complexes. Its detection mechanism is mainly based on the spatial position and geometric relationship between atoms. In the embodiments of the present disclosure, the correct interaction relationship between proteins and ligand molecules is obtained by the PLIP tool, so as to supervise the learning of the interaction graph obtained by the prediction model, so as to obtain more accurate predictions.

[0035] Docking: A method that predicts the most likely conformation of a small molecule when it binds to a target protein to form a stable complex through physical simulation or computational chemistry.

[0036] Protein-Ligand Complex: A co-crystal structure of a ligand molecule bound to a protein or a three-dimensional complex structure generated from a protein and a ligand molecule by a Docking method. In an embodiment of the present disclosure, the protein-ligand complex can be constructed in the form of a graph with atoms as nodes as input to the affinity prediction method of the protein and ligand molecule of the present disclosure.

[0037] Pocket: A structure on the surface or internal cavity of a protein that is used to bind molecules or peptides to produce biological reactions. For example, a protein pocket may include a small molecule 5 angstroms away from the protein. Protein amino acids.

[0038] In summary, the solutions provided by the embodiments of the present disclosure involve technologies such as artificial intelligence and graph neural networks. The embodiments of the present disclosure will be further described below in conjunction with the accompanying drawings.

[0039] Figure 1 is a schematic diagram of a scenario showing the processing of an affinity prediction request initiated from a user terminal according to an embodiment of the present disclosure.

[0040] exist Figure 1 In the application, a user can initiate an affinity prediction request through his user terminal, for example, by uploading a three-dimensional structure diagram of a bound protein and ligand molecule through a specific interface on his user terminal. The user terminal can then transmit these three-dimensional structure data to the server of the application through the network (or directly) for processing.

[0041] Optionally, the user terminal may specifically include a smart phone, a tablet computer, a laptop, a vehicle-mounted terminal, a wearable device, and the like. The user terminal may also be a client that installs a browser or various applications (including system applications and third-party applications). The network may be an Internet of Things based on the Internet and / or telecommunications network, which may be a wired network or a wireless network, for example, it may be an electronic network that can realize information exchange functions such as a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a cellular data communication network, etc. The server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.

[0042] like Figure 1 As shown, the server can perform affinity prediction in real time based on the received data (e.g., the three-dimensional structure diagram data of the protein and the ligand molecule). Subsequently, the server can return the obtained affinity prediction result (e.g., the affinity value of the protein and the ligand molecule or the activity of the ligand molecule) to the user terminal through the network as a response to the user's affinity prediction request.

[0043] In fact, the user's affinity prediction request can usually be used in many tasks such as actual drug development. For example, the affinity prediction task can be used for ligand molecule screening for a specific protein, and the ligand molecule screening can be based on the interaction strength between the ligand molecule and the protein (that is, the activity of the ligand molecule). The interaction between protein and ligand molecules often occurs in many basic biological activities. Understanding the interaction between protein and ligand is of great significance for understanding many biological systems and assisting drug development. For example, in actual drug development, only a small number of molecules in a large number of molecules in the molecular library may have therapeutic significance for the target protein, and it is very challenging to find drugs that specifically bind to the target from a large number of molecules. Although high-throughput screening experimental technology can test a large number of molecules on the target protein, it takes a lot of time and cost. Therefore, the hit rate can be improved by predicting highly active molecules with strong interactions between proteins and ligands based on molecular structure.

[0044] At present, researchers have accumulated a lot of effective experience in the field of protein-ligand binding affinity calculation and proposed many calculation methods, but there are still some shortcomings. Among them, the existing technology for predicting the affinity of proteins and ligand molecules using a three-dimensional convolutional neural network (CNN) model divides the three-dimensional structure of the protein and ligand molecule into three-dimensional rectangular grids, and uses various chemical information blocks encoded in each grid as the input of the three-dimensional convolutional neural network model to determine the affinity value of the protein and the ligand molecule. The deep learning method provides a new idea for the affinity prediction method through end-to-end learning of complex neural networks. This type of research currently focuses on the combination of different molecular encoding methods with different convolutional neural network models. Its prediction effect is higher than that of traditional scoring functions and machine learning methods, but there is still room for improvement in prediction accuracy. Therefore, in order to further improve the accuracy and generalization ability of deep learning-based methods to predict the interaction between proteins and small molecules, an algorithm based on a molecular graph with three-dimensional structural information (i.e., a graph neural network (GNN) algorithm) has also emerged to achieve protein-small molecule affinity prediction. However, these technologies have some significant disadvantages. For example, for 3D CNN trained on a 3D structure grid, the 3D rectangular grid points are some high-dimensional sparse 3D matrices, which leads to low computational efficiency and makes it difficult to capture key interactions. In addition, although the existing GNN model can determine the key atoms based on a single protein and ligand molecule, it is not interpretable and cannot reflect the key interactions between proteins and small molecules. For example, for Figure 1 The affinity prediction results shown in the figure show that the output of the existing GNN model can only reflect the interaction strength (or affinity value) between the protein and the ligand molecule, but cannot specifically reflect the atomic-level interaction between the protein and the ligand molecule. In addition, the output results are difficult to verify their accuracy due to their poor interpretability.

[0045] Based on this, the present disclosure provides a method for predicting the affinity of proteins and ligand molecules, which determines the affinity of proteins and ligand molecules and generates an interaction map based on the three-dimensional structural diagram of the binding of proteins and ligand molecules, thereby achieving efficient and accurate affinity prediction, and the generated interaction map makes the model interpretable.

[0046] Compared with the existing protein and small molecule affinity prediction methods, the method provided by the embodiments of the present disclosure can achieve protein and small molecule affinity prediction more efficiently and accurately, while generating an interaction map that can reflect the atomic interaction between proteins and small molecules, making the prediction results of the method provided by the embodiments of the present disclosure explainable.

[0047] The method provided by the embodiment of the present disclosure is based on the node features, edge features and geometric features extracted from the three-dimensional structure diagram of the protein and ligand molecule binding, and obtains the affinity of the protein and the ligand molecule through a pre-trained prediction model, and obtains an interaction graph for indicating the interaction between the atoms of the protein and the ligand molecule. On the basis of improving the affinity prediction performance, it is possible to judge whether the predicted atomic interaction between the protein and the ligand molecule is correct, so that the prediction result is interpretable. Among them, the prediction model is obtained by error correction of affinity prediction and interaction graph prediction, so the method of the embodiment of the present disclosure can learn more accurate affinity prediction and interaction graph prediction on the basis of error correction.

[0048] Figure 2 is a flow chart showing a method 200 for predicting affinity between a protein and a ligand molecule according to an embodiment of the present disclosure. Figure 3 It is a schematic flowchart showing a method for predicting the affinity between a protein and a ligand molecule according to an embodiment of the present disclosure.

[0049] like Figure 2 As shown, in step 201, a three-dimensional structure diagram of the binding of the protein and the ligand molecule can be obtained, wherein the three-dimensional structure diagram uses the atoms of the protein and the ligand molecule as nodes.

[0050] As described above, the three-dimensional structure diagram of the protein binding to the ligand molecule can be a three-dimensional structure diagram of the protein-ligand complex generated based on the protein and the ligand molecule using a molecular docking method, which can be an atomic topology diagram obtained by conformation of the protein-ligand complex. The affinity prediction method of the protein and the ligand molecule disclosed in the present invention can use the three-dimensional structure diagram as input, such as Figure 3 As shown. Optionally, in the three-dimensional structure diagram, each atom of the protein or ligand molecule can be used as each node in the structure, and the atomic pairs (i.e., node pairs) composed of two atoms in the three-dimensional structure constitute the two vertices of the edge of the three-dimensional structure diagram. Optionally, the co-crystal structure of the protein-ligand complex can be a suitable pose selected from multiple docking poses by evaluating the accuracy of the docking pose, and the selection of the pose can be included in the prediction model optimization of the affinity prediction method of the protein and ligand molecule disclosed in the present invention.

[0051] like Figure 3 As shown, based on the three-dimensional structure map, various spatial features of protein binding to ligand molecules can be determined. These spatial features can be obtained in different ways and used in the affinity prediction process in different ways.

[0052] In step 202, node features of atoms of each of the protein and the ligand molecule may be determined from the three-dimensional structure graph, and edge features and geometric features of the three-dimensional structure graph may be determined based on each node in the three-dimensional structure graph.

[0053] Optionally, the features extracted from the three-dimensional spatial structure of the protein-ligand protein binding may include node features (i.e., atomic features), edge features, and geometric features. Among them, the node features may be independently determined based on the properties of each atom, while the edge features and geometric features may be determined based on the connections between atom pairs (node ​​pairs) and the spatial relationship between these connections.

[0054] like Figure 3 As shown, the node features extracted from the three-dimensional structure graph may include multi-dimensional features of atoms corresponding to each component (such as a ligand molecule and a protein molecule) in the three-dimensional structure, for example, including but not limited to the type of atom, whether the atom is an aromatic ring atom, whether the atom is a chiral atom, etc. The node features extracted from each component may be spliced ​​together to form a one-dimensional feature vector of the three-dimensional structure graph, which is input into a prediction model for affinity prediction.

[0055] According to an embodiment of the present disclosure, determining the edge features and geometric features of the three-dimensional structure graph based on each node in the three-dimensional structure graph may include: establishing a distance graph of the three-dimensional structure graph based on each node in the three-dimensional structure graph, the distance graph indicating the distance between each node pair in the three-dimensional structure graph; and determining the edge features and geometric features of the three-dimensional structure graph based on the three-dimensional structure graph and its distance graph, wherein the edge features may indicate the covalent bond features between corresponding node pairs in the three-dimensional structure graph, and the geometric features may indicate the spatial relationship between the edges constituted by each node pair in the three-dimensional structure graph.

[0056] Optionally, the edge features and geometric features of the three-dimensional structure graph can be based on Figure 3The distance graph (distance matrix) shown in the figure can be determined, and the distance graph can give the corresponding atomic pair distance for each atomic pair in the three-dimensional structure, so that the edge features and geometric features in the three-dimensional structure can be extracted according to part or all of the information in the distance graph. For example, the edge features of the three-dimensional structure graph may include the features of the covalent bonds between the corresponding atomic pairs, such as the type of covalent bonds and whether the covalent bonds are in a ring, etc., and the geometric features of the three-dimensional structure graph may include such as the covalent bond angle (for example, the angle between the covalent bonds formed by each atom and the two nearest atoms), the interaction angle (for example, with each atom as the center (denoted as B), find the nearest covalently bonded atom (denoted as A), and then find the nearest atom (denoted as C) of another relative molecule (for example, if the atom is from a small molecule compound, find the nearest amino acid molecule) to form an angle ∠ABC), local charge (for example, the local charge value corresponding to the atoms of the ligand molecule), etc.

[0057] Of course, the specific contents of the node features, edge features and geometric features of the above-mentioned three-dimensional structure graph are only used as examples in the method disclosed in the present invention. The present invention does not limit the specific features extracted, and other more or fewer different features can also be used in the affinity prediction of the present invention to make the affinity prediction results more accurate.

[0058] like Figure 3 As shown, the edge features extracted from the distance graph generated according to the three-dimensional structure graph can be combined with the geometric features transformed by the RBF kernel (Radial basis function kernel) as auxiliary features for affinity prediction based only on node features, and input into the prediction model for affinity prediction.

[0059] In step 203, based on the node features, edge features and geometric features of the three-dimensional structure graph, the affinity between the protein and the ligand molecule can be determined through a pre-trained prediction model, and an interaction map between the protein and the ligand molecule can be obtained. The interaction map can be used to indicate the interaction between the atoms of the protein and the ligand molecule.

[0060] Optionally, the similarity between each atom and other atoms in the three-dimensional structure of the protein and ligand molecule can be preliminarily determined based on the atomic characteristics of each atom in the three-dimensional structure of the protein and ligand molecule. Then, the combination of the above-mentioned edge characteristics and geometric characteristics can be used as an aid to jointly determine the importance of each atom to the affinity of the protein and ligand molecule binding with the determined similarities between atoms, that is, the importance of each atom in the atomic interaction between the protein and the ligand molecule.

[0061] According to an embodiment of the present disclosure, the prediction model can adopt a self-attention mechanism. Therefore, the importance of each atom in the atomic interaction between the protein and the ligand molecule can be determined by the attention weight determined based on the self-attention mechanism, and the similarity between atoms can be determined by the input of the prediction model (i.e., a one-dimensional feature vector composed of the atomic features of each atom in the binding of the protein molecule and the ligand molecule, such as Figure 3 As shown in Figure 2, the query vector, key vector and value vector of the self-attention mechanism are used as the query vector, key vector and value vector, supplemented by the edge features and geometric features of the three-dimensional structure graph.

[0062] Specifically, according to an embodiment of the present disclosure, step 203 may include: based on the node features, edge features and geometric features of the three-dimensional structure graph, determining the affinity between the protein and the ligand molecule, and the attention vectors of the protein and the ligand molecule respectively through a self-attention mechanism, wherein each element in the attention vector indicates the contribution of the corresponding atom to the affinity between the protein and the ligand molecule; and based on the attention vectors of the protein and the ligand molecule respectively, obtaining an interaction graph between the protein and the ligand molecule, wherein each element in the interaction graph indicates the possibility of interaction between the corresponding atomic pairs of the protein and the ligand molecule.

[0063] Optionally, in a pre-trained prediction model, the similarity between each atom in the three-dimensional structure graph and other atoms can be determined based on the node features, edge features and geometric features of the three-dimensional structure graph through a self-attention mechanism, that is, the attention weight of each atom in each component (such as ligand molecules and protein molecules) in the three-dimensional structure graph, which indicates the contribution of the corresponding atom to the binding of the protein and the ligand molecule, that is, the strength of the interaction between the corresponding atom and other atoms in the three-dimensional structure graph.

[0064] Therefore, for each component in the three-dimensional structure graph, the attention weights corresponding to all its atoms can constitute its attention vector. According to an embodiment of the present disclosure, based on the node features, edge features and geometric features of the three-dimensional structure graph, determining the attention vectors of the protein and the ligand molecule through a self-attention mechanism can include: splicing the node features of the atoms of the protein and the ligand molecule into a one-dimensional feature vector; and based on the self-attention mechanism, using the one-dimensional feature vector as a query vector, a key vector and a value vector, and combining the edge features and geometric features of the three-dimensional structure graph, determining the attention weight of each node in the three-dimensional structure graph, wherein the attention weight can indicate the contribution of the node to the affinity between the protein and the ligand molecule; wherein the attention weights of the atoms of the protein and the ligand molecule can constitute their respective attention vectors.

[0065] like Figure 3 As shown, the node features of all atoms in the three-dimensional structure diagram can be spliced ​​into a one-dimensional feature vector and input into a pre-trained prediction model. Optionally, the prediction model can be a Transformer model based on the self-attention mechanism, and the one-dimensional feature vector can therefore be used as the query vector (query, Q), key vector (key, K) and value vector (value, V) of the prediction model to learn the relationship between the atomic features within the one-dimensional feature vector. In order to ensure the diversity of features, different multiple linear transformation layers can be applied to process Q, K and V.

[0066] Next, the determination of attention weights may be performed based on the obtained Q, K, and V vectors. Optionally, the attention weights may be determined based on the dot-product of the query vector Q and the key vector K, and the determined attention weights may be normalized (e.g., using a normalization function such as softmax). For example, Figure 3 As shown, the obtained query vector Q and key vector K can be matrix multiplied, and the result of the matrix dot multiplication can be scaled to avoid the gradient tending to 0 (gradient vanishing) due to the large input order of the normalization function when the order of magnitude of the matrix dot multiplication result is too large. The scaling process can make the distribution of the normalized attention weights more uniform. Optionally, before the normalization process is performed to obtain the attention weights, the scaled matrix dot multiplication result can be combined with the auxiliary features composed of the edge features and geometric features of the three-dimensional structure graph (for example, the scaled matrix dot multiplication result is added to the auxiliary feature matrix determined based on the edge features and geometric features of the three-dimensional structure graph) to consider more possible influencing factors for the affinity prediction process, so that the affinity prediction is more accurate.

[0067] Therefore, as described above, by normalizing the combination of the scaled matrix dot product result and the auxiliary feature matrix (determined based on the edge features and geometric features of the three-dimensional structure graph), an attention vector for predicting the affinity between the protein and the ligand molecule can be determined, which includes an attention weight corresponding to each atom in the three-dimensional structure graph, indicating the strength of the interaction of the corresponding atom with other atoms in the binding of the protein to the ligand molecule.

[0068] Optionally, the features of each atom can be updated based on the attention weight of each atom in the determined three-dimensional feature map, so that the updated atomic features are more beneficial for determining the interaction strength between the atom and other atoms, that is, more beneficial for predicting the affinity of the protein binding to the ligand molecule. That is, in the prediction model, the process of determining the attention weight based on the obtained Q, K and V vectors and determining a new one-dimensional feature vector based on the determined attention weight can be performed multiple times to obtain better attention weight and affinity prediction results.

[0069] According to an embodiment of the present disclosure, determining the affinity of the protein to the ligand molecule through a self-attention mechanism based on the node features, edge features and geometric features of the three-dimensional structure graph may include: based on the node features, edge features and geometric features of the three-dimensional structure graph, updating the one-dimensional feature vector multiple times, wherein in each update: using the last updated one-dimensional feature vector as the query vector, key vector and value vector, and combining the edge features and geometric features of the three-dimensional structure graph, determining the attention weight of each node in the three-dimensional structure graph; determining an updated one-dimensional feature vector based on the attention weight of each node in the three-dimensional structure graph and the value vector; and determining the affinity of the protein to the ligand molecule based on the one-dimensional feature vector that has been updated multiple times.

[0070] Optionally, the number of updates to the feature vector can be determined based on actual needs (e.g. Figure 3 6 updates as shown). In each feature vector update, as described above, after the attention vector for affinity prediction between the protein and the ligand molecule is determined, the attention vector can be used to update the value vector V (for example, matrix multiplication of the attention vector and the value vector V) to obtain an updated one-dimensional feature vector, and each element in the one-dimensional feature vector can still be in a one-to-one correspondence with each atom in the three-dimensional structure diagram.

[0071] Therefore, after multiple updates as mentioned above, the attention weights and one-dimensional feature vectors used for the final affinity prediction and interaction graph determination can be determined.

[0072] Optionally, for the final affinity prediction, the determined one-dimensional feature vector may be input into the task layer to output the affinity prediction result. For example, the task layer may perform a linear transformation (e.g., weighted summation) on the one-dimensional feature vector based on the trained weights to obtain the affinity prediction result (e.g., a one-dimensional affinity prediction value).

[0073] Regarding the final interaction graph determination, according to an embodiment of the present disclosure, based on the respective attention vectors of the protein and the ligand molecule, obtaining the interaction graph of the protein and the ligand molecule may include: for any atomic pair of the protein and the ligand molecule, based on the product of the corresponding attention weights in the respective attention vectors of the protein and the ligand molecule, determining the corresponding element in the interaction graph, and the corresponding element may correspond to the atomic pair.

[0074] like Figure 3As shown, the attention weights output from the prediction model can be combined into different attention vectors based on the various components in the three-dimensional structure diagram. For example, the attention weights corresponding to the atoms belonging to the protein molecule constitute the attention vector of the protein molecule, while the attention weights corresponding to the atoms belonging to the ligand molecule constitute the attention vector of the ligand molecule.

[0075] Therefore, based on the attention vectors of the protein molecule and the ligand molecule, the interaction graph of the binding of the protein and the ligand molecule can be determined. For example, the interaction graph can be determined based on the product of the attention vectors of the protein molecule and the ligand molecule, where each element is the product of the corresponding element in the attention vector of the protein molecule and the corresponding element in the attention vector of the ligand molecule, such as Figure 3 shown.

[0076] Specifically, Figure 4 is a schematic diagram showing the determination of an interaction graph based on the respective attention vectors of a protein and a ligand molecule and the error between the interaction graph and the true interaction graph according to an embodiment of the present disclosure.

[0077] like Figure 4 As shown, the attention vectors of the protein molecule and the ligand molecule are shown respectively, where each rectangular grid corresponds to an atom. Therefore, each element in the interaction graph can be the product of the attention weight of the protein atom at the corresponding position and the attention weight of the ligand atom. For example, for the element in the second row and second column of the interaction graph, it corresponds to the atom pair composed of the second atom of the protein molecule and the ligand molecule respectively (assuming that the atoms in the protein molecule and the ligand molecule are sorted in advance, that is, the atomic sorting of each component in the input one-dimensional feature vector), the value of the element is the product of the attention weights of the atom pair, and the value of the element belongs to the range [0,1].

[0078] As described above, the protein-ligand affinity prediction method disclosed herein can not only output accurate affinity prediction results, but also output an interaction map between the protein and the ligand molecule, which indicates the atomic-level interaction between the protein and the ligand molecule.

[0079] Alternatively, in the real interaction diagram between the protein and the ligand molecule, the elements corresponding to the atomic pairs with non-covalent interactions can be assigned a value of 1 ( Figure 4 ), otherwise 0 ( Figure 4 Therefore, the error between the true interaction map and the interaction map obtained above can be expressed as:

[0080] -(z j log(p(z j))+(1-z j )log(1-p(z j ))) (1)

[0081] Among them, z j represents the value of the jth element in the true interaction graph, and p(z j ) represents the value of the jth element in the interaction graph obtained above, where the jth element in the true interaction graph and the jth element in the interaction graph obtained above correspond to the same atomic pair of the protein and the ligand molecule (e.g. Figure 4 (shown in the bold dashed box in the figure).

[0082] Therefore, as described above, the interaction map can predict whether there is an interaction between the protein and the ligand molecule, check whether important atomic interactions are found, and explain the correctness of the affinity prediction results (for example, whether the prediction results are reasonable), so that the affinity prediction model is interpretable. In addition, when the true interaction map between the protein and the ligand molecule can be obtained (for example, calculated by the PLIP tool), the correctness of the prediction results can be evaluated based on the error between the true interaction map and the interaction map obtained above, and the above prediction model can be error corrected. According to an embodiment of the present disclosure, the prediction model can be trained by error correction of affinity prediction and interaction map prediction, which will be specifically explained in the following description of prediction model training.

[0083] According to an embodiment of the present disclosure, the method for predicting the affinity between a protein and a ligand molecule of the present disclosure further includes a step 204 for training a prediction model, wherein the step 204 may include: Figure 5 Steps 2041-2045 are shown. Figure 5 is a flow chart illustrating a training prediction model according to an embodiment of the present disclosure.

[0084] like Figure 5 As shown, in step 2041, a plurality of samples of three-dimensional structure diagrams of proteins binding to ligand molecules may be obtained. Optionally, different three-dimensional structure diagrams formed by different proteins binding to ligand molecules may be obtained for training the prediction model of the present disclosure, so that the prediction model of the present disclosure may be applicable to a wider range of application scenarios.

[0085] In step 2042, for each of the three-dimensional structure graph samples of the plurality of three-dimensional structure graph samples of the binding of proteins to ligand molecules, node features, edge features, and geometric features of the three-dimensional structure graph sample may be determined. As described above, step 2042 may determine the node features, edge features, and geometric features of the three-dimensional structure graph samples in the same manner as described with reference to step 202 as inputs to the prediction model to be trained.

[0086] In step 2043, the real affinity and real interaction map corresponding to the three-dimensional structure map sample can be obtained, wherein each element in the real interaction map indicates whether there is an interaction between the corresponding atomic pairs of the protein in the three-dimensional structure map sample and the ligand molecule.

[0087] As described above, the true affinities and true interactions of these different three-dimensional structure maps can be predetermined to serve as prior information for supervised learning of the affinity prediction results and interaction maps obtained by the prediction model.

[0088] In step 2044, the affinity and interaction graph corresponding to the three-dimensional structure graph sample can be determined through a prediction model based on the node features, edge features and geometric features of the three-dimensional structure graph sample, wherein the determined affinity and interaction graph contains parameters to be optimized of the prediction model.

[0089] Similar to the description of the reference method 200 above, the affinity and interaction graph corresponding to the three-dimensional structure graph sample can be determined by the prediction model based on the node features, edge features and geometric features of the three-dimensional structure graph sample. At this time, the determined affinity and interaction graph is used to compare with the real affinity and interaction graph and optimize the parameters of the prediction model based on error correction, wherein the parameters can also include Figure 3 Parameters of the task layer are shown.

[0090] In step 2045, the parameters to be optimized of the prediction model can be determined by optimizing the affinity prediction error between the true affinity and the determined affinity corresponding to each three-dimensional structure graph sample of the multiple protein-ligand molecule binding three-dimensional structure graph samples, and the interaction graph prediction error between the true interaction graph and the determined interaction graph, so as to obtain the pre-trained prediction model.

[0091] Optionally, the error between the determined affinity and interaction graph and the true affinity and interaction graph can be used as the loss objective function for prediction model optimization, that is, the objective function of prediction model optimization can include a combination of the objective function of affinity prediction and the objective function of interaction prediction.

[0092] For example, for N 3D structure graph samples, the objective function of affinity prediction is It can be expressed as:

[0093]

[0094] Among them, for the i-th three-dimensional structure graph sample, y i is the true affinity value between protein and ligand molecule, f(x i ) is the affinity prediction value, where x i A one-dimensional feature vector representing the input.

[0095] The objective function L for interaction prediction I It can be expressed as:

[0096]

[0097] Wherein, M represents the M atom pairs between the protein and the ligand molecule, and the error of each atom pair is shown in the above formula (1).

[0098] Therefore, the objective function of the prediction model in the method for predicting the affinity between a protein and a ligand molecule disclosed in the present invention can be expressed as:

[0099] L=L A +λL I (4)

[0100] Where λ represents the objective function L for interaction prediction I The parameter λ can be used to control the influence of the prediction error of the interaction graph on the overall prediction error.

[0101] Therefore, by training the prediction model for the above prediction error objective function based on multiple different three-dimensional structure graph samples, the atomic interaction relationship between the binding of various proteins and ligand molecules can be learned, so that the prediction model can see whether the interaction between the protein and the ligand molecule learned by the prediction model is correct on the basis of improving the affinity prediction performance, so that the prediction model has interpretability. In addition, during the training process of the prediction model, the affinity prediction error between the true affinity corresponding to each three-dimensional structure graph sample and the affinity determined by the prediction model, and the interaction graph prediction error between the true interaction graph and the interaction graph determined by the prediction model are used to perform error correction for affinity prediction and interaction graph prediction, so as to adjust the parameters to be optimized of the prediction model, so that the prediction model disclosed in the present invention can learn more accurate affinity prediction and interaction graph prediction based on the error correction.

[0102] In addition, in the embodiments of the present disclosure, it is also possible to consider adding cofactor molecules (if any) that play a key role in the interaction between the protein and the ligand molecule into the affinity prediction model. Fig. 6A Schematic diagram showing the role of cofactor molecules in the binding of proteins and ligand molecules according to an embodiment of the present disclosure.

[0103] like Fig. 6A As shown in the figure, the co-crystal structure of cytochrome P450 (CYP450) is taken as an example. The protein contains a cofactor molecule (ferroporphyrin) which is connected to the protein backbone by binding to the proximal cysteine ​​residue. Fig. 6A As can be seen from the figure, the six-membered ring of the ligand molecule forms a π-π stacking with the five-membered ring of the iron porphyrin molecule (eg Fig. 6A The cofactor molecule has strong interactions with both the protein and the ligand molecule. Therefore, when predicting the affinity between the protein and the ligand molecule, the cofactor molecule can be added to the prediction model in order to learn the real interaction between the protein and the ligand molecule.

[0104] Therefore, according to an embodiment of the present disclosure, the method for predicting the affinity of a protein and a ligand molecule of the present disclosure may further include: in a case where the binding of the protein to the ligand molecule requires the participation of a cofactor molecule, the acquired three-dimensional structure graph may further include the cofactor molecule, and the nodes of the three-dimensional structure graph may further include atoms of the cofactor molecule; and determining the node features of the atoms of the cofactor molecule from the three-dimensional structure graph, and determining the edge features and geometric features of the three-dimensional structure graph based on each node in the three-dimensional structure graph that includes the atoms of the cofactor molecule, so as to determine the affinity of the protein and the ligand molecule based on the node features, edge features and geometric features of the three-dimensional structure graph, and obtain an interaction graph between the protein and the ligand molecule.

[0105] Optionally, in the case where a cofactor molecule is involved in the binding of a protein to a ligand molecule, the obtained three-dimensional structure graph may also include the cofactor molecule, and the nodes of the three-dimensional structure graph may also include the atoms of the cofactor molecule. Therefore, the above one-dimensional feature vector may include not only the atomic features of the ligand molecule and the protein molecule, but also the atomic features of the atoms in the cofactor molecule, and similarly, the edge features and geometric features of the three-dimensional structure graph may also take into account the cofactor molecule, such as Figure 3In addition, based on the node features, edge features, and geometric features of the three-dimensional structure graph considering the cofactor molecule, the obtained attention vector can also include the attention weights of the atoms corresponding to the cofactor molecule. Since the cofactor molecule can be located on the protein molecule, the attention weight of the cofactor molecule can be merged into the attention vector of the protein molecule, as shown in Figure 3 As shown in the attention vector of the protein molecule in the upper right corner of , the last two rectangular grids correspond to the atoms of the cofactor molecule. When considering the true interaction graph and the interaction graph prediction error, the cofactor molecule can also be considered similarly, which will not be repeated in this article.

[0106] The following is a comparison between the actual binding results of proteins and ligand molecules and the prediction results of the corresponding mutual prediction graphs as an example to present the affinity prediction method of proteins and ligand molecules disclosed in the present invention. Figure 6B It is a schematic diagram showing the interaction map prediction results according to an embodiment of the present disclosure and the actual binding structure of the protein ligand molecule.

[0107] like Figure 6B As shown, Figure 6B (a) is the predicted result of the interaction diagram between protein and ligand molecule. Figure 6B (b) is the actual binding structure of the protein and ligand molecule. Figure 6B In (a), the horizontal axis corresponds to each atom in the protein molecule, while the vertical axis corresponds to each atom in the ligand molecule. Each square corresponds to the interaction strength between the corresponding atom in the protein molecule and the corresponding atom in the ligand molecule. The darker the color, the stronger the interaction, while the lighter the color, the weaker the interaction. Figure 6B In (b), the co-crystal structure shows that the nitrogen (N) in the five-membered heterocyclic ring of histidine in the protein forms a hydrogen bond with the oxygen atom of the ligand molecule (shown by a dotted line), and the hydrogen (H) and oxygen (O) of glutamine also form hydrogen bond interactions with the N and H of the ligand molecule, which can be seen from Figure 6B As can be seen from the interaction diagram shown in (a), the interaction values ​​of histidine and glutamate (as indicated by the solid line boxes) are higher.

[0108] Therefore, the above examples can show that the affinity prediction method of proteins and ligand molecules disclosed in the present invention can not only improve the accuracy of affinity prediction of proteins and ligand molecules, but also make the prediction model have a certain degree of interpretability, thereby ensuring the quality of virtual screening of drug molecules in, for example, the above-mentioned actual drug development field, discovering better and more accurate lead compounds, and thus performing subsequent lead compound optimization.

[0109] In addition, in the embodiments of the present disclosure, the affinity prediction method of the present disclosure can also be compared with other prediction models (such as the Gnina model and the S-MAN model) used to achieve the same purpose, so as to show the effectiveness of the affinity prediction method of the protein and the ligand molecule of the present disclosure. The test prediction results of these methods on these test sets are shown in the following table for the PDBbind core set test set and the internal test set (including the Normal data set and the Novel data set). Among them, the Pearson correlation coefficient r can be used to represent the accuracy of the prediction. The larger the r, the more accurate the affinity prediction result, and the better the prediction model performance. Among them, the Normal data set can contain 6776 protein-ligand molecule pairs, and the target protein is composed of common protein families (such as Kinase, GPCR, and Protease families), while the Novel data set can contain 773 protein-ligand molecule pairs, which are data collected from the most recent literature and contain multiple novel protein family data. As Figure 6B As shown, the prediction effect of the prediction method of the present invention (ie, the Pearson correlation coefficient r determined based on the scores on various data sets) is significantly better than that of other prediction methods.

[0110]

[0111] Figure 7 is a schematic diagram showing an apparatus 700 for predicting affinity between a protein and a ligand molecule according to an embodiment of the present disclosure.

[0112] According to an embodiment of the present disclosure, the protein-ligand affinity prediction device 700 may include a data acquisition module 701 , a feature extraction module 702 and a prediction module 703 .

[0113] The data acquisition module 701 may be configured to acquire a three-dimensional structure diagram of the binding of a protein to a ligand molecule, wherein the three-dimensional structure diagram uses atoms of the protein and the ligand molecule as nodes. Optionally, the data acquisition module 701 may perform the operations described above with reference to step 201 .

[0114] For example, the three-dimensional structure diagram of the protein combined with the ligand molecule can be a three-dimensional structure diagram of the protein-ligand complex generated based on the protein and the ligand molecule using methods such as molecular docking, which can be an atomic topology diagram obtained by conforming the protein-ligand complex. The affinity prediction method of the protein and the ligand molecule disclosed in the present invention can use the three-dimensional structure diagram as input. Optionally, in the three-dimensional structure diagram, each atom of the protein or ligand molecule can be used as each node in the structure, and the atomic pairs (i.e., node pairs) composed of two atoms in the three-dimensional structure constitute the two vertices of the edge of the three-dimensional structure diagram.

[0115] The feature extraction module 702 can be configured to determine node features of atoms of each of the protein and the ligand molecule from the three-dimensional structure graph, and determine edge features and geometric features of the three-dimensional structure graph based on each node in the three-dimensional structure graph. Optionally, the feature extraction module 702 can perform the operations described above with reference to step 202.

[0116] Optionally, the features extracted from the three-dimensional spatial structure of the protein and ligand protein binding may include node features (i.e., atomic features), edge features, and geometric features. Among them, the node features may be independently determined based on the properties of each atom, while the edge features and geometric features may be determined based on the connection between atom pairs (node ​​pairs) and the spatial relationship between these connections. The node features extracted from the three-dimensional structure diagram may include the multidimensional features of the atoms corresponding to each component (such as a ligand molecule and a protein molecule) in the three-dimensional structure, for example, including but not limited to the type of atom, whether the atom is an aromatic ring atom, whether the atom is a chiral atom, etc. The node features extracted from each component may be spliced ​​together to form a one-dimensional feature vector of the three-dimensional structure diagram, which is input into a prediction model for affinity prediction. Optionally, the edge features and geometric features of the three-dimensional structure diagram may be determined based on a distance map (distance matrix) generated from the three-dimensional structure, and the distance map may give corresponding atom pair distances for each atom pair in the three-dimensional structure, so that the edge features and geometric features in the three-dimensional structure may be extracted based on part or all of the information in the distance map.

[0117] The prediction module 703 can be configured to determine the affinity of the protein and the ligand molecule based on the node features, edge features and geometric features of the three-dimensional structure graph through a pre-trained prediction model, and obtain an interaction map between the protein and the ligand molecule, wherein the interaction map is used to indicate the interaction between the atoms of the protein and the ligand molecule; wherein the prediction model is trained by error correction of affinity prediction and interaction map prediction. Optionally, the prediction module 703 can perform the operations described above with reference to step 203.

[0118] For example, the similarity between each atom and other atoms in the three-dimensional structure of the protein and ligand molecule can be preliminarily determined based on the atomic features of each atom in the three-dimensional structure of the protein and ligand molecule. Then, the importance of each atom for the affinity of the protein and ligand molecule binding can be determined jointly with the determined similarity between atoms, using the combination of the above edge features and geometric features as an aid, that is, the importance of each atom in the atomic interaction between the protein and the ligand molecule. Optionally, the importance of each atom in the atomic interaction between the protein and the ligand molecule can be determined based on the attention weight determined based on the self-attention mechanism, and the similarity between atoms can be determined using the input of the prediction model (i.e., a one-dimensional feature vector composed of the atomic features of each atom in the binding of the protein molecule and the ligand molecule) as the query vector, key vector and value vector of the self-attention mechanism and supplemented by the edge features and geometric features of the three-dimensional structure graph.

[0119] Optionally, in the pre-trained prediction model, the similarity between each atom in the three-dimensional structure graph and other atoms can be determined based on the node features, edge features and geometric features of the three-dimensional structure graph through the self-attention mechanism, that is, the attention weight of each atom in each component (such as ligand molecules and protein molecules) in the three-dimensional structure graph, which indicates the contribution of the corresponding atom to the binding of the protein to the ligand molecule, that is, the intensity of the interaction between the corresponding atom and other atoms in the three-dimensional structure graph. For example, the prediction model can be a Transformer model based on the self-attention mechanism, and the one-dimensional feature vector can therefore be used as the query vector (query, Q), key vector (key, K) and value vector (value, V) of the prediction model to learn the relationship between the atomic features inside the one-dimensional feature vector. Among them, in order to ensure the diversity of features, different multiple linear transformation layers can be applied to process Q, K and V. Next, the determination of the attention weight can be performed based on the obtained Q, K and V vectors. Optionally, the attention weight can be determined based on the dot-product of the query vector Q and the bond vector K, and the determined attention weight is normalized (for example, using a normalization function such as softmax) to determine an attention vector for predicting the affinity between the protein and the ligand molecule, which includes an attention weight corresponding to each atom in the three-dimensional structure diagram, indicating the strength of the interaction of the corresponding atom with other atoms in the binding of the protein and the ligand molecule.

[0120] Optionally, the features of each atom can be updated based on the attention weight of each atom in the determined three-dimensional feature map, so that the updated atomic features are more favorable for determining the interaction strength between it and other atoms, that is, more favorable for predicting the affinity of the protein and the ligand molecule. That is to say, in the prediction model, the process of determining the attention weight based on the obtained Q, K and V vectors and determining the new one-dimensional feature vector based on the determined attention weight can be performed multiple times to obtain better attention weights and affinity prediction results. Among them, in each feature vector update, after determining the attention vector for affinity prediction of the protein and the ligand molecule, the attention vector can be used to update the value vector V (for example, matrix multiplying the attention vector and the value vector V) to obtain an updated one-dimensional feature vector, and each element in the one-dimensional feature vector can still be one-to-one with each atom in the three-dimensional structure map. Therefore, after the above multiple updates, the attention weight and one-dimensional feature vector for the final affinity prediction and interaction map determination can be determined.

[0121] Optionally, for the final affinity prediction, the determined one-dimensional feature vector can be input into the task layer to output the affinity prediction result. For example, the task layer can linearly transform the one-dimensional feature vector based on the weight obtained by training (e.g., weighted summation) to obtain the affinity prediction result (e.g., one-dimensional affinity prediction value). For the final interaction graph determination, the attention weights output from the prediction model can be combined into different attention vectors based on the various components in the three-dimensional structure graph. For example, the attention weights corresponding to the atoms belonging to the protein molecule constitute the attention vector of the protein molecule, while the attention weights corresponding to the atoms belonging to the ligand molecule constitute the attention vector of the ligand molecule. Therefore, based on the attention vectors of the protein molecule and the ligand molecule, the interaction graph of the binding of the protein and the ligand molecule can be determined. For example, the interaction graph can be determined based on the product of the attention vectors of the protein molecule and the ligand molecule, wherein each element is the product of the corresponding element in the attention vector of the protein molecule and the corresponding element in the attention vector of the ligand molecule.

[0122] Therefore, by outputting the interaction map between the protein and the ligand molecule, it is possible to predict whether there is an interaction between the protein and the ligand molecule, check whether important atomic interactions are found, and explain the correctness of the affinity prediction results (for example, whether the prediction results are reasonable), so that the affinity prediction model is interpretable, and when the true interaction map between the protein and the ligand molecule can be obtained (for example, calculated by the PLIP tool), the correctness of the prediction results can be evaluated based on the error between the true interaction map and the obtained interaction map, and the above prediction model can be error corrected, as shown in reference. Figure 5 described.

[0123] According to yet another aspect of the present disclosure, a device for predicting affinity between a protein and a ligand molecule is provided. Figure 8 FIG. 2 is a schematic diagram showing a device 2000 for predicting affinity between a protein and a ligand molecule according to an embodiment of the present disclosure.

[0124] like Figure 8 As shown, the protein-ligand molecule affinity prediction device 2000 may include one or more processors 2010 and one or more memories 2020. The memory 2020 stores a computer-readable code, and when the computer-readable code is run by the one or more processors 2010, the protein-ligand molecule affinity prediction method described above may be executed.

[0125] The processor in the embodiments of the present disclosure may be an integrated circuit chip having signal processing capabilities. The processor may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present application may be implemented or executed. The general-purpose processor may be a microprocessor or the processor may be any conventional processor, etc., and may be an X86 architecture or an ARM architecture.

[0126] In general, various example embodiments of the present disclosure may be implemented in hardware or dedicated circuits, software, firmware, logic, or any combination thereof. Certain aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device. When various aspects of the embodiments of the present disclosure are illustrated or described as block diagrams, flow charts, or using some other graphical representation, it will be understood that the blocks, devices, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, dedicated circuits or logic, general purpose hardware or controllers or other computing devices, or some combination thereof as non-limiting examples.

[0127] For example, the method or device according to the embodiment of the present disclosure may also be implemented by Fig. 9 The architecture of the computing device 3000 shown in FIG. Fig. 9As shown, the computing device 3000 may include a bus 3010, one or more CPUs 3020, a read-only memory (ROM) 3030, a random access memory (RAM) 3040, a communication port 3050 connected to a network, an input / output component 3060, a hard disk 3070, etc. The storage device in the computing device 3000, such as the ROM 3030 or the hard disk 3070, may store various data or files used for processing and / or communication of the method for predicting affinity between a protein and a ligand molecule provided in the present disclosure, as well as program instructions executed by the CPU. The computing device 3000 may also include a user interface 3080. Of course, Figure 8 The architecture shown is only exemplary and can be omitted according to actual needs when implementing different devices. Fig. 9 One or more components of a computing device are shown.

[0128] According to yet another aspect of the present disclosure, a computer-readable storage medium is provided. Fig.10 A schematic diagram 4000 of a storage medium according to the present disclosure is shown.

[0129] like Fig.10 As shown, the computer storage medium 4020 stores computer readable instructions 4010. When the computer readable instructions 4010 are executed by the processor, the affinity prediction method of the protein and ligand molecule according to the embodiment of the present disclosure described with reference to the above figures can be executed. The computer readable storage medium in the embodiment of the present disclosure can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM) or a flash memory. Volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus random access memory (DR RAM). It should be noted that the memory of the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory. It should be noted that the memory of the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0130] The embodiments of the present disclosure also provide a computer program product or a computer program, which includes a computer instruction stored in a computer-readable storage medium. The processor of the computer device reads the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer device executes the method for predicting the affinity between the protein and the ligand molecule according to the embodiments of the present disclosure.

[0131] Embodiments of the present disclosure provide a method, apparatus, device and computer-readable storage medium for predicting affinity between a protein and a ligand molecule.

[0132] Compared with the existing protein and small molecule affinity prediction methods, the method provided by the embodiments of the present disclosure can achieve protein and small molecule affinity prediction more efficiently and accurately, while generating an interaction map that can reflect the atomic interaction between proteins and small molecules, making the prediction results of the method provided by the embodiments of the present disclosure explainable.

[0133] The method provided by the embodiment of the present disclosure is based on the node features, edge features and geometric features extracted from the three-dimensional structure diagram of the protein and ligand molecule binding, and obtains the affinity of the protein and the ligand molecule through a pre-trained prediction model, and obtains an interaction graph for indicating the interaction between the atoms of the protein and the ligand molecule. On the basis of improving the affinity prediction performance, it is possible to judge whether the predicted atomic interaction between the protein and the ligand molecule is correct, so that the prediction result is interpretable. Among them, the prediction model is obtained by error correction of affinity prediction and interaction graph prediction, so the method of the embodiment of the present disclosure can learn more accurate affinity prediction and interaction graph prediction on the basis of error correction.

[0134] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, a program segment, or a part of a code, and the module, program segment, or a part of the code contains at least one executable instruction for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0135] In general, various example embodiments of the present disclosure may be implemented in hardware or dedicated circuits, software, firmware, logic, or any combination thereof. Certain aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device. When various aspects of the embodiments of the present disclosure are illustrated or described as block diagrams, flow charts, or using some other graphical representation, it will be understood that the blocks, devices, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, dedicated circuits or logic, general purpose hardware or controllers or other computing devices, or some combination thereof as non-limiting examples.

[0136] The exemplary embodiments of the present disclosure described in detail above are merely illustrative and not restrictive. It should be understood by those skilled in the art that various modifications and combinations may be made to these embodiments or their features without departing from the principles and spirit of the present disclosure, and such modifications should fall within the scope of the present disclosure.

Claims

1. A method for predicting the affinity between a protein and a ligand molecule, comprising: Obtaining a three-dimensional structural diagram of the binding of the protein to the ligand molecule, wherein the three-dimensional structural diagram uses atoms of the protein and the ligand molecule as nodes; Determine node features of atoms of each of the protein and the ligand molecule from the three-dimensional structure graph, and determine edge features and geometric features of the three-dimensional structure graph based on each node in the three-dimensional structure graph; as well as Based on the node features, edge features and geometric features of the three-dimensional structure graph, the affinity of the protein to the ligand molecule is determined by a pre-trained prediction model, and an interaction map between the protein and the ligand molecule is obtained, wherein the interaction map is used to indicate the interaction between the atoms of the protein and the ligand molecule; Wherein, the prediction model adopts a self-attention mechanism, and the affinity between the protein and the ligand molecule is determined through a pre-trained prediction model based on the node features, edge features and geometric features of the three-dimensional structure graph, and the interaction graph between the protein and the ligand molecule is obtained, including: based on the node features, edge features and geometric features of the three-dimensional structure graph, the affinity between the protein and the ligand molecule and the attention vectors of the protein and the ligand molecule are determined through a self-attention mechanism, each element in the attention vector indicates the contribution of the corresponding atom to the affinity between the protein and the ligand molecule; and based on the attention vectors of the protein and the ligand molecule, the interaction graph between the protein and the ligand molecule is obtained, each element in the interaction graph indicates the possibility of interaction between the corresponding atomic pairs of the protein and the ligand molecule.

2. The method of claim 1, further comprising: Obtain samples of three-dimensional structure images of multiple proteins binding to ligand molecules; For each three-dimensional structure graph sample among the three-dimensional structure graph samples of the plurality of proteins and ligand molecules binding, determining the node features, edge features and geometric features of the three-dimensional structure graph sample; Obtaining a true affinity and a true interaction map corresponding to the three-dimensional structure map sample, wherein each element in the true interaction map indicates whether there is an interaction between the corresponding atomic pair of the protein in the three-dimensional structure map sample and the ligand molecule; Based on the node features, edge features and geometric features of the three-dimensional structure graph sample, determining the affinity and interaction graph corresponding to the three-dimensional structure graph sample through a prediction model, wherein the determined affinity and interaction graph contains parameters to be optimized of the prediction model; and By optimizing the affinity prediction error between the true affinity and the determined affinity corresponding to each three-dimensional structure graph sample of the plurality of protein-ligand molecule binding three-dimensional structure graph samples, and the interaction graph prediction error between the true interaction graph and the determined interaction graph, the parameters to be optimized of the prediction model are determined to obtain the pre-trained prediction model.

3. The method of claim 1, wherein: Based on the node features, edge features and geometric features of the three-dimensional structure graph, determining the attention vectors of the protein and the ligand molecule respectively through the self-attention mechanism includes: Concatenate the node features of the atoms of the protein and the ligand molecule into a one-dimensional feature vector; and Based on the self-attention mechanism, the one-dimensional feature vector is used as a query vector, a key vector and a value vector, and the attention weight of each node in the three-dimensional structure graph is determined in combination with the edge features and geometric features of the three-dimensional structure graph, wherein the attention weight indicates the contribution of the node to the affinity between the protein and the ligand molecule; The attention weights of the atoms of the protein and the ligand molecule respectively constitute their respective attention vectors.

4. The method of claim 3, wherein: Based on the attention vectors of the protein and the ligand molecule, respectively, obtaining an interaction graph between the protein and the ligand molecule comprises: For any atom pair of the protein and the ligand molecule, a corresponding element in the interaction graph is determined based on the product of corresponding attention weights in the respective attention vectors of the protein and the ligand molecule, and the corresponding element corresponds to the atom pair.

5. The method of claim 3, wherein: Based on the node features, edge features and geometric features of the three-dimensional structure graph, determining the affinity between the protein and the ligand molecule through a self-attention mechanism includes: Based on the node features, edge features and geometric features of the three-dimensional structure graph, the one-dimensional feature vector is updated multiple times, wherein in each update: Using the last updated one-dimensional feature vector as the query vector, key vector and value vector, and combining the edge features and geometric features of the three-dimensional structure graph, determining the attention weight of each node in the three-dimensional structure graph; Determining an updated one-dimensional feature vector based on the attention weight of each node in the three-dimensional structure graph and the value vector; and Based on the one-dimensional feature vectors that have been updated multiple times, the affinity between the protein and the ligand molecule is determined.

6. The method of claim 1, wherein: Determining the edge features and geometric features of the three-dimensional structure graph based on each node in the three-dimensional structure graph includes: According to each node in the three-dimensional structure graph, establishing a distance graph of the three-dimensional structure graph, wherein the distance graph indicates the distance between each pair of nodes in the three-dimensional structure graph; and Based on the three-dimensional structure graph and its distance graph, the edge features and geometric features of the three-dimensional structure graph are determined, wherein the edge features indicate the covalent bond features between corresponding node pairs in the three-dimensional structure graph, and the geometric features indicate the spatial relationship between the edges formed by each node pair in the three-dimensional structure graph.

7. The method of claim 1, wherein: The method further comprises: In the case where the binding of the protein to the ligand molecule requires the participation of a cofactor molecule, the acquired three-dimensional structure graph further includes the cofactor molecule, and the nodes of the three-dimensional structure graph further include atoms of the cofactor molecule; and Node features of the atoms of the cofactor molecule are determined from the three-dimensional structure graph, and edge features and geometric features of the three-dimensional structure graph are determined based on each node in the three-dimensional structure graph that includes the atoms of the cofactor molecule, so as to determine the affinity of the protein to the ligand molecule based on the node features, edge features and geometric features of the three-dimensional structure graph, and obtain an interaction graph between the protein and the ligand molecule.

8. A device for predicting the affinity between a protein and a ligand molecule, comprising: A data acquisition module is configured to acquire a three-dimensional structure diagram of the binding of the protein and the ligand molecule, wherein the three-dimensional structure diagram uses atoms of the protein and the ligand molecule as nodes; A feature extraction module is configured to determine node features of atoms of each of the protein and the ligand molecule from the three-dimensional structure graph, and determine edge features and geometric features of the three-dimensional structure graph based on each node in the three-dimensional structure graph; as well as A prediction module is configured to determine the affinity between the protein and the ligand molecule based on the node features, edge features and geometric features of the three-dimensional structure graph through a pre-trained prediction model, and obtain an interaction map between the protein and the ligand molecule, wherein the interaction map is used to indicate the interaction between the atoms of the protein and the ligand molecule; The prediction model adopts a self-attention mechanism, and the node features, edge features and geometric features of the three-dimensional structure graph are based on the pre-trained prediction model to determine the affinity between the protein and the ligand molecule, and obtain the interaction graph between the protein and the ligand molecule, including: Based on the node features, edge features and geometric features of the three-dimensional structure graph, determine the affinity between the protein and the ligand molecule, and the attention vectors of the protein and the ligand molecule respectively through a self-attention mechanism, wherein each element in the attention vector indicates the contribution of the corresponding atom to the affinity between the protein and the ligand molecule; and Based on the respective attention vectors of the protein and the ligand molecule, an interaction graph between the protein and the ligand molecule is obtained, wherein each element in the interaction graph indicates the possibility of interaction between the corresponding atomic pair of the protein and the ligand molecule.

9. The device of claim 8, wherein: Based on the node features, edge features and geometric features of the three-dimensional structure graph, determining the attention vectors of the protein and the ligand molecule respectively through the self-attention mechanism includes: Concatenate the node features of the atoms of the protein and the ligand molecule into a one-dimensional feature vector; and Based on the self-attention mechanism, the one-dimensional feature vector is used as a query vector, a key vector and a value vector, and the attention weight of each node in the three-dimensional structure graph is determined in combination with the edge features and geometric features of the three-dimensional structure graph, wherein the attention weight indicates the contribution of the node to the affinity between the protein and the ligand molecule; The attention weights of the atoms of the protein and the ligand molecule respectively constitute their respective attention vectors.

10. A device for predicting affinity between a protein and a ligand molecule, comprising: one or more processors; as well as One or more memories, wherein a computer executable program is stored, and when the computer executable program is executed by the processor, the method according to any one of claims 1 to 7 is performed.

11. A computer program product, stored on a computer-readable storage medium, and comprising computer instructions which, when executed by a processor, cause a computer device to perform the method of any one of claims 1 to 7.

12. A computer-readable storage medium having computer-executable instructions stored thereon, wherein the instructions are used to implement the method according to any one of claims 1 to 7 when executed by a processor.

Citation Information

Patent Citations

  • Method and apparatus for training predictive model for determining molecular binding force

    CN113241126A