A method for predicting the affinity of a protein ligand, related device, and equipment

By generating target topology maps and extracting features using attention networks, the problem of limited prediction accuracy of drug ligand-protein target affinity model in the prior art is solved, and more accurate affinity prediction is achieved.

CN115116538BActive Publication Date: 2025-06-24TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210360448.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-07
Publication Date
2025-06-24
Estimated Expiration
2042-04-07

AI Technical Summary

Technical Problem

The prediction accuracy of existing drug ligand-protein target affinity models is limited, and there is room for further optimization.

Method used

By obtaining the conformation of protein ligands, generating target topology maps, extracting atomic features in drug molecules and binding pockets, using internal attention networks and interactive attention networks to obtain molecular interaction characteristics, and then predicting affinity.

Benefits of technology

Improves the accuracy of affinity prediction of protein ligands and allows more fully to learn the characteristics within and between molecules.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115116538B_ABST
    Figure CN115116538B_ABST
Patent Text Reader

Abstract

The present application discloses a method for predicting the affinity of protein-ligand based on artificial intelligence technology. The present application includes: obtaining the conformation of the protein-ligand; generating a target topological graph according to the conformation of the protein-ligand; generating the first atomic feature of each drug atom in the drug molecule according to the target topological graph; generating the second atomic feature of each receptor atom in the binding pocket according to the target topological graph; obtaining the molecular interaction feature through the internal attention network included in the affinity prediction model based on the first atomic feature of each drug atom and the second atomic feature of each receptor atom; obtaining the affinity through the interaction attention network included in the affinity prediction model based on the molecular interaction feature. The present application also provides a device. When predicting the affinity of the protein-ligand, the present application can fully learn the internal features of the molecule and the features between molecules, which is beneficial to predicting a more accurate affinity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical fields of drugs and artificial intelligence, and in particular, to a method for predicting the affinity of a protein ligand, a related device, and a device. Background Art

[0002] Virtual screening (VS), also known as computer screening, is to simulate the interaction between a target protein and a candidate drug using molecular docking software on a computer before performing a bioactivity screening, and then calculate the affinity between the two through a drug ligand-protein target affinity model. In this way, the activity ranking and screening of the candidate drug molecule library are carried out, so as to speed up the drug development process and reduce the R & D cost.

[0003] With the booming development of machine learning (ML) and deep learning (DL), the drug ligand-protein target affinity model based on ML and DL can predict the affinity more quickly and accurately. Currently, in the Random Forests Score (RF Score) method based on ML, a random forest regression model can be constructed through the interaction features of proteins and drug ligands, and thus, the affinity of protein ligand binding can be predicted.

[0004] However, the inventors found that there are at least the following problems in the existing solutions. Although the existing drug ligand-protein target affinity model can learn without relying on experts for feature selection, the prediction accuracy of the affinity is limited and there is still room for further optimization. Summary of the Invention

[0005] The embodiments of the present application provide a method for predicting the affinity of a protein ligand, a related device, and a device. When predicting the affinity of a protein ligand, the present application can fully learn the internal features of molecules and the features between molecules, which is beneficial to predicting more accurate affinity.

[0006] In view of this, on the one hand, the present application provides a method for predicting the affinity of a protein ligand, including:

[0007] Obtain a protein ligand conformation, where the protein ligand conformation is constructed based on a protein and a drug molecule;

[0008] Generate a target topological graph according to the protein ligand conformation, where the target topological graph includes a node set and an edge set. The node set is used to represent the drug atoms in the drug molecule and the receptor atoms in the binding pocket, and the edge set is used to represent the chemical bonds connecting the atoms. The binding pocket is generated based on the protein;

[0009] Generate the first atomic feature of each drug atom in the drug molecule according to the target topological graph;

[0010] Generate the second atomic feature of each receptor atom in the binding pocket according to the target topological graph;

[0011] Based on the first atomic feature of each drug atom and the second atomic feature of each receptor atom, obtain the molecular interaction feature through the internal attention network included in the affinity prediction model;

[0012] Based on the molecular interaction feature, obtain the affinity for the protein-ligand conformation through the interaction attention network included in the affinity prediction model.

[0013] On the other hand, the present application provides an affinity prediction device, including:

[0014] An acquisition module, configured to acquire the protein-ligand conformation, where the protein-ligand conformation is constructed based on a protein and a drug molecule;

[0015] A generation module, configured to generate a target topological graph according to the protein-ligand conformation, where the target topological graph includes a node set and an edge set, the node set is used to represent drug atoms in the drug molecule and receptor atoms in the binding pocket, the edge set is used to represent chemical bonds connecting atoms, and the binding pocket is generated based on the protein;

[0016] The generation module is further configured to generate the first atomic feature of each drug atom in the drug molecule according to the target topological graph;

[0017] The generation module is further configured to generate the second atomic feature of each receptor atom in the binding pocket according to the target topological graph;

[0018] The acquisition module is further configured to obtain the molecular interaction feature through the internal attention network included in the affinity prediction model based on the first atomic feature of each drug atom and the second atomic feature of each receptor atom;

[0019] The acquisition module is further configured to obtain the affinity for the protein-ligand conformation through the interaction attention network included in the affinity prediction model based on the molecular interaction feature.

[0020] In a possible design, in another implementation manner of the other aspect of the embodiments of the present application,

[0021] The acquisition module is specifically configured to acquire the first format file corresponding to the protein;

[0022] Acquire the second format file corresponding to the drug molecule;

[0023] Based on the first format file and the second format file, X candidate protein-ligand conformations and the scores of each candidate protein-ligand conformation are generated through a molecular docking application, where X is an integer greater than or equal to 1;

[0024] Among the X candidate protein-ligand conformations, the candidate protein-ligand conformation with the highest score is selected as the protein-ligand conformation.

[0025] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,

[0026] The generation module is specifically configured to determine drug atoms belonging to a drug molecule according to the protein-ligand conformation;

[0027] Determine M amino acid molecules belonging to the protein according to the protein-ligand conformation, and determine the receptor atoms included in each amino acid molecule, where the M amino acid molecules belong to the protein and M is an integer greater than or equal to 1;

[0028] Taking each drug atom in the drug molecule as the center, determine at least one receptor atom whose atomic distance is less than or equal to the distance threshold;

[0029] Obtain the distance between each receptor atom in the at least one receptor atom and the drug molecule;

[0030] Sort the at least one receptor atom in ascending order of distance;

[0031] For the at least one sorted receptor atom, add the amino acid molecule to which the receptor atom belongs in sequence until the graph construction stop condition is met, and obtain the target topological graph.

[0032] In a possible design, in another implementation manner of another aspect of the embodiments of the present application, the first atomic feature includes a first node feature and a first distance feature;

[0033] The generation module is specifically configured to obtain the atomic association data corresponding to each drug atom in the drug molecule, where the atomic association data includes at least one of atomic number, number of neighbors, formal charge, radical electrons, hybrid orbitals, aromatic ring correlation, number of attached hydrogen atoms, chiral center, chiral type, and molecular type;

[0034] For each drug atom in the drug molecule, perform feature extraction on the atomic association data corresponding to the drug atom to obtain the first node feature;

[0035] According to the target topological graph, obtain the distance data between every two drug atoms in the drug molecule;

[0036] Characterize the distance data between every two drug atoms in the drug molecule to obtain the first distance feature.

[0037] In a possible design, in another implementation of another aspect of the embodiments of the present application, the first atomic feature further includes a first edge feature;

[0038] The acquisition module is further configured to obtain the edge connection data between every two drug atoms in the drug molecule according to the target topological graph, where the edge connection data includes at least one of the covalent bond type and the covalent bond position relationship;

[0039] The generation module is further configured to characterize the edge connection data between every two drug atoms in the drug molecule to obtain the first edge feature.

[0040] In a possible design, in another implementation of another aspect of the embodiments of the present application, the first atomic feature further includes a first quantization feature;

[0041] The acquisition module is further configured to, for each drug atom in the drug molecule, obtain the quantization data corresponding to the drug atom according to the target topological graph, where the quantization data includes at least one of the covalent bond angle, the interaction angle, and the local charge. The covalent bond angle is the angle formed by the drug atom as the vertex and the two closest atoms. The interaction angle is the angle formed by the drug atom as the vertex, the first atom, and the second atom. The first atom is the atom with the closest edge connection relationship and the shortest distance to the drug atom, and the second atom is the receptor atom with the shortest distance to the drug atom;

[0042] The generation module is further configured to, for each drug atom in the drug molecule, characterize the quantization data corresponding to the drug atom to obtain the first quantization feature.

[0043] In a possible design, in another implementation of another aspect of the embodiments of the present application, the second atomic feature includes a second node feature and a second distance feature;

[0044] The generation module is specifically configured to obtain the atomic association data corresponding to each receptor atom in the binding pocket, where the atomic association data includes at least one of the element serial number, the number of neighbors, the formal charge, the radical electrons, the hybrid orbitals, the aromatic ring correlation, the number of connected hydrogen atoms, the chiral center, the chiral type, and the molecular type;

[0045] For each receptor atom in the binding pocket, characterize the atomic association data corresponding to the receptor atom to obtain the second node feature;

[0046] Obtain the distance data between every two receptor atoms in the binding pocket according to the target topological graph;

[0047] Characterize the distance data between pairwise receptor atoms in the binding pocket to obtain the second distance feature.

[0048] In a possible design, in another implementation of another aspect of the embodiments of the present application, the second atomic feature further includes a second edge feature;

[0049] The acquisition module is further configured to obtain the edge connection data between pairwise receptor atoms in the binding pocket according to the target topological graph, where the edge connection data includes at least one of the covalent bond type and the covalent bond position relationship;

[0050] The generation module is further configured to characterize the edge connection data between pairwise receptor atoms in the binding pocket to obtain the second edge feature.

[0051] In a possible design, in another implementation of another aspect of the embodiments of the present application, the second atomic feature further includes a second quantization feature;

[0052] The acquisition module is further configured to, for each receptor atom in the binding pocket, obtain the quantization data corresponding to the receptor atom according to the target topological graph, where the quantization data includes at least one of the covalent bond angle, the interaction angle, and the local charge. The covalent bond angle is the angle formed by the receptor atom as the vertex and the two closest atoms. The interaction angle is the angle formed by the receptor atom as the vertex, the third atom, and the fourth atom. The third atom is the atom with the closest edge connection relationship to the receptor atom, and the fourth atom is the closest drug atom to the receptor atom;

[0053] The generation module is further configured to, for each receptor atom in the binding pocket, characterize the quantization data corresponding to the receptor atom to obtain the second quantization feature.

[0054] In a possible design, in another implementation of another aspect of the embodiments of the present application, the drug molecule includes L drug atoms, and the binding pocket includes P receptor atoms, where both L and P are integers greater than or equal to 1;

[0055] The acquisition module is specifically configured to perform feature embedding processing on the first atomic features of each drug atom to obtain L first input embedding features and H first bias embedding features, where H represents the number of attention heads, and H is an integer greater than or equal to 1;

[0056] Perform feature embedding processing on the second atomic features of each receptor atom to obtain P second input embedding features and H second bias embedding features;

[0057] Based on L first input embedding features and H first bias embedding features, obtain drug molecule interaction features through the first attention network included in the internal attention network, where the internal attention network belongs to the affinity prediction model;

[0058] Based on P second input embedding features and H second bias embedding features, obtain binding pocket interaction features through the second attention network included in the internal attention network;

[0059] Concatenate the drug molecule interaction features and the binding pocket interaction features to obtain molecular interaction features.

[0060] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,

[0061] The acquisition module is specifically configured to obtain L first query vectors, L first key vectors, and L first value vectors through the linear network layer included in the first attention network based on L first input embedding features, where the first attention network belongs to the internal attention network;

[0062] Perform matrix multiplication on the L first query vectors and the L first key vectors to obtain H first intermediate matrices;

[0063] Based on the H first intermediate matrices and the H first bias embedding features, obtain H first attention matrices through the normalization exponential function layer included in the first attention network;

[0064] Perform matrix multiplication on the H first attention matrices and the L first value vectors to obtain H first target matrices;

[0065] Based on the H first target matrices, obtain drug molecule interaction features through the target neural network included in the first attention network;

[0066] The acquisition module is specifically configured to obtain P second query vectors, P second key vectors, and P second value vectors through the linear network layer included in the second attention network based on P second input embedding features, where the second attention network belongs to the internal attention network;

[0067] Perform matrix multiplication on the P second query vectors and the P second key vectors to obtain H second intermediate matrices;

[0068] Based on the H second intermediate matrices and the H second bias embedding features, obtain H second attention matrices through the normalization exponential function layer included in the second attention network;

[0069] Perform matrix multiplication on the H second attention matrices and the P second value vectors to obtain H second target matrices;

[0070] Based on H second target matrices, the binding pocket interaction features are obtained through the target neural network included in the second attention network.

[0071] In a possible design, in another implementation manner of another aspect of the embodiments of the present application, the drug molecule includes L drug atoms, and the binding pocket includes P receptor atoms, where both L and P are integers greater than or equal to 1;

[0072] The obtaining module is specifically configured to splice the molecular interaction features and the target embedding features of the virtual nodes to obtain N input embedding features, where the N input embedding features include drug molecule interaction features, binding pocket interaction features, and target embedding features, and N is equal to (L + P + 1);

[0073] Based on the N input embedding features, N target query vectors, N target key vectors, and N target value vectors are obtained through the linear network layer included in the interaction attention network, where the interaction attention network belongs to the affinity prediction model;

[0074] The matrix multiplication of the N target query vectors and the N target key vectors is performed to obtain H target intermediate matrices, where H represents the number of attention heads, and H is an integer greater than or equal to 1;

[0075] Based on the H target intermediate matrices and the H target bias embedding features, H target attention matrices are obtained through the normalization exponential function layer included in the interaction attention network, where the H target bias embedding features are generated according to the first atomic features of each drug atom, the second atomic features of each receptor atom, and the atomic features of the virtual node;

[0076] The matrix multiplication of the H target attention matrices and the N target value vectors is performed to obtain H target matrices;

[0077] Based on the H target matrices, the to-be-processed features are obtained through the target neural network included in the interaction attention network;

[0078] The target to-be-processed features are obtained according to the to-be-processed features;

[0079] Based on the target to-be-processed features, the affinity for the protein-ligand conformation is obtained through the output layer.

[0080] In a possible design, in another implementation manner of another aspect of the embodiments of the present application, the affinity prediction device further includes a processing module;

[0081] The obtaining module is further configured to obtain a first weight parameter and a second weight parameter from the gamma matrix, where the gamma matrix is used to describe the feature relationship between the drug molecule and the binding pocket;

[0082] The generation module is further configured to generate H to-be-processed bias embedding features according to the first atomic feature of each drug atom, the second atomic feature of each receptor atom, and the atomic feature of the virtual node.

[0083] The processing module is configured to adjust the first numerical set in each to-be-processed bias embedding feature by using the first weight parameter, and adjust the second numerical set in each to-be-processed bias embedding feature by using the second weight parameter, so as to obtain H target bias embedding features.

[0084] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,

[0085] The processing module is further configured to, based on the H target intermediate matrices and the H target bias embedding features, after obtaining the H target attention matrices through the normalization exponential function layer included in the interactive attention network, perform an averaging process on the values at the same positions in the H target attention matrices to obtain an average attention matrix.

[0086] The acquisition module is further configured to acquire an association distribution matrix from the average attention matrix, where the association distribution matrix represents the association between L drug atoms and P receptor atoms.

[0087] In a possible design, in another implementation manner of another aspect of the embodiments of the present application, the affinity prediction device further includes a training module;

[0088] The acquisition module is further configured to acquire an original conformation sample, where the original conformation sample is constructed based on an original protein sample and an original drug molecule sample.

[0089] The generation module is further configured to generate T training topology graph samples according to the original conformation sample, where each training topology graph sample includes a training drug molecule sample, and each training topology graph sample has a corresponding true affinity value, and T is an integer greater than or equal to 1.

[0090] The acquisition module is further configured to, based on the T training topology graph samples, obtain the affinity prediction value corresponding to each training topology graph sample through the affinity prediction model.

[0091] The training module is configured to update the model parameters in the affinity prediction model according to the affinity prediction value and the true affinity value corresponding to each training topology graph sample.

[0092] In a possible design, in another implementation manner of another aspect of the embodiments of the present application, the affinity prediction device further includes a determination module;

[0093] A determination module, configured to determine, for each of the T training topological graph samples, the molecular structure error between the training drug molecule sample included in the training topological graph sample and the original drug molecule sample;

[0094] The determination module is further configured to determine the label category corresponding to each training topological graph sample according to the molecular structure error corresponding to each training topological graph sample, wherein the label category corresponding to the training topological graph sample with a molecular structure error less than the error threshold is a positive sample label, and the label category corresponding to the training topological graph sample with a molecular structure error greater than or equal to the error threshold is a negative sample label;

[0095] An acquisition module is further configured to obtain the predicted category probability corresponding to each training topological graph sample through an affinity prediction model based on the T training topological graph samples;

[0096] A training module is specifically configured to determine a first loss value by using a first loss function according to the affinity predicted value and the true affinity value corresponding to each training topological graph sample;

[0097] Determine a second loss value by using a second loss function according to the predicted category probability and the label category corresponding to each training topological graph sample;

[0098] Update the model parameters in the affinity prediction model according to the first loss value and the second loss value.

[0099] On the other hand, this application provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the methods in the above aspects are implemented.

[0100] On the other hand, this application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the methods in the above aspects are implemented.

[0101] In another aspect of this application, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the methods in the above aspects are implemented.

[0102] As can be seen from the above technical solutions, the embodiments of this application have the following advantages:

[0103] In the embodiments of the present application, a method for predicting the affinity of a protein ligand is provided. First, the conformation of the protein ligand is obtained, and a target topological graph is generated using the conformation of the protein ligand. The target topological graph includes a node set and an edge set. The node set is used to represent the drug atoms in the drug molecule and the receptor atoms in the binding pocket, and the edge set is used to represent the chemical bonds connecting the atoms. Based on this, on the one hand, the first atomic feature of each drug atom in the drug molecule can be generated according to the target topological graph. On the other hand, the second atomic feature of each receptor atom in the binding pocket can be generated according to the target topological graph. Thus, based on the first atomic feature of each drug atom and the second atomic feature of each receptor atom, the molecular interaction feature is obtained through the internal attention network included in the affinity prediction model. Furthermore, based on the molecular interaction feature, the affinity for the protein ligand conformation is obtained through the interaction attention network included in the affinity prediction model. In the above manner, the target topological graph constructed using the rich structural information of the protein ligand conformation is combined, and the atomic features of each drug atom in the drug molecule and the atomic features of each receptor atom in the amino acid molecule can be extracted respectively. Based on this, first, based on the atomic features of each drug atom and each receptor atom, the internal features of the drug molecule and the amino acid molecule are learned respectively, thereby obtaining the molecular interaction feature. Then, the interaction attention network is used to learn the molecular interaction feature, thereby obtaining the affinity. It can be seen that when predicting the affinity of a protein ligand, the internal features of the molecule and the features between molecules can be fully learned, which is beneficial to predicting a more accurate affinity. Description of the Drawings

[0104] Figure 1 It is a schematic architecture diagram of a protein ligand affinity prediction system in the embodiments of the present application;

[0105] Figure 2 It is a schematic working flow diagram of predicting the affinity of a protein ligand in the embodiments of the present application;

[0106] Figure 3 It is a schematic flow diagram of a method for predicting the affinity of a protein ligand in the embodiments of the present application;

[0107] Figure 4 It is a schematic diagram of generating the conformation of a protein ligand based on molecular docking application in the embodiments of the present application;

[0108] Figure 5 It is a schematic diagram of a target topological graph in the embodiments of the present application;

[0109] Figure 6 It is a schematic diagram of the distance data between two atoms in the embodiments of the present application;

[0110] Figure 7 It is a schematic diagram of the covalent bond angle in the embodiments of the present application;

[0111] Figure 8 It is a schematic diagram of the interaction angle in the embodiment of the present application;

[0112] Figure 9 It is a schematic diagram of the network structure of the affinity prediction model in the embodiment of the present application;

[0113] Figure 10 It is a schematic diagram of the network structure of the internal attention network in the embodiment of the present application;

[0114] Figure 11 It is a schematic diagram of the network structure of the attention bias module in the embodiment of the present application;

[0115] Figure 12 It is a schematic diagram of the network structure of the interaction attention network in the embodiment of the present application;

[0116] Figure 13 It is a schematic diagram of the network structure of the output layer in the embodiment of the present application;

[0117] Figure 14 It is a schematic diagram of the to-be-processed bias embedding feature in the embodiment of the present application;

[0118] Figure 15 It is a schematic diagram of the average attention matrix in the embodiment of the present application;

[0119] Figure 16 It is a schematic diagram of realizing data enhancement based on the original conformation sample in the embodiment of the present application;

[0120] Figure 17 It is a schematic diagram of the affinity prediction device in the embodiment of the present application;

[0121] Figure 18 It is a schematic diagram of the structure of the computer device in the embodiment of the present application. Detailed implementation manners

[0122] The embodiment of the present application provides a method for predicting the affinity of a protein ligand, a related device, and a device. When predicting the affinity of a protein ligand, the present application can fully learn the internal features of the molecule and the features between molecules, which is beneficial to predicting a more accurate affinity.

[0123] In the description, claims and the above-mentioned drawings of this application, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented, for example, in an order other than those illustrated or described herein. In addition, the terms "comprising" and "corresponding to" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units need not be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0124] Drug discovery is the process of identifying drug compounds with potential therapeutic effects. In this process, predicting the interaction between drug molecules and target proteins is an essential step. Proteins are important drug targets, and drugs play important roles in the human body by interacting with various targets, which can enhance or inhibit their functions and play a regulatory role to achieve the purpose of treating a certain disease. Applying artificial intelligence (AI) technology to drug research and development can obtain some candidate drug molecules through the reasoning and decision-making functions of machines, thereby reducing the time cost of manual drug research and development.

[0125] Among them, AI is to use digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results in theory, methods, technologies and application systems. In other words, AI is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. AI is also to study the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning and decision-making. AI technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. AI basic technologies generally include technologies such as sensors, dedicated AI chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. AI software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning (ML) / deep learning.

[0126] ML is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills, and reorganize the existing knowledge structure to continuously improve their own performance. ML is the core of AI and the fundamental way to make computers intelligent, and its applications cover all fields of AI. ML and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.

[0127] With the rapid development of disciplines such as molecular biology and structural biology, the three-dimensional structures and functions of a large number of biological macromolecules have been determined, and computer-aided drug design has made breakthrough progress. By means of molecular docking and molecular dynamics research, the binding conformations of drug molecules and target proteins are determined, and then the binding activities of drug molecules and receptor proteins are evaluated, so as to effectively discover new chemical entities, greatly improving the efficiency of drug development and reducing the R & D costs.

[0128] Based on this, the present application provides a method for predicting the binding affinity between a drug molecule and a receptor protein, and this method is applied to Figure 1 the protein-ligand affinity prediction system shown as follows. As shown in the figure, the protein-ligand affinity prediction system includes at least one of a server and a terminal, and the client is deployed on the terminal. Among them, the client can run on the terminal in the form of a browser, or can run on the terminal in the form of an independent application (APP), etc. For the specific display form of the client, no limitation is made here. The server involved in the present application can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a notebook computer, a handheld computer, a personal computer, a smart TV, a smart watch, a vehicle-mounted device, a wearable device, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and no limitation is made here in the present application. The number of the server and the terminal is also not limited. The solution provided by the present application can be completed independently by the terminal, or can be completed independently by the server, or can also be completed by the cooperation of the terminal and the server. In this regard, no specific limitation is made in the present application.

[0129] The binding affinity between a drug molecule and a receptor protein can be applied to different scenarios, which will be introduced separately below.

[0130] I. Structured Based Virtual Screening;

[0131] Applied in the drug virtual screening process. That is, through model prediction and ranking from a vast library of small molecule candidate drugs, small molecule compounds that are most likely to become hit compounds are screened out. The screening of hit compounds is a very important step in the pharmaceutical industry. Experimental screening of thousands of random candidate molecules is extremely resource-consuming. Therefore, by predicting the binding affinity between drug molecules and receptor proteins through models, the proportion of active molecules screened out can be effectively increased, thus saving R & D costs.

[0132] II. Structured Based Lead Optimization;

[0133] Applied in the lead compound optimization process. That is, it can help drug developers refer to molecules with higher predicted binding affinity when designing small molecule compounds, and design and optimize novel small molecules in a more efficient way.

[0134] In view of the fact that this application involves some terms related to the professional field, for the sake of easy understanding, explanations will be given below.

[0135] (1) Computer Aided Drug Design (CADD): A method based on computer chemistry to design and optimize lead compounds by simulating, calculating, and predicting the relationship between drugs and receptor biomacromolecules through computers.

[0136] (2) Ligand: Usually refers to candidate drug molecules, including small molecules and biological agents. They can interact with proteins in biological processes and act as agonists or inhibitors to treat diseases.

[0137] (3) Binding affinity: Defined as the strength of the binding interaction between a protein and a ligand (e.g., a drug molecule), which can be measured by experimental methods.

[0138] (4) Machine learning (ML): A class of algorithms that automatically analyze and obtain rules from data and use the rules to predict unknown data.

[0139] (5) Deep learning (DL): A branch of machine learning, an algorithm that attempts to use multiple processing layers with complex structures or composed of multiple non-linear transformations to perform high-level abstraction on data.

[0140] (6) Neural Network (NN): A deep learning model that mimics the structure and function of biological neural networks in the fields of machine learning and cognitive science.

[0141] (7) Attention Mechanism: A weight mechanism that models the relative importance of queries and keys in a neural network.

[0142] (8) Graph Neural Network (GNN): A neural network model structure based on topological graph data as input.

[0143] (9) Docking: A theoretical simulation method that predicts the binding mode and affinity by the characteristics of the receptor and the interaction mode between the receptor and the drug molecule.

[0144] (10) Complex: The co-crystal structure of the binding of a drug molecule to a protein, or the binding model generated by molecular docking software.

[0145] (11) Pocket: A cavity structure formed on the surface or inside of a protein structure that is suitable for binding drug molecules.

[0146] (12) GammaMatrix: A weight matrix in the model that can control the characteristics of different types of edges.

[0147] (13) Pose-Sensitive: The sensitivity of the model prediction to the quality of different binding poses of the same drug molecule.

[0148] For ease of explanation, the workflow for predicting protein-ligand affinity will be described below in conjunction with Figure 2 Please refer to Figure 2 , Figure 2 FIG. is a schematic diagram of a workflow for predicting protein-ligand affinity in an embodiment of the present application. As shown in the figure, a protein and a chemical molecule are obtained. Among them, the protein can form an initial pocket, and the chemical molecule is represented as a drug molecule. Based on this, the Docking application is used to combine the initial pocket and the drug molecule to obtain several complex poses. Then, one of the complex poses is selected as the protein-ligand conformation. Based on the protein-ligand conformation, the binding affinity of the protein-ligand conformation is predicted by an affinity prediction model (Interformer).

[0149] Combined with the above introduction, the method for predicting the affinity of protein ligands in the present application will be introduced below. Please refer toFigure 3 , in the embodiments of the present application, the method for predicting the affinity of a protein ligand can be executed by a computer device, which can be a terminal or a server, and includes:

[0150] 110. Obtain a protein ligand conformation, where the protein ligand conformation is constructed based on a protein and a drug molecule;

[0151] In one or more embodiments, first determine the protein and drug molecule to be synthesized, and then generate a protein ligand conformation based on the protein and drug scores. Among them, the protein ligand conformation is a complex, and the protein ligand conformation includes the initial pocket corresponding to the protein and the drug molecule.

[0152] Specifically, there are multiple ways to obtain the protein ligand conformation. Exemplarily, one way is to directly obtain the co-crystal conformation of the protein ligand (i.e., the protein ligand conformation) from the Protein Data Bank (PDB). Exemplarily, another way is to synthesize the protein ligand conformation through a Docking application.

[0153] 120. Generate a target topology graph according to the protein ligand conformation, where the target topology graph includes a node set and an edge set. The node set is used to represent the drug atoms in the drug molecule and the receptor atoms in the binding pocket, and the edge set is used to represent the chemical bonds connecting the atoms. The binding pocket is generated based on the protein;

[0154] In one or more embodiments, a corresponding target topology graph can be generated according to the protein ligand conformation. The target topology graph includes a node set and an edge set. Among them, the node set includes multiple atoms, and each atom corresponds to a node. The edge set includes at least one connecting edge, and each connecting edge represents a chemical bond and is used to connect two different nodes.

[0155] Specifically, the protein ligand conformation includes an initial pocket and a drug molecule. The initial pocket includes multiple receptor atoms, and the drug molecule includes multiple drug atoms. The protein includes several amino acid molecules. Based on this, in one case, all the multiple receptor atoms in the initial pocket can be amino acid atoms. In another case, the multiple receptor atoms in the initial pocket can include amino acid atoms and cofactors. Therefore, the present application does not limit the atomic type constituting the initial pocket.

[0156] It can be understood that the number of receptor atoms included in the pocket in the target topology graph is less than or equal to the number of receptor atoms included in the initial pocket in the protein ligand conformation.

[0157] It should be noted that the drug atoms in this application can be understood as ligand atoms, and the amino acid atoms in this application can be understood as the closest amino atoms.

[0158] 130. Generate the first atomic feature of each drug atom in the drug molecule according to the target topology map;

[0159] In one or more embodiments, the first atomic feature of each drug atom in the drug molecule is generated according to the target topology map. Among them, the first atomic feature includes the first node feature and the first distance feature, or the first atomic feature includes the first node feature, the first distance feature and the first edge feature, or the first atomic feature includes the first node feature, the first distance feature and the first quantization feature, or the first atomic feature includes the first node feature, the first distance feature, the first edge feature and the first quantization feature.

[0160] 140. Generate the second atomic feature of each receptor atom in the binding pocket according to the target topology map;

[0161] In one or more embodiments, the second atomic feature of each receptor atom in the pocket is generated according to the target topology map. Among them, the second atomic feature includes the second node feature and the second distance feature, or the second atomic feature includes the second node feature, the second distance feature and the second edge feature, or the second atomic feature includes the second node feature, the second distance feature and the second quantization feature, or the second atomic feature includes the second node feature, the second distance feature, the second edge feature and the second quantization feature.

[0162] 150. Based on the first atomic feature of each drug atom and the second atomic feature of each receptor atom, obtain the molecular interaction feature through the internal attention network included in the affinity prediction model;

[0163] In one or more embodiments, the affinity prediction model is called. Among them, the affinity prediction model includes an internal attention network and an interaction attention network. Based on this, the first atomic feature of each drug atom is used as the input of the internal attention network, and the second atomic feature of each receptor atom is used as the input of the internal attention network, thereby obtaining the molecular interaction feature.

[0164] 160. Based on the molecular interaction feature, obtain the affinity for the protein-ligand conformation through the interaction attention network included in the affinity prediction model.

[0165] In one or more embodiments, molecular interaction features are used as the input of the interaction attention network in the affinity prediction model, and the affinity for the protein-ligand conformation is obtained through the interaction attention network. The affinity can be expressed as molecular activity data (pIC50), where pIC50 = -log(IC50). Here, IC50 is the half-inhibitory concentration or half-inhibitory rate, which represents the concentration of the drug or inhibitor required to inhibit a specified biological process by half.

[0166] In the embodiments of the present application, a method for predicting the affinity of a protein-ligand is provided. In the above manner, by using the target topological graph constructed with the rich structural information of the protein-ligand conformation, the atomic features of each drug atom in the drug molecule and the atomic features of each receptor atom in the amino acid molecule can be respectively extracted in combination with the target topological graph. Based on this, first, based on the atomic features of each drug atom and each receptor atom, the internal features of the drug molecule and the amino acid molecule are respectively learned, thereby obtaining molecular interaction features. Then, the interaction attention network is used to learn the molecular interaction features to obtain the affinity. It can be seen that when predicting the affinity of a protein-ligand, the internal features of the molecule and the features between molecules can be fully learned, which is beneficial to predicting a more accurate affinity.

[0167] Optionally, on the basis of the corresponding respective embodiments above, in another optional embodiment provided by the embodiments of the present application, obtaining the protein-ligand conformation may specifically include: Figure 3 Obtaining the first format file corresponding to the protein;

[0168] Obtaining the second format file corresponding to the drug molecule;

[0169] Based on the first format file and the second format file, X candidate protein-ligand conformations and the score of each candidate protein-ligand conformation are generated through a molecular docking application, where X is an integer greater than or equal to 1;

[0170] From the X candidate protein-ligand conformations, the candidate protein-ligand conformation with the highest score is selected as the protein-ligand conformation.

[0171] In one or more embodiments, a method for generating a protein-ligand conformation based on a Docking application is combined. As can be seen from the foregoing embodiments, before synthesizing the protein-ligand conformation, it is necessary to select and prepare the protein and the ligand (i.e., the drug molecule). The selection of the protein plays an important role in the docking result. Due to the improvement of structure determination technology, the three-dimensional structure of a specific receptor can be obtained in the PDB. Thus, the first format file corresponding to the protein is obtained. The ligand can be drawn in software or downloaded from a chemical database. Thus, the second format file corresponding to the drug molecule is obtained.

[0172] ​

[0173] It is understandable that the first format file can be a PDB format file. The second format file can be in the format of Simplified Molecular Input Line Entry System (Smiles), or in the format of Structural Data File (SDF).

[0174] Specifically, for the sake of easy understanding, please refer to Figure 4 , Figure 4 FIG. is a schematic diagram of generating a protein-ligand conformation based on molecular docking application in the embodiments of the present application. As shown in the figure, first, obtain the first format file corresponding to the protein and the second format file corresponding to the drug molecule. Use the docking algorithm in the Docking application to process the first format file and the second format file, and thus X candidate protein-ligand conformations can be obtained. Based on this, use the molecular docking scoring function to evaluate the binding activity of each candidate protein-ligand conformation, that is, obtain the corresponding score. Assuming X is 5, a schematic mapping relationship between the candidate protein-ligand conformations and the scores as shown in Table 1 can be obtained.

[0175] Table 1

[0176] Candidate protein-ligand conformation Score Candidate protein-ligand conformation 1 85 Candidate protein-ligand conformation 2 14 Candidate protein-ligand conformation 3 66 Candidate protein-ligand conformation 4 98 Candidate protein-ligand conformation 5 50

[0177] Among them, the molecular docking scoring function is an approximate method for estimating the binding affinity. Generally, the higher the score, the better the binding affinity of the candidate protein-ligand conformation. Therefore, the "candidate protein-ligand conformation 3" with the highest score can be selected as the protein-ligand conformation.

[0178] Secondly, in the embodiments of the present application, a method for generating a protein-ligand conformation based on the Docking application is provided. Through the above method, since the Docking application predicts the binding conformation of the small molecule ligand to the target binding site, therefore, this Docking application can be used to screen bioactive compounds in the early stage of drug development, thereby increasing the flexibility and diversity of the synthesized protein-ligand conformations.

[0179] Optionally, on the basis of the above Figure 3 corresponding embodiments, in another optional embodiment provided by the embodiments of the present application, generating a target topology map according to the protein-ligand conformation may specifically include:

[0180] Determine the drug atoms belonging to the drug molecule according to the protein-ligand conformation;

[0181] Determine M amino acid molecules according to the protein-ligand conformation, and determine the receptor atoms included in each amino acid molecule, where the M amino acid molecules belong to the protein and M is an integer greater than or equal to 1;

[0182] Taking each drug atom in the drug molecule as the center, determine at least one receptor atom whose atomic distance is less than or equal to the distance threshold;

[0183] Obtain the distance between each receptor atom in at least one receptor atom and the drug molecule;

[0184] Sort at least one receptor atom in ascending order of distance;

[0185] For at least one receptor atom after sorting, sequentially add the amino acid molecule to which the receptor atom belongs until the graph construction stop condition is met, and obtain the target topological graph.

[0186] In one or more embodiments, a method for constructing a target topological graph is introduced. As can be seen from the foregoing embodiments, the protein-ligand conformation includes an initial pocket and a drug molecule. Among them, the drug molecule includes L drug atoms. The initial pocket includes M amino acid molecules derived from the protein, or the initial pocket includes M amino acid molecules derived from the protein and several cofactors. Each amino acid molecule includes several amino acid atoms. Therefore, the amino acid atoms and cofactors can be collectively referred to as "receptor atoms".

[0187] Specifically, taking each drug atom in the drug molecule as the center, respectively retain the receptor atoms that are less than or equal to the distance threshold (for example, ) from the drug atom. And take the distance between the receptor atom and the connected drug atom as the distance between the receptor atom and the drug molecule. For the convenience of introduction, please refer to Figure 2 , Figure 2 as a schematic of the distance relationship between the receptor atom and the drug atom.

[0188] Table 2

[0189]

[0190]

[0191] It can be seen that the 1st drug atom is connected to 4 receptor atoms, the 2nd drug atom is connected to 3 receptor atoms, the 3rd drug atom is connected to 2 receptor atoms, and the 4th drug atom is connected to 1 receptor atom. Based on this, sort these receptor atoms in ascending order of distance. For the convenience of understanding, please refer to Figure 3 , Figure 3It is a schematic of the sorted distance relationship between the receptor atoms and the drug atoms.

[0192] Table 3

[0193]

[0194] Based on this, the entire amino acid residue is completed starting from the receptor atom with the shortest distance. That is, first, the entire amino acid residue is completed starting from the 15th receptor atom. Among them, the 15th receptor atom belongs to amino acid molecule C. Therefore, amino acid molecule C is completed based on the 15th receptor atom. Next, the entire amino acid residue is completed starting from the 6th receptor atom. Among them, the 6th receptor atom belongs to amino acid molecule A. Therefore, amino acid molecule A is completed based on the 6th receptor atom. And so on, until the graph construction stop condition is met, and the target topological graph is obtained. The number of amino acid molecules included in the target topological graph is less than or equal to the number of amino acid molecules included in the protein, and the target topological graph includes a pocket and a drug molecule.

[0195] It should be noted that, in one case, when the number of atoms exceeds the number threshold (for example, 400), it means that the graph construction stop condition is met. In another case, when the number of atoms reaches the number threshold (for example, 400), and the distance between the atoms is less than the preset distance threshold (for example, )), it means that the graph construction stop condition is met.

[0196] For ease of understanding, please refer to Figure 5 , Figure 5 which is a schematic diagram of the target topological graph in the embodiment of the present application. As shown in the figure, the target topological graph includes 4 amino acid molecules and a drug molecule. Among them, the 4 amino acid molecules are aspartic acid (ASP), methionine (MET), glycine (GLY), and valine (VAL) respectively. The drug molecule includes hydrogenium (H) atoms, nitrogen (N) atoms, carbon (C) atoms, and oxygen (O) atoms. Among them, two connecting edges represent double bond covalent bonds, a single thin connecting edge represents a covalent bond, and a single thick connecting edge represents a non-covalent bond.

[0197] It can be understood that Figure 5 the shown target topological graph is only a schematic and should not be construed as a limitation of the present application.

[0198] Secondly, in the embodiments of the present application, a method for constructing a target topological graph is provided. Through the above method, considering the problem of over-smoothing in GNN, it can only consist of one or two layers, and more layers cannot bring better performance. The affinity prediction model takes the topological graph as the input, can have up to 30 layers, and has better and deeper expression ability.

[0199] Optionally, based on the above Figure 3 In another optional embodiment provided by the embodiments of the present application on the basis of the corresponding respective embodiments, the first atomic feature includes a first node feature and a first distance feature;

[0200] According to the target topological graph, generating the first atomic feature of each drug atom in the drug molecule may specifically include:

[0201] Obtaining the atomic association data corresponding to each drug atom in the drug molecule, where the atomic association data includes at least one of atomic number, number of neighbors, formal charge, radical electrons, hybrid orbitals, aromatic ring correlation, number of connected hydrogen atoms, chiral center, chiral type, and molecular type;

[0202] For each drug atom in the drug molecule, performing feature processing on the atomic association data corresponding to the drug atom to obtain the first node feature;

[0203] According to the target topological graph, obtaining the distance data between every two drug atoms in the drug molecule;

[0204] Performing feature processing on the distance data between every two drug atoms in the drug molecule to obtain the first distance feature.

[0205] In one or more embodiments, a method for generating the first node feature and the first distance feature is introduced. As can be seen from the foregoing embodiments, for each drug atom in the drug molecule, the corresponding first node feature and first distance feature of the drug atom can be obtained. Hereinafter, any drug atom will be taken as an example for illustration.

[0206] I. Atomic association data:

[0207] Specifically, obtaining the atomic association data corresponding to each drug atom in the drug molecule, and the atomic association data includes one or more of the following:

[0208] (1) Atomic number: That is, each atom is assigned a unique identifier (identity document, ID). For example, the atomic number of a hydrogen atom is "1", and the atomic number of an oxygen atom is "5".

[0209] (2) Neighbor count (atom_degree): That is, it represents the number of atoms that have an edge connection relationship with the drug atom. For example, if a certain drug atom is connected to 2 other atoms, then the neighbor count of this drug atom is "2".

[0210] (3) Formal charge: That is, the formal charge of an atom is expressed as "FC = V - N - B / 2", where FC represents the formal charge of the atom, V represents the number of valence electrons of an isolated neutral atom, N represents the number of non-bonding valence electrons of this drug atom in the drug molecule, and B represents the total number of electrons shared in bonding with other drug atoms in the drug molecule.

[0211] (4) Radical electrons (atom_num_radical_electrons): That is, under external conditions such as light and heat, the number of atoms or groups with unpaired electrons formed by the homolytic cleavage of covalent bonds in the drug molecule.

[0212] (5) Hybrid orbital (atom_hybridization): That is, it represents the type of hybrid orbital. For example, "SP", "SP2", "SP3", "SP3D", and "SP3D2".

[0213] (6) Aromatic ring correlation (atom_is_aromatic): That is, it represents whether it belongs to an aromatic ring. For example, "1" means it belongs to an aromatic ring, and "0" means it does not belong to an aromatic ring.

[0214] (7) Number of hydrogen atoms connected (atom_total_num_H): That is, it represents the number of hydrogen atoms connected to the drug atom.

[0215] (8) Chiral center (atom_is_chiral_center): That is, it represents whether it belongs to a chiral center. For example, "1" means it belongs to a chiral center, and "0" means it does not belong to a chiral center.

[0216] (9) Chiral type (atom_chirality_type): That is, the chiral type can be a left-handed type or a right-handed type.

[0217] (10) Molecular type (amino_type): That is, it represents the amino acid type or the drug molecule type. For example, "0" represents a drug molecule, and "1" represents aspartic acid.

[0218] Perform feature extraction on atomic association data based on feature engineering. Among them, feature extraction can be performed on atomic number, number of neighbors, formal charge, radical electrons, aromatic ring correlation, number of attached hydrogen atoms, chiral centers, and molecular type to obtain 1D feature values. Feature extraction can be performed on chiral types to obtain 2D one-hot encoded vectors. Feature extraction can be performed on hybrid orbitals to obtain 5D one-hot encoded vectors.

[0219] Based on this, the first node feature of a drug atom can be represented as a 16D vector, that is, [16, 1].

[0220] II. Distance data:

[0221] Specifically, according to the target topological graph, obtain the distance data between pairwise drug atoms in the drug molecule.

[0222] For ease of understanding, please refer to Figure 6 , Figure 6 , which is a schematic diagram of the distance data between pairwise atoms in the embodiments of the present application. As shown in the figure, assume that the drug molecule includes 5 drug atoms, namely drug atom v1, drug atom v2, drug atom v3, drug atom v4, and drug atom v5. The first distance feature of each drug atom includes the distance values between this drug atom and each other drug atom. Among them, the darker the color of the square, the closer the distance between the atoms. Thus, a 5×5 distance feature matrix is obtained.

[0223] Based on this, the first distance feature corresponding to a drug atom can be represented as an L-dimensional vector, that is, [L, 1]. Where L represents the total number of drug atoms in the drug molecule. According to the first distance features of each drug atom, the distance feature matrix of the drug molecule can be obtained, that is, [L, L, 1]. Where the "1" in [L, L, 1] represents 1 matrix.

[0224] Secondly, in the embodiments of the present application, a method for generating the first node feature and the first distance feature is provided. Through the above method, the node features corresponding to drug atoms can be constructed based on atomic association data, and the distance features of drug atoms can be constructed based on distance data. Thus, drug atoms can be described more comprehensively, thereby improving the accuracy of subsequent model prediction.

[0225] Optionally, on the basis of the above Figure 3 corresponding embodiments, in another optional embodiment provided by the embodiments of the present application, the first atomic feature further includes a first edge feature;

[0226] It may further include:

[0227] According to the target topological graph, obtain the edge data between pairwise drug atoms in the drug molecule, where the edge data includes at least one of the covalent bond type and the covalent bond position relationship;

[0228] Perform feature processing on the edge data between pairwise drug atoms in the drug molecule to obtain the first edge feature.

[0229] In one or more embodiments, a method for generating the first edge feature is introduced. As can be seen from the foregoing embodiments, for each drug atom in the drug molecule, the corresponding first edge feature of the drug atom can be obtained. Hereinafter, any drug atom will be taken as an example for illustration.

[0230] Specifically, according to the target topological graph, obtain the edge data between pairwise drug atoms in the drug molecule. The edge data includes one or more of the following:

[0231] (1) Covalent bond type (bond_type): It can be divided into polar covalent bonds and non-polar covalent bonds. For example, "0" indicates no covalent bond, "1" indicates a non-polar covalent bond, and "2" indicates a polar covalent bond.

[0232] (2) Covalent bond position relationship (bond_is_in_ring): That is, it indicates whether the covalent bond is in a ring. For example, "0" indicates that the covalent bond is in a ring, and "1" indicates that the covalent bond is not in a ring.

[0233] Based on this, the first edge feature corresponding to a drug atom can be represented as a vector of dimension L, that is, [L,1]. Where L represents the total number of drug atoms in the drug molecule. According to the first edge features of each drug atom, the edge feature matrix of the drug molecule can be obtained, that is, [L,L,2]. Where the "2" in [L,L,2] represents 2 matrices, one represents the matrix corresponding to the covalent bond type, and the other represents the matrix corresponding to the covalent bond position relationship.

[0234] Again, in the embodiments of the present application, a method for generating the first edge feature is provided. Through the above method, the edge feature corresponding to the drug atom can be constructed based on the edge data. Thus, the drug atom can be more comprehensively described, thereby further improving the accuracy of subsequent model prediction.

[0235] Optionally, on the basis of the corresponding embodiments described above, in another optional embodiment provided by the embodiments of the present application, the first atomic feature further includes a first quantization feature; Figure 3 It may further include:

[0236]

[0237] ​For each drug atom in the drug molecule, according to the target topological graph, obtain the quantization data corresponding to the drug atom, where the quantization data includes at least one of the covalent bond angle, the interaction angle, and the local charge. The covalent bond angle is the angle formed by the drug atom as the vertex and the two closest atoms. The interaction angle is the angle formed by the drug atom as the vertex, the first atom, and the second atom. The first atom is the atom that has an edge connection with the drug atom and is the closest one, and the second atom is the closest receptor atom to the drug atom.

[0238] For each drug atom in the drug molecule, perform feature extraction on the quantization data corresponding to the drug atom to obtain the first quantization feature.

[0239] In one or more embodiments, a method for generating the first quantization feature is introduced. As can be seen from the foregoing embodiments, for each drug atom in the drug molecule, the corresponding first quantization feature can be obtained. Hereinafter, an example of any drug atom will be used for illustration.

[0240] Specifically, obtain the quantization data corresponding to each drug atom in the drug molecule. The quantization data includes one or more of the following:

[0241] (1) Covalent bond angle: That is, the angle formed by the drug atom as the vertex and the two closest atoms.

[0242] For ease of understanding, please refer to Figure 7 , Figure 7 which is a schematic diagram of the covalent bond angle in the embodiment of the present application. As shown in the figure, taking the 0th drug atom as an example, the 0th drug atom is connected to the 1st receptor atom, the 2nd receptor atom, the 3rd receptor atom, and the 4th drug atom. Among them, the distance between the 0th drug atom and the 1st receptor atom is The distance between the 0th drug atom and the 2nd receptor atom is The distance between the 0th drug atom and the 3rd receptor atom is The distance between the 0th drug atom and the 4th drug atom is

[0243] It can be seen that the two closest atoms to the 0th drug atom are the 2nd receptor atom and the 4th drug atom. Therefore, the covalent bond angle is the angle θ.

[0244] (2) Interaction angle: That is, the angle formed by the drug atom as the vertex, the first atom, and the second atom. Here, the first atom is the atom that has an edge connection with the drug atom and is the closest one, and the second atom is the closest receptor atom to the drug atom.

[0245] For ease of understanding, please refer toFigure 8 , Figure 8 is a schematic diagram of the interaction angle in the embodiment of the present application. As shown in the figure, taking the 0th drug atom as an example, the 0th drug atom is connected to the 1st receptor atom, the 2nd receptor atom, the 3rd receptor atom, and the 4th drug atom. Among them, the distance between the 0th drug atom and the 1st receptor atom is The distance between the 0th drug atom and the 2nd receptor atom is The distance between the 0th drug atom and the 3rd receptor atom is The distance between the 0th drug atom and the 4th drug atom is

[0246] It can be seen that first, find the atom closest to the 0th drug atom, that is, the first atom is the 4th drug atom. Then find the atom in the nearest relative molecule. For a drug atom, its relative molecule is an amino acid molecule. Therefore, find the receptor atom included in the amino acid molecule closest to the 0th drug atom, that is, the second atom is the 2nd drug atom. Therefore, the interaction angle is the angle α.

[0247] (3) Partial charge: that is, the local charge value corresponding to the small molecule atom obtained during the pretreatment using the Merck Molecular Force Field 94 (MMFF94).

[0248] Based on feature engineering, the quantitative data is characterized. Among them, the covalent bond angle and the local charge can be characterized to obtain 1-dimensional feature values. It should be noted that for a drug atom, at most 4 covalent bonds can be connected. Therefore, at most 4 interaction angles can be formed. Exemplarily, assume that a certain drug atom is only connected to 1 covalent bond. Then, after characterizing the interaction angle, 4-dimensional feature values are obtained, where 3 feature values are "0", and the other feature value is the degree of the interaction angle.

[0249] Based on this, the first quantitative feature of a drug atom can be represented as a 6-dimensional vector, that is, [6,1].

[0250] Furthermore, in the embodiment of the present application, a method for generating the first quantitative feature is provided. Through the above method, the quantitative feature corresponding to the drug atom can be constructed based on the quantitative data. Thus, the drug atom can be more comprehensively described, thereby further improving the accuracy of subsequent model prediction.

[0251] Optionally, based on the above Figure 3 corresponding embodiments, in another optional embodiment provided by the embodiment of the present application, the second atomic feature includes a second node feature and a second distance feature;

[0252] According to the target topological graph, generate the second atomic features for each receptor atom in the binding pocket, which may specifically include:

[0253] Obtain the atomic association data corresponding to each receptor atom in the binding pocket, where the atomic association data includes at least one of atomic number, number of neighbors, formal charge, radical electrons, hybrid orbitals, aromatic ring correlation, number of connected hydrogen atoms, chiral center, chiral type, and molecular type;

[0254] For each receptor atom in the binding pocket, perform feature extraction on the atomic association data corresponding to the receptor atom to obtain the second node features;

[0255] According to the target topological graph, obtain the distance data between every two receptor atoms in the binding pocket;

[0256] Perform feature extraction on the distance data between every two receptor atoms in the binding pocket to obtain the second distance features.

[0257] In one or more embodiments, a method for generating the second node features and the second distance features is introduced. As can be seen from the foregoing embodiments, for each receptor atom in the pocket, the corresponding second node features and the second distance features can be obtained. Hereinafter, any one receptor atom will be taken as an example for illustration.

[0258] I. Atomic association data:

[0259] Specifically, obtain the atomic association data corresponding to each receptor atom in the pocket, and the atomic association data includes one or more of the following:

[0260] (1) Atomic number: That is, each atom is assigned a unique identifier (identity document, ID). For example, the atomic number of a hydrogen atom is "1", and the atomic number of an oxygen atom is "5".

[0261] (2) Number of neighbors: That is, it represents the number of atoms having an edge connection relationship with the receptor atom. For example, if a certain receptor atom is connected to 2 other atoms, the number of neighbors of the receptor atom is "2".

[0262] (3) Formal charge: That is, the formal charge of an atom is expressed as "FC = V - N - B / 2", where FC represents the formal charge of the atom, V represents the number of valence electrons of an isolated neutral atom, N represents the number of non-bonding valence electrons of the receptor atom in the pocket, and B represents the total number of electrons shared in bonding with other receptor atoms in the pocket.

[0263] (4) Radical electrons (atom_num_radical_electrons): That is, under external conditions such as photothermal conditions, the number of atoms or groups with unpaired electrons formed by the homolytic cleavage of covalent bonds in the pocket.

[0264] (5) Hybrid orbitals (atom_hybridization): That is, it represents the type of hybrid orbitals. For example, "SP", "SP2", "SP3", "SP3D", and "SP3D2".

[0265] (6) Aromatic ring correlation (atom_is_aromatic): That is, it represents whether it belongs to an aromatic ring. For example, "1" means it belongs to an aromatic ring, and "0" means it does not belong to an aromatic ring.

[0266] (7) Number of hydrogen atoms connected (atom_total_num_H): That is, it represents the number of hydrogen atoms connected to the receptor atom.

[0267] (8) Chiral center (atom_is_chiral_center): That is, it represents whether it belongs to a chiral center. For example, "1" means it belongs to a chiral center, and "0" means it does not belong to a chiral center.

[0268] (9) Chiral type (atom_chirality_type): That is, the chiral type can be a left-handed type or a right-handed type.

[0269] (10) Molecular type (amino_type): That is, it represents the amino acid type or the drug molecule type. For example, "0" means a drug molecule, and "1" means aspartic acid.

[0270] Based on feature engineering, the atomic association data is characterized. Among them, the atomic number, number of neighbors, formal charge, radical electrons, aromatic ring correlation, number of hydrogen atoms connected, chiral center, and molecular type can be characterized to obtain 1D feature values. The chiral type can be characterized to obtain a 2D one-hot encoded vector. The hybrid orbitals can be characterized to obtain a 5D one-hot encoded vector.

[0271] Based on this, the second node feature of a receptor atom can be represented as a 16D vector, that is, [16,1].

[0272] II. Distance data:

[0273] Specifically, according to the target topological graph, the distance data between pairwise receptor atoms in the pocket is obtained.

[0274] It is understandable that, as described in the foregoing embodiments, it is assumed that the pocket includes 5 receptor atoms, namely receptor atom y1, receptor atom y2, receptor atom y3, receptor atom y4, and receptor atom y5. The second distance feature of each receptor atom includes the distance values between this receptor atom and each other receptor atom. Thus, a 5×5 distance feature matrix is obtained.

[0275] Based on this, the second distance feature corresponding to one receptor atom can be represented as a vector of P dimensions, that is, [P,1]. Where P represents the total number of receptor atoms in the pocket. According to the second distance features of each receptor atom, the distance feature matrix of the pocket can be obtained, that is, [P,P,1]. Where the "1" in [P,P,1] represents 1 matrix.

[0276] Secondly, in the embodiments of the present application, a method for generating the second node feature and the second distance feature is provided. Through the above method, the node feature corresponding to the receptor atom can be constructed based on the atomic association data, and the distance feature of the receptor atom can be constructed based on the distance data. Thus, the receptor atom can be described more comprehensively, thereby improving the accuracy of subsequent model prediction.

[0277] Optionally, on the basis of the corresponding various embodiments above, in another optional embodiment provided by the embodiments of the present application, the second atomic feature further includes a second edge feature; Figure 3 It may further include:

[0278] According to the target topological graph, obtain the edge connection data between pairwise receptor atoms in the binding pocket, where the edge connection data includes at least one of the covalent bond type and the covalent bond position relationship;

[0279] Perform feature processing on the edge connection data between pairwise receptor atoms in the binding pocket to obtain the second edge feature.

[0280] In one or more embodiments, a method for generating the second edge feature is introduced. As can be seen from the foregoing embodiments, for each receptor atom in the pocket, the second edge feature corresponding to the receptor atom can be obtained. Hereinafter, any one receptor atom will be taken as an example for illustration.

[0281] Specifically, according to the target topological graph, obtain the edge connection data between pairwise receptor atoms in the pocket. The edge connection data includes one or more of the following:

[0282] (1) Covalent bond type (bond_type): It can be divided into polar covalent bonds and non-polar covalent bonds. For example, "0" indicates no covalent bond, "1" indicates a non-polar covalent bond, and "2" indicates a polar covalent bond.

[0283] (1) Covalent bond type (bond_type): It can be divided into polar covalent bonds and non-polar covalent bonds. For example, "0" indicates no covalent bond, "1" indicates a non-polar covalent bond, and "2" indicates a polar covalent bond.

[0284] (2) Covalent bond position relationship (bond_is_in_ring): That is, it indicates whether the covalent bond is in a ring. For example, "0" indicates that the covalent bond is in the ring, and "1" indicates that the covalent bond is not in the ring.

[0285] Based on this, the second edge feature corresponding to a receptor atom can be represented as an L-dimensional vector, that is, [P, 1]. Where P represents the total number of receptor atoms in the pocket. According to the second edge features of each receptor atom, the edge feature matrix of the pocket can be obtained, that is, [P, P, 2]. Where the "2" in [P, P, 2] represents two matrices, one represents the matrix corresponding to the covalent bond type, and the other represents the matrix corresponding to the covalent bond position relationship.

[0286] Again, in the embodiments of the present application, a method for generating the second edge feature is provided. Through the above method, the edge feature corresponding to the receptor atom can be constructed based on the edge connection data. Thus, the receptor atom can be more comprehensively described, thereby further improving the accuracy of subsequent model prediction.

[0287] Optionally, based on the above Figure 3 In another optional embodiment provided by the embodiments of the present application on the basis of the corresponding various embodiments, the second atomic feature further includes a second quantization feature;

[0288] It may further include:

[0289] For each receptor atom in the binding pocket, according to the target topological graph, obtain the quantization data corresponding to the receptor atom, where the quantization data includes at least one of the covalent bond angle, interaction angle, and local charge. The covalent bond angle is the angle formed by the receptor atom as the vertex and the two closest atoms. The interaction angle is the angle formed by the receptor atom as the vertex and the third atom and the fourth atom. The third atom is the atom with the closest edge connection relationship to the receptor atom, and the fourth atom is the closest drug atom to the receptor atom;

[0290] For each receptor atom in the binding pocket, perform feature processing on the quantization data corresponding to the receptor atom to obtain the second quantization feature.

[0291] In one or more embodiments, a method for generating the second quantization feature is introduced. As can be seen from the foregoing embodiments, for each receptor atom in the pocket, the second quantization feature corresponding to the receptor atom can be obtained. Hereinafter, any one receptor atom will be taken as an example for illustration.

[0292] Specifically, obtain the quantization data corresponding to each receptor atom in the pocket. The quantization data includes one or more of the following:

[0293] (1) Covalent bond angle: That is, with the receptor atom as the vertex, the angle formed by the two closest atoms.

[0294] It should be noted that the method for determining the covalent bond angle can be referred to the foregoing embodiments and will not be elaborated here.

[0295] (2) Interaction angle: That is, with the receptor atom as the vertex, the angle formed by the third atom and the fourth atom. Here, the third atom is the atom with the closest edge connection relationship with the receptor atom, and the fourth atom is the closest drug atom to the receptor atom.

[0296] It should be noted that the method for determining the interaction angle can be referred to the foregoing embodiments and will not be elaborated here.

[0297] (3) Partial charge: That is, the local charge value corresponding to the small molecule atom obtained during MMFF94 preprocessing.

[0298] Based on feature engineering, the quantization data is characterized. Among them, the covalent bond angle and the local charge can be characterized to obtain 1D feature values. It should be noted that for a receptor atom, at most 4 covalent bonds can be connected. Therefore, at most 4 interaction angles can be formed. Exemplarily, assume that a certain receptor atom is only connected to 1 covalent bond. Then, after characterizing the interaction angle, 4D feature values are obtained, among which 3 feature values are "0", and the other feature value is the degree of the interaction angle.

[0299] Based on this, the second quantization feature of a receptor atom can be represented as a 6D vector, that is, [6,1].

[0300] Furthermore, in the embodiments of the present application, a method for generating the second quantization feature is provided. Through the above method, the quantization feature corresponding to the receptor atom can be constructed based on the quantization data. Thus, the receptor atom can be more comprehensively described, thereby further improving the accuracy of subsequent model prediction.

[0301] Optionally, on the basis of the foregoing Figure 3 corresponding embodiments, in another optional embodiment provided by the embodiments of the present application, the drug molecule includes L drug atoms, and the binding pocket includes P receptor atoms, where both L and P are integers greater than or equal to 1;

[0302] Based on the first atomic feature of each drug atom and the second atomic feature of each receptor atom, the molecular interaction feature is obtained through the internal attention network included in the affinity prediction model, which specifically may include:

[0303] Perform feature embedding processing on the first atomic features of each drug atom to obtain L first input embedding features and H first bias embedding features, where H represents the number of attention heads and H is an integer greater than or equal to 1;

[0304] Perform feature embedding processing on the second atomic features of each receptor atom to obtain P second input embedding features and H second bias embedding features;

[0305] Based on the L first input embedding features and the H first bias embedding features, obtain drug molecule interaction features through the first attention network included in the internal attention network, where the internal attention network belongs to the affinity prediction model;

[0306] Based on the P second input embedding features and the H second bias embedding features, obtain binding pocket interaction features through the second attention network included in the internal attention network;

[0307] Concatenate the drug molecule interaction features and the binding pocket interaction features to obtain molecule interaction features.

[0308] In one or more embodiments, a method for generating molecule interaction features is introduced. As can be seen from the foregoing embodiments, the first atomic features of each drug atom may include first node features, first distance features, first edge features, and first quantization features, and the second atomic features of each receptor atom may include second node features, second distance features, second edge features, and second quantization features. This will be illustrated with examples below.

[0309] (1) The first node features of L drug atoms are represented as [L, 16];

[0310] (2) The first distance features of L drug atoms are represented as [L, L, 1];

[0311] (3) The first edge features of L drug atoms are represented as [L, L, 2];

[0312] (4) The first quantization features of L drug atoms are represented as [L, 6];

[0313] (5) The second node features of P receptor atoms are represented as [P, 16];

[0314] (6) The second distance features of P receptor atoms are represented as [P, P, 1];

[0315] (7) The second edge features of P receptor atoms are represented as [P, P, 2];

[0316] (8) The second quantization features of P receptor atoms are represented as [P, 6];

[0317] Exemplarily, the first node features of L drug atoms are subjected to feature embedding processing in the following manner to obtain L node embedding features:

[0318] Embedding -> [L, 16] -> Index Embedding [L, 16, 256] -> Sum [dim = 1] -> [L, 256];

[0319] Exemplarily, the first distance features of L drug atoms are subjected to feature embedding processing in the following manner to obtain L distance embedding features:

[0320] Embedding -> [L, L, 1] -> Gaussian radial basis function (RBF) layer -> [H, L, L];

[0321] Exemplarily, the first edge features of L drug atoms are subjected to feature embedding processing in the following manner to obtain L edge embedding features:

[0322] Embedding -> [L, L, 2] -> Linear [L, L, 2, H] -> Sum [dim = 2] -> [H, L, L];

[0323] Exemplarily, the first quantization features of L drug atoms are subjected to feature embedding processing in the following manner to obtain L quantization embedding features:

[0324] Embedding -> [L, 6] -> Index Embedding [L, 6, 256] -> Sum [dim = 1] -> [L, 256];

[0325] Exemplarily, the second node features of P receptor atoms are subjected to feature embedding processing in the following manner to obtain P node embedding features:

[0326] Embedding -> [P, 16] -> Index Embedding [P, 16, 256] -> Sum [dim = 1] -> [P, 256];

[0327] Exemplarily, the second distance features of P receptor atoms are subjected to feature embedding processing in the following manner to obtain P distance embedding features:

[0328] Embedding -> [P, P, 1] -> Gaussian RBF layer -> [H, P, P];

[0329] Exemplarily, the second edge features of P receptor atoms are subjected to feature embedding processing in the following manner to obtain P edge embedding features:

[0330] Embedding -> [P, P, 2] -> Linear[P, P, 2, H] -> Sum[dim = 2] -> [H, P, P];

[0331] Exemplarily, the second quantization features of P receptor atoms are subjected to feature embedding processing in the following manner to obtain P quantization embedding features:

[0332] Embedding -> [P, 6] -> Index Embedding[P, P, 256] -> Sum[dim = 1] -> [P, 256];

[0333] Among them, H represents the number of attention heads (multi - head). Based on this, the following results can be further obtained:

[0334] L node embedding features [L, 256] + L quantization embedding features [L, 256] -> L first input embedding features [L, 256];

[0335] L distance embedding features [H, L, L] + L edge embedding features [H, L, L] -> H first bias embedding features [H, L, L];

[0336] P node embedding features [P, 256] + P quantization embedding features [P, 256] -> P second input embedding features [P, 256];

[0337] P distance embedding features [H, P, P] + P edge embedding features [H, P, P] -> H second bias embedding features [H, P, P];

[0338] For ease of understanding, please refer to Figure 9 , Figure 9This is a schematic diagram of a network structure of the affinity prediction model in the embodiments of the present application. As shown in the figure, the affinity prediction model includes an intra-interaction network and an inter-interaction network. The intra-interaction network includes a first attention network and a second attention network. Among them, the first attention network and the second attention network can share weights. Based on this, L first input embedding features and H first bias embedding features are used as the input of the first attention network, and the drug molecule interaction features are output through the first attention network. Similarly, P second input embedding features and H second bias embedding features are used as the input of the second attention network, and the binding pocket interaction features are output through the second attention network. Thus, the drug molecule interaction features and the binding pocket interaction features are concatenated to obtain the molecular interaction features. Then, the molecular interaction features and the target embedding features are used as the input of the inter-interaction network, and the predicted affinity is output through the inter-interaction network.

[0339] It can be understood that the basic module of the affinity prediction model is the multi-head attention from the Transformer. Among them, the intra-interaction network only allows the model to interact with the information inside the small molecule or inside the pocket. The inter-interaction network is to learn the interaction features between the drug molecule and the pocket through the molecular interaction features output by the previous layer.

[0340] Secondly, in the embodiments of the present application, a method for generating molecular interaction features is provided. Through the above method, using the first atomic features of L drug atoms and the second atomic features of P receptor atoms, based on this, the internal features of the drug molecule and the amino acid molecule can be learned respectively to obtain the molecular interaction features. It can be seen that the molecular interaction features can better represent the internal relationship of the molecule, thus helping to improve the accuracy of affinity prediction.

[0341] Optionally, based on the corresponding embodiments above Figure 3 In another optional embodiment provided by the embodiments of the present application, based on L first input embedding features and H first bias embedding features, obtaining the drug molecule interaction features through the first attention network included in the intra-interaction network may specifically include:

[0342] Based on the L first input embedding features, L first query vectors, L first key vectors, and L first value vectors are obtained through the linear network layer included in the first attention network, where the first attention network belongs to the intra-interaction network;

[0343] Perform matrix multiplication on L first query vectors and L first key vectors to obtain H first intermediate matrices;

[0344] Based on the H first intermediate matrices and H first bias embedding features, obtain H first attention matrices through the normalization exponential function layer included in the first attention network;

[0345] Perform matrix multiplication on the H first attention matrices and L first value vectors to obtain H first target matrices;

[0346] Based on the H first target matrices, obtain the drug molecule interaction features through the target neural network included in the first attention network;

[0347] Based on P second input embedding features and H second bias embedding features, obtain the binding pocket interaction features through the second attention network included in the internal attention network, which may specifically include:

[0348] Based on the P second input embedding features, obtain P second query vectors, P second key vectors, and P second value vectors through the linear network layer included in the second attention network, where the second attention network belongs to the internal attention network;

[0349] Perform matrix multiplication on the P second query vectors and P second key vectors to obtain H second intermediate matrices;

[0350] Based on the H second intermediate matrices and H second bias embedding features, obtain H second attention matrices through the normalization exponential function layer included in the second attention network;

[0351] Perform matrix multiplication on the H second attention matrices and P second value vectors to obtain H second target matrices;

[0352] Based on the H second target matrices, obtain the binding pocket interaction features through the target neural network included in the second attention network.

[0353] In one or more embodiments, a method for realizing intramolecular interaction based on an internal attention network is introduced. As can be seen from the foregoing embodiments, the internal attention network includes a first attention network and a second attention network, and the first attention network and the second attention network each include an attention bias module, and the attention bias module is used to provide an adjustment of the attention weights between nodes.

[0354] Specifically, for ease of understanding, please refer to Figure 10 , Figure 10This is a schematic diagram of a network structure of the internal attention network in an embodiment of the present application. As shown in the figure, assuming that the drug molecule includes L drug atoms and the pocket includes P receptor atoms, according to the foregoing embodiments, L first input embedding features corresponding to the L drug atoms and H first bias embedding features can be obtained, and P second input embedding features corresponding to the P receptor atoms and H second bias embedding features can be obtained. The following will be combined with Figure 10 for illustration.

[0355] I. Process of obtaining drug molecule interaction features:

[0356] (1) Input the L first input embedding features into the linear network layer (linear) of the first attention network. Thus, L first query vectors, L first key vectors, and L first value vectors are obtained, that is:

[0357] QKV = [L, 256] -> [L, 256](Q)[L, 256](K)[L, 256](V);

[0358] Among them, Q represents the query vector. K represents the key vector. V represents the value vector. [L, 256] represents the L first input embedding features. [L, 256](Q) represents the L first query vectors. [L, 256](K) represents the L first key vectors. [L, 256](V) represents the L first value vectors.

[0359] (2) Perform matrix multiplication on the L first query vectors and the L first key vectors. Thus, H first intermediate matrices are obtained, that is:

[0360] [L, 256](Q) + [L, 256](K) -> Matmul(Q * K^T) -> [H, L, L];

[0361] Among them, [H, L, L] represents the H first intermediate matrices. H represents the number of attention heads. [L, 256](Q) represents the L first query vectors. [L, 256](K) represents the L first key vectors. T represents transpose. Matmul represents matrix multiplication.

[0362] (3) Add the H first intermediate matrices and the H first bias embedding features. Thus, H first summation matrices are obtained. Then, call the softmax layer included in the first attention network to calculate the H first summation matrices, and H first attention matrices can be output, that is:

[0363] [H, L, L] + Intra - Attention Bias[H, L, L] -> [H, L, L] -> Softmax -> [H, L, L];

[0364] Among them, the first [H, L, L] represents H first intermediate matrices. Intra - Attention Bias[H, L, L] represents H first bias embedding features. The second [H, L, L] represents H first summation matrices. The third [H, L, L] represents H first attention matrices. Softmax represents the normalized exponential function.

[0365] (4) Multiply the H first attention matrices and the L first value vectors, thereby obtaining H first target matrices, that is:

[0366] [H, L, L] + [L, 256](V) -> Matmul(Q * K^T * V) -> [H, L, 256 / H];

[0367] Among them, [H, L, L] represents H first attention matrices. [L, 256](V) represents L first value vectors. Matmul represents matrix multiplication. [H, L, 256 / H] represents H first target matrices.

[0368] (5) Input the H first target matrices into the target neural network included in the first attention network, and output the drug - molecule interaction features through the target neural network, that is:

[0369] [H, L, 256 / H] -> Transition - FFN -> [L, 256 * 4] -> [L, 256];

[0370] Among them, [H, L, 256 / H] represents H first target matrices. The target neural network can be a Transition Feed - Forward Network (Transition - FFN), and Transition - FFN includes two layers of linear, for example, 256 -> 1024 -> 256, and there are two weights, namely [256, 1024] and [1024, 256]. [L, 256 * 4] represents L interaction features to be processed. [L, 256] represents the drug - molecule interaction features.

[0371] II. Process of obtaining the binding - pocket interaction features:

[0372] (1) Input the P second input embedding features into the linear of the second attention network, thereby obtaining P second query vectors, P second key vectors, and P second value vectors, that is:

[0373] QKV = [P, 256] -> [P, 256](Q)[P, 256](K)[P, 256](V);

[0374] Among them, Q represents the query vector. K represents the key vector. V represents the value vector. [P, 256] represents P second input embedding features. [P, 256](Q) represents P second query vectors. [P, 256](K) represents P second key vectors. [P, 256](V) represents P second value vectors.

[0375] (2) Perform matrix multiplication on the P second query vectors and the P second key vectors. Thus, H second intermediate matrices are obtained, that is:

[0376] [P, 256](Q) + [P, 256](K) -> Matmul(Q * K^T) -> [H, P, P];

[0377] Among them, [H, P, P] represents H second intermediate matrices. H represents the number of attention heads. [P, 256](Q) represents P second query vectors. [P, 256](K) represents P second key vectors. T represents transpose. Matmul represents matrix multiplication.

[0378] (3) Add the H second intermediate matrices and the H second bias embedding features. Thus, H second sum matrices are obtained. Then, call the softmax layer included in the second attention network to calculate the H second attention matrices, that is:

[0379] [H, P, P] + Intra - AttentionBias[H, P, P] -> [H, P, P] -> Softmax -> [H, P, P];

[0380] Among them, the first [H, P, P] represents H second intermediate matrices. Intra - AttentionBias[H, P, P] represents H second bias embedding features. The second [H, P, P] represents H second sum matrices. The third [H, P, P] represents H second attention matrices. Softmax represents the normalized exponential function.

[0381] (4) Perform matrix multiplication on the H second attention matrices and the P second value vectors. Thus, H second target matrices are obtained, that is:

[0382] [H, P, P] + [P, 256](V) -> Matmul(Q * K^T * V) -> [H, P, 256 / H];

[0383] Among them, [H, P, P] represents H second attention matrices. [P, 256](V) represents P second value vectors. Matmul represents matrix multiplication. [H, L, 256 / H] represents H second target matrices.

[0384] (5) Input the H second target matrices into the target neural network included in the second attention network, and output the combined pocket interaction features through the target neural network, that is:

[0385] [H, P, 256 / H]->Transition-FFN->[P, 256*4]->[P, 256];

[0386] Among them, [H, P, 256 / H] represents H second target matrices. The target neural network can be Transition-FFN. Transition-FFN includes two layers of linear, for example, 256->1024->256, and there are two weights, namely [256, 1024] and [1024, 256]. [P, 256*4] represents P to-be-processed interaction features. [P, 256] represents the combined pocket interaction features.

[0387] Based on this, for the sake of easy understanding, please refer to Figure 11 , Figure 11 which is a schematic diagram of a network structure of the attention bias module in the embodiment of the present application. As shown in the figure, after adding the edge feature matrix and the distance feature matrix of the molecule and inputting them into the attention bias network, the corresponding H bias embedding features can be output.

[0388] Again, in the embodiment of the present application, a method for realizing the internal interaction of molecules based on the internal attention network is provided. Through the above method, the first attention network and the second attention network in the internal attention network are used to process the drug molecule and the pocket respectively, thereby realizing the internal interaction of the molecule. In addition, based on the attention bias module, the attention weights between nodes can be adjusted, thereby improving the feasibility and operability of the solution.

[0389] Optionally, on the basis of the above Figure 3 corresponding embodiments, in another optional embodiment provided by the embodiment of the present application, the drug molecule includes L drug atoms, and the binding pocket includes P receptor atoms, where L and P are both integers greater than or equal to 1;

[0390] Based on the molecular interaction features, the affinity for the protein-ligand conformation is obtained through the interaction attention network included in the affinity prediction model, which specifically may include:

[0391] Concatenate the molecular interaction features and the target embedding features of the virtual nodes to obtain N input embedding features, where the N input embedding features include drug molecular interaction features, binding pocket interaction features, and target embedding features, and N is equal to (L + P + 1);

[0392] Based on the N input embedding features, obtain N target query vectors, N target key vectors, and N target value vectors through the linear network layer included in the interactive attention network, where the interactive attention network belongs to the affinity prediction model;

[0393] Perform matrix multiplication on the N target query vectors and the N target key vectors to obtain H target intermediate matrices, where H represents the number of attention heads, and H is an integer greater than or equal to 1;

[0394] Based on the H target intermediate matrices and the H target bias embedding features, obtain H target attention matrices through the normalization exponential function layer included in the interactive attention network, where the H target bias embedding features are generated according to the first atomic features of each drug atom, the second atomic features of each receptor atom, and the atomic features of the virtual node;

[0395] Perform matrix multiplication on the H target attention matrices and the N target value vectors to obtain H target matrices;

[0396] Based on the H target matrices, obtain the feature to be processed through the target neural network included in the interactive attention network;

[0397] Obtain the target feature to be processed according to the feature to be processed;

[0398] Based on the target feature to be processed, obtain the affinity for the protein-ligand conformation through the output layer.

[0399] In one or more embodiments, a method for realizing the interaction between molecules based on the interactive attention network is introduced. As can be seen from the foregoing embodiments, the interactive attention network includes an attention bias module, and the attention bias module is used to provide the attention weights for adjusting the nodes.

[0400] Specifically, for the sake of understanding, please refer to Figure 12 , Figure 12 which is a schematic diagram of a network structure of the interactive attention network in the embodiments of the present application. As shown in the figure, assume that the drug molecule includes L drug atoms, and the pocket includes P receptor atoms. Combining the foregoing embodiments, the molecular interaction features can be obtained. The following will be described in conjunction with Figure 12 for illustration.

[0401] (1) Concatenate the molecular interaction features and the target embedding features of the virtual node. Thus, N input embedding features are obtained. A virtual node is a node that connects to other nodes.

[0402] [N,256] = [1,256] || [L,256] || [P,256];

[0403] Among them, [N,256] represents N input embedding features. [1,256] represents the target embedding features of the virtual node. [L,256] represents the molecular interaction features of the drug. [P,256] represents the binding pocket interaction features. "||" represents concatenation. N is equal to L + P + 1, and the dimension of the hidden state hidden_dim = 256.

[0404] (2) Input the N input embedding features into the linear layer of the interactive attention network. Thus, N target query vectors, N target key vectors, and N target value vectors are obtained, that is:

[0405] QKV = [1,256] || [L,256] || [P,256] -> [N,256](Q)[N,256](K)[N,256](V);

[0406] Among them, Q represents the query vector. K represents the key vector. V represents the value vector. [N,256] represents N input embedding features. [N,256](Q) represents N target query vectors. [N,256](K) represents N target key vectors. [N,256](V) represents N target value vectors.

[0407] (3) Perform matrix multiplication on the N target query vectors and the N target key vectors. Thus, H target intermediate matrices are obtained, that is:

[0408] [N,256](Q) + [N,256](K) -> Matmul(Q * K^T) -> [H,N,N];

[0409] Among them, [H,N,N] represents H target intermediate matrices. H represents the number of attention heads. [N,256](Q) represents N target query vectors. [N,256](K) represents N target key vectors. T represents transpose. Matmul represents matrix multiplication.

[0410] (4) Add the H target intermediate matrices and the H target bias embedding features. Thus, H target sum matrices are obtained. Then, call the softmax layer included in the interactive attention network to calculate the H target sum matrices, and H target attention matrices can be output, that is:

[0411] [H, N, N] + Inter - AttentionBias(complex)[H, N, N] -> [H, N, N] -> Softmax -> [H, N, N];

[0412] Among them, the first [H, N, N] represents H target intermediate matrices. Inter - AttentionBias(complex)[H, N, N] represents H target bias embedding features. The second [H, N, N] represents H target summation matrices. The third [H, N, N] represents H target attention matrices. Softmax represents the normalized exponential function.

[0413] It should be noted that L distance embedding features [H, L, L] + L edge embedding features [H, L, L] + P distance embedding features [H, P, P] + P edge embedding features [H, P, P] + 1 distance embedding feature [H, 1, 1] of the virtual node + 1 edge embedding feature [H, 1, 1] of the virtual node -> H target bias embedding features [H, N, N].

[0414] (5) Multiply the H target attention matrices and the N target value vectors, thereby obtaining H target matrices, that is:

[0415] [H, N, N] + [N, 256](V) -> Matmul(Q * K^T * V) -> [H, N, 256 / H];

[0416] Among them, [H, L, L] represents H target attention matrices. [N, 256](V) represents N target value vectors. Matmul represents matrix multiplication. [H, N, 256 / H] represents H target matrices.

[0417] (6) Input the H target matrices into the target neural network included in the interactive attention network, and output the to - be - processed features through the target neural network, that is:

[0418] [H, N, 256 / H] -> Transition - FFN -> [N, 256 * 4] -> [N, 256];

[0419] Among them, [H, N, 256 / H] represents H target matrices. The target neural network can be Transition - FFN. Transition - FFN includes two layers of linear, for example, 256 -> 1024 -> 256, and there are two weights, namely [256, 1024] and [1024, 256]. [N, 256 * 4] represents N to - be - processed interactive features. [N, 256] represents the to - be - processed features.

[0420] (7) Based on this, the target features to be processed can be obtained according to the features to be processed, that is:

[0421] For easier understanding, see Figure 13 , Figure 13 This is a schematic diagram of a network structure of the output layer in an embodiment of the present application. As shown in the figure, the virtual node in the first column of the feature to be processed is taken as the target feature to be processed.

[0422] [N,256]->get the first column of virtual node->[1,256];

[0423] Among them, [N, 256] represents the features to be processed, and [1, 256] represents the target features to be processed.

[0424] (8) Inputting the target feature to be processed into the output layer included in the affinity prediction model, thereby obtaining the affinity of the protein ligand conformation, that is:

[0425] [1,256]->Linear->[1];

[0426] Among them, [1,256] represents the target feature to be processed. Linear represents the linear network layer in the output layer. [1] represents the affinity of the protein ligand conformation.

[0427] Secondly, in the embodiment of the present application, a method for realizing interaction between molecules based on an interactive attention network is provided. In the above manner, the interactive attention network is used to process the interaction between drug molecules and pockets, thereby realizing interaction between molecules. In addition, based on the attention bias module, the attention weights between nodes can be adjusted, thereby improving the feasibility and operability of the solution.

[0428] Optionally, in the above Figure 3 On the basis of the corresponding embodiments, another optional embodiment provided by the embodiment of the present application may further include:

[0429] Obtaining a first weight parameter and a second weight parameter from a gamma matrix, wherein the gamma matrix is ​​used to describe a characteristic relationship between a drug molecule and a binding pocket;

[0430] Generate H biased embedding features to be processed according to the first atomic feature of each drug atom, the second atomic feature of each receptor atom, and the atomic features of the virtual node;

[0431] A first weight parameter is used to adjust a first set of values ​​in each bias embedding feature to be processed, and a second weight parameter is used to adjust a second set of values ​​in each bias embedding feature to be processed, so as to obtain H target bias embedding features.

[0432] In one or more embodiments, a method of processing a drug molecule and a pocket using a gamma matrix is introduced. As can be seen from the foregoing embodiments, first, H to-be-processed bias embedding features can be constructed based on the distance embedding features and edge embedding features corresponding to the drug molecule, and the distance embedding features and edge embedding features corresponding to the pocket. Among them, the first distance embedding feature and the first edge embedding feature are obtained based on the first atomic feature, and the second distance embedding feature and the second edge embedding feature are obtained based on the second atomic feature.

[0433] Specifically, the gamma matrix is represented as a 2×2 matrix, that is, the gamma matrix includes 2×2 learnable weight parameters. The gamma matrix can describe the feature relationship between the drug molecule and the pocket. Among them, the first weight parameter corresponding to the [0,1] position in the gamma matrix is used to adjust the features from the drug molecule to the pocket, and the second weight parameter corresponding to the [1,0] position in the gamma matrix is used to adjust the features from the pocket to the drug molecule. Based on this, the following method is used for feature adjustment:

[0434] AttentionBias(Ligand)*w1->Inter-AttentionBias(Ligand);

[0435] AttentionBias(Pocket)*w2->Inter-AttentionBias(Pocket);

[0436] Among them, AttentionBias(Ligand) represents the features from the drug molecule to the pocket among the H to-be-processed bias embedding features. w1 represents the first weight parameter. AttentionBias(Pocket) represents the features from the pocket to the drug molecule among the H to-be-processed bias embedding features. w2 represents the second weight parameter.

[0437] For ease of understanding, please refer to Figure 14 , Figure 14 which is a schematic diagram of the to-be-processed bias embedding features in the embodiments of the present application. As shown in the figure, taking one to-be-processed bias embedding feature as an example, where z1 represents a virtual node, v1 to v5 represent drug atoms in the drug molecule, and y1 to y9 represent receptor atoms in the pocket. It can be seen that the features from the drug molecule to the pocket are composed of white squares, and the features from the pocket to the drug molecule are composed of black squares.

[0438] Again, in the embodiments of the present application, a method for processing drug molecules and pockets using a gamma matrix is provided. Through the above method, since the weight parameters in the gamma matrix are learnable, these weight parameters can be used to adjust the distance features and edge features between drug molecules and pockets, thereby achieving a better interaction effect.

[0439] Optionally, based on the respective corresponding embodiments above, in another optional embodiment provided by the embodiments of the present application, after obtaining H target attention matrices through the normalization exponential function layer included in the interactive attention network based on H target intermediate matrices and H target bias embedding features, it may further include: Figure 3 After obtaining H target attention matrices through the normalization exponential function layer included in the interactive attention network based on H target intermediate matrices and H target bias embedding features, it may further include:

[0440] Performing an averaging process on the values at the same positions in the H target attention matrices to obtain an average attention matrix;

[0441] Obtaining an association distribution matrix from the average attention matrix, where the association distribution matrix represents the association between L drug atoms and P receptor atoms.

[0442] In one or more embodiments, a method for displaying the interaction relationship between molecules is introduced. As can be seen from the foregoing embodiments, adding the H target intermediate matrices and the H target bias embedding features results in H target sum matrices. Then, by calling the softmax layer included in the interactive attention network to calculate the H target sum matrices, H target attention matrices can be output. Based on this, an averaging process is performed on the H target attention matrices, that is, the values at the corresponding positions are averaged, thereby obtaining the average values at each position. The average values at each position constitute the average attention matrix, and the value range of the average value is from 0 to 1. The larger the value, the better the interaction between atoms.

[0443] Specifically, for ease of understanding, please refer to Figure 15 , Figure 15 is a schematic diagram of the average attention matrix in the embodiments of the present application. As shown in the figure, z1 represents a virtual node, v1 to v5 represent drug atoms in the drug molecule, and y1 to y9 represent receptor atoms in the pocket. Taking out the association distribution matrix from the average attention matrix, where Figure 15 the one circled by the thick black frame in is the association distribution matrix, that is:

[0444] 2D attention matrix = [ligand_len, pocket_len];

[0445] Among them, the 2D attention matrix represents the correlation distribution matrix. ligand_len represents L drug atoms. pocket_len represents P receptor atoms.

[0446] Again, in the embodiments of the present application, a method for displaying the intermolecular interaction relationship is provided. Through the above method, the correlation distribution matrix can be used to reflect the interpretability of the prediction result, enabling relevant personnel to understand and analyze the prediction result, thus helping to expand the application prospect.

[0447] Optionally, based on the above Figure 3 In another optional embodiment provided by the embodiments of the present application on the basis of the corresponding respective embodiments, it may further include:

[0448] Obtain the original conformation sample, where the original conformation sample is constructed based on the original protein sample and the original drug molecule sample;

[0449] Generate T training topology graph samples according to the original conformation sample, where each training topology graph sample includes the training drug molecule sample, and each training topology graph sample has a corresponding true affinity value, and T is an integer greater than or equal to 1;

[0450] Based on the T training topology graph samples, obtain the affinity prediction value corresponding to each training topology graph sample through the affinity prediction model;

[0451] Update the model parameters in the affinity prediction model according to the affinity prediction value and the true affinity value corresponding to each training topology graph sample.

[0452] In one or more embodiments, a method for training an affinity prediction model is introduced. As can be seen from the foregoing embodiments, before the affinity prediction model is put into use, it also needs to be trained. The training of the affinity prediction model mainly includes two steps. The first step only involves the interaction of the internal information of the molecule, and the second step is the interaction between the drug molecule and the pocket. Exemplarily, the hyperparameters of the affinity prediction model are as follows:

[0453] The dimension of the hidden state is hidden_dim = 256;

[0454] The dropout rate is dropout = 0.2;

[0455] The attention dropout rate is attention_dropout = 0.2;

[0456] The number of attention heads Wienum_heads = 8;

[0457] The number of network layers is n_layers = 12;

[0458] The learning rate is lr = 0.0006;

[0459] The batch size is batch_size = 20;

[0460] The weight decay is L2 - decay = 1e - 3;

[0461] The early stopping count is early - stop - patience = 10;

[0462] During training, input the atomic features corresponding to each training topological graph sample among T training topological graph samples, and obtain the affinity prediction value corresponding to each training topological graph sample through the affinity prediction model. Then compare it with the true affinity value, calculate the loss value using a loss function (such as mean absolute error (MAE) or mean square error (MSE), etc.), and use the loss value for backpropagation to update the model parameters.

[0463] It can be understood that the method of extracting each atomic feature based on the training topological graph sample is similar to that based on the target topological graph, so it will not be elaborated here.

[0464] Secondly, in the embodiments of the present application, a method for training an affinity prediction model is provided. Through the above method, data augmentation can be achieved using molecular docking applications to cope with the lack of training sets, and better networks can be trained by improving the quantity and quality of existing data. In addition, constructing negative samples for training can enable the model to have the characteristics of Pose - Sensitive, thereby improving the prediction ability of the model and bringing more possibilities for the realization of model applications.

[0465] Optionally, on the basis of the above Figure 3 In another optional embodiment provided by the embodiments of the present application corresponding to each of the above - mentioned embodiments, it may further include:

[0466] For each training topological graph sample among the T training topological graph samples, determine the molecular structure error between the training drug molecule sample included in the training topological graph sample and the original drug molecule sample;

[0467] According to the molecular structure error corresponding to each training topological graph sample, determine the label category corresponding to each training topological graph sample. Among them, the label category corresponding to the training topological graph sample with a molecular structure error less than the error threshold is a positive sample label, and the label category corresponding to the training topological graph sample with a molecular structure error greater than or equal to the error threshold is a negative sample label;

[0468] Based on T training topological graph samples, obtain the predicted class probabilities corresponding to each training topological graph sample through an affinity prediction model;

[0469] According to the affinity prediction value and the true affinity value corresponding to each training topological graph sample, update the model parameters in the affinity prediction model, which may specifically include:

[0470] According to the affinity prediction value and the true affinity value corresponding to each training topological graph sample, determine a first loss value using a first loss function;

[0471] According to the predicted class probability and the label class corresponding to each training topological graph sample, determine a second loss value using a second loss function;

[0472] Update the model parameters in the affinity prediction model according to the first loss value and the second loss value.

[0473] In one or more embodiments, a method for increasing the classification function to train an affinity prediction model is introduced. As can be seen from the foregoing embodiments, using the enhanced data to train the affinity prediction model can make the predicted value of the affinity prediction model for the correct pose higher than that for the non-ideal pose, greatly improving the sensitivity of the model to the pose.

[0474] Specifically, for easy understanding, please refer to Figure 16 , Figure 16 This is a schematic diagram of data enhancement based on the original conformation sample in the embodiments of the present application. As shown in the figure, taking the original conformation sample as the eutectic data as an example, use the Docking application to perform molecular reconstruction (redock) to generate T training topological graph samples. T is an integer greater than or equal to 1. For example, T = 20. By this means, assuming there are 10,000 original conformation samples, 200,000 training topological graph samples can be constructed. For easy introduction, the following will take 1 original conformation sample as an example for illustration.

[0475] For the T training topological graph samples, in addition to obtaining the corresponding true affinity values through experiments, it is also necessary to label the tags of each training topological graph sample. For easy understanding, please refer to the following method:

[0476]

[0477] Among them, lable represents the true affinity value of the training topological graph sample. pValue represents the true affinity value of the training topological graph sample. RMSD is the root mean square deviation, that is, it represents the molecular structure error between the redocked training topological graph sample and the original conformation sample.

[0478] It can be seen that for the training topological graph samples with molecular structure errors less than the error threshold, the corresponding label category is the positive sample label, that is, pValue > 0. For the training topological graph samples with molecular structure errors greater than or equal to the error threshold, the corresponding label category is the negative sample label, that is, pValue < 0. The affinity between atoms in positive samples is higher than that in negative samples.

[0479] It should be noted that in this application, the error threshold is In actual situations, the error threshold can also be set to other values. This is only an illustration here and should not be construed as a limitation to this application.

[0480] On the one hand, the first loss value can be calculated in the following way. According to the affinity prediction value, the true affinity value, and the label category corresponding to each training topological graph sample, the first loss value is calculated in the following way:

[0481]

[0482] Among them, diff represents the first loss value. pred represents the affinity prediction value. lable represents the true affinity value of the training topological graph sample.

[0483] When the prediction is a positive sample label, the MSE loss function can be used as the first loss function. Thus, the first loss value is calculated. That is, MSE = mean(diff^2), where diff = pred - lable.

[0484] When the prediction is a negative sample label, if the affinity prediction value is less than the true affinity value, the first loss value is 0.

[0485] When the prediction is a negative sample label, if the affinity prediction value is greater than or equal to the true affinity value, then the MSE loss function can be used as the first loss function. Thus, the first loss value is calculated. That is, MSE = mean(diff^2), where diff = pred + lable.

[0486] On the other hand, a classification layer can be added to the output layer of the affinity prediction model, that is, a linear layer is connected to the virtual node of the output layer to output the predicted class probability in one dimension. Among them, the predicted class probability can represent the probability that the training topological graph sample is predicted as the positive sample label. Based on this, the second loss function is used to judge whether the prediction result is correct, that is, pose scoring is performed to determine whether the prediction result is a positive sample or a negative sample. The second loss function can be Binary Cross Entropy (BCE), and thus, the second loss value is calculated.

[0487] Finally, the first loss value and the second loss value are summed to obtain the total loss value. The total loss value is used to update the model parameters in the affinity prediction model.

[0488] Again, in the embodiment of the present application, a method for adding a classification function to train the affinity prediction model is provided. Through the above method, multi-task training will improve the Pose-Sensitive of the output layer, which is beneficial to be applied to the virtual screening scenario (that is, applied when the proportion of active molecules is very low, such as 1:200).

[0489] The affinity prediction model provided by the present application has obtained better performance on the following three test sets, which are respectively:

[0490] Test set one: A test set composed of 3400 data points of G Protein-Coupled Receptors (GPCR), kinase, and protease family targets, called the docking_test, and evaluated using the Pearson Correlation R.

[0491] Test set two: The activity data collected from research papers, including 111 targets and 7554 pairs of test data, called v3-test, and evaluated using Pearson Correlation R.

[0492] Test set three: 4 representative targets, with approximately 200 activity data and 20000 decoy data for each target, called VS-Test, and the virtual screening performance is evaluated using the enrichment score (EF).

[0493] The affinity prediction device in the present application will be described in detail below. Please refer to Figure 17 , Figure 17A schematic diagram of an embodiment of the affinity prediction device in the embodiments of the present application. The affinity prediction device 20 includes:

[0494] An acquisition module 210, configured to acquire a protein-ligand conformation, where the protein-ligand conformation is constructed based on a protein and a drug molecule;

[0495] A generation module 220, configured to generate a target topological graph according to the protein-ligand conformation, where the target topological graph includes a node set and an edge set. The node set is used to represent drug atoms in the drug molecule and receptor atoms in the binding pocket, and the edge set is used to represent chemical bonds connecting atoms. The binding pocket is generated based on the protein;

[0496] The generation module 220 is further configured to generate a first atomic feature of each drug atom in the drug molecule according to the target topological graph;

[0497] The generation module 220 is further configured to generate a second atomic feature of each receptor atom in the binding pocket according to the target topological graph;

[0498] The acquisition module 210 is further configured to obtain molecular interaction features through an internal attention network included in the affinity prediction model based on the first atomic feature of each drug atom and the second atomic feature of each receptor atom;

[0499] The acquisition module 210 is further configured to obtain the affinity for the protein-ligand conformation through an interaction attention network included in the affinity prediction model based on the molecular interaction features.

[0500] In the embodiments of the present application, an affinity prediction device is provided. By using the above device, a target topological graph constructed by using rich structural information of the protein-ligand conformation is used. Combining the target topological graph, the atomic features of each drug atom in the drug molecule and the atomic features of each receptor atom in the amino acid molecule can be respectively extracted. Based on this, first, based on the atomic features of each drug atom and each receptor atom, the internal features of the drug molecule and the amino acid molecule are respectively learned, thereby obtaining molecular interaction features. Then, the interaction attention network is used to learn the molecular interaction features to obtain the affinity. It can be seen that when predicting the affinity of a protein-ligand, the internal features of the molecule and the features between molecules can be fully learned, which is beneficial to predicting a more accurate affinity.

[0501] Optionally, based on the corresponding embodiment above, Figure 17 In another embodiment of the affinity prediction device 20 provided in the embodiments of the present application,

[0502] The acquisition module 210 is specifically configured to acquire a first format file corresponding to the protein;

[0503] Obtain the second format file corresponding to the drug molecule;

[0504] Based on the first format file and the second format file, generate X candidate protein-ligand conformations and the scores of each candidate protein-ligand conformation through a molecular docking application, where X is an integer greater than or equal to 1;

[0505] Select the candidate protein-ligand conformation with the highest score among the X candidate protein-ligand conformations as the protein-ligand conformation.

[0506] In the embodiments of the present application, an affinity prediction device is provided. Using the above device, since the Docking application predicts the binding conformation of the small molecule ligand to the target binding site, it is possible to screen bioactive compounds at the early stage of drug development with the help of this Docking application, thereby increasing the flexibility and diversity of the synthesized protein-ligand conformations.

[0507] Optionally, on the basis of the above Figure 17 In another embodiment of the affinity prediction device 20 provided in the embodiments of the present application,

[0508] The generation module 220 is specifically configured to determine the drug atoms belonging to the drug molecule according to the protein-ligand conformation;

[0509] Determine M amino acid molecules belonging to the protein according to the protein-ligand conformation, and determine the receptor atoms included in each amino acid molecule, where the M amino acid molecules belong to the protein and M is an integer greater than or equal to 1;

[0510] Taking each drug atom in the drug molecule as the center, determine at least one receptor atom whose atomic distance is less than or equal to the distance threshold;

[0511] Obtain the distance between each receptor atom in the at least one receptor atom and the drug molecule;

[0512] Sort the at least one receptor atom in ascending order of distance;

[0513] For the at least one receptor atom after sorting, sequentially add the amino acid molecules to which the receptor atoms belong until the graph construction stop condition is met to obtain the target topological graph.

[0514] In the embodiments of the present application, an affinity prediction device is provided. Using the above device, considering that the GNN has the problem of over-smoothing and can only consist of one or two layers, more layers cannot bring better performance. And the affinity prediction model is based on the topological graph as the input, which can have up to 30 layers and has better and deeper expression capabilities.

[0515] Optionally, in the above Figure 17On the basis of the corresponding embodiment, in another embodiment of the affinity prediction device 20 provided in the embodiment of the present application, the first atomic feature includes a first node feature and a first distance feature;

[0516] A generating module 220 is specifically used to obtain atomic association data corresponding to each drug atom in the drug molecule, wherein the atomic association data includes at least one of an element number, a number of neighbors, a formal charge, a free radical electron, a hybridization orbital, an aromatic ring correlation, a number of connected hydrogen atoms, a chiral center, a chiral type, and a molecular type;

[0517] For each drug atom in the drug molecule, the atomic association data corresponding to the drug atom is characterized to obtain a first node feature;

[0518] According to the target topological map, the distance data between each pair of drug atoms in the drug molecule is obtained;

[0519] The distance data between each pair of drug atoms in the drug molecule are characterized to obtain a first distance feature.

[0520] In an embodiment of the present application, an affinity prediction device is provided. Using the above device, node features corresponding to drug atoms can be constructed based on atom association data, and distance features of drug atoms can be constructed based on distance data, thereby being able to more comprehensively describe drug atoms, thereby improving the accuracy of subsequent model predictions.

[0521] Optionally, in the above Figure 17 On the basis of the corresponding embodiment, in another embodiment of the affinity prediction device 20 provided in the embodiment of the present application, the first atomic feature further includes a first edge feature;

[0522] The acquisition module 210 is further used to acquire edge data between two drug atoms in the drug molecule according to the target topological graph, wherein the edge data includes at least one of a covalent bond type and a covalent bond position relationship;

[0523] The generation module 220 is further used to perform characterization processing on the edge data between each pair of drug atoms in the drug molecule to obtain a first edge feature.

[0524] In an embodiment of the present application, an affinity prediction device is provided. By using the above device, edge features corresponding to drug atoms can be constructed based on edge data, thereby being able to more comprehensively describe drug atoms, thereby further improving the accuracy of subsequent model predictions.

[0525] Optionally, in the above Figure 17 On the basis of the corresponding embodiment, in another embodiment of the affinity prediction device 20 provided in the embodiment of the present application, the first atomic feature further includes a first quantitative feature;

[0526] The obtaining module 210 is further configured to obtain, for each drug atom in the drug molecule, quantization data corresponding to the drug atom according to the target topological graph, where the quantization data includes at least one of a covalent bond angle, an interaction angle, and a local charge. The covalent bond angle is an angle formed by the drug atom as a vertex and the two closest atoms. The interaction angle is an angle formed by the drug atom as a vertex, a first atom, and a second atom. The first atom is an atom that has an edge relationship with the drug atom and is the closest one, and the second atom is a receptor atom that is the closest to the drug atom.

[0527] The generating module 220 is further configured to perform feature extraction on the quantization data corresponding to each drug atom in the drug molecule to obtain a first quantization feature.

[0528] In an embodiment of the present application, an affinity prediction device is provided. By using the above device, quantization features corresponding to drug atoms can be constructed based on quantization data. Thus, drug atoms can be more comprehensively described, further improving the accuracy of subsequent model prediction.

[0529] Optionally, based on the above Figure 17 In another embodiment of the affinity prediction device 20 provided in the embodiment of the present application, on the basis of the corresponding embodiment, the second atom feature includes a second node feature and a second distance feature.

[0530] The generating module 220 is specifically configured to obtain atom association data corresponding to each receptor atom in the binding pocket, where the atom association data includes at least one of an element serial number, a neighbor number, a formal charge, a radical electron, a hybrid orbital, an aromatic ring correlation, a number of connected hydrogen atoms, a chiral center, a chiral type, and a molecular type.

[0531] For each receptor atom in the binding pocket, perform feature extraction on the atom association data corresponding to the receptor atom to obtain a second node feature.

[0532] According to the target topological graph, obtain distance data between every two receptor atoms in the binding pocket.

[0533] Perform feature extraction on the distance data between every two receptor atoms in the binding pocket to obtain a second distance feature.

[0534] In an embodiment of the present application, an affinity prediction device is provided. By using the above device, node features corresponding to receptor atoms can be constructed based on atom association data, and distance features corresponding to receptor atoms can be constructed based on distance data. Thus, receptor atoms can be more comprehensively described, further improving the accuracy of subsequent model prediction.

[0535] Optionally, based on the aboveFigure 17 Based on the corresponding embodiment, in another embodiment of the affinity prediction device 20 provided in the embodiments of the present application, the second atomic feature further includes a second edge feature;

[0536] The acquisition module 210 is further configured to obtain the edge connection data between pairwise receptor atoms in the binding pocket according to the target topological graph, where the edge connection data includes at least one of the covalent bond type and the covalent bond position relationship;

[0537] The generation module 220 is further configured to perform feature processing on the edge connection data between pairwise receptor atoms in the binding pocket to obtain a second edge feature.

[0538] In the embodiments of the present application, an affinity prediction device is provided. By using the above device, edge features corresponding to receptor atoms can be constructed based on the edge connection data. Thus, receptor atoms can be more comprehensively described, thereby further improving the accuracy of subsequent model prediction.

[0539] Optionally, based on the corresponding embodiment above, in another embodiment of the affinity prediction device 20 provided in the embodiments of the present application, the second atomic feature further includes a second quantization feature; Figure 17 Based on the corresponding embodiment above, in another embodiment of the affinity prediction device 20 provided in the embodiments of the present application, the second atomic feature further includes a second quantization feature;

[0540] The acquisition module 210 is further configured to, for each receptor atom in the binding pocket, obtain the quantization data corresponding to the receptor atom according to the target topological graph, where the quantization data includes at least one of the covalent bond angle, the interaction angle, and the local charge. The covalent bond angle is the angle formed by the receptor atom as the vertex and the two closest atoms. The interaction angle is the angle formed by the receptor atom as the vertex, the third atom, and the fourth atom. The third atom is the closest atom having an edge connection relationship with the receptor atom, and the fourth atom is the closest drug atom to the receptor atom;

[0541] The generation module 220 is further configured to, for each receptor atom in the binding pocket, perform feature processing on the quantization data corresponding to the receptor atom to obtain a second quantization feature.

[0542] In the embodiments of the present application, an affinity prediction device is provided. By using the above device, quantization features corresponding to receptor atoms can be constructed based on the quantization data. Thus, receptor atoms can be more comprehensively described, thereby further improving the accuracy of subsequent model prediction.

[0543] Optionally, based on the corresponding embodiment above, Figure 17 In another embodiment of the affinity prediction device 20 provided in the embodiments of the present application, the drug molecule includes L drug atoms, and the binding pocket includes P receptor atoms, where both L and P are integers greater than or equal to 1;

[0544] An acquisition module 210, specifically configured to perform feature embedding processing on the first atomic features of each drug atom to obtain L first input embedding features and H first bias embedding features, where H represents the number of attention heads, and H is an integer greater than or equal to 1;

[0545] Perform feature embedding processing on the second atomic features of each receptor atom to obtain P second input embedding features and H second bias embedding features;

[0546] Based on the L first input embedding features and the H first bias embedding features, obtain drug molecule interaction features through a first attention network included in an internal attention network, where the internal attention network belongs to an affinity prediction model;

[0547] Based on the P second input embedding features and the H second bias embedding features, obtain binding pocket interaction features through a second attention network included in the internal attention network;

[0548] Concatenate the drug molecule interaction features and the binding pocket interaction features to obtain molecule interaction features.

[0549] In an embodiment of the present application, an affinity prediction device is provided. By using the above device, based on the first atomic features of L drug atoms and the second atomic features of P receptor atoms, the internal features of drug molecules and amino acid molecules can be learned respectively to obtain molecule interaction features. It can be seen that the molecule interaction features can better represent the internal relationship of molecules, thereby helping to improve the accuracy of affinity prediction.

[0550] Optionally, on the basis of the corresponding embodiment above, in another embodiment of the affinity prediction device 20 provided in the embodiment of the present application, Figure 17 The acquisition module 210 is specifically configured to, based on the L first input embedding features, obtain L first query vectors, L first key vectors, and L first value vectors through a linear network layer included in the first attention network, where the first attention network belongs to the internal attention network;

[0551] Perform matrix multiplication on the L first query vectors and the L first key vectors to obtain H first intermediate matrices;

[0552] Based on the H first intermediate matrices and the H first bias embedding features, obtain H first attention matrices through a normalization exponential function layer included in the first attention network;

[0553] Perform matrix multiplication on the H first attention matrices and the L first value vectors to obtain H first target matrices;

[0554] Perform matrix multiplication on the H first attention matrices and the L first value vectors to obtain H first target matrices;

[0555] Based on H first target matrices, obtain the drug molecule interaction features through the target neural network included in the first attention network;

[0556] The obtaining module 210 is specifically configured to, based on P second input embedding features, obtain P second query vectors, P second key vectors, and P second value vectors through the linear network layer included in the second attention network, where the second attention network belongs to the internal attention network;

[0557] Perform matrix multiplication on the P second query vectors and the P second key vectors to obtain H second intermediate matrices;

[0558] Based on the H second intermediate matrices and H second bias embedding features, obtain H second attention matrices through the normalization exponential function layer included in the second attention network;

[0559] Perform matrix multiplication on the H second attention matrices and the P second value vectors to obtain H second target matrices;

[0560] Based on the H second target matrices, obtain the binding pocket interaction features through the target neural network included in the second attention network.

[0561] In the embodiments of the present application, an affinity prediction device is provided. By using the above device and utilizing the first attention network and the second attention network in the internal attention network to process the drug molecule and the pocket respectively, the internal interaction of the molecule is realized. In addition, based on the attention bias module, the attention weights between nodes can be adjusted, thereby improving the feasibility and operability of the solution.

[0562] Optionally, on the basis of the corresponding embodiments above, in another embodiment of the affinity prediction device 20 provided by the embodiments of the present application, the drug molecule includes L drug atoms, and the binding pocket includes P receptor atoms, where both L and P are integers greater than or equal to 1; Figure 17 The obtaining module 210 is specifically configured to splice the molecule interaction features and the target embedding features of the virtual node to obtain N input embedding features, where the N input embedding features include the drug molecule interaction features, the binding pocket interaction features, and the target embedding features, and N is equal to (L + P + 1);

[0563] Based on the N input embedding features, obtain N target query vectors, N target key vectors, and N target value vectors through the linear network layer included in the interaction attention network, where the interaction attention network belongs to the affinity prediction model;

[0564] Based on the N input embedding features, obtain N target query vectors, N target key vectors, and N target value vectors through the linear network layer included in the interaction attention network, where the interaction attention network belongs to the affinity prediction model;

[0565] Perform matrix multiplication on N target query vectors and N target key vectors to obtain H target intermediate matrices, where H represents the number of attention heads and H is an integer greater than or equal to 1;

[0566] Based on the H target intermediate matrices and H target bias embedding features, obtain H target attention matrices through the normalization exponential function layer included in the interactive attention network, where the H target bias embedding features are generated according to the first atomic feature of each drug atom, the second atomic feature of each receptor atom, and the atomic feature of the virtual node;

[0567] Perform matrix multiplication on the H target attention matrices and N target value vectors to obtain H target matrices;

[0568] Based on the H target matrices, obtain the feature to be processed through the target neural network included in the interactive attention network;

[0569] Obtain the target feature to be processed according to the feature to be processed;

[0570] Based on the target feature to be processed, obtain the affinity for the protein-ligand conformation through the output layer.

[0571] In the embodiments of the present application, an affinity prediction device is provided. By using the above device, the interactive attention network is used to perform interactive processing on the drug molecule and the pocket, thereby realizing the interaction between molecules. In addition, based on the attention bias module, the attention weights between nodes can be adjusted, thereby improving the feasibility and operability of the solution.

[0572] Optionally, on the basis of the corresponding embodiments described above Figure 17 In another embodiment of the affinity prediction device 20 provided in the embodiments of the present application, the affinity prediction device 20 further includes a processing module 230;

[0573] The acquisition module 210 is further configured to acquire a first weight parameter and a second weight parameter from the gamma matrix, where the gamma matrix is used to describe the feature relationship between the drug molecule and the binding pocket;

[0574] The generation module 220 is further configured to generate H to-be-processed bias embedding features according to the first atomic feature of each drug atom, the second atomic feature of each receptor atom, and the atomic feature of the virtual node;

[0575] The processing module 230 is configured to adjust the first numerical set in each to-be-processed bias embedding feature by using the first weight parameter, and adjust the second numerical set in each to-be-processed bias embedding feature by using the second weight parameter to obtain H target bias embedding features.

[0576] In an embodiment of the present application, an affinity prediction device is provided. By using the above device, since each weight parameter in the gamma matrix can be learned, these weight parameters can be used to adjust the distance features and edge features between the drug molecule and the pocket, thereby achieving a better interaction effect.

[0577] Optionally, based on the corresponding embodiment above, in another embodiment of the affinity prediction device 20 provided in the embodiment of the present application, Figure 17 the processing module 230 is further configured to, after obtaining H target attention matrices through the normalization exponential function layer included in the interactive attention network based on the H target intermediate matrices and the H target bias embedding features, perform an averaging process on the values at the same positions in the H target attention matrices to obtain an average attention matrix;

[0578] The obtaining module is further configured to obtain an association distribution matrix from the average attention matrix, where the association distribution matrix represents the association between L drug atoms and P receptor atoms.

[0579]

[0580] In an embodiment of the present application, an affinity prediction device is provided. By using the above device, the use of the association distribution matrix can reflect the interpretability of the prediction result, enabling relevant personnel to understand and analyze the prediction result, thereby helping to expand the application prospect.

[0581] Figure 17 Optionally, based on the corresponding embodiment above, in another embodiment of the affinity prediction device 20 provided in the embodiment of the present application, the affinity prediction device 20 further includes a training module 240;

[0582] The obtaining module 210 is further configured to obtain an original conformation sample, where the original conformation sample is constructed based on an original protein sample and an original drug molecule sample;

[0583] The generating module 220 is further configured to generate T training topology graph samples according to the original conformation sample, where each training topology graph sample includes a training drug molecule sample, and each training topology graph sample has a corresponding true affinity value, and T is an integer greater than or equal to 1;

[0584]

[0585] The obtaining module 210 is further configured to obtain an affinity prediction value corresponding to each training topology graph sample through the affinity prediction model based on the T training topology graph samples;

[0586] The training module 240 is configured to update the model parameters in the affinity prediction model according to the affinity prediction value and the true affinity value corresponding to each training topology graph sample.

[0586] In an embodiment of the present application, an affinity prediction device is provided. By using the above device, data augmentation can be achieved by means of molecular docking applications, so as to cope with the situation of lack of training sets, and better networks can be trained by improving the quantity and quality of existing data. In addition, constructing negative samples for training can enable the model to have the characteristics of Pose-Sensitive, thereby improving the prediction ability of the model and bringing more possibilities for the realization of model applications.

[0587] Optionally, on the basis of the corresponding embodiment above, in another embodiment of the affinity prediction device 20 provided in the embodiment of the present application, the affinity prediction device 20 further includes a determination module 250; Figure 17 For each training topology graph sample among the T training topology graph samples, the determination module 250 is configured to determine the molecular structure error between the training drug molecule sample included in the training topology graph sample and the original drug molecule sample;

[0588] The determination module 250 is further configured to determine the label category corresponding to each training topology graph sample according to the molecular structure error corresponding to each training topology graph sample, wherein the label category corresponding to the training topology graph sample with a molecular structure error less than the error threshold is a positive sample label, and the label category corresponding to the training topology graph sample with a molecular structure error greater than or equal to the error threshold is a negative sample label;

[0589] The acquisition module 210 is further configured to obtain the predicted category probability corresponding to each training topology graph sample through the affinity prediction model based on the T training topology graph samples;

[0590] The training module 240 is specifically configured to determine a first loss value by using a first loss function according to the affinity prediction value and the true affinity value corresponding to each training topology graph sample;

[0591] Determine a second loss value by using a second loss function according to the predicted category probability and the label category corresponding to each training topology graph sample;

[0592] Update the model parameters in the affinity prediction model according to the first loss value and the second loss value.

[0593]

[0594] In an embodiment of the present application, an affinity prediction device is provided. By using the above device, multi-task training can improve the Pose-Sensitive of the output layer, which is beneficial to be applied to the virtual screening scenario (that is, applied in the case where the proportion of active molecules is very low, such as 1:200).

[0595] Figure 18FIG. 0 is a schematic structural diagram of a computer device provided by an embodiment of the present application. The computer device 300 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 322 (for example, one or more processors) and a memory 332, and one or more storage media 330 (for example, one or more mass storage devices) storing application programs 342 or data 344. Among them, the memory 332 and the storage media 330 may be transient storage or persistent storage. The programs stored in the storage media 330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the computer device. Further, the central processing unit 322 may be configured to communicate with the storage media 330 and execute a series of instruction operations in the storage media 330 on the computer device 300.

[0596] The computer device 300 may further include one or more power supplies 326, one or more wired or wireless network interfaces 350, one or more input / output interfaces 358, and / or one or more operating systems 341, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM and so on.

[0597] The steps performed by the computer device in the above embodiments may be based on the Figure 18 shown computer device structure.

[0598] An embodiment of the present application further provides a computer device, including a memory and a processor. The memory stores a computer program. When the processor executes the computer program, the steps of the methods described in the foregoing embodiments are implemented.

[0599] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the methods described in the foregoing embodiments are implemented.

[0600] An embodiment of the present application further provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the methods described in the foregoing embodiments are implemented.

[0601] It can be understood that in the specific embodiments of the present application, data related to user information, etc. is involved. When the above embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0602] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0603] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling, direct coupling, or communication connection can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms.

[0604] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0605] In addition, the functional units in each embodiment of the present application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0606] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a server, a terminal, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media that can store computer programs, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0607] As described above, the above embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of various embodiments of this application.

Claims

1. A method for predicting the affinity of a protein ligand, characterized in that, Comprising: Obtaining a protein-ligand conformation, wherein the protein-ligand conformation is constructed based on a protein and a drug molecule; Generating a target topological graph according to the protein-ligand conformation, wherein the target topological graph includes a node set and an edge set, the node set is used to represent drug atoms in the drug molecule and receptor atoms in the binding pocket, and the edge set is used to represent chemical bonds connecting atoms, and the binding pocket is generated based on the protein; Generating a first atomic feature for each drug atom in the drug molecule according to the target topological graph; Generating a second atomic feature for each receptor atom in the binding pocket according to the target topological graph; Based on the first atomic feature of each drug atom and the second atomic feature of each receptor atom, obtaining a molecular interaction feature through an internal attention network included in the affinity prediction model; Based on the molecular interaction feature, obtaining the affinity for the protein-ligand conformation through an interaction attention network included in the affinity prediction model.

2. The affinity prediction method according to claim 1, wherein The obtaining of the protein-ligand conformation includes: Obtaining a first format file corresponding to the protein; Obtaining a second format file corresponding to the drug molecule; Based on the first format file and the second format file, generating X candidate protein-ligand conformations and the score of each candidate protein-ligand conformation through a molecular docking application, wherein X is an integer greater than or equal to 1; Selecting the candidate protein-ligand conformation with the highest score from the X candidate protein-ligand conformations as the protein-ligand conformation.

3. The affinity prediction method according to claim 1, characterized in that The generating of the target topological graph according to the protein-ligand conformation includes: Determining drug atoms belonging to the drug molecule according to the protein-ligand conformation; Determining M amino acid molecules belonging to the protein according to the protein-ligand conformation and determining receptor atoms included in each amino acid molecule, wherein the M amino acid molecules belong to the protein and M is an integer greater than or equal to 1; Taking each drug atom in the drug molecule as a center, determining at least one receptor atom with an atomic distance less than or equal to a distance threshold; Obtaining the distance between each receptor atom in the at least one receptor atom and the drug molecule; Sorting the at least one receptor atom in ascending order of distance; For the sorted at least one receptor atom, sequentially adding the amino acid molecule to which the receptor atom belongs until a graph construction stop condition is met, to obtain the target topological graph.

4. The affinity prediction method according to claim 1, wherein The first atomic feature includes a first node feature and a first distance feature; The generating of the first atomic feature for each drug atom in the drug molecule according to the target topological graph includes: Obtaining atomic association data corresponding to each drug atom in the drug molecule, wherein the atomic association data includes at least one of atomic number, number of neighbors, formal charge, radical electrons, hybrid orbitals, aromatic ring correlation, number of connected hydrogen atoms, chiral center, chiral type, and molecular type; For each drug atom in the drug molecule, perform feature extraction on the atom correlation data corresponding to the drug atom to obtain the first node feature; According to the target topological graph, obtain the distance data between every two drug atoms in the drug molecule; Perform feature extraction on the distance data between every two drug atoms in the drug molecule to obtain the first distance feature.

5. The affinity prediction method according to claim 4, wherein The first atomic feature further includes a first edge feature; The method further includes: According to the target topological graph, obtain the edge connection data between every two drug atoms in the drug molecule, where the edge connection data includes at least one of the covalent bond type and the covalent bond position relationship; Perform feature extraction on the edge connection data between every two drug atoms in the drug molecule to obtain the first edge feature.

6. The affinity prediction method according to claim 4 or 5, characterized in that The first atomic feature further includes a first quantization feature; The method further includes: For each drug atom in the drug molecule, according to the target topological graph, obtain the quantization data corresponding to the drug atom, where the quantization data includes at least one of the covalent bond angle, the interaction angle, and the local charge. The covalent bond angle is the angle formed by the drug atom as the vertex and the two closest atoms. The interaction angle is the angle formed by the drug atom as the vertex, the first atom, and the second atom. The first atom is the atom with the closest edge connection relationship and the shortest distance to the drug atom, and the second atom is the receptor atom with the shortest distance to the drug atom; For each drug atom in the drug molecule, perform feature extraction on the quantization data corresponding to the drug atom to obtain the first quantization feature.

7. The affinity prediction method according to claim 1, characterized in that The second atomic feature includes a second node feature and a second distance feature; Generating the second atomic feature of each receptor atom in the binding pocket according to the target topological graph includes: Obtain the atom correlation data corresponding to each receptor atom in the binding pocket, where the atom correlation data includes at least one of the element serial number, the number of neighbors, the formal charge, the radical electrons, the hybrid orbitals, the aromatic ring correlation, the number of connected hydrogen atoms, the chiral center, the chiral type, and the molecular type; For each receptor atom in the binding pocket, perform feature extraction on the atom correlation data corresponding to the receptor atom to obtain the second node feature; According to the target topological graph, obtain the distance data between every two receptor atoms in the binding pocket; Perform feature extraction on the distance data between every two receptor atoms in the binding pocket to obtain the second distance feature.

8. The affinity prediction method according to claim 7, wherein The second atomic feature further includes a second edge feature; The method further includes: According to the target topological graph, obtain the edge connection data between every two receptor atoms in the binding pocket, where the edge connection data includes at least one of the covalent bond type and the covalent bond position relationship; Perform feature extraction on the edge connection data between every two receptor atoms in the binding pocket to obtain the second edge feature.

9. The affinity prediction method according to claim 7 or 8, characterized in that The second atomic feature further includes a second quantization feature; The method further includes: For each receptor atom in the binding pocket, according to the target topology map, obtain the quantization data corresponding to the receptor atom, where the quantization data includes at least one of a covalent bond angle, an interaction angle, and a local charge. The covalent bond angle is an angle formed by the receptor atom as the vertex and the two closest atoms. The interaction angle is an angle formed by the receptor atom as the vertex, a third atom, and a fourth atom. The third atom is an atom that has an edge connection relationship with the receptor atom and is the closest one, and the fourth atom is a drug atom that is the closest to the receptor atom. For each receptor atom in the binding pocket, perform feature extraction on the quantization data corresponding to the receptor atom to obtain the second quantization feature.

10. The affinity prediction method according to claim 1, characterized in that, The drug molecule includes L drug atoms, and the binding pocket includes P receptor atoms, where both L and P are integers greater than or equal to 1. Based on the first atomic feature of each drug atom and the second atomic feature of each receptor atom, obtain the molecular interaction feature through the internal attention network included in the affinity prediction model, including: Perform feature embedding on the first atomic feature of each drug atom to obtain L first input embedding features and H first bias embedding features, where H represents the number of attention heads, and H is an integer greater than or equal to 1. Perform feature embedding on the second atomic feature of each receptor atom to obtain P second input embedding features and H second bias embedding features. Based on the L first input embedding features and the H first bias embedding features, obtain the drug molecule interaction feature through the first attention network included in the internal attention network, where the internal attention network belongs to the affinity prediction model. Based on the P second input embedding features and the H second bias embedding features, obtain the binding pocket interaction feature through the second attention network included in the internal attention network. Concatenate the drug molecule interaction feature and the binding pocket interaction feature to obtain the molecular interaction feature.

11. The affinity prediction method according to claim 10, characterized in that, Based on the L first input embedding features and the H first bias embedding features, obtain the drug molecule interaction feature through the first attention network included in the internal attention network, including: Based on the L first input embedding features, obtain L first query vectors, L first key vectors, and L first value vectors through the linear network layer included in the first attention network, where the first attention network belongs to the internal attention network. Perform matrix multiplication on the L first query vectors and the L first key vectors to obtain H first intermediate matrices. Based on the H first intermediate matrices and the H first bias embedding features, obtain H first attention matrices through the normalization exponential function layer included in the first attention network. Perform matrix multiplication on the H first attention matrices and the L first value vectors to obtain H first target matrices. Based on the H first target matrices, obtain the drug molecule interaction features through the target neural network included in the first attention network; The obtaining of the binding pocket interaction features based on the P second input embedding features and the H second bias embedding features through the second attention network included in the internal attention network includes: Based on the P second input embedding features, obtain P second query vectors, P second key vectors, and P second value vectors through the linear network layer included in the second attention network, where the second attention network belongs to the internal attention network; Perform matrix multiplication on the P second query vectors and the P second key vectors to obtain H second intermediate matrices; Based on the H second intermediate matrices and the H second bias embedding features, obtain H second attention matrices through the normalization exponential function layer included in the second attention network; Perform matrix multiplication on the H second attention matrices and the P second value vectors to obtain H second target matrices; Based on the H second target matrices, obtain the binding pocket interaction features through the target neural network included in the second attention network.

12. The affinity prediction method according to claim 1, wherein The drug molecule includes L drug atoms, and the binding pocket includes P receptor atoms, where both L and P are integers greater than or equal to 1; The obtaining of the affinity for the protein-ligand conformation based on the molecule interaction features through the interaction attention network included in the affinity prediction model includes: Concatenate the molecule interaction features and the target embedding features of the virtual node to obtain N input embedding features, where the N input embedding features include drug molecule interaction features, binding pocket interaction features, and the target embedding features, and N is equal to (L + P + 1); Based on the N input embedding features, obtain N target query vectors, N target key vectors, and N target value vectors through the linear network layer included in the interaction attention network, where the interaction attention network belongs to the affinity prediction model; Perform matrix multiplication on the N target query vectors and the N target key vectors to obtain H target intermediate matrices, where H represents the number of attention heads, and H is an integer greater than or equal to 1; Based on the H target intermediate matrices and H target bias embedding features, obtain H target attention matrices through the normalization exponential function layer included in the interaction attention network, where the H target bias embedding features are generated according to the first atomic features of each drug atom, the second atomic features of each receptor atom, and the atomic features of the virtual node; Perform matrix multiplication on the H target attention matrices and the N target value vectors to obtain H target matrices; Based on the obtained H target matrices, obtain the feature to be processed through the target neural network included in the interaction attention network; Obtain the target feature to be processed according to the feature to be processed; Based on the target feature to be processed, obtain the affinity for the protein-ligand conformation through the output layer.

13. The affinity prediction method according to claim 12, characterized in that, The method further includes: Obtain a first weight parameter and a second weight parameter from the gamma matrix, where the gamma matrix is used to describe the feature relationship between the drug molecule and the binding pocket; Generate H to-be-processed bias embedding features according to the first atomic feature of each drug atom, the second atomic feature of each receptor atom, and the atomic feature of the virtual node; Use the first weight parameter to adjust the first numerical set in each to-be-processed bias embedding feature, and use the second weight parameter to adjust the second numerical set in each to-be-processed bias embedding feature to obtain the H target bias embedding features.

14. The affinity prediction method according to claim 12, wherein After obtaining the H target attention matrices through the normalization exponential function layer included in the interactive attention network based on the H target intermediate matrices and the H target bias embedding features, the method further includes: Perform an averaging process on the values at the same positions in the H target attention matrices to obtain an average attention matrix; Obtain an association distribution matrix from the average attention matrix, where the association distribution matrix represents the association between the L drug atoms and the P receptor atoms.

15. The affinity prediction method according to claim 1, wherein The method further includes: Obtain an original conformation sample, where the original conformation sample is constructed based on an original protein sample and an original drug molecule sample; Generate T training topology graph samples according to the original conformation sample, where each training topology graph sample includes a training drug molecule sample, and each training topology graph sample has a corresponding true affinity value, and T is an integer greater than or equal to 1; Based on the T training topology graph samples, obtain the affinity prediction value corresponding to each training topology graph sample through an affinity prediction model; Update the model parameters in the affinity prediction model according to the affinity prediction value and the true affinity value corresponding to each training topology graph sample.

16. The affinity prediction method according to claim 15, wherein The method further includes: For each training topology graph sample among the T training topology graph samples, determine the molecular structure error between the training drug molecule sample included in the training topology graph sample and the original drug molecule sample; According to the molecular structure error corresponding to each training topology graph sample, determine the label category corresponding to each training topology graph sample, where the label category corresponding to the training topology graph sample with a molecular structure error less than the error threshold is a positive sample label, and the label category corresponding to the training topology graph sample with a molecular structure error greater than or equal to the error threshold is a negative sample label; Based on the T training topology graph samples, obtain the predicted category probability corresponding to each training topology graph sample through the affinity prediction model; The updating the model parameters in the affinity prediction model according to the affinity prediction value and the true affinity value corresponding to each training topology graph sample includes: Determine a first loss value using a first loss function according to the predicted affinity value and the true affinity value corresponding to each training topology graph sample; Determine a second loss value using a second loss function according to the predicted class probability and the label class corresponding to each training topology graph sample; Update the model parameters in the affinity prediction model according to the first loss value and the second loss value.

17. An affinity prediction device, characterized in that Includes: An acquisition module for acquiring a protein-ligand conformation, where the protein-ligand conformation is constructed based on a protein and a drug molecule; A generation module for generating a target topology graph according to the protein-ligand conformation, where the target topology graph includes a node set and an edge set, the node set is used to represent drug atoms in the drug molecule and receptor atoms in the binding pocket, and the edge set is used to represent chemical bonds connecting atoms, and the binding pocket is generated based on the protein; The generation module is further configured to generate first atomic features of each drug atom in the drug molecule according to the target topology graph; The generation module is further configured to generate second atomic features of each receptor atom in the binding pocket according to the target topology graph; The acquisition module is further configured to obtain molecular interaction features through an internal attention network included in the affinity prediction model based on the first atomic features of each drug atom and the second atomic features of each receptor atom; The acquisition module is further configured to obtain the affinity for the protein-ligand conformation through an interaction attention network included in the affinity prediction model based on the molecular interaction features.

18. The affinity prediction device according to claim 17, wherein The acquisition module is specifically configured to: Obtain a first format file corresponding to the protein; Obtain a second format file corresponding to the drug molecule; Based on the first format file and the second format file, generate X candidate protein-ligand conformations and the score of each candidate protein-ligand conformation through a molecular docking application, where X is an integer greater than or equal to 1; Select the candidate protein-ligand conformation with the highest score among the X candidate protein-ligand conformations as the protein-ligand conformation.

19. The affinity prediction device according to claim 17, wherein The generation module is specifically configured to: Determine drug atoms belonging to the drug molecule according to the protein-ligand conformation; Determine M amino acid molecules belonging to the protein according to the protein-ligand conformation, and determine the receptor atoms included in each amino acid molecule, where the M amino acid molecules belong to the protein and M is an integer greater than or equal to 1; Taking each drug atom in the drug molecule as the center, determine at least one receptor atom with an atomic distance less than or equal to a distance threshold; Obtain the distance between each receptor atom in the at least one receptor atom and the drug molecule; Sort the at least one receptor atom in ascending order of distance; For the sorted at least one receptor atom, sequentially add the amino acid molecule to which the receptor atom belongs until the graph construction stop condition is satisfied to obtain the target topology graph.

20. The affinity prediction device according to claim 17, wherein The first atomic feature includes a first node feature and a first distance feature; The generation module is specifically configured to: Obtain the atomic association data corresponding to each drug atom in the drug molecule, where the atomic association data includes at least one of atomic number, number of neighbors, formal charge, radical electrons, hybridization orbitals, aromatic ring correlation, number of connected hydrogen atoms, chiral center, chiral type, and molecular type; For each drug atom in the drug molecule, perform feature extraction on the atomic association data corresponding to the drug atom to obtain the first node feature; According to the target topological graph, obtain the distance data between every two drug atoms in the drug molecule; Perform feature extraction on the distance data between every two drug atoms in the drug molecule to obtain the first distance feature.

21. The affinity prediction device according to claim 20, wherein The first atomic feature further includes a first edge feature; The acquisition module is further configured to obtain the edge connection data between every two drug atoms in the drug molecule according to the target topological graph, where the edge connection data includes at least one of covalent bond type and covalent bond position relationship; The generation module is further configured to perform feature extraction on the edge connection data between every two drug atoms in the drug molecule to obtain the first edge feature.

22. The affinity prediction device according to claim 20 or 21, characterized in that The first atomic feature further includes a first quantization feature; The acquisition module is further configured to, for each drug atom in the drug molecule, obtain the quantization data corresponding to the drug atom according to the target topological graph, where the quantization data includes at least one of covalent bond angle, interaction angle, and local charge. The covalent bond angle is the angle formed by the drug atom as the vertex and the two closest atoms. The interaction angle is the angle formed by the drug atom as the vertex, the first atom, and the second atom. The first atom is the atom with the closest edge connection relationship and the shortest distance to the drug atom, and the second atom is the receptor atom with the shortest distance to the drug atom; The generation module is further configured to, for each drug atom in the drug molecule, perform feature extraction on the quantization data corresponding to the drug atom to obtain the first quantization feature.

23. The affinity prediction device according to claim 17, wherein The second atomic feature includes a second node feature and a second distance feature; The generation module is specifically configured to: Obtain the atomic association data corresponding to each receptor atom in the binding pocket, where the atomic association data includes at least one of atomic number, number of neighbors, formal charge, radical electrons, hybridization orbitals, aromatic ring correlation, number of connected hydrogen atoms, chiral center, chiral type, and molecular type; For each receptor atom in the binding pocket, perform feature extraction on the atomic association data corresponding to the receptor atom to obtain the second node feature; According to the target topological graph, obtain the distance data between every two receptor atoms in the binding pocket; Perform feature extraction on the distance data between every two receptor atoms in the binding pocket to obtain the second distance feature.

24. The affinity prediction device according to claim 23, characterized in that, The second atomic feature further includes a second edge feature; The obtaining module is further configured to obtain edge connection data between every two receptor atoms in the binding pocket according to the target topological graph, where the edge connection data includes at least one of a covalent bond type and a covalent bond positional relationship; The generating module is further configured to perform feature extraction on the edge connection data between every two receptor atoms in the binding pocket to obtain the second edge feature.

25. The affinity prediction device according to claim 23 or 24, characterized in that The second atomic feature further includes a second quantization feature; The obtaining module is further configured to, for each receptor atom in the binding pocket, obtain quantization data corresponding to the receptor atom according to the target topological graph, where the quantization data includes at least one of a covalent bond angle, an interaction angle, and a local charge. The covalent bond angle is an angle formed by the receptor atom as a vertex and the two closest atoms. The interaction angle is an angle formed by the receptor atom as a vertex, a third atom, and a fourth atom. The third atom is an atom that has an edge connection relationship with the receptor atom and is the closest one, and the fourth atom is a drug atom that is the closest to the receptor atom; The generating module is further configured to, for each receptor atom in the binding pocket, perform feature extraction on the quantization data corresponding to the receptor atom to obtain the second quantization feature.

26. The affinity prediction device according to claim 17, wherein The drug molecule includes L drug atoms, and the binding pocket includes P receptor atoms, where both L and P are integers greater than or equal to 1; The obtaining module is specifically configured to: Perform feature embedding on the first atomic feature of each drug atom to obtain L first input embedding features and H first bias embedding features, where H represents the number of attention heads, and H is an integer greater than or equal to 1; Perform feature embedding on the second atomic feature of each receptor atom to obtain P second input embedding features and H second bias embedding features; Based on the L first input embedding features and the H first bias embedding features, obtain drug molecule interaction features through the first attention network included in the internal attention network, where the internal attention network belongs to the affinity prediction model; Based on the P second input embedding features and the H second bias embedding features, obtain binding pocket interaction features through the second attention network included in the internal attention network; Concatenate the drug molecule interaction features and the binding pocket interaction features to obtain the molecular interaction features.

27. The affinity prediction device according to claim 26, characterized in that The obtaining module is specifically configured to: Based on the L first input embedding features, obtain L first query vectors, L first key vectors, and L first value vectors through the linear network layer included in the first attention network, where the first attention network belongs to the internal attention network; Perform matrix multiplication on the L first query vectors and the L first key vectors to obtain H first intermediate matrices; Based on the H first intermediate matrices and the H first bias embedding features, obtain H first attention matrices through the normalization exponential function layer included in the first attention network; Perform matrix multiplication on the H first attention matrices and the L first value vectors to obtain H first target matrices; Based on the H first target matrices, obtain the drug molecule interaction feature through the target neural network included in the first attention network; The obtaining module is specifically configured to: Based on the P second input embedding features, obtain P second query vectors, P second key vectors, and P second value vectors through the linear network layer included in the second attention network, where the second attention network belongs to the internal attention network; Perform matrix multiplication on the P second query vectors and the P second key vectors to obtain H second intermediate matrices; Based on the H second intermediate matrices and the H second bias embedding features, obtain H second attention matrices through the normalization exponential function layer included in the second attention network; Perform matrix multiplication on the H second attention matrices and the P second value vectors to obtain H second target matrices; Based on the H second target matrices, obtain the binding pocket interaction feature through the target neural network included in the second attention network.

28. The affinity prediction device according to claim 17, wherein The drug molecule includes L drug atoms, and the binding pocket includes P receptor atoms, where both L and P are integers greater than or equal to 1; The obtaining module is specifically configured to: Concatenate the molecule interaction feature and the target embedding feature of the virtual node to obtain N input embedding features, where the N input embedding features include the drug molecule interaction feature, the binding pocket interaction feature, and the target embedding feature, and N is equal to (L + P + 1); Based on the N input embedding features, obtain N target query vectors, N target key vectors, and N target value vectors through the linear network layer included in the interaction attention network, where the interaction attention network belongs to the affinity prediction model; Perform matrix multiplication on the N target query vectors and the N target key vectors to obtain H target intermediate matrices, where H represents the number of attention heads, and H is an integer greater than or equal to 1; Based on the H target intermediate matrices and H target bias embedding features, obtain H target attention matrices through the normalization exponential function layer included in the interaction attention network, where the H target bias embedding features are generated according to the first atomic feature of each drug atom, the second atomic feature of each receptor atom, and the atomic feature of the virtual node; Perform matrix multiplication on the H target attention matrices and the N target value vectors to obtain H target matrices; Based on the H target matrices, obtain the feature to be processed through the target neural network included in the interaction attention network; Obtain the target feature to be processed according to the feature to be processed; Based on the target feature to be processed, obtain the affinity for the protein-ligand conformation through the output layer.

29. The affinity prediction device according to claim 28, wherein The apparatus further includes: a processing module; The obtaining module is further configured to obtain a first weight parameter and a second weight parameter from the gamma matrix, where the gamma matrix is used to describe the feature relationship between the drug molecule and the binding pocket; The generating module is further configured to generate H to-be-processed bias embedding features according to the first atomic feature of each drug atom, the second atomic feature of each receptor atom, and the atomic feature of the virtual node; The processing module is configured to adjust the first numerical set in each to-be-processed bias embedding feature by using the first weight parameter, and adjust the second numerical set in each to-be-processed bias embedding feature by using the second weight parameter to obtain the H target bias embedding features.

30. The affinity prediction device according to claim 28, characterized in that, The apparatus further includes: a processing module; The processing module is configured to, after obtaining H target attention matrices through the normalization exponential function layer included in the interactive attention network based on the H target intermediate matrices and the H target bias embedding features, perform an averaging process on the values at the same positions in the H target attention matrices to obtain an average attention matrix; The obtaining module is further configured to obtain an association distribution matrix from the average attention matrix, where the association distribution matrix represents the association between the L drug atoms and the P receptor atoms.

31. The affinity prediction device according to claim 17, wherein The apparatus further includes: a training module; The obtaining module is further configured to obtain an original conformation sample, where the original conformation sample is constructed based on an original protein sample and an original drug molecule sample; The generating module is further configured to generate T training topology graph samples according to the original conformation sample, where each training topology graph sample includes a training drug molecule sample, and each training topology graph sample has a corresponding true affinity value, and T is an integer greater than or equal to 1; The obtaining module is further configured to obtain, based on the T training topology graph samples, an affinity prediction value corresponding to each training topology graph sample through an affinity prediction model; The training module is configured to update the model parameters in the affinity prediction model according to the affinity prediction value and the true affinity value corresponding to each training topology graph sample.

32. The affinity prediction device according to claim 31, wherein The apparatus further includes: a determination module; The determination module is configured to, for each training topology graph sample among the T training topology graph samples, determine the molecular structure error between the training drug molecule sample included in the training topology graph sample and the original drug molecule sample; The determination module is further configured to determine the label category corresponding to each training topology graph sample according to the molecular structure error corresponding to each training topology graph sample, where the label category corresponding to the training topology graph sample with a molecular structure error less than the error threshold is a positive sample label, and the label category corresponding to the training topology graph sample with a molecular structure error greater than or equal to the error threshold is a negative sample label; The obtaining module is further configured to obtain the predicted class probability corresponding to each training topology graph sample through the affinity prediction model based on the T training topology graph samples; The training module is specifically configured to: Determine a first loss value by using a first loss function according to the affinity prediction value and the true affinity value corresponding to each training topology graph sample; Determine a second loss value by using a second loss function according to the predicted class probability and the label class corresponding to each training topology graph sample; Update the model parameters in the affinity prediction model according to the first loss value and the second loss value.

33. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the affinity prediction method according to any one of claims 1 to 16 are implemented.

34. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the affinity prediction method according to any one of claims 1 to 16 are implemented.

35. A computer program product comprising a computer program, characterized in that, When the computer program is executed by the processor, the steps of the affinity prediction method according to any one of claims 1 to 16 are implemented.