Affinity prediction method, related apparatus and medium
By determining the molecular and atomic structure maps of biomolecules, extracting binding region maps, and fusing various map features into molecular interaction features, the problem of inaccurate predictions caused by reliance on sequence information in existing technologies is solved, achieving more accurate affinity prediction.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2026-01-06
- Publication Date
- 2026-07-30
AI Technical Summary
Existing technologies that use artificial intelligence to predict the binding affinity of biomolecules rely on the sequence information of the biomolecules, leading to inaccurate predictions.
By determining the molecular and atomic structure diagrams of biomolecules, capturing the binding region diagram, extracting various diagram features, and fusing them into molecular interaction features, affinity prediction is performed.
It improves the accuracy of predicting the binding affinity between biomolecules and provides comprehensive information on molecular representation and interactions.
Smart Images

Figure CN2026070735_30072026_PF_FP_ABST
Abstract
Description
Affinity prediction methods, related devices and media
[0001] This application claims priority to Chinese Patent Application No. 202510120735.6, filed on January 24, 2025, entitled “Affinity Prediction Method, Related Apparatus and Medium”, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This disclosure relates to the field of big data technology, and more specifically to affinity prediction. Background Technology
[0003] Currently, in various biomedical scenarios such as drug development and immunotherapy, understanding the interaction relationships and strengths between proteins and biomolecules such as peptides, antibodies, and antigens is crucial for drug design and target localization. For example, the interaction between proteins and peptides plays a key role in signal transduction and gene expression regulation. Research on the interaction relationships between proteins and biomolecules such as peptides, antibodies, and antigens is often closely related to the binding affinity of these biomolecules, as binding affinity can clearly and intuitively reflect the interactions between these biomolecules.
[0004] Although methods have emerged in related technologies to predict the binding affinity between proteins and biomolecules such as peptides, antibodies, and antigens using artificial intelligence, such as predicting binding affinity based on the sequence information of biomolecules through trained neural network models, this approach often relies heavily on the sequence information of biomolecules, which may lead to inaccurate predictions of the binding affinity between biomolecules. Summary of the Invention
[0005] This disclosure provides an affinity prediction method, related apparatus, and medium that can improve the accuracy of predicting the binding affinity between biomolecules.
[0006] According to one aspect of this disclosure, an affinity prediction method is provided, the method comprising:
[0007] Determine the first molecular structure diagram and atomic structure diagram corresponding to the first biomolecule, and the second molecular structure diagram corresponding to the second biomolecule;
[0008] A binding region diagram is extracted from the second molecular structure diagram, wherein the binding region diagram is used to indicate the residues in the second biomolecule that bind to the first biomolecule.
[0009] The first image features of the first molecular structure map, the second image features of the atomic structure map, the third image features of the second molecular structure map, and the fourth image features of the binding region map are extracted respectively.
[0010] Based on the features of the first graph, the features of the second graph, the features of the third graph, and the features of the fourth graph, molecular interaction features are determined.
[0011] Based on the molecular interaction characteristics, the affinity of the first biomolecule and the second biomolecule is predicted to obtain the binding affinity prediction results.
[0012] According to one aspect of this disclosure, an affinity prediction device is provided, the device comprising:
[0013] The first determining unit is used to determine the first molecular structure diagram and atomic structure diagram corresponding to the first biomolecule, and the second molecular structure diagram corresponding to the second biomolecule;
[0014] The cutting unit is used to cut out a binding region map from the second molecular structure diagram, wherein the binding region map is used to indicate the residues in the second biomolecule that bind to the first biomolecule.
[0015] The extraction unit is used to extract the first image features of the first molecular structure diagram, the second image features of the atomic structure diagram, the third image features of the second molecular structure diagram, and the fourth image features of the binding region diagram, respectively.
[0016] The second determining unit is used to determine molecular interaction features based on the features of the first graph, the features of the second graph, the features of the third graph, and the features of the fourth graph;
[0017] The prediction unit is used to predict the affinity of the first biomolecule and the second biomolecule based on the molecular interaction characteristics, and obtain the binding affinity prediction result.
[0018] According to one aspect of this disclosure, an electronic device is provided, including a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the affinity prediction method as described above.
[0019] According to one aspect of this disclosure, a computer-readable storage medium is provided, the storage medium storing a computer program that, when executed by a processor, implements the affinity prediction method as described above.
[0020] According to one aspect of this disclosure, a computer program product is provided, comprising a computer program that is read and executed by a processor of a computer device, causing the computer device to perform the affinity prediction method as described above.
[0021] In this embodiment, the first molecular structure diagram and atomic structure diagram corresponding to the first biomolecule can comprehensively represent the chemical and physical properties of the first biomolecule, as well as its overall structural information. Simultaneously, the first molecular structure diagram corresponding to the first biomolecule can also comprehensively characterize its overall structural information. The second biomolecule can be considered as a ligand of the first biomolecule, and the binding region diagram extracted from the second molecular structure diagram can clearly characterize the binding sites and interactions between the first and second biomolecules. Based on this, molecular structural features (features from the first and third diagrams), atomic features (features from the second diagram), and binding region features (features from the fourth diagram) are extracted from the determined multiple structure diagrams. These extracted features are then integrated to form a comprehensive molecular interaction feature, achieving multimodal feature extraction and fusion, which is beneficial for obtaining a more comprehensive understanding of the molecular representation and interaction information of the two biomolecules. Finally, affinity prediction based on the molecular interaction features can improve the accuracy of predicting the binding affinity between biomolecules.
[0022] Other features and advantages of this disclosure will be set forth in the following description and will be apparent in part from the description or may be learned by practicing the disclosure. The objectives and other advantages of this disclosure may be realized and obtained by means of the structures particularly pointed out in the description, claims and drawings. Attached Figure Description
[0023] Figure 1 is a system architecture diagram of an affinity prediction method applied according to an embodiment of the present disclosure;
[0024] Figure 2 illustrates a schematic diagram of the affinity prediction method according to an embodiment of the present disclosure applied in a molecular affinity prediction scenario;
[0025] Figure 3 is a flowchart of an affinity prediction method according to an embodiment of the present disclosure;
[0026] Figure 4 is a schematic diagram of the process of generating a first molecular structure diagram according to an embodiment of the present disclosure;
[0027] Figure 5 is a schematic diagram illustrating the process of extracting and combining region maps according to an embodiment of this disclosure;
[0028] Figure 6 is a flowchart of generating a first graph feature according to an embodiment of the present disclosure;
[0029] Figure 7 is a flowchart of the node features of a generated node according to an embodiment of this disclosure;
[0030] Figure 8 is a flowchart of generating first graph features according to another embodiment of this disclosure;
[0031] Figure 9 is a flowchart of generating third graph features according to an embodiment of this disclosure;
[0032] Figure 10 is a schematic diagram illustrating the process of generating third graph features according to an embodiment of this disclosure;
[0033] Figure 11 is a schematic diagram of the process of generating binding affinity prediction results according to an embodiment of the present disclosure;
[0034] Figure 12 is a flowchart of a generative affinity prediction model according to an embodiment of the present disclosure;
[0035] Figure 13 is a flowchart of determining affinity pseudo-labels according to an embodiment of the present disclosure;
[0036] Figure 14 is a schematic diagram illustrating the implementation process of a generative affinity prediction model according to an embodiment of the present disclosure;
[0037] Figure 15 is a flowchart of a pre-trained feature extraction network according to an embodiment of the present disclosure;
[0038] Figure 16 is a flowchart of generating first sample molecular map features and second sample molecular map features according to an embodiment of the present disclosure;
[0039] Figure 17 is a flowchart of training an original network according to an embodiment of the present disclosure;
[0040] Figure 18 is a schematic diagram illustrating the implementation process of a pre-trained feature extraction network according to an embodiment of the present disclosure;
[0041] Figure 19 is a comparison chart of experimental data based on different sample partitioning methods according to an embodiment of the present disclosure;
[0042] Figures 20A-20D are comparative diagrams of ablation experimental data according to an embodiment of the present disclosure;
[0043] Figure 21 is a block diagram of an affinity prediction device according to an embodiment of the present disclosure;
[0044] Figure 22 is a terminal structure diagram of an affinity prediction method according to an embodiment of the present disclosure;
[0045] Figure 23 is a server structure diagram of an affinity prediction method according to an embodiment of the present disclosure. Detailed Implementation
[0046] To make the objectives, technical solutions, and advantages of this disclosure clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and are not intended to limit the scope of this disclosure.
[0047] The system architecture and scenarios in which this disclosure is applied are described below.
[0048] Figure 1 is a system architecture diagram of the affinity prediction method applied according to an embodiment of the present disclosure. It includes a target terminal 140, an Internet 130, a gateway 120, a server 110, and a biomolecular database 150, etc.
[0049] The target terminal 140 can take various forms, including desktop computers, laptops, PDAs (personal digital assistants), tablets, mobile phones, in-vehicle terminals, home theater terminals, smart TVs, and dedicated terminals. Furthermore, it can be a single device or a collection of multiple devices. The target terminal 140 can communicate with the Internet 130 via wired or wireless means to exchange data. The target can specify biomolecular pairs for affinity prediction through the target terminal 140, where biomolecular pairs include, but are not limited to, protein-peptide pairs, antigen-antibody pairs, and protein-protein pairs.
[0050] Server 110 refers to a computer system capable of providing certain services to object terminal 140. Compared to ordinary object terminal 140, server 110 has higher requirements in terms of stability, security, and performance. Server 110 can be a high-performance computer in a network platform, a cluster of multiple high-performance computers, a portion of a high-performance computer (e.g., a virtual machine), a combination of portions of multiple high-performance computers (e.g., virtual machines), or a cloud server, etc. Server 110 includes various types of services, and the implementation of each service of server 110 is often associated with some intermediate databases or storage media. Server 110 is used to construct sample data based on existing biomolecules in biomolecular database 150 and to train an affinity prediction model using the sample data. At the same time, server 110 is also used to predict the binding affinity of two biomolecules in a specified biomolecular pair using the trained affinity prediction model.
[0051] The biomolecule database 150 is used to store various types of biomolecules. The biomolecule database 150 can be set up independently or integrated into the server 110 or other electronic devices.
[0052] Gateway 120, also known as an internetwork connector or protocol converter, is a computer system or device that acts as a translator, enabling network interconnection at the transport layer. It bridges the gap between two systems using different communication protocols, data formats, languages, or even completely different architectures. Gateways can also provide filtering and security functions. Messages sent from target terminal 140 to server 110 are forwarded to the corresponding server via gateway 120. Messages sent from server 110 to target terminal 140 are also forwarded to the corresponding target terminal 140 via gateway 120.
[0053] The embodiments disclosed herein can be applied to various scenarios, such as the molecular affinity prediction scenario shown in Figure 2.
[0054] The implementation environment for this molecular affinity prediction scenario may include a model training device 10 and a model usage device 20.
[0055] The model training device 10 can be an electronic device such as a mobile phone, desktop computer, tablet computer, laptop computer, vehicle terminal, server, intelligent robot, smart TV, multimedia playback device, or other electronic devices with strong computing power; this application does not limit this. The model training device 10 is used to train the affinity prediction model 30.
[0056] In this embodiment, the affinity prediction model 30 is a machine learning model. Optionally, the model training device 10 can use machine learning to train the affinity prediction model 30 to achieve better performance.
[0057] Optionally, the affinity prediction model 30 includes: a feature extraction network, an interaction feature extraction layer, and a binding affinity prediction layer. Its training process is as follows (this is only a brief description; the specific training process is described in the following embodiments and will not be elaborated here): The feature extraction network extracts features from the first molecular structure diagram and atomic structure diagram of the first biomolecule, obtaining graph features corresponding to the first molecular structure diagram and the atomic structure diagram; simultaneously, the feature extraction network extracts features from the second molecular structure diagram and binding region diagram of the second biomolecule, obtaining graph features corresponding to the second molecular structure diagram and the binding region diagram. The interaction feature extraction layer, based on the graph features of the first molecular structure diagram, the atomic structure diagram, the second molecular structure diagram, and the binding region diagram, obtains molecular interaction features, which characterize the interaction relationship between the first and second biomolecules. The binding affinity prediction layer, based on the molecular interaction features, obtains the binding affinity prediction result, which is used to determine the magnitude of the affinity between the first and second biomolecules. Based on the combined affinity prediction results, the parameters of the interaction feature extraction layer and the combined affinity prediction layer are adjusted to obtain the trained affinity prediction model.
[0058] In some embodiments, the model-using device 20 may be an electronic device such as a mobile phone, desktop computer, tablet computer, laptop computer, in-vehicle terminal, server, intelligent robot, smart TV, multimedia playback device, or other electronic devices with strong computing power; this application does not limit this. The trained affinity prediction model can perform the task of predicting the binding affinity of biomolecules. Exemplarily, the trained affinity prediction model can be used to perform the task of predicting the binding affinity of protein-peptide, or the task of predicting the binding affinity of protein-protein, or the task of predicting the binding affinity of antigen-antibody.
[0059] It should be noted that, in this context (i.e., the embodiment corresponding to Figure 2), the first and second biomolecules, distinct from the first and second biomolecules described below, constitute the sample data for model training. The first and second biomolecules described below, however, are the biomolecule data required for binding affinity prediction during model application.
[0060] The embodiments of this disclosure are described in general below.
[0061] According to one embodiment of this disclosure, an affinity prediction method is provided.
[0062] This affinity prediction method is generally applied in business scenarios that require the study of biomolecular interactions, such as immunoassay and drug development, for example, the molecular affinity prediction scenario shown in Figure 2. This disclosure provides a scheme for predicting binding affinity based on the graph structures of biomolecules at the amino acid scale (e.g., molecular structure diagrams), the graph structures at the atomic scale (e.g., atomic structure diagrams), and the binding region diagrams of biomolecules, which can improve the accuracy of predicting the binding affinity between biomolecules.
[0063] As shown in Figure 3, the affinity prediction method according to an embodiment of the present disclosure can be executed by an electronic device, which may be the server or object terminal shown in Figure 1. The affinity prediction method according to an embodiment of the present disclosure may include:
[0064] Step 310: Determine the first molecular structure diagram and atomic structure diagram corresponding to the first biomolecule, and the second molecular structure diagram corresponding to the second biomolecule;
[0065] Step 320: Extract the binding region diagram from the second molecular structure diagram;
[0066] Step 330: Extract the first image features of the first molecular structure diagram, the second image features of the atomic structure diagram, the third image features of the second molecular structure diagram, and the fourth image features of the binding region diagram, respectively;
[0067] Step 340: Determine molecular interaction features based on the features of the first, second, third, and fourth graphs;
[0068] Step 350: Based on molecular interaction characteristics, predict the affinity of the first biomolecule and the second biomolecule to obtain the binding affinity prediction results.
[0069] Steps 310-350 are described in detail below.
[0070] In step 310, the first molecular structure diagram and atomic structure diagram corresponding to the first biomolecule and the second molecular structure diagram corresponding to the second biomolecule are determined.
[0071] The first biomolecule is a long-chain molecule polymerized from amino acid monomers. This includes, but is not limited to, proteins, peptides, and antibodies. The second biomolecule can also be a long-chain molecule polymerized from amino acid monomers, or it can be other types of molecules. For example, the second biomolecule can be a protein, peptide, carbohydrate, or nucleic acid.
[0072] For example, when the second biomolecule is a protein, the first biomolecule is a polypeptide, and the binding affinity between the protein and the polypeptide is predicted; as another example, when the second biomolecule is an antigen, the first biomolecule is an antibody, and the binding affinity between the antigen and the antibody is predicted.
[0073] The first molecular structure diagram is used to indicate the results of structural diagram modeling of the first biomolecule at the amino acid scale.
[0074] Atomic structure diagrams are used to indicate the results of structural modeling of the first biomolecule at the atomic scale.
[0075] The second molecular structure diagram is used to indicate the results of structural modeling of the second biomolecule at the amino acid scale.
[0076] To save space, the specific process of determining the first molecular structure diagram, atomic structure diagram, and second molecular structure diagram in the embodiments of this disclosure will be described in detail below, and will not be repeated here.
[0077] In step 320, the binding region diagram is extracted from the second molecular structure diagram.
[0078] Binding region maps are used to indicate the residues in a second biomolecule that bind molecularly to a first biomolecule. A binding region map is a sub-map extracted from the second molecular structure map that contains only the residues in contact with the first biomolecule and the edges between these residues, indicating the local structural arrangement of the residues in the second biomolecule that are in contact with the first biomolecule.
[0079] To save space, the specific process of extracting the binding region map from the second molecular structure diagram in this embodiment will be described in detail below, and will not be repeated here.
[0080] In step 330, the first map features of the first molecular structure map, the second map features of the atomic structure map, the third map features of the second molecular structure map, and the fourth map features of the binding region map are extracted respectively.
[0081] The first graph features are used to learn higher-order representation information from the first molecular structure graph of the first biomolecule at the amino acid scale.
[0082] The second graph features are used to learn higher-order representation information from the atomic structure graph of the first biomolecule at the atomic scale.
[0083] The third graph features are used to learn higher-order representation information from the second molecular structure graph of the second biomolecule at the amino acid scale.
[0084] The fourth feature is used to learn higher-order representation information from the binding region map of the second biomolecule at the amino acid scale.
[0085] One possible implementation is to use a feature extraction network based on a pre-defined affinity prediction model to extract features from the first image of the first molecular structure, the second image of the atomic structure, the third image of the second molecular structure, and the fourth image of the binding region. Here, the affinity prediction model refers to a neural network model used to predict the binding affinity between biomolecules. Affinity prediction models typically take the three-dimensional structures of two biomolecules at the amino acid scale and the three-dimensional structures at the atomic scale as input, and the binding affinity between the two biomolecules as output.
[0086] Feature extraction networks refer to neural networks that extract three-dimensional structural features from two biomolecules. These networks can be based on graph convolutional neural networks, graph attention neural networks, or graph transformers (a type of Transformer model for graph-structured data). The four graph features mentioned above can indicate the features extracted by the feature extraction network. For example, the first graph feature indicates the high-order representation information learned by the feature extraction network from the first molecular structure diagram of the first biomolecule at the amino acid scale; the others are similar and will not be elaborated further here.
[0087] To save space, the specific process of extracting the features of the first image, the second image, the third image, and the fourth image in this embodiment will be described in detail below, and will not be repeated here.
[0088] In step 340, molecular interaction features are determined based on the features of the first graph, the second graph, the third graph, and the fourth graph.
[0089] Molecular interaction features are used to characterize the interaction relationship between the first and second biomolecules. Four types of graph features can reflect the characteristics of the first biomolecule, the characteristics of the second biomolecule, and the characteristics of the interaction between the first and second biomolecules. Therefore, based on these four graph features, molecular interaction features can be obtained.
[0090] As one possible implementation, this embodiment first concatenates the features of the first, second, third, and fourth images to obtain a feature concatenation result. Then, self-attention calculation is performed on the feature concatenation result to obtain molecular interaction features.
[0091] In the process of calculating self-attention on the feature concatenation results, the query features are obtained by linearly projecting the results through the query channel of the affinity prediction model, the key features are obtained by linearly projecting the results through the key channel of the affinity prediction model, and the value features are obtained by linearly projecting the results through the value channel of the affinity prediction model. Next, the query features and key features are dot-producted to obtain the dot product result, which is then normalized using the softmax function to obtain the attention weights. Finally, the attention weights and value features are multiplied to obtain the molecular interaction features.
[0092] In step 350, based on molecular interaction characteristics, the affinity of the first biomolecule and the second biomolecule is predicted to obtain the binding affinity prediction results.
[0093] The binding affinity prediction results are used to indicate the magnitude of the affinity between the first and second biomolecules. A larger predicted binding affinity indicates a more stable interaction between the first and second biomolecules.
[0094] In this specific implementation, the molecular interaction information contained in the molecular interaction features is captured by the affinity prediction model, and affinity is predicted based on the captured molecular interaction information, and the binding affinity prediction result is output.
[0095] Through steps 310-350 above, in this embodiment of the present disclosure, the first molecular structure diagram and atomic structure diagram corresponding to the first biomolecule can comprehensively represent the chemical and physical properties of the first biomolecule, as well as its overall structural information. Simultaneously, the first molecular structure diagram corresponding to the first biomolecule can also comprehensively characterize its overall structural information. The second biomolecule can be considered as a ligand of the first biomolecule, and the binding region diagram extracted from the second molecular structure diagram can clearly characterize the binding sites and interactions between the first and second biomolecules. Based on this, molecular structural features (features from the first and third diagrams), atomic features (features from the second diagram), and binding region features (features from the fourth diagram) are extracted from the determined multiple structure diagrams. These extracted features are then integrated to form comprehensive molecular interaction features, achieving multimodal feature extraction and fusion, which is beneficial for obtaining a more comprehensive understanding of the molecular representation and interaction information of the two biomolecules. Finally, affinity prediction based on molecular interaction features can improve the accuracy of predicting the binding affinity between biomolecules.
[0096] The above is a general description of steps 310-350. Since steps 340 and 350 have already been described in detail above, the specific implementations of steps 310, 320, and 330 will be described in detail below.
[0097] Step 310 will be described in detail below.
[0098] In step 310, the first molecular structure diagram and atomic structure diagram corresponding to the first biomolecule and the second molecular structure diagram corresponding to the second biomolecule are determined.
[0099] In one embodiment, the first molecular structure diagram corresponding to the first biomolecule is determined in the following manner:
[0100] To determine the residues contained in the first biomolecule and the contact relationships between the residues;
[0101] Each residue is identified as a node, and edges are constructed between nodes corresponding to residues with contact relationships to generate the first molecular structure diagram.
[0102] Here, residues refer to the portion remaining after each amino acid monomer in the first biomolecule forms a peptide bond. Amino acid monomers are linked together in the first biomolecule by peptide bonds. Residues are the basic structural and functional units of the first biomolecule. Different residues determine the three-dimensional structure and biological function of the first biomolecule through different interactions (such as hydrogen bonds, hydrophobic interactions, ionic bonds, etc.).
[0103] Contact relationship is used to indicate whether residues in a first biomolecule are linked together by peptide bonds. If two residues in a first biomolecule are linked by peptide bonds, then a contact relationship is considered to exist between these two residues.
[0104] Specifically, firstly, the first biomolecule is broken down into individual amino acids using a series of chemical and physical methods. These amino acids are then separated and quantitatively analyzed using chromatographic techniques to determine the types and relative amounts of each amino acid in the first biomolecule. Based on these types and relative amounts, the residues contained in the first biomolecule are identified. Next, nuclear magnetic resonance (NMR) technology is used to determine the interatomic distances, secondary structure, and overall conformation of the first biomolecule, thereby identifying peptide bonds. Based on the identified peptide bonds, the contact relationships between residues are determined. Further, each residue is identified as a node, and edges are constructed between nodes corresponding to residues with contact relationships, forming a graph structure. This graph structure serves as the first molecular structure diagram corresponding to the first biomolecule.
[0105] It should be noted that each node in the first molecular structure diagram has node features used to characterize the amino acid type, intrinsic disorder properties, etc., of the residues corresponding to the node. To save space, the specific method for determining the node features of each node in the first molecular structure diagram in this embodiment will be described in detail below. It will not be repeated here. Furthermore, since this embodiment only considers whether the residues in the first biomolecule are connected, and does not focus on the relative distance and orientation information between residues, it is not necessary to determine the edge features of the edges in the first molecular structure diagram.
[0106] As shown in Figure 4, the first biomolecule is a polypeptide with the sequence [TSAFEASDL]. A structural diagram of this polypeptide is modeled at the amino acid scale. The structural diagram of this first molecule has 9 nodes, each corresponding to one of the 9 amino acids in the sequence. Since adjacent amino acids in the polypeptide sequence are linked by peptide bonds, during structural diagram modeling, an edge is constructed between the node corresponding to amino acid T and the node corresponding to amino acid S, and between the node corresponding to amino acid S and the node corresponding to amino acid A, and so on, to form a complete structural diagram of the first molecule.
[0107] The advantage of this embodiment is that by modeling the structure of the first biomolecule at the amino acid scale, and applying the three-dimensional structural features of the first biomolecule reflected in the structure diagram of the first molecule to affinity prediction, the accuracy of affinity prediction can be improved.
[0108] In one embodiment, the atomic structure diagram corresponding to the first biomolecule is determined in the following way:
[0109] Convert the sequence of the first biomolecule into the target string;
[0110] Extract atoms from the target string as nodes, and extract the bonds between atoms in the target string as edges to obtain the atomic structure graph.
[0111] Here, the target string refers to the SMILES string, which is a symbolic representation of a first biomolecule used to provide a detailed description of its chemical species structure. The target string allows for the capture of atomic-level details of the first biomolecule, which are crucial for understanding its properties and potential interactions with second biomolecules.
[0112] Specifically, when performing atomic-scale modeling of the first biomolecule, the amino acid sequence of the first biomolecule is converted into the string "SMILES", thus obtaining the target string. This target string can be viewed as a graph of the interatomic interactions of the first biomolecule. Next, atoms are extracted from the target string as nodes, and the bonds between atoms in the target string are extracted as edges, resulting in an atomic structure graph, thereby converting the target string into a molecular graph.
[0113] Furthermore, for each node in the atomic structure diagram, a set of features is determined based on the atomic symbol of the atom corresponding to that node, the number of adjacent atoms, the number of adjacent hydrogen atoms, the implicit value of the atom, and whether the atom is in an aromatic structure. Each feature in this set corresponds to a property of the atom. For example, one feature in the set indicates the atomic symbol of the atom corresponding to that node. Further, the features of each node can be aggregated to form the node feature corresponding to that node.
[0114] The advantage of this embodiment is that the SMILES string (i.e., the target string) can describe complex molecular structures in a concise textual form. Converting the sequence of the first biomolecule into a SMILES string achieves a standardized representation of the molecular structure, making the molecular structure data easier for computers to process. Furthermore, the SMILES string not only contains the types and connections of atoms in the first biomolecule, but also represents ring structures, stereochemical information, etc. This allows the SMILES string obtained from the sequence of the first biomolecule to more comprehensively reflect the chemical properties of the first biomolecule. Further, constructing the target string into an atomic structure diagram visually displays the connections between atoms in the first biomolecule. This atomic structure diagram can also serve as input features for an affinity prediction model, enabling the model to effectively process the graph structure data of the atomic structure diagram during prediction. By learning the feature representations of atomic nodes and bond edges, it can accurately predict the binding affinity between the first and second biomolecules.
[0115] In one embodiment, the second molecular structure diagram corresponding to the second biomolecule is modeled based on the structure diagram of the second biomolecule at the amino acid scale. The specific process for determining the second molecular structure diagram corresponding to the second biomolecule is basically the same as the specific process for determining the first molecular structure diagram described above. The difference lies in that the node features of each node in the second molecular structure diagram are different from those of each node in the first molecular structure diagram. Specifically, the node features of each node in the second molecular structure diagram are used to characterize information such as the amino acid type and geometric structure of the residues corresponding to the nodes. The specific method for determining the node features of each node in the second molecular structure diagram in this embodiment will be described in detail below. Simultaneously, the second molecular structure diagram needs to consider the relative distance and orientation information between residues. Therefore, it is necessary to determine the edge features of each edge in the second molecular structure diagram based on the relative distance and orientation information between the residues corresponding to the nodes connected by the edges, so that the edge features represent the local relationships around the residues. For the sake of brevity, this will not be elaborated here.
[0116] Step 320 will be described in detail below.
[0117] In step 320, a binding region map is extracted from the second molecular structure diagram, wherein the binding region map is used to indicate the residues in the second biomolecule that bind to the first biomolecule.
[0118] In one embodiment, step 320 specifically includes, but is not limited to, the following steps:
[0119] Based on the binding sites of the first and second biomolecules, and the sequence of the second biomolecule, a mask sequence corresponding to the second biomolecule is generated.
[0120] Based on the mask sequence, the second molecular structure map is segmented to obtain the binding region map.
[0121] The binding site indicates the specific location where the first biomolecule binds to the second biomolecule, and the sequence of the second biomolecule indicates the specific types and order of the amino acids constituting the second biomolecule. The mask sequence of the second biomolecule uses 0 and 1 to identify the residues involved in the binding of the second biomolecule to the first biomolecule. Specifically, in the mask sequence of the second biomolecule, residues not involved in the binding of the second biomolecule to the first biomolecule are represented as 0, and residues involved in the binding site of the second biomolecule to the first biomolecule are represented as 1.
[0122] Specifically, firstly, the pocket locations of the second biomolecule are identified, where a pocket location refers to a specific region on or inside the surface of the first biomolecule that can accommodate the second biomolecule. The binding sites of the first and second biomolecules are determined based on the pocket locations. Next, based on the binding sites, residues involved in the binding of the second and first biomolecules are found in the sequence of the second biomolecule; these found residues are identified as pocket residues. Further, for the sequence of the second biomolecule, the values corresponding to the residues involved in the binding of the second and first biomolecules are replaced with 1, and the values corresponding to the residues not involved in the binding of the second and first biomolecules are replaced with 0, resulting in a mask sequence corresponding to the second biomolecule.
[0123] Furthermore, based on the positions with a value of 1 in the mask sequence, the nodes in the second molecular structure diagram are segmented, retaining nodes whose node indices match the indices corresponding to positions with a value of 1 in the mask sequence, as well as the edges connecting these nodes, to obtain a binding region diagram. Alternatively, based on pocket residues and the connections between pocket residues, the second molecular structure diagram is extracted to contain only pocket residues and the edges between them, and the extracted local graph structure is determined as the binding region diagram. The binding region diagram can also be called a pocket diagram. The binding region diagram is a subgraph of the second molecular structure diagram.
[0124] Since each node in the second molecular structure diagram has node features and each edge has edge features, for each node in the extracted binding region diagram, the node features and edge features corresponding to the pocket residues are retained so that these node features and edge features provide information about the local conformation, evolutionary characteristics and spatial relationships of the pocket residues.
[0125] As shown in Figure 5, for a second biomolecule with the amino acid sequence [LPGEGTQVHPRAPLLLK], which is a protein, the following steps are taken: First, the pocket positions of the second biomolecule are identified, locating the binding pockets (i.e., binding sites) for binding with the first biomolecule. Simultaneously, graph generation is performed on the second biomolecule to obtain its structural map. Next, based on the binding pockets, the amino acid sequence of the second biomolecule is masked, resulting in the mask sequence [00000000011111000]. Finally, based on the positions with a value of 1 in the mask sequence, the local structure corresponding to the node from amino acid P to the node corresponding to amino acid L is extracted from the second molecular structure map, thus obtaining the binding region map.
[0126] The advantage of this embodiment is that by generating a mask sequence using binding sites and sequence information, the region in the second biomolecule that binds to the first biomolecule can be accurately located. Simultaneously, the mask sequence filters out irrelevant sequence information, allowing for greater focus on the features of the binding region when segmenting the second molecule's structural map, thus improving segmentation accuracy. Furthermore, the extracted binding region map provides richer structural information, which is beneficial for enhancing the feature extraction network's learning of the feature representation of the binding region. This enables the feature extraction network to more accurately capture the geometric and topological features of the binding region, thereby improving the accuracy of subsequent feature extraction and affinity prediction.
[0127] Step 330 will be described in detail below.
[0128] In step 330, the first map features of the first molecular structure map, the second map features of the atomic structure map, the third map features of the second molecular structure map, and the fourth map features of the binding region map are extracted respectively.
[0129] The first molecular structure diagram in this embodiment is generated using residues in the first biomolecule as nodes and the relationships between residues as edges.
[0130] Referring to Figure 6, in one embodiment, the specific process of extracting the first map features of the first molecular structure map in step 330 may include, but is not limited to, the following steps 610-630:
[0131] Step 610: Determine the node characteristics of each node in the first molecular structure diagram;
[0132] Step 620: For each node, perform feature fusion on the node features of the node's neighboring nodes to obtain fused node features;
[0133] Step 630: Generate features based on the individual node features of multiple nodes and the fused node features to obtain the first graph features.
[0134] Steps 610-630 are described in detail below.
[0135] In step 610, node features are used to indicate the amino acid type, intrinsic disorder properties, and evolutionary characteristics of the residues corresponding to each node in the first molecular structure diagram.
[0136] To save space, the specific process of determining the node features of each node in the first molecular structure diagram according to the embodiments of this disclosure will be described in detail below. It will not be repeated here.
[0137] In step 620, a neighboring node refers to any node in the first molecular structure graph that is directly connected to the node by an edge, except for the node in question.
[0138] The merged node features are used to indicate the overall node feature information of all neighboring nodes of a node.
[0139] In the specific implementation of this embodiment, for each node in the first molecular structure diagram, the node features of the neighboring nodes corresponding to that node are spliced together to achieve feature fusion of the node features of the neighboring nodes and obtain fused node features.
[0140] In step 630, for each node in the first molecular structure graph, feature fusion is performed based on the node's features and the corresponding fused node features to obtain the first graph features. For example, cross-attention calculation is performed using the feature extraction network of the affinity prediction model to obtain the attention calculation results. Then, the feature extraction network generates feature representations based on the attention calculation results of multiple nodes to obtain the first graph features.
[0141] It should be noted that this feature extraction network includes, but is not limited to, graph convolutional neural networks, graph attention neural networks, and graph transformers.
[0142] For example, in this embodiment of the disclosure, a Graph Transformer is used to construct the feature extraction network. The feature extraction network can be constructed from one or more Graph Transformer structures.
[0143] The advantage of this embodiment is that by extracting the node features of each node in the first molecular structure graph and fusing the node features of the corresponding neighboring nodes, the inherent attributes and local structural information of each node in the first molecular structure graph can be captured, and the feature representation of each node can be enhanced, reducing information loss. Furthermore, feature generation based on the individual node features of multiple nodes and the fused node features can generate global graph features, improving the overall representation of the first molecular structure graph, enhancing the feature comprehensiveness and accuracy of the first graph features, and thus improving the accuracy of affinity prediction based on the first graph features.
[0144] Referring to Figure 7, in one embodiment, step 610 specifically includes, but is not limited to, the following steps 710-740:
[0145] Step 710: For each node, perform one-hot encoding on the residues corresponding to the node to obtain the sequence features of the node;
[0146] Step 720: Determine the intrinsic disorder characteristics of the nodes based on the intrinsic disorder fraction of the residues corresponding to the nodes;
[0147] Step 730: Embed the residues corresponding to the nodes to obtain the evolutionary features of the nodes;
[0148] Step 740: Concatenate the sequence features, inherent disorder features, and evolutionary features to obtain the node features.
[0149] Steps 710-740 are described in detail below.
[0150] In step 710, sequence features are used to indicate the amino acid types of each residue that constitutes the first biomolecule, as well as the sequence of each residue.
[0151] In this specific implementation, since proteins are composed of 20 different amino acid residues, the specific types of amino acid residues constituting the first and second biomolecules are a subset of these 20. Based on this, for each node, one-hot encoding can represent each residue as a unique vector. The generated vector indicates the amino acid type, and the sequence characteristics of the node are formed based on the one-hot encoded vectors corresponding to all nodes. For example, for a node corresponding to residue A, its one-hot encoded vector has only the position corresponding to A set to 1, while the remaining positions are 0.
[0152] In step 720, the intrinsic disorder fraction is used to quantify the degree of disorder of each residue contained in the first biomolecule, that is, whether the residue is in a state of random curling or lack of fixed three-dimensional structure in the structure of the first biomolecule.
[0153] The inherent disorder features are used to provide information about the dynamics and flexibility of the structure of the first biomolecule.
[0154] In a specific implementation of this embodiment, the determination of the intrinsic disorder fraction of the residues corresponding to the node includes the following steps:
[0155] The sequence of the first biomolecule is input into a preset online tool for disorder prediction, and a disorder tendency graph of each residue in the first biomolecule is obtained.
[0156] For the residues corresponding to the nodes, the intrinsic disorder fraction of the residues is determined based on the disorder tendency graph of the residues.
[0157] The online tool refers to the IUpred2A tool. Disorder tendency graphs are used to represent the inherent disordered nature of each residue in the first biomolecule.
[0158] Specifically, firstly, the sequence of the first biomolecule is input into an online tool, and the disorder prediction type is selected as long disorder. The online tool is used to predict the long disorder regions of the first biomolecule, obtaining a disorder tendency graph for each residue in the first biomolecule. Next, for the residues corresponding to the nodes, the disorder tendency graph is converted into a numerical form to obtain the intrinsic disorder score of the residues. The intrinsic disorder score is typically between 0 and 1; residues with an intrinsic disorder score greater than 0.5 are considered disordered, and residues with an intrinsic disorder score no greater than 0.5 are considered ordered.
[0159] Therefore, a disorder tendency graph is predicted using online tools. This graph allows for the rapid identification of functional regions such as binding sites to determine the interactions between biomolecules. Furthermore, the intrinsic disorder fraction is determined based on the disorder tendency graph, improving its accuracy and consequently enhancing the accuracy of node features. In step 730, evolutionary features are used to indicate the evolutionary properties of the first biomolecule.
[0160] The first biomolecule belongs to the first type. The first type is used to indicate the molecular type of the first biomolecule. The first type can be a polypeptide or a protein, etc.
[0161] In a specific implementation of this embodiment, step 730 may include, but is not limited to, the following steps:
[0162] The first extraction network, which is pre-trained, is used to embed the residues corresponding to the nodes to obtain the evolutionary features of the nodes.
[0163] The first extraction network is obtained by training the second extraction network based on the first type of biomolecular samples; the second extraction network is used to extract the evolutionary features of each node of the second molecular structure diagram corresponding to the second biomolecule.
[0164] The first extraction network was trained in the following way:
[0165] The probability distribution is obtained from the sequence of the biomolecular sample through the second extraction network. The sequence of the biomolecular sample includes the identifiers corresponding to each sequentially arranged component element in the biomolecular sample. In the sequence of the biomolecular sample, the identifiers corresponding to some components of the biomolecular sample are masked. The probability distribution is used to reflect the probability distribution of the components corresponding to the masked identifiers in the sequence of the biomolecular sample across the components of the first type of biomolecular sample.
[0166] Based on the probability distribution, the parameters of the second extraction network are adjusted to obtain the first extraction network.
[0167] The second feature extraction network outputs the constituent element corresponding to the highest probability in the probability distribution.
[0168] Specifically, when obtaining the probability distribution based on the sequences of biomolecular samples using the second extraction network, the sequences of biomolecular samples are first obtained from the transfer learning dataset. Next, the sequences of the first type of biomolecules are clustered according to similarity; for example, sequences with a similarity greater than 50% are grouped into one class. Furthermore, the sequences in the transfer learning dataset are drawn from multiple clusters, rather than being concentrated in a few. This ensures the diversity of sequence features in the transfer learning dataset.
[0169] For example, the sequence of a biomolecule sample can be represented in the form of "MA-DJ--LM", where the identifiers corresponding to three constituent elements are masked. Furthermore, if the first type of biomolecule is a polypeptide, then the identifiers corresponding to three amino acids in the sequence of the biomolecule sample are masked. Accordingly, the probability distribution can be used to reflect the probability distribution of these three amino acids across the constituent amino acids of the polypeptide, as predicted by the second feature extraction network.
[0170] In some embodiments, the second feature extraction network is derived based on the Transformer model. The Transformer model is a sequence-to-sequence model for natural language processing, based on an encoder-decoder structure, which learns the relationships between different positions in a sequence through a self-attention mechanism. In some embodiments, the probability distribution can be obtained based on Attention Maps generated by feature extraction from the sequence of biomolecular samples using the self-attention mechanism.
[0171] Furthermore, when adjusting the parameters of the second feature extraction network based on the probability distribution, a first probability can be obtained based on the probability distribution. The first probability refers to the probability predicted by the second feature extraction network that the constituent element corresponding to the masked identifier is one of the aforementioned constituent elements. If the biomolecule of the first type is a polypeptide, and the sample label of the biomolecule sample sequence indicates that the masked identifiers in the biomolecule sample sequence are A, Q, and J, then the first probability is the probability predicted by the second feature extraction network that the amino acid corresponding to A, Q, and J is the amino acid corresponding to the masked identifier in the biomolecule sample sequence. In some embodiments, the parameters of the second feature extraction network are adjusted based on the first probability. In some embodiments, the parameters of the second feature extraction network are adjusted to maximize the first probability, resulting in the first feature extraction network.
[0172] In some embodiments, the parameters of the second feature extraction network can be adjusted based on the cross-entropy loss of the probability distribution and the true probability distribution to obtain the first feature extraction network. In the true probability distribution, for each component element corresponding to the first type of biomolecule, the probability corresponding to the aforementioned component elements is 1, and the probability corresponding to the component elements other than the aforementioned component elements is 0. In some embodiments, the parameters of the second feature extraction network are adjusted with the aim of minimizing the aforementioned cross-entropy loss to obtain the first feature extraction network.
[0173] The above method partially masks the biomolecular sequence and then uses a second feature extraction network to predict the constituent elements corresponding to the masked portion of the sequence. This allows the trained first feature extraction network to learn the interaction relationships between the constituent elements of the first type of biomolecule after sequential arrangement, thus enabling more accurate feature extraction. Furthermore, by embedding the residues corresponding to nodes through the first extraction network, the conservation and variability of the residues during evolution can be captured, adding the evolutionary properties of the residues to the node features.
[0174] In step 740, the sequence features, intrinsic disorder features, and evolutionary features are spliced together to form a whole from the sequence information, intrinsic disorder information, and evolutionary property information of the node, thus obtaining the node features.
[0175] The advantages of this embodiment are that, for the first biomolecule, one-hot encoding of the residues corresponding to each node clearly distinguishes different residues in the first biomolecule while preserving the complete sequence information of the first biomolecule. Simultaneously, one-hot encoding transforms discrete categories into numerical vectors, making it easier for affinity prediction models to process this data and mine potential information in the sequence of the first biomolecule (such as interactions between residues, associations between specific sequence patterns and structure or function, etc.). Furthermore, considering that different degrees of intrinsic disorder in different biomolecules lead to differences in their function and structure, incorporating intrinsic disorder features into node features provides a more comprehensive description of the properties of the first biomolecule. Furthermore, embedding the residues corresponding to the nodes captures the conservation and variability of the residues during evolution, adding the evolutionary properties of the residues to the node features. Finally, concatenating sequence features, intrinsic disorder features, and evolutionary features fuses information from different dimensions to form a comprehensive node feature representation. This integration of multi-source information makes the node features of each node richer and more comprehensive, enabling the description of the properties of the first biomolecule from multiple perspectives. This provides more sufficient input information for the affinity prediction model, thereby improving the accuracy of affinity prediction.
[0176] Referring to Figure 8, in one embodiment, step 630 specifically includes, but is not limited to, the following steps 810-840:
[0177] Step 810: For each node, perform cross-attention calculation based on node features and fused node features to obtain node interaction features;
[0178] Step 820: Perform feature weighting on the node interaction features and node features to obtain the node representation vector;
[0179] Step 830: Concatenate the node representation vectors of multiple nodes to obtain the molecular graph representation vector of the first molecular structure graph;
[0180] Step 840: Perform global average pooling on the molecular graph representation vector to obtain the first graph features.
[0181] Steps 810-840 are described in detail below.
[0182] In step 810, node interaction features are used to indicate the interdependence between each node and its neighboring nodes.
[0183] As one possible implementation, for each node, based on node features and fused node features, cross-attention calculation can be performed through a feature extraction network to obtain node interaction features. In this specific implementation, for each node, firstly, the node features are linearly projected through the query channel of the feature extraction network to obtain a query vector. Then, the fused node features are linearly projected through the key and value channels of the feature extraction network to obtain a key vector projected by the key channel and a value vector projected by the value channel, respectively. Next, attention calculation is performed based on the key vector, value vector, and query vector to obtain the node interaction features.
[0184] In step 820, the node representation vector is used to indicate the key feature information of the node.
[0185] In the specific implementation of this embodiment, firstly, a first weight and a second weight are determined. The first weight is used to indicate the importance of the feature information of the node interaction features; the second weight is used to indicate the importance of the feature information of the node features. The sum of the first weight and the second weight is 1. Next, the product of the first weight and the node interaction features, and the product of the second weight and the node features are added together to achieve feature weighting of the node interaction features and node features, thereby obtaining the node representation vector of the node.
[0186] In step 830, the molecular graph representation vector is used to indicate global information about the entire first molecular structure graph.
[0187] In the specific implementation of this embodiment, the node representation vectors of each node are sequentially spliced together according to the order of the residues corresponding to each node in the sequence of the first biomolecule to obtain the molecular graph representation vector of the first molecular structure diagram.
[0188] In step 840, global average pooling is performed on the molecular graph representation vector to sum the spatial information in the molecular graph representation vector and obtain the first graph feature.
[0189] The advantages of this embodiment are that the cross-attention mechanism can learn the interrelationships and dependencies between nodes, thereby enhancing the feature representation of each node. The cross-attention mechanism can capture long-distance dependencies between nodes, leading to a more comprehensive understanding of the node's contextual information; simultaneously, cross-attention computation can improve processing performance on graph-structured data. Furthermore, feature weighting of node interaction features and node features can highlight important features and suppress unimportant features, which helps to more accurately capture the key information of nodes. Concatenating the node representation vectors of multiple nodes and performing global mean pooling on the concatenated molecular graph representation vector can integrate the global information of the entire first molecular structure graph while retaining its overall features and spatial information, thereby improving the feature comprehensiveness and accuracy of the first graph features.
[0190] In one embodiment, when extracting features from the atomic structure graph using a feature extraction network, since each node in the atomic structure graph also contains a set of features as described above, the set of features corresponding to each node can be concatenated into a whole, and the concatenated result can be used as the node feature. Then, for the atomic structure graph, based on the node features of each node, operations similar to steps 620-630 above are performed to obtain the second graph features corresponding to the atomic structure graph. For the sake of brevity, this will not be elaborated further here.
[0191] The second molecular structure diagram in this embodiment is generated using residues in the second biomolecule as nodes and the relationships between residues as edges.
[0192] Referring to Figure 9, in one embodiment, the specific process of extracting the third map features of the second molecular structure map in step 330 may include, but is not limited to, the following steps 910-930:
[0193] Step 910: For each node in the second molecular structure diagram, determine the sequence characteristics, evolutionary characteristics, and geometric structure characteristics of the residues corresponding to the node, and determine the node characteristics based on the splicing results of the sequence characteristics, evolutionary characteristics, and geometric structure characteristics.
[0194] Step 920: For each node, perform feature fusion on the node features of the node's neighboring nodes to obtain fused node features;
[0195] Step 930: Generate features based on the individual node features of multiple nodes and the fused node features to obtain the features of the third graph.
[0196] Steps 910-930 are described in detail below.
[0197] In step 910, sequence features are used to indicate the amino acid types of each residue constituting the second biomolecule, as well as the order in which these residues are arranged. Evolutionary features are used to indicate the evolutionary properties of the second biomolecule. Geometric features are used to indicate the atomic-level geometric characteristics of the second biomolecule.
[0198] In the specific implementation of this embodiment, the process of determining the sequence features for each node in the second molecular structure diagram is similar to step 710 described above. To save space, it will not be repeated here.
[0199] When determining the evolutionary characteristics of each node, the residues corresponding to the node are embedded based on the first feature extraction network to generate the embedded representation of the residues corresponding to the node, and this embedded representation is determined as the evolutionary characteristic of the node. The first feature extraction network is a protein language model trained on multiple protein sequences, and this protein language model can be a neural network model with an ESM-2 structure.
[0200] To determine the geometric structural features of each node, firstly, for each node, the atomic coordinates of the main chain atoms related to the residues corresponding to that node in the second biomolecule are determined according to a preset three-dimensional coordinate system. Next, based on the atomic coordinates of the main chain atoms, dihedral angles are calculated through mathematical operations such as vector cross products and dot products. These dihedral angles include, but are not limited to, the Phi (φ) angle, Psi (ψ) angle, and Omega (ω) angle. These three dihedral angles describe the torsion angles between peptide bond planes. Further, for each dihedral angle, the cosine and sine values are calculated and encoded into a feature vector. Finally, the feature vectors corresponding to each of the three dihedral angles are concatenated to obtain the geometric structural features of the node.
[0201] Furthermore, for each node in the second molecular structure diagram, sequence features, evolutionary features, and geometric structure features are spliced together to obtain the splicing result, and the splicing result is determined as the node feature of that node.
[0202] In step 920, the specific process of feature fusion of the node features of the node's neighboring nodes is similar to that in step 620 above. To save space, it will not be described in detail again.
[0203] In step 930, the specific process of generating features based on the individual node features of multiple nodes and the fused node features to obtain the features of the third graph is similar to steps 810-840 above. To save space, it will not be elaborated here.
[0204] Figure 10 illustrates the specific implementation of feature extraction using a feature extraction network for the second molecular structure graph. Specifically, the feature extraction network is constructed from a Graph Transformer structure, including a graph attention computation layer and a global average pooling layer. The graph attention computation layer comprises a linear layer, a scaled dot product attention layer, a feature mixing layer, and a concatenation layer. First, the second molecular structure graph is input into the feature extraction network. The graph attention computation layer sequentially uses each node of the second molecular structure graph as a center node. The fused node features, consisting of the node features of the determined center node and the node features of its neighboring nodes, are input into the linear layer for projection transformation. A query vector Q is obtained based on the node features of the center node, and a key vector K and a value vector V are obtained based on the fused node features. Further, the query vector Q, key vector K, and value vector V are input into the scaled dot product attention layer for cross-attention computation to obtain node interaction features. Finally, the feature mixing layer weights the node interaction features and node features to obtain the node representation vector. Next, a concatenation layer is used to concatenate the node representation vectors of multiple nodes to obtain the molecular graph representation vector of the first molecular structure graph. Finally, a global average pooling layer is used to perform global average pooling on the molecular graph representation vector to obtain the first graph features. It should be noted that when performing cross-attention calculation in the scaling dot product attention layer, a dot product operation is first performed on the query vector Q and the key vector V, and a score is calculated for the dot product operation to obtain the attention score. Then, the attention score is normalized into attention weights using the softmax function, where the attention weights are values between 0 and 1. Finally, a dot product operation is performed on the attention weights and the value vector V to obtain the node interaction features.
[0205] The advantage of this embodiment is that, for each node of the second molecular structure diagram, the evolutionary features, geometric features, and sequence features of the residues corresponding to the node are combined to form node features. This allows the node features to comprehensively characterize the amino acid type, atomic-level geometric characteristics, and evolutionary properties of the residues, improving the comprehensiveness and accuracy of the node features. Furthermore, by utilizing a feature extraction network to generate features based on the node features of each node and the fused node features, a third graph feature that comprehensively characterizes the global information of the second molecular structure can be obtained. This third graph feature is then applied to subsequent affinity prediction, allowing the introduction of sequence and three-dimensional structural information of the biomolecule into affinity prediction, thereby improving the accuracy of affinity prediction.
[0206] In one embodiment, since the fusion region map is extracted from the second molecular structure map, this extraction operation does not change the node features of each node in the second molecular structure map, nor the connection relationships between nodes. Therefore, when using a feature extraction network to extract features from the fusion region map, the process of determining the node features of each node in the fusion region map is consistent with step 910 above, and the method of determining the fusion node features for each node in the fusion region map is also consistent with step 920 above; the specific method of generating feature representations based on the node features and fusion node features of each node in the fusion region map to obtain the features of the fourth map is also similar to step 930 above. For the sake of brevity, further details are omitted.
[0207] In steps 340-350, molecular interaction features are determined based on the features of the first, second, third, and fourth graphs; based on the molecular interaction features, the affinity of the first and second biomolecules is predicted to obtain the binding affinity prediction results.
[0208] As described above, the feature extraction network in this embodiment of the disclosure can be a neural network based on graph convolutional neural networks, graph attention neural networks, graph transformers, etc. When the feature extraction network is based on a graph transformer structure, one or more graph transformer structures can be used.
[0209] However, when a feature extraction network uses a Graph Transformer structure, it needs to extract features from the first molecular structure graph, the second molecular structure graph, the atomic structure graph, and the binding region graph. In this case, the feature extraction network tends to be generalized and cannot specifically extract features for various graph structures, leading to low feature extraction accuracy. Therefore, this disclosure provides a scheme for constructing a feature extraction network based on multiple Graph Transformer structures, enabling the use of corresponding Graph Transformer structures for feature extraction across various graph structures, thereby improving feature extraction accuracy.
[0210] In this embodiment of the disclosure, the feature extraction network includes a first extraction module for extracting features from the second molecular structure diagram, a second extraction module for extracting features from the binding region diagram, a third extraction module for extracting features from the first molecular structure diagram, and a fourth extraction module for extracting features from the atomic structure diagram. The first extraction module, the second extraction module, the third extraction module, and the fourth extraction module are all Graph Transformer structures (the specific structure can be referred to Figure 10 described above).
[0211] As shown in Figure 11, the affinity prediction model of this embodiment includes a self-attention layer, a regression prediction layer, and a feature extraction network composed of a first extraction module, a second extraction module, a third extraction module, and a fourth extraction module. The self-attention layer corresponds to the interactive feature extraction layer in Figure 2, and the regression prediction layer corresponds to the binding affinity prediction layer in Figure 2. Specifically, the first biomolecule is a peptide residue sequence, and the second biomolecule is a protein structure. First, the second molecular structure diagram corresponding to the protein structure is input to the first extraction module (Protein Graph Transformer) for feature extraction. The binding region diagram corresponding to the binding pocket extracted from the protein structure is input to the second extraction module (Pocket Graph Transformer) for feature extraction. The first molecular structure diagram corresponding to the peptide residue sequence is input to the third extraction module (Peptide Graph Transformer) for feature extraction. Finally, the atomic structure diagram corresponding to the atomic structure extracted from the peptide residue sequence is input to the fourth extraction module (SMILES Graph Transformer) for feature extraction. The specific process is similar to steps 810-840 described above. Next, the features from the first image output by the third extraction module, the third image output by the first extraction module, the fourth image output by the second extraction module, and the second image output by the fourth extraction module are concatenated to obtain the feature concatenation result. Then, a self-attention layer is used to perform self-attention calculation on the feature concatenation result to obtain the molecular interaction features. Finally, based on the molecular interaction features, an affinity prediction layer is used to predict the binding affinity of the first and second biomolecules to obtain the binding affinity prediction result. The specific process is similar to steps 340-350 above. For the sake of brevity, it will not be elaborated further.
[0212] The process of generating an affinity prediction model according to an embodiment of this disclosure will be described in detail below.
[0213] Because the amount of labeled data available for adjusting the pre-trained model during actual training is relatively limited, this disclosure provides a scheme for adjusting the pre-trained model based on a self-distillation strategy to obtain an affinity prediction model, which can improve the prediction performance and generalization ability of the trained affinity prediction model.
[0214] In this embodiment of the disclosure, the affinity prediction model is obtained by adjusting a pre-trained model.
[0215] Among them, the pre-trained model refers to a neural network model that includes the aforementioned feature extraction network, self-attention network, and regression prediction network, and is capable of predicting the binding affinity of biomolecules. However, the affinity prediction performance of the pre-trained model does not meet the expected requirements.
[0216] Referring to Figure 12, in one embodiment, the specific process of adjusting the pre-trained model into an affinity prediction model may include, but is not limited to, the following steps 1210-1250:
[0217] Step 1210: Obtain multiple first sample pairs and multiple second sample pairs;
[0218] Step 1220: Input the first sample pair and the second sample pair into the pre-trained model to predict affinity, and obtain the first sample prediction result of the first sample pair and the second sample prediction result of the second sample pair;
[0219] Step 1230: Determine the first loss function based on the affinity reference labels of multiple first samples and the prediction results of the first samples;
[0220] Step 1240: Determine the second loss function based on the affinity pseudo-labels and prediction results of multiple second samples for their respective second samples;
[0221] Step 1250: Based on the first loss function and the second loss function, adjust the parameters of the pre-trained model to obtain the affinity prediction model.
[0222] Steps 1210-1250 are described in detail below.
[0223] In step 1210, the first sample pair includes a first sample biomolecule, a second sample biomolecule, and an affinity reference tag between the first sample biomolecule and the second sample biomolecule, and each second sample pair includes a third sample biomolecule, a fourth sample biomolecule, and an affinity pseudo-tag between the third sample biomolecule and the fourth sample biomolecule.
[0224] For example, the first and second sample pairs refer to protein-peptide pairs used for model training. The first and third sample biomolecules refer to proteins; the second and fourth sample biomolecules refer to peptides.
[0225] Affinity reference labels are used to indicate the actual binding affinity between the first and second sample biomolecules.
[0226] Affinity pseudo-labels are used to indicate labels determined by predictions of second-sample biomolecules based on a model fine-tuned from a pre-trained model.
[0227] In this specific implementation, since the preset biomolecular database often stores various types of biomolecules and the interactions between them, with authorization, proteins and peptides with recorded true binding affinity are extracted from the preset biomolecular database to form a first sample pair, and the true binding affinity between the proteins and peptides recorded in the biomolecular database is used as an affinity reference label. Similarly, proteins and peptides without recorded true binding affinity are extracted from the preset biomolecular database to form a second sample pair. The biomolecular database used to extract the first and second sample pairs can be the PCSB PDB database.
[0228] To save space, the specific method for determining the affinity pseudo-label of the second sample pair in this embodiment will be described in detail below. It will not be repeated here.
[0229] In step 1220, the prediction result of the first sample is used to indicate the binding affinity predicted by the pre-trained model for the first and second sample biomolecules. The prediction result of the second sample is used to indicate the binding affinity predicted by the pre-trained model for the third and fourth sample biomolecules.
[0230] In the specific implementation of this embodiment, the process of step 1220 is similar to that of steps 310-350 described above. To save space, it will not be described again.
[0231] In step 1230, the first loss function is used to indicate the overall degree of difference between the predicted binding affinity and the actual binding affinity of the pre-trained model for all first sample pairs.
[0232] In the specific implementation of this embodiment, since both the first sample prediction result and the affinity reference label are numerical data, the sample pair error can be obtained by calculating the mean squared error (MSE) or root mean square error (RMSE) of the first sample prediction result and affinity reference label for each first sample pair. The sample pair error of all first sample pairs is then averaged to obtain the first loss function.
[0233] In step 1240, the second loss function is used to indicate the overall degree of difference between the predicted binding affinity and the actual binding affinity of the pre-trained model for all second sample pairs.
[0234] In this specific implementation, step 1240 is similar to step 1230 described above. The difference lies in that step 1230 determines the overall prediction accuracy of the pre-trained model for all first sample pairs, while step 1240 determines the overall prediction accuracy of the pre-trained model for all second sample pairs. For the sake of brevity, these details will not be elaborated further.
[0235] In step 1250, the first loss function and the second loss function can be weighted and calculated to obtain the total loss function. Then, the parameters of the pre-trained model are adjusted based on the total loss function to obtain the affinity prediction model. The specific implementation process of weighting the first and second loss functions is similar to the feature weighting process of node interaction features and node features in step 820 above. For the sake of brevity, it will not be elaborated further.
[0236] When adjusting the parameters of the pre-trained model based on the total loss function, the model that minimizes the total loss function is selected as the training model. The parameters of the pre-trained model are continuously adjusted, and iterative training is performed according to steps 1210-1250 above. The parameters of the pre-trained model that minimizes the total loss function are taken as the final parameters, and the pre-trained model with the final parameters is taken as the trained affinity prediction model.
[0237] The advantage of this embodiment is that by combining labeled and unlabeled data, all available data resources can be fully utilized. Especially when labeled data is scarce, the pseudo-label generation method effectively utilizes unlabeled data, significantly increasing the amount of sample data for adjusting the pre-trained model. Furthermore, inputting the first and second sample pairs into the pre-trained model for affinity prediction allows the prediction results of the first and second samples to reflect the current performance level of the pre-trained model. Simultaneously, employing supervised and semi-supervised learning methods to construct the first and second loss functions enables comprehensive optimization of the pre-trained model, reducing overfitting on a single dataset and improving the model's robustness and generalization ability. By simultaneously considering labeled and unlabeled data, the model can learn more comprehensive feature representations, thereby improving the prediction accuracy of the trained affinity prediction model.
[0238] Referring to Figure 13, in one embodiment, the specific process for determining the affinity pseudo-label of the second sample pair includes, but is not limited to, the following steps 1310-1340:
[0239] Step 1310: Group the multiple first sample pairs to determine multiple benchmark sample pair groups;
[0240] Step 1320: Based on multiple benchmark sample pairs, train pre-trained models respectively to obtain intermediate models corresponding to each of the multiple benchmark sample pairs;
[0241] Step 1330: For each second sample pair, input the second sample pair into each intermediate model to predict affinity, and obtain multiple affinity prediction results;
[0242] Step 1340: Based on multiple affinity prediction results, determine the affinity pseudo-label of the second sample pair.
[0243] Steps 1310-1340 are described in detail below.
[0244] In step 1310, the reference sample pair group is used to indicate multiple first sample pairs that are grouped in the same group.
[0245] Among them, the multiple first sample pairs contained in each benchmark sample pair group are not completely identical.
[0246] In this specific implementation, firstly, according to a preset number of iterations, multiple first sample pairs are initially grouped to obtain several initial sample pair groups, each with a different first sample pair. Next, based on the number of iterations, in each iteration, one initial sample pair group is used as a validation dataset. All initial sample pair groups other than the one used as the validation dataset are integrated together to form a benchmark sample pair group. Based on this, a corresponding benchmark sample pair group is formed for each iteration round.
[0247] For example, multiple first sample pairs can be considered as a dataset. This dataset can be divided into five distinct subsets. In each iteration, four subsets are used as training data, while the remaining subset is used as validation data. This process is repeated five times, selecting a different subset as the validation set each time, thus ensuring that each subset has a chance to serve as the test set. Based on this, five benchmark sample pair groups can be obtained.
[0248] In step 1320, the intermediate model refers to a model obtained by adjusting the parameters of the pre-trained model based on the prediction results of the pre-trained model for the benchmark sample pair group.
[0249] The total number of intermediate models is equal to the number of multiple benchmark sample pairs.
[0250] In the specific implementation of this embodiment, pre-trained models are trained based on multiple benchmark sample pairs to obtain intermediate models corresponding to each benchmark sample pair, including but not limited to the following steps:
[0251] For each benchmark sample pair group, the affinity prediction of each first sample pair in the benchmark sample pair group is performed based on the pre-trained model to obtain the predicted binding affinity. The specific process is basically the same as the first sample pair input into the pre-trained model for affinity prediction in step 1220 above to obtain the first sample prediction result of the first sample pair.
[0252] Based on the affinity reference labels and predicted binding affinity of multiple first samples from the baseline sample pair group, the parameters of the pre-trained model are adjusted to obtain an intermediate model. The specific process is basically the same as steps 1, 2, and 30 above. To save space, it will not be described in detail.
[0253] In step 1330, the affinity prediction result is used to indicate the magnitude of the binding affinity predicted by the intermediate model for the second sample pair.
[0254] In the specific implementation of this embodiment, step 1330 is similar to steps 310-350 described above. To save space, it will not be described again.
[0255] In step 1340, for each second sample pair, the average of the multiple affinity prediction results for the second sample pair is calculated, and the average value of the affinity prediction results is used as the affinity pseudo-label for the second sample pair.
[0256] The advantage of this embodiment is that grouping multiple first sample pairs ensures diversity and representativeness in each benchmark sample pair group, which improves the generalization ability of the pre-trained model on different datasets and allows it to better adapt to different data distributions, reducing overfitting. Furthermore, by training the pre-trained model on different benchmark sample pairs, multiple intermediate models with different parameters and feature representations can be obtained, enabling parallel training of the pre-trained model and improving training efficiency. By inputting the second sample pair into multiple intermediate models, affinity prediction can be performed from multiple different perspectives. Each intermediate model may capture different features and patterns, thus providing more comprehensive affinity prediction results. Integrating the prediction results of multiple intermediate models can generate more accurate pseudo-labels. In addition, constructing pseudo-labels for the second sample pair addresses the training of affinity prediction models with limited labeled data, making full use of a large amount of unlabeled data and improving model performance. This approach combines grouping, multi-model training, multi-perspective prediction, and pseudo-label generation, improving the performance and robustness of the affinity prediction model in affinity prediction tasks.
[0257] Figure 14 illustrates the specific steps of adjusting a pre-trained model into an affinity prediction model. Specifically, first, data is collected from a biomolecular database to obtain labeled complexes, which can be protein-peptide pairs with binding affinity tags. Next, the pre-trained model is fine-tuned using the labeled complexes to obtain multiple intermediate models, similar to steps 1310-1320 above. Further, the multiple intermediate models obtained from the fine-tuned pre-trained model are used to predict the affinity of unlabeled complexes extracted from the biomolecular database, resulting in pseudo-labels for the unlabeled complexes, similar to steps 1330-1340 above. Here, the unlabeled complexes can be unlabeled protein-peptide pairs. Further, a pseudo-dataset is constructed based on the unlabeled complexes and their corresponding pseudo-labels, and a labeled dataset is constructed based on the labeled complexes. The pre-trained model is then fine-tuned using both the labeled and pseudo-datasets to obtain the affinity prediction model, similar to steps 1220-1250 above. For brevity, these steps will not be elaborated further.
[0258] The process of training a feature extraction network for a pre-trained model according to an embodiment of this disclosure will be described in detail below.
[0259] Referring to Figure 15, in one embodiment, the specific process of training the feature extraction network of the pre-trained model may include, but is not limited to, the following steps 1510-1560:
[0260] Step 1510: Obtain multiple third sample pairs;
[0261] Step 1520: Determine the reference label for the binding free energy of the third and fourth sample biomolecules in each second sample pair;
[0262] Step 1530: Input multiple second sample pairs and third sample pairs into the original network to generate feature representations, and obtain the first sample molecular map features of each of the multiple second sample pairs and the second sample molecular map features of each of the multiple third sample pairs.
[0263] Step 1540: For each second sample pair, perform free energy prediction based on the molecular graph features of the first sample to obtain the first prediction result;
[0264] Step 1550: For each third sample pair, perform affinity prediction based on the molecular map features of the second sample to obtain the second prediction result;
[0265] Step 1560: Based on the first prediction results of multiple second samples and the combined free energy reference label, and the second prediction results of multiple third samples and the affinity reference label, train the original network to obtain the feature extraction network.
[0266] Steps 1510-1560 are described in detail below.
[0267] In step 1510, each third sample pair includes a fifth sample biomolecule, a sixth sample biomolecule, and an affinity reference label for the fifth and sixth sample biomolecules, wherein the fifth and sixth sample biomolecules are of different molecular types.
[0268] For example, the third sample pair refers to the protein-compound pair used for model training, the fifth sample biomolecule is a protein, and the sixth sample biomolecule is a compound. Affinity reference labels are used to indicate the true binding affinity between the fifth and sixth sample biomolecules.
[0269] In this specific implementation, since pre-defined biomolecular databases often store various types of biomolecules and their interactions, proteins and compounds are extracted from these databases to form a third sample pair, with authorization. The actual binding affinity between the proteins and compounds recorded in the biomolecular database is used as an affinity reference label. The biomolecular database used to extract the third sample pair can be the ChEMBL+CrossDocked database.
[0270] In step 1520, a binding free energy reference label is used to indicate the true binding free energy between the third sample biomolecule and the fourth sample biomolecule.
[0271] In this specific implementation, firstly, with authorization, a computational tool is invoked to calculate the free energies of the third and fourth sample biomolecules, thereby obtaining their binding free energies. The calculated binding free energies are then designated as binding free energy reference labels. The computational tool includes, but is not limited to, software used for molecular dynamics simulations such as GROMACS or APBS.
[0272] In step 1530, the original network refers to a neural network structure that has not yet been trained and has the ability to extract features from the input graph structure. The original network includes, but is not limited to, graph convolutional neural networks, graph attention neural networks, and graph transformers. The feature extraction performance of the feature extraction network is superior to that of the original network.
[0273] The first sample molecular graph features are used to indicate the combined feature information of the graph structure of the two sample biomolecules in the second sample pair. The second sample molecular graph features are used to indicate the combined feature information of the graph structure of the two sample biomolecules in the third sample pair.
[0274] To save space, the specific process of generating feature representations for the second and third sample pairs in this embodiment will be described in detail below, and will not be repeated here.
[0275] In step 1540, the first prediction result is used to indicate the predicted binding free energy of the third sample biomolecule and the fourth sample biomolecule in the second sample pair.
[0276] In the specific implementation of this embodiment, the process of step 1540 is similar to that of step 350 described above. To save space, it will not be described again.
[0277] In step 1550, the second prediction result is used to indicate the predicted binding affinity between the fifth sample biomolecule and the sixth sample biomolecule of the third sample pair.
[0278] In the specific implementation of this embodiment, the process of step 1550 is similar to that of step 350 described above. To save space, it will not be described again.
[0279] In step 1560, the specific process of training the original network based on the first prediction results of multiple second samples and their respective binding free energy reference labels, and the second prediction results of multiple third samples and their respective affinity reference labels, is similar to steps 1230-1250 above. To save space, it will not be described in detail here.
[0280] The advantage of this embodiment is that, during the training of the original network, on the one hand, it considers the feature extraction of biomolecules from the second sample pair, enabling the original network to learn the structure and physicochemical properties of biomolecules of the same type as the two biomolecules in the first sample pair, and to learn the complex interactions between the third sample biomolecule (e.g., protein) and the fourth sample biomolecule (e.g., peptide), and trains the original network by estimating the binding free energy of the third and fourth sample biomolecules. On the other hand, it considers the feature extraction of biomolecules from the third sample pair, enabling the original network to capture a wide range of molecular interactions at the atomic scale during training, and trains the original network by estimating the binding affinity of the fifth sample biomolecule (e.g., protein) and the sixth biomolecule (e.g., compound). Comprehensive training of the original network is achieved through two pre-training tasks (estimating the binding free energy of protein-peptide complexes and predicting the affinity of protein-compound complexes), improving the overall feature extraction accuracy of the trained feature extraction network.
[0281] It's important to note that the task of estimating the binding free energy of protein-peptide complexes focuses on protein-peptide interactions at the residue (amino acid) scale, enabling pristine networks to capture the role of amino acid motifs and residue-related properties (e.g., inherent disorder tendencies and evolutionary characteristics) in determining binding affinity. The task of predicting the affinity of protein-compound complexes allows models to capture a wide range of molecular interactions at the atomic scale. Compounds, typically small molecules, interact with proteins through various mechanisms, such as hydrogen bonding, hydrophobic interactions, and electrostatic forces. Peptides can also interact with proteins through similar mechanisms. Although they differ in molecular size, knowledge gained from protein-compound affinity can be applied to protein-peptide affinity prediction because some key features and patterns contributing to binding affinity may be preserved in both types of interactions.
[0282] In this embodiment of the disclosure, the original network includes a first extraction module, a second extraction module, a third extraction module, and a fourth extraction module.
[0283] Referring to Figure 16, in one embodiment, step 1530 specifically includes, but is not limited to, the following steps 1610-1670:
[0284] Step 1610: For the second sample pair, determine the first sample molecular structure diagram and the first sample binding region diagram of the third sample biomolecule that has the same molecular type as the second biomolecule, and determine the second sample molecular structure diagram of the fourth sample biomolecule that has the same molecular type as the first biomolecule.
[0285] Step 1620: For the third sample pair, determine the third sample molecular structure diagram and the second sample binding region diagram of the fifth sample biomolecule, which has the same molecular type as the second biomolecule, and determine the sample atomic structure diagram of the sixth sample biomolecule.
[0286] Step 1630: Based on the first extraction module, extract the features of the first sample image from the molecular structure image of the first sample, and extract the features of the second sample image from the molecular structure image of the third sample.
[0287] Step 1640: Based on the second extraction module, extract the features of the third sample map from the first sample combined with the region map, and extract the features of the fourth sample map from the second sample combined with the region map.
[0288] Step 1650: Based on the third extraction module, extract features from the molecular structure diagram of the second sample to obtain features from the fifth sample diagram;
[0289] Step 1660: Based on the fourth extraction module, extract the features of the sixth sample image from the sample atomic structure image;
[0290] Step 1670: Integrate the first sample map features, the third sample map features, and the fifth sample map features into the first sample molecular map features, and integrate the second sample map features, the fourth sample map features, and the sixth sample map features into the second sample molecular map features.
[0291] Steps 1610-1670 are described in detail below.
[0292] In step 1610, the first sample molecular structure map is used to indicate the result of structural diagram modeling of the third sample biomolecule at the amino acid scale. The first sample binding region map is a sub-map extracted from the first sample molecular structure map that contains only the residues in contact with the fourth sample biomolecule and the edges between these residues. The molecular type of the second biomolecule can be a protein or an antigen, etc.
[0293] The second sample molecular structure diagram is used to indicate the results of structural modeling of the fourth sample biomolecule at the amino acid scale. The molecular types of the fourth sample biomolecule and the first biomolecule can be peptides or antibodies.
[0294] In the specific implementation of this embodiment, the process of step 1610 is similar to that of steps 310-320 described above. To save space, it will not be described again.
[0295] In step 1620, the sample atomic structure diagram is used to indicate the results of the atomic-scale structural modeling of the sixth sample biomolecule. The sixth sample biomolecule is a compound.
[0296] The third sample molecular structure map is used to indicate the results of structural modeling of the fifth sample biomolecule at the amino acid scale. The second sample binding region map is a subgraph extracted from the fifth sample molecular structure map, containing only the residues that contact the sixth sample biomolecule and the edges between these residues. The fifth sample biomolecule is a protein.
[0297] In the specific implementation of this embodiment, the process of step 1620 is similar to that of steps 310-320 described above. To save space, it will not be described again.
[0298] In step 1630, the first sample map features are used to indicate the higher-order representation information learned by the first extraction module from the first sample molecular structure map.
[0299] The features of the second sample image are used to indicate the higher-order representation information learned by the first extraction module from the molecular structure image of the second sample.
[0300] In the specific implementation of this embodiment, step 1630 is similar to steps 810-840 described above. To save space, it will not be described again.
[0301] In step 1640, the third sample map features are used to indicate the higher-order representation information learned by the second extraction module from the first sample combined region map.
[0302] The fourth sample map features are used to indicate the higher-order representation information learned by the second extraction module from the second sample combined region map.
[0303] In the specific implementation of this embodiment, step 1640 is similar to steps 810-840 described above. To save space, it will not be described again.
[0304] In step 1650, the fifth sample map features are used to indicate the higher-order representation information learned by the third extraction module from the second sample molecular structure map.
[0305] In the specific implementation of this embodiment, step 1650 is similar to steps 810-840 described above. To save space, it will not be described again.
[0306] In step 1660, the sixth sample map feature is used to indicate the higher-order representation information learned by the fourth extraction module from the sample atomic structure map.
[0307] In the specific implementation of this embodiment, step 1660 is similar to steps 810-840 described above. To save space, it will not be described again.
[0308] In step 1670, firstly, the features of the first sample image, the third sample image, and the fifth sample image are concatenated to obtain the first sample molecular image features. Next, the features of the second sample image, the fourth sample image, and the sixth sample image are concatenated to obtain the second sample molecular image features.
[0309] The advantage of this embodiment is that by setting up multiple feature modules corresponding to various types in the feature extraction network, features are extracted from different types of biomolecular structure maps through different extraction modules, and these features are integrated into a more comprehensive molecular map feature. This can gradually enhance the feature extraction network's ability to understand molecular structure and function. This approach can more accurately capture and utilize the interaction information between molecules, reduce the mutual interference of extracting graph structure feature information of different types of biomolecules at the same time, and improve the performance and robustness of the feature extraction network in affinity prediction.
[0310] Referring to Figure 17, in one embodiment, step 1560 specifically includes, but is not limited to, the following steps: 1710-1730:
[0311] Step 1710: Determine the third loss function based on the first prediction results of multiple second samples and the combined free energy reference label;
[0312] Step 1720: Determine the fourth loss function based on the second prediction results and affinity reference labels of multiple third samples for their respective samples;
[0313] Step 1730: Based on the third and fourth loss functions, train the original network to obtain the feature extraction network.
[0314] Steps 1710-1730 are described in detail below.
[0315] In step 1710, the third loss function is used to indicate the overall degree of difference between the predicted binding free energy and the actual binding free energy for all second sample pairs.
[0316] In this specific implementation, since both the first prediction result and the binding free energy reference label are numerical data, the specific process of step 1710 is similar to that of step 1230 described above. To save space, it will not be repeated here.
[0317] In step 1720, the fourth loss function is used to indicate the overall degree of difference between the predicted binding affinity and the actual binding affinity for all third sample pairs.
[0318] In this specific implementation, since both the second prediction result and the affinity reference label are numerical data, the specific process of step 1720 is similar to that of step 1230 described above. To save space, it will not be repeated here.
[0319] In step 1730, the specific process of training the original network based on the third and fourth loss functions is similar to that in step 1250 above. To save space, it will not be described again.
[0320] The advantage of this embodiment is that the third loss function can more accurately evaluate the model's performance in predicting binding free energy. The fourth loss function can more accurately evaluate the model's performance in predicting binding affinity. By using both the third and fourth loss functions simultaneously, multi-task learning is achieved in the original network during model training. This allows for comprehensive optimization of the original network, enabling it to learn how to extract feature information useful for multiple tasks from the input data, thereby improving the feature extraction capability of the original model and increasing the accuracy of feature extraction in the trained feature extraction network.
[0321] In one embodiment, step 1560 specifically includes, but is not limited to, the following steps:
[0322] Step 1: Obtain multiple fourth sample pairs, wherein each fourth sample pair includes a seventh sample biomolecule, an eighth sample biomolecule, and an affinity reference label for the seventh sample biomolecule and the eighth sample biomolecule, and the seventh sample biomolecule and the eighth sample biomolecule have the same molecular type.
[0323] Step 2: Input multiple fourth sample pairs into the original network to generate feature representations, and obtain the third sample molecular map features of each of the multiple fourth sample pairs;
[0324] Step 3: For each fourth sample pair, perform affinity prediction based on the molecular map features of the third sample to obtain the third prediction result;
[0325] Step 4: Based on the first prediction results of multiple second samples and their respective binding free energy reference labels, the second prediction results of multiple third samples and their respective affinity reference labels, and the third prediction results of multiple fourth samples and their respective affinity reference labels, train the original network to obtain the feature extraction network.
[0326] The fourth sample pair refers to the protein-protein pair used for model training, while the seventh and eighth sample biomolecules are both proteins. The affinity reference label indicates the true binding affinity between the seventh and eighth sample biomolecules.
[0327] The third sample molecular graph features are used to indicate the combined feature information of the graph structure of the two sample biomolecules in the fourth sample pair.
[0328] The third prediction result is used to indicate the predicted binding affinity for the seventh and eighth sample biomolecules.
[0329] Specifically, step 1 is similar to step 1510 above, step 2 is consistent with the method of feature representation of the second sample pair in step 1530 above, and steps 3 and 1550 are similar. Step 4 is similar to steps 1710-1730 above. To save space, it will not be described in detail.
[0330] The advantage of this embodiment is that it introduces sample biomolecules of the same molecular type, such as protein-protein pairs, as training samples, which can effectively expand the number of training samples. When pre-training the original network, a task for predicting the binding affinity of protein-protein pairs is added, which changes the original network from learning two tasks to learning three tasks. This can further improve the learning breadth of the original network, thereby improving the feature extraction accuracy of the trained feature extraction network.
[0331] The implementation details of the overall training method and application method of the affinity prediction model according to the present disclosure will be described in detail below with reference to Figures 18, 14 and 11.
[0332] In this embodiment, the affinity prediction model is designed to predict the binding affinity between proteins and peptides.
[0333] The overall training of the affinity prediction model can be divided into three stages. The first stage is to pre-train the feature extraction network, the second stage is to build a pre-trained model based on the feature extraction network, and the third stage is to fine-tune the pre-trained model based on a combination of supervised and semi-supervised learning.
[0334] The first stage:
[0335] Figure 18 illustrates the pre-training of a feature extraction network based on the original network. Specifically, the pre-trained feature extraction network includes two pre-training tasks: predicting the binding free energy of proteins and peptides, and predicting the binding affinity of proteins and compounds.
[0336] This paper addresses the task of predicting the binding free energy of proteins and peptides. First, protein-peptide pairs are extracted from the PCSB PDB biomolecular database, and the binding free energy ΔG between the protein and peptide pairs is calculated, where the protein sequence is [LPGEGTQVHPRAPLLLK] and the peptide sequence is [TSAFEASDL]. Next, based on the sequence characteristics, intrinsic disorder fraction, and evolutionary characteristics of each residue in the peptide, a peptide molecular structure map and node features of each node in the peptide molecular structure map are generated. Simultaneously, based on the geometric structure characteristics, sequence characteristics, and evolutionary characteristics of each residue in the protein, a protein molecular structure map and node features of each node in the protein molecular structure map are generated. Next, binding pockets in the protein are identified and masked, and based on the generated mask sequence, a binding region map is extracted from the protein molecular structure map. Further, a third extraction module is used to extract graph features from the peptide molecular structure map, a first extraction module is used to extract graph features from the protein molecular structure map, and a second extraction module is used to extract graph features from the binding region map. Finally, the extracted three types of graphical features are input into the regression prediction layer to predict the binding free energy, thus obtaining the predicted binding free energy of the protein-peptide pair. Furthermore, for proteins without binding pockets, only the graphical features of the peptide and protein molecular structures can be extracted and input into the regression prediction layer to predict the binding free energy. Based on this, the third loss function mentioned above can be generated according to the difference between the predicted and calculated binding free energies of each protein-peptide pair.
[0337] This paper addresses the task of predicting the binding affinity of proteins and compounds. First, protein-compound pairs are extracted from the biomolecular database ChEMBL+CrossDocked, and the binding affinity between these pairs is extracted. Next, binding pockets in the proteins are identified and masked, and binding region maps are extracted from the protein molecular structure map based on the generated mask sequences. Simultaneously, the atomic structure of the compounds is converted into a SMILES string representation, and atomic structure maps and node features of each node in the atomic structure maps are generated based on the SMILES string representation. Further, graph features of the atomic structure maps are extracted using a fourth extraction module, graph features of the protein molecular structure maps are extracted using a first extraction module, and graph features of the binding region maps are extracted using a second extraction module. Finally, the extracted graph features are input into a separate regression prediction layer (opposite to the regression prediction layer used to predict binding free energy) to predict the binding affinity of the protein-compound pairs. Furthermore, for proteins without binding pockets, only the graph features of the atomic structure maps and the graph features of the protein molecular structure maps can be extracted and input into the regression prediction layer for binding affinity prediction. Based on this, the fourth loss function mentioned above can be generated by comparing the predicted binding affinity of each protein-compound pair with the binding affinity extracted from the biomolecular database.
[0338] Finally, the parameters of the first, second, third, and fourth extraction modules in the feature extraction network are adjusted using the third and fourth loss functions to obtain the trained feature extraction network.
[0339] The second stage:
[0340] Discard the regression prediction layer used in Figure 18, and connect the preset self-attention layer, regression prediction layer and trained feature extraction network in sequence to construct a complete pre-trained model. The specific structure of the complete pre-trained model can be seen in Figure 11.
[0341] It should be noted that the regression prediction layer in the pre-trained model can be a neural network model with the same structure as the regression prediction layer used in Figure 18, or it can be a neural network model with a different structure. The regression prediction layer in the pre-trained model of this embodiment is preset separately and is not applied to the regression prediction layer in the pre-training task of Figure 18.
[0342] The third stage:
[0343] The pre-trained model is adjusted to obtain an affinity prediction model that meets the expected requirements. The specific process of pre-training the model can be referred to Figure 14 and steps 1210-1250 above.
[0344] Model application phase:
[0345] When applying the trained affinity prediction model to affinity prediction, graph structure generation and graph feature generation are performed for the proteins and peptides to be predicted. The specific methods for graph structure generation can be found in the section on generating peptide molecular structure diagrams, protein molecular structure diagrams, binding region diagrams, and original structure diagrams in Figure 18. The specific process of graph feature generation, and the specific process of affinity prediction based on multiple extracted graph features, can be found in Figure 11 and steps 340-350 above. For the sake of brevity, these will not be elaborated further.
[0346] Referring to Figure 19 below, the performance metrics of the affinity prediction model trained under different sample partitioning methods according to the embodiments of this disclosure, and the experimental data differences between the performance metrics of other models, will be described in detail and by way of example.
[0347] To verify the effectiveness of the affinity prediction model in this disclosure, the target protein-peptide affinity prediction performance was tested using 1532 protein-peptide affinity datasets from the biomolecular database (PDBbind). In this test, the prediction performance of different models under different partitioning methods was compared. Models 1 to 7 are other neural network models that differ from the affinity prediction model in this disclosure. To comprehensively evaluate the generalization ability of different methods, experiments were conducted in this disclosure under different partitioning methods, including novel protein partitioning (no duplicate proteins in the training and test sets), novel peptide partitioning (no duplicate peptides in the training and test sets), and novel peptide and protein partitioning (both proteins and peptides in the training and test sets are novel). Simultaneously, multiple indicators were used for evaluation, including mean-square error (MSE), Pearson correlation coefficient, Spearman correlation coefficient, and concordance index (CI). A lower MSE is better, while higher Pearson correlation coefficient, Spearman correlation coefficient, and CI are better. As shown in Figure 19, the affinity prediction model of this disclosure embodiment outperforms other models in all MSE, Pearson correlation coefficient, Spearman correlation coefficient, and CI across the three partitioning methods, demonstrating the superiority of the affinity prediction model trained in this disclosure embodiment.
[0348] The differences in experimental data in ablation experiments using the affinity prediction model of the present disclosure embodiments are illustrated in detail below with reference to Figures 20A-20D.
[0349] To examine the contribution of each module of the affinity prediction model to improving the accuracy of affinity prediction, different modules were removed, and the affinity prediction performance of the model after removing the module was evaluated. Similarly, experiments were conducted using mean squared error, Pearson correlation coefficient, Spearman correlation coefficient, and consistency index as evaluation dimensions. The elimination strategies in the ablation experiments included, but were not limited to, the following: not removing any part, removing binding free energy pre-training, removing protein-compound pre-training, removing peptide sequence features, removing the intrinsic disorder fraction of peptides, removing peptide evolutionary features, removing protein sequence features, removing protein geometric features, removing protein evolutionary features, removing self-attention mechanisms, removing the fourth extraction module, removing the third extraction module, removing the second extraction module, and removing the first extraction module. As shown in Figures 20A-20D, all modules of the affinity prediction model in this embodiment of the present disclosure contribute to the affinity prediction performance. Among them, the four most important modules are the third feature module for feature extraction of the polypeptide's graph structure, pre-training of the binding free energy of the target protein-peptide pair, evolutionary features of the protein, and pre-training of the affinity of the protein-compound pair. Furthermore, the ablation experiment further demonstrates that the transfer effect of large-scale pre-training in the affinity prediction of the target protein-peptide pair in this embodiment of the present disclosure (i.e., learning general knowledge from large-scale protein-ligand binding data and then fine-tuning it for protein-peptide affinity prediction) helps improve the accuracy of affinity prediction.
[0350] The apparatus and device according to embodiments of this disclosure will now be described.
[0351] It is understood that although the steps in the above flowcharts are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated in this embodiment, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the above flowcharts may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps.
[0352] It should be noted that in various specific embodiments of this application, when processing is required based on data related to the characteristics of the target object, such as target object attribute information or a set of attribute information, the permission or consent of the target object will be obtained first. Furthermore, the collection, use, and processing of this data will comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require obtaining target object attribute information, separate permission or consent from the target object will be obtained through pop-ups or redirection to a confirmation page. Only after obtaining the target object's separate permission or consent will the necessary target object-related data for the normal operation of the embodiments of this application be obtained.
[0353] Figure 21 is a schematic diagram of the affinity prediction device 2100 provided in an embodiment of this disclosure. The affinity prediction device 2100 includes:
[0354] The first determining unit 2110 is used to determine the first molecular structure diagram and atomic structure diagram corresponding to the first biomolecule, and the second molecular structure diagram corresponding to the second biomolecule;
[0355] The cutting unit 2120 is used to cut out a binding region map from the second molecular structure diagram, wherein the binding region map is used to indicate the residues in the second biomolecule that bind to the first biomolecule.
[0356] Extraction unit 2130 is used to extract the first image features of the first molecular structure map, the second image features of the atomic structure map, the third image features of the second molecular structure map, and the fourth image features of the binding region map, respectively.
[0357] The second determining unit 2140 is used to determine molecular interaction features based on the features of the first graph, the features of the second graph, the features of the third graph, and the features of the fourth graph;
[0358] The prediction unit 2150 is used to predict the affinity of a first biomolecule and a second biomolecule based on molecular interaction characteristics, and obtain the binding affinity prediction result.
[0359] Optionally, the first molecular structure diagram is generated using residues in the first biomolecule as nodes and the relationships between residues as edges;
[0360] Extraction unit 2130 is used for:
[0361] The first determining module (not shown) is used to determine the node characteristics of each node in the first molecular structure diagram;
[0362] The fusion module (not shown) is used to fuse the node features of the neighboring nodes of each node to obtain fused node features.
[0363] The first generation module (not shown) is used to generate features based on the individual node features of multiple nodes and the fused node features to obtain the first graph features.
[0364] Optionally, the first generation module (not shown) is used for:
[0365] For each node, cross-attention calculation is performed based on node features and fused node features to obtain node interaction features;
[0366] By weighting the node interaction features and node features, we obtain the node representation vector of the node.
[0367] The node representation vectors of multiple nodes are concatenated to obtain the molecular graph representation vector of the first molecular structure graph.
[0368] Global average pooling is performed on the molecular graph representation vectors to obtain the features of the first graph.
[0369] Optionally, the first determining module (not shown) is used for:
[0370] For each node, the residues corresponding to the node are one-hot encoded to obtain the sequence features of the node;
[0371] The intrinsic disorder characteristics of nodes are determined based on the intrinsic disorder fraction of the residues corresponding to the nodes.
[0372] The evolutionary characteristics of the nodes are obtained by embedding the residues corresponding to the nodes.
[0373] The node features are obtained by concatenating sequence features, inherent disorder features, and evolutionary features.
[0374] Optionally, the extraction unit 2130 is used for:
[0375] Based on a feature extraction network of a preset affinity prediction model, the first map features of the first molecular structure map, the second map features of the atomic structure map, the third map features of the second molecular structure map, and the fourth map features of the binding region map are extracted respectively.
[0376] The affinity prediction device 2100 also includes an adjustment unit (not shown), which includes:
[0377] A first acquisition module (not shown) is used to acquire multiple first sample pairs and multiple second sample pairs, wherein each first sample pair includes a first sample biomolecule, a second sample biomolecule, and an affinity reference label between the first sample biomolecule and the second sample biomolecule, and each second sample pair includes a third sample biomolecule, a fourth sample biomolecule, and an affinity pseudo-label between the third sample biomolecule and the fourth sample biomolecule.
[0378] The first prediction module (not shown) is used to input the first sample pair and the second sample pair into the pre-trained model to predict affinity, and obtain the first sample prediction result of the first sample pair and the second sample prediction result of the second sample pair.
[0379] The second determining module (not shown) is used to determine the first loss function based on the affinity reference labels of multiple first samples and the prediction results of the first samples.
[0380] The third determination module (not shown) is used to determine the second loss function based on the affinity pseudo-labels and prediction results of multiple second samples for their respective second samples;
[0381] An adjustment module (not shown) is used to adjust the parameters of the pre-trained model based on the first loss function and the second loss function to obtain the affinity prediction model.
[0382] Optionally, the affinity pseudo-labels for the second sample pair are determined as follows:
[0383] Multiple first sample pairs are grouped to determine multiple benchmark sample pair groups, wherein the multiple first sample pairs contained in each benchmark sample pair group are not completely identical.
[0384] Based on multiple benchmark sample pairs, pre-trained models are trained separately to obtain intermediate models corresponding to each of the multiple benchmark sample pairs. The total number of intermediate models is equal to the number of multiple benchmark sample pairs.
[0385] For each second sample pair, the second sample pair is input into each intermediate model to predict affinity, resulting in multiple affinity prediction results;
[0386] Based on multiple affinity prediction results, the affinity pseudo-labels for the second sample pair were determined.
[0387] Optionally, the affinity prediction device 2100 further includes a training unit (not shown), which includes:
[0388] The second acquisition module (not shown) is used to acquire multiple third sample pairs, wherein each third sample pair includes a fifth sample biomolecule, a sixth sample biomolecule, and an affinity reference label between the fifth sample biomolecule and the sixth sample biomolecule, and the fifth sample biomolecule and the sixth sample biomolecule have different molecular types.
[0389] The fourth determining module (not shown) is used to determine the binding free energy reference label of the third sample biomolecule and the fourth sample biomolecule in each second sample pair;
[0390] The second generation module (not shown) is used to input multiple second sample pairs and third sample pairs into the original network to generate feature representations, thereby obtaining the first sample molecular map features of each of the multiple second sample pairs and the second sample molecular map features of each of the multiple third sample pairs.
[0391] The second prediction module (not shown) is used to predict the free energy for each second sample pair based on the molecular graph features of the first sample, and obtain the first prediction result.
[0392] The third prediction module (not shown) is used to predict affinity for each third sample pair based on the molecular map features of the second sample to obtain the second prediction result.
[0393] The training module (not shown) is used to train the original network based on the first prediction results of multiple second samples and the combined free energy reference label, as well as the second prediction results of multiple third samples and the affinity reference label, to obtain the feature extraction network.
[0394] Optionally, the original network includes a first extraction module, a second extraction module, a third extraction module, and a fourth extraction module;
[0395] The second generation module (not shown) is used for:
[0396] For the second sample pair, determine the first sample molecular structure diagram and the first sample binding region diagram of the third sample biomolecule that has the same molecular type as the second biomolecule, and determine the second sample molecular structure diagram of the fourth sample biomolecule that has the same molecular type as the first biomolecule.
[0397] For the third sample pair, determine the molecular structure diagram of the fifth sample biomolecule, which has the same molecular type as the second biomolecule, and the binding region diagram of the second sample, and determine the atomic structure diagram of the sixth sample biomolecule.
[0398] Based on the first extraction module, the features of the first sample image are extracted from the molecular structure image of the first sample, and the features of the second sample image are extracted from the molecular structure image of the third sample.
[0399] Based on the second extraction module, the features of the third sample map are extracted from the first sample combined with the region map, and the features of the fourth sample map are extracted from the second sample combined with the region map.
[0400] Based on the third extraction module, features of the fifth sample image are extracted from the molecular structure image of the second sample.
[0401] Based on the fourth extraction module, features of the sixth sample map are extracted from the sample atomic structure map;
[0402] The features of the first sample map, the third sample map, and the fifth sample map are integrated into the first sample molecular map features, and the features of the second sample map, the fourth sample map, and the sixth sample map are integrated into the second sample molecular map features.
[0403] Optionally, the training module (not shown) is used for:
[0404] A third loss function is determined based on the first prediction results of multiple second samples and the combined free energy reference labels.
[0405] A fourth loss function is determined based on the second prediction results and affinity reference labels of multiple third samples for their respective third samples;
[0406] The original network is trained based on the third and fourth loss functions to obtain the feature extraction network.
[0407] Optionally, the training module (not shown) is used for:
[0408] Multiple fourth sample pairs are obtained, wherein each fourth sample pair includes a seventh sample biomolecule, an eighth sample biomolecule, and an affinity reference label for the seventh sample biomolecule and the eighth sample biomolecule, wherein the seventh sample biomolecule and the eighth sample biomolecule have the same molecular type.
[0409] Multiple fourth sample pairs are input into the original network to generate feature representations, resulting in the third sample molecular map features of each fourth sample pair.
[0410] For each fourth sample pair, affinity prediction is performed based on the molecular map features of the third sample to obtain the third prediction result;
[0411] The original network is trained based on the first prediction results of multiple second samples and their respective binding free energy reference labels, the second prediction results of multiple third samples and their respective affinity reference labels, and the third prediction results of multiple fourth samples and their respective affinity reference labels, to obtain the feature extraction network.
[0412] Optionally, the interception unit 2120 is used for:
[0413] Based on the binding sites of the first and second biomolecules, and the sequence of the second biomolecule, a mask sequence corresponding to the second biomolecule is generated.
[0414] Based on the mask sequence, the second molecular structure map is segmented to obtain the binding region map.
[0415] Optionally, the first molecular structure diagram corresponding to the first biomolecule is determined in the following way:
[0416] To determine the residues contained in the first biomolecule and the contact relationships between the residues;
[0417] Each residue is identified as a node, and edges are constructed between nodes corresponding to residues with contact relationships to generate the first molecular structure diagram.
[0418] Optionally, the atomic structure diagram corresponding to the first biomolecule is determined in the following way:
[0419] Convert the sequence of the first biomolecule into the target string;
[0420] Extract atoms from the target string as nodes, and extract the bonds between atoms in the target string as edges to obtain the atomic structure graph.
[0421] Optionally, the intrinsic disorder fraction of the residues corresponding to the node is determined in the following way:
[0422] The sequence of the first biomolecule is input into a preset online tool for disorder prediction, and a disorder tendency graph of each residue in the first biomolecule is obtained.
[0423] For the residues corresponding to the nodes, the intrinsic disorder fraction of the residues is determined based on the disorder tendency graph of the residues.
[0424] Optionally, the first biomolecule belongs to the first type;
[0425] Embedding of the residues corresponding to the nodes yields the evolutionary characteristics of the nodes, including:
[0426] The evolutionary features of nodes are obtained by embedding the residues corresponding to the nodes based on the pre-trained first extraction network.
[0427] The first extraction network is obtained by training the second extraction network based on the first type of biomolecular samples; the second extraction network is used to extract the evolutionary features of each node of the second molecular structure diagram corresponding to the second biomolecule.
[0428] The first extraction network was trained in the following way:
[0429] The probability distribution is obtained from the sequence of the biomolecular sample through the second extraction network. The sequence of the biomolecular sample includes the identifiers corresponding to each sequentially arranged component element in the biomolecular sample. In the sequence of the biomolecular sample, the identifiers corresponding to some components of the biomolecular sample are masked. The probability distribution is used to reflect the probability distribution of the components corresponding to the masked identifiers in the sequence of the biomolecular sample across the components of the first type of biomolecular sample.
[0430] Based on the probability distribution, the parameters of the second extraction network are adjusted to obtain the first extraction network.
[0431] Optionally, the second molecular structure diagram is generated using residues in the second biomolecule as nodes and the relationships between residues as edges;
[0432] The features of the third diagram of the second molecular structure diagram were extracted in the following way:
[0433] For each node in the second molecular structure diagram, the sequence characteristics, evolutionary characteristics, and geometric structural characteristics of the residues corresponding to the node are determined, and the node characteristics are determined based on the splicing results of the sequence characteristics, evolutionary characteristics, and geometric structural characteristics.
[0434] For each node, the node features of the node's neighboring nodes are fused to obtain fused node features;
[0435] The third graph features are generated based on the individual node features of multiple nodes and the fused node features.
[0436] Referring to Figure 22, which is a partial structural block diagram of a terminal implementing the affinity prediction method of this disclosure, the terminal includes: a radio frequency (RF) circuit 2210, a memory 2215, an input unit 2230, a display unit 2240, a sensor 2250, an audio circuit 2260, a wireless fidelity (WiFi) module 2270, a processor 2280, and a power supply 2290, etc. Those skilled in the art will understand that the terminal structure shown in Figure 22 does not constitute a limitation on a mobile phone or computer, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0437] The RF circuit 2210 can be used to receive and transmit signals during information transmission or calls. In particular, it receives downlink information from the base station and processes it with the processor 2280; in addition, it transmits uplink data to the base station.
[0438] The memory 2215 can be used to store software programs and modules. The processor 2280 executes various functional applications and data processing of the target terminal by running the software programs and modules stored in the memory 2215.
[0439] The input unit 2230 can be used to receive input numeric or character information, and to generate key signal inputs related to the settings and function control of the target terminal. Specifically, the input unit 2230 may include a touch panel 2231 and other input devices 2232.
[0440] Display unit 2240 can be used to display input or provided information, as well as various menus of the target terminal. Display unit 2240 may include display panel 2241.
[0441] Audio circuitry 2260, speaker 2261, and microphone 2262 provide an audio interface.
[0442] In this embodiment, the processor 2280 included in the terminal can execute the affinity prediction method of the previous embodiment.
[0443] The terminals disclosed in this embodiment include, but are not limited to, mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle terminals, and aircraft. The embodiments of this invention can be applied to various scenarios, including but not limited to data security, blockchain, data storage, and information technology.
[0444] Figure 23 is a partial structural block diagram of a server implementing the affinity prediction method of this disclosure. The server can vary considerably due to different configurations or performance, and may include one or more central processing units (CPUs) 2322 (e.g., one or more processors) and memory 2332, and one or more storage media 2330 (e.g., one or more mass storage devices) for storing application programs 2342 or data 2344. The memory 2332 and storage media 2330 may be temporary or persistent storage. The program stored in the storage media 2330 may include one or more modules (not shown in the figure), each module including a series of instruction operations on the server. Furthermore, the CPU 2322 may be configured to communicate with the storage media 2330 and execute the series of instruction operations in the storage media 2330 on the server.
[0445] The server may also include one or more power supplies 2326, one or more wired or wireless network interfaces 2350, one or more input / output interfaces 2358, and / or one or more operating systems 2341, such as Windows Server. TM Mac OS X TM Unix TM Linux TM FreeBSD TM etc.
[0446] The central processing unit 2322 in the server can be used to execute the affinity prediction method of the embodiments of this disclosure.
[0447] This disclosure also provides a computer-readable storage medium for storing program code for executing the affinity prediction methods of the foregoing embodiments.
[0448] This disclosure also provides a computer program product comprising a computer program. A processor of a computer device reads and executes the computer program, causing the computer device to perform the affinity prediction method described above.
[0449] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in this disclosure and the foregoing drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “including,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatuses.
[0450] It should be understood that in this disclosure, "at least one item" means one or more, and "more than one" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0451] It should be understood that in the description of the embodiments disclosed herein, "multiple" means two or more, "greater than", "less than", "exceeding" etc. are understood to exclude the number itself, and "above", "below", "within" etc. are understood to include the number itself.
[0452] In the several embodiments provided in this disclosure, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.
[0453] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0454] Furthermore, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0455] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0456] It should also be understood that the various implementation methods provided in this disclosure can be combined arbitrarily to achieve different technical effects.
[0457] The above is a detailed description of the embodiments of this disclosure. However, this disclosure is not limited to the above embodiments. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of this disclosure. All such equivalent modifications or substitutions are included within the scope defined by the claims of this disclosure.
Claims
1. An affinity prediction method, the method being performed by an electronic device, the method comprising: Determine the first molecular structure diagram and atomic structure diagram corresponding to the first biomolecule, and the second molecular structure diagram corresponding to the second biomolecule; A binding region diagram is extracted from the second molecular structure diagram, wherein the binding region diagram is used to indicate the residues in the second biomolecule that bind to the first biomolecule. The first image features of the first molecular structure map, the second image features of the atomic structure map, the third image features of the second molecular structure map, and the fourth image features of the binding region map are extracted respectively. Based on the features of the first graph, the features of the second graph, the features of the third graph, and the features of the fourth graph, molecular interaction features are determined. Based on the molecular interaction characteristics, the affinity of the first biomolecule and the second biomolecule is predicted to obtain the binding affinity prediction results.
2. The affinity prediction method according to claim 1, wherein the first molecular structure diagram is generated using residues in the first biomolecule as nodes and the relationships between the residues as edges; The first feature of the first molecular structure diagram was extracted in the following way: Determine the node characteristics of each node in the first molecular structure diagram; For each node, the node features of the neighboring nodes of the node are fused to obtain fused node features; The first graph features are obtained by generating features based on the individual node features of multiple nodes and the fused node features.
3. The affinity prediction method according to claim 2, wherein the step of generating features based on the node features of multiple nodes and the fused node features to obtain the first graph features includes: For each node, cross-attention calculation is performed based on the node features and the fused node features to obtain node interaction features; The node interaction features and the node features are weighted to obtain the node representation vector of the node; The node representation vectors of the plurality of nodes are concatenated to obtain the molecular graph representation vector of the first molecular structure graph. Global average pooling is performed on the molecular graph representation vector to obtain the features of the first graph.
4. The affinity prediction method according to claim 2 or 3, wherein determining the node characteristics of each node in the first molecular structure diagram includes: For each node, the residues corresponding to the node are one-hot encoded to obtain the sequence features of the node; The intrinsic disorder characteristics of the node are determined based on the intrinsic disorder fraction of the residues corresponding to the node. The evolutionary characteristics of the node are obtained by embedding the residues corresponding to the node. The node features are obtained by concatenating the sequence features, the inherent disorder features, and the evolutionary features.
5. The affinity prediction method according to any one of claims 1 to 4, wherein the extraction of the first map features of the first molecular structure map, the second map features of the atomic structure map, the third map features of the second molecular structure map, and the fourth map features of the binding region map respectively comprises: Based on the feature extraction network of the preset affinity prediction model, the first map features of the first molecular structure map, the second map features of the atomic structure map, the third map features of the second molecular structure map, and the fourth map features of the binding region map are extracted respectively. The affinity prediction model is obtained by adjusting the pre-trained model in the following way: Multiple first sample pairs and multiple second sample pairs are obtained, wherein each first sample pair includes a first sample biomolecule, a second sample biomolecule, and an affinity reference label between the first sample biomolecule and the second sample biomolecule, and each second sample pair includes a third sample biomolecule, a fourth sample biomolecule, and an affinity pseudo-label between the third sample biomolecule and the fourth sample biomolecule. The first sample pair and the second sample pair are input into the pre-trained model to predict affinity, and the first sample prediction result of the first sample pair and the second sample prediction result of the second sample pair are obtained. Based on the affinity reference labels of the multiple first samples and the prediction results of the first samples, a first loss function is determined; Based on the affinity pseudo-labels of the multiple second samples and the prediction results of the second samples, a second loss function is determined; Based on the first loss function and the second loss function, the parameters of the pre-trained model are adjusted to obtain the affinity prediction model.
6. The affinity prediction method according to claim 5, wherein the affinity pseudo-label of the second sample pair is determined in the following manner: The plurality of first sample pairs are grouped to determine a plurality of baseline sample pair groups, wherein, Each of the aforementioned baseline sample pair groups contains multiple first sample pairs that are not entirely identical; Based on the multiple benchmark sample pairs, the pre-trained models are trained respectively to obtain intermediate models corresponding to each of the multiple benchmark sample pairs, wherein the total number of intermediate models is equal to the number of the multiple benchmark sample pairs; For each of the second sample pairs, the second sample pairs are respectively input into each of the intermediate models to predict affinity, resulting in multiple affinity prediction results; Based on the multiple affinity prediction results, the affinity pseudo-label of the second sample pair is determined.
7. The affinity prediction method according to claim 5 or 6, wherein the feature extraction network of the pre-trained model is trained in the following manner: Obtain multiple third sample pairs, among which, Each of the third sample pairs includes a fifth sample biomolecule, a sixth sample biomolecule, and an affinity reference label for the fifth sample biomolecule and the sixth sample biomolecule, wherein the fifth sample biomolecule and the sixth sample biomolecule are of different molecular types; Determine the binding free energy reference label between the third sample biomolecule and the fourth sample biomolecule in each of the second sample pairs; The plurality of second sample pairs and the plurality of third sample pairs are input into the original network to generate feature representations, thereby obtaining the first sample molecular graph features of each of the plurality of second sample pairs and the second sample molecular graph features of each of the plurality of third sample pairs. For each second sample pair, the free energy is predicted based on the molecular graph features of the first sample to obtain the first prediction result; For each third sample pair, affinity prediction is performed based on the molecular map features of the second sample to obtain a second prediction result; Based on the first prediction results of the multiple second samples and the binding free energy reference label, and the second prediction results of the multiple third samples and the affinity reference label, the original network is trained to obtain the feature extraction network.
8. The affinity prediction method according to claim 7, wherein the original network comprises a first extraction module, a second extraction module, a third extraction module, and a fourth extraction module; The step of inputting the plurality of second sample pairs and the third sample pairs into the original network for feature representation generation, to obtain the first sample molecular map features of each of the plurality of second sample pairs and the second sample molecular map features of each of the plurality of third sample pairs, includes: For the second sample pair, determine the first sample molecular structure diagram and the first sample binding region diagram of the third sample biomolecule that has the same molecular type as the second biomolecule, and determine the second sample molecular structure diagram of the fourth sample biomolecule that has the same molecular type as the first biomolecule. For the third sample pair, determine the third sample molecular structure diagram and the second sample binding region diagram of the fifth sample biomolecule, which has the same molecular type as the second biomolecule, and determine the sample atomic structure diagram of the sixth sample biomolecule. Based on the first extraction module, first sample map features are extracted from the first sample molecular structure map, and second sample map features are extracted from the third sample molecular structure map; Based on the second extraction module, the third sample map features are extracted from the first sample combined region map, and the fourth sample map features are extracted from the second sample combined region map; Based on the third extraction module, features of the fifth sample image are extracted from the molecular structure image of the second sample. Based on the fourth extraction module, features of the sixth sample map are extracted from the sample atomic structure map; The first sample map feature, the third sample map feature, and the fifth sample map feature are integrated into the first sample molecular map feature, and the second sample map feature, the fourth sample map feature, and the sixth sample map feature are integrated into the second sample molecular map feature.
9. The affinity prediction method according to claim 7 or 8, wherein training the original network based on the first prediction results of the plurality of second samples and the binding free energy reference label, and the second prediction results of the plurality of third samples and the affinity reference label, to obtain the feature extraction network, comprises: Based on the first prediction results of the multiple second samples and the binding free energy reference label, a third loss function is determined. Based on the second prediction results of the multiple third samples and the affinity reference label, a fourth loss function is determined; The original network is trained based on the third loss function and the fourth loss function to obtain the feature extraction network.
10. The affinity prediction method according to claim 7 or 8, wherein training the original network based on the first prediction results of the plurality of second samples and the binding free energy reference label, and the second prediction results of the plurality of third samples and the affinity reference label, to obtain the feature extraction network, comprises: Multiple fourth sample pairs are obtained, wherein each fourth sample pair includes a seventh sample biomolecule, an eighth sample biomolecule, and an affinity reference label for the seventh sample biomolecule and the eighth sample biomolecule, wherein the seventh sample biomolecule and the eighth sample biomolecule have the same molecular type. The multiple fourth sample pairs are input into the original network to generate feature representations, thereby obtaining the third sample molecular graph features of each of the multiple fourth sample pairs. For each fourth sample pair, affinity prediction is performed based on the molecular graph features of the third sample to obtain the third prediction result; Based on the first prediction results of the multiple second samples and the binding free energy reference label, the second prediction results of the multiple third samples and the affinity reference label, and the third prediction results of the multiple fourth samples and the affinity reference label, the original network is trained to obtain the feature extraction network.
11. The affinity prediction method according to any one of claims 1 to 10, wherein extracting the binding region map from the second molecular structure map comprises: Based on the binding sites of the first biomolecule and the second biomolecule, and the sequence of the second biomolecule, a mask sequence corresponding to the second biomolecule is generated; Based on the mask sequence, the second molecular structure map is segmented to obtain the binding region map.
12. The affinity prediction method according to any one of claims 1 to 11, wherein the first molecular structure diagram corresponding to the first biomolecule is determined by the following method: Determine the residues contained in the first biomolecule and the contact relationships between the residues; Each residue is identified as a node, and edges are constructed between nodes corresponding to residues with the contact relationship to generate the first molecular structure diagram.
13. The affinity prediction method according to any one of claims 1 to 12, wherein the atomic structure diagram corresponding to the first biomolecule is determined in the following manner: The sequence of the first biomolecule is converted into the target string; Atoms are extracted from the target string as nodes, and the bonds between the atoms in the target string are extracted as edges to obtain the atomic structure diagram.
14. The affinity prediction method according to any one of claims 4 to 13, wherein the intrinsic disorder fraction of the residues corresponding to the node is determined by the following manner: The sequence of the first biomolecule is input into a preset online tool for disorder prediction to obtain a disorder tendency graph of each residue in the first biomolecule. For the residues corresponding to the nodes, the intrinsic disorder fraction of the residues is determined based on the disorder tendency graph of the residues.
15. The affinity prediction method according to any one of claims 4 to 13, wherein the first biomolecule belongs to the first type; The process of embedding the residues corresponding to the node to obtain the evolutionary characteristics of the node includes: The evolutionary features of the node are obtained by embedding the residues corresponding to the node based on the pre-trained first extraction network. The first extraction network is obtained by training a second extraction network based on the first type of biomolecular samples; the second extraction network is used to extract the evolutionary features of each node of the second molecular structure diagram corresponding to the second biomolecule. The first extraction network was trained in the following way: The second extraction network obtains a probability distribution based on the sequence of the biomolecular sample; wherein the sequence of the biomolecular sample includes identifiers corresponding to each sequentially arranged component element in the biomolecular sample; in the sequence of the biomolecular sample, the identifiers corresponding to some component elements of the biomolecular sample are masked; the probability distribution is used to reflect the probability distribution of the component elements corresponding to the masked identifiers in the sequence of the biomolecular sample across the component elements of the first type of biomolecular sample. Based on the probability distribution, the parameters of the second extraction network are adjusted to obtain the first extraction network.
16. The affinity prediction method according to any one of claims 1 to 15, wherein the second molecular structure diagram is generated using residues in the second biomolecule as nodes and the relationships between the residues as edges; The features of the third diagram of the second molecular structure diagram were extracted in the following way: For each node in the second molecular structure diagram, the sequence characteristics, evolutionary characteristics, and geometric structural characteristics of the residues corresponding to the node are determined, and the node characteristics of the node are determined based on the splicing result of the sequence characteristics, the evolutionary characteristics, and the geometric structural characteristics. For each node, the node features of the neighboring nodes of the node are fused to obtain fused node features; The third graph features are obtained by generating features based on the individual node features of multiple nodes and the fused node features.
17. An affinity prediction device, the affinity prediction device comprising: The first determining unit is used to determine the first molecular structure diagram and atomic structure diagram corresponding to the first biomolecule, and the second molecular structure diagram corresponding to the second biomolecule; The extraction unit is used to extract a binding region map from the second molecular structure diagram, wherein the binding region map is used to indicate the residues in the second biomolecule that bind to the first biomolecule. The extraction unit is used to extract the first image features of the first molecular structure map, the second image features of the atomic structure map, the third image features of the second molecular structure map, and the fourth image features of the binding region map, respectively. The second determining unit is used to determine molecular interaction features based on the features of the first graph, the features of the second graph, the features of the third graph, and the features of the fourth graph; The prediction unit is used to predict the affinity of the first biomolecule and the second biomolecule based on the molecular interaction characteristics, and obtain the binding affinity prediction result.
18. An electronic device comprising a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the affinity prediction method according to any one of claims 1 to 16.
19. A computer-readable storage medium storing a computer program that, when executed by a processor, implements the affinity prediction method according to any one of claims 1 to 16.
20. A computer program product comprising a computer program that is read and executed by a processor of a computer device, causing the computer device to perform the affinity prediction method according to any one of claims 1 to 16.