Binding analysis of protein molecules to ligand molecules
By acquiring structural feature representations of ligands and protein molecules, constructing complex structural features, and utilizing attention models, the problem of time-consuming and labor-intensive binding analysis in traditional drug discovery is solved, achieving efficient and accurate molecular binding screening and improving the efficiency of drug discovery.
Patent Information
- Application Number
- CN202111013709.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-31
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2041-08-31
AI Technical Summary
In traditional drug discovery, the analysis of the binding of protein molecules and ligand molecules relies on time-consuming and labor-intensive manual experiments, and machine learning-based methods have limited accuracy, making it difficult to efficiently and accurately screen out molecules that bind effectively.
By acquiring structural feature representations of ligand molecules and protein molecules, constructing complex structural feature representations, and using an attention model to generate aggregation feature representations, the effectiveness and affinity of binding are determined, thereby improving the accuracy of molecular binding analysis.
It enables more efficient and accurate molecular binding analysis, effectively screening out ligand molecules that bind to target protein molecules and improving the efficiency of drug discovery.
Smart Images

Figure CN115732038B_ABST
Abstract
Description
BACKGROUND
[0001] In drug discovery, an important task is to analyze whether a protein molecule (e.g., a target protein) can effectively bind with a small drug molecule (also referred to as a ligand molecule, L). Traditional drug discovery processes rely on chemical experiments to study the binding between molecules.
[0002] In recent years, with the development of computer technology, some machine learning techniques have been gradually applied to predict the binding between protein molecules and ligand molecules. Predicting the binding between molecules through machine learning techniques can greatly reduce the cost of drug discovery, and people are increasingly concerned about how to improve the accuracy of molecular ensemble prediction based on machine learning techniques. SUMMARY
[0003] According to implementations of the present disclosure, a scheme of molecular binding analysis is provided. In the analysis scheme, a first feature representation determined based on a structure of a ligand molecule can be obtained, and a second feature representation determined based on a structure of a protein molecule can be obtained. Further, a third feature representation of a complex structure can also be determined, where the complex structure is constructed based on the protein molecule and the ligand molecule. Further, the first feature representation, the second feature representation, and the third feature representation can be used to generate an aggregated feature representation, and determine evaluation information about the binding between the ligand molecule and the protein molecule. The evaluation information can indicate the effectiveness of the binding, or can also indicate the affinity of the binding pose of the binding. In this way, more efficient and accurate binding analysis can be achieved.
[0004] The summary is provided to introduce a selection of concepts in a simplified form that are further described below in the DETAILED DESCRIPTION. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF DRAWINGS
[0005] Figure 1 A block diagram of an example computing device according to some implementations of the present disclosure is shown;
[0006] Figure 2 A flowchart of a process of molecular binding analysis according to some implementations of the present disclosure is shown;
[0007] Figure 3 A schematic block diagram of a molecular binding analysis module according to some implementations of the present disclosure is shown;
[0008] Figure 4 A schematic diagram of an attention model according to some implementations of the present disclosure is shown;
[0009] Figure 5A comparison diagram incorporating a binding affinity prediction scheme with other schemes is shown in accordance with some implementations of the present disclosure; and
[0010] Figure 6 A comparison diagram incorporating a pose prediction scheme with other schemes is shown in accordance with some implementations of the present disclosure.
[0011] In these drawings, like or similar elements are referred to with like or similar reference numerals. DETAILED DESCRIPTION
[0012] The present disclosure will now be discussed with respect to several example implementations. It should be appreciated that these implementations are discussed solely for the purposes of exemplifying the present disclosure and to enable those with ordinary skill in the art to better understand and thus implement the present disclosure, and are not intended to limit the scope of the subject matter in any way.
[0013] As used herein, the term “includes” and its variants are to be read to mean “comprising, but not limited to.” The term “based on” is to be construed as “based at least in part on.” The terms “a” and “an” are to be construed to mean “at least one” The term “another” is to be construed to mean “at least one other.” The terms “first,” “second,” and the like can refer to different or identical objects. Also, the terms “comprising,” “having,” and the like are to be construed to be open-ended and to mean “including, but not limited to.”
[0014] As discussed above, traditional drug discovery processes rely on chemical experiments to detect binding affinity between proteins and ligand molecules. For a particular targeted protein, an experimenter can need to expend significant human and time costs to screen through a vast number of small molecules to find small molecules that can effectively bind to the targeted protein.
[0015] In recent years, computer-aided drug discovery (CADD) has been increasingly applied to improve the cost of drug discovery. However, the accuracy of CADD techniques, such as those based on machine learning techniques, is limited, and there is a desire to improve the accuracy of molecular binding analysis to assist drug discovery efforts.
[0016] According to implementations of the present disclosure, a scheme of molecular binding analysis is provided. In the analysis scheme, a first feature representation determined based on a structure of a ligand molecule can be obtained, and a second feature representation determined based on a structure of a protein molecule can be obtained. Further, a third feature representation of a complex structure can also be determined, where the complex structure is constructed based on the protein molecule and the ligand molecule. Further, the first feature representation, the second feature representation, and the third feature representation can be used to generate an aggregated feature representation, and determine evaluation information about a binding between the ligand molecule and the protein molecule. The evaluation information can indicate an effectiveness of the binding, or can also indicate an affinity of a binding pose of the binding. Thereby, a more efficient and accurate binding analysis can be achieved.
[0017] By aggregating the features of the ligand molecule, the features of the protein molecule, and the features of the complex structure, embodiments of the present disclosure can take into account interactions between the molecules sufficiently, thereby improving accuracy of the molecular binding analysis.
[0018] The basic principles and several example implementations of the present disclosure are explained below with reference to the accompanying drawings.
[0019] Example device
[0020] Figure 1 A schematic block diagram of an example device 100 that can be used to implement embodiments of the present disclosure is shown. It should be understood that Figure 1 The device 100 shown is merely exemplary and should not be construed as limiting the scope of functionality, or the scope of implementations, as described herein. As shown, Figure 1 As shown, components of the device 100 can include, but are not limited to, one or more processors or processing units 110, a memory 120, a storage device 130, one or more communication units 140, one or more input devices 150, and one or more output devices 160.
[0021] In some implementations, the device 100 can be implemented as various user terminals or service terminals. The service terminals can be servers, large computing devices, etc. provided by various service providers. The user terminals are, such as, any type of mobile terminal, fixed terminal, or portable terminal including a mobile handset, a multimedia computer, a multimedia tablet, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an electronic book device, a game device, or any combinations thereof, including an accessory or peripheral device thereof or any combinations thereof. It is also contemplated that the device 100 can support any type of interface to the user (such as "wearable" circuitry, etc.).
[0022] The processing unit 110 can be a real or virtual processor and capable of performing various processing according to programs stored in the memory 120. In a multi-processor system, multiple processing units perform computer-executable instructions in parallel to improve parallel processing capability of the device 100. The processing unit 110 can also be referred to as a central processing unit (CPU), a microprocessor, a controller, a microcontroller.
[0023] The device 100 typically includes a plurality of computer storage media. Such media can be any available media that is accessible by the device 100 and includes both volatile and non-volatile media, removable and non-removable media. The memory 120 can be volatile memory (such as a register, a cache, a random access memory (RAM)), non-volatile memory (such as a read-only memory (ROM), an electrically erasable programmable read only memory (EEPROM), a flash memory), or some combination thereof. The memory 120 can include one or more molecular binding analysis modules 125 configured to perform the encoding / decoding functions of various implementations described herein. The encoding / decoding modules 125 can be accessed and run by the processing unit 110 to implement the respective functions. The storage device 130 can be a removable or non-removable media and can include machine readable media that is capable of storing information and / or data and can be accessed within the device 100.
[0024] The functionality of the components of device 100 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, device 100 can operate in a networked environment using logical connections to one or more other servers, personal computers (PCs), or other general network nodes. Device 100 can also communicate as needed via communication unit 140 with one or more external devices (not shown), such as database 110, other storage devices, servers, display devices, etc., with one or more devices that enable user interaction with device 100, or with any device that enables device 100 to communicate with one or more other computing devices (e.g., network interface cards, modems, etc.). Such communication can be performed via input / output (I / O) interfaces (not shown).
[0025] Input device 150 can be one or more various input devices, such as a mouse, keyboard, trackball, voice input device, camera, etc. Output device 160 can be one or more output devices, such as a monitor, speaker, printer, etc.
[0026] In some implementations, such as Figure 1 As shown, device 100 can, for example, receive identifiers corresponding to protein molecule 172 and ligand molecule 174 via input device 150. For example, a user can input a PDB file via input device 150 to indicate the corresponding protein molecule 172. The user can also input a string in SMILES (Simplified Molecular Input Line Entry System) format via input device 150 to indicate the corresponding ligand molecule 174.
[0027] In some implementations, the analysis module 125 can determine evaluation information 180 regarding the binding between protein molecule 172 and ligand molecule 174 based on the structure of protein molecule 172 and ligand molecule 174. This evaluation information 180 may, for example, include prediction results output by the binding analysis module 125. Alternatively or supplementarily, in other implementations, evaluation information can be further constructed based on the prediction results output by the analysis module 125 to meet the evaluation requirements of specific work scenarios. The process for determining the evaluation information 180 will be described in detail below.
[0028] Molecular binding analysis
[0029] Figure 2 A flowchart of a process 200 for molecular binding analysis according to some implementations of this disclosure is shown. Process 200 can, for example, be performed by... Figure 1The device 100 in some implementations can be configured to determine the structure of the protein molecule 172 based on user input. The protein molecule 172 can be a target protein to be analyzed in a drug discovery process. The user, for example, can provide a PDB file corresponding to the protein molecule 172 to be analyzed to the device 100.
[0030] As shown in FIG. 2, at block 202, the device 100 obtains a first feature representation of the ligand molecule 174 and a second feature representation of the protein molecule 172, where the first feature representation is determined based on the structure of the ligand molecule 174, and the second feature representation is determined based on the structure of the protein molecule 172. Figure 2
[0031] In some implementations, the device 100 can determine the structure of the protein molecule 172 based on user input for generating the second feature representation. The protein molecule 172 can be a target protein to be analyzed in a drug discovery process. The user, for example, can provide a PDB file corresponding to the protein molecule 172 to be analyzed to the device 100.
[0032] In some implementations, the device 100 can also determine the structure of the ligand molecule 174 based on user input. For example, the user can provide information of the protein molecule 172 and the ligand molecule 174 to be analyzed to the device 100. As an example, the user can provide a SMILES string corresponding to the ligand molecule 174, and the device 100 can further determine the structure of the ligand molecule 174 based on the SMILES string.
[0033] In some implementations, the molecule binding analysis module 125, for example, can be used to screen ligand molecules from a set of candidate ligand molecules that are capable of binding to the protein molecule 172. In this case, the device 100, for example, can automatically select a candidate ligand molecule from the set of candidate ligand molecules as the ligand molecule 174 to be analyzed, and further determine the structure information of the ligand molecule 174.
[0034] In some implementations, the device 100, for example, can determine the first feature representation based on the structure information of the ligand molecule 174. In some implementations, the first feature representation can include a feature graph that can include a plurality of nodes and edges between the nodes. In some implementations, each node in the feature graph can correspond to an atom in the ligand molecule 174, and the edges can correspond to the binding between the atoms.
[0035] In some implementations, the first feature representation can include node features for characterizing the atoms. The node features, for example, can be used to characterize: the symbol of the atom (e.g., C, N, O, F, P, Cl, Br, B, H, etc.), the number of covalent bonds, the charge state, the number of free radical electrons, the hybridization state, the aromaticity, the number of attached hydrogen atoms, whether it is a chiral center, the type of chirality, the amino acid type, or a combination of one or more of the above.
[0036] In some implementations, the first feature representation can further include edge features for characterizing bonding relationships between atoms. The edge features, for example, can be used to characterize: bonding type (e.g., single bond, double bond, aromatic bond, etc.), conjugation status, whether a ring, stereo, connection type (e.g., protein-protein, protein-ligand, ligand-ligand), distance information, or a combination of one or more of the above.
[0037] The device 100 can determine a second feature representation based on the structure of the protein molecule 172. Similarly, the second feature representation can also include a feature map, and include node features for characterizing atoms in the protein molecule 172 and edge features for characterizing bonding relationships between atoms. The information characterized by the node features and the edge features can be similar to the first feature representation discussed above.
[0038] In some implementations, for the second feature representation, the edge features included therein can further characterize, for example, whether a pair of atoms corresponding to an edge are located on the same amino acid.
[0039] At block 204, the device 100 can determine a third feature representation of a complex structure, where the complex structure is constructed based on the protein molecule 172 and the ligand molecule 174. Depending on the specific use of the binding analysis module 125, the complex structure can be determined based on different manners.
[0040] In some embodiments, the binding analysis module 125 can be used, for example, to analyze whether the protein molecule 172 and the ligand molecule 174 can effectively bind. In this case, the device 100 can first construct a complex structure based on the protein molecule 172 and the ligand molecule 174, for example.
[0041] Specifically, the device 100 can first determine a binding pocket of the protein molecule 172. In some implementations, the information of the binding pocket of the protein molecule 172 can be determined based on user input. Illustratively, a user can characterize the binding pocket of the protein molecule 172 through a PDB file.
[0042] Further, the device 100 can determine a set of candidate complex structures based on the binding pocket of the protein molecule 172. Specifically, the device 100 can place the ligand molecule 174 into the binding pocket in an appropriate manner, thereby generating a candidate complex structure.
[0043] Additionally, the device 100 can determine the target complex structure from the set of candidate complex structures based on a binding free energy (BFE) corresponding to the set of candidate complex structures. Illustratively, the device 100 can select the candidate complex structure with the minimum binding free energy from the set of candidate complex structures as the target complex structure.
[0044] In some implementations, the binding analysis module 125, for example, can be used to evaluate a binding pose between the protein molecule 172 and the ligand molecule 174. In this case, the complex structure, for example, can be constructed by a user or determined from a set of candidate complex structures automatically constructed by the device 100.
[0045] Further, after determining the complex structure, the device 100 can further generate a third feature representation of the complex structure. Similar to the first feature representation and the second feature representation, the third feature representation, for example, can also include a feature graph and include node features for characterizing atoms in the complex structure and edge features for characterizing binding relationships between the atoms. The information characterized by the node features and the edge features can be similar to the first feature representation discussed above.
[0046] With continued reference to Figure 2 At block 206, the device 100 generates an aggregated feature representation based on the first feature representation, the second feature representation, and the third feature representation. The process for generating the aggregated feature representation will be discussed with reference to Figure 3 in detail, Figure 3 A schematic block diagram 300 of the molecular binding analysis module 125 according to some implementations of the present disclosure is shown.
[0047] As Figure 3 shown, the binding analysis module 125 can utilize a binding feature determination module 302 to determine the edge features of the ligand molecule 174 and an atom feature determination module 304 to determine the node features of the ligand molecule 174 to obtain the first feature representation.
[0048] In addition, the binding analysis module 125 can utilize a binding feature determination module 310 to determine the edge features of the protein molecule 172 and an atom feature determination module 308 to determine the node features of the protein molecule 172 to obtain the second feature representation.
[0049] As Figure 3 shown, the binding analysis module 125 can utilize a binding feature determination module 305 to determine the edge features of the complex structure 340 and a combination of the node features of the protein molecule 172 and the node features of the ligand molecule 174 as the node features of the complex structure 340 to obtain the third feature representation.
[0050] In some implementations, asFigure 3 As shown, the binding analysis module 125 can also include an attention model 330-1 configured to generate an intermediate feature representation corresponding to a set of inputs of the ligand molecule 174, the complex structure 340, and the protein molecule 172.
[0051] In some implementations, the binding analysis module 125 can determine the first input to the attention model 330-1 based on the first feature representation. In other implementations, the binding analysis module 125 can further process the first feature representation with the first graph model 315-1 to generate the first input to the attention model 330-1.
[0052] In some implementations, the first graph model 315-1 can include a Graph Transformer Model configured to update the first feature representation based on edge features corresponding to edges in the first feature representation and node features corresponding to pairs of nodes of the edges, using a multi-head attention mechanism. In this way, the node features and the edge features can be updated accordingly based on the binding between the nodes.
[0053] Similarly, the binding analysis module 125 can process the second feature representation and the third feature representation with the second graph model 325-1 and the third graph model 320-1, respectively, to determine the second input and the third input to the attention model 330-1, respectively.
[0054] Figure 4 Further shown is a schematic diagram 400 of an attention model according to some implementations of the present disclosure. As shown, the second input 410 can be represented as Figure 4 The third input 420 can be represented as The first input 410 can be represented as where l corresponds to the sequence number of the attention model 330-1.
[0055] As shown, the attention model 330-1 can process the second input 410 with a padding module 440 to make the number of its feature dimensions consistent with the number of feature dimensions of the third input 420. For example, the attention model 330-1 can expand the second input 410 to the feature dimensions corresponding to the third input 420 by a padding zero operation. Figure 4 Similarly, the attention model 330-1 can process the first input 430 with a padding module 445 to make the number of its feature dimensions consistent with the number of feature dimensions of the third input 420.
[0056]
[0057] Furthermore, note that model 330-1 can utilize multipliers 450 and 460 and adder 470 to compute a weighted sum of the padded first input, padded second input, and third input 420, thereby determining the intermediate feature representation 480. For example, the intermediate feature representation 480 can be represented as:
[0058]
[0059] Here, contact represents the join operation, and r and l represent the weight coefficients.
[0060] In some implementations, such as Figure 3 As shown, the combined analysis module 125 may include multiple layers (also called transfer layers) consisting of a first graph model, a second graph model, a third graph model, and an attention model. In each layer, the first graph model can be configured to receive the output of the corresponding first graph model from the previous layer, the second graph model can be configured to receive the output of the corresponding second graph model from the previous layer, and the third graph model can be configured to receive the output of the attention model from the previous layer. Further, the outputs of the first, second, and third graph models are used as a set of inputs to the attention model in that layer to determine new intermediate feature representations.
[0061] exist Figure 3 In the example, the analysis module 125 may include N layers, where the Nth layer includes a first graph model 315-N, a second graph model 325-N, a third graph model 320-N, and an attention model 330-N. The attention model 330-N may have the following characteristics as shown in the reference... Figure 4 The discussed structure is used to generate aggregated feature representations.
[0062] Continue to refer to Figure 2 In 208, device 100 determines evaluation information 180 regarding the binding between ligand molecule 174 and protein molecule 172 based on the polymerization feature representation, wherein the evaluation information 180 indicates the effectiveness of the binding or the affinity of the binding posture.
[0063] In some embodiments, such as Figure 3 As shown, the analysis module 125 may also include a prediction model 335, which is configured to determine the evaluation information 180 based on the aggregated feature representation. That is, the evaluation information 180 is the prediction result output by the prediction model 335.
[0064] In some embodiments, the parameters in the prediction model 335, the one or more first graph models 315-1 to 315-N, the one or more second graph models 325-1 to 325-N, the one or more third graph models 320-1 to 320-N, and the one or more attention models 330-1 to 330-N are determined through co-training.
[0065] In some implementations, the binding analysis module 125 can include a first analysis model, which can have a structure as described in Figure 3 for determining whether a binding between the protein molecule 172 and the ligand molecule 174 can be effective.
[0066] Specifically, the binding analysis module 125 can obtain a set of candidate ligand molecules, which can be, for example, drug molecules for determining whether they can bind to a target protein effectively. Further, the binding analysis module 125 can process pairs of protein molecules-candidate ligand molecules using the first analysis model, and determine the effectiveness of the binding between each pair of protein molecules-candidate ligand molecules based on the evaluation information generated by the first analysis model.
[0067] Further, the binding analysis module 125 can screen target ligand molecules that can bind to the target protein effectively from the set of candidate ligand molecules based on the determined binding effectiveness. In this way, embodiments of the present disclosure can efficiently implement the screening of bindable ligand molecules with respect to a specified protein molecule.
[0068] Figure 5 A comparison diagram 500 of the binding effectiveness prediction scheme according to some implementations of the present disclosure and other schemes is shown. As Figure 5 shown, when performing effective prediction using the scheme of the present disclosure, the scheme of the present disclosure can obtain a higher AUROC (Area Under the Receiver Operating Characteristic curve) compared to other schemes, i.e., improve the accuracy of the prediction of the binding effectiveness.
[0069] In some implementations, the binding analysis module 125 can include a second analysis model, which can have a structure as described in Figure 3 for determining the affinity of a binding pose in which the protein molecule 172 and the ligand molecule 174 bind in a complex structure.
[0070] Specifically, the binding analysis module 125 can construct a set of candidate complex structures based on, for example, the protein molecule 172 and the ligand molecule 174, where the set of candidate complex structures corresponds to a set of candidate binding poses.
[0071] Further, the binding analysis module 125 can determine evaluation information corresponding to the set of candidate complex structures by utilizing the second analysis model, to determine the affinity of the set of candidate binding poses.
[0072] Additionally, the binding analysis module 125 can determine a target binding pose from the set of candidate binding poses based on the affinity indicated by the evaluation information. In some implementations, the evaluation information can indicate two states (e.g., by two labels) to indicate whether the candidate binding pose is a valid binding pose or not. Alternatively, the evaluation information can also be used to indicate a score of the affinity (e.g., can be represented by a value between 0 and 1), for example.
[0073] Exemplarily, the binding analysis module 125 can filter out the candidate binding poses that are identified as valid binding poses, or the candidate binding poses that have an affinity score greater than a threshold, from the set of candidate binding poses as the target binding pose. Thus, embodiments of the present disclosure can efficiently determine the binding pose of a specified protein molecule and a ligand molecule.
[0074] Figure 6 A comparison diagram of the binding pose prediction scheme according to some implementations of the present disclosure and other schemes is shown. Figure 6 The vertical axis of the table represents the RMSD (root-mean-square deviation) of the complex structure corresponding to the binding pose output by the model is less than 2.5 A The horizontal axis represents the proportion of the top-1, top-2, and top-3 binding poses. The top-1 represents the best one binding pose, the top-2 represents the best two binding poses, and the top-3 represents the best three binding poses. From the table, it can be seen that embodiments of the present disclosure can provide more accurate prediction of the binding pose. Figure 6
[0075] In some embodiments, the first analysis model and the second analysis model can be included in the binding analysis module 125 simultaneously, for example. The binding analysis module 125 can utilize the first analysis model to filter out at least one target ligand molecule from a set of candidate ligand molecules, and utilize the second analysis model to further determine a valid binding pose of the protein molecule and the at least one target ligand molecule, for example.
[0076] It should be understood that, as shown in Figure 3 the first analysis model and the second analysis model can have similar model structures, but the number of transfer layers can be different, for example, or the dimension of the features used in each model can also be different, for example.
[0077] In addition, in the training process, the first analysis model and the second analysis model can utilize different training data. The first analysis model for determining the binding effectiveness may, for example, utilize a training data set constructed based on experimental data in reality. The second analysis model for determining the affinity of the binding pose may, for example, utilize the RMSD of the crystal corresponding to the complex structure, for example, the binding pose with RMSD less than 2.0 A can be determined as a positive sample, and the binding pose with RMSD greater than 3.0 A can be determined as a negative sample. In addition, in the training process, the first analysis model and the second analysis model can utilize different training data. The first analysis model for determining the binding effectiveness may, for example, utilize a training data set constructed based on experimental data in reality. The second analysis model for determining the affinity of the binding pose may, for example, utilize the RMSD of the crystal corresponding to the complex structure, for example, the binding pose with RMSD less than 2.0 A can be determined as a positive sample, and the binding pose with RMSD greater than 3.0 A can be determined as a negative sample. In addition, in the training process, the first analysis model and the second analysis model can utilize different training data. The first analysis model for determining the binding effectiveness may, for example, utilize a training data set constructed based on experimental data in reality. The second analysis model for determining the affinity of the binding pose may, for example, utilize the RMSD of the crystal corresponding to the complex structure, for example, the binding pose with RMSD less than 2.0 A can be determined as a positive sample, and the binding pose with RMSD greater than 3.0 A can be determined as a negative sample.
[0078] In this way, the embodiments of the present disclosure can effectively screen out ligand molecules capable of effectively binding to protein molecules, and provide information about the effective binding pose between the protein molecules and the ligand molecules.
[0079] Example implementations
[0080] The following lists some example implementations of the present disclosure.
[0081] In a first aspect, the present disclosure provides a method of molecular binding molecule. The method comprises: obtaining a first feature representation of a ligand molecule and a second feature representation of a protein molecule, the first feature representation being determined based on the structure of the ligand molecule, and the second feature representation being determined based on the structure of the protein molecule; determining a third feature representation of a complex structure, the complex structure being constructed based on the protein molecule and the ligand molecule; generating an aggregated feature representation based on the first feature representation, the second feature representation and the third feature representation; and determining evaluation information (also referred to as prediction information) about the binding between the ligand molecule and the protein molecule based on the aggregated feature representation, the evaluation information indicating the effectiveness of the binding or indicating the affinity of the binding pose of the binding.
[0082] In some implementations, at least one of the first feature representation, the second feature representation and the third feature representation is a feature graph comprising node features and edge features, the node features being used to characterize atoms in the molecule, and the edge features being used to characterize the binding relationship between the atoms.
[0083] In some implementations, generating the aggregated feature representation based on the first feature representation, the second feature representation and the third feature representation comprises: determining a first set of inputs based on the first feature representation, the second feature representation and the third feature representation; processing the first set of inputs using an attention model to determine a first intermediate feature representation; and determining the aggregated feature representation based on the first intermediate feature representation.
[0084] In some implementations, the attention model is a first attention model, and determining the aggregated feature representation based on the first intermediate feature representation includes: determining a second set of inputs based on the first feature representation, the second feature representation, and the first intermediate feature representation; processing the second set of inputs using a second attention model to determine a second intermediate feature representation; and determining the aggregated feature representation based on the second intermediate feature representation.
[0085] In some implementations, the parameters of the attention model are determined based on training a binding analysis model, the binding analysis model including the at least one attention model and a prediction model configured to output the evaluation information based on the aggregated feature representation.
[0086] In some implementations, determining the first set of inputs includes: updating the first feature representation using a multi-head attention mechanism based on edge features in the first feature representation corresponding to the edge and node features of the pair of nodes corresponding to the edge; and determining the first set of inputs based on the updated first feature representation.
[0087] In some implementations, the complex structure is a target complex structure, and the method further includes: determining a binding pocket of the protein molecule; determining a first set of candidate complex structures based on the binding pocket of the protein molecule; and determining the target complex structure from the set of candidate complex structures based on binding free energies corresponding to the first set of candidate complex structures, wherein determining the evaluation information regarding the binding between the ligand molecule and the protein molecule includes determining effectiveness of the binding between the ligand molecule and the protein molecule.
[0088] In some implementations, the ligand molecule is a target ligand molecule from a set of candidate ligand molecules, and the method further includes: determining a first set of evaluation information indicating binding effectiveness between the protein molecule and the set of candidate ligand molecules; and determining the target ligand molecule from the set of candidate ligand molecules based on the binding effectiveness indicated by the first set of evaluation information.
[0089] In some implementations, the complex structure is a target complex structure from a second set of candidate complex structures, the second set of candidate complex structures corresponding to a set of candidate binding poses, and the method further includes: determining a second set of evaluation information corresponding to the second set of candidate complex structures, the second set of evaluation information indicating affinity of the set of candidate binding poses; and determining a target binding pose from the set of candidate binding poses based on the second set of evaluation information.
[0090] In a second aspect of the disclosure, a device is provided. The device includes: a processing unit; and a memory coupled to the processing unit and containing instructions stored thereon that, when executed by the processing unit, cause the device to perform the following actions: obtaining a first feature representation of a ligand molecule and a second feature representation of a protein molecule, the first feature representation being determined based on a structure of the ligand molecule, the second feature representation being determined based on a structure of the protein molecule; determining a third feature representation of a complex structure, the complex structure being constructed based on the protein molecule and the ligand molecule; generating an aggregated feature representation based on the first feature representation, the second feature representation, and the third feature representation; and determining, based on the aggregated feature representation, evaluation information (which can also be referred to as prediction information) about a binding between the ligand molecule and the protein molecule, the evaluation information indicating an effectiveness of the binding or indicating an affinity of a binding pose of the binding.
[0091] In some implementations, at least one of the first feature representation, the second feature representation, and the third feature representation is a feature graph including node features and edge features, the node features being used to characterize atoms in a molecule, the edge features being used to characterize binding relationships between the atoms.
[0092] In some implementations, generating the aggregated feature representation based on the first feature representation, the second feature representation, and the third feature representation includes: determining a first set of inputs based on the first feature representation, the second feature representation, and the third feature representation; processing the first set of inputs using an attention model to determine a first intermediate feature representation; and determining the aggregated feature representation based on the first intermediate feature representation.
[0093] In some implementations, the attention model is a first attention model, and determining the aggregated feature representation based on the first intermediate feature representation includes: determining a second set of inputs based on the first feature representation, the second feature representation, and the first intermediate feature representation; processing the second set of inputs using a second attention model to determine a second intermediate feature representation; and determining the aggregated feature representation based on the second intermediate feature representation.
[0094] In some implementations, parameters of the attention model are determined based on training a binding analysis model, the binding analysis model including at least the attention model and a prediction model, the prediction model being used to output the evaluation information based on the aggregated feature representation.
[0095] In some implementations, determining the first set of inputs includes: updating the first feature representation based on edge features in the first feature representation corresponding to edge pairs and node features of pairs of nodes corresponding to the edge pairs using a multi-head attention mechanism; and determining the first set of inputs based on the updated first feature representation.
[0096] In some implementations, the complex structure is a target complex structure, and the actions further include: determining a binding pocket of the protein molecule; determining a first set of candidate complex structures based on the binding pocket of the protein molecule; and determining the target complex structure from the set of candidate complex structures based on binding free energies corresponding to the first set of candidate complex structures, wherein determining the evaluation information about the binding between the ligand molecule and the protein molecule includes determining effectiveness of the binding between the ligand molecule and the protein molecule.
[0097] In some implementations, the ligand molecule is a target ligand molecule from a set of candidate ligand molecules, and the actions further include: determining a first set of evaluation information indicating binding effectiveness between the protein molecule and the set of candidate ligand molecules; and determining the target ligand molecule from the set of candidate ligand molecules based on the binding effectiveness indicated by the first set of evaluation information.
[0098] In some implementations, the complex structure is a target complex structure from a second set of candidate complex structures, the second set of candidate complex structures corresponding to a set of candidate binding poses, and the actions further include: determining a second set of evaluation information corresponding to the second set of candidate complex structures, the second set of evaluation information indicating affinity of the set of candidate binding poses; and determining a target binding pose from the set of candidate binding poses based on the second set of evaluation information.
[0099] In a third aspect of the present disclosure, a computer program product is provided. The computer program product is tangibly stored in a non-transitory computer storage medium and includes machine executable instructions that, when executed by a device, cause the device to perform the following actions: obtaining a first feature representation of a ligand molecule and a second feature representation of a protein molecule, the first feature representation being determined based on a structure of the ligand molecule, the second feature representation being determined based on a structure of the protein molecule; determining a third feature representation of a complex structure, the complex structure being constructed based on the protein molecule and the ligand molecule; generating an aggregated feature representation based on the first feature representation, the second feature representation, and the third feature representation; and determining evaluation information (which can also be referred to as prediction information) about a binding between the ligand molecule and the protein molecule based on the aggregated feature representation, the evaluation information indicating effectiveness of the binding or indicating affinity of a binding pose of the binding.
[0100] In some implementations, at least one of the first feature representation, the second feature representation, and the third feature representation is a feature graph including node features and edge features, the node features being used to characterize atoms in the molecules, the edge features being used to characterize binding relationships between the atoms.
[0101] In some implementations, generating an aggregated feature representation based on a first feature representation, a second feature representation, and a third feature representation includes: determining a first set of inputs based on the first feature representation, the second feature representation, and the third feature representation; processing the first set of inputs using an attention model to determine a first intermediate feature representation; and determining the aggregated feature representation based on the first intermediate feature representation.
[0102] In some implementations, the attention model is a first attention model, and determining the aggregated feature representation based on the first intermediate feature representation includes: determining a second set of inputs based on the first feature representation, the second feature representation, and the first intermediate feature representation; processing the second set of inputs using a second attention model to determine a second intermediate feature representation; and determining the aggregated feature representation based on the second intermediate feature representation.
[0103] In some implementations, the parameters of the attention model are determined based on training a combined analysis model, which includes at least one attention model and a prediction model used to output evaluation information based on aggregated feature representations.
[0104] In some implementations, determining the first set of inputs includes: using a multi-head attention mechanism to update the first feature representation based on the edge features corresponding to the edge and the node features of a pair of nodes corresponding to the edge in the first feature representation; and determining the first set of inputs based on the updated first feature representation.
[0105] In some implementations, the complex structure is the target complex structure, and the actions further include: determining the binding pocket of the protein molecule; determining a first set of candidate complex structures based on the binding pocket of the protein molecule; and determining the target complex structure from a set of candidate complex structures based on the binding free energy corresponding to the first set of candidate complex structures, wherein determining evaluation information regarding the binding between the ligand molecule and the protein molecule includes: determining the effectiveness of the binding between the ligand molecule and the protein molecule.
[0106] In some implementations, the ligand molecule is one of a set of candidate ligand molecules, and the action further includes: determining a first set of evaluation information indicating the binding effectiveness between the protein molecule and the set of candidate ligand molecules; and determining the target ligand molecule from the set of candidate ligand molecules based on the binding effectiveness indicated by the first set of evaluation information.
[0107] In some implementations, the composite structure is one of the candidate composite structures in a second set of candidate composite structures, which corresponds to a set of candidate bonding poses. The action also includes: determining a second set of evaluation information corresponding to the second set of candidate composite structures, which indicates the affinity of the set of candidate bonding poses; and determining the target bonding pose from the set of candidate bonding poses based on the second set of evaluation information.
[0108] The functionality described above in this document can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, non- limiting example types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip (SOCs), Complex Programmable Logic Devices (CPLDs), etc.
[0109] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, causes the machine to perform the functions / acts specified in the flowcharts and / or block diagrams. The program code can execute entirely on a machine, partly on a machine, as a stand-alone software package, partly on a machine and partly on a remote machine or entirely on a remote machine or server.
[0110] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0111] Further, while operations are depicted in a particular order, this should not be understood as requiring such an order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing can be advantageous. Likewise, while several specific implementation details are contained in the above discussion, these should not be construed as limiting the scope of the disclosure in any way. Certain features that are described in the context of separate implementations can also be implemented in combination in a single implementation. Conversely, various features that are described in the context of a single implementation can also be implemented in subcombination or as separate implementations.
[0112] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.
Claims
1. A method for molecular binding analysis, comprising: A first feature representation of a ligand molecule and a second feature representation of a protein molecule are obtained, wherein the first feature representation is determined based on the structure of the ligand molecule and the second feature representation is determined based on the structure of the protein molecule, wherein the first feature representation includes a first feature map, which includes first node features characterizing atoms in the ligand molecule and first edge features characterizing the binding relationship between atoms, wherein the second feature representation includes a second feature map, which includes second node features characterizing atoms in the protein molecule and second edge features characterizing the binding relationship between atoms; A third feature representation of the composite structure is determined, the composite structure being constructed based on the protein molecule and the ligand molecule, wherein the third feature representation includes a third feature map, which includes third node features characterizing the atoms in the composite structure and third edge features characterizing the bonding relationships between the atoms; Based on the first feature representation, the second feature representation, and the third feature representation, an aggregated feature representation is generated; as well as Based on the aggregation feature representation, evaluation information regarding the binding between the ligand molecule and the protein molecule is determined. This evaluation information indicates the effectiveness of the binding or the affinity of the binding posture. The evaluation information is output by a prediction model based on the aggregation feature representation.
2. The method according to claim 1, wherein generating the aggregated feature representation based on the first feature representation, the second feature representation, and the third feature representation comprises: The first set of inputs is determined based on the first feature representation, the second feature representation, and the third feature representation; The first set of inputs is processed using an attention model to determine a first intermediate feature representation; and The aggregated feature representation is determined based on the first intermediate feature representation.
3. The method according to claim 2, wherein the attention model is a first attention model, and determining the aggregated feature representation based on the first intermediate feature representation includes: Based on the first feature representation, the second feature representation, and the first intermediate feature representation, a second set of inputs is determined; The second set of inputs is processed using a second attention model to determine the second intermediate feature representation; and The aggregated feature representation is determined based on the second intermediate feature representation.
4. The method of claim 2, wherein the parameters of the attention model are determined based on training a combined analysis model, the combined analysis model comprising at least one attention model and the prediction model.
5. The method of claim 2, wherein determining the first set of inputs comprises: Using a multi-head attention mechanism, the first feature representation is updated based on the edge features corresponding to the edge and the node features of a pair of nodes corresponding to the edge in the first feature representation; as well as The first set of inputs is determined based on the updated first feature representation.
6. The method according to claim 1, wherein the composite structure is a target composite structure, and the method further comprises: Determine the binding pocket of the protein molecule; Based on the binding pocket of the protein molecule, a first group of candidate complex structures is determined; as well as Based on the binding free energy corresponding to the first group of candidate composite structures, the target composite structure is determined from the group of candidate composite structures. The evaluation information determining the binding between the ligand molecule and the protein molecule includes: To determine the effectiveness of the binding between the ligand molecule and the protein molecule, and Determining the target composite structure from the group of candidate composite structures based on the binding free energy corresponding to the first group of candidate composite structures includes: selecting the candidate composite structure with the smallest binding free energy from the group of candidate composite structures as the target composite structure.
7. The method of claim 1, wherein the ligand molecule is one of a group of candidate ligand molecules, and the method further comprises: Determine a first set of evaluation information indicating the binding effectiveness between the protein molecule and the set of candidate ligand molecules; as well as Based on the binding effectiveness indicated by the first set of evaluation information, the target ligand molecule is determined from the set of candidate ligand molecules.
8. The method according to claim 1, wherein the composite structure is one of the candidate composite structures in a second group of candidate composite structures, the second group of candidate composite structures corresponding to a group of candidate bonding postures, and the method further comprises: A second set of evaluation information is determined corresponding to the second set of candidate composite structures, wherein the second set of evaluation information indicates the affinity of the candidate bonding postures. as well as Based on the second set of evaluation information, the target combination posture is determined from the set of candidate combination postures.
9. An apparatus comprising: Processing unit; as well as A memory, coupled to the processing unit and containing instructions stored thereon, which, when executed by the processing unit, cause the device to perform the following actions: A first feature representation of a ligand molecule and a second feature representation of a protein molecule are obtained. The first feature representation is determined based on the structure of the ligand molecule, and the second feature representation is determined based on the structure of the protein molecule. The first feature representation includes a first feature map, which includes first node features characterizing atoms in the ligand molecule and first edge features characterizing the binding relationships between atoms. The second feature representation includes a second feature map, which includes second node features characterizing atoms in the protein molecule and second edge features characterizing the binding relationships between atoms; A third feature representation of the composite structure is determined, the composite structure being constructed based on the protein molecule and the ligand molecule, wherein the third feature representation includes a third feature map, which includes third node features characterizing the atoms in the composite structure and third edge features characterizing the bonding relationships between the atoms; Based on the first feature representation, the second feature representation, and the third feature representation, an aggregated feature representation is generated; as well as Based on the aggregation feature representation, evaluation information regarding the binding between the ligand molecule and the protein molecule is determined. This evaluation information indicates the effectiveness of the binding or the affinity of the binding posture. The evaluation information is output by a prediction model based on the aggregation feature representation.
10. The device of claim 9, wherein generating the aggregated feature representation based on the first feature representation, the second feature representation, and the third feature representation comprises: The first set of inputs is determined based on the first feature representation, the second feature representation, and the third feature representation; The first set of inputs is processed using an attention model to determine a first intermediate feature representation; and The aggregated feature representation is determined based on the first intermediate feature representation.
11. The device of claim 10, wherein the attention model is a first attention model, and determining the aggregated feature representation based on the first intermediate feature representation comprises: Based on the first feature representation, the second feature representation, and the first intermediate feature representation, a second set of inputs is determined; The second set of inputs is processed using a second attention model to determine the second intermediate feature representation; and The aggregated feature representation is determined based on the second intermediate feature representation.
12. The device of claim 10, wherein the parameters of the attention model are determined based on training a combined analysis model, the combined analysis model comprising at least one attention model and the prediction model.
13. The device of claim 10, wherein determining the first set of inputs includes: Using a multi-head attention mechanism, the first feature representation is updated based on the edge features corresponding to the edge and the node features of a pair of nodes corresponding to the edge in the first feature representation; as well as The first set of inputs is determined based on the updated first feature representation.
14. The device according to claim 9, wherein the composite structure is a target composite structure, and the action further includes: Determine the binding pocket of the protein molecule; Based on the binding pocket of the protein molecule, a first group of candidate complex structures is determined; as well as Based on the binding free energy corresponding to the first group of candidate composite structures, the target composite structure is determined from the group of candidate composite structures. The evaluation information determining the binding between the ligand molecule and the protein molecule includes: To determine the effectiveness of the binding between the ligand molecule and the protein molecule, and Determining the target composite structure from the group of candidate composite structures based on the binding free energy corresponding to the first group of candidate composite structures includes: selecting the candidate composite structure with the smallest binding free energy from the group of candidate composite structures as the target composite structure.
15. The device of claim 9, wherein the ligand molecule is one of a group of candidate ligand molecules, and the action further includes: Determine a first set of evaluation information indicating the binding effectiveness between the protein molecule and the set of candidate ligand molecules; as well as Based on the binding effectiveness indicated by the first set of evaluation information, the target ligand molecule is determined from the set of candidate ligand molecules.
16. The device according to claim 9, wherein the composite structure is one of the candidate composite structures in a second group of candidate composite structures, the second group of candidate composite structures corresponding to a group of candidate combination postures, and the action further includes: A second set of evaluation information is determined corresponding to the second set of candidate composite structures, wherein the second set of evaluation information indicates the affinity of the candidate bonding postures. as well as Based on the second set of evaluation information, the target combination posture is determined from the set of candidate combination postures.
17. A computer program product tangibly stored in a non-transitory computer storage medium and comprising machine-executable instructions that, when executed by a device, cause the device to perform the following actions: A first feature representation of a ligand molecule and a second feature representation of a protein molecule are obtained, wherein the first feature representation is determined based on the structure of the ligand molecule and the second feature representation is determined based on the structure of the protein molecule, wherein the first feature representation includes a first feature map, which includes first node features characterizing atoms in the ligand molecule and first edge features characterizing the binding relationship between atoms, wherein the second feature representation includes a second feature map, which includes second node features characterizing atoms in the protein molecule and second edge features characterizing the binding relationship between atoms; A third feature representation of the composite structure is determined, the composite structure being constructed based on the protein molecule and the ligand molecule, wherein the third feature representation includes a third feature map, which includes third node features characterizing the atoms in the composite structure and third edge features characterizing the bonding relationships between the atoms; Based on the first feature representation, the second feature representation, and the third feature representation, an aggregated feature representation is generated; as well as Based on the aggregation feature representation, evaluation information regarding the binding between the ligand molecule and the protein molecule is determined. This evaluation information indicates the effectiveness of the binding or the affinity of the binding posture. The evaluation information is output by a prediction model based on the aggregation feature representation.
18. The computer program product of claim 17, wherein generating the aggregated feature representation based on the first feature representation, the second feature representation, and the third feature representation comprises: The first set of inputs is determined based on the first feature representation, the second feature representation, and the third feature representation; The first set of inputs is processed using an attention model to determine a first intermediate feature representation; and The aggregated feature representation is determined based on the first intermediate feature representation.
19. The computer program product of claim 18, wherein the attention model is a first attention model, and determining the aggregated feature representation based on the first intermediate feature representation comprises: Based on the first feature representation, the second feature representation, and the first intermediate feature representation, a second set of inputs is determined; The second set of inputs is processed using a second attention model to determine the second intermediate feature representation; and The aggregated feature representation is determined based on the second intermediate feature representation.
20. The computer program product of claim 18, wherein the parameters of the attention model are determined based on training a combined analysis model, the combined analysis model comprising at least one attention model and the prediction model.
21. The computer program product of claim 18, wherein determining the first set of inputs includes: Using a multi-head attention mechanism, the first feature representation is updated based on the edge features corresponding to the edge and the node features of a pair of nodes corresponding to the edge in the first feature representation; as well as The first set of inputs is determined based on the updated first feature representation.
22. The computer program product of claim 17, wherein the composite structure is a target composite structure, and the action further includes: Determine the binding pocket of the protein molecule; Based on the binding pocket of the protein molecule, a first group of candidate complex structures is determined; as well as Based on the binding free energy corresponding to the first group of candidate composite structures, the target composite structure is determined from the group of candidate composite structures. The evaluation information determining the binding between the ligand molecule and the protein molecule includes: To determine the effectiveness of the binding between the ligand molecule and the protein molecule, and Determining the target composite structure from the group of candidate composite structures based on the binding free energy corresponding to the first group of candidate composite structures includes: selecting the candidate composite structure with the smallest binding free energy from the group of candidate composite structures as the target composite structure.
23. The computer program product of claim 17, wherein the ligand molecule is one of a group of candidate ligand molecules, and the action further includes: Determine a first set of evaluation information indicating the binding effectiveness between the protein molecule and the set of candidate ligand molecules; as well as Based on the binding effectiveness indicated by the first set of evaluation information, the target ligand molecule is determined from the set of candidate ligand molecules.
24. The computer program product of claim 17, wherein the composite structure is one of the candidate composite structures in a second group of candidate composite structures, the second group of candidate composite structures corresponding to a group of candidate combination postures, and the action further includes: A second set of evaluation information is determined corresponding to the second set of candidate composite structures, wherein the second set of evaluation information indicates the affinity of the candidate bonding postures. as well as Based on the second set of evaluation information, the target combination posture is determined from the set of candidate combination postures.
Citation Information
Patent Citations
Method and device for predicting binding free energy of protein and ligand molecules
CN112466410A
Method and apparatus for training predictive model for determining molecular binding force
CN113241126A