Protein docking method, device, electronic device and storage medium
Through multiple iterative protein docking methods, combined with rigid and flexible docking and graph neural network, the problem of low protein docking accuracy is solved, and more accurate protein complex conformation prediction is achieved, and it is applied to multiple biological research fields.
Patent Information
- Application Number
- CN202311443053.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-01
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2043-11-01
AI Technical Summary
The protein docking methods in the prior art have low accuracy and are difficult to effectively predict the conformation of the protein complex.
A multi-round iteration protein docking method is used to combine at least two molecular docking methods, such as rigid and flexible docking, and a complex conformation is generated through the graph neural network until the iteration end condition is met.
It improves the accuracy of protein docking and can more accurately predict the conformation of protein complexes. It is used in fields such as antigen-antibody docking, polypeptide protein docking, disease mechanism research, protein function research, structural biology, etc.
Smart Images

Figure CN117497041B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to the fields of biological computing, molecular docking, and deep learning technology, and in particular to a protein docking method, device, electronic device, storage medium, and computer program product. Background Art
[0002] Protein docking, the process by which two or more protein molecules interact to form a complex (also called a complex), plays an important role in biology and can include, for example, antigen-antibody docking and peptide-protein docking. However, existing protein docking methods suffer from low docking accuracy. Summary of the Invention
[0003] The present disclosure provides a protein docking method, device, electronic device, storage medium and computer program product.
[0004] According to a first aspect of the present disclosure, a protein docking method is proposed, comprising: in each round of iteration, docking a first protein and a second protein according to at least two molecular docking methods to generate a complex conformation; identifying that an iteration end condition is not met, continuing the next round of iteration until the iteration end condition is met, and obtaining a final complex conformation.
[0005] According to a second aspect of the present disclosure, a protein docking device is proposed, comprising: a docking module for docking a first protein and a second protein according to at least two molecular docking methods in each round of iteration to generate a complex conformation; and a processing module for identifying when an iteration end condition is not met, continuing the next round of iteration until the iteration end condition is met, and obtaining a final complex conformation.
[0006] According to a third aspect of the present disclosure, an electronic device is proposed, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the protein docking method proposed in the first aspect above.
[0007] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is proposed, wherein the computer instructions are used to enable the computer to execute the protein docking method proposed in the first aspect.
[0008] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the protein docking method provided in the first aspect.
[0009] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0011] Figure 1 Schematic diagram of the process of a protein docking method according to an embodiment of the present disclosure;
[0012] Figure 2 Schematic diagram of a process for docking a protein according to another embodiment of the present disclosure;
[0013] Figure 3 Schematic diagram of a process for docking a protein according to another embodiment of the present disclosure;
[0014] Figure 4 Schematic diagram of a process for docking a protein according to another embodiment of the present disclosure;
[0015] Figure 5 Schematic diagram of the structure of a protein docking device according to one embodiment of the present disclosure;
[0016] Figure 6 A schematic block diagram of an electronic device according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0017] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0018] Artificial Intelligence (AI) is a discipline that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. Currently, AI technology has been widely used due to its high degree of automation, high precision, and low cost.
[0019] Biocomputing refers to computational models that use biological macromolecules as "data". It is mainly divided into three types: protein computing, RNA (Ribonucleic Acid) computing, and DNA (DeoxyriboNucleic Acid) computing. It also refers to a subfield of computer science and computer engineering that uses bioengineering and biology to build computers. However, similar to bioinformatics, it is an interdisciplinary science that uses computers to store and process biological data.
[0020] Molecular docking is a method for drug design that leverages the characteristics of receptors and the interactions between them and drug molecules. It is a theoretical simulation method that primarily studies interactions between molecules (such as ligands and receptors) and predicts their binding patterns and affinities. In recent years, molecular docking has become a key technology in computer-assisted drug discovery.
[0021] DL (Deep Learning) is a new research direction in the field of ML (Machine Learning). It is a science that learns the inherent laws and representation levels of sample data, enabling machines to have analytical and learning capabilities like humans and recognize data such as text, images, and sounds. It is widely used in speech and image recognition.
[0022] Figure 1 FIG. 1 is a flow chart of a protein docking method according to an embodiment of the present disclosure. Figure 1 As shown, the method includes:
[0023] S101, in each round of iteration, docking the first protein and the second protein according to at least two molecular docking methods to generate a complex conformation.
[0024] It should be noted that the protein docking method of the present embodiment can be executed by a hardware device with data processing capabilities and / or the necessary software to drive the operation of the hardware device. Alternatively, the execution entity may include a workstation, server, computer, user terminal, and other intelligent device. User terminals include, but are not limited to, mobile phones, computers, intelligent voice interaction devices, smart appliances, and in-vehicle terminals.
[0025] It should be noted that the at least two molecular docking methods are not particularly limited and may include, for example, rigid docking, flexible docking, and semi-flexible docking. During rigid docking, neither the monomeric conformation of the first protein nor the monomeric conformation of the second protein changes. During flexible docking, at least one of the monomeric conformation of the first protein and the monomeric conformation of the second protein changes.
[0026] It should be noted that there are no excessive restrictions on the execution order of at least two molecular docking methods. For example, taking the at least two molecular docking methods including rigid docking and flexible docking as an example, the first protein and the second protein can be docked in the order of rigid docking and flexible docking to generate a complex conformation, or the first protein and the second protein can be docked in the order of flexible docking and rigid docking to generate a complex conformation.
[0027] It can be understood that each molecular docking method can generate a complex conformation. If a complex conformation does not currently exist, a complex conformation can be generated according to a molecular docking method. If a complex conformation currently exists, a complex conformation can be regenerated according to a molecular docking method, that is, the complex conformation is updated.
[0028] For example, in each round of iteration, the first protein and the second protein are docked according to the rigid docking method to generate a complex conformation, and the complex conformation is updated according to the flexible docking method.
[0029] For example, in each round of iteration, the first protein and the second protein are docked according to the flexible docking method to generate a complex conformation, and the complex conformation is updated according to the rigid docking method.
[0030] S102, identifying that the iteration end condition is not met, continuing the next round of iteration process until the iteration end condition is met, and obtaining the final complex conformation.
[0031] It should be noted that there are no specific restrictions on the conditions for ending the iterations. For example, the conditions may include the number of iterations reaching a set threshold, the parameters of the complex conformation being within a set range, and the successful docking of the first and second proteins. The parameters of the complex conformation are used to characterize the docking accuracy.
[0032] In one embodiment, obtaining the final complex conformation includes taking the complex conformation obtained for the last time as the final complex conformation.
[0033] In one embodiment, obtaining the final complex conformation includes screening the final complex conformation from multiple complex conformations generated by multiple rounds of iterative processes.
[0034] It should be noted that the protein docking method proposed in this disclosure is applicable to at least one of the following scenarios:
[0035] Scenario 1: Antigen-antibody docking: The present disclosure can be used to predict the conformation of antigen-antibody complexes, thereby assisting in antibody design.
[0036] Scenario 2, peptide-protein docking: The present disclosure can be used to predict the conformation of peptide-protein complexes, thereby assisting in peptide drug design.
[0037] Scenario 3: Disease Mechanism Research: The occurrence and development of many diseases are related to abnormal interactions between proteins. Protein docking can help researchers understand the molecular mechanisms of these abnormal interactions, thereby providing new ideas for disease diagnosis and treatment.
[0038] Scenario 4: Protein Function Research: Protein docking can reveal interactions between proteins, helping scientists understand protein functions and regulatory mechanisms. This is crucial for uncovering biological processes such as cell signaling and gene regulation.
[0039] Scenario 5, Structural Biology: In protein structural biology, protein docking can be used to predict the binding patterns of protein complexes, thereby helping researchers explain the structure and function of protein complexes.
[0040] Scenario 6: Disease Mechanism Research: The occurrence and development of many diseases are related to abnormal interactions between proteins. Protein docking can help researchers understand the molecular mechanisms of these abnormal interactions, thereby providing new insights into disease diagnosis and treatment.
[0041] Scenario 7: Protein interaction network analysis: By predicting the interactions between proteins, a protein interaction network can be constructed to help researchers understand the interconnections and regulatory networks of proteins within cells.
[0042] Scenario 8, Protein Engineering: In the field of biotechnology, protein docking can be used to design new protein structures, such as fusion proteins, antibodies, enzymes, etc., to achieve specific functions.
[0043] The protein docking method proposed in this disclosure docks a first protein and a second protein using at least two molecular docking methods in each iterative process, generating a complex conformation. If an iterative termination criterion is not met, the next iterative process continues until the criterion is met, resulting in a final complex conformation. This allows for protein docking to be achieved by comprehensively considering at least two molecular docking methods, and multiple rounds of protein docking, i.e., multiple iterations of complex conformation, can be performed, thereby improving the accuracy of protein docking.
[0044] Based on any of the above embodiments, a target map can be constructed, and the target map is used to generate a conformation image of the complex.
[0045] Regarding building the target graph, you can combine Figure 2 Further understanding, Figure 2 FIG. 1 is a flow chart of a protein docking method according to another embodiment of the present disclosure, as shown in FIG. Figure 2 As shown, the method includes:
[0046] S201, obtaining information on the residues of the first protein based on the monomer conformation of the first protein.
[0047] S202, obtaining information about the residues of the second protein based on the monomer conformation of the second protein.
[0048] It should be noted that there are no excessive restrictions on the information of the residue. For example, it may include the type of residue and the position of the residue (such as the position of the α-carbon atom of the residue), where the α-carbon atom refers to the first carbon atom of the residue.
[0049] It can be understood that the monomer conformation of a protein carries information about the residues of the protein, and the information about the residues of the protein can be extracted from the monomer conformation of the protein.
[0050] For example, information about the residues of the first protein can be extracted from the monomer conformation of the first protein.
[0051] S203, constructing a target graph based on the information of the residues of the first protein and the information of the residues of the second protein, wherein the nodes in the target graph are used to represent the residues of the first protein or the residues of the second protein, and the edges in the target graph are used to represent the edges between the two residues, and the target graph is used to generate a conformation of the complex.
[0052] It should be noted that the target graph is a heterogeneous graph, consisting of two types of nodes: the first type of node represents residues of the first protein, and the second type of node represents residues of the second protein. Nodes in the target graph correspond one-to-one with residues. Edges in the target graph represent the edges between residues corresponding to the two nodes connected by the edge.
[0053] It can be understood that the number of nodes in the target graph is the sum of the number of residues in the first protein and the number of residues in the second protein.
[0054] It is understandable that there may or may not be an edge between any two nodes in the target graph.
[0055] In the embodiments of the present disclosure, a target graph is constructed based on the information of the residues of the first protein and the residues of the second protein, including the following possible implementations:
[0056] Method 1: constructing nodes corresponding one-to-one to residues to generate a target graph, wherein the residues include residues of the first protein and residues of the second protein.
[0057] Therefore, in this method, nodes corresponding one to one with residues can be constructed to achieve node construction.
[0058] Method 2: Based on information of the first residue and information of the second residue, determine the distance between the feature of the first residue and the feature of the second residue, wherein the first residue and the second residue belong to a target protein, wherein the target protein is the first protein or the second protein, sort the multiple second residues in ascending order according to the distance, and use the top N second residues in the sort as target residues, wherein N is a positive integer, and add connecting edges between the nodes corresponding to the first residue and the nodes corresponding to the target residue to generate a target graph.
[0059] Therefore, in this method, based on the information of the first residue and the information of the second residue, the distance between the characteristics of the first residue and the second residue can be determined, and the N second residues closest to the first residue can be used as target residues. Connecting edges can be added between the nodes corresponding to the first residue and the nodes corresponding to the target residues to realize the construction of edges between residues of the same protein.
[0060] It should be noted that the first residue and the second residue belong to the target protein, that is, the first residue and the second residue belong to the same protein, for example, the first residue and the second residue belong to the first protein, or the first residue and the second residue belong to the second protein.
[0061] It should be noted that there is no excessive limitation on N, for example, it can be 10.
[0062] In one embodiment, based on the information of the first residue and the information of the second residue, the distance between the feature of the first residue and the feature of the second residue is determined, including performing feature extraction on the information of the first residue to obtain the feature of the first residue, performing feature extraction on the information of the second residue to obtain the feature of the second residue, and obtaining the distance between the feature of the first residue and the feature of the second residue.
[0063] In one embodiment, the information of the first residue and the information of the second residue can be input into a KNN (K-Nearest Neighbor) model, and the KNN model outputs the target residue.
[0064] It should be noted that there are not too many restrictions on distance, for example, it can include Euclidean distance, Manhattan distance, etc.
[0065] Method 3: The residue of the first protein is used as the third residue, and the residue of the second protein is used as the fourth residue. A connecting edge is added between the node corresponding to the third residue and the node corresponding to the fourth residue to generate the target graph.
[0066] Therefore, in this method, a connecting edge can be added between the node corresponding to the third residue and the node corresponding to the fourth residue to achieve the construction of edges between residues of different proteins.
[0067] It should be noted that the third residue is any residue of the first protein, and the fourth residue is any residue of the second protein.
[0068] Method 4: Determine the characteristics of the node based on the information of the residue corresponding to the node.
[0069] It should be noted that there are no excessive restrictions on the characteristics of the nodes. For example, they may include the category of the residue corresponding to the node, the position of the residue corresponding to the node, etc.
[0070] Method 5: Determine the characteristics of the edge based on the information of the two residues corresponding to the edge.
[0071] It should be noted that there are not too many restrictions on the characteristics of the edge. For example, it may include the length of the edge, the angle difference between the two residues corresponding to the edge, etc.
[0072] In one embodiment, the feature of the edge is determined based on information of the two residues corresponding to the edge, including determining the distance between the two residues corresponding to the edge as the length of the edge based on the positions of the two residues corresponding to the edge.
[0073] In one embodiment, determining the characteristics of the edge based on information of the two residues corresponding to the edge includes obtaining the difference in angles between the two residues corresponding to the edge as the angle difference between the two residues corresponding to the edge.
[0074] Based on any of the above embodiments, the method further includes updating the residue information of the first protein and the residue information of the second protein based on the most recently obtained complex conformation, and returning to the step of constructing a target graph based on the residue information of the first protein and the residue information of the second protein to update the target graph. In this way, the residue information can be updated using the most recent complex conformation, and the target graph can be reconstructed using the updated residue information, thereby achieving real-time updating of the target graph during the protein docking process.
[0075] It can be understood that the complex conformation carries information about the residues of the first protein and the residues of the second protein. The information about the residues of the first protein and the residues of the second protein can be extracted from the most recently obtained complex conformation, and the original information about the residues of the first protein can be replaced with the extracted information about the residues of the first protein, and the original information about the residues of the second protein can be replaced with the extracted information about the residues of the second protein.
[0076] It is understandable that updating the target graph may include deleting or adding edges between nodes, which is not limited here.
[0077] In some examples, after updating the information about the residues of the first protein and the residues of the second protein, the features of the nodes and edges in the target graph are also updated based on the information about the residues of the first protein and the residues of the second protein. The features of the nodes are determined based on the information about the residues corresponding to the nodes, and the features of the edges are determined based on the information about the two residues corresponding to the edge. Thus, the updated residue information can be used to update the features of the nodes and edges in the target graph, enabling real-time updates of the node and edge features during the docking process.
[0078] The protein docking method proposed in the present disclosure can obtain information about the residues of the first protein based on the monomer conformation of the first protein, obtain information about the residues of the second protein based on the monomer conformation of the second protein, and construct a target graph based on the information about the residues of the first protein and the information about the residues of the second protein to generate a complex conformation.
[0079] In the above embodiment, at least two molecular docking methods include rigid docking. In step S102, the first protein and the second protein are docked according to at least two molecular docking methods to generate a complex conformation. Figure 3 Further understanding, Figure 3 FIG. 1 is a flow chart of a protein docking method according to another embodiment of the present disclosure, as shown in FIG. Figure 3 As shown, the method includes:
[0080] S301, input the target graph into the first graph neural network, and obtain the position of the key point in the complex conformation based on the target graph through the first graph neural network, wherein the key point is the position point on the contact surface of the first protein and the second protein in the complex conformation.
[0081] It should be noted that during the process of obtaining the position of the key point in the complex conformation, the monomer conformation of the first protein and the monomer conformation of the second protein did not change.
[0082] It should be noted that there are not too many restrictions on the first graph neural network. For example, it can include GCN (Graph Convolutional Network), GRN (Graph Recurrent Network), GAT (Graph Attention Network), etc.
[0083] It should be noted that there is no excessive restriction on the number of key points.
[0084] In one embodiment, the position of the key point in the complex conformation is obtained based on the target graph through the first graph neural network, including updating the features of the nodes and the features of the edges in the target graph through the first graph neural network, and obtaining the position of the key point in the complex conformation based on the features of the nodes and the features of the edges through the first graph neural network, so as to achieve the acquisition of the position of the key point in the complex conformation.
[0085] It should be noted that the features of the nodes and edges in the target graph are updated by the first graph neural network, and this can be achieved by using any graph neural network feature updating method in the relevant technology, without further limitation here.
[0086] In one embodiment, obtaining the position of a key point in the complex conformation based on the target graph using a first graph neural network includes obtaining the position of a residue in the complex conformation based on the target graph using the first graph neural network, obtaining the distance between the third residue and the fourth residue based on the position of the third residue and the position of the fourth residue in the complex conformation, and if the distance between the third residue and the fourth residue is less than a set threshold, using the position of the midpoint of the edge between the third residue and the fourth residue in the complex conformation as the position of the key point in the complex conformation. In this embodiment, if the distance between the third residue and the fourth residue is less than the set threshold, using the midpoint of the edge between the third residue and the fourth residue as the key point.
[0087] S302: Generate a complex conformation based on the position of the key point in the complex conformation.
[0088] In one embodiment, generating a complex conformation based on the position of a key point in the complex conformation includes determining a receptor and a ligand from a first protein and a second protein, obtaining a rotation-translation matrix based on the position of the key point in the monomeric conformation of the ligand and the position of the key point in the complex conformation, performing a global spatial transformation on the monomeric conformation of the ligand based on the rotation-translation matrix, and generating a complex conformation based on the transformed monomeric conformation of the ligand and the monomeric conformation of the receptor. Thus, a rotation-translation matrix can be obtained based on the position of the key point in the complex conformation to perform a global spatial transformation on the monomeric conformation of the ligand, and the complex conformation can be generated by combining the transformed monomeric conformation of the ligand and the monomeric conformation of the receptor.
[0089] It should be noted that performing an overall spatial transformation on the monomeric conformation of the ligand refers to performing an overall rotation and an overall translation on the monomeric conformation of the ligand.
[0090] In some examples, determining a receptor and a ligand from a first protein and a second protein includes determining the first protein as a receptor and the second protein as a ligand, or determining the second protein as a receptor and the first protein as a ligand.
[0091] In some examples, a rotation and translation matrix is obtained based on the positions of key points in the monomeric conformation of the ligand and the positions of key points in the complex conformation, including generating a first matrix based on the positions of multiple key points in the monomeric conformation of the ligand, generating a second matrix based on the positions of multiple key points in the complex conformation, the second matrix being the product of the rotation and translation matrix and the first matrix, and performing matrix decomposition on the second matrix based on the first matrix to obtain the rotation and translation matrix. It should be noted that matrix decomposition can be implemented using any matrix decomposition method in the relevant art, for example, it can include SVD (Singular Value Decomposition).
[0092] In some examples, a rotation-translation matrix is obtained based on the position of the key point in the monomeric conformation of the ligand and the position of the key point in the complex conformation, including taking the elements in the rotation-translation matrix as unknowns, constructing a system of equations based on the position of the key point in the monomeric conformation of the ligand and the position of the key point in the complex conformation, and the elements in the rotation-translation matrix, solving the system of equations, and obtaining the solutions of the system of equations as the elements in the rotation-translation matrix to generate the rotation-translation matrix.
[0093] In some examples, the monomeric conformation of the ligand is spatially transformed as a whole based on the rotation and translation matrix, including spatially transforming the positions of the residues of the ligand in the monomeric conformation based on the rotation and translation matrix, and obtaining the transformed monomeric conformation of the ligand based on the positions of multiple residues of the ligand in the monomeric conformation.
[0094] The protein docking method proposed in the present disclosure includes at least two molecular docking methods, including rigid docking, and inputting a target graph into a first graph neural network. The first graph neural network obtains the position of the key point in the complex conformation based on the target graph, wherein the key point is a position point on the contact surface of the first protein and the second protein in the complex conformation. Based on the position of the key point in the complex conformation, a complex conformation is generated to achieve rigid docking of the protein.
[0095] In the above embodiment, at least two molecular docking methods include flexible docking. In step S102, the first protein and the second protein are docked according to at least two molecular docking methods to generate a complex conformation. Figure 4 Further understanding, Figure 4 FIG. 1 is a flow chart of a protein docking method according to another embodiment of the present disclosure, as shown in FIG. Figure 4 As shown, the method includes:
[0096] S401, inputting the target graph into the second graph neural network, and obtaining the position of the residue in the complex conformation based on the target graph through the second graph neural network, wherein, in the process of obtaining the position of the residue in the complex conformation, at least one of the monomer conformation of the first protein and the monomer conformation of the second protein changes.
[0097] It should be noted that the second graph neural network can refer to the relevant content of the first graph neural network, which will not be repeated here.
[0098] It should be noted that the residues include residues of the first protein and residues of the second protein.
[0099] In one embodiment, the position of the residue in the complex conformation is obtained based on the target graph by a second graph neural network, including updating the features of the nodes and the features of the edges in the target graph by the second graph neural network, and obtaining the position of the residue in the complex conformation based on the features of the nodes and the features of the edges by the second graph neural network, so as to achieve the acquisition of the position of the residue in the complex conformation.
[0100] It should be noted that the features of the nodes and edges in the target graph are updated by the second graph neural network, and this can be achieved by using any graph neural network feature updating method in the relevant technology, without further limitation here.
[0101] S402: generating a complex conformation based on positions of the plurality of residues in the complex conformation.
[0102] In an embodiment of the present disclosure, a complex conformation is generated based on the positions of multiple residues in the complex conformation, including generating a complex conformation based on the positions of multiple residues of a first protein in the complex conformation and the positions of multiple residues of a second protein in the complex conformation.
[0103] The protein docking method proposed in the present disclosure includes at least two molecular docking methods including flexible docking, wherein a target graph is input into a second graph neural network, and the positions of the residues in the complex conformation are obtained based on the target graph by the second graph neural network, wherein, in the process of obtaining the positions of the residues in the complex conformation, at least one of the monomer conformation of the first protein and the monomer conformation of the second protein changes, and a complex conformation is generated based on the positions of multiple residues in the complex conformation to achieve flexible docking of proteins.
[0104] Based on any of the above embodiments, in each round of iteration, the first protein and the second protein may be docked in the order of rigid docking and flexible docking to generate a complex conformation.
[0105] For example, a target graph can be constructed based on the information of the residues of the first protein and the residues of the second protein, and the target graph can be input into the first graph neural network. The position of the key point in the complex conformation can be obtained through the first graph neural network, and the complex conformation A can be generated based on the position of the key point in the complex conformation.
[0106] Based on the complex conformation A, the information of the residues of the first protein and the information of the residues of the second protein are updated, and the step of constructing the target graph based on the information of the residues of the first protein and the information of the residues of the second protein is returned to update the target graph. The characteristics of the nodes and the characteristics of the edges in the target graph can also be updated based on the information of the residues of the first protein and the information of the residues of the second protein.
[0107] The target graph is input into the second graph neural network, and the second graph neural network obtains the positions of the residues in the complex conformation based on the target graph, and generates the complex conformation B based on the positions of multiple residues in the complex conformation.
[0108] Based on any of the above embodiments, in each round of iteration, the first protein and the second protein may be docked in the order of flexible docking and rigid docking to generate a complex conformation.
[0109] For example, a target graph can be constructed based on the information of the residues of the first protein and the residues of the second protein, and the target graph can be input into the second graph neural network. The positions of the residues in the complex conformation can be obtained through the second graph neural network, and the complex conformation C can be generated based on the positions of multiple residues in the complex conformation.
[0110] Based on the complex conformation C, the information of the residues of the first protein and the information of the residues of the second protein are updated, and the step of constructing a target graph based on the information of the residues of the first protein and the information of the residues of the second protein is returned to update the target graph. The features of the nodes and the features of the edges in the target graph can also be updated based on the information of the residues of the first protein and the information of the residues of the second protein.
[0111] The target graph is input into the first graph neural network, and the first graph neural network obtains the position of the key point in the complex conformation based on the target graph, and generates the complex conformation D based on the position of the key point in the complex conformation.
[0112] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0113] According to an embodiment of the present disclosure, the present disclosure also provides a protein docking device for implementing the above-mentioned protein docking method.
[0114] Figure 5 4 is a block diagram of a protein docking apparatus according to an embodiment of the present disclosure.
[0115] like Figure 5 As shown, the protein docking device 500 includes a docking module 501 and a processing module 502.
[0116] A docking module 501 is configured to dock the first protein with the second protein according to at least two molecular docking methods in each iteration process to generate a complex conformation;
[0117] The processing module 502 is used to identify that the iteration end condition is not met, and continue the next round of iteration process until the iteration end condition is met to obtain the final complex conformation.
[0118] In one embodiment of the present disclosure, the device further includes: a construction module, wherein the construction module is used to: obtain information about the residues of the first protein based on the monomer conformation of the first protein; obtain information about the residues of the second protein based on the monomer conformation of the second protein; and construct a target graph based on the information about the residues of the first protein and the information about the residues of the second protein, wherein the nodes in the target graph are used to represent the residues of the first protein or the residues of the second protein, the edges in the target graph are used to represent the edges between two residues, and the target graph is used to generate the complex conformation.
[0119] In one embodiment of the present disclosure, the at least two molecular docking methods include rigid docking, and the docking module 501 is further used to: input the target graph into a first graph neural network, and obtain the position of the key point in the complex conformation based on the target graph through the first graph neural network, wherein the key point is the position point on the contact surface of the first protein and the second protein in the complex conformation; based on the position of the key point in the complex conformation, generate the complex conformation.
[0120] In one embodiment of the present disclosure, the docking module 501 is further used to: update the features of the nodes and the features of the edges in the target graph through the first graph neural network; and obtain the position of the key point in the complex conformation based on the features of the nodes and the features of the edges through the first graph neural network.
[0121] In one embodiment of the present disclosure, the docking module 501 is further used to: determine the receptor and the ligand from the first protein and the second protein; obtain a rotation-translation matrix based on the position of the key point in the monomer conformation of the ligand and the position of the key point in the complex conformation; perform an overall spatial transformation on the monomer conformation of the ligand based on the rotation-translation matrix; and generate the complex conformation based on the transformed monomer conformation of the ligand and the monomer conformation of the receptor.
[0122] In one embodiment of the present disclosure, the at least two molecular docking methods include flexible docking, and the docking module 501 is further used to: input the target graph into a second graph neural network, and obtain the position of the residue in the complex conformation based on the target graph through the second graph neural network, wherein, in the process of obtaining the position of the residue in the complex conformation, at least one of the monomer conformation of the first protein and the monomer conformation of the second protein changes; and generate the complex conformation based on the positions of multiple residues in the complex conformation.
[0123] In one embodiment of the present disclosure, the construction module is further used to: update the information of the residues of the first protein and the information of the residues of the second protein based on the most recently obtained complex conformation; return to execute the step of constructing the target graph based on the information of the residues of the first protein and the information of the residues of the second protein to update the target graph.
[0124] In one embodiment of the present disclosure, after the information of the residues of the first protein and the information of the residues of the second protein are updated, the construction module is further used to: update the features of the nodes and the features of the edges in the target graph based on the information of the residues of the first protein and the information of the residues of the second protein, wherein the features of the nodes are determined based on the information of the residues corresponding to the nodes, and the features of the edges are determined based on the information of the two residues corresponding to the edges.
[0125] In one embodiment of the present disclosure, the construction module is further used to: determine the distance between the feature of the first residue and the feature of the second residue based on information of the first residue and information of the second residue, wherein the first residue and the second residue belong to a target protein, wherein the target protein is the first protein or the second protein; sort the multiple second residues in ascending order according to the distance, and take the first N second residues in the sort as target residues, wherein N is a positive integer; add connecting edges between the node corresponding to the first residue and the node corresponding to the target residue to generate the target graph.
[0126] In one embodiment of the present disclosure, the building module is further used to: use the residue of the first protein as the third residue; use the residue of the second protein as the fourth residue; and add a connecting edge between the node corresponding to the third residue and the node corresponding to the fourth residue to generate the target graph.
[0127] The protein docking device proposed in this disclosure docks a first protein and a second protein according to at least two molecular docking methods during each iteration, generating a complex conformation. If an iteration termination criterion is not met, the next iteration is continued until the criterion is met, resulting in a final complex conformation. This allows for protein docking by comprehensively considering at least two molecular docking methods, and allows for multiple rounds of protein docking, i.e., multiple iterations of complex conformation generation, thereby improving the accuracy of protein docking.
[0128] According to embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0129] Figure 6 A schematic block diagram of an example electronic device 600 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0130] like Figure 6 As shown, the device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. Various programs and data required for the operation of the device 600 can also be stored in the RAM 603. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0131] Various components in device 600 are connected to I / O interface 605, including an input unit 606, such as a keyboard, mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, optical disk, etc.; and a communication unit 609, such as a network card, modem, wireless communication transceiver, etc. The communication unit 609 allows device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0132] The computing unit 601 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units for running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 601 performs the various methods and processes described above, such as the docking method of the protein. For example, in some embodiments, the docking method of the protein can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by the computing unit 601, one or more steps of the docking method of the protein described above can be performed. Alternatively, in other embodiments, the computing unit 601 can be configured to perform the docking method of the protein in any other appropriate manner (e.g., by means of firmware).
[0133] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0134] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. Such program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0135] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0136] To initiate interaction with a user account, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user account; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user account can provide input to the computer. Other types of devices can also be used to initiate interaction with the user account; for example, feedback provided to the user account can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user account can be received in any form (including acoustic input, voice input, or tactile input).
[0137] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user account computer with a graphical user account interface or a web browser through which a user account can interact with embodiments of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0138] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0139] According to an embodiment of the present disclosure, the present disclosure further provides a computer program product, including a computer program, wherein when the computer program is executed by a processor, the steps of the protein docking method described in the above embodiment of the present disclosure are implemented.
[0140] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.
[0141] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A protein docking method comprising: Based on the monomer conformation of the first protein, obtaining information about the residues of the first protein; Based on the monomer conformation of the second protein, obtaining information about the residues of the second protein; constructing a target graph based on information about the residues of the first protein and information about the residues of the second protein, wherein nodes in the target graph are used to represent residues of the first protein or residues of the second protein, edges in the target graph are used to represent edges between two residues, and the target graph is used to generate a conformation of the complex; In each iteration, docking the first protein and the second protein based on the target graph according to at least two molecular docking methods to generate a complex conformation, wherein the at least two molecular docking methods include rigid docking and flexible docking; Identify that the iteration end condition is not met, and continue the next round of iteration process until the iteration end condition is met to obtain the final complex conformation; For the flexible docking method, the first protein and the second protein are docked to generate a complex conformation, including: Inputting the target graph into a second graph neural network, and obtaining, by the second graph neural network, positions of the residues in the complex conformation based on the target graph, wherein in the process of obtaining the positions of the residues in the complex conformation, at least one of the monomer conformation of the first protein and the monomer conformation of the second protein changes; generating a conformation of the complex based on positions of a plurality of residues in the conformation of the complex; The method further comprises: extracting information about the residues of the first protein and information about the residues of the second protein from the most recently obtained complex conformation, and updating the information about the residues of the first protein and the information about the residues of the second protein; The step of constructing the target graph based on the information of the residues of the first protein and the information of the residues of the second protein is returned to update the target graph.
2. The method according to claim 1, wherein For the rigid docking method, the first protein and the second protein are docked to generate a complex conformation, including: Inputting the target graph into a first graph neural network, and obtaining positions of key points in the complex conformation based on the target graph by the first graph neural network, wherein the key points are positions on the contact surface between the first protein and the second protein in the complex conformation; Based on the positions of the key points in the complex conformation, the complex conformation is generated.
3. The method according to claim 2, wherein: The obtaining, based on the target graph and using the first graph neural network, positions of key points in the complex conformation includes: Updating the features of the nodes and edges in the target graph through the first graph neural network; The position of the key point in the complex conformation is obtained through the first graph neural network based on the characteristics of the nodes and the characteristics of the edges.
4. The method according to claim 2, wherein: Generating the complex conformation based on the position of the key point in the complex conformation includes: determining a receptor and a ligand from the first protein and the second protein; Obtaining a rotation-translation matrix based on the position of the key point in the monomer conformation of the ligand and the position of the key point in the complex conformation; Based on the rotation and translation matrix, performing an overall spatial transformation on the monomer conformation of the ligand; The complex conformation is generated based on the transformed monomeric conformation of the ligand and the monomeric conformation of the receptor.
5. The method according to any one of claims 1 to 4, wherein After updating the information of the residues of the first protein and the information of the residues of the second protein, the method further includes: Based on the information of the residues of the first protein and the information of the residues of the second protein, the features of the nodes and the features of the edges in the target graph are updated, wherein the features of the nodes are determined based on the information of the residues corresponding to the nodes, and the features of the edges are determined based on the information of the two residues corresponding to the edges.
6. The method according to any one of claims 1 to 4, wherein The constructing of a target graph based on the information of the residues of the first protein and the information of the residues of the second protein comprises: determining a distance between a feature of the first residue and a feature of the second residue based on information of a first residue and information of a second residue, wherein the first residue and the second residue belong to a target protein, wherein the target protein is the first protein or the second protein; sorting the plurality of second residues in ascending order according to the distance, and taking the top N second residues in the sorting as target residues, where N is a positive integer; Adding a connecting edge between the node corresponding to the first residue and the node corresponding to the target residue to generate the target graph.
7. The method according to any one of claims 1 to 4, wherein The constructing of a target graph based on the information of the residues of the first protein and the information of the residues of the second protein comprises: using a residue of the first protein as the third residue; using a residue of the second protein as the fourth residue; A connecting edge is added between the node corresponding to the third residue and the node corresponding to the fourth residue to generate the target graph.
8. A protein docking device, comprising: A construction module for obtaining information about residues of the first protein based on the monomer conformation of the first protein; Based on the monomer conformation of the second protein, obtaining information about the residues of the second protein; constructing a target graph based on information about the residues of the first protein and information about the residues of the second protein, wherein nodes in the target graph are used to represent residues of the first protein or residues of the second protein, edges in the target graph are used to represent edges between two residues, and the target graph is used to generate a conformation of the complex; a docking module configured to dock the first protein and the second protein based on the target graph in each iteration according to at least two molecular docking methods to generate a complex conformation, wherein the at least two molecular docking methods include rigid docking and flexible docking; A processing module, configured to identify that an iteration end condition is not satisfied, and continue the next round of iteration process until the iteration end condition is satisfied, thereby obtaining a final complex conformation; Wherein, for the flexible docking mode, the docking module is further used to: Inputting the target graph into a second graph neural network, and obtaining, by the second graph neural network, positions of the residues in the complex conformation based on the target graph, wherein in the process of obtaining the positions of the residues in the complex conformation, at least one of the monomer conformation of the first protein and the monomer conformation of the second protein changes; generating a conformation of the complex based on positions of a plurality of residues in the conformation of the complex; The building block is further configured to: extracting information about the residues of the first protein and information about the residues of the second protein from the most recently obtained complex conformation, and updating the information about the residues of the first protein and the information about the residues of the second protein; The step of constructing the target graph based on the information of the residues of the first protein and the information of the residues of the second protein is returned to update the target graph.
9. The device according to claim 8, wherein For the rigid docking mode, the docking module is further used to: Inputting the target graph into a first graph neural network, and obtaining positions of key points in the complex conformation based on the target graph by the first graph neural network, wherein the key points are positions on the contact surface between the first protein and the second protein in the complex conformation; Based on the positions of the key points in the complex conformation, the complex conformation is generated.
10. The device according to claim 9, wherein The docking module is further used for: Updating the features of the nodes and edges in the target graph through the first graph neural network; The position of the key point in the complex conformation is obtained through the first graph neural network based on the characteristics of the nodes and the characteristics of the edges.
11. The device according to claim 9, wherein The docking module is further used for: determining a receptor and a ligand from the first protein and the second protein; Obtaining a rotation-translation matrix based on the position of the key point in the monomer conformation of the ligand and the position of the key point in the complex conformation; Based on the rotation and translation matrix, performing an overall spatial transformation on the monomer conformation of the ligand; The complex conformation is generated based on the transformed monomeric conformation of the ligand and the monomeric conformation of the receptor.
12. The device according to any one of claims 8 to 11, wherein: After updating the information of the residues of the first protein and the information of the residues of the second protein, the building module is further configured to: Based on the information of the residues of the first protein and the information of the residues of the second protein, the features of the nodes and the features of the edges in the target graph are updated, wherein the features of the nodes are determined based on the information of the residues corresponding to the nodes, and the features of the edges are determined based on the information of the two residues corresponding to the edges.
13. The device according to any one of claims 8 to 11, wherein: The building block is further used to: determining a distance between a feature of the first residue and a feature of the second residue based on information of a first residue and information of a second residue, wherein the first residue and the second residue belong to a target protein, wherein the target protein is the first protein or the second protein; sorting the plurality of second residues in ascending order according to the distance, and taking the top N second residues in the sorting as target residues, where N is a positive integer; Adding a connecting edge between the node corresponding to the first residue and the node corresponding to the target residue to generate the target graph.
14. The device according to any one of claims 8 to 11, wherein: The building block is further used to: using a residue of the first protein as the third residue; using a residue of the second protein as the fourth residue; A connecting edge is added between the node corresponding to the third residue and the node corresponding to the fourth residue to generate the target graph.
15. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.
16. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 7.
17. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method for calculating and simulating protein-protein docking
CN102314560A
Protein data processing method and device, electronic equipment and storage medium
CN116978450A