Protein docking prediction method and device, computer equipment and readable storage medium
By acquiring and processing the sequence and structural characteristics of similar proteins, combining key structural characteristics for feature fusion, the characterization of protein docking prediction is solved, and the problem of low accuracy of docking prediction in the existing technology is achieved and more accurate protein docking prediction is achieved.
Patent Information
- Application Number
- CN202411716230.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-27
- Publication Date
- 2025-05-02
AI Technical Summary
The existing protein docking prediction methods adopt rigid docking settings, and cannot accurately predict the conformational changes of proteins during docking, resulting in limited understanding of the dynamic properties of protein interactions.
By obtaining the first similar protein and the second similar protein corresponding to the target protein, sequence enhancement processing and structural enhancement processing were performed respectively, enhanced sequence characteristics and enhanced spatial characteristics were obtained, and feature fusion was performed in combination with key structural characteristics to generate the characterization of the proteome to be connected, and finally docking prediction was performed.
It improves the accuracy of protein docking prediction, can more accurately describe the microstructure changes in the protein docking process, and enhances the predictive ability of complex interactions.
Smart Images

Figure CN119920301A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of protein docking, and in particular to a protein docking prediction method, device, computer equipment and readable storage medium. Background Art
[0002] Protein docking technology plays a vital role in biological processes, including DNA replication, signal transduction, immune defense, etc. However, most existing docking methods use rigid docking settings, which cannot accurately predict the conformational changes of proteins during the docking process, limiting the understanding of the dynamic nature of protein interactions.
[0003] There is currently no effective solution to the problem of low accuracy in protein docking prediction in existing technologies. Summary of the invention
[0004] Based on this, it is necessary to provide a protein docking prediction method, device, computer equipment and readable storage medium to address the above technical problems.
[0005] In a first aspect, the present application provides a protein docking prediction method, the method comprising:
[0006] Obtaining a first similar protein and a second similar protein corresponding to a target protein; the target protein is any protein in a protein group to be docked;
[0007] According to the first similar protein, performing sequence enhancement processing on the target protein to obtain enhanced sequence features;
[0008] According to the second similar protein, the target protein is subjected to structural enhancement processing to obtain enhanced spatial features;
[0009] Extracting key structural region features of the target protein to obtain key structural features;
[0010] Perform feature fusion according to the enhanced sequence features, the enhanced spatial features and the key structural features corresponding to the target protein to obtain a representation of the proteome to be docked;
[0011] According to the characterization of the proteome to be docked, docking prediction is performed on the proteome to be docked to obtain a protein docking prediction result.
[0012] In one embodiment, obtaining a first similar protein and a second similar protein corresponding to the target protein includes:
[0013] Performing a similar sequence search on the target protein, and determining a first similar protein matching the target protein sequence from a preset database;
[0014] A similar structure search is performed on the target protein, and a second similar protein matching the structure of the target protein is determined from a preset database.
[0015] In one embodiment, performing sequence enhancement processing on the target protein according to the first similar protein to obtain enhanced sequence features includes:
[0016] Performing sequence encoding processing on the target protein to obtain a first sequence embedding vector;
[0017] Performing sequence encoding processing on the first similar protein to obtain a second sequence embedding vector;
[0018] According to the second sequence embedding vector, sequence enhancement processing is performed on the first sequence embedding vector to obtain enhanced sequence features.
[0019] In one embodiment, the step of performing structural enhancement processing on the target protein according to the second similar protein to obtain enhanced spatial features includes:
[0020] Performing spatial encoding processing on the target protein to obtain a first spatial feature;
[0021] performing spatial encoding processing on the second similar protein to obtain a second spatial feature;
[0022] According to the second spatial feature, the first spatial feature is subjected to structural enhancement processing to obtain an enhanced spatial feature.
[0023] In one embodiment, the performing spatial encoding processing on the target protein to obtain the first spatial feature includes:
[0024] Determine a plurality of amino acids corresponding to the target protein, and take each of the amino acids as a graph node;
[0025] Determining the distance between two of the graph nodes;
[0026] If the distance between the two graph nodes is less than a preset distance threshold, an edge is established between the two graph nodes;
[0027] According to the graph nodes and the edges between the graph nodes, a first data structure graph is obtained;
[0028] Performing a spatial invariance transformation on the first data structure graph to obtain a second data structure graph;
[0029] The second data structure graph is subjected to spatial encoding processing to obtain a first spatial feature.
[0030] In one embodiment, extracting key structural region features of the target protein to obtain key structural features includes:
[0031] Perform structural encoding processing on the target protein to obtain global structural features;
[0032] The global structural features are decoupled to obtain key structural features.
[0033] In one embodiment, the step of performing feature fusion according to the enhanced sequence features, the enhanced spatial features, and the key structural features corresponding to the target protein to obtain the proteome representation to be docked includes:
[0034] According to the enhanced spatial features corresponding to the target protein, feature fusion is performed on the enhanced sequence features to obtain fusion features;
[0035] According to the fusion features, the enhanced spatial features and the key structural features corresponding to the target protein, a characterization of the proteome to be docked is obtained.
[0036] In one embodiment, performing docking prediction on the proteome to be docked according to the characterization of the proteome to be docked to obtain a protein docking prediction result includes:
[0037] For each of the target proteins in the to-be-docking protein group, geometrically adjusting the target protein according to the characterization of the to-be-docking protein group to obtain a corresponding structure-adjusted protein;
[0038] The structure-adjusted protein is subjected to docking prediction to obtain a protein docking prediction result.
[0039] In a second aspect, the present application also provides a protein docking prediction device, the device comprising:
[0040] An acquisition module, used to acquire a first similar protein and a second similar protein corresponding to a target protein; the target protein is any protein in the protein group to be docked;
[0041] A sequence enhancement module, used for performing sequence enhancement processing on the target protein according to the first similar protein to obtain enhanced sequence features;
[0042] A structure enhancement module, used for performing structure enhancement processing on the target protein according to the second similar protein to obtain enhanced spatial features;
[0043] A feature extraction module is used to extract key structural region features of the target protein to obtain key structural features;
[0044] A protein characterization determination module is used to perform feature fusion according to the enhanced sequence features, enhanced spatial features and key structural features corresponding to the target protein to obtain the characterization of the proteome to be docked;
[0045] The prediction module is used to perform docking prediction on the proteome to be docked according to the characterization of the proteome to be docked, and obtain a protein docking prediction result.
[0046] In a third aspect, the present application further provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the method described in any one of the embodiments of the first aspect are implemented.
[0047] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the method described in any one of the embodiments of the first aspect above are implemented.
[0048] In a fifth aspect, the present application further provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the method described in any one of the embodiments of the first aspect above.
[0049] The above-mentioned protein docking prediction method, device, computer equipment and readable storage medium obtain the first similar protein and the second similar protein corresponding to the target protein, and perform sequence enhancement processing and structure enhancement processing on the target protein respectively to obtain enhanced sequence features and enhanced spatial features, which lays a foundation for more accurately describing the microstructural changes in the protein docking process; at the same time, the key structural region features of the target protein are extracted to obtain key structural features, which can effectively eliminate the representations irrelevant to the docking process and reduce unnecessary noise; further, the enhanced sequence features, enhanced spatial features and key structural features are feature fused to obtain the representation of the protein group to be docked; based on the representation of the protein group to be docked, the docking prediction of the protein group to be docked is performed, which not only improves the accuracy of the docking prediction, but also can more accurately describe the microstructural changes in the protein docking process. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the drawings required for use in the embodiments of the present application or related technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0051] Figure 1An application environment diagram of a protein docking prediction method in one embodiment;
[0052] Figure 2 A schematic diagram of a protein docking prediction method in one embodiment;
[0053] Figure 3 is a schematic flow chart of a sequence enhancement processing step in one embodiment;
[0054] Figure 4 A schematic flow chart of a structural enhancement processing step in one embodiment;
[0055] Figure 5 is a schematic flow chart of a structural enhancement processing step in a specific embodiment;
[0056] Figure 6 A schematic diagram of key structural region feature extraction in one embodiment;
[0057] Figure 7 A schematic diagram of protein docking in one embodiment;
[0058] Figure 8 is a structural block diagram of a protein docking prediction device in one embodiment;
[0059] Fig. 9 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0060] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0061] At present, most protein docking methods are based on the concept of rigid docking. Rigid Docking is one of the earliest computational methods used to predict protein interactions and has certain effectiveness in some simple docking scenarios. However, rigid docking has many obvious shortcomings and limitations, especially when dealing with complex interactions in real biological systems.
[0062] Problems with rigid docking include: Rigid Docking assumes that protein monomers maintain their original conformation during the docking process, which means that it ignores the conformational changes that may occur when proteins bind. This assumption is inconsistent with the protein interactions in actual biological systems, because most proteins will undergo some degree of conformational adjustment during the docking process (such as side chain rotation, main chain movement, etc.). Protein binding is often accompanied by the induced fit phenomenon, and the binding of proteins to ligands leads to conformational adjustments to adapt to a more stable binding state. However, rigid docking only performs rigid translation and rotation on proteins, which makes it difficult to handle binding situations that require conformational changes, reducing the accuracy of predictions for complex interactions. Based on this, the embodiments of the present application aim to provide a protein docking prediction method to solve the problem of low protein docking prediction accuracy in the above-mentioned prior art.
[0063] The protein docking prediction method provided in the present application example can be applied to Figure 1 In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or it can be placed on the cloud or other network servers. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptops, tablet computers, and Internet of Things devices. The server 104 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services.
[0064] In an exemplary embodiment, Figure 2 As shown, Figure 2 : is a flow chart of a protein docking prediction method in an embodiment; this embodiment is illustrated by applying the method to a terminal, and it is understandable that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:
[0065] Step S201, obtaining a first similar protein and a second similar protein corresponding to the target protein.
[0066] The target protein is any protein in the proteome to be docked. The proteome to be docked refers to a group of proteins that need to be analyzed for interaction in protein docking research; the proteome to be docked includes multiple proteins; for example, the proteome to be docked includes enzymes and substrates; or, the proteome to be docked includes antibodies and antigens.
[0067] Among them, the first similar protein refers to a protein whose sequence similarity with the target protein is higher than a preset sequence similarity threshold; sequence similarity refers to the percentage of identical and similar residues at corresponding positions between sequences in the total length; among them, the preset sequence similarity threshold needs to be set according to actual conditions and is not specifically limited here.
[0068] In an exemplary embodiment, the method for obtaining the first similar protein corresponding to the target protein may be: determining the first similar protein corresponding to the target protein based on multiple sequence alignment (MSA).
[0069] The second similar protein refers to a protein whose structural similarity with the target protein is higher than a preset structural similarity threshold; structural similarity refers to the degree of similarity between the three-dimensional structures of proteins. The preset structural similarity threshold needs to be set according to actual conditions and is not specifically limited here.
[0070] In an exemplary embodiment, the method for obtaining the second similar protein corresponding to the target protein may be: based on the three-dimensional structure of the target protein and a preset structural similarity threshold, the second similar protein is retrieved from a preset database. The preset database may include, but is not limited to, a protein structure database (Protein Data Bank, PDB).
[0071] Step S202: performing sequence enhancement processing on the target protein according to the first similar protein to obtain enhanced sequence features.
[0072] It should be noted that since the sequence similarity between the sequence of the first similar protein and the sequence of the target protein is higher than the preset sequence similarity threshold, sequence enhancement processing of the target protein through the first similar protein can effectively enhance the sequence features corresponding to the target protein, so that in the downstream protein docking prediction, it is possible to better capture the subtle conformational changes in the protein docking process, thereby laying the foundation for improving the accuracy of protein docking prediction.
[0073] In an exemplary embodiment, based on the first similar protein, the target protein is subjected to sequence enhancement processing, and the method for obtaining enhanced sequence features may be: based on a pre-trained protein language model, the sequence embedding vectors corresponding to the first similar protein and each group of target proteins are determined respectively; the sequence embedding vector corresponding to the first similar protein is superimposed with the sequence embedding vector corresponding to the target protein to obtain the enhanced sequence features corresponding to the target protein.
[0074] In another exemplary embodiment, the method for performing sequence enhancement processing on the target protein according to the first similar protein to obtain enhanced sequence features may also be: based on a pre-trained protein language model, determining the sequence embedding vectors corresponding to the first similar protein and the target protein respectively; through an attention mechanism, integrating the sequence embedding vector corresponding to the first similar protein into the sequence embedding vector corresponding to the target protein, to obtain enhanced sequence features corresponding to the target protein.
[0075] Step S203: Based on the second similar protein, the target protein is subjected to structural enhancement processing to obtain enhanced spatial features.
[0076] It should be noted that since the structural similarity between the structure of the second similar protein and the structure of the target protein is higher than the preset structural similarity threshold, structural enhancement processing of the target protein through the second similar protein can effectively enhance the spatial characteristics corresponding to the target protein, so as to obtain a variety of possible spatial conformation information, thereby laying the foundation for improving the accuracy of protein docking prediction.
[0077] In an exemplary embodiment, the method of performing structural enhancement processing on the target protein based on the second similar protein to obtain enhanced spatial features may be: performing spatial encoding processing on the second similar protein and the target protein respectively to obtain the spatial features corresponding to the second similar protein and the target protein respectively; fusing the spatial features of the second similar protein with the spatial features of the target protein to obtain enhanced spatial features of the target protein.
[0078] Step S204, extract key structural region features of the target protein to obtain key structural features.
[0079] Among them, the key structural region refers to the characteristic region of the target protein that participates in the binding effect during docking; for example, the docking pocket region. By extracting key structural features, the accuracy of structural prediction near the binding site can be improved, thereby improving the prediction accuracy during the dynamic binding process.
[0080] In an exemplary embodiment, a method for extracting key structural region features of a target protein to obtain key structural features may be: extracting key structural region features of the target protein through a trained geometric graph neural network (GNN) to obtain key structural features.
[0081] Step S205 , performing feature fusion according to the enhanced sequence features, enhanced spatial features and key structural features corresponding to the target protein to obtain a representation of the proteome to be docked.
[0082] The characterization of the proteome to be docked includes the enhanced sequence features, enhanced spatial features and key structural features corresponding to each target protein in the proteome to be docked. The enhanced sequence features are one-dimensional feature information; the enhanced spatial features and key structural features are three-dimensional feature information.
[0083] It should be noted that feature fusion based on the enhanced sequence features, enhanced spatial features and key structural features corresponding to the target protein can effectively enhance the expressiveness of the features and ensure that the sequence information, spatial information and key structural information can be fully utilized in the prediction process. It is understandable that when each amino acid in the target protein is taken as a graph node, the spatial information can include, but is not limited to, the three-dimensional structural information of each graph node and the spatial distance information between two graph nodes.
[0084] Step S206, performing docking prediction on the proteome to be docked according to the characterization of the proteome to be docked, and obtaining a protein docking prediction result.
[0085] In an exemplary embodiment, a representation of a proteome to be docked is obtained based on a trained geometric graph neural network, and further, the representation of the proteome to be docked is input into a trained prediction branch, and docking prediction is performed on the proteome to be docked through the trained prediction branch to obtain a protein docking prediction result. The trained prediction branch may include, but is not limited to, a multi-layer perceptron.
[0086] Preferably, when training the prediction branch, a multi-task learning strategy is used to improve the accuracy of the model's prediction of protein docking; specifically, the prediction of binding sites is added as an auxiliary task during the training process to help the model better learn the local features of the binding interface and enhance the model's understanding of the binding region. In addition, in terms of loss function design, the loss of protein docking prediction and the loss of binding site prediction are designed separately, and optimized in a weighted manner to balance the importance of the two.
[0087] In this embodiment, by obtaining the first similar protein and the second similar protein corresponding to the target protein, the target protein is subjected to sequence enhancement processing and structure enhancement processing respectively, and enhanced sequence features and enhanced spatial features are obtained, which lays a foundation for more accurately describing the microstructural changes in the protein docking process; at the same time, key structural region features of the target protein are extracted to obtain key structural features, which can effectively eliminate representations irrelevant to the docking process and reduce unnecessary noise; further, the enhanced sequence features, enhanced spatial features and key structural features are feature fused to obtain the representation of the protein group to be docked; based on the representation of the protein group to be docked, docking prediction is performed on the proteome to be docked, which not only improves the accuracy of the docking prediction, but also can more accurately describe the microstructural changes in the protein docking process.
[0088] In one embodiment, obtaining a first similar protein and a second similar protein corresponding to a target protein comprises the following steps:
[0089] Step 1: perform similar sequence search on the target protein and determine the first similar protein matching the target protein sequence from a preset database.
[0090] It should be noted that similar sequence retrieval helps capture conserved functional regions in sequences, which are usually closely related to the structure and function of proteins, thereby performing data enhancement at the one-dimensional sequence level, enabling the model to better understand the evolutionary background and biological semantics of the sequence.
[0091] Step 2: perform a similar structure search on the target protein and determine a second similar protein that matches the target protein structure from a preset database.
[0092] It should be noted that similar structure retrieval can not only help the model capture more structurally conserved folding methods, but also enable the model to have a stronger understanding of similar spatial folding through this data enhancement, especially in the binding interface area that requires docking prediction.
[0093] Among them, the preset databases may include but are not limited to PDB, UniProt, and InterPro; among them, PDB is the largest biological macromolecule structure database, containing a large amount of three-dimensional structure information of proteins and nucleic acids; UniProt can provide rich protein sequence and function annotation information; InterPro can provide annotation information of protein families, domains and conserved sites.
[0094] In an exemplary embodiment, based on a preset database, similar sequence retrieval and similar structure retrieval are performed on the target protein respectively to determine a first similar protein matching the target protein sequence and a second similar protein matching the target protein structure from the preset database.
[0095] In this embodiment, based on the preset database, the first similar protein matching the target protein sequence and the second similar protein matching the target protein structure can be accurately matched, thereby helping the downstream prediction task to capture the subtle spatial changes that occur during the protein docking process.
[0096] It lays the foundation for achieving sequence enhancement and structure enhancement processing of target proteins.
[0097] In one embodiment, Figure 3 As shown, Figure 3The flowchart of the sequence enhancement processing step in one embodiment is as follows; according to the first similar protein, the target protein is subjected to sequence enhancement processing to obtain enhanced sequence features, including the following steps:
[0098] Step S301, performing sequence encoding processing on the target protein to obtain a first sequence embedding vector.
[0099] The sequence encoding process refers to converting the protein sequence into a numerical vector representation. It should be noted that the amino acid sequence of the target protein is the basis for realizing the sequence encoding process.
[0100] The first sequence embedding vector refers to a numerical vector corresponding to the target protein sequence; it can be understood that the semantic information of the sequence can be captured based on the first sequence embedding vector.
[0101] Step S302: performing sequence encoding processing on the first similar protein to obtain a second sequence embedding vector.
[0102] The second sequence embedding vector refers to a numerical vector corresponding to the first similar protein. The second sequence embedding vector is used to perform sequence enhancement processing on the first sequence embedding vector.
[0103] Step S303: performing sequence enhancement processing on the first sequence embedding vector according to the second sequence embedding vector to obtain enhanced sequence features.
[0104] In an exemplary embodiment, based on a pre-trained protein language model, sequence encoding processing is performed on the target protein to obtain a first sequence embedding vector; based on the pre-trained protein language model, sequence encoding processing is performed on the first similar protein to obtain a second sequence embedding vector; further, based on a multi-layer perceptron, the second sequence embedding vector is integrated into the first sequence embedding vector to obtain enhanced sequence features.
[0105] In this embodiment, sequence encoding processing is performed on the target protein and the first similar protein respectively to obtain a first sequence embedding vector and a second sequence embedding vector. Furthermore, according to the second sequence embedding vector, sequence enhancement processing is performed on the first sequence embedding vector to obtain enhanced sequence features. Based on this, the quality of the target protein sequence features can be significantly improved, helping downstream prediction tasks to accurately capture subtle spatial changes that occur during protein docking.
[0106] In one embodiment, Figure 4 As shown, Figure 4 FIG. 1 is a flow chart of a structural enhancement processing step in an embodiment; based on the second similar protein, the target protein is subjected to structural enhancement processing to obtain enhanced spatial features, including the following steps:
[0107] Step S401, performing spatial encoding processing on the target protein to obtain a first spatial feature.
[0108] Among them, spatial encoding processing refers to converting the three-dimensional structural information of protein into a numerical vector representation.
[0109] Among them, the first spatial feature is used to characterize the three-dimensional structural information of amino acids in the target protein and the spatial distance information between amino acids.
[0110] Step S402: perform spatial encoding processing on the second similar protein to obtain a second spatial feature.
[0111] The second spatial feature is used to characterize the three-dimensional structural information of amino acids in the second similar protein and the spatial distance information between the amino acids.
[0112] Step S403: Perform structural enhancement processing on the first spatial feature according to the second spatial feature to obtain an enhanced spatial feature.
[0113] In an exemplary embodiment, the target protein is spatially encoded by a trained geometric graph neural network to obtain a first spatial feature; the second similar protein is spatially encoded by a trained geometric graph neural network to obtain a second spatial feature; further, the second spatial feature and the first spatial feature are feature fused to obtain an enhanced spatial feature.
[0114] In this embodiment, the first spatial feature and the second spatial feature are obtained by performing spatial encoding processing on the target protein and the second similar protein respectively. Furthermore, the first spatial feature is structurally enhanced according to the second spatial feature to obtain the enhanced spatial feature. Based on this, the quality of the spatial feature of the target protein can be significantly improved, helping downstream prediction tasks to accurately capture more structurally conservative folding methods.
[0115] In one embodiment, Figure 5 As shown, Figure 5 Schematic diagram of a flow chart of a structure enhancement processing step in a specific embodiment; performing spatial encoding processing on a target protein to obtain a first spatial feature includes the following steps:
[0116] Step S501, determining multiple amino acids corresponding to the target protein, and taking each amino acid as a graph node.
[0117] It should be noted that by using amino acids as graph nodes, the graph structure can be used to capture the spatial and topological information of proteins.
[0118] Step S502, determining the distance between two graph nodes.
[0119] Among them, the distance between two graph nodes is the Euclidean distance.
[0120] Step S503: if the distance between any two graph nodes is less than a preset distance threshold, edges are established between any two graph nodes.
[0121] Among them, the preset distance threshold needs to be set according to actual prediction needs and is not specifically limited here.
[0122] Among them, edges represent the relationship or interaction between graph nodes.
[0123] It can be understood that if the distance between two graph nodes is less than a preset distance threshold, it indicates that there is a strong correlation between the two graph nodes.
[0124] Step S504: Obtain a first data structure graph according to the graph nodes and the edges between the graph nodes.
[0125] The first data structure graph is a mathematical structure composed of multiple graph nodes and multiple edges, and is used to represent the relationship between amino acids.
[0126] Step S505: perform space-invariance transformation on the first data structure graph to obtain a second data structure graph.
[0127] Among them, spatial invariance transformation refers to a series of operations such as random translation and rotation of the first data structure graph, so that the model output encoding of the protein structure remains unchanged for a specific type of transformation (such as translation, rotation, etc.); it can be understood that by performing spatial invariance transformation on the first data structure graph, the robustness of the model can be effectively enhanced, so that it can still maintain good performance when facing data of different perspectives, positions or scales.
[0128] In an exemplary embodiment, a trained geometric graph neural network with SE(3) invariance is used to perform spatial invariance transformation on the first data structure graph; wherein SE(3) invariance means that the output of the model remains unchanged under rotation and translation transformations.
[0129] Step S506: perform spatial encoding processing on the second data structure graph to obtain a first spatial feature.
[0130] Exemplarily, multiple amino acids corresponding to the target protein are determined, and each amino acid is taken as a graph node; further, the Euclidean distance between the two graph nodes is calculated, and it is determined whether the Euclidean distance between the two graph nodes is less than a preset distance threshold; if the distance between the two graph nodes is less than the preset distance threshold, an edge is established between the two graph nodes; if the distance between the two graph nodes is greater than or equal to the preset distance threshold, no edge is established between the two graph nodes; and so on, according to the two graph nodes and the edges between the two graph nodes, a first data structure graph is obtained; further, based on the SE(3)-transform, the first data structure graph is spatially invariantly transformed to obtain a second data structure graph; wherein the SE(3)-transform is used to achieve SE(3) invariance; further, the second data structure graph is spatially encoded through a trained geometric graph neural network to obtain a first spatial feature.
[0131] It should be noted that the specific implementation method of performing spatial encoding processing on the second similar protein to obtain the second spatial feature is the same as the specific implementation method of performing spatial encoding processing on the target protein to obtain the first spatial feature, and will not be repeated here.
[0132] In this embodiment, each amino acid in the target protein is taken as a graph node, and by comparing the distance between each pair of graph nodes with a preset distance threshold, corresponding edges are established between each pair of graph nodes to form a first data structure graph; then, the first data structure graph is subjected to a spatial invariance transformation to obtain a second data structure graph; this can ensure that the encoding of the protein structure remains invariant under transformations such as rotation and translation; further, the second data structure graph is subjected to spatial encoding processing to obtain a first spatial feature, which can accurately characterize the three-dimensional structural information of the amino acids in the target protein and the spatial distance information between the amino acids.
[0133] In one embodiment, extracting key structural region features of a target protein to obtain key structural features comprises the following steps:
[0134] Step 1: Perform structural encoding processing on the target protein to obtain global structural features.
[0135] Step 2: Decouple the global structural features to obtain key structural features.
[0136] It should be noted that by extracting the key structural region features of the target protein and obtaining the key structural features, targeted characterization enhancement can be performed, allowing the model to obtain more targeted knowledge in these key areas, thereby improving the prediction accuracy in the dynamic docking process.
[0137] For example, see Figure 6, through the trained geometric graph neural network (GNN), the target protein is structurally encoded to obtain the global structural features; further, the global structural features are decoupled to obtain key structural features and irrelevant information; among them, the key structural features include pocket representation. Figure 6 In , Scorer is a scorer used to evaluate the effectiveness of the extracted features.
[0138] In this embodiment, the target protein is subjected to structural encoding processing to obtain global structural features; the global structural features are then decoupled to obtain key structural feature points. Based on this, representations irrelevant to the docking process are eliminated, unnecessary noise is reduced, and the model's learning of key binding areas is enhanced, making the model more refined and accurate in predicting structures near the binding site.
[0139] In one embodiment, feature fusion is performed according to the enhanced sequence features, enhanced spatial features, and key structural features corresponding to the target protein to obtain the representation of the proteome to be docked, including the following steps:
[0140] Step 1: According to the enhanced spatial features corresponding to the target protein, the enhanced sequence features are fused to obtain the fused features.
[0141] Step 2: According to the fusion features, enhanced spatial features and key structural features corresponding to the target protein, the proteome characterization to be docked is obtained.
[0142] It should be noted that the enhanced spatial feature is three-dimensional feature information, and the enhanced sequence feature is one-dimensional feature information.
[0143] For example, based on the intra-monomer spatial attention module, the enhanced spatial features corresponding to the target protein are integrated into the enhanced sequence features to obtain fused features. The intra-monomer spatial attention module realizes the fusion of enhanced spatial features and enhanced sequence features through the attention mechanism. The attention mechanism can help the model focus on important spatial and sequence features, thereby improving the expressiveness of the features.
[0144] In this embodiment, based on the enhanced spatial features corresponding to the target protein, the enhanced sequence features are feature fused to obtain the fusion features; then, based on the fusion features, enhanced spatial features and key structural features corresponding to the target protein, the representation of the proteome to be docked is obtained, which can more accurately describe the microstructural changes in the protein docking process and lay the foundation for improving the accuracy of docking prediction.
[0145] In one embodiment, according to the characterization of the proteome to be docked, docking prediction is performed on the proteome to be docked to obtain a protein docking prediction result, including the following steps:
[0146] Step 1: For each target protein in the proteome to be docked, according to the characterization of the proteome to be docked, geometrically adjust the target protein to obtain the corresponding structure-adjusted protein.
[0147] The geometric adjustment may include but is not limited to translation, rotation, dynamic change of structural sites, etc.
[0148] Step 2: Perform docking prediction on the structure-adjusted protein to obtain protein docking prediction results.
[0149] For example, see Figure 7 Based on the trained prediction branch, for each target protein in the to-be-docked proteome, according to the to-be-docked proteome characterization, the target protein is geometrically adjusted to obtain the corresponding structure-adjusted protein; further, docking prediction is performed on the structure-adjusted protein to obtain the protein docking prediction result. The trained prediction branch may include, but is not limited to, a multi-layer perceptron. Figure 7 It can be seen that when docking prediction is performed on the docking proteome, the dynamic binding process of proteins can be accurately reflected, thus improving the accuracy of protein docking prediction.
[0150] In this embodiment, for each target protein in the proteome to be docked, the target protein is geometrically adjusted according to the characterization of the proteome to be docked to obtain the corresponding structure-adjusted protein, and then the structure-adjusted protein is docked and predicted to obtain a protein docking prediction result, which effectively improves the accuracy of protein docking prediction and can accurately reflect the process of dynamic protein binding.
[0151] The above-mentioned protein docking prediction method obtains the first similar protein and the second similar protein corresponding to the target protein, and performs sequence enhancement processing and structure enhancement processing on the target protein respectively, so as to obtain enhanced sequence features and enhanced spatial features, which lays a foundation for more accurately describing the microstructural changes in the docking process; at the same time, the key structural region features of the target protein are extracted to obtain key structural features, which can effectively eliminate the representations irrelevant to the docking process and reduce unnecessary noise; further, the enhanced sequence features, enhanced spatial features and key structural features are feature fused to obtain the representation of the protein group to be docked; based on the representation of the proteome to be docked, the docking prediction of the proteome to be docked is performed, which not only improves the accuracy of the docking prediction, but also can more accurately describe the microstructural changes in the protein docking process.
[0152] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.
[0153] Based on the same inventive concept, the embodiment of the present application also provides a protein docking prediction device for implementing the protein docking prediction method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more protein docking prediction device embodiments provided below can refer to the limitations of the protein docking prediction method above, and will not be repeated here.
[0154] In an exemplary embodiment, Figure 8 As shown, Figure 8 80 is a structural block diagram of a protein docking prediction device in an embodiment; the protein docking prediction device includes an acquisition module 801, a sequence enhancement module 802, a structure enhancement module 803, a feature extraction module 804, a protein characterization determination module 805 and a prediction module 806;
[0155] An acquisition module 801 is used to acquire a first similar protein and a second similar protein corresponding to a target protein; the target protein is any protein in the protein group to be docked;
[0156] A sequence enhancement module 802 is used to perform sequence enhancement processing on the target protein according to the first similar protein to obtain enhanced sequence features;
[0157] The structure enhancement module 803 is used to perform structure enhancement processing on the target protein according to the second similar protein to obtain enhanced spatial features;
[0158] The feature extraction module 804 is used to extract key structural region features of the target protein to obtain key structural features;
[0159] The protein characterization determination module 805 is used to perform feature fusion according to the enhanced sequence features, enhanced spatial features and key structural features corresponding to the target protein to obtain the characterization of the proteome to be docked;
[0160] The prediction module 806 is used to perform docking prediction on the to-be-docked proteome according to the characterization of the to-be-docked proteome, and obtain a protein docking prediction result.
[0161] The above-mentioned protein docking prediction device obtains the first similar protein and the second similar protein corresponding to the target protein, and performs sequence enhancement processing and structure enhancement processing on the target protein respectively, so as to obtain enhanced sequence features and enhanced spatial features, which lays a foundation for more accurately describing the microstructural changes in the docking process; at the same time, the key structural region features of the target protein are extracted to obtain key structural features, which can effectively eliminate the representations irrelevant to the docking process and reduce unnecessary noise; further, the enhanced sequence features, enhanced spatial features and key structural features are feature fused to obtain the representation of the protein group to be docked; based on the representation of the protein group to be docked, the docking prediction of the protein group to be docked is performed, which not only improves the accuracy of the docking prediction, but also can more accurately describe the microstructural changes in the protein docking process.
[0162] In one embodiment, the acquisition module 801 is also used to
[0163] Performing a similar sequence search on the target protein, and determining a first similar protein matching the target protein sequence from a preset database;
[0164] A similar structure search is performed on the target protein, and a second similar protein matching the target protein structure is determined from a preset database.
[0165] In one embodiment, the sequence enhancement module 802 is also used to
[0166] Perform sequence encoding processing on the target protein to obtain a first sequence embedding vector;
[0167] Performing sequence encoding processing on the first similar protein to obtain a second sequence embedding vector;
[0168] According to the second sequence embedding vector, sequence enhancement processing is performed on the first sequence embedding vector to obtain enhanced sequence features.
[0169] In one embodiment, the structural reinforcement module 803 is also used to
[0170] Performing spatial encoding processing on the target protein to obtain the first spatial feature;
[0171] Performing spatial encoding processing on the second similar protein to obtain a second spatial feature;
[0172] According to the second spatial feature, the first spatial feature is subjected to structural enhancement processing to obtain an enhanced spatial feature.
[0173] In one embodiment, the structural reinforcement module 803 is also used to
[0174] Determine multiple amino acids corresponding to the target protein, and treat each amino acid as a graph node;
[0175] Determine the distance between two graph nodes;
[0176] If the distance between two graph nodes is less than the preset distance threshold, an edge is established between the two graph nodes;
[0177] According to the graph nodes and the edges between the graph nodes, a first data structure graph is obtained;
[0178] Performing a space-invariant transformation on the first data structure graph to obtain a second data structure graph;
[0179] The second data structure graph is spatially encoded to obtain a first spatial feature.
[0180] In one embodiment, the feature extraction module 804 is also used to
[0181] Perform structural encoding processing on the target protein to obtain global structural features;
[0182] The global structural features are decoupled to obtain the key structural features.
[0183] In one embodiment, the protein characterization determination module 805 is further configured to
[0184] According to the enhanced spatial features corresponding to the target protein, the enhanced sequence features are fused to obtain fused features;
[0185] According to the fusion features, enhanced spatial features and key structural features corresponding to the target protein, the characterization of the proteome to be docked is obtained.
[0186] In one embodiment, the prediction module 806 is also used to
[0187] For each target protein in the proteome to be docked, geometric adjustment is performed on the target protein according to the characterization of the proteome to be docked to obtain the corresponding structure-adjusted protein;
[0188] Docking prediction is performed on the structure-adjusted protein to obtain protein docking prediction results.
[0189] Each module in the above-mentioned protein docking prediction device can be implemented in whole or in part by software, hardware and a combination thereof. Each of the above-mentioned modules can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.
[0190] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in FIG. Fig. 9 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store protein docking prediction related data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a protein docking prediction method is implemented.
[0191] Those skilled in the art will understand that Fig. 9 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0192] In one embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above method embodiments when executing the computer program.
[0193] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0194] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0195] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., but are not limited to this.
[0196] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0197] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A protein docking prediction method, characterized in that: The method comprises: Obtaining a first similar protein and a second similar protein corresponding to a target protein; the target protein is any protein in a protein group to be docked; According to the first similar protein, performing sequence enhancement processing on the target protein to obtain enhanced sequence features; According to the second similar protein, the target protein is subjected to structural enhancement processing to obtain enhanced spatial features; Extracting key structural region features of the target protein to obtain key structural features; Perform feature fusion according to the enhanced sequence features, the enhanced spatial features and the key structural features corresponding to the target protein to obtain a representation of the proteome to be docked; According to the characterization of the proteome to be docked, docking prediction is performed on the proteome to be docked to obtain a protein docking prediction result.
2. The method according to claim 1, characterized in that The step of obtaining a first similar protein and a second similar protein corresponding to the target protein includes: Performing a similar sequence search on the target protein, and determining a first similar protein matching the target protein sequence from a preset database; A similar structure search is performed on the target protein, and a second similar protein matching the structure of the target protein is determined from a preset database.
3. The method according to claim 1, characterized in that The step of performing sequence enhancement processing on the target protein according to the first similar protein to obtain enhanced sequence features includes: Performing sequence encoding processing on the target protein to obtain a first sequence embedding vector; Performing sequence encoding processing on the first similar protein to obtain a second sequence embedding vector; According to the second sequence embedding vector, sequence enhancement processing is performed on the first sequence embedding vector to obtain enhanced sequence features.
4. The method according to claim 1, characterized in that: The step of performing structural enhancement processing on the target protein according to the second similar protein to obtain enhanced spatial features includes: Performing spatial encoding processing on the target protein to obtain a first spatial feature; performing spatial encoding processing on the second similar protein to obtain a second spatial feature; According to the second spatial feature, the first spatial feature is subjected to structural enhancement processing to obtain an enhanced spatial feature.
5. The method according to claim 4, characterized in that The performing spatial encoding processing on the target protein to obtain a first spatial feature includes: Determine a plurality of amino acids corresponding to the target protein, and take each of the amino acids as a graph node; Determining the distance between two of the graph nodes; If the distance between the two graph nodes is less than a preset distance threshold, an edge is established between the two graph nodes; According to the graph nodes and the edges between the graph nodes, a first data structure graph is obtained; Performing a spatial invariance transformation on the first data structure graph to obtain a second data structure graph; The second data structure graph is subjected to spatial encoding processing to obtain a first spatial feature.
6. The method according to claim 1, characterized in that The key structural region feature extraction of the target protein to obtain the key structural features includes: Perform structural encoding processing on the target protein to obtain global structural features; The global structural features are decoupled to obtain key structural features.
7. The method according to claim 1, characterized in that The step of performing feature fusion according to the enhanced sequence features, the enhanced spatial features and the key structural features corresponding to the target protein to obtain a representation of the proteome to be docked comprises: According to the enhanced spatial features corresponding to the target protein, feature fusion is performed on the enhanced sequence features to obtain fusion features; According to the fusion features, the enhanced spatial features and the key structural features corresponding to the target protein, a characterization of the proteome to be docked is obtained.
8. The method according to claim 1, characterized in that The step of performing docking prediction on the proteome to be docked according to the characterization of the proteome to be docked to obtain a protein docking prediction result comprises: For each of the target proteins in the to-be-docking protein group, geometrically adjusting the target protein according to the characterization of the to-be-docking protein group to obtain a corresponding structure-adjusted protein; The structure-adjusted protein is subjected to docking prediction to obtain a protein docking prediction result.
9. A protein docking prediction device, characterized in that: The device comprises: An acquisition module, used to acquire a first similar protein and a second similar protein corresponding to a target protein; the target protein is any protein in the protein group to be docked; A sequence enhancement module, used for performing sequence enhancement processing on the target protein according to the first similar protein to obtain enhanced sequence features; A structure enhancement module, used for performing structure enhancement processing on the target protein according to the second similar protein to obtain enhanced spatial features; A feature extraction module is used to extract key structural region features of the target protein to obtain key structural features; A protein characterization determination module is used to perform feature fusion according to the enhanced sequence features, enhanced spatial features and key structural features corresponding to the target protein to obtain the characterization of the proteome to be docked; The prediction module is used to perform docking prediction on the proteome to be docked according to the characterization of the proteome to be docked, and obtain a protein docking prediction result.
10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
12. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Protein-small molecule compound docking method based on local template similarity
CN116884505A
Protein structure prediction method and device, electronic equipment and storage medium
CN117558337A
Protein binding site prediction method, system, medium, equipment and product
CN118522346A
Training method for protein complex structure prediction model, device, and medium
WO2024153242A1