A method, apparatus, device and storage medium for predicting the structure of a protein complex

By using trained position prediction model and target map prediction network in protein structure prediction, the optimal translation and rotation relationship between protein chains is determined, and the problems of inaccurate energy model and complex conformation space in protein structure prediction are solved, and efficient protein complex structure prediction is achieved.

CN116052758BActive Publication Date: 2025-05-30BIOMAP (BEIJING) INTELLIGENCE TECH LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310085702.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-19
Publication Date
2025-05-30
Estimated Expiration
2043-01-19

AI Technical Summary

Technical Problem

The energy model in protein structure prediction is inaccurate, and the protein conformation space is huge and complex, resulting in large calculation amounts and increased time consumption.

Method used

By avoiding the use of complex network structure prediction networks to predict the precise distance between different protein chains that make up the complex at one time, using trained position prediction models and target map prediction networks, the optimal translation rotation relationship between protein chains is determined to reduce computational volume and time consumption.

Benefits of technology

It effectively reduces the calculation amount and time of protein complex structure prediction, improves the accuracy of prediction results, and avoids high consumption of complex networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116052758B_ABST
    Figure CN116052758B_ABST
Patent Text Reader

Abstract

The present application provides a method, apparatus, device and storage medium for predicting the structure of a protein complex. Among them, the amino acid sequences of multiple protein chains in the protein complex are input into a trained position prediction model to obtain the predicted atomic positions of the first representative atoms of each of the protein chains; according to the predicted atomic positions and the predicted distance map corresponding to the multiple protein chains, an optimal translational and rotational relationship between the multiple protein chains is determined, so that the distance difference between the calculated distance map of the multiple protein chains under the optimal translational and rotational relationship and the predicted distance map is minimized; according to the optimal translational and rotational relationship, the predicted structure of the protein complex is obtained. By adopting the above method, the computational amount in predicting the structure of a protein complex can be reduced, and the time required for predicting the structure of a protein complex can be shortened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of bioinformatics, and in particular, to a method, apparatus, device, and storage medium for predicting the structure of a protein complex. Background Art

[0002] Protein molecules play a great role in the cellular activities of organisms, and many activities of organisms are based on the activity of proteins. The structure of protein molecules determines their functions. Therefore, modeling the structure and bioactive state of biomolecules is very helpful for understanding and treating protein-related diseases, and is also instructive for the manufacture of engineered proteins.

[0003] The view that the amino acid sequence information of proteins determines their three-dimensional structure is widely accepted, and it is also the theoretical basis for using computers to predict protein structures.

[0004] The inventors found in their research that using the computing power of computers and optimization algorithms to predict the three-dimensional structure of proteins through their sequence information, that is, the protein folding problem remains a difficult problem. The difficulties in protein structure prediction are mainly in two aspects. First, the energy model used for protein structure prediction is inaccurate. Second, the conformational space of proteins is extremely large and complex, which greatly increases the time required to accurately predict the structure of the components of the complex at one time using a complex network structure prediction network. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide a method, apparatus, device, and storage medium for predicting the structure of a protein complex, which can reduce the computational amount and shorten the time required for predicting the structure of a protein complex by avoiding using a complex network structure prediction network to predict the accurate distances between different protein chains that make up the complex at one time.

[0006] In a first aspect, an embodiment of the present application provides a method for predicting the structure of a protein complex, the method comprising:

[0007] Inputting the amino acid sequences of multiple protein chains in the protein complex into a trained position prediction model to obtain the predicted atomic positions of the first representative atoms of each protein chain;

[0008] Determine the optimal translational and rotational relationship between multiple protein chains based on the predicted atomic positions and predicted distance maps corresponding to the multiple protein chains, such that the distance difference between the calculated distance map of the multiple protein chains under the optimal translational and rotational relationship and the predicted distance map is minimized; wherein, the predicted distance map is determined based on the amino acid sequences of the multiple protein chains and is used to characterize the approximate distances between the second representative atoms of the multiple protein chains; the calculated distance map is determined based on the predicted atomic positions corresponding to the multiple protein chains and the translational and rotational relationship between the multiple protein chains and is used to characterize the distances between the first representative atoms of the multiple protein chains under this translational and rotational relationship; the second representative atoms and the first representative atoms are the same or different;

[0009] Obtain the predicted structure of the protein complex according to the optimal translational and rotational relationship.

[0010] Optionally, before determining the optimal translational and rotational relationship between multiple protein chains based on the predicted atomic positions and predicted distance maps corresponding to the multiple protein chains, the method further includes:

[0011] Input the amino acid sequences of multiple protein chains into a trained target map prediction network to obtain a quasi-predicted distance map, where the quasi-predicted distance map includes at least one M*N*P matrix, where P is the number of distance classifications, and the element at the mnth position of the pth layer of the matrix represents the probability that the predicted distance between the mth first representative atom of the first protein chain and the nth first representative atom of the second protein chain falls into the pth distance classification, and different distance classifications correspond to different distance intervals; where p is an integer selected from 1 to P; m is an integer selected from 1 to M, and n is an integer selected from 1 to N;

[0012] Obtain the predicted distance map according to the quasi-predicted distance map; wherein, the predicted distance map includes the approximate distance between the mth first representative atom of the first protein chain and the nth first representative atom of the second protein chain, and the approximate distance is represented by the category corresponding to the maximum probability of the mth first representative atom of the first protein chain and the nth first representative atom of the second protein chain.

[0013] Optionally, the target map prediction network includes an evolutionary network module and a map generation module;

[0014] The target map prediction network is trained through the following steps:

[0015] Input the amino acid sequences of multiple training protein chains included in a protein complex with a known structure into the evolutionary network module to obtain paired features;

[0016] Input the paired features into the map generation module to obtain the predicted approximate atomic distances;

[0017] Adjust the parameters of the evolutionary network module and / or the map generation module according to the difference value between the predicted approximate atomic distances and the actual approximate atomic distances to obtain the target map prediction network.

[0018] Optionally, when the second representative atom is different from the first representative atom, the determining of the optimal translational and rotational relationship between the multiple protein chains according to the predicted atomic positions and the predicted distance map corresponding to the multiple protein chains includes:

[0019] Convert the calculated distance map into a converted calculated distance map, where the converted calculated distance map is characterized by the distances between the second representative atoms; the distance difference between the converted calculated distance maps of the multiple protein chains under the optimal translational and rotational relationship and the predicted distance map is the smallest;

[0020] Alternatively, convert the predicted distance map into a converted predicted distance map, where the converted predicted distance map is characterized by the distances between the first representative atoms; the distance difference between the calculated distance maps of the multiple protein chains under the optimal translational and rotational relationship and the converted predicted distance map is the smallest.

[0021] Optionally, the determining of the optimal translational and rotational relationship between the multiple protein chains according to the predicted atomic positions and the predicted distance map corresponding to the multiple protein chains includes:

[0022] Obtain the optimal translational and rotational relationship according to the following formula:

[0023]

[0024] where, x A is the predicted atomic position of one protein chain, x B is the predicted atomic position of another protein chain, C is the predicted distance map, T is the translational and rotational relationship, R is the rotation matrix, and t is the translation matrix.

[0025] Optionally, when the protein complex includes W protein chains, the optimal translational and rotational relationship includes W - 1 translational and rotational relationships, where W is greater than or equal to 3;

[0026] The W - 1 translational and rotational relationships minimize the weighted sum of the distance differences between the calculated distance maps of the multiple protein chains pairwise and the predicted distance map;

[0027] Alternatively, each of the W-1 translational and rotational relationships enables the computational distance map between the two protein chains representing the relative relationship to have the smallest distance difference from the predicted distance map.

[0028] Optionally, when the protein complex includes a third protein chain, a fourth protein chain, and a fifth protein chain, the optimal translational and rotational relationship includes a first optimal translational and rotational relationship between the third protein chain and the fourth protein chain and a second optimal translational and rotational relationship between the complex and the fifth protein chain; the complex is the complex formed by the third protein chain and the fourth protein chain under the first optimal translational and rotational relationship;

[0029] The step of inputting the amino acid sequences of multiple protein chains in the protein complex into the trained position prediction model to obtain the predicted atomic positions of the first representative atoms in the multiple protein chains includes: inputting the amino acid sequences of the third protein chain, the fourth protein chain, and the fifth protein chain into the trained position prediction model to obtain the predicted atomic positions of the first representative atoms in the third protein chain, the predicted atomic positions of the first representative atoms in the fourth protein chain, and the predicted atomic positions of the first representative atoms in the fifth protein chain;

[0030] The step of determining the optimal translational and rotational relationship between the multiple protein chains according to the predicted atomic positions corresponding to the multiple protein chains and the predicted distance map includes: determining the first optimal translational and rotational relationship between the third protein chain and the fourth protein chain according to the predicted atomic position of the third protein chain, the predicted atomic position in the fourth protein chain, and the approximate distance between the second representative atoms of the third protein chain and the fourth protein chain included in the predicted distance map;

[0031] The method further includes:

[0032] Obtaining the predicted atomic position of the first representative atom in the complex;

[0033] Obtaining a composite predicted distance map corresponding to the complex and the fifth protein chain; the composite predicted distance map is determined according to the amino acid sequence of the complex and the amino acid sequence of the fifth protein chain, and the composite predicted distance map is used to represent the approximate distance between the second representative atoms of the complex and the second representative atoms of the fifth protein chain;

[0034] Determining the second optimal translational and rotational relationship between the complex and the fifth protein chain according to the predicted atomic position corresponding to the complex, the predicted atomic position corresponding to the fifth protein chain, and the composite predicted distance map.

[0035] Optionally, obtaining the predicted structure of the protein complex according to the optimal translation and rotation relationship includes:

[0036] Performing coordinate transformation on the predicted atomic positions corresponding to the protein chains according to the optimal translation and rotation relationship to obtain the predicted structure of the protein complex.

[0037] In a second aspect, an embodiment of the present application provides a protein complex structure prediction device, and the device includes:

[0038] A predicted atomic position determination module, configured to input the amino acid sequences of multiple protein chains in a protein complex into a trained position prediction model to obtain the predicted atomic positions of the first representative atoms of each protein chain;

[0039] An optimal translation and rotation relationship determination module, configured to determine the optimal translation and rotation relationship between multiple protein chains according to the predicted atomic positions corresponding to the multiple protein chains and the predicted distance map, so that the distance difference between the calculated distance map of the multiple protein chains under the optimal translation and rotation relationship and the predicted distance map is minimized; wherein, the predicted distance map is determined according to the amino acid sequences of the multiple protein chains and is used to characterize the approximate distances between the second representative atoms of the multiple protein chains; the calculated distance map is determined according to the predicted atomic positions corresponding to the multiple protein chains and the translation and rotation relationship between the multiple protein chains; the second representative atoms and the first representative atoms are the same or different;

[0040] A protein complex structure prediction module, configured to obtain the predicted structure of the protein complex according to the optimal translation and rotation relationship.

[0041] Optionally, the device further includes:

[0042] A quasi-predicted distance map determination module, configured to input the amino acid sequences of the multiple protein chains into a trained target map prediction network to obtain a quasi-predicted distance map before determining the optimal translation and rotation relationship between the multiple protein chains according to the predicted atomic positions corresponding to the multiple protein chains and the predicted distance map, wherein the quasi-predicted distance map includes at least one tensor of M*N*P, where P is the number of distance classifications, and the element at the mnth position of the pth layer of the matrix represents the probability that the predicted distance between the mth first representative atom of the first protein chain and the nth first representative atom of the second protein chain falls into the pth distance classification, and different distance classifications correspond to different distance intervals; where p is an integer selected from 1 to P; m is an integer selected from 1 to M, and n is an integer selected from 1 to N;

[0043] A predicted distance map determination module, configured to obtain the predicted distance map according to the quasi-predicted distance map; wherein, the predicted distance map includes the approximate distance between the m-th first representative atom of the first protein chain and the n-th first representative atom of the second protein chain, and the approximate distance is represented by the category corresponding to the maximum probability of the m-th first representative atom of the first protein chain and the n-th first representative atom of the second protein chain.

[0044] Optionally, the target map prediction network includes an evolutionary network module and a map generation module;

[0045] A paired feature determination module, configured to input the amino acid sequences of multiple training protein chains included in a protein complex with a known structure into the evolutionary network module during the training of the target map prediction network to obtain paired features;

[0046] A predicted atomic approximate distance determination module, configured to input the paired features into the map generation module to obtain a predicted atomic approximate distance;

[0047] A target map prediction network determination module, configured to adjust the parameters of the evolutionary network module and / or the map generation module according to the difference value between the predicted atomic approximate distance and the actual atomic approximate distance to obtain the target map prediction network.

[0048] Optionally, when the optimal translation and rotation relationship determination module is used to determine the optimal translation and rotation relationship between multiple protein chains according to the predicted atomic positions and the predicted distance map corresponding to the multiple protein chains, it is specifically configured to:

[0049] When the second representative atom is different from the first representative atom, convert the calculated distance map into a converted calculated distance map, which is characterized by the distance between the second representative atoms; the distance difference between the converted calculated distance map of the multiple protein chains under the optimal translation and rotation relationship and the predicted distance map is the smallest;

[0050] Or, convert the predicted distance map into a converted predicted distance map, which is characterized by the distance between the first representative atoms; the distance difference between the calculated distance map of the multiple protein chains under the optimal translation and rotation relationship and the converted predicted distance map is the smallest.

[0051] Optionally, when the optimal translation and rotation relationship determination module is used to determine the optimal translation and rotation relationship between multiple protein chains according to the predicted atomic positions and the predicted distance map corresponding to the multiple protein chains, it is specifically configured to:

[0052] The optimal translation and rotation relationship is obtained according to the following formula:

[0053]

[0054] Wherein, x A is the predicted atomic position of a protein chain, and x B is the predicted atomic position of another protein chain, C is the predicted distance map, T is the translation and rotation relationship, R is the rotation matrix, and t is the translation matrix.

[0055] Optionally, when the protein complex includes W protein chains, the optimal translation and rotation relationship includes W - 1 translation and rotation relationships, where W is greater than or equal to 3;

[0056] The W - 1 translation and rotation relationships minimize the weighted sum of the distance differences between the calculated distance maps between pairs of the multiple protein chains and the predicted distance map;

[0057] Alternatively, each of the W - 1 translation and rotation relationships minimizes the distance difference between the calculated distance map between the two protein chains representing the relative relationship and the predicted distance map.

[0058] Optionally, when the protein complex includes a third protein chain, a fourth protein chain, and a fifth protein chain, the optimal translation and rotation relationship includes a first optimal translation and rotation relationship between the third protein chain and the fourth protein chain and a second optimal translation and rotation relationship between the complex and the fifth protein chain; the complex is the complex formed by the third protein chain and the fourth protein chain under the first optimal translation and rotation relationship;

[0059] When the predicted atomic position determination module is used to input the amino acid sequences of multiple protein chains in the protein complex into the trained position prediction model to obtain the predicted atomic positions of the first representative atoms in the multiple protein chains, it is specifically used for:

[0060] Input the amino acid sequences of the third protein chain, the fourth protein chain, and the fifth protein chain into the trained position prediction model to obtain the predicted atomic positions of the first representative atoms in the third protein chain, the predicted atomic positions of the first representative atoms in the fourth protein chain, and the predicted atomic positions of the first representative atoms in the fifth protein chain;

[0061] When the optimal translation and rotation relationship determination module is used to determine the optimal translation and rotation relationship between multiple protein chains according to the predicted atomic positions and the predicted distance map corresponding to the multiple protein chains, it is specifically used for:

[0062] Determine the first optimal translational rotation relationship between the third protein chain and the fourth protein chain based on the predicted atomic positions of the third protein chain, the predicted atomic positions in the fourth protein chain, and the approximate distances between the second representative atoms of the third protein chain and the fourth protein chain included in the predicted distance map;

[0063] The apparatus further includes:

[0064] A complex predicted atomic position determination module, configured to obtain the predicted atomic positions of the first representative atoms in the complex;

[0065] A complex predicted distance map determination module, configured to obtain a complex predicted distance map corresponding to the complex and the fifth protein chain; the complex predicted distance map is determined according to the amino acid sequence of the complex and the amino acid sequence of the fifth protein chain, and the complex predicted distance map is used to characterize the approximate distances between the second representative atoms of the complex and the second representative atoms of the fifth protein chain;

[0066] A second optimal translational rotation relationship determination module, configured to determine a second optimal translational rotation relationship between the complex and the fifth protein chain according to the predicted atomic positions corresponding to the complex, the predicted atomic positions corresponding to the fifth protein chain, and the complex predicted distance map.

[0067] Optionally, when the protein complex structure prediction module is used to obtain the predicted structure of the protein complex according to the optimal translational rotation relationship, it is specifically configured to:

[0068] Perform coordinate transformation on the predicted atomic positions corresponding to the protein chain according to the optimal translational rotation relationship to obtain the predicted structure of the protein complex.

[0069] In a third aspect, an embodiment of the present application provides a computer device, including: a processor, a memory, and a bus, where the memory stores machine-readable instructions executable by the processor. When the computer device runs, the processor communicates with the memory through the bus, and when the machine-readable instructions are executed by the processor, the steps of the protein complex structure prediction method described in any optional implementation manner in the first aspect above are executed.

[0070] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, the steps of the protein complex structure prediction method described in any optional implementation manner in the first aspect above are executed.

[0071] The technical solutions provided by the embodiments of the present application include but are not limited to the following beneficial effects:

[0072] In the embodiments of the present application, the predicted atomic positions of the representative atoms of each protein chain and the predicted distance map including the approximate distances between the representative atoms of different protein chains are first obtained, and then the translational and rotational relationships between different protein chains are solved, so that the difference between the calculated distance map determined according to the predicted atomic positions and the translational and rotational relationships and the predicted distance map is minimized under the translational and rotational relationships. Then, the predicted structure of the protein complex is obtained according to the translational and rotational relationships. In this way, the docking relationship of the protein complex can be accurately predicted without using a complex structure prediction network, and the accuracy of the prediction result can be improved while greatly reducing the model complexity and computing power consumption. To make the above objects, features, and advantages of the present invention more obvious and understandable, the following specific embodiments are given in conjunction with the accompanying drawings and are described in detail as follows. Description of the Drawings

[0073] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0074] Figure 1 Shows the flowchart of a method for predicting the structure of a protein complex provided in Embodiment 1 of the present invention;

[0075] Figure 2 Shows the flowchart of a method for determining a predicted distance map provided in Embodiment 1 of the present invention;

[0076] Figure 3 Shows the flowchart of a method for determining a target map prediction network provided in Embodiment 1 of the present invention;

[0077] Figure 4 Shows the flowchart of a method for determining the second optimal translational and rotational relationship provided in Embodiment 1 of the present invention;

[0078] Figure 5 Shows the structural schematic diagram of a device for predicting the structure of a protein complex provided in Embodiment 2 of the present invention;

[0079] Figure 6 Shows the structural schematic diagram of a computer device provided in Embodiment 3 of the present invention. Detailed Embodiments

[0080] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only a part rather than all of the embodiments of the present invention. The components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0081] Embodiment 1

[0082] To facilitate the understanding of this application, the following will combine Figure 1 the content described in the flowchart of a method for predicting the structure of a protein complex provided in Embodiment 1 of the present invention shown in

[0083] Refer to Figure 1 as shown, Figure 1 which shows the flowchart of a method for predicting the structure of a protein complex provided in Embodiment 1 of the present invention. Among them, the method includes steps S101 to S103:

[0084] S101: Input the amino acid sequences of multiple protein chains in the protein complex into the trained position prediction model to obtain the predicted atomic positions of the first representative atoms of each protein chain.

[0085] Specifically, a protein complex is a complex formed by two or more functionally related polypeptide chains through disulfide bonds or other protein interactions. A protein complex includes multiple protein chains; protein chains are, for example, the heavy and light chains of an antibody.

[0086] The position prediction model is used to predict the atomic positions of the first representative atoms in each protein chain of the protein complex to obtain the predicted atomic positions of the first representative atoms. For example, if a protein complex includes 3 protein chains, the 3 protein chains are sent into the position prediction model separately or simultaneously, and the predicted atomic positions of the first representative atoms corresponding to the 3 protein chains can be obtained separately or simultaneously. The position prediction model can be pre-trained and can be any existing protein monomer structure prediction model, such as alphafold2. The prediction result of the position prediction model can include the positions of all atoms in the protein chain. For the convenience of subsequent calculations, the positions of the first representative atoms are selected as representative points.

[0087] S102: Determine the optimal translational and rotational relationship among multiple protein chains based on the predicted atomic positions and predicted distance maps corresponding to the multiple protein chains, such that the distance difference between the calculated distance map of the multiple protein chains under the optimal translational and rotational relationship and the predicted distance map is minimized; wherein, the predicted distance map is determined based on the amino acid sequences of the multiple protein chains and is used to characterize the approximate distances between the second representative atoms of the multiple protein chains; the calculated distance map is determined based on the predicted atomic positions corresponding to the multiple protein chains and the translational and rotational relationship among the multiple protein chains and is used to characterize the distances between the first representative atoms of the multiple protein chains under this translational and rotational relationship; the second representative atoms and the first representative atoms are the same or different.

[0088] It can be understood that when the predicted distance map includes the approximate distances of the second representative atoms of N protein chains, the approximate distances in the predicted distance map are determined based on the amino acid sequences of these N protein chains.

[0089] The predicted distance map contains the approximate distances between the second representative atoms in at least two different protein chains. For example, if a protein complex includes protein chain A, protein chain B, and protein chain C, the predicted distance map can include the approximate distances between the second representative atoms in protein chain A and protein chain B, or can also include the approximate distances between the second representative atoms in protein chain A, protein chain B, and protein chain C.

[0090] For example, if a protein complex includes protein chain A, protein chain B, and protein chain C, and protein chain A includes A 1 、A 2 and A 3 in total three second representative atoms, protein chain B includes B 1 、B 2 and B 3 in total three second representative atoms, and protein chain C includes C 1 、C 2 in total two second representative atoms, then the predicted distance map can be as shown in Table 1 (the distances are only for illustration).

[0091] Table 1

[0092]

[0093] When predicting the approximate distances of the second-generation atoms in more than three different protein chains in the distance map, in addition to the two-dimensional matrix form in Table 1, the predicted distance map can also be a multi-dimensional tensor including the approximate distances of the second-generation atoms in pairwise protein chains. For example, when predicting the approximate distances of the second-generation atoms in three different protein chains A, B, and C in the distance map, one layer represents the approximate distances of the second-generation atoms in AB, and one layer represents the approximate distances of the second-generation atoms in BC. Optionally, there can also be one layer in the predicted distance map representing the approximate distances of the second-generation atoms in AC.

[0094] The translation-rotation relationship represents the relative translation amount / rotation amount between multiple protein chains. The calculated distance map is determined based on the predicted atomic positions corresponding to multiple protein chains and the translation-rotation relationship between multiple protein chains, and is used to characterize the distances between the first-generation atoms of multiple protein chains under this translation-rotation relationship. Assuming that the protein chains are rigid bodies, when the predicted atomic positions of the first-generation atoms of multiple protein chains have been determined, under each translation-rotation relationship, a calculated distance map corresponding to this translation-rotation relationship can be determined, which is used to characterize the distances between the first-generation atoms of multiple protein chains under this translation-rotation relationship.

[0095] The first-generation atoms and the second-generation atoms can be α-carbon atoms, β-carbon atoms (denoted as Cb hereinafter), and can also be hydrogen atoms or oxygen atoms. The first-generation atoms and the second-generation atoms can be the atoms that only exist one in an amino acid.

[0096] The first-generation atoms and the second-generation atoms can be the same or different. When the first-generation atoms and the second-generation atoms are different, the distances in the calculated distance map are characterized by the first-generation atoms, and the predicted distance map is characterized by the second-generation atoms. Since the distance characterization criteria of the calculated distance map and the predicted distance map are different, the difference between the two cannot be calculated. At this time, both need to be uniformly converted to be characterized by the first-generation atoms or the second-generation atoms. Assuming that the protein chains are rigid bodies, the relative positional relationship between the first-generation atoms and the second-generation atoms is determined, so the distances characterized by one kind of representative atom can be converted to be characterized by the other kind of representative atom.

[0097] Exemplarily, the calculated distance map can be converted into a converted calculated distance map, and the converted calculated distance map is characterized by the distances between the second-generation atoms; the distance difference between the converted calculated distance map of multiple protein chains under the optimal translation-rotation relationship and the predicted distance map is the smallest;

[0098] Alternatively, convert the predicted distance map into a converted predicted distance map, which is characterized by the distances between the first representative atoms; the distance difference between the calculated distance maps of multiple protein chains under the optimal translation and rotation relationship and the converted predicted distance map is minimized.

[0099] It can be understood that under each translation and rotation relationship, a calculated distance map corresponding to this translation and rotation relationship can be determined, and thus a distance difference between the calculated distance map and the predicted distance map can be determined. The translation and rotation relationship that minimizes the distance difference is the optimal translation and rotation relationship. The method for the optimal translation and rotation relationship can use methods in the prior art, such as the gradient descent method, etc.

[0100] S103: Obtain the predicted structure of the protein complex according to the optimal translation and rotation relationship.

[0101] Specifically, the optimal translation and rotation relationship is used to indicate the relative translation / rotation amount between multiple protein chains. According to the translation amount and rotation amount, the positions of the atoms in the protein complex are converted to obtain the positions of the atoms under the optimal translation and rotation relationship, and thus the predicted structure of the protein complex can be obtained.

[0102] The predicted distance map can be determined according to a neural network. In a feasible implementation, see Figure 2 shown, Figure 2 shows a flowchart of a method for determining a predicted distance map provided in the first embodiment of the present invention. Among them, before determining the optimal translation and rotation relationship between multiple protein chains according to the predicted atomic positions and the predicted distance map corresponding to the multiple protein chains, the method further includes steps S201 - S202:

[0103] S201: Input the amino acid sequences of multiple protein chains into a trained target map prediction network to obtain a quasi-predicted distance map. Among them, the quasi-predicted distance map includes at least one tensor of M * N * P, where P is the number of distance classifications. The element value of the element at the p-th layer, the m-th row, and the n-th column (i.e., the m-th row and the n-th column or the n-th row and the m-th column) of the matrix represents the probability that the predicted distance between the m-th first representative atom of the first protein chain and the n-th first representative atom of the second protein chain falls into the p-th distance classification. Different distance classifications correspond to different distance intervals; where p is an integer selected from 1 to P; m is an integer selected from 1 to M, and n is an integer selected from 1 to N;

[0104] Specifically, the target map prediction network is used to determine a distance prediction map according to the amino acid sequences of protein chains.

[0105] For example, common distances can be divided into 64 categories. Different distance classifications correspond to different distance intervals, and distances greater than the common distances are uniformly classified as the 64th category. The quasi-predicted distance map can be a tensor of N*N*64. For each element in N*N, it is a 64-classification problem. If the model predicts that the distance between the first representative atom A of protein chain A 1 and the first representative atom B of protein chain B 1 corresponds to the 53rd category, then the value of the 53rd layer in the 64 layers corresponding to A 1 -B 1 is relatively large (e.g., 0.9), and the values of other layer elements are relatively small.

[0106] S202: Obtain the predicted distance map according to the quasi-predicted distance map; wherein, the predicted distance map includes the approximate distance between the mth first representative atom of the first protein chain and the nth first representative atom of the second protein chain, and the approximate distance is represented by the category corresponding to the maximum probability of the mth first representative atom of the first protein chain and the nth first representative atom of the second protein chain.

[0107] Specifically, obtaining the predicted distance map according to the quasi-predicted distance map includes:

[0108] Directly using the quasi-predicted distance map as the predicted distance map; or, compressing the dimension of each layer matrix in the quasi-predicted distance map to obtain a two-dimensional matrix for indicating the distance between the first representative atoms in different protein chains.

[0109] Compared with a complex network that can predict the exact distance between protein chains at one time, the target map prediction network in the embodiments of the present invention only predicts the classification corresponding to the distance, which can greatly reduce the model size and the complexity of training.

[0110] In a feasible implementation, the target map prediction network includes an evolutionary network module (evoformer module) and a map generation module;

[0111] See Figure 3 shown in Figure 3 shows a flowchart of a method for determining a target map prediction network provided in Embodiment 1 of the present invention. Before inputting the amino acid sequences of multiple protein chains into the trained target map prediction network to obtain a quasi-predicted distance map, the method includes steps S301 to S303 of training the target map prediction network:

[0112] S301: Input the amino acid sequences of multiple training protein chains included in a protein complex with a known structure into the evolutionary network module to obtain paired features.

[0113] It is understandable that the amino acid sequences of the training protein chains in the protein complex are known, and the distances between the second-generation atoms between different protein chains are known (determined, for example, by experimental methods or other reliable computational methods), and the distance map (approximate distances of actual atoms) is also known. For example, the actual atomic distances can be classified to obtain approximate distances of actual atoms.

[0114] Specifically, the pairing feature is pair embedding (a matrix whose dimension may be different from that of the predicted distance map); among them, the evolutionary network module may also have protein homologous sequence features, which are not specifically limited here.

[0115] S302: Input the pairing feature into the map generation module to obtain a predicted approximate atomic distance.

[0116] Specifically, the shape of the predicted approximate atomic distance is the same as that of the quasi-predicted distance map or the predicted distance map.

[0117] S303: Adjust the parameters of the evolutionary network module and / or the map generation module according to the difference value between the predicted approximate atomic distance (i.e., the predicted classification) and the actual approximate atomic distance (i.e., the classification corresponding to the actual distance) to obtain the target map prediction network.

[0118] It is understandable that the shapes of the actual approximate atomic distance and the predicted approximate atomic distance are the same.

[0119] Specifically, calculate the difference value between the predicted approximate atomic distance and the actual approximate atomic distance of every two first-generation atoms, and adjust the parameters of the evolutionary network module and / or the map generation module according to this difference value to obtain the target map prediction network.

[0120] In a feasible implementation, when the second-generation atoms are different from the first-generation atoms, the determining of the optimal translation and rotation relationship between multiple protein chains according to the predicted atomic positions and the predicted distance map corresponding to the multiple protein chains includes:

[0121] Convert the calculated distance map into a converted calculated distance map, which is characterized by the distances between the second-generation atoms; the distance difference between the converted calculated distance maps of the multiple protein chains under the optimal translation and rotation relationship and the predicted distance map is the smallest.

[0122] Alternatively, convert the predicted distance map into a converted predicted distance map, which is characterized by the distances between the first representative atoms; the distance difference between the calculated distance maps of multiple protein chains under the optimal translation and rotation relationship and the converted predicted distance map is minimized.

[0123] Specifically, since the first representative atoms and the second representative atoms can be converted into each other according to their positions, if it is necessary to characterize with the second representative atoms, convert the calculated distance map into a converted calculated distance map, which is characterized by the distances between the second representative atoms; if it is necessary to characterize with the first representative atoms, convert the predicted distance map into a converted predicted distance map, which is characterized by the distances between the first representative atoms; whether characterized by the first representative atoms or the second representative atoms, it is necessary to minimize the distance difference between the calculated distance maps of multiple protein chains under the optimal translation and rotation relationship and the converted predicted distance map.

[0124] In a feasible embodiment, determining the optimal translation and rotation relationship between multiple protein chains according to the predicted atomic positions and predicted distance maps corresponding to the multiple protein chains includes:

[0125] Obtain the optimal translation and rotation relationship according to the following formula:

[0126]

[0127] where x A is the predicted atomic position of one protein chain, x B is the predicted atomic position of another protein chain, C is the predicted distance map, T is the translation and rotation relationship, R is the rotation matrix, and t is the translation matrix.

[0128] Specifically, the gradient descent method can be used to solve the minimum value of the difference value.

[0129] is the calculated distance map. The conversion of the positions of the first representative atoms can be carried out at the levels of x A , x B and C, or can also be carried out after obtaining the calculated distance map.

[0130] In a feasible embodiment, when the protein complex includes W protein chains, the optimal translation and rotation relationship includes W - 1 translation and rotation relationships, where W is greater than or equal to 3.

[0131] It is understandable that when there are W protein chains in the protein complex, there is a translational and rotational relationship corresponding to each pair of protein chains, but no loop is formed. Therefore, there are a total of W - 1 translational and rotational relationships. That is to say, the optimal translational and rotational relationship can be a set composed of W - 1 translational and rotational relationships.

[0132] The W - 1 translational and rotational relationships minimize the weighted sum of the distance differences between the calculated distance maps of multiple pairs of the protein chains and the predicted distance map.

[0133] For example, for a protein complex composed of three protein chains ABC, the optimal translational and rotational relationship T1 between AB and the optimal translational and rotational relationship T2 between BC need to be found. T1 and T2 constitute the optimal translational and rotational relationship T. When the distance difference d1 between the calculated distance map (here referring to the calculated distance map corresponding to A and B, calculated using the predicted atomic positions of A and B and T11) and the predicted distance map (here referring to the predicted distance map corresponding to A and B, which can be obtained by inputting the amino acid sequences of A and B into the model) under a certain translational and rotational relationship T11 between AB, and the distance difference d2 between the calculated distance map (here referring to the calculated distance map corresponding to B and C, calculated using the predicted atomic positions of B and C and T21) and the predicted distance map (here referring to the predicted distance map corresponding to B and C, which can be obtained by inputting the amino acid sequences of B and C into the model) under a certain translational and rotational relationship T21 between BC, the weighted sum is minimized, then T11 is considered as T1 and T21 is considered as T2. That is, T1 and T2 minimize the weighted sum of d1 and d2.

[0134] Alternatively, each of the W - 1 translational and rotational relationships minimizes the distance difference between the calculated distance map of the two protein chains representing the relative relationship and the predicted distance map.

[0135] In this case, T1 and T2 are decoupled. That is, T1 minimizes d1 and T2 minimizes d2.

[0136] In this embodiment, when there are 4 protein chains ABCD, the translational and rotational relationship T1 between AB can be determined first, then the translational and rotational relationship T2 between BC can be determined, and then the translational and rotational relationship T3 between CD can be determined. At this time, the change brought to B after the combination of AB is not considered.

[0137] In a feasible embodiment, when the protein complex includes a third protein chain, a fourth protein chain, and a fifth protein chain, the optimal translational and rotational relationship includes a first optimal translational and rotational relationship between the third protein chain and the fourth protein chain and a second optimal translational and rotational relationship between the complex and the fifth protein chain; the complex is a complex formed by the third protein chain and the fourth protein chain under the first optimal translational and rotational relationship.

[0138] Step S101 includes: inputting the amino acid sequences of the third protein chain, the fourth protein chain, and the fifth protein chain into the trained position prediction model to obtain the predicted atomic positions of the first representative atoms in the third protein chain, the predicted atomic positions of the first representative atoms in the fourth protein chain, and the predicted atomic positions of the first representative atoms in the fifth protein chain.

[0139] Specifically, when the protein complex includes three protein chains, such as protein chain A, protein chain B, and protein chain C, input the amino acid sequences corresponding to protein chain A, protein chain B, and protein chain C into the trained position prediction model to obtain the predicted atomic positions of the first representative atoms of protein chain A, the predicted atomic positions of the first representative atoms of protein chain B, and the predicted atomic positions of the first representative atoms of protein chain C.

[0140] Step S102 includes: determining the first optimal translational and rotational relationship between the third protein chain and the fourth protein chain according to the predicted atomic positions of the third protein chain, the predicted atomic positions in the fourth protein chain, and the approximate distances between the second representative atoms of the third protein chain and the fourth protein chain included in the predicted distance map.

[0141] Continuing with the previous example, according to the predicted atomic positions of protein chain A, the predicted atomic positions of protein chain B, and the predicted distance map, determine the first optimal translational and rotational relationship T1 between the third protein chain and the fourth protein chain in the manner of step S102. The predicted distance map here only needs to include the approximate distances between the second representative atoms corresponding to protein chain A and protein chain B, and does not need to include the approximate distances between the second representative atoms of AC or BC. Protein chain A and protein chain B form complex AB under T1.

[0142] See Figure 4 as shown Figure 4 shows a flowchart of a method for determining a second optimal translational and rotational relationship provided in the first embodiment of the present invention. The method further includes steps S401 to S403:

[0143] S401: Obtain the predicted atomic positions of the first representative atoms in the complex.

[0144] For example, given the predicted atomic positions of protein chain A, the predicted atomic positions of protein chain B, and T1, the predicted atomic positions of the first representative atoms in the complex can be calculated.

[0145] S402: Obtain the composite predicted distance map corresponding to the complex and the fifth protein chain; the composite predicted distance map is determined based on the amino acid sequence of the complex and the amino acid sequence of the fifth protein chain, and the composite predicted distance map is used to characterize the approximate distance between the second representative atoms of the complex and the second representative atoms of the fifth protein chain.

[0146] For example, input the amino acid sequence of the complex and the amino acid sequence of protein chain C into the target map prediction network to obtain a composite predicted distance map, which is used to characterize the approximate distance between the second representative atoms of the complex and the second representative atoms of protein chain C.

[0147] S403: Determine the second optimal translational and rotational relationship between the complex and the fifth protein chain based on the predicted atomic positions corresponding to the complex, the predicted atomic positions corresponding to the fifth protein chain, and the composite predicted distance map.

[0148] Based on the predicted atomic positions of complex AB, the predicted atomic positions of protein chain C, and the composite predicted distance map, determine the second optimal translational and rotational relationship in the manner of step 102.

[0149] In this way, the calculation of the translational and rotational relationship between more than three protein chains can be decomposed into multiple calculations of the translational and rotational relationship between two by two. For example, when there are 4 protein chains including A, B, C, and D, the translational and rotational relationship T1 between A and B can be determined first, then the translational and rotational relationship T2 between AB and C can be determined, and then the translational and rotational relationship T3 between ABC and D can be determined. In this way, the changes in the protein chain after combination are considered.

[0150] In a feasible implementation, the obtaining of the predicted structure of the protein complex according to the optimal translational and rotational relationship includes:

[0151] Perform coordinate transformation on the predicted atomic positions corresponding to the protein chain according to the optimal translational and rotational relationship to obtain the predicted structure of the protein complex.

[0152] Specifically, the optimal translational and rotational relationship includes the translation amount and rotation amount that the atoms in the protein chain need to move. Perform coordinate transformation on the predicted atomic positions corresponding to the protein chain according to the translation amount and rotation amount to obtain the target atomic positions, and combine each atom according to its target atomic positions to construct the predicted structure of the protein complex.

[0153] Example 2

[0154] See Figure 5 as shown Figure 5 As shown, a schematic structural diagram of a protein complex structure prediction device provided in Example 2 of the present invention is shown. Among them, as Figure 5 shown, a protein complex structure prediction device provided in Example 2 of the present invention includes:

[0155] A predicted atom position determination module 501, configured to input the amino acid sequences of multiple protein chains in a protein complex into a trained position prediction model, and obtain the predicted atom positions of the first representative atoms of each of the protein chains;

[0156] An optimal translation and rotation relationship determination module 502, configured to determine the optimal translation and rotation relationship between multiple protein chains according to the predicted atom positions and predicted distance maps corresponding to the multiple protein chains, so that the distance difference between the calculated distance map of the multiple protein chains under the optimal translation and rotation relationship and the predicted distance map is minimized; wherein, the predicted distance map is determined according to the amino acid sequences of the multiple protein chains and is used to characterize the approximate distances between the second representative atoms of the multiple protein chains; the calculated distance map is determined according to the predicted atom positions corresponding to the multiple protein chains and the translation and rotation relationship between the multiple protein chains; the second representative atoms and the first representative atoms are the same or different;

[0157] A protein complex structure prediction module 503, configured to obtain the predicted structure of the protein complex according to the optimal translation and rotation relationship.

[0158] In a feasible implementation, the device further includes:

[0159] A quasi-predicted distance map determination module, configured to input the amino acid sequences of the multiple protein chains into a trained target map prediction network to obtain a quasi-predicted distance map before determining the optimal translation and rotation relationship between the multiple protein chains according to the predicted atom positions and predicted distance maps corresponding to the multiple protein chains. The quasi-predicted distance map includes at least one M*N*P matrix, where P is the number of distance classifications, and the element at the mnth position of the pth layer of the matrix represents the probability that the predicted distance between the mth first representative atom of the first protein chain and the nth first representative atom of the second protein chain falls into the pth distance classification. Different distance classifications correspond to different distance intervals; where p is an integer selected from 1 to P; m is an integer selected from 1 to M, and n is an integer selected from 1 to N;

[0160] A predicted distance map determination module, configured to obtain the predicted distance map based on the quasi-predicted distance map; wherein, the predicted distance map includes the approximate distance between the m-th first representative atom of the first protein chain and the n-th first representative atom of the second protein chain, and the approximate distance is represented by the category corresponding to the maximum probability of the m-th first representative atom of the first protein chain and the n-th first representative atom of the second protein chain.

[0161] In a feasible implementation, the target map prediction network includes an evolutionary network module and a map generation module;

[0162] A paired feature determination module, configured to input the amino acid sequences of multiple training protein chains included in a protein complex with a known structure into the evolutionary network module to obtain paired features when training the target map prediction network;

[0163] A predicted atom approximate distance determination module, configured to input the paired features into the map generation module to obtain a predicted atom approximate distance;

[0164] A target map prediction network determination module, configured to adjust the parameters of the evolutionary network module and / or the map generation module according to the difference value between the predicted atom approximate distance and the actual atom approximate distance to obtain the target map prediction network.

[0165] In a feasible implementation, the optimal translation and rotation relationship determination module is configured to determine the optimal translation and rotation relationship between multiple protein chains according to the predicted atom positions and the predicted distance map corresponding to the multiple protein chains, and specifically is configured to:

[0166] When the second representative atom is different from the first representative atom, convert the calculated distance map into a converted calculated distance map, which is characterized by the distance between the second representative atoms; the distance difference between the converted calculated distance map of the multiple protein chains under the optimal translation and rotation relationship and the predicted distance map is the smallest;

[0167] Or, convert the predicted distance map into a converted predicted distance map, which is characterized by the distance between the first representative atoms; the distance difference between the calculated distance map of the multiple protein chains under the optimal translation and rotation relationship and the converted predicted distance map is the smallest.

[0168] In a feasible implementation, the optimal translation and rotation relationship determination module is configured to determine the optimal translation and rotation relationship between multiple protein chains according to the predicted atom positions and the predicted distance map corresponding to the multiple protein chains, and specifically is configured to:

[0169] The optimal translation and rotation relationship is obtained according to the following formula:

[0170]

[0171] where x A is the predicted atomic position of a protein chain, and x B is the predicted atomic position of another protein chain, C is the predicted distance map, T is the translation and rotation relationship, R is the rotation matrix, and t is the translation matrix.

[0172] In a feasible embodiment, when the protein complex includes W protein chains, the optimal translation and rotation relationship includes W - 1 translation and rotation relationships, where W is greater than or equal to 3;

[0173] The W - 1 translation and rotation relationships minimize the weighted sum of the distance differences between the calculated distance maps between pairs of the multiple protein chains and the predicted distance map;

[0174] Alternatively, each of the W - 1 translation and rotation relationships minimizes the distance difference between the calculated distance map between the two protein chains representing the relative relationship and the predicted distance map.

[0175] In a feasible embodiment, when the protein complex includes a third protein chain, a fourth protein chain, and a fifth protein chain, the optimal translation and rotation relationship includes a first optimal translation and rotation relationship between the third protein chain and the fourth protein chain and a second optimal translation and rotation relationship between the complex and the fifth protein chain; the complex is the complex formed by the third protein chain and the fourth protein chain under the first optimal translation and rotation relationship;

[0176] When the predicted atomic position determination module is used to input the amino acid sequences of multiple protein chains in the protein complex into the trained position prediction model to obtain the predicted atomic positions of the first representative atoms in the multiple protein chains, it is specifically used for:

[0177] Input the amino acid sequences of the third protein chain, the fourth protein chain, and the fifth protein chain into the trained position prediction model to obtain the predicted atomic positions of the first representative atoms in the third protein chain, the predicted atomic positions of the first representative atoms in the fourth protein chain, and the predicted atomic positions of the first representative atoms in the fifth protein chain;

[0178] When the optimal translation and rotation relationship determination module is used to determine the optimal translation and rotation relationship between multiple protein chains according to the predicted atomic positions and the predicted distance map corresponding to the multiple protein chains, it is specifically used for:

[0179] Determine the first optimal translational and rotational relationship between the third protein chain and the fourth protein chain based on the predicted atomic positions of the third protein chain, the predicted atomic positions in the fourth protein chain, and the approximate distances between the second representative atoms of the third protein chain and the fourth protein chain included in the predicted distance map.

[0180] The apparatus further includes:

[0181] A complex predicted atomic position determination module, configured to obtain the predicted atomic positions of the first representative atoms in the complex;

[0182] A complex predicted distance map determination module, configured to obtain a complex predicted distance map corresponding to the complex and the fifth protein chain; the complex predicted distance map is determined according to the amino acid sequence of the complex and the amino acid sequence of the fifth protein chain, and the complex predicted distance map is used to characterize the approximate distances between the second representative atoms of the complex and the second representative atoms of the fifth protein chain.

[0183] A second optimal translational and rotational relationship determination module, configured to determine a second optimal translational and rotational relationship between the complex and the fifth protein chain according to the predicted atomic positions corresponding to the complex, the predicted atomic positions corresponding to the fifth protein chain, and the complex predicted distance map.

[0184] In a feasible implementation, when the protein complex structure prediction module is used to obtain the predicted structure of the protein complex according to the optimal translational and rotational relationship, it is specifically configured to:

[0185] Perform coordinate transformation on the predicted atomic positions corresponding to the protein chain according to the optimal translational and rotational relationship to obtain the predicted structure of the protein complex.

[0186] Embodiment III

[0187] Based on the same application concept, see Figure 6 as shown Figure 6 shows a schematic structural diagram of a computer device provided in Embodiment III of the present invention, where, as Figure 6 shown, a computer device 600 provided in Embodiment III of the present application includes:

[0188] A processor 601, a memory 602, and a bus 603. The memory 602 stores machine-readable instructions executable by the processor 601. When the computer device 600 runs, communication is performed between the processor 601 and the memory 602 through the bus 603. When the machine-readable instructions are run by the processor 601, the steps of the protein complex structure prediction method shown in Embodiment I above are executed.

[0189] Example 4

[0190] Based on the same inventive concept, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is run by a processor, it executes the steps of the protein complex structure prediction method described in any one of the above embodiments.

[0191] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems and devices described above can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.

[0192] The computer program product for predicting the structure of a protein complex provided by an embodiment of the present invention includes a computer-readable storage medium storing program codes, and the instructions included in the program codes can be used to execute the methods described in the foregoing method embodiments. For specific implementation, reference can be made to the method embodiments, and details will not be described herein again.

[0193] The protein complex structure prediction device provided by an embodiment of the present invention can be specific hardware on a device or software or firmware installed on the device. The implementation principle and the technical effects produced by the device provided by the embodiment of the present invention are the same as those of the foregoing method embodiments. For a brief description, for the parts not mentioned in the device embodiment, reference can be made to the corresponding content in the foregoing method embodiments. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can all refer to the corresponding processes in the above method embodiments, and will not be described herein again.

[0194] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some communication interfaces. The indirect couplings or communication connections of the devices or units can be electrical, mechanical, or other forms.

[0195] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0196] In addition, each functional unit in the embodiments provided by the present invention may be integrated into one processing unit, may exist physically alone for each unit, or two or more units may be integrated into one unit.

[0197] If the above-mentioned function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0198] It should be noted that: similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. In addition, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0199] Finally, it should be noted that: the above-mentioned embodiments are only specific implementation manners of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements for some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention. All should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for predicting the structure of a protein complex, characterized in that, the method comprises: Inputting the amino acid sequences of multiple protein chains in the protein complex into a trained position prediction model to obtain the predicted atomic positions of the first representative atoms of each of the protein chains; Determining the optimal translational and rotational relationship between the multiple protein chains according to the predicted atomic positions corresponding to the multiple protein chains and the predicted distance map, such that the distance difference between the calculated distance map of the multiple protein chains under the optimal translational and rotational relationship and the predicted distance map is minimized; wherein, the predicted distance map is determined according to the amino acid sequences of the multiple protein chains and is used to characterize the approximate distances between the second representative atoms of the multiple protein chains; the calculated distance map is determined according to the predicted atomic positions corresponding to the multiple protein chains and the translational and rotational relationship between the multiple protein chains and is used to characterize the distances between the first representative atoms of the multiple protein chains under this translational and rotational relationship; the second representative atoms and the first representative atoms are the same or different; Obtaining the predicted structure of the protein complex according to the optimal translational and rotational relationship.

2. The method according to claim 1, characterized in that, before determining the optimal translational and rotational relationship between the multiple protein chains according to the predicted atomic positions corresponding to the multiple protein chains and the predicted distance map, the method further comprises: Inputting the amino acid sequences of the multiple protein chains into a trained target map prediction network to obtain a quasi-predicted distance map, wherein the quasi-predicted distance map includes at least one M*N*P matrix, where P is the number of distance classifications, and the element at the p-th layer, the m-th row, and the n-th column of the matrix represents the probability that the predicted distance between the m-th first representative atom of the first protein chain and the n-th first representative atom of the second protein chain falls into the p-th distance classification, and different distance classifications correspond to different distance intervals; wherein p is an integer selected from 1 to P; m is an integer selected from 1 to M, and n is an integer selected from 1 to N; Obtaining the predicted distance map according to the quasi-predicted distance map; wherein, the predicted distance map includes the approximate distance between the m-th first representative atom of the first protein chain and the n-th first representative atom of the second protein chain, and the approximate distance is represented by the category corresponding to the maximum probability of the m-th first representative atom of the first protein chain and the n-th first representative atom of the second protein chain.

3. The method according to claim 2, characterized in that, the target map prediction network includes an evolutionary network module and a map generation module; the target map prediction network is trained through the following steps: Inputting the amino acid sequences of multiple training protein chains included in a protein complex with a known structure into the evolutionary network module to obtain paired features; Inputting the paired features into the map generation module to obtain predicted atomic approximate distances; Adjust the parameters of the evolutionary network module and / or the map generation module according to the difference value between the predicted atomic approximate distance and the actual atomic approximate distance, so as to obtain the target map prediction network.

4. The method according to any one of claims 1-3, wherein, when the second representative atom is different from the first representative atom, the determining the optimal translational and rotational relationship between the plurality of protein chains according to the predicted atomic positions and the predicted distance map corresponding to the plurality of protein chains includes: Converting the calculated distance map into a converted calculated distance map, which is characterized by the distances between the second representative atoms; the distance difference between the converted calculated distance maps of the plurality of protein chains under the optimal translational and rotational relationship and the predicted distance map is the smallest; Alternatively, converting the predicted distance map into a converted predicted distance map, which is characterized by the distances between the first representative atoms; the distance difference between the calculated distance maps of the plurality of protein chains under the optimal translational and rotational relationship and the converted predicted distance map is the smallest.

5. The method according to any one of claims 1-4, wherein, the determining the optimal translational and rotational relationship between the plurality of protein chains according to the predicted atomic positions and the predicted distance map corresponding to the plurality of protein chains includes: Obtaining the optimal translational and rotational relationship according to the following formula: where x A is the predicted atomic position of one protein chain, x B is the predicted atomic position of another protein chain, C is the predicted distance map, T is the translational rotation relationship, R is the rotation matrix, and t is the translation matrix.

6. The method according to any one of claims 1-5, wherein, when the protein complex includes W protein chains, the optimal translational and rotational relationship includes W-1 translational and rotational relationships, where W is greater than or equal to 3; the W-1 translational and rotational relationships minimize the weighted sum of the distance differences between the calculated distance maps and the predicted distance maps between the plurality of protein chains pairwise; Alternatively, each of the W-1 translational and rotational relationships minimizes the distance difference between the calculated distance map and the predicted distance map between the two protein chains representing the relative relationship.

7. The method according to any one of claims 1-5, wherein, when the protein complex includes a third protein chain, a fourth protein chain, and a fifth protein chain, the optimal translational and rotational relationship includes a first optimal translational and rotational relationship between the third protein chain and the fourth protein chain and a second optimal translational and rotational relationship between the complex and the fifth protein chain; the complex is a complex formed by the third protein chain and the fourth protein chain under the first optimal translational and rotational relationship; Inputting the amino acid sequences of multiple protein chains in the protein complex into the trained position prediction model to obtain the predicted atomic positions of the first representative atoms in the multiple protein chains, including: inputting the amino acid sequence of the third protein chain, the amino acid sequence of the fourth protein chain, and the amino acid sequence of the fifth protein chain into the trained position prediction model to obtain the predicted atomic positions of the first representative atoms in the third protein chain, the predicted atomic positions of the first representative atoms in the fourth protein chain, and the predicted atomic positions of the first representative atoms in the fifth protein chain; Determining the optimal translational and rotational relationship between the multiple protein chains according to the predicted atomic positions corresponding to the multiple protein chains and the predicted distance map, including: determining the first optimal translational and rotational relationship between the third protein chain and the fourth protein chain according to the predicted atomic position of the third protein chain, the predicted atomic position in the fourth protein chain, and the approximate distance between the second representative atoms of the third protein chain and the fourth protein chain included in the predicted distance map; The method further includes: Obtaining the predicted atomic positions of the first representative atoms in the complex; Obtaining the composite predicted distance map corresponding to the complex and the fifth protein chain; the composite predicted distance map is determined according to the amino acid sequence of the complex and the amino acid sequence of the fifth protein chain, and the composite predicted distance map is used to characterize the approximate distance between the second representative atoms of the complex and the second representative atoms of the fifth protein chain; Determining the second optimal translational and rotational relationship between the complex and the fifth protein chain according to the predicted atomic position corresponding to the complex, the predicted atomic position corresponding to the fifth protein chain, and the composite predicted distance map.

8. The method according to any one of claims 1-7, wherein, Obtaining the predicted structure of the protein complex according to the optimal translational and rotational relationship, including: Performing coordinate transformation on the predicted atomic positions corresponding to the protein chain according to the optimal translational and rotational relationship to obtain the predicted structure of the protein complex.

9. A protein complex structure prediction device, wherein, The device includes; A predicted atomic position determination module, configured to input the amino acid sequences of multiple protein chains in the protein complex into the trained position prediction model to obtain the predicted atomic positions of the first representative atoms of each protein chain; An optimal translation and rotation relationship determination module, configured to determine an optimal translation and rotation relationship between multiple protein chains according to the predicted atomic positions and predicted distance maps corresponding to the multiple protein chains, so that the distance difference between the calculated distance map of the multiple protein chains under the optimal translation and rotation relationship and the predicted distance map is minimized; wherein, the predicted distance map is determined according to the amino acid sequences of the multiple protein chains and is used to characterize the approximate distances between the second-generation atoms of the multiple protein chains; the calculated distance map is determined according to the predicted atomic positions corresponding to the multiple protein chains and the translation and rotation relationship between the multiple protein chains; the second-generation atoms and the first-generation atoms are the same or different; A protein complex structure prediction module, configured to obtain a predicted structure of the protein complex according to the optimal translation and rotation relationship.

10. A computer device, characterized in that, it includes: A processor, a memory and a bus, the memory stores machine-readable instructions executable by the processor, when the computer device runs, the processor communicates with the memory through the bus, and when the machine-readable instructions are executed by the processor, the steps of the protein complex structure prediction method according to any one of claims 1-8 are executed.

11. A computer-readable storage medium, characterized in that, a computer program is stored on the computer-readable storage medium, and when the computer program is run by a processor, the steps of the protein complex structure prediction method according to any one of claims 1-8 are executed.

Citation Information

Patent Citations

  • Template-based multi-domain protein structure assembling method

    CN107180164A

  • Gray BP neural network based protein interaction prediction method

    CN108427867A