Method and apparatus for predicting the structure of protein complexes

The N-stage folding iterative network layer enhances protein complex structure prediction by considering monomer chain independence, improving accuracy and efficiency in predicting multiple chain scenarios.

JP7857363B2Active Publication Date: 2026-05-12BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
BEIJING BAIDU NETCOM SCI & TECH CO LTD
Filing Date
2024-08-27
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing methods for predicting protein complex structures are inefficient and lack accuracy, particularly in scenarios involving multiple protein chains.

Method used

A method and apparatus using an N-stage folding iterative network layer to predict protein complex structures by inputting initial coordinates and multiple sequence alignment features, considering relative independence of monomer chains through residue-level and monomer chain-level positional transformations.

Benefits of technology

Improves the efficiency and accuracy of protein complex structure prediction, especially in cases with multiple chains, by accurately updating residue positions and maintaining interchain docking relationships.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007857363000008
    Figure 0007857363000008
  • Figure 0007857363000009
    Figure 0007857363000009
  • Figure 0007857363000010
    Figure 0007857363000010
Patent Text Reader

Abstract

To provide a method and a device for predicting the structure of a protein complex that accurately predict the structure of the protein to be predicted and improve efficiency of structure prediction of the protein complex.SOLUTION: A method comprises the following steps: acquiring an initial coordinate of each amino acid residue in a target protein complex, and acquiring a target residue pair feature, a first multi-sequence alignment feature and a second multi-sequence alignment feature of each protein monomer in the target protein complex; inputting the initial coordinate of each amino acid residue, the target residue pair feature of each protein monomer, the first multi-sequence alignment feature and the second multi-sequence alignment feature into an N-level folding iterative network layer, and predicting the torsion angle of each amino acid residue, the position transformation of the residue level and the position transformation of the monomer chain level by the N-level folded iterative network layer to obtain a target coordinate of each amino acid residue and obtain a prediction structure of the protein complex.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This disclosure relates to the field of artificial intelligence technology, particularly to technologies such as natural language processing and biocomputing. [Background technology]

[0002] Protein complexes are stable macromolecular complexes formed by the interaction of two or more protein molecules, and they play important roles in various biological functions such as enzymatic reactions, cell signaling, metabolic regulation, and gene expression. Here, the function of a protein is largely determined by the spatial structure of the protein itself, and the technique of predicting the three-dimensional structure (tertiary structure) of a protein in space based on the amino acid category (primary structure) of the protein chain has extremely high research value in the life sciences field.

[0003] Therefore, accurately predicting protein structures, improving the efficiency of protein complex structure prediction, and addressing various biological applications has become a key research area. [Overview of the project] [Problems that the invention aims to solve]

[0004] This disclosure provides a method and apparatus for predicting the structure of protein complexes. [Means for solving the problem]

[0005] According to one aspect of this disclosure, a method for predicting the structure of a protein complex is provided, and this method is: The steps include obtaining the initial coordinates of each amino acid residue in the target protein complex, and obtaining the target residue pair features, first multiplex alignment features, and second multiplex alignment features of each protein monomer in the target protein complex, Inputting the initial coordinates of each amino acid residue, the target residue pair features of each protein monomer, the first multiple sequence alignment feature, and the second multiple sequence alignment feature into an N-stage folding iterative network layer, and using the N-stage folding iterative network layer to predict the torsion angle of each amino acid residue, the residue-level position transformation, and the monomer chain-level position transformation, to obtain the target coordinates of each amino acid residue and obtain the predicted structure of the protein complex, and where the first multiple sequence alignment feature is a regularized multiple sequence alignment feature, the second multiple sequence alignment feature is a mapped multiple sequence alignment feature, and N is an integer greater than 1.

[0006] The present disclosure considers the relative independence of each monomer chain in the protein complex, and by adding the monomer chain-level position transformation based on the residue-level position transformation, updates the coordinates of each amino acid residue, accurately predicts the structure of the predicted protein, improves the efficiency of the structure prediction of the protein complex, and can be better applied to application scenarios where the protein complex contains multiple chains.

[0007] According to another aspect of the present disclosure, a protein complex structure prediction device is provided, and the device includes an acquisition module for acquiring the initial coordinates of each amino acid residue in the target protein complex and acquiring the target residue pair features, the first multiple sequence alignment feature, and the second multiple sequence alignment feature of each protein monomer in the target protein complex, and a structure prediction module for inputting the initial coordinates of each amino acid residue, the target residue pair features of each protein monomer, the first multiple sequence alignment feature, and the second multiple sequence alignment feature into an N-stage folding iterative network layer, and using the N-stage folding iterative network layer to predict the torsion angle of each amino acid residue, the residue-level position transformation, and the monomer chain-level position transformation, to obtain the target coordinates of each amino acid residue and obtain the predicted structure of the protein complex, and Here, the first multiple sequence alignment feature is a regularized multiple sequence alignment feature, the second multiple sequence alignment feature is a mapped multiple sequence alignment feature, and N is an integer greater than 1.

[0008] According to another aspect of the present disclosure, an electronic device is provided, including at least one processor and a memory communicably connected to the at least one processor, where the memory stores instructions executed by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is caused to execute the method for predicting the structure of a protein complex according to an embodiment of the first aspect of the present disclosure.

[0009] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, where the computer instructions cause a computer to execute the method for predicting the structure of a protein complex according to an embodiment of the first aspect of the present disclosure.

[0010] According to another aspect of the present disclosure, a computer program is provided, and when the computer program is executed by a processor, the steps of the method for predicting the structure of a protein complex according to an embodiment of the first aspect of the present disclosure are realized.

[0011] Note that the content described in this part does not identify essential or important features of the embodiments of the present disclosure, nor does it limit the scope of the present disclosure. Other features of the present disclosure will be more easily understood from the following description.

Brief Description of the Drawings

[0012] The drawings are for a better understanding of the solution and do not limit the present disclosure. [Figure 1] It is a flowchart of a method for predicting the structure of a protein complex according to an embodiment of the present disclosure. [Figure 2]This is a flowchart of a method for predicting the structure of a protein complex according to one embodiment of the present disclosure. [Figure 3] This is a structural diagram of a protein complex structure prediction method according to one embodiment of the present disclosure. [Figure 4] This is a flowchart of a method for predicting the structure of a protein complex according to one embodiment of the present disclosure. [Figure 5] This is a structural diagram of a protein complex structure prediction method according to one embodiment of the present disclosure. [Figure 6] This is a structural diagram of a protein complex structure prediction device according to one embodiment of the present disclosure. [Figure 7] This is a block diagram of an electronic device that implements the method of the embodiment of this disclosure. [Modes for carrying out the invention]

[0013] The following description, in conjunction with the drawings, illustrates exemplary embodiments of the present disclosure and includes various details of the embodiments for the sake of clarity, which should be considered illustrative. Therefore, as those skilled in the art will see, various changes and modifications can be made to the embodiments, provided they do not deviate from the scope and spirit of the disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0014] The embodiments described herein relate to the fields of artificial intelligence technologies such as computer vision and deep learning.

[0015] Artificial intelligence, abbreviated as AI in English, is a scientific field that researches and develops theories, methods, techniques, and applied systems that simulate and extend human intelligence.

[0016] Natural Language Processing (NLP) is a significant field in computer science and artificial intelligence. It studies various theories and methods to enable effective communication between humans and computers using natural language. NLP is a science that integrates linguistics, computer science, and mathematics. Therefore, while research in this field is closely related to linguistics, as it deals with natural language—that is, the language people use in everyday life—there is a crucial difference. NLP does not study natural language in general, but rather studies computer systems, particularly software systems, that can effectively realize natural language communication. It is a part of computer science.

[0017] Biocomputing refers to a new mode of computing developed through research and development utilizing the unique information processing mechanisms of biological systems. Biocomputing research encompasses two aspects: devices and systems. It provides basic units that use ordered systems constructed from organic (or biological) materials at the molecular scale to detect, process, transmit, and store information through physical and chemical processes at the molecular level.

[0018] The method and apparatus for predicting the structure of protein complexes according to this disclosure will be described below in conjunction with the drawings.

[0019] Figure 1 is a flowchart of a method for predicting the structure of a protein complex according to one embodiment of the present disclosure, and as shown in Figure 1, the method includes the following steps S101 to S102.

[0020] In S101, the initial coordinates of each amino acid residue in the target protein complex are obtained, and the target residue pair features, first multiple sequence alignment features, and second multiple sequence alignment features of each protein monomer in the target protein complex are obtained.

[0021] A protein complex has multiple protein monomers, each protein monomer having one amino acid sequence, and when amino acids bind to each other to form a peptide bond, they lose one molecule of water, thus the amino acid unit in a polypeptide / protein is an amino acid residue. In the embodiments of this disclosure, in order to adapt to the rotational invariance of the protein structure, the coordinates of each residue are represented using relative positional transformations, and the spatial structure of the protein complex is initialized at the origin of the coordinates, that is, the coordinates of each amino acid residue in the target protein complex are initialized, and the initial coordinates are T i =(I,0(→)) is obtained, and here T i =(I,0(→)) represents the apocentric coordinate system expressed by rotation / translation, where I is the identity matrix indicating no rotation, the 0(→) vector indicates no translation, and i represents the i-th amino acid residue.

[0022] In the embodiments of this disclosure, template features of each protein monomer are obtained, pair features of the amino acid sequence of each protein monomer are constructed, and for each protein monomer, target residue pair features of the protein monomer are obtained based on the template features and pair features of the protein monomer.

[0023] In some embodiments, for each protein monomer, homologous sequences of the protein monomer are searched and obtained from multiple gene sequence databases based on the target amino acid sequence of the protein monomer, multiple sequence alignment is performed on the homologous sequences of the protein monomer to obtain multiple sequence alignment features of the protein monomer, and then, based on the multiple sequence alignment features, different processing is performed to obtain a first multiple sequence alignment feature and a second multiple sequence alignment feature, which can be selected, such that the first multiple sequence alignment feature is a regularized multiple sequence alignment feature and the second multiple sequence alignment feature is a mapped multiple sequence alignment feature.

[0024] In S102, the initial coordinates of each amino acid residue, the target residue pair features of each protein monomer, the first multiple sequence alignment features, and the second multiple sequence alignment features are input to an N-level folding repeat network layer. The N-level folding repeat network layer predicts the twist angle of each amino acid residue, the positional transformation at the residue level, and the positional transformation at the monomer chain level, thereby obtaining the target coordinates of each amino acid residue and obtaining the predicted structure of the protein complex. Here, N is an integer greater than 1.

[0025] The torsion angle in the residue side chains can be selected based on a side chain and torsion angle predictor in an N-level folded repeat network layer.

[0026] In some embodiments, when predicting the structure of a protein complex, the residue codes of multiple chains in the complex are directly mapped to coordinate transformations, and these transformations act only on residues, which the present disclosure refers to as residue-level positional transformations.

[0027] In the embodiments of this disclosure, the relative independence of each monomer chain in the protein complex is taken into consideration. By adding monomer chain-level positional transformations to residue-level positional transformations, the coordinates of each amino acid residue are updated, thereby enabling separation of intrachain residue position predictions from daughter chain-wide position predictions and improving the overall effectiveness of the protein structure prediction model.

[0028] In the embodiments of this disclosure, the initial coordinates of each amino acid residue, the target residue pair features of each protein monomer, the first multiple sequence alignment features, and the second multiple sequence alignment features are input to an N-level folding repeat network layer. The N-level folding repeat network layer predicts the twist angle of each amino acid residue, the positional transformation at the residue level, and the positional transformation at the monomer chain level. By obtaining the target coordinates of each amino acid residue, the relative independence of each monomer chain in the protein complex is taken into consideration. By adding the positional transformation at the monomer chain level to the positional transformation at the residue level, the coordinates of each amino acid residue are updated, the structure of the predicted protein is accurately predicted, the efficiency of protein complex structure prediction is improved, and the method can be better applied to application scenarios in which the protein complex contains multiple chains.

[0029] Figure 2 is a flowchart of a method for predicting the structure of a protein complex according to one embodiment of the present disclosure, and as shown in Figure 2, the method includes the following steps S201 to S204.

[0030] In S201, the initial coordinates of each amino acid residue in the target protein complex are obtained, and the target residue pair features, first multiple sequence alignment features, and second multiple sequence alignment features of each protein monomer in the target protein complex are obtained.

[0031] For an introduction to step S201, please refer to the relevant explanation in the above embodiment; a detailed explanation is omitted here.

[0032] In S202, the initial coordinates, target residue pair features, and second multiple sequence alignment features are input to the first-stage folded repeat network layer to predict the positional transformation at the residue level and the positional transformation at the monomer chain level for each amino acid residue, thereby obtaining the target residue code 1 and candidate positional transformation 1 of the first-stage folded repeat network layer.

[0033] In the first stage, a folded repeating network layer performs an invariant point attention mechanism on the initial coordinates, target residue pair features, and second multiple sequence alignment features to obtain a rotationally invariant residue code. Furthermore, a linear network performs a mapping process to obtain target residue code 1.

[0034] Based on the backbone update algorithm, the target residue code of each amino acid residue is mapped, the residue-level positional transformation is predicted for target residue code 1 to obtain the first positional transformation 1 for each amino acid residue, and the monomer chain-level positional transformation is predicted for target residue code 1 to obtain the second positional transformation 1 for each amino acid residue.

[0035] As shown in Figure 3, in some embodiments, the monomer chain-level repositioning update (Chain Affine Update) process includes splitting two or more adjacent amino acid residues into different monomer chains based on the target residue code of each amino acid residue, for example, three adjacent amino acid residues may be split into the same monomer chain, i.e., residue codes [s1, s2, s3, ..., s i ,…,s r Specify ] and, according to [1,2,3,…,i,…,r] in the table below, the monomer chain before splicing (for example, s1~s3 belong to monomer chain 1, s r-2 ~s r The target residue codes of each monomer chain are localized to the monomer chain n, and the chain-level representation of the candidate residue codes is obtained by averaging the target residue codes of each monomer chain.

[0036] In the embodiments of this disclosure, for any one monomer chain, the mean value of the target residue codes of the target amino acid residue is calculated to obtain a chain-level candidate residue code. The candidate residue codes are then mapped based on a multilayer neural network structure (the multilayer linear network shown in Figure 3), and a second positional transformation of each amino acid residue in the monomer chain is obtained.

[0037] As shown in Figure 3, in some embodiments, the multilayer neural network structure includes a three-layer linear network structure, in which candidate residue codes are input to the first linear network for mapping to obtain a first transformation representation. The first transformation representation is input to the second linear network for mapping to obtain a second transformation representation. The first and second transformation representations are input to the third linear network for mapping to obtain a second positional transformation of each amino acid residue in the monomer chain. The structure of the three-layer linear network may be the same or different, and the embodiments of this application are not limited thereto.

[0038] Based on the first position transformation 1, the second position transformation 1, and the initial coordinates, a position update is performed to obtain a candidate position transformation 1 for the first stage of the folded network layer.

[0039] In S203, target residue pair features, target residue code m-1 and candidate positional transformation m-1 from the m-th stage folding repeat network layer are input to the m-th stage folding repeat network layer. Residue-level positional transformations and monomer chain-level positional transformations are predicted for each amino acid residue, and the target residue code m and candidate positional transformation m from the m-th stage folding repeat network layer are obtained, with the value of m ranging from 2 to N.

[0040] An invariance point attention mechanism is applied to the candidate positional transformation m-1 and target residue pair features of the m-1-th stage folded repeat network layer input to the m-th stage folded network layer to obtain a residue code with rotational invariance, and a mapping process is performed by a linear network to obtain the target residue code m.

[0041] Similarly, using the method in step S202, predict the positional transformation at the residue level for the target residue code m until a candidate positional transformation N and a target residue code N for the Nth-th stage folded network layer are obtained, thereby obtaining the first positional transformation m for each amino acid residue, predict the positional transformation at the monomer chain level for the target residue code m, thereby obtaining the second positional transformation m for each amino acid residue, and based on the first positional transformation m and the second positional transformation m, obtain the candidate positional transformation m for the m-th stage folded network layer.

[0042] In S204, the Nth-th stage folded repeat network layer predicts the side chain and twist angle for the first multiple sequence alignment feature and the target residue code N of the Nth-th stage folded repeat network layer, thereby obtaining the twist angle in the side chain of each amino acid residue. Based on the twist angle in the side chain of each amino acid residue and the candidate position transformation N of the Nth-th stage folded repeat network layer, the target coordinates of each amino acid residue are obtained.

[0043] The first multiple sequence alignment feature and the target residue code N of the Nth-stage folded repeat network layer are input to the side chain and twist angle predictor of the Nth-stage folded repeat network layer to obtain the twist angle in the side chain of the amino acid residue. Based on the twist angle in the side chain of each amino acid residue, the candidate position transformation N of the Nth-stage folded repeat network layer, and the retrograde position update, the target coordinates of the amino acid residue are obtained.

[0044] In the embodiments of this invention, the relative independence of each monomer chain in the protein complex is taken into consideration. By adding monomer chain-level positional transformations to residue-level positional transformations, the coordinates of each amino acid residue can be updated, enabling separation of intrachain residue position predictions and daughter chain overall position predictions. This improves the overall effectiveness of the protein structure prediction model, allowing for more appropriate overall adjustment of interchain docking relationships while maintaining the relative positions of residues within a single chain, and thus better fitting the protein complex structure prediction.

[0045] Figure 4 is a flowchart of a method for predicting the structure of a protein complex according to one embodiment of the present disclosure, and as shown in Figure 4, the method includes the following steps S401 to S408.

[0046] In S401, template features of each protein monomer are obtained, and pair features of the amino acid sequences of each protein monomer are constructed.

[0047] In some embodiments, the target amino acid sequence of each protein monomer is matched against multiple first amino acid sequences in a protein structure database to obtain second amino acid sequences whose similarity exceeds a preset threshold. The distances between the coordinates of amino acid residues in the second amino acid sequence are then extracted and used as template features for each protein monomer. In other words, a protein structure database with a similar sequence to the amino acid sequence of a protein monomer is searched for, and the distances between residues are extracted based on a protein sequence analysis tool, such as a Hidden Markov Model (HMM) search method (HHSearch), and used as template features.

[0048] In some embodiments, the amino acid sequence of each protein monomer is input into two pre-configured linear networks to obtain candidate sequence coding features. One empty dimension is added to each of the candidate sequence coding features in different directions to obtain the first sequence coding feature and the second sequence coding feature, and the first and second sequence coding features are added together to obtain the pair features of each protein monomer. Multiple sequences, complex sequences with a spliced ​​length of r, are coded in two linear network Linear layers to obtain sequence coding features with shape [r,c], i.e., the first sequence coding feature z1 and the second sequence coding feature z2, where c is the hidden layer depth (hyperparameter) of the Linear network. Then, one empty dimension is added to z1 and z2 (transforming the shape of z1 to [r,1,c] and the shape of z2 to [1,r,c]), and they are added together to obtain the pair features z pair Obtain z pair The shape is [r,r,c], zpair = z1 + z2.

[0049] In S402, after inputting and mapping the template features of each protein monomer into a linear network, the candidate residue pair features are obtained by adding them to the pair features of each protein monomer.

[0050] In the embodiment of the present disclosure, the shape of the spliced Template feature is [r, r]. After being encoded by the Linear layer, the feature z temp is obtained, which is consistent with the shape of the Piar feature. Then, the candidate residue pair features are obtained by adding the feature z temp and the pair features.

[0051] In S403, the candidate residue pair features are input into a preset encoder for encoding to obtain the target residue pair features of each protein monomer.

[0052] In the embodiment of the present disclosure, the candidate residue pair features are input into an encoder (Evofomer Encoder) for encoding to obtain the target residue pair features of each protein monomer.

[0053] In S404, based on the target amino acid sequence of each protein monomer, the homologous sequences of each protein monomer are retrieved from multiple gene sequence databases.

[0054] The present invention first uses the amino acid sequence of each protein monomer in the complex as a query request and retrieves homologous sequences from multiple gene sequence databases. By using existing tools JackHMMER and HHblits, a deeper analysis and annotation of the protein sequence can be realized. By using JackHMMER, a fast heuristic search of the hidden Markov model (HMM) can be performed, and by using HHblits, more detailed annotation of the discovered protein sequence can be performed, thereby obtaining the homologous sequences of each protein monomer.

[0055] In S405, multiple sequence alignment is performed on the homologous sequences of each protein monomer to obtain candidate multiple sequence alignment features for each protein monomer.

[0056] Using the obtained homologous sequences, multiple sequence alignment features (MSA) of each monomer are obtained.

[0057] In S406, candidate multiple sequence alignment features of each protein monomer are input into a pre-configured encoder and encoded to obtain target multiple sequence alignment features of each protein monomer.

[0058] The candidate multiple sequence alignment features of each protein monomer are input into an encoder (Evofomer Encoder) and encoded to obtain the target multiple sequence alignment features of each protein monomer.

[0059] S407: The target multiple sequence alignment features of each protein monomer are regularized to obtain the first multiple sequence alignment features of each protein monomer, and the target multiple sequence alignment features of each protein monomer are mapped to obtain the second multiple sequence alignment features of each protein monomer.

[0060] Regularization (Norm processing) is performed on the target multiple sequence alignment features of each protein monomer to obtain the first multiple sequence alignment features of each protein monomer. Then, based on a linear network, the target multiple sequence alignment features of each protein monomer are mapped to obtain the second multiple sequence alignment features of each protein monomer.

[0061] In S408, the initial coordinates of each amino acid residue, the target residue pair characteristics of each protein monomer, the first multiple sequence alignment characteristics, and the second multiple sequence alignment characteristics are input into an N-level folding and repeating network layer. The N-level folding and repeating network layer predicts the twist angle of each amino acid residue, the positional transformation at the residue level, and the positional transformation at the monomer chain level, thereby obtaining the target coordinates of each amino acid residue and obtaining the predicted structure of the protein complex.

[0062] For an explanation of step S408, please refer to the relevant information in the above embodiment, and a detailed explanation will be omitted here.

[0063] In the embodiments of this invention, the initial coordinates of each amino acid residue in the target protein complex are obtained, and the target residue pair features, first multiple sequence alignment features, and second multiple sequence alignment features of each protein monomer in the target protein complex are obtained, thereby effectively promoting the development of the protein monomer structure prediction task and enhancing the overall effectiveness of protein structure prediction.

[0064] Figure 5 is a structural diagram of a protein complex structure prediction method according to one embodiment of the present disclosure. As shown in Figure 5, in the embodiment of the present disclosure, homologous sequences are searched from multiple gene sequence databases (Sequence Data Bases) based on the amino acid sequences [Sequence 1, ..., Sequence N] of each protein monomer in the target protein complex, and multiple sequence alignment is performed to obtain multiple sequence alignment features [MSA 1, ..., MSA N] for each protein monomer. These multiple sequence alignment features [MSA 1, ..., MSA N] are input into an Evofomer Encoder and encoded to obtain target multiple sequence alignment features. The amino acid sequences of each protein monomer are input into two pre-configured linear networks (Linear) to generate pair features. Protein structures with similar sequences to the target amino acid sequence of each protein monomer are searched from a protein structure database (Structure Data Base), and the distance between residues is extracted as a template feature. Selectable options include Pair and Merge, which indicate merging, i.e., MSA representation features and pair features. This method demonstrates merging representation features, inputting the template features of each protein monomer into a linear network for mapping, adding them to the pair features of each protein monomer, inputting them into an Evofomer encoder for encoding, and obtaining the target residue pair features of each protein monomer. The Evofomer encoder can extract hidden layer codes for each residue from MSA, Pair, and Template data. In the decoding stage, this disclosure supports training tasks by performing MSA shielding prediction (Mask MSA), LDDT prediction (LDDT is a pre-existing measurement method used in the field of protein structure prediction), residue distance prediction, etc., based on the Structure Module in the AF2Multimer protein complex structure prediction model. Here, initial frame represents the initial coordinates.

[0065] As shown in Figure 5, a single protein complex containing r residues is input and processed by the Evofomer Encoder. The model then obtains the MSA feature codes for each residue i of the target protein complex, i.e., the target multi-sequence alignment features.

[0066]

number

[0067]

number

[0068] In embodiments of this disclosure, the structure of a target protein complex is predicted using an N-step folded iteration network layer (Fold Iteration module), and the value of N may be 8, and in other realizations, N may be any other value, and is not limited to the embodiments of this application.

[0069] To adapt to the rotational indenaturation of protein structures, the present invention provides relative positional transformation T i =(R i ,t(→) i The coordinates of each residue are represented using ), and the coordinate origin T iThe spatial structure of the protein complex is initialized with =(I,0(→)). This invention first uses two regularized Norm layer networks

[0070]

number

[0071]

number

[0072] Here, each layer Fold Iteration is s i , Z i,j , T i After accessing it, first, the Invariant Point Attention module processes the residue codes s that have rotational invariance. i Obtain the following. After passing through network layers such as Linear, Norm, and Dropout layers, this disclosure uses the obtained code to obtain the twist angle af i∈R of the side chain of each residue. 2 And predict the coordinates of each residue TC k = (RC k, tC k), where the Dropout layer is used to randomly discard network parameters and plays a small role, and is not shown in Figure 5; this layer may be omitted or deleted.

[0073] As shown in Figure 5, in the embodiments of this disclosure, in each layer Fold Iteration, a Side Chain and torsion angle predictor is used to determine the torsion angle af i∈R in the side chain of residue i. 2 We predict that, here

[0074]

number

[0075] In predicting the protein complex backbone network structure, this disclosure first encodes residue features using a shallow neural network (Linear and Norm) structure, and then uses the BackboneUpdate algorithm to perform the Euclidean transformation of each residue i T i , that is, predict the first positional transformation. BackboneUpdate maps the hidden layer features to a single 6-dimensional representation, where the first three dimensions b i , c i d i The rotation matrix R of residue i is obtained through the equation. i Used as the last three dimensions t(→) i This directly represents the transfer transformation of residues. Here, the input for Fold Iteration considers only the main chain positional transformation affine and residue code, without considering the side chains and twist angles predicted in the previous step.

[0076] Predicted TR i=(R i ,t(→) i After obtaining the model T i =T i O Based on TR i, the relative positions of the residues are updated to complete the residue-level positional transformation, and in the equation, the first T i This is the updated T i This shows the second T i This is the T before the update. i This indicates.

[0077] Based on the above transformation, this disclosure introduces a monomer chain-level repositioning Chainaffine module to predict chain-level residue transformations, and the Chainaffine module aims to predict the overall transformation TC k of the protein complex chain k. The module structure can be realized in various forms, for example, as shown in Figure 3, which is a method for realizing one Chainaffine module. i Taking this as input, the Chainaffine module first processes the residue position code, s i The sequence is divided into different monomer chains, and the average value of all residue representations in the same monomer is calculated, followed by the chain-level hidden layer representation.

[0078]

number

[0079] After this, the Chainaffine module uses a multilayer neural network structure

[0080]

number

[0081] As shown in Figure 5, when updating the intrachain residue positions, all residues i in chain k share the transformed TC k, and each residue acquires a chain-level second positional transformed TC i. Finally, the model T i =T i O Based on TC i, the position of each residue in the complex is updated. In the formula, the first T i The updated T i This represents the second T i This is the T before the update. i This represents TC i = T iIt acts on the front side, causing the residues within the chain to be reoriented around the origin.

[0082] After obtaining the Euclidean transformation of the backbone network and the side chain twist angle, the embodiments of this application use the residue update module and the frame update module to perform T i The `af i` is then converted into the 3D coordinates of each residue to complete the prediction of the protein's tertiary structure. Here, the Angles module predicts the side-chain twist angles, and the Coordinates convert module receives the main chain transformation and side-chain twist angles and then outputs the transformed spatial coordinates.

[0083] Figure 6 is a structural diagram of a protein complex structure prediction device according to one embodiment of the present disclosure. As shown in Figure 6, the protein complex structure prediction device 600 is: An acquisition module 610 for obtaining the initial coordinates of each amino acid residue in the target protein complex, and for obtaining the target residue pair features, first multiplex alignment features, and second multiplex alignment features of each protein monomer in the target protein complex, The system includes a structure prediction module 620 for obtaining the predicted structure of a protein complex by inputting the initial coordinates of each amino acid residue, the target residue pair features of each protein monomer, the first multiple sequence alignment features, and the second multiple sequence alignment features into an N-level folding and repeating network layer, and by using the N-level folding and repeating network layer to predict the twist angle of each amino acid residue, the positional transformation at the residue level, and the positional transformation at the monomer chain level, thereby obtaining the target coordinates of each amino acid residue. Here, the first multi-array alignment feature is a regularized multi-array alignment feature, the second multi-array alignment feature is a mapped multi-array alignment feature, and N is an integer greater than 1.

[0084] In some embodiments, the structural prediction module 620 further, The initial coordinates, target residue pair features, and second multiple sequence alignment features are input to the first-stage folded repeat network layer, and the positional transformation at the residue level and monomer chain level is predicted for each amino acid residue to obtain the target residue code 1 and candidate positional transformation 1 of the first-stage folded repeat network layer. For the m-th folding repeat network layer, the target residue pair features, the target residue code m-1 and candidate positional transformation m-1 from the m-1-th folding repeat network layer are input to the m-th folding repeat network layer to predict the positional transformation at the residue level and the positional transformation at the monomer chain level for each amino acid residue, thereby obtaining the target residue code m and candidate positional transformation m from the m-th folding repeat network layer, where the value of m is between 2 and N. The Nth-th stage folded repeat network layer predicts the side chain and twist angle for the first multiple sequence alignment features and the target residue code N of the Nth-th stage folded repeat network layer, thereby obtaining the twist angle in the side chain of each amino acid residue. Based on the twist angle in the side chain of each amino acid residue and the candidate position transformation N of the Nth-th stage folded repeat network layer, the target coordinates of each amino acid residue are obtained.

[0085] In some embodiments, the structural prediction module 620 further, The first stage of the folded repeating network layer performs invariant point attention mechanism and mapping processing on the initial coordinates, target residue pair features, and second multiple sequence alignment features to obtain target residue code 1. Predicting the positional transformation at the residue level for target residue code 1 and obtaining the first positional transformation 1 for each amino acid residue, and predicting the positional transformation at the monomer chain level for target residue code 1 and obtaining the second positional transformation 1 for each amino acid residue, Based on the first position transformation 1, the second position transformation 1, and the initial coordinates, a position update is performed to obtain a candidate position transformation 1 for the first stage of the folded network layer.

[0086] In some embodiments, the structural prediction module 620 further, The candidate positional transformation m-1 of the m-1-th folded repeating network layer input to the m-th folded network layer and the target residue pair features are subjected to an invariance attention mechanism and mapping process to obtain the target residue code m. The first positional transformation m of each amino acid residue is obtained by predicting the positional transformation at the residue level for the target residue code m, and the second positional transformation m of each amino acid residue is obtained by predicting the positional transformation at the monomer chain level for the target residue code m. Based on the first position transformation m and the second position transformation m, a candidate position transformation m for the m-th stage of the folded network layer is obtained.

[0087] In some embodiments, the structural prediction module 620 further, Based on the backbone update algorithm, the target residue code of each amino acid residue is mapped to obtain the first positional transformation of each amino acid residue.

[0088] In some embodiments, the structural prediction module 620 further, For each amino acid residue, two or more adjacent amino acid residues are split into different monomer chains based on the target residue code of the amino acid residue. For any one monomer chain, the average value is calculated for the target amino acid residue code of the target amino acid residue to obtain a chain-level candidate residue code. The candidate residue codes are then mapped based on a multilayer neural network structure to obtain the second positional transformation of each amino acid residue in the monomer chain.

[0089] In some embodiments, the multilayer neural network structure includes a three-layer linear network. The structural prediction module 620 further, The candidate residue codes are input into the first linear network and mapped to obtain the first transformed representation. The first transformed representation is input to the second linear network and mapped to obtain the second transformed representation. The first and second transformation representations are input into a third linear network and mapped to obtain the second positional transformation of each amino acid residue in the monomer chain.

[0090] In some embodiments, the acquisition module 610 further, We obtain template features for each protein monomer and construct pair features of the amino acid sequence for each protein monomer. After inputting the template features of each protein monomer into a linear network and mapping them, the pair features of each protein monomer are added together to obtain candidate residue pair features. The candidate residue pair features are input into a pre-configured encoder and encoded to obtain the target residue pair features of each protein monomer.

[0091] In some embodiments, the acquisition module 610 is The target amino acid sequence of each protein monomer is matched against multiple first amino acid sequences in the protein structure database, and the second amino acid sequence whose similarity is greater than a predetermined threshold is obtained. The distances between the coordinates of amino acid residues in the second amino acid sequence are extracted and used as template features for each protein monomer.

[0092] In some embodiments, the acquisition module 610 further, The amino acid sequences of each protein monomer are input into two pre-configured linear networks to obtain candidate sequence coding features. By adding one empty dimension in each of the candidate sequence coding features in different directions, we obtain the first sequence coding feature and the second sequence coding feature. The first sequence-encoded feature and the second sequence-encoded feature are added together to obtain the pair features of each protein monomer.

[0093] In some embodiments, the acquisition module 610 further, Based on the target amino acid sequence of each protein monomer, homologous sequences of each protein monomer are searched and obtained from multiple gene sequence databases. Multiple sequence alignment is performed on the homologous sequences of each protein monomer to obtain candidate multiple sequence alignment features for each protein monomer. The candidate multiple sequence alignment features of each protein monomer are input into a pre-configured encoder and encoded to obtain the target multiple sequence alignment features of each protein monomer. The target multiple sequence alignment features of each protein monomer are regularized to obtain the first multiple sequence alignment features of each protein monomer, and the target multiple sequence alignment features of each protein monomer are mapped to obtain the second multiple sequence alignment features of each protein monomer.

[0094] This disclosure considers the relative independence of each monomer chain in a protein complex and updates the coordinates of each amino acid residue by adding monomer chain-level positional transformations based on residue-level positional transformations, thereby accurately predicting the structure of the predicted protein, improving the efficiency of protein complex structure prediction, and making it more applicable to application scenarios in which the protein complex contains multiple chains.

[0095] According to embodiments of the present disclosure, the present disclosure further provides electronic devices and readable storage media. According to embodiments of the present disclosure, the present disclosure provides a computer program, and when the computer program is executed by a processor, a method for predicting the structure of a protein complex provided by the present disclosure is realized.

[0096] Figure 7 is a block diagram of an electronic device that implements an embodiment of the present disclosure. The electronic device can implement the protein complex structure prediction method of the embodiment of the present disclosure, and the electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, mobile phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure as described and / or required herein.

[0097] As shown in Figure 7, device 700 includes a computing unit 701 that can perform various appropriate operations and processes based on computer programs stored in read-only memory (ROM) 702 or computer programs loaded from storage unit 708 into random access memory (RAM) 703. RAM 703 may also store various programs and data necessary for the operation of device 700. The computing unit 701, ROM 702, and RAM 703 are connected to each other via bus 704. An input / output (I / O) interface 705 is also connected to bus 704.

[0098] Multiple components of device 700 are connected to the I / O interface 705 and include input units 706 such as a keyboard and mouse, output units 707 such as various types of displays and speakers, storage units 708 such as magnetic disks and optical disks, and communication units 709 such as a network card, modem, and wireless communication transceiver. The communication units 709 enable device 700 to exchange information / data with other devices via computer networks such as the Internet and / or various telecommunication networks.

[0099] The computing unit 701 is a variety of general-purpose and / or dedicated processing components having processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, computing units that execute various machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 performs each of the methods and processes described above, for example, the text correction method. For example, in some embodiments, the text correction method is implemented and obtained as a computer software program tangibly embedded in a machine-readable medium such as a storage unit 708. In some embodiments, part or all of the computer program is loaded and / or installed into device 700 via ROM 702 and / or communication unit 709 and obtained. When the computer program is loaded into RAM 703 and executed by the computing unit 701, one or more steps of the text correction method described above are performed and obtained. Optionally, in other embodiments, the computing unit 701 may be configured to perform the text correction method in any other suitable manner (e.g., via firmware).

[0100] Various embodiments of the systems and technologies described herein are implemented and obtained by digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SOCs), load-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include being implemented by one or more computer programs, which may run and / or be interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, and which may include receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, at least one input device, and at least one output device.

[0101] Program code for carrying out the methods of this disclosure can be programmed in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, an application-specific computer, or other programmable data processing device, so that when the program code is executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagrams are performed. The program code may run entirely on the machine, partially on the machine, as a standalone software package, with part running on the machine and part on a remote machine, or entirely on a remote machine or server.

[0102] In the context of this disclosure, machine-readable media are tangible media that can contain or store programs used by or in combination with instruction execution systems, devices, or equipment. Machine-readable media are machine-readable signal media or machine-readable storage media. Machine-readable media include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or equipment, or any suitable combination thereof. More specific examples of machine-readable storage media include, but are not limited to, electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0103] To provide user interaction, the systems and technologies described herein can be implemented on a computer, which has a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) and a keyboard and pointing device (e.g., a mouse or trackball), and the user can provide input to the computer using the keyboard and pointing device. Other types of devices can also provide user interaction; for example, the feedback provided to the user may be any form of sensing feedback (e.g., vision feedback, auditory feedback, or haptic feedback), and input from the user may be received in any form (including sound input, voice input, or haptic input).

[0104] The systems and technologies described herein can be implemented in computing systems including backend components (e.g., as data servers), computing systems including middleware components (e.g., application servers), computing systems including frontend components (e.g., user computers having a graphical user interface or web browser, through which users interact with embodiments of the systems and technologies described herein), or computing systems including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by digital data communication in any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0105] A computer system may include clients and servers. Clients and servers are generally geographically separated and typically interact via a communication network. The client-server relationship is generated by computer programs running on corresponding computers that have a client-server relationship with each other. The servers may be cloud servers, servers in a distributed system, or servers combined with blockchain technology.

[0106] Furthermore, the steps can be rearranged, added, or deleted using the various forms of flows shown above. For example, each step described herein may be performed in parallel, sequentially, or in a different order, as long as it achieves the desired result of the technical solution disclosed herein.

[0107] The specific embodiments described above do not limit the scope of protection of this disclosure. Those skilled in the art can make various modifications, combinations, subcombinations, and substitutions based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for predicting the structure of a protein complex, The steps include obtaining the initial coordinates of each amino acid residue in the target protein complex, and obtaining the target residue pair characteristics, first multiple sequence alignment characteristics, and second multiple sequence alignment characteristics of each protein monomer in the target protein complex, The process includes the steps of: inputting the initial coordinates of each amino acid residue, the target residue pair features of each protein monomer, the first multiple sequence alignment features, and the second multiple sequence alignment features into an N-level folding repeat network layer; predicting the twist angle of each amino acid residue, the residue-level positional transformation, and the monomer chain-level positional transformation using the N-level folding repeat network layer; obtaining the target coordinates of each amino acid residue; and obtaining the predicted structure of the protein complex. The target residue pair feature includes a template feature of the corresponding protein monomer and an amino acid sequence pair feature, the first multiple sequence alignment feature is a multiple sequence alignment feature in which the target multiple sequence alignment feature of the corresponding protein monomer is regularized, and the second multiple sequence alignment feature is a multiple sequence alignment feature to which the target multiple sequence alignment feature of the corresponding protein monomer is mapped, where N is an integer greater than 1. Methods for predicting the structure of protein complexes.

2. The steps include inputting the initial coordinates, target residue pair features, and second multiple sequence alignment features into the first-stage folded repeat network layer, predicting the positional transformation at the residue level and the positional transformation at the monomer chain level for each amino acid residue, and obtaining the target residue code 1 and candidate positional transformation 1 of the first-stage folded repeat network layer, A step of inputting the target residue pair features, the target residue code m-1 of the m-1 stage folded repeat network layer, and the candidate positional transformation m-1 to the m-th stage folded repeat network layer, predicting the positional transformation at the residue level and the positional transformation at the monomer chain level for each amino acid residue, thereby obtaining the target residue code m and the candidate positional transformation m of the m-th stage folded repeat network layer, wherein the value of m is between 2 and N. The process further includes: using an N-th folded repeating network layer to predict side chains and twist angles for the first multiple sequence alignment feature and the target residue code N of the N-th folded repeating network layer, obtaining the twist angle in the side chain of each amino acid residue, and obtaining the target coordinates of each amino acid residue based on the twist angle in the side chain of each amino acid residue and the candidate position transformation N of the N-th folded repeating network layer. A method for predicting the structure of a protein complex according to claim 1.

3. The steps of inputting the initial coordinates, the target residue pair features, and the second multiple sequence alignment features into the first-stage folded repeat network layer, predicting the positional transformation at the residue level and the positional transformation at the monomer chain level for each amino acid residue, and obtaining the target residue code 1 and candidate positional transformation 1 of the first-stage folded repeat network layer are as follows: The first step involves performing an invariant point attention mechanism and mapping process on the initial coordinates, the target residue pair features, and the second multiple sequence alignment features using the first folded repeat network layer to obtain the target residue code 1. The steps include: predicting the positional transformation at the residue level for the target residue code 1 to obtain a first positional transformation 1 for each amino acid residue; and predicting the positional transformation at the monomer chain level for the target residue code 1 to obtain a second positional transformation 1 for each amino acid residue; The process includes the step of performing a position update based on the first position transformation 1, the second position transformation 1, and the initial coordinates to obtain a candidate position transformation 1 for the first stage of the folded network layer, The method for predicting the structure of a protein complex according to claim 2.

4. The steps of inputting the target residue pair features, the target residue code m-1 and candidate positional transformation m-1 of the m-1 folding repeat network layer into the m-stage folding repeat network layer, predicting the positional transformation at the residue level and the positional transformation at the monomer chain level for each amino acid residue, and obtaining the target residue code m and candidate positional transformation m of the m-stage folding repeat network layer are as follows: The steps include: applying an invariance point attention mechanism and mapping process to the candidate position transformation m-1 of the m-1 stage folded repeat network layer input to the m-stage folded network layer and the target residue pair feature to obtain the target residue code m; The steps include: predicting the positional transformation at the residue level for the target residue code m to obtain a first positional transformation m for each amino acid residue; and predicting the positional transformation at the monomer chain level for the target residue code m to obtain a second positional transformation m for each amino acid residue. The process includes the step of obtaining a candidate position transformation m for the m-stage folded network layer based on the first position transformation m and the second position transformation m, The method for predicting the structure of a protein complex according to claim 2.

5. The process of predicting the positional transformation at the residue level for the target residue code of each amino acid residue and obtaining the first positional transformation of each amino acid residue is as follows: The step of mapping the target residue code of each amino acid residue based on a backbone update algorithm to obtain the first positional transformation of each amino acid residue, A method for predicting the structure of a protein complex according to claim 3.

6. The process of predicting monomer chain-level positional transformations for the target residue code of each amino acid residue and obtaining a second positional transformation for each amino acid residue is as follows: For each amino acid residue, the step of splitting two or more adjacent amino acid residues into different monomer chains based on the target residue code of the amino acid residue, The method includes the steps of: for any one monomer chain, calculating the average value of the target amino acid residue code of the target amino acid residue to obtain a chain-level candidate residue code; mapping the candidate residue code based on a multilayer neural network structure to obtain the second positional transformation of each amino acid residue in the monomer chain; A method for predicting the structure of a protein complex according to claim 3.

7. The aforementioned multilayer neural network structure includes a three-layer linear network. The step of mapping the candidate residue codes based on the multilayer neural network structure to obtain the second repositional transformation of each amino acid residue in the monomer chain is: The steps include inputting the candidate residue codes into a first linear network and mapping them to obtain a first transformed representation, The steps include: inputting the first transformation representation into a second linear network and mapping it to obtain a second transformation representation; The process includes the step of inputting the first transformation representation and the second transformation representation into a third linear network and mapping them to obtain the second positional transformation of each amino acid residue in the monomer chain, A method for predicting the structure of a protein complex according to claim 6.

8. The step of obtaining the target residue pair characteristics of each protein monomer in the target protein complex is: The steps include obtaining template features for each protein monomer and constructing pair features of amino acid sequences for each protein monomer, The steps include: inputting the template features of each protein monomer into a linear network and mapping them, then adding them to the pair features of each protein monomer to obtain candidate residue pair features; The process includes the step of inputting the candidate residue pair features into a pre-configured encoder and encoding them to obtain the target residue pair features of each protein monomer. A method for predicting the structure of a protein complex according to claim 1.

9. Obtaining template characteristics for each of the aforementioned protein monomers is possible. The target amino acid sequence of each protein monomer is matched against multiple first amino acid sequences in a protein structure database to obtain a second amino acid sequence whose similarity is greater than a predetermined threshold. This includes extracting the distances between the coordinates of amino acid residues in the second amino acid sequence and using them as template features for each protein monomer. A method for predicting the structure of a protein complex according to claim 8.

10. Constructing the pair characteristics of the amino acid sequence of each protein monomer is The amino acid sequences of each protein monomer are input into two pre-configured linear networks to obtain candidate sequence coding features. The process involves adding one empty dimension in each of the candidate sequence coding features in different directions to obtain the first sequence coding feature and the second sequence coding feature, This includes adding the first sequence encoding feature and the second sequence encoding feature to obtain the pair features of each protein monomer, A method for predicting the structure of a protein complex according to claim 8.

11. The step of obtaining the first multiple sequence alignment features and the second multiple sequence alignment features of each protein monomer in the target protein complex is: The steps include: searching for and obtaining homology sequences of each protein monomer from multiple gene sequence databases based on the target amino acid sequence of each protein monomer; The steps include performing multiple sequence alignment on the homologous sequences of each protein monomer to obtain candidate multiple sequence alignment features for each protein monomer, The steps include inputting the candidate multiple sequence alignment features of each protein monomer into a pre-configured encoder and encoding them to obtain the target multiple sequence alignment features of each protein monomer, The target multiple sequence alignment features of each protein monomer are regularized to obtain the first multiple sequence alignment features of each protein monomer, and each protein monomer The process includes the step of mapping the target multiple sequence alignment features of the body to obtain the second multiple sequence alignment features of each protein monomer, A method for predicting the structure of a protein complex according to claim 1.

12. A protein complex structure prediction device, An acquisition module for obtaining the initial coordinates of each amino acid residue in a target protein complex, and for obtaining target residue pair characteristics, first multiple sequence alignment characteristics, and second multiple sequence alignment characteristics of each protein monomer in the target protein complex, The system includes a structure prediction module that inputs the initial coordinates of each amino acid residue, the target residue pair features of each protein monomer, the first multiple sequence alignment features, and the second multiple sequence alignment features into an N-level folding and repeating network layer, and uses the N-level folding and repeating network layer to predict the twist angle of each amino acid residue, the positional transformation at the residue level, and the positional transformation at the monomer chain level, thereby obtaining the target coordinates of each amino acid residue and obtaining the predicted structure of the protein complex. The target residue pair feature includes a template feature of the corresponding protein monomer and an amino acid sequence pair feature, the first multiple sequence alignment feature is a multiple sequence alignment feature in which the target multiple sequence alignment feature of the corresponding protein monomer is regularized, and the second multiple sequence alignment feature is a multiple sequence alignment feature to which the target multiple sequence alignment feature of the corresponding protein monomer is mapped, where N is an integer greater than 1. A device for predicting the structure of protein complexes.

13. The aforementioned structural prediction module further, The initial coordinates, target residue pair features, and second multiple sequence alignment features are input to the first-stage folded repeat network layer, and the positional transformation at the residue level and the positional transformation at the monomer chain level are predicted for each amino acid residue to obtain the target residue code 1 and candidate positional transformation 1 of the first-stage folded repeat network layer. For the m-th folding repeat network layer, the target residue pair features, the target residue code m-1 and candidate positional transformation m-1 of the m-1-th folding repeat network layer are input to the m-th folding repeat network layer, and the positional transformation at the residue level and the positional transformation at the monomer chain level are predicted for each amino acid residue to obtain the target residue code m and candidate positional transformation m of the m-th folding repeat network layer, the value of m being between 2 and N. The Nth-stage folded repeating network layer performs side chain and twist angle predictions for the first multiple sequence alignment feature and the target residue code N of the Nth-stage folded repeating network layer, thereby obtaining the twist angle in the side chain of each amino acid residue, and based on the twist angle in the side chain of each amino acid residue and the candidate position transformation N of the Nth-stage folded repeating network layer, the target coordinates of each amino acid residue are obtained. The protein complex structure prediction device according to claim 12.

14. The aforementioned structural prediction module further, The first-stage folded repeat network layer performs an invariant point attention mechanism and mapping process on the initial coordinates, the target residue pair features, and the second multiple sequence alignment features to obtain the target residue code 1. Predicting the positional transformation at the residue level for the target residue code 1 and obtaining the first positional transformation 1 for each amino acid residue, and predicting the positional transformation at the monomer chain level for the target residue code 1 and obtaining the second positional transformation 1 for each amino acid residue, Based on the first position transformation 1, the second position transformation 1, and the initial coordinates, a position update is performed to obtain the candidate position transformation 1 for the first stage of the folded network layer. A protein complex structure prediction device according to claim 13.

15. The aforementioned structural prediction module The candidate position transformation m-1 of the m-1 stage folded repeat network layer input to the m-stage folded network layer and the target residue pair features are subjected to an invariance point attention mechanism and mapping process to obtain the target residue code m. Predicting the positional transformation at the residue level for the target residue code m is performed to obtain the first positional transformation m for each amino acid residue, and predicting the positional transformation at the monomer chain level for the target residue code m is performed to obtain the second positional transformation m for each amino acid residue, Based on the first position transformation m and the second position transformation m, a candidate position transformation m for the m-th stage of the folded network layer is obtained. A protein complex structure prediction device according to claim 13.

16. The aforementioned structural prediction module further, The target residue code of each amino acid residue is mapped based on the backbone update algorithm to obtain the first positional transformation of each amino acid residue. A protein complex structure prediction device according to claim 14.

17. The aforementioned structural prediction module further, For each of the amino acid residues, two or more adjacent amino acid residues are split into different monomer chains based on the target residue code of the amino acid residue. For any one monomer chain, the average value of the target amino acid residue code of the target amino acid residue is calculated to obtain a chain-level candidate residue code, and the candidate residue codes are mapped based on a multilayer neural network structure to obtain the second positional transformation of each amino acid residue in the monomer chain. A protein complex structure prediction device according to claim 14.

18. The aforementioned multilayer neural network structure includes a three-layer linear network. The aforementioned structural prediction module further, The candidate residue codes are input into the first linear network and mapped to obtain the first transformed representation. The first transformation representation is input to a second linear network and mapped to obtain a second transformation representation. The first transformation representation and the second transformation representation are input to a third linear network and mapped to obtain the second positional transformation of each amino acid residue in the monomer chain. A protein complex structure prediction device according to claim 17.

19. The aforementioned acquisition module further, Template features of each protein monomer are obtained, and pair features of the amino acid sequence of each protein monomer are constructed. After inputting the template features of each protein monomer into a linear network and mapping them, candidate residue pair features are obtained by adding them to the pair features of each protein monomer. The candidate residue pair features are input into a pre-configured encoder and encoded to obtain the target residue pair features of each protein monomer. The protein complex structure prediction device according to claim 12.

20. The acquisition module, The target amino acid sequence of each protein monomer is matched against a plurality of first amino acid sequences in a protein structure database to obtain a second amino acid sequence whose similarity is greater than a predetermined threshold. The distances between the coordinates of amino acid residues in the second amino acid sequence are extracted and used as template features for each protein monomer. A protein complex structure prediction device according to claim 19.

21. The aforementioned acquisition module further, The amino acid sequences of each protein monomer are input into two pre-configured linear networks to obtain candidate sequence coding features. By adding one empty dimension in each of the candidate sequence coding features in different directions, the first sequence coding feature and the second sequence coding feature are obtained. The first sequence encoding feature and the second sequence encoding feature are added together to obtain the pair features of each protein monomer. A protein complex structure prediction device according to claim 19.

22. The aforementioned acquisition module further, Based on the target amino acid sequence of each protein monomer, homologous sequences of each protein monomer are searched for and obtained from multiple gene sequence databases. Multiple sequence alignment is performed on the homologous sequences of each protein monomer to obtain candidate multiple sequence alignment characteristics for each protein monomer. The candidate multiple sequence alignment features of each of the protein monomers are input into a pre-configured encoder and encoded to obtain the target multiple sequence alignment features of each of the protein monomers. The target multiple sequence alignment features of each protein monomer are regularized to obtain the first multiple sequence alignment features of each protein monomer, and the target multiple sequence alignment features of each protein monomer are mapped to obtain the second multiple sequence alignment features of each protein monomer. The protein complex structure prediction device according to claim 12.

23. It is an electronic device, At least one processor, Includes memory that is communicably connected to at least one processor, The memory stores instructions that can be executed by the at least one processor, and the execution of these instructions by the at least one processor causes the at least one processor to perform the method according to any one of claims 1 to 11. Electronic devices.

24. A non-temporary, computer-readable storage medium in which computer instructions are stored, The computer instruction causes the computer to perform the method according to any one of claims 1 to 11. A non-temporary, computer-readable storage medium.

25. It is a computer program, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 11 are realized. Computer program.