Method and apparatus for predicting the structure of a protein complex

Through N-level folding iterative network layer and multi-sequence alignment features, combined with relative position transformation, the accuracy and efficiency of protein complex structure prediction are solved, and more efficient protein structure prediction is achieved.

CN117577170BActive Publication Date: 2025-07-22BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311477801.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-08
Publication Date
2025-07-22
Estimated Expiration
2043-11-08

AI Technical Summary

Technical Problem

The prior art is difficult to accurately predict the structure of protein complexes, resulting in inefficiency in biological applications.

Method used

The N-level folding iterative network layer is adopted, combining multi-sequence alignment features and relative position transformations to predict the torsion angle and position transformation of amino acid residues in protein complexes, taking into account the relative independence of monomer chains, and improving the accuracy of structural prediction.

Benefits of technology

It improves the efficiency and accuracy of protein complex structure prediction, adapts to the application scenarios of multiple chains, and enhances the overall effect of the protein structure prediction model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117577170B_ABST
    Figure CN117577170B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method and apparatus for predicting the structure of a protein complex, which relates to technical fields such as artificial intelligence, natural language processing, and biological computing. The method includes: obtaining the initial coordinates of each amino acid residue in a target protein complex, and obtaining the target residue pair features, the first multiple sequence alignment features, and the second multiple sequence alignment features of each protein monomer in the target protein complex; inputting the initial coordinates of each amino acid residue, the target residue pair features, the first multiple sequence alignment features, and the second multiple sequence alignment features of each protein monomer into an N-level folding iterative network layer, and predicting the torsion angle, residue-level position transformation, and monomer chain-level position transformation of each amino acid residue by the N-level folding iterative network layer to obtain the target coordinates of each amino acid residue, thereby obtaining the predicted structure of the protein complex. The present disclosure can accurately predict the structure of a protein and improve the efficiency of predicting the structure of a protein complex.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to the fields of natural language processing, biological computing, etc. Background Art

[0002] A protein complex is a stable macromolecular complex formed by the interaction of two or more protein molecules, and plays an important role in different biological functions, such as enzyme reactions, cell signal transduction, metabolic regulation, and gene expression. Among them, the function of a protein is largely determined by its own spatial structure, and the technology of predicting the three-dimensional structure (tertiary structure) of a protein in space based on the amino acid category (primary structure) of the protein chain has extremely high research value in the field of life science.

[0003] Therefore, how to accurately predict the structure of a protein, improve the efficiency of predicting the structure of a protein complex, and meet various biological applications has become one of the important research directions. Summary of the Invention

[0004] The present disclosure provides a method and apparatus for predicting the structure of a protein complex.

[0005] According to one aspect of the present disclosure, there is provided a method for predicting the structure of a protein complex, the method comprising:

[0006] Obtaining the initial coordinates of each amino acid residue in the target protein complex, and obtaining the target residue pair feature, the first multiple sequence alignment feature, and the second multiple sequence alignment feature of each protein monomer in the target protein complex;

[0007] Inputting the initial coordinates of each amino acid residue, the target residue pair feature, the first multiple sequence alignment feature, and the second multiple sequence alignment feature of each protein monomer into an N-level folding iterative network layer, and predicting the torsion angle, residue-level position transformation, and monomer chain-level position transformation of each amino acid residue by the N-level folding iterative network layer to obtain the target coordinates of each amino acid residue, thereby obtaining the predicted structure of the protein complex;

[0008] Wherein, the first multiple sequence alignment feature is the regularized multiple sequence alignment feature, the second multiple sequence alignment feature is the mapped multiple sequence alignment feature, and N is an integer greater than 1.

[0009] The present disclosure considers the relative independence of each monomer chain in the protein complex, adds monomer chain-level position transformation on the basis of residue-level position transformation to update the coordinates of each amino acid residue, can accurately predict the structure of the protein, improve the efficiency of predicting the structure of the protein complex, and better adapt to the application scenario where the protein complex contains multiple chains.

[0010] According to another aspect of the present disclosure, there is provided a structure prediction device for a protein complex, including:

[0011] An acquisition module, configured to acquire the initial coordinates of each amino acid residue in the target protein complex, and acquire the target residue pair features, the first multiple sequence alignment feature, and the second multiple sequence alignment feature of each protein monomer in the target protein complex;

[0012] A structure prediction module, configured to input the initial coordinates of each amino acid residue, the target residue pair features, the first multiple sequence alignment feature, and the second multiple sequence alignment feature of each protein monomer into an N-level folding iterative network layer, and the N-level folding iterative network layer predicts the torsion angle of each amino acid residue, the residue-level position transformation, and the monomer chain-level position transformation, so as to obtain the target coordinates of each amino acid residue and obtain the predicted structure of the protein complex;

[0013] Wherein, the first multiple sequence alignment feature is a regularized multiple sequence alignment feature, the second multiple sequence alignment feature is a mapped multiple sequence alignment feature, and N is an integer greater than 1.

[0014] According to another aspect of the present disclosure, there is provided an electronic device, including at least one processor, and

[0015] A memory communicatively connected to the at least one processor; wherein,

[0016] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the structure prediction method of the protein complex in the first aspect embodiment of the present disclosure.

[0017] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the structure prediction method of the protein complex in the first aspect embodiment of the present disclosure.

[0018] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program, and the computer program realizes the steps of the structure prediction method of the protein complex in the first aspect embodiment of the present disclosure when executed by a processor.

[0019] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0021] Figure 1 is a flowchart of a method for predicting the structure of a protein complex according to an embodiment of the present disclosure;

[0022] Figure 2 is a flowchart of a method for predicting the structure of a protein complex according to an embodiment of the present disclosure;

[0023] Figure 3 is a structural diagram of a method for predicting the structure of a protein complex according to an embodiment of the present disclosure;

[0024] Figure 4 is a flowchart of a method for predicting the structure of a protein complex according to an embodiment of the present disclosure;

[0025] Figure 5 is a structural diagram of a method for predicting the structure of a protein complex according to an embodiment of the present disclosure;

[0026] Figure 6 is a structural diagram of a device for predicting the structure of a protein complex according to an embodiment of the present disclosure;

[0027] Figure 7 is a block diagram of an electronic device for implementing the method of the embodiment of the present disclosure. Detailed Embodiments

[0028] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist in understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.

[0029] The embodiments of the present disclosure relate to the fields of artificial intelligence technologies such as computer vision and deep learning.

[0030] Artificial Intelligence (AI) is abbreviated in English as AI. It is a new technical science that studies, develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence.

[0031] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can enable effective communication between humans and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life. So it has a close connection with the research of linguistics, but there are also important differences. Natural language processing does not generally study natural language, but aims to develop computer systems, especially software systems, that can effectively achieve natural language communication. Thus, it is a part of computer science.

[0032] Bio-computation refers to a new computing model developed by leveraging the inherent information processing mechanism of biological systems. Bio-computation research includes two aspects: devices and systems. It uses an ordered system composed of organic (or biological) materials at the molecular scale to provide basic units for information detection, processing, transmission, and storage through physical and chemical processes at the molecular level.

[0033] The following describes the method and apparatus for predicting the structure of a protein complex according to the present disclosure with reference to the accompanying drawings.

[0034] Figure 1 is a flowchart of a method for predicting the structure of a protein complex according to an embodiment of the present disclosure, as Figure 1 shown. The method includes the following steps:

[0035] S101, obtain the initial coordinates of each amino acid residue in the target protein complex, and obtain the target residue pair features, the first multiple sequence alignment features, and the second multiple sequence alignment features of each protein monomer in the target protein complex.

[0036] A protein complex has multiple protein monomers, and each protein monomer has an amino acid sequence. When amino acids combine to form peptide bonds, they lose a molecule of water. Therefore, the amino acid units in polypeptides / proteins are called amino acid residues. In the embodiments of the present disclosure, in order to adapt to the rotational invariance of the protein structure, the relative position transformation is used to represent the coordinates of each residue, and the spatial structure of the protein complex is initialized with the coordinate origin. That is to say, the coordinates of each amino acid residue in the target protein complex are initialized to obtain the initial coordinates as where represents the far point coordinates represented by rotation / translation, I is the identity matrix indicating no rotation, the vector indicates no translation, and i represents the i-th amino acid residue.

[0037] In the embodiments of the present disclosure, the template features of each protein monomer are obtained, and the paired features of the amino acid sequence of each protein monomer are constructed. For each protein monomer, the target residue pair features of the protein monomer are obtained according to the template features and paired features of the protein monomer.

[0038] In some embodiments, for each protein monomer, according to the target amino acid sequence of the protein monomer, homologous sequences of the protein monomer are retrieved from multiple gene sequence databases, multiple sequence alignment of the homologous sequences of the protein monomer is performed, and the multiple sequence alignment features of the protein monomer are obtained. Then, different processes are performed based on the multiple sequence alignment features to obtain the first multiple sequence alignment feature and the second multiple sequence alignment feature. Optionally, the first multiple sequence alignment feature is the regularized multiple sequence alignment feature, and the second multiple sequence alignment feature is the mapped multiple sequence alignment feature.

[0039] S102, input the initial coordinates of each amino acid residue, the target residue pair features of each protein monomer, the first multiple sequence alignment feature, and the second multiple sequence alignment feature into the N-level folding iterative network layer. The N-level folding iterative network layer predicts the torsion angle of each amino acid residue, the residue-level position transformation, and the monomer chain-level position transformation to obtain the target coordinates of each amino acid residue, and the predicted structure of the protein complex is obtained.

[0040] Wherein, N is an integer greater than 1.

[0041] Optionally, the torsion angle on the residue side chain can be predicted based on the Side Chain and torsion angle predictor in the N-level folding iterative network layer.

[0042] In some embodiments, when performing structure prediction on a protein complex, the residue encoding of multiple chains in the complex is directly mapped to a coordinate transformation, and these transformations only act on the residues. The present disclosure refers to this transformation as the residue-level position transformation.

[0043] In the embodiments of the present disclosure, considering the relative independence of each monomer chain in the protein complex, a monomer chain-level position transformation is added on the basis of the residue-level position transformation to update the coordinates of each amino acid residue, which can decouple the prediction of the residue positions within the chain and the prediction of the overall position of the sub-chain, and enhance the overall effect of the protein structure prediction model.

[0044] In the embodiments of the present disclosure, the initial coordinates of each amino acid residue, the target residue pair features of each protein monomer, the first multiple sequence alignment feature, and the second multiple sequence alignment feature are input into the N-level folding iterative network layer. The N-level folding iterative network layer predicts the torsion angle of each amino acid residue, the residue-level position transformation, and the monomer chain-level position transformation to obtain the target coordinates of each amino acid residue. Considering the relative independence of each monomer chain in the protein complex, the monomer chain-level position transformation is added on the basis of the residue-level position transformation to update the coordinates of each amino acid residue, so as to accurately predict the structure of the protein, improve the efficiency of the structure prediction of the protein complex, and better adapt to the application scenario where the protein complex contains multiple chains.

[0045] Figure 2 is a flowchart of a method for predicting the structure of a protein complex according to an embodiment of the present disclosure, as Figure 2 shown. The method includes the following steps:

[0046] S201, obtain the initial coordinates of each amino acid residue in the target protein complex, and obtain the target residue pair features, the first multiple sequence alignment feature, and the second multiple sequence alignment feature of each protein monomer in the target protein complex.

[0047] For the introduction of step S201, reference can be made to the relevant introduction in the above embodiments, which will not be elaborated here.

[0048] S202, input the initial coordinates, the target residue pair features, and the second multiple sequence alignment feature into the first-level folding iterative network layer to predict the residue-level position transformation and the monomer chain-level position transformation for each amino acid residue, and obtain the target residue encoding 1 and the candidate position transformation 1 of the first-level folding iterative network layer.

[0049] The first-level folding iterative network layer performs an invariant point attention mechanism on the initial coordinates, the target residue pair features, and the second multiple sequence alignment feature to obtain a residue encoding with rotational invariance, and then performs a mapping process through a linear network to obtain the target residue encoding 1.

[0050] Based on the backbone update algorithm, the target residue encoding of each amino acid residue is mapped to implement the prediction of the residue-level position transformation of the target residue encoding 1, obtain the first position transformation 1 of each amino acid residue, and perform the prediction of the monomer chain-level position transformation on the target residue encoding 1 to obtain the second position transformation 1 of each amino acid residue.

[0051] As Figure 3As shown, in some embodiments, the Chain AffineUpdate process at the monomer chain level includes: for each amino acid residue, two or more adjacent amino acid residues are divided into different monomer chains according to the target residue encoding of the amino acid residue. For example, three adjacent amino acid residues can be divided into the same monomer chain. That is to say, given the residue encoding [s1, s2, s3, …, s i , …, s r , the monomer chain before splicing can be located through the following table [1, 2, 3, …, i, …, r] (such as s1 to s3 belong to monomer chain 1, s r-2 to s r belong to monomer chain n). After averaging the target residue encodings of each monomer chain, a chain-level representation, that is, a candidate residue encoding, is obtained.

[0052] In the embodiments of the present disclosure, for the target amino acid residue on any monomer chain, the mean value of the target residue encoding of the target amino acid residue is calculated to obtain a chain-level candidate residue encoding, and based on a multi-layer neural network structure (such as the multi-layer linear network Linear shown in Figure 3 ), the candidate residue encoding is mapped to obtain the second position transformation of each amino acid residue on the monomer chain.

[0053] As shown in Figure 3 , in some embodiments, the multi-layer neural network structure includes a three-layer linear network structure. The candidate residue encoding is input into the first linear network for mapping to obtain a first transformation representation. The first transformation representation is input into the second linear network for mapping to obtain a second transformation representation. The first transformation representation and the second transformation representation are input into the third linear network for mapping to obtain the second position transformation of each amino acid residue on the monomer chain. The structures of the three-layer linear networks can be the same or different, and the embodiments of the present application do not limit this.

[0054] Based on the first position transformation 1, the second position transformation 1, and the initial coordinates, the position is updated to obtain the candidate position transformation 1 of the first-level folding iteration network layer.

[0055] S203. For the mth-level folding iteration network layer, the target residue pair feature, the target residue encoding m-1 of the (m - 1)th-level folding iteration network layer, and the candidate position transformation m-1 are input into the mth-level folding iteration network layer to predict the residue-level position transformation and the monomer chain-level position transformation for each amino acid residue, and the target residue encoding m and the candidate position transformation m of the mth-level folding iteration network layer are obtained, where m takes values from 2 to N.

[0056] Apply the invariant point attention mechanism to the candidate position transformation m-1 and the target residue pair features of the (m-1)-th level folded iterative network layer in the m-th level folded iterative network layer of the input, to obtain residue encodings with rotational invariance, and then perform mapping processing through a linear network to obtain the target residue encoding m.

[0057] Similarly, continue to predict the residue-level position transformation of the target residue encoding m using the method in step S202, to obtain the first position transformation m of each amino acid residue, predict the monomer chain-level position transformation of the target residue encoding m, to obtain the second position transformation m of each amino acid residue, and obtain the candidate position transformation m of the m-th level folded iterative network layer based on the first position transformation m and the second position transformation m, until the candidate position transformation N of the N-th level folded iterative network layer and the target residue encoding N are obtained.

[0058] S204. The N-th level folded iterative network layer predicts the side chains and torsion angles of the first multiple sequence alignment feature and the target residue encoding N of the N-th level folded iterative network layer, to obtain the torsion angles on the side chains of each amino acid residue, and obtain the target coordinates of each amino acid residue based on the torsion angles on the side chains of each amino acid residue and the candidate position transformation N of the N-th level folded iterative network layer.

[0059] Input the first multiple sequence alignment feature and the target residue encoding N of the N-th level folded iterative network layer into the side chain and torsion angle predictor of the N-th level folded iterative network layer, to obtain the torsion angles on the side chains of amino acid residues, and obtain the target coordinates of amino acid residues based on the torsion angles on the side chains of each amino acid residue, the candidate position transformation N of the N-th level folded iterative network layer, and retrograde position update.

[0060] In the embodiments of the present application, considering the relative independence of each monomer chain in the protein complex, the monomer chain-level position transformation is added on the basis of the residue-level position transformation to update the coordinates of each amino acid residue, which can decouple the prediction of the in-chain residue positions and the prediction of the overall position of the sub-chain, enhance the overall effect of the protein structure prediction model, and can better globally adjust the docking relationship between chains while retaining the relative positions of the internal residues in a single chain, and is more suitable for the prediction of the structure of protein complexes.

[0061] Figure 4 is a flowchart of a method for predicting the structure of a protein complex according to an embodiment of the present disclosure, as Figure 4 shown, the method includes the following steps:

[0062] S401. Obtain the template features of each protein monomer and construct the paired features of the amino acid sequences of each protein monomer.

[0063] In some embodiments, the target amino acid sequence of each protein monomer is matched and queried against multiple first amino acid sequences in a protein structure database to obtain second amino acid sequences with a similarity greater than a preset threshold, and the distances between the coordinates of the amino acid residues of the second amino acid sequences are extracted as the template features of each protein monomer. That is to say, the amino acid sequence of the protein monomer is searched in the resolved protein structure database for protein structures with similar sequences, and based on tools for protein sequence analysis, such as the search method based on the hidden Markov model (HMM) (HHSearch), the distances between residues are extracted as the template features.

[0064] In some embodiments, the amino acid sequence of each protein monomer is input into two preset linear networks to obtain candidate sequence encoding features. An empty dimension is added to different directions of the candidate sequence encoding features to obtain a first sequence encoding feature and a second sequence encoding feature, and the first sequence encoding feature and the second sequence encoding feature are added together to obtain the pairing feature of each protein monomer. The complex sequence with a length of r after splicing multiple sequences is encoded by two linear network Linear layers to obtain a sequence encoding feature with a shape of [r, c], that is, the first sequence encoding feature z1 and the second sequence encoding feature z2, where c is the depth of the hidden layer of the Linear network (hyperparameter). Thereafter, an empty dimension is added to z1 and z2 (the shape of z1 is converted to [r, 1, c], and the shape of z2 is converted to [1, r, c]), and added together to obtain the matching pair feature z pair , where z pair has a shape of [r, r, c], and z pair = z1 + z2.

[0065] S402, after mapping the template features of each protein monomer through a linear network, add them to the pairing features of each protein monomer to obtain candidate residue pair features;

[0066] In the embodiments of the present disclosure, the Template features have a shape of [r, r] after splicing, and then are encoded through a Linear layer to obtain a feature z with the same shape as the Piar feature temp , and the feature z temp is added to the pairing feature to obtain candidate residue pair features.

[0067] S403, input the candidate residue pair features into a preset encoder for encoding to obtain the target residue pair features of each protein monomer.

[0068] In the embodiments of the present disclosure, the candidate residue pair features are input into an encoder (Evofomer Encoder) for encoding to obtain the target residue pair features of each protein monomer.

[0069] S404. Query and obtain the homologous sequences of each protein monomer from multiple gene sequence databases according to the target amino acid sequence of each protein monomer.

[0070] In the present invention, first, the amino acid sequences of each protein monomer in the complex are used as query requests, and homologous sequences are searched in multiple gene sequence databases. Using existing tools JackHMMER and HHblits can achieve a more in-depth analysis and annotation of protein sequences. Using JackHMMER, a fast heuristic search of the hidden Markov model (HMM) can be performed, and using HHblits, a more in-depth annotation of the already discovered protein sequences can be carried out, so as to obtain the homologous sequences of each protein monomer.

[0071] S405. Perform multiple sequence alignment on the homologous sequences of each protein monomer to obtain the candidate multiple sequence alignment features of each protein monomer.

[0072] On the obtained homologous sequences, multiple sequence alignment is used to obtain the multiple sequence alignment features (MSA) of each monomer.

[0073] S406. Input the candidate multiple sequence alignment features of each protein monomer into a preset encoder for encoding to obtain the target multiple sequence alignment features of each protein monomer.

[0074] Input the candidate multiple sequence alignment features of each protein monomer into the encoder (Evofomer Encoder) for encoding to obtain the target multiple sequence alignment features of each protein monomer.

[0075] S407. Regularize the target multiple sequence alignment features of each protein monomer to obtain the first multiple sequence alignment features of each protein monomer, and map the target multiple sequence alignment features of each protein monomer to obtain the second multiple sequence alignment features of each protein monomer.

[0076] Perform normalization (Norm) processing on the target multiple sequence alignment features of each protein monomer to obtain the first multiple sequence alignment features of each protein monomer, and map the target multiple sequence alignment features of each protein monomer based on a linear network to obtain the second multiple sequence alignment features of each protein monomer.

[0077] S408. Input the initial coordinates of each amino acid residue, the target residue pair features of each protein monomer, the first multiple sequence alignment features, and the second multiple sequence alignment features into the N-level folding iterative network layer. The N-level folding iterative network layer predicts the torsion angle of each amino acid residue, the residue-level position transformation, and the monomer chain-level position transformation to obtain the target coordinates of each amino acid residue, and obtain the predicted structure of the protein complex.

[0078] For the introduction of step S408, reference can be made to the relevant content in the above embodiments, which will not be elaborated here.

[0079] In the embodiments of the present application, obtaining the initial coordinates of each amino acid residue in the target protein complex, and obtaining the target residue pair features, the first multiple sequence alignment features, and the second multiple sequence alignment features of each protein monomer in the target protein complex effectively promotes the development of the protein monomer structure prediction task and can enhance the overall effect of protein structure prediction.

[0080] Figure 5 is a structural diagram of a method for predicting the structure of a protein complex according to an embodiment of the present disclosure. As Figure 5 shown, in the embodiments of the present disclosure, homologous sequences are searched in multiple gene sequence databases (Sequence Data Base) according to the amino acid sequences [Sequence 1,..., Sequence N] of each protein monomer in the target protein complex, and multiple sequence alignment is performed to obtain the multiple sequence alignment features [MSA 1,..., MSAN] of each protein monomer. The multiple sequence alignment features [MSA 1,..., MSA N] are input into the Evofomer Encoder for encoding to obtain the target multiple sequence alignment features. The amino acid sequence of each protein monomer is input into two preset linear networks Linear to generate paired Pair features. The target amino acid sequence of each protein monomer is used to search for protein structures with sequence similarity in the protein structure database (Structure Data Base), and the distance between residues is extracted as the template (Template) feature. Optionally, Pair and Merge means merging, that is, the MSA representation feature and the pair representation feature are merged. After the template feature of each protein monomer is input into the linear network for mapping, it is added to the paired feature of each protein monomer and input into the Evofomer encoder for encoding to obtain the target residue pair feature of each protein monomer. The Evofomer encoder can extract the hidden layer encoding of each residue from the MSA, Pair, and Template data. In the decoding stage, the present disclosure can be based on the structure prediction module (Structure Module) in the AF2Multimer model for protein complex structure prediction, auxiliary training tasks such as MSA masking prediction (Mask MSA), LDDT prediction (LDDT is a pre-existing metric method in the field of protein structure prediction), and residue distance prediction. Among them, initial frame represents the initial coordinates.

[0081] As shown Figure 5 in the figure, a protein complex containing r residues is input. After passing through the Evofomer Encoder, the model obtains the MSA feature encoding of each residue i of the target protein complex, that is, the target multiple sequence alignment feature i ∈ [0, 1,..., r] (c s is the hyperparameter hidden layer size), and the residue pair encoding containing Pair features and Template features, that is, Z i,j ∈ R 1×c , i, j ∈ [0, 1,..., r]. The target multiple sequence alignment features of each protein monomer are input into the regularization Norm layer network to obtain the first multiple sequence alignment features of each protein monomer. The target multiple sequence alignment features of each protein monomer are input into the linear network Linear to obtain the second multiple sequence alignment features of each protein monomer. Among them, R represents the shape, which here expresses that the shape of is [1, c], where c is the hidden layer depth of the model

[0082] In the embodiments of the present disclosure, an N-level folding iteration network layer (FoldIteration module) is used to predict the structure of the target protein complex. Optionally, N can take the value of 8. In other implementations, N can take other values, and the embodiments of the present application do not limit this

[0083] To adapt to the rotational invariance of the protein structure, the present invention adopts relative position transformation represents the coordinates of each residue, and the spatial structure of the protein complex is initialized with the coordinate origin The present invention first updates with two regularization Norm layer networks and Z i,j encoding, and maps with the linear Linear layer to the hidden layer representation s i on, where s i represents the encoding of the i-th residue, Z i,j represents the pair encoding from residue i to residue j, T i represents the rotation and translation of residue i. R i is the rotation transformation of residue i, t i is the translation transformation of residue i. Based on the AlphaFold model, the absolute coordinates are converted into relative rotation and translation to represent the residue coordinates to achieve rotational invariance

[0084] Among them, each layer of FoldIteration accesses s i , Z i,j , and T iAfter that, first, the invariant point attention module is used to obtain the residue encoding s with rotational invariance. i After passing through network layers such as the Linear layer, Norm layer, and Dropout layer, the present disclosure uses the obtained encoding to predict the torsion angles of the side chains of each residue and the coordinates of each residue Among them, the Dropout layer is used to randomly discard network parameters and has a relatively small effect. Figure 5 It is not shown in [the figure] and this layer can be omitted or deleted.

[0085] As Figure 5 shown, in the embodiments of the present disclosure, in each Fold Iteration layer, the side chain and torsion angle predictor is used to predict the torsion angle on the side chain of residue i, where f ∈ {ωφ ψ χ1 χ2 χ3 χ4} represents the 7 components that can be twisted on each residue.

[0086] In the prediction of the backbone network structure of the protein complex, the present disclosure first encodes the residue features using a shallow neural network (Linear and Norm) structure, and then uses the BackboneUpdate algorithm to predict the Euclidean transformation T i of each residue i, that is, the first position transformation. In BackboneUpdate, the hidden layer features are mapped to a 6-dimensional representation. Among them, the first three dimensions b i c i d i are used as the rotation matrix R of residue i through an equation. i The last three dimensions directly represent the translational transformation of the residue. Among them, the input of FoldIterationd does not consider the side chain and torsion angle predicted in the previous step, but only considers the position transformation affine of the backbone and the residue encoding.

[0087] After obtaining the prediction of the spatial position transformation of residue i, the model updates the relative position of the residue according to to complete the residue-level position transformation. In the formula, the first T i represents the updated T i and the second T i represents the T before update. i .

[0088] Based on the above transformations, the present disclosure introduces a position transformation Chainaffine module at the monomer chain level to predict residue transformations at the chain level. The Chainaffine module aims to predict the overall transformation of chain k of the protein complex. The module structure can be implemented by various methods, such as Figure 3 As shown, it is one of the implementation methods of the Chainaffine module. Taking s i as the input, the Chainaffine module first divides s i into different monomer chains according to the residue position encoding, and obtains the hidden layer representation at the chain level after calculating the mean of all residue representations of the same monomer. k ∈ [0, 1,..., n], (n is the number of sub-chains of the protein complex, and d is the hidden layer size of the Chainaffine network).

[0089] After that, the Chainaffine module uses a multi-layer neural network structure to map to a transformation representation containing 6 dimensions, and obtains the spatial position transformation of each chain using the same method as BackboneUpdate, that is, the second position transformation.

[0090] As Figure 5 shown, when updating the residue positions within the chain, all residues i on chain k share the transformation to obtain the second position transformation at the chain level for each residue. Finally, the model updates the positions of each residue in the complex according to The first T in the formula i represents the updated T i , and the second T i represents the T before update i . Here, acts on the front side of T i to transform the residues within the chain around the origin.

[0091] After obtaining the Euclidean transformation of the backbone network and the torsion angles of the side chains, in the embodiments of the present application, based on the residue update module residue update and the update frame Update frame module, T i and are converted into the three-dimensional coordinates of each residue to complete the prediction of the protein tertiary structure. Among them, the Angles module predicts the side chain torsion angles, and the Coordinates convert module outputs the converted spatial coordinates after receiving the backbone transformation and the side chain torsion angles.

[0092] Figure 6The structural diagram of a protein complex structure prediction device according to an embodiment of the present disclosure is as follows: Figure 6 As shown, the protein complex structure prediction device 600 includes:

[0093] An acquisition module 610, configured to acquire the initial coordinates of each amino acid residue in the target protein complex, and acquire the target residue pair features, the first multiple sequence alignment feature, and the second multiple sequence alignment feature of each protein monomer in the target protein complex;

[0094] A structure prediction module 620, configured to input the initial coordinates of each amino acid residue, the target residue pair features, the first multiple sequence alignment feature, and the second multiple sequence alignment feature of each protein monomer into an N-level folding iterative network layer, and the N-level folding iterative network layer predicts the torsion angle, residue-level position transformation, and monomer chain-level position transformation of each amino acid residue to obtain the target coordinates of each amino acid residue, and obtain the predicted structure of the protein complex;

[0095] Wherein, the first multiple sequence alignment feature is the regularized multiple sequence alignment feature, the second multiple sequence alignment feature is the mapped multiple sequence alignment feature, and N is an integer greater than 1.

[0096] In some embodiments, the structure prediction module 620 is further configured to:

[0097] Input the initial coordinates, the target residue pair features, and the second multiple sequence alignment feature into the first-level folding iterative network layer to predict the residue-level position transformation and monomer chain-level position transformation of each amino acid residue, and obtain the target residue encoding 1 and the candidate position transformation 1 of the first-level folding iterative network layer;

[0098] For the m-level folding iterative network layer, input the target residue pair features, the target residue encoding m-1 of the m-1 level folding iterative network layer, and the candidate position transformation m-1 into the m-level folding iterative network layer to predict the residue-level position transformation and monomer chain-level position transformation of each amino acid residue, and obtain the target residue encoding m and the candidate position transformation m of the m-level folding iterative network layer, where m ranges from 2 to N;

[0099] The N-level folding iterative network layer performs side chain and torsion angle prediction on the first multiple sequence alignment feature and the target residue encoding N of the N-level folding iterative network layer, obtains the torsion angle on the side chain of each amino acid residue, and obtains the target coordinates of each amino acid residue based on the torsion angle on the side chain of each amino acid residue and the candidate position transformation N of the N-level folding iterative network layer.

[0100] In some embodiments, the structure prediction module 620 is further configured to:

[0101] The initial coordinates, target residue pair features, and second multiple sequence alignment features are processed by the first-level folded iterative network layer through an invariant point attention mechanism and mapping to obtain the target residue encoding 1;

[0102] Predict the residue-level position transformation for the target residue encoding 1 to obtain the first position transformation 1 for each amino acid residue, and predict the monomer chain-level position transformation for the target residue encoding 1 to obtain the second position transformation 1 for each amino acid residue;

[0103] Perform position update based on the first position transformation 1, the second position transformation 1, and the initial coordinates to obtain the candidate position transformation 1 of the first-level folded iterative network layer.

[0104] In some embodiments, the structure prediction module 620 is further configured to:

[0105] Process the candidate position transformation m-1 of the (m-1)-th level folded iterative network layer and the target residue pair features in the input m-th level folded iterative network layer through an invariant point attention mechanism and mapping to obtain the target residue encoding m;

[0106] Predict the residue-level position transformation for the target residue encoding m to obtain the first position transformation m for each amino acid residue, and predict the monomer chain-level position transformation for the target residue encoding m to obtain the second position transformation m for each amino acid residue;

[0107] Obtain the candidate position transformation m of the m-th level folded iterative network layer based on the first position transformation m and the second position transformation m.

[0108] In some embodiments, the structure prediction module 620 is further configured to:

[0109] Map the target residue encoding of each amino acid residue based on the backbone update algorithm to obtain the first position transformation of each amino acid residue.

[0110] In some embodiments, the structure prediction module 620 is further configured to:

[0111] For each amino acid residue, divide two or more adjacent amino acid residues into different monomer chains according to the target residue encoding of the amino acid residue;

[0112] For the target amino acid residue on any monomer chain, calculate the mean value of the target residue encoding of the target amino acid residue to obtain the candidate residue encoding at the chain level, and map the candidate residue encoding based on a multi-layer neural network structure to obtain the second position transformation of each amino acid residue on this monomer chain.

[0113] In some embodiments, the multi-layer neural network structure includes a three-layer linear network, and the structure prediction module 620 is further configured to:

[0114] Input the candidate residue encoding into the first linear network for mapping to obtain the first transformed representation;

[0115] Input the first transformed representation into the second linear network for mapping to obtain the second transformed representation;

[0116] Input the first transformed representation and the second transformed representation into the third linear network for mapping to obtain the second position transformation of each amino acid residue on the monomer chain.

[0117] In some embodiments, the obtaining module 610 is further configured to:

[0118] Obtain the template features of each protein monomer and construct the paired features of the amino acid sequence of each protein monomer;

[0119] After inputting the template features of each protein monomer into the linear network for mapping, add them to the paired features of each protein monomer to obtain the candidate residue pair features;

[0120] Input the candidate residue pair features into a preset encoder for encoding to obtain the target residue pair features of each protein monomer.

[0121] In some embodiments, the obtaining module 610 is further configured to:

[0122] Match and query the target amino acid sequence of each protein monomer against multiple first amino acid sequences in the protein structure database to obtain a second amino acid sequence with a similarity greater than a preset threshold;

[0123] Extract the distances between the coordinates of the amino acid residues of the second amino acid sequence as the template features of each protein monomer;

[0124] In some embodiments, the obtaining module 610 is further configured to:

[0125] Input the amino acid sequence of each protein monomer into two preset linear networks to obtain candidate sequence encoding features;

[0126] Add an empty dimension to different directions of the candidate sequence encoding features to obtain a first sequence encoding feature and a second sequence encoding feature;

[0127] Add the first sequence encoding feature and the second sequence encoding feature to obtain the paired features of each protein monomer.

[0128] In some embodiments, the obtaining module 610 is further configured to:

[0129] Query and obtain the homologous sequences of each protein monomer from multiple gene sequence databases according to the target amino acid sequence of each protein monomer;

[0130] Perform a multiple sequence alignment on the homologous sequences of each protein monomer to obtain the candidate multiple sequence alignment features of each protein monomer;

[0131] Input the candidate multiple sequence alignment features of each protein monomer into a preset encoder for encoding to obtain the target multiple sequence alignment features of each protein monomer;

[0132] Regularize the target multiple sequence alignment features of each protein monomer to obtain the first multiple sequence alignment features of each protein monomer, and map the target multiple sequence alignment features of each protein monomer to obtain the second multiple sequence alignment features of each protein monomer.

[0133] The present disclosure takes into account the relative independence of each monomer chain in a protein complex, adds a position transformation at the monomer chain level on the basis of the position transformation at the residue level to update the coordinates of each amino acid residue, can accurately predict the structure of a protein, improve the efficiency of the structure prediction of a protein complex, and better adapt to the application scenario where a protein complex contains multiple chains.

[0134] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0135] Figure 7 is a block diagram of an electronic device for implementing the embodiments of the present disclosure. The electronic device can implement the method for predicting the structure of a protein complex in the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0136] As Figure 7 shown, the device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other through a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0137] Multiple components in device 700 are connected to I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, an optical disc, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0138] The computing unit 701 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 executes the various methods and processes described above, such as the method for predicting the structure of a protein complex. For example, in some embodiments, the method for predicting the structure of a protein complex can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the method for predicting the structure of a protein complex described above can be executed. Alternatively, in other embodiments, the computing unit 701 can be configured to execute the method for predicting the structure of a protein complex by any other suitable means (e.g., by means of firmware).

[0139] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), systems-on-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs, which can be executed and / or interpreted on a programmable system including at least one programmable processor, and the programmable processor can be a special or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0140] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, a special purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0141] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0142] In order to provide interaction with a user, the systems and techniques described herein may be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).

[0143] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0144] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, a server of a distributed system, or a server incorporating blockchain.

[0145] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.

[0146] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A method for predicting the structure of a protein complex, wherein, Including: Obtain the initial coordinates of each amino acid residue in the target protein complex, and obtain the target residue pair features, first multiple sequence alignment features, and second multiple sequence alignment features of each protein monomer in the target protein complex; Input the initial coordinates of each amino acid residue, the target residue pair features, first multiple sequence alignment features, and second multiple sequence alignment features of each protein monomer into an N-level folding iterative network layer, and the N-level folding iterative network layer predicts the torsion angle, residue-level position transformation, and monomer chain-level position transformation of each amino acid residue to obtain the target coordinates of each amino acid residue, thereby obtaining the predicted structure of the protein complex; Wherein, the first multiple sequence alignment feature is a regularized multiple sequence alignment feature, the second multiple sequence alignment feature is a mapped multiple sequence alignment feature, and N is an integer greater than 1; Wherein, the process of predicting the monomer chain-level position transformation for the target residue encoding of each amino acid residue and obtaining the second position transformation of each amino acid residue includes: For each amino acid residue, according to the target residue encoding of the amino acid residue, divide two or more adjacent amino acid residues into different monomer chains; For the target amino acid residue on any monomer chain, calculate the mean value of the target residue encoding of the target amino acid residue to obtain the candidate residue encoding at the chain level, and map the candidate residue encoding based on a multi-layer neural network structure to obtain the second position transformation of each amino acid residue on this monomer chain.

2. The method according to claim 1, wherein The method further includes: inputting the initial coordinates, the target residue pair features, and the second multiple sequence alignment features into the first-level folding iterative network layer to predict the residue-level position transformation and monomer chain-level position transformation of each amino acid residue, and obtaining the target residue encoding 1 and candidate position transformation 1 of the first-level folding iterative network layer; For the m-th level folding iterative network layer, input the target residue pair features, the target residue encoding m-1 and candidate position transformation m-1 of the (m-1)-th level folding iterative network layer into the m-th level folding iterative network layer to predict the residue-level position transformation and monomer chain-level position transformation of each amino acid residue, and obtain the target residue encoding m and candidate position transformation m of the m-th level folding iterative network layer, where m takes values from 2 to N; The N-th level folding iterative network layer predicts the side chain and torsion angle for the first multiple sequence alignment features and the target residue encoding N of the N-th level folding iterative network layer, obtains the torsion angle on the side chain of each amino acid residue, and obtains the target coordinates of each amino acid residue based on the torsion angle on the side chain of each amino acid residue and the candidate position transformation N of the N-th level folding iterative network layer.

3. The method according to claim 2, wherein, Inputting the initial coordinates, the target residue pair features, and the second multiple sequence alignment features into the first-level folding iterative network layer to predict residue-level position transformation and monomer chain-level position transformation for each amino acid residue, and obtaining the target residue encoding 1 and candidate position transformation 1 of the first-level folding iterative network layer, includes: The first-level folding iterative network layer performs invariant point attention mechanism and mapping processing on the initial coordinates, the target residue pair features, and the second multiple sequence alignment features to obtain the target residue encoding 1; Performing residue-level position transformation prediction on the target residue encoding 1 to obtain the first position transformation 1 of each amino acid residue, and performing monomer chain-level position transformation prediction on the target residue encoding 1 to obtain the second position transformation 1 of each amino acid residue; Based on the first position transformation 1, the second position transformation 1, and the initial coordinates, perform position update to obtain the candidate position transformation 1 of the first-level folding iterative network layer.

4. The method according to claim 2, wherein, Inputting the target residue pair features, the target residue encoding m-1 and candidate position transformation m-1 of the (m-1)-th level folding iterative network layer into the m-th level folding iterative network layer to predict residue-level position transformation and monomer chain-level position transformation for each amino acid residue, and obtaining the target residue encoding m and candidate position transformation m of the m-th level folding iterative network layer, includes: Performing invariant point attention mechanism and mapping processing on the candidate position transformation m-1 of the (m-1)-th level folding iterative network layer and the target residue pair features input into the m-th level folding iterative network layer to obtain the target residue encoding m; Performing residue-level position transformation prediction on the target residue encoding m to obtain the first position transformation m of each amino acid residue, and performing monomer chain-level position transformation prediction on the target residue encoding m to obtain the second position transformation m of each amino acid residue; Based on the first position transformation m and the second position transformation m, obtain the candidate position transformation m of the m-th level folding iterative network layer.

5. The method according to claim 2, wherein, The process of performing residue-level position transformation prediction on the target residue encoding of each amino acid residue to obtain the first position transformation of each amino acid residue, includes: Based on the backbone update algorithm, map the target residue encoding of each amino acid residue to obtain the first position transformation of each amino acid residue.

6. The method according to claim 1, wherein, The multi-layer neural network structure includes three layers of linear networks. The process of mapping the candidate residue encoding based on the multi-layer neural network structure to obtain the second position transformation of each amino acid residue on the monomer chain, includes: Input the candidate residue encoding into the first linear network for mapping to obtain the first transformation representation; Input the first transformation representation into the second linear network for mapping to obtain the second transformation representation; Input the first transformation representation and the second transformation representation into the third linear network for mapping to obtain the second position transformation of each amino acid residue on the monomer chain.

7. The method according to claim 1, wherein Obtain the target residue pair features of each protein monomer in the target protein complex, including: Obtain the template features of each protein monomer and construct the paired features of the amino acid sequence of each protein monomer; After inputting the template features of each protein monomer into a linear network for mapping, add them to the paired features of each protein monomer to obtain candidate residue pair features; Input the candidate residue pair features into a preset encoder for encoding to obtain the target residue pair features of each protein monomer.

8. The method according to claim 7, wherein The obtaining of the template features of each protein monomer includes: Match and query the target amino acid sequence of each protein monomer against multiple first amino acid sequences in a protein structure database to obtain second amino acid sequences with a similarity greater than a preset threshold; Extract the distances between the coordinates of the amino acid residues of the second amino acid sequence as the template features of each protein monomer.

9. The method according to claim 7, wherein The constructing of the paired features of the amino acid sequence of each protein monomer includes: Input the amino acid sequence of each protein monomer into two preset linear networks to obtain candidate sequence encoding features; Add an empty dimension to different directions of the candidate sequence encoding features to obtain first sequence encoding features and second sequence encoding features; Add the first sequence encoding features and the second sequence encoding features to obtain the paired features of each protein monomer.

10. The method according to claim 1, wherein Obtain the first multiple sequence alignment feature and the second multiple sequence alignment feature of each protein monomer in the target protein complex, including: Query and obtain the homologous sequences of each protein monomer from multiple gene sequence databases according to the target amino acid sequence of each protein monomer; Perform multiple sequence alignment on the homologous sequences of each protein monomer to obtain the candidate multiple sequence alignment features of each protein monomer; Input the candidate multiple sequence alignment features of each protein monomer into a preset encoder for encoding to obtain the target multiple sequence alignment features of each protein monomer; Regularize the target multiple sequence alignment features of each protein monomer to obtain the first multiple sequence alignment feature of each protein monomer, and map the target multiple sequence alignment features of each protein monomer to obtain the second multiple sequence alignment feature of each protein monomer.

11. A structural prediction device for a protein complex, wherein, Including: An obtaining module for obtaining the initial coordinates of each amino acid residue in the target protein complex and obtaining the target residue pair features, the first multiple sequence alignment feature, and the second multiple sequence alignment feature of each protein monomer in the target protein complex; A structure prediction module for inputting the initial coordinates of each amino acid residue, the target residue pair features, the first multiple sequence alignment feature, and the second multiple sequence alignment feature of each protein monomer into an N-level folding iterative network layer, and the N-level folding iterative network layer predicts the torsion angle, residue-level position transformation, and monomer chain-level position transformation of each amino acid residue to obtain the target coordinates of each amino acid residue and obtain the predicted structure of the protein complex; Wherein, the first multiple sequence alignment feature is the regularized multiple sequence alignment feature, the second multiple sequence alignment feature is the mapped multiple sequence alignment feature, and N is an integer greater than 1; Wherein, the structure prediction module is further configured to: For each of the amino acid residues, divide two or more adjacent amino acid residues into different monomer chains according to the target residue encoding of the amino acid residue; For a target amino acid residue on any monomer chain, calculate the mean value of the target residue encoding of the target amino acid residue to obtain a candidate residue encoding at the chain level, and map the candidate residue encoding based on a multi-layer neural network structure to obtain a second position transformation of each amino acid residue on the monomer chain.

12. The apparatus according to claim 11, wherein, The structure prediction module is further configured to: Input the initial coordinates, the target residue pair feature, and the second multiple sequence alignment feature into the first-level folding iteration network layer to predict the residue-level position transformation and the monomer-chain-level position transformation for each amino acid residue, and obtain the target residue encoding 1 and the candidate position transformation 1 of the first-level folding iteration network layer; For the m-th level folding iteration network layer, input the target residue pair feature, the target residue encoding m-1 and the candidate position transformation m-1 of the (m-1)-th level folding iteration network layer into the m-th level folding iteration network layer to predict the residue-level position transformation and the monomer-chain-level position transformation for each amino acid residue, and obtain the target residue encoding m and the candidate position transformation m of the m-th level folding iteration network layer, where m ranges from 2 to N; The N-th level folding iteration network layer predicts the side chain and torsion angle for the first multiple sequence alignment feature and the target residue encoding N of the N-th level folding iteration network layer, obtains the torsion angle on the side chain of each amino acid residue, and obtains the target coordinates of each amino acid residue based on the torsion angle on the side chain of each amino acid residue and the candidate position transformation N of the N-th level folding iteration network layer.

13. The apparatus according to claim 12, wherein, The structure prediction module is further configured to: The first-level folding iteration network layer performs invariant point attention mechanism and mapping processing on the initial coordinates, the target residue pair feature, and the second multiple sequence alignment feature to obtain the target residue encoding 1; Predict the residue-level position transformation for the target residue encoding 1 to obtain a first position transformation 1 of each amino acid residue, and predict the monomer-chain-level position transformation for the target residue encoding 1 to obtain a second position transformation 1 of each amino acid residue; Update the position based on the first position transformation 1, the second position transformation 1, and the initial coordinates to obtain the candidate position transformation 1 of the first-level folding iteration network layer.

14. The device according to claim 12, wherein, The structure prediction module is further configured to: Perform invariant point attention mechanism and mapping processing on the candidate position transformation m-1 of the (m-1)-th level folding iteration network layer input into the m-th level folding iteration network layer and the target residue pair feature to obtain the target residue encoding m; Predict the residue-level position transformation for the target residue encoding m to obtain the first position transformation m for each amino acid residue, and predict the monomer-chain-level position transformation for the target residue encoding m to obtain the second position transformation m for each amino acid residue; Obtain the candidate position transformation m of the m-th folding iteration network layer based on the first position transformation m and the second position transformation m.

15. The device according to claim 12, wherein The structure prediction module is further configured to: Map the target residue encoding of each amino acid residue based on the backbone update algorithm to obtain the first position transformation of each amino acid residue.

16. The device according to claim 11, wherein, The multi-layer neural network structure includes three linear networks, and the structure prediction module is further configured to: Input the candidate residue encoding into the first linear network for mapping to obtain the first transformation representation; Input the first transformation representation into the second linear network for mapping to obtain the second transformation representation; Input the first transformation representation and the second transformation representation into the third linear network for mapping to obtain the second position transformation of each amino acid residue on the monomer chain.

17. The apparatus according to claim 11, wherein, The obtaining module is further configured to: Obtain the template features of each protein monomer and construct the paired features of the amino acid sequence of each protein monomer; After inputting the template features of each protein monomer into a linear network for mapping, add them to the paired features of each protein monomer to obtain the candidate residue pair features; Input the candidate residue pair features into a preset encoder for encoding to obtain the target residue pair features of each protein monomer.

18. The device according to claim 17, wherein, The obtaining module is further configured to: Match and query the target amino acid sequence of each protein monomer against multiple first amino acid sequences in the protein structure database to obtain the second amino acid sequence with a similarity greater than a preset threshold; Extract the distances between the coordinates of the amino acid residues of the second amino acid sequence as the template features of each protein monomer.

19. The apparatus according to claim 17, wherein, The obtaining module is further configured to: Input the amino acid sequence of each protein monomer into two preset linear networks to obtain the candidate sequence encoding features; Add an empty dimension to different directions of the candidate sequence encoding features to obtain the first sequence encoding feature and the second sequence encoding feature; Add the first sequence encoding feature and the second sequence encoding feature to obtain the paired features of each protein monomer.

20. The apparatus according to claim 11, wherein The obtaining module is further configured to: Query and obtain the homologous sequences of each protein monomer from multiple gene sequence databases according to the target amino acid sequence of each protein monomer; Perform multiple sequence alignment on the homologous sequences of each protein monomer to obtain the candidate multiple sequence alignment features of each protein monomer; Input the candidate multiple sequence alignment features of each protein monomer into a preset encoder for encoding to obtain the target multiple sequence alignment features of each protein monomer; Regularize the target multiple sequence alignment features of each of the protein monomers to obtain the first multiple sequence alignment features of each of the protein monomers, and map the target multiple sequence alignment features of each of the protein monomers to obtain the second multiple sequence alignment features of each of the protein monomers.

21. An electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method according to any one of claims 1-10.

22. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-10.

23. A computer program product, comprising a computer program which, when executed by a processor, implements the steps of the method according to any one of claims 1-10.

Citation Information

Patent Citations

  • Double-layer mutual enhancement protein three-dimensional structure prediction method and system

    CN113223608A

  • Method for predicting distance between protein residues based on self-attention mechanism

    CN114708903A