Protein structure prediction method and device, equipment and storage medium
By using a protein model containing multiple sub-models, the problem of insufficient prediction accuracy of protein side chain structure and multi-chain protein full atomic structure in the prior art is solved, and higher prediction accuracy and natural protein proximity are achieved.
Patent Information
- Application Number
- CN202510272990.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-06
AI Technical Summary
Existing protein structure prediction methods have insufficient accuracy in predicting side chain structure and full atomic structure of multi-chain proteins, especially for the atomic-level side chain structure prediction of single-chain proteins.
The protein model containing the first sub-model, the second sub-model and the third sub-model is used to predict the following steps: the first sub-model is used to determine the amino acid sequence and the backbone structure; the second sub-model is used to predict the side chain structure; the third sub-model is used to update the sequence and the backbone structure based on the side chain conformation, avoiding conflicts and making the prediction results closer to the natural protein.
Improve the accuracy of protein structure prediction, especially in the prediction of side chain structure and full-atom structure of multi-chain proteins, the generated protein structure is closer to natural proteins.
Smart Images

Figure CN120108486A_ABST
Abstract
Description
Technical Field
[0001] The exemplary embodiments of the present disclosure generally relate to the computer field, and in particular, to a protein structure prediction method, apparatus, device, computer-readable storage medium, and computer program product. Background Art
[0002] Proteins are biological molecules or macromolecules composed of long chains of amino acid residues. Proteins perform many important life activities in organisms, and the functions of proteins are mainly determined by their three-dimensional (3D) structures. Knowing the structure of proteins is very important in the fields of medicine and biotechnology. For example, if a protein plays a key role in a disease, drug molecules can be designed based on the structure of the protein to treat the disease. With the development of artificial intelligence technology, the use of machine learning models to predict protein structures is becoming more and more common. Summary of the invention
[0003] In the first aspect of the present disclosure, a protein structure prediction method is provided. The method includes: based on input information associated with the target protein, using the first sub-model in the protein model, determining the sequence information and backbone structure information of the target protein, the sequence information indicates the amino acid sequence of at least one peptide chain of the target protein, and the backbone structure information indicates the backbone conformation of at least one peptide chain; based on the sequence information and the backbone structure information, using the second sub-model in the protein model, determining the side chain structure information of at least one peptide chain, the side chain structure information indicates the side chain conformation of each residue in at least one peptide chain; and based on the side chain structure information, the sequence information and the backbone structure information, using the third sub-model in the protein model, determining the structure of the target protein.
[0004] In the second aspect of the present disclosure, a device for protein structure prediction is provided. The device includes: a first determination module, configured to determine the sequence information and backbone structure information of the target protein based on the input information associated with the target protein, using the first sub-model in the protein model, the sequence information indicates the amino acid sequence of at least one peptide chain of the target protein, and the backbone structure information indicates the backbone conformation of at least one peptide chain; a second determination module, configured to determine the side chain structure information of at least one peptide chain based on the sequence information and the backbone structure information, using the second sub-model in the protein model, the side chain structure information indicates the side chain conformation of each residue in the at least one peptide chain; and an improvement module, configured to determine the structure of the target protein based on the side chain structure information, the sequence information and the backbone structure information, using the third sub-model in the protein model.
[0005] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processor; and at least one memory, the at least one memory is coupled to the at least one processor and stores instructions for execution by the at least one processor. When the instructions are executed by the at least one processor, the device executes the method of the first aspect.
[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein computer-executable instructions are stored on the computer-readable storage medium, and the computer-executable instructions can be executed by a processor to implement the method of the first aspect.
[0007] In a fifth aspect of the present disclosure, a computer program product is provided, comprising computer executable instructions, wherein when the computer executable instructions are executed by a processor, the method according to the first aspect of the present disclosure is implemented.
[0008] It should be understood that the contents described in this content section are not intended to limit the key features or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:
[0010] Figure 1 A schematic diagram showing an example environment in which embodiments according to the present disclosure may be implemented;
[0011] Figure 2 A schematic diagram showing an example architecture for protein structure prediction according to some embodiments of the present disclosure;
[0012] Figure 3 A schematic diagram showing an example architecture of a first sub-model according to some embodiments of the present disclosure;
[0013] Figure 4 A schematic diagram showing an example architecture of a second sub-model according to some embodiments of the present disclosure;
[0014] Figure 5 A schematic diagram showing an example architecture of a third sub-model according to some embodiments of the present disclosure;
[0015] Fig. 6A A schematic diagram showing an example architecture of a sub-model training phase of a protein model according to some embodiments of the present disclosure;
[0016] Figure 6BA schematic diagram showing an example architecture of a joint training phase of a protein model according to some embodiments of the present disclosure;
[0017] Figure 7 A flowchart showing an example process of protein structure prediction according to some embodiments of the present disclosure;
[0018] Figure 8 A schematic structural block diagram showing an example apparatus for protein structure prediction according to some embodiments of the present disclosure; and
[0019] Fig. 9 A block diagram of an electronic device capable of implementing various embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0020] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.
[0021] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below.
[0022] Herein, unless explicitly stated, executing a step “in response to A” does not mean executing the step immediately after “A” but may include one or more intermediate steps.
[0023] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions.
[0024] It is understandable that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, scope of use, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0025] For example, in response to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested to be performed will require obtaining and using the user's personal information, so that the user can independently choose whether to provide personal information to software or hardware such as electronic devices, applications, servers or storage media that execute operations of the technical solution of the present disclosure based on the prompt message.
[0026] As an optional but non-limiting implementation, in response to receiving an active request from the user, the prompt information is sent to the user in a manner such as a pop-up window, in which the prompt information can be presented in text form. In addition, the pop-up window can also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0027] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that meet relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0028] As used herein, the term "model" can learn the association between the corresponding input and output from the training data, so that after the training is completed, the corresponding output can be generated for a given input. The generation of the model can be based on machine learning technology. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multi-layer processing units. A neural network model is an example of a model based on deep learning. In this article, "model" may also be referred to as "machine learning model", "learning model", "machine learning network" or "learning network", and these terms are used interchangeably in this article.
[0029] A "neural network" is a machine learning network based on deep learning. A neural network is capable of processing inputs and providing corresponding outputs, and typically includes an input layer and an output layer and one or more hidden layers between the input layer and the output layer. Neural networks used in deep learning applications typically include many hidden layers, thereby increasing the depth of the network. The layers of a neural network are connected in sequence so that the output of the previous layer is provided as input to the next layer, where the input layer receives the input of the neural network and the output of the output layer serves as the final output of the neural network. Each layer of a neural network includes one or more nodes (also called processing nodes or neurons), each of which processes input from the previous layer.
[0030] Generally, machine learning can be roughly divided into three stages, namely the training stage, the testing stage, and the application stage (also called the inference stage). In the training stage, a given model can be trained using a large amount of training data, and the parameter values are continuously updated iteratively until the model can obtain consistent inferences that meet the expected goals from the training data. Through training, the model can be considered to be able to learn the association from input to output (also called the mapping of input to output) from the training data. The parameter values of the trained model are determined. In the testing stage, the test input is applied to the trained model to test whether the model can provide the correct output, thereby determining the performance of the model. In the application stage, the model can be used to process the actual input based on the parameter values obtained from the training to determine the corresponding output.
[0031] As used herein, the term "weight tensor" may be a tensor having any suitable number of dimensions. For example, a weight tensor may be a weight matrix. Hereinafter, the weight matrix will be mainly used as an example of a weight tensor for description.
[0032] As mentioned above, proteins are biological molecules or macromolecules composed of long chains of amino acid residues. Proteins perform many important life activities in organisms, and the function of proteins is mainly determined by their three-dimensional (3D) structure. Knowing the structure of proteins is very important in the fields of medicine and biotechnology. For example, if a certain protein plays a key role in a certain disease, drug molecules can be designed based on the structure of the protein to treat the disease. In traditional technology, the prediction of protein structure mainly relies on experimental methods, but the prediction of protein structure using implementation methods is not only costly, but also time-consuming.
[0033] With the development of artificial intelligence technology, the use of machine learning models to predict protein structure is becoming more and more common. Conventional protein structure prediction models usually show relatively good results in predicting the amino acid sequence and backbone structure of proteins, but the accuracy of protein side chain structure, especially the prediction of atomic-level side chain structure, still needs to be improved. Moreover, conventional protein structure prediction usually has a relatively high accuracy in predicting the structure of single-chain proteins, but the accuracy of predicting the structure of multi-chain proteins needs to be improved, especially the accuracy of predicting the full atomic structure of multi-chain proteins.
[0034] In view of this, an embodiment of the present disclosure proposes an improved scheme for protein structure prediction. In this scheme, based on the input information associated with the target protein, the sequence information and backbone structure information of the target protein are determined using the first sub-model in the protein model. The sequence information indicates the amino acid sequence of at least one peptide chain of the target protein, and the backbone structure information indicates the backbone conformation of at least one peptide chain. Based on the sequence information and the backbone structure information, the side chain structure information of at least one peptide chain is determined using the second sub-model in the protein model, and the side chain structure information indicates the side chain conformation at each residue in at least one peptide chain. Then, based on the side chain structure information, the sequence information and the backbone structure information, the structure of the target protein is determined using the third sub-model in the protein model.
[0035] In an embodiment of the present disclosure, a protein model is designed to include a first sub-model, a second sub-model, and a third sub-model. The first sub-model can be used to predict the amino acid sequence and backbone structure at the residue level, and the second sub-model can be used to predict the side chain conformation of the target protein based on the amino acid sequence and backbone structure generated by the first sub-model. The third sub-model updates the amino acid sequence and backbone structure based on the side chain conformation, which can not only avoid conflicts, but also make the generated target protein more similar to the natural protein. Thus, it is beneficial to improve the accuracy of protein structure prediction.
[0036] Various example implementations of the solution are described in detail below in conjunction with the accompanying drawings.
[0037] Example Environment
[0038] Figure 1 1 is a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. Figure 1 In the environment 100 of the present invention, it is desirable to train and use such a model 130. In some examples, the model 130 can be configured to predict the structure of a protein, such as predicting the amino acid sequence, backbone structure, side chain structure, etc. of a protein. The model 130 can be a model of various suitable types, such as a flow matching model, a diffusion model, a transformer model, and the like.
[0039] like Figure 1 As shown, the environment 100 includes a training sample set 110 , a model training system 120 , and a model application system 140 . Figure 1 The upper part shows the process of the model training phase, and the lower part shows the process of the model application phase. Before training, the parameter values of the model 130 may have initial values, or may have pre-trained parameter values obtained through a pre-training process.
[0040] In the model training stage, the model 130 can be trained based on a training sample set 110 including a plurality of training samples 111 and using a model training system 120. Here, each training sample 111 can involve a bigram format. For example, for a protein structure prediction task, the training sample 111 can include a model input 112 and a model output 113. The model input 112 can include a partial backbone structure, a partial amino acid sequence, a partial side chain structure, or other protein-related reference information of a protein. The model output 113 can include the structure of a protein corresponding to the model input 112, such as the backbone structure, amino acid sequence, side chain structure, etc. of a protein. The training sample 111 including the model input 112 and the model output 113 can be used to train the model 130. For example, a training process can be iteratively performed using a large number of training samples. The model 130 can be trained via forward propagation and back propagation, and the parameter values of the model 130 can be updated and adjusted during the training process.
[0041] After the training is completed, the model 130' can be obtained. At this time, the parameter values of the model 130' have been updated, and based on the updated parameter values, the model 130' can be used to implement the protein structure prediction task in the model application stage. In the model application stage, the model 130' (at this time, the model 130' has the trained parameter values) can be used to perform the corresponding task through the model application system 140. For example, a model input 141 including at least one word unit can be received, and a corresponding model output 142 can be output.
[0042] exist Figure 1 In the present invention, the model training system 120 and the model application system 140 may include any computing system with computing capabilities, such as various computing devices / systems, terminal devices, servers, etc. The terminal device may involve any type of mobile terminal, fixed terminal or portable terminal, including a mobile phone, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. Servers include but are not limited to mainframes, edge computing nodes, computing devices in cloud environments, and the like.
[0043] It should be understood that the structure and functionality of environment 100 are described for exemplary purposes only and does not imply any limitation on the scope of the present disclosure.
[0044] Example Scenario
[0045] Some example embodiments of the present disclosure will be described below with continued reference to the accompanying drawings. Figure 2A schematic diagram of an example architecture 200 for protein structure prediction according to the following embodiment of the present disclosure is shown. In the following, for the convenience of discussion, the example architecture 200 is described from the perspective of the model application system 140 by taking the model application link as an example, but this is only exemplary.
[0046] In some embodiments of the present disclosure, the model application system 140 can determine the sequence information and backbone structure information 204 of the target protein based on the input information 202 associated with the target protein and using the first submodel 212 in the protein model 210. The sequence information indicates the amino acid sequence of at least one peptide chain of the target protein, and the backbone structure information indicates the backbone conformation of at least one peptide chain. In some embodiments, if the target protein is a single-chain protein, the sequence information may include the amino acid sequence of one peptide chain, and the backbone structure information may indicate the backbone conformation of one peptide chain. If the target protein is a multi-chain protein, the sequence information may indicate the amino acid sequence of multiple peptide chains, and the backbone structure information may indicate the backbone conformation of multiple peptide chains.
[0047] In some examples, for a single protein chain of length N The amino acid sequence of a peptide chain can be expressed as S = [S 1 , S 2 , …, S N ],in represents a set of standard amino acids. The backbone conformation of the peptide chain can be expressed as T = [T 1 , T 2 , …, T N ], where T i ∈SE(3), contains a rotation matrix R i ∈SO(3) and a translation vector
[0048] In other examples, targeting multi-chain proteins Sequence information can indicate the amino acid sequence of multiple peptide chains. The amino acid sequence can be expressed as S j =[S j 1 , S j 2 , …, S j N ], where j represents the peptide chain number (Chain ID). The backbone structure information can indicate the backbone conformation of multiple peptide chains, and the backbone conformation can be expressed as T j =[T j 1 , T j 2 , …, T j N ], where Tj i ∈SE(3), contains a rotation matrix E j i ∈SO(3) and a translation vector
[0049] In practical applications, the target protein structure prediction task may include multiple types of tasks, and the specific content included in the input information 202 is related to the task type of the structure prediction task. In some embodiments, the input information 202 includes at least one of the following: at least partial sequence information of the target protein, at least partial backbone structure information of the target protein, an indication of each peptide chain in at least one peptide chain, or an indication of a binding region between peptide chains in at least one peptide chain.
[0050] In some examples, the structure prediction task may indicate predicting the backbone structure information 204 and the side chain structure information 206 of the target protein. The input information 202 may include at least part of the sequence information of the target protein. Of course, the input information 202 may also include part of the backbone structure information. In some examples, the structure prediction task may indicate predicting the sequence information of the target protein. The input information may include at least part of the backbone structure information of the target protein. Of course, the input information 202 may also include part of the sequence information.
[0051] In some examples, the structure prediction task may indicate predicting the structure of a multi-chain target protein. The input information 202 may include an indication of each peptide chain in the multiple peptide chains, and an indication of the binding region between the peptide chains in the multiple peptide chains. As an example, the input information 202 may include a peptide chain number and a label indicating whether the residue in each peptide chain is located in the binding region between the peptide chains, for example, if a residue is located in the binding region, the label corresponding to the residue may be 1. If a residue is not in the binding region, the label corresponding to the residue may be 0. It is understandable that for the structure prediction task of a multi-chain protein, the input information may also include at least part of the backbone structure information or at least part of the sequence information of the target protein.
[0052] In some examples, the structure prediction task may indicate tasks related to protein structure prediction, such as antibody protein structure prediction, peptide chain structure prediction, or zero-sample prediction. For example, for antibody protein structure prediction, the input information may include a partial amino acid sequence of the antibody protein, information related to the antibody type, information related to homologous proteins, etc. For another example, for a zero-sample prediction task, the input information may include the category of the target protein (e.g., enzyme protein, transport protein, structural protein, etc.), a functional description of the target protein, etc.
[0053] It should be understood that for different structure prediction tasks, the input information may be the same or different. Moreover, the input information is not limited to the above information, but may also include any other appropriate information related to the target protein, which is not limited in the embodiments of the present disclosure.
[0054] In some embodiments, the model application system 140 may perform encoding on the input information to obtain an input feature representation. The model application system 140 may obtain a first model output based on the input feature representation using the first sub-model. The first model output may include a probability representation and backbone structure information, and the probability representation may include a probability matrix. For example, for a protein of length N, the probability representation may include an N×20 probability matrix, and each element in the probability matrix indicates the probability of belonging to a corresponding standard protein. Thereafter, the model application system 140 may determine the amino acid sequence of at least one peptide chain of the target protein based on the probability representation.
[0055] The following combination Figure 3 and an example architecture of the first sub-model 212, and an exemplary explanation of the reasoning process of the first sub-model 212 is given. However, it should not be understood that the first sub-model 212 is limited to using Figure 3 In the example architecture shown, in actual application, the first sub-model 212 can adopt any appropriate model structure, which is not limited in the embodiments of the present disclosure.
[0056] Figure 3 FIG. 3 is a schematic diagram showing an example architecture 300 of the first sub-model 212 according to some embodiments of the present disclosure. Figure 3 As shown, the input information may include at least one of sequence information 302, peptide chain number 304, indication of the binding region between peptide chains in at least one peptide chain 303, and backbone structure information 304. Figure 3 The above information is shown in , but in actual application, the input information may lack one or more of the above information, for example, the sequence information 302, the backbone structure information 304, etc. may be missing.
[0057] The first submodel 212 may include a trained protein language model 312, an encoder 314, and an encoder 316. The first submodel 212 may encode the sequence information 302 using the protein language model 312, and the first submodel 212 may also encode the sequence information 302, the peptide chain number 304, and the indication 306 of the binding region using the encoder 314. Then, based on the encoding result of the protein language model 312 and the encoding result of the encoder 322, an input feature representation 322 (also referred to as a residue representation) may be obtained. The first submodel 212 may also encode the peptide chain code 304, the indication 306 of the binding region, and the backbone structure 308 using the encoder 316 to obtain an input feature representation 324 (also referred to as a residue pair representation). Encoding the sequence information 302 using the protein language model 312 can enhance the representation capability of the sequence information 302, especially for the representation capability of the sequence information of proteins with relatively complex structures.
[0058] The first sub-model 212 may include a plurality of first model units 318, such as a first model unit 318-1, a first model unit 318-2, ..., a first model unit 318-K, where K is a positive integer. The model application system 140 may provide an input feature representation 322 and an input feature representation 324 to the first model unit 318-1, and the first model unit 318-1 may output a pair of intermediate feature representations. The first model unit 318-2 may generate another pair of intermediate feature representations for the next model unit based on a pair of intermediate feature representations generated by the previous model unit (i.e., the first model unit 318-1), thereby inferring. The last model unit (i.e., the first model unit 318-K) among the plurality of first model units 318 generates an output feature representation 326 based on a pair of intermediate feature representations generated by the previous model unit. The first sub-model 212 may generate a probability representation 332 based on the output feature representation 326, for example, through a sequence prediction network. The model application system 140 can also generate sequence information 336 based on the probability representation 332, such as a string of letters or numbers that can indicate an amino acid sequence. In some examples, the probability representation 332 can also be used as the sequence information 336 generated by the first sub-model 212.
[0059] During the reasoning process, each first model unit 318 will also generate trunk structure correction information, and use the trunk structure correction information in turn to correct the trunk structure information 308. For example, the trunk structure correction information may include a rotation matrix and a translation vector. The first sub-model 212 can use the rotation matrix and translation vector output by the first model unit 318-1 to correct the trunk structure information and obtain the corrected trunk structure information. Afterwards, the first sub-model 212 can also use the rotation matrix and translation vector output by the first model unit 318-2 to update the corrected trunk structure information and obtain another corrected trunk structure information. And so on, the first sub-model 212 can finally generate the trunk structure information 334. It can be understood that, Figure 3 The input information in includes the backbone structure information 308. In actual application, the input information may not include the backbone structure information, or may only include part of the backbone structure information. In this case, the first sub-model 212 may also gradually generate the backbone structure information 334 of the target protein through the above process.
[0060] In some embodiments of the present disclosure, Figure 2 As shown, the model application system 140 determines the side chain structure information 206 of at least one peptide chain based on the sequence information and the backbone structure information 204 using the second submodel 214 in the protein model 210. The side chain structure information 206 indicates the side chain conformation at each residue in the at least one peptide chain.
[0061] In some embodiments, the side chain conformation includes the torsion angle of each bond in the side chain group of the corresponding residue. Here, the bond of the side chain group may refer to a rotatable bond in the side chain group. As an example, for a protein of length N, the side chain structure information 206 may include a set of torsion angles of the protein. The set of torsion angles may be expressed as in Indicates the torsion angle of the rotatable bond of the side chain group of residue i. Each amino acid side chain group has up to four rotatable bonds, and these bonds maintain a basically consistent atomic structure. Therefore, combining the amino acid type (indicated by the amino acid sequence) and the torsion angle of the side chain group can not only provide a comprehensive representation of the side chain conformation, but also provide rich structural information. Moreover, the amount of data is small, which is conducive to maintaining calculation efficiency. Of course, the above-mentioned representation method of side chain conformation is only exemplary, and any other appropriate method can be selected according to actual needs to represent the side chain conformation of the protein.
[0062] In some embodiments, the model application system 140 may encode the sequence information and the backbone structure information to obtain at least one input feature representation. The model application system 140 may generate an output feature representation based on the at least one input feature representation using the second submodel 214. Then, the side chain structure information of the target protein is determined based on the output feature representation.
[0063] The following combination Figure 4 The second sub-model 214 is an example of an architecture, and the reasoning process of the second sub-model 214 is exemplarily described. However, it should not be understood that the second sub-model 214 is limited to using Figure 4 The example architecture shown, in actual application, the second sub-model 214 can adopt any appropriate model architecture, and the embodiments of the present disclosure are not limited to this.
[0064] Figure 4 FIG. 4 is a schematic diagram showing an example architecture 400 of the second sub-model 214 according to some embodiments of the present disclosure. Figure 4 As shown, the second sub-model 214 may include a trained protein language model 412, an encoder 414, and an encoder 416. The second sub-model 214 may use the protein language model 412 to encode the sequence information 336, and the second sub-model 214 may also use the encoder 414 to encode the sequence information 336 and the backbone structure information 334. Afterwards, an input feature representation 422 (which may be referred to as a residue feature representation) may be generated based on the encoding result of the protein language model 412 and the encoding result of the encoder 414. The second sub-model 214 may also use the encoder 416 to encode the sequence information 336 and the backbone structure information 334 to generate an input feature representation 424 (which may be referred to as a residue pair feature representation).
[0065] and Figure 3Similar to the first sub-model 212 shown, the second sub-model 214 may include a plurality of second model units 418, such as a second model unit 318-1, a second model unit 318-2, ..., a second model unit 318-M, where M is a positive integer. The first model unit (i.e., the second model unit 318-1) among the plurality of second model units 418 may generate a pair of intermediate feature representations based on the input feature representation 422 and the input feature representation 424. The second model unit 318-2 may generate a pair of intermediate feature representations for the next model unit based on the pair of intermediate feature representations output by the previous model unit (i.e., the second model unit 318-1). By analogy, the last model unit (i.e., the second model unit 318-M) among the plurality of second model units 418 may generate an output feature representation 426 based on a pair of intermediate feature representations generated by the previous model unit. The second sub-model 214 may generate side chain structure information 428, such as a set of side chain torsion angles of the target protein, based on the output feature representation 426 using a side chain structure prediction network.
[0066] In some embodiments of the present disclosure, continue to combine Figure 2 As shown, the model application system 140 determines the structure 208 of the target protein based on the side chain structure information 206, the sequence information and the backbone structure information 204 using the third sub-model 216 in the protein model 210. In some embodiments, the third sub-model 216 can be trained to modify the sequence information and the backbone structure information 204 generated by the first sub-model 212 to obtain the corrected sequence information and the corrected backbone structure information. In this case, the structure 208 of the target protein can include the side chain structure information 206 generated by the second sub-model 214, the corrected sequence information and the corrected backbone structure information.
[0067] In some examples, the structure 208 of the target protein may include an amino acid sequence of at least one peptide chain (e.g., at least one set of strings), a rotation matrix and a translation vector of the backbone structure of at least one peptide chain, and at least one set of torsion angles corresponding to at least one peptide chain. In some examples, the structure 208 of the target protein may also include an indication of at least one peptide chain of the target protein, such as a peptide chain number. In other examples, the structure 208 of the target protein may also include an indication of a binding region between peptide chains, such as a label indicating whether a residue in a peptide chain is located in a binding region.
[0068] In some embodiments, the model application system 140 can use the third submodel 216 to determine the sequence correction information corresponding to the sequence information and the backbone structure correction information corresponding to the backbone structure information based on the side chain structure information, the sequence information and the backbone structure correction information. Afterwards, the sequence information and the backbone structure information are corrected based on the sequence correction information and the backbone structure correction information, respectively, to determine the structure of the target protein. In some examples, the model application system 140 can also use the sequence correction information to correct the corresponding probability representation. For example, the third submodel 216 can generate probability correction information (e.g., a probability matrix), and use the probability correction information to correct the probability representation generated by the first submodel 212 to obtain a corrected probability representation. Afterwards, the corrected sequence information of the target protein can be generated based on the corrected probability representation.
[0069] The following combination Figure 5 and an example architecture of the third sub-model 216, and an exemplary explanation of the reasoning process of the third sub-model 216 is given. However, it should not be understood that the third sub-model 216 is limited to using Figure 5 In the example architecture shown, in actual application, the third sub-model 216 can adopt any appropriate model structure, which is not limited in the embodiments of the present disclosure.
[0070] Figure 5 FIG. 5 is a schematic diagram showing an example architecture 500 of the third sub-model 216 according to some embodiments of the present disclosure. Figure 5 As shown, the third sub-model 216 may include a trained protein language model 512, an encoder 514, and an encoder 516. The input of the third sub-model 216 may include the probability representation 332, the sequence information 336, the side chain structure information 428, the peptide chain number 304, the indication of the binding region 306, and the backbone structure information 334. The third sub-model 216 may encode the sequence information 336 using the protein language model 512, and the third sub-model 216 may also encode the sequence information 336, the side chain structure information 428, the peptide chain number 304, and the indication of the binding region 306 using the encoder 514. Afterwards, the input feature representation 522 may be generated based on the encoding result of the protein language model 512 and the encoding result of the encoder 514. The third sub-model 216 may also encode the side chain structure 428, the peptide chain number 304, the indication of the binding region 306, and the backbone structure information 334 using the encoder 516 to generate the input feature representation 524.
[0071] The third sub-model 216 may include a plurality of third model units 518, such as a third model unit 518-1, a third model unit 518-2, ..., a third model unit 518-L, where L is a positive integer. Here, the processing process of the input feature representation 522 and the input feature representation 524 by the plurality of third model units 518 is similar to the processing process of the plurality of first model units in the first sub-model 216, which will not be described here, and the introduction of the first sub-model 216 in the foregoing content may be referred to. Finally, the plurality of third model units 518 may generate an output feature representation 526, and the third sub-model 216 may generate a probability correction representation 528 based on the output feature representation 526 using, for example, a sequence correction network. Afterwards, the probability correction representation 528 may be used to correct the probability representation 332 generated by the second sub-model 214 to generate a corrected probability representation 532. The model application system 140 may determine the corrected sequence information based on the corrected probability representation 532. In addition, each third model unit 518 may also generate trunk structure correction information, and use the trunk structure correction information to correct the trunk structure information generated by the second sub-model 214 to obtain the corrected trunk structure information 534 .
[0072] The specific process of protein structure prediction of the embodiment of the present disclosure is introduced above. In the embodiment of the present disclosure, by using the protein model 210 including the first sub-model 212, the second sub-model 214 and the third sub-model 216, the generated target protein can be more similar to the natural protein, which is beneficial to improve the accuracy of protein structure prediction. The torsion angle of the side chain group can not only accurately indicate the side chain conformation, but also has a small amount of data, which is beneficial to improve the reasoning efficiency. By introducing the indication of the peptide chain (such as the peptide chain code) and the indication of the binding area between the peptide chains (such as the label indicating whether the residue is located in the binding area), the protein model 210 can be used not only for the structure prediction task of single-chain protein, but also for the structure prediction task of multi-chain protein.
[0073] The following will be combined Fig. 6A and Figure 6B Introducing the training process of the protein model. For ease of discussion, the example architecture 600A and the example architecture 600B will be described from the perspective of the model training system 120 hereinafter, but this is merely exemplary.
[0074] In some embodiments of the present disclosure, the training process of the protein model may include a sub-model training process and a joint training process. In the sub-model training process, the model training system 120 may train the first sub-model 212 and the second sub-model 214 separately. Fig. 6A The training process of the first sub-model 212 and the second sub-model 214 during the sub-model training process is introduced.
[0075] like Fig. 6A As shown, the model training system 120 can obtain a sample set 602. The sample set 602 includes a plurality of sample structures 604 corresponding to a plurality of sample proteins, each of which includes at least sample sequence information and sample backbone structure information of the corresponding sample protein. The sample sequence information may indicate the amino acid sequence of at least one peptide chain of the sample protein, and the sample backbone structure information may indicate the backbone conformation of at least one peptide chain of the sample protein. In some embodiments, each sample structure 604 may also include sample side chain structure information of the corresponding sample protein. The sample side chain structure information may indicate the side chain conformation of each residue in at least one peptide chain of the sample protein. In some embodiments, each sample structure 604 may also include an indication of each peptide chain in the sample protein, or an indication of the binding region between peptide chains.
[0076] In some embodiments of the present disclosure, the model training system 120 obtains multiple sample inputs corresponding to multiple sample proteins by adding noise to at least a portion of the sample structures in the sample set. The main task of the first submodel 212 is to predict the amino acid sequence and backbone structure of the protein. In some embodiments, the tasks of the first submodel 212 mainly include unconditional generation tasks, conditional generation tasks, folding tasks, and reverse folding tasks. The unconditional generation task refers to generating the structure of the protein from scratch by the first submodel 212 without providing sequence information and backbone structure information to the first submodel 212. The conditional generation task refers to generating a protein structure that meets the constraints by the first submodel 212 under given constraints. The folding task refers to instructing the first submodel 212 to predict the backbone structure of the protein given the amino acid sequence (i.e., sequence information). The reverse folding task refers to instructing the first submodel 212 to predict the amino acid sequence of the protein given the backbone structure of the protein.
[0077] The purpose of adding noise to the sample structure here is to simulate the above-mentioned task of the first sub-model 212, and train the first sub-model 212 to reconstruct the amino acid sequence or the backbone structure from the noisy input, or to reconstruct the amino acid sequence and the backbone structure. The following will exemplify the noise adding method in combination with the different tasks of the first sub-model 212. However, it should be understood that the first sub-model 212 is not limited to performing the above-mentioned tasks, and correspondingly is not limited to adding noise to the sample structure in the manner shown below, and any other appropriate manner can be selected to add noise to the sample structure according to the actual task requirements. The embodiments of the present disclosure are not limited to this.
[0078] In some embodiments, for a first proportion of sample proteins in the sample set, the model training system 120 can add noise to the sample sequence information and the sample backbone structure information in the corresponding sample structure, respectively, to obtain a sample input corresponding to the corresponding sample protein. In this way, the sample sequence information and the sample backbone structure information in the sample input are both noisy, which can simulate an unconditional generation task. Such sample input can train the first sub-model to predict the structure of the protein in input information containing incomplete sequence information and backbone structure information, or even input information without sequence information and backbone structure information.
[0079] In some examples, the first ratio may be any appropriate ratio, for example, the first ratio may be 50%, that is, noise may be added to the sample sequence information and sample trunk structure information of 50% of the sample proteins in the sample set to form sample inputs corresponding to the 50% sample proteins. Of course, the above first ratio is only exemplary.
[0080] In some embodiments, the model training system 120 may select a sample protein from a sample set according to a first ratio. The model training system 120 may determine a first time step parameter indicating the noise intensity of the sample sequence information, and the model training system 120 may add noise to the sample sequence information of the selected sample protein based on the first time step parameter to obtain the sample sequence information after adding noise as part of the sample input. The model training system 120 may also determine a second time step parameter indicating the noise intensity of the sample backbone structure information, and the model training system 120 may add noise to the sample backbone structure information of the selected sample protein based on the second time step parameter to obtain the sample backbone structure information after adding noise as part of the sample input. The folding task and the reverse folding task of the protein mainly depend on the dependency between the amino acid sequence and the backbone structure. Using the first time step parameter and the second time step parameter, the noise adding process to the amino acid sequence and the backbone structure can be decoupled, so that the noise intensity between the sample sequence information and the sample backbone structure information is not completely aligned, which can avoid destroying the dependency between the amino acid sequence and the backbone structure to a certain extent.
[0081] In some examples, model training system 120 may be configured to: Within the value range, randomly determine the time step parameter t corresponding to the sample sequence information S and the time step parameter t corresponding to the sample backbone structure information T The model training system 120 can be based on the time step parameter t S To the sample sequence information S 1 (that is, the original sequence information without adding noise) Add noise to obtain the sample sequence information S with noise tSThe model training system 120 can also be based on the time step parameter t T To the sample backbone structure information T 1 (that is, the original backbone structure information without adding noise) adds noise, and the obtained sample backbone structure information with noise is expressed as T tT .
[0082] In some embodiments, for a second proportion (e.g., 50%) of sample proteins in the sample set, the model training system 120 may add noise to one of the sample sequence information or sample trunk structures in the corresponding sample structure to obtain a sample input corresponding to the corresponding sample protein.
[0083] In some examples, for the sample proteins of the third proportion (e.g., 25%) in the sample set, the model training system 120 can add noise to the sample sequence information in the corresponding sample structure. In this way, the reverse folding task can be simulated, and the ability of the first sub-model 212 to predict the amino acid sequence based on the backbone structure can be trained by using such sample input.
[0084] In some examples, for the fourth proportion of sample proteins in the sample set, the model training system 120 can add noise to the sample backbone structure information in the corresponding sample structure. In this way, the folding task can be simulated, and the first sub-model 212 can be trained to predict the backbone structure of the protein based on the amino acid sequence using such sample input.
[0085] In some embodiments, the plurality of sample proteins include a plurality of multi-chain proteins, each multi-chain protein includes a plurality of peptide chains, each item of sample sequence information includes multiple items of peptide chain sequence information corresponding to the multiple peptide chains, and each item of sample backbone structure information includes multiple items of peptide chain backbone structure information corresponding to the multiple peptide chains. For the multi-chain proteins among the plurality of multi-chain proteins, the model training system 120 can add noise to the peptide chain sequence information and / or peptide chain backbone structure information corresponding to some of the peptide chains among the plurality of peptide chains to obtain the corresponding sample input. In this way, the noise intensity between the plurality of peptide chains can be made incompletely aligned, which is beneficial to retaining the correlation between the plurality of peptide chains, and further beneficial to improving the performance of the first sub-model 212 for the structural prediction task of the multi-chain protein.
[0086] In some embodiments of the present disclosure, Fig. 6A As shown, for a first sample input 606 among multiple sample inputs, the model training system 120 determines a prediction result 608 of the structure of a first sample protein corresponding to the first sample input based on the first sample input 606 and using the first sub-model 212. The prediction result 608 includes first predicted sequence information and first predicted backbone structure information of the first sample protein.
[0087] In some embodiments of the present disclosure, the model training system 120 updates the parameters of the first sub-model 212 based on the first difference between the prediction result 608 and the sample structure 604 of the first sample protein. In some embodiments, the model training system 120 may update the parameters of the first sub-model 212 based on the flow matching loss function as shown below:
[0088]
[0089] in represents the probability representation of residue i, represents the sample sequence information of residue i, CrossEntropy(·) represents the cross entropy loss, represents the noisy backbone structure corresponding to residue i, represents the predicted backbone structure corresponding to residue i, MSE(·) represents mean square error, and VF(·) represents vector field.
[0090] In some embodiments, Figure 3 As shown, the first sub-model 212 includes a plurality of first model units 318 arranged in sequence. The model training system 120 may update the parameters of the first sub-model based on the first difference and the second difference between the intermediate prediction results of two adjacent first model units in the plurality of first model units. In some examples, a consistency loss may be added to the flow matching loss function shown in formula (1): For example, the model training system 120 may update the parameters of the first sub-model 212 based on the loss function shown below:
[0091]
[0092] In some embodiments of the present disclosure, Fig. 6A As shown, the model training system 120 can obtain a sample set 602. The sample set 602 includes a plurality of sample structures 604 corresponding to a plurality of sample proteins, each sample structure 604 including sample sequence information, sample backbone structure information, and sample side chain structure information of the corresponding sample protein. In box 612, the model training system 120 can select sample sequence information and sample backbone structure information of a second sample protein from the sample set 602. The model training system 120 can determine the predicted side chain structure information 612 of at least one peptide chain of the second sample protein using the second sub-model 214 based on the sample sequence information and sample backbone structure information of the second sample protein. Afterwards, the model training system 120 can update the parameters of the second sub-model based on the third difference between the sample side chain structure information and the predicted side chain structure information 612 of the second sample protein.
[0093] In the joint training process of the protein model 210, the model training system 120 may jointly train the first sub-model 212, the second sub-model 214, and the third sub-model 216. Figure 6B The joint training process is introduced.
[0094] In some embodiments of the present disclosure, the model training system 120 may obtain a sample set, which includes a sample structure 622 corresponding to a plurality of sample proteins. The sample structure 622 includes sample sequence information, sample backbone structure information, and sample side chain structure information. The model training system 120 may obtain a plurality of sample inputs corresponding to a plurality of sample proteins by adding noise to at least one of the sample sequence information or sample backbone structure information in at least a portion of the sample structure 622 in the sample set. For the noise adding process, please refer to the introduction of the aforementioned sub-model training process, which will not be repeated here.
[0095] In some embodiments of the present disclosure, the model training system 120 may select a second sample input 624 from a plurality of sample inputs. The model training system may provide the second sample input to the first sub-model 212. At block 626, the model training system may determine second predicted sequence information and second predicted backbone structure information of a third sample protein corresponding to the second sample input.
[0096] In some embodiments of the present disclosure, the model training system 120 may determine the second predicted side chain structure information 630 of the third sample protein using the second sub-model 214 based on the second predicted sequence information and the second predicted backbone structure information. The model training system 120 may determine the predicted structure 632 of the third sample protein using the third sub-model 216 based on the second predicted side chain structure information, the second predicted sequence information, and the second predicted backbone structure information. Thereafter, the model training system 120 may update the parameters of at least one of the first sub-model 212, the second sub-model 214, or the third sub-model 216 based on the fourth difference between the predicted structure 632 and the sample structure 622 of the third sample protein.
[0097] In some embodiments, the model training system 120 may update the parameters of one of the first sub-model 212, the second sub-model 214, or the third sub-model 216 in a predetermined order based on the fourth difference to train the corresponding sub-model. By training the three sub-models 216 in an iterative manner, the three sub-models may be updated in a targeted manner. In some examples, in view of the fact that the first sub-model 212 and the second sub-model 214 have been trained in the sub-model training phase, in the joint training process, the first sub-model 212 and the second sub-model 214 may update the parameters of the first number of times in each iteration cycle, and the third sub-model 216 may update the parameters of the second number of times. For example, each iteration cycle may include twelve parameter update operations, and the model training system may perform two parameter update operations on the first sub-model 212, two parameter update operations on the second sub-model 214, and eight parameter update operations on the third sub-model 216. This is conducive to improving training efficiency.
[0098] In some embodiments, Figure 6B As shown, the model training system 120 can determine a third time step parameter indicating the noise intensity in the second sample input 624. In box 628, the model training system 120 can determine whether the third time parameter exceeds a predetermined parameter threshold. If the third time step parameter exceeds the predetermined parameter threshold, the model training system 120 can determine the second predicted side chain structure information of the third sample protein using the second sub-model 214 based on the second predicted sequence information and the second predicted backbone structure information. If the third time step parameter does not exceed the predetermined parameter threshold, the model training system can update the parameters of the first sub-model 212 based on the second predicted sequence information and the second predicted backbone structure information. However, the second predicted sequence information and the second predicted backbone structure information are not provided to the second sub-model 214 and the third sub-model 216. In the case of a high noise level, the second predicted sequence information and the second predicted backbone structure information are significantly different from the sequence information and backbone structure information of the real protein, which is insufficient to support the second sub-model 214 to perform high-quality side chain structure prediction, resulting in the optimization operation of the third sub-model 216 having no practical significance. Thus, unnecessary system resource consumption can be reduced and training efficiency can be improved.
[0099] In some embodiments, in the joint training stage of the protein model 210, the second sub-model 214 is also trained in the following manner: for a fourth sample protein among the multiple sample proteins, the model training system 120 determines the third predicted side chain structure information of the fourth sample protein using the second sub-model 214 based on the sample sequence information and sample backbone structure information of the fourth sample protein. Afterwards, the model training system 120 updates the parameters of the second sub-model based on the fifth difference between the sample side chain structure information of the fourth sample protein and the third predicted side chain structure information. That is, in the joint training stage, a part of the input of the second sub-model 214 is the second predicted sequence information and the second predicted backbone structure information generated by the first sub-model 212, and another part of the input is the sample sequence information and sample backbone structure information selected from the sample set. In this way, it is conducive to maintaining the ability of the second sub-model 214 to make high-quality predictions of side chain conformations based on the real amino acid sequence and backbone structure.
[0100] In this way, in the embodiments of the present disclosure, a two-stage training process is adopted. In the sub-model training stage, the first sub-model 212 and the second sub-model 214 are trained separately, so that the first sub-model 212 can quickly and high-quality grasp the amino acid sequence and backbone structure prediction capabilities, and the second sub-model 214 can quickly grasp the high-quality side chain structure prediction capabilities. In the joint training stage, by jointly training the first sub-model 212, the second sub-model 214 and the third sub-model 216, the overall training effect of the protein model 210 can be ensured.
[0101] Example procedures, devices and equipment
[0102] Figure 7 A flowchart of an example process 700 for protein structure prediction according to some embodiments of the present disclosure is shown. The process 700 may be implemented in the model training system 120 or the model application system 140.
[0103] In box 710, the model application system 140 determines the sequence information and backbone structure information of the target protein based on the input information associated with the target protein and using the first sub-model in the protein model, wherein the sequence information indicates the amino acid sequence of at least one peptide chain of the target protein, and the backbone structure information indicates the backbone conformation of at least one peptide chain.
[0104] In block 720 , the model application system 140 determines the side chain structure information of at least one peptide chain based on the sequence information and the backbone structure information using the second sub-model in the protein model, wherein the side chain structure information indicates the side chain conformation at each residue in the at least one peptide chain.
[0105] In block 730 , the model application system 140 determines the structure of the target protein using the third sub-model in the protein model based on the side chain structure information, the sequence information, and the backbone structure information.
[0106] In some embodiments, determining the structure of the target protein includes: based on the side chain structure information, the sequence information and the backbone structure information, using a third sub-model, determining sequence correction information corresponding to the sequence information and backbone structure correction information corresponding to the backbone structure information; and correcting the sequence information and the backbone structure information based on the sequence correction information and the backbone structure correction information, respectively, to determine the structure of the target protein.
[0107] In some embodiments, the input information includes at least one of the following: at least partial sequence information of the target protein, at least partial backbone structure information of the target protein, an indication of each peptide chain in at least one peptide chain, or an indication of the binding region between peptide chains in at least one peptide chain.
[0108] In some embodiments, side chain conformations include the torsion angles of various bonds in the side chain groups of the corresponding residues.
[0109] In some embodiments, in the sub-model training stage of the protein model, the first sub-model is trained in the following manner: a sample set is obtained, the sample set includes a plurality of sample structures corresponding to a plurality of sample proteins, each sample structure includes at least sample sequence information and sample backbone structure information of the corresponding sample protein; a plurality of sample inputs corresponding to the plurality of sample proteins are obtained by adding noise to at least a portion of the sample structures in the sample set; for a first sample input among the plurality of sample inputs, based on the first sample input, using the first sub-model, a prediction result of the structure of the first sample protein corresponding to the first sample input is determined, the prediction result including first predicted sequence information and first predicted backbone structure information of the first sample protein; and parameters of the first sub-model are updated based on a first difference between the prediction result and the sample structure of the first sample protein.
[0110] In some embodiments, obtaining multiple sample inputs corresponding to multiple sample proteins respectively includes at least one of the following: for a first proportion of sample proteins in the sample set, adding noise to the sample sequence information and the sample trunk structure information in the corresponding sample structure respectively to obtain the sample input corresponding to the corresponding sample protein, or for a second proportion of sample proteins in the sample set, adding noise to one of the sample sequence information or the sample trunk structure in the corresponding sample structure to obtain the sample input corresponding to the corresponding sample protein.
[0111] In some embodiments, for a first proportion of sample proteins in a sample set, adding noise to the sample sequence information and the sample trunk structure information in the corresponding sample structure respectively includes: selecting sample proteins from the sample set according to a first proportion; adding noise to the sample sequence information of the selected sample protein based on a first time step parameter indicating the noise intensity corresponding to the sample sequence information of the selected sample protein, so as to obtain the sample sequence information after the noise is added as a part of the sample input corresponding to the selected sample protein; and adding noise to the sample trunk structure information of the selected sample protein based on a second time step parameter indicating the noise intensity corresponding to the sample trunk structure information of the selected sample protein, so as to obtain the sample trunk structure information after the noise is added as a part of the sample input corresponding to the selected sample protein.
[0112] In some embodiments, a plurality of sample proteins include a plurality of multi-chain proteins, each multi-chain protein includes a plurality of peptide chains, each item of sample sequence information includes multiple items of peptide chain sequence information corresponding to the multiple peptide chains, each item of sample backbone structure information includes multiple items of peptide chain backbone structure information corresponding to the multiple peptide chains, and wherein by adding noise to at least a portion of the sample structures in the sample set, obtaining a plurality of sample inputs corresponding to the plurality of sample proteins respectively includes: for a multi-chain protein in the plurality of multi-chain proteins, adding noise to the peptide chain sequence information and peptide chain backbone structure information corresponding to a portion of the peptide chains in the plurality of peptide chains to obtain the corresponding sample input.
[0113] In some embodiments, the first sub-model includes a plurality of model units arranged in sequence, each of the plurality of model units is configured to generate an intermediate prediction result of the model unit for the structure of the first sample protein based on an output of a previous model unit or an input of the first sub-model, and wherein updating the parameters of the first sub-model includes: updating the parameters of the first sub-model based on a first difference and a second difference between the intermediate prediction results of two adjacent model units of the plurality of model units.
[0114] In some embodiments, in the sub-model training stage of the protein model, the second sub-model is trained in the following manner: a sample set is obtained, the sample set includes a plurality of sample structures corresponding to a plurality of sample proteins, each sample structure includes sample sequence information, sample backbone structure information and sample side chain structure information of the corresponding sample protein; for a second sample protein among the plurality of sample proteins, based on the sample sequence information and sample backbone structure information of the second sample protein, using the second sub-model, first predicted side chain structure information of at least one peptide chain of the second sample protein is determined; and based on a third difference between the sample side chain structure information of the second sample protein and the first predicted side chain structure information, parameters of the second sub-model are updated.
[0115] In some embodiments, in the joint training stage of the protein model, the first sub-model, the second sub-model and the third sub-model are jointly trained in the following manner: a sample set is obtained, the sample set includes sample structures corresponding to multiple sample proteins, and the sample structures include sample sequence information, sample backbone structure information and sample side chain structure information; a plurality of sample inputs corresponding to the multiple sample proteins are obtained by adding noise to at least one of the sample sequence information or the sample backbone structure information in at least a portion of the sample structures in the sample set; for a second sample input among the multiple sample inputs, based on the second sample input, using the first sub-model, second predicted sequence information and second predicted backbone structure information of a third sample protein corresponding to the second sample input are determined; based on the second predicted sequence information and the second predicted backbone structure information, using the second sub-model, second predicted side chain structure information of the third sample protein is determined; based on the second predicted side chain structure information and the second predicted backbone structure information, using the third sub-model, a predicted structure of the third sample protein is determined; and based on a fourth difference between the predicted structure and the sample structure of the third sample protein, a parameter of at least one of the first sub-model, the second sub-model or the third sub-model is updated.
[0116] In some embodiments, updating the parameters of at least one of the first sub-model, the second sub-model, or the third sub-model includes: based on the fourth difference, updating the parameters of one of the first sub-model, the second sub-model, or the third sub-model in a predetermined order to train the corresponding sub-model.
[0117] In some embodiments, determining second predicted side chain structure information of a third sample protein using a second sub-model based on the second predicted sequence information and the second predicted backbone structure information includes: determining a third time step parameter indicating noise intensity in the second sample input; and in response to the third time step parameter exceeding a predetermined parameter threshold, determining second predicted side chain structure information of the third sample protein using the second sub-model based on the second predicted sequence information and the second predicted backbone structure information.
[0118] In some embodiments, in the joint training stage of the protein model, the second sub-model is also trained in the following manner: for a fourth sample protein among the multiple sample proteins, based on the sample sequence information and the sample backbone structure information of the fourth sample protein, using the second sub-model, determining the third predicted side chain structure information of the fourth sample protein; and based on the fifth difference between the sample side chain structure information of the fourth sample protein and the third predicted side chain structure information, updating the parameters of the second sub-model.
[0119] The embodiments of the present disclosure also provide corresponding devices for implementing the above methods or processes. Figure 8A schematic structural block diagram of an example apparatus 800 for protein structure prediction according to certain embodiments of the present disclosure is shown. The apparatus 800 may be implemented as or included in the model training system 120 or the model application system 140. Each module / component in the apparatus 800 may be implemented by hardware, software, firmware, or any combination thereof.
[0120] like Figure 8 As shown, the device 800 includes: a first determination module 810, configured to determine the sequence information and backbone structure information of the target protein based on the input information associated with the target protein and using the first sub-model in the protein model, the sequence information indicates the amino acid sequence of at least one peptide chain of the target protein, and the backbone structure information indicates the backbone conformation of the at least one peptide chain; a second determination module 820, configured to determine the side chain structure information of at least one peptide chain based on the sequence information and the backbone structure information and using the second sub-model in the protein model, the side chain structure information indicates the side chain conformation at each residue in the at least one peptide chain; and an improvement module 830, configured to determine the structure of the target protein based on the side chain structure information, the sequence information and the backbone structure information and using the third sub-model in the protein model.
[0121] In some embodiments, the improvement module 830 is further configured to: determine the sequence correction information corresponding to the sequence information and the backbone structure correction information corresponding to the backbone structure information using a third sub-model based on the side chain structure information, the sequence information and the backbone structure information; and correct the sequence information and the backbone structure information based on the sequence correction information and the backbone structure correction information, respectively, to determine the structure of the target protein.
[0122] In some embodiments, the input information includes at least one of the following: at least partial sequence information of the target protein, at least partial backbone structure information of the target protein, an indication of each peptide chain in at least one peptide chain, or an indication of the binding region between peptide chains in at least one peptide chain.
[0123] In some embodiments, side chain conformations include the torsion angles of various bonds in the side chain groups of the corresponding residues.
[0124] In some embodiments, the device 800 also includes: a first training module, configured to train the first sub-model in the sub-model training stage of the protein model in the following manner: obtaining a sample set, the sample set including a plurality of sample structures corresponding to a plurality of sample proteins, each sample structure including at least sample sequence information and sample backbone structure information of the corresponding sample protein; obtaining a plurality of sample inputs corresponding to the plurality of sample proteins by adding noise to at least a portion of the sample structures in the sample set; for a first sample input among the plurality of sample inputs, based on the first sample input, using the first sub-model, determining a prediction result of the structure of the first sample protein corresponding to the first sample input, the prediction result including first predicted sequence information and first predicted backbone structure information of the first sample protein; and updating the parameters of the first sub-model based on a first difference between the prediction result and the sample structure of the first sample protein.
[0125] In some embodiments, the first training module is further configured to: for a first proportion of sample proteins in the sample set, add noise to the sample sequence information and the sample trunk structure information in the corresponding sample structure, respectively, to obtain a sample input corresponding to the corresponding sample protein; or for a second proportion of sample proteins in the sample set, add noise to one of the sample sequence information or the sample trunk structure in the corresponding sample structure, to obtain a sample input corresponding to the corresponding sample protein.
[0126] In some embodiments, the first training module is further configured to: select a sample protein from a sample set according to a first ratio; add noise to the sample sequence information of the selected sample protein based on a first time step parameter indicating the noise intensity corresponding to the sample sequence information of the selected sample protein to obtain the sample sequence information after the noise is added as a part of the sample input corresponding to the selected sample protein; and add noise to the sample trunk structure information of the selected sample protein based on a second time step parameter indicating the noise intensity corresponding to the sample trunk structure information of the selected sample protein to obtain the sample trunk structure information after the noise is added as a part of the sample input corresponding to the selected sample protein.
[0127] In some embodiments, multiple sample proteins include multiple multi-chain proteins, each multi-chain protein includes multiple peptide chains, each sample sequence information includes multiple peptide chain sequence information corresponding to the multiple peptide chains, each sample trunk structure information includes multiple peptide chain trunk structure information corresponding to the multiple peptide chains, and the first training module is further configured to: for the multi-chain proteins among the multiple multi-chain proteins, add noise to the peptide chain sequence information and peptide chain trunk structure information corresponding to some peptide chains among the multiple peptide chains to obtain corresponding sample input.
[0128] In some embodiments, the first sub-model includes a plurality of model units arranged in sequence, each of the plurality of model units is configured to generate an intermediate prediction result of the model unit for the structure of the first sample protein based on the output of the previous model unit or the input of the first sub-model, and the first training module is further configured to update the parameters of the first sub-model based on the first difference and the second difference between the intermediate prediction results of two adjacent model units in the plurality of model units.
[0129] In some embodiments, the device 800 also includes: a second training module, configured to train the second sub-model in the sub-model training stage of the protein model in the following manner: obtaining a sample set, the sample set including a plurality of sample structures corresponding to a plurality of sample proteins, each sample structure including sample sequence information, sample backbone structure information and sample side chain structure information of the corresponding sample protein; for a second sample protein among the plurality of sample proteins, based on the sample sequence information and sample backbone structure information of the second sample protein, using the second sub-model, determining first predicted side chain structure information of at least one peptide chain of the second sample protein; and updating the parameters of the second sub-model based on a third difference between the sample side chain structure information of the second sample protein and the first predicted side chain structure information.
[0130] In some embodiments, the apparatus 800 further includes: a third training module configured to jointly train the first sub-model, the second sub-model, and the third sub-model in the joint training stage of the protein model in the following manner: obtaining a sample set, the sample set including sample structures corresponding to a plurality of sample proteins, the sample structures including sample sequence information, sample backbone structure information, and sample side chain structure information; obtaining a plurality of sample inputs corresponding to the plurality of sample proteins by adding noise to at least one of the sample sequence information or the sample backbone structure information in at least a portion of the sample structures in the sample set; for a second sample input among the plurality of sample inputs, determining, based on the second sample input, using the first sub-model, second predicted sequence information and second predicted backbone structure information of a third sample protein corresponding to the second sample input; determining, based on the second predicted sequence information and the second predicted backbone structure information, using the second sub-model, second predicted side chain structure information of the third sample protein; determining, based on the second predicted side chain structure information, the second predicted sequence information, and the second predicted backbone structure information, using the third sub-model, a predicted structure of the third sample protein; and updating a parameter of at least one of the first sub-model, the second sub-model, or the third sub-model based on a fourth difference between the predicted structure and the sample structure of the third sample protein.
[0131] In some embodiments, the third training module is further configured to: based on the fourth difference, update the parameters of one of the first sub-model, the second sub-model or the third sub-model in a predetermined order to train the corresponding sub-model.
[0132] In some embodiments, the third training module is further configured to: determine a third time step parameter indicating the noise intensity in the second sample input; and in response to the third time step parameter exceeding a predetermined parameter threshold, determine second predicted side chain structure information of the third sample protein based on the second predicted sequence information and the second predicted backbone structure information using the second sub-model.
[0133] In some embodiments, the third training module is further configured to: for a fourth sample protein among the multiple sample proteins, based on the sample sequence information and the sample backbone structure information of the fourth sample protein, use the second sub-model to determine the third predicted side chain structure information of the fourth sample protein; and update the parameters of the second sub-model based on the fifth difference between the sample side chain structure information of the fourth sample protein and the third predicted side chain structure information.
[0134] The units and / or modules included in the device 800 can be implemented in various ways, including software, hardware, firmware or any combination thereof. In some embodiments, one or more units and / or modules can be implemented using software and / or firmware, such as machine executable instructions stored on a storage medium. In addition to or as an alternative to machine executable instructions, some or all of the units and / or modules in the device 800 can be implemented at least in part by one or more hardware logic components. As an example and not limitation, exemplary types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0135] Fig. 9 900 is a block diagram of an electronic device in which one or more embodiments of the present disclosure may be implemented. It should be understood that Fig. 9 The electronic device 900 shown is merely exemplary and should not constitute any limitation on the functionality and scope of the embodiments described herein. Fig. 9 The electronic device 900 shown may include or be implemented as Figure 1 The model training system 120 or the model application system 140 in Figure 8 Device 800.
[0136] like Fig. 9As shown, the electronic device 900 is in the form of a general electronic device. The components of the electronic device 900 may include, but are not limited to, one or more processors 910, a memory 920, a storage device 930, one or more communication units 940, one or more input devices 950, and one or more output devices 960. The processor 910 may be an actual or virtual processor and is capable of performing various processes according to executable instructions stored in the memory 920. In a multi-processor system, multiple processors execute computer executable instructions in parallel to improve the parallel processing capability of the electronic device 900.
[0137] The electronic device 900 typically includes a plurality of computer storage media. Such media can be any accessible media that can be obtained by the electronic device 900, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 920 can be a volatile memory (e.g., a register, a cache, a random access memory (RAM)), a non-volatile memory (e.g., a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), flash memory) or some combination thereof. The storage device 930 can be a removable or non-removable medium, and can include a machine-readable medium, such as a flash drive, a disk, or any other medium, which can be used to store information and / or data and can be accessed within the electronic device 900.
[0138] The electronic device 900 may further include additional removable / non-removable, volatile / non-volatile storage media. Fig. 9 As shown in , a disk drive for reading or writing from a removable, non-volatile disk (e.g., a "floppy disk") and an optical drive for reading or writing from a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to the bus (not shown) by one or more data media interfaces. The memory 920 may include a computer program product 925 having one or more executable instruction modules that are configured to perform various methods or actions of various embodiments of the present disclosure.
[0139] The communication unit 940 implements communication with other electronic devices through a communication medium. Additionally, the functions of the components of the electronic device 900 can be implemented with a single computing cluster or multiple computing machines that can communicate through a communication connection. Therefore, the electronic device 900 can operate in a networked environment using a logical connection with one or more other servers, a network personal computer (PC), or another network node.
[0140] The input device 950 may be one or more input devices, such as a mouse, a keyboard, a tracking ball, etc. The output device 960 may be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 900 may also communicate with one or more external devices (not shown) through the communication unit 940 as needed, such as a storage device, a display device, etc., communicate with one or more devices that allow a user to interact with the electronic device 900, or communicate with any device that allows the electronic device 900 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).
[0141] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer-executable instruction product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.
[0142] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, devices, equipment, and computer-executable instruction products implemented according to the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of boxes in the flowchart and / or block diagram can be implemented by computer-readable executable instructions.
[0143] These computer executable instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device that implements the functions / actions specified in one or more boxes in the flowchart and / or block diagram is generated. These computer executable instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other equipment to work in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured product, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0144] Computer-executable instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0145] The flowcharts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer executable instruction products according to multiple implementations of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, an executable instruction or a part of an instruction, and a module, an executable instruction or a part of an instruction contains one or more executable instructions for realizing the specified logical function. In some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.
[0146] The above descriptions of various implementations of the present disclosure are exemplary, non-exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described implementations. The selection of terms used herein is intended to best explain the principles of the implementations, practical applications, or improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the various implementations disclosed herein.
Claims
1. A protein structure prediction method, comprising: Based on input information associated with the target protein, using a first submodel in the protein model, determining sequence information and backbone structure information of the target protein, wherein the sequence information indicates an amino acid sequence of at least one peptide chain of the target protein, and the backbone structure information indicates a backbone conformation of the at least one peptide chain; Based on the sequence information and the backbone structure information, using a second submodel in the protein model, determining side chain structure information of the at least one peptide chain, wherein the side chain structure information indicates a side chain conformation at each residue in the at least one peptide chain; as well as Based on the side chain structure information, the sequence information and the backbone structure information, the structure of the target protein is determined using a third submodel in the protein model.
2. The method according to claim 1, wherein determining the structure of the target protein comprises: Based on the side chain structure information, the sequence information and the backbone structure information, using the third sub-model, determining sequence correction information corresponding to the sequence information and backbone structure correction information corresponding to the backbone structure information; as well as The sequence information and the backbone structure information are corrected based on the sequence correction information and the backbone structure correction information, respectively, to determine the structure of the target protein.
3. The method according to claim 1, wherein the input information includes at least one of the following: at least a portion of the sequence information of the target protein, At least part of the backbone structure information of the target protein, an indication of each of the at least one peptide chain, or An indication of a binding region between peptide chains in the at least one peptide chain. The method according to claim 1 , wherein the side chain conformation comprises the torsion angle of each bond in the side chain groups of the corresponding residues.
5. The method according to claim 1, wherein in the sub-model training stage of the protein model, the first sub-model is trained in the following manner: Acquire a sample set, wherein the sample set includes a plurality of sample structures corresponding to a plurality of sample proteins, each sample structure including at least sample sequence information and sample trunk structure information of the corresponding sample protein; By adding noise to at least a portion of the sample structures in the sample set, a plurality of sample inputs respectively corresponding to the plurality of sample proteins are obtained; For a first sample input among the multiple sample inputs, based on the first sample input, using the first sub-model, determining a prediction result of the structure of a first sample protein corresponding to the first sample input, wherein the prediction result includes first predicted sequence information and first predicted backbone structure information of the first sample protein; as well as Based on a first difference between the prediction result and the sample structure of the first sample protein, the parameters of the first sub-model are updated.
6. The method according to claim 5, wherein obtaining a plurality of sample inputs respectively corresponding to the plurality of sample proteins comprises at least one of the following: For a first proportion of sample proteins in the sample set, respectively adding noise to the sample sequence information and the sample trunk structure information in the corresponding sample structure to obtain a sample input corresponding to the corresponding sample protein, or For a second proportion of sample proteins in the sample set, noise is added to one of the sample sequence information or the sample trunk structure in the corresponding sample structure to obtain a sample input corresponding to the corresponding sample protein.
7. The method according to claim 6, wherein for a first proportion of sample proteins in the sample set, adding noise to the sample sequence information and the sample trunk structure information in the corresponding sample structure respectively comprises: selecting sample proteins from the sample set according to the first ratio; Based on a first time step parameter indicating a noise intensity corresponding to the sample sequence information of the selected sample protein, adding noise to the sample sequence information of the selected sample protein to obtain the sample sequence information after adding noise as a part of the sample input corresponding to the selected sample protein; as well as Based on a second time step parameter indicating the noise intensity of the sample backbone structure information of the selected sample protein, noise is added to the sample backbone structure information of the selected sample protein to obtain the sample backbone structure information after noise is added as a part of the sample input corresponding to the selected sample protein.
8. The method according to claim 5, wherein the plurality of sample proteins include a plurality of multi-chain proteins, each multi-chain protein includes a plurality of peptide chains, each item of sample sequence information includes a plurality of peptide chain sequence information corresponding to the plurality of peptide chains, each item of sample backbone structure information includes a plurality of peptide chain backbone structure information corresponding to the plurality of peptide chains, and Wherein, by adding noise to at least a part of the sample structures in the sample set, obtaining a plurality of sample inputs respectively corresponding to the plurality of sample proteins comprises: For a multi-chain protein among the plurality of multi-chain proteins, Noise is added to peptide chain sequence information and peptide chain backbone structure information corresponding to some of the peptide chains in the plurality of peptide chains to obtain corresponding sample inputs.
9. The method according to claim 5, wherein the first sub-model comprises a plurality of model units arranged in sequence, each of the plurality of model units is configured to generate an intermediate prediction result of the model unit on the structure of the first sample protein based on an output of a previous model unit or an input of the first sub-model, and wherein updating the parameters of the first sub-model comprises: Based on the first difference and a second difference between intermediate prediction results of two adjacent model units among the multiple model units, the parameters of the first sub-model are updated.
10. The method according to claim 1, wherein in the sub-model training stage of the protein model, the second sub-model is trained in the following manner: Acquire a sample set, wherein the sample set includes a plurality of sample structures corresponding to a plurality of sample proteins, each sample structure includes sample sequence information, sample backbone structure information, and sample side chain structure information of the corresponding sample protein; For a second sample protein among the plurality of sample proteins, based on the sample sequence information and the sample backbone structure information of the second sample protein, using the second sub-model, determine first predicted side chain structure information of at least one peptide chain of the second sample protein; as well as Based on a third difference between the sample side chain structure information of the second sample protein and the first predicted side chain structure information, the parameters of the second sub-model are updated.
11. The method according to claim 1, wherein in the joint training stage of the protein model, the first sub-model, the second sub-model and the third sub-model are jointly trained in the following manner: Acquire a sample set, wherein the sample set includes sample structures corresponding to a plurality of sample proteins, wherein the sample structures include sample sequence information, sample backbone structure information, and sample side chain structure information; Obtaining a plurality of sample inputs corresponding to the plurality of sample proteins by adding noise to at least one of the sample sequence information or the sample trunk structure information in at least a portion of the sample structures in the sample set; For a second sample input among the multiple sample inputs, based on the second sample input, using the first sub-model, determine second predicted sequence information and second predicted backbone structure information of a third sample protein corresponding to the second sample input; Based on the second predicted sequence information and the second predicted backbone structure information, using the second sub-model, determining the second predicted side chain structure information of the third sample protein; Based on the second predicted side chain structure information, the second predicted sequence information and the second predicted backbone structure information, using the third sub-model, determining the predicted structure of the third sample protein; as well as Based on a fourth difference between the predicted structure and the sample structure of the third sample protein, a parameter of at least one of the first sub-model, the second sub-model, or the third sub-model is updated.
12. The method according to claim 11, wherein updating the parameters of at least one of the first sub-model, the second sub-model or the third sub-model comprises: Based on the fourth difference, the parameters of one of the first sub-model, the second sub-model or the third sub-model are updated in a predetermined order to train the corresponding sub-model.
13. The method according to claim 11, wherein determining the second predicted side chain structure information of the third sample protein using the second sub-model based on the second predicted sequence information and the second predicted backbone structure information comprises: determining a third time step parameter indicative of noise intensity in the second sample input; as well as In response to the third time step parameter exceeding a predetermined parameter threshold, second predicted side chain structure information of the third sample protein is determined using the second sub-model based on the second predicted sequence information and the second predicted backbone structure information.
14. The method according to claim 11, wherein in the joint training phase of the protein model, the second sub-model is also trained by: For a fourth sample protein among the plurality of sample proteins, based on the sample sequence information and the sample backbone structure information of the fourth sample protein, using the second sub-model, determine third predicted side chain structure information of the fourth sample protein; and Based on a fifth difference between the sample side chain structure information of the fourth sample protein and the third predicted side chain structure information, the parameters of the second sub-model are updated.
15. A device for protein structure prediction, comprising: A first determination module is configured to determine, based on input information associated with the target protein, sequence information and backbone structure information of the target protein using a first submodel in the protein model, wherein the sequence information indicates an amino acid sequence of at least one peptide chain of the target protein, and the backbone structure information indicates a backbone conformation of the at least one peptide chain; A second determination module is configured to determine the side chain structure information of the at least one peptide chain based on the sequence information and the backbone structure information using a second sub-model in the protein model, wherein the side chain structure information indicates the side chain conformation at each residue in the at least one peptide chain; as well as The improvement module is configured to determine the structure of the target protein based on the side chain structure information, the sequence information and the backbone structure information using the third sub-model in the protein model.
16. An electronic device, comprising: at least one processor; as well as At least one memory, the at least one memory is coupled to the at least one processor and stores instructions for execution by the at least one processor, the instructions causing the electronic device to perform the method according to any one of claims 1 to 14 when executed by the at least one processor.
17. A computer-readable storage medium having computer-executable instructions stored thereon, wherein the computer-executable instructions can be executed by a processor to implement the method according to any one of claims 1 to 14.
18. A computer program product comprising computer executable instructions, wherein the computer executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 14.