Protein multi-conformation structure prediction method and electronic equipment thereof
Through the artificially generated MSA information optimization on the global energy landscape, the problem that the existing technology is difficult to effectively explore the multi-conformation state of proteins is solved, and more efficient multi-conformation structure prediction is achieved.
Patent Information
- Application Number
- CN202510282253.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art has limitations in the prediction of protein multi-conformational states, and it is difficult to effectively explore multiple local optimal conformations that a single sequence may correspond to.
By using artificially generated MSA information as search variables, it is directly optimized on the global energy landscape, breaking away from the dependence on natural sequence information, and improving the exploration efficiency of multi-conformational structures.
It improves the flexibility of exploring the multi-conformational state of proteins, and can more effectively discover multiple protein folding forms corresponding to a single sequence, improving the accuracy and reliability of predictions.
Smart Images

Figure CN120108487A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of protein design, and more specifically to a method for predicting protein multi-conformation structures and an electronic device thereof. Background Art
[0002] Proteins are the main executors of life activities, and the correct execution of their functions depends on specific three-dimensional structures. However, proteins are not static entities, but dynamically switch between multiple conformational states under physiological conditions to perform specific biological functions, such as signal transduction, enzyme catalysis, molecular recognition, etc. Therefore, accurately predicting and designing the multiple conformational states of proteins is of great significance for a deep understanding of their functional mechanisms and regulatory principles, as well as for directed evolution and drug design.
[0003] In recent years, artificial intelligence methods have made breakthrough progress in the prediction of single conformational structures of proteins, but there are still limitations in the prediction of multiple conformational states. Summary of the invention
[0004] In view of the shortcomings of the prior art, one of the purposes of the present invention is to propose a new method, which aims to use artificially generated MSA as a search variable to directly optimize on the global energy landscape, thereby getting rid of the dependence on natural sequence information and improving the exploration efficiency of multi-conformational structures.
[0005] To achieve the above object, the present invention provides the following technical solution: a method for predicting protein multi-conformation structure, comprising the following steps: 1) Generation of loss values: randomly generate MSA information, the MSA information is used as initial MSA information, the initial MSA information and sequence information are input into the AlphaFold2 model to generate an AF2 output loss value, the AF2 output loss value represents the output information of the AlphaFold2 model, the AF2 output loss value is converted into a loss value, the loss value is a computable value, and the loss value includes a contact loss value or / and a pLDDT loss value or / and a pAE loss value or / and a collision loss value or / and a gyration radius loss value or / and a secondary structure loss value; The contact loss value reflects the difference between the contact situation of amino acids in the predicted structure and the set ideal contact state; The pLDDT loss value reflects the self-assessed error in the accuracy of the local structure in the prediction; The pAE loss value reflects the self-assessment error of the model in predicting structural alignment without reference; the collision loss value reflects the degree of spatial conflict between non-bonded atoms in the predicted structure, which is used to measure the physical rationality and collision-free nature of the predicted structure; The gyration radius loss value reflects the difference between the gyration radius of the predicted structure and the target value; The secondary structure loss value reflects the difference between the secondary structure of the predicted structure and the target secondary structure; If the loss value includes more than two values, a weighted sum is performed as a new loss value; 2) Generation of AF2 output structure: Calculate the gradient of the loss value to the MSA information through back propagation, optimize the MSA information according to the gradient, obtain the optimized MSA information, repeatedly iterate the optimization to obtain the final optimized MSA information, input the final optimized MSA information and the sequence information into the AlphaFold2 model to obtain the AF2 output structure; 3) Generation of multi-conformational structures: repeat step 1) M times to obtain M AF2 output structures, input the M AF2 output structures as templates and the sequence information into the AlphaFold2 model, obtain the average pLDDT value of the M AF2 output structures, sort the M AF2 output structures according to the average pLDDT value of each AF2 output structure, and select the optimal m AF2 output structures as the multi-conformational structure.
[0006] Preferably, the AF2 output loss value includes the distance map logits Z con , the distance map logits Z con Includes all residue pairs C β -C β Atomic distance information, missing C β Glycine C α Instead, the conversion formula for contact loss value is:
[0007] Where b represents the distance interval; λ bin is a constant; R b ij Indicates whether the distance between the residue pair (i, j) exceeds the set λth bin Distance interval; Z ij,con represents the C between residue pair (i, j) β- C β logits vector of atomic distances; P ij,con represents the C between residue pair (i, j) β -C β The probability distribution of atomic distances falling in different distance intervals; P' ij,con Indicates that after 0-λ bin The C between residue pair (i, j) after region constraint adjustment β- C βThe probability distribution of atomic distances falling in different distance intervals; p b ij,con represents the C between residue pair (i, j) β -C β The probability that the atomic distance falls in the bth distance interval; p' b ij,con Indicates that after 0-λ bin The C between residue pair (i, j) after region constraint adjustment β -C β The probability that the atomic distance falls in the bth distance interval; contact is the contact loss value; L ij represents the loss of each residue pair (i, j); sort(L ij ) k For all L ij Sort and select the smallest N contact The loss items are averaged.
[0008] Preferably, the AF2 output loss value includes the logits value Z plddt , the logits value Z plddt Including pLDDT information of all residues, the pLDDT loss value conversion formula is:
[0009] Among them, P i,pLDDT represents the probability distribution of the pLDDT value of a single residue i falling within different confidence intervals; Z i,pLDDT The logits vector representing the pLDDT values of a single residue i falling into different confidence intervals; pLDDT is the pLDDT loss value; L is the protein sequence length; p b i,plddt represents the probability of residue i in the bth confidence interval.
[0010] Preferably, the AF2 output loss value includes a pairwise error matrix e, and the pairwise error matrix e includes alignment error information of all residue pairs. The conversion formula of the pAE loss value is:
[0011] Among them, P ij,pAE represents the probability distribution of residue pairs (i, j) falling into different error intervals, e ij The logits vector representing the residue pair (i, j) falling into different error intervals; pAE is the pAE loss value; L is the protein sequence length; p bij,pAE represents the probability of the residue pair (i, j) in the bth error interval; b is the center value of the bth distance interval.
[0012] Preferably, the AF2 output loss value includes N nbpairs and d pred(i) , the N nbpairs and d pred(i) Converted into collision loss value, the conversion formula is:
[0013] in, clash is the collision loss value; d lit(i) is the minimum distance that should be physically maintained between atoms; N nbpairs is the number of all non-bonded atom pairs in the structure; N nbpairs is the distance between the ith pair of non-bonded atoms in the predicted structure; is the tolerance.
[0014] Preferably, the AF2 output loss value includes the coordinate information r of all residues, and the coordinate information r of all residues is converted into the gyration radius loss value, and the conversion formula is:
[0015] in, rg is the loss value of the radius of gyration; R g is the radius of gyration; r i represents the i-th residue C α The coordinates of the atoms, L is the length of the protein sequence, R g is the loss value of the radius of gyration, R 0 is the preset turning radius threshold.
[0016] Preferably, the AF2 output loss value includes the distance map logits Z con , the distance map logits Z con Includes all residue pairs C β -C β Atomic distance information, missing C β Glycine C α Instead, assume that the distance range corresponding to the bth interval is [ ], the conversion formula of the secondary structure loss value is:
[0017] in, sse Indicates the secondary structure loss value; Pij,sse represents the C between residue pair (i, j) β The probability distribution of atomic distances falling in different distance intervals; Zi j,con represents the C between residue pair (i, j) β The logits vector of atoms whose distances fall within a certain distance interval; P ’ ij,sse represents the corrected C β The probability distribution of atomic distances falling in different distance intervals; L is the length of the protein sequence; b is the distance interval.
[0018] Preferably, the loss values at least include a contact loss value, and the contact loss value has the highest weight.
[0019] Preferably, the method for generating the MSA information in step 1) is: using Gumbel distributed noise to initialize the logits, and then processing the logits through a temperature-regulated softmax function to obtain the initial MSA information.
[0020] In view of the deficiencies in the prior art, a second object of the present invention is to provide a device capable of running the above algorithm.
[0021] To achieve the above object, the present invention provides the following technical solution: an electronic device, comprising: Processor and A memory storing executable codes, which, when executed by the processor, enable the processor to execute an algorithm corresponding to the method for predicting the multi-conformational structure of a protein.
[0022] Compared with the prior art, the advantages of the present invention are: many proteins have multiple local optimal conformations, and MSA information is an important basis for AlphaFold2 to predict protein structure. However, the MSA information composed of natural sequences usually obtained through sequence search can often only guide the model to predict a local optimal conformation, and it is difficult to effectively explore multiple local optimal conformations that may correspond to a single sequence. The present invention explores a larger conformational space by artificially randomly generating initial MSA information, thereby increasing the possibility of discovering multiple protein folding forms corresponding to a single sequence. At the same time, by evaluating the self-consistency and physical rationality of the predicted structure in multiple dimensions, the MSA information is directly optimized on the global energy landscape to improve the efficiency of exploring multi-conformational structures. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 A simplified flowchart of the method for predicting protein multi-conformational structures. DETAILED DESCRIPTION
[0024] The present invention is further described in detail below in conjunction with the accompanying drawings and embodiments.
[0025] Gumble noise is a type of random noise derived from a Gumble distribution.
[0026] The AlphaFold2 model is a deep learning model developed by Google's DeepMind, specifically for protein structure prediction.
[0027] MSA (Multiple Sequence Alignment) is a widely used method in bioinformatics to align multiple biomolecule sequences (such as protein or DNA sequences) to identify their similarities and conserved regions. The sequence conservation and co-evolution information provided by MSA is an important input feature for deep learning methods (such as AlphaFold2) to predict protein structure. This information helps the model capture the interactions and structural relationships between sequences more accurately.
[0028] One-hot encoding is a commonly used encoding method, which is mainly used to convert categorical data into a numerical form that can be input into machine / deep learning models. In this encoding method, each category is represented as a vector of all zeros, with a value of 1 only at the position corresponding to the category.
[0029] Logits is a term in deep learning and machine learning, usually referring to the raw scores output by the last layer of the model in classification tasks without being processed by an activation function (such as softmax). These scores represent the relative confidence of each category.
[0030] MLM (Masked Language Model) loss is a loss function used to train language models by randomly masking some elements in the input sequence and requiring the model to predict these masked parts. In protein design, it helps the model learn the contextual information of the amino acid sequence and generate reasonable sequences.
[0031] DSSP (Dictionary of Protein Secondary Structure) loss is a loss function used to predict protein secondary structure, which optimizes model predictions by comparing with real secondary structures (such as α-helix, β-sheet). It ensures that the generated protein structure conforms to the actual biophysical properties.
[0032] The contact map entropy loss is a loss function used to measure the contact pattern between residues in the three-dimensional structure of a protein. By controlling the entropy of the contact map, the model is encouraged to generate diverse contact patterns between residue pairs, thereby improving the predicted structural diversity. Example
[0033] like Figure 1 As shown, a method for predicting multi-conformation structures of proteins comprises the following steps: 1) Generation of loss values: Randomly generate MSA information, and the MSA information is used as the initial MSA information. Many proteins have multiple local optimal conformations, and MSA information is an important basis for AlphaFold2 to predict protein structure. However, the MSA information composed of natural sequences usually obtained through sequence search can often only guide the model to predict a local optimal conformation, and it is difficult to effectively explore multiple local optimal conformations that may correspond to a single sequence. By modifying the MSA information, a larger conformational space can be explored, and the possibility of discovering multiple protein folding forms corresponding to a single sequence can be increased. The initial MSA information and sequence information are input into the AlphaFold2 model to generate an AF2 output loss value, and the AF2 output loss value represents the output information of the AlphaFold2 model. The AF2 output loss value is converted into a loss value, and the loss value is a computable value. The loss value includes a contact loss value or / and a pLDDT loss value or / and a pAE loss value or / and a clash loss value or / and R g loss value and / or secondary structure loss value; The contact loss value reflects the difference between the contact situation of amino acids in the predicted structure and the set ideal contact state; The pLDDT loss value reflects the self-assessed error in the accuracy of the local structure in the prediction; The pAE loss value reflects the self-assessment error of the model in predicting structural alignment without reference; the collision loss value reflects the degree of spatial conflict between non-bonded atoms in the predicted structure, which is used to measure the physical rationality and collision-free nature of the predicted structure; The gyration radius loss value reflects the difference between the gyration radius of the predicted structure and the target value; The secondary structure loss value reflects the difference between the secondary structure of the predicted structure and the target secondary structure; If the loss value includes more than two values, a weighted sum is performed as a new loss value. By using the artificially generated MSA as the search variable, optimization is performed directly on the global energy landscape, thereby getting rid of the dependence on natural sequence information and exploring the conformational space of proteins more flexibly.
[0034] The loss value may also include other loss values used to reflect the difference between the predicted structure and the target structure, such as MLM (Masked Language Model) loss, DSSP (Dictionary of Protein Secondary Structure) loss, contact map entropy loss, etc. The above loss values have specific calculation methods in the prior art and will not be described in detail here.
[0035] There are other calculation methods for the contact loss value, pLDDT loss value, pAE loss value, collision loss value, gyration radius loss value and secondary structure loss value in the prior art, which will not be elaborated here.
[0036] 2) Generation of AF2 output structure: Calculate the gradient of the loss value to the MSA information through back propagation, optimize the MSA information according to the gradient, obtain the optimized MSA information, repeatedly iterate the optimization to obtain the final optimized MSA information, input the final optimized MSA information and the sequence information into the AlphaFold2 model to obtain the AF2 output structure. The gradient back propagation method is a prior art and will not be described in detail here; 3) Generation of multi-conformational structures: repeat step 1) M times to obtain M AF2 output structures, input the M AF2 output structures as templates and the sequence information into the AlphaFold2 model, obtain the average pLDDT value of the M AF2 output structures, sort the M AF2 output structures according to the average pLDDT value of each AF2 output structure, and select the optimal m AF2 output structures as the multi-conformational structure.
[0037] The present invention improves the flexibility of exploring the multi-conformational states of proteins by generating and optimizing MSA information; can flexibly combine various loss functions to adapt to different scenarios; introduces physical rationality constraints to ensure that the predicted structure meets the actual physical conditions; uses a confidence assessment mechanism for adaptive optimization to improve the reliability of prediction; and supports a wide range of application scenarios, such as studying protein functional diversity and kinetic behavior. Example
[0038] The difference from Example 1 is that the method for generating the initial MSA information is to initialize the logits using Gumbel distributed noise. This step introduces a certain amount of randomness into the sequence matrix. Then, these logits are processed by a temperature-regulated softmax function to obtain an initial MSA information, which represents the probability distribution of selecting each amino acid at different amino acid positions. The temperature parameter is used to control the smoothness of the distribution, a lower temperature makes the distribution sharper, and a higher temperature makes the distribution smoother. Gumble noise is a random noise derived from a Gumble distribution. Adding gumble noise to logits and then taking softmax with a temperature parameter is also called the "Gumble-softmax method" (prior art), which can convert continuous logits into a nearly discrete one-hot encoding representation through a gumble-softmax operation. Example
[0039] The difference from the first embodiment is that the loss value includes the contact loss value.
[0040] The calculation method of the contact loss value specifically listed in the present invention is different from that in the prior art, specifically: The AF2 output loss values include the distance map logits Z con , distance map logits Z con is a L×L×64 tensor representing the pairwise contacts of the residues in the predicted structure. The distance map logits Z con Includes all residue pairs C β -C β Atomic distance information, missing C β Glycine C α Instead, the conversion formula for contact loss value is:
[0041] Where b represents the distance interval; λ bin is a constant; R b ij Indicates whether the distance between the residue pair (i, j) exceeds the set λth bin Distance interval; Z ij,con represents the C between residue pair (i, j) β- C β logits vector of atomic distances; P ij,con represents the C between residue pair (i, j) β -C β The probability distribution of atomic distances falling in different distance intervals; P' ij,con Indicates that after 0-λ binThe C between residue pair (i, j) after region constraint adjustment β- C β The probability distribution of atomic distances falling in different distance intervals; p b ij,con represents the C between residue pair (i, j) β -C β The probability that the atomic distance falls in the bth distance interval; p' b ij,con Indicates that after 0-λ bin The C between residue pair (i, j) after region constraint adjustment β -C β The probability that the atomic distance falls in the bth distance interval; sort(L ij ) k For all L ij Sort and select the smallest N contact The loss items are averaged.
[0042] Distance map logits Z con is a three-dimensional tensor whose last dimension contains 64 logits representing different distance intervals between residue pairs, which cover the range of 2 to 22 Å. In order to ensure that the output of the model has a reasonable number of residue contacts, we ij,con The regional constraint correction was performed and the corrected probability distribution P' was obtained. ij,con The correction process is achieved by introducing a region matrix R, whose elements R b ij Mark whether the distance between residue pairs (i, j) falls within λ bin Within the critical interval. bin is a constant, usually 14Å (i.e. the first 32 distance intervals). Then R b ij Multiply by a very large value and subtract it to significantly reduce the probability of not meeting the region conditions. The above contact loss value calculation method is a new method for calculating the contact loss value invented by Liwen. In this way, the model's ability to identify and optimize key residue contacts is further optimized, encouraging the model to generate high-quality protein structures with a reasonable number of contacts. Example
[0043] The difference from Example 1 is that the loss value includes the pLDDT loss value.
[0044] The calculation method of the pLDDT loss value specifically listed in the present invention is different from the prior art, specifically: LDDT (Local Distance Difference Test) is a score used to evaluate the local accuracy of protein structure models. It directly measures the degree of preservation of the local environment by comparing the distance differences between all atom pairs in the model and the reference structure without aligning the model to the reference structure. In AF2, in order to evaluate the accuracy of model prediction, pLDDT (predicted Local Distance Difference Test) was introduced, that is, the predicted LDDT-C α The AF2 output loss value includes the predicted LDDT-C α score, the LDDT-C α The score is converted into the pLDDT loss value, the LDDT-C α The score is linearly projected to 50 confidence intervals (0-1) through a single representation in the AlphaFold2 model to obtain the logits value Z pLDDT (This calculation process is performed inside the AlphaFold2 model), logits value Z pLDDT is a L×50 matrix. The logits value Z pLDDT Including pLDDT information of all residues, the pLDDT loss value conversion formula is:
[0045] Among them, P i,pLDDT represents the probability distribution of the pLDDT value of a single residue i falling within different confidence intervals; Z i,pLDDT The logits vector representing the pLDDT value of a single residue i falling into different confidence intervals; L is the length of the protein sequence; p b i,plddt represents the probability of residue i in the bth confidence interval, and (1-b / 50) is the weight used to calculate the weighted average, which aims to impose a greater penalty on lower pLDDT values. By weighted summing and averaging the pLDDT distributions of all residues, the pLDDT loss can effectively penalize residue predictions with lower pLDDT values, thereby encouraging the model to generate higher confidence predicted structures.
[0046] The above pLDDT loss value calculation method is a new pLDDT loss value calculation method invented by Liwen. In this way, the pLDDT loss value helps guide the model to predict conformations with higher confidence, thereby identifying which conformations are more likely to reflect the real situation. Example
[0047] The difference from Example 1 is that the loss value includes the pAE loss value.
[0048] The calculation method of the pAE loss value specifically listed in the present invention is different from that in the prior art, specifically: The predicted alignment error (pAE) reflects the self-assessment error of the model in predicting structural alignment without reference. The AF2 output loss value includes a pairwise representation, which is converted into the pAE loss value. The pairwise representation is linearly projected into 64 distance intervals (0~31.5 Å) to obtain the pairwise error matrix e (this calculation process is performed inside the AlphaFold2 model). The pairwise error matrix e includes the alignment error information of all residue pairs. The conversion formula of the pAE loss value is:
[0049] Among them, P ij,pAE represents the probability distribution of residue pairs (i, j) falling into different error intervals, e ij The logits vector representing the residue pair (i, j) falling into different error intervals; L is the length of the protein sequence; p b ij,pAE represents the probability of the residue pair (i, j) in the bth error interval; is the center value of the bth distance interval, which is used to represent the typical distance within this distance interval. By summing and averaging the weighted errors of all residue pairs, pAE loss can effectively penalize predictions with large errors, thereby improving the prediction accuracy of the model. This process ensures that the model can align residues more accurately during structure prediction, thereby reducing the error between the prediction and the actual structure.
[0050] The above-mentioned pAE loss value calculation method is a new pAE loss value calculation method invented by Liwen. In this way, it helps to evaluate the relative accuracy of different conformations and guide the model to predict conformations that are closer to the real structure. Example
[0051] The difference from the first embodiment is that the loss value includes the collision loss value.
[0052] The calculation method of the collision loss value specifically listed in the present invention is different from that in the prior art, specifically: In AF2, collision loss is used to avoid generating protein structures with unreasonable atomic collisions. It ensures that the structure generated by the model is physically reasonable by penalizing the short interatomic distances in the predicted structure. The calculation of collision loss is based on a one-sided flat-bottom potential, which only imposes penalties on atom pairs below a certain distance threshold. The AF2 output loss value includes N nbpairsand d pred(i) , the N nbpairs and d pred(i) Converted into collision loss value, the conversion formula is:
[0053] Among them, d lit(i) is the minimum distance that should be physically maintained between atoms; N nbpairs is the number of all non-bonded atom pairs in the structure; N nbpairs is the distance between the ith pair of non-bonded atoms in the predicted structure; is the tolerance, typically 1.5 Å. This results in a higher collision penalty when an atom pair causes an unreasonable collision. By calculating this distance penalty for all non-bonded atom pairs, the collision penalty effectively avoids the model from generating structures with atomic collisions, thereby ensuring the physical rationality of the predicted structure.
[0054] The above collision loss value calculation method is a new collision loss value calculation method invented by Liwen. The model may generate multiple conformations with different physical properties, so the collision loss value can guide the model to generate a more physically reasonable conformation. Example
[0055] The difference from Example 1 is that the loss value includes the gyration radius loss value.
[0056] The calculation method of the gyration radius loss value specifically listed in the present invention is different from that in the prior art, specifically: Radius of Gyration g ) is a physical quantity used to describe the mass distribution of a protein around its center of mass. It measures the compactness of a structure by calculating the square root of the average square distance of all components relative to the center of mass. The AF2 output loss value includes the coordinate information r of all residues, and the coordinate information r of all residues is converted into the gyration radius loss value. The conversion formula is:
[0057] Among them, R g is the radius of gyration; r i represents the i-th residue C α The coordinates of the atoms, L is the length of the protein sequence, R g is the loss value of the radius of gyration, R 0 is a preset gyration radius threshold used to constrain the compactness of the structure. This formula describes the spatial distribution of all residues relative to the center of mass. In order to avoid generating long helices that deviate from the actual structure, the gyration radius is set to constrain the generation of the structure. g Less than or equal to R0 , use the exponential function to fine-tune the radius of gyration. g More than R 0 , directly use R 0 As a loss, a stronger penalty is imposed on the generated structure with too large a radius of gyration. By imposing this form of constraint on the radius of gyration, it is possible to avoid the generation of unreasonable long helices when generating protein structures, thereby improving the physical rationality and accuracy of structure prediction.
[0058] The above calculation method of the gyration radius loss value is a new calculation method invented by Liwen. The compactness or ductility of the conformation is an important feature, and the gyration radius loss value can guide the model to predict a conformation that is more in line with the target characteristics. Example
[0059] The difference from Example 1 is that the loss value includes the secondary structure loss value.
[0060] The calculation method of the secondary structure loss value specifically listed in the present invention is different from that in the prior art, specifically: The AF2 output loss value includes a pair representation, the pair representation is converted into the secondary structure loss value, and the pair representation is symmetric and linearly projected to obtain a distance map logits Z con (This calculation process is performed inside the AlphaFold2 model), the distance map logits Z con Includes all residue pairs C β -C β Atomic distance information, missing C β Glycine C α Instead, in order to identify the secondary structure of proteins, we define the indicator function of the secondary structure based on the boundary conditions of the distance interval, assuming that the distance range corresponding to the bth interval is [ ].
[0061] α-helix: When the distance between residue pairs Below 6.5 Å, the region may correspond to an α-helix.
[0062] β-fold: When the distance between residue pairs Above 8.5 Å, this region may correspond to β-sheets.
[0063]
[0064] Based on these criteria, the indicator functions of α-helix and β-sheet are combined to obtain the secondary structure indicator matrix B. , generate the modified probability distribution and calculate the original probability distribution at the same time:
[0065] Among them, P ij,sse represents the C between residue pair (i, j) β The probability distribution of atomic distances falling in different distance intervals; Zi j,con represents the C between residue pair (i, j) β The logits vector of atoms whose distances fall within a certain distance interval; P ’ ij,sse represents the corrected C β The probability distribution of atomic distances falling into different distance intervals (a total of 64 distance intervals, representing the distance range of 2~22Å).
[0066] In this way, the corrected probability distribution will significantly reduce the probability of residue pairs that do not conform to the secondary structure indicator matrix B.
[0067]
[0068] L is the length of the protein sequence; b is the distance interval. The summation symbol condition j=i+3 in the formula has been written directly into the expression, so the loss function is only calculated for the matrix elements that are 3 residues apart. This means that we are concerned about the impact of 3 residues on the structure - a common geometric feature in α-helical and β-pleated structures. By calculating the cross entropy loss of the corrected probability distribution and the true distribution, the model can effectively optimize its prediction results to better match the actual secondary structure.
[0069] The above-mentioned method for calculating the secondary structure loss value is a new method for calculating the secondary structure loss value invented by Liwen. This secondary structure loss value can help guide the model to capture these conformation-specific secondary structure features.
[0070] Embodiment 9: The difference from the first embodiment is that the loss value at least includes the contact loss value, and the contact loss value has the highest weight.
[0071] Embodiment 10: An electronic device, comprising: Processor and A memory storing executable codes, which, when executed by the processor, cause the processor to execute an algorithm corresponding to the method for predicting protein multi-conformational structures disclosed in Examples 1-9.
[0072] The above is only a preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary researchers in this technical field, some improvements and modifications without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.
Claims
1. A method for predicting protein multi-conformational structures, characterized in that The following steps are involved: 1) Generation of loss values: randomly generate MSA information, the MSA information is used as initial MSA information, the initial MSA information and sequence information are input into the AlphaFold2 model to generate an AF2 output loss value, the AF2 output loss value represents the output information of the AlphaFold2 model, the AF2 output loss value is converted into a loss value, the loss value is a computable value, and the loss value includes a contact loss value or / and a pLDDT loss value or / and a pAE loss value or / and a collision loss value or / and a gyration radius loss value or / and a secondary structure loss value; The loss values reflect the self-consistency and physical plausibility of the predicted structure; The contact loss value reflects the difference between the contact situation of amino acids in the predicted structure and the set ideal contact state; The pLDDT loss value reflects the self-assessed error in the accuracy of the local structure in the prediction; The pAE loss value reflects the self-assessment error of the model in predicting structural alignment without reference; the collision loss value reflects the degree of spatial conflict between non-bonded atoms in the predicted structure, which is used to measure the physical rationality and collision-free nature of the predicted structure; The gyration radius loss value reflects the difference between the gyration radius of the predicted structure and the target value; The secondary structure loss value reflects the difference between the secondary structure of the predicted structure and the target secondary structure; If the loss value includes more than two values, a weighted sum is performed as a new loss value; 2) Generation of AF2 output structure: Calculate the gradient of the loss value to the MSA information through back propagation, optimize the MSA information according to the gradient, obtain the optimized MSA information, repeatedly iterate the optimization to obtain the final optimized MSA information, input the final optimized MSA information and the sequence information into the AlphaFold2 model to obtain the AF2 output structure; 3) Generation of multi-conformational structures: repeat step 1) M times to obtain M AF2 output structures, input the M AF2 output structures as templates and the sequence information into the AlphaFold2 model, obtain the average pLDDT value of the M AF2 output structures, sort the M AF2 output structures according to the average pLDDT value of each AF2 output structure, and select the optimal m AF2 output structures as the multi-conformational structure.
2. A method for predicting protein multi-conformational structures according to claim 1, characterized in that: The AF2 output loss values include the distance map logits Z con , the distance map logits Z con Includes all residue pairs C β -C β Atomic distance information, missing C β Glycine C α Instead, the conversion formula for contact loss value is: Where b represents the distance interval; λ bin is a constant; R b ij Indicates whether the distance between the residue pair (i, j) exceeds the set λth bin Distance interval; Z ij,con represents the C between residue pair (i, j) β- C β logits vector of atomic distances; P ij,con represents the C between residue pair (i, j) β -C β The probability distribution of atomic distances falling in different distance intervals; P' ij,con Indicates that after 0-λ bin The C between residue pair (i, j) after region constraint adjustment β- C β The probability distribution of atomic distances falling in different distance intervals; p b ij,con represents the C between residue pair (i, j) β -C β The probability that the atomic distance falls in the bth distance interval; p' b ij,con Indicates that after 0-λ bin The C between residue pair (i, j) after region constraint adjustment β -C β The probability that the atomic distance falls in the bth distance interval; contact is the contact loss value; L ij represents the loss of each residue pair (i, j); sort(L ij ) k For all L ij Sort and select the smallest N contact The loss items are averaged.
3. A method for predicting protein multi-conformational structures according to claim 1, characterized in that: The AF2 output loss value includes the logits value Z plddt , the logits value Z plddt Including pLDDT information of all residues, the pLDDT loss value conversion formula is: Among them, P i,pLDDT represents the probability distribution of the pLDDT value of a single residue i falling within different confidence intervals; Z i,pLDDT The logits vector representing the pLDDT values of a single residue i falling into different confidence intervals; pLDDT is the pLDDT loss value; L is the protein sequence length; p b i,plddt represents the probability of residue i in the bth confidence interval.
4. The method for predicting a protein multi-conformational structure according to claim 1, characterized in that: The AF2 output loss value includes a pairwise error matrix e, which includes the alignment error information of all residue pairs. The conversion formula of the pAE loss value is: Among them, P ij,pAE represents the probability distribution of residue pairs (i, j) falling into different error intervals, e ij The logits vector representing the residue pair (i, j) falling into different error intervals; pAE is the pAE loss value; L is the protein sequence length; p b ij,pAE represents the probability of the residue pair (i, j) in the bth error interval; is the center value of the bth distance interval.
5. The method for predicting a protein multi-conformational structure according to claim 1, characterized in that: The AF2 output loss value includes N nbpairs and d pred(i) , the N nbpairs and d pred(i) Converted into collision loss value, the conversion formula is: in, clash is the collision loss value; d lit(i) is the minimum distance that should be physically maintained between atoms; N nbpairs is the number of all non-bonded atom pairs in the structure; N nbpairs is the distance between the ith pair of non-bonded atoms in the predicted structure; is the tolerance.
6. A method for predicting protein multi-conformational structures according to claim 1, characterized in that: The AF2 output loss value includes the coordinate information r of all residues. The coordinate information r of all residues is converted into the gyration radius loss value. The conversion formula is: in, rg is the loss value of the radius of gyration; R g is the radius of gyration; r i represents the i-th residue C α The coordinates of the atoms, L is the length of the protein sequence, R g is the turning radius loss value, and R0 is the preset turning radius threshold.
7. A method for predicting protein multi-conformational structures according to claim 1, characterized in that: The AF2 output loss values include the distance map logits Z con , the distance map logits Z con Includes all residue pairs C β -C β Atomic distance information, missing C β Glycine C α Instead, assume that the distance range corresponding to the bth interval is [ ], the conversion formula of the secondary structure loss value is: in, sse represents the secondary structure loss value; P ij,sse represents the C between residue pair (i, j) β The probability distribution of atomic distances falling in different distance intervals; Zi j,con represents the C between residue pair (i, j) β The logits vector of atoms whose distances fall within a certain distance interval; P ’ ij,sse represents the corrected C β The probability distribution of atomic distances falling in different distance intervals; L is the length of the protein sequence; b is the distance interval.
8. The method for predicting protein multi-conformational structures according to claim 1, characterized in that: The loss values at least include a contact loss value, and the contact loss value has the highest weight.
9. A method for predicting protein multi-conformational structures according to claim 3, characterized in that: The method for generating the MSA information in step 1) is: using Gumbel distributed noise to initialize logits, and then processing the logits through a temperature-regulated softmax function to obtain initial MSA information.
10. An electronic device, characterized in that: include: Processor and A memory storing executable codes, which, when executed by the processor, enable the processor to execute a method for predicting a protein multi-conformational structure according to any one of claims 1 to 9.