Protein structure prediction method based on the mutual attraction between low-entropy hydration layers of residue side chains

By using water entropy force to predict the protein structure based on the mutual attraction relationship between the low-entropy hydrated layers of the side chain of amino acid residues, the problem of poor prediction effect in the prior art is solved, and accurate prediction and kinetic analysis of the protein structure are achieved.

CN119207542BActive Publication Date: 2025-08-19HARBIN INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411219964.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-02
Publication Date
2025-08-19
Estimated Expiration
2044-09-02

AI Technical Summary

Technical Problem

In the prior art, in the prediction of protein structure, methods based on the effects of hydrogen bonds, electrostatic forces, van der Waals forces, etc. have poor prediction effect, and it is impossible to accurately predict the molecular structure of a protein.

Method used

Using a method based on the mutual attraction relationship between the low-entropy hydrated layers of amino acid residue side chains, the protein structure is predicted by using water entropy force, and the secondary and tertiary structures of the protein are predicted by determining the preferred pairing and entropy appreciation between the amino acid residue side chains.

Benefits of technology

Accurately analyzing the folding dynamics mechanism of protein structures can accurately predict the secondary and tertiary structures of proteins, helping scientists design new proteins and perform purposeful mutations for important biological applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119207542B_ABST
    Figure CN119207542B_ABST
Patent Text Reader

Abstract

The protein structure prediction method based on the mutual attraction relationship between the low entropy hydration layers of the side chains of the residues belongs to the field of structural biology technology. In order to solve the problem that the current prediction method of protein folding structure has poor prediction effect. The protein structure prediction method based on the mutual attraction relationship between the low entropy hydration layers of the side chains of the residues described in the present invention uses water entropy force to realize protein structure prediction; the prediction is based on the primary structure of the protein to predict the secondary structure and tertiary structure of the protein; water entropy force refers to the interaction attraction between the low entropy hydration layers of the side chains of the amino residues; the interaction attraction refers to the entropy increase of the low entropy water molecules in the low entropy hydration layer driving the lateral adhesion between the side chains of the amino acid residues; lateral adhesion refers to the state in which the side chains of the two residues are in a nearly parallel state. The present invention is used for protein structure prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of structural biology, and in particular relates to a protein structure prediction method. Background Art

[0002] Protein structure prediction is crucial for biological research. Protein function is largely determined by its structure, so determining a protein's three-dimensional structure is crucial for understanding its functions, interactions, and pathological mechanisms. Currently, protein structure prediction primarily utilizes methods such as deep learning, comparative modeling, and molecular dynamics simulations. Deep learning, also known as artificial intelligence and machine learning, has played a key role in protein structure prediction, but advancements in this area can only be attributed to advances in artificial intelligence. Artificial intelligence methods currently lack clear physical and chemical mechanisms for directly deriving protein folding rules and predicting protein structure. Furthermore, molecular dynamics simulations, which attempt to simulate the protein folding process in a computer, are considered a physics-based approach to solving the protein folding problem. This approach requires significant high-performance computing resources to simulate the motion and interactions of polypeptide chains over time to infer the most stable protein structure. However, current protein folding dynamics simulations remain unable to accurately predict protein molecular structures.

[0003] The prevailing view in academia today is that the physical driving forces for protein folding include hydrophobic interactions, hydrogen bonding, electrostatic forces, van der Waals forces, and ionic bonding forces. Predictions based on this theory are not ideal for predicting protein folding structures. Accurate protein folding only occurs in an aqueous environment. In any other non-aqueous solvent, the vast majority of proteins cannot fold correctly. It must be emphasized that water molecules have extremely strong electrical polarity and a very high dielectric constant, which creates a significant electrostatic shielding effect. Therefore, it can be seen that the aforementioned prediction methods based on hydrogen bonding, electrostatic forces, van der Waals forces, and ionic bonding forces cannot be well explained theoretically. This also explains why the aforementioned prediction methods have poor results, and confirms that the aforementioned prediction methods have theoretically exhibited unsatisfactory results. Currently, there is no method for predicting protein structure based solely on the hydrophobic interactions between amino acid residue side chains, and this is urgently needed. Summary of the Invention

[0004] The present invention aims to solve the problem that current methods for predicting protein folding structures have poor prediction effects.

[0005] A protein structure prediction method based on the mutual attraction relationship between the low-entropy hydration layers of residue side chains, the prediction method uses water entropy force to achieve protein structure prediction; the prediction is based on the primary structure of the protein to predict the secondary structure and tertiary structure of the protein; the water entropy force refers to the interaction attraction between the low-entropy hydration layers of the amino residue side chains; the interaction attraction refers to the entropy increase of low-entropy water molecules in the low-entropy hydration layer driving the lateral adhesion between the amino acid residue side chains; the lateral adhesion refers to the two residue side chains being in a nearly parallel state.

[0006] Furthermore, the process of using water entropy force to predict protein structure includes:

[0007] 1. Preliminary preparation for forecasting:

[0008] First, an entropy value is assigned to the side chains of amino acid residues in the primary structure of the protein. The entropy value is a relative measure of the entropy increase potential of the low-entropy hydration layer around the side chains of the amino acid residues.

[0009] Then, the amino acid residue side chains in the primary structure of the protein after entropy assignment are divided into five residue forms:

[0010] Residues with micro-entropy increase potential: serine S, threonine T, aspartic acid D, asparagine N;

[0011] Hydrophilic residues: histidine H, arginine R, lysine K, glutamic acid E, glutamine Q, aspartic acid D, asparagine N, tyrosine Y, tryptophan W, serine S;

[0012] Turn residues: glycine G, proline P;

[0013] Hydrophobic residues: isoleucine I, valine V, leucine L, phenylalanine F, tyrosine Y, tryptophan W, alanine A, methionine M, cysteine C, histidine H;

[0014] Residues with high entropy increase potential: isoleucine I, valine V, leucine L, lysine K, arginine R, phenylalanine F, tyrosine Y, tryptophan W, alanine A, methionine M, cysteine C, glutamic acid E, glutamine Q, histidine H;

[0015] Finally, determine the preferred pairing between amino acid residue side chains

[0016] The preferred pairing between amino acid residue side chains is that the lateral attachment between multiple pairs of amino acid residue side chains results in sufficient entropy increase in the low-entropy hydration layer of the multiple pairs of amino acid residue side chains. The pair relationship between the multiple pairs of amino acid residues is the preferred pairing between amino acid residue side chains.

[0017] The preferred pairing between amino acid residue side chains includes preferred pairing between amino acid residue side chains in a 1-4 / 5 sequence position relationship, preferred pairing between amino acid residue side chains in a 1-3 sequence position relationship, and preferred pairing between amino acid residue side chains in a 1-2 sequence position relationship;

[0018] The 1-4 / 5 sequence position relationship refers to the relationship between the side chain of one residue and the side chain of another residue in the primary structure of the protein, which are separated by 2 or 3 amino acid residues.

[0019] The 1-3 sequence position relationship refers to the relationship between the side chain of one residue and the side chain of another residue in the primary structure of the protein, which are separated by one amino acid residue.

[0020] The 1-2 sequence position relationship refers to the sequence position relationship between a side chain of one residue and the side chain of another adjacent residue in the primary structure of the protein;

[0021] 2. Protein secondary structure prediction:

[0022] Based on the sequence position of the residue side chains, the classification and entropy of the amino acid residues, and the preferred pairing between the residue side chains in step 1, the following process is used to predict the protein secondary structure:

[0023] S100, starting a new line below the amino acid sequence marked with entropy values, repeatedly listing the amino acid residues in the amino acid sequence that are high entropy increase potential residues as a high entropy increase potential residue connected segment; then, starting a new line, repeatedly listing the amino acid residues in the amino acid sequence that are low entropy residues and turn residues;

[0024] S200. List the unblocked connected segments of residues with high entropy increase potential, and assume that the listed unblocked connected segments of residues with high entropy increase potential are α-helices or β-sheets in the secondary structure of the protein:

[0025] S201. Assuming that a connected segment of residues with high entropy increase potential forms an α-helix, if a side chain of a residue in the connected segment of residues with high entropy increase potential achieves preferred pairing between side chains of amino acid residues, the amino acid residue in the connected segment of residues with high entropy increase potential is marked with an *.

[0026] S202. Assuming that the connected segment of residues with high entropy increase potential forms a β-sheet, if the connected segment of residues with high entropy increase potential is a preferred pairing between the side chains of amino acid residues in a 1-3 sequence position relationship, then the amino acid residues in the connected segment of residues with high entropy increase potential are marked with &;

[0027] S203, analyze the marked * and &:

[0028] S2031. If none of the connected segments of residues with high entropy increase potential in the primary structure of the protein are marked with &-&, and more than 80% of the amino acid residues in the connected segment of residues with high entropy increase potential are marked with *, then the connected segment of residues with high entropy increase potential is predicted to form an α-helix;

[0029] S2032. If a connected segment of residues with high entropy increase potential in the primary structure of the protein is marked with &-&, it is predicted that the connected segment of residues with high entropy increase potential forms a β-sheet; if the number of * marked in a connected segment of residues with high entropy increase potential in the primary structure of the protein is less than 50% of the number of amino acid residues in the connected segment of residues with high entropy increase potential, it is predicted that the connected segment of residues with high entropy increase potential forms a β-sheet;

[0030] S300, searching the amino acid side chains on both sides of the connected segment of residues with high entropy increase potential that has been predicted in S200 to form an α-helix, if the amino acid residues in the amino acid side chains on both sides have a preferred pairing between the side chains of the amino acid residues in a 1-4 / 5 sequence position relationship with the amino acid residues in the connected segment of residues with high entropy increase potential, then it is predicted that the side chains of the amino acid residues in the 1-4 / 5 sequence position relationship preferably pair to form an α-helix;

[0031] S400, determine whether the S200 hypothesis predicts the formation of α-helix or β-sheet in the protein secondary structure:

[0032] Counting the entropy increments of amino acid residues in the primary structure of the protein predicted to be α-helix and β-sheet respectively, and comparing the entropy increments of the amino acid residues predicted to be α-helix with the entropy increments of the amino acid residues predicted to be β-sheet; when the value of the relatively larger entropy increment of the two is greater than the value of the other entropy increment by more than 20%, the prediction result of the relatively larger entropy value is determined as the final prediction result; when the value of the relatively larger entropy increment of the two is less than 20%, the prediction of the formation of α-helix or β-sheet in the secondary structure of the protein is determined as a potential result in S200;

[0033] Among them, the entropy increment statistical method of β-folding: there is a preferred pairing between the side chains of amino acid residues with a 1-3 sequence position relationship in the primary structure of the protein, and the entropy value of the amino acid residues in the preferred pairing is used as the entropy increment;

[0034] Statistical method for entropy increment of α-helix: there is a preferred pairing between the side chains of amino acid residues with a 1-4 / 5 sequence position relationship in the primary structure of the protein, and the entropy value of the amino acid residues in the preferred pairing is used as the entropy increment;

[0035] S500: If there are two consecutive amino acid residues with low entropy increasing potential in the segment connected by residues with high entropy increasing potential that are not predicted to be α-helices in the primary structure of the protein, the two amino acid residues are predicted to form a turn and are marked as T;

[0036] A segment of 4 or more consecutive hydrophilic amino acid residues in the primary structure of a protein is named a "continuous hydrophilic residue segment"; if there is a segment with more than 7 amino acid residues in the connected segment of high entropy increase potential residues that are not predicted to be α-helices in the primary structure of a protein, and the number of residues with low entropy increase potential and / or turn residues and / or amino acid residue A and / or residues in the continuous hydrophilic residue segment accounts for more than or equal to 50% of the number of residues in the segment with more than 7 amino acid residues, then the segment with more than 7 amino acid residues is predicted to form a random coil and is used as a potential prediction result;

[0037] S600. A fragment of a protein primary structure consisting of five or more consecutive hydrophilic amino acid residues, if the fragment has been predicted to be a connected fragment of residues with high entropy increase potential, then the fragment consisting of five or more consecutive hydrophilic amino acid residues is predicted to form a random curl, and used as a potential prediction result, thereby completing the protein structure prediction method based on the mutual attraction relationship between the low-entropy hydration layers of the residue side chains.

[0038] Furthermore, preferred pairings between amino acid residue side chains include:

[0039] FF, QQ, KK, EE, TT, RR, AA, SS, MM;

[0040] MA, QE, TS, RE, KR, EK, FY, WY, WE, RW, RY, KY, HV, DE, NE, DQ, NQ;

[0041] QR / K, I / V / L / FI / V / L / F, I / V / L / F / YI / V / L / F / Y, I / V / LY, I / V / LK, I / V / LR, I / V / LA, I / V / LW, I / V / L / F / YW, I / V / LM, CM / Y, MI / V / L, A / TD / N.

[0042] Furthermore, the preferred pairing between amino acid residue side chains in step S201 refers to the preferred pairing between amino acid residue side chains with a 1-4 / 5 sequence position relationship; if there is a preferred pairing between amino acid residue side chains with a 1-2 sequence position relationship in the connected segment of residues with high entropy increase potential, or there is a sequence position relationship of I / V / L 1-4 / 5G / A1-4 / 5I / V / L / F in the connected segment of residues with high entropy increase potential, then it is marked as * below the amino acid.

[0043] Furthermore, the preferred pairing between the side chains of the amino acid residues in the 1-3 sequence position relationship described in S202 is I / V / L / FI / V / L / F; if, taking the amino acids in I / V / L / FI / V / L / F as the starting point, there is a preferred pairing between the side chains of the amino acid residues in the 1-4 / 5 sequence position relationship in the adjacent connected fragment of high entropy increase potential residues, then the amino acids in the preferred pairing between the side chains of the amino acid residues in the 1-3 sequence position relationship are not marked as &.

[0044] Furthermore, as described in S2032, if the connected segment of residues with high entropy increase potential in the primary structure of the protein is marked with &-&, and more than 85% of the amino acid residues in the amino acid side chain region of the primary structure of the protein away from the connected segment of residues with high entropy increase potential marked with &-& are marked with *, and there is a preferred pairing between the side chains of amino acid residues that has a 1-4 / 5 sequence position relationship of I / V / L / F / YI / V / L / F / Y with the irrelevant residues marked with &-&, then it is predicted that the connected segment of residues with high entropy increase potential marked with &-& forms an α-helix; wherein, the distance refers to the continuous spacing of two low-entropy residues or corner residues, or the spacing of at least 3 hydrophilic residues, or the spacing of one amino acid residue P; the irrelevant residues refer to the other amino acid residues in the primary structure of the protein except those marked with &-&; the amino acid side chain region refers to the region composed of amino acid side chains with the number of amino acid residues greater than 5.

[0045] Furthermore, the connected segment of residues with high entropy increase potential described in S2032 is predicted to form a β-fold, and there are two consecutive low entropy residues aspartic acid D or asparagine N in the connected segment of residues with high entropy increase potential, then the two consecutive low entropy residues aspartic acid D or asparagine N are predicted to form a turn as a prediction result.

[0046] Furthermore, in the entropy increment statistics of the β-sheet described in S400, if the same amino acid residue in the primary structure of the protein realizes a preferred pairing between the side chains of two amino acid residues with a 1-3 sequence position relationship, the entropy increment of the amino acid residue is accumulated twice;

[0047] In the entropy increment statistics of α-helix, if the same amino acid residue in the primary structure of the protein realizes multiple preferred pairings between the side chains of amino acid residues with 1-4 / 5 sequence position relationships, the entropy increment of the amino acid residue is accumulated twice.

[0048] Furthermore, in the entropy increase statistics of the α-helix described in S400, if there is a preferred pairing between the side chains of amino acid residues with a 1-2 sequence position relationship in the primary structure of the protein, and the preferred pairing is I / V / L / MI / V / L / M, then the two pairs of I / V / L / M are recorded as one entropy increase.

[0049] Furthermore, the amino acid residues in the two consecutive amino acid residues in S600 are one or two of S, T, N, D, and G.

[0050] Furthermore, the relative measurement value of the entropy increase potential of the low entropy hydration layer outside the side chains of amino acid residues in the primary structure of the protein is a measurement indicator that uses the length of the hydrophobic portion on the side chains of amino acid residues in the primary structure of the protein as the relative measurement value.

[0051] Furthermore, the relative measure of the entropy increase potential of the low-entropy hydration layer on the periphery of the amino acid residue side in the primary structure of the protein is specifically to take the number of carbon-carbon bonds on the amino acid residue chain in the primary structure of the protein as 1 entropy value.

[0052] Furthermore, the entropy values of the 20 amino acids are as follows:

[0053] Serine S=0.5, threonine T=0.5, aspartic acid D=0.5, asparagine N=0.5, glycine G=0, histidine H=3, arginine R=3, glutamic acid E=3, lysine K=4, tryptophan W=5, isoleucine I=6, valine V=6, leucine L=6, phenylalanine F=6, proline P=0, cysteine C=4, glutamine Q=3, alanine A=1, methionine M=4, tyrosine Y=5.

[0054] Furthermore, the sufficient entropy increase described in step one means that when the side chains of two amino acid residues in the primary structure of the protein are similar, the two residue side chains are laterally attached, so that the low entropy hydration layer of the two side chains undergoes sufficient entropy increase; the similarity of the amino acid residue side chains means that the smaller the difference in the entropy values of the two hydrophilic amino acid residues, the more similar the two hydrophilic residues are; and the smaller the difference in the entropy values of the two hydrophobic amino acid residues, the more similar the two hydrophobic residues are.

[0055] Furthermore, in the process of determining the unblocked high entropy increase potential residue connected segment, first find the segment consisting of consecutive high entropy increase potential residues in the sequence, name the original "high entropy increase potential residue connected segment", then determine the non-high entropy increase potential residues that do not block the high entropy increase potential residue connected segment, and then determine the unblocked high entropy increase potential residue connected segment, specifically including:

[0056] (1) When there is only one amino acid residue S or T between two high entropy increase potential residues, the S or T residue participates in forming an unblocked high entropy increase potential residue connected segment;

[0057] (2) When there is only one isolated N, D or G residue between two high entropy increase potential residues in the amino acid sequence, and there is a high entropy pairing between I, V, L, F, and Y with a 1-3 or 1-4 / 5 relationship across the D, N or G, then the D, N or G is considered to participate in the formation of an unblocked high entropy increase potential residue connected segment;

[0058] (3) When any one of the amino acid residues S, T, N, D, and G appears in the amino acid sequence of the polypeptide chain and is adjacent to any one of the amino acid residues S, T, N, D, and G, the 1-4 / 5 sequence position relationship across the two adjacent amino acid residues is checked. When the two amino acid residues at the 1-4 / 5 position are both I, V, L, F, and Y residues, and at least one of the amino acid residues is I, V, or L, it is predicted that the two adjacent amino acid residues participate in forming an unblocked high entropy increase potential residue connected segment;

[0059] (4) When a P residue appears in the amino acid sequence of a polypeptide chain, and the P residue is not in a 1-2 or 1-3 sequence position relationship with a D or N residue, and there is a 1-4 / 5 sequence position relationship across the P residue, and the amino acid residues on 1-4 / 5 are I, V, L, or F, then the P residue is predicted to participate in the formation of an unblocked high entropy increase potential residue connected segment;

[0060] (5) When a G or A residue appears in the amino acid sequence of a polypeptide chain, if an amino acid residue that forms a 1-4 / 5 sequence position relationship with the residue can be found on one side of the sequence where the G or A residue is located, and an amino acid residue that forms a 1-4 / 5 sequence position relationship with the G or A residue can also be found on the other side of the sequence where the G or A residue is located, and if there is a single D, N, G or P residue that is not adjacent to other low-entropy residues or corner residues between the I, V, L residues, or there are two adjacent S, T, N, D residues, then it is predicted that these single D, N, G or P residues or two adjacent S, T, N, D, G residues participate in the formation of an unblocked high-entropy potential residue connected segment.

[0061] Furthermore, the length of the unblocked connected segment of residues with high entropy increase potential is greater than or equal to 3 amino acid residues.

[0062] Furthermore, if a fragment of 5 or more amino acid residues exists in a sequence whose primary structure is not predicted to be an α-helix or a β-turn, when the fragment is in the general initial thermodynamic metastable state of unfolded proteins, if there is no hydrophilic amino acid residue side chain in the fragment whose height of the hydrophilic atom minus the height of the adjacent hydrophobic residue side chain is less than or equal to 1, then the fragment is predicted to be a β-sheet; wherein the height of the residue side chain refers to the height of the residue side chain, which is the number of covalent bonds within the shortest path between the top atom of the amino acid residue side chain and the main chain carbon atom in the fragment.

[0063] A protein structure prediction method based on the mutual attraction between low-entropy hydration layers of residue side chains, wherein the prediction method utilizes water entropy force to achieve protein structure prediction; the prediction is based on the secondary structure of the protein predicted by the protein structure prediction method based on the mutual attraction between low-entropy hydration layers of residue side chains to predict the tertiary structure of the protein, including the following steps:

[0064] Definition of the similarity of amino acid residue side chains: the smaller the difference in the "entropy values" of the side chains of two hydrophilic residues, the more similar the two hydrophilic residues are; the smaller the difference in the "entropy values" of the side chains of two hydrophobic residues, the more similar the two hydrophobic residues are;

[0065] The amino acid residues that achieve preferred pairing between I / V / L / F in the predicted secondary structure are marked, and the hydrophobic surface of the hydrophobic side chain cluster that achieves lateral adhesion between the side chains of the I / V / L / F amino acid residues is marked and named "cluster hydrophobic surface". The protein structure is predicted to be the one where the hydrophobic surface of an I / V / L / F cluster and the hydrophobic surface of another adjacent I / V / L / F cluster are surface-adapted.

[0066] When there is no preferred pairing relationship between I / V / L / F in a β-sheet or α-helical secondary structure, the secondary structure is searched for a preferred pairing relationship between I / V / LY. If so, the hydrophobic surface of the hydrophobic side chain clusters that are laterally attached between the side chains of the I / V / LY amino acid residues are also marked as "clustered hydrophobic surface"; the surface attachment of an I / V / L / Y cluster hydrophobic surface to the adjacent hydrophobic surface of another cluster is predicted as a protein structure;

[0067] When an α-helical structure has two clustered hydrophobic surfaces and the two clustered hydrophobic surfaces are not connected in the axial direction of the α-helix, it is predicted that one of the clustered hydrophobic surfaces with a smaller number of I / V / L / F / Y will not align with the other I / V / LF / clustered hydrophobic surface in an adjacent β-sheet or α-helical secondary structure as a possible prediction result.

[0068] Furthermore, the protein structure prediction method based on the mutual attraction relationship between the low entropy hydration layers of the residue side chains further includes the following steps:

[0069] For the case where the current prediction is β-sheet, it is necessary to mark the strong hydrophobic residues of the 1-3 position relationship preferred pairing connected segments in the β-sheet, i.e. I / V / L / F / W / Y, and at the same time mark these hydrophobic segments connected by the 1-3 position relationship, compare the lengths of adjacent β-sheets in the sequence, and predict that α-helices of similar lengths will fold into a sheet structure. Then, based on the side chain distribution of the strong hydrophobic residues of the two β-sheets, the corresponding position relationship between the two is based on the most strong hydrophobic residues on one β-sheet, which can achieve preferred pairing with the side chains of the amino acid residues on the other α-helix. The state of the two β-sheets being attached is the predicted sheet folding structure.

[0070] The present invention has the following beneficial effects:

[0071] After research, the present inventors believe that only hydrophobic interactions between amino acid residue side chains are the core driving force of protein folding, and that hydrogen bonding, electrostatic forces, van der Waals forces, ionic bonding, and other interactions are byproducts of hydrophobic interactions, that is, products of entropy and enthalpy compensation. Furthermore, the present inventors propose a method for predicting protein structure based solely on the laws governing hydrophobic interactions between amino acid residue side chains. The present invention can precisely analyze the folding dynamics and laws governing protein structure, enabling accurate prediction of protein folding structures. Based on these laws, the present inventors provide a method for predicting protein secondary and tertiary structures. This method can aid scientists in designing new proteins and purposefully mutating existing protein structures for important biological applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] Figure 1 Schematic diagram of the amino acid residues of the low entropy hydration layer of 20 amino acids. The atoms in the dotted box are hydrophobic atoms;

[0073] Figure 2 It is a diagram of the lateral assembly state of two amino acid residues E and one amino acid residue K;

[0074] Figure 3 This is a diagram of the side chain alignment of adjacent amino acid residues in a typical strand structure.

[0075] Figure 4 A schematic diagram of the relationship between the side chain of an amino acid residue and the side chain of another amino acid residue separated by 2 or 3 amino acid residues, i.e., the 1-4 / 5 sequence position relationship;

[0076] Figure 5 Schematic diagram of the positional relationship of 1-2 sequences;

[0077] Figure 6 There are three carbon-carbon bonds in the side chain of amino acid residue E;

[0078] Figure 7 is the optimal entropy increase pairing graph;

[0079] Figure 8 A schematic diagram of the alignment of the hydrophobic end of an amino acid residue side chain and the hydrophobic portion of another amino acid residue side chain;

[0080] Figure 9 Schematic diagram of the hydrophobic surface of the cluster;

[0081] Figure 10 This is a schematic diagram of the folding structure of the sheet;

[0082] Figure 11 Schematic diagram of the configuration of the "cluster hydrophobic surfaces" fully fitting together;

[0083] Figure 12 Schematic diagram of helix structure;

[0084] Figure 13 Schematic diagram of the sheet structure (beta strand);

[0085] Figure 14 Schematic diagram of the corner structure (turn);

[0086] Figure 15 Schematic diagram of amorphous structure (random coil);

[0087] Figure 16 Schematic diagram of the protein structure in the surface adhesion state. DETAILED DESCRIPTION

[0088] The currently prevalent theory regarding entropy and enthalpy changes during protein folding holds that as proteins transition from an unfolded to a folded state, their conformational freedom decreases significantly, leading to a decrease in entropy. This suggests that protein folding is accompanied by a decrease in entropy. Furthermore, the prevailing theory holds that as protein folding begins, hydrogen bonds, hydrophobic interactions, and van der Waals forces gradually form within the molecule. These interactions release energy, leading to a decrease in enthalpy. Therefore, the prevailing theory holds that entropy gradually decreases during protein folding, and enthalpy decreases during the folding process. Research related to the present invention suggests that this is problematic because it fails to account for the entropy increase of the aqueous solvent in the system. Specifically, protein folding occurs only in an aqueous solvent environment; proteins cannot fold correctly in any other non-aqueous solvent environment. Therefore, the present invention proposes a novel theory: the protein folding process is a process in which both entropy and enthalpy increase. This new theory, the basis of the present invention, posits that protein folding is driven by the entropy increase of a low-entropy hydration layer surrounding the hydrophobic atoms on the side chains of polypeptide residues. That is, the hydrophobic interaction between the hydrophobic atoms of the side chains of different residues causes the entropy of the low-entropy hydration layer to increase to free solvent water, and this entropy increase process drives the folding of the protein. In addition, during the protein folding process, the polypeptide chain molecules need to first get rid of the strong hydrogen bonds, electrostatic forces and van der Waals forces formed between the surrounding strongly charged polar water molecules and the polypeptide chain before they can form the weaker hydrogen bonds, electrostatic forces and van der Waals forces within the protein molecules. Therefore, the protein folding process is an enthalpy increase process, not an enthalpy decrease process. In other words, the new concept believes that the protein folding process is a systematic entropy increase process that compensates for the enthalpy increase, which is usually achieved through a balance between entropy increase and enthalpy increase. Based on this theory, the present invention carried out research and carried out simulation verification, which can very accurately predict protein structure. Therefore, the method of the present invention is an original technical invention, which proposes a new theory of the protein structure formation process and proposes a prediction method based on the corresponding theory. The core technology of the present invention is that: the present invention discovered that the folding process of proteins is dominated by the hydrophobic interaction between the side chains of amino acid residues, that is, it is believed that the hydrogen bonding force, electrostatic and van der Waals forces formed within the molecule during the folding process of proteins do not dominate the folding process of proteins. The method described in the present invention ignores physical forces such as hydrogen bonding, electrostatic and van der Waals forces, and only prioritizes the analysis of the hydrophobic interaction between the side chains of amino acid residues to predict the structure of proteins, which is different from existing methods for predicting the structure of proteins. The method of the present invention believes that the protein folding process depends only on "water entropy force", that is, the hydrophobic interaction relationship between the low entropy hydration layers of the side chains of amino acid residues determines the precise folding rules of proteins, and the interaction attraction between the low entropy hydration layers of the side chains of amino acid residues is defined as "water entropy force" by the present invention.

[0089] The present invention believes that the folding process of protein is dominated by the mutual attraction between the low-entropy hydration layers of the side chains of amino acid residues. The carbon atoms and sulfur atoms in the side chains of amino acid residues are hydrophobic, which causes the water molecules around these hydrophobic atoms to be in a low-entropy state, that is, there is a low-entropy hydration layer around the hydrophobic atoms. The entropy increase trend of the low-entropy hydration layer will cause mutual attraction and collapse between such low-entropy hydration layers, and at the same time prevent the low-entropy hydration layer and the hydrophilic amino acid residues from getting close, thereby causing lateral attraction and lateral adhesion between specific amino acid residue side chains, becoming one of the core mechanical action mechanisms that dominate protein folding. In the folding process of protein, hydrophobic interaction occurs between the side chains of hydrophobic amino acid residues, that is, spontaneous aggregation between the side chains of hydrophobic amino acid residues drives the folding of protein. The present invention points out that there is also a low entropy hydration layer around the hydrophobic atoms in the side chains of hydrophilic amino acid residues, and the water molecules therein also have an entropy increase effect. This entropy increase effect will cause some hydrophilic amino acid residue side chains to laterally fit with adjacent hydrophilic amino acid residue side chains. This entropy increase effect is also an important driving force for protein folding. It should be emphasized that this is the starting point of the present invention. Figure 2 As shown,

[0090] Driven by the entropy increase of water molecules in the environment, some hydrophilic amino acid residue side chains will spontaneously align laterally with adjacent hydrophilic amino acid residue side chains (that is, the amino acid residue side chains are in a state parallel to each other). In this state, a state of aggregation of hydrophobic atoms in multiple hydrophilic amino acid residue side chains is formed, which means that the entropy increase of the low-entropy hydration layer is achieved. The present invention believes that the main physical mechanism of protein folding is that the entropy increase of low-entropy water molecules drives the lateral aligning between amino acid residue side chains (lateral aligning refers to the state where two amino acid residue side chains are in a state of approximately parallel). The method of the present invention mainly predicts which two amino acid residue side chains in an amino acid residue sequence will align laterally. This state is commonly found in the secondary structure strand and helix of proteins, and the relative position relationship of the sequence position of the amino acid residue side chains is used to represent the sequence position of the amino acid residue side chains. In other words, the entropy increase of the entropy hydration layer around the hydrophilic amino acid residue side chains causes the hydrophobic parts of some hydrophilic amino acid residue side chains to align with each other, and also drives the hydrophilic top atoms of the amino acid residue side chains to align with each other. This is a protein folding dynamic mechanism revealed for the first time by the present invention.

[0091] For example, in the secondary structure of a strand, the side chain of an amino acid residue is usually in an approximately parallel fit with the side chain of the adjacent amino acid residue. In the present invention, the sequence position relationship of the side chains of two amino acid residues separated by one amino acid residue in the amino acid residue sequence of the primary structure is named as a "1-3" sequence position relationship, such as Figure 3 shown.

[0092] Similarly, in the helix secondary structure, the side chains of two other amino acid residues separated by 2 and 3 amino acid residues in an amino acid residue side chain sequence are in an approximately parallel fit state, such as Figure 4 As shown, the present invention names the relationship between the sequence position of an amino acid residue side chain and the sequence position of a certain amino acid residue side chain in the primary structure of the amino acid residue sequence as a "1-4 / 5" sequence position relationship.

[0093] In addition, in the helix structure, the side chain of an amino acid residue is in an approximately parallel rotation state with the side chain of an adjacent amino acid residue in the sequence. The present invention names the sequence position relationship of two adjacent amino acid residues in the amino acid residue sequence of the primary structure as the "1-2" sequence position relationship, such as Figure 5 shown.

[0094] from Figure 1 It can be seen that the side chains of many hydrophilic amino acid residues also contain hydrophobic carbon atoms and sulfur atoms. Therefore, it is believed that these hydrophobic carbon atoms and sulfur atoms lead to the existence of a low entropy hydration layer around them, that is, there is also a low entropy hydration layer outside the side chains of many hydrophilic amino acid residues. In addition, since the roots of almost all amino acid residue side chains are composed of hydrophobic carbon atoms (such as Figure 1 As shown), it is believed that these hydrophobic carbon atoms obscure the hydrophilicity of the main chain structure of the polypeptide chain. Therefore, the present invention regards the main chain structure of the polypeptide chain as hydrophobic, that is, in the process of predicting protein structure, the hydrophilicity of the carbonyl oxygen and amide hydrogen groups of the main chain is ignored. Specifically, the water molecules around the polypeptide chain must pass through the root of the branch chain to approach the main chain. Since the root of the branch chain is hydrophobic, the water molecules around the hydrophobic root are low entropy, so most of the water molecules that have hydrophilic effects with the main chain of the polypeptide chain are these low entropy water molecules, which leads to the fact that the hydrophilicity of the main chain cannot be fully expressed, that is, there are not a large number of free water molecules that can frequently generate and break hydrogen bonds with the main chain. Therefore, when predicting protein folding, the hydrophilicity of the main chain is ignored. It should be pointed out that the side chains of amino acid residues such as E, Q, R, K, and H have relatively long hydrophobic roots, see Figure 1 Therefore, it is believed that there is also a low-entropy hydration layer outside the side chains of residues such as E, Q, R, K, and H.

[0095] In order to make the objectives, technical solutions and advantages of the embodiments of the present invention more clearly understood, the spirit of the contents disclosed in the present invention will be described in detail below. After understanding the embodiments of the contents of the present invention, any technician in the relevant technical field can change and modify the contents of the present invention based on the techniques taught by the contents of the present invention without departing from the spirit and scope of the contents of the present invention.

[0096] The exemplary embodiments and descriptions of the present invention are used to explain the present invention, but are not intended to limit the present invention. The present invention will be further described below in conjunction with specific embodiments. Specific implementation method one:

[0098] This embodiment is a protein structure prediction method based on the mutual attraction relationship between the low-entropy hydration layers of the side chains of amino acid residues. This embodiment mainly predicts the secondary structure of the spatial structure of the amino acid sequence of the protein. First, it predicts which fragments in the amino acid residue sequence of the primary structure of the protein (polypeptide chain) are folded into secondary structures such as α-helix (referred to as helix in this invention), β-fold (referred to as strand in this invention), β-turn (referred to as turn in this invention) and random coil (referred to as randomcoil in this invention), then predicts the part between the helix and strand structure: turn or randomcoil, and finally predicts the tertiary structure of the protein.

[0099] In the prediction method of this embodiment, the common abbreviations of amino acid residues are used to represent each amino acid: alanine (A); valine (V); leucine (L); isoleucine (I); proline (P); phenylalanine (F); tryptophan (W); methionine (M); glycine (G); serine (S); threonine (T); cysteine (C); tyrosine (Y); asparagine (N); glutamine (Q); aspartic acid (D); glutamic acid (E); histidine (H); lysine (K); arginine (R). Figure 1 . Figure 1 The molecular structures of 20 amino acid residues are shown, and the hydrophobic carbon atoms and sulfur atoms in the side chains of each amino acid residue are marked with wireframes.

[0100] The following is a detailed description of a protein folding dynamics mechanism disclosed by the present invention:

[0101] 1. The preliminary work for forecasting is as follows:

[0102] First, an entropy value is assigned to the side chains of amino acid residues in the primary structure of the protein. The entropy value is a relative measure of the entropy increase potential of the low-entropy hydration layer around the side chains of the amino acid residues. Specifically, the number of carbon-carbon bonds in the hydrophobic part of the side chains of the amino acid residues is used as a measurement indicator. The indicator is named "entropy value". The present invention defines the number of carbon-carbon bonds on the side chain of an amino acid residue as 1 entropy value. The "entropy value" refers to the relative measure of the entropy increase potential of the low-entropy hydration layer around the side chains of the amino acid residues. For example, the length of the hydrophobic part of the side chain of the amino acid E is 3, containing 3 carbon-carbon bonds. Therefore, the entropy value of the amino acid E residue is defined as 3, see Figure 6 .

[0103] The "entropy value" of each amino acid residue is marked below the amino acid residues on the polypeptide chain. Considering that amino acid residues I, V, L, and F are strongly hydrophobic residues, the present invention defines the entropy value of I, V, L, and F as 6. At the same time, considering the masking effect of hydrophilic atoms on hydrophobic atoms, the present invention modifies the "entropy value". Specifically, the entropy value of the amino acid residue side chain is defined as follows: S = 0.5, T = 0.5, D = 0.5, N = 0.5, G = 0, H = 3, R = 3, E = 3, K = 4, W = 5, I = 6, V = 6, L = 6, F = 6, P = 0, C = 4, Q = 3, A = 1, M = 4, Y = 5.

[0104] Then, the amino acid residues are classified as follows:

[0105] Residues with micro-entropy increase potential include: S, T, D, and N;

[0106] Hydrophilic residues include: H, R, K, E, Q, D, N, H, Y, W, S;

[0107] Common turn residues (referred to as turn residues) include: G, P;

[0108] Hydrophobic residues include: I, V, L, F, Y, W, A, M, C, H;

[0109] Residues with high entropy increase potential include: I, V, L, K, R, F, Y, W, A, M, C, E, Q, and H.

[0110] Among them, a low-entropy residue refers to a residue in which the number of carbon atoms in the side chain is less than or equal to 2, and the side chain contains at least one oxygen atom or hydrogen atom.

[0111] Next, determine the preferred pairing between amino acid residue side chains:

[0112] Protein folding realizes the lateral fitting of multiple pairs of amino acid residue side chains in its polypeptide sequence. The present invention finds that many of these fitting states fully realize the entropy increase potential of the low-entropy hydration layer of each two amino acid residue side chains in multiple pairs of amino acid residue side chains. In other words, it is believed that the lateral fitting of certain pairs of amino acid residue side chains will cause the low-entropy hydration layer to disappear, realizing that the low-entropy hydration layer of each of the two amino acid residue side chains fully realizes its entropy increase potential. Therefore, the present invention defines that the lateral fitting between multiple pairs of specific amino acid residue side chains will result in a sufficient entropy increase in the low-entropy hydration layer of the two amino acid residue side chains, and these amino acid residue pair relationships are defined as the "preferred entropy increase pairing" of the two amino acid residues, see Figure 7 .

[0113] The present invention believes that when the side chains of two amino acid residues are similar, the lateral adhesion of the side chains of the two amino acid residues will lead to a sufficient entropy increase in the low-entropy hydration layer outside the two side chains. The similarity between the side chains of amino acid residues is defined as follows: "When the difference in the "entropy value" of the side chains of two hydrophilic amino acid residues is smaller, the two hydrophilic residues are considered to be more similar. When the difference in the "entropy value" of the side chains of two hydrophobic amino acid residues is smaller, the two hydrophobic residues are considered to be more similar." For example, when the entropy values of two amino acid residues are both less than or equal to 4, and the difference in their entropy values is less than or equal to 1, the side chains of the two amino acid residues are considered to be similar. When the entropy values of two amino acid residues are both greater than or equal to 4, and the difference in their entropy values is less than or equal to 2, the side chains of the two amino acid residues are considered to be similar. For example, amino acid residue L is similar to amino acid residue W. For example, the side chain of an amino acid residue E is similar to the side chain of another amino acid residue Q. Because the entropy values of the two amino acid residues are the same and both are hydrophilic residue side chains, the lateral adhesion of the side chains of the two amino acid residues will lead to sufficient mutual attraction and collapse between the low entropy water layers of the side chains of the two amino acid residues, and will not cause the hydrophilic end of one amino acid residue side chain to adhere to the hydrophobic part of the other amino acid residue side chain, see Figure 8 .

[0114] Based on the similarity between amino acid residue side chains, the present invention defines a list of preferred entropy-increasing pairings between amino acid residue side chains as follows:

[0115] FF, QQ, KK, EE, TT, RR, AA, SS, MM;

[0116] MA, QE, TS, RE, KR, EK, FY, WY, WE, RW, RY, KY, HV, DE, NE, DQ, NQ;

[0117] QR / K, I / V / L / FI / V / L / F, I / V / LY, I / V / LK, I / V / LR, I / V / LA, I / V / LW, I / V / L / F / YW, I / V / LM, CM / Y, MI / V / L; A / TD / N. “ / ” represents the relationship of “or”.

[0118] For example, in the helix structure, the side chains of two amino acid residues separated by 2 and 3 amino acid residues in the sequence are in an approximately parallel fitting state. When the lateral fitting between the side chains of amino acid residues occurs between the side chains of amino acid residues in the "preferred entropy increasing pairing", it is believed that these amino acid residue side chains have achieved sufficient entropy increase in their low entropy hydration layer, that is, they have achieved "preferred entropy increasing pairing".

[0119] 2. Specific prediction of protein structure

[0120] Based on the sequence position of amino acid residue side chains, amino acid residue classification and entropy value, and the preferred pairing between amino acid residue side chains, the following process is used to predict protein structure;

[0121] S100. Mark the entropy value under each amino acid in the amino acid sequence of the primary structure of the protein whose structure needs to be predicted to form a table.

[0122] S200. In another line below the amino acid sequence of the protein's primary structure, repeatedly write the amino acids of these "high entropy increase potential residues" below each high entropy amino acid residue. This serves as a segment consisting of adjacent connected segments of high entropy increase potential residues in the protein's primary structure amino acid sequence, which is the original "high entropy increase potential residue connected segment." Additionally, in a separate line, repeatedly write the amino acids below the corresponding positions of the low entropy residues and turn residues in the protein's primary structure amino acid sequence.

[0123] In the protein structure prediction method described in the present invention, it is first predicted which fragments of the amino acid sequence will fold into the typical strand structure and helix structure in the protein secondary structure: that is, it is first predicted which sequence fragments in the amino acid sequence of the protein primary structure will fold into the typical strand secondary structure or helix secondary structure, and then further predicted whether the fragment will fold into the strand structure or the helix structure.

[0124] The amino acid sequence for predicting the primary structure of a protein described in the present invention is folded into a strand secondary structure or a helix secondary structure, which is achieved based on an "unblocked connected segment of residues with high entropy increase potential". The definition of "unblocked connected segment of residues with high entropy increase potential" in the present invention is: two adjacent high entropy increase potential residues in the amino acid sequence of the primary structure of a protein and a non-high entropy increase potential residue that does not block the connection between the high entropy increase potential residues. Then this segment of the amino acid residue sequence is the "unblocked connected segment of residues with high entropy increase potential", and the length of the "unblocked connected segment of residues with high entropy increase potential" is greater than or equal to 3 amino acid residues.

[0125] The method for determining the “unblocked connected fragment of residues with high entropy increase potential” is as follows:

[0126] First, find the segment consisting of consecutive high entropy increase potential residues in the sequence, which is named the original "high entropy increase potential residue connected segment"; then determine the non-high entropy increase potential residues that do not block the high entropy increase potential residue connected segment:

[0127] (1) When there is only one amino acid residue S or T between two high entropy increase potential residues, it is considered that the S or T does not block the high entropy increase potential residue connected segment, that is, it is considered that the S or T amino acid residue does not block the high entropy increase potential residue connected segment, and participates in the formation of a continuous high entropy increase potential residue segment. The formation of a continuous high entropy increase potential residue segment means participating in the high entropy increase potential residue connected segment.

[0128] (2) When there is only one isolated N, D or G amino acid residue between two high entropy increase potential residues in an amino acid sequence, and there is a high entropy pairing between I, V, L, F, and Y with a 1-3 or 1-4 / 5 relationship across the D, N or G (at least one of the paired amino acid residues is I, V or L), it is considered that the D, N or G does not block the connected segment of high entropy increase potential residues, that is, it is considered that the N, D or G constitutes a continuous segment of high entropy increase potential residues.

[0129] If the D or N is predicted to be a helix by subsequent methods, and if the amino acid residue side entropy value with a 1-4 / 5 relationship with the D or N in the helix structure is greater than or equal to 4, the D or N will be predicted to be a turn structure.

[0130] (3) When any one of the amino acid residues S, T, N, D, and G appears in the amino acid sequence of the polypeptide chain and is adjacent to any one of the amino acid residues S, T, N, D, and G, the 1-4 / 5 sequence position relationship across the two adjacent amino acid residues is checked. When the two amino acid residues at the 1-4 / 5 position are both I, V, L, F, and Y, and at least one of the amino acid residues is I, V, or L, it is predicted that the two adjacent amino acid residues do not block the formation of a high entropy increase potential residue connected segment. If the high entropy increase potential residue connected segment where the two adjacent amino acid residues are located is not predicted as a helix by the following method, the two consecutive adjacent amino acid residues S, T, N, D, and G are predicted to be a turn structure T as a prediction result.

[0131] (4) When a P amino acid residue appears in the amino acid sequence of a polypeptide chain, and the P amino acid residue is not in a 1-2 or 1-3 sequence position relationship with a D or N amino acid residue (e.g., PQD, PCCN), and there is a 1-4 / 5 sequence position relationship across the P amino acid residue, and the amino acid residues on 1-4 / 5 are I, V, L, or F, then the predicted P amino acid residue does not block the high entropy increase potential residue-connected fragment. However, if the fragment containing P is ultimately not predicted to be a helix, then the P is predicted to be a turn structure.

[0132] (5) When a G or A amino acid residue appears in the amino acid sequence of a polypeptide chain, if an amino acid residue that forms a 1-4 / 5 sequence position relationship with the amino acid residue can be found on one side of the sequence where the G or A amino acid residue is located, and an amino acid residue that forms a 1-4 / 5 sequence position relationship with the G or A amino acid residue can also be found on the other side of the sequence where the G or A amino acid residue is located, and if there is a single D, N, G or P amino acid residue that is not adjacent to other low-entropy residues or corner residues, or two adjacent S, T, N, D amino acid residues between the above-mentioned I, V, L residues, then it is predicted that these single D, N, G or P amino acid residues or two adjacent S, T, N, D, G amino acid residues will not block the high entropy increase potential residues in the fragment from connecting the fragment.

[0133] For example: IVL 1-4 / 5G / A1-4 / 5IVLF (for example, VAAQGRARL), IVL 1-4 / 5A1-4 / 5A1-4 / 5IVLF, IVL 1-4 / 5A1-4 / 5A 1-4 / 5A (for example, IMQDAGVTANTRA).

[0134] S300. Find a fragment in the amino acid sequence of a polypeptide chain that is less than or equal to 5 amino acid residues in length and contains 3 or more microentropy residues or turn residues. Predict that the fragment will fold into a turn structure. Do not make this prediction for the ends of the amino acid sequence of the polypeptide chain. Find a fragment in the amino acid sequence of a polypeptide chain that is less than or equal to 5 amino acid residues in length and contains 3 or more microentropy residues or turn residues. If the fragment contains at least one turn residue, that is, a G or P amino acid residue, predict that these amino acid residues will fold into a turn structure.

[0135] Based on the above judgment, the predicted high entropy increase potential residue connected segments are marked below the original amino acid sequence;

[0136] S400: Predict whether the connected segments of high entropy increase potential residues in the amino acid sequence of a polypeptide chain fold into a helix structure or a strand structure:

[0137] S401, assuming that the connected segments of residues with high entropy increase potential in the amino acid sequence of the polypeptide chain are helix or strand structures;

[0138] (A) Assuming that a certain connected segment of residues with high entropy increase potential folds into a helix structure, mark whether the side chain of each amino acid residue in the segment achieves a preferred pairing relationship when folded into a helix. If the side chain of an amino acid residue in the connected segment of residues with high entropy increase potential achieves preferred pairing, mark it with * below the amino acid, otherwise do not mark it. The preferred pairing here refers to the preferred pairing achieved by the "1-4 / 5" sequence position relationship. In addition, if there is a preferred pairing of I / V / L and I / V / L achieved by a 1-2 sequence position relationship in the connected segment of residues with high entropy increase potential (for example, EAQRLL), or if there is an I / V / L 1-4 / 5G / A 1-4 / 5I / V / L / F series position relationship in the connected segment of residues with high entropy increase potential, mark * below these IVFGA as well.

[0139] (B) Assuming that a connected segment of residues with high entropy increase potential is predicted to be a strand structure, mark whether the side chain of each amino acid residue achieves a preferred pairing of I / V / L / F and I / V / L / F in a 1-3 sequence position relationship (e.g., IDV) in the strand state. The two I / V / L / F amino acid residues that achieve the preferred pairing relationship of I / V / L / F and I / V / L / F in the 1-3 sequence position relationship are marked with & below.

[0140] In addition, when a pair of I / V / L and I / V / L with a 1-3 sequence position relationship is marked with &, it is called marked with &-&, and these two I / V / L can find I / V / L / F / Y or A with a 1-4 / 5 sequence position relationship in the connected segment of residues with high entropy increase potential to form a preferred pair, then the &-& marked with these two I / V / L is considered invalid, that is, this &-& is no longer marked.

[0141] S402, perform the following comparative analysis on the * and & marked below the amino acid residue sequence:

[0142] When none of the amino acid residues in a high entropy increase potential residue connected segment are marked with &-&, and more than 80% of the amino acid residues are marked with *, the segment is predicted to be a helix;

[0143] When a connected segment of a high entropy-increasing residue segment is marked with &-& (excluding I / V / LSI / V / L marked with &-&, such as VSV), the segment is usually predicted to be a strand. Exceptionally, in a region of a high entropy-increasing residue segment marked "away" from the segment (where "away" is defined as being separated by two consecutive low entropy residues or turn residues, or separated by at least three hydrophilic residues, or separated by one amino acid residue P), a subsegment of length greater than 5 can be found within the connected segment of a high entropy-increasing residue segment, where more than 85% of the amino acid residues are marked with * and there is a 1-4 / 5 sequence position relationship between I / V / L / F / Y and I / V / L / F / Y for the amino acid residues other than those marked with &-&. In this case, the subsegment is predicted to fold into a helix structure, and the rest of the connected segment of the high entropy-increasing residue segment is predicted to be a strand. When the number of “connected segments with high entropy increase potential residues” marked with * is less than 50% of the number of amino acid residues in the segment, the segment is predicted to be a strand.

[0144] In a segment predicted to be a strand, if two low-entropy residues D or N appear consecutively in the amino acid residue sequence, then these two consecutive low-entropy residues D or N are predicted to be a turn structure as a prediction result.

[0145] S403: Search for amino acid residues immediately adjacent to the amino acid residue sequence fragment predicted to be a helix, and determine whether the side chains of the searched amino acid residues have a preferred pairing relationship of 1-4 / 5 with the amino acid residues in the predicted helix structure. If a preferred pairing relationship exists, the residues with the preferred pairing are predicted to fold into a helix structure. Repeat this operation until the first adjacent amino acid residue is found that does not have a preferred pairing relationship of 1-4 / 5 with the amino acid residue in the originally predicted helix structure.

[0146] S404. Calculate the entropy gain achieved by each amino acid residue in the helix and strand states (two activation scenarios).

[0147] The calculation method of strand entropy value is as follows: when a preferred pairing appears in the 1-3 sequence position relationship in the amino acid sequence of the polypeptide chain, the entropy value of the preferred paired amino acid residue is marked below the amino acid residue as the entropy value of the amino acid residue. When an amino acid residue has two preferred pairs in the 1-3 sequence position relationship, the entropy value of the amino acid residue is accumulated twice.

[0148] The calculation method for helix entropy increment is as follows: when a preferred pairing occurs in the 1-4 / 5 sequence position relationship in the sequence, the entropy value of the amino acid residue with the preferred pairing is marked below the amino acid residue as the entropy increment of the amino acid residue. When an amino acid residue achieves multiple preferred pairings in the 1-4 / 5 sequence position relationship, the entropy increment of the amino acid residue is accumulated twice. In the helix entropy increment statistics, when a certain I / V / L / M and another I / V / L / M are adjacent in the sequence, that is, a "1-2" sequence position relationship occurs, then these two I / V / L / Ms are considered to have achieved one entropy increment, and the entropy increment is accumulated once.

[0149] The entropy gains achieved for each amino acid residue side chain in the two scenarios for folding into a helix and a strand are compared. If the larger of the two calculated entropy gains is 20% greater than the smaller, the one with the larger overall entropy gain is selected as the predicted outcome. If the difference in the overall entropy between the predicted strand and helix does not exceed 20% of the smaller of the two, the fragment is predicted to be a strand or a helix as two possible outcomes. When calculating the entropy of the strand, I / V / LGI / V / L is considered to achieve the entropy-optimal pairing of two IVL residues, where "-" represents an arbitrary residue.

[0150] S405. When a fragment of 5 residues or more in length is in a strand state, and there is no situation where the hydrophilic atoms (nitrogen and oxygen atoms) of an amino acid residue side chain and the hydrophobic atoms of the adjacent amino acid residue side chain are at the same height (the height of the residue side chain) (in this method, amino acid residues S and T are regarded as hydrophobic, that is, the oxygen atoms at the top of the side chains of S and T amino acid residues are regarded as hydrophobic atoms), then the fragment is finally predicted to be a strand. When the hydrophilic atoms (nitrogen and oxygen atoms) of an amino acid residue side chain and the hydrophobic atoms of the adjacent amino acid residue side chain are at the same height, hydrophilic and hydrophobic repulsion will occur between the hydrophilic atoms and the hydrophobic atoms, and the strand state will be destroyed. The height of the residue side chain refers to the number of covalent bonds in the shortest path between the top atom of the amino acid residue side chain and the carbon atom of its main chain in the fragment.

[0151] When the predicted length of a strand is greater than or equal to 12 amino acid residues, a search is performed to determine whether there are three consecutive hydrophilic residues in the middle of the strand. If so, the three amino acid residues are predicted to form a turn structure.

[0152] S406. When a certain amino acid residue segment in the amino acid sequence of a polypeptide chain is predicted to be a helix, the IVLF groups of connected segments in the helix structure that are laterally attached to each other through a 1-4 / 5 sequence position relationship or a 1-2 sequence position relationship are marked, and the segment sequence is searched to see whether there is an isolated I\V\LF amino acid residue that is not laterally attached to any I\V\LF residue side chain in the group of IVLF amino acid residues. If there is, and the I\V\LF amino acid residue segment is adjacent to the S\T\N\D amino acid residue segment or is adjacent to each other with only one amino acid residue in between, then the segment containing I\V\L\F and S\T\N\D is predicted to no longer be a helix.

[0153] When the length of the predicted helix structure is less than or equal to 6 amino acid residues, predicting the fragment as a randomcoil is also a possible prediction result.

[0154] S500. Two consecutive amino acid residues in a connected segment of high entropy increase potential residues that are not predicted as helix are predicted as a turn structure, marked as T; the two consecutive amino acid residues are any one or two of S, T, N, D, and G, such as SS or ST.

[0155] A segment of four or more consecutive hydrophilic amino acid residues in the primary structure of a protein is designated a "continuous hydrophilic residue segment," for example, EEQQK. When a segment greater than seven amino acid residues exists within a connected segment of high entropy-increasing residues that is not predicted to be a helix, and the proportion of residues in the segment containing low entropy-increasing residues and / or turn residues and / or amino acid residue A and / or continuous hydrophilic residues is greater than or equal to 50%, the segment is predicted to fold into a random coil structure.

[0156] S600: Searching for a segment consisting of five or more consecutive hydrophilic residues in the amino acid sequence of the polypeptide chain; if the segment has been predicted as a segment connected by residues with high entropy increase potential, predicting the segment consisting of five or more consecutive hydrophilic residues as a random coil as a prediction result.

[0157] Based on the above judgment, the fragments predicted to have turn structures are marked below the original amino acid sequence.

[0158] In addition, the hydrophilic residues with entropy values greater than or equal to 3 in the current turn structure and random coil structure as a prediction result are marked, and the preferred pairing relationship between these hydrophilic residues is marked, and the lateral fitting state of the amino acid residue side chains with preferred pairing relationship is used as a prediction result; when the turn structure as a prediction result cannot be determined to be the turn structure, the micro-entropy residues in the predicted turn structure are marked, and the preferred pairing relationship between these micro-entropy residues is marked, and the lateral fitting state of the amino acid residue side chains with preferred pairing relationship is used as a prediction result.

[0159] At the same time, the hydrophobic residues I / V / L / F / M in the turn structure and random coil structure are marked, and the side chains of the hydrophobic residues will be aligned with the side chains of the adjacent I / V / L / F / Y / W / K amino acid residues as a prediction result; the I / V / L / F amino acid residues in the two secondary structures linked by the turn structure and the random coil structure are marked, and the alignment of I / V / L / F in the two secondary structures is used as a prediction result.

[0160] After the above determination, if a fragment of 5 or more amino acid residues exists in the protein's primary structure, and if the fragment is in the universal initial thermodynamic metastable state of unfolded proteins and there is no hydrophilic amino acid residue whose side chain height minus the height of the adjacent hydrophobic residue is less than or equal to 1, then the fragment is predicted to be a strand. The universal initial thermodynamic metastable state of unfolded proteins can be determined according to Yang Lin et al.'s "Universal Initial Thermodynamic Metastable State of Unfolded Proteins." Specific implementation method 2:

[0162] This embodiment is a protein structure prediction method based on the mutual attraction relationship between the low-entropy hydration layers of the side chains of amino acid residues. This embodiment predicts the spatial tertiary structure of the protein on the basis of the prediction of the spatial structure of the amino acid sequence of the protein in the specific embodiment one, that is, the secondary structure is first predicted using the specific embodiment one, and then the tertiary structure is further predicted based on the secondary structure.

[0163] First, the similarity of amino acid residue side chains is defined: the smaller the difference in the "entropy values" of the side chains of two hydrophilic residues, the more similar the two hydrophilic residues are; the smaller the difference in the "entropy values" of the side chains of two hydrophobic residues, the more similar the two hydrophobic residues are.

[0164] The amino acid residues that achieve the preferred pairing between I / V / L / F in the predicted secondary structure are marked, and the hydrophobic surface of the hydrophobic side chain cluster that achieves lateral fit between the side chains of the I / V / L / F amino acid residues is marked and named "cluster hydrophobic surface", such as Figure 9 As shown, it is predicted that the surface adhesion between the hydrophobic surface of one I / V / L / F cluster and the hydrophobic surface of another adjacent I / V / L / F cluster is a protein structure;

[0165] When there is no preferred pairing relationship between I / V / L / F in a strand or helix secondary structure, the secondary structure is searched for a preferred pairing relationship between I / V / LY. If so, the hydrophobic surface of the hydrophobic side chain clusters that are laterally attached between the side chains of the I / V / LY amino acid residues are also marked as "cluster hydrophobic surface"; the surface attachment of an I / V / L / Y cluster hydrophobic surface to the adjacent cluster hydrophobic surface is predicted to form a protein structure.

[0166] When a helix structure has two cluster hydrophobic surfaces and the two cluster hydrophobic surfaces are not connected in the axial direction of the helix, it is predicted as a possible prediction result that one of the cluster hydrophobic surfaces with a smaller number of I / V / L / F / Y will not fit together with the other I / V / LF / cluster hydrophobic surface in an adjacent strand or helix secondary structure.

[0167] For the current prediction of a strand, it is necessary to mark the strong hydrophobic residues (I / V / L / F / W / Y) of the preferred paired connected segments in the 1-3 position relationship in the strand, and at the same time mark these hydrophobic segments connected by the 1-3 position relationship, compare the lengths of adjacent strands in the sequence, and predict that strands of similar lengths will fold into a sheet structure. Then, based on the side chain distribution of the strong hydrophobic residues of the two strands, the most strong hydrophobic residues (I / V / L / F / W / Y) on one strand in the corresponding position relationship between the two strands are laterally aligned in a way that can achieve preferred pairing with the side chains of the amino acid residues on the other strand. The state of the two strands being aligned is the predicted sheet folding structure (in fact, β sheet is composed of β strands), such as Figure 10 shown.

[0168] In the process of predicting the protein structure where the cluster hydrophobic surface undergoes surface adhesion, we can actually start with a prediction based on a secondary structure predicted to be a turn structure. For the two predicted secondary structures connected to the turn structure, the hydrophobic surface of the hydrophobic side chain cluster is marked and named "cluster hydrophobic surface". The relative position of the two secondary structures connected to the turn structure is used to determine the local tertiary structure of the successfully predicted secondary structure. At the same time, the adhesion state of the two secondary structures and the adhesion state of the "cluster hydrophobic surfaces" of other secondary structures must also be considered to ensure that the "cluster hydrophobic surfaces" are fully adhered to each other. The corresponding configuration of the fully adhered state is the prediction result, such as Figure 11 shown.

[0169] Example 1

[0170] For ease of description, in the examples, amino acid residues folding into a helix structure are labeled H, amino acid residues folding into a strand structure are labeled E, turn structures are labeled T, and amorphous structures (random coils) are labeled C. It should be noted that the present invention includes but is not limited to these labeling methods, and other labeling methods may also be used in other embodiments.

[0171] The method of gradually predicting the tertiary structure of a protein is achieved by predicting four secondary structures:

[0172] 1. The four secondary structures are as follows:

[0173] ① Helix (see Figure 12 );

[0174] ② Beta strand structure (see Figure 13 );

[0175] ③ Turn structure (see Figure 14 );

[0176] ④Amorphous structure (random coil) (see Figure 15 ).

[0177] 2. Preprocessing steps:

[0178] The amino acid residue segments predicted to be helix are marked as H, the amino acid residue segments predicted to be strand are marked as E, and the amino acid residue segments predicted to be turn structures are marked as T.

[0179] This example uses the protein with sequence number 1054 in the PDB database to predict its structure:

[0180] 1. Mark the entropy value under each amino acid in the protein amino acid sequence whose structure needs to be predicted to form a table, where a represents an entropy value of 0.5. As shown below:

[0181]

[0182] 2. On a new line below the protein amino acid sequence, repeat these "high entropy increase potential residues" below each high entropy increase potential residue. This marks the segments consisting of adjacent connected segments of high entropy increase potential residues in the sequence, which are the original "high entropy increase potential residue connected segments". In addition, on a new line below the protein amino acid sequence, repeat these low entropy residues and corner residues below the corresponding positions of the amino acid residue sequence, as follows:

[0183]

[0184] 3. Find the segments in the sequence that are composed of consecutive high entropy increase potential residues, which are named the original "high entropy increase potential residue connected segment". In addition, when there is only one amino acid residue S or T between two high entropy increase potential residues, it is considered that the S or T does not block the high entropy increase potential residue connected segment, and the high entropy increase potential residue connected segment still exists; that is, the S or T amino acid residue is considered to participate in the formation of a continuous high entropy increase potential residue segment.

[0185] After removing these S and T residues, the remaining low-entropy residues and corner residues that may block the connection between the high-entropy-increasing potential residues are marked in the fourth row of the table below.

[0186]

[0187] 4. When there is only one isolated N, D or G amino acid residue between two high entropy increase potential residues in the amino acid sequence, and there is a high entropy pairing between I / V / L / F / Y with a 1-3 or 1-4 / 5 sequence position relationship across the D, N or G (at least one of the paired amino acid residues is I / V / L), then it is considered that the D, N or G does not block the high entropy increase potential residue connected segment, and the high entropy increase potential residue connected segment still exists; that is, it is considered that the N, D or G constitutes a continuous high entropy increase potential residue segment. The remaining low entropy residues and corner residues that may block the high entropy increase potential residue connected segment are marked in the last row (add one row below the above table, which is the last row when processing this step).

[0188] When two adjacent S / T / N / D / G residues appear in an amino acid sequence, if there is a 1-4 / 5 relationship between the I / V / L / F amino acid residues across these two amino acid residues, then it is predicted that the two adjacent S / T / N / D / G residues do not block the formation of a connected segment of residues with high entropy increase potential. If the connected segment of residues with high entropy increase potential where the two adjacent S / T / N / D / G residues are located is not predicted as a helix by the following method, then the two consecutive S / T / N / D / G residues are predicted as a turn structure T as a prediction result.

[0189] When a P amino acid residue appears in the amino acid sequence, and the P is not in a 1-2 or 1-3 sequence position relationship with D / N (e.g., PQD, PCCN), and there is a 1-4 / 5 sequence position relationship between the I / V / L / F amino acid residues across the P, then the predicted P amino acid residue does not block the high entropy increase potential residue-connected fragment, but if the fragment containing P is ultimately not predicted to be a helix, then the P is predicted to be a turn structure.

[0190] The remaining low entropy residues and corner residues that may block the connection segment of high entropy potential residues are marked in the last row:

[0191]

[0192] 5. When the amino acid residue position relationship of IVL 1-4 / 5 G / A1-4 / 5 IVLF (for example, VAAQGRARL), IVL1-4 / 5 A1-4 / 5 A1-4 / 5 IVLF (for example, IMQDAGVTANTRV), and IVL 1-4 / 5 A1-4 / 5 A 1-4 / 5 A (for example, IMQDAGVTANTRA) appears in the amino acid sequence, if a single D, N, G, P or two adjacent S / T / N / D amino acid residues exists in the fragment, it is predicted that these single D, N, G or P two adjacent S / T / N / D / G amino acid residues will not block the high entropy increase potential residues in the fragment from connecting the fragment.

[0193] When a fragment with the amino acid residue position relationship of IVL 1-4 / 5 G / A 1-4 / 5 IVLF (for example, VAAQGRARL), IVL 1-4 / 5A1-4 / 5 A 1-4 / 5 IVLF (for example, ), IVL 1-4 / 5 A 1-4 / 5 A 1-4 / 5 A appears in the amino acid sequence, if there is a single D, N, G or P amino acid residue or two adjacent S / T / N / D amino acid residues that are not adjacent to other low entropy residues or turn residues in the fragment, then it is predicted that these single D, N, G or P amino acid residues or two adjacent S / T / N / D / G amino acid residues do not block the high entropy increase potential residues in the fragment from connecting the fragment:

[0194]

[0195] 6. According to S300: Find a fragment in the sequence that is less than or equal to 5 amino acid residues in length and contains 3 or more microentropy residues or corner residues, and predict that the fragment folds into a turn structure. If there are two adjacent microentropy residues or corner residues in the fragment and they are not adjacent to another microentropy residue or corner residue, then accurately predict that the two adjacent microentropy residues or corner residues fold into a turn structure. For specific prediction methods of the turn structure, see Sections 30, 31, and 32. Find a fragment in the sequence that is less than or equal to 5 amino acid residues in length and contains 3 or more microentropy residues or corner residues. If the fragment contains at least one turn residue, that is, a G or P amino acid residue, then predict that these amino acid residues fold into a turn structure T.

[0196] Based on this, the predicted turn structure of the protein is marked in the third row:

[0197]

[0198] 7. In summary, the fragments that are initially predicted to have corner structures are marked in the third row of the table below:

[0199]

[0200] 8. The corresponding predicted microentropy residue connected segments are as follows:

[0201]

[0202] 9. The method for predicting whether a connected segment of residues with high entropy increase potential in an amino acid residue sequence is folded into a helix structure or a strand structure is as follows: First, assume that a connected segment of residues with high entropy increase potential is folded into a helix structure, and mark whether the side chain of each amino acid residue in the segment has achieved a preferred pairing relationship when folded into a helix. If an amino acid residue side chain has achieved preferred pairing, it is marked as * below the amino acid, otherwise it is not marked. The preferred pairing here refers to the preferred pairing achieved through the "1-4 / 5" sequence position relationship. In addition, if there is a preferred pairing of IVL and IVL achieved by a 1-2 sequence position relationship in the sequence (for example, EAQRLL), if there is an IVL 1-4 / 5 G / A1-4 / 5IVLF series position relationship in the sequence, * is also marked below these IVFGA.

[0203]

[0204] 10. Assuming that a connected segment of residues with high entropy increase potential is predicted to be a strand structure, mark whether the side chain of each amino acid residue achieves a preferred pairing of IVLF and IVLF at a 1-3 position relationship (e.g., IDV) when folded into a strand state. The two IVLF amino acid residues that achieve the preferred pairing of IVLF and IVLF at the 1-3 position relationship are marked with & below. In addition, when a pair of IVL and IVL at the 1-3 position relationship are marked with & (called marked with &-&), and both IVLs can find IVLFY or A at the 1-4 / 5 relationship sequence position relationship in the connected segment of residues with high entropy increase potential to form a preferred pairing, the &-& marked on the two IVLs is considered invalid, that is, the &-& is no longer marked.

[0205]

[0206] 11. Compare the * and & in the first two tables, and perform the following comparative analysis on the * and & marked below the amino acid residue sequence: When no amino acid residues in a high entropy increase potential residue connected segment are marked with &-&, and more than 80% of the residues are marked with *, the segment is predicted to be a helix.

[0207]

[0208] 12. When a segment of a high entropy-increasing potential connected segment is marked with &-& (excluding I / V / LSI / V / L marked with &-&, such as the sequence VSV), the segment is generally predicted to be a strand. Exceptionally, if a subsegment of a high entropy-increasing potential connected segment is found in the region marked "away" from the segment (where "away" is defined as being separated by two consecutive low entropy residues or turn residues, or separated by at least three hydrophilic residues) and contains more than 85% of the amino acid residues marked with *, and a 1-4 / 5 sequence position relationship exists between IVLFY and IVLFY that is unrelated to the amino acid residues marked with &-&, then the subsegment is predicted to fold into a helix structure, and the rest of the high entropy-increasing potential connected segment is predicted to be a strand. When the number of residues marked with * in a "high entropy-increasing potential connected segment" is less than 50% of the total number of amino acid residues in the segment, the segment is predicted to be a strand.

[0209]

[0210] 13. When a segment of a high entropy-increasing potential connected segment is marked with &-& (excluding I / V / LSI / V / L marked with &-&, such as VSV), the segment is generally predicted to be a strand. Exceptionally, if a subsegment of a high entropy-increasing potential connected segment is found in the region marked "away" from the segment (where "away" is defined as being separated by two consecutive low entropy residues or turn residues, or separated by at least three hydrophilic residues) and contains more than 85% of the amino acid residues marked with *, and a 1-4 / 5 sequence relationship exists between IVLFY and IVLFY that is unrelated to the amino acid residues marked with &-&, then the subsegment is predicted to fold into a helix structure, and the rest of the high entropy-increasing potential connected segment is predicted to be a strand. When the number of *-marked segments within a "high entropy-increasing potential connected segment" is less than 50% of the total number of amino acid residues in the segment, the segment is predicted to be a strand.

[0211]

[0212] 14. If the connected segment of the high entropy increase potential residues where the two adjacent S / T / N / D / G are located is not predicted as a helix by the following method, then the two consecutive S / T / N / D / G are predicted to be a turn structure T as a prediction result.

[0213]

[0214] 15. Search the amino acid residues adjacent to the predicted helix amino acid residue sequence fragment to determine whether the side chains of these amino acid residues have a 1-4 / 5 relationship of preferred pairing with the amino acid residues in the predicted helix structure. If so, the amino acid residues are predicted to fold into the helix structure. Repeat this process until the first adjacent amino acid residue is found that does not have a 1-4 / 5 relationship of preferred pairing with the amino acid residues in the predicted helix structure.

[0215]

[0216] 16. The prediction results are listed in the third row of the table below, and the results of the software's analysis of the true structure of the protein are predicted in the fourth row. It can be seen that the prediction is accurate.

[0217]

[0218] 17. When the turn structure cannot be predicted, mark the micro-entropy residues in the predicted turn structure, and then mark the preferred pairing relationship between these micro-entropy residues. The lateral fitting state of these amino acid residues with preferred pairing relationship is a prediction result.

[0219]

[0220] 18. Mark the amino acid residues that achieve optimal alignment between IVLF in the predicted secondary structure:

[0221]

[0222] The hydrophobic surface of the hydrophobic side chain cluster that achieves lateral adhesion between the side chains of IVLF amino acid residues is marked and named as "cluster hydrophobic surface". It is predicted that the state in which the hydrophobic surface of an IVLF cluster and the hydrophobic surface of another adjacent IVLF cluster in the amino acid residue sequence are adhered is a protein structure. When there is no preferred pairing relationship between IVLFs in a secondary structure, it is searched whether there is a preferred pairing relationship between IVLYs in the secondary structure, and the hydrophobic surface of the hydrophobic side chain cluster that is laterally adhered between the side chains of IVLY amino acid residues is marked as "cluster hydrophobic surface", and the state in which the hydrophobic surface of an IVLY cluster and the hydrophobic surface of another adjacent cluster in the amino acid residue sequence are adhered is a protein structure. When there are two cluster hydrophobic surfaces in the manifestation of a helix structure, and the two cluster hydrophobic surfaces do not have a connecting segment in the axial direction of the helix, it is predicted that one of the clusters with a smaller number of IVLFYs will not adhere to the hydrophobic surface of another adjacent IVLF cluster as a possible prediction result, such as Figure 16 shown.

[0223] The above examples are merely illustrative of the calculation model and process of the present invention and are not intended to limit the embodiments of the present invention. Persons skilled in the art will readily appreciate that other variations or modifications based on the above description are possible. This list of embodiments is not exhaustive; however, any obvious variations or modifications derived from the technical solution of the present invention remain within the scope of protection of the present invention.

Claims

1. A protein structure prediction method based on the mutual attraction between low-entropy hydration layers of residue side chains, characterized by: The prediction method is to use water entropy force to predict protein structure; the prediction is to predict the secondary structure of the protein based on the primary structure of the protein; the water entropy force refers to the interaction attraction between the low entropy hydration layers of the amino residue side chains; The process of using water entropy force to predict protein structure includes:

1. Preliminary preparation for forecasting: First, an entropy value is assigned to the side chains of amino acid residues in the primary structure of the protein. The entropy value is a relative measure of the entropy increase potential of the low-entropy hydration layer around the side chains of the amino acid residues. Then, the amino acid residue side chains in the primary structure of the protein after entropy assignment are divided into five residue forms: Residues with micro-entropy increase potential: serine S, threonine T, aspartic acid D, asparagine N; Hydrophilic residues: histidine H, arginine R, lysine K, glutamic acid E, glutamine Q, aspartic acid D, asparagine N, tyrosine Y, tryptophan W, serine S; Turn residues: glycine G, proline P; Hydrophobic residues: isoleucine I, valine V, leucine L, phenylalanine F, tyrosine Y, tryptophan W, alanine A, methionine M, cysteine C, histidine H; Residues with high entropy increase potential: isoleucine I, valine V, leucine L, lysine K, arginine R, phenylalanine F, tyrosine Y, tryptophan W, alanine A, methionine M, cysteine C, glutamic acid E, glutamine Q, histidine H; Finally, determine the preferred pairing between amino acid residue side chains The preferred pairing between amino acid residue side chains is that the lateral attachment between multiple pairs of amino acid residue side chains results in sufficient entropy increase in the low-entropy hydration layer of the multiple pairs of amino acid residue side chains. The pair relationship between the multiple pairs of amino acid residues is the preferred pairing between amino acid residue side chains. The preferred pairing between amino acid residue side chains includes preferred pairing between amino acid residue side chains in a 1-4 / 5 sequence position relationship, preferred pairing between amino acid residue side chains in a 1-3 sequence position relationship, and preferred pairing between amino acid residue side chains in a 1-2 sequence position relationship; The preferred pairing between amino acid residue side chains is determined as follows: FF, QQ, KK, EE, TT, RR, AA, SS, MM; MA, QE, TS, RE, KR, EK, FY, WY, WE, RW, RY, KY, HV, DE, NE, DQ, NQ; QR / K, I / V / L / FI / V / L / F, I / V / L / F / YI / V / L / F / Y, I / V / LY, I / V / LK, I / V / L- R, I / V / LA, I / V / LW, I / V / L / F / YW, I / V / LM, CM / Y, MI / V / L, A / TD / N; The 1-4 / 5 sequence position relationship refers to the relationship between the side chain of one residue and the side chain of another residue in the primary structure of the protein, which are separated by 2 or 3 amino acid residues. The 1-3 sequence position relationship refers to the relationship between the sequence positions of one residue side chain and another residue side chain in the primary structure of the protein, with an interval of one amino acid residue; The 1-2 sequence position relationship refers to the sequence position relationship between a side chain of one residue and the side chain of another adjacent residue in the primary structure of the protein; 2. Protein secondary structure prediction: Based on the sequence position of the residue side chains, the classification and entropy of the amino acid residues, and the preferred pairing between the residue side chains in step 1, the following process is used to predict the protein secondary structure: S100, starting a new line below the amino acid sequence marked with entropy values, repeatedly listing the amino acid residues in the amino acid sequence that are high entropy increase potential residues as a high entropy increase potential residue connected segment; then, starting a new line, repeatedly listing the amino acid residues in the amino acid sequence that are low entropy residues and turn residues; S200. List the unblocked connected segments of residues with high entropy increase potential, and assume that the listed unblocked connected segments of residues with high entropy increase potential are α-helices or β-sheets in the secondary structure of the protein: S201. Assuming that a connected segment of residues with high entropy increase potential forms an α-helix, if a side chain of a residue in the connected segment of residues with high entropy increase potential achieves preferred pairing between side chains of amino acid residues, the amino acid residue in the connected segment of residues with high entropy increase potential is marked with an *. S202. Assuming that the connected segment of residues with high entropy increase potential forms a β-sheet, if the connected segment of residues with high entropy increase potential is a preferred pairing between the side chains of amino acid residues in a 1-3 sequence position relationship, then the amino acid residues in the connected segment of residues with high entropy increase potential are marked with &; S203, analyze the marked * and &: S2031. If none of the connected segments of residues with high entropy increase potential in the primary structure of the protein are marked with &-&, and more than 80% of the amino acid residues in the connected segment of residues with high entropy increase potential are marked with *, then the connected segment of residues with high entropy increase potential is predicted to form an α-helix; S2032. If the connected segment of residues with high entropy increase potential in the primary structure of the protein is marked with &-&, it is predicted that the connected segment of residues with high entropy increase potential forms a β-sheet; if the number of * marked in the connected segment of residues with high entropy increase potential in the primary structure of the protein is less than 50% of the number of amino acid residues in the connected segment of residues with high entropy increase potential, it is predicted that the connected segment of residues with high entropy increase potential forms a β-sheet; S300, searching the amino acid side chains on both sides of the connected segment of residues with high entropy increase potential that has been predicted in S200 to form an α-helix, if the amino acid residues in the amino acid side chains on both sides have a preferred pairing between the side chains of the amino acid residues in a 1-4 / 5 sequence position relationship with the amino acid residues in the connected segment of residues with high entropy increase potential, then it is predicted that the side chains of the amino acid residues in the 1-4 / 5 sequence position relationship preferably pair to form an α-helix; S400, determine whether the S200 hypothesis predicts the formation of α-helix or β-sheet in the protein secondary structure: The entropy increments of the amino acid residues in the primary structure of the protein predicted to be α-helix and β-sheet are respectively counted, and the entropy values of the amino acid residues predicted to be α-helix and β-sheet are compared; when the value of the relatively larger entropy increment of the two is greater than the value of the other entropy increment, the prediction result of the relatively larger entropy value is determined as the final prediction result; when the value of the relatively larger entropy increment of the two is less than 20% of the value of the other entropy increment, the prediction of the formation of α-helix or β-sheet in the secondary structure of the protein by S200 is determined as a potential result; Among them, the entropy increment statistical method of β-folding: there is a preferred pairing between the side chains of amino acid residues with a 1-3 sequence position relationship in the primary structure of the protein, and the entropy value of the amino acid residues in the preferred pairing is used as the entropy increment; Statistical method for entropy increment of α-helix: there is a preferred pairing between the side chains of amino acid residues with a 1-4 / 5 sequence position relationship in the primary structure of the protein, and the entropy value of the amino acid residues in the preferred pairing is used as the entropy increment; S500: If there are two consecutive amino acid residues with low entropy increasing potential in the segment connected by residues with high entropy increasing potential that are not predicted to be α-helices in the primary structure of the protein, the two amino acid residues are predicted to form a turn and are marked as T; A segment of 4 or more consecutive hydrophilic amino acid residues in the primary structure of a protein is named a "continuous hydrophilic residue segment." If a segment of a protein with more than 7 amino acid residues in a connected segment of high entropy-increasing residues that is not predicted to be an α-helix contains residues with low entropy-increasing potential and / or turn residues and / or amino acid residue A and / or residues in a continuous hydrophilic residue segment, and the number of residues in the segment with more than 7 amino acid residues accounts for more than or equal to 50% of the total number of residues in the segment with more than 7 amino acid residues, then the segment with more than 7 amino acid residues is predicted to form a random coil and is considered as a potential prediction result. S600, if a fragment of a protein primary structure consisting of five or more consecutive hydrophilic amino acid residues has been predicted to be a connected fragment of residues with high entropy increase potential, predicting the fragment consisting of five or more consecutive hydrophilic amino acid residues to form a random coil as a potential prediction result, thereby completing the protein structure prediction method based on the mutual attraction between low-entropy hydration layers of residue side chains; In the process of determining the unblocked high entropy increase potential residue connected segment, S200 first finds a segment consisting of consecutive high entropy increase potential residues in the sequence and names it the original "high entropy increase potential residue connected segment". Then, non-high entropy increase potential residues that do not block the high entropy increase potential residue connected segment are determined, and then the unblocked high entropy increase potential residue connected segment is determined, specifically including: (1) When there is only one amino acid residue S or T between two high entropy increase potential residues, the S or T residue participates in forming an unblocked high entropy increase potential residue connected segment; (2) When there is only one isolated N, D or G residue between two high entropy increase potential residues in the amino acid sequence, and there is a high entropy pairing between I, V, L, F, and Y with a 1-3 or 1-4 / 5 relationship across D, N or G, it is considered that D, N or G participates in the formation of an unblocked high entropy increase potential residue connected segment; (3) When any amino acid residue among S, T, N, D, and G appears in the amino acid sequence of the polypeptide chain and is adjacent to any other amino acid residue among S, T, N, D, and G, the 1-4 / 5 sequence position relationship across the two adjacent amino acid residues is checked. When the two amino acid residues at the 1-4 / 5 positions are both I, V, L, F, and Y residues, and at least one of the amino acid residues is I, V, or L, it is predicted that the two adjacent amino acid residues participate in forming an unblocked high entropy increase potential residue connected segment; (4) When a P residue appears in the amino acid sequence of a polypeptide chain, and the P residue is not in a 1-2 or 1-3 sequence position relationship with a D or N residue, and there is a 1-4 / 5 sequence position relationship across P, and the amino acid residues on 1-4 / 5 are I, V, L, or F, then it is predicted that the P residue participates in the formation of an unblocked high entropy increase potential residue connected segment; (5) When a G or A residue appears in the amino acid sequence of a polypeptide chain, if an amino acid residue that forms a 1-4 / 5 sequence position relationship with the G or A residue can be found on one side of the sequence where the G or A residue is located, and an amino acid residue that forms a 1-4 / 5 sequence position relationship with the G or A residue can also be found on the other side of the sequence where the G or A residue is located, and if there is a single D, N, G or P residue that is not adjacent to other low-entropy residues or corner residues between the I, V, L residues, or there are two adjacent S, T, N, D residues, then it is predicted that these single D, N, G or P residues or two adjacent S, T, N, D, G residues participate in the formation of an unblocked high-entropy potential residue connected segment.

2. The prediction method according to claim 1, wherein: The preferred pairing between the side chains of amino acid residues in step S201 refers to the preferred pairing between the side chains of amino acid residues with a 1-4 / 5 sequence position relationship; if there is a preferred pairing between the side chains of amino acid residues with a 1-2 sequence position relationship in the connected segment of residues with high entropy increase potential, or there is a sequence position relationship of IVL 1-4 / 5 G / A 1-4 / 5 IVLF in the connected segment of residues with high entropy increase potential, then it is marked as * below the amino acid.

3. The prediction method according to claim 1, wherein: The preferred pairing between the side chains of the amino acid residues in the 1-3 sequence position relationship described in S202 is I / V / L / FI / V / L / F; if, starting from the amino acids in I / V / L / FI / V / L / F, there is a preferred pairing between the side chains of the amino acid residues in the 1-4 / 5 sequence position relationship in the adjacent connected fragments of residues with high entropy increase potential, then the amino acids in the preferred pairing between the side chains of the amino acid residues in the 1-3 sequence position relationship are not marked as &.

4. The prediction method according to claim 1, wherein: As described in S2032, if the connected segment of residues with high entropy increase potential in the primary structure of the protein is marked with &-&, and more than 85% of the amino acid residues in the amino acid side chain region of the primary structure of the protein far away from the connected segment of residues with high entropy increase potential marked with &-& are marked with *, and there is a preferred pairing between the side chains of amino acid residues that has a 1-4 / 5 sequence position relationship of I / V / L / F / YI / V / L / F / Y with the irrelevant residues marked with &-&, then it is predicted that the connected segment of residues with high entropy increase potential marked with &-& forms an α-helix; wherein, the distant refers to being separated by two consecutive low entropy residues or turn residues, or being separated by at least 3 hydrophilic residues, or being separated by one amino acid residue P; the irrelevant residues refer to other amino acid residues in the primary structure of the protein except those marked with &-&; the amino acid side chain region refers to the region composed of amino acid side chains with the number of amino acid residues greater than 5.

5. The prediction method according to claim 1 or 4, characterized in that: The connected segment of high entropy increase potential residues described in S2032 is predicted to form a β-fold, and the connected segment of high entropy increase potential residues contains two consecutive low entropy residues aspartic acid D or asparagine N, then the two consecutive low entropy residues aspartic acid D or asparagine N are predicted to form a turn as a prediction result.

6. The prediction method according to claim 1, wherein: In the entropy increment statistics of the β-sheet described in S400, if the same amino acid residue in the primary structure of the protein realizes a preferred pairing between the side chains of two amino acid residues with a 1-3 sequence position relationship, the entropy increment of the amino acid residue is accumulated twice; In the entropy increment statistics of α-helix, if the same amino acid residue in the primary structure of the protein realizes multiple preferred pairings between the side chains of amino acid residues with 1-4 / 5 sequence position relationships, the entropy increment of the amino acid residue is accumulated twice.

7. The prediction method according to claim 1 or 6, characterized in that: In the entropy increment statistics of the α-helix described in S400, if there is a preferred pairing between the side chains of amino acid residues with a 1-2 sequence position relationship in the primary structure of the protein, and the preferred pairing is I / V / L / MI / V / L / M, then the two pairs of I / V / L / M are recorded as one entropy increment.

8. The prediction method according to claim 1, wherein: The amino acid residues in the two consecutive amino acid residues in S600 are one or two of S, T, N, D, and G.

9. The prediction method according to claim 1, wherein: The relative measurement value of the entropy increase potential of the low entropy hydration layer outside the side chains of amino acid residues in the primary structure of the protein is a measurement indicator that takes the length of the hydrophobic portion on the side chains of amino acid residues in the primary structure of the protein as the relative measurement value.

10. The prediction method according to claim 1 or 9, characterized in that: The relative measure of the entropy increase potential of the low-entropy hydration layer on the periphery of the amino acid residue side in the primary structure of the protein is specifically to take the length of a carbon-carbon bond on the amino acid residue chain in the primary structure of the protein as one entropy value.

11. The prediction method according to claim 1, wherein: The entropy values of the 20 amino acids are as follows: Serine S=0.5, threonine T=0.5, aspartic acid D=0.5, asparagine N=0.5, glycine G=0, histidine H=3, arginine R=3, glutamic acid E=3, lysine K=4, tryptophan W=5, isoleucine I=6, valine V=6, leucine L=6, phenylalanine F=6, proline P=0, cysteine C=4, glutamine Q=3, alanine A=1, methionine M=4, tyrosine Y=5.

12. The prediction method according to claim 1, wherein: The sufficient entropy increase described in step 1 means that when the side chains of two amino acid residues in the primary structure of the protein are similar, the side chains of the two residues are laterally attached, so that the low entropy hydration layer of the two side chains undergoes sufficient entropy increase; the similarity of the amino acid residue side chains means that the smaller the difference in the entropy values of the two hydrophilic amino acid residues, the more similar the two hydrophilic residues are; and the smaller the difference in the entropy values of the two hydrophobic amino acid residues, the more similar the two hydrophobic residues are.

13. The prediction method according to claim 1, wherein: The length of the unblocked connected segment of residues with high entropy increase potential is greater than or equal to 3 amino acid residues.

14. The prediction method according to claim 1, wherein: If a protein sequence whose primary structure is not predicted to be an α-helix or β-turn contains a segment of 5 or more amino acid residues, and when the segment is in the general initial thermodynamic metastable state of unfolded proteins, and if there is no hydrophilic amino acid residue in the segment whose side chain height is less than or equal to 1 minus the height of the adjacent hydrophobic residue side chain, then the segment is predicted to be a β-sheet; where the height of the residue side chain refers to the number of covalent bonds in the shortest path between the top atom of the amino acid residue side chain and the main chain carbon atom of the amino acid residue in the segment.

15. A protein structure prediction method based on the mutual attraction between low-entropy hydration layers of residue side chains, characterized by: The prediction method is to predict the protein structure by using water entropy force; the prediction is based on the secondary structure of the protein predicted by the protein structure prediction method based on the mutual attraction relationship between the low entropy hydration layers of the residue side chains according to any one of claims 6 to 14 to predict the tertiary structure of the protein, comprising the following steps: Definition of the similarity of amino acid residue side chains: the smaller the difference in the "entropy value" of the side chains of two hydrophilic residues, the more similar the two hydrophilic residues are; the smaller the difference in the "entropy value" of the side chains of two hydrophobic residues, the more similar the two hydrophobic residues are; The amino acid residues that achieve preferred pairing between I / V / L / F in the predicted secondary structure are marked, and the hydrophobic surfaces of the hydrophobic side chain clusters that achieve lateral adhesion between the side chains of the I / V / L / F amino acid residues are marked and named "cluster hydrophobic surfaces". The protein structure is predicted to be where the hydrophobic surface of an I / V / L / F cluster is surface-adapted to the hydrophobic surface of another adjacent I / V / L / F cluster. When there is no preferred pairing relationship between I / V / L / F in a β-sheet or α-helical secondary structure, the secondary structure is searched for a preferred pairing relationship between I / V / LY. If so, the hydrophobic surface of the hydrophobic side chain clusters that are laterally attached between the side chains of the I / V / LY amino acid residues are also marked as "clustered hydrophobic surface"; the surface attachment of an I / V / L / Y cluster hydrophobic surface to the adjacent cluster hydrophobic surface is predicted as a protein structure.

16. The protein structure prediction method based on the mutual attraction between low-entropy hydration layers of residue side chains according to claim 15, characterized in that: It also includes the following steps: For the case where the current prediction is β-sheet, it is necessary to mark the strong hydrophobic residues of the 1-3 position relationship preferred pairing connected segments in the β-sheet, i.e. I / V / L / F / W / Y, and at the same time mark these hydrophobic segments connected by the 1-3 position relationship, compare the lengths of adjacent β-sheets in the sequence, and predict that α-helices of similar lengths will fold into a sheet structure. Then, based on the side chain distribution of the strong hydrophobic residues of the two β-sheets, the corresponding position relationship between the two is based on the most strong hydrophobic residues on one β-sheet, which can achieve preferred pairing with the side chains of the amino acid residues on the other α-helix. The state of the two β-sheets being attached is the predicted sheet folding structure.

Citation Information

Patent Citations

  • Protein-protein docking method and device based on protein surface low-entropy hydration layer recognition

    CN114512180A

  • Interaction entropy calculation method and application thereof in protein dynamic behavior analysis

    CN118506850A