Protein structure prediction method based on protein network
By building a protein network and extracting topological features, combining machine learning models and optimization algorithms, the problems of insufficient accuracy and high computational cost in complex sequences are solved, and efficient and accurate large-scale protein structure prediction is achieved.
Patent Information
- Application Number
- CN202510181795.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-10
AI Technical Summary
Traditional methods of protein structure prediction are insufficient in prediction accuracy and excessive in calculation costs when facing complex protein sequences.
By building a protein network, topological features, such as edge weights, node degrees, clustering coefficients, etc., and combining three-dimensional structural data, the machine learning model is trained, the prediction model is constructed, the mapping relationship between topological features and protein structure is identified, and the preliminary prediction structure is optimized to obtain the final predicted three-dimensional structure.
It improves the efficiency and accuracy of protein structure prediction, reduces calculation costs, and can complete large-scale protein structure prediction tasks in a short time, meeting the needs of biomedical research and drug development.
Smart Images

Figure CN120126547A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of bioinformatics and computational biotechnology, and in particular to a protein structure prediction method based on protein network. Background Art
[0002] Proteins are the main executors of life activities and undertake various key biological functions in organisms, such as catalyzing metabolic reactions, transmitting signals, and forming cell structures. The function of proteins is closely related to their unique three-dimensional structure. Therefore, accurate prediction of the three-dimensional structure of proteins is extremely important for understanding their biological functions, disease mechanisms, and drug development.
[0003] Protein structure prediction has always been one of the core issues in the field of bioinformatics. Traditional protein structure prediction methods mainly include two categories: physical and chemical simulation and homology modeling. The physical and chemical simulation method is based on the physical and chemical properties of proteins, such as the interaction force between amino acids, hydrogen bonds, van der Waals forces, etc., and simulates the folding process of proteins through molecular dynamics simulation or Monte Carlo simulation to predict their final three-dimensional structure. However, this method requires detailed calculations of each atom of the protein, and the computational complexity is extremely high. The prediction of large or complex protein structures often requires a lot of computing resources and time, and it is difficult to promote on a large scale in practical applications. The homology modeling method is based on the homology between the protein sequence of known structure and the protein sequence to be predicted. By comparing the sequence similarity, the protein of known structure is used as a template to model the structure of the target protein. This method can improve the accuracy of prediction to a certain extent, but it depends on the protein template of known structure. When the similarity between the protein to be predicted and the protein sequence of known structure is low, the accuracy of the prediction result will be greatly reduced, and for those proteins without homology templates, the homology modeling method is powerless.
[0004] In summary, traditional protein structure prediction methods mainly rely on techniques such as homology modeling and ab initio prediction. However, these methods often have problems of insufficient prediction accuracy and high computational cost when faced with some complex protein sequences.
[0005] Therefore, there is an urgent need for a protein structure prediction method that can effectively improve the efficiency and accuracy of protein structure prediction and reduce the computational cost. Summary of the invention
[0006] Based on this, it is necessary to provide a protein structure prediction method based on protein network to address the above technical problems.
[0007] A protein structure prediction method based on a protein network, comprising the following steps: obtaining protein sequence data and three-dimensional structure data, and constructing a protein network; extracting topological features from the protein network, the topological features including edge weights, node degrees, clustering coefficients, betweenness centrality, path lengths, and modularity coefficients; training at least three machine learning models based on the topological features and the three-dimensional structure data to construct a prediction model, the prediction model being capable of identifying the mapping relationship between the topological features and the protein structure; obtaining the topological features of a protein to be measured, inputting same into the trained prediction model, and outputting a preliminary predicted three-dimensional structure of the protein to be measured; and optimizing the preliminary predicted three-dimensional structure through an optimization algorithm to output a predicted three-dimensional structure of the protein to be measured.
[0008] In one embodiment, the obtaining protein sequence data and three-dimensional structure data, and constructing a protein network includes: obtaining protein sequence data and corresponding three-dimensional structure data, taking amino acids as nodes and the interactions between amino acids as edges to obtain basic nodes and connection information, the interactions including hydrogen bond relationships and hydrophobic effect relationships; and constructing a protein network based on the basic nodes and the connection information according to protein sequence similarity, physical interactions, or functional association information.
[0009] In one embodiment, the extracting the topological features of the protein network includes:
[0010] Calculating the edge weight according to the protein network, the formula being:
[0011]
[0012] wherein, w ij represents the edge weight between node i and node j, and d ij represents the Euclidean distance between node i and node j;
[0013] The formula for calculating the node degree is:
[0014]
[0015] wherein, k i represents the node degree, and N(i) represents the set of neighbor nodes of node i;
[0016] The formula for calculating the clustering coefficient is:
[0017]
[0018] wherein, C i represents the clustering coefficient, and E i represents the number of edges actually existing between the neighbor nodes of node i;
[0019] The formula for calculating the betweenness centrality is as follows:
[0020]
[0021] In the formula, B i represents the betweenness centrality, σ st represents the total number of shortest paths between node s and node t, and σ st (i) represents the number of shortest paths passing through node i;
[0022] The formula for calculating the modularity coefficient is as follows:
[0023]
[0024] In the formula, Q represents the modularity coefficient, l c represents the number of edges within community c, d c represents the sum of the degrees of the nodes within community c, and m represents the total number of edges in the network.
[0025] In one embodiment, training at least three machine learning models based on the topological features and three-dimensional structure data to construct a prediction model includes: using the extracted topological features as input and using the three-dimensional structure data as a training set to train at least three machine learning models; adjusting the model parameters and structure, and having the machine learning models learn the mapping relationship between the topological features and the protein structure to construct a prediction model; and using a cross-validation method to evaluate and optimize the prediction model to obtain a trained prediction model.
[0026] In one embodiment, the machine learning models include a support vector machine, a random forest model, and a deep neural network model.
[0027] In one embodiment, the optimization algorithms include a genetic algorithm, a simulated annealing algorithm, and a particle swarm optimization algorithm.
[0028] In one embodiment, optimizing the preliminary predicted three-dimensional structure through an optimization algorithm and outputting the predicted three-dimensional structure of the protein to be tested includes: when performing the initial three-dimensional structure optimization, using the van der Waals force or hydrogen bond interaction as the objective function of the optimization, and using the physical properties, chemical properties, spatial constraints, and geometric constraint information of the protein to constrain the optimization algorithm.
[0029] Compared with the prior art, the advantages and beneficial effects of the present invention are as follows: By obtaining protein sequence data and three-dimensional structure data, constructing a protein network, and extracting topological features from the protein network, including edge weight, node degree, clustering coefficient, betweenness centrality, path length, and modularity coefficient, and based on the topological structure and three-dimensional structure data, training at least three machine learning models to construct a prediction model, which can identify the mapping relationship between topological features and protein structures; obtaining the topological features of the protein to be tested, inputting them into the trained prediction model, outputting the preliminary predicted three-dimensional structure of the protein to be tested, and using an optimization algorithm to optimize the preliminary predicted three-dimensional structure, and outputting the predicted three-dimensional structure of the protein to be tested. Through the topological structure of the protein network, considering the related functions and functional associations of proteins from a global perspective, it is possible to capture the information of protein structures more comprehensively, improve the accuracy of prediction results. At the same time, combining machine learning models and optimization algorithms makes the prediction process more efficient, capable of completing the prediction task of large-scale protein structures in a short time, and having a low prediction cost, which can meet the large-scale protein structure prediction needs in fields such as biomedical research and drug development. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 It is a schematic flowchart of a protein structure prediction method based on a protein network in an embodiment;
[0031] Figure 2 It is a flowchart of a protein structure prediction method based on a protein network in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0032] Before describing the specific embodiments of the present invention, the overall concept of the present invention is described as follows:
[0033] The present invention is mainly developed based on the protein structure prediction process. Currently, there are problems of insufficient prediction accuracy and too high computational cost in the prediction methods.
[0034] The inventors analyzed and found that proteins do not exist in isolation in cells, but interact with other biomolecules such as other proteins, nucleic acids, and small molecules to form complex protein networks. By constructing a protein network, the interaction relationships and functional associations of proteins can be presented in the form of a network. Using the topological features of the network, such as node degree, clustering coefficient, betweenness centrality, etc., the potential connections and functional modules between proteins can be mined, providing new ideas and methods for protein structure prediction. Since there is a certain correlation between the topological structure of the protein network and the three-dimensional structure of the protein, by analyzing the topological features of the network, the efficiency and accuracy of protein structure prediction can be effectively improved, providing a new way to solve the problems existing in traditional methods.
[0035] Therefore, the present invention proposes a protein structure prediction method based on a protein network. By obtaining protein sequence data and three-dimensional structure data, a protein network is constructed. Topological features, including edge weights, node degrees, clustering coefficients, betweenness centrality, path lengths, and modularity coefficients, are extracted from the protein network. Based on the topological structure and three-dimensional structure data, at least three machine learning models are trained to construct a prediction model that can identify the mapping relationship between topological features and protein structures. The topological features of the protein to be tested are obtained and input into the trained prediction model, and the preliminary predicted three-dimensional structure of the protein to be tested is output. Then, an optimization algorithm is used to optimize the preliminary predicted three-dimensional structure, and the predicted three-dimensional structure of the protein to be tested is output. Through the topological structure of the protein network, the related functions and functional associations of proteins are considered from a global perspective, which can more comprehensively capture the information of protein structures and improve the accuracy of prediction results. At the same time, the combination of machine learning models and optimization algorithms makes the prediction process more efficient, capable of completing the prediction task of large-scale protein structures in a relatively short time, with low prediction costs, and can meet the large-scale protein structure prediction requirements in fields such as biomedical research and drug development.
[0036] After introducing the overall concept of the present invention, in order to make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below through specific embodiments in conjunction with the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0037] In one embodiment, as Figure 1 shown, a protein structure prediction method based on a protein network is provided, including the following steps:
[0038] Step S110, obtain protein sequence data and three-dimensional structure data, and construct a protein network.
[0039] Specifically, when obtaining protein sequence data and three-dimensional structure data, protein sequence data can be obtained from a public biological database such as UniProt, and three-dimensional structure data can be obtained from the Protein Data Bank, so that a protein network can be constructed based on the protein sequence data and three-dimensional structure data.
[0040] For example, when studying a specific type of protein, protein sequences with similar functions can be screened out from UniProt, and then known structure proteins that match them can be found in the Protein Data Bank, providing basic nodes and connection information for constructing the network.
[0041] Among them, step S110 includes: obtaining a protein sequence and corresponding three-dimensional structure data, taking amino acids as nodes and the interactions between amino acids as edges to obtain basic nodes and connection information, where the interactions include hydrogen bond relationships and hydrophobic effect relationships; based on the basic nodes and connection information, constructing a protein network based on protein sequence similarity or physical interaction or functional association information.
[0042] Specifically, obtain a protein sequence and corresponding known three-dimensional structure data, represent each amino acid as a node in the network, and construct edges according to the interactions between amino acids. The interactions can be hydrogen bond relationships, hydrophobic interaction relationships, etc., to obtain basic nodes and connection information.
[0043] When constructing a protein network, based on the basic nodes and connection information, the network can be constructed according to protein sequence similarity. By calculating the similarity score between protein sequences, for example, using tools such as BLAST for sequence alignment, when the sequence similarity exceeds a certain threshold, it is considered that two proteins may be evolutionarily related, and thus a connection is established in the protein network. In addition, the network can also be constructed according to information such as the physical interactions and functional associations of proteins. For example, the protein interaction data obtained through techniques such as yeast two-hybrid experiments is used as the basis for network construction, and the interacting proteins are connected in the network to form a protein network.
[0044] Step S120, extract topological features from the protein network. The topological features include edge weight, node degree, clustering coefficient, betweenness centrality, path length, and modularity coefficient.
[0045] Specifically, according to the constructed protein network, extract its key topological features, including but not limited to edge weight, node degree, clustering coefficient, betweenness centrality, path length, and modularity coefficient. The complex structure information of the protein network is converted into quantifiable data through the extracted topological features, providing input data for subsequent machine learning model training.
[0046] Among them, step S120 includes: calculating the edge weight according to the protein network, and the formula is:
[0047]
[0048] In the formula, w ij represents the edge weight between node i and node j, and d ij represents the Euclidean distance between node i and node j; the formula for calculating the node degree is:
[0049]
[0050] In the formula, k iDenote the node degree, and \(N(i)\) represents the set of neighbor nodes of node \(i\); the formula for calculating the clustering coefficient is:
[0051]
[0052] In the formula, \(C\) i represents the clustering coefficient, and \(E\) i represents the actual number of edges existing between the neighbor nodes of node \(i\); the formula for calculating the betweenness centrality is:
[0053]
[0054] In the formula, \(B\) i represents the betweenness centrality, \(\sigma\) st represents the total number of shortest paths between node \(s\) and node \(t\), and \(\sigma\) st (i) represents the number of shortest paths passing through node \(i\); the formula for calculating the modularity coefficient is:
[0055]
[0056] In the formula, \(Q\) represents the modularity coefficient, \(l\) c represents the number of edges within community \(c\), \(d\) c represents the sum of the degrees of the nodes within community \(c\), and \(m\) represents the total number of edges in the network.
[0057] Specifically, the edge weight refers to the relevant value or measure of the edges connecting the various nodes of the chain, reflecting the strength of the relationship between the nodes.
[0058] The node degree refers to the number of other nodes directly connected to a protein node. In a protein network, proteins with a high node degree often have important biological functions, and they may be involved in multiple biological processes or have extensive interactions with other proteins.
[0059] The clustering coefficient reflects the degree of tight connection between the neighbor nodes of a node in a protein network. A higher clustering coefficient means that there are more local cluster structures in the network, which may imply the modular characteristics of proteins in terms of function.
[0060] The betweenness centrality measures the degree to which a protein node serves as an intermediary for the shortest paths between other nodes in the network. Proteins with a high betweenness centrality play a key role in information transmission and network stability.
[0061] The modularity coefficient measures the degree of clustering in the network and is used to detect the community structure in the network. The higher the modularity coefficient, the better the community structure division of the network.
[0062] The above topological structure reveals the position and role of proteins in the network from different perspectives, which is closely related to the structure and function of proteins and facilitates the subsequent determination of the relationship between protein structure and topological features.
[0063] Step S130: Based on the topological features and three-dimensional structure data, train at least three machine learning models to construct a prediction model that can identify the mapping relationship between topological features and protein structure.
[0064] Specifically, after obtaining the topological structure of the protein network, multiple machine learning models can be used for training, such as support vector machines, neural networks, and random forests. Taking the neural network as an example, by constructing a multi-layer neural network structure, the preprocessed network topological features are input into the network. After multiple non-linear transformations and weight updates, the predicted protein structure information is finally output. During the training process, protein data with known structures is used as the training set. By adjusting the parameters of the network, the model can learn the mapping relationship between network topological features and protein structure. For example, using a convolutional neural network can effectively extract the local correlation information between features and may have a better effect on predicting the local structure features of proteins.
[0065] After training the machine learning model using topological features and three-dimensional structure data, a prediction model is constructed that can identify the mapping relationship between topological features and three-dimensional structure data.
[0066] Among them, step S130 includes: using the extracted topological features as input and the three-dimensional structure data as the training set to train at least three machine learning models; adjusting the model parameters and structure, enabling the machine learning model to learn the mapping relationship between topological features and protein structure, and constructing a prediction model; using the cross-validation method to evaluate and optimize the prediction model to obtain the trained prediction model.
[0067] Specifically, use the extracted topological features as input and the known three-dimensional structure data as the training set to train at least three machine learning models; by adjusting the parameters and structure of the model, enable the model to learn the mapping relationship between topological features and protein structure, and establish a prediction model. During the training process, methods such as cross-validation can be used to evaluate and optimize the performance of the prediction model to obtain the trained prediction model, improve the generalization ability and prediction accuracy of the model, and enable more accurate prediction of the unknown structure of proteins.
[0068] Among them, the machine learning models include support vector machines, random forest models, and deep neural network models.
[0069] Step S140: Obtain the topological features of the protein to be measured, input them into the trained prediction model, and output the preliminary predicted three-dimensional structure of the protein to be measured.
[0070] Specifically, obtain the topological structure of the protein to be measured, input it into the trained prediction model, and output the preliminary predicted three-dimensional structure of the protein to be measured. The preliminary predicted three-dimensional structure contains the basic features and information of the protein structure, such as the distribution of secondary structure elements, folding patterns, etc.
[0071] Step S150: Optimize the preliminary predicted three-dimensional structure through an optimization algorithm, and output the predicted three-dimensional structure of the protein to be measured.
[0072] Specifically, in order to improve the accuracy of the predicted structure, an optimization algorithm can be used to optimize the preliminary three-dimensional structure. By performing a series of operations such as mutation, crossover, and selection on the preliminary three-dimensional structure, search for the optimal solution space of the protein structure, make the predicted structure closer to the real structure, and output the predicted three-dimensional structure of the protein to be measured. Through the combination of the machine learning model and the optimization algorithm, the prediction process is made more efficient, capable of completing the prediction task of large-scale protein structures in a shorter time, and improving the prediction accuracy.
[0073] Among them, the optimization algorithms include genetic algorithms, simulated annealing algorithms, and particle swarm optimization algorithms.
[0074] Among them, Step S150 includes: when performing three-dimensional structure optimization, use van der Waals forces or hydrogen bond interactions as the objective function of optimization, and use the physical properties, chemical properties, spatial constraints, and geometric constraint information of the protein to constrain the optimization algorithm.
[0075] Specifically, during the optimization process, some physicochemical constraint conditions, such as van der Waals forces and hydrogen bond interactions, can be used as the objective function of optimization. By setting appropriate objective functions of optimization, such as energy functions and structure similarity functions, etc., iteratively optimize the predicted structure to make it gradually approach the real three-dimensional structure to ensure the rationality and stability of the predicted structure. In addition, the physical properties, chemical properties, spatial constraints, and geometric constraints of the protein can also be combined to guide and constrain the optimization algorithm, improving the optimization efficiency and effect.
[0076] In this embodiment, by obtaining protein sequence data and three-dimensional structure data, a protein network is constructed, and topological features are extracted from the protein network, including edge weights, node degrees, clustering coefficients, betweenness centrality, path lengths, and modularity coefficients. Based on the topological structure and three-dimensional structure data, at least three machine learning models are trained to construct a prediction model that can identify the mapping relationship between topological features and protein structures; the topological features of the protein to be tested are obtained and input into the trained prediction model, and the preliminary predicted three-dimensional structure of the protein to be tested is output. Then, an optimization algorithm is used to optimize the preliminary predicted three-dimensional structure, and the predicted three-dimensional structure of the protein to be tested is output. Through the topological structure of the protein network, the relevant functions and functional associations of the protein are considered from a global perspective, which can more comprehensively capture the information of the protein structure and improve the accuracy of the prediction results. At the same time, the combination of machine learning models and optimization algorithms makes the prediction process more efficient, can complete the prediction task of large-scale protein structures in a short time, and has a low prediction cost, which can meet the large-scale protein structure prediction needs in the fields of biomedical research, drug development, etc.
[0077] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. Among them, the storage medium can be a magnetic disk, an optical disc, a read-only memory (ROM), or a random access memory (RAM), etc.
[0078] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device, so that they can be stored in a computer storage medium (ROM / RAM, magnetic disk, optical disc) and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps of them can be made into a single integrated circuit module to implement. Therefore, the present invention is not limited to any specific combination of hardware and software.
[0079] The above content is a further detailed description of the present invention in combination with specific embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.
Claims
1. A protein structure prediction method based on protein network, characterized in that: The following steps are involved: Obtain protein sequence data and three-dimensional structure data to construct protein networks; Extracting topological features according to the protein network, wherein the topological features include edge weight, node degree, clustering coefficient, betweenness centrality, path length and modularity coefficient; Based on the topological features and the three-dimensional structure data, at least three machine learning models are trained to construct a prediction model, wherein the prediction model can identify a mapping relationship between topological features and protein structures; Obtain the topological features of the protein to be tested, input them into the trained prediction model, and output the preliminary predicted three-dimensional structure of the protein to be tested; The preliminary predicted three-dimensional structure is optimized by an optimization algorithm, and the predicted three-dimensional structure of the protein to be tested is output.
2. The protein structure prediction method based on protein network according to claim 1, characterized in that: The step of obtaining protein sequence data and three-dimensional structure data and constructing a protein network comprises: Obtain protein sequence data and corresponding three-dimensional structure data, use amino acids as nodes, and interactions between amino acids as edges to obtain basic nodes and connection information, wherein the interactions include hydrogen bonding relationships and hydrophobic effect relationships; According to the basic nodes and connection information, a protein network is constructed based on protein sequence similarity, physical interaction or functional association information.
3. The protein structure prediction method based on protein network according to claim 2, characterized in that: The extracting the topological features of the protein network comprises: The edge weight is calculated according to the protein network, and the formula is: In the formula, w ij represents the edge weight between node i and node j, d ij represents the Euclidean distance between node i and node j; The formula for calculating the node degree is: In the formula, k i represents the node degree, N(i) represents the set of neighbor nodes of node i; The formula for calculating the clustering coefficient is: In the formula, C i represents the clustering coefficient, E i Indicates the actual number of edges between neighbor nodes of node i; The formula for calculating the betweenness center is: In the formula, B i represents the betweenness center, σ st represents the total number of shortest paths between node s and node t, σ st (i) represents the number of shortest paths passing through node i; The formula for calculating the modularity coefficient is: In the formula, Q represents the modularity coefficient, l c represents the number of edges within community c, d c represents the sum of the degrees of the nodes in community c, and m represents the total number of edges in the network.
4. The protein structure prediction method based on protein network according to claim 1, characterized in that: The method of training at least three machine learning models based on the topological features and the three-dimensional structure data to construct a prediction model includes: Using the extracted topological features as input and the three-dimensional structure data as a training set, training at least three machine learning models; Adjusting the model parameters and structure, learning the mapping relationship between the topological features and the protein structure through a machine learning model, and constructing a prediction model; The prediction model is evaluated and optimized using a cross-validation method to obtain a trained prediction model.
5. The protein structure prediction method based on protein network according to claim 1, characterized in that: The machine learning models include support vector machines, random forest models and deep neural network models.
6. The protein structure prediction method based on protein network according to claim 1, characterized in that: The optimization algorithms include genetic algorithm, simulated annealing algorithm and particle swarm optimization algorithm.
7. The protein structure prediction method based on protein network according to claim 1, characterized in that: The step of optimizing the preliminary predicted three-dimensional structure by an optimization algorithm and outputting the predicted three-dimensional structure of the protein to be tested comprises: When performing the initial three-dimensional structure optimization, van der Waals force or hydrogen bonding is used as the optimization objective function, and the optimization algorithm is constrained by the physical properties, chemical properties, spatial constraints and geometric constraints of the protein.