Protein mutation site prediction method based on graph neural network
Through a graph neural network-based method, the amino acid residues are encoded in combination with protein local structural information, which solves the existing methods' dependence on homologous information and high training costs, and achieves more efficient and accurate prediction of protein mutation sites.
Patent Information
- Application Number
- CN202510249455.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-03
AI Technical Summary
The existing protein mutation site prediction methods are highly dependent on homologous information data, have high training costs, and the model lacks utilization of local environmental characteristics of protein amino acid residues.
Using a graph neural network-based method, protein amino acid sequence characteristics are extracted, and each amino acid residue is encoded based on protein local structural information, thereby learning the potential residue mutation sites in the protein structure. Specific steps include extracting binding site information, building a protein map, updating node characteristics using a messaging mechanism, and predicting amino acid residues through a classifier.
By accurately extracting binding site information and making full use of protein structure information, the accuracy and efficiency of mutation site prediction are improved, training costs are reduced, and the model makes more fully utilized local environmental characteristics.
Smart Images

Figure CN120089199A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of neural networks, and in particular, relates to a method for predicting protein mutation sites based on graph neural networks. Background Art
[0002] Proteins are the material basis and the only form for all life activities. However, natural proteins existing in organisms often have various defects in practical applications and need to be modified. Protein engineering, such as de novo protein design or molecular modification of natural proteins, has become an important way to solve this problem. Protein engineering can synthesize proteins with specific amino acid sequences and spatial structures according to needs. Based on determining the relationship between protein amino acid sequences and biological functions, it designs and synthesizes proteins with specific functions, providing support for fields such as synthetic biology and new drug research and development.
[0003] Specifically, protein engineering is to modify existing proteins or create a new protein by means of chemistry, physics, molecular biology, etc. Most traditional protein engineering methods modify target proteins by introducing random mutations. With the rapid development of computer technology and bioinformatics technology, computer-aided design has been widely applied to protein engineering, eliminating a large number of repetitive experiments. Among them, an emerging application is to train deep learning models to predict the mutation effects of each site in proteins or the sites allowing mutations, and to explore protein variants with specific functions. In recent years, prediction methods based on deep learning have been proposed and verified and applied in practical applications. Most of these methods extract features from protein sequences based on multiple sequence alignment (MSA) and protein language models (PLMs). The results of the former highly depend on the amount of data of protein homologous information. However, in reality, not all protein sequences can be subjected to homologous alignment. The latter is based on natural language processing, and building a model usually requires a large amount of training data, complex model design, and generally high training costs.
[0004] Deep learning is an important branch of machine learning. It constructs a deep neural network model to learn the internal laws and representation levels of sample data, and then realizes the recognition of data such as text, images, and sounds, endowing the computer with the ability to analyze and learn analogous to the human brain. The basis of deep learning is the neural network, which consists of multiple neurons. A neural network usually includes an input layer, a hidden layer, and an output layer. The number of hidden layers determines the depth of the network. Each layer of neurons performs linear transformation and non-linear transformation on the input data, and then passes the result to the next layer of neurons, and finally outputs the result of the model.
[0005] The backpropagation algorithm plays an important role in the training process of deep learning neural networks. During the backpropagation process, the error between the model output and the actual label is first calculated, and then the error is propagated backward through the network. According to the chain rule, the gradient of each layer is calculated. This process can optimize the parameters of the neural network (such as weights and biases). Repeat the above process multiple times until the error no longer changes and the network can produce accurate outputs.
[0006] Compared with traditional machine learning, deep learning neural networks can automatically learn features and patterns from data, form more abstract high-level representation attribute categories or features by combining low-level features, and avoid the incompleteness of manually designed feature representations. By designing and establishing an appropriate number of neuron computing nodes and multi-layer operation hierarchies, deep learning neural networks can approximate the real model to the greatest extent, meet the automation requirements for processing complex things, and drive technological breakthroughs in multiple fields.
[0007] The powerful learning ability, reasoning ability, and adaptability of deep learning neural networks enable them to learn the mapping relationship between protein amino acid sequences and functions. Based on data support, by learning known protein structures, the model can provide guidance for the design of target protein molecules with specific functions, providing technical support for fields such as synthetic biology and new drug research and development.
[0008] Proteins are biological macromolecules formed by the connection of 20 kinds of amino acids. Each protein has its own unique spatial structure, and its molecular structure can be divided into four levels. The primary structure is the amino acid sequence that composes the polypeptide chain of the protein. The secondary structure is the local spatial conformation formed by the main chain of the polypeptide chain folding and coiling in a certain direction through hydrogen bond interactions. The tertiary structure is a regular three-dimensional structure formed by further coiling and folding on the basis of the secondary structure. The regular spatial structure formed by the association of two or more subunits or subunits with tertiary structures through secondary bonds is the quaternary structure of the protein.
[0009] The higher-level structure of a protein determines its function, and the primary structure of a protein determines its higher-level structure. Therefore, the primary structure of a protein, i.e., the amino acid sequence, determines the function of the protein. Graph Neural Networks (GNNs) are a type of deep learning neural network model, and their application in protein data has developed rapidly, especially in dealing with data with complex topological relationships. The amino acid sequence and spatial structure of a protein constitute graph data with complex dependencies, where each amino acid residue can be regarded as a node in the graph, and the chemical bonds or spatial proximity relationships connecting these residues serve as the edges of the graph. The graph neural network model captures global and local graph structure information by designing a message passing mechanism to exchange information between nodes (amino acid residues).
[0010] The advantage of graph neural networks in protein modeling is that it can not only process the primary sequence (amino acid sequence) of a protein but also combine the secondary, tertiary, and quaternary structure information of the protein. In this case, the secondary structure of a protein can be represented by the adjacency matrix of the graph, and the graph neural network gradually updates the feature representation of each node through multiple stacked message passing processes. In this update process, the initial state of the node (residue) is dynamically adjusted to learn more expressive features. This method makes up for the deficiency of traditional sequence models in not being able to fully utilize protein spatial structure information.
[0011] Traditional sequence-based prediction models cannot comprehensively capture important information in the local structure of proteins. In contrast, graph neural networks can use the spatial neighborhood of each residue to judge the possible effects of its mutations. For example, in practical applications, the active sites of some enzymes usually consist of several residues that are spatially close. By modeling this local environment, graph neural networks can identify important residues related to function, thereby improving the accuracy of mutation prediction. Summary of the Invention
[0012] In view of this, in response to the defects of existing prediction methods, such as high dependence on homologous information data and high training costs, and the problem that the model inadequately utilizes the local environmental characteristics of protein amino acid residues, the present invention proposes a method for predicting protein mutation sites based on graph neural networks, which applies graph neural networks to extract features of protein amino acid sequences, encodes each amino acid residue by combining protein local structure information, and thus learns potential residue mutation sites in the protein structure.
[0013] To achieve the above object, the technical solution of the present invention is realized as follows:
[0014] The first aspect of the present invention provides a method for predicting protein mutation sites based on graph neural networks, including the following steps:
[0015] Step 1: Extract the binding site information corresponding to each protein in the dataset;
[0016] Step 2: According to the PDB file of the protein and the binding site information, convert each protein in the dataset into graph-structured data to obtain a protein graph;
[0017] Step 3: Build a graph neural network model, input the obtained protein graph into the model. In the graph neural network, use the message passing mechanism to update the features of the nodes, and input the features of the nodes into a classifier for amino acid residue prediction;
[0018] Step 4: Compare the amino acid residue types predicted by the model with the wild-type residue labels, and calculate the loss value;
[0019] Step 5: Backpropagation, further learn the model parameters until the loss value is stable, and save the best parameters of the model after training;
[0020] Step 6: Input the target protein into the trained model, compare the amino acid residues predicted by the model with the wild-type amino acid residues at the corresponding sites. If they are different from the wild-type and there are property differences, it is considered that a feasible mutation may occur at this site.
[0021] Furthermore, in the above Step 1, each protein node constructs its feature vector through binding site information, spatial geometric structure, and chemical background, including amino acid type, temperature factor, occupancy, atomic coordinates, binding site information, and neighboring amino acid types.
[0022] Furthermore, the above Step 2 includes the following steps:
[0023] Step 201: Read the PDB-format protein file in the dataset, including the whole protein and its binding pocket information, use a chemical information processing tool to read the sequence and structure information contained therein, convert it into graph-structured data, map the amino acids to the nodes in the graph, and establish the edges in the graph according to the positional relationship;
[0024] Step 202: Map the three-dimensional structure data of the protein to the graph-structured data. One protein generates a DGL graph, and each amino acid corresponds to a node in the graph. The node features contain the chemical and physical information of the amino acid residue, including amino acid type, C α atomic coordinates, temperature factor, occupancy, binding site information, using one-hot encoding, and the position of the residue is determined by the three-dimensional coordinates of the C α atom. The distance between two residues is calculated according to the Euclidean distance between two C α atoms, as shown in formula (1), where d is the distance between two residues C αThe Euclidean distance between atoms, where x, y, and z are three spatial coordinates. When the condition that the distance is less than 7 angstroms is satisfied, an undirected edge is established in the graph;
[0025]
[0026] Step 203: Collect the graph representations of each protein into a graph dataset. Each graph contains nodes, edges, and corresponding feature information.
[0027] Furthermore, in the above-mentioned step 3, using the message passing mechanism to update the features of nodes includes:
[0028] For a specific node v, the message passing formula at time step t + 1 is:
[0029]
[0030] Among them, is the information of all neighbor nodes received by node v at time step t + 1, N(v) is the set of neighbor nodes of node v, and represent the features of node v and neighbor node ω at time step t respectively, e vω is the edge feature between node v and ω, M t is the message function, which is used to convert the information of neighbor nodes into messages that can be received by the current node;
[0031] After receiving the messages from neighbor nodes, node v will update its own features according to the received messages, as shown in formula (3):
[0032]
[0033] Among them, U t is the update function. Through the above message passing mechanism, node v will continuously receive the status information from neighbor nodes and update its own status according to this information.
[0034] The second aspect of the present invention provides a protein mutation site prediction device based on a graph neural network, including:
[0035] The first processing module is used to extract the binding site information corresponding to each protein in the dataset;
[0036] The second processing module is used to convert each protein in the dataset into graph structure data according to the PDB file of the protein and the binding site information, and obtain a protein graph;
[0037] The third processing module is used to construct a graph neural network model, input the obtained protein graph into the model, in the graph neural network, use the message passing mechanism to update the features of nodes, and input the features of nodes into a classifier for amino acid residue prediction;
[0038] The fourth processing module is used to compare the amino acid residue types predicted by the model with the wild-type residue labels and calculate the loss value;
[0039] The fifth processing module is used for backpropagation, further learning the model parameters until the loss value is stable, and saving the best parameters of the model after the training ends;
[0040] The sixth processing module is used to input the target protein into the trained model, compare the amino acid residues predicted by the model with the wild-type amino acid residues at the corresponding sites. If they are different from the wild-type and there are differences in properties, it is considered that a feasible mutation may occur at this site.
[0041] In the third aspect of the present invention, an electronic device is provided, including a processor and a memory communicatively connected to the processor and used for storing instructions executable by the processor. The processor is used to execute the above-mentioned method for predicting protein mutation sites based on a graph neural network.
[0042] In the fourth aspect of the present invention, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the above-mentioned method for predicting protein mutation sites based on a graph neural network is implemented.
[0043] Compared with the prior art, the method for predicting protein mutation sites based on a graph neural network of the present invention has the following advantages:
[0044] 1. Precise extraction of binding site information: Binding sites are usually key regions where proteins interact with other molecules. By extracting the binding site information of proteins, the characteristics of regions related to protein functions can be captured more precisely, which helps to evaluate the impact of a certain mutation on protein functions.
[0045] 2. Structural modeling of proteins: By extracting node features and edge features from PDB files, converting protein structure information into graph structure data, and making full use of its three-dimensional spatial information and the interactions between amino acids. Compared with traditional methods based on one-dimensional sequences, it can capture the spatial relationships and functional regions of amino acid residues in proteins more precisely, enabling the model to better understand the folding characteristics and functional sites of proteins.
[0046] 3. Biological interpretability of the model: The introduction of binding site information can help the model focus on regions of the protein that are crucial for its function. When the predicted site is near the binding site, it may affect the binding of the protein to the ligand. Additionally, the information transfer mechanism of the graph neural network allows node features to propagate in the network and interact with the information of neighboring nodes, enabling the analysis of which neighboring nodes have a significant impact during the prediction process of a certain mutation site. Description of the Drawings
[0047] The drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0048] Figure 1 It is a schematic diagram of the protein modeling graph structure data of the present invention;
[0049] Figure 2 It is a structural diagram of the protein mutation site prediction network model based on the graph neural network of the present invention;
[0050] Figure 3 It is a flowchart of the model's operation. Detailed Embodiments
[0051] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0052] In the description of the present invention, it should be understood that the orientation or positional relationships indicated by the terms "center", "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. are based on the orientation or positional relationships shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present invention. In addition, the terms "first", "second", etc. are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise specified, the meaning of "a plurality" is two or more.
[0053] In the description of the present invention, it should be noted that unless otherwise clearly specified and defined, the terms "installation", "connection", and "linkage" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood through specific circumstances.
[0054] The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0055] Embodiment 1:
[0056] The present invention proposes a new idea for extracting protein node features. Proteins interact with other molecules through binding sites. Therefore, the binding site region usually involves important parts of protein functions and has evolutionary conservation. To comprehensively consider the influence of this region, binding site information is added to the node features. The specific method is to query the PDB file of the binding site corresponding to each protein in the dataset, and query whether the residue belongs to the binding site region by comparing the information such as the three-dimensional coordinates of C α atoms, amino acid types, polypeptide chains, etc. In addition, when extracting the node features of amino acid residues, the local environment of the target residue is added, including geometric features (such as the relative positions of surrounding residues) and chemical background (amino acid types of neighboring residues, etc.) information to describe the dependence of amino acid residues on their local environment. Generally speaking, each protein node (i.e., amino acid residue) constructs its feature vector through binding site information, spatial geometric structure, and chemical background, including amino acid type, temperature factor, occupancy, atomic coordinates, binding site information, and neighboring amino acid types. Among them, the node feature extraction method of the present invention provides necessary biological information for the following graph structure construction. These information are the feature vectors of each node, helping the model understand the chemical and physical properties of proteins.
[0057] The present invention converts the proteins in the dataset into graph structure data. Using graph structure data to represent proteins can better capture the complex interactions and dependencies between residues in the protein structure. Among them, amino acid residues are used as nodes, and edges are established in the graph according to the Euclidean distance between residues, retaining the spatial information between residues. As Figure 1 shown, it shows an example of the local environment and graph structure of the 148th amino acid ALA of protein 1A1B. By converting proteins into graph structure data, it helps the model understand the sequence and structure information of proteins and perform tasks such as amino acid prediction and protein function prediction more effectively. Specifically, it includes the following:
[0058] Read the protein file in PDB format in the dataset, including the whole protein and its binding pocket information. Use chemical information processing tools to read the sequence and structure information contained therein, convert it into graph-structured data, map amino acids to nodes in the graph, and establish edges in the graph according to the positional relationship.
[0059] Map the three-dimensional protein structure data to graph-structured data, and generate a DGL graph for each protein. Each amino acid corresponds to a node in the graph, and the node features contain the chemical and physical information of the amino acid residue, including the amino acid type, C α atom coordinates, temperature factor, occupancy, binding site information, and adopt one-hot encoding. The position of the residue is determined by the three-dimensional coordinates of the C α atom. The distance between two residues is calculated according to the Euclidean distance between two C α atoms, as shown in formula (1), where d is the Euclidean distance between two C α atoms, and x, y, and z are the three spatial coordinates. When the condition that the distance is less than 7 Å is satisfied, an undirected edge is established in the graph.
[0060]
[0061] Collect the graph representations of each protein into a graph dataset. Each graph contains nodes, edges, and corresponding feature information, providing data support for subsequent model training.
[0062] The graph-structured data of the present invention provides input for the prediction method, enabling the graph neural network to learn based on this structural information and capture the spatial relationships in the protein and the interactions between amino acids.
[0063] Based on the above content, the present invention proposes a Figure 2 protein mutation site prediction method based on a graph neural network as shown. The output of this method is the predicted structure of the protein mutation site, which is the core output of the whole method and is based on the structural information obtained in the first two steps to identify mutations that may affect protein function.
[0064] The specific steps of this method are as follows:
[0065] 1) According to the PDB file of the protein and the binding site information, use chemical information processing tools to perform structural analysis on the three-dimensional structure of each protein, convert it into the form of a DGL graph, and calculate the node features and edge features therein to form a dataset.
[0066] 2) Input the obtained protein graph into the model. In the graph neural network, use the Message Passing mechanism to update the features of nodes. Specifically, for a specific node v, the message passing formula at time step t+1 is:
[0067]
[0068] Where, is the information of all neighbor nodes received by node v at time step t+1, N(v) is the set of neighbor nodes of node v, and represent the features of node v and neighbor node ω at time step t respectively, e vω is the edge feature between node v and ω, M t is the message function, which is used to convert the information of neighbor nodes into messages that can be received by the current node.
[0069] 3) After receiving the messages from neighbor nodes, node v will update its own features according to the received messages, as shown in formula (3):
[0070]
[0071] Where, U t is the update function. Through the above message passing mechanism, node v will continuously receive the status information from neighbor nodes and update its own status according to this information.
[0072] 4) After the above process, the network extracts the features of nodes and inputs them into the classifier for amino acid residue prediction.
[0073] 5) Compare the type of amino acid residue predicted by the model with the wild-type residue label and calculate the loss value.
[0074] 6) Backpropagation, further learn the model parameters until the loss value is stable, and save the best parameters of the model after training.
[0075] 7) Input the target protein into the trained model. Considering the powerful learning ability of the model and the dependence of residues on the surrounding environment, compare the amino acid residues predicted by the model with the wild-type amino acid residues at the corresponding sites. If they are different from the wild-type and there are differences in properties (for example, the wild-type amino acid is a neutral amino acid while the predicted amino acid is an acidic amino acid), then it is considered that a feasible mutation may occur at this site.
[0076] The entire model consists of a total of 5 layers of networks, including an input layer that receives protein graph data, which contains node features and edge features. There are 3 graph convolutional layers that process the information from the input layer, aggregate the information of neighboring nodes, and update the node features, gradually enhancing the network's understanding of the global structure. The last layer is the output layer, which predicts whether each node is a mutation site and the possible direction of the mutation through a classifier. Figure 3 It shows the running process of the entire model.
[0077] The method of the present invention is trained and tested on the PDBbind dataset, and the network model parameters are set as shown in Table 1.
[0078] Table 1
[0079]
[0080] The present invention uses an NVIDIA 3090 graphics card with a cuda version above 11.4 for training. The total number of training times is set to 50, the learning rate is 1e-2, 16 proteins are input in one batch, the optimizer uses Adam, the node feature dimension is 30, and the number of classifications is 21, including 20 common amino acid types and unknown amino acids.
[0081] The dataset used in the training and testing of the present invention is the PDBbind dataset, which contains the binding data of protein-ligand complexes. Each protein contains the information of the protein source PDB file, binding site information, and ligand small molecule information. 14,129 valid data are extracted from the dataset as the basis of the dataset, and the dataset is finally divided into a training set, a validation set, and a test set according to the ratio of 8:1:1.
[0082] The obtained test accuracy reaches 99.26%. When inputting the target protein, the model can output the predicted potential mutation sites, and an example is shown in Table 2.
[0083] Table 2
[0084]
[0085] Example two:
[0086] A device for predicting protein mutation sites based on a graph neural network, comprising:
[0087] A first processing module for extracting the binding site information corresponding to each protein in the dataset;
[0088] A second processing module for converting each protein in the dataset into graph structure data according to the PDB file of the protein and the binding site information to obtain a protein graph;
[0089] The third processing module is used to construct a graph neural network model, input the obtained protein graph into the model. In the graph neural network, a message passing mechanism is used to update the features of nodes, and the features of the nodes are input into a classifier for amino acid residue prediction;
[0090] The fourth processing module is used to compare the amino acid residue types predicted by the model with the wild-type residue labels and calculate the loss value;
[0091] The fifth processing module is used for backpropagation, further learning the model parameters until the loss value is stable, and saving the best parameters of the model after the training is completed;
[0092] The sixth processing module is used to input a target protein into the trained model, compare the amino acid residues predicted by the model with the wild-type amino acid residues at the corresponding sites. If they are different from the wild-type and there are differences in properties, it is considered that a feasible mutation may occur at this site.
[0093] Example 3:
[0094] An electronic device includes a processor and a memory communicatively connected to the processor and used to store instructions executable by the processor. The processor is used to execute the above-mentioned method for predicting protein mutation sites based on a graph neural network.
[0095] Example 4:
[0096] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the above-mentioned method for predicting protein mutation sites based on a graph neural network.
[0097] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.
Claims
1. A protein mutation site prediction method based on graph neural network, characterized in that: The steps include: Step 1: Extract the binding site information corresponding to each protein in the dataset; Step 2: According to the protein's PDB file and binding site information, each protein in the data set is converted into graph structure data to obtain a protein graph; Step 3: Build a graph neural network model and input the obtained protein graph into the model. In the graph neural network, use the message passing mechanism to update the features of the nodes, and input the features of the nodes into the classifier for amino acid residue prediction. Step 4: Compare the amino acid residue types predicted by the model with the wild-type residue labels and calculate the loss value; Step 5: Back propagation, further learning the model parameters until the loss value is stable, and saving the optimal parameters of the model after training; Step 6: Input the target protein into the trained model and compare the amino acid residues predicted by the model with the wild-type amino acid residues at the corresponding sites. If they are different from the wild type and have different properties, it is considered that the site may have a feasible mutation.
2. A method for predicting protein mutation sites based on graph neural network according to claim 1, characterized in that: In step 1, each protein node constructs its feature vector through binding site information, spatial geometric structure and chemical background, including amino acid type, temperature factor, occupancy, atomic coordinates, binding site information, and adjacent amino acid type.
3. The method for predicting protein mutation sites based on graph neural network according to claim 1, characterized in that: The step 2 comprises the following steps: Step 201: read the PDB format protein file in the data set, including the protein as a whole and its binding pocket information, use chemical information processing tools to read the sequence and structure information contained therein, convert it into graph structure data, map amino acids to nodes in the graph, and establish edges in the graph according to positional relationships; Step 202: Map the protein three-dimensional structure data to graph structure data. One protein corresponds to a DGL graph. Each amino acid corresponds to a node in the graph. The node features contain chemical and physical information about the amino acid residues, including amino acid type, physicochemical properties, C α Atomic coordinates, temperature factors, occupancy, binding site information, one-hot encoding, residue positions in C α The three-dimensional coordinates of the atoms are determined, and the distance between two residues is determined by the two C α The Euclidean distance between atoms is calculated as shown in formula (1), where d is the distance between two residues C α The Euclidean distance between atoms, x, y, and z are three spatial coordinates. When the distance is less than 7 angstroms, an undirected edge is established in the graph; Step 203: Collect the graph representation of each protein into a graph dataset, where each graph contains nodes, edges and corresponding feature information.
4. The method for predicting protein mutation sites based on graph neural network according to claim 1, characterized in that: In step 3, the features of updating the node using the message passing mechanism include: For a specific node v, the message passing formula at time t+1 is: in, is the information of all neighbor nodes received by node v at time step t+1, N(v) is the set of neighbor nodes of node v, and denote the features of node v and neighbor node ω at time step t, respectively, and e vω is the edge feature between nodes v and ω, M t It is a message function, which is used to convert the information of neighbor nodes into messages that can be received by the current node; After receiving the message from the neighbor node, node v will update its own features according to the received message, as shown in formula (3): Among them, U t is an update function. Through the above message passing mechanism, node v will continuously receive status information from neighboring nodes and update its own status based on this information.
5. A protein mutation site prediction device based on graph neural network, characterized in that: include: The first processing module is used to extract the binding site information corresponding to each protein in the data set; The second processing module is used to convert each protein in the data set into graph structure data according to the PDB file and binding site information of the protein to obtain a protein graph; The third processing module is used to build a graph neural network model, input the obtained protein graph into the model, use the message passing mechanism to update the features of the nodes in the graph neural network, and input the features of the nodes into the classifier for amino acid residue prediction; The fourth processing module is used to compare the amino acid residue type predicted by the model with the wild-type residue label and calculate the loss value; The fifth processing module is used for back propagation to further learn the model parameters until the loss value is stable. After the training is completed, the optimal parameters of the model are saved; The sixth processing module is used to input the target protein into the trained model, and compare the amino acid residues predicted by the model with the wild-type amino acid residues at the corresponding sites. If they are different from the wild type and have different properties, it is considered that the site is likely to produce a feasible mutation.
6. An electronic device, comprising a processor and a memory connected to the processor for storing instructions executable by the processor, characterized in that: The processor is used to execute a protein mutation site prediction method based on a graph neural network as described in any one of claims 1 to 4.
7. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for predicting protein mutation sites based on a graph neural network according to any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Protein and nucleic acid binding site prediction method based on graph neural network characterization
CN114765063A
Rigid body protein docking method based on isotropic graph neural network
CN116312752A
Data processing method and device for virus protein mutation prediction
CN116631507A
Data processing method and apparatus for virus protein mutation prediction
US20240404629A1
Cited By
Efficient protein stability prediction method for selective state space modeling
CN120496640A