Prediction Method for Mutation Effects of Metal Binding Sites Based on Graph Panoramic Attention Enhancement
By constructing a neural network framework with panoramic attention enhancement, fusing protein language model and metalloprotein structure information, the problem of low prediction accuracy in the existing methods is solved, and more accurate prediction of mutation effects of metal binding sites is achieved.
Patent Information
- Application Number
- CN202510368369.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-03-27
AI Technical Summary
The existing prediction methods for the mutation effect of metal binding sites fail to fully explore the deep semantic information in protein language models, lack learning of metal protein structure, and it is difficult to capture the complexity of different metal coordination structure patterns, resulting in low prediction accuracy.
A neural network framework based on panoramic attention enhancement is constructed, and the protein language model embedding and metalloprotein structure information are fused through a two-way feedback mechanism, and the metalloprotein structure is deeply perceived by the panoramic attention network, and the panoramic message delivery path is constructed by combining multiple attention-driven units, deep fusion structure and sequence listing.
It improves the accuracy and generalization ability of predicting mutation effects of metal binding sites, enhances the learning of metal protein structure, can better capture complex structural patterns, and inhibits prediction results that do not conform to biological common sense.
Smart Images

Figure CN119889441B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of bioinformatics, and in particular, to a method for predicting the mutation effect of metal binding sites based on enhanced graph panoramic attention. Background Art
[0002] It is estimated that more than 25% of proteins need to interact with one or more metal ions to perform their biological functions within cells, such as electron transfer, molecular recognition, gene transcription, and expression regulation. Missense mutations at metal binding sites caused by single nucleotide changes can hinder protein-metal interactions and disrupt the structure of metalloproteins, thereby impairing the function of the cellular protein network and potentially leading to disease development and affecting drug resistance. Therefore, accurately predicting disease-related mutations at metal binding sites is crucial for understanding the molecular mechanisms of many diseases and developing effective treatment strategies.
[0003] Currently, existing methods for predicting the mutation effect of metal binding sites are mainly divided into two categories: traditional methods and deep learning methods. Traditional methods combine various machine learning algorithms, including Naive Bayes, Support Vector Machine, and Random Forest, etc. However, these methods rely heavily on manually extracted features, and due to their limited ability to learn from given data, it becomes challenging to further improve.
[0004] In recent years, deep learning methods have been widely applied in the research direction of mutation effect prediction due to their efficient feature extraction ability and good feature expression ability. They mainly include predicting by extracting sequence information in protein language model embeddings through simple CNN or Transformer encoders, and predicting by learning the structural context around mutation sites through neural networks.
[0005] The above two deep learning models mainly have the following problems: 1. The first model does not fully exploit the deeper semantic information in the protein language model and lacks the learning of the structure of metalloproteins, while the structure of proteins is precisely an important factor affecting the realization of protein biological functions. 2. The second model is difficult to capture the inherent complexity of different metal coordination structure patterns due to the lack of comprehensive message aggregation of the environment around metal binding sites.
[0006] Therefore, how to better integrate the sequence information in protein language model embeddings and the complex structural information of metal protein coordination to endow deep learning models with more powerful feature expression ability to improve the accuracy of mutation effect prediction remains a difficult problem before us. Summary of the Invention
[0007] To solve the above technical problems existing in the prior art, the present invention provides a method for predicting the mutation effect of metal binding sites based on enhanced graph panoramic attention. This method deeply explores the rich semantic information contained in the protein language model embeddings, drives the construction of a graph panoramic attention network to deeply perceive the metal protein structure, and through a two-way feedback guided fusion mechanism, can effectively fuse the structural and sequence representations, realize comprehensive information utilization, make the final feature representation powerful and generalizable, and improve the accuracy of predicting the mutation effect of metal binding sites.
[0008] The technical solution of the present invention to solve the above technical problems is: A method for predicting the mutation effect of metal binding sites based on enhanced graph panoramic attention, comprising the following steps:
[0009] S1: Obtain disease-related mutation data of metal binding sites from the database;
[0010] S2: Divide the obtained disease-related mutation data set of metal binding sites into a training sample set, a validation sample set, and a test sample set;
[0011] S3: Perform data preprocessing operations on the sample sets divided in step S2; The data preprocessing operations include constructing a metal coordination network graph, encoding biological feature information, and obtaining deep semantic embeddings of the protein language model;
[0012] S4: Use the preprocessed training sample set and validation sample set obtained in step S3 to train and validate the neural network framework, and save the neural network framework after training is completed;
[0013] S5: Input the preprocessed test sample set obtained in step S3 into the neural network framework that has been trained in step S4 to obtain the prediction result of the mutation effect of metal binding sites.
[0014] Further, when constructing the metal coordination network graph in step S3, for the sample sets divided in step S2, construct an amino acid graph of the metal binding site network according to the graphing method of metal ions, ligands, and ligand neighbors, and construct an atomic graph according to a threshold of distance 3;
[0015] When encoding biological feature information in step S3, encode biological feature information including residue depth, secondary structure, relative solvent accessibility surface area, coordination bond, and hydrogen bond;
[0016] When obtaining the deep semantic embeddings of the protein language model in step S3, obtain the embeddings of the 33rd layer of the ESM2 protein language model esm2_t33_650M_UR50D() on the long sequence and map them to the corresponding amino acids on the metalpdb structure to obtain the ESM2 semantic embeddings.
[0017] Furthermore, the neural network framework in step S4 is a protein language model-embedded enhanced graph panoramic attention learning framework, specifically as follows:
[0018] In the atomic view attention aggregation module part of the framework, the preprocessed atomic graph and amino acid graph are input, and the structural features of the neighborhood around the metal binding site are extracted from the perspectives of atoms and amino acids, and the feature embeddings of atoms and amino acids are integrated through the self-attention mechanism of atomic clusters;
[0019] In the graph panoramic attention perception enhancement module part of the framework, the network weaves four attention-driven modules: node-to-node, node-to-edge, edge-to-edge, and edge-to-node to construct a panoramic message passing path. At the same time, each driving module injects deep semantic information from the protein language model to enhance the learning of the graph structure, and the structural information is iteratively fed back to the ESM encoding module to continuously optimize the representation of the sequence information;
[0020] In the self-distillation knowledge part of the framework, the graph structure features with mutated sites masked and the ESM2 embeddings obtained from the amino acid sequence after the amino acids at the mutated sites are replaced with "X" are input into the atomic view attention aggregation module and the graph panoramic attention perception enhancement module with normal gradient backpropagation; and a target feature generator is constructed, and the input of the target feature generator is the graph structure features and the ESM2 embeddings obtained from the normal sequence; the parameters of the target feature generator are updated through exponential moving average;
[0021] In the mutated effect prediction part of the framework, the feature representation after fusing the sequence and structure obtained in the encoding part passes through a linear layer of (D, 21) to obtain the predicted probabilities of amino acid types before and after mutation, calculate their maximum likelihood values, and finally pass through a linear layer and a sigmoid function to obtain the probability of mutated pathogenicity.
[0022] Furthermore, in the graph panoramic attention perception enhancement module of the protein language model-embedded enhanced graph panoramic attention learning framework:
[0023] In the node-to-node module, use the formula to perform preliminary feature fusion on the node feature matrix updated by the atomic view attention aggregation module and the downsampled ESM2 semantic embedding; where, represents element-wise addition at the corresponding positions, represents the node feature matrix preliminarily updated by the atomic view attention aggregation module; represents the downsampled ESM2 protein language model semantic embedding;
[0024] Use the formula Concatenate the three-hop information of the node matrix, where the three-hop information is zero-hop, one-hop, and two-hop information respectively; among them, represents the node matrix of the first-order adjacency matrix, represents the node matrix of the second-order adjacency matrix, represents the fully connected layer plus the ReLU linear rectifier activation function;
[0025] Construct a complete connection graph for the three-hop information of all nodes, and use the formula to obtain the updated three-hop feature representation of node i; among them, represents the unupdated three-hop feature representation of node i, represents the dimension of the node feature vector, represents the learnable weight matrix, represents the transpose of the matrix;
[0026] Use the formula ; Get the node feature matrix after attention aggregation; among them, represents the zero-hop feature matrix obtained by multi-hop information interaction, represents the first-hop and second-hop feature matrices obtained by multi-hop information interaction, represents the learnable weight matrix;
[0027] Furthermore, in the graph panoramic attention perception enhancement module of the graph panoramic attention learning framework enhanced by protein language model embedding:
[0028] In the ESM encoding module one, use the formula to obtain the Q, K, and V matrices of the initial semantic embedding of ESM2; where E represents the ESM2 semantic embedding after dimensionality reduction by the fully connected layer; represents the learnable weight matrix;
[0029] Use the formula to induce the structural bias and update the ESM2 feature representation; where represents the node structure feature matrix updated by the node-to-node module, represents the dimension of the feature vector in the matrix, represents the transpose of the matrix;
[0030] Furthermore, in the graph panoramic attention perception enhancement module of the graph panoramic attention learning framework enhanced by protein language model embedding:
[0031] In the node-to-edge module, use the formula to obtain the self-attention weight coefficient after concatenating the features of node i and the outgoing edge ij; where Denotes the feature concatenation operation, Denotes the feature representation of node i, Denotes the edge Feature representation of, Denotes a learnable weight matrix;
[0032] Using the formula Obtain the updated edge Feature representation; Denotes a learnable weight matrix;
[0033] Using the formula Weight the semantic information in the ESM2 embedding to guide the encoding of edge features in the graph structure; where, Denotes the feature representations of nodes i and j updated via the ESM encoding module 1, Denotes the outer product of vectors, Denotes taking the average row by row;
[0034] Using the formula Obtain the edge Feature representation after the node-to-edge module update, and the feature matrix after all edges are updated is denoted as .
[0035] Furthermore, in the graph panoramic attention perception enhancement module of the graph panoramic attention learning framework enhanced by the protein language model embedding:
[0036] In the ESM encoding module 2, using the formula Obtain the Q, K, and V matrices of the ESM feature representation after the first update ; where, Denotes a learnable weight matrix;
[0037] Using the formula Generalize the structural information of all incoming edges of the corresponding node; where, Denotes the edge Feature representation after being updated by the node-to-edge module, Denotes the set of all incoming edges of node i, Denotes concatenation by row, Denotes a learnable weight matrix;
[0038] Using the formula Generalize the edge structure information and feedback it to the sequence encoding branch for attention update of the ESM2 feature representation to obtain the ESM feature representation after the second update; where Denotes The dimension of the feature vector in the matrix, Denotes the transpose of the matrix.
[0039] Furthermore, in the graph panoramic attention perception enhancement module based on the protein language model embedding enhanced graph panoramic attention learning framework:
[0040] In the edge-to-edge module, using the formula , , obtain the K and V matrices of all incoming edges of node i , and the Q matrix of a certain outgoing edge ; where, represents the relative angle between adjacent edges and , Onehot() means dividing the relative angle range of into 6 small intervals, corresponding to 6-dimensional one-hot encoding, represents the learnable weight matrix, represents the feature representation of the edge after being updated by the node-to-edge module;
[0041] Using the formula plan the features of the ESM2 embedded virtual edge to construct , and formed triangular perception; where, and respectively represent the feature representations of nodes x and j updated by the ESM encoding module two;
[0042] Using the formula inject the triangular perception into the calculation of the attention coefficient between edges and ; represents the dimension of the matrix eigenvector;
[0043] Using the formula and obtain the feature representation of the aggregated and updated edge ; where, represents the set of all incoming edges of the node, and the feature matrix after updating all edges is represented as .
[0044] Furthermore, in the graph panoramic attention perception enhancement module based on the protein language model embedding enhanced graph panoramic attention learning framework:
[0045] In the edge-to-node module, using the formula obtain the Q matrix of node j, the K matrix of node i, and the V matrix of edge ; where, Denote the feature representations of node j and node i after being updated via the node-to-node module; Denote the edge after being updated via the edge-to-edge module Feature representation;
[0046] Using the formula and Obtain the importance of the incoming edge for node j; where, Denote The dimension of the eigenvector in the matrix, Denote the set of all neighbor nodes of node i.
[0047] Using the formula The updated semantic features of the ESM encoding module three To guide the convergence update of the features of node j; where, Denote the set of all incoming edges of node j, and the updated feature matrix of all nodes is denoted as ;
[0048] In the bidirectional cross-attention fusion module, using the formula and Obtain the structure features and sequence features after cross-attention update; where, Denote the node features after being updated via the edge-to-node module, Denote the sequence semantic features after being updated via the ESM encoding module three, Denote the learnable weight matrix, Denote the transpose of the matrix, Denote The dimension of the eigenvector after the matrix undergoes a linear transformation;
[0049] Using the formula Enhance the non-linear expression of the sequence features. Similarly, using the formula Enhance the non-linearity of the structure representation; Using the formula Obtain the final representation after fusing the structure and sequence features; where Denote the learnable weight matrix; Denote the bias term.
[0050] Furthermore, in the prediction part of the graph panoramic attention learning framework based on protein language model embedding enhancement, using the loss formula Maximize the probability that the predicted amino acid at the masked mutation site is the wild-type amino acid type, and minimize the probability that the predicted amino acid at the mutation site is the pathogenic amino acid type; where represents the probability that the amino acid type is predicted to be the wild-type amino acid, represents the probability that the amino acid at the same mutation site in the training set is predicted to be the pathogenic amino acid type, represents the number of mutant samples at the same site in the same metalpdb structure in the training set.
[0051] The beneficial effects of the present invention are as follows:
[0052] 1. The graph panoramic attention perception enhancement module creatively constructed by the present invention based on the protein language model and the graph panoramic attention mechanism fully exploits the rich semantic features contained in the protein language model embedding to enhance the learning of the metal protein structure.
[0053] 2. By combining multiple attention-driven units in the present invention to construct a panoramic message passing path, the network can capture complex structural patterns under different metal-binding conformations. The framework seamlessly integrates structural information at the residue and atomic levels, deeply fuses structural and sequence representations, and the comprehensive information induction better solves the problem of difficult characterization of the complex coordination patterns between metal ions and proteins in different metal coordination environments.
[0054] 3. Through the newly designed loss in the present invention, the probability that the mutation site is predicted to be the wild-type amino acid type is maximized, and the probability that the site is predicted to be the pathogenic amino acid type is minimized, thereby enhancing the network's reasoning ability regarding the true amino acid identity of the mutation site, suppressing prediction results that are contrary to biological common sense, and better solving the problem of inaccurate prediction of the mutation effect at the metal-binding site. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 is the overall flowchart of the present invention;
[0056] Figure 2 is the structural schematic diagram of the graph panoramic attention learning framework based on protein language model embedding enhancement in the present invention;
[0057] Figure 3 is the schematic diagram of the graph panoramic attention perception enhancement module of the graph panoramic attention learning framework based on protein language model embedding enhancement in the present invention;
[0058] Figure 4 is the schematic diagram of the graph panoramic attention module of the graph panoramic attention perception enhancement module part of the graph panoramic attention learning framework based on protein language model embedding enhancement in the present invention;
[0059] Figure 5 is the schematic diagram of the bidirectional cross-attention fusion module of the graph panoramic attention perception enhancement module part of the graph panoramic attention learning framework based on protein language model embedding enhancement in the present invention;
[0060] Figure 6 is Figure 2 an enlarged view of the masked structural information in
[0061] Figure 7 is Figure 2 an enlarged view of the unmasked structural information in Detailed implementation manners
[0062] Next, in combination with the attached Figures 1-7 drawings and specific embodiments, the implementation manners of the present invention will be further described in detail.
[0063] A method for predicting the mutation effect of metal binding sites based on graph panoramic attention enhancement, specifically a method for predicting the mutation effect of metal binding sites based on protein language model-guided graph structure learning, as Figure 1 shown in the flowchart in
[0064] S1: Obtain metal binding site disease-related mutation data from the Clinvar and Uniprot databases. Specifically:
[0065] Download the variant summary.txt.gz file in the Clinvar database, download humsavar.txt in the Uniprot Humsavar database, obtain all single missense variants therein, download the xml file on the sifts website that records the correspondence between Uniprot and pdb positions, and after processing, obtain a complete mapping file. Using this mapping file, from the obtained disease-related mutation data, screen out the samples with determined metalpdb structures as the final metal binding site disease-related mutation dataset. All single missense variants are obtained through the ID mapping online conversion function on the Uniprot website to obtain the corresponding UniprotID and pdbID.
[0066] Then, through the total mapping file, find the corresponding position in the pdb of the mutation position on Uniprot. In this step, multiple pdb files may be found, and each pdb file may correspond to multiple metalpdb files. Finally, according to the pdb resolution (the lower the value, the higher the priority) and the metalpdb coverage rate (the higher the value, the higher the priority), find the metalpdb with the highest priority, that is, find the local metal protein structure of the mutation sample.
[0067] S2: Divide the obtained disease-related mutation dataset on the metal binding sites into a training sample set, a validation sample set, and a test sample set. Specifically:
[0068] After balancing the positive and negative samples in the metal binding site mutation data set obtained in step S1, it is randomly divided into a training sample set, a validation sample set, and a test sample set according to the ratio of 0.8:0.1:0.1.
[0069] S3: Perform data preprocessing operations on the sample set divided in step S2. The data preprocessing operations include constructing a metal coordination network diagram, encoding biological feature information, and obtaining the deep semantic embedding of the protein language model.
[0070] When constructing the metal coordination network diagram, for the sample set divided in step S2, construct an amino acid diagram of the metal binding site network according to the graph construction method of metal ions, ligands, and ligand neighbors, and construct an atomic diagram according to a threshold of a distance of 3. Specifically, for the amino acid diagram, the metal ion is represented as the central node, and edges are added between the central node and all metal ligands, where the metal ligand is defined as a residue or cofactor containing at least one non-hydrogen atom within 3 Å of the metal ion. In addition, edges are added between the metal ligand and the ligand neighbor, and the ligand neighbor is defined as a residue or cofactor containing at least one non-hydrogen atom within 5 Å of the metal ligand. For the atomic diagram, all atoms are used as nodes, and edges are added between two atoms within a distance of 3 Å to simulate the interaction between atoms.
[0071] When encoding biological feature information, encode biological feature information including residue depth, secondary structure, relative solvent accessibility surface area, coordination bond, and hydrogen bond. Specifically, the node features of the amino acid diagram include residue type, residue depth, secondary structure, relative solvent accessibility surface area, whether it is a coordination or acceptor of a hydrogen bond, and whether it is a ligand or acceptor of a coordination bond; the edge features include the length and angle of the edge (composed of the polar angle and azimuth angle in spherical coordinates); the node features of the atomic diagram include atom type, the number of covalent bonds formed by the atom, the number of electron charges assigned to the atom, chirality state, hybridization state, van der Waals radius, and atomic mass; the encoding of the edge features is the same as that of the amino acid Figure 1 consistent.
[0072] When obtaining the deep semantic embedding of the protein language model, obtain the embedding of the 33rd layer of the ESM2 protein language model esm2_t33_650M_UR50D() on the long sequence and map it to the corresponding amino acids on the metalpdb structure to obtain the ESM2 semantic embedding. Specifically, first traverse the metalpdb to obtain several pdb chains contained therein, and then use the mapping file constructed in step S1 to find the corresponding complete uniprot single chain; according to the pre-trained esm2_t33_650M_UR50D() model weights provided by facebookresearch, obtain the semantic feature embeddings of all unipro chains at the 33rd layer of the network, and finally map them back to the amino acids on the incomplete pdb chains in the metalpdb to obtain the ESM2 semantic embedding.
[0073] S4: Use the preprocessed training sample set and validation sample set obtained in step S3 to train and validate the neural network model, and save the neural network model framework after training. Input the preprocessed training sample set obtained in step S3 into the protein language model-embedded enhanced graph panoramic attention learning framework for training. Adopt an early stopping strategy. Stop the training process when the model's loss does not decrease for 30 consecutive iterations on the validation set or when the number of iterations reaches the set value of 300, and save the protein language model-embedded enhanced graph panoramic attention learning framework with the best mutation effect prediction performance on the validation sample set.
[0074] The neural network framework in step S4 is the protein language model-embedded enhanced graph panoramic attention learning framework, specifically:
[0075] In the atomic view attention aggregation module part of the framework, input the preprocessed atomic graph and amino acid graph, and conduct detailed structural feature extraction on the neighborhood around the metal binding site from the perspectives of atoms and amino acids. Through the self-attention mechanism of the atomic cluster, the problem of inconsistent dimensions between atoms and amino acids can be compensated, thus seamlessly integrating the feature embeddings of both.
[0076] In the graph panoramic attention perception enhancement module part of the framework, the network will weave four attention-driven modules: node-to-node, node-to-edge, edge-to-edge, and edge-to-node to construct a panoramic message passing path to comprehensively represent the complex structure of the neighborhood around the metal binding site. During this process, deep semantic sequence information from the protein language model will be injected into each driving module to guide the learning of the graph structure. At the same time, the iteration of the structural information will also be fed back to the ESM encoding module to continuously optimize the representation of the sequence information. Through the bidirectional guidance and fusion mechanism, it can help the network better learn the graph structure and deeply fuse the sequence and structure representations, ensuring more accurate prediction using more comprehensive information.
[0077] In the self-distillation knowledge part of the framework, input the graph structure features with the mutation site masked and the ESM2 embedding obtained after replacing the amino acid at the mutation site with "X" into the atomic view attention aggregation module and the graph panoramic attention perception enhancement module with normal gradient backpropagation. In addition, an additional target feature generator is constructed, and its input is the complete graph structure feature and the ESM2 embedding obtained from the normal sequence. After message passing, alignment will be performed in the feature space to enable the encoder to obtain stronger information reasoning ability. The target feature generator updates its parameters through exponential moving average (EMA).
[0078] In the mutation effect prediction part of the framework, the feature representation obtained by fusing the sequences and structures acquired in the coding part is passed through a linear layer of (D, 21) to obtain the predicted probabilities of amino acid types before and after mutation, calculate their maximum likelihood values, and finally, through a linear layer and a sigmoid function, the probability of mutation pathogenicity can be obtained.
[0079] See Figure 2 , in the atomic view attention aggregation module part of the framework for enhancing graph panoramic attention learning based on protein language model embedding, first, the preprocessed training sample set obtained in step S3 is input into the atomic view attention aggregation module part, and the formula is used to obtain the attention weights between adjacent atoms; where represents the initial feature representation of atoms i and j, represents the initial feature representation of the atomic graph edge , represents the dimension of the atomic feature representation, represents the attention coefficient between atoms i and j, represents matrix transpose, represents the learnable weight matrix.
[0080] For atom i, the feature representation after attention aggregation is obtained through the formula ; where represents the number of attention heads; represents the learnable weight matrix.
[0081] All updated atomic features are represented by .
[0082] The feature matrix of amino acids is preliminarily updated using the formula , where represents the initial amino acid biological feature encoding matrix; is the learnable weight matrix.
[0083] The self-attention weight coefficients of atoms are obtained using the formula among all atoms belonging to the same amino acid; where is the learnable weight matrix.
[0084] Using the formula to align the atomic scale and the amino acid scale to generate the atom graph feature representation optimized by attention guidance;
[0085] Using the formula to integrate the structural information of different scales of amino acids and atoms.
[0086] See Figure 3, in the graph panoramic attention perception enhancement module of the encoding part of the graph panoramic attention learning framework enhanced by protein language model embedding in the present invention, specifically: see Figure 4 (a), in the node-to-node module of the graph panoramic attention perception enhancement module of the graph panoramic attention learning framework enhanced by protein language model embedding, using the formula to perform preliminary feature fusion on the node feature matrix preliminarily updated by the atomic view attention aggregation module and the ESM2 semantic embedding after dimensionality reduction, and strengthen the overall information representation. Among them, represents element-wise addition at the corresponding position, represents the node feature matrix preliminarily updated by the atomic view attention aggregation module; represents the ESM2 protein language model semantic embedding after dimensionality reduction.
[0087] Using the formula to splice the three-hop information of the node matrix, and the three-hop information is the zero-hop, one-hop, and two-hop information respectively; among them, represents the node matrix 's first-order adjacency matrix; represents the node matrix 's second-order adjacency matrix; represents the fully connected layer plus the ReLU linear rectifier activation function.
[0088] Construct a fully connected graph with the three-hop information of all nodes, and use the formula to obtain the updated three-hop feature representation of node i; among them, represents the unupdated three-hop feature representation of node i, represents the dimension of the node feature vector, represents the learnable weight matrix, represents the transpose of the matrix. By replacing the interaction between nodes with the interaction between multi-hop information, the defect that the node features are overly smoothed as the number of network layers increases in the graph neural network is alleviated.
[0089] Using the formula ; to obtain the node feature matrix after attention aggregation; among them, represents the zero-hop feature matrix obtained by multi-hop information interaction, represents the first-hop and second-hop feature matrices obtained by multi-hop information interaction, represents the learnable weight matrix. By using the non-linear interaction between multi-hop neighbor representations to drive the information transmission between nodes, the network can obtain representative embeddings of complex patterns around metal ions and metal binding sites.
[0090] In the ESM encoding module 1, using the formula Obtain the Q, K, and V matrices of the initial semantic embedding of ESM2; where E represents the semantic embedding of ESM2 after dimensionality reduction by the fully connected layer; represents the learnable weight matrix.
[0091] Use the formula Induce the structural bias and update the ESM2 feature representation; where represents the node structure feature matrix updated by the node-to-node module, represents the dimension of the eigenvector in the matrix, represents the transpose of the matrix.
[0092] See Figure 4 (b), the node-to-edge module in the panoramic attention perception enhancement module of the graph, use the formula to obtain the self-attention weight coefficient after concatenating the features of node i and the outgoing edge ij; where represents the feature concatenation operation, represents the feature representation of node i, represents the edge 's feature representation, represents the learnable weight matrix.
[0093] Use the formula to obtain the updated edge feature representation; represents the learnable weight matrix.
[0094] Use the formula to weight the rich semantic information in the ESM2 embedding for guiding the optimization of the edge feature encoding in the graph structure; where, represents the feature representation of nodes i and j updated by the ESM encoding module one, represents the outer product of vectors, represents taking the average row by row.
[0095] Use the formula to obtain the edge feature representation updated by the node-to-edge module. This module effectively models the information interaction between nodes and edges without losing the inherent information of the edges. The feature matrix of all updated edges is represented as .
[0096] In the ESM encoding module two, use the formula to obtain the Q, K, and V matrices of the ESM feature representation after the first update; where, represents the learnable weight matrix.
[0097] Use the formula Summarize the structural information of all incoming edges of the corresponding node; among them, represents the edge after being updated by the node-to-edge module Feature representation, represents the set of all incoming edges of node i, represents splicing by row, represents the learnable weight matrix.
[0098] Use the formula Summarize the edge structure information and feedback it to the sequence encoding branch for attention update of the ESM2 feature representation to obtain the second updated ESM feature representation; among them represents the dimension of the eigenvector in the matrix, represents the transpose of the matrix.
[0099] See Figure 4 (c), the edge-to-edge module in the panoramic attention perception enhancement module of the graph,
[0100] Use the formula , , to obtain the K and V matrices of all incoming edges of node i and the Q matrix of a certain outgoing edge ; among them, represents the adjacent edge and the relative angle between them, Onehot() means dividing the relative angle range of into 6 small intervals, corresponding to a 6-dimensional one-hot encoding, represents the learnable weight matrix, represents the edge after being updated by the node-to-edge module Feature representation.
[0101] Use the formula to plan the features of the ESM2 embedded virtual edge to construct , and the formed triangular perception; among them, and respectively represent the feature representations of nodes x and j updated by the ESM encoding module two.
[0102] Use the formula to inject the triangular perception bias into the calculation of the attention coefficient between the edges and ; represents the dimension of the eigenvector of the matrix.
[0103] Using the formula and to obtain the edges after aggregated update Feature representation; where represents the set of all incoming edges of the node, and the feature matrix of all edges after update is represented as . This module conducts message passing and update from the perspective of edges, which helps to identify structural details that may be overlooked when learning graph representation only from the node perspective.
[0104] See Figure 4 (d), the edge-to-node module in the graph panoramic attention perception enhancement module. In the edge-to-node module, using the formula to obtain the Q matrix of node j, the K matrix of node i, and the V matrix of edge ; where, represents the feature representations of node j and node i after being updated via the node-to-node module; represents the feature representation of edge after being updated via the edge-to-edge module; represents the learnable weight matrix.
[0105] Using the formula and to obtain the importance degree of the incoming edge for node j; where, represents the dimension of the eigenvector in the matrix; represents the set of all neighbor nodes of node i.
[0106] Using the formula to guide the convergent update of the features of node j with the updated semantic features from the ESM encoding module III; where, represents the set of all incoming edges of node j, and the feature matrix of all nodes after update is represented as . This module effectively aggregates the key features from different edge structure representations to generate more robust graph embeddings.
[0107] See Figure 5 , the bidirectional cross-attention fusion module in the graph panoramic attention perception enhancement module. Specifically: using the formula and to obtain the structure features and sequence features after cross-attention update; where, represents the node features after being updated via the edge-to-node module, represents the sequence semantic features after being updated via the ESM encoding module III, represents the learnable weight matrix, denotes the transpose of a matrix, denotes the dimension of the eigenvector after the matrix undergoes a linear transformation;
[0108] Using the formula to enhance the non - linear expression of sequence features. Similarly, using the formula to enhance the non - linearity of structural representations; Using the formula to obtain the final representation after fusing the structure and sequence features; where denotes a learnable weight matrix; denotes the bias term, further enhancing the non - linearity of the structure and sequence features to better fit the complex correlations between different types of features.
[0109] Based on the protein language model embedding to enhance the self - distillation module part in the graph panoramic attention learning framework. Specifically: Input the graph structure features with masked mutation sites, and the ESM2 embedding obtained from the amino acid sequence where the amino acids at the mutation sites are replaced by 'X' into the framework encoding block composed of the atomic view attention aggregation module and the graph panoramic attention perception enhancement module. Input the unmasked graph structure features and the ESM2 embedding obtained from the normal sequence into the target feature generator network, which shares the same architecture as the encoding block.
[0110] Using the formula to align the feature outputs of the encoding block and the target feature generator network in the feature space, aiming to strengthen the information reasoning ability of the network and the learning ability regarding the global feature distribution; where, represents the feature output of the encoding block, represents the feature output of the target feature generator; N represents the batch size; denotes the sensitivity hyperparameter. Using the formula to update the parameters of the target feature generator network; where, are the parameters of the atomic view attention aggregation module of the framework encoder and the guided graph panoramic attention learning module; is the weight decay coefficient, and the weights of the encoding block are updated by normal gradient backpropagation.
[0111] In the mutation effect prediction part of the graph panoramic attention learning framework enhanced by protein language model embedding, input the sequence and structure fusion features obtained from the encoding block into a fully - connected layer of dimension (d, 21) for subsequent prediction. The specific prediction process is as follows:
[0112] Using the formula to obtain the probabilities that the amino acid type at the mutation position i is predicted to be the wild - type amino acid and the mutated amino acid. The log - likelihood ratio between the two is used for variant effect prediction, where, represents the probability that the amino acid type at position i is predicted to be the wild-type amino acid, represents the probability that the amino acid type at position i is predicted to be the mutated amino acid type a.
[0113] For the calculated log-likelihood ratio score a fully connected layer and a sigmoid activation layer are used to obtain the probability of developing a disease after the amino acid at position i mutates from the wild-type amino acid to amino acid a.
[0114] Using the loss formula maximize the probability that the predicted amino acid type at the masked mutation site is the wild-type amino acid, and minimize the probability that the predicted amino acid type at the mutation site is the pathogenic amino acid type, where represents the probability that the amino acid type is predicted to be the wild-type amino acid, represents the probability that the amino acid at the same mutation site in the training set is predicted to be the pathogenic amino acid type, represents the number of mutant samples at the same site in the same metalpdb structure in the training set, which is in line with the evolutionary conservation of the amino acid sequence. Long-term natural selection determines that the protein sequence distribution has a certain statistical law. By designing this loss function, the network can be optimized in a direction more consistent with the biological background setting.
[0115] Obtain the weighted binary cross-entropy loss of the predicted probability value , and combine the weighted scaled cosine error loss and the weighted loss designed, obtain the overall loss value of the neural network model, and optimize the model parameters through the adaptive gradient method (AdamW optimizer) to make the loss continuously approach the minimum value to train the neural network model.
[0116] Input the preprocessed validation sample set obtained in step S3 into the above-mentioned trained protein language model embedding enhanced graph panoramic attention learning framework, and use the common classification metrics AUROC, AUPRC, ACC, F1-score, MCC, PPV, Recall to calculate the classification performance of the validation sample set, and save the information related to the parameters of the protein language model embedding enhanced graph panoramic attention learning framework when the classification performance on the validation sample set is the best.
[0117] S5: Input the preprocessed test sample set obtained in step S3 into the neural network model trained in step S4 to obtain the prediction result of the mutation effect of the metal binding site.
[0118] The above is only used to illustrate the design concept and implementation plan of the present invention, rather than limiting it. Those skilled in the art should understand that other solutions obtained by modifying the technical solution of the present invention or making equivalent substitutions are still within the scope defined by the claims of this application.
Claims
1. A method for predicting the effect of metal binding site mutations based on enhanced image panorama attention, characterized in that: The following steps are involved: S1: Obtain disease-related mutation data of metal binding sites from the database; S2: Divide the acquired disease-related mutation dataset on metal binding sites into a training sample set, a validation sample set, and a test sample set; S3: performing data preprocessing operations on the sample set divided in step S2; the data preprocessing operations include constructing a metal coordination network diagram, encoding biological feature information, and obtaining deep semantic embedding of a protein language model; S4: Use the preprocessed training sample set and verification sample set obtained in step S3 to train and verify the neural network framework, and save the trained neural network framework; S5: inputting the preprocessed test sample set obtained in step S3 into the neural network framework trained in step S4 to obtain the prediction result of the metal binding site mutation effect; In step S3, when constructing the metal coordination network diagram, the sample set divided in step S2 is constructed according to the composition mode of metal ions, ligands and ligand neighbors to construct an amino acid diagram of the metal binding site network, and an atomic diagram is constructed according to a threshold value of 3; In step S3, when encoding the biometric information, the biometric information including residue depth, secondary structure, relative solvent accessible surface area, coordination bonds, and hydrogen bonds is encoded; In step S3, when obtaining the deep semantic embedding of the protein language model, the embedding of the 33rd layer of the ESM2 protein language model esm2_t33_650M_UR50D() on the long sequence is obtained and mapped to the corresponding amino acids on the metalpdb structure to obtain the ESM2 semantic embedding; The neural network framework in step S4 is based on the protein language model embedding enhanced graph panoramic attention learning framework, specifically: In the atomic view attention aggregation module of the framework, the pre-processed atomic graph and amino acid graph are input to extract the structural features of the neighborhood around the metal binding site from both atomic and amino acid perspectives, and the feature embedding of atoms and amino acids is integrated through the self-attention mechanism of atomic clusters. In the graph-wide attention perception enhancement module of the framework, the network weaves four attention-driven modules, namely node-to-node, node-to-edge, edge-to-edge, and edge-to-node, to build a panoramic message delivery path. At the same time, each driver module injects deep semantic information from the protein language model to enhance the learning of graph structure, and iteratively feeds back structural information to the ESM encoding module to continuously optimize the representation of sequence information. In the self-distillation knowledge part of the framework, the ESM2 embedding obtained from the masked graph structure features of the mutation site and the amino acid sequence after the amino acid at the mutation site is replaced with "X" is input into the atomic view attention aggregation module and the graph panoramic attention perception enhancement module of the normal gradient feedback; and a target feature generator is constructed, and the input of the target feature generator is the graph structure features and the ESM2 embedding obtained from the normal sequence; The target feature generator updates parameters through exponential moving average; In the mutation effect prediction part of the framework, the feature representation of the fusion of the sequence and structure obtained in the coding part is passed through a linear layer (D, 21) to obtain the predicted probability of the amino acid type before and after the mutation, and its maximum likelihood value is calculated. Finally, it passes through a linear layer and a sigmoid function to obtain the probability of the pathogenicity of the mutation.
2. The method for predicting the effect of metal binding site mutations based on enhanced image panorama attention according to claim 1, characterized in that: In the graph panoramic attention perception enhancement module of the enhanced graph panoramic attention learning framework based on protein language model embedding: In the node-to-node module, use the formula The node feature matrix updated by the atomic view attention aggregation module and the ESM2 semantic embedding after dimensionality reduction are used for preliminary feature fusion; Indicates the addition of elements at corresponding positions, G n represents the node feature matrix initially updated by the atomic view attention aggregation module; E represents the semantic embedding of the ESM2 protein language model after dimensionality reduction; Using the formula The three-hop information of the concatenated node matrix is zero-hop, one-hop, and two-hop information; A represents the node matrix The first-order adjacency matrix, A 2 Represents the node matrix The second-order adjacency matrix of , f() represents the fully connected layer plus the ReLU linear rectification activation function; The three-hop information of all nodes constructs a fully connected graph, using the formula Get the updated three-hop feature representation of node i; where F i represents the three-hop feature representation of the unupdated node i, d1 represents the dimension of the node feature vector, and W Q , W K , W V represents the learnable weight matrix, and T represents the transpose of the matrix; Using the formula α 1,2 =Softmax(LeakyRelu(F′ 0 W0+F′ 1,2 W 1,2 )); Get the node feature matrix after attention aggregation; where F′ 0 Indicates the zeroth hop feature matrix of multi-hop information interaction, F′ 1,2 It represents the first and second hop feature matrices of multi-hop information interaction, W0, W 1,2 represents the learnable weight matrix.
3. The method for predicting the effect of metal binding site mutations based on enhanced image panorama attention according to claim 2, characterized in that: In the graph panoramic attention perception enhancement module of the enhanced graph panoramic attention learning framework based on protein language model embedding: In ESM encoding module 1, using formula Q 1 =EW q1 , K 1 =EW k1 , V 1 =EW v1 The Q, K, and V matrices of the initial semantic embedding of ESM2 are obtained; among them, E represents the semantic embedding of ESM2 after dimensionality reduction by the fully connected layer; W q1 , W k1 , W v1 represents the learnable weight matrix; Using the formula Inductive structural bias and update ESM2 feature representation; represents the node structure feature matrix after being updated by the node-to-node module, and d2 represents Q 1 , K 1 , V 1 The dimension of the eigenvectors in the matrix. T represents the transpose of the matrix.
4. The method for predicting the effect of metal binding site mutations based on enhanced image panorama attention according to claim 3, characterized in that: In the graph panoramic attention perception enhancement module of the enhanced graph panoramic attention learning framework based on protein language model embedding: In the node-to-edge module, using the formula Get the self-attention weight coefficient after concatenating the features of node i and outgoing edge ij; where || represents the feature concatenation operation, h i represents the feature representation of node i, Represents edge The feature representation of W ne represents the learnable weight matrix; Using the formula Get the updated edge Feature representation; W ne2 represents the learnable weight matrix; Using the formula The semantic information in the ESM2 embedding is weighted to guide the encoding of edge features in the graph structure; represents the feature representation of node i,j after being updated by ESM encoding module 1, represents the vector outer product, and mean() represents the row-wise average; Using the formula Get the edge updated by the node-to-edge module Feature representation, the feature matrix after all edges are updated is expressed as 5. The method for predicting the effect of metal binding site mutations based on enhanced image panorama attention according to claim 4, characterized in that: In the graph panoramic attention perception enhancement module of the enhanced graph panoramic attention learning framework based on protein language model embedding: In ESM encoding module 2, using formula O 2 =E 1 W q2 , K 2 =E 1 W k2 , V 2 =E 1 W v2 Get the ESM feature representation E after the first update 1 Q, K, V matrices; where W q2 , W k2 , W v2 represents the learnable weight matrix; Using the formula Summarize the structural information of all incoming edges of the corresponding node; among them, Represents the edge after the node-to-edge module updates Feature representation, YI represents the set of all incoming edges of node i, Concat() represents row concatenation, W β represents the learnable weight matrix; Using the formula The summarized edge structure information is fed back to the sequence encoding branch for attention update of ESM2 feature representation to obtain the second updated ESM feature representation; where d3 represents Q 2 ,K 2 ,V 2 The dimension of the eigenvectors in the matrix. T represents the transpose of the matrix.
6. The method for predicting the effect of metal binding site mutations based on enhanced image panorama attention according to claim 5, characterized in that: In the graph panoramic attention perception enhancement module of the enhanced graph panoramic attention learning framework based on protein language model embedding: In the edge-to-edge module, Using the formula Get all incoming edges of node i The K, V matrix of an outgoing edge The Q matrix of xij Represents adjacent edges and Onehot() means dividing the relative angle range of 0 to 180° into 6 small intervals, corresponding to the 6-dimensional onehot encoding, W qe ,W ke ,W ve ,W angle represents the learnable weight matrix, Represents the edge after the node-to-edge module updates Feature representation; Using the formula Planning ESM2 Embedded Virtual Edge features to construct and The triangular perception formed; among them, and They respectively represent the feature representations of nodes x and j after being updated by ESM encoding module 2; Using the formula Injecting triangle-aware bias into edges and Calculation of attention coefficient between; d4 represents The dimensions of the matrix eigenvectors; Using the formula and Get the updated edge after aggregation Feature representation; where XI represents the set of all incoming edges of the node, and the updated feature matrix of all edges is expressed as 7. The method for predicting the effect of metal binding site mutations based on enhanced image panorama attention according to claim 6, characterized in that: In the graph panoramic attention perception enhancement module of the enhanced graph panoramic attention learning framework based on protein language model embedding: In the edge-to-node module, use the formula Get the Q matrix of node j, the K matrix of node i, and the edge V matrix; where h j ,h i represents the feature representation of node j and node i after being updated by the node-to-node module; Represents the edge updated by the edge-to-edge module The characteristic representation of W qn , W kn , W ve represents the learnable weight matrix; Using the formula and Get the edge The importance of node j; d5 represents Q j , K i The dimension of the eigenvectors in the matrix, V nei Represents the set of all neighbor nodes of node i; Using the formula The updated semantic features of ESM encoding module 3 To guide the convergence update of node j features; where IJ represents the set of all incoming edges of node j, and the updated feature matrix of all nodes is expressed as In the bidirectional cross attention fusion module, using formula and Get the updated structural features after cross attention and sequence feature E 3* ;in, represents the node features updated by the edge-to-node module, E 3 represents the updated sequence semantic features after ESM encoding module 3, W c1 ,W c2 ,W c3 ,W c4 ,W c5 ,W c6 represents the learnable weight matrix, T represents the transpose of the matrix, and d6 represents E 3 The dimension of the eigenvector of the matrix after linear transformation; Using Formula E 3*′ =LayerNorm(E 3 +max(0,E 3 *W c7 +b1)W c8 +b2) enhances the nonlinear expression of sequence features. Similarly, Using the formula Enhance the nonlinearity of structural representation; using the formula The final structure and sequence feature fusion representation is obtained; where W c7 , W c8 , W c9 , W c10 , W c11 represents the learnable weight matrix; b1, b2, b3, b4 represent bias terms.
8. The method for predicting the effect of metal binding site mutations based on enhanced image panorama attention according to claim 1, characterized in that: In the prediction part of the protein language model embedding-enhanced graph panoramic attention learning framework, the loss formula is used Maximize the probability that the amino acid predicted by the masked mutation site is the wild-type amino acid type, and minimize the probability that the amino acid predicted by the mutation site is the pathogenic amino acid type; where p ref represents the probability that the amino acid type is predicted to be a wild-type amino acid, represents the probability that the amino acid at the same mutation site in the training set is predicted to be the pathogenic amino acid type, and k represents the number of mutation samples at the same site in the same Metalpdb structure in the training set.
Citation Information
Patent Citations
Protein torsion angle prediction method based on embedded features and dynamic convolutional network
CN117577169A
Methods and systems for predicting protein function
CN1328601A