Epidermal growth factor receptor pathogenicity prediction method, device, server and medium
By using graph neural networks to hierarchically aggregate EGFR mutation sites and generate protein structure feature vectors, the problem of inaccurate assessment of the pathogenicity of multiple mutation sites in existing technologies is solved, enabling more precise pathogenicity assessment and personalized treatment plans.
Patent Information
- Application Number
- CN202411923792.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2044-12-25
AI Technical Summary
Existing technologies cannot accurately capture the spatial relationships between multiple EGFR mutation sites and their combined impact on protein structure, resulting in an inability to accurately assess the pathogenicity of mutation sites.
By extracting amino acid residue features, a graph neural network is used to hierarchically aggregate amino acid residues, amino acid clusters, and protein subunits to generate protein structural feature vectors, thereby assessing the pathogenicity of mutation sites.
It enables accurate pathogenicity assessment of multiple EGFR mutation sites, improves the accuracy and comprehensiveness of prediction, and provides a scientific and reliable basis for clinical treatment.
Smart Images

Figure CN119832984B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of epidermal growth factor receptor technology, and more particularly to a method, device, server, and medium for predicting the pathogenicity of epidermal growth factor receptor. Background Technology
[0002] Epidermal growth factor receptor (EGFR) is an important member of the tyrosine kinase receptor family, playing a crucial role in maintaining cellular physiological functions. However, EGFR is susceptible to mutations due to multiple factors, including environmental, genetic, and lifestyle factors, becoming a contributing factor to various cancers and posing a threat to human health. Currently, tyrosine kinase inhibitors (TKIs) are the main targeted drugs for cancer patients with EGFR mutations, controlling tumor growth by inhibiting EGFR activity. However, the diversity and continuous dynamic changes of EGFR mutations present challenges to targeted therapy.
[0003] Predicting potential EGFR mutation sites and assessing their pathogenicity can provide more precise decision support for clinical treatment. Currently, models such as graph neural networks are used to predict the protein structure of a single mutation site and assess its pathogenicity. However, when predicting protein structural variations caused by complex mutations involving multiple mutation sites, it is impossible to accurately capture and predict the spatial relationships between multiple mutation sites and their combined impact on protein structure, thus failing to accurately assess the pathogenicity of the mutation sites. Summary of the Invention
[0004] This invention provides a method, device, server, and medium for predicting the pathogenicity of epidermal growth factor receptor (EGFR), which can accurately predict the impact of multiple EGFR mutation sites on protein structure, thereby accurately assessing the pathogenicity of mutation sites.
[0005] In a first aspect, embodiments of the present invention provide a method for predicting the pathogenicity of epidermal growth factor receptors, comprising:
[0006] Based on the three-dimensional spatial conformation of a complex mutant amino acid sequence containing multiple mutation sites, amino acid residue features are extracted. These features include: amino acid residue type features, amino acid residue centroid coordinate features, key residue features, and dihedral features. The amino acid residue features are in matrix form.
[0007] Based on the amino acid residue features, an initial amino acid residue structure map is generated, wherein the features of each node in the initial amino acid residue structure map are the corresponding amino acid residue features; the initial amino acid residue structure map is input into a first graph neural network that has been trained, and the amino acid residue features of each node are updated using the convolutional layer of the first graph neural network to obtain the predicted amino acid residue feature matrix of each node.
[0008] Based on the predicted amino acid residue feature matrix of each node, the amino acid residues are clustered by DBSCAN based on spatial proximity. The predicted amino acid residue feature matrices of all the predicted amino acid residues in each category are weighted and aggregated to generate the amino acid cluster feature matrix of each category.
[0009] An initial amino acid cluster structure map is generated based on the amino acid cluster feature matrix. The features of each node in the initial amino acid cluster structure map are the corresponding amino acid cluster feature matrix. The initial amino acid cluster structure map is input into the trained second graph neural network. The convolutional layer of the second graph neural network is used to update the amino acid cluster feature matrix of each node to obtain the predicted amino acid cluster feature matrix of each node.
[0010] Based on the predicted amino acid cluster feature matrix, the amino acid cluster combination threshold is determined. The predicted amino acid cluster feature matrix is then hierarchically clustered using the amino acid cluster combination threshold. All predicted amino acid cluster feature matrices in each category are weighted and aggregated to generate multiple protein subunit feature matrices.
[0011] An initial protein subunit structure map is generated based on the protein subunit feature matrix. The features of each node in the initial protein subunit structure map are the corresponding protein subunit feature matrix. The initial protein subunit structure map is input into a trained third graph neural network. The convolutional layer of the third graph neural network is used to update the protein subunit features of each node to obtain the predicted protein subunit feature matrix of each node.
[0012] Feature mapping is performed on all predicted protein subunit feature matrices to generate predicted protein structure feature vectors. These predicted protein structure feature vectors are then used to assess the pathogenicity of mutation sites.
[0013] Furthermore, based on the predicted amino acid residue feature matrix of each node, the amino acid residues are clustered according to spatial proximity, and a weighted aggregation is performed on all predicted amino acid residue feature matrices in each category to generate an amino acid cluster feature matrix for each category, including:
[0014] Based on the predicted amino acid residue feature matrix of each node, the three-dimensional coordinates of the amino acid residues of each node are extracted;
[0015] DBSCAN clustering was performed based on the three-dimensional coordinates of amino acid residues at each node, with a neighborhood radius of 12 angstroms and a minimum number of points of 6.
[0016] Based on the spatial attention mechanism, amino acid residues located in the tyrosine kinase region are weighted, and the sum of the feature matrices of all predicted amino acid residues under each category is calculated to obtain the amino acid cluster feature matrix for each category.
[0017] Furthermore, based on the predicted amino acid cluster feature matrix, an amino acid cluster combination threshold is determined. This threshold is then used to perform hierarchical clustering of the predicted amino acid cluster feature matrix. Weighted aggregation of all predicted amino acid cluster feature matrices within each category generates multiple protein subunit feature matrices, including:
[0018] Based on the predicted amino acid cluster feature matrix, the maximum distance between the centroid coordinates of all amino acid clusters is calculated, and 12% of the maximum distance is used as the amino acid cluster combination threshold.
[0019] Based on the amino acid cluster combination threshold and the centroid coordinates of the amino acid clusters, hierarchical clustering is performed on the amino acid clusters of each node. Each category contains multiple predicted amino acid cluster feature matrices where the distance between the centroid coordinates of the amino acid clusters is less than the amino acid cluster combination threshold.
[0020] Based on the spatial attention mechanism, the feature matrices of all predicted amino acid clusters under each category are weighted, and the sum of the features of all predicted amino acid clusters under each category is calculated to obtain the protein subunit feature matrix of each category.
[0021] Furthermore, the extraction of amino acid residue features based on the three-dimensional spatial conformation of a complex mutant amino acid sequence containing multiple mutation sites includes:
[0022] The predicted highly pathogenic EGFR single-point mutation sites were obtained, and the highly pathogenic EGFR mutation sites were combined with sensitive mutation sites to obtain a set of predicted complex mutant amino acid sequences.
[0023] Each compound mutant amino acid sequence in the predicted set of compound mutant amino acid sequences is subjected to three-dimensional structure transformation to obtain the three-dimensional spatial conformation of the compound mutant amino acid sequence.
[0024] Traverse the three-dimensional spatial conformation of the complex mutant amino acid sequence, and extract the sequence features of all amino acid residues, the type features of amino acid residues, and the three-dimensional coordinates of all atoms in the complex mutant amino acid sequence;
[0025] Based on the amino acid residue type characteristics and amino acid residue sequence characteristics of each amino acid residue, the key residue characteristics of each amino acid residue are determined;
[0026] The centroid coordinate characteristics of each amino acid residue are determined based on the mean of the three-dimensional coordinates of all atoms of each amino acid residue.
[0027] The dihedral features of each amino acid residue are determined based on the three-dimensional coordinates of all atoms in each amino acid residue.
[0028] Furthermore, the method also includes:
[0029] The key residue features of each amino acid residue are identified, and the amino acid residue features of each node are weighted in the convolutional layer of the first graph neural network using an attention mechanism based on the key residue features.
[0030] Furthermore, the pathogenicity assessment of the predicted protein structure feature vector includes:
[0031] The predicted protein structure feature vectors are clustered, and the pathogenicity and drug sensitivity of unknown mutation sites in the same category as known mutation sites are assessed using protein structure features containing known mutation sites.
[0032] Furthermore, the method also includes:
[0033] Using the key amino acid residue features in the predicted amino acid residue features, the predicted amino acid residue feature matrix is clustered. Using the predicted amino acid residue features that include known mutation sites, the pathogenicity, pathogenic mechanism and drug sensitivity of unknown mutation sites in the same category are evaluated.
[0034] Secondly, embodiments of the present invention also provide an epidermal growth factor receptor pathogenicity prediction device, comprising:
[0035] The clinical mutation intersection analysis module is used to extract amino acid residue features based on the three-dimensional spatial conformation of the compound mutant amino acid sequence.
[0036] A point-to-surface conformation flexible fusion module is used to predict protein structural features based on amino acid residue characteristics;
[0037] The point-to-surface conformation flexible fusion module includes a first prediction unit, which is used to predict amino acid residue features using the convolutional layer of a first graph neural network;
[0038] An amino acid cluster feature generation unit is used to generate amino acid cluster features based on predicted amino acid residue features.
[0039] The second prediction unit is used to predict amino acid cluster features using the convolutional layers of the second graph neural network;
[0040] The protein subunit feature generation unit is used to generate protein subunit features based on predicted amino acid cluster features.
[0041] The third graph convolutional unit is used to predict protein subunit features using the convolutional layers of the third graph neural network;
[0042] The protein structure feature generation unit is used to generate protein structure features based on protein subunit features.
[0043] The point-to-surface conformation clustering dynamic evaluation module is used to assess the pathogenicity of mutation sites based on protein structural characteristics.
[0044] Thirdly, embodiments of the present invention also provide a server, comprising:
[0045] One or more processors;
[0046] Storage device for storing one or more programs.
[0047] When the one or more programs are executed by the one or more processors, the one or more processors implement the epidermal growth factor receptor pathogenicity prediction method as provided in the above embodiments.
[0048] Fourthly, embodiments of the present invention also provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the epidermal growth factor receptor pathogenicity prediction method provided in the above embodiments.
[0049] This invention provides a method, device, server, and medium for predicting the pathogenicity of epidermal growth factor receptor (EGFR). It extracts amino acid residue features from the three-dimensional spatial conformation of a complex mutant amino acid sequence, uses a graph neural network to predict these features, and aggregates the influence of mutation sites at the amino acid residue level. Then, based on the principle of spatial proximity, the amino acid residues are aggregated to form higher-level amino acid clusters. A graph neural network is used to predict these clusters, aggregating the influence of mutations into the higher-level amino acid cluster structure. Finally, the distance between the amino acid clusters is used to aggregate them into higher-level protein structural units—protein subunits. A graph neural network is used again to predict the protein subunits, mapping the prediction results to feature vectors of complete protein structures for protein structure classification. This technology can predict the impact of mutations at different protein structural levels, progressively aggregating the effects of mutation sites to the entire protein structure. It can more accurately simulate the impact of mutations on the overall protein structure, improving the accuracy and comprehensiveness of predictions. Furthermore, it can use the classified protein structures to assess the pathogenicity and drug sensitivity of mutation sites, providing an important reference for the study of the biological characteristics of rare or unknown mutations. This can provide a more scientific and reliable basis for clinical treatment, and is expected to bring more personalized treatment plans and better treatment outcomes to patients with EGFR-mutant cancers. Attached Figure Description
[0050] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an undue limitation of the invention. In the drawings:
[0051] Figure 1This is a flowchart of a method for predicting the pathogenicity of epidermal growth factor receptors according to Embodiment 1 of the present invention;
[0052] Figure 2 This is a flowchart of a method for predicting the pathogenicity of epidermal growth factor receptors according to Embodiment 2 of the present invention;
[0053] Figure 3 This is a flowchart of a method for predicting the pathogenicity of epidermal growth factor receptors according to Embodiment 3 of the present invention;
[0054] Figure 4 This is a flowchart of a method for predicting the pathogenicity of epidermal growth factor receptors according to Embodiment 4 of the present invention;
[0055] Figure 5 This is a schematic diagram of the structure of an epidermal growth factor receptor pathogenicity prediction device according to Embodiment Six of the present invention;
[0056] Figure 6 This is a schematic diagram of the structure of a server according to Embodiment 7 of the present invention. Detailed Implementation
[0057] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, the accompanying drawings show only the parts relevant to the present invention, and not all of the structures.
[0058] Example 1
[0059] Figure 1 This is a flowchart of a method for predicting the pathogenicity of epidermal growth factor receptor according to Embodiment 1 of the present invention. This embodiment predicts the impact of mutation sites on protein structure through a graph neural network, and aggregates the data stepwise from the amino acid residue level to the protein subbase level, ultimately forming a protein structure feature vector for evaluating the pathogenicity of mutation sites. The specific steps include:
[0060] S101. Based on the three-dimensional spatial conformation of a complex mutant amino acid sequence containing multiple mutation sites, amino acid residue features are extracted. The amino acid residue features include: amino acid residue type features, amino acid residue centroid coordinate features, key residue features, and dihedral features. The amino acid residue features are in matrix form.
[0061] Statistical analysis of mutation information revealed a specific mutation pattern: compound mutations tend to occur on top of two sensitive mutation sites, L858R and Del19 (exon 19 deletion), resulting in compound mutant amino acid sequences containing multiple mutation sites. An amino acid sequence represents the arrangement of multiple amino acids. When a compound mutation occurs, the amino acids at the corresponding mutation sites in the sequence change. Three-dimensional conformation conversion software can be used to generate a three-dimensional conformation of the amino acid sequence, i.e., the three-dimensional structure of the amino acid sequence. This three-dimensional conformation contains the three-dimensional structure of all amino acid sequences. By traversing this conformation, the spatial structural features of all amino acid residues can be extracted. These features can identify the type of amino acid residue and whether its location is a critical residue. The three-dimensional coordinates of all atoms in each amino acid residue can be used to derive the centroid coordinates and dihedral angles of the residues. These features collectively describe the characteristics of the amino acid residues, forming amino acid residue features.
[0062] For ease of calculation, the amino acid residue features are expressed in matrix form, including amino acid residue type features representing the type of amino acid residue, amino acid residue centroid coordinate features representing the centroid coordinates of amino acids, key residue features representing whether an amino acid residue is a key residue, and dihedral features representing the conformation of amino acid residues.
[0063] S102, Based on the amino acid residue features, an initial amino acid residue structure map is generated, wherein the features of each node in the initial amino acid residue structure map are the corresponding amino acid residue features; the initial amino acid residue structure map is input into the trained first graph neural network, and the convolutional layer of the first graph neural network is used to update the amino acid residue features of each node to obtain the predicted amino acid residue feature matrix of each node.
[0064] Amino acid residues are arranged according to their amino acid sequences. In the three-dimensional conformation generated from the amino acid sequences, there are specific positional relationships between the amino acid residues. Based on these positional relationships, an initial amino acid residue structure diagram is generated, with each amino acid residue serving as a node. The node positions in the diagram correspond to the positions of the amino acid residues, and the edges represent the interactions between them. The initial amino acid residue structure diagram is then updated using a first graph neural network. Convolutional layers are used to aggregate neighboring nodes, as shown in the following formula:
[0065]
[0066] in, It is the updated node feature matrix, σ is the activation function, such as ReLU, N(i) is the set of neighboring nodes of node i, and cij It is a normalization constant, which can be the edge feature e. ij The function can also be a constant 1, W (0) and b (0) It is a learnable weight matrix and bias vector, x j Let be the node feature matrix of the neighboring nodes of node i.
[0067] The updated nodes incorporate features from neighboring nodes, reflecting changes in amino acid residues caused by interacting adjacent amino acid residues, resulting in a predicted amino acid residue feature matrix for each node. During node updates, the features of mutated amino acid residues are aggregated into the features of adjacent amino acid residues, demonstrating the influence of the mutated site on other adjacent sites.
[0068] S103. Based on the predicted amino acid residue feature matrix of each node, the amino acid residues are clustered by DBSCAN based on spatial proximity. The predicted amino acid residue feature matrices of all categories are weighted and aggregated to generate the amino acid cluster feature matrix of each category.
[0069] Multiple spatially related amino acid residues often share a common characteristic. The predicted amino acid residue feature matrix, including those influenced by mutation sites, is classified based on spatial proximity. The DBSCAN clustering algorithm is used to classify the centroid coordinates of spatially similar amino acid residues according to their centroid coordinates, thus classifying the predicted amino acid residue features. All amino acid residues in each category are aggregated into amino acid clusters, which contain mutated or mutated amino acid residues. When aggregating the predicted amino acid residue features under each category, the tyrosine kinase region is crucial for EGFR signal transduction and is also a key target for tyrosine kinase inhibitors (TKIs). The three-dimensional structure of amino acid residues in this region has a critical impact on EGFR biological activity. Therefore, the structural features of amino acid residues located in the tyrosine kinase region from position 682 to 958 need to be emphasized. Thus, a weighting operation is used to increase the weight of amino acid residues in the tyrosine kinase region from position 682 to 958, enhancing their influence and generating the aggregated amino acid cluster feature matrix.
[0070] S104. Based on the amino acid cluster feature matrix, an initial amino acid cluster structure map is generated, wherein the features of each node in the initial amino acid cluster structure map are the corresponding amino acid cluster feature matrix; the initial amino acid cluster structure map is input into the trained second graph neural network, and the convolutional layer of the second graph neural network is used to update the amino acid cluster feature matrix of each node to obtain the predicted amino acid cluster feature matrix of each node.
[0071] Different amino acid clusters formed by the aggregation of multiple spatially proximate amino acid residues exhibit different characteristics, and each cluster is affected differently by mutations. The aggregated amino acid clusters also have relative spatial relationships. Based on these relationships, an initial amino acid cluster structure diagram is generated, with each cluster as a node. The node positions in the diagram correspond to the positions of the amino acid clusters, and the edges represent the interactions between the clusters. This initial amino acid cluster structure diagram is then updated using a second graph neural network. Convolutional layers are used to aggregate neighboring nodes for each node, calculated using the following formula:
[0072]
[0073] in, This represents the predicted amino acid cluster feature matrix, where σ is the activation function, such as ReLU, and N... cluster(k) Let c represent the set of neighboring nodes of node k. km It is the normalization constant, W (1) and b (1) It consists of a learnable weight matrix and a bias vector. Let be the node feature matrix of the neighboring nodes of node k.
[0074] The updated nodes incorporate features from neighboring nodes, reflecting the changes that occur when amino acid clusters are influenced by interacting neighboring amino acid clusters, resulting in a predicted amino acid cluster feature matrix for each node. During node updates, features from multiple mutation-affected amino acid clusters can be aggregated into the features of adjacent amino acid clusters, demonstrating the impact of mutated amino acid clusters on other neighboring amino acid clusters. By aggregating mutation-affected amino acid residues to form mutation-affected amino acid clusters, the influence between these clusters can be used to further realistically simulate the impact of mutations on a larger scale.
[0075] S105. Based on the predicted amino acid cluster feature matrix, determine the amino acid cluster combination threshold, use the amino acid cluster combination threshold to perform hierarchical clustering on the predicted amino acid cluster feature matrix, and perform weighted aggregation on all predicted amino acid cluster feature matrices in each category to generate multiple protein subunit feature matrices.
[0076] To account for the impact of mutations on protein structures at different levels, the centroid coordinates of amino acid clusters were extracted from the predicted amino acid cluster feature matrix. Hierarchical clustering was then performed using these centroid coordinates, with a distance threshold set to 12% of the maximum distance between the predicted amino acid cluster centroid coordinates. The predicted amino acid cluster feature matrix was then categorized, and all amino acid clusters within each category were aggregated to form larger structural units—protein subunits. Different weights were assigned to amino acid clusters based on their distances, and these weights were used to aggregate all predicted amino acid cluster feature matrices within each category into a protein subunit feature matrix.
[0077] S106. Based on the protein subunit feature matrix, an initial protein subunit structure map is generated, wherein the features of each node in the initial protein subunit structure map are the corresponding protein subunit feature matrices; the initial protein subunit structure map is input into the trained third graph neural network, and the protein subunit features of each node are updated using the convolutional layers of the third graph neural network to obtain the predicted protein subunit feature matrix of each node.
[0078] To comprehensively consider the impact of mutations on higher-level protein structures, an initial protein subunit structure map is generated using protein subunits as nodes, based on the positional relationships between the aggregated protein subunits. The node positions in the map correspond to the positions of the protein subunits, and the edges represent the interactions between protein subunits. The initial protein subunit structure map is then updated using a third-graph neural network. Convolutional layers are used to aggregate neighboring nodes for each node, and the calculation formula is as follows:
[0079]
[0080] in, This represents the predicted protein subunit feature matrix, where σ is the activation function, such as ReLU, and N is the value of N. Domain (n) represents the set of neighboring nodes, d np It is the normalization constant, W (2) and b (2) It consists of a learnable weight matrix and a bias vector. Let be the node feature matrix of the neighboring nodes of node n.
[0081] The updated nodes incorporate the features of their neighboring nodes, resulting in a predicted protein subunit feature matrix for each node. During node updates, features of protein subunits affected by mutations can be aggregated into the features of adjacent protein subunits, reflecting the higher-level impact of mutations on the overall protein structure.
[0082] S107, perform feature mapping on all predicted protein subunit feature matrices to generate predicted protein structure feature vectors, and use the predicted protein structure feature vectors to assess the pathogenicity of mutation sites.
[0083] The predicted protein subunit features are subjected to mean pooling to obtain a global feature vector g. Then, a fully connected layer is used for feature mapping to form a predicted protein structure feature vector y that expresses the characteristics of the complete protein structure, as shown in the following formula:
[0084] y=σ(W (3) g+b (3) )
[0085] Among them, W (3) It is the learnable weight matrix of the fully connected layer, b (3) σ is the bias vector of the fully connected layer, and σ is the sigmoid activation function.
[0086] Predicted protein structure feature vectors are clustered, and proteins within the same category exhibit similar structures. Predicted protein structures containing known mutation sites are identified using these feature vectors. Based on the characteristics of known mutations, the pathogenicity of predicted protein structures within the same category is assessed. Proteins with identical or similar structures often exhibit conformational consistency or high similarity. This structural homology leads to similar performance in biological characteristics such as drug sensitivity and pathogenicity. Therefore, known mutation site information can be used to effectively predict and assess the biological characteristics of unknown or rare mutation sites within the same category.
[0087] This embodiment utilizes a first graph neural network to update and predict amino acid residues, simulating the impact of mutations on amino acid residues. Based on spatial proximity, the predicted amino acid residue feature matrix is aggregated into an amino acid cluster feature matrix. A second graph neural network is then used to update and predict amino acid clusters, aggregating amino acid residues affected by mutations into amino acid clusters, simulating the impact of mutations at a higher level. The predicted amino acid cluster feature matrix affected by mutations is then aggregated into a higher-level protein subunit feature matrix. A third graph neural network is used to predict and update protein subunits, more accurately simulating the impact of mutations on different levels of protein structure. Finally, the predicted protein subunit feature matrices are integrated to form a complete protein feature vector. This complete protein feature vector is used for classification, using known mutation sites of the same type as references to assess the pathogenicity and drug sensitivity of unknown mutation sites of the same type. This approach more comprehensively considers the impact of mutations on adjacent structures at different levels of protein structure, resulting in more accurate mutation classification. It can serve as an important reference when studying the pathogenicity and drug sensitivity of unknown mutations.
[0088] In one optional implementation of this embodiment, to enhance the importance of key residues, the attention given to key residues is increased through weighting during the update of the convolutional layers of the first graph neural network. This is achieved by identifying key residue features for each amino acid residue and then weighting the amino acid residue features of each node in the convolutional layers of the first graph neural network using an attention mechanism based on these key residue features.
[0089] Key residues are amino acid residues located in critical functional regions of the EGFR kinase structure. Changes (mutations) in these amino acid residues can significantly affect kinase function and are closely related to the occurrence and development of disease. Key residue features are used to identify whether an amino acid residue is a key residue. When at least one of node i or its neighboring node j is a key residue, an attention mechanism is added for weighting when the nodes representing the two amino acid residues are aggregated. The formula for calculating the attention weight of the two nodes is as follows:
[0090]
[0091] Where exp represents the exponential function, LeakyReLU represents the activation function, and a is the learnable attention vector. T W is the transpose of vector a, ∥ denotes vector concatenation, and W (0) Let represent the learnable weight matrix, where K represents the set of key amino acid residues, and N(i) represents the set of neighboring nodes of node i. `if i∈K or j∈K` indicates that at least one of node i or its neighboring node j is a key residue; `0, otherwise` indicates other cases. That is, the node weight is determined by whether node i or its neighboring node j is a key residue; otherwise, the node weight is set to 'a'. ij It is 0.
[0092] Accordingly, the update formula for the convolutional layer of the neural network in the first graph is:
[0093]
[0094] By weighting key amino acid residues using an attention mechanism, we can more realistically simulate the impact of mutations on amino acid residues, especially when key amino acid residues change, and more accurately predict the effects of mutations.
[0095] Example 2
[0096] Figure 2This is a flowchart of a method for predicting the pathogenicity of epidermal growth factor receptor as described in Embodiment 2 of the present invention. This embodiment is an optimization based on the above embodiment. In this embodiment, based on the predicted amino acid residue feature matrix of each node, the amino acid residues are clustered by DBSCAN based on spatial proximity. The predicted amino acid residue feature matrices of all predicted amino acid residues in each category are weighted and aggregated to generate an amino acid cluster feature matrix for each category. Specifically, the optimization is as follows:
[0097] Based on the predicted amino acid residue feature matrix of each node, the three-dimensional coordinates of the amino acid residues of each node are extracted;
[0098] DBSCAN clustering was performed based on the three-dimensional coordinates of amino acid residues at each node, with a neighborhood radius of 12 angstroms and a minimum number of points of 6.
[0099] Based on the spatial attention mechanism, amino acid residues located in the tyrosine kinase region are weighted, and the sum of the feature matrices of all predicted amino acid residues under each category is calculated to obtain the amino acid cluster feature matrix for each category.
[0100] Accordingly, the epidermal growth factor receptor pathogenicity prediction method provided in this embodiment specifically includes:
[0101] S201, based on the three-dimensional spatial conformation of a complex mutant amino acid sequence containing multiple mutation sites, extract amino acid residue features, which include: amino acid residue type features, amino acid residue centroid coordinate features, key residue features and dihedral features, and the amino acid residue features are in matrix form.
[0102] S202, Based on the amino acid residue features, an initial amino acid residue structure map is generated, wherein the features of each node in the initial amino acid residue structure map are the corresponding amino acid residue features; the initial amino acid residue structure map is input into the trained first graph neural network, and the convolutional layer of the first graph neural network is used to update the amino acid residue features of each node to obtain the predicted amino acid residue feature matrix of each node.
[0103] S203, based on the predicted amino acid residue feature matrix of each node, extract the three-dimensional coordinates of the amino acid residues of each node.
[0104] The predicted amino acid residue feature matrix includes the aggregated and updated amino acid residue centroid coordinate features of each node. Based on the amino acid residue centroid coordinate features, the three-dimensional coordinates of the amino acid centroid are determined.
[0105] S204 uses DBSCAN clustering based on the three-dimensional coordinates of amino acid residues at each node, setting the neighborhood radius of DBSCAN to 12 angstroms and the minimum number of points to 6.
[0106] The distance between the centroids of any two amino acid residues can be calculated using the Euclidean formula to obtain the amino acid residue distance matrix. The calculation formula is as follows:
[0107]
[0108] Where distance is the distance between the centroids of two amino acid residues, (X i Y i Z i ) and (X j Y j Z j ) are the centroid coordinates of the two amino acid residues.
[0109] The DBSCAN algorithm was used to perform cluster analysis on amino acid residues in the protein structure. A distance matrix between amino acid residues was constructed, and the neighborhood radius (eps) was set to... (Å), with a minimum number of samples (min_samples) parameter of 6, to identify and classify spatially close and densely packed groups of amino acid residues.
[0110] S205, based on the spatial attention mechanism, weights the amino acid residues located in the tyrosine kinase region and calculates the sum of the feature matrices of all predicted amino acid residues under each category to obtain the feature matrix of amino acid clusters for each category.
[0111] Aggregate all amino acid residue features under each category into an amino acid cluster feature c. k Calculate the spatial attention weights between any two nodes i and j in each category. The calculation formula is as follows:
[0112]
[0113] Where, N spatial (k) represents the set of nodes (i.e., the set of amino acid residues) under each category, exp represents the exponential function, LeakyReLU scaled,l Let a be the activation function, and a be a learnable attention vector. T W is the transpose of vector a, ∥ denotes vector concatenation, and W (1) This represents the learnable weight matrix. and Let i and j represent the node feature matrices, respectively. This represents the node characteristics of all neighboring nodes.
[0114] Based on the amino acid residue sequence characteristics, the node features of the tyrosine kinase region located at positions 682 to 958 are scaled using the LeakyReLU function. A scaling factor β (β>1) is defined to increase the node weight of the amino acid residue features in the tyrosine kinase region, emphasizing the importance of these amino acid residues in the amino acid clusters.
[0115]
[0116] Where LeakyReLU(x) = max(0,x) + α·min(0,x) is the standard LeakyReLU activation function, and α is a positive number less than 1, usually taken as 0.01, used to control the slope when negative inputs are received.
[0117] Weights are assigned to nodes in each category based on the interaction between the two nodes corresponding to the distance between them. Multiple amino acid residues may exhibit new properties after combining into clusters. The tyrosine kinase region is a key site for TKI-targeted drug action, and the structure of amino acid residues significantly affects drug affinity. Therefore, it is necessary to increase the amino acid residue weights in the tyrosine kinase region during aggregation, utilizing spatial attention weights. To aggregate the feature vectors of all nodes within amino acid cluster k The characteristic matrix of amino acid clusters was obtained.
[0118]
[0119] in, It is the total attention weight of node j within cluster k.
[0120] S206. Based on the amino acid cluster feature matrix, an initial amino acid cluster structure map is generated, wherein the features of each node in the initial amino acid cluster structure map are the corresponding amino acid cluster feature matrix; the initial amino acid cluster structure map is input into the trained second graph neural network, and the convolutional layer of the second graph neural network is used to update the amino acid cluster feature matrix of each node to obtain the predicted amino acid cluster feature matrix of each node.
[0121] S207. Based on the predicted amino acid cluster feature matrix, determine the amino acid cluster combination threshold, use the amino acid cluster combination threshold to perform hierarchical clustering on the predicted amino acid cluster feature matrix, and perform weighted aggregation on all predicted amino acid cluster feature matrices in each category to generate multiple protein subunit feature matrices.
[0122] S208. Based on the protein subunit feature matrix, an initial protein subunit structure map is generated, wherein the features of each node in the initial protein subunit structure map are the corresponding protein subunit feature matrices; the initial protein subunit structure map is input into the trained third graph neural network, and the protein subunit features of each node are updated using the convolutional layers of the third graph neural network to obtain the predicted protein subunit feature matrix of each node.
[0123] S209 involves feature mapping of all predicted protein subunit feature matrices to generate predicted protein structure feature vectors, which are then used to assess the pathogenicity of mutation sites.
[0124] This embodiment classifies the updated predicted amino acid residue feature matrix using DBSCAN based on spatial proximity, aggregating all amino acid residues in the same category into amino acid clusters. By utilizing the amino acid residues affected by mutations to form a higher protein hierarchical structure, the impact of mutations is predicted from a higher protein structural level. At the same time, the spatial attention mechanism is used to weight the amino acid residues in the tyrosine kinase region, emphasizing the importance of amino acid residues in the tyrosine kinase region when aggregating amino acid residues into amino acid clusters, which can more realistically simulate the characteristics exhibited by amino acid clusters affected by mutations.
[0125] Example 3
[0126] Figure 3 This is a flowchart of a method for predicting the pathogenicity of epidermal growth factor receptor as described in Embodiment 3 of the present invention. This embodiment is an optimization based on the above embodiment. In this embodiment, an amino acid cluster combination threshold is determined according to the predicted amino acid cluster feature matrix. The predicted amino acid cluster feature matrix is then hierarchically clustered using the amino acid cluster combination threshold. All predicted amino acid cluster feature matrices in each category are weighted and aggregated to generate multiple protein subunit feature matrices. Specifically, the optimization is as follows:
[0127] Based on the predicted amino acid cluster feature matrix, the maximum distance between the centroid coordinates of all amino acid clusters is calculated, and 12% of the maximum distance is used as the amino acid cluster combination threshold.
[0128] Based on the amino acid cluster combination threshold and the centroid coordinates of the amino acid clusters, hierarchical clustering is performed on the amino acid clusters of each node. Each category contains multiple predicted amino acid cluster feature matrices where the distance between the centroid coordinates of the amino acid clusters is less than the amino acid cluster combination threshold.
[0129] Based on the spatial attention mechanism, the feature matrices of all predicted amino acid clusters under each category are weighted, and the sum of the features of all predicted amino acid clusters under each category is calculated to obtain the protein subunit feature matrix of each category.
[0130] Accordingly, the epidermal growth factor receptor pathogenicity prediction method provided in this embodiment specifically includes:
[0131] S301, based on the three-dimensional spatial conformation of a complex mutant amino acid sequence containing multiple mutation sites, extract amino acid residue features, which include: amino acid residue type features, amino acid residue centroid coordinate features, key residue features and dihedral features, and the amino acid residue features are in matrix form.
[0132] S302, Based on the amino acid residue features, an initial amino acid residue structure map is generated, wherein the features of each node in the initial amino acid residue structure map are the corresponding amino acid residue features; the initial amino acid residue structure map is input into the trained first graph neural network, and the convolutional layer of the first graph neural network is used to update the amino acid residue features of each node to obtain the predicted amino acid residue feature matrix of each node.
[0133] S303: Based on the predicted amino acid residue feature matrix of each node, the amino acid residues are clustered by DBSCAN based on spatial proximity. The predicted amino acid residue feature matrices of all categories are weighted and aggregated to generate the amino acid cluster feature matrix of each category.
[0134] S304. Based on the amino acid cluster feature matrix, an initial amino acid cluster structure map is generated, wherein the features of each node in the initial amino acid cluster structure map are the corresponding amino acid cluster feature matrix; the initial amino acid cluster structure map is input into the trained second graph neural network, and the convolutional layer of the second graph neural network is used to update the amino acid cluster feature matrix of each node to obtain the predicted amino acid cluster feature matrix of each node.
[0135] S305, Based on the predicted amino acid cluster feature matrix, calculate the maximum distance between the centroid coordinates of all amino acid clusters, and use 12% of the maximum distance as the amino acid cluster combination threshold.
[0136] The centroid coordinates of amino acid clusters are determined based on the centroid coordinate features in the predicted amino acid cluster feature matrix. The distance between the centroids of any two amino acid clusters is calculated using the Euclidean formula, resulting in the amino acid cluster distance matrix. Using 12% of the maximum distance between the centroids of any two amino acid clusters as the amino acid cluster combination threshold can help identify amino acid clusters with significant interactions.
[0137] S306: Based on the amino acid cluster combination threshold and the centroid coordinates of the amino acid clusters, hierarchical clustering is performed on the amino acid clusters of each node. Each category contains multiple predicted amino acid cluster feature matrices where the distance between the centroid coordinates of the amino acid clusters is less than the amino acid cluster combination threshold.
[0138] Using a threshold of 12% of the maximum distance between the centroids of amino acid clusters, hierarchical clustering is performed on the distance matrix of amino acid clusters, and the amino acid clusters in each category are combined to form protein subunits.
[0139] S307, based on the spatial attention mechanism, weights all predicted amino acid cluster feature matrices under each category and calculates the sum of all predicted amino acid cluster features for each category to obtain the protein subunit feature matrix for each category.
[0140] Aggregate all amino acid cluster features under each category into a protein subunit feature. k Calculate the spatial attention weights between any two amino acid clusters m and n in each category. The calculation formula is as follows:
[0141]
[0142] Where m, n∈N Domain (D k ), N Domain (D k ) represents the set of nodes for each category, exp represents the exponential function, LeakyReLU is the activation function, and a is the learnable attention vector. T W is the transpose of vector a, ∥ denotes vector concatenation, and W (2) This represents the learnable weight matrix. and These represent the node characteristic matrices of nodes m and n, respectively.
[0143] Using spatial attention weights Aggregate the feature matrices of all amino acid clusters under each category into protein subunit feature matrices.
[0144]
[0145] in, This indicates that the amino acid cluster m is located in structural unit D. k Total attention weight within, N Domain (D k ) represents the cluster of amino acids that make up a protein subunit.
[0146] S308, Based on the protein subunit feature matrix, an initial protein subunit structure map is generated, wherein the features of each node in the initial protein subunit structure map are the corresponding protein subunit feature matrix; the initial protein subunit structure map is input into the trained third graph neural network, and the protein subunit features of each node are updated using the convolutional layer of the third graph neural network to obtain the predicted protein subunit feature matrix of each node.
[0147] S309 performs feature mapping on all predicted protein subunit feature matrices to generate predicted protein structure feature vectors, and uses these predicted protein structure feature vectors to assess the pathogenicity of mutation sites.
[0148] This embodiment classifies the updated predicted amino acid cluster feature matrix using hierarchical clustering, aggregating all amino acid clusters in the same category into protein subunits. It utilizes the affected amino acid clusters to form a higher protein hierarchical structure, and uses spatial attention mechanism to weight the amino acid clusters, emphasizing different importance for amino acid clusters with different characteristics. It predicts the impact of mutations from a higher protein structural level, and can more comprehensively express the impact of mutations on different levels of protein structure.
[0149] Example 4
[0150] Figure 4 This is a flowchart of a method for predicting the pathogenicity of epidermal growth factor receptor as described in Embodiment 4 of the present invention. This embodiment is an optimization based on the above embodiment. In this embodiment, amino acid residue features are extracted based on the three-dimensional spatial conformation of a complex mutant amino acid sequence containing multiple mutation sites. Specifically, the optimization is as follows:
[0151] The predicted highly pathogenic EGFR single-point mutation sites were obtained, and the highly pathogenic EGFR mutation sites were combined with sensitive mutation sites to obtain a set of predicted complex mutant amino acid sequences.
[0152] Each compound mutant amino acid sequence in the predicted set of compound mutant amino acid sequences is subjected to three-dimensional structure transformation to obtain the three-dimensional spatial conformation of the compound mutant amino acid sequence.
[0153] Traverse the three-dimensional spatial conformation of the complex mutant amino acid sequence, and extract the sequence features of all amino acid residues, the type features of amino acid residues, and the three-dimensional coordinates of all atoms in the complex mutant amino acid sequence;
[0154] Based on the amino acid residue type characteristics and amino acid residue sequence characteristics of each amino acid residue, the key residue characteristics of each amino acid residue are determined;
[0155] The centroid coordinate characteristics of each amino acid residue are determined based on the mean of the three-dimensional coordinates of all atoms of each amino acid residue.
[0156] The dihedral features of each amino acid residue are determined based on the three-dimensional coordinates of all atoms in each amino acid residue.
[0157] Accordingly, the epidermal growth factor receptor pathogenicity prediction method provided in this embodiment specifically includes:
[0158] S401, obtain the predicted highly pathogenic EGFR single-point mutation sites, and combine the highly pathogenic EGFR mutation sites with sensitive mutation sites to obtain a set of predicted complex mutant amino acid sequences.
[0159] The AlphaMissense prediction model was used to predict all potential missense mutation sites in EGFR. Clinical data and EGFR mutation samples from publicly available gene databases were obtained. Using Venn diagrams and the obtained mutation sample data, the predicted potential missense mutation sites can be classified into known mutations and unknown mutations. Known mutation sites have relatively mature treatment methods in clinical practice. By analyzing the correlation or similarity between unknown mutation sites and known mutation sites, treatment or medication can be used to treat unknown mutation sites by referring to the treatment methods for known mutation sites.
[0160] The predicted potential missense mutation sites include potentially benign mutations and key pathogenic mutations. Key pathogenic mutations were classified using SIFT (based on sequence conservation), PROVEAN (based on machine learning), PolyPhen-2 (based on machine learning), and AlphaMissense (integrating multiple principles). The prediction results of each model were used to classify mutation sites as pathogenic or benign based on their cutoff values. Specifically, the SIFT model had a cutoff value of 0.05, indicating high pathogenicity; the PROVEAN model had a cutoff value of -2.5, indicating pathogenicity; the PolyPhen-2 model had a cutoff value of 0.99, indicating a high probability of pathogenicity; and the AlphaMissense model had a cutoff value of 0.99, indicating high pathogenicity. If a mutation site is predicted to be pathogenic in all models, it is classified as highly pathogenic; if three models predict it as pathogenic, it is classified as relatively highly pathogenic; if two models predict it as pathogenic, it is classified as moderately pathogenic; and otherwise, it is classified as low pathogenic. Mutation sites predicted as highly pathogenic by all models are selected and combined with two sensitive mutations, L858R and Del19 (exon 19 deletion), to obtain a predicted set of compound mutant amino acid sequences consisting of multiple compound mutant amino acid sequences.
[0161] S402, perform three-dimensional structural transformation on each compound mutant amino acid sequence in the predicted set of compound mutant amino acid sequences to obtain the three-dimensional spatial conformation of the compound mutant amino acid sequence.
[0162] To better simulate the impact of mutation sites on the overall protein structure, a three-dimensional conformational transformation tool was used to convert complex mutated amino acid sequences into their corresponding three-dimensional conformations. Specifically, the AlphaFold3 tool can convert complex mutated amino acid sequences into three-dimensional conformations. This model uses an improved Transformer architecture and a conditional diffusion model to predict and output the three-dimensional structure of the protein.
[0163] S403, traverses the three-dimensional spatial conformation of the complex mutant amino acid sequence, and extracts the sequence features of all amino acid residues, the type features of amino acid residues, and the three-dimensional coordinates of all atoms in the complex mutant amino acid sequence.
[0164] The three-dimensional spatial conformation of the complex mutant amino acid sequence was traversed using tools. Specifically, the PDBParser and PDBIO modules from the Python machine learning library Biopython were used to traverse all amino acid residues and each amino acid chain in the three-dimensional spatial conformation (PDB file or CIF file), extracting the three-dimensional coordinates (x, y, z) of all atoms. i ,y i ,z i ).
[0165] The type of amino acid residue can be determined by the atomic arrangement of each residue. Based on 20 amino acid types, a 20-dimensional binary vector representing the amino acid residue type is generated using one-hot encoding. Different amino acid residue types are distinguished by one element being 1 and the others being 0. For example, alanine is encoded as [1,0,0,…,0], cysteine as [0,1,0,…,0], and so on. Furthermore, the order of the amino acid residues can be determined based on the traversal sequence.
[0166] S404 determines the key residue characteristics of each amino acid residue based on the amino acid residue type characteristics and amino acid residue sequence characteristics.
[0167] Certain amino acid residues are located in key functional regions of the EGFR kinase structure. These residues play crucial roles in the kinase's catalytic reactions, participating directly or indirectly in ATP binding and hydrolysis, substrate recognition and phosphorylation, and the regulation of kinase activity. Changes (mutations) in these amino acid residues can affect kinase function and are closely related to disease development. Exemplary key amino acid residues include Lys.K745, Glu.E762, Met.M766, Lys.K777, Val.V834, His.H835, Arg.R836, Thr.T854, Asp.D855, Phe.F856, Gly.G857, Lys.K860, Val.V876, and Val.V884. Key amino acid residues are specific types of amino acid residues located at specific positions in an amino acid sequence. By traversing the amino acid residue type features and amino acid residue sequence features obtained, we can identify whether an amino acid residue is a key residue and generate corresponding amino acid residue features. Amino acid residue features can be represented using one-dimensional binary elements, for example: 0 represents a non-key amino acid residue and 1 represents a key amino acid residue.
[0168] S405, based on the mean of the three-dimensional coordinates of all atoms of each amino acid residue, determines the centroid coordinate characteristics of each amino acid residue.
[0169] By averaging the coordinates of all atoms in each amino acid residue, the centroid coordinates (Xc, Yc, Zc) of the amino acid residue are determined, representing its position in the three-dimensional conformation.
[0170] Xc=(x1+x2+...+x n ) / n
[0171] Yc=(y1+y2+...+y n ) / n
[0172] Zc=(z1+z2+...+z n ) / n
[0173] Where n represents the total number of atoms in the residue, x i y i z i These represent the coordinates of the i-th atom along the X, Y, and Z axes, respectively. A three-element feature of the amino acid residue's centroid coordinates is generated using the centroid coordinates of the amino acid residue.
[0174] S406 determines the dihedral features of each amino acid residue based on the three-dimensional coordinates of all atoms of each amino acid residue.
[0175] The dihedral angles phi (φ) and psi (ψ) of the main chain are two important parameters used to describe the torsion angle of the amino acid main chain. These two angles define the relative position and orientation of amino acid residues in the protein chain. These angles are obtained by calculating the three-dimensional coordinates of five key atoms on the amino acid main chain, including the carbon atom C′ (the C atom of the previous amino acid), the amino nitrogen atom N, the central carbon atom Cα, the carboxyl carbon atom C, and the nitrogen atom N″ (the N atom of the next amino acid) on the main chain.
[0176] Let the three-dimensional coordinates of the five atoms be C′(x1,y1,z1), N(x2,y2,z2), Cα(x3,y3,z3), C(x4,y4,z4), and N″(x5,y5,z5). Then the vectors between the atoms can be represented as follows:
[0177] Vector AB = NC′ = (x2-x1, y2-y1, z2-z1)
[0178] Vector BC = Cα - N = (x³ - x², y³ - y², z³ - z²)
[0179] Vector CD = C - Cα = (x⁴ - x³, y⁴ - y³, z⁴ - z³)
[0180] Vector DE = N″ - C = (x5 - x4, y5 - y4, z5 - z4)
[0181] The normal vector is calculated using the cross product: n1 = AB × BC, n2 = BC × CD;
[0182] Calculate the cosine of the angle between the vectors using the dot product: dot phi =n1·n2; dot psi =n2·DE;
[0183] The magnitude of a vector is represented as: norm_n1 = |n1|, norm_n2 = |n2|, norm_DE = |DE|;
[0184] Therefore, a dihedral angle can be expressed as:
[0185] Calculate the size of the dihedral angle using the arccos function in Python's NumPy library:
[0186] phi=arccos(cos_phi), psi=arccos(cos_psi).
[0187] The dihedral feature consisting of two elements is generated by utilizing the size of the dihedral angle.
[0188] The amino acid residue type features, amino acid residue centroid coordinate features, key residue features, and dihedral features are spliced together to form a matrix of amino acid residue features.
[0189] S407, Based on the amino acid residue features, an initial amino acid residue structure map is generated, wherein the features of each node in the initial amino acid residue structure map are the corresponding amino acid residue features; the initial amino acid residue structure map is input into the trained first graph neural network, and the convolutional layer of the first graph neural network is used to update the amino acid residue features of each node to obtain the predicted amino acid residue feature matrix of each node.
[0190] S408: Based on the predicted amino acid residue feature matrix of each node, the amino acid residues are clustered by DBSCAN based on spatial proximity. The predicted amino acid residue feature matrices of all categories are weighted and aggregated to generate the amino acid cluster feature matrix of each category.
[0191] S409, Based on the amino acid cluster feature matrix, an initial amino acid cluster structure map is generated, wherein the features of each node in the initial amino acid cluster structure map are the corresponding amino acid cluster feature matrix; the initial amino acid cluster structure map is input into the trained second graph neural network, and the convolutional layer of the second graph neural network is used to update the amino acid cluster feature matrix of each node to obtain the predicted amino acid cluster feature matrix of each node.
[0192] S410: Based on the predicted amino acid cluster feature matrix, determine the amino acid cluster combination threshold, use the amino acid cluster combination threshold to perform hierarchical clustering on the predicted amino acid cluster feature matrix, and perform weighted aggregation on all predicted amino acid cluster feature matrices in each category to generate multiple protein subunit feature matrices.
[0193] S411, Based on the protein subunit feature matrix, an initial protein subunit structure map is generated, wherein the features of each node in the initial protein subunit structure map are the corresponding protein subunit feature matrix; the initial protein subunit structure map is input into the trained third graph neural network, and the protein subunit features of each node are updated using the convolutional layer of the third graph neural network to obtain the predicted protein subunit feature matrix of each node.
[0194] S412 performs feature mapping on all predicted protein subunit feature matrices to generate predicted protein structure feature vectors, and uses these predicted protein structure feature vectors to assess the pathogenicity of mutation sites.
[0195] This embodiment uses clinical sample data and somatic mutation database data to predict mutation sites, obtaining potential missense mutation sites including known and unknown mutations. Different pathogenicity analysis models are used to assess the pathogenicity of all potential missense mutations. The obtained highly pathogenic mutations are combined with common sensitive mutation sites to form a complex mutant amino acid sequence. A three-dimensional spatial conformation is generated using the complex mutant amino acid sequence, and the three-dimensional spatial conformation of the complex mutant amino acid sequence is traversed to obtain the characteristics of all its amino acid residues. This is used to predict and analyze the impact of multiple mutations on different levels of protein structure.
[0196] Example 5
[0197] This embodiment is an optimization based on the above embodiment. In this embodiment, the pathogenicity of the predicted protein structure feature vector will be evaluated. Specifically, the optimization is as follows:
[0198] The predicted protein structure feature vectors are clustered, and the pathogenicity and drug sensitivity of unknown mutation sites in the same category as known mutation sites are assessed using protein structure features containing known mutation sites.
[0199] Hierarchical clustering is used to perform cluster analysis on predicted protein structure feature vectors, dividing them into different categories. The structural features of proteins containing known mutations in each category are identified. Since the structural features of proteins in the same category are highly similar, the three-dimensional conformational change patterns they exhibit after binding to the same drug also tend to be consistent. The pathogenicity and drug sensitivity information of known mutations in clinical data can be referenced to evaluate the pathogenicity and drug sensitivity of protein structures with unknown mutations in the same category. This provides scientific and accurate guidance for clinical treatment and drug matching of unknown mutation sites, and helps to achieve personalized drug matching.
[0200] Optionally, the predicted amino acid residue feature matrix can be clustered using key amino acid residue features from the predicted amino acid residue features, and pathogenicity, pathogenic mechanism and drug sensitivity can be evaluated for unknown mutation sites in the same category using predicted amino acid residue features containing known mutation sites.
[0201] Analysis of the impact of mutations on key residue features in predicted amino acid residue characteristics revealed that the predicted amino acid residue features within each category exhibit similar conformations and functions. For example, based on the predicted amino acid residue feature matrix, amino acid residues are classified into nine categories: α, β, γ, δ, ε, ζ, η, θ, and ι. The three-dimensional conformational features of the tyrosine kinase domains of amino acid residues within each category are similar. Specifically, α-class mutants involve classic clinical drug resistance mutations such as T790M, C797S, G719D, and the S768I_G719C complex mutation. Based on known mutations, mutants in this category may exhibit drug insensitivity and can therefore serve as a reference for pathogenicity. β-class mutants exhibit a kinase central domain conformational difference of less than 3.60 Å between Lys.K745 and Glu.E762, which can form salt bridges. Thr.T854, Asp.D855, Phe.F856, and Gly.G857 are also mentioned. The dihedral angle of the amino acid residues determines the orientation of phenylalanine. At this point, aspartic acid can form a coordinate bond with Mg-ATP. This conformational change provides the necessary spatial structure and chemical environment for ATP binding. Based on known mutations, it can be known that mutants in this category may have abnormal protein activation functions and can be used as a reference for pathogenicity. Similarly, γ, δ, ε, ζ, and θ mutants can be used as a reference for potential pathogenicity and need to be monitored. The pathogenicity of η and ι mutants remains questionable and requires further analysis and verification.
[0202] Example 6
[0203] Figure 5 This is a schematic diagram of the structure of an epidermal growth factor receptor pathogenicity prediction device according to Embodiment Six of the present invention, as shown below. Figure 5 As shown, the device includes:
[0204] The clinical mutation intersection analysis module 810 is used to extract amino acid residue features based on the three-dimensional spatial conformation of the compound mutant amino acid sequence.
[0205] The point-to-surface conformation flexible fusion module 820 is used to predict protein structural features based on amino acid residue characteristics;
[0206] The point-to-surface conformation flexible fusion module 820 includes a first prediction unit 821, which is used to predict amino acid residue features using the convolutional layer of the first graph neural network.
[0207] Amino acid cluster feature generation unit 822 is used to generate amino acid cluster features based on predicted amino acid residue features;
[0208] The second prediction unit 823 is used to predict amino acid cluster features using the convolutional layer of the second graph neural network;
[0209] Protein subunit feature generation unit 824 is used to generate protein subunit features based on predicted amino acid cluster features;
[0210] The third graph convolutional unit 825 is used to predict protein subunit features using the convolutional layers of the third graph neural network;
[0211] Protein structure feature generation unit 826 is used to generate protein structure features based on protein subunit features;
[0212] The point-to-surface conformation clustering dynamic evaluation module 830 is used to evaluate the pathogenicity of mutation sites based on protein structural characteristics.
[0213] The epidermal growth factor receptor pathogenicity prediction system described in this embodiment extracts amino acid residue features and predicts mutation sites according to the protein hierarchy, from low-level to high-level structures. It then uses clustering and spatial attention mechanisms to aggregate low-level structures into higher-level structures, ultimately forming complete protein structural features. These complete protein structural features are then classified, and the pathogenicity and drug sensitivity of unknown mutation sites are assessed using the characteristics of known mutation sites within the same category. This system can more comprehensively consider the impact of mutations on adjacent structures at different protein hierarchies and can serve as an important reference for research on unknown mutations.
[0214] Based on the above embodiments, the amino acid cluster feature generation unit includes:
[0215] The amino acid residue coordinate extraction subunit is used to extract the three-dimensional coordinates of amino acid residues of each node based on the predicted amino acid residue feature matrix of each node.
[0216] DBSCAN clustering subunits are used to classify the predicted amino acid residue feature matrix of each node.
[0217] The weighted subunit of the tyrosine kinase region is used to weight the amino acid residues located in the tyrosine kinase region.
[0218] Based on the above embodiments, the protein subunit feature generation unit includes:
[0219] The amino acid cluster combination threshold calculation subunit is used to calculate the amino acid cluster combination threshold based on the maximum distance between the centroid coordinates of all amino acid clusters in the predicted amino acid cluster feature matrix.
[0220] Hierarchical clustering subunits are used to classify the predicted amino acid cluster feature matrix of each node;
[0221] The amino acid cluster weighting subunit is used to weight all predicted amino acid cluster feature matrices under each category.
[0222] The epidermal growth factor receptor pathogenicity prediction system provided in this embodiment of the invention can execute the epidermal growth factor receptor pathogenicity prediction method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method.
[0223] Example 7
[0224] Figure 6 This is a schematic diagram of the structure of a server according to Embodiment 7 of the present invention. Figure 6 A block diagram of an exemplary server 12 suitable for implementing embodiments of the present invention is shown. Figure 6 The server 12 shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of the present invention.
[0225] like Figure 6 As shown, server 12 is presented in the form of a general-purpose computing device. Components of server 12 may include, but are not limited to, one or more processors or processing units;
[0226] System memory, bus 18 connecting different system components (including system memory 28 and processing unit 16).
[0227] Bus 18 represents one or more of several bus architectures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus architectures. For example, these architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus.
[0228] Server 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by server 12, including volatile and non-volatile media, removable and non-removable media.
[0229] System memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. Server 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 34 may be used to read and write non-removable, non-volatile magnetic media (… Figure 6 Not shown; usually referred to as a "hard drive"). Although Figure 6Not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk") and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to bus 18 via one or more data media interfaces. Memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of the present invention.
[0230] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. Program modules 42 typically perform the functions and / or methods described in the embodiments of the present invention.
[0231] Server 12 can also communicate with one or more external devices 14 (e.g., keyboard, pointing device, display 24, etc.), and with one or more devices that enable users to interact with the device / system / server 12, and / or with any device that enables the server 12 to communicate with one or more other computing devices (e.g., network card, modem, etc.). This communication can be performed via input / output (I / O) interface 22. Furthermore, server 12 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 20. As shown, network adapter 20 communicates with other modules of server 12 via bus 18. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with server 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0232] The processing unit 16 executes various functional applications and data processing by running programs stored in the system memory 28, such as implementing the epidermal growth factor receptor pathogenicity prediction method provided in the embodiments of the present invention.
[0233] Example 8
[0234] Embodiment 8 of the present invention also provides a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the epidermal growth factor receptor pathogenicity prediction method provided in the above embodiments.
[0235] The computer storage medium of this invention can be any combination of one or more computer-readable media. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0236] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.
[0237] The program code contained on a computer-readable medium may be transmitted using any suitable medium, including—but not limited to—wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0238] Computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0239] Note that the above description is merely a preferred embodiment of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of the present invention, the scope of which is determined by the scope of the appended claims.
Claims
1. A method for predicting the pathogenicity of epidermal growth factor receptor, characterized in that, include: Based on the three-dimensional spatial conformation of a complex mutant amino acid sequence containing multiple mutation sites, amino acid residue features are extracted. These features include: amino acid residue type features, amino acid residue centroid coordinate features, key residue features, and dihedral features. The amino acid residue features are in matrix form. Based on the amino acid residue features, an initial amino acid residue structure map is generated, wherein the features of each node in the initial amino acid residue structure map are the corresponding amino acid residue features; the initial amino acid residue structure map is input into a first graph neural network that has been trained, and the amino acid residue features of each node are updated using the convolutional layer of the first graph neural network to obtain the predicted amino acid residue feature matrix of each node. Based on the predicted amino acid residue feature matrix of each node, the amino acid residues are clustered by DBSCAN based on spatial proximity. The predicted amino acid residue feature matrices of all the predicted amino acid residues in each category are weighted and aggregated to generate the amino acid cluster feature matrix of each category. An initial amino acid cluster structure map is generated based on the amino acid cluster feature matrix. The features of each node in the initial amino acid cluster structure map are the corresponding amino acid cluster feature matrix. The initial amino acid cluster structure map is input into the trained second graph neural network. The convolutional layer of the second graph neural network is used to update the amino acid cluster feature matrix of each node to obtain the predicted amino acid cluster feature matrix of each node. Based on the predicted amino acid cluster feature matrix, the amino acid cluster combination threshold is determined. The predicted amino acid cluster feature matrix is then hierarchically clustered using the amino acid cluster combination threshold. All predicted amino acid cluster feature matrices in each category are weighted and aggregated to generate multiple protein subunit feature matrices. An initial protein subunit structure map is generated based on the protein subunit feature matrix. The features of each node in the initial protein subunit structure map are the corresponding protein subunit feature matrix. The initial protein subunit structure map is input into a trained third graph neural network. The convolutional layer of the third graph neural network is used to update the protein subunit features of each node to obtain the predicted protein subunit feature matrix of each node. Feature mapping is performed on all predicted protein subunit feature matrices to generate predicted protein structure feature vectors. These predicted protein structure feature vectors are then used to assess the pathogenicity of mutation sites.
2. The method according to claim 1, characterized in that, The process involves DBSCAN clustering of amino acid residues based on spatial proximity, using the predicted amino acid residue feature matrix of each node. Weighted aggregation of all predicted amino acid residue feature matrices within each category generates an amino acid cluster feature matrix for each category, including: Based on the predicted amino acid residue feature matrix of each node, the three-dimensional coordinates of the amino acid residues of each node are extracted; DBSCAN clustering was performed based on the three-dimensional coordinates of amino acid residues at each node, with a neighborhood radius of 12 angstroms and a minimum number of points of 6. Based on the spatial attention mechanism, amino acid residues located in the tyrosine kinase region are weighted, and the sum of the feature matrices of all predicted amino acid residues under each category is calculated to obtain the amino acid cluster feature matrix for each category.
3. The method according to claim 1, characterized in that, The process involves determining an amino acid cluster combination threshold based on the predicted amino acid cluster feature matrix, performing hierarchical clustering on the predicted amino acid cluster feature matrix using the amino acid cluster combination threshold, and weighting and aggregating all predicted amino acid cluster feature matrices in each category to generate multiple protein subunit feature matrices, including: Based on the predicted amino acid cluster feature matrix, the maximum distance between the centroid coordinates of all amino acid clusters is calculated, and 12% of the maximum distance is used as the amino acid cluster combination threshold. Based on the amino acid cluster combination threshold and the centroid coordinates of the amino acid clusters, hierarchical clustering is performed on the amino acid clusters of each node. Each category contains multiple predicted amino acid cluster feature matrices where the distance between the centroid coordinates of the amino acid clusters is less than the amino acid cluster combination threshold. Based on the spatial attention mechanism, the feature matrices of all predicted amino acid clusters under each category are weighted, and the sum of the features of all predicted amino acid clusters under each category is calculated to obtain the protein subunit feature matrix of each category.
4. The method according to claim 1, characterized in that, The extraction of amino acid residue features based on the three-dimensional spatial conformation of a complex mutant amino acid sequence containing multiple mutation sites includes: The predicted highly pathogenic EGFR single-point mutation sites were obtained, and the highly pathogenic EGFR mutation sites were combined with sensitive mutation sites to obtain a set of predicted complex mutant amino acid sequences. Each compound mutant amino acid sequence in the predicted set of compound mutant amino acid sequences is subjected to three-dimensional structure transformation to obtain the three-dimensional spatial conformation of the compound mutant amino acid sequence. Traverse the three-dimensional spatial conformation of the complex mutant amino acid sequence, and extract the sequence features of all amino acid residues, the type features of amino acid residues, and the three-dimensional coordinates of all atoms in the complex mutant amino acid sequence; Based on the amino acid residue type characteristics and amino acid residue sequence characteristics of each amino acid residue, the key residue characteristics of each amino acid residue are determined; The centroid coordinate characteristics of each amino acid residue are determined based on the mean of the three-dimensional coordinates of all atoms of each amino acid residue. The dihedral features of each amino acid residue are determined based on the three-dimensional coordinates of all atoms in each amino acid residue.
5. The method according to claim 1, characterized in that, The method further includes: The key residue features of each amino acid residue are identified, and the amino acid residue features of each node are weighted in the convolutional layer of the first graph neural network using an attention mechanism based on the key residue features.
6. The method according to claim 1, characterized in that, The assessment of pathogenicity of mutation sites using predicted protein structure feature vectors includes: The predicted protein structure feature vectors are clustered, and the pathogenicity and drug sensitivity of unknown mutation sites in the same category as known mutation sites are assessed using protein structure features containing known mutation sites.
7. The method according to claim 1, characterized in that, The method further includes: Using the key amino acid residue features in the predicted amino acid residue features, the predicted amino acid residue feature matrix is clustered. Using the predicted amino acid residue features that include known mutation sites, the pathogenicity, pathogenic mechanism and drug sensitivity of unknown mutation sites in the same category are evaluated.
8. An epidermal growth factor receptor pathogenicity prediction device for implementing the epidermal growth factor receptor pathogenicity prediction method as described in claim 1, characterized in that, include: The clinical mutation intersection analysis module is used to extract amino acid residue features based on the three-dimensional spatial conformation of the compound mutant amino acid sequence. A point-to-surface conformation flexible fusion module is used to predict protein structural features based on amino acid residue characteristics; The point-to-surface conformation flexible fusion module includes a first prediction unit, which is used to predict amino acid residue features using the convolutional layer of a first graph neural network; An amino acid cluster feature generation unit is used to generate amino acid cluster features based on predicted amino acid residue features. The second prediction unit is used to predict amino acid cluster features using the convolutional layers of the second graph neural network; The protein subunit feature generation unit is used to generate protein subunit features based on predicted amino acid cluster features. The third graph convolutional unit is used to predict protein subunit features using the convolutional layers of the third graph neural network; The protein structure feature generation unit is used to generate protein structure features based on protein subunit features. The point-to-surface conformation clustering dynamic evaluation module is used to assess the pathogenicity of mutation sites based on protein structural characteristics.
9. A server, characterized in that, The server includes: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the epidermal growth factor receptor pathogenicity prediction method as described in any one of claims 1-7.
10. A storage medium comprising computer-executable instructions, which, when executed by a computer processor, are used to perform the epidermal growth factor receptor pathogenicity prediction method as described in any one of claims 1-7.
Citation Information
Patent Citations
Protein-protein interaction site prediction method based on deep map convolutional network
CN113192559A
Gene missense mutation pathogenicity prediction system based on deep learning
CN117672382A