Phosphorylation site prediction method and system for regulating protein interaction by fusing large language model and graph neural network

By integrating large language models and graph neural networks, combining protein sequence and structure information, and constructing a multi-level prediction model, the problems of difficult feature acquisition and high computational cost in existing technologies are solved, and efficient and accurate prediction of phosphorylation sites is achieved, which has application potential in disease diagnosis and drug discovery.

CN120748477APending Publication Date: 2025-10-03SUZHOU UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510594073.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing technologies find it difficult to effectively utilize protein sequence, structure, and PPI information to accurately predict phosphorylation sites. They have problems such as difficulty in feature acquisition, weak model generalization ability, and high computational cost.

Method used

By integrating a large language model and graph neural network, pre-training and dynamically adjusting the model, combining protein sequence and structure information, building a multi-level prediction model, extracting multi-dimensional features and performing feature fusion, and using AlphaFold protein structure data, we can overcome computational challenges.

Benefits of technology

It improves the prediction accuracy and computational efficiency of the regulatory effects of phosphorylation sites on PPIs, provides in-depth insights into disease mechanisms and drug discovery, and shows great application potential in the study of functional phosphorylation sites.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120748477A_ABST
    Figure CN120748477A_ABST
Patent Text Reader

Abstract

The invention discloses a method and a system for predicting phosphorylation sites for regulating protein interaction by fusing a large language model and a graph neural network. The method comprises the following steps: acquiring protein sequence data and phosphorylation site multi-dimensional information; constructing a first prediction model, pre-training the first prediction model, and dynamically adjusting the first prediction model based on a pre-training result; the first prediction model is used for primary prediction of phosphorylation site regulation and control of protein-protein interaction, and a second prediction model is constructed in combination with obtained structural information of a protein monomer; and based on the second prediction model, carrying out secondary prediction on the phosphorylation site regulation protein-protein interaction. According to the method, the prediction precision of the PPI regulation effect of the phosphorylation sites is improved by combining information of two important levels of sequences and structures, the calculation and storage bottlenecks in large-scale data processing are effectively solved, and a new thought and method are provided for research of the phosphorylation sites.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of prediction of key sites regulating protein function, and in particular to a method and system for predicting phosphorylation sites that regulate protein interactions by integrating a large language model and a graph neural network. Background Art

[0002] Proteins significantly expand their chemical space and functional diversity through post-translational modifications (PTMs), of which phosphorylation is one of the most extensively studied PTMs. Phosphorylation can regulate various cellular signaling pathways by triggering specific protein-protein interactions (PPIs), affecting protein conformation, subcellular localization, and function within signaling networks. Abnormal phosphorylation is closely associated with major diseases such as cancer, neurological disorders, and cardiovascular disease, as it can have a critical impact by remodeling PPIs within signaling networks.

[0003] Although many experimental strategies have been developed to identify phosphorylation-dependent protein-protein interactions, such as mass spectrometry- and proteomics-based methods, these approaches are often time-consuming, expensive, and cannot be rapidly expanded to newly discovered phosphorylation sites. Furthermore, while database resources have facilitated the study of the association between PTMs and PPIs, they are primarily based on known data and struggle to predict the functions of unknown phosphorylation sites. Machine learning methods have been used to predict PTM sites, but are still in their infancy in the study of functionally relevant phosphorylation sites. While recently proposed sequence-based models (such as PhosPPI and FuncPhos-SEQ) and structure-based models (such as FuncPhos-STR) can predict the functions of phosphorylation sites to a certain extent, they primarily rely on artificially constructed features (such as PPI network features or structural features) and cannot fully utilize large-scale sequence and structural data. With the development of pre-trained protein language models (pLMs) and geometric deep learning technology, although some pLMs-based models (such as Phosformer) and graph neural networks (such as MIND-S) have performed well in predicting PTM functions, existing models still have problems such as limitations in feature extraction, bottlenecks in structural information processing, and lack of cross-modal integration. Summary of the Invention

[0004] In view of the above-mentioned problems, the present invention is proposed.

[0005] Therefore, the technical problem solved by the present invention is: how to make full use of protein sequence, structure and PPI information to accurately predict the phosphorylation sites that regulate PPI, so as to solve the problems of difficult feature acquisition, weak model generalization ability and high computational cost.

[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0007] In a first aspect, an embodiment of the present invention provides a method for predicting phosphorylation sites regulating protein interactions by integrating a large language model and a graph neural network, comprising:

[0008] Obtain protein sequence data and multi-dimensional information of phosphorylation sites;

[0009] Constructing a first prediction model, pre-training the first prediction model, and dynamically adjusting the first prediction model based on the pre-training results;

[0010] The first prediction model is used for the primary prediction of phosphorylation sites regulating protein-protein interactions, and the second prediction model is constructed in combination with the structural information of the protein monomer;

[0011] Based on the second prediction model, a secondary prediction is performed on the phosphorylation site regulating protein-protein interaction.

[0012] As a preferred solution of the method for predicting phosphorylation sites regulating protein interactions by integrating a large language model and a graph neural network, the pre-training of the first prediction model includes:

[0013] Extracting amino acid sequence fragments of target length from the multi-dimensional information of the phosphorylation sites;

[0014] A specific model is used for pre-training, with the extracted amino acid sequence fragment of the target length as input, and a specific embedding method is used to represent the input fragment to identify the phosphorylation site motif features.

[0015] The beneficial effects of this preferred technical solution are: extracting amino acid sequence fragments based on multi-dimensional information of phosphorylation sites, being able to focus on key data for model training, and the specific embedding method can effectively represent the input fragments, helping the model to better learn sequence features and provide a more accurate basis for subsequent predictions.

[0016] As a preferred solution for the phosphorylation site prediction method for regulating protein interactions by integrating a large language model and a graph neural network, the pre-training of the first prediction model also includes: using the sequence embedding vector of a specific dimension obtained after pre-training the specific model as the input of the first prediction model.

[0017] The beneficial effect of this preferred technical solution is that the specific dimensional sequence embedding vector obtained by pre-training is used as the input of the fine-tuning model, which can fully utilize the common features learned in the pre-training stage, reduce the training time and cost of the fine-tuning stage, and enable the fine-tuning model to converge to better performance more quickly.

[0018] As a preferred solution of the method for predicting phosphorylation sites regulating protein interactions by integrating a large language model and a graph neural network, dynamically adjusting the first prediction model based on the pre-training results includes:

[0019] Dynamic adjustment is performed based on the output of pre-training to adjust the pre-trained general features into specific features required to predict functional phosphorylation sites that regulate protein-protein interactions.

[0020] The beneficial effect of this preferred technical solution is that by fine-tuning the pre-trained general features into specific features, the model can be more focused on predicting functional phosphorylation sites that regulate protein-protein interactions, thereby improving the accuracy and specificity of the model for this specific task.

[0021] As a preferred solution of the method for predicting phosphorylation sites regulating protein interactions by integrating a large language model and a graph neural network, the first prediction model is used for a single prediction of phosphorylation sites regulating protein-protein interactions, including:

[0022] If the probability value output by the activation function of the dynamically adjusted first prediction model is greater than 0.5, the predicted phosphorylation site is considered to regulate protein-protein interaction;

[0023] If the probability value output by the activation function of the dynamically adjusted first prediction model is not greater than 0.5, it is considered that the predicted phosphorylation site does not regulate protein-protein interaction.

[0024] As a preferred solution for the method of predicting phosphorylation sites regulating protein interactions by integrating a large language model and a graph neural network, the second prediction model is constructed by combining the structural information of the protein monomer and comprising: three sub-networks of the second prediction model;

[0025] Subnetwork 1 is used to extract the features of each protein residue to form a feature set matrix, and perform feature dimensionality reduction processing on the feature set matrix; the feature set includes protein sequence features and sequence evolution information features;

[0026] Subnetwork 2 is used to construct a graphical model to describe the three-dimensional structure and molecular properties of specific amino acids. Subnetwork 2 takes the protein node feature matrix learned by subnetwork 1 and the adjacency matrix constructed from the subgraph of the graphical model as input, respectively. Through two different computational networks, spatial information of phosphorylation sites is extracted and combined as feature vectors learned by subnetwork 2.

[0027] Subnetwork three is used to encode phosphorylated fragments of a specific length and extract feature vectors of a specific dimension.

[0028] The beneficial effects of this preferred technical solution are as follows: the three subnetworks of the second prediction model extract protein features from different perspectives. Subnetwork 1 integrates sequence and evolutionary information, subnetwork 2 extracts spatial information, and subnetwork 3 encodes phosphorylated fragments. This multi-dimensional feature extraction enables the model to more comprehensively capture protein properties, thereby improving prediction accuracy. Furthermore, linear transformation reduces the dimensionality of the input features, which helps reduce computational complexity and prevent overfitting.

[0029] As a preferred solution for the method of predicting phosphorylation sites regulating protein interactions by integrating a large language model and a graph neural network, the method comprises: concatenating and fusion-enhancing the features learned by the second and third sub-networks, and transmitting the concatenated features to a multi-layer perceptron;

[0030] The probability that the phosphorylation site has the function of regulating protein-protein interaction is represented according to the output of the multi-layer perceptron activation function, so as to predict the phosphorylation site information that regulates protein-protein interaction.

[0031] The beneficial effects of this preferred technical solution are: the features learned by different sub-networks are spliced ​​and fused, which can integrate various information. The multi-layer perceptron can perform more complex nonlinear processing on the fused features. The activation function normalizes the output into a probability distribution, making the prediction results more intuitive and easy to understand, further improving the accuracy of predicting phosphorylation sites regulating protein-protein interactions from the sequence and structural levels.

[0032] In a second aspect, an embodiment of the present invention provides a phosphorylation site prediction system for regulating protein interactions that integrates a large language model and a graph neural network, including:

[0033] Data acquisition module, used to obtain protein sequence data and multi-dimensional information of phosphorylation sites;

[0034] A first prediction model dynamic adjustment module is used to construct a first prediction model, pre-train the first prediction model, and dynamically adjust the first prediction model based on the pre-training result;

[0035] A second prediction model construction module is used to construct a second prediction model based on the first prediction model for the primary prediction of phosphorylation sites regulating protein-protein interactions and the obtained structural information of protein monomers;

[0036] A prediction module is used to perform a secondary prediction on the phosphorylation site regulating protein-protein interaction based on the second prediction model.

[0037] In a third aspect, an embodiment of the present invention provides an electronic device, including:

[0038] memory and processor;

[0039] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the one or more programs are executed by the one or more processors, the one or more processors implement the phosphorylation site prediction method for regulating protein interactions by integrating a large language model and a graph neural network as described in any embodiment of the present invention.

[0040] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the phosphorylation site prediction method for regulating protein interactions that integrates a large language model and a graph neural network.

[0041] The beneficial effects of the present invention are as follows: the present invention combines a pre-trained protein language model and a graph neural network, which can directly extract task-specific features from protein sequences and structures, avoiding the limitations of traditional feature engineering methods and improving the prediction accuracy of the regulatory effects of phosphorylation sites on PPIs; by integrating sequence-based pre-training models and graph neural network structural features, combined with AlphaFold protein structure data, the protein sequence information and structure information are effectively combined, thereby improving the comprehensiveness and accuracy of the prediction results; the kNN algorithm is used to focus on the spatial features around the target phosphorylation site, and combined with the graph neural network to effectively perform convolution operations on the subgraph molecular representation of the protein, overcoming the challenges of storing and calculating large-scale protein 3D structures and significantly improving computational efficiency; by predicting the PPI regulatory effects of phosphorylation sites, it can provide in-depth insights into disease mechanisms, disease diagnosis and drug discovery, especially in the remodeling of PPI networks related to phosphorylation sites, helping to identify potential therapeutic targets; compared with traditional machine learning methods based on manual features, the present invention can directly extract task-related features from protein sequence and structure data, has higher versatility and adaptability, and shows good application potential in the study of functional phosphorylation sites. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0043] Figure 1 This is an overall flow chart of the method for predicting phosphorylation sites regulating protein interactions that integrates a large language model and a graph neural network provided by the present invention;

[0044] Figure 2 This is a schematic structural diagram of the first prediction model of the method for predicting phosphorylation sites regulating protein interactions that integrates a large language model and a graph neural network provided by the present invention;

[0045] Figure 3 This is a flow chart of the first prediction model of the method for predicting phosphorylation sites regulating protein interactions that integrates a large language model and a graph neural network provided by the present invention;

[0046] Figure 4 This is a schematic structural diagram of a second prediction model of the method for predicting phosphorylation sites regulating protein interactions that integrates a large language model and a graph neural network provided by the present invention;

[0047] Figure 5 It is a flow chart of the second prediction model of the method for predicting phosphorylation sites regulating protein interactions that integrates a large language model and a graph neural network provided by the present invention. DETAILED DESCRIPTION

[0048] To make the above-mentioned objects, features, and advantages of the present invention more clearly understood, the following detailed description of the specific embodiments of the present invention is given in conjunction with the accompanying drawings. It is obvious that the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in this field without creative work should fall within the scope of protection of the present invention.

[0049] Example 1, reference Figure 1 , which is the first embodiment of the present invention, provides a method for predicting phosphorylation sites regulating protein interactions by integrating a large language model and a graph neural network, comprising:

[0050] S100: Acquire protein sequence data and multi-dimensional information of phosphorylation sites;

[0051] S200: Constructing a first prediction model, pre-training the first prediction model, and dynamically adjusting the first prediction model based on the pre-training result;

[0052] S300: The first prediction model is used to predict the phosphorylation sites regulating protein-protein interactions. Combined with the structural information of the protein monomer, the second prediction model is constructed.

[0053] S400: Based on the second prediction model, secondary prediction of phosphorylation sites regulating protein-protein interactions is performed.

[0054] It should be noted that most existing methods rely on artificially designed features for function prediction. However, these features may not be available or this method is difficult to apply to newly sequenced proteins or less studied organisms. This example directly extracts important features from protein sequence and structure. Existing prediction methods based on protein 3D structure require the storage and processing of high-resolution structural data, which consumes a large amount of computing resources and storage space when processing long sequences or large protein datasets. This example introduces an efficient subgraph feature extraction method to extract spatial features efficiently and at low cost, reducing computational overhead. Sequence information and structural information are important sources for predicting phosphorylation site function, but existing methods typically process them separately or only use single-modality data, resulting in limited prediction accuracy. This example integrates sequence features extracted by a pre-trained protein language model with structural features captured by a graph convolutional network and a graph attention network to achieve deep fusion of multimodal information and improve prediction performance. Existing models mostly aim for generalized function prediction and are not specifically optimized for specific functions, making them difficult to meet practical application needs. This example addresses this shortcoming by optimizing the deep learning framework and focusing on predicting functional phosphorylation sites directly related to PPI regulation.

[0055] Example 2, reference Figures 2 to 5 , which is an embodiment of the present invention, provides a method for predicting phosphorylation sites regulating protein interactions by integrating a large language model and a graph neural network based on the previous embodiment, including:

[0056] In this embodiment, obtaining protein sequence data and multi-dimensional information of phosphorylation sites in step S100 includes:

[0057] Obtain functional phosphorylation sites, molecular functions of phosphorylation sites, regulatory annotations of biological processes and protein-protein interactions, and their disease-related supplementary data, as well as protein sequence information, to construct a sequence-based prediction model of the regulatory function of phosphorylation sites on protein-protein interactions;

[0058] In a preferred embodiment, functional phosphorylation sites, i.e., phosphorylation sites that regulate PPIs, can be obtained from the PhosphoSitePlus (PSP) and PTMint databases, regulatory annotations of phosphorylation sites' molecular functions, biological processes, and protein-protein interactions, and supplementary data of disease-related information can be obtained from the PSP, PTMint, iPTMnet, and PTMD databases, and protein sequence information can be obtained from the UniProt database to construct a sequence-based prediction model for phosphorylation sites that regulate PPIs (PhosPPI-SEQ), i.e., the first prediction model.

[0059] like Figure 2 and Figure 3As shown in Figure 3, the PhosPPI-SEQ prediction model consists of a pre-trained encoder, followed by a feed-forward neural network layer, and finally a binary classification layer.

[0060] In another possible implementation, when constructing the first prediction model, a construction method based on a support vector machine (SVM) may be used;

[0061] Specifically, data processing involves obtaining functional phosphorylation sites, their molecular functions, regulatory annotations of biological processes and protein-protein interactions, and supplementary disease-related data, along with protein sequence information. Feature extraction can be performed on protein sequence information, such as amino acid composition, dipeptide composition, and amino acid physicochemical properties. These features are then combined into a feature vector, which serves as the input for the SVM.

[0062] Model construction: Select an appropriate kernel function (such as linear kernel, radial basis kernel, etc.) to build an SVM classifier. Use methods such as cross-validation to tune the SVM parameters (such as penalty factor C, kernel function parameters, etc.) to improve model performance.

[0063] Model training: Using data on known functional phosphorylation sites and non-functional phosphorylation sites as a training set, the SVM model is trained to enable it to learn the relationship between phosphorylation sites and protein-protein interaction regulation.

[0064] In this embodiment, pre-training the first prediction model in step S200 includes:

[0065] Extract target-length amino acid sequence fragments from multi-dimensional information of phosphorylation sites;

[0066] A specific model is used for pre-training, with the extracted amino acid sequence fragment of the target length as input, and a specific embedding method is used to represent the input fragment to identify the phosphorylation site motif features.

[0067] Specifically, based on the obtained functional phosphorylation site, 64 residues on each side of the functional phosphorylation site were extracted, totaling 129 amino acid sequence fragments.

[0068] Use the Transformer model to pre-train the first prediction model;

[0069] It should be noted that the Transformer model takes 129 amino acid sequence segments as input and uses three embedding methods to represent the input segments: token embedding, segment embedding, and position embedding. The Transformer model consists of a six-layer encoder and decoder. The encoder is responsible for converting the tokenized sequence into a 768-dimensional sequence embedding vector, while the decoder restores the embedding vector to the tokenized sequence. Each encoder consists of six multi-head attention layers and six feedforward neural network (FNN) layers. Each attention layer contains 12 attention heads and outputs a 768-dimensional embedding tensor.

[0070] Pre-training is performed based on the masked language model (MLM) objective function;

[0071] Specifically, 214,082 phosphorylation site fragments were selected from the UniProt database as training data; when constructing the masked language model, 15% of the amino acid tags in the sequence were randomly selected and replaced with <mask>tags, and use the available contextual information to predict the similarity of these tags to the original amino acid tags, using the masked language model loss function L MLM Calculate the difference between the predicted mask mark and the true mark, expressed as:

[0072]

[0073] in, is the set of masked tokens in the input, is the set of unmasked tokens in the input, N is the number of masked tokens in the input, and the indices of the masked tokens are v1,v2,…,v n , Represents the nth masked token, and the difference between the masked token predicted by the model and the true token is evaluated by calculating the cross entropy loss, and is used for backpropagation to optimize the model parameters.

[0074] After pre-training, the encoder uses the obtained 768-dimensional sequence embedding vector as the input of the subsequent fine-tuning model.

[0075] In this embodiment, dynamically adjusting the first prediction model based on the pre-training result in step S200 includes:

[0076] Dynamic adjustment is performed based on the output of pre-training to adjust the pre-trained general features into specific features required to predict functional phosphorylation sites that regulate protein-protein interactions.

[0077] Specifically, a pre-trained Transformer encoder is used as a feature extraction module; protein motif features and protein sequence evolution information features are added and input into a feedforward neural network layer, which is used to learn the feature mapping of specific tasks; the data processed by the feedforward neural network layer is input into a binary classification layer, and a Sigmoid activation function is used to output a probability value between 0 and 1, indicating whether the input phosphorylation site has PPI regulatory function.

[0078] In this embodiment, the use of the dynamically adjusted first prediction model in step S200 for a prediction of phosphorylation sites regulating protein-protein interactions includes:

[0079] If the probability value output by the Sigmoid activation function is greater than 0.5, the predicted site is considered to regulate PPI; if it is less than or equal to 0.5, the predicted site is considered not to regulate PPI.

[0080] In this embodiment, step S300 combines the acquisition of structural information of protein monomers to construct the second prediction model, including: three sub-networks of the second prediction model;

[0081] Subnetwork 1 is used to extract the features of each protein residue to form a feature set matrix and perform feature dimensionality reduction on the feature set matrix; the feature set includes protein sequence features and sequence evolution information features;

[0082] Subnetwork 2 is used to build a graphical model to describe the three-dimensional structure and molecular properties of specific amino acids. Subnetwork 2 takes the protein node feature matrix learned by subnetwork 1 and the adjacency matrix constructed from the subgraph of the graphical model as input. Through two different computational networks, it extracts the spatial information of phosphorylation sites and combines this spatial information as the feature vector learned by subnetwork 2.

[0083] Subnetwork three is used to encode phosphorylated fragments of a specific length and extract feature vectors of a specific dimension.

[0084] In a preferred embodiment, protein structure information can be obtained from the AlphaFold database to construct a prediction model for the regulatory function of phosphorylation sites on protein-protein interactions (PhosPPI-Alpha) based on sequence and structure, i.e., the second prediction model;

[0085] like Figure 4 and Figure 5 As shown in Figure 2, the PhosPPI-Alpha prediction model mainly consists of three sub-networks: SeqNet sub-network, graph neural sub-network, and MotifNet sub-network;

[0086] Specifically, the SeqNet subnetwork is used to extract features from each residue of the protein to form a feature set matrix. The feature set contains protein sequence features and protein sequence evolution information features. Protein sequence features are extracted from the Encoder layer for each residue using a pre-trained model; protein sequence evolution information features are obtained by comparing the target sequence with evolutionarily related sequences to obtain the probability of occurrence or conservation of amino acids at each site. The final protein sequence features are represented as a matrix of L*788, where L is the length of the protein sequence and 788 represents the total dimension of the node features obtained by combining the PhosPPI-SEQ model and PSSM (Position-Specific Scoring Matrix). A linear transformation is applied to reduce the input feature dimension, expressed as:

[0087] H (0) =[WX PhosPPI-SEQ ,WX PSSM ]

[0088] Among them, W is a learnable parameter, H (0) are the node features input to the model.

[0089] The graph neural subnetwork is used to construct the k-nearest neighbor (kNN) graph G(A,E), which describes the three-dimensional structure and molecular properties of the k amino acids closest to the phosphorylation site.

[0090] Specifically, the graph neural network sub-network takes two inputs: the protein node feature matrix learned by the SeqNet sub-network and the adjacency matrix constructed from the kNN sub-graph. These two matrices are fed into a residue-based graph convolutional network (GCN) and a graph attention network (GAT) consisting of L layers, respectively, to extract spatial information about phosphorylation sites. The results of these two methods are then combined to form the feature vector learned by the graph neural network.

[0091] The contact graph of a protein with a length of L can be represented as a square matrix C with an order of L = {c p,q }, expressed as:

[0092]

[0093] Among them, δ p,q represents the Euclidean distance between the Cα atoms of residues p and q.

[0094] Furthermore, the input of GCN is the node feature matrix H and the adjacency matrix A. The initial residual and identity mapping method are used to solve the traditional GCN transition smoothing problem. The output node features are calculated by the improved formula. The improved GCN calculation formula is as follows:

[0095] H (l+1) =σ((1-α)PH (1) +αH (0) )((1-β l )I n +β l W (l) )

[0096] Among them, H (l+1) It is the output node feature of the improved GCN; H (l) and H (0) are the input node features and initial node features of the lth layer respectively; W (l) is the learnable weight matrix of the lth layer; σ is the nonlinear activation function, usually ReLU, α is set to 0.6; P = D -1 / 2 AD -1 / 2 , A is the adjacency matrix that represents the connection relationship between nodes in the graph, and D is the degree matrix which is also the diagonal matrix of A;

[0097] Furthermore, β l Expressed as:

[0098]

[0099] Here, λ is set to 0.1.

[0100] The input of GAT is also the adjacency matrix A and the node feature matrix H. Each node applies a shared linear transformation layer, which is expressed as:

[0101] z (l) =W (l) H (l)

[0102] Among them, W (l) is the learnable parameter of layer l, H (l) is the node feature input to the lth layer.

[0103] By applying a shared linear transformation layer, the input features are transformed into higher-level features and the inter-node attention scores are calculated, which are expressed as:

[0104]

[0105] Among them, the i-th amino acid is the central amino acid, W l and a l is a learnable parameter matrix, || represents concatenation, is the set of neighbor nodes of the ith amino acid, are the features of the i-th, j-th, and k-th amino acids, respectively. LeakyRelu is the ReLU activation function with a negative input slope set to 0.2.

[0106] After calculating the attention score between each neighbor node and the target node, the target node features are updated as follows:

[0107]

[0108] Among them, σ represents the ReLU activation function, P (l+1) is the output node feature of the i-th GAT layer.

[0109] Combine the node features obtained by GCN and GAT and update the node features in the following way:

[0110]

[0111] H (l+1) =H mid W l +H 0

[0112] It should be noted that H mid Contains the results of GCN and GAT, H mid and W l The dot product operation is used to reduce H mid dimensions so that it can be compared with the matrix H 0 Add, in the last step H mid W l and H 0 Add them together to build a residual network, making the network deeper.

[0113] It should be noted that the introduction of graph neural networks can improve the ability of the second prediction model to identify phosphorylation sites that regulate PPI.

[0114] The MotifNet sub-network is used to encode phosphorylation fragments with a sequence length of 129 using the protein sequence encoding tool and extract a 768-dimensional feature vector.

[0115] In another possible implementation, when constructing the second prediction model, a method based on a combination of a deep residual network (ResNet) and a graph neural network may also be used;

[0116] Specifically, sequence feature extraction: A deep residual network (ResNet) is used to extract features from protein sequences. ResNet effectively addresses the vanishing and exploding gradient problems in deep neural networks and can learn more complex sequence features. The protein sequence is input into ResNet, where sequence features are extracted through multiple convolutional layers and residual blocks.

[0117] Structural feature extraction: A graph neural network (GNN) is used to process protein monomer structural information. Similar to the example, a k-nearest neighbor graph (kNN) can be constructed to describe the protein's three-dimensional structure and molecular properties. Graph convolutional layers and graph attention layers are used to extract features from the graph structure.

[0118] Feature fusion and model building: The sequence features extracted by ResNet and the structural features extracted by GNN are concatenated or weighted and fused, and then input into the fully connected layer for classification prediction. Batch normalization layers and activation functions (such as ReLU) can be added between the fully connected layers to improve model performance and training stability.

[0119] In this embodiment, the secondary prediction of the phosphorylation site regulating protein-protein interaction based on the second prediction model in step S400 includes:

[0120] The features learned by sub-network 2 and sub-network 3 are concatenated and fused, and the concatenated features are transmitted to the multi-layer perceptron;

[0121] The probability that the phosphorylation site has the function of regulating protein-protein interaction is represented according to the output of the multi-layer perceptron activation function, so as to predict the phosphorylation site information that regulates protein-protein interaction.

[0122] Specifically, the features learned by the graph neural network and the MotifNet sub-network are concatenated and fused, and then passed to the MLP (Multi-layer Perceptron), which is expressed as:

[0123] Y=Softmax(H L w+b+X)

[0124] Among them, H L is the output of the last layer, W is the learnable parameter, b is the bias term, and X is the feature learned by the sub-network MotifNet.

[0125] Similar to the first prediction model, the Softmax activation function is used to normalize the output into a probability distribution to predict whether the phosphorylation site has PPI regulatory function.

[0126] It should be noted that the PhosPPI-Alpha model includes the PhosPPI-SEQ model; the PhosPPI-SEQ model only predicts whether the phosphorylation site regulates PPI from the sequence level; the PhosPP-Alpha model not only considers the structural level using AlphaFold protein monomer structure information, but also retains the sequence capture of the PhosPPI-SEQ model.

[0127] Example 3. The above is a schematic scheme of the method for predicting phosphorylation sites that regulate protein interactions by integrating a large language model and a graph neural network in this embodiment. It should be noted that the technical solution of the system for predicting phosphorylation sites that regulate protein interactions by integrating a large language model and a graph neural network and the technical solution of the method for predicting phosphorylation sites that regulate protein interactions by integrating a large language model and a graph neural network belong to the same concept. For details not described in detail in the technical solution of the system for predicting phosphorylation sites that regulate protein interactions by integrating a large language model and a graph neural network in this embodiment, please refer to the description of the technical solution of the method for predicting phosphorylation sites that regulate protein interactions by integrating a large language model and a graph neural network.

[0128] This embodiment also provides a phosphorylation site prediction system for regulating protein interactions that integrates a large language model and a graph neural network, including:

[0129] Data acquisition module, used to obtain protein sequence data and multi-dimensional information of phosphorylation sites;

[0130] A first prediction model dynamic adjustment module is used to construct a first prediction model, pre-train the first prediction model, and dynamically adjust the first prediction model based on the pre-training results;

[0131] The second prediction model construction module is used to predict the phosphorylation site regulation of protein-protein interaction based on the first prediction model, and to construct the second prediction model in combination with the information obtained from the protein monomer structure;

[0132] The prediction module is used to perform secondary prediction on the phosphorylation site regulating protein-protein interaction based on the second prediction model.

[0133] This embodiment further provides an electronic device applicable to a method for predicting phosphorylation sites regulating protein interactions that integrates a large language model and a graph neural network, including:

[0134] Memory and processor; the memory is used to store computer-executable instructions, and the processor is used to execute computer-executable instructions to implement the phosphorylation site prediction method for regulating protein interactions that integrates a large language model and a graph neural network as proposed in the above embodiment.

[0135] This embodiment also provides a storage medium on which a computer program is stored. When the program is executed by a processor, it implements the phosphorylation site prediction method for regulating protein interactions by integrating a large language model and a graph neural network as proposed in the above embodiment.

[0136] The storage medium proposed in this embodiment and the phosphorylation site prediction method for regulating protein interactions by integrating a large language model and a graph neural network proposed in the above embodiment belong to the same inventive concept. Technical details not fully described in this embodiment can be found in the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.

[0137] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.< / mask>

Claims

1. A method for predicting phosphorylation sites regulating protein interactions by integrating a large language model and a graph neural network, characterized in that: include: Obtain protein sequence data and multi-dimensional information of phosphorylation sites; Constructing a first prediction model, pre-training the first prediction model, and dynamically adjusting the first prediction model based on the pre-training results; The first prediction model is used for the primary prediction of phosphorylation sites regulating protein-protein interactions, and the second prediction model is constructed in combination with the structural information of the protein monomer; Based on the second prediction model, a secondary prediction is performed on the phosphorylation site regulating protein-protein interaction.

2. The method for predicting phosphorylation sites regulating protein interactions by integrating a large language model and a graph neural network according to claim 1, wherein: Pre-training the first prediction model includes: Extracting amino acid sequence fragments of target length from the multi-dimensional information of the phosphorylation sites; A specific model is used for pre-training, with the extracted amino acid sequence fragment of the target length as input, and a specific embedding method is used to represent the input fragment to identify the phosphorylation site motif features.

3. The method for predicting phosphorylation sites regulating protein interactions by integrating a large language model and a graph neural network according to claim 2, wherein: The pre-training of the first prediction model further includes: using the sequence embedding vector of the specific dimension obtained after pre-training the specific model as the input of the first prediction model.

4. The method for predicting phosphorylation sites regulating protein interactions by integrating a large language model and a graph neural network as claimed in claim 3, characterized in that: Dynamically adjusting the first prediction model based on the pre-training result includes: Dynamic adjustment is performed based on the output of pre-training to adjust the pre-trained general features into specific features required to predict functional phosphorylation sites that regulate protein-protein interactions.

5. The method for predicting phosphorylation sites regulating protein interactions by integrating a large language model and a graph neural network as claimed in claim 4, characterized in that: The first prediction model is used for a prediction of phosphorylation sites regulating protein-protein interactions, including: If the probability value output by the activation function of the dynamically adjusted first prediction model is greater than 0.5, the predicted phosphorylation site is considered to regulate protein-protein interaction; If the probability value output by the activation function of the dynamically adjusted first prediction model is not greater than 0.5, it is considered that the predicted phosphorylation site does not regulate protein-protein interaction.

6. The method for predicting phosphorylation sites regulating protein interactions by integrating a large language model and a graph neural network according to claim 5, wherein: Combined with the obtained structural information of protein monomers, the second prediction model is constructed including: three sub-networks of the second prediction model; Subnetwork 1 is used to extract the features of each protein residue to form a feature set matrix, and perform feature dimensionality reduction processing on the feature set matrix; the feature set includes protein sequence features and sequence evolution information features; Subnetwork 2 is used to construct a graphical model to describe the three-dimensional structure and molecular properties of specific amino acids. Subnetwork 2 takes the protein node feature matrix learned by subnetwork 1 and the adjacency matrix constructed from the subgraph of the graphical model as input, respectively. Through two different computational networks, spatial information of phosphorylation sites is extracted and combined as feature vectors learned by subnetwork 2. Subnetwork three is used to encode phosphorylated fragments of a specific length and extract feature vectors of a specific dimension.

7. The method for predicting phosphorylation sites regulating protein interactions by integrating a large language model and a graph neural network according to claim 6, wherein: Based on the second prediction model, performing a secondary prediction on the phosphorylation site regulating protein-protein interaction includes: The features learned by the sub-network 2 and the sub-network 3 are concatenated and fused, and the concatenated features are transmitted to a multi-layer perceptron; The probability that the phosphorylation site has the function of regulating protein-protein interaction is represented according to the output of the multi-layer perceptron activation function, so as to predict the phosphorylation site information that regulates protein-protein interaction.

8. A system for predicting phosphorylation sites regulating protein interactions that integrates a large language model and a graph neural network, using the method according to any one of claims 1 to 7, characterized in that: include: Data acquisition module, used to obtain protein sequence data and multi-dimensional information of phosphorylation sites; A first prediction model dynamic adjustment module is used to construct a first prediction model, pre-train the first prediction model, and dynamically adjust the first prediction model based on the pre-training result; A second prediction model construction module is used to construct a second prediction model based on the first prediction model for the primary prediction of phosphorylation sites regulating protein-protein interactions and the obtained structural information of protein monomers; A prediction module is used to perform a secondary prediction on the phosphorylation site regulating protein-protein interaction based on the second prediction model.

9. An electronic device comprising: memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the steps of the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Large model and multi-mode fused lysine modification site prediction method and device

    CN120954516A