Protein phosphorylation site prediction method based on federal learning
Through a method based on federated learning, preprocessing and characterizing multi-party biomedical data is preprocessed and characterized, and the Federal Learning Committee is established for model training and aggregation, which solves the problems of small data scale and poor generalization of the model in protein phosphorylation site prediction, achieving efficient and accurate prediction and data sharing.
Patent Information
- Application Number
- CN202510029209.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-08
AI Technical Summary
The prior art has problems in the prediction of protein phosphorylation site, such as small data scale, low data isomerism, and poor model generalization, and it is difficult to effectively share and utilize multi-party biomedical data.
Using a federated learning-based method, a federal learning committee is established by preprocessing and characterizing biomedical data from multiple medical institutions and laboratories, a federal learning committee is established, model training and aggregation is carried out, and global models are generated to achieve efficient prediction of protein phosphorylation sites.
Overcome the problem of data imbalance, improve prediction accuracy and model generalization capabilities, realize the sharing and utilization of multi-party data, and effectively protect data privacy.
Smart Images

Figure CN119943140A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of bioinformatics analysis, and mainly to a protein phosphorylation site prediction method based on federated learning. Background Art
[0002] With the completion of the Human Genome Project and the advent of the post-genomic era, researchers have a deeper understanding of the importance of protein phosphorylation in biological regulation. Protein phosphorylation is an important type of protein post-translational modification, which adds phosphate groups to specific amino acid residues such as serine, threonine and tyrosine through kinase catalysis. When external stimuli reach the cell surface, they trigger a series of signal transduction events, ultimately affecting the biological response within the cell.
[0003] As an important regulatory mechanism of signal transduction, phosphorylation can convert external signals into biological effects within cells, such as cell proliferation, apoptosis, and differentiation. By controlling the phosphorylation state of proteins, cells can respond promptly and accurately to changes in the external environment. Phosphorylation is also involved in the regulation of protein interactions, which can change the conformation and charge state of proteins, affect their binding ability and specificity with other proteins, and thus regulate the establishment and stability of protein interaction networks. During cell proliferation and differentiation, phosphorylation can ensure that cells divide and proliferate at appropriate times and conditions, regulate the activity of cell cycle proteins, and control the progress of the cell cycle, thereby maintaining the balance and stability of cell proliferation. Therefore, accurate prediction of protein phosphorylation sites is of great significance for understanding cell regulatory mechanisms, disease diagnosis, and drug design.
[0004] At present, the identification of phosphorylation sites mainly relies on experimental methods and computational methods. Among them, experimental methods such as mass spectrometry and antibody detection can directly quantify and locate phosphorylation sites, but they are limited by resources and costs and are not suitable for large-scale application. Computational methods have advantages in prediction speed and cost efficiency, but traditional computational methods are often limited by the scale and quality of data.
[0005] With the development of deep learning technology, deep learning based on multi-layer neural networks has made significant progress in the field of bioinformatics. Deep learning can automatically generate complex patterns and capture high-level abstractions, providing new ideas for the prediction of protein phosphorylation sites. However, deep learning usually requires a large amount of data for training, and data in the medical and disease fields are often restricted by privacy protection, making data acquisition difficult. These methods are usually trained on independent databases, and there are problems such as small data set size, low data heterogeneity, and poor model generalization.
[0006] Therefore, it is necessary to provide a method suitable for the task of multi-dimensional unbalanced phosphorylation site prediction, which can overcome the problems of data and scale imbalance, effectively protect data privacy, and realize the sharing and utilization of multi-party biomedical data. Summary of the invention
[0007] The purpose of the present invention is to provide a protein phosphorylation site prediction method based on federated learning. The method comprehensively utilizes the biomedical data of different medical institutions and laboratories, overcomes the imbalance problem of data of each institution, realizes efficient prediction of protein phosphorylation sites based on deep learning, and can effectively protect the data privacy of each medical institution and laboratory.
[0008] To achieve the above object, the present invention provides a method for predicting protein phosphorylation sites based on federated learning, comprising the following steps:
[0009] S1. Preprocess the protein sequence data to obtain positive samples and negative samples. Make the positive and negative samples independent and identically distributed through random sampling, and divide them into training sets and test sets. At the same time, perform statistics on the size, label ratio and site ratio of the client's local data to form a feature description of the local data set.
[0010] S2. Establish a federated learning committee, which consists of fixed members and temporary members. Fixed members own the global model test data set and are responsible for generating the initial global model. Temporary members are selected by the clients participating in federated learning and are responsible for testing the prediction accuracy of the local model.
[0011] S3. Each client participates in federated learning, obtains the initial global model, and trains the model using the local data set to obtain the local model and local prediction accuracy of each client;
[0012] S4. The local models, prediction accuracy and data set feature descriptions trained by each client are passed to the Federated Learning Committee for testing. The selection criteria are determined based on the prediction accuracy of the Federated Learning Committee, and local models that meet the selection criteria are selected to participate in aggregation to generate a global model.
[0013] S5. The federated learning committee uses the global model test data set to test the generated global model, and decides to terminate federated learning or start the next round of federated learning based on the test results until convergence.
[0014] Preferably, the selection criteria for temporary members of the Federal Learning Committee are generated, including:
[0015] During the first round of federated learning, the local model of each client is tested by a fixed member to obtain the prediction accuracy of the federated learning committee, and the difference between the prediction accuracy of the committee and the local model is calculated. The absolute value of the difference is used as the reliability of this round, and then the temporary members of the federated learning committee of this round are selected from the clients based on the reliability.
[0016] In subsequent rounds of federated learning, the local models of each client are tested through the temporary members selected in the previous round to obtain the prediction accuracy of each temporary member. The difference with the prediction accuracy of the local model is calculated respectively, and then the absolute value of the difference is summed as the reliability of subsequent rounds. The temporary members of the federated learning committee in subsequent rounds are selected based on the reliability.
[0017] Preferably, local model training includes determining the window size, performing feature extraction using multiple feature blocks, and concatenating feature maps along the feature dimension, and then generating a phosphorylation prediction score through a fully connected layer to output the prediction accuracy of the local model.
[0018] Preferably, local models that meet the selection criteria are selected to participate in the aggregation, including:
[0019] First, the local model of the non-Federated Learning Committee client is tested using the local test set through the temporary members of the FLC to obtain the prediction accuracy of each temporary member;
[0020] Secondly, the difference between the prediction accuracy of the temporary member and the prediction accuracy of the local model is calculated, and the absolute value of each difference is summed up as the reliability of the corresponding local model;
[0021] Then, the selection standard line is determined according to the reliability distribution, and the local models that are below the standard are eliminated, and then the local models that meet the selection criteria are globally aggregated.
[0022] Preferably, determining the selection standard line includes sorting the reliability of each client local model, calculating the upper and lower quartiles and the interquartile range, and determining the selection standard line k as follows:
[0023] k=q 1 -t*iqr;
[0024] In the formula, q 1 is the upper quartile, iqr is the interquartile range, and t is the ratio of iqr.
[0025] Preferably, the local models that meet the selection criteria are globally aggregated, including assigning corresponding weights to the local models participating in the aggregation according to the feature description of the data set of each client, and then obtaining the global model, as follows:
[0026]
[0027] In the formula, is the parameter of the i-th local model in the l-th round of federated learning, n is the number of clients participating in the global model aggregation, m i is the weight of the i-th local model.
[0028] Therefore, the present invention adopts the above-mentioned protein phosphorylation site prediction method based on federated learning, which has the following technical effects:
[0029] (1) The federated learning method solves the data “island” problem caused by only analyzing and training highly private data locally.
[0030] (2) Multiple local models are screened and aggregated after performance screening to improve training efficiency, save time and computing power costs, achieve more accurate prediction of phosphorylation sites, and provide a more reliable data prediction tool.
[0031] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 It is a general flow chart of the model federated learning training in an embodiment of a protein phosphorylation site prediction method based on federated learning;
[0033] Figure 2 The present invention is a data processing flow chart in an embodiment of a protein phosphorylation site prediction method based on federated learning. DETAILED DESCRIPTION
[0034] The present invention can be explained in more detail by the following examples. The purpose of disclosing the present invention is to protect all changes and improvements within the scope of the present invention. The present invention is not limited to the following examples.
[0035] like Figure 1 As shown, the present invention provides a protein phosphorylation site prediction method based on federated learning, comprising the following steps:
[0036] S1. Each client preprocesses the local protein sequence data, such as Figure 2 shown.
[0037] First, for the prediction of general sites, all repeated sequences were deleted, and then the protein data with the number of amino acid residue sequences below the window size were removed.
[0038] Secondly, from the screened protein site data, the phosphorylation sites and nearby fragment data of all verified amino acid sites are extracted as positive samples; for negative samples, a subset of a specific number of other sites and their nearby fragment information are randomly selected and fused with the positive samples to form a local dataset on the client, which is then divided into a training dataset and a test dataset.
[0039] In order to construct independent and identically distributed data, for a data set with m samples, it is necessary to mix its positive and negative samples and use a random algorithm to permute the positions of the sample data in the data set.
[0040] Analyze the imbalance of the client's local data set itself (covering label imbalance and site imbalance) and the imbalance between clients (covering data scale imbalance, label imbalance and site imbalance) to form a data imbalance feature description.
[0041] S2. To support federated learning, a federated learning committee needs to be established, which includes a fixed member and several temporary members to evaluate the effectiveness of federated learning. Specifically, if 20 clients participate in federated learning and temporary members are selected at a ratio of 30%, a 7-member federated learning committee will be formed.
[0042] The selection rules for temporary members are as follows: define the reliability scoring standard e and filter out unreliable clients. Specifically, after the client undergoes federated learning training, each client is tested to obtain a local prediction accuracy. During the initial selection, the local model and the locally measured accuracy are passed to the committee, and the fixed members detect the local model of each client to obtain the prediction accuracy. The local accuracy of each local model is subtracted from the accuracy measured by the fixed members, and the absolute value is taken as the reliability of this round. Six temporary members are selected from low to high according to the reliability. In subsequent rounds, the six temporary members selected in the previous round will be responsible for testing the local models of all other clients in this round. In this way, each local model obtains six prediction accuracy rates, which are subtracted from the phosphorylation prediction accuracy of the local model to obtain the absolute value, and then the sum is calculated to obtain the reliability. Six temporary members for the next round are selected from low to high according to the reliability.
[0043] The six temporary members selected in each round will form a seven-member federated learning committee together with the fixed members, which will determine the selection criteria for the local models of each client participating in the global model aggregation.
[0044] S3. For each round of federated learning, the federated learning committee will provide a starting global model. In the first round, fixed members will generate the initial global model. In other rounds, the global model gathered by the committee will be used as the starting global model for the new round. Each client will then use local data to train the starting global model and test the prediction accuracy.
[0045] First, the local protein sequence is input into the model with a total of k DC-CNN blocks to form sequence feature clusters Among them, I and L k I represents the size of the amino acid symbol dictionary and the local window size of the corresponding phosphorylation site, respectively. In this embodiment, a hot encoding scheme for encoding protein sequences is adopted, so I is set to 21.
[0046] Secondly, by analyzing the window size configurations of different DC-CNN blocks, three effective window sizes were determined to be 15, 33, and 51. Each DC-CNN block contains a one-dimensional convolutional layer that performs convolution operations along the length of the protein sequence and applies activation functions such as ReLU to achieve nonlinear transformations.
[0047] For the kth DC-CNN block, the first convolutional layer generates the feature map for the block for:
[0048]
[0049] In the formula, f k is the activation function, W k is the weight matrix, is the bias term.
[0050] In addition, in order to enhance the generalization ability of the model and reduce the threat of overfitting during training, the dropout algorithm is used after the convolutional layer to randomly discard a part of the neurons.
[0051] Afterwards, Intra-BCL is introduced to strengthen the propagation of phosphorylation information in the DC-CNN block, connecting all the front convolutional layers with the back convolutional layers, that is, concatenating the feature maps along the feature dimension to enhance the model's ability to capture phosphorylation information. Accordingly, the sequence feature is formed by concatenating the feature map and the output of all previous layers as the input of the i-th convolutional layer, as follows:
[0052]
[0053] In the formula, is the feature map generated by the i-th convolutional layer, C is the number of convolutional layers in each DC-CNN block, is the bias term of the i-th convolutional layer. When the network depth increases, the abstract features of the previous layer at different levels are more easily transferred to the current layer.
[0054] The protein phosphorylation site sequences created by each DC-CNN block are further integrated along the first dimension through intra-BCL:
[0055]
[0056] In the formula, h f The concatenation of the feature maps generated by the Cth convolutional layer, is the bias term of the Cth convolutional layer. In this way, multiple feature maps are connected to each other through the plane layer and converted into a one-dimensional tensor.
[0057] Subsequently, a fully connected neural network is used to create a softmax function to obtain the phosphorylation prediction score and output the prediction accuracy of the trained model.
[0058] S4. The Federated Learning Committee evaluates the reliability of the local models trained by each client and selects local models that meet the standards to participate in global aggregation, as follows:
[0059] First, each client packages the trained local model, prediction accuracy, and dataset feature description and passes it to the Federated Learning Committee.
[0060] Secondly, temporary members of the federated learning committee test the local models of other clients based on their local test data sets. Each client local model will receive 6 test values, and the reliability of the client local model will be obtained by combining these 6 test values with the prediction accuracy measured by the client itself. The committee will establish standards based on the reliability distribution and select client local models to participate in the convergence of the global model.
[0061] Next, the reliability of each client's local model is ranked, the upper and lower quartiles and the interquartile range are calculated, and the selection standard line k is determined as follows:
[0062] k=q 1 -t*iqr;
[0063] In the formula, q 1 is the upper quartile, iqr is the interquartile range, and t is the ratio of iqr, which is 0.5 in this embodiment.
[0064] After that, the client local models with accuracy lower than k are eliminated, and the parameters of the remaining local models are aggregated. During the aggregation process, the imbalance of the dataset size, label distribution, and site type distribution between the clients is weighted to obtain the global model, as follows:
[0065]
[0066] In the formula, is the i-th model parameter of the l-th round of federated learning, n is the number of clients, m i is the weight of the ith model.
[0067] S5. The committee tests the global model based on the global model test data set to obtain the prediction accuracy. If the goal is achieved, the federated learning will end and the global model will be passed to the model demander. If the goal is not achieved, a new round of federated learning will continue until the global model achieves the goal.
[0068] Therefore, the protein phosphorylation site prediction method based on federated learning provided by the present invention can overcome the problem of data imbalance, effectively protect data privacy, and realize the sharing and utilization of multi-party data.
[0069] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solution of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solution to deviate from the spirit and scope of the technical solution of the present invention.
Claims
1. A protein phosphorylation site prediction method based on federated learning, characterized in that: The following steps are involved: S1. Preprocess the protein sequence data to obtain positive samples and negative samples. Make the positive and negative samples independent and identically distributed through random sampling, and divide them into training sets and test sets. At the same time, perform statistics on the size, label ratio and site ratio of the client's local data to form a feature description of the local data set. S2. Establish a federated learning committee, which consists of fixed members and temporary members. Fixed members own the global model test data set and are responsible for generating the initial global model. Temporary members are selected by the clients participating in federated learning and are responsible for testing the prediction accuracy of the local model. S3. Each client participates in federated learning, obtains the initial global model, and trains the model using the local data set to obtain the local model and local prediction accuracy of each client; S4. The local models, prediction accuracy and data set feature descriptions trained by each client are passed to the Federated Learning Committee for testing. The selection criteria are determined based on the prediction accuracy of the Federated Learning Committee, and local models that meet the selection criteria are selected to participate in aggregation to generate a global model. S5. The federated learning committee uses the global model test data set to test the generated global model, and decides to terminate federated learning or start the next round of federated learning based on the test results until convergence.
2. A method for predicting protein phosphorylation sites based on federated learning according to claim 1, characterized in that: The selection criteria for temporary members of the Federal Learning Committee are generated, including: During the first round of federated learning, the local model of each client is tested by a fixed member to obtain the prediction accuracy of the federated learning committee, and the difference between the prediction accuracy of the committee and the local model is calculated. The absolute value of the difference is used as the reliability of this round, and then the temporary members of the federated learning committee of this round are selected from the clients based on the reliability. In subsequent rounds of federated learning, the local models of each client are tested through the temporary members selected in the previous round to obtain the prediction accuracy of each temporary member. The difference with the prediction accuracy of the local model is calculated respectively, and then the absolute value of the difference is summed as the reliability of subsequent rounds. The temporary members of the federated learning committee in subsequent rounds are selected based on the reliability.
3. A method for predicting protein phosphorylation sites based on federated learning according to claim 1, characterized in that: Local model training includes determining the window size, using multiple feature blocks to extract features, and concatenating feature maps along the feature dimension, and then generating phosphorylation prediction scores through a fully connected layer to output the prediction accuracy of the local model.
4. The method for predicting protein phosphorylation sites based on federated learning according to claim 1, characterized in that: Select local models that meet the selection criteria to participate in the aggregation, including: First, the local model of the non-Federated Learning Committee client is tested using the local test set through the temporary members of the FLC to obtain the prediction accuracy of each temporary member; Secondly, the difference between the prediction accuracy of each temporary member and the prediction accuracy of the local model is calculated, and the absolute value of each difference is summed up as the reliability of the corresponding local model; Then, the selection standard line is determined according to the reliability distribution, and the local models that are below the standard are eliminated, and then the local models that meet the selection criteria are globally aggregated.
5. A method for predicting protein phosphorylation sites based on federated learning according to claim 4, characterized in that: Determine the selection standard line, including sorting the reliability of each client's local model, calculating the upper and lower quartiles and the interquartile range, and determining the selection standard line k, as follows: k = q1 - t * iqr; Where q1 is the upper quartile, iqr is the interquartile range, and t is the ratio of iqr.
6. A method for predicting protein phosphorylation sites based on federated learning according to claim 4, characterized in that: The local models that meet the selection criteria are globally aggregated, including assigning corresponding weights to the local models participating in the aggregation according to the feature description of the data set of each client, and then obtaining the global model as follows: In the formula, is the parameter of the i-th local model in the l-th round of federated learning, n is the number of clients participating in the global model aggregation, m i is the weight of the i-th local model.
Citation Information
Patent Citations
Client selection and adaptive model aggregation method and system based on reinforcement learning in federated learning
CN117459570A
Dimension-driven decentralized federated learning data deviation processing method
CN118395191A
Robust efficient decentralized federated learning method and system based on committee consensus
CN118940865A
Decentralized federated learning method and system and related equipment
CN118982064A