A method for predicting kinase-specific substrate proteins that regulate yeast autophagy
By integrating multi-omics data and deep learning algorithms, we screened and predicted the kinase-specific substrate proteins of yeast autophagy, solving the problems of long identification cycle and high cost in traditional methods, and achieving efficient and accurate prediction of yeast autophagy regulatory substrates.
Patent Information
- Application Number
- CN202210700281.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-20
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2042-06-20
AI Technical Summary
Existing technologies for identifying functional phosphorylated substrate proteins that regulate yeast autophagy have long experimental cycles, high costs, and low efficiency. Traditional methods are unable to accurately predict kinase-specific substrate proteins.
By integrating multi-omics data of transcriptome, quantitative proteome and phosphoproteomics, and using deep learning algorithms and transfer learning methods, we screened out genes that interact with autophagy core genes, extracted their transcriptional, protein and phosphorylation expression changes and amino acid frequency characteristics, constructed a prediction model, optimized the screening range and accurately predicted kinase-specific substrate proteins.
It effectively narrowed the screening scope of kinase-specific substrate proteins, reduced the workload of experimental verification, accurately predicted the kinase-specific substrate proteins that regulate yeast autophagy, and improved the identification efficiency and accuracy.
Smart Images

Figure CN115171791B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of bioinformation technology, and more specifically, relates to a method for predicting kinase-specific substrate proteins that regulate yeast autophagy. Background Art
[0002] Autophagy is a degradation pathway based on lysosomes (in animals) and vacuoles (in yeast and plants). Through the formation of autophagosomes, damaged organelles, misfolded proteins, and other cellular material are engulfed and transported to lysosomes / vacuoles for degradation, meeting metabolic needs and replacing some organelles. Autophagy can be divided into two types: selective autophagy and non-selective autophagy. Under conditions of nutrient deprivation and various stimuli, autophagic activity increases significantly, exerting a protective function for the cell. The occurrence of cellular autophagy involves the formation and extension of autophagic vacuoles, the maturation of autophagosomes, and the fusion of autophagosomes with lysosomes. In the field of autophagy research, Saccharomyces cerevisiae is a classic and important model organism for studying the molecular regulation of autophagy. To date, over 40 autophagy core genes have been identified in yeast, approximately half of which have orthologs in mammals. Different stages of autophagy are tightly regulated by the autophagy core. Currently, research has shown that a total of 18 autophagy core proteins are crucial for autophagosome formation during cellular autophagy. Although several core autophagy proteins, including atg1, have been identified, the molecular mechanisms of autophagy regulation involved in these core autophagy proteins still need further study. Therefore, it is particularly important to further explore and identify new functional genes involved in the regulation of yeast autophagy. However, the main limitations of traditional experimental methods for verifying and discovering important functional phosphorylated substrate proteins involved in the regulation of cellular autophagy are: (1) they are not combined with quantitative proteomics technology based on chemical characterization methods to provide valuable information for the screening and identification of phosphorylated substrate proteins; (2) they do not use machine learning methods to learn the information of known functional substrate proteins to improve the identification efficiency of phosphorylated proteins; (3) the use of experimental identification methods alone identifies too wide a range of candidate functional substrates, resulting in a long phosphorylated substrate experimental verification cycle and huge human and material costs.
[0003] The research group of the inventor of the present invention has previously developed a method for predicting functional genes that regulate yeast autophagy (Chinese patent application CN113077841). Although this method can accurately predict functional genes that regulate yeast autophagy to a certain extent, it is still unable to accurately predict functional phosphorylated substrate proteins. Summary of the Invention
[0004] In response to the above defects or improvement needs of traditional experimental methods in the prior art for identifying functional genes that regulate yeast autophagy, which require long experimental cycles and high costs, the purpose of the present invention is to provide a method for predicting kinase-specific substrate proteins that regulate yeast autophagy. By improving the overall process design of the method and selecting four key gene characteristics, based on multi-omics data of transcriptome, quantitative proteome and phosphoproteomics, the prediction of autophagy-related kinase-specific substrate proteins is achieved by integrating multi-omics data. Using the prediction method of the present invention, the screening range of kinase-specific substrate proteins can be effectively narrowed before traditional experimental methods are verified, thereby reducing the workload of experimental verification and accurately predicting kinase-specific substrate proteins that regulate yeast autophagy.
[0005] To achieve the above objectives, according to the present invention, a method for predicting kinase-specific substrate proteins that regulate yeast autophagy is provided, characterized in that it comprises the following steps:
[0006] S1: Based on pre-selected key autophagy regulatory genes, time-series transcriptome, proteome, and phosphoproteomic analyses were performed on normal yeast samples before and after starvation treatment, mutant yeast samples with knockout of the key autophagy regulatory gene, and mutant yeast samples with plasmid complementation of the key autophagy regulatory gene. Gene expression data, protein expression data, and phosphorylated protein expression data were obtained for yeast samples before and after starvation treatment.
[0007] S2: Based on a pre-selected protein interaction database, genes that interact with the key autophagy regulatory genes are screened out and recorded as Set A; based on a pre-selected autophagy gene database, known autophagy genes with known autophagy functions in Set A are marked and recorded as Set B, which is a subset of Set A; and at the same time, based on the correspondence between autophagy genes and phosphorylated substrates collected in advance, autophagy genes with known phosphorylated substrates in Set A are marked and recorded as Set C, which is also a subset of Set A;
[0008] S3: Extracting the transcriptional expression changes, protein expression changes, phosphorylated protein expression changes, and protein sequence composition corresponding to each gene in the set A before and after the knockout of the key autophagy regulatory gene and after the plasmid complementation of the key autophagy regulatory gene, as well as the four types of features for each gene in the set A;
[0009] S4: Establish a prediction model, based on the four types of features determined in step S3, using a deep learning algorithm, using the genes in set B as the corresponding training positive dataset, and using the genes in set A but excluding the genes in set B as the corresponding training negative dataset, to train the prediction model;
[0010] S5: Through the transfer learning method, the trained prediction model obtained in the step S4 is optimized with the genes in the set C corresponding to the positive training data set and the genes in the set A excluding the genes in the set C corresponding to the negative training data set, and the positive data set is enhanced by the meta-learning method to obtain the final prediction model; then, the genes in the set A are scored using the final prediction model, and the genes whose scores meet the pre-set requirements are encoded by the corresponding proteins, which are predicted to be kinase-specific substrate proteins that regulate yeast autophagy.
[0011] As a further preferred embodiment of the present invention, in step S1, the autophagy inducer used in the starvation induction treatment is selected from nitrogen source-deficient medium, sugar source-deficient medium and rapamycin.
[0012] As a further preferred embodiment of the present invention, in step S1, the yeast cells in the normal yeast sample and the mutant yeast sample both correspond to Saccharomyces cerevisiae cells.
[0013] As a further preferred embodiment of the present invention, in step S1, time-series transcriptome, proteome and phosphoproteomic analysis is performed to obtain gene expression data, protein expression data and phosphoproteomic expression data of yeast samples before and after starvation induction treatment, specifically:
[0014] Time-sequential transcriptome analysis: Transcriptome sequencing of yeast samples was performed using a second-generation gene sequencer. After obtaining test data, the data was searched and quantitatively analyzed using the Trimmomatic-STAR-RSEM series software to obtain gene expression information and information on differentially expressed genes;
[0015] Proteomic analysis: Yeast samples were analyzed for proteome by liquid chromatography-mass spectrometry. After obtaining the test data, the data were searched and quantitatively analyzed using MaxQuant software to obtain protein distribution and intensity information. The intensity information was then filled with missing values and normalized using Perseus software.
[0016] Phosphoproteome analysis: Yeast samples were analyzed for phosphoproteome by liquid chromatography-mass spectrometry. After obtaining the test data, the parameters of the MaxQuant software were adjusted according to the requirements of the phosphorylation data analysis. Subsequently, the MaxQuant software was used to search the database and quantitatively analyze the data to obtain protein distribution and intensity information. The intensity information was then filled with missing values and normalized using Perseus software.
[0017] As a further preferred embodiment of the present invention, in step S1, gene expression data, protein expression data, and phosphorylated protein expression data of the yeast sample after starvation induction treatment are obtained, specifically, gene expression data, protein expression data, and phosphorylated protein expression data of the yeast sample after starvation induction treatment are obtained at different times;
[0018] Correspondingly, in step S3, the transcriptional expression changes, protein expression changes and phosphorylated protein expression changes of each gene in the set A before and after the knockout of the key regulatory gene for autophagy, and after the key regulatory gene for autophagy is complemented by the plasmid are extracted. Specifically, the transcriptional expression changes, protein expression changes and phosphorylated protein expression changes of each gene in the set A at different times before and after the knockout of the key regulatory gene for autophagy, as well as the transcriptional expression changes, protein expression changes and phosphorylated protein expression changes at different times after the key regulatory gene for autophagy is complemented by the plasmid are extracted.
[0019] As a further preferred embodiment of the present invention, in step S2, the protein interaction database is selected from the group consisting of BioGRID, DIP, HINT, IID, IntAct, iRefIndex, Mentha, MINT and STRING databases;
[0020] The autophagy gene database is THANATOS.
[0021] As a further preferred embodiment of the present invention, in step S3, the composition of the protein sequence corresponding to each gene is specifically calculated by calculating the frequencies of occurrence of 20 typical amino acids in the protein sequence.
[0022] As a further preferred embodiment of the present invention, in step S4, the deep learning algorithm is specifically constructed based on a neural network model.
[0023] As a further preferred embodiment of the present invention, in step S5, the optimization and data enhancement are specifically: using transfer learning to use the model to be processed as a pre-training model, using a meta-learning method to repeatedly sample positive data sets and negative data sets, and obtaining a final prediction model after multiple iterations.
[0024] Through the above technical solutions conceived by the present invention, compared with the prior art, the method of the present invention first screens yeast genes that interact with autophagy core genes (i.e., key autophagy regulatory genes) based on a public database of protein-protein interactions (the specific protein interaction database used can be pre-selected), and marks the known autophagy genes therein; then, based on the transcriptome, proteome and phosphoproteomic data, extracts the changes in the transcriptional expression levels, protein expression levels and phosphorylated protein expression levels of each gene before and after the knockout of the autophagy core gene (transcriptional expression level changes, protein expression level changes and phosphorylated protein expression level changes). The changes in expression levels are taken as the first, second and third characteristics of the gene respectively), and the frequency of occurrence of amino acids related to the gene is counted (as the fourth characteristic of the gene); further, based on the four types of characteristics, the above genes are trained using deep algorithms (such as deep neural networks) to make the known autophagy genes score higher; then, the previous model is used as a pre-training model using transfer learning, and the positive and negative data sets are repeatedly sampled using the meta-learning method, and the optimized model is obtained after multiple iterations; finally, the genes that rank high in the optimized model score are predicted to be important functional genes mediated by the core autophagy genes and serve as their phosphorylation substrate proteins.
[0025] This method combines transcriptomic, quantitative proteomic, and phosphoproteomic data to comprehensively consider changes in genes at the transcriptional, protein, and phosphoprotein levels. Furthermore, it introduces new features by statistically analyzing protein amino acid composition and uses a deep learning algorithm to rapidly predict kinase-specific substrate proteins of yeast autophagy. This method effectively narrows the screening range of kinase-specific substrate proteins, reduces the workload of experimental verification, and accurately predicts kinase-specific substrate proteins that regulate yeast autophagy.
[0026] If a functional phosphorylated protein participates in the regulation of a biological process, the kinase and its phosphorylated substrate that constitute the regulatory event must simultaneously play a regulatory function in the process, and phosphorylation modification is the main mechanism for regulating substrate function. Functional phosphorylation sites that are specifically regulated by the kinase need to exist on the substrate protein. Unlike the previous study CN113077841 by the research group of the inventor of the present invention, which can only predict functional genes that regulate yeast autophagy but cannot predict functional phosphorylated substrate proteins, the present invention integrates transcriptome, quantitative protein, and quantitative phosphorylated proteome data, and utilizes an artificial intelligence method based on deep learning to learn the characteristic information of known functional phosphorylated substrates, efficiently predict functional proteins specifically modified by kinase proteins, effectively narrow the scope of gene screening, reduce the workload of experimental verification, and discover kinase-specific, autophagy-regulated, functional phosphorylated substrate proteins. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 This is a flow chart of the method for predicting functional proteins involved in regulating autophagy based on the integration of multi-omics data.
[0028] Figure 2 Schematic diagram of the method for predicting functional proteins involved in regulating autophagy based on the integration of multi-omics data.
[0029] Figure 3 This is an experiment on the regulation of autophagy by Rgd1 and Whi5; Figure 3 A in the figure corresponds to the GFP-Atg8 cleavage detected by immunoblotting for WT and Rgd1Δ yeast strains (Δ in this application indicates knockout, for example, Rgd1Δ yeast strain represents a yeast strain with Rgd1 knockout) under nitrogen starvation treatment for 1 hour and 2 hours; Figure 3 Panel B corresponds to the ratio of free GFP detected in WT and Rgd1Δ yeast strains to the total amount of free GFP and GFP-Atg8 under nitrogen starvation conditions for 1 and 2 hours; Figure 3 Panel C corresponds to WT and Rgd1Δ yeast strains expressing GFP-Atg8. After 0, 1, and 2 hours of culture in SD-N medium, the yeast strains were stained with FM 4-64, and the GFP signals retained in the vacuoles were observed using confocal microscopy. Figure 3 D in the figure corresponds to the results of DeepPhagy analysis of WT and Rgd1△ yeast strains under nitrogen starvation for 1 hour and 2 hours. Figure 3 The autophagy activity obtained after analysis of the picture in C; Figure 3 E in the figure corresponds to the ALP activities of WT, Atg1△ and Rgd1△ yeast strains detected after 0 and 4 hours of nitrogen starvation treatment (the columns in the figure indicate that they are grouped in pairs from left to right, where group 1 corresponds to WT, group 2 corresponds to Atg1△, and group 3 corresponds to Rgd1△); Figure 3 F in the figure corresponds to the GFP-Atg8 cleavage detected by immunoblotting in WT and Whi5△ yeast strains under conditions of nitrogen starvation for 1 hour and 2 hours; Figure 3 G in the figure corresponds to the ratio of free GFP detected in WT and Whi5Δ yeast strains to the total amount of free GFP and GFP-Atg8 under nitrogen starvation conditions for 1 and 2 hours; Figure 3 H in the figure corresponds to WT and Whi5△ yeast strains expressing GFP-atg8. After 0, 1, and 2 hours of culture in SD-N medium, the cells were stained with FM 4-64, and the GFP signals retained in the vacuoles of the yeast strains were observed using confocal microscopy. Figure 3 The I in the figure corresponds to the WT and Whi5△ yeast strains treated with nitrogen starvation for 1 hour and 2 hours, which were detected by DeepPhagy. Figure 3The autophagy activity was obtained after analyzing the pictures in H; Figure 3 The J in the figure corresponds to the ALP activity detected in WT, Atg1△ and Whi5△ yeast strains after 0 and 4 hours of nitrogen starvation treatment (the columns in the figure are grouped in pairs from left to right, where group 1 corresponds to WT, group 2 corresponds to Atg1△, and group 3 corresponds to Whi5△).
[0030] Figure 4 The results show that Atg1 interacts with Rgd1 and Whi5 and may be involved in the phosphorylation of Whi5. Figure 4 Panel A corresponds to yeast strains expressing FLAG-tagged Atg1 and yeast strains expressing HA-tagged Rgd1, which were harvested and lysed, respectively, and subjected to immunoprecipitation analysis using anti-HA beads, and the immunoprecipitates were detected by immunoblotting using anti-FLAG and anti-HA antibodies; Figure 4 Panel B corresponds to the relative expression intensity of HA-tagged Rgd1 protein in the co-immunoprecipitation reaction; Figure 4 Panel C corresponds to yeast strains expressing FLAG-tagged Atg1 and yeast strains expressing HA-tagged Whi5, which were harvested and lysed, respectively, and subjected to co-immunoprecipitation analysis using anti-HA beads, and the immunoprecipitates were detected by immunoblotting using anti-FLAG and anti-HA antibodies; Figure 4 D in the figure corresponds to the relative expression intensity of HA-tagged Whi5 protein in the co-immunoprecipitation reaction; Figure 4 The E in the figure corresponds to the potential phosphorylation site on the Whi5 protein targeted by Atg1 predicted by the GPS algorithm; Figure 4 F in the figure corresponds to the expression intensity of phosphorylation sites on serine residues 78 and 149 of the Whi5 protein according to the phosphoproteomic data (pS indicates phosphorylated serine residues); Figure 4 G in the figure corresponds to immunoblot analysis of GFP-Atg8 cleavage in yeast cells expressing empty vector Ctrl, Whi5-WT, and Whi5-2A(S78 / 149A), respectively; Figure 4 H in the figure corresponds to the autophagic activity of yeast cells expressing empty vector Ctrl, Whi5-WT and Whi5-2A (S78 / 149A), respectively (the columns in the figure correspond to Ctrl, Whi5-WT and Whi5-2A from left to right). DETAILED DESCRIPTION
[0031] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only intended to illustrate the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below may be combined with each other as long as they do not conflict with each other.
[0032] Taking Saccharomyces cerevisiae as an example, the present invention may include the following steps to predict the kinase-specific substrate protein that regulates yeast autophagy:
[0033] (1) using nitrogen-reduced medium to treat normal Saccharomyces cerevisiae, Saccharomyces cerevisiae after knocking out the autophagy core gene, and Saccharomyces cerevisiae after the knocked-out autophagy core gene is complemented with a plasmid to obtain yeast samples before and after autophagy occurs; wherein the autophagy core gene (i.e., the key regulatory gene for autophagy) can be pre-selected, for example, it can be any known autophagy gene with a known autophagy function in an existing autophagy gene database (such as a public autophagy gene database);
[0034] (2) performing transcriptome, quantitative proteome, and phosphoproteomic identification on each yeast sample to obtain mRNA sequencing information, proteome identification information, and phosphoproteomic identification information of the yeast sample;
[0035] (3) Transcriptome data can be processed using the Trimmomatic-STAR-RSEM series software to obtain quantitative information of genes and information on significant changes in genes;
[0036] (4) Quantitative proteomic data can be searched and quantitatively analyzed using MaxQuant analysis software to obtain protein abundance information; Perseus software can be used to fill missing values and normalize the quantitative information;
[0037] (5) The MaxQuant analysis software is also used to search and quantitatively analyze the mass spectrometry data for the phosphorylated proteomic data. However, since the phosphorylated proteomic data contains phosphorylation site information, the MaxQuant software analysis parameters for conventional proteomic data cannot extract the phosphorylation site information in the mass spectrometry data. Therefore, the parameters of the MaxQuant software need to be adjusted. The specific adjustment step can be to add Phospho(STY) in the Variable modifications in the Modifications option in the group-specific parameters. The purpose is to allow the MaxQuant software to collect relevant information about phosphorylation sites during the search step; then the adjusted MaxQuant software is used to search and quantitatively analyze the data to obtain protein distribution and intensity information. Then, the intensity information is also filled with missing values and normalized using the Perseus software.
[0038] (6) Screening out genes that interact with autophagy core genes based on information from pre-selected protein interaction databases (e.g., BioGRID, DIP, HINT, IID, IntAct, iRefIndex, Mentha, MINT, and STRING);
[0039] Among them, the pre-selected protein interaction databases are all existing databases;
[0040] (7) Based on a pre-selected autophagy gene database (e.g., THANATOS), mark the known autophagy genes among the genes screened out in step (6); wherein the THANATOS database is an existing database;
[0041] (8) Based on the transcriptome, quantitative proteome and phosphorylated proteome data, the expression changes of the gene transcription level, protein level and phosphorylated protein level in step (6) are extracted, and the number of known autophagy genes that interact with the selected autophagy core genes (the number of known autophagy genes corresponding to each gene, i.e., protein interaction information) is collected by counting the number of known autophagy genes in step (6) and the pre-selected autophagy gene database (i.e., THANATOS database) using pre-selected protein interaction databases (i.e., BioGRID, DIP, HINT, IID, IntAct, iRefIndex, Mentha, MINT and STRING databases);
[0042] (9) Based on the known autophagy genes that interact with the selected autophagy core genes collected in step (8), the frequency of occurrence of amino acids in their corresponding proteins is counted as the fourth feature.
[0043] (10) Using the deep neural network model in the deep learning algorithm and the logistic regression model in the machine learning algorithm, the genes that are known to be involved in regulating autophagy genes and interacting with the autophagy core genes screened and marked in step (7) are used as the positive data set for training, and the remaining unknown genes that interact with the autophagy core genes screened in step (6) are used as the negative data set for training. Based on the gene expression information, protein expression information and phosphorylated protein expression information extracted in step (8), and based on the protein amino acid occurrence frequency extracted in step (9) as the training features of the model, the genes in step (6) are pre-trained with the deep neural network model separately to form four deep neural network models; the scores of each gene in the training data set are used as new features by the four feature pre-training models, and are integrated using the logistic regression model. The finally obtained logistic regression model can make the known autophagy genes rank high;
[0044] The Deep Neural Networks (DNN) algorithm used for model training can be directly called from the Keras open source package. The specific parameters used for training the model are as follows: loss = 'categorical_crossentropy', optimizer = 'adam', metrics = ['accuracy'], dropout = 0.1, input layer and hidden layer nodes = 200, activation = 'relu', output layer nodes = 2, activation = 'sigmoid';
[0045] The logistic regression (LR) algorithm used for model training can be directly called from the scikit-learn open source package. The specific parameters used for training the model are as follows: penalty = l2 (ridgeregression), C = 6.0, solver = 'lbfgs', multi_class = 'ovr', class_weight = 'balanced'.
[0046] (11) Based on published literature, mark the genes corresponding to the phosphorylated protein substrates of the autophagy core genes among the genes screened in step (6);
[0047] (12) Based on the four deep neural network models obtained in step (10), the genes whose corresponding proteins are known to be phosphorylated protein substrates of autophagy core genes and interact with autophagy core genes, which are screened and marked in step (11), are used as positive training data sets, and the remaining unknown genes that interact with autophagy core genes, which are screened in step (6), are used as negative training data sets. The positive data sets and negative data sets are repeatedly sampled using the meta-learning method, and the optimized model is obtained after multiple iterations;
[0048] (13) Using the optimized model obtained in step (12), each gene in step (6) is scored. The genes with higher scores (such as genes with scores greater than 0.9; the standard of scores greater than 0.9 can be flexibly adjusted according to actual conditions. Of course, the standard can also be based on the top percentage of the gene's position among all genes) and their corresponding proteins are predicted to be phosphorylated substrate proteins of the autophagy core gene.
[0049] In addition, the above framework that integrates multiple deep neural network models and logistic regression models can be called the hybrid meta-learning framework HUST-Atg1.
[0050] Most Atg proteins and autophagy regulators that interact with Atg1 undergo only modest changes at the mRNA, protein, or phosphorylation levels during autophagy. Therefore, traditional differential expression analysis makes it difficult to directly identify new Atg1-interacting proteins and substrates from multiple data sets. Considering that if other molecules exhibit similar molecular features to known Atg proteins and / or autophagy regulators, they may have similar functions and may also interact with Atg1, based on this idea, the present invention designed a method to predict kinase-specific substrate proteins that regulate yeast autophagy.
[0051] Preliminary collection of correspondences between autophagy genes and phosphorylation substrates can be done by consulting relevant published literature. For example, during our research, we collected eight known Atg1 phosphorylation substrates from existing literature, but directly training a computational model with such a small dataset would result in high overfitting and error-prone. To accurately and robustly predict Atg1 interacting partners and substrates, we preferentially designed the hybrid meta-learning framework HUST-Atg1. This hybrid meta-learning framework first integrates transcriptomics, proteomics, phosphoproteomics, and sequence features, using multiple deep neural network models for pre-training. These models are then integrated into a pre-trained model using a logistic regression model to predict Atg1 interacting partner proteins. Meta-learning methods are then used to transfer learning from the pre-trained model and fine-tune parameters to predict Atg1 phosphorylation substrates that may be involved in autophagy.
[0052] The following are specific embodiments:
[0053] Example 1
[0054] This example provides a method for predicting functional proteins involved in regulating autophagy based on multi-omics integration. Figure 1 and Figure 2 As shown, the following steps are included:
[0055] To prepare a benchmark dataset, 666 Atg1-interacting partners were identified from nine public databases. After mapping with autophagy-functional genes from THANATOS, 65 known Atg1-interacting Atg proteins and autophagy regulators were considered positive data, while the remaining 601 Atg1-interacting partners were considered negative data. For each Atg1-interacting partner, we considered sequence features including mRNA, protein, and phosphorylated protein expression levels, and amino acid composition as four types of informative features. In a hybrid learning model for predicting Atg1-interacting partners, we trained an initial DNN model for each feature using a DNN framework and employed the four generated scores as auxiliary features, trained by PLR, to obtain an optimized prediction score for each Atg1-interacting partner potentially involved in autophagy, including the Rgd1 protein.
[0056] To predict Atg1 substrates, we collected eight known Atg1 substrates, including Atg1, Atg2, Atg4, Atg6, Atg9, Atg13, Atg29, and Ykt6, as a benchmark dataset. To fully learn the characteristics of the positive dataset and mitigate the effects of data imbalance and overfitting during training, we employed meta-learning, a widely used small-shot learning strategy, to fine-tune the second model. Transfer learning was then used to transfer the new benchmark dataset to the hybrid learning model. The resulting optimized model was used to predict the potential involvement of Atg1 phosphorylated substrates in autophagy, including the Whi5 protein.
[0057] Example 2
[0058] We verified the function of Rgd1 in autophagy induced by nitrogen reduction ( Figure 3AE in). WT and rgd1Δ cells were transformed with a plasmid expressing GFP-Atg8, respectively, and then treated with SD-N medium for 0, 1, and 2 hours. Immunoblotting analysis showed that the ratio of free GFP to the total amount of free GFP and GFP-Atg8 was significantly lower in the rgd1Δ strain than in the WT strain. In addition, the accumulation of free GFP generated by GFP-Atg8 within the vacuole of the WT and rgd1Δ strains was observed by confocal fluorescence microscopy, and we found that the number of free GFP molecules retained within the vacuole was greater in the WT cells than in the rgd1Δ cells. The autophagy activity in the WT and rgd1Δ cells can then be quantitatively measured using our recently developed autophagy deep learning tool DeepPhagy[1]. Our analysis showed that the autophagy activity of rgd1Δ cells was significantly reduced compared with that of WT cells. In addition, to evaluate the effect of Rgd1 on yeast autophagy activity, the Pho8Δ60 assay was used as a quantitative readout. After 4 hours of nitrogen starvation, alkaline phosphatase (ALP) activity was detected in WT cells (as a positive control), atg1Δ cells (as a negative control), and rgd1Δ cells. The analysis showed that ALP activity in rgd1Δ cells was much lower than that in WT cells. Figure 3 (FJ in Figures 5 and 6). Using GFP-Atg8 immunoblot analysis, we found that GFP-Atg8 cleavage was significantly attenuated in WHI5-deficient yeast cells compared to WT cells. Furthermore, we observed that less free GFP was retained within the vacuole of whi5Δ cells than in WT cells. Consistent with this finding, autophagic activity quantified by DeepPhagy was lower in whi5Δ cells than in WT cells. Furthermore, using Pho8Δ60 analysis, we found that ALP activity was lower in whi5Δ cells than in WT cells under nitrogen starvation. Together, our results support the idea that both Rgd1 and Whi5 may be crucial for autophagy in Saccharomyces cerevisiae.
[0059] Furthermore, to verify whether the two proteins Rgd1 and Whi5 interact with Atg1, we generated yeast cells expressing FLAG-tagged Atg1 and glutathione S-transferase (GST)- and hemagglutinin (HA)-tagged Rgd1 or Whi5, respectively. By using Co-IP assays, we observed that both Rgd1 and Whi5 were potentially associated with Atg1 in vivo ( Figure 4 Furthermore, computational prediction analysis using our previously developed GPS algorithm [2] indicated that the serine residues at S78 and S149 of Whi5 are potential Atg1-mediated phosphorylation sites, so we conducted systematic experimental verification ( Figure 4EH in ). In addition, we carefully examined the intensity values of the p site on Whi5 from the phosphoproteomic dataset in this study. The omics data showed that pS78 of Whi5 was identifiable in WT cells at 1 hour and 4 hours, and in atg1Δ cells at 4 hours, while pS149 of Whi5 was only identifiable in WT cells at 1 hour. In addition, at 4 hours, the intensity value of pS78 of Whi5 in WT cells was higher than that in atg1Δ cells. Therefore, our results suggest that the loss of ATG1 may reduce the phosphorylation of pS78 and pS149 of Whi5. To explore whether these two serine residues on Whi5 play a role in nitrogen starvation-induced autophagy, we constructed a WHI5 mutant (Whi5-2A) using a site-directed mutagenesis method. Yeast cells expressing GFP-Atg8 were transformed with a control plasmid, a plasmid expressing complete Whi5, and a plasmid expressing the Whi5-2A mutant, respectively, and then cultured in SD-N medium for 1 hour. Using immunoblotting analysis, we observed that expression of intact Whi5 significantly increased autophagic activity, whereas expression of the Whi5 mutant had little effect on autophagy. These results suggest that two serine residues on Whi5 may be critical for Whi5's function in autophagy. Collectively, our results suggest that Rgd1 and Whi5 may interact with Atg1 and that Whi5 may be an Atg1 substrate and participate in nitrogen starvation-induced autophagy.
[0060] Based on the above results, it is proved that this method can accurately predict new functional proteins involved in regulating cellular autophagy, and therefore has important application value in the field of biological research.
[0061] The above examples are based on Saccharomyces cerevisiae. In addition to Saccharomyces cerevisiae, the method of the present invention is also applicable to other yeasts. The database (including the protein interaction database and the autophagy gene database) can also be flexibly adjusted and pre-selected according to the actual situation. In addition, regarding the establishment of the prediction model, other details not described in detail can be directly referred to related existing technologies, such as GPS-Palm[3], cMAK[4], etc.
[0062] The detailed information of the references used in the above article is as follows:
[0063] [1] Zhang, Y., et al., DeepPhagy: a deep learning framework for quantitatively measuring autophagy activity in Saccharomycescerevisiae. Autophagy, 2020.16(4):p.626-640.
[0064] [2]Xue, Y., et al., GPS 2.0, a tool to predict kinase-specific phosphorylation sites in hierarchy. Mol Cell Proteomics, 2008.7(9):p.1598-608.
[0065] [3]Ning, W., et al., GPS-Palm: a deep learning-based graphic presentation system for the prediction of S-palmitoylation sites in proteins. BriefBioinform, 2021.22(2):p.1836-1847.
[0066] [4]Peng, D., et al., Atg9-centered multi-omics integration reveals new autophagy regulators in Saccharomyces cerevisiae. Autophagy, 2021: p.1-24.
[0067] It will be easily understood by those skilled in the art that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for predicting kinase-specific substrate proteins that regulate yeast autophagy, characterized in that: The following steps are involved: S1: Based on pre-selected key autophagy regulatory genes, time-series transcriptome, proteome, and phosphoproteomic analyses were performed on normal yeast samples before and after starvation treatment, mutant yeast samples with knockout of the key autophagy regulatory gene, and mutant yeast samples with plasmid complementation of the key autophagy regulatory gene. Gene expression data, protein expression data, and phosphorylated protein expression data were obtained for yeast samples before and after starvation treatment. S2: Based on a pre-selected protein interaction database, genes that interact with the key autophagy regulatory genes are screened out and recorded as Set A; based on a pre-selected autophagy gene database, known autophagy genes with known autophagy functions in Set A are marked and recorded as Set B, which is a subset of Set A; and at the same time, based on the correspondence between autophagy genes and phosphorylated substrates collected in advance, autophagy genes with known phosphorylated substrates in Set A are marked and recorded as Set C, which is also a subset of Set A; S3: Extracting the transcriptional expression changes, protein expression changes, phosphorylated protein expression changes, and protein sequence composition corresponding to each gene in the set A before and after the knockout of the key autophagy regulatory gene and after the plasmid complementation of the key autophagy regulatory gene, as well as the four types of features for each gene in the set A; S4: Establish a prediction model, based on the four types of features determined in step S3, using a deep learning algorithm, using the genes in set B as the corresponding training positive dataset, and using the genes in set A but excluding the genes in set B as the corresponding training negative dataset, to train the prediction model; S5: Through the transfer learning method, the trained prediction model obtained in the step S4 is optimized with the genes in the set C corresponding to the positive training data set and the genes in the set A excluding the genes in the set C corresponding to the negative training data set, and the positive data set is enhanced by the meta-learning method to obtain the final prediction model; then, the genes in the set A are scored using the final prediction model, and the genes whose scores meet the pre-set requirements are encoded by the corresponding proteins, which are predicted to be kinase-specific substrate proteins that regulate yeast autophagy.
2. The method for predicting kinase-specific substrate proteins that regulate yeast autophagy according to claim 1, wherein: In the step S1, the autophagy inducer used in the starvation induction treatment is selected from nitrogen source-deficient medium, sugar source-deficient medium and rapamycin.
3. The method for predicting kinase-specific substrate proteins that regulate yeast autophagy according to claim 1, wherein: In step S1, the yeast cells in the normal yeast sample and the mutant yeast sample both correspond to Saccharomyces cerevisiae cells.
4. The method for predicting kinase-specific substrate proteins that regulate yeast autophagy according to claim 1, wherein: In step S1, sequential transcriptome, proteome, and phosphoproteomic analysis is performed to obtain gene expression data, protein expression data, and phosphoproteomic expression data of yeast samples before and after starvation induction treatment, specifically: Time-sequential transcriptome analysis: Transcriptome sequencing of yeast samples was performed using a second-generation gene sequencer. After obtaining test data, the data was searched and quantitatively analyzed using the Trimmomatic-STAR-RSEM series software to obtain gene expression information and information on differentially expressed genes; Proteomic analysis: Yeast samples were analyzed for proteome by liquid chromatography-mass spectrometry. After obtaining the test data, the data were searched and quantitatively analyzed using MaxQuant software to obtain protein distribution and intensity information. The intensity information was then filled with missing values and normalized using Perseus software. Phosphoproteome analysis: Yeast samples were analyzed for phosphoproteome by liquid chromatography-mass spectrometry. After obtaining the test data, the parameters of the MaxQuant software were adjusted according to the requirements of the phosphorylation data analysis. Subsequently, the MaxQuant software was used to search the database and quantitatively analyze the data to obtain protein distribution and intensity information. The intensity information was then filled with missing values and normalized using Perseus software.
5. The method for predicting kinase-specific substrate proteins that regulate yeast autophagy according to claim 1, wherein: In step S1, gene expression data, protein expression data, and phosphorylated protein expression data of the yeast sample after starvation induction treatment are obtained, specifically, gene expression data, protein expression data, and phosphorylated protein expression data of the yeast sample after starvation induction treatment are obtained at different times; Correspondingly, in step S3, the transcriptional expression changes, protein expression changes and phosphorylated protein expression changes of each gene in the set A before and after the knockout of the key regulatory gene for autophagy, and after the key regulatory gene for autophagy is complemented by the plasmid are extracted. Specifically, the transcriptional expression changes, protein expression changes and phosphorylated protein expression changes of each gene in the set A at different times before and after the knockout of the key regulatory gene for autophagy, as well as the transcriptional expression changes, protein expression changes and phosphorylated protein expression changes at different times after the key regulatory gene for autophagy is complemented by the plasmid are extracted.
6. The method for predicting kinase-specific substrate proteins that regulate yeast autophagy according to claim 1, wherein: In step S2, the protein interaction database is selected from nine databases: BioGRID, DIP, HINT, IID, IntAct, iRefIndex, Mentha, MINT and STRING; The autophagy gene database is THANATOS.
7. The method for predicting kinase-specific substrate proteins that regulate yeast autophagy according to claim 1, wherein: In step S3, the protein sequence composition corresponding to each gene is specifically calculated by calculating the frequencies of occurrence of 20 typical amino acids in the protein sequence.
8. The method for predicting kinase-specific substrate proteins that regulate yeast autophagy according to claim 1, wherein: In step S4, the deep learning algorithm is specifically constructed based on a neural network model.
9. The method for predicting kinase-specific substrate proteins that regulate yeast autophagy according to claim 1, wherein: In step S5, the optimization and data enhancement are specifically as follows: using transfer learning to use the model to be processed as a pre-training model, using a meta-learning method to repeatedly sample positive data sets and negative data sets, and obtaining a final prediction model after multiple iterations.
Citation Information
Patent Citations
Method for identifying functional kinase for regulating cell autophagy
CN110970087A
Method for predicting functional gene for regulating and controlling yeast autophagy
CN113077841A