Method for identifying key protein based on ViT model
Through the ViT model-based method, the characteristics of PPI network data and subcellular localization information data are combined, and the problem of insufficient subjective selection and global information perception ability when identifying key proteins in the prior art is solved, achieving a more efficient and accurate recognition effect.
Patent Information
- Application Number
- CN202510008836.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-06
AI Technical Summary
In the prior art, when identifying key proteins, there are problems such as subjective selection of subcellular locations, weak model perception of global information, and low recognition rate.
Using a ViT model-based method, the topological features of PPI network data were extracted through node2vec technology, and the subcellular location information data were encoded into index vectors of continuous values, the characteristics were fused through external product operations, and training was used to identify key proteins.
It improves the recognition rate of key proteins and has strong scalability, which can be expanded to other biomultiomic data, solves the problem of subjective selection of subcellular locations, and achieves more efficient and accurate identification.
Smart Images

Figure CN119943154A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of bioinformatics, and in particular to a method for identifying key proteins based on a ViT model. Background Art
[0002] Key proteins are the basis for cell reproduction and survival. If they are killed, the cells will stop reproducing or even die. Identifying key proteins is very important for revealing the molecular mechanisms of cells and discovering new drug targets.
[0003] With the development of computing technology and high-throughput technology, researchers have extracted features from multi-omics data to identify key proteins. For the identification method of a single feature, people have proposed a series of centrality methods mainly based on PPI network data. However, these methods have low ability to identify key proteins. Subsequently, researchers found that fusing multiple features helps to obtain a higher recognition rate. Therefore, researchers extracted multiple features from multi-omics data and combined them to identify key proteins. With the increase in the number of features, combined models such as random walk models and linear models become no longer applicable. Machine learning-based methods reflect the advantages of fusing a large number of features. Support vector machines (SVM), decision trees, and ensemble learning are widely used. However, these methods require manual extraction of multiple features from multi-omics data, and then input these features into some machine learning models, which requires a lot of manpower costs and adds subjective factors in feature selection. Then, deep learning-based methods are applied to the task of identifying key proteins. Usually, the deep learning-based method is an end-to-end framework. Multiple omics data are input into the model framework, and the model automatically extracts (learns) multiple features, then fuses these features, trains the model with training data, and finally tests the test data, that is, classifies proteins into two categories: key / non-key. Although these deep learning-based methods have better performance in the task of identifying key proteins, there is still room for further improvement in the recognition rate.
[0004] A large number of methods that use subcellular localization information data subjectively select multiple subcellular locations. For example, some researchers subjectively selected 11 common subcellular locations as the feature space and encoded each protein into an 11-dimensional vector using only 0 and 1. The MBIEP model selected more subcellular locations (1,024), but sorted the subcellular locations in descending order according to the number of proteins in each subcellular location. These methods extract features by subjectively selecting certain subcellular locations, which does not conform to the biological background of the protein.
[0005] In the method of identifying key proteins based on deep learning models, convolutional neural network (CNN) is mainly used. CNN can effectively extract features through multiple convolutional layers, but its perception ability is mainly strong for local information and weak for global information. Summary of the invention
[0006] The purpose of the present invention is to address the deficiencies in the prior art and to propose a method for identifying key proteins based on the ViT model. This method improves the recognition rate of key proteins and has strong scalability. It can be extended to other biological multi-omics data, providing new ideas and methods for the identification of key proteins and solving the problem of subjective selection of subcellular locations.
[0007] The technical solution for achieving the purpose of the present invention is:
[0008] The method for identifying key proteins based on the ViT model includes the following steps:
[0009] Step 1. Use node2vec technology to extract the topological structure features of PPI network data; first, load the PPI network data: PPI network data includes edge data and node label data of the PPI network, and use the networkx tool to process the PPI network data into a Graph object, then use node2vec technology to extract features from the PPI network data, encode each node in the PPI network data into a d-dimensional embedded vector representation, and then use the preprocessing.MinMaxScaler function in the scikit-learn tool to normalize the obtained embedded vector into a matrix Where n represents the number of nodes in the PPI network data, and the nodes in the PPI network data represent proteins;
[0010] Step 2. Encode each protein in the subcellular localization information data into an indicator vector with continuous values; use the preprocessing.MinMaxScaler function in the scikit-learn tool to normalize the confidence values of the subcellular localization information data to obtain a matrix with rows representing proteins and columns representing subcellular locations.
[0011] Step 3. The features extracted from the PPI network data and subcellular localization information data are fused through the outer product operation and input into the ViT model;
[0012] First, feature fusion is performed, and the formula is as follows:
[0013]
[0014] Where D(j) is the new feature matrix of protein j, P S (j) T represents the transposed embedding vector of protein j in the PPI network data after node2vec processing, S s (j) represents the index vector obtained from the subcellular localization information data of protein j, V represents the node set of PPI network data, and the formula (1) The symbol represents the outer product operation. The new feature matrix D(j) contains richer feature representations because D(j) combines the feature information of PPI network data and subcellular localization information data. H=W=d,C=1 is regarded as an image, (H,W) is the size of the image, i.e., the feature matrix D(j), C represents the number of channels, assuming that the feature matrix D(j) is a single-channel image, the number of channels C of the image is set to a constant 1, and the feature matrices of n proteins are concatenated as D=Concat(D(1),…,D(n));
[0015] Then use the utils.data.Dataset function in the PyTorch framework to customize the dataset and divide it into 80% training set and 20% test set. Then, use the tensor function in the PyTorch framework to convert the dataset into a PyTorch tensor, and then use the utils.data.DataLoader function in the PyTorch framework to build a DataLoader object.
[0016] Step 4. Train the ViT model and classify proteins. In order to avoid overfitting, K-fold cross validation is applied to train the ViT model to improve the robustness and reliability of the ViT model. In addition, since unbalanced data sets often affect the performance of the ViT model, the Focal Loss loss function is used to calculate the loss value during the training process. After the ViT model is trained, only the category token is extracted, and then the category token is input into the classifier to obtain the prediction score of the protein. For proteins with a prediction score greater than 0.5, the classification result is a key protein, and for proteins with a prediction score less than 0.5, the classification result is a non-key protein.
[0017] Compared with the prior art, the contributions of this technical solution are as follows:
[0018] The method of this technical solution aims to solve or alleviate the problems of subjective selection of subcellular locations, weak model perception of global information, and low recognition rate in existing methods, so as to achieve more efficient and accurate recognition of key proteins. When compared with 7 representative methods on three PPI network data, this technical solution showed better recognition ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 is a flow chart of an embodiment;
[0020] Figure 2 A flowchart of an embodiment for extracting features from subcellular data;
[0021] Figure 3 A distribution diagram of proteins in subcellular localization data appearing in each subcellular location;
[0022] Figure 4 The flowchart of the embodiment is to fuse the PPI network data features and the subcellular localization information data features through the outer product operation. DETAILED DESCRIPTION
[0023] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.
[0024] Reference Figure 1 , a method for identifying key proteins based on the ViT model, comprising the following steps:
[0025] Step 1. Use node2vec technology to extract the topological structure features of PPI network data. First, load the PPI network data: PPI network data includes edge data and node label data of the PPI network, and use the networkx tool to process the PPI network data into a Graph object. Then, use node2vec technology to extract features from the PPI network data, and encode each node (protein) in the PPI network data into a d-dimensional embedding vector representation. Then, use the preprocessing.MinMaxScaler function in the scikit-learn tool to normalize the obtained embedding vector into a matrix Where n represents the number of nodes (proteins) in the PPI network data, and the matrix P is obtained S The specific process is: n d-dimensional embedding vectors are concatenated into a matrix with n rows and d columns, and then the matrix is normalized using the MinMaxScaler function to obtain the matrix P S ;
[0026] Step 2. Encode each protein in the subcellular localization information data into an indicator vector with continuous values, such as Figure 2 , Figure 3 As shown, the confidence values of the subcellular localization information data are normalized using the preprocessing.MinMaxScaler function in the scikit-learn tool to obtain a matrix in which rows represent proteins and columns represent subcellular locations.
[0027] Step 3. Design a feature fusion method for biological multi-omics data, such as Figure 4 As shown, the features of PPI network data and subcellular localization information data are fused through the outer product operation to meet the shape of the ViT model input data, including the following:
[0028] First, feature fusion is performed, and the formula is as follows:
[0029]
[0030] Where D(j) is the new feature matrix of protein j, P S (j) T represents the transposed embedding vector of protein j in the PPI network data after node2vec processing, S S (j) represents the index vector obtained from the subcellular localization information data of protein j, V represents the node set of PPI network data, and the formula (1) The symbol represents the outer product operation. The new feature matrix D(j) contains richer feature representations because D(j) combines the feature information of PPI network data and subcellular localization information data. It can be regarded as an image, (H, W) is the size of the image (feature matrix D(j)), C represents the number of channels, the feature matrix D(j) is assumed to be a single-channel image, and the number of channels C of the image is set to a constant 1. In this way, the feature matrices of n proteins are concatenated as D = Concat (D(1), ..., D(n));
[0031] Then use the utils.data.Dataset function in the PyTorch framework to customize the dataset and divide the dataset into 80% training set and 20% test set. Then, use the tensor function in the PyTorch framework to convert the dataset into a PyTorch tensor, and then use the utils.data.DataLoader function in the PyTorch framework to build a DataLoader object.
[0032] Step 4. Train the ViT model and classify proteins into key proteins or non-key proteins, including the following:
[0033] First, the training set is used to train the ViT model by applying K-fold cross validation, and the performance of the ViT model is evaluated by an independent test set. Secondly, since unbalanced data sets will affect the performance of the ViT model, when the distribution of samples is unbalanced, the distribution of the loss function will also be skewed. In order to solve this problem, this example introduces the Focal Loss loss function in the process of training the ViT model to calculate the loss value. After the ViT model is trained, only the category token is extracted, and then the category token is input into the classifier to obtain the prediction score of the protein. For proteins with a prediction score greater than 0.5, the classification result is a key protein, and for proteins with a prediction score less than 0.5, the classification result is a non-key protein.
[0034] In order to verify the performance of the EPViT method proposed in this example, EPOC, ION, JDC, NCCO, MBIEP, DeepEP and deep learning framework are used as comparison methods, and comparative experiments are conducted on three PPI network datasets of yeast. The PPI network datasets used in this example are shown in Table 1. The parameters used in this example are shown in Table 2.
[0035] Table 1 Description of the PPI network dataset used in this example
[0036]
[0037] Table 2 Parameters used in this example
[0038]
[0039]
[0040] At this point, the entire processing process of this example is completed.
[0041] The comparison results of this example with the existing methods on three yeast PPI network datasets are shown below. The recognition rate of key proteins using this method shows excellent performance in common classification indicators: Accuracy, Precision, Recall, F1, AUC, and AUPR.
[0042] Compared with non-deep learning methods for identifying key proteins (EPOC, ION, JDC and NCCO):
[0043] The Accuracy, Precision, Recall, and F1 of the EPViT method proposed in this example on the PPI_5093 dataset are 0.836, 0.653, 0.656, and 0.654, respectively, which are better than EPOC (0.761, 0.496, 0.456, and 0.475), ION (0.769, 0.514, 0.444, and 0.477), JDC (0.751, 0.472, 0.456, and 0.464), and NCCO (0.761, 0.496, 0.456, and 0.475) on the PPI 5093 dataset.
[0044] The Accuracy, Precision, Recall, and F1 of the EPViT method proposed in this example on the PPI_3672 dataset are 0.807, 0.576, 0.663, and 0.616, respectively, which are better than EPOC (0.741, 0.452, 0.488, 0.469), ION (0.764, 0.497, 0.471, 0.484), JDC (0.706, 0.375, 0.384, 0.379), and NCCO (0.616, 0.214, 0.238, 0.225) on the PPI_3672 dataset.
[0045] The Accuracy, Precision, Recall, and F1 of the EPViT method proposed in this example on the PPI_2708 dataset are 0.806, 0.704, 0.614, and 0.656, respectively, which are better than EPOC (0.730, 0.550, 0.571, 0.560), ION (0.726, 0.548, 0.528, 0.538), JDC (0.673, 0.456, 0.448, 0.452), and NCCO (0.551, 0.256, 0.258, 0.257) on the PPI_2708 dataset.
[0046] Compared with the deep learning-based methods for identifying key proteins (MBIEP, DeepEP and deep learning framework), the performance comparison of the EPViT method proposed in this example and the deep learning-based methods on three PPI datasets is shown in Table 3:
[0047] Table 3. Performance comparison of this example and deep learning-based methods on three PPI datasets
[0048]
[0049] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, in order to highlight the advantages and benefits of the technical solution provided by the present invention, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as belonging to the protection scope of the present invention.
Claims
1. A method for identifying key proteins based on the ViT model, characterized in that: The steps include: Step 1. Use node2vec technology to extract topological structure features from PPI network data. First, load the PPI network data: PPI network data includes edge data and node label data of the PPI network, and use the networkx tool to process the PPI network data into a Graph object. Then use node2vec technology to extract features from the PPI network data, and encode each node in the PPI network data into a d-dimensional embedded vector representation. Then, use the preprocessing.MinMaxScaler function in the scikit-learn tool to normalize the obtained embedded vector into a matrix. Where n represents the number of nodes in the PPI network data, and the nodes in the PPI network data represent proteins; Step 2. Encode each protein in the subcellular localization information data into an indicator vector with continuous values, and use the preprocessing.MinMaxScaler function in the scikit-learn tool to normalize the confidence values of the subcellular localization information data to obtain a matrix with rows representing proteins and columns representing subcellular locations. Step 3. Design a feature fusion method for biological multi-omics data, fuse the features of PPI network data and subcellular localization information data through outer product operation, and satisfy the shape of ViT model input data, including the following: First, feature fusion is performed, and the formula is as follows: Where D(j) is the new feature matrix of protein j, P S (j) T represents the transposed embedding vector of protein j in the PPI network data after node2vec processing, S S (j) represents the index vector obtained from the subcellular localization information data of protein j, V represents the node set of PPI network data, and the formula (1) The symbol represents the outer product operation. The new feature matrix D(j) contains richer feature representations because D(j) combines the feature information of PPI network data and subcellular localization information data. H=W=d,C=1 is regarded as an image, (H,W) is the size of the image, i.e., the feature matrix D(j), C represents the number of channels, the feature matrix D(j) is assumed to be a single-channel image, the number of channels C of the image is set to a constant 1, and the feature matrices of n proteins are concatenated as D=Concat(D(1),…,D(n)); Then use the utils.data.Dataset function in the PyTorch framework to customize the dataset and divide it into 80% training set and 20% test set. Then, use the tensor function in the PyTorch framework to convert the dataset into a PyTorch tensor, and then use the utils.data.DataLoader function in the PyTorch framework to build a DataLoader object. Step 4. Train the ViT model and classify proteins into key proteins or non-key proteins, including the following: First, the training set is used to train the ViT model by applying K-fold cross validation, and the performance of the trained ViT model is evaluated by an independent test set. Secondly, since unbalanced data sets will affect the performance of the ViT model, when the distribution of samples is unbalanced, the distribution of the loss function will also be skewed. In order to solve this problem, the Focal Loss loss function is introduced in the process of training the ViT model to calculate the loss value. After the ViT model is trained, only the category token is extracted, and then the category token is input into the classifier to obtain the prediction score of the protein. For proteins with a prediction score greater than 0.5, the classification result is a key protein, and for proteins with a prediction score less than 0.5, the classification result is a non-key protein.