Key protein identification method based on protein network
By building a protein interaction network and using central algorithms and machine learning models, key proteins are identified, which solves the cost and time-consuming problems of traditional methods, and achieves efficient and economical identification of key proteins, providing unique advantages for disease mechanism analysis and drug target discovery.
Patent Information
- Application Number
- CN202510181746.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-10
AI Technical Summary
Traditional experimental methods are costly and time-consuming to identify key proteins, making it difficult to effectively solve the identification problems of biological processes and disease mechanisms.
By building a protein interaction network, the importance of proteins is predicted and optimized using central algorithms and machine learning models to identify key proteins.
It realizes key protein recognition from a global perspective, multi-data integration, efficient computing and interpretability perspectives, significantly reduces cost and time-consuming, and provides unique advantages in disease mechanism analysis, drug target discovery and synthetic biology design.
Smart Images

Figure CN120126543A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of bioinformatics and systems biology, and in particular to a key protein identification method based on protein network. Background Art
[0002] Proteins are the main molecules that perform various functions in organisms, and their interaction networks play an important role in biological processes and disease mechanisms. Key proteins refer to proteins with important functions in the network, and their identification is of great significance for understanding biological processes and disease mechanisms. Traditional experimental methods for identifying key proteins are costly and time-consuming. Summary of the invention
[0003] Based on this, it is necessary to provide a key protein identification method based on protein network to address the above technical problems.
[0004] A method for identifying key proteins based on protein networks, comprising the following steps:
[0005] Acquiring protein interaction data; wherein the protein interaction data includes: proteins and corresponding interactions;
[0006] constructing a protein interaction network according to the protein interaction data;
[0007] Acquire a centrality algorithm, analyze the protein interaction network according to the centrality algorithm, and obtain a centrality value of the protein;
[0008] Obtaining protein features, and inputting the protein features into an optimized machine learning model to obtain a protein importance prediction result; wherein the protein features include: a protein centrality value, a protein sequence feature, and a protein functional annotation;
[0009] Optimizing the protein importance prediction results by a mathematical optimization algorithm to obtain key proteins;
[0010] A protein analysis result is obtained based on the key protein.
[0011] In one embodiment, obtaining protein interaction data comprises:
[0012] Protein interaction source data are obtained, and the protein interaction source data are preprocessed to remove low-confidence interactions to obtain protein interaction data.
[0013] In one embodiment, constructing a protein interaction network based on the protein interaction data comprises:
[0014] The proteins are used as nodes and the corresponding interactions as edges to construct the initial protein interaction network;
[0015] The initial protein interaction network is subjected to denoising processing, and the network is converted into an undirected graph or a directed graph model to obtain a protein interaction network.
[0016] In one embodiment, the centrality algorithm includes:
[0017] Such as degree centrality, betweenness centrality, closeness centrality and eigenvector centrality.
[0018] In one embodiment, obtaining protein features, inputting the protein features into an optimized machine learning model, and obtaining a protein importance prediction result further includes:
[0019] A test protein feature is obtained, a training set and a test set are constructed using the test protein feature, and a machine learning model is trained using a cross-validation method to obtain an optimized machine learning model.
[0020] In one embodiment, the machine learning model includes:
[0021] Random forests, support vector machines, and protein networks.
[0022] In one embodiment, the mathematical optimization algorithm includes:
[0023] Linear programming and graph cut algorithms.
[0024] In one embodiment, the protein importance prediction results are optimized by a mathematical optimization algorithm, and the key proteins obtained include:
[0025] The proteins are ranked according to their centrality values and the predicted importance results of the proteins;
[0026] The sorting results are optimized by mathematical optimization algorithms to obtain optimized protein structures;
[0027] Acquire experimental verification data, and obtain key proteins based on the experimental verification data and the optimized protein structure.
[0028] A key protein identification system based on protein network, used to implement the key protein identification method based on protein network as described above, comprising:
[0029] A data acquisition module, used to acquire protein interaction data; wherein the protein interaction data includes: proteins and corresponding interactions;
[0030] A network construction module, used for constructing a protein interaction network according to the protein interaction data;
[0031] A numerical calculation module, used for obtaining a centrality algorithm, analyzing the protein interaction network according to the centrality algorithm, and obtaining a centrality value of the protein;
[0032] A result prediction module is used to obtain protein features, input the protein features into the optimized machine learning model, and obtain protein importance prediction results; wherein the protein features include: protein centrality value, protein sequence features and protein functional annotations;
[0033] A result optimization module, used to optimize the protein importance prediction results through a mathematical optimization algorithm to obtain key proteins;
[0034] The feature analysis module is used to obtain protein analysis results based on the key protein.
[0035] A device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of a key protein identification method based on a protein network described in each of the above embodiments are implemented.
[0036] Compared with the existing technology, the advantages and beneficial effects of the present invention are: the present invention shows unique advantages in disease mechanism analysis, drug target discovery and synthetic biology design through global perspective, multi-data integration, efficient calculation and interpretability. It converts the complexity of biological systems into computable network models, providing powerful tools for precision medicine and functional genomics. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 A schematic diagram of a process of a key protein identification method based on a protein network in one embodiment;
[0038] Figure 2 A schematic diagram of the structure of a key protein recognition system based on a protein network in one embodiment;
[0039] Figure 3 Schematic diagram of the internal structure of a device in one embodiment. DETAILED DESCRIPTION
[0040] Before describing the specific embodiments of the present invention, the overall concept of the present invention is described as follows:
[0041] The present invention is mainly developed for the key protein identification process. Currently, the experimental identification of key proteins is costly and time-consuming.
[0042] Therefore, the present invention proposes a key protein identification method based on protein network, which constructs a protein interaction network and converts it into a graph model, and uses specific graph theory algorithms and machine learning methods to analyze the network to identify key proteins.
[0043] After introducing the overall concept of the present invention, in order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below through specific implementation methods in conjunction with the accompanying drawings.
[0044] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in one or more embodiments of this specification should be understood by people with ordinary skills in the field to which the present invention belongs. The words "first", "second" and similar words used in one or more embodiments of this specification do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0045] In one embodiment, Figure 1 As shown, a key protein identification method based on protein network is provided, comprising the following steps:
[0046] Step S101, obtaining protein interaction data; wherein the protein interaction data includes: proteins and corresponding interactions.
[0047] Specifically, protein interaction source data are obtained from public databases, including but not limited to STRING, BioGRID and IntAct.
[0048] On this basis, protein interaction data are obtained including:
[0049] Protein interaction source data are obtained, and the protein interaction source data are preprocessed to remove low-confidence interactions to obtain protein interaction data.
[0050] Specifically, the obtained protein interaction source data are preprocessed to remove low-confidence interactions to obtain protein interaction data.
[0051] Step S102: constructing a protein interaction network according to the protein interaction data.
[0052] Specifically, a protein interaction network is constructed based on the protein interaction data and converted into a graph model.
[0053] On this basis, constructing a protein interaction network according to the protein interaction data includes:
[0054] The proteins are used as nodes and the corresponding interactions as edges to construct the initial protein interaction network;
[0055] The initial protein interaction network is subjected to denoising processing, and the network is converted into an undirected graph or a directed graph model to obtain a protein interaction network.
[0056] Specifically, proteins are used as nodes and corresponding interactions as edges to construct an initial protein interaction network; the initial protein interaction network is denoised to remove low-confidence interactions, and the network is converted into an undirected graph or directed graph model to obtain a protein interaction network.
[0057] Step S103, obtaining a centrality algorithm, analyzing the protein interaction network according to the centrality algorithm, and obtaining a centrality value of the protein.
[0058] Specifically, a centrality algorithm is used to perform a preliminary analysis of the protein interaction network to obtain the centrality value of the protein.
[0059] On this basis, centrality algorithms include:
[0060] Such as degree centrality, betweenness centrality, closeness centrality and eigenvector centrality.
[0061] Specifically, the degree centrality calculation formula is as follows:
[0062]
[0063] Among them, C D (v) is the calculated value of degree centrality, deg(v) is the degree of node v, and n is the total number of network nodes.
[0064] The calculation formula of betweenness centrality is as follows:
[0065]
[0066] Among them, C B (v) is the calculated value of betweenness centrality, σ st is the total number of shortest paths from node s to node t, σ st is the number of shortest paths passing through node v.
[0067] The formula for calculating closeness centrality is as follows:
[0068]
[0069] Among them, C C (v) is the calculated value of closeness centrality, d(u,v) is the shortest path length from node u to node v, and n is the total number of network nodes.
[0070] The formula for calculating eigenvector centrality is as follows:
[0071]
[0072] Among them, C E (v) is the calculated value of eigenvector centrality, λ is the eigenvalue, and N(v) is the set of neighbor nodes of node v.
[0073] Step S104, obtaining protein features, inputting the protein features into the optimized machine learning model to obtain protein importance prediction results; wherein the protein features include: protein centrality value, protein sequence features and protein functional annotations.
[0074] Specifically, the sequence features of proteins are obtained through databases such as UniProt, NCBI Protein Database, Pfam, InterPro, ProtParam, etc. The functional annotations of proteins are obtained through databases such as UniProt, Gene Ontology (GO) Database, KEGG Pathway, AlphaFold Protein Structure Database, etc.
[0075] The protein features such as protein centrality value, protein sequence characteristics and protein functional annotation are input into the trained optimized machine learning model to obtain the protein importance prediction results.
[0076] On this basis, obtaining protein features and inputting the protein features into the optimized machine learning model to obtain the protein importance prediction results also includes:
[0077] A test protein feature is obtained, a training set and a test set are constructed using the test protein feature, and a machine learning model is trained using a cross-validation method to obtain an optimized machine learning model.
[0078] On this basis, the machine learning model includes:
[0079] Random forests, support vector machines, and protein networks.
[0080] Specifically, the test protein features are obtained, and the test protein feature data set is divided into a training set and a test set. Usually 80% of the data is used as the training set and 20% of the data is used as the test set to ensure that the data set is evenly distributed and avoid bias. Use cross-validation methods (such as k-fold cross-validation) to optimize model parameters. Common cross-validation methods include: k-Fold Cross-Validation: Divide the data set into k subsets, and use one of the subsets as the validation set in turn, and the rest as the training set. Leave-One-Out Cross-Validation (LOOCV): Use one sample as the validation set each time, and the rest as the training set.
[0081] Machine learning models include: Random Forest: Applicable to high-dimensional data and can handle nonlinear relationships. Support Vector Machine (SVM): Applicable to small sample data and can handle high-dimensional features. Protein Interaction Networks: A method based on Graph Neural Networks (GNN), applicable to protein interaction networks.
[0082] Use Grid Search or Random Search to optimize model hyperparameters.
[0083] The model performance was evaluated using metrics such as accuracy, precision, recall, F1 score, and area under the ROC curve.
[0084] Models such as random forests can provide feature importance scores, which can be used to interpret model predictions through SHAP values, assessing the contribution of each feature to the prediction to predict the importance of proteins.
[0085] Step S105, optimizing the protein importance prediction result by a mathematical optimization algorithm to obtain key proteins.
[0086] Specifically, the mathematical optimization algorithm includes:
[0087] Linear programming and graph cut algorithms.
[0088] Specifically, the linear programming calculation formula is as follows:
[0089] minimize c T xsubject to Ax≤b,x≥0
[0090] Among them, c is the objective function coefficient vector, A is the constraint matrix, and b is the constraint vector.
[0091] The calculation formula of the graph cut algorithm is as follows:
[0092]
[0093] Among them, w ij is the edge weight, x i and x j is the node label.
[0094] On this basis, the protein importance prediction results were optimized through mathematical optimization algorithms, and the key proteins obtained included:
[0095] The proteins are ranked according to their centrality values and the predicted importance results of the proteins;
[0096] The sorting results are optimized by mathematical optimization algorithms to obtain optimized protein structures;
[0097] Acquire experimental verification data, and obtain key proteins based on the experimental verification data and the optimized protein structure.
[0098] Specifically, the proteins are first ranked, and the ranking methods include but are not limited to: weighted ranking fusion, assigning weights to each feature (such as centrality value, prediction score) and calculating the comprehensive score; non-dominated ranking, for multi-objective optimization problems (such as maximizing centrality and prediction score at the same time), using the Pareto frontier to select the optimal solution set.
[0099] Then, the sorting results are optimized by mathematical optimization methods to obtain the optimized protein structure.
[0100] Finally, the functional importance of the key protein is verified by experiments, experimental verification data is obtained, and the key protein is obtained according to the experimental verification data and the optimized protein structure. The experiment includes: wet experiment and dry experiment.
[0101] Step S106, obtaining protein analysis results according to the key protein.
[0102] Specifically, functional annotation and pathway analysis of key proteins are performed to further understand their roles in biological processes and disease mechanisms and obtain protein analysis results.
[0103] This method is applicable to a variety of organisms, including humans, mice, and yeast, and can be applied to the identification of disease-related proteins and the discovery of drug targets.
[0104] The present invention provides a method for identifying key proteins based on protein networks, which shows unique advantages in disease mechanism analysis, drug target discovery and synthetic biology design through global perspective, multi-data integration, efficient calculation and interpretability. The protein interaction network reflects the collaborative relationship of proteins in cells. By analyzing the network topology (such as centrality and modularity), proteins that are critical to network stability or function can be discovered. Network features (such as centrality values), sequence features (such as conservation), functional annotations (GO, KEGG pathways) and experimental data (such as expression profiles) can be integrated and weighted through machine learning models (such as random forests and graph neural networks) to reduce the deviation of a single data source. The complexity of biological systems is converted into computable network models, providing a powerful tool for precision medicine and functional genomics.
[0105] It should be noted that the method of the embodiment of the present invention can be performed by a single device, such as a computer or a server. The method of this embodiment can also be applied in a distributed scenario and completed by multiple devices cooperating with each other. In the case of such a distributed scenario, one of the multiple devices can only perform one or more steps in the method of the embodiment of the present invention, and the multiple devices will interact with each other to complete the described method.
[0106] It should be noted that some embodiments of the present invention are described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the above embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0107] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present invention also provides a key protein identification system based on a protein network.
[0108] refer to Figure 2 , the key protein recognition system based on protein network comprises:
[0109] The data acquisition module 201 is used to acquire protein interaction data; wherein the protein interaction data includes: proteins and corresponding interactions;
[0110] A network construction module 202, used to construct a protein interaction network according to the protein interaction data;
[0111] A numerical calculation module 203 is used to obtain a centrality algorithm, analyze the protein interaction network according to the centrality algorithm, and obtain a centrality value of the protein;
[0112] The result prediction module 204 is used to obtain protein features, input the protein features into the optimized machine learning model, and obtain protein importance prediction results; wherein the protein features include: protein centrality value, protein sequence features and protein functional annotations;
[0113] A result optimization module 205 is used to optimize the protein importance prediction results by a mathematical optimization algorithm to obtain key proteins;
[0114] The feature analysis module 206 is used to obtain protein analysis results based on the key proteins.
[0115] For the convenience of description, the above system is described as being divided into various modules according to their functions. Of course, when implementing the present invention, the functions of each module can be implemented in the same or multiple software and / or hardware.
[0116] The system of the above embodiment is used to implement a corresponding protein network-based key protein identification method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.
[0117] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, a key protein identification method based on a protein network as described in any of the above embodiments is implemented.
[0118] Figure 3 A more specific schematic diagram of the hardware structure of an electronic device provided in this embodiment is shown, and the device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are connected to each other through the bus 1050 in the device.
[0119] The processor 1010 can be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0120] The memory 1020 may be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 may store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.
[0121] The input / output interface 1030 is used to connect the input / output module to realize information input and output. The input / output module can be configured in the device as a component (not shown in the figure), or it can be externally connected to the device to provide corresponding functions. The input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.
[0122] The communication interface 1040 is used to connect a communication module (not shown) to realize communication interaction between the device and other devices. The communication module can realize communication through a wired mode (such as USB, network cable, etc.) or a wireless mode (such as mobile network, WIFI, Bluetooth, etc.).
[0123] The bus 1050 includes a path that transmits information between the various components of the device (eg, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).
[0124] It should be noted that, although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040 and the bus 1050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it can be understood by those skilled in the art that the above device may also only include the components necessary for implementing the embodiments of the present specification, and does not necessarily include all the components shown in the figure.
[0125] The electronic device of the above embodiment is used to implement a corresponding key protein identification method based on a protein network in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.
[0126] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present invention (including the claims) is limited to these examples. Under the concept of the present invention, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present invention as described above, which are not provided in detail for the sake of simplicity.
[0127] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code that includes one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present invention includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present invention belong.
[0128] In addition, to simplify the description and discussion, and in order not to obscure the embodiments of the present invention, known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided figures. In addition, the system may be shown in the form of a block diagram so as to avoid obscuring the embodiments of the present invention, and this also takes into account the fact that the details of the implementation of these block diagram systems are highly dependent on the platform on which the embodiments of the present invention will be implemented (i.e., these details should be fully within the scope of understanding of those skilled in the art). Where specific details (e.g., circuits) are set forth to describe exemplary embodiments of the present invention, it will be apparent to those skilled in the art that embodiments of the present invention may be implemented without these specific details or with variations in these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0129] Although the invention has been described in conjunction with specific embodiments of the invention, many replacements, modifications and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may use the embodiments discussed.
[0130] The embodiments of the present invention are intended to cover all such substitutions, modifications and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present invention should be included in the protection scope of the present invention.
Claims
1. A key protein identification method based on protein network, characterized in that: include: Acquiring protein interaction data; wherein the protein interaction data includes: proteins and corresponding interactions; constructing a protein interaction network according to the protein interaction data; Acquire a centrality algorithm, analyze the protein interaction network according to the centrality algorithm, and obtain a centrality value of the protein; Obtaining protein features, and inputting the protein features into an optimized machine learning model to obtain a protein importance prediction result; wherein the protein features include: a protein centrality value, a protein sequence feature, and a protein functional annotation; Optimizing the protein importance prediction results by a mathematical optimization algorithm to obtain key proteins; A protein analysis result is obtained based on the key protein.
2. A method for identifying key proteins based on protein networks according to claim 1, characterized in that: The obtaining of protein interaction data comprises: Protein interaction source data are obtained, and the protein interaction source data are preprocessed to remove low-confidence interactions to obtain protein interaction data.
3. A key protein identification method based on protein network according to claim 1, characterized in that: The constructing a protein interaction network according to the protein interaction data comprises: The proteins are used as nodes and the corresponding interactions as edges to construct the initial protein interaction network; The initial protein interaction network is subjected to denoising processing, and the network is converted into an undirected graph or a directed graph model to obtain a protein interaction network.
4. The method for identifying key proteins based on protein network according to claim 1, characterized in that: The centrality algorithm includes: Such as degree centrality, betweenness centrality, closeness centrality and eigenvector centrality.
5. The method for identifying key proteins based on protein network according to claim 1, characterized in that: The method of obtaining protein features and inputting the protein features into the optimized machine learning model to obtain protein importance prediction results also includes: A test protein feature is obtained, a training set and a test set are constructed using the test protein feature, and a machine learning model is trained using a cross-validation method to obtain an optimized machine learning model.
6. A method for identifying key proteins based on protein networks according to claim 5, characterized in that: The machine learning model includes: Random forests, support vector machines, and protein networks.
7. The method for identifying key proteins based on protein network according to claim 1, characterized in that: The mathematical optimization algorithm includes: Linear programming and graph cut algorithms.
8. The method for identifying key proteins based on protein network according to claim 1, characterized in that: The protein importance prediction results are optimized by a mathematical optimization algorithm to obtain key proteins including: The proteins are ranked according to their centrality values and the predicted importance results of the proteins; The sorting results are optimized by mathematical optimization algorithms to obtain optimized protein structures; Acquire experimental verification data, and obtain key proteins based on the experimental verification data and the optimized protein structure.
9. A key protein identification system based on protein network, characterized in that: A method for realizing a key protein identification method based on a protein network as claimed in any one of claims 1 to 8, comprising: A data acquisition module, used to acquire protein interaction data; wherein the protein interaction data includes: proteins and corresponding interactions; A network construction module, used for constructing a protein interaction network according to the protein interaction data; A numerical calculation module, used for obtaining a centrality algorithm, analyzing the protein interaction network according to the centrality algorithm, and obtaining a centrality value of the protein; A result prediction module is used to obtain protein features, input the protein features into the optimized machine learning model, and obtain protein importance prediction results; wherein the protein features include: protein centrality value, protein sequence features and protein functional annotations; A result optimization module, used to optimize the protein importance prediction results through a mathematical optimization algorithm to obtain key proteins; The feature analysis module is used to obtain protein analysis results based on the key protein.
10. A device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.