Biological signal transduction pathway analysis method based on protein network

By constructing a cell-specific protein interaction network and using machine learning algorithms, the problem that traditional analytical methods are difficult to discover new signaling mechanisms is solved, achieving higher prediction accuracy and wider application scope.

CN120126541APending Publication Date: 2025-06-10CHONGQING RUANJIANG TURING ARTIFICIAL INTELLIGENCE TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510181647.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

Traditional signaling pathway analysis methods rely on known biological knowledge, making it difficult to discover new signaling mechanisms, and are limited by data singularity, narrow analysis perspective and inefficient computing.

Method used

Using a biological signaling path analysis method based on protein networks, a cell-specific protein interaction network is constructed by acquiring and preprocessing multiomics data, topological features are extracted, and signaling pathways are predicted using machine learning algorithms.

Benefits of technology

It realizes a global perspective to reveal complex signaling mechanisms, improve prediction accuracy, and is suitable for disease mechanism research, drug target discovery and biomarker screening.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126541A_ABST
    Figure CN120126541A_ABST
Patent Text Reader

Abstract

The invention provides a biological signal transduction pathway analysis method based on a protein network, which comprises the following steps: acquiring omics data, and preprocessing the omics data to obtain preprocessed omics data; constructing a cell specific protein interaction network according to the preprocessing omics data; extracting topological structure characteristics from the cell specific protein interaction network; obtaining a signal transduction path according to the topological structure features and a network topological feature prediction model; and obtaining a signal transduction path analysis result according to the signal transduction path. According to the method, multi-omics data and a machine learning algorithm are combined, a complex signal transduction mechanism can be revealed from a global perspective, and the prediction accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of bioinformatics technology, and in particular to a biological signal transduction pathway analysis method based on protein network. Background Art

[0002] Biological signal transduction pathways are important mechanisms for cells to respond to external stimuli, and their abnormalities are closely related to the occurrence and development of a variety of diseases. Traditional signal pathway analysis methods mostly rely on known biological knowledge and are difficult to discover new signal transduction mechanisms. With the development of high-throughput sequencing technology, a large amount of protein interaction data has been generated, providing new opportunities for analyzing signal transduction pathways at the system level. Summary of the invention

[0003] Based on this, it is necessary to provide a biological signal transduction pathway analysis method based on protein networks to address the above technical problems.

[0004] A biological signal transduction pathway analysis method based on a protein network comprises the following steps:

[0005] Acquiring omics data, and preprocessing the omics data to obtain preprocessed omics data;

[0006] constructing a cell-specific protein interaction network based on the preprocessed omics data;

[0007] Extracting topological structural features from the cell-specific protein interaction network; wherein the topological features include: node degree centrality, betweenness centrality, closeness centrality and modularity;

[0008] Obtaining a signal transduction pathway according to the topological structure characteristics and the network topological characteristics prediction model;

[0009] A signal transduction pathway analysis result is obtained according to the signal transduction pathway.

[0010] In one embodiment, obtaining omics data, and preprocessing the omics data to obtain preprocessed omics data includes:

[0011] Performing data cleaning on the omics data to obtain cleaned omics data;

[0012] The cleaned omics data are subjected to data standardization processing to obtain pre-processed omics data.

[0013] In one embodiment, constructing a cell-specific protein interaction network based on the pre-processed omics data comprises:

[0014] Obtain protein interactome data, and obtain an initial cell-specific protein interaction network according to the protein interactome data;

[0015] Dynamically enhance the initial cell-specific protein interaction network according to the preprocessed omics data to obtain a cell-specific protein interaction network.

[0016] In one embodiment, before obtaining the signal transduction pathway according to the topological structure features and the network topological feature prediction model, it further includes:

[0017] Obtain a feature training sample set, construct a training set and a test set with the feature training sample set, and use the cross-validation method to train a machine learning model to obtain a network topological feature prediction model; wherein, the feature training sample set includes: sample cell-specific protein features and corresponding signal transduction pathways.

[0018] In one embodiment, the machine learning model includes:

[0019] Random forest, support vector machine and protein network.

[0020] In one embodiment, obtaining the signal transduction pathway analysis result according to the signal transduction pathway includes:

[0021] Perform a pathway prediction operation on the signal transduction pathway to obtain a pathway prediction result;

[0022] Perform a functional annotation enrichment analysis on the railway prediction result to obtain a signal transduction pathway analysis result.

[0023] In one embodiment, performing a pathway prediction operation on the signal transduction pathway to obtain a pathway prediction result includes:

[0024] Score the signal transduction pathway to obtain a prediction score:

[0025]

[0026] Among them, S(p) represents the prediction score of pathway p, w i represents the weight of the i-th feature, and f i (p) represents the value of pathway p on the i-th feature;

[0027] Sort the prediction scores, and use the pathways with prediction scores higher than the preset threshold as the prediction score results.

[0028] A biological signal transduction pathway analysis system based on a protein network, which is used to implement a biological signal transduction pathway analysis method based on a protein network as described above, includes:

[0029] An omics data acquisition module, configured to acquire omics data, preprocess the omics data, and obtain preprocessed omics data;

[0030] An interaction acquisition module, configured to construct a cell-specific protein interaction network according to the preprocessed omics data;

[0031] A topological feature calculation module, configured to extract topological structure features from the cell-specific protein interaction network; wherein, the topological features include: node degree centrality, betweenness centrality, and closeness centrality;

[0032] A signal conduction pathway acquisition module, configured to obtain a signal conduction pathway according to the topological structure features and a network topological feature prediction model;

[0033] A pathway analysis acquisition module, configured to obtain a signal conduction pathway analysis result according to the signal conduction pathway.

[0034] A device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, the steps of a method for analyzing a biological signal conduction pathway based on a protein network described in each of the above embodiments are implemented.

[0035] A storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of a method for analyzing a biological signal conduction pathway based on a protein network described in each of the above embodiments are implemented.

[0036] Compared with the prior art, the advantages and beneficial effects of the present invention are as follows: The present invention combines multi-omics data and machine learning algorithms, can reveal complex signal conduction mechanisms from a global perspective, and improves prediction accuracy. The present invention is applicable to disease mechanism research, drug target discovery, and biomarker screening. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 It is a schematic flowchart of a method for analyzing a biological signal conduction pathway based on a protein network in an embodiment;

[0038] Figure 2 It is a schematic structural diagram of a system for analyzing a biological signal conduction pathway based on a protein network in an embodiment;

[0039] Figure 3 It is a schematic internal structure diagram of a device in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] Before describing the specific embodiments of the present invention in detail, the overall concept of the present invention is described as follows:

[0041] The present invention is mainly developed in the process of biological signal transduction pathway analysis. At present, most signal pathway analysis methods rely on known biological knowledge and it is difficult to discover new signal transduction mechanisms. Traditional methods are often limited by problems such as data singularity, narrow analysis perspective, or low computational efficiency, and it is difficult to comprehensively and deeply explore the complex and delicate signal transduction network in organisms.

[0042] Therefore, the present invention proposes a method for analyzing biological signal transduction pathways based on protein networks, integrating multi-omics data, constructing a cell-specific protein interaction network, and using network topology structure features and machine learning algorithms to predict potential signal transduction pathways, providing new ideas for disease mechanism research and drug target discovery.

[0043] After introducing the overall concept of the present invention, in order to make the purpose, technical solution and advantages of the present invention clearer, the following further details the present invention through specific embodiments in conjunction with the accompanying drawings.

[0044] It should be noted that unless otherwise defined, the technical terms or scientific terms used in one or more embodiments of this specification should be the ordinary meanings understood by those of ordinary skill in the art to which the present invention belongs. The "first", "second" and similar terms used in one or more embodiments of this specification do not denote any order, quantity or importance, but are only used to distinguish different components. Words such as "including" or "comprising" mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Words such as "connected" or "linked" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right", etc. are only used to represent relative position relationships, and when the absolute position of the object being described changes, the relative position relationship may also change accordingly.

[0045] In one embodiment, as Figure 1 shown, a method for analyzing biological signal transduction pathways based on protein networks is provided, including the following steps:

[0046] Step S101, obtain omics data and preprocess the omics data to obtain preprocessed omics data.

[0047] Specifically, collect multiple biological information data sources and integrate high-quality multi-omics data. The multi-omics data includes, but is not limited to, proteomics, transcriptomics, metabolomics and other data, which come from different experimental platforms or public databases. Preprocess the multi-omics data to obtain preprocessed omics data. Data preprocessing is the starting step of the entire analysis method, aiming to clean and standardize the multi-omics data to ensure the quality and consistency of the data.

[0048] On this basis, obtain omics data and preprocess the omics data to obtain preprocessed omics data, including:

[0049] Perform data cleaning on the omics data to obtain cleaned omics data;

[0050] Perform data standardization on the cleaned omics data to obtain preprocessed omics data.

[0051] Specifically, the data preprocessing steps include: data cleaning and data standardization.

[0052] Data cleaning is an important part of data preprocessing, aiming to remove noise and outliers in the data. Noise data may come from experimental errors, interference during data collection, etc., while outliers may mislead subsequent analysis. Data cleaning methods include, but are not limited to, data filtering, data smoothing and other techniques. For example, set thresholds to identify and remove data points outside the normal range.

[0053] Data standardization is to normalize the data to meet the requirements of subsequent analysis. Multi-omics data from different sources may have different dimensions and distribution characteristics, so standardization is needed to eliminate these differences. Common standardization methods include Z-score standardization, Min-Max standardization, etc. Z-score standardization makes the data comparable by converting the data into a distribution with a mean of 0 and a standard deviation of 1; Min-Max standardization scales the data to the [0,1] interval for subsequent calculation and analysis.

[0054] Step S102, construct a cell-specific protein interaction network according to the preprocessed omics data.

[0055] Specifically, a cell-specific protein interaction network is constructed based on protein interaction data and preprocessed omics data. The protein interaction data can be obtained from public databases (such as STRING, BioGRID, etc.) or through experimental verification. Cell specificity refers to optimizing the protein interaction network according to a specific cell type or tissue source to make it more in line with the actual biological background. After the network construction is completed, proteins serve as nodes and the interactions between proteins serve as edges, forming a complex network structure.

[0056] On this basis, constructing a cell-specific protein interaction network according to the preprocessed omics data includes:

[0057] Obtain protein interactome data and obtain an initial cell-specific protein interaction network according to the protein interactome data;

[0058] Dynamically enhance the initial cell-specific protein interaction network according to the preprocessed omics data to obtain an enhanced cell-specific protein interaction network.

[0059] Specifically, protein interaction data integration is the basis of network construction. Since protein interaction data may come from multiple different databases, these data may vary in quality, coverage, and format. Therefore, integrate these data to construct a comprehensive and accurate protein interaction dataset. The integration method can include steps such as data format conversion, duplicate data removal, and data quality assessment. For example, by comparing protein identifiers in different databases, combine interaction data from different databases into a unified dataset.

[0060] Use the integrated protein interaction data to construct an initial cell-specific protein interaction network. Cell specificity refers to optimizing the network according to a specific cell type or tissue source to make it more in line with the actual biological background. For example, by combining cell type-specific gene expression data, screen out protein interactions that are active in a specific cell type, thereby constructing an initial cell-specific protein interaction network. After the initial cell-specific protein interaction network is constructed, proteins serve as nodes and the interactions between proteins serve as edges, forming a complex network structure, providing a basis for subsequent analysis.

[0061] Dynamically enhance the initial cell-specific protein interaction network based on the preprocessed omics data to obtain an enhanced cell-specific protein interaction network. Adjust the node attributes, adding protein expression levels, fold changes (log2FC), and modification levels. Adjust the edge weights. Through a weighting strategy, adjust the weights of the network edges according to factors such as protein expression levels and interaction reliability to enhance the biological significance of the network. For example: enhance interaction reliability. If both interacting partners are significantly differentially expressed under specific conditions, increase the weight (e.g., weight × 1.5). For functional synergy, based on GO / KEGG annotations, if the interacting proteins are involved in the same pathway, increase the weight.

[0062] Step S103, extract topological structure features from the cell-specific protein interaction network; wherein, the topological features include: node degree centrality, betweenness centrality, closeness centrality, and modularity.

[0063] Specifically, extract a series of network topological features from the constructed cell-specific protein interaction network, such as node degree, modularity, etc. Topological features refer to the distribution rules and mutual relationships of nodes and edges in the network, and these features can reflect the importance of nodes (proteins) in the network and the complexity of the network connection pattern. Apply complex network analysis techniques, such as network module division, key path identification, etc., to further reveal the topological characteristics of the network.

[0064] Common topological features include node degree centrality, betweenness centrality, closeness centrality, and modularity, etc. By calculating these features, key nodes and potential functional modules in the network can be identified.

[0065] Node degree centrality is one of the network topological features, used to measure the connectivity of each node in the network. Node degree refers to the number of direct connections between a node and other nodes. The higher the degree centrality, the closer the connection of the node with other nodes in the network, and it may play a key role in the signal transduction process. Calculating the degree centrality of each node can be achieved by a simple counting method, that is, counting the number of adjacent nodes of each node.

[0066] Betweenness centrality is one of the network topological features, used to measure the mediating role of each node in the network. Betweenness centrality refers to the frequency of a node appearing in all shortest paths in the network. The higher the betweenness centrality, the stronger the mediating role of the node in the network, and it may be a key node for signal transduction. Calculating betweenness centrality requires traversing the shortest paths between all node pairs in the network and counting the number of times each node appears in these paths.

[0067] Closeness centrality is one of the network topological features, which is used to measure the centrality of each node in the network. Closeness centrality refers to the reciprocal of the average shortest path length from a node to all other nodes in the network. The higher the closeness centrality, the closer the position of the node to the center of the network, and the faster it can communicate with other nodes. Calculating closeness centrality requires calculating the shortest path length from each node to all other nodes and taking the reciprocal of the average value.

[0068] Step S104, obtain the signal transduction pathway according to the topological structure feature and the network topological feature prediction model.

[0069] Specifically, apply the trained network topological feature prediction model to predict new or unknown signal transduction pathways.

[0070] On this basis, before obtaining the signal transduction pathway according to the topological structure feature and the network topological feature prediction model, it further includes:

[0071] Obtain a feature training sample set, construct a training set and a test set with the feature training sample set, and use the cross-validation method to train a machine learning model to obtain a network topological feature prediction model; wherein, the feature training sample set includes: sample cell-specific protein features and corresponding signal transduction pathways.

[0072] Specifically, obtain a feature training sample set and select the features that have the greatest impact on the model prediction performance from the obtained feature training sample set. Since the feature training sample set may contain redundant information or noise, it is necessary to screen out the most valuable features through feature selection methods. Common feature selection methods include statistics-based methods (such as correlation analysis), model-based methods (such as Lasso regression), and information theory-based methods (such as mutual information). Through feature selection, the training efficiency and prediction performance of the model can be improved.

[0073] Use the selected features to train a machine learning model. Divide the selected feature training samples into a training set and a validation set, train the model through the training set, and adjust the model parameters through the validation set to obtain the optimal model performance. Usually, 80% of the data is used as the training set and 20% of the data is used as the validation set to ensure the balanced distribution of the data set and avoid bias.

[0074] Evaluate the performance of the model through cross-validation. Cross-validation is a commonly used model evaluation method. By dividing the dataset into multiple subsets, taking one subset as the test set in turn, and the remaining subsets as the training set, training and testing the model repeatedly, and finally evaluating the performance of the model by synthesizing multiple test results. Common cross-validation methods include: k-Fold Cross-Validation: Divide the dataset into k subsets, and use one subset as the validation set in turn, and the rest as the training set. Leave-One-Out Cross-Validation (LOOCV): Use one sample as the validation set each time, and the rest as the training set. Through cross-validation, the generalization ability and stability of the model can be effectively evaluated.

[0075] Machine learning models include: Random Forest: Suitable for high-dimensional data and can handle non-linear relationships. Support Vector Machine (SVM): Suitable for small-sample data and can handle high-dimensional features. Protein Interaction Networks: A method based on Graph Neural Networks (GNN), suitable for protein interaction networks.

[0076] Use Grid Search or Random Search to optimize the model hyperparameters.

[0077] Use metrics such as accuracy, precision, recall, F1-score, and the area under the ROC curve to evaluate the model performance.

[0078] Models such as Random Forest can provide feature importance scores, interpret the model prediction results through SHAP values, and evaluate the contribution of each feature to the prediction to predict the importance of the signal transduction pathway.

[0079] Step S105, obtain the signal transduction pathway analysis result according to the signal transduction pathway.

[0080] Specifically, perform biological function annotation on the predicted pathway, and use databases such as GO (Gene Ontology) and KEGG (Kyoto Encyclopedia of Genes and Genomes) to analyze the biological processes and molecular functions involved in the pathway.

[0081] Combine experimental verification or other independent data sources to further confirm the reliability of the prediction results, and provide a scientific basis for subsequent disease mechanism research and drug target discovery.

[0082] On this basis, obtaining the signal transduction pathway analysis result according to the signal transduction pathway includes:

[0083] Perform a pathway prediction operation on the signal transduction pathway to obtain a pathway prediction result;

[0084] Perform a functional annotation enrichment analysis on the railway prediction result to obtain a signal transduction pathway analysis result.

[0085] On this basis, perform a pathway prediction operation on the signal transduction pathway, and the obtained pathway prediction result includes:

[0086] Score the signal transduction pathway to obtain a prediction score:

[0087]

[0088] Among them, S(p) represents the prediction score of pathway p, w i represents the weight of the i-th feature, f i (p) represents the value of pathway p on the i-th feature;

[0089] Sort the prediction scores, and use the pathways with prediction scores higher than the preset threshold as the prediction score result.

[0090] Specifically, score each signal transduction pathway, and select the pathways with prediction scores higher than a certain threshold as the prediction score result.

[0091] Perform functional annotation on the prediction score result to analyze its biological significance. Use the GO and KEGG databases to perform functional enrichment analysis on the predicted pathways, and analyze the potential role of the predicted pathways in the occurrence and development of diseases.

[0092] Use the trained machine learning model to predict the signal transduction pathway of the unknown protein interaction network. The prediction process is based on the model's analysis of the topological features of the protein nodes in the network to identify potential signal transduction pathways. The prediction result can be represented in the form of a path, that is, a set of interconnected protein nodes that may be functionally relevant in the signal transduction process.

[0093] Perform functional annotation on the predicted pathways to analyze their biological significance. Functional annotation can be achieved by combining the functional annotation information of proteins (such as gene ontology annotation, disease correlation annotation, etc.). For example, by querying the Gene Ontology (GO) database, understand the functional categories of the proteins in the predicted pathway; or by querying the disease correlation database, analyze the relationship between the predicted pathway and specific diseases. The purpose of functional annotation is to help researchers understand the biological functions of the predicted pathways in cells and their potential application values.

[0094] Pathway prediction scoring function:

[0095]

[0096] Where S(p) represents the prediction score of pathway p, and w i represents the weight of the i-th feature, f i (p) represents the value of path p on the i-th feature.

[0097] Functional annotation enrichment analysis formula:

[0098]

[0099] Among them, P represents the P value of enrichment analysis, K represents the size of the gene set of interest, N represents the total gene set size, n represents the size of the differentially expressed gene set, and k represents the number of differentially expressed genes in the gene set of interest.

[0100] This paper proposes an innovative protein network-based biological signal transduction pathway analysis method, which overcomes the limitations of traditional analytical methods in revealing new biological signal transduction mechanisms. It integrates multi-omics data (including but not limited to genomics, transcriptomics, proteomics and metabolomics data), and cleverly uses the theories and methods of bioinformatics and network science to innovatively construct a cell-specific protein interaction network analysis framework.

[0101] It should be noted that the method of the embodiment of the present invention can be performed by a single device, such as a computer or a server. The method of this embodiment can also be applied in a distributed scenario and completed by multiple devices cooperating with each other. In the case of such a distributed scenario, one of the multiple devices can only perform one or more steps in the method of the embodiment of the present invention, and the multiple devices will interact with each other to complete the described method.

[0102] It should be noted that some embodiments of the present invention are described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the above embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0103] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present invention also provides a biological signal transduction pathway analysis system based on a protein network.

[0104] refer to Figure 2 , the biological signal transduction pathway analysis system based on protein network comprises:

[0105] An omics data acquisition module 201, configured to acquire omics data, preprocess the omics data, and obtain preprocessed omics data;

[0106] An interaction acquisition module 202, configured to construct a cell-specific protein interaction network according to the preprocessed omics data;

[0107] A topological feature calculation module 203, configured to extract topological structure features from the cell-specific protein interaction network; wherein, the topological features include: node degree centrality, betweenness centrality, and closeness centrality;

[0108] A signal conduction pathway acquisition module 204, configured to obtain a signal conduction pathway according to the topological structure features and a network topological feature prediction model;

[0109] A pathway analysis acquisition module 205, configured to obtain a signal conduction pathway analysis result according to the signal conduction pathway.

[0110] For the convenience of description, when describing the above system, various modules are described separately according to their functions. Of course, when implementing the present invention, the functions of each module can be implemented in one or more software and / or hardware.

[0111] The system of the above embodiment is used to implement a corresponding method for analyzing a biological signal conduction pathway based on a protein network in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0112] Based on the same inventive concept, corresponding to the method in any of the above embodiments, the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the program, it implements a method for analyzing a biological signal conduction pathway based on a protein network in any of the above embodiments.

[0113] Figure 3 FIG. shows a more specific schematic diagram of the hardware structure of the electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. Among them, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other inside the device through the bus 1050.

[0114] The processor 1010 can be implemented in the form of a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0115] The memory 1020 can be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.

[0116] The input / output interface 1030 is used to connect to the input / output module to achieve information input and output. The input / output module can be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Among them, the input device can include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device can include a display, a speaker, a vibrator, an indicator light, etc.

[0117] The communication interface 1040 is used to connect to a communication module (not shown in the figure) to achieve communication interaction between this device and other devices. Among them, the communication module can achieve communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as a mobile network, WIFI, Bluetooth, etc.).

[0118] The bus 1050 includes a path for transmitting information between various components of the device (such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).

[0119] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in the specific implementation process, this device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solutions of the embodiments of this specification, and does not necessarily include all the components shown in the figure.

[0120] The electronic device of the above embodiment is used to implement the corresponding method for analyzing a biological signal transduction pathway based on a protein network in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated herein.

[0121] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present invention also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute a method for analyzing a biological signal transduction pathway based on a protein network as described in any of the above embodiments.

[0122] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.

[0123] The computer instructions stored in the storage medium of the above embodiment are used to cause the computer to execute a method for analyzing a biological signal transduction pathway based on a protein network as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be elaborated herein.

[0124] Those of ordinary skill in the art should understand that: the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the present invention (including the claims) is limited to these examples; under the concept of the present invention, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the embodiments of the present invention as described above, which are not provided in detail for the sake of brevity.

[0125] Any process or method description depicted in a flowchart or otherwise described herein can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a specific logical function or process. The scope of the preferred embodiments of the present invention includes additional implementations, where functions may be performed in a substantially simultaneous manner or in a reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.

[0126] In addition, for simplicity of illustration and discussion, and so as not to make the embodiments of the present invention difficult to understand, well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Further, the systems may be shown in block diagram form in order to avoid making the embodiments of the present invention difficult to understand, and this also takes into account the fact that details regarding the implementation of these block diagram systems are highly dependent on the platform on which the embodiments of the present invention are to be implemented (i.e., these details should be entirely within the understanding of those skilled in the art). In cases where specific details (such as circuits) are set forth to describe exemplary embodiments of the present invention, it will be apparent to those skilled in the art that the embodiments of the present invention may be practiced without these specific details or with variations of these specific details. Accordingly, these descriptions should be considered illustrative rather than restrictive.

[0127] Although the present invention has been described in connection with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art based on the foregoing description. For example, other memory architectures (such as dynamic RAM (DRAM)) may be used with the embodiments discussed.

[0128] Embodiments of the present invention are intended to cover all such alternatives, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the embodiments of the present invention shall be included within the protection scope of the present invention.

Claims

1. A biological signal transduction pathway analysis method based on protein network, characterized in that: include: Acquiring omics data, and preprocessing the omics data to obtain preprocessed omics data; constructing a cell-specific protein interaction network based on the preprocessed omics data; Extracting topological structural features from the cell-specific protein interaction network; wherein the topological features include: node degree centrality, betweenness centrality, closeness centrality and modularity; Obtaining a signal transduction pathway according to the topological structure characteristics and the network topological characteristics prediction model; A signal transduction pathway analysis result is obtained according to the signal transduction pathway.

2. A biological signal transduction pathway analysis method based on protein network according to claim 1, characterized in that: The acquiring of omics data and preprocessing of the omics data to obtain preprocessed omics data comprises: Performing data cleaning on the omics data to obtain cleaned omics data; The cleaned omics data are subjected to data standardization processing to obtain pre-processed omics data.

3. The method for analyzing biological signal transduction pathways based on protein networks according to claim 1, characterized in that: The constructing of a cell-specific protein interaction network according to the pre-processed omics data comprises: Acquiring protein interaction data, and obtaining an initial cell-specific protein interaction network based on the protein interaction data; The initial cell-specific protein interaction network is dynamically enhanced according to the pre-processed omics data to obtain a cell-specific protein interaction network.

4. The method for analyzing biological signal transduction pathways based on protein networks according to claim 1, characterized in that: Before obtaining the signal transduction pathway according to the topological structure characteristics and the network topological characteristics prediction model, the method further includes: A feature training sample set is obtained, a training set and a test set are constructed with the feature training sample set, a machine learning model is trained using a cross-validation method, and a network topology feature prediction model is obtained; wherein the feature training sample set includes: sample cell-specific protein features and corresponding signal transduction pathways.

5. A biological signal transduction pathway analysis method based on protein network according to claim 4, characterized in that: The machine learning model includes: Random forests, support vector machines, and protein networks.

6. The method for analyzing biological signal transduction pathways based on protein networks according to claim 1, characterized in that: The obtaining of the signal transduction pathway analysis result according to the signal transduction pathway comprises: performing a pathway prediction operation on the signal transduction pathway to obtain a pathway prediction result; Functional annotation enrichment analysis was performed on the railway prediction results to obtain signal transduction pathway analysis results.

7. A biological signal transduction pathway analysis method based on protein network according to claim 6, characterized in that: The performing a pathway prediction operation on the signal transduction pathway to obtain a pathway prediction result comprises: The signal transduction pathway is scored to obtain a prediction score: Where S(p) represents the prediction score of pathway p, and w i represents the weight of the i-th feature, f i (p) represents the value of path p on the i-th feature; The prediction scores are sorted, and pathways with prediction scores higher than a preset threshold are taken as prediction score results.

8. A biological signal transduction pathway analysis system based on protein network, characterized in that: A method for analyzing a biological signal transduction pathway based on a protein network according to any one of claims 1 to 8, comprising: An omics data acquisition module, used for acquiring omics data, and preprocessing the omics data to obtain preprocessed omics data; An interaction acquisition module, used for constructing a cell-specific protein interaction network based on the preprocessed omics data; A topological feature calculation module, used to extract topological structural features from the cell-specific protein interaction network; wherein the topological features include: node degree centrality, betweenness centrality and closeness centrality; A conduction pathway acquisition module, used to obtain a signal conduction pathway according to the topological structure characteristics and the network topological characteristics prediction model; The pathway analysis acquisition module is used to obtain a signal transduction pathway analysis result according to the signal transduction pathway.

9. A device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • PRAK signal path analysis method and system

    CN121350496A