Path database expansion method, system, storage medium and electronic device
Through the pathway database expansion method based on lipid level information, the lipid pathway database is expanded, which solves the problem of insufficient accuracy in lipid recognition and mapping in the prior art, and achieves a more comprehensive evaluation of the effect of lipid effects.
Patent Information
- Application Number
- CN202411488692.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-24
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2044-10-24
AI Technical Summary
The lack of methods for expanding the lipid pathway database in the prior art leads to insufficient accuracy of lipid recognition and mapping, and ignores the overall structure of the metabolic network.
The pathway database expansion method based on lipid level information is adopted to expand the lipid-path annotation relationship by obtaining lipid-path annotation relationship and lipid level information to ensure that all the daughter lipids of the parent lipid are annotated into the corresponding pathway.
Improve the accuracy of lipid recognition and mapping, and more comprehensively evaluate the role of lipids in biological systems by taking into account the topology of metabolic networks.
Smart Images

Figure CN119007827B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the technical field of lipidomics, and relates to a database expansion method, and in particular to a pathway database expansion method, system, storage medium and electronic device. Background Art
[0002] The core analysis of lipidomics is to link the detected lipids with their biological functions to gain a deeper understanding of the roles and effects of these lipids in biological systems. A common strategy is to use metabolic pathway databases such as KEGG (Kyoto Encyclopedia of Genes and Genomes) and Reactome, or in genome-scale metabolic networks (GSMN), to map the list of lipids of interest to metabolic datasets through various statistical algorithms. In this way, researchers can identify and explain the synergistic effects of lipids in biological contexts and reveal changes in lipid metabolism under different physiological and pathological conditions. However, the existing technology lacks methods to expand lipid pathway databases. Summary of the invention
[0003] The purpose of the present application is to provide a pathway database expansion method, system, storage medium and electronic device for expanding a lipid pathway database.
[0004] In a first aspect, the present application provides a pathway database expansion method based on lipid hierarchical information, the method comprising: obtaining a lipid-pathway annotation relationship of a lipid pathway database; obtaining lipid hierarchical information according to a lipid classification system; and expanding the lipid-pathway annotation relationship according to the lipid hierarchical information, wherein, for any parent lipid, if the parent lipid is annotated to a first pathway, all child lipids of the parent lipid are annotated to the first pathway.
[0005] In an implementation of the first aspect, for any second pathway in the lipid pathway database, the method further includes: obtaining a node importance factor of the second pathway corresponding to the detected lipid; obtaining a coverage adjustment factor and a topology adjustment factor of the second pathway corresponding to the detected lipid; and obtaining a pathway score of the second pathway based on the node importance factor, the coverage adjustment factor and the topology adjustment factor.
[0006] In an implementation of the first aspect, obtaining the node importance factor of the second pathway corresponding to the detected lipid includes: obtaining the importance adjustment factor of the node in the second pathway; for each node in the second pathway, obtaining the detected lipid matching it; obtaining the individual influence factor of each matching lipid; and obtaining the node importance factor based on the individual influence factor and the importance adjustment factor.
[0007] In an implementation of the first aspect, obtaining the coverage adjustment factor of the second pathway corresponding to the detected lipid includes: obtaining the quotient between the number of the detected lipids and the number of nodes in the second pathway as an initial coverage; and obtaining the coverage adjustment factor based on the initial coverage.
[0008] In an implementation of the first aspect, obtaining the topological adjustment factor of the second pathway corresponding to the detected lipid includes: identifying the linear path and branch nodes in the second pathway; and configuring the topological adjustment factor according to changes in the nodes in the linear path and / or changes in the branch nodes.
[0009] In an implementation of the first aspect, the method further includes: performing a permutation test to obtain a permutation score; calculating an original p-value based on the pathway score and the permutation score; and correcting the original p-value to obtain a corrected p-value.
[0010] In an implementation of the first aspect, the method further includes: performing enrichment analysis based on the expanded pathway database.
[0011] In a second aspect, an embodiment of the present application provides a pathway database expansion system based on lipid hierarchical information, the system comprising: an annotation relationship acquisition module for acquiring lipid-pathway annotation relationships of a lipid pathway database; a hierarchical information acquisition module for acquiring lipid hierarchical information according to a lipid classification system; an expansion module for expanding the lipid-pathway annotation relationship according to the lipid hierarchical information, wherein, for any parent lipid, if the parent lipid is annotated to a first pathway, all child lipids of the parent lipid are annotated to the first pathway.
[0012] In a third aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the methods described in the first aspect of the embodiment of the present application.
[0013] In a fourth aspect, an embodiment of the present application provides an electronic device, comprising: a memory storing a computer program; and a processor communicatively connected to the memory, for executing any of the methods described in the first aspect of the embodiment of the present application when the computer program is called.
[0014] As described above, the path database expansion method, system, storage medium and electronic device provided by the embodiments of the present application have the following beneficial effects:
[0015] The embodiment of the present application can expand the lipid-pathway annotation relationship according to the lipid hierarchy information. Among them, for any parent lipid, if the parent lipid is annotated to a certain pathway, all the child lipids of the parent lipid are annotated to the pathway. The expanded lipid pathway database can improve the accuracy of lipid identification and mapping. In addition, the topological structure of the metabolic network is taken into account during the expansion, which is conducive to a more comprehensive assessment of the role of lipids in biological systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 Shown is a flow chart of a path database expansion method provided in an embodiment of the present application.
[0017] Figure 2 Shown is a flow chart of obtaining the pathway score of the second pathway in an embodiment of the present application.
[0018] Figure 3 Shown is a flowchart for obtaining node importance factors in an embodiment of the present application.
[0019] Figure 4 Shown is a flow chart of obtaining a coverage adjustment factor in an embodiment of the present application.
[0020] Figure 5 Shown is a flow chart of obtaining a topology adjustment factor in an embodiment of the present application.
[0021] Figure 6 Shown is a flow chart of obtaining and correcting the p-value in an embodiment of the present application.
[0022] Figure 7 Shown is a flow chart of enrichment analysis in the examples of this application.
[0023] Figure 8 Shown is a schematic diagram of the structure of the pathway database expansion system provided in an embodiment of the present application.
[0024] Fig. 9 Shown is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application.
[0025] Component number description
[0026] 800 Channel Database Expansion System 810 Annotation relationship acquisition module 820 Hierarchical information acquisition module 830 Extension Module 900 Electronic devices 801 processor 902 Memory 9021 operating system 9022 app 903 Network Interface 904 System Bus 905 User Interface DETAILED DESCRIPTION
[0027] The following describes the embodiments of the present application through specific examples, and those skilled in the art can easily understand other advantages and effects of the present application from the contents disclosed in this specification. The present application can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict.
[0028] It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present application, and thus the drawings only show components related to the present application rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component may be changed at will, and the component layout may also be more complicated.
[0029] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of this application, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.
[0030] In this application, unless otherwise clearly specified and limited, the terms "connected" and "connection" should be understood in a broad sense, for example, it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, it can be the internal connection of two elements or the interaction relationship between two elements. For ordinary technicians in this field, the specific meanings of the above terms in this application can be understood according to specific circumstances.
[0031] One of the key steps in pathway analysis is to match lipid molecules detected in the laboratory with metabolites recorded in the pathway database. However, the mapping process of lipids is much more complicated than that of other types of metabolites. On the one hand, common nodes are usually used in the database to represent certain lipid categories, and it is difficult for high-precision experimental data to find exact matches in the database. On the other hand, common nodes in the database, such as phosphatidylcholine or triglycerides, may correspond to multiple specific lipid molecules detected in the laboratory. For example, when analyzing the PC (diacylphosphocholine) family, mass spectrometry detected PC (34:1) and PC (34:2), but if all these molecules correspond to a node PC in the pathway database, the use of an exact match will result in a match failure, which will prevent some pathways from being enriched, thus seriously affecting the downstream interpretation of experimental data.
[0032] In addition, related technologies focus on simple lipid-pathway mapping when performing lipid analysis, while ignoring the overall structure of the metabolic network. This simplified approach cannot capture the true role and relationship of lipids in complex biological systems. In addition, most enrichment analysis algorithms assume that metabolites are independent of each other, while the metabolic pathway is actually a complete network, an assumption that is clearly inconsistent with biological realities.
[0033] At least for the above problems, the embodiments of the present application provide a pathway database expansion method based on lipid hierarchy information. The pathway database expansion method can expand the lipid-pathway annotation relationship according to the lipid hierarchy information. Among them, for any parent lipid, if the parent lipid is annotated to a certain pathway, all the child lipids of the parent lipid are annotated to the pathway. The expanded lipid pathway database can improve the accuracy of lipid identification and mapping. In addition, the topological structure of the metabolic network is taken into account during the expansion, which is conducive to a more comprehensive assessment of the role of lipids in biological systems.
[0034] The technical solutions in the embodiments of the present application will be described in detail below in conjunction with the drawings in the embodiments of the present application.
[0035] Figure 1 The flowchart of the path database expansion method provided by the embodiment of the present application is shown. Figure 1 As shown, the pathway database expansion method includes the following steps S11 to S13.
[0036] S11, obtain lipid-pathway annotation relationships from lipid pathway databases. Lipid pathway databases are information resources for metabolic pathways, functions, and interactions, designed to help researchers understand the biological effects and mechanisms of lipids inside and outside cells. Examples of lipid pathway databases include KEGG, Reactome, LipidMap, SABIO-RK, MetaCyc, BioCyc, etc.
[0037] Metabolic pathways refer to a series of interrelated biochemical reactions and processes involving the synthesis, conversion, decomposition and regulation of lipids. Metabolic pathways play an important role in the normal function of cells and organisms, including the construction of cell membranes, energy storage, signal transmission and regulation of cell functions.
[0038] The lipid-pathway annotation relationship refers to the correspondence between lipid molecules and metabolic pathways. This annotation relationship can be used to understand the role of specific lipids in metabolic pathways and their impact on biological processes.
[0039] S12, obtaining lipid hierarchy information according to the lipid classification system. The lipid hierarchy information includes parent and subset relationships of lipids.
[0040] In some implementations, LIPIDMAP can be downloaded to extract information such as the main category, subcategory, specific molecular species, molecular structure, etc. of each lipid molecule to construct lipid level information.
[0041] S13, expanding the lipid-pathway annotation relationship according to the lipid hierarchy information. In the expansion, for any parent lipid, if the parent lipid is annotated to a certain pathway (hereinafter referred to as the first pathway), all the child lipids of the parent lipid are annotated to the first pathway. The child lipids of the parent lipid include direct child lipids and indirect child lipids.
[0042] For example, if lipid a is annotated to pathway b, the next-level sub-lipids of lipid a are a1 and a2, and the next-level sub-lipids of lipid a1 are a11 and a12. When the lipid detected in the laboratory (referred to as the detection lipid) corresponds to a11, the detection lipid cannot be mapped to the original lipid pathway database. After adopting the database expansion method provided in the embodiment of the present application, lipids a1, a2, a11 and a12 are all annotated to pathway b. At this time, the detection lipid can correspond to the expanded lipid pathway database, thereby improving the accuracy of lipid identification and mapping.
[0043] In some implementations, the pathway database expansion method may also include: for each detected lipid, locating its position in the ontology; expanding each detected lipid to all its parent lipids according to lipid hierarchy information, the parent lipids including direct parent lipids and indirect parent lipids; calculating the weights of the detected lipid and its parent lipids at each level, and the weights are used to evaluate the distance between the detected lipid and the parent lipid.
[0044] Exemplarily, the weight calculation formula can be W = 1-0.2d, where d is the hierarchical distance between a parent lipid and the detection lipid or its direct parent class. For example, for the detection lipid and its direct parent lipid, the weight W is 1; for the upper two parent lipids of the detection lipid, the weight W=0.8; for the upper three parent lipids of the detection lipid, the weight W=0.6, and so on.
[0045] See also Figure 2 In some implementations, for any pathway in the lipid pathway database (hereinafter referred to as the second pathway), the pathway database expansion method further includes the following steps S21 to S23.
[0046] S21, obtaining the node importance factor of the second pathway corresponding to the detected lipid.
[0047] S22, obtaining a coverage adjustment factor and a topology adjustment factor corresponding to the detected lipid in the second pathway.
[0048] S23, obtaining a path score of the second path according to the node importance factor, the coverage adjustment factor, and the topology adjustment factor.
[0049] See also Figure 3 In some implementations, obtaining the node importance factor of the second pathway corresponding to the detected lipid includes the following steps S31 to S34.
[0050] S31, obtaining the importance adjustment factor of the node in the second path.
[0051] S32, for each node in the second pathway, obtain the detection lipids matching the node. For a certain node, the detection lipids matching the node include the detection lipids directly matching the node and the detection lipids matching the parent class (including the direct parent class and the indirect parent class) of the node.
[0052] S33, obtaining the individual impact factor of each matching lipid. The matching lipid refers to the detection lipid obtained in step S32 that matches the node in the second pathway.
[0053] Exemplarily, for matching lipid i, its individual impact factor IF_i can be obtained by the following formula:
[0054] IF_i =log2FC_i*(-log10(p_value_i))*W_i;
[0055] Among them, log2FC_i is used to measure the degree of change in the expression level of matching lipid i under different experimental conditions, and p_value_i is the p-value of matching lipid i, which is used to evaluate the significance level between the observed experimental data and the statistical hypothesis. Both can be obtained by differential expression analysis of matching lipid i. W_i is the matching weight of matching lipid i. If matching lipid i directly matches a node, W_i=1, otherwise W_i=0.8.
[0056] S34, obtaining the node importance factor according to the individual impact factor and the importance adjustment factor. Specifically, for any node c in the second pathway, the node importance factor IF_node_c of the node c is: IF_node_c=(ΣIF_i) / N*importance adjustment factor, where N is the number of detection lipids matching the node c. The node importance factor of the second pathway corresponding to the detection lipid is the sum of the node importance factors of all nodes in the second pathway.
[0057] In some implementations, obtaining the importance adjustment factor of the node in the second path includes S311 to S312.
[0058] S311, obtain the normalized degree centrality of the node in the second pathway. The degree centrality of a node refers to the number of edges directly connected to the node, including in-degree centrality and out-degree centrality. Among them, the in-degree centrality is the number of edges pointing to the node, the out-degree centrality is the number of edges starting from the node, and the total degree centrality = in-degree centrality + out-degree centrality. In order to compare between metabolic networks of different sizes, the total degree centrality can be normalized to obtain the normalized degree centrality. Exemplarily, normalized degree centrality = total degree centrality / (N-1), where N is the total number of nodes in the metabolic network.
[0059] S312, mapping the normalized degree centrality to the target interval to obtain an importance adjustment factor.
[0060] For example, the target interval is, for example, [0.8, 1.2]. In the embodiment of the present application, the following formula can be used for mapping:
[0061] Importance adjustment factor = 0.8 + 0.4 * (S-S_min) / (S_max-S_min);
[0062] Where S is the normalized degree centrality, S_min and S_max are the minimum and maximum values of S among all nodes, respectively.
[0063] The importance adjustment factor can be obtained through the above steps S311 and S312. The importance adjustment factor of the least important node is 0.8, and its influence will be slightly reduced. The importance adjustment factor of the most important node is 1.2, and its influence will be slightly increased. The importance adjustment factor of most nodes is about 1, keeping their original influence unchanged.
[0064] See also Figure 4 In some implementations, obtaining the coverage adjustment factor of the second pathway corresponding to the detected lipid includes S41 and S42.
[0065] S41, obtaining the quotient between the number of detected lipids and the number of nodes in the second pathway as an initial coverage rate.
[0066] S42, obtaining a coverage adjustment factor according to the initial coverage. For example, the coverage adjustment factor may be the square root of the initial coverage, but the present application is not limited thereto.
[0067] See also Figure 5 In some implementations, obtaining a topological adjustment factor of the second pathway corresponding to the detected lipid includes S51 and S52.
[0068] S51, identifying linear paths and branch nodes in the second path.
[0069] S52. Configure a topology adjustment factor according to the changes of nodes in the linear path and / or the changes of branch nodes.
[0070] Exemplarily, an initial value can be configured for the topology adjustment factor, such as 1, and the initial value is adjusted according to the changes of nodes in the linear path and / or the changes of branch nodes to obtain the final topology adjustment factor.
[0071] Exemplarily, since the experimental data is unknown, the lipids of the nodes may increase or decrease compared to the reference group. For the linear path, if the changes of three or more consecutive nodes are the same, such as all up-regulated or all down-regulated, an additional reward is given, such as 1.5 times. For branch nodes, if the changes of the branch node and its upstream and downstream nodes are the same, an additional reward is given, such as 1.2 times.
[0072] In some implementation manners, obtaining the path score of the second path according to the node importance factor, the coverage adjustment factor, and the topology adjustment factor includes: obtaining the product of the node importance factor, the coverage adjustment factor, and the topology adjustment factor corresponding to the second path for detecting lipids as the path score of the second path. This path score can be used to determine whether the second path is an important path under the corresponding experimental conditions (such as a certain disease state). Based on this path score, researchers can identify which biological pathways are significantly important under specific experimental conditions, providing a direction for subsequent research. In addition, in disease research, the path score can help reveal the potential pathological mechanisms of the disease. By comparing the path scores of the disease group and the control group, researchers can discover which pathways play important roles in the occurrence and development of the disease. In drug development, by identifying the second path that is active under specific conditions, potential drug targets can be discovered, which helps to design small molecule drugs or biological agents that can intervene in these pathways.
[0073] Please refer to Figure 6 , in some implementation manners, the method for expanding the path database may further include the following steps S61 to S63.
[0074] S61. Conduct a permutation test to obtain a permutation score. For example, 1000 permutations can be conducted, and the expression value labels of the original lipids are randomly shuffled each time. For each permutation, the path score is recalculated, and these path scores constitute the permutation score.
[0075] S62. Calculate the original p-value according to the path score of the second path and the permutation score. For example, the proportion of the path score of the second path being greater than the permutation score can be calculated as the p-value.
[0076] S63, correcting the original p-value to obtain a corrected p-value. For example, the Benjamini-Hochberg method can be used to perform FDR correction to obtain the corrected p-value, but the present application is not limited thereto.
[0077] In some implementations, the pathway database expansion method may further include: generating a table including the name of the second pathway, the pathway score, the original p-value, and the corrected p-value.
[0078] The following is an example of how to obtain the pathway score. In this example, the experimental data include: PC (16:0 / 18:1), log2FC=1.5, p-value=0.001; PC (18:0 / 20:4), log2FC= -0.8, p-value=0.01; LPC (16:0), log2FC=0.5, p-value=0.05; PA (16:0 / 18:1), log2FC=1.2, p-value=0.008.
[0079] Taking the second pathway as the glycerophospholipid metabolic pathway as an example, the results of lipid-level expansion of the detected lipids are as follows: PC (16:0 / 18:1)-phosphatidylcholine (PC)-glycerophospholipid;
[0080] PC (18:0 / 20:4)-phosphatidylcholine (PC)-glycerophospholipids;
[0081] LPC (16:0) - lysophosphatidylcholine (LPC) -> phosphatidylcholine (PC);
[0082] PA (16:0 / 18:1) - Phosphatidic acid (PA) > glycerophospholipids.
[0083] Based on the expanded pathway database, PC (16:0 / 18:1), PC (18:0 / 20:4) and LPC (16:0) are the detection lipids matching the PC node, and PA (16:0 / 18:1) is the detection lipid matching the PA node.
[0084] For the original detected lipids and their direct parent classes, their weights were configured as 1.0; for higher-level parent classes, such as PC, their weights were configured as 0.8.
[0085] Assuming that there are 100 nodes in the glycerophospholipid metabolic pathway, the nodes matching the detected lipids are PC nodes and PA nodes. Obtaining the node importance factors corresponding to the detected lipids in the glycerophospholipid metabolic pathway includes:
[0086] For the PC node, the individual impact factors of the detected lipids matched with it are:
[0087] PC(16:0 / 18:1):IF_1=1.5*(-log10(0.001))*1=4.5;
[0088] PC(18:0 / 20:4):IF_2= - 0.8*(-log10(0.01))*1= - 1.6;
[0089] LPC(16:0):IF_3=0.5*(-log10(0.05))*0.8=0.5*1.3*0.8=0.52.
[0090] If the importance adjustment factor of the PC node is 1.1, the node importance factor of the PC node is: IF_PC = (4.5 + (-1.6 + 0.52) / 3 * 1.1 = 1.254.
[0091] For the PA node, the matching individual impact factor of the detected lipids is: PA (16:0 / 18:1): IF_4 = 1.2*(-log10(0.008))*1 = 2.5164.
[0092] If the importance adjustment factor of the PA node is 1.05, the node importance factor of the PA node is: IF_PA=2.5164*1.05=2.64222.
[0093] The node importance factor (or initial score) of the glycerophospholipid metabolic pathway relative to the detected lipids is: IF_PC+IF_PA=1.254+2.64222=3.89622.
[0094] Initial coverage = 4 / 100 = 0.04, coverage adjustment factor is 0.2. PC to PA forms a consistent change path, giving a 1.2-fold reward, topology adjustment factor = 1.2. The pathway score of the glycerophospholipid metabolism pathway is: 3.89622*0.2*1.2=0.9350928. After permutation test, the p-value is 0.003, and the p-value after FDR correction is 0.01.
[0095] In some implementations, the pathway database expansion method further includes: performing enrichment analysis based on the expanded pathway database.
[0096] For example, see Figure 7 , the enrichment analysis may include the following steps S71 to S73.
[0097] S71, obtaining a list of lipids of interest. Exemplarily, the process may include:
[0098] i. Lipidomics data were acquired using liquid chromatography-mass spectrometry (LC-MS / MS).
[0099] ii. Perform data normalization and / or missing value processing on the collected data. The data normalization method is, for example, total ion current intensity normalization or internal standard normalization. The missing value processing method is, for example, minimum value substitution or multiple interpolation.
[0100] iii. Matching using a database, such as LIPIDMAPS, for lipid identification.
[0101] iv. Using statistical methods, such as t-test or ANOVA, identify significantly changed lipids. Set screening criteria, such as p-value < 0.05 and fold change > 1.5 or 2-fold.
[0102] v. Summarize the lipids that meet statistical and biological significance into a lipid list.
[0103] S72, using the lipid list of interest to intersect with the compounds of a certain pathway in the expanded pathway database, find out the common compounds therein and count them. Exemplarily, the process may include:
[0104] i. Obtain metabolic pathway information based on the expanded pathway database. Parse the pathway file and extract the list of compounds in each pathway. Establish a mapping relationship from compound ID to pathway.
[0105] ii. Convert experimentally detected lipid IDs to standard IDs used by pathway databases.
[0106] iii. For each pathway, calculate the intersection of its compound set and the lipid list of interest. Record the size of the intersection (i.e., the number of common compounds) for each pathway.
[0107] iv. Define a background set, which can be all lipids detected in the experiment or all lipids in the database. Calculate the size of the intersection between the background set and each pathway.
[0108] S73, using statistical tests to evaluate whether the observed count value is higher than random, so as to determine whether the pathway is significantly enriched. This process, for example, includes:
[0109] i. Select a statistical model, for example, you can use the hypergeometric distribution model, Fisher's exact test, or the chi-square test.
[0110] ii. For each pathway, calculate the p-value based on the following four numbers: the number of lipids associated with that pathway in the lipid list of interest (n), the total size of the lipid list of interest (N), the number of lipids of that pathway in the background set (M), the total size of the background set (T): P(X ≥ n) = ∑(x = n to min(N, M))C(M, x)*C(TM, Nx) / C(T, N).
[0111] iii. Set a significance threshold. For example, you can choose a corrected p-value < 0.05 as the significance threshold.
[0112] Since the child lipids of the parent lipids in the expanded pathway database are all annotated to the pathway corresponding to the parent lipids, even if the exact matching method is adopted, the exact matching can be achieved in the embodiments of the present application. In addition, the metabolic network topology is fully considered in the embodiments of the present application, so that the role of lipids in biological systems can be more comprehensively evaluated.
[0113] The protection scope of the path database expansion method provided in the embodiment of the present application is not limited to the execution order of the steps listed in this embodiment. All solutions implemented by adding, reducing or replacing steps in the prior art based on the principles of the present application are included in the protection scope of the present application.
[0114] The embodiment of the present application also provides a pathway database expansion system, which can implement the pathway database expansion method provided in the embodiment of the present application. However, the implementation device of the pathway database expansion method provided in the embodiment of the present application includes but is not limited to the structure of the pathway database expansion system listed in the embodiment of the present application. All structural deformations and replacements of the prior art made according to the principles of the embodiment of the present application are included in the protection scope of the present application.
[0115] Figure 8 The schematic diagram of the structure of the path database expansion system provided by the embodiment of the present application is shown. Figure 8 As shown, the pathway database expansion system 800 includes an annotation relationship acquisition module 810, a hierarchical information acquisition module 82 and an expansion module 830. The annotation relationship acquisition module 810 is used to obtain the lipid-pathway annotation relationship of the lipid pathway database. The hierarchical information acquisition module 820 is used to obtain lipid hierarchical information according to the lipid classification system. The expansion module 830 is used to expand the lipid-pathway annotation relationship according to the lipid hierarchical information. Among them, for any parent lipid, if the parent lipid is annotated to the first pathway, all the child lipids of the parent lipid are annotated to the first pathway.
[0116] It should be noted that the annotation relationship acquisition module 810, the level information acquisition module 82 and the expansion module 830 in the pathway database expansion system 800 are Figure 1Steps S11 to S13 in the pathway database expansion method are in one-to-one correspondence and are not described in detail here.
[0117] The embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by the processor, the path database expansion method provided by the embodiment of the present application is implemented. A person of ordinary skill in the art can understand that all or part of the steps in the method for implementing the above embodiment can be completed by instructing the processor through a program, and the program can be stored in a computer-readable storage medium, and the storage medium is a non-transitory medium, such as a random access memory, a read-only memory, a flash memory, a hard disk, a solid-state hard disk, a magnetic tape, a floppy disk, an optical disc, and any combination thereof. The above storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more available media integrated. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a tape), an optical medium (for example, a digital video disc (DVD)), or a semiconductor medium (for example, a solid-state disk (SSD)), etc.
[0118] An embodiment of the present application also provides an electronic device. Fig. 9 is a schematic block diagram of an electronic device provided in an embodiment of the present application. Fig. 9 As shown, the electronic device 900 includes: at least one processor 901, a memory 902, at least one network interface 903 and a user interface 905. The various components in the electronic device 900 are coupled together through a bus system 904. It can be understood that the bus system 904 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 904 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, the bus system 904 is not used in the following examples. Fig. 9 In the specification, various buses are labeled as bus systems.
[0119] The user interface 905 may include a display, a keyboard, a mouse, a trackball, a click gun, keys, buttons, a touch pad or a touch screen.
[0120] It is understood that the memory 902 can be a volatile memory or a non-volatile memory, and can also include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM). The memory described in the embodiments of the present application is intended to include but is not limited to these and any other suitable categories of memory.
[0121] The memory 902 in the embodiment of the present application is used to store various categories of data to support the operation of the electronic device 900. Examples of these data include: any executable program for operating on the electronic device 900, such as an operating system 9021 and an application 9022; the operating system 9021 includes various system programs, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks. The application 9022 may include various applications, such as a media player (Media Player), a browser (Browser), etc., for implementing various application services. The path database expansion method provided in the embodiment of the present application may be included in the application 9022.
[0122] The method disclosed in the above embodiment of the present application can be applied to the processor 901, or implemented by the processor 901. The processor 901 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware in the processor 901 or instructions in the form of software. The above processor 901 may be a general-purpose processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The processor 901 can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor 901 may be a microprocessor or any conventional processor, etc. In combination with the steps of the accessory optimization method provided in the embodiment of the present application, it can be directly embodied as a hardware decoding processor to execute, or it can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium, which is located in a memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware.
[0123] In an exemplary embodiment, the electronic device 900 may be implemented by one or more application specific integrated circuits (ASIC), DSP, programmable logic device (PLD), complex programmable logic device (CPLD) to execute the aforementioned method.
[0124] The embodiment of the present application may also provide a computer program product, which includes one or more computer instructions. When the computer instructions are loaded and executed on a computing device, the process or function described in the embodiment of the present application is generated in whole or in part. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer or data center to another website, computer or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means.
[0125] When the computer program product is executed by a computer, the computer executes the method described in the above method embodiment. The computer program product may be a software installation package, and when the above method is required, the computer program product may be downloaded and executed on a computer.
[0126] The terms "component", "module", "system", etc. used in this specification are used to represent computer-related entities, hardware, firmware, a combination of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process, a processor, an object, an executable file, an execution thread, a program and / or a computer running on a processor. By way of illustration, both applications and computing devices running on a computing device can be components. One or more components may reside in a process and / or an execution thread, and a component may be located on a computer and / or distributed between two or more computers. In addition, these components may be executed from various computer-readable media having various data structures stored thereon. Components may, for example, communicate through local and / or remote processes according to signals having one or more data packets (e.g., data from two components interacting with another component between a local system, a distributed system and / or a network, such as the Internet interacting with other systems through signals).
[0127] Those of ordinary skill in the art will appreciate that the various illustrative logical blocks and steps described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0128] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0129] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0130] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0131] In the above embodiments, the functions of each functional unit can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions (programs). When the computer program instructions (programs) are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. Computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, computer instructions can be transmitted from one website site, computer, server or data center to another website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode.
[0132] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory, random access memory, disk or optical disk, etc., various media that can store program codes.
[0133] In summary, the embodiments of the present application provide a pathway database expansion method, system, storage medium and electronic device. The embodiments of the present application can expand the lipid-pathway annotation relationship according to the lipid level information. Among them, for any parent lipid, if the parent lipid is annotated to a certain pathway, all the child lipids of the parent lipid are annotated to the pathway. The expanded lipid pathway database can improve the accuracy of lipid identification and mapping. In addition, the topological structure of the metabolic network is taken into account during the expansion, which is conducive to a more comprehensive assessment of the role of lipids in biological systems. Therefore, the present application effectively overcomes the various defects in the prior art and has a higher industrial value.
[0134] The above embodiments are merely illustrative of the principles and effects of the present application and are not intended to limit the present application. Anyone familiar with the technology may modify or change the above embodiments without violating the spirit and scope of the present application. Therefore, all equivalent modifications or changes made by a person of ordinary skill in the art without departing from the spirit and technical ideas disclosed in the present application shall still be covered by the claims of the present application.
Claims
1. A pathway database expansion method based on lipid level information, characterized in that: include: Obtain lipid-pathway annotation relationships from the lipid pathway database; Obtain lipid hierarchy information based on the lipid classification system; The lipid-pathway annotation relationship is expanded according to the lipid level information, wherein, for any parent lipid, if the parent lipid is annotated to the first pathway, all the child lipids of the parent lipid are annotated to the first pathway; For any second pathway in the lipid pathway database, obtaining a node importance factor of the second pathway corresponding to the detected lipid; Obtaining a coverage adjustment factor and a topology adjustment factor of the second pathway corresponding to the detected lipid; Obtaining a path score of the second path according to the node importance factor, the coverage adjustment factor, and the topology adjustment factor; Wherein, obtaining the node importance factor of the second pathway corresponding to the detected lipid includes: obtaining the importance adjustment factor of the node in the second pathway; for each node in the second pathway, obtaining the detected lipid matched therewith; obtaining the individual influence factor of each matched lipid; obtaining the node importance factor according to the individual influence factor and the importance adjustment factor; Obtaining the coverage adjustment factor of the second pathway corresponding to the detected lipid comprises: obtaining a quotient between the number of the detected lipid and the number of nodes in the second pathway as an initial coverage; obtaining the coverage adjustment factor according to the initial coverage; Acquiring the topological adjustment factor of the second pathway corresponding to the detected lipid comprises: identifying a linear path and a branch node in the second pathway; configuring the topological adjustment factor according to a change in a node in the linear path and / or a change in a branch node; Obtaining the importance adjustment factor of the node in the second path includes: obtaining the normalized degree centrality of the node in the second path; mapping the normalized degree centrality to a target interval to obtain the importance adjustment factor; For matching lipid i, its individual impact factor IF_i can be obtained by the following formula: IF_i=log2FC_i*(-log10(p_value_i))*W_i; Among them, log2FC_i is used to measure the degree of change in the expression level of the matching lipid i under different experimental conditions, p_value_i is the p-value of the matching lipid i, which is used to evaluate the significance level between the observed experimental data and the statistical hypothesis, and W_i is the matching weight of the matching lipid i. If the matching lipid i directly matches a node, W_i=1, otherwise W_i=0.
8.
2. The method for expanding the channel database according to claim 1, characterized in that: Also includes: Perform a permutation test to obtain a permutation score; Calculating a raw p-value based on the pathway score and the permutation score; The raw p-values were corrected to obtain corrected p-values.
3. The method for expanding the channel database according to claim 1, characterized in that: Also includes: Enrichment analysis was performed based on the expanded pathway database.
4. A pathway database expansion system based on lipid level information, characterized in that: For implementing the pathway database expansion method according to any one of claims 1 to 3, the system comprises: An annotation relationship acquisition module, used to obtain lipid-pathway annotation relationships of a lipid pathway database; A hierarchical information acquisition module, used to obtain lipid hierarchical information according to a lipid classification system; An expansion module is used to expand the lipid-pathway annotation relationship according to the lipid hierarchy information, wherein, for any parent lipid, if the parent lipid is annotated to the first pathway, all the child lipids of the parent lipid are annotated to the first pathway.
5. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 3 is implemented.
6. An electronic device, characterized in that: The electronic device comprises: A memory storing a computer program; A processor is communicatively connected to the memory, and executes the method according to any one of claims 1 to 3 when calling the computer program.
Citation Information
Patent Citations
Protein function annotation method and system thereof
CN109712669A
Disease correlation evolution system and method based on lipidomics method
CN116313155A