Bioinformation analysis system, method, electronic device and medium
The topological analysis and optimization of biological networks through the biological information analysis system has been solved, and the problem of incomplete bioinformatic analysis in the existing technology has been achieved, and the detailed display of biological molecules' interactions and functions and the revelation of causal relationships are achieved.
Patent Information
- Application Number
- CN202411876625.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2044-12-19
AI Technical Summary
The existing technology cannot fully and accurately analyze biological information, the construction of biological networks depends on simple correlations, and the prior knowledge network does not match the analytical data, and biological networks are decentralized and redundant. Multi-omics data sets can only display part of the information, and biological database information is redundant and difficult to explain, so different biological networks need to be compared or aligned.
It provides a biological information analysis system, including a biological pathway acquisition module, a network topology analysis module, a network mining inference module and a network update optimization module. It uses biological information statistics to obtain biological pathway data sets, optimize the biological network through topology analysis and graph neural network, and combines basic information and potential information for visual display.
It realizes detailed, efficient, clear and specific display of the interactions and biological functions between biological molecules, and can integrate and optimize biological networks and reveal complex causal relationships and biological mechanisms.
Smart Images

Figure CN119763669B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the technical field of systems biology and network biology, and relates to a biological information analysis system, and in particular to a biological information analysis system, method, electronic equipment and medium. Background Art
[0002] Biomolecules, including proteins, nucleic acids, carbohydrates, lipids, and small inorganic molecules, are the fundamental units of living organisms. They participate in and control various biological activities within living systems and perform essential life functions. Biological pathways are ordered sequences of molecular interactions that accomplish specific biological functions, describing how biomolecules interact within cells to realize specific life processes. Biological networks are systems that describe the regulatory, communication, binding, and reaction relationships between biomolecules. These biological connections form complex network structures, including protein interaction networks, gene regulatory networks, and metabolic networks. They help reveal the holistic nature and complexity of living systems. First, the development of high-throughput technologies has enabled life science research to obtain large-scale biological datasets of diverse modalities at low cost, greatly expanding the availability and quantity of biomolecules. Second, biological networks are inherently decentralized and redundant, meaning that the failure of individual nodes does not paralyze the entire network, nor does any single node dominate communication across the entire network. Third, biomolecules interact with each other in a variety of ways, and biological phenotypes at all levels operate independently. These interactions form complex networks rich in information. This has also driven the understanding and interpretation of biological phenotypes to evolve from simple biomolecules to complex biological networks, thereby obtaining more comprehensive and systematic biological information. Bioinformatics has gradually shifted its focus from individual genes, proteins, and search algorithms to large-scale networks such as biomes, interactomes, genomes, and proteomes.
[0003] Biological networks are crucial for understanding biological processes beyond the analysis of individual biomolecules. However, they currently present several challenges. First, many biological networks are constructed solely based on correlations between organisms, but simple correlations cannot represent complex causal relationships. Second, prior knowledge networks often do not align with the analytical data. Third, biological networks are attribute-dependent networks, often with dependencies between nodes and edges. These networks often exhibit "rewiring" capabilities, meaning the emergence or disappearance of one biological relationship is often accompanied by the disappearance or emergence of others. Fourth, biological networks contain a variety of biological interaction relationships and nodes. These include both ubiquitous nodes and edges and those related to biological perturbation conditions or phenotypes, though many studies tend to focus on the latter. Fifth, thanks to the development of high-throughput technologies, multi-omics datasets are increasingly being used in life science research. However, an omics network in a given dimension often only reveals a subset of the overall biological network. Sixth, knowledge-driven approaches rely on biological databases constructed from prior knowledge. Some databases contain redundant information. For example, in the Gene Ontology, a single analysis can yield dozens, hundreds, or even thousands of significant entries, often with various inclusion relationships between them. The functions of some entries are too general and difficult to explain their biological mechanisms. Seventh, some life science research requires the comparison or alignment of different biological networks. Summary of the Invention
[0004] In view of the above-mentioned shortcomings of the prior art, the purpose of this application is to provide a biological information analysis system, method, electronic device and medium to solve the problem of being unable to comprehensively and accurately analyze biological information.
[0005] In a first aspect, the present application provides a bioinformatics analysis system, which includes: a biological pathway acquisition module, which is used to map or transform a biological molecule data set using a bioinformatics statistical method to obtain a corresponding biological pathway data set; a network topology analysis module, which is used to perform topological analysis on the biological network formed by the biological pathway data set to obtain an integrated biological network; a network mining and inference module, which is used to perform inference analysis on the integrated biological network to obtain basic information and potential information of the integrated biological network; a network update and optimization module, which is used to mine the integrated biological network using a graph neural network to optimize the integrated biological network by combining the basic information and the potential information; and a content display module, which is used to visually display the content obtained by the biological pathway acquisition module, the network topology analysis module, the network mining and inference module, and / or the network update and optimization module.
[0006] In this application, biological pathway analysis is performed on biomolecules, and topological analysis is performed on the biomolecule bionetwork to obtain an integrated bionetwork. A network mining and inference module is used to mine and infer the integrated bionetwork, and the network update and optimization module is used to update and optimize the acquired information within the integrated bionetwork. Finally, a visual display of various content is performed. This bioinformatics analysis system can utilize biomolecules and their biopathways in a detailed, efficient, clear, and specific manner, fully demonstrating the interactions and biological functions between biomolecules.
[0007] In an implementation of the first aspect, the network topology analysis module is further used to: perform topological analysis on the biological network formed by the biological pathway dataset to obtain the overall characteristics of the biological network; use a difference network to perform comparative analysis on the biological networks of any class and any perturbation condition to obtain the correlation relationship between the biological networks; and integrate the biological networks according to the correlation relationship to obtain the integrated biological network.
[0008] In an implementation manner of the first aspect, the overall features include average degree, average distance, graph density, clustering coefficient, modularity, power law distribution and / or natural connectivity.
[0009] In an implementation of the first aspect, the network topology analysis module is further used to: obtain multi-omics or multimodal biological network data of the same biological phenotype; and use the multi-omics or multimodal biological network data to infer the biological interaction relationship between unknown biological molecules.
[0010] In one implementation of the first aspect, the network mining inference module is further used to: use an algorithm to mine the target attributes of the core nodes or quantified nodes of the integrated biological network, use an algorithm to mine the modular structure of the integrated biological network to obtain the sub-networks in the integrated biological network; and / or predict or quantify the target attributes of the edges to mine the path information between the target nodes.
[0011] In an implementation of the first aspect, the network mining and inference module is further used to: predict whether there is a causal relationship between edges based on the integrated biological network, and use the causal relationship to perform multi-dimensional inference analysis on the integrated biological network.
[0012] In one implementation of the first aspect, the network update optimization module is further used to: use a graph neural network to perform comparative analysis on integrated biological networks under different disturbance conditions or different omics to obtain biological connections between biological networks; and use the biological connections to mine and update information of the integrated biological network.
[0013] In a second aspect, the present application provides a bioinformatics analysis method, which includes: using bioinformatics statistical methods to map or transform biological molecule data sets to obtain corresponding biological pathway data sets; performing topological analysis on the biological network composed of the biological pathway data sets to obtain an integrated biological network; performing inference analysis on the integrated biological network to obtain basic information and potential information of the integrated biological network; using graph neural networks to mine the integrated biological network to optimize the integrated biological network by combining the basic information and the potential information; and visually displaying the various contents obtained.
[0014] In a third aspect, the present application provides an electronic device, comprising: a memory for storing a computer program; and a processor for executing the computer program stored in the memory, so that the electronic device performs the bioinformation analysis method as described in the first aspect.
[0015] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the bioinformatics analysis method described in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1A Shown is a schematic diagram of an application scenario of the bioinformatics analysis method described in this application.
[0017] Figure 1B Shown are structural diagrams of the client-cloud interaction scenarios in these implementations.
[0018] Figure 2 Shown is a structural schematic diagram of the biological information analysis system described in an embodiment of the present application.
[0019] Figure 3 Shown is a schematic diagram of the biological pathway described in the examples of this application.
[0020] Figure 4 Displays the gene interaction scores of the string database output by the visualization module.
[0021] Figure 5 Shown is a schematic diagram of the OTU association network described in an embodiment of the present application.
[0022] Figure 6 Shown is a chord diagram of the correlation described in the embodiments of the present application.
[0023] Figure 7A Shown is a schematic diagram of the per-cancer correlation network described in the Examples of this application.
[0024] Figure 7BShown is a schematic diagram of the per-cancer causal network described in the Examples of this application.
[0025] Figure 7C Shown is a schematic diagram of the cancer correlation network described in the Examples of this application.
[0026] Figure 7D Shown is a schematic diagram of the cancer causal network described in the Examples of this application.
[0027] Figure 8 Shown is a schematic flow chart of the bioinformatics analysis method described in an embodiment of the present application.
[0028] Figure 9 Shown is a structural schematic diagram of an electronic device described in an embodiment of the present application.
[0029] Component number description
[0030] 1 Bioinformatics analysis device
[0031] 11 Storage Devices
[0032] 12 Local Processor
[0033] 13 Display Terminal
[0034] 2-end-cloud interactive system
[0035] 20 Terminal
[0036] 21 Cloud Servers
[0037] 100 Bioinformatics Analysis System
[0038] 110 Biological Pathway Acquisition Module
[0039] 120 Network Topology Analysis Module
[0040] 130 Network Mining Inference Module
[0041] 140 Network Update Optimization Module
[0042] 150 Content Display Module
[0043] 900 Electronic Equipment
[0044] 910 Memory
[0045] 920 processor
[0046] 930 Display
[0047] Steps S11 to S15 DETAILED DESCRIPTION
[0048] The following describes the embodiments of the present application through specific examples. Those skilled in the art can easily understand the other advantages and effects of the present application from the content disclosed in this specification. The present application can also be implemented or applied through other different specific embodiments. The details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that the following embodiments and features in the embodiments can be combined with each other unless they conflict.
[0049] It should be noted that in the embodiments of this application, words such as "optionally" or "for example" represent examples, illustrations, or descriptions. Any embodiment or design described in this application as "optionally" or "for example" should not be interpreted as being more preferred or advantageous than other embodiments or designs. Rather, the use of words such as "optionally" or "for example" is intended to present the relevant concepts in a concrete manner.
[0050] In the embodiments of the present application, "at least one" refers to one or more, and "more" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: the existence of A alone, the existence of A and B at the same time, and the existence of B alone, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can represent: a, b, c, ab, ac, bc or abc, where a, b, c can be single or multiple.
[0051] It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present application. Therefore, the illustrations only show components related to the present application and are not drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component can be changed at will, and the component layout type may also be more complicated.
[0052] Figure 1A The diagram shows an application scenario of the bioinformation analysis method described in this application. The bioinformation analysis device 1 can be used to implement the bioinformation analysis method provided in the embodiment of this application, but the application scenario of the bioinformation analysis method provided in the embodiment of this application is not limited to Figure 1A The biological information analysis device 1 shown in FIG. Figure 1A As shown, the biological information analysis apparatus 1 includes a storage device 11, a local processor 12, and a display terminal 13. The biological information analysis method provided in the embodiment of the present application can be applied to the local processor 12.
[0053] in, Figure 1A The local processor 12 in the embodiment can be a local processor or a local processor cluster composed of multiple local processors or a cloud computing center, etc., which is not limited here. Figure 1A Only one storage device 11, one local processor 12 and one display terminal 13 are shown, but it should be understood that Figure 1A The examples are only used to understand this solution, and the specific numbers of local processors 12 and display terminals 13 should be flexibly determined based on actual conditions.
[0054] In other implementations, the bioinformation analysis device 1 may not include the display terminal 13, but may instead include only a local processor 12 with a display function and a storage device 11. The bioinformation analysis method provided in the embodiments of this application can be applied to the local processor 12. The local processor 12 with a display function may include a tablet computer, a laptop, a PDA, a mobile phone, a personal computer, or a voice interaction device, without limitation.
[0055] In still other implementations, the bioinformatics analysis method described in this application can be applied to end-cloud interaction scenarios. Figure 1B Shown is a schematic diagram of the structure of the end-cloud interaction scenario in these implementation methods. Figure 1B As shown, the terminal-cloud interaction system 2 includes a terminal 20 and a cloud server 21. The terminal 20 and the cloud server 21 can communicate with each other, and the communication method is not limited to wired or wireless.
[0056] The terminal 20 may be mobile or fixed. For example, the terminal 20 may be a wireless terminal or a wired terminal. A wireless terminal may refer to a device with wireless transceiver capabilities that can be deployed indoors, outdoors, and in industrial workshops. The terminal 20 may be a mobile phone, a tablet computer, a laptop computer, and the like, which are not limited here. The cloud server 21 may include one or more servers, or one or more processing nodes, or one or more virtual machines running on the server. The cloud server 21 may also be referred to as a server cluster, a management platform, a data processing center, and the like, which are not limited in the embodiments of the present application.
[0057] The technical solutions in the embodiments of the present application will be described in detail below with reference to the accompanying drawings in the embodiments of the present application.
[0058] Figure 2 The structure diagram of the biological information analysis system described in the embodiment of the present application is shown as follows: Figure 2 As shown, the biological analysis system 100 includes a biological pathway acquisition module 110 , a network topology analysis module 120 , a network mining and inference module 130 , a network update and optimization module 140 , and a content display module 150 .
[0059] The biological pathway acquisition module 110 is used to map or transform the biomolecule dataset using a bioinformatics statistical method to obtain a corresponding biological pathway dataset.
[0060] Optionally, the bioinformatics statistical methods include ORA (Over-Representation Analysis), GSEA (Gene Set Enrichment Analysis), GSVA (Gene Set Variation Analysis) and other methods for modeling the relationship between biomolecules and biological pathways, but the present application is not limited thereto.
[0061] The network topology analysis module 120 is used to perform topology analysis on the biological network formed by the biological pathway dataset to obtain an integrated biological network.
[0062] Optionally, the network topology analysis module 120 analyzes the entire biological network and analyzes the overall topological characteristics of the network or performs comparative analysis between multiple networks. This includes performing network topology analysis on biomolecular networks and their derived biological pathway networks to explore the overall characteristics of the biological network. The overall characteristics may include indicators such as average degree, average distance, graph density, clustering coefficient, modularity, power-law distribution, and / or natural connectivity, which measure the network topology structure. This application is not limited to these indicators.
[0063] In some possible implementations, the network topology analysis module 120 is further used to perform a topological analysis on the biological network formed by the biological pathway data set to obtain the overall characteristics of the biological network; use a differential network to compare and analyze the biological networks of any type and any perturbation condition to obtain the correlation between the biological networks; integrate the biological networks according to the correlation to obtain the integrated biological network. Use a differential network to compare the same type of biological networks under different perturbation conditions. Further, explore the unique mechanisms of different biological phenotype networks, such as paying more attention to rheumatoid arthritis than osteoarthritis. Explore the common biological molecules or their biological effects of different biological phenotypes, such as the common basic biological molecular networks related to autoimmune diseases such as systemic lupus erythematosus and rheumatoid arthritis.
[0064] The network topology analysis module 120 is also used to obtain multi-omics or multimodal biological network data of the same biological phenotype; and to infer the biological interaction relationship between unknown biological molecules using the multi-omics or multimodal biological network data. Differential networks are used to integrate different types of biological networks under the same perturbation conditions or to infer homogeneous specific biological networks. Further, multi-omics or multimodal biological networks that integrate the same biological phenotype are explored, for example, molecular networks that represent biological interaction relationships between multi-omics node types such as genomes, transcriptomes, proteomes, and metabolomes. Topological analysis also includes inferring biological interaction relationships between unknown or unmeasured biological molecules using multi-omics or multimodal data, and inferring networks between cells using single-cell multi-omics data, for example, protein sequence similarity, functional relationships, and physical interactions to infer unknown edges of the PPI (Protein-Protein Interaction Network) network, such as predicting edges between two nodes with similar sequences in terms of their functional relationships. Topological analysis also includes comparing biological networks under different conditions, matching nodes or edges in different biological networks or comparing network similarities in order to discover commonalities, similarities or associations between biological networks. Furthermore, based on the idea of embedded representation, the biological network is mapped to a low-dimensional vector space, and the entities of the biological network are aligned in the low-dimensional space, for example, cell type alignment of single-cell multimodal data. Based on feature measurement methods, the features of the biological network are screened, and the network comparison is converted into a feature similarity measurement method by combining attribute information such as expression levels. For example, "drug repositioning" compares the expression spectrum of the drug perturbation biological network with the disease biological network, and screens the diseases that are most negatively correlated with drug perturbation as potential treatment targets for new uses of old drugs.
[0065] The network mining and inference module 130 is used to perform inference analysis on the integrated biological network to obtain basic information and potential information of the integrated biological network.
[0066] In some possible implementations, the network mining inference module 130 is further used to mine the target attributes of the core nodes or quantified nodes of the integrated biological network using an algorithm, and to mine the modular structure of the integrated biological network using an algorithm to obtain the subnetworks in the integrated biological network; and / or predict or quantify the target attributes of the edges to mine the path information between the target nodes. Specifically, the network mining inference module 130 can obtain the target attributes of the core nodes or quantified nodes of the integrated biological network, the path information between the subnetworks and / or the target nodes separately or in combination, and the present application is not limited to this. The target attributes of the core nodes or quantified nodes of the biological network are mined by relevant algorithms, and further, the centrality of the nodes is quantified and sorted, and the hub nodes are screened, such as biomarkers of diseases screened by biological networks, and node-related attributes are predicted, such as biological network modeling of a certain disease, and whether the gene represented by the node is predicted to be mutated or whether it is a driver mutation. It is also used to mine the modular structure of the biological network through relevant algorithms to reveal functionally related subnetworks in the biological network. Predict or quantify the attributes of the edges, such as quantifying the strength of protein interactions in the PPI network. Mining the path information between nodes of interest, such as the biological network of the disease, analyzing the path between upstream products and downstream products, and exploring the path differences between different subtypes of the disease to assist in explaining the molecular mechanism of disease occurrence.
[0067] In some other possible implementations, the network mining and inference module 130 is also used to predict whether there is a causal relationship between edges based on the integrated biological network, and to perform multi-dimensional inference analysis on the integrated biological network using the causal relationship. A causal network is inferred for a biological network or its sub-network using causal inference related methods. Preferably, the network includes but is not limited to related networks, literature mining networks, physical networks, and non-causal networks such as protein sequence binding prediction networks. Optionally, the functions include determining whether there is a correlation or causality between two nodes, combining prior knowledge with non-causal networks to construct a phenotype-related causal network, exploring the similarities and differences of causal networks under different disturbances, and combining multimodal networks to form a larger causal network under the same phenotype.
[0068] The network update optimization module 140 is used to mine the integrated biological network using a graph neural network, so as to optimize the integrated biological network by combining the basic information and the potential information, and to visualize the content obtained by each module.
[0069] In some possible implementations, the algorithms of the graph neural network include mainstream algorithms such as graph convolutional neural networks, graph autoencoders, graph generative networks, graph recurrent networks, and graph attention networks. Preferably, the functions include mining node information and / or edge information of biological networks and / or their changes under different perturbations, such as predicting essential cancer genes, predicting new protein interactions, biological connections between biological networks based on dual graphs, such as drug response prediction based on drug graphs and cell line genetic maps, reasoning and prediction of biological networks, such as inferring cell type-specific biological networks from single-cell multi-omics data, simulating and generating graph structured data that conform to the characteristics of biological systems, such as simulating the interaction relationships between social species, including food chains, competition, and cooperation.
[0070] The content display module 150 is used to visually display the content obtained by the biological pathway acquisition module, the network topology analysis module, the network mining and inference module, and / or the network update and optimization module.
[0071] The visualization method includes visualization of biological network analysis and mining results, including but not limited to static images such as network diagrams, dendrograms, chord diagrams, and word clouds, as well as interactive files such as HTML and corresponding tables. The content of the visualization is not limited to content obtained by the biological pathway acquisition module 110, the network topology analysis module 120, the network mining and inference module 130, and the network update and optimization module 140. The content of each module can be displayed separately or combined.
[0072] The network update optimization module 140 is also used to use graph neural networks to perform comparative analysis on integrated biological networks under different disturbance conditions or different omics to obtain biological connections between biological networks; and use the biological connections to mine and update information of the integrated biological network.
[0073] In other possible implementations, the bioinformatics analysis system is also extended to biological networks of non-biological molecular types and biological networks derived therefrom. Optionally, the biological networks include but are not limited to cell-cell networks based on single-cell omics data, microbial networks based on sequencing methods such as 16S, 18S, and metagenomics, ecological networks that study interactions between species, evolutionary networks that focus on evolutionary relationships between species, physiological system networks related to physiological effects within individuals, such as immune system networks, nervous system networks, and other cells, as well as biological networks at various levels of life above cells and below the ecosystem and other networks derived therefrom. It should be clearly stated that the method is extended to all biological networks that can perform network analysis.
[0074] In the embodiments of the present application, biological pathway analysis is performed on biomolecules, and topological analysis is performed on the biomolecule biological network to obtain an integrated biological network. A network mining and inference module 130 is used to mine and infer the integrated biological network, and the network update and optimization module 140 is used to update and optimize the acquired information within the integrated biological network. Finally, a visual display of various contents is performed. The bioinformation analysis system 100 can utilize biomolecules and their biological pathways in a detailed, efficient, clear, and specific manner, fully demonstrating the interactions and biological functions between biomolecules.
[0075] Example 1: Division of subnetwork functional modules of the biological network of the GO pathway and its application in further visualization and mining
[0076] We are currently conducting GO (Gene Ontology) ORA (Over-Representation Analysis) enrichment analysis on 2520 differentially expressed genes in a specific disease medication topic. We used the Python library Scipy to perform a Fisher one-tailed test to statistically analyze the mapped pathways and determine the significance of enrichment. After correcting the p-value, we still found over 1000 significant pathways. For pathways with complex or broad paths, we used the GO database's documentation of relationships between entries, i.e., prior knowledge, to construct a biological network of GO pathways. The GO pathway network revealed a high degree of clustering and short average distances between nodes. Therefore, we applied the Louvain community detection algorithm to the biological network and divided the pathways into 15 community modules. Figure 3 Shown is a schematic diagram of the biological pathway described in the examples of this application. Figure 3 As shown, the 15 modules have numerous and dense intra-module connections and fewer inter-module connections, indicating that each module contains a group of entries with closely related or similar functions. Because one module is closely related to mitochondria, we analyzed the GO pathways of differentially expressed genes following different drug concentration treatments. In light of the research context, we focused on mitophagy, focusing on mitophagy entries. The GO pathways enriched at high and low concentrations are significantly different, and the information displayed by the significantly enriched entries is complementary. When studying the mechanisms of mitophagy, we prioritize the functional benefits of drugs on autophagy rather than on mitochondrial structure, and visualize these effects.
[0077] Example 2: Obtaining hub genes of the PPI network
[0078] In a study of lung adenocarcinoma, 28 target genes were identified, and the interactions between these target genes needed to be understood. The network mining inference module used prior knowledge from the string database to obtain a PPI network. Figure 4Displays the gene interaction scores of the string database output by the visualization module. Figure 4 As shown, for this network, the algorithm's plug-in functionality was used to quantify the centrality of each gene. For example, the top five hub genes in the MCC (Matthews Correlation Coefficient) algorithm, GAPDH, TPI1, PGAM1, LDHA, and PKM, were selected for subsequent analysis, narrowing down the 28 target genes to investigate their functions. A table was used to visualize the quantified centrality of each gene. Specifically, the algorithm's plug-in functionality was rewritten using Python and R to create the Cytohubba plug-in for Cytoscape software. The rewritten script was integrated into the network mining and inference module, allowing for automated integration into the workflow. This eliminates the need for manual data organization, file import into the software for analysis, parameter adjustment, and export of results before proceeding to the next stage of analysis.
[0079] Example 3: Differential network analysis of 16S microbial networks
[0080] Based on the network diagram of 57 genera of 16S sequencing of a certain autoimmune disease and normal control, it was found that there were significant differences between the network diagrams of the disease and normal groups. For example, there were significantly fewer associations between the disease group and the normal group. The network topology analysis module used the R programming package to perform network comparison of microbiome data. The diffnet function was used to construct the differential association network. Fisher' s The z-test was used to identify differential associations, connecting nodes with associations that were significantly different between the two groups. Figure 5 The diagram shows the OTU association network described in the embodiment of the present application. Figure 5 As shown, ou There was a strong positive correlation between t7 and out31 in the normal group but no correlation in the disease group. ou There was a strong positive correlation between t29 and out31 in the disease group but no correlation in the normal group. ou There was a strong positive correlation between t19 and out48 in the normal group and a negative correlation in the disease group.
[0081] Example 4: Overall topological analysis of metabolite autocorrelation networks
[0082] Figure 6 Shown is a chord diagram of the correlation described in the embodiment of this application. Figure 6As shown, a chord diagram visually displays the correlations between the five major metabolite categories of an immunodeficiency disease and a normal control, with black indicating positive correlations and gray indicating negative correlations. As can be seen from the comparison, immunodeficiency has sparser connections than the normal control group, meaning that immunodeficiency has fewer correlations and a more pronounced negative correlation trend. The network topology analysis module 120 is used to calculate the average degree, average distance, graph density, clustering coefficient, modularity, power-law distribution, and natural connectivity of two networks using the R programming package, although this application is not limited thereto. The relevant data is shown in the table below.
[0083] Topology Normal control Immunodeficiency Average 24.03614458 16.37804878 Average distance 2.650363516 2.858518208 Graph density 0.145673604 0.100478827 Clustering coefficient 0.660452734 0.60084009 Modularity 0.21726459 0.284356047 Power law distribution_alpha 14.67653148 39.67437649 Power law distribution xmin 56 47 Natural connectivity 0.939759036 0.945121951
[0084] As can be seen from the above table: from the topological analysis indicators such as average degree, average distance, graph density, and clustering coefficient, it can be seen that the metabolite correlation network diagram of the normal control has a closer connection than that of the immunodeficiency.
[0085] Example 5: Causal Inference of Protein-Metabolism Networks
[0086] The network mining inference module 130 is used to predict whether there is a causal relationship between the edges based on the integrated biological network, and to perform multi-dimensional inference analysis on the integrated biological network using the causal relationship. Specifically, based on the metabolome and protein data of two groups of samples of patients with a certain type of cancer and adjacent tissues, preliminary studies have found that 6 metabolites and 4 proteins play a significant role in the occurrence and prognosis of cancer. The mechanism of action of these 10 biological molecules is mined, and the expression data of metabolites and proteins are combined to construct correlation networks on cancer and adjacent tissues respectively. However, correlation does not represent causality, so the network mining inference module 130 reuses the Bayesian network and adopts the HC (Hill-Climbing) algorithm to construct the structure of the Bayesian network, infer causal relationships, and construct a protein-metabolism causal network. Figure 7A Shown is a schematic diagram of the per-cancer correlation network described in the Examples of this application. Figure 7B Shown is a schematic diagram of the per-cancer causal network described in the Examples of this application. Figure 7C Shown is a schematic diagram of the cancer correlation network described in the Examples of this application. Figure 7D Shown is a schematic diagram of the cancer causal network described in the embodiment of this application. Figures 7A to 7D As shown in the figure, the color of the edge between nodes indicates the credibility of the correlation or causal relationship direction, the thickness of the edge indicates the significance of the correlation or the strength of the causal relationship, and the arrow indicates the causal relationship between molecules.
[0087] Figure 8 The flow chart of the bioinformatics analysis method described in the embodiment of the present application is shown as follows: Figure 8 As shown, the biological information analysis method includes the following steps S11 to S15.
[0088] Step S11 : Mapping or transforming the biomolecule dataset using a bioinformatics statistical method to obtain a corresponding biological pathway dataset.
[0089] Step S12: performing a topological analysis on the biological network formed by the biological pathway dataset to obtain an integrated biological network.
[0090] Step S13: performing inference analysis on the integrated biological network to obtain basic information and potential information of the integrated biological network.
[0091] Step S14: mining the integrated biological network using a graph neural network to optimize the integrated biological network by combining the basic information and the potential information.
[0092] Step S15: Visually display the various acquired contents.
[0093] In the examples of this application, biological pathway analysis is performed on biomolecules, and topological analysis is performed on the biomolecule biological network to obtain an integrated biological network. Mining and inference analysis are performed on the integrated biological network, and the acquired information is updated and displayed in the integrated biological network. This bioinformatics analysis method can comprehensively, efficiently, clearly and specifically utilize biomolecules and their biological pathways, fully demonstrating the interactions and biological functions between biomolecules.
[0094] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices or methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of modules / units is only a logical function division. There may be other division methods in actual implementation. For example, multiple modules or units can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules or units, which can be electrical, mechanical or other forms.
[0095] The modules / units described as separate components may or may not be physically separate, and the components displayed as modules / units may or may not be physical modules, that is, they may be located in one place or distributed across multiple network elements. Some or all of the modules / units may be selected according to actual needs to achieve the purpose of the embodiments of the present application. For example, the functional modules / units in the various embodiments of the present application may be integrated into a processing module, or each module / unit may exist physically separately, or two or more modules / units may be integrated into a single module / unit.
[0096] Those skilled in the art should further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the composition and steps of each example according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0097] An embodiment of the present application also provides an electronic device. Figure 9 The structure diagram of the electronic device 900 according to the embodiment of the present application is shown. Figure 9 As shown, in this embodiment, the electronic device 900 includes a memory 910 and a processor 920 .
[0098] The memory 910 is used to store computer programs; preferably, the memory 910 includes: ROM, RAM, magnetic disk, USB flash drive, memory card or optical disk, etc., various media that can store program codes.
[0099] Specifically, the memory 910 may include a computer system readable medium in the form of a volatile memory, such as a random access memory (RAM) and / or a cache memory. The electronic device 900 may further include other removable / non-removable, volatile / non-volatile computer system storage media. The memory 910 may include at least one program product having a set (e.g., at least one) program modules that are configured to perform the functions of the various embodiments of the present application. It will be understood that the memory 910 may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM) or a programmable read-only memory (PROM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM) and synchronous static random access memory (SSRAM). The memory described in the embodiments of the present invention is intended to include, but is not limited to, these and any other suitable categories of memory.
[0100] The processor 920 is connected to the memory 910 and is configured to execute the computer program stored in the memory 910 so that the electronic device 900 executes the bioinformation analysis method described in any embodiment of the present application.
[0101] Optionally, the processor 920 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, and discrete hardware components.
[0102] Optionally, the electronic device 900 in this embodiment may further include a display 930. The display 930 is communicatively connected to the memory 910 and the processor 920, and is used to display a graphical user interface (GUI) interactive interface related to the bioinformatics analysis method described in the embodiment of the present application.
[0103] The present application also provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the bioinformatics analysis method described in any embodiment of the present application is implemented.
[0104] As used in this specification, the terms "component," "module," "system," and the like are used to represent computer-related entities, hardware, firmware, a combination of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or a computer. By way of illustration, both an application running on a computing device and a computing device can be a component. One or more components can reside in a process and / or an execution thread, and a component can be located on a computer and / or distributed between two or more computers. In addition, these components can be executed from various computer-readable media having various data structures stored thereon. Components can communicate, for example, via local and / or remote processes based on signals having one or more data packets (e.g., data from two components interacting with another component on a local system, a distributed system, and / or a network, such as the Internet interacting with other systems via signals).
[0105] The descriptions of the processes or structures corresponding to the above figures have different emphases. For parts that are not described in detail in a certain process or structure, please refer to the relevant descriptions of other processes or structures.
[0106] The above embodiments are merely illustrative of the principles and effects of this application and are not intended to limit this application. Anyone skilled in the art may modify or alter the above embodiments without departing from the spirit and scope of this application. Therefore, all equivalent modifications or alterations made by one of ordinary skill in the art without departing from the spirit and technical concepts disclosed in this application shall be covered by the claims of this application.
Claims
1. A biological information analysis system, characterized in that: include: A biological pathway acquisition module is used to map or transform biological molecule datasets using bioinformatics statistical methods to obtain corresponding biological pathway datasets; A network topology analysis module is configured to perform topological analysis on the biological network formed by the biological pathway dataset to obtain an integrated biological network; the network topology analysis module is further configured to: perform topological analysis on the biological network formed by the biological pathway dataset to obtain overall characteristics of the biological network; use a differential network to compare and analyze biological networks of any class and any perturbation condition to obtain correlations between biological networks; integrate the biological networks based on the correlations to obtain the integrated biological network; and obtain multi-omics or multimodal biological network data of the same biological phenotype; and use the multi-omics or multimodal biological network data to infer biological interaction relationships between unknown biomolecules; a network mining and inference module, configured to perform inference analysis on the integrated biological network to obtain basic information and potential information of the integrated biological network; the network mining and inference module is further configured to: predict whether there is a causal relationship between edges based on the integrated biological network, and perform multi-dimensional inference analysis on the integrated biological network using the causal relationship; a network update and optimization module for mining the integrated biological network using a graph neural network to optimize the integrated biological network by combining the basic information and the potential information; the network update and optimization module is further configured to: use a graph neural network to perform comparative analysis on the integrated biological networks under different perturbation conditions or different omics to obtain biological connections between biological networks; and use the biological connections to mine and update information of the integrated biological network; A content display module is used to visually display the content obtained by the biological pathway acquisition module, the network topology analysis module, the network mining and inference module, and / or the network update and optimization module.
2. The biological information analysis system according to claim 1, characterized in that The overall features include average degree, average distance, graph density, clustering coefficient, modularity, power law distribution and / or natural connectivity.
3. The biological information analysis system according to claim 1, wherein The network mining inference module is also used to: Using an algorithm to mine the core nodes or target attributes of the quantified nodes of the integrated biological network, and using an algorithm to mine the modular structure of the integrated biological network to obtain a subnetwork in the integrated biological network; and / or Predict or quantify the target attributes of edges to mine path information between target nodes.
4. A bioinformatics analysis method, characterized in that: include: Use bioinformatics statistical methods to map or transform biological molecular datasets to obtain corresponding biological pathway datasets; Performing a topological analysis on the biological network formed by the biological pathway dataset to obtain an integrated biological network; obtaining the integrated biological network includes: performing a topological analysis on the biological network formed by the biological pathway dataset to obtain overall characteristics of the biological network; using a differential network to compare and analyze biological networks of any class and any perturbation condition to obtain correlation relationships between biological networks; integrating the biological networks based on the correlation relationships to obtain the integrated biological network; and obtaining multi-omics or multimodal biological network data of the same biological phenotype; and inferring biological interaction relationships between unknown biological molecules using the multi-omics or multimodal biological network data. Performing inference analysis on the integrated biological network to obtain basic information and potential information of the integrated biological network; obtaining the basic information and potential information of the integrated biological network includes: predicting whether there is a causal relationship between edges based on the integrated biological network, and performing multi-dimensional inference analysis on the integrated biological network using the causal relationship; Mining the integrated biological network using a graph neural network to optimize the integrated biological network by combining the basic information and the potential information; mining the integrated biological network using a graph neural network includes: performing comparative analysis on the integrated biological networks under different perturbation conditions or different omics using the graph neural network to obtain biological connections between biological networks; mining and updating information of the integrated biological network using the biological connections; Provide a visual display of the various acquired contents.
5. An electronic device, characterized in that: The electronic device comprises: memory for storing computer programs; A processor, configured to execute the computer program stored in the memory, so that the electronic device executes the biological information analysis method according to claim 4.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the biological information analysis method according to claim 4 is implemented.
Citation Information
Patent Citations
Biosynthesis path prediction method based on machine learning and user platform
CN117409872A
Real-time metering data processing platform
CN117725537A