Method and system for mining functional modules of non-control population-dependent microbial network analysis

By constructing and analyzing microbial symbiosis networks and identifying key bacteria and functional modules, the problem that the existing technology cannot effectively simulate microbial relationships and community changes is solved, and a deep understanding of microbial communities and the identification of key species are achieved.

CN117352063BActive Publication Date: 2025-06-17SHANDONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311398569.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-26
Publication Date
2025-06-17
Estimated Expiration
2043-10-26

AI Technical Summary

Technical Problem

The existing statistical analysis process for metagenomic data of microbiomes cannot effectively simulate the relationship between microorganisms and changes in microbial communities, and most network analysis methods require a control group, which cannot identify subnetworks and central bacteria in the overall microbial interaction network.

Method used

A microbial network analysis method that is dependent on non-control populations is used to construct a microbial symbiotic network through microbiomic data, and a topological structure difference analysis, community division algorithm and key node identification are used to mine functional modules and identify key bacteria.

Benefits of technology

The disclosure of the necessary microbial relationships in the microbial community construction process and stable maintenance was achieved, the impact of mutual relations on host health was inferred, key species and functional modules in the community were identified, and the dependence of the control group was avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117352063B_ABST
    Figure CN117352063B_ABST
Patent Text Reader

Abstract

The present invention relates to a method and system for mining functional modules of microbial network analysis independent of control populations, including: (1) data preparation: obtaining the relative abundances of bacteria and the grouping information of samples through microbiomics data; (2) constructing a microbial symbiotic network; (3) analyzing the microbial symbiotic network; (4) mining functional modules. The present invention combines network analysis with metagenomic sequencing technology, providing rich research means for discovering the microbial relationships necessary for the construction process of microbial communities and the stable maintenance of communities, and inferring the impacts of various interactions on host health, etc. Network analysis can reveal non-random species co-occurrence patterns in the microbial community, enabling better reproduction of direct interactions or niche-sharing characteristics between species, which is crucial for understanding the assembly mechanism of microbial communities, the functions of ecosystems, and identifying key species in the community.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and system for mining functional modules of microbial network analysis independent of control populations, and belongs to the technical field of microbial determination and inspection. Background Art

[0002] Human microbiota and environmental microbiota play crucial roles in maintaining human health and biogeochemical cycles, respectively. A microbial community is an aggregate of many microorganisms, and the vast majority of microorganisms (more than 99%) cannot be cultured by traditional culture methods. To fully reveal the roles of microorganisms in the microbial community, high-throughput sequencing technology is used to obtain the sequence information of the microbial community. Currently, a statistical analysis process based on microbial metagenomic data has been established to achieve rapid and effective processing of big data of the microbiome, and to analyze the species composition and functional composition of the microbial community as well as the differential bacteria between different groups.

[0003] In the natural environment, microorganisms do not exist in the form of isolated individuals, but form complex co-occurrence networks through direct or indirect interactions. The interactions between microorganisms include types such as mutualism, commensalism, parasitism, predation, amensalism, and competition, and the interactions will have three effects on the participants: positive, negative, and neutral. Similarly, the environment affects the composition of the community, so there is also a close interaction between microorganisms and the environment. Microorganisms are small in size, large in number, have strong metabolism, and reproduce rapidly, resulting in extremely high complexity in the composition and function of the microbial community. However, it is very difficult to directly explore these different types of interactions in the ecosystem. High-throughput sequencing data contains a large amount of population information and has great advantages in simulating the symbiotic relationship of microorganisms. However, the existing statistical analysis process of microbial metagenomic data cannot well simulate the relationships between microorganisms and the changes of the microbial community. Networks can well describe the relationships between microorganisms and the overall community. Combining network analysis with metagenomic sequencing technology provides rich research means for discovering the microbial relationships necessary for the construction process of the microbial community and the stable maintenance of the community, and inferring the impacts of various interactions on host health. Network analysis can reveal non-random species co-occurrence patterns in the microbial community, enabling better reproduction of the direct interactions or niche-sharing characteristics between species, and is crucial for understanding the assembly mechanism of the microbial community, the function of the ecosystem, and identifying key species in the community.

[0004] The understanding and comprehension of the qualitative and quantitative characteristics of complex networks are important and challenging topics in the network era. As an important feature in complex networks, modular structure (or community structure) is an important and ubiquitous structural feature. Accurately mining and analyzing the modular structure has theoretical and practical significance for understanding the evolution, structure, and dynamics of complex networks. As the organizational form of functional modules in biological complex networks, modular structure has important significance in the field of life sciences. Although many effective algorithms have been proposed to analyze functional modules, these methods have certain deficiencies both at the algorithm level and in terms of the interpretability of biological networks.

[0005] Previous studies mainly relied on statistical methods to help identify significantly different microorganisms by comparing the microbial abundances between different study cohorts and control groups. However, no microorganism exists independently, and each microorganism is closely interconnected with other microorganisms, forming a well-balanced ecological network. This network creates better living conditions through mutual symbiosis and cooperative metabolism, or interacts with each other to avoid the host's immune surveillance. Microbiome-wide association studies (MWAS) mainly use differential abundance analysis of each individual microorganism to identify characteristic microorganisms with significantly different distributions in disease groups. However, this method usually ignores the complex interactions between microorganisms, resulting in relatively low accuracy and being unable to truly reflect the microbial interrelationships within the ecological network.

[0006] To address this problem, a variety of network-based methods have been developed to explore the composition and interactions within the microecosystem. Such as NetShift, NetMoss, MetagenoNets, NetCoMi, etc., which construct, analyze, and compare microbial association networks from high-throughput sequencing data. However, most of these network analysis methods require a control group and are unable to identify subnetworks and central bacteria in the overall microbial interaction network.

[0007] Therefore, in the face of these problems, it is necessary to conduct targeted research and propose methods for mining functional modules. Summary of the Invention

[0008] Aiming at the deficiencies of the prior art, the present invention provides a method for mining functional modules of non-control-population-dependent microbial network analysis;

[0009] The present invention also provides a system for mining functional modules of non-control-population-dependent microbial network analysis.

[0010] The present invention discloses non-random species co-occurrence patterns in microbial communities, which can better reproduce the characteristics of direct interactions or niche sharing between species, thereby understanding the assembly mechanism of microbial communities, the structure and function of microecosystems, and identifying key communities in microecosystems and the driving bacteria in key communities.

[0011] The technical solution of the present invention is as follows:

[0012] A method for mining functional modules of microbial network analysis independent of control populations, comprising the following steps:

[0013] (1) Data preparation: Obtain the relative abundances of bacteria and the grouping information of samples through metagenomic data.

[0014] (2) Construct a microbial co-occurrence network (association network): A microbial co-occurrence network is a graph that includes two basic elements: nodes and edges. Abstract the microecosystem as a network, with each species of bacteria as a node in the microbial co-occurrence network and the interaction relationship between bacteria as an edge in the microbial co-occurrence network.

[0015] (3) Analyze the microbial co-occurrence network, including:

[0016] 3.1 Analyze the microbial co-occurrence network based on topological differences, including:

[0017] 3.1.1 Direct numerical comparison of global indicators of the microbial co-occurrence network, including describing the scale, cohesion, and connectivity of the microbial co-occurrence network.

[0018] 3.1.2 Compare the overall distributions of the characteristics of each point in the microbial co-occurrence network through the Kolmogorov-Smirnov test (KS test).

[0019] 3.2 Identify the key bacteria in the microbial co-occurrence network, including:

[0020] 3.2.1 Divide the subgroups (modules) of the microbial co-occurrence network, including: Divide the constructed microbial co-occurrence network through a community detection algorithm based on random walks into different sub-networks, that is, divide the microbial community into different subgroups.

[0021] 3.2.2 According to the topological properties of the microbial co-occurrence network, including cohesion, stability, and connectivity, including: Select the three sub-networks with the highest cohesion, stability, and connectivity, that is, the indicators describing the scale, cohesion, and connectivity of the microbial co-occurrence network, as modules, regarded as more important microbial subgroups; further analysis will be carried out later.

[0022] 3.2.3 Finding key nodes in the module, including:

[0023] Select the points with relatively high comprehensive index rankings of node degree, closeness centrality, betweenness centrality, and eigenvector centrality as key nodes. The key nodes in the module include:

[0024] ① Some nodes with the highest number of connecting edges in the module; ② Some nodes in the central position in the module, that is, the nodes with the highest centrality. Centrality includes closeness centrality, betweenness centrality, and eigenvector centrality; ③ Connecting nodes, module center points, and network center points. Specifically, all points in the module are classified into four categories according to the intra-module connectivity Zi and inter-module connectivity Pi of the points: peripheral nodes, connectors, module hubs, and network hubs;

[0025] (4) Mining functional modules: After dividing the network into modules, perform correlation analysis between the modules and physiological and metabolic indicators to mine the functions of important modules.

[0026] According to the preference of the present invention, in step (2), the interaction relationship between bacteria refers to the significant correlation relationship between bacteria, and the weight of the edge refers to the spearman correlation coefficient between two bacteria;

[0027] Based on the relative abundances of bacteria, calculate the spearman correlation coefficient between two bacteria, that is, the spearman rank correlation coefficient;

[0028] The calculation of the Spearman rank correlation coefficient is as follows:

[0029]

[0030] Among them, ρ refers to the Spearman rank correlation coefficient; d i refers to the difference in ranks of the corresponding variables, that is, the difference in the positions (ranks) of the paired variables after the two variables are sorted respectively; n refers to the number of observation objects;

[0031] Given the relative abundances of each bacterium in each sample within the group, that is, the relative abundance of one bacterium in a group of samples corresponds to a group of variables, calculate the spearman correlation coefficient between any two bacteria;

[0032] If the absolute value of the Spearman correlation coefficient between two bacteria is greater than or equal to 0.8 and the p1 value is less than or equal to 0.05, it is considered that there is a significant correlation between these two bacteria. An edge is connected between these two bacteria, and the weight is the Spearman correlation coefficient; otherwise, no edge is connected. By analogy, a microbial co-occurrence network is constructed. For the significance test result of the Spearman correlation coefficient, the p1 value is the probability of the result when the null hypothesis is true. The null hypothesis H0: the Spearman correlation coefficient between the two bacteria is 0, that is, the two bacteria are not related; the alternative hypothesis H1: the Spearman correlation coefficient between the two bacteria is not 0.

[0033] Preferably according to the present invention, in step 3.1.1, by comparing the global indicators of the microbial co-occurrence network, the differences in the interaction relationships between the sample microbiota of each group are evaluated, including:

[0034] The indicators describing the scale of the microbial co-occurrence network include distance and diameter. The distance refers to the shortest distance between two points in the microbial co-occurrence network, and the diameter refers to the longest distance between two points in the microbial co-occurrence network.

[0035] The indicators describing the cohesion of the microbial co-occurrence network include graph density and clustering coefficient. The graph density refers to the ratio of the number of actually existing edges to the number of potentially existing edges in the microbial co-occurrence network; the actually existing edges refer to all the edges in a microbial co-occurrence network; the potentially existing edges refer to the sum of the number of edges that would exist if there were an edge between any two points in the microbial co-occurrence network; the clustering coefficient refers to the relative frequency of connected triples closing to form triangles, and a connected triple is three points connected by two edges.

[0036] The indicators describing the connectivity of the microbial co-occurrence network include average degree and average weighted degree. The average degree refers to the degree of all points in the microbial co-occurrence network, and the degree is the average value of the number of edges associated with the node; the average weighted degree is the average value of the weighted degrees of all points in the microbial co-occurrence network.

[0037] Compare the values of the above global indicators of the microbial co-occurrence networks constructed by different groups. Calculate the indicators describing the scale of the microbial co-occurrence network, including distance and diameter. Calculate the indicators describing the cohesion of the network, including graph density and clustering coefficient. Calculate the indicators describing the connectivity of the network, including average degree and average weighted degree. Obtain the values of these indicators for the microbial co-occurrence networks constructed by different groups, and compare the properties of the networks based on the magnitudes of these indicator values.

[0038] Preferably according to the present invention, in step 3.1.2,

[0039] The point characteristics include: node degree, closeness centrality, betweenness centrality, eigenvector centrality.

[0040] The node degree refers to: in a microbial symbiotic network G = (V, E), the node degree of point v is the number of edges associated with point v; V refers to the set of all points in the microbial symbiotic network, and E refers to the set of all edges in the microbial symbiotic network;

[0041] The closeness centrality refers to: the reciprocal of the sum of the distances from a certain point in the microbial symbiotic network to all other points;

[0042] The betweenness centrality refers to: the number of shortest paths passing through a certain point between any two points other than that point divided by the sum of the number of shortest paths between any two points;

[0043] The eigenvector centrality refers to: obtained by performing eigenvalue decomposition on the adjacency matrix of the network;

[0044] Calculate the numerical values of the characteristics (node degree, closeness centrality, betweenness centrality, eigenvector centrality) of each point in the microbial symbiotic network, obtain the numerical distribution of the characteristics of each point, perform a KS test, obtain the significance p2 value of the test result, and set p2 <= 0.05. The two correlation networks in different parts constructed are significantly different.

[0045] According to the preference of the present invention, in step 3.2.1, the constructed microbial symbiotic network is divided into different sub-networks through a community partitioning algorithm of random walk, including:

[0046] The probability of walking from point i to point j is defined as distance similarity. If i and j are in the same subgroup, obviously the probability is higher. Therefore, a hierarchical clustering method is adopted to establish the subgroup structure.

[0047] According to the preference of the present invention, in step 3.2.3, the within-module connectivity Z i (Within-module connectivity, Z i ) and the among-module connectivity P i (Among-module connectivity, P i ) are calculated as follows (Guimerà and Nunes 2005):

[0048]

[0049]

[0050] For Z i , K i is the number of edges connecting node i to other nodes in module S i , K Si is the average value of the K values of all nodes in module S i , and the K value is the value of this node in module S iThe number of connections with other nodes in the KSi It is module S i The standard deviation of the K values ​​of all nodes in;

[0051] For P i , K is is the node i and module S i The number of edges between nodes, k i is the degree of node i, M represents the module, N M That represents all modules;

[0052] According to the formula, if the edges related to a node are evenly distributed in all modules, the P of the node i The value is close to 1; if all the edges related to a node are within the module to which it belongs, the P of the node is i The value is 0;

[0053] Node attributes are divided into four types according to the topological characteristics of the nodes, including:

[0054] Z i >2.5 and P i When <0.62, the node is judged as the module center (a node with high connectivity inside the module);

[0055] Z i <2.5 and P i When >0.62, the node is judged as a connection node (a node with high connectivity between two modules);

[0056] Z i >2.5 and P i When >0.62, the node is judged as a network center (a node with high connectivity in the entire network);

[0057] Z i <2.5 and P i When <0.62, the node is judged as a peripheral node (a node that does not have high connectivity both within the module and between modules);

[0058] Connection nodes, module centers and network centers are key nodes, i.e., key bacteria in the microbial symbiotic network (Deng et al. 2012).

[0059] Preferably, according to the present invention, in step (4), the functional modules are mined by using the WGCNA weighted gene co-expression network analysis method to find the bacterial flora related to the physiological and metabolic indicators.

[0060] A mining system for functional modules of non-control-population-dependent microbial network analysis, comprising a data preparation unit, a microbial symbiotic network construction unit, a microbial symbiotic network analysis unit and a functional module mining unit;

[0061] The data preparation unit is used to execute the step (1); the microbial symbiotic network construction unit is used to execute the step (2); the microbial symbiotic network analysis unit is used to execute the step (3); the functional module mining unit is used to execute the step (4).

[0062] The beneficial effects of the present invention are as follows:

[0063] 1. The existing statistical analysis processes of metagenomic data cannot well simulate the relationships between microorganisms and the changes of microbial communities. Networks can well describe the relationships between microorganisms and the overall community. Combining network analysis with metagenomic sequencing technology provides rich research means for discovering the microbial relationships necessary for the construction process of microbial communities and the stable maintenance of communities, and inferring the impacts of various interactions on host health. Network analysis can reveal non-random species co-occurrence patterns in the microbial community, enabling better reproduction of the direct interactions or niche-sharing characteristics between species, which is crucial for understanding the assembly mechanism of microbial communities, the functions of ecosystems, and identifying key species in the community.

[0064] 2. There are few existing microbial network construction methods for the characteristics of metagenomic sequencing data, and the research problems are relatively scattered, unable to form a complete network analysis process from network construction to network difference analysis, then to the discovery of key bacteria, and finally to the mining of functional modules. In view of the characteristics of metagenomic sequencing data, integrate and optimize algorithms to form a complete network analysis process. Realize the comparison of different microecosystem structures and functions simulated by networks. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 It is a schematic diagram of the framework of the mining method for functional modules of non-control-population-dependent microbial network analysis of the present invention;

[0066] Figure 2 It is a specific display schematic diagram of different sub-networks obtained after dividing the microbial correlation network by the random walk algorithm;

[0067] Figure 3 It is a specific display schematic diagram of the sub-network with the best topological properties selected according to the comprehensive topological properties of different sub-networks;

[0068] Figure 4 It is a schematic diagram of clustering based on sample abundance similarity;

[0069] Figure 5(a) is a schematic diagram of setting a soft threshold for scale-free features;

[0070] Figure 5(b) is a schematic diagram of setting a soft threshold according to the average connectivity;

[0071] Figure 6 It is a schematic diagram of hierarchical clustering according to the correlation coefficient of the sample distribution of different bacteria;

[0072] Figure 7 It is a correlation heatmap of the association between bacteria within the module and phenotypic data; Specific implementation manners

[0073] The present invention will be further limited below in conjunction with the accompanying drawings of the specification and embodiments, but not limited thereto.

[0074] Embodiment 1

[0075] A method for mining a mathematical framework and functional modules based on network analysis of microbiome data (metagenomic data), comprising the following steps:

[0076] (1) Data preparation: Obtain the relative abundance of bacteria and the grouping information of samples through microbiome data;

[0077] The relative abundance of a certain species: The relative content of a species in a sample, that is, the number of a certain species in a sample / the sum of the numbers of all species in the sample. The abundance information of the species is obtained from the data downloaded by metagenomic sequencing technology. The sequences of all sample sequencing are obtained, and through quality control, clustering, and annotation to the microbial database of these sequences, the absolute abundance of microorganisms (mainly bacteria here) in the sample is obtained. The absolute abundance of a species in a certain sample divided by the sum of the absolute abundances of all species in this sample is the relative abundance of this species in this sample. The grouping information of the samples includes, for example, a certain number of healthy group and disease group samples collected in the experimental design.

[0078] Finally, two tables are obtained for the data information, as shown in Table 1 and Table 2 (partial extraction):

[0079] Table 1

[0080]

[0081] Table 2

[0082] SampleID part B01 B B02 B B03 B B04 B B05 B

[0083] Table 1 is the relative abundance table of microorganisms in different samples. In the columns, B01, B02, B03, B04, B05 represent different samples; the rows represent different bacteria (species). Table 2 is the grouping table of samples. SampleID represents different sample names; part represents different groupings. Here, B (buccal mucosa site) is taken as an example.

[0084] (2) Construct a microbial co-occurrence network (association network): A microbial co-occurrence network is a graph that includes two basic elements: nodes and edges. The microecosystem is abstracted as a network, with each species of bacteria as a node in the microbial co-occurrence network and the interaction relationship between bacteria as an edge in the microbial co-occurrence network. This microbial network model is the microbial co-occurrence network. The interaction relationship is inferred through the covariation of the relative abundances of species, and the calculation method of the spearman correlation coefficient is used to infer the correlation between bacteria.

[0085] (3) Analyze the microbial co-occurrence network, including:

[0086] 3.1 Analyze the differences in the microbial co-occurrence network based on topological structure, including:

[0087] 3.1.1 Direct numerical comparison of the global indicators of the microbial co-occurrence network, including describing the scale, cohesion, and connectivity of the microbial co-occurrence network;

[0088] 3.1.2 Compare the overall distribution of the characteristics of each point in the microbial co-occurrence network through the Kolmogorov-Smirnov test (KS test);

[0089] The purpose is to characterize a certain characteristic of the entire microbial co-occurrence network by describing the overall distribution of point characteristics. Use the Kolmogorov-Smirnov test (KS test) to compare the differences in the distribution states of point characteristics, so as to find the differences in the interaction relationships of the microbial communities in different grouped samples. The KS test is a method used to determine whether there is a bias between the distribution of a set of observed values and the theoretical distribution, or to compare whether there are significant differences in the distributions of two sets of observed values.

[0090] 3.2 Find the key bacteria in the microbial co-occurrence network, including:

[0091] 3.2.1 Divide the subgroups (modules) of the microbial co-occurrence network, including: Through the community detection algorithm based on random walk, divide the constructed microbial co-occurrence network into different sub-networks, that is, divide the microbial community into different subgroups;

[0092] 3.2.2 Based on the topological properties of the microbial co-occurrence network, including cohesiveness, stability, and connectivity, the following steps are taken: Select the three sub-networks with the highest cohesiveness, stability, and connectivity, that is, the indicators describing the scale of the microbial co-occurrence network, the indicator describing the cohesiveness of the microbial co-occurrence network, and the indicator describing the connectivity of the microbial co-occurrence network, as modules, which are regarded as more important microbial subgroups; further analysis will be carried out subsequently.

[0093] Cohesiveness: Cohesiveness refers to the degree of close interaction between microorganisms in the network. In a network with high cohesiveness, the connections between microorganisms are closer, forming more tightly-knit communities or groups. This can help identify microorganisms with similar functions or interdependencies.

[0094] Robustness: The robustness of the network refers to the ability of the network to maintain its function when nodes or edges are randomly deleted or disturbed. In the microbial co-occurrence network, robustness is a key property because the interactions between microorganisms may be disturbed by external environmental factors, and the robustness of the network helps to maintain the persistence of symbiotic relationships.

[0095] Connectivity: Connectivity refers to the degree of connection between nodes in the network. High connectivity indicates that there are more interactions between microorganisms in the network, thus forming a more complex ecological network. This is very important for understanding information transfer and resource sharing between microorganisms.

[0096] Select the following topological property indicators of the network to comprehensively measure the cohesiveness, stability, and connectivity of the microbial co-occurrence network:

[0097] Average Degree: Describes the number of other nodes that each node is connected to on average, which helps to understand the connectivity of the network.

[0098] Average Path Length: Measures the average shortest distance between nodes in the network, which is very important for the stability and information transfer of the network.

[0099] Network Density: Represents the ratio of the actual number of edges to the possible number of edges, and is used to measure the connectivity and tightness of the network.

[0100] Betweenness Centrality: Measures the degree to which a node acts as a bridge in the network, which helps to identify key nodes for information transfer.

[0101] Clustering Coefficient: Measures the degree to which nodes in a network form tight clusters and helps understand the cohesion in the network.

[0102] Modularity: Used to detect subgraphs or modules in a network and helps understand the functional regions in the network.

[0103] Network Diameter: Describes the shortest distance between the farthest nodes in a network and is crucial for the information transmission speed.

[0104] Connectivity Metrics: Include the number of connected components, the size of the largest connected component, etc., and are used to measure the overall connectivity of the network.

[0105] 3.2.3 Finding key nodes in the module, including:

[0106] Select important nodes according to the topological properties of nodes in the microbial symbiotic network. These nodes play important roles in the scale, cohesion, and connectivity of the microbial symbiotic network, that is, the bacterial species that play important roles in the microbial community, namely key bacteria or driver bacteria;

[0107] Select the points with the top rankings in the comprehensive indicators of node degree, closeness centrality, betweenness centrality, and eigenvector centrality as key nodes. The key nodes in the module include:

[0108] ① Some nodes with the highest number of connecting edges in the module; Node degree refers to the number of edges associated with the node. One point or several points with the highest number of connecting edges in the module.

[0109] ②Those nodes that are in the central position within the module, i.e., the nodes with the highest centrality; centrality includes closeness centrality, betweenness centrality, and eigenvector centrality; for closeness centrality, the measurement idea is: if a node is "close" to many other nodes, then the node is in the central position of the network. It is defined as the reciprocal of the sum of the distances from a certain node to all other nodes (the concept of "distance" between nodes in a network graph is defined as the length of the shortest path between nodes, and if there is no path, it is positive infinity; so it is the reciprocal of the sum of the lengths of the shortest paths between a node and all other nodes in the graph). Betweenness centrality, the measurement attempts to generalize to what extent a certain node is "between" other node pairs. This centrality is based on the view that the "importance" of a node is related to its position in the network path. If these paths are regarded as the channels required for communication, then the nodes on multiple paths are the key links in the communication process. It is defined as the number of shortest paths passing through a certain node. Eigenvector centrality, the measurement idea is: if the centrality of a node's neighbors is higher, the centrality of the node itself is also higher. Bonacich (1972) defined the following form of centrality measurement based on previous work: vector c Ei =(c Ei (1),…,c Ei (N v )) T is the solution to the eigenvalue problem Ac Ei =α -1 c Ei where A is the adjacency matrix of the network graph G. Bonacich claims that the optimal value of α -1 is the largest eigenvalue of A, and c Ei is the corresponding eigenvector.

[0110] ③Connect the nodes, module center points, and network center points, specifically referring to: classifying all points in the module into four categories according to the within-module connectivity Zi and between-module connectivity Pi of the points: peripheral nodes (Peripherals), connector nodes (Connectors), module center points (Module hubs), and network center points (Network hubs);

[0111] (4) Mining functional modules: After dividing the network into modules, perform a correlation analysis between the modules and physiological and metabolic indicators to mine the functions of important modules.

[0112] Example 2

[0113] A mathematical framework for network analysis of microbiome data (metagenomic data) and a method for mining functional modules according to Embodiment 1, wherein:

[0114] In step (2), the interaction relationship between bacteria refers to the significant correlation relationship between bacteria, and the weight of the edge refers to the spearman correlation coefficient between two bacteria;

[0115] Based on the relative abundances of bacteria, calculate the spearman correlation coefficient between two bacteria, that is, the spearman rank correlation coefficient; the spearman rank correlation coefficient (Spearman's rank correlation) is a non-parametric statistic, whose value has nothing to do with the specific values of two sets of correlated variables, but only with the magnitude relationship between their values. The spearman rank correlation is calculated based on the differences between the ranks of each pair of two columns of paired ranks, so it is also called the "method of rank differences".

[0116] The calculation of the Spearman rank correlation coefficient is as follows:

[0117]

[0118] Among them, ρ refers to the Spearman rank correlation coefficient; d i refers to the difference in ranks of the corresponding variables, that is, the difference in the positions (ranks) of the paired variables after the two variables are sorted respectively; n refers to the number of observation objects;

[0119] Given the relative abundances of each bacterium in each sample within a group, that is, the relative abundance of a bacterium in a group of samples corresponds to a set of variables, calculate the spearman correlation coefficient between any two bacteria;

[0120] If the absolute value of the spearman correlation coefficient between two bacteria is greater than or equal to 0.8 and the p1 value is less than or equal to 0.05, then it is considered that these two bacteria have a significant correlation relationship, and a connection is made between these two bacteria, with the weight being the spearman correlation coefficient, otherwise no connection is made, and so on, to construct a microbial symbiotic network; for the significance test result of the spearman correlation coefficient, the p1 value is the probability of the result occurring when the null hypothesis is true. The null hypothesis H0: the spearman correlation coefficient between two bacteria is 0, that is, the two bacteria are not correlated; the alternative hypothesis H1: the spearman correlation coefficient between two bacteria is not 0.

[0121] In step 3.1.1, by comparing the global indicators of the microbial symbiotic network, evaluate the differences in the interaction relationships between the microbial communities of each group of samples, including:

[0122] The metrics describing the scale of the microbial symbiotic network include distance and diameter. Distance refers to the shortest distance between two points in the microbial symbiotic network, and diameter refers to the longest distance between two points in the microbial symbiotic network. The larger the scale of the microbial symbiotic network, the more bacteria with correlative relationships there are.

[0123] The metrics describing the cohesion of the microbial symbiotic network include graph density and clustering coefficient. Graph density is the ratio of the number of actually existing edges to the number of potentially existing edges in the microbial symbiotic network. The actually existing edges refer to all the edges in a microbial symbiotic network. The potentially existing edges refer to the sum of the number of edges that would exist if there were an edge between any two points in the microbial symbiotic network. For a network graph with n nodes, the number of potentially existing edges is n(n - 1) / 2. For example, for a network graph with 5 nodes, the number of potentially existing edges is 10. The clustering coefficient refers to the relative frequency of connected triples closing to form triangles. A connected triple is three points connected by two edges. The stronger the cohesion of the microbial symbiotic network, the greater the average correlation between each bacterium and other bacteria in this group.

[0124] The metrics describing the connectivity of the microbial symbiotic network include average degree and average weighted degree. Average degree refers to the degree of all points in the microbial symbiotic network. Degree is the average value of the number of edges associated with a node. Average weighted degree is the average value of the weighted degrees of all points in the microbial symbiotic network. For each edge e in the network graph, a real number W(e) is corresponding, and W(e) is called the "weight" of e, and such a network graph is called a weighted network. For a weighted network, the sum of the weights of the edges associated with a certain node becomes the weighted degree. The higher the connectivity of the microbial symbiotic network, the stronger the correlation between the sample bacteria in this group. A change in the abundance of one bacterium may cause changes in the abundances of many other bacteria.

[0125] Compare the numerical values of the above global metrics of the microbial symbiotic networks constructed by different groups. Calculate the metrics describing the scale of the microbial symbiotic network, including distance and diameter. Calculate the metrics describing the cohesion of the network, including graph density and clustering coefficient. Calculate the metrics describing the connectivity of the network, including average degree and average weighted degree, to obtain the values of these metrics for the microbial symbiotic networks constructed by different groups, and compare the properties of the networks according to the magnitudes of these metric values.

[0126] Its calculation formulas are shown in Table 3:

[0127] Global and node topological property metrics of an undirected weighted network. Suppose there is a simple, finite, undirected, weighted graph, which contains V vertices and E edges.

[0128] Table 3

[0129]

[0130]

[0131]

[0132]

[0133]

[0134] In step 3.1.2,

[0135] The point features include: node degree, closeness centrality, betweenness centrality, and eigenvector centrality;

[0136] The node degree refers to: in a microbial co-occurrence network G = (V, E), the node degree of point v refers to the number of edges associated with point v; V refers to the set of all points in the microbial co-occurrence network, and E refers to the set of all edges in the microbial co-occurrence network; in different types of microbial co-occurrence networks, the node degree (frequency) distribution has different characteristics, that is, it follows different distribution patterns. The network degree distribution pattern reflects the special structural characteristics of the network. For example, the degree of the microbial co-occurrence network generally conforms to the power-law distribution, most species have a small number of connections, and very few species have a very large number of connections, indicating that the construction method of the microbial community is a non-random process.

[0137] The closeness centrality refers to: the reciprocal of the sum of the distances from a certain point in the microbial co-occurrence network to all other points; if a node is "close" to many other nodes, then the node is in the center position of the network. The distribution of this feature of the node reflects the cohesion and connectivity characteristics of the network.

[0138] The betweenness centrality refers to: the number of shortest paths passing through a certain point between any two points other than this point divided by the sum of the number of shortest paths between any two points; it measures to what extent a certain node is "between" other node pairs. The distribution of this node feature reflects the connectivity degree of the network.

[0139] The eigenvector centrality refers to: a measure method used to measure the importance of a node in the network. Its calculation is obtained by performing eigenvalue decomposition on the adjacency matrix of the network; where the eigenvalue (λ) represents the eigenvalue of the adjacency matrix, and the eigenvector (v) is the eigenvector of the adjacency matrix. The calculation of the eigenvector centrality depends on the components of the eigenvector corresponding to the largest eigenvalue of the adjacency matrix; if a point has a higher closeness centrality, the eigenvector centrality of the point itself is also higher; the distribution of this feature of the node reflects the connectivity characteristics of the network.

[0140] Calculate the numerical values of the characteristics of each point (node degree, closeness centrality, betweenness centrality, eigenvector centrality) in the microbial co-occurrence network, obtain the numerical distribution of the characteristics of each point, conduct a KS test, obtain the significance p2 value of the test result, and set p2 <= 0.05. The two correlation networks in different parts constructed are significantly different.

[0141] In step 3.2.1, through the community detection algorithm of random walk, the constructed microbial co-occurrence network is divided into different sub-networks, including:

[0142] Suppose a walker randomly walks in a microbial co-occurrence network (similar to Brownian motion). The walker is likely to be trapped in a densely connected area, and this densely connected area is the community discovered by the walker. To measure the behavior of the walker, the probability of walking from point i to point j is defined as the distance similarity. If i and j are in the same subgroup, obviously the probability is higher. Therefore, a hierarchical clustering method is adopted to establish the subgroup structure.

[0143] In step 3.2.3, the within-module connectivity Z i (Within-module connectivity, Z i ) and the among-module connectivity P i (Among-module connectivity, P i ) are calculated as follows (Guimerà and Nunes 2005):

[0144]

[0145]

[0146] For Z i , K i is the number of edges connecting node i to other nodes in module S i , K Si is the average value of the K values of all nodes in module S i , the K value is the number of connections of this node to other nodes in module S i , σ KSi is the standard deviation of the K values of all nodes in module S i ;

[0147] For P i , K is is the number of edges connecting node i to the nodes in module S i , k i is the degree of node i, M represents the module, and N M represents all the modules;

[0148] According to the formula, if the edges related to a certain node are evenly distributed among all modules, the P value of this node is close to 1; if all the edges related to a certain node are within the module to which it belongs, the P value of this node is 0. i When Z > 2.5 and P < 0.62, it is determined that this node is the module center point (a node with high connectivity inside the module); i

[0149] According to the topological characteristics of the nodes, the node attributes are divided into 4 types, including:

[0150] Z i > 2.5 and P i < 0.62, it is determined that this node is the module center point (a node with high connectivity inside the module);

[0151] Z i < 2.5 and P i > 0.62, it is determined that this node is the connecting node (a node with high connectivity between two modules);

[0152] Z i > 2.5 and P i > 0.62, it is determined that this node is the network center point (a node with high connectivity in the whole network);

[0153] Z i < 2.5 and P i < 0.62, it is determined that this node is the peripheral node (a node with low connectivity both inside and between modules);

[0154] The connecting node, the module center point and the network center point are key nodes, that is, the key bacteria in the microbial co-occurrence network (Deng et al. 2012). Attributed to these nodes being in the hub position, it is not difficult to see that the absence of these key nodes may cause the decomposition of modules and the network, and they play an important role in maintaining the stability of the network structure. The microbial co-occurrence network can generally be divided into multiple modules. Modules are highly connected regions in the network. Modules may reflect the heterogeneity of the habitat, the aggregation of phylogenetically closely related species, the overlap of ecological niches and the co-evolution of species, and are considered to be phylogenetically, evolutionarily or functionally independent units (Olesen et al. 2007). In the above steps, the microbial co-occurrence network has been constructed, the network modules have been divided, and the most important modules in the network have been found. The key nodes identified in this step in the ecological network modules often represent the key species that may play an important role in maintaining the stability of the microbial community structure.

[0155] In step (4), the weighted gene co-expression network analysis method (WGCNA) is used to mine functional modules and find the bacterial communities related to physiological and metabolic indicators.

[0156] Example 3

[0157] A method for mining a mathematical framework and functional modules based on network analysis of metagenomic data, according to Embodiment 2, is characterized in that:

[0158] Samples are taken from five parts: gingival crevicular fluid (GCF), dental plaque (P), buccal mucosa (B), saliva (S), and tongue coating (T) respectively, and high-throughput sequencing technology is used to analyze the microbial flora in different parts of the oral cavity. It includes:

[0159] 1. Construct a multi-site microbial symbiotic network (association network) of the oral cavity:

[0160] Based on the relative abundance information of each sample species obtained by metagenomic sequencing technology, a spearman correlation network between bacteria in five parts: gingival crevicular fluid (GCF), dental plaque (P), buccal mucosa (B), saliva (S), and tongue coating (T) of the oral cavity is constructed. Points represent bacteria, edges represent that the two bacteria connected by the edge have a significant correlation relationship, and the weight of the edge represents the spearman correlation coefficient between the two bacteria. Based on the relative abundance of bacteria, we calculate the spearman correlation between bacteria in each group of samples. If the absolute value of the correlation coefficient between two bacteria is greater than or equal to 0.8 and the p-value is less than or equal to 0.05, it is considered that the two bacteria have a significant correlation relationship, and an edge is connected between the two bacteria, with the weight being the correlation coefficient, otherwise no edge is connected. And so on, a correlation network between bacteria is constructed.

[0161] 2.1 Differential analysis of the network based on topological structure

[0162] 2.1.1 Direct numerical comparison of network global indicators

[0163] Calculate important topological structure features of the network: the number of nodes, the number of edges, average degree, average weighted degree, node and edge connectivity, average path length, network diameter, graph density, clustering coefficient, betweenness centrality, degree centrality, modularity;

[0164] 2.1.2 Kolmogorov-Smirnov test to compare the overall distribution of each node feature

[0165] A method called Kolmogorov-Smirnov test (abbreviated as KS test) is used to evaluate whether there are significant differences in the overall distribution patterns of node attributes (including indicators such as node degree, betweenness centrality, and closeness centrality) between the bacterial correlation networks in different parts of the oral cavity, thereby indicating that there are different microbial communities in different parts of the oral cavity;

[0166] 2.2 Finding the key bacteria in the network

[0167] The overall graph is as Figure 1 shown;Figure 1 Among them, the left figure is the different sub-networks obtained after the random walk algorithm divides the microbial correlation network; the middle figure is the sub-network with the best topological properties selected according to the comprehensive topological properties of different sub-networks; the following figure shows the nodes with the best graph properties selected according to the comprehensive topological properties of different nodes in the selected sub-network and the nodes related to this node.

[0168] 2.2.1 Network subgroup (module) division (select the random walk subgroup division algorithm)

[0169] The rule of random walk is to assume that a rover randomly walks in a network (similar to Brownian motion). The rover is likely to be trapped in the densely connected area, and this dense area is the community discovered by the rover. To measure the behavior of the rover, the probability of walking from point i to point j is defined as the distance similarity. If i and j are in the same subgroup, obviously the probability is higher. Therefore, the hierarchical clustering method can be used to establish the subgroup structure. Taking the dorsal tongue as an example, Figure 2 It is a specific display schematic diagram of the different sub-networks obtained after the random walk algorithm divides the microbial correlation network;

[0170] 2.2.2 Select modules according to the topological properties (cohesion, stability, connectivity) of the network

[0171] Sub-network topological properties

[0172] nodes_num, # Number of nodes

[0173] edges_num, # Number of edges

[0174] average_degree, # Average degree

[0175] nodes_connectivity, # Vertex connectivity

[0176] edges_connectivity, # Edges connectivity

[0177] average_path_length, # Average path length

[0178] graph_diameter, # Network diameter

[0179] graph_density, # Graph density

[0180] clustering_coefficient, # Clustering coefficient

[0181] betweenness_centralization, # Betweenness centralization

[0182] degree_centralization, # Degree centralization

[0183] modularity # Modularity

[0184] The properties of the sub - networks are shown in Table 1;

[0185] Select modules with the number of nodes greater than or equal to 10 and ranking top three in other network topology properties for key research.

[0186] Figure 3 It is a specific display schematic diagram of the sub - network with the best topological properties selected according to the comprehensive topological properties of different sub - networks; it is the sub - network with the best comprehensive performance selected.

[0187] 2.2.3 Finding key nodes in the module

[0188] The topological properties of the sub - networks are shown in Table 2;

[0189] Select the bacteria with high rankings in the comprehensive indexes of node degree, closeness centrality, betweenness centrality, and eigenvector centrality as key bacteria for key research.

[0190] 3 Mining of functional modules

[0191] Based on the relative abundance information of species in each sample obtained by metagenomic sequencing technology, select the WGCNA weighted gene co - expression network analysis method to mine functional modules and find bacteria related to the site

[0192] 3.1 Data screening

[0193] Screen genes with the top 75% of the median absolute deviation, with at least MAD greater than 0.01; N(otus)=266, N(samples)=189

[0194] 3.2 Checking for outlier samples

[0195] Figure 4Schematic diagram for clustering based on sample abundance similarity; check for outlier samples and then remove them to ensure reasonable subsequent analysis; where "height" is the clustering step size;

[0196] 3.3 The principle for screening the soft threshold is to make the constructed network more conform to the characteristics of a scale-free network. (Most connections are concentrated in a few centers);

[0197] Screening criteria: R-square = 0.85, Power = 6 (optimal beta value), available.

[0198] For an undirected network, when the power is less than 15 or for a directed network when the power is less than 30, there is no power value that can make the R^2 of the scale-free network graph structure reach 0.8, and the average connectivity is relatively high, such as above 100. This may be because some samples are too different from other samples. Figure 5(a) is a schematic diagram for setting the soft threshold for scale-free characteristics; the principle for screening the soft threshold is to make the constructed network more conform to the characteristics of a scale-free network (most connections are concentrated in a few centers). The screening criteria are: R-square = 0.85, Power = 6 (optimal beta value), available. Figure 5(b) is a schematic diagram for setting the soft threshold according to the average connectivity; screening criteria: for an undirected network, when the power is less than 15 or for a directed network when the power is less than 30, there is no power value that can make the R^2 of the scale-free network graph structure reach 0.8, and the average connectivity is relatively high, such as above 100. This may be because some samples are too different from other samples.

[0199] 3.4 Network construction (constructing the co-expression matrix in one step)

[0200] The general idea is: calculate the adjacency between genes, calculate the similarity between genes based on the adjacency, then deduce the dissimilarity coefficient between genes, and obtain the hierarchical clustering tree of genes accordingly. Then, according to the standard of the hybrid dynamic shear tree, set the minimum number of genes in each gene module to 30.

[0201] After determining the gene modules according to the dynamic shear method, analyze again, calculate the eigenvector values of each module in turn, and then perform clustering analysis on the modules, and merge the modules with closer distances into new modules. The clustering results are as Figure 6 shown. Figure 6 Schematic diagram for hierarchical clustering according to the correlation coefficient of sample distribution of different bacteria; clustering into different modules, and different branches of the number of clusters represent different modules. Different module grayscales represent different modules. Among them, "height" represents the clustering step size, and the value range is (0, 1);

[0202] 3.5 Association between bacteria within the module and phenotypic data (select the module most relevant to the trait and find the most important bacteria in the module)

[0203] Figure 7 It is a correlation heatmap of the association between bacteria in the module and phenotypic data (selecting the module most relevant to the phenotype); different gray levels on the vertical axis represent different modules obtained by clustering, and the horizontal axis represents different groups (taking different oral parts as an example). Different gray levels of the squares represent different correlation coefficients between the module and the phenotype.

[0204] 3.6 Mining key genes from the module

[0205] Mining key genes from the module;

[0206] 3.6.1 Connectivity

[0207] Connectivity is similar to the concept of the degree of a node in a network. In a weighted co-expression network, since each edge represents the magnitude of the correlation between two genes and corresponds to a numerical value, the connectivity of a gene in the co-expression network is defined as the sum of the numerical values of all the edges connected to this gene.

[0208] In addition, according to whether the connected genes are in the same module as this gene, the edges can be divided into two categories. Those within the same module as this gene are defined as within, and those in different modules are defined as out.

[0209] The hub gene is the gene with the largest connectivity under this module. Note that only the edges within this module are considered at this time.

[0210] 3.6.2 Module membership

[0211] Abbreviated as MM, the MM value can be obtained by performing a correlation analysis on the expression level of this gene and the first principal component of the module, that is, the module eigengene. Therefore, the MM value is essentially a correlation coefficient.

[0212] 3.6.3 Gene significance

[0213] Abbreviated as GS, the expression level of this gene is correlated with the corresponding phenotypic value. The final value of the correlation coefficient is GS, and GS reflects the correlation between the gene expression level and the phenotypic data.

[0214] The correlation between the selected brown module and T, screening criteria: MM >= 0.8, and GS >= 0.2. The screening results are shown in Tables 3 and 4.

[0215] Table 4 shows the topological properties of the sub-network.

[0216] Table 4

[0217]

[0218] Table 5 shows the topological properties of each node in the sub-network.

[0219] Table 5

[0220]

[0221] Table 6 shows three indicators of OTUs included in the module with the greatest correlation with the dorsal tongue.

[0222] Table 6

[0223]

[0224]

[0225] Table 7 shows the bacterial classification information corresponding to the key OTUs.

[0226] Table 7

[0227]

[0228] Example 4

[0229] A mining system for functional modules based on network analysis of microbiome data, which executes a mathematical framework and a mining method for functional modules based on network analysis of microbiome data (metagenomic data) described in any one of Examples 1-3, including a data preparation unit, a microbial co-occurrence network construction unit, a microbial co-occurrence network analysis unit, and a functional module mining unit;

[0230] The data preparation unit is used to execute step (1) described in Example 1; the microbial co-occurrence network construction unit is used to execute step (2) described in Example 1; the microbial co-occurrence network analysis unit is used to execute step (3) described in Example 1; and the functional module mining unit is used to execute step (4) described in Example 1.

Claims

1. A method for mining functional modules of non-control-population-dependent microbial network analysis, characterized in that, It includes the following steps: (1) Data preparation: Obtain the relative abundances of bacteria and the grouping information of samples from microbiome data; (2) Construct a microbial co-occurrence network: A microbial co-occurrence network is a graph that includes two basic elements: nodes and edges; Abstract the microecosystem as a network, take each species of bacteria as a node in the microbial co-occurrence network, and take the interaction relationship between bacteria as an edge in the microbial co-occurrence network; (3) Analyze the microbial co-occurrence network, including: 3.1 Analyze the microbial co-occurrence network based on topological differences, including: 3.1.1 Direct numerical comparison of global indicators of the microbial co-occurrence network, including describing the scale, cohesion, and connectivity of the microbial co-occurrence network; 3.1.2 Compare the overall distributions of the characteristics of each node in the microbial co-occurrence network through the Kolmogorov-Smirnov test; 3.2 Identify key bacteria in the microbial co-occurrence network, including: 3.2.1 Divide subgroups of the microbial co-occurrence network, including: Divide the constructed microbial co-occurrence network through a community detection algorithm based on random walk, and divide it into different sub-networks, that is, divide the microbial community into different subgroups; 3.2.2 According to the topological properties of the microbial co-occurrence network, including cohesion, stability, and connectivity, including: Select the three sub-networks with the highest cohesion, stability, and connectivity, that is, the indicators describing the scale of the microbial co-occurrence network, the indicator describing the cohesion of the microbial co-occurrence network, and the indicator describing the connectivity of the microbial co-occurrence network as modules, regarded as more important microbial subgroups; Further analysis will be carried out later; 3.2.3 Identify key nodes in the module, including: Select the points with the top rankings in the comprehensive indicators of node degree, closeness centrality, betweenness centrality, and eigenvector centrality as key nodes. The key nodes in the module include: ① Some nodes with the highest number of connecting edges in the module; ② Some nodes in the central position in the module, that is, the nodes with the highest centrality. Centrality includes closeness centrality, betweenness centrality, and eigenvector centrality; ③ Connecting nodes, module central points, and network central points. Specifically, it means: Classify all points in the module into four categories according to the within-module connectivity Zi and between-module connectivity Pi of the points: peripheral nodes, connecting nodes, module central points, and network central points; (4) Mine functional modules: After dividing the network into modules, perform correlation analysis between the modules and physiological and metabolic indicators to mine the functions of important modules; In step (2), the interaction relationship between bacteria refers to the significant correlation relationship between bacteria, and the weight of the edge refers to the spearman correlation coefficient between two bacteria; Based on the relative abundances of bacteria, calculate the spearman correlation coefficient between two bacteria, that is, the spearman rank correlation coefficient; The calculation of the Spearman rank correlation coefficient is as follows: where ρ refers to the Spearman rank correlation coefficient; d i refers to the difference in ranks of the corresponding variables, that is, the difference in the positions of the paired variables after the two variables are sorted respectively; n refers to the number of observation objects; Given the relative abundances of each bacterium in each sample within a group, that is, the relative abundances of a bacterium in the samples within a group correspond to a set of variables, calculate the spearman correlation coefficient between any two bacteria; If the absolute value of the Spearman correlation coefficient between two bacteria is greater than or equal to 0.8 and the p1 value is less than or equal to 0.05, it is considered that there is a significant correlation between these two bacteria. An edge is connected between these two bacteria, and the weight is the Spearman correlation coefficient; otherwise, no edge is connected. And so on, a microbial co-occurrence network is constructed. For the significance test result of the Spearman correlation coefficient, the p1 value is the probability of the result occurring when the null hypothesis is true. The null hypothesis H0: the Spearman correlation coefficient between the two bacteria is 0, that is, the two bacteria are not related; the alternative hypothesis H1: the Spearman correlation coefficient between the two bacteria is not 0. In step 3.2.3, the within-module connectivity Zi and the between-module connectivity Pi are calculated as follows: For Zi, Ki is the number of edges connecting node i to other nodes in module Si, and K S i is the average value of the K values of all nodes in module Si, where the K value is the number of connections of the node to other nodes in module Si, and σ KS i is the standard deviation of the K values of all nodes in module Si; For Pi, Kis is the number of edges between node i and the nodes in module S i where k i is the degree of node i, M represents a module, and N M represents all modules; According to the formula, if the edges related to a certain node are evenly distributed among all modules, then the P value of this node i is close to 1; if all the edges related to a certain node are within the module to which it belongs, then the P i value of this node is 0. According to the topological characteristics of the nodes, the node attributes are divided into 4 types, including: Z i > 2.5 and P i < 0.62, determine that this node is the module center point; Z i <2.5 and P i > 0.62, determine that this node is a connection node; Z i > 2.5 and P i > 0.62, determine that this node is the network center point; Z i <2.5 and P i <0.62, determine that this node is a peripheral node; Connecting nodes, module central points, and network central points are key nodes, that is, the key bacteria in the microbial co-occurrence network.

2. The method for mining functional modules of non-control-population-dependent microbial network analysis according to claim 1, characterized in that, In step 3.1.1, by comparing the global indicators of the microbial co-occurrence network, the differences in the interaction relationships between the sample flora of each group are evaluated, including: The indicators describing the scale of the microbial co-occurrence network include distance and diameter. The distance refers to the shortest distance between two points in the microbial co-occurrence network, and the diameter refers to the longest distance between two points in the microbial co-occurrence network. The indicators describing the cohesion of the microbial co-occurrence network include graph density and clustering coefficient. The graph density refers to the ratio of the number of actually existing edges to the number of possible edges in the microbial co-occurrence network; the actually existing edges refer to all the edges in a microbial co-occurrence network; the possible edges refer to the sum of the number of edges that would occur if there were an edge between any two points in the microbial co-occurrence network; the clustering coefficient refers to the relative frequency of the formation of triangles by the closure of connected triples, and a connected triple is three points connected by two edges. The indicators describing the connectivity of the microbial co-occurrence network include average degree and average weighted degree. The average degree refers to the degree of all points in the microbial co-occurrence network, and the degree is the average value of the number of edges associated with the node; the average weighted degree is the average value of the weighted degrees of all points in the microbial co-occurrence network. Compare the values of the above global indicators of the microbial co-occurrence networks constructed by different groups. Calculate the indicators describing the scale of the microbial co-occurrence network, including distance and diameter. Calculate the indicators describing the cohesion of the network, including graph density and clustering coefficient. Calculate the indicators describing the connectivity of the network, including average degree and average weighted degree. Obtain the values of these indicators for the microbial co-occurrence networks constructed by different groups, and compare the properties of the networks according to the magnitudes of these indicator values.

3. The method for mining functional modules of non-control-population-dependent microbial network analysis according to claim 1, wherein, In step 3.1.2, The point characteristics include: node degree, closeness centrality, betweenness centrality, and eigenvector centrality. The node degree refers to: in a microbial co-occurrence network G=(V, E), the node degree of point v is the number of edges associated with point v; V refers to the set of all points in the microbial co-occurrence network, and E refers to the set of all edges in the microbial co-occurrence network. The closeness centrality refers to: the reciprocal of the sum of the distances from a certain point in the microbial co-occurrence network to all other points. Betweenness centrality refers to: the number of shortest paths passing through a certain point for any two points other than that point divided by the sum of the number of shortest paths for any two points; Eigenvector centrality refers to: obtained by performing eigenvalue decomposition on the adjacency matrix of the network; Calculate the characteristic values of each point in the microbial co-occurrence network, obtain the numerical distribution of the characteristics of each point, perform a KS test, obtain the significance p2 value of the test result, and set p2 <= 0.

05. The two correlation networks in different parts constructed are significantly different.

4. The method for mining functional modules of non-control-population-dependent microbial network analysis according to claim 1, wherein, In step 3.2.1, the constructed microbial co-occurrence network is divided into different sub-networks through the community partitioning algorithm of random walk, including: The probability of walking from point i to point j is defined as distance similarity. If i and j are in the same subgroup, obviously the probability is higher. Therefore, a hierarchical clustering method is adopted to establish the subgroup structure.

5. The method for mining functional modules of non-control-population-dependent microbial network analysis according to any one of claims 1-4, wherein, In step (4), functional modules are mined through the WGCNA weighted gene co-expression network analysis method to find the microbial communities related to physiological and metabolic indicators.

6. A system for mining functional modules of non-control-population-dependent microbial network analysis, wherein, Implement a method for mining functional modules of a non-control-population-dependent microbial network analysis according to any one of claims 1-5, including a data preparation unit, a microbial co-occurrence network construction unit, a microbial co-occurrence network analysis unit, and a functional module mining unit; The data preparation unit is used to execute step (1); the microbial co-occurrence network construction unit is used to execute step (2); the microbial co-occurrence network analysis unit is used to execute step (3); the functional module mining unit is used to execute step (4).

Citation Information

Patent Citations

  • Method for mining relevance of species in micro-organisms

    CN110136778A

  • Method for identifying corn straw carbon assimilation key microorganisms in soil

    CN113658637A