Core flora identification method for large-scale lung micro-ecological flora interaction network
By standardizing and clustering the metagenomic sequencing samples of bronchoalveolar lavage fluid, a microbial interaction network was constructed, and core microbial communities with stable interactions were identified. This solved the problem of difficulty in analyzing the lung microecological interaction network in existing technologies, and achieved the identification of core microbial communities with high stability and wide applicability, providing technical support for the diagnosis of lung diseases.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-23
- Publication Date
- 2026-03-27
AI Technical Summary
Existing technologies are insufficient to accurately analyze the pulmonary microecological interaction network, resulting in poor pathogen reproducibility and low transformability. Furthermore, the handling of bronchoalveolar lavage fluid is highly invasive and has a limited sample size, which affects the size of the research cohort and statistical power.
By collecting metagenomic sequencing samples from bronchoalveolar lavage fluid, performing standardized processing, unsupervised clustering and population stratification, constructing a microbial community interaction network, identifying stable interoperating core microbial communities, and performing module identification to determine the final core microbial community.
It achieves large-scale, highly stable identification of core microbial communities, overcomes the statistical bias caused by sample type limitations and small-scale cohorts, and provides a reliable theoretical basis for the diagnosis and intervention of lung diseases.
Smart Images

Figure CN121747902A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of microbial community identification technology, and more specifically, to a core microbial community identification method for large-scale lung microecological microbial community interaction networks. Background Technology
[0002] Recent studies have confirmed that the lungs are not a sterile environment in the traditional sense, but a dynamic and complex microbial ecosystem. This ecological network is closely related to lung health and the development of diseases. However, most clinical studies still rely on easily obtainable sputum or throat swab samples for lung microbiome analysis. These samples mainly reflect the upper respiratory tract microbiota and are often contaminated by oral colonizing bacteria, making it difficult to accurately represent the microbial structure of the lower respiratory tract. Although bronchoalveolar lavage fluid (BALF) can more accurately characterize the lower respiratory tract microbiota, its collection requires bronchoscopy and specialized equipment, which is not only highly invasive and burdensome for patients, but also often results in limited sample sizes. This further leads to small study cohort sizes and high heterogeneity between cohorts, thus severely limiting the statistical power and generalizability of lower respiratory tract microecological studies.
[0003] Numerous studies have reported significant changes in the composition and abundance of the lung microbiome under pathological conditions such as chronic obstructive pulmonary disease (COPD), acute respiratory distress syndrome (ARDS), and various types of pneumonia. However, existing research largely focuses on differential abundance analyses of single or a few pathogens, failing to reveal the complex interactions among microbial communities, including synergistic, antagonistic, and metabolic complementarities, from a holistic microecological perspective. In fact, the interaction networks formed by nutrient competition, metabolite exchange, and microbial community regulation are key to determining community stability, functional redundancy, and their impact on the host immune response. However, current techniques struggle to accurately analyze these dynamic network structures, resulting in poor reproducibility and low translatability of many pathogens or candidate biomarkers across different populations or research platforms.
[0004] Given the aforementioned shortcomings, there is an urgent need to develop a core microbial community identification method based on a large-scale pulmonary microbial interaction network. This method aims to stably capture core microbial communities that play a crucial role in maintaining pulmonary homeostasis and resisting pathogen invasion from massive samples and multi-center cohorts. The method should overcome limitations in sample type, contamination interference, and statistical bias caused by small-scale cohorts. By systematically reconstructing the microbial interaction network and combining it with robust screening algorithms, it can identify core microorganisms with high ecological stability and functional representativeness, providing a new theoretical foundation and technical support for the diagnosis, prognostic assessment, and development of intervention strategies for respiratory pathogen infections. Summary of the Invention
[0005] The present invention provides a core microbial community identification method for large-scale lung microbial community interaction networks, aiming to solve the problems of limited sample types, contamination interference, and statistical bias caused by small-scale cohorts.
[0006] This invention provides a method for identifying core microbiota in large-scale lung microbiota interaction networks, comprising: Bronchoalveolar lavage fluid metagenomic sequencing samples were collected and the bronchoalveolar lavage fluid metagenomic sequencing samples were standardized to obtain standardized bronchoalveolar lavage fluid metagenomic sequencing samples. Based on the standardized bronchoalveolar lavage fluid metagenomic sequencing samples, unsupervised clustering and population stratification were performed on the microbial profiles of multiple centers to obtain the microbial clusters existing in multiple centers. Based on the microbial clusters existing in multiple centers, a microbial community interaction network is constructed through different subtypes, and the microbial community interaction network is used to identify microbial communities that have stable interactions in different microecological subtype clusters, thus obtaining the core microbial community identification results. The core microbial community identification results are integrated and module identification is performed to determine the final core microbial community identification results.
[0007] In one possible implementation, collecting bronchoalveolar lavage fluid metagenomic sequencing samples includes: collecting bronchoalveolar lavage fluid metagenomic sequencing samples corresponding to patients of different ages, and said patients having different underlying diseases and / or lung diseases.
[0008] In one possible implementation, the bronchoalveolar lavage fluid metagenomic sequencing sample is standardized to obtain a standardized bronchoalveolar lavage fluid metagenomic sequencing sample, including: The host sequence is removed from the bronchoalveolar lavage fluid metagenomic sequencing sample to obtain a bronchoalveolar lavage fluid metagenomic sequencing sample with the host sequence removed. The host sequence-removed bronchoalveolar lavage fluid metagenomic sequencing samples were compared using a pre-set pathogen reference genome database to achieve abundance species screening and whole genome comparison verification, and to obtain screening verification results. The screening and verification results were subjected to background bacteria filtration and false positive bacteria filtration to determine the filtration results; Based on the filtering results, the number of reads for microbial species was normalized to relative abundance, and the samples were screened based on the relative abundance to determine the metagenomic sequencing samples of bronchoalveolar lavage fluid after standardization.
[0009] In one possible implementation, based on the standardized bronchoalveolar lavage fluid metagenomic sequencing samples, unsupervised clustering and population stratification are performed on the microbial profiles of multiple centers to obtain the microbial clusters present in multiple centers, including: Obtain the relative abundance matrix of metagenomic sequencing samples of bronchoalveolar lavage fluid after standardization, and obtain the Bray-Curtis dissimilarity based on the relative abundance matrix; Based on the Bray-Curtis dissimilarity, a dissimilarity matrix is constructed, and the PAM clustering algorithm is used to cluster the dissimilarity matrix to obtain the clustering results; Based on the clustering results, the silhouette coefficient method is used to evaluate the clustering quality in order to select the optimal number of clusters with the best silhouette coefficient. Microbial clusters existing in multiple centers are screened based on the optimal number of clusters with the best profile coefficient.
[0010] In one possible implementation, the PAM clustering algorithm is used to cluster the dissimilarity matrix to obtain the clustering results, including: Multiple standardized bronchoalveolar lavage fluid metagenomic sequencing samples were randomly selected to initialize centroid points; After standardization, the metagenomic sequencing samples of bronchoalveolar lavage fluid are assigned to the cluster of the centroid with the lowest dissimilarity. All samples in the cluster of the centroid are traversed, and the sample with the lowest total dissimilarity within the cluster is selected as the new centroid. This step is repeated until the centroid no longer changes or the maximum number of iterations is reached to obtain the clustering result.
[0011] In one possible implementation, based on the clustering results, the silhouette coefficient method is used to evaluate the clustering quality in order to select the optimal number of clusters with the best silhouette coefficient, including: For any cluster in the clustering results, obtain the sample contour coefficients corresponding to the samples in that cluster, and obtain the total contour coefficients of the cluster based on the sample contour coefficients. Based on the total profile coefficient of the cluster to which it belongs, select the optimal number of clusters with the best profile coefficient.
[0012] In one possible implementation, selecting the optimal number of clusters with the best profile coefficient based on the total profile coefficient of the cluster includes: Set the range of clusters to be evaluated; For any cluster number value within the cluster number range, the Bray-Curtis dissimilarity matrix is used as input to perform PAM clustering on each candidate cluster number and calculate the average profile coefficient corresponding to each cluster number to obtain the average profile coefficient corresponding to the cluster number value. The cluster number with the highest average profile coefficient is selected as the optimal cluster number.
[0013] In one possible implementation, screening microbial clusters existing in multiple centers based on the optimal number of clusters with the best profile coefficient includes: screening multiple microbial clusters existing in the largest number of centers based on the optimal number of clusters with the best profile coefficient.
[0014] In one possible implementation, based on the aforementioned microbial clusters existing in multiple centers, a microbial community interaction network is constructed through different subtype hierarchies. This network is then used to identify stable interoperating microbial communities within different microecological subtype clusters, yielding core microbial community identification results, including: For microbial clusters present in multiple centers, species present in less than 20% of the samples are removed to obtain the microbial clusters after removal; The Fastspar algorithm was used to obtain the abundance correlation between species in the microbial clusters after removal, and the abundance correlation was subjected to a permutation test. Only significant interaction relationships with p-value < 0.01 were retained to construct a microbial cluster-specific community interaction network. Identify the interacting microbial groups that can maintain stable interactions in all microbial clusters to obtain the core taxa; Subnetworks containing only core taxa are extracted from the microbial community interaction networks corresponding to all microbial clusters, and subnetworks from different cluster sources are weighted and fused to obtain the core microbial community identification results.
[0015] In one possible implementation, the core microbial community identification results are integrated and module identification is performed to determine the final core microbial community identification results, including: Based on the core microbial community identification results of all centers, species that are identified as core taxa in at least two center datasets are selected to obtain the selection results; The core bacterial interaction networks of each queue corresponding to the screening results are weighted and fused, and then network modules are identified based on hierarchical clustering to obtain the final core bacterial community identification results.
[0016] Beneficial effects: This invention provides a method for identifying core microbiota in large-scale pulmonary microbiome interaction networks. Using standardized bronchoalveolar lavage fluid metagenomic samples, multi-center unsupervised clustering and population stratification are employed to identify stable microbial clusters across centers. Based on this, a hierarchical microbiome interaction network is constructed, screening for stable interactions among different microbiome subtypes. Finally, modular integration is used to determine the final core microbiota. This method overcomes the shortcomings of previous techniques, such as sample contamination, low statistical power in small-scale cohorts, and neglect of microbiome interactions. It achieves large-scale, highly stable core microbiota identification, providing a reliable theoretical foundation and technical support for the diagnosis and intervention of lung diseases. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart of a core microbial community identification method for large-scale lung microbial community interaction networks proposed in an embodiment of the present invention. Figure 2 This is a schematic diagram of the silhouette coefficients and optimal cluster number of five central datasets proposed in an embodiment of the present invention; Figure 3 This is a schematic diagram of five different microbial communities obtained based on unsupervised clustering, as proposed in an embodiment of the present invention; Figure 4 This is a schematic diagram showing the distribution of two antagonistic core bacterial groups (N-Core & P-Core) in critically ill and non-critically ill patients according to an embodiment of the present invention. Figure 5 This is a schematic diagram illustrating the relationship between two sets of mutually antagonistic core microbial communities (N-Core & P-Core) proposed in an embodiment of the present invention and inflammatory response, age, microecological diversity, and pathogen load. Detailed Implementation The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0019] The theoretical basis of the lung core microbiota identification method originates from the concept of "core microorganisms" in community ecology: under varying host and environmental conditions, only those strains that consistently maintain stable interaction patterns and act as key nodes in the ecological network can be considered the "pillars" maintaining the overall structure and function of the community. Based on this, this invention first aggregated metagenomic sequencing data from tens of thousands of bronchoalveolar lavage fluid (BALF) samples from dozens of medical centers across the country. Using short reads obtained through high-throughput sequencing, this invention employed a well-constructed pathogen reference genome database for comparison, accurately annotating the microbial species composition of each sample. Subsequently, the self-developed ScanTrack algorithm was used to remove environmental and reagent background bacteria, and the abundance of the remaining species was deeply corrected to minimize false positive interference.
[0020] Building upon this foundation, unsupervised clustering techniques were applied to the corrected microbiome profiles, dividing the samples into several microbial clusters with similar community structure characteristics to achieve subtype stratification of the population. Each microbial cluster can reflect the differences in clinical phenotypes or pathological states of the corresponding host. For samples within each cluster, a co-occurrence and co-loss network between species was further constructed. By calculating indicators such as node degree centrality, intra-module connectivity, and network edge weights, the stability of interactions between species was quantitatively assessed. Finally, through comparison and fusion of cross-cluster network patterns, key species exhibiting high-confidence interactions in all independent co-occurrence networks were screened as the core pulmonary microbiota. Ultimately, this invention identified two antagonistic groups of microbiota in the lungs. One group (N-Core, 88 species) was associated with the host's health status, while the other antagonistic group (P-Core, 48 species) was associated with disease states and more severe inflammatory states. By integrating the core microbiota features into a machine learning model, this invention significantly improved the diagnostic performance of lung diseases.
[0021] like Figure 1 As shown, this embodiment of the invention provides a method for identifying core microbiota in large-scale lung microbiota interaction networks, including: S1. Collect metagenomic sequencing samples of bronchoalveolar lavage fluid and standardize the metagenomic sequencing samples of bronchoalveolar lavage fluid to obtain standardized metagenomic sequencing samples of bronchoalveolar lavage fluid. S2. Based on the metagenomic sequencing samples of the bronchoalveolar lavage fluid after the standardization process, unsupervised clustering and population stratification are performed on the microbial profiles of multiple centers to obtain the microbial clusters existing in multiple centers. In this embodiment, the "center" refers to a regional laboratory center or medical center. Bronchoalveolar lavage fluid metagenomic sequencing samples also need to be collected from multiple centers. It is important to note that this "center" is different from a clustering center.
[0022] S3. Based on the microbial clusters existing in multiple centers, construct a microbial community interaction network through different subtypes, and identify stable interoperating microbial communities in different microecological subtype clusters through the microbial community interaction network to obtain the core microbial community identification result; S4. Integrate the core microbial community identification result and perform module identification to determine the final core microbial community identification result.
[0023] In one possible implementation, collecting bronchoalveolar lavage fluid metagenomic sequencing samples includes: collecting bronchoalveolar lavage fluid metagenomic sequencing samples corresponding to patients of different ages, and said patients having different underlying diseases and / or lung diseases.
[0024] For example, a total of 16,319 metagenomic sequencing samples of bronchoalveolar lavage fluid from Shanghai (SH), Zhengzhou (ZZ), Chengdu (CD), Tianjin (TJ), and Guangzhou (GZ) were collected, spanning different age groups and including hospitalized patients with various underlying diseases or lung diseases.
[0025] In one possible implementation, the bronchoalveolar lavage fluid metagenomic sequencing sample is standardized to obtain a standardized bronchoalveolar lavage fluid metagenomic sequencing sample, including: The host sequence is removed from the bronchoalveolar lavage fluid metagenomic sequencing sample to obtain a bronchoalveolar lavage fluid metagenomic sequencing sample with the host sequence removed. The host sequence-removed bronchoalveolar lavage fluid metagenomic sequencing samples were compared using a pre-set pathogen reference genome database to achieve abundance species screening and whole genome comparison verification, and to obtain screening verification results. The screening and verification results were subjected to background bacteria filtration and false positive bacteria filtration to determine the filtration results; Based on the filtering results, the number of reads for microbial species was normalized to relative abundance, and the samples were screened based on the relative abundance to determine the metagenomic sequencing samples of bronchoalveolar lavage fluid after standardization.
[0026] For example, a) First, strict quality control is performed using the bioinformatics quality control software fastp (v0.22.0), requiring the output data to have at least 1 Mb of clean reads and a Q30 score >85%. Then, the quality-controlled reads are aligned to the human genome (T2T-CHM13c2.0) using Bowtie2 (v2.3.5.1) to remove host sequences from the samples.
[0027] b. The host-de-hosted sequences were aligned to a self-built pathogen reference genome database using bwa-men. This database integrates NCBI RefSeq and GenBank, with only a single high-quality representative strain from RefSeq or GenBank included for each species, covering 9855 bacteria, 6926 viruses, 1582 fungi, 312 parasites, 177 mycobacteria, and 184 mycoplasma / chlamydia species. The Least Common Ancestor (LCA) algorithm was applied to reads cross-aligned across multiple species, including only species-specific sequences. SAMtools and BEDtools were used to calculate the genome coverage (similarity) and average sequencing depth (number of reads covered) of each species. For low-abundance species with <3 unique aligned reads, BLASTn whole-genome alignment was used for verification, retaining only results with high consistency (>90% similarity) and full-length matching.
[0028] c. Based on the ScanTrack algorithm, compare the SNP spectra of bronchoalveolar lavage fluid samples and paired negative control samples to identify background bacteria in the samples, quantify the abundance of colonization-contamination mixed bacteria, and then perform correction.
[0029] d. For the removal of false positives, the microbial reference genome is divided into 1kb continuous intervals, and the number of intervals covered by reads for each species in the sample (S_bin_num) is counted. Based on the assumption of uniform sequencing coverage, a random forest regression model is constructed to predict S_bin_num, using the total alignment reads and sequencing depth as predictions. Subsequently, based on the predicted values and actual observations, a coverage uniformity index bin_FC is defined as: bin_FC = measured S_bin_num / model predicted S_bin_num. False positive species with bin_FC > 2 and S_bin_num < 10 are excluded.
[0030] e. Normalize the number of reads for microbial species in the samples to relative abundance, and remove species appearing in <10 samples.
[0031] In one possible implementation, based on the standardized bronchoalveolar lavage fluid metagenomic sequencing samples, unsupervised clustering and population stratification are performed on the microbial profiles of multiple centers to obtain the microbial clusters present in multiple centers, including: Obtain the relative abundance matrix of metagenomic sequencing samples of bronchoalveolar lavage fluid after standardization, and obtain the Bray-Curtis dissimilarity based on the relative abundance matrix; For example, for each queue, samples are first obtained. x Species relative abundance matrix: Where n is the number of samples and p is the number of microbial species. For the sample i In species j The relative abundance (the sum of relative abundance is 1).
[0032] Subsequently, the Bray-Curtis dissimilarity between sample pairs is calculated using the following formula: , in, These represent the relative abundance of species j in samples A and B, respectively. Next, we construct... Dissimilarity matrix. Here, the Bray-Curtis dissimilarity between samples is calculated, and the results are filled into the matrix: This matrix satisfies: = 0 (The sample is completely similar to itself); (0: completely similar; 1: completely different); Based on the Bray-Curtis dissimilarity, a dissimilarity matrix is constructed, and the PAM clustering algorithm is used to cluster the dissimilarity matrix to obtain the clustering results; Based on the clustering results, the silhouette coefficient method is used to evaluate the clustering quality in order to select the optimal number of clusters with the best silhouette coefficient. Microbial clusters existing in multiple centers are screened based on the optimal number of clusters with the best profile coefficient.
[0033] In one possible implementation, the PAM clustering algorithm is used to cluster the dissimilarity matrix to obtain the clustering results, including: Multiple standardized bronchoalveolar lavage fluid metagenomic sequencing samples were randomly selected to initialize centroid points; After standardization, the metagenomic sequencing samples of bronchoalveolar lavage fluid are assigned to the cluster of the centroid with the lowest dissimilarity. All samples in the cluster of the centroid are traversed, and the sample with the lowest total dissimilarity within the cluster is selected as the new centroid. This step is repeated until the centroid no longer changes or the maximum number of iterations is reached to obtain the clustering result.
[0034] For example, the PAM clustering algorithm is used to... Clustering based on dissimilarity matrix: (1) Initialize centroids: randomly select k One sample was used as the initial center point ( k (2) Sample allocation: Assign each sample to the cluster to which the centroid with the lowest dissimilarity belongs. (3) Centroid update: Traverse all samples in the current cluster and select the sample with the lowest total dissimilarity in the cluster as the new centroid, i.e.: ;in, Let c be the sample set of the c-th cluster. (4) Iterative optimization: Repeat steps (2) and (3) until the center point no longer changes or the maximum number of iterations is reached.
[0035] In one possible implementation, based on the clustering results, the silhouette coefficient method is used to evaluate the clustering quality in order to select the optimal number of clusters with the best silhouette coefficient, including: For any cluster in the clustering results, obtain the sample contour coefficients corresponding to the samples in that cluster, and obtain the total contour coefficients of the cluster based on the sample contour coefficients. Based on the total profile coefficient of the cluster to which it belongs, select the optimal number of clusters with the best profile coefficient.
[0036] In one possible implementation, selecting the optimal number of clusters with the best profile coefficient based on the total profile coefficient of the cluster includes: Set the range of clusters to be evaluated; For any cluster number value within the cluster number range, the Bray-Curtis dissimilarity matrix is used as input to perform PAM clustering on each candidate cluster number and calculate the average profile coefficient corresponding to each cluster number to obtain the average profile coefficient corresponding to the cluster number value. The cluster number with the highest average profile coefficient is selected as the optimal cluster number.
[0037] In one possible implementation, screening microbial clusters existing in multiple centers based on the optimal number of clusters with the best profile coefficient includes: screening multiple microbial clusters existing in the largest number of centers based on the optimal number of clusters with the best profile coefficient.
[0038] For example, the silhouette index is used to evaluate clustering quality, and the optimal number of clusters with the best silhouette index is selected. The silhouette index measures the compactness of a sample within its own cluster and its separation from its nearest neighbor clusters. For each sample i, its silhouette index is defined as: Where: a(i) represents the average dissimilarity (average distance within the cluster) between sample i and other samples in the same cluster; b(i) represents the average dissimilarity (minimum distance between clusters) between sample i and its nearest neighbor cluster; s(i) ranges from [-1, 1], the closer to 1, the more clearly the sample belongs to this cluster (good clustering effect); approximately 0 indicates that the sample is located at the boundary between two clusters; the closer to -1, the more likely the sample is to be misclassified into a cluster. The average of all samples yields the overall silhouette coefficient of the cluster: Next, the optimal number of clusters k is selected through the following steps: Set the range of clusters to be evaluated (here, k = 2 to k = 10); use the previously calculated Bray-Curtis dissimilarity matrix as input, and perform PAM clustering on each candidate cluster number; calculate the average silhouette coefficient S(k) corresponding to each cluster number; select the cluster number with the highest average silhouette coefficient as the optimal cluster number. Ultimately, the optimal number of clusters for different multicenter datasets fluctuated between 3 and 4. The results of silhouette coefficient calculation and optimal cluster number selection using the above method are as follows: Figure 2 As shown.
[0039] Ultimately, five distinct microbial clusters were identified that were widely distributed across multiple centers. These were: P1 (P. mela) driven by *Prevotelamelaninogenica*, P2 (R. muci) driven by *Rothia mucilaginosa*, P3 (L. theo) driven by *Lasiodiplodia theobromae*, P4 (P. aeru) driven by *Pseudomonaes aeruginosa*, and P5 (A. bau) driven by *Acinetobacter baumannii*. The driving bacteria were defined as the bacteria with the highest relative abundance in each cluster. The overall clustering results for multiple datasets are shown below. Figure 3 As shown.
[0040] In one possible implementation, based on the aforementioned microbial clusters existing in multiple centers, a microbial community interaction network is constructed through different subtype hierarchies. This network is then used to identify stable interoperating microbial communities within different microecological subtype clusters, yielding core microbial community identification results, including: For microbial clusters present in multiple centers, species present in less than 20% of the samples are removed to obtain the microbial clusters after removal; The Fastspar algorithm was used to obtain the abundance correlation between species in the microbial clusters after removal, and the abundance correlation was subjected to a permutation test. Only significant interaction relationships with p-value < 0.01 were retained to construct a microbial cluster-specific community interaction network. Identify the interacting microbial groups that can maintain stable interactions in all microbial clusters to obtain the core taxa; Subnetworks containing only core taxa are extracted from the microbial community interaction networks corresponding to all microbial clusters, and subnetworks from different cluster sources are weighted and fused to obtain the core microbial community identification results.
[0041] For example, species present in more than 20% of the samples in each central dataset are retained to ensure high detection rates of microbial community members in the analysis. Then, Fastspar (a fast, optimized version of the SparCC algorithm) is used to calculate the abundance correlations between species for each microbial cluster. This algorithm effectively avoids compositional bias in microbiome data through logarithmic ratio transformation of component data. 1000 permutation tests are performed on each pair of species correlations, retaining only significant interactions with p-values <0.01 to construct a microbial cluster-specific co-occurrence network. Subsequently, interacting microbial communities that maintain stable interactions across all clusters (i.e., whose association strength and direction are highly consistent across multiple resampling attempts) are identified and labeled as the core taxa of that cluster. Then, a sub-network containing only the core taxa is extracted from the microbial interaction network of each cluster, and these core microbial interaction sub-networks from different clusters are weighted and fused. Due to the varying sample sizes of different clusters (…),… Differences may introduce statistical power bias; therefore, a univariate weighting scheme is used to weight and fuse the subnetworks. Conditional bias is calculated as follows: , in, r i The intra-cluster correlation coefficient. v i Reflects the probability of cluster estimates ( v i The smaller the value, the higher the weight. Then, the weighted correlation coefficients of the edges in the fused network are calculated: in, k The number of clusters within the queue, and the final edge weights of the fused network ( The weighted average contribution of each cluster is represented by . Subsequently, the fused network is transformed into a similarity matrix between species, and a hierarchical clustering algorithm (Ward's method) is applied to determine the optimal number of modules by maximizing the average silhouette coefficient. This invention finds that the core microbial community network of all centers can be divided into two mutually antagonistic core modules.
[0042] In one possible implementation, the core microbial community identification results are integrated and module identification is performed to determine the final core microbial community identification results, including: Based on the core microbial community identification results of all centers, species that are identified as core taxa in at least two center datasets are selected to obtain the selection results; The core bacterial interaction networks of each cohort corresponding to the screening results are weighted and fused, and then network modules are identified based on hierarchical clustering to obtain the final core bacterial community identification results. Specifically, species identified as core taxa in at least two central datasets are first screened to ensure their universality across populations. Subsequently, the core bacterial interaction networks of each cohort are integrated again using the weighted fusion network method described above, and then network modules are identified based on hierarchical clustering. This invention ultimately identified two groups of mutually antagonistic core bacterial communities. One group (N-Core, 88 bacteria) is dominated by symbiotic bacteria (such as *Prevotella* and *Veillonella*), and is significantly associated with reduced inflammation levels, increased microbial diversity, and enhanced colonization resistance. The other group (P-Core, 48 bacteria) is enriched with opportunistic pathogens (such as *Pseudomonas* and *Acinetobacter*), which are significantly amplified in critically ill patients and are associated with exacerbated inflammatory responses and pathogen invasion and infection, such as... Figure 4-5 As shown.
[0043] The present invention has the following technical effects: 1. This invention overcomes the limitations of single-network analysis by constructing a cross-subtype fusion network (integrating co-occurrence patterns of different microbial clusters) to screen core microbial communities that maintain stable interaction relationships across multiple host states. Simultaneously, it introduces network module identification to reveal the role of core microbial communities in maintaining community functional redundancy or resistance to disturbances, rather than focusing solely on abundance changes of individual species.
[0044] 2. This invention integrates bronchoalveolar lavage fluid (BALF) samples to construct a population cohort covering all age groups and various underlying disease states. This overcomes the limitations of single-center studies, which have a single sample source and high population homogeneity. It can improve the robustness and universality of the results and reduce the bias of single-center studies.
[0045] 3. This invention significantly improves the accuracy of lung disease diagnosis and prognostic assessment by integrating two antagonistic groups of core microbiota in the lung microbiome (such as the N-Core microbiota with immune protection and the P-Core microbiota with pro-inflammatory and pathogenic effects) as key features into a machine learning model. In the diagnostic model, the abundance ratio of N-Core / P-Core effectively distinguishes between severe and mild cases by quantifying the dynamic balance of "microbiome defense-attack". Furthermore, the prognostic model shows that patients with log(N / P)>0 (i.e., the N-Core dominant group) have a lower 30-day mortality rate and a shorter average hospital stay.
[0046] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0047] Embodiments of the present invention are described with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, electronic devices, and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0048] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0049] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0050] Although preferred embodiments of the present invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present invention.
[0051] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.
[0052] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A method for identifying core microbiota in large-scale pulmonary microbial interaction networks, characterized in that, include: Bronchoalveolar lavage fluid metagenomic sequencing samples were collected and the bronchoalveolar lavage fluid metagenomic sequencing samples were standardized to obtain standardized bronchoalveolar lavage fluid metagenomic sequencing samples. Based on the standardized bronchoalveolar lavage fluid metagenomic sequencing samples, unsupervised clustering and population stratification were performed on the microbial profiles of multiple centers to obtain the microbial clusters existing in multiple centers. Based on the microbial clusters existing in multiple centers, a microbial community interaction network is constructed through different subtypes, and the microbial community interaction network is used to identify microbial communities that have stable interactions in different microecological subtype clusters, thus obtaining the core microbial community identification results. The core microbial community identification results are integrated and module identification is performed to determine the final core microbial community identification results.
2. The method for identifying core microbiota in a large-scale pulmonary microbial community interaction network according to claim 1, characterized in that, Collecting metagenomic sequencing samples of bronchoalveolar lavage fluid, including collecting metagenomic sequencing samples of bronchoalveolar lavage fluid from patients of different ages, and the patients having different underlying diseases and / or lung diseases.
3. The method for identifying core microbiota in a large-scale pulmonary microbial community interaction network according to claim 1, characterized in that, The bronchoalveolar lavage fluid metagenomic sequencing samples were standardized to obtain standardized bronchoalveolar lavage fluid metagenomic sequencing samples, including: The host sequence is removed from the bronchoalveolar lavage fluid metagenomic sequencing sample to obtain a bronchoalveolar lavage fluid metagenomic sequencing sample with the host sequence removed. The host sequence-removed bronchoalveolar lavage fluid metagenomic sequencing samples were compared using a pre-set pathogen reference genome database to achieve abundance species screening and whole genome comparison verification, and to obtain screening verification results. The screening and verification results were subjected to background bacteria filtration and false positive bacteria filtration to determine the filtration results; Based on the filtering results, the number of reads for microbial species was normalized to relative abundance, and the samples were screened based on the relative abundance to determine the metagenomic sequencing samples of bronchoalveolar lavage fluid after standardization.
4. The method for identifying core microbiota in a large-scale pulmonary microbial community interaction network according to claim 1, characterized in that, Based on the standardized bronchoalveolar lavage fluid metagenomic sequencing samples, unsupervised clustering and population stratification were performed on the microbial profiles of multiple centers to obtain the microbial clusters present in multiple centers, including: Obtain the relative abundance matrix of metagenomic sequencing samples of bronchoalveolar lavage fluid after standardization, and obtain the Bray-Curtis dissimilarity based on the relative abundance matrix; Based on the Bray-Curtis dissimilarity, a dissimilarity matrix is constructed, and the PAM clustering algorithm is used to cluster the dissimilarity matrix to obtain the clustering results; Based on the clustering results, the silhouette coefficient method is used to evaluate the clustering quality in order to select the optimal number of clusters with the best silhouette coefficient. Microbial clusters existing in multiple centers are screened based on the optimal number of clusters with the best profile coefficient.
5. The method for identifying core microbiota in a large-scale pulmonary microbial community interaction network according to claim 4, characterized in that, The PAM clustering algorithm is used to cluster the dissimilarity matrix, and the clustering results are obtained, including: Multiple standardized bronchoalveolar lavage fluid metagenomic sequencing samples were randomly selected to initialize centroid points; After standardization, the metagenomic sequencing samples of bronchoalveolar lavage fluid are assigned to the cluster of the centroid with the lowest dissimilarity. All samples in the cluster of the centroid are traversed, and the sample with the lowest total dissimilarity within the cluster is selected as the new centroid. This step is repeated until the centroid no longer changes or the maximum number of iterations is reached to obtain the clustering result.
6. The method for identifying core microbiota in a large-scale pulmonary microbial community interaction network according to claim 5, characterized in that, Based on the clustering results, the silhouette coefficient method is used to evaluate the clustering quality in order to select the optimal number of clusters with the best silhouette coefficient, including: For any cluster in the clustering results, obtain the sample contour coefficients corresponding to the samples in that cluster, and obtain the total contour coefficients of the cluster based on the sample contour coefficients. Based on the total profile coefficient of the cluster to which it belongs, select the optimal number of clusters with the best profile coefficient.
7. The method for identifying core microbiota in a large-scale pulmonary microbial community interaction network according to claim 6, characterized in that, Based on the total profile coefficient of the cluster, the optimal number of clusters with the best profile coefficient is selected, including: Set the range of the number of clusters to be evaluated; For any cluster number value within the cluster number range, the Bray-Curtis dissimilarity matrix is used as input to perform PAM clustering on each candidate cluster number and calculate the average profile coefficient corresponding to each cluster number to obtain the average profile coefficient corresponding to the cluster number value. The cluster number with the highest average profile coefficient is selected as the optimal cluster number.
8. The method for identifying core microbiota in a large-scale pulmonary microbial community interaction network according to claim 6, characterized in that, Based on the optimal number of clusters with the best profile coefficient, microbial clusters existing in multiple centers are screened, including: based on the optimal number of clusters with the best profile coefficient, multiple microbial clusters existing in the largest number of centers are screened.
9. The method for identifying core microbiota in a large-scale pulmonary microbial community interaction network according to claim 1, characterized in that, Based on the aforementioned microbial clusters existing in multiple centers, a microbial community interaction network is constructed through different subtype hierarchies. This network is then used to identify stable interoperating microbial communities within different microecological subtype clusters, yielding core microbial community identification results, including: For microbial clusters present in multiple centers, species present in less than 20% of the samples are removed to obtain the microbial clusters after removal; The Fastspar algorithm was used to obtain the abundance correlation between species in the microbial clusters after removal, and the abundance correlation was subjected to a permutation test. Only significant interaction relationships with p-value < 0.01 were retained to construct a microbial cluster-specific community interaction network. Identify the interacting microbial groups that can maintain stable interactions in all microbial clusters to obtain the core taxa; Subnetworks containing only core taxa are extracted from the microbial community interaction networks corresponding to all microbial clusters, and subnetworks from different cluster sources are weighted and fused to obtain the core microbial community identification results.
10. The method for identifying core microbiota in a large-scale pulmonary microbial community interaction network according to claim 9, characterized in that, The core microbial community identification results are integrated and module-based to determine the final core microbial community identification results, including: Based on the core microbial community identification results of all centers, species that are identified as core taxa in at least two center datasets are selected to obtain the selection results; The core bacterial interaction networks of each queue corresponding to the screening results are weighted and fused, and then network modules are identified based on hierarchical clustering to obtain the final core bacterial community identification results.