A method and related equipment for microbial grouping analysis
By combining central logarithmic ratio transformation and Euclidean distance quantization with density clustering, the subjective and nonlinear analytical problems in microbial community analysis were solved, realizing automated and standardized microbial grouping analysis, and improving exploration efficiency and consistency of results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-18
- Publication Date
- 2026-03-10
AI Technical Summary
Existing technologies in microbial community analysis suffer from several problems, including strong subjectivity in grouping, weak ability to analyze complex nonlinear relationships, bias in biomarker screening, and difficulty in standardization due to reliance on manual processes. These issues affect the accurate identification of microbial community changes and the efficiency of exploration.
By acquiring microbial data from the target study area, central logarithmic ratio transformation and Euclidean distance quantification are employed. Combined with density clustering and multi-level analysis, sample types are automatically identified, subjective biases are eliminated, natural microbial ecotypes are discovered, and correlation analysis is performed with methane leakage data to achieve standardized and automated grouping analysis.
This approach achieves objectivity and consistency in microbial grouping results, effectively identifies nonlinear distribution patterns and ecological transition zones, improves analytical efficiency and result reproducibility, and provides a foundation for large-scale exploration.
Smart Images

Figure CN121350509B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a method and related equipment for microbial grouping analysis. Background Technology
[0002] Submarine natural gas hydrates represent an important potential energy source, and their exploration and development depend on the accurate identification and assessment of methane leakage activities. Microbial communities are extremely sensitive to environmental changes caused by methane leakage, and their composition and structure can serve as effective biomarkers indicating the distribution of methane leakage and hydrate deposits.
[0003] Currently, the most commonly used method in this field is a comparative analysis of microbial communities based on human experience and traditional statistics. The typical process is as follows: First, researchers divide samples into pre-defined groups based on sampling location (e.g., mining areas versus non-mining areas) or geochemical parameters (e.g., methane concentration thresholds). Second, they compare differences between groups by calculating the α-diversity index and β-diversity distance matrix, and use statistical methods such as ANOSIM to test the significance of differences in community structure between groups. Finally, in biomarker screening, sample groupings are often manually interpreted using visualization methods, and rank-sum tests are used to identify known functional species with significant differences in abundance between groups as candidate biomarkers. Correlation analysis is then used to establish their relationship with environmental parameters.
[0004] However, existing technologies have the following main drawbacks:
[0005] The grouping is highly subjective and masks the true pattern: the entire analysis begins with "artificially pre-defined groupings," such as rigidly dividing the data into "shallow" and "deep" layers based on depth. This subjective division may not conform to the true, continuous, or non-linear natural grouping boundaries of the microbial community, making it impossible to discover ecotypes or transitional states in the data that were not pre-defined.
[0006] Weak ability to interpret complex nonlinear relationships: Traditional multivariate statistical methods (such as PCA) have limited explanatory power for the complex nonlinear interactions between microbial communities and environmental factors, which may lead to misinterpretation of community change drivers or omission of weak correlation signals.
[0007] Biomarker screening is biased: methods tend to focus on "star species" with high abundance and known functions, while ignoring low abundance but high specificity, or synergistic indicator combinations of multiple species, which reduces the sensitivity and specificity of biomarkers.
[0008] The process relies on manual labor and is difficult to standardize and promote: The analysis process involves a lot of manual judgment (grouping, interpretation, species selection), which leads to different results from person to person. The analysis process is difficult to standardize and automate, which limits its application potential in large-scale exploration. Summary of the Invention
[0009] The main objective of this invention is to provide a method, apparatus, electronic device, storage medium, and program product for microbial grouping analysis, aiming to solve at least one problem of the prior art.
[0010] To achieve the above objectives, one aspect of the present invention provides a method for microbial grouping analysis, the method comprising:
[0011] Acquire microbial data for the target study area; the microbial data includes microbial abundance data and methane leakage data from multiple samples;
[0012] Based on microbial abundance data, the transformed value of each sample is obtained through central log ratio transformation, and then the distance between any samples is obtained by Euclidean distance quantization.
[0013] The minimum sample size and neighborhood radius are adaptively set based on the number of samples and species corresponding to the microbial data.
[0014] Density clustering is performed on samples in the microbial data to obtain clustering results. Based on the clustering results, the sample type of each sample is labeled by combining the minimum number of samples and the neighborhood radius.
[0015] The clustering results are analyzed at multiple levels based on the sample type to obtain the analytical results of the clustering results;
[0016] Correlation analysis was performed based on the analysis results and methane leakage data to obtain the microbial grouping analysis results of the target study area.
[0017] In some embodiments, the microbial abundance data includes the raw abundance value of each species in the sample. Based on the microbial abundance data, a transformed value for each sample is obtained through a central log ratio transformation, including the following steps:
[0018] The original abundance values of all species in a single sample are multiplied together to obtain the cumulative result;
[0019] The geometric mean is obtained by taking the number of species as the root of the sum of the products.
[0020] Logarithmically transform the ratio of the original abundance value to the geometric mean to obtain the transformed value for each species in each sample.
[0021] In some embodiments, Euclidean distance quantization is used to obtain the distance value between any samples, including the following steps:
[0022] The Euclidean distance matrix is calculated based on the transformation value of each species in the sample to obtain the distance value between any samples;
[0023] The expression for the distance value is:
[0024] ;
[0025] In the formula, The distance between the i-th sample and the j-th sample is represented by m; m represents the number of species; k represents the k-th species. This represents the transformation value of the k-th species in the i-th sample; This represents the transformation value of the k-th species in the j-th sample.
[0026] In some embodiments, the minimum number of samples and the neighborhood radius are adaptively set based on the number of samples and species corresponding to the microbial data, including the following steps:
[0027] When the number of species is less than or equal to the first threshold and the number of species at a preset multiple does not exceed a preset proportion of the number of samples, the minimum number of samples is set to the number of species at a preset multiple; otherwise, the minimum number of samples is set based on a preset fixed value.
[0028] The value of k is determined based on the minimum number of samples.
[0029] Based on the k value and the distance between all samples, a distance-ranked line chart is drawn using the k-distance plot method.
[0030] The neighborhood radius is determined by the inflection point of the distance-sorted line graph.
[0031] In some embodiments, the clustering result includes multiple sample clusters, and the sample types include core samples, boundary samples, and noise samples. Based on the clustering result, the sample type of each sample is labeled by combining the minimum number of samples and the neighborhood radius, including the following steps:
[0032] If there exists a minimum number of samples within the neighborhood radius of the first sample, then the first sample is determined as the core sample.
[0033] If the second sample is within the neighborhood radius of the core sample and does not meet the conditions of the core sample, the second sample is determined to be a boundary sample.
[0034] If the third sample does not belong to any sample cluster, the third sample is determined to be a noise sample.
[0035] In some embodiments, the sample types include core samples, boundary samples, and noise samples. The clustering result includes multiple sample clusters, each containing core samples and boundary samples while excluding noise samples. The clustering result is then analyzed at multiple levels based on the sample types to obtain the analytical result of the clustering result, including the following steps:
[0036] In-depth analysis of the core ecotype was conducted on each sample cluster to identify the key species driving ecotype differentiation;
[0037] The in-depth analysis of core ecotypes includes differential abundance analysis methods and methods for identifying significantly enriched species.
[0038] Based on the Brectis distance between each boundary sample and the core samples of each sample cluster, a similarity matrix between the boundary sample and each sample cluster is constructed.
[0039] Based on the similarity matrix, the distribution of boundary samples among sample clusters is determined by principal coordinate analysis, thereby identifying transitional state samples;
[0040] Transition state verification is performed on transition state samples based on environmental gradients to obtain analytical results of transition state samples.
[0041] Technical errors are identified based on sequencing depth, sequence quality indicators, and DNA extraction quality of noisy samples, and low biomass samples are then identified after excluding technical anomalies.
[0042] Ecological specificity analysis and biological significance assessment of noise samples were conducted to obtain the results of ecological significance mining of noise samples.
[0043] In some embodiments, the clustering results include multiple sample clusters. Based on the analysis results and methane leakage data, correlation analysis is performed to obtain the microbial grouping analysis results of the target study area, including the following steps:
[0044] The clustering results are spatially correlated with geochemical data, and then the differences in geochemical parameters among different sample clusters are examined through statistical methods.
[0045] Based on the first average distance from the target sample to the same sample cluster and the second average distance from the target sample to other sample clusters, the silhouette coefficient of the target sample is quantified as a clustering quality evaluation index.
[0046] Based on the clustering quality assessment indicators and analysis results, a standardized analysis report was compiled as the microbial grouping analysis results for the target study area.
[0047] To achieve the above objectives, another aspect of the present invention provides a microbial grouping analysis device, the device comprising:
[0048] The first module is used to acquire microbial data for the target study area; the microbial data includes microbial abundance data and methane leakage data from multiple samples.
[0049] The second module is used to obtain the transformed value of each sample based on the microbial abundance data through the central log ratio transformation, and then use Euclidean distance quantization to obtain the distance value between any samples.
[0050] The third module is used to adaptively set the minimum number of samples and the neighborhood radius based on the number of samples and species corresponding to the microbial data.
[0051] The fourth module is used to perform density clustering on samples in the microbial data, obtain clustering results, and label the sample type of each sample based on the clustering results, combined with the minimum number of samples and the neighborhood radius.
[0052] The fifth module is used to perform multi-level analysis of the clustering results based on sample type to obtain the analysis results of the clustering results;
[0053] The sixth module is used to perform correlation analysis based on the analysis results and methane leakage data to obtain the microbial grouping analysis results of the target study area.
[0054] To achieve the above objectives, another aspect of the present invention provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the aforementioned method.
[0055] To achieve the above objectives, another aspect of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned method.
[0056] To achieve the above objectives, another aspect of the present invention provides a computer program product, including a computer program that, when executed by a processor, implements the aforementioned method.
[0057] The embodiments of the present invention include at least the following beneficial effects: The present invention provides a method, apparatus, electronic device, storage medium, and program product for microbial grouping analysis. This scheme acquires microbial data of a target study area; wherein, the microbial data includes microbial abundance data and methane leakage data of multiple samples; based on the microbial abundance data, a transformation value for each sample is obtained through central logarithmic ratio transformation, and then Euclidean distance quantization is used to obtain the distance value between any samples; a minimum sample number and neighborhood radius are adaptively set based on the number of samples and species corresponding to the microbial data; density clustering is performed on the samples in the microbial data to obtain clustering results; based on the clustering results, the sample type of each sample is labeled in combination with the minimum sample number and neighborhood radius; multi-level analysis is performed on the clustering results based on the sample type to obtain the analytical results of the clustering results; correlation analysis is performed based on the analytical results and the methane leakage data to obtain the microbial grouping analysis results of the target study area. This invention automatically groups preprocessed microbial abundance data using density clustering, eliminating the manual pre-setting grouping steps that rely on prior knowledge in traditional methods. This ensures that the grouping results are entirely determined by the distribution characteristics of the data itself, effectively eliminating subjective bias. Specifically, this invention uses clustering based on the density relationships between samples, effectively discovering natural microbial ecotypes of arbitrary shapes and densities within the data. It is particularly suitable for capturing nonlinear distribution patterns and ecological transition zones formed along environmental gradients (such as methane concentration). In detail, this invention standardizes and automates the analysis process from data preprocessing, adaptive parameter optimization, density clustering execution, to multi-level result analysis and geochemical correlation, significantly improving analysis efficiency, result consistency, and repeatability, laying the foundation for large-scale exploration applications. Attached Figure Description
[0058] Figure 1 This is a schematic diagram of an implementation environment for the microbial grouping analysis method provided in this embodiment of the invention;
[0059] Figure 2 This is a schematic flowchart of a microbial grouping analysis method provided in an embodiment of the present invention;
[0060] Figure 3 This is a schematic diagram of the expansion process of the central logarithmic ratio transformation processing provided in an embodiment of the present invention;
[0061] Figure 4 This is a schematic diagram of the unfolding process of step S300 provided in an embodiment of the present invention;
[0062] Figure 5 This is a schematic diagram of the sample type expansion process for marking each sample according to an embodiment of the present invention;
[0063] Figure 6This is a schematic diagram of the unfolding process of step S500 provided in the embodiment of the present invention;
[0064] Figure 7 This is a schematic diagram of the unfolding process of step S600 provided in an embodiment of the present invention;
[0065] Figure 8 This is a schematic diagram of the overall process of the microbial grouping analysis method provided in the embodiments of the present invention;
[0066] Figure 9 This is a schematic diagram of the structure of a microbial grouping analysis device provided in an embodiment of the present invention;
[0067] Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0068] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of this invention; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this invention as detailed in the appended claims.
[0069] It is understood that the terms “first,” “second,” etc., used in this invention may be used herein to describe various concepts, but unless specifically stated otherwise, these concepts are not limited by these terms. These terms are used only to distinguish one concept from another. For example, first information may also be referred to as second information without departing from the scope of embodiments of the invention, and similarly, second information may also be referred to as first information. Depending on the context, the words “if,” “when,” or “in response to determination” as used herein may be interpreted as “when…” or “when…” or “in response to determination.”
[0070] The terms “at least one,” “multiple,” “each,” “any,” etc., used in this invention, “at least one” includes one, two, or more than two; “multiple” includes two or more than two; “each” refers to each of the corresponding multiple; and “any” refers to any one of the multiple.
[0071] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein is for the purpose of describing embodiments of the invention only and is not intended to limit the invention.
[0072] Among related technologies, the most commonly used method is a comparative analysis of microbial communities based on human experience and traditional statistics. However, existing technologies have many drawbacks: for example, manual grouping is highly subjective, has a weak ability to analyze complex nonlinear relationships, has biases in biomarker screening, and the process relies on manual intervention, making it difficult to standardize and promote.
[0073] In view of this, this invention provides a method and related equipment for microbial grouping analysis. This method acquires microbial data of a target study area; the microbial data includes microbial abundance data and methane leakage data from multiple samples; based on the microbial abundance data, a central logarithmic ratio transformation is performed to obtain the transformation value of each sample, and then Euclidean distance quantization is used to obtain the distance value between any two samples; a minimum sample size and neighborhood radius are adaptively set based on the number of samples and species corresponding to the microbial data; density clustering is performed on the samples in the microbial data to obtain clustering results; based on the clustering results, the sample type of each sample is labeled using the minimum sample size and neighborhood radius; multi-level analysis is performed on the clustering results based on the sample type to obtain analytical results of the clustering results; and correlation analysis is performed based on the analytical results and the methane leakage data to obtain the microbial grouping analysis results of the target study area. This invention automatically groups preprocessed microbial abundance data using density clustering, eliminating the manual pre-setting grouping steps that rely on prior knowledge in traditional methods. This ensures that the grouping results are entirely determined by the distribution characteristics of the data itself, effectively eliminating subjective bias. Specifically, this invention uses clustering based on the density relationships between samples, effectively discovering natural microbial ecotypes of arbitrary shapes and densities within the data. It is particularly suitable for capturing nonlinear distribution patterns and ecological transition zones formed along environmental gradients (such as methane concentration). In detail, this invention standardizes and automates the analysis process from data preprocessing, adaptive parameter optimization, density clustering execution, to multi-level result analysis and geochemical correlation, significantly improving analysis efficiency, result consistency, and repeatability, laying the foundation for large-scale exploration applications.
[0074] It is understood that the microbial grouping analysis method provided by this invention can be applied to any computer device with data processing and computing capabilities, and this computer device can be various terminals or servers. When the computer device in the embodiment is a server, the server is an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. Optionally, the terminal can be a smartphone, tablet computer, laptop computer, or desktop computer, but it is not limited to these.
[0075] like Figure 1 The diagram shown is a schematic representation of an implementation environment provided by an embodiment of the present invention. (Refer to...) Figure 1 The implementation environment includes at least one terminal 102 and a server 101. The terminal 102 and the server 101 can be connected via a network, either wirelessly or via a wired connection, to complete data transmission and exchange.
[0076] Server 101 can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0077] Additionally, server 101 can also be a node server in a blockchain network. Blockchain is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms.
[0078] Terminal 102 can be a smartphone, tablet computer, laptop computer, desktop computer, smart speaker, smartwatch, etc., but is not limited to these. Terminal 102 and server 101 can be directly or indirectly connected via wired or wireless communication, and this embodiment of the invention does not impose any limitations.
[0079] For example, based on Figure 1 The implementation environment shown in this embodiment of the invention provides a microbial grouping analysis method. The following description uses the application of this microbial grouping analysis method in server 101 as an example. It can be understood that this microbial grouping analysis method can also be applied in terminal 102.
[0080] Reference Figure 2 , Figure 2 This is an optional flowchart of the microbial grouping analysis method provided in the embodiments of the present invention. The subject executing the microbial grouping analysis method can be any of the aforementioned computer devices (including servers or terminals). Figure 2 The method may include, but is not limited to, steps S100 to S600.
[0081] Step S100: Obtain microbial data for the target study area;
[0082] The microbial data includes microbial abundance data and methane leakage data from multiple samples.
[0083] For example, in some specific implementations, firstly, the abundance data of microbial OTUs in the study area are collected, which can be in the form of a microbial OTU / ASV abundance table, where rows represent samples, columns represent microbial species (OTU or ASV), and values are abundance (which can be absolute abundance or relative abundance).
[0084] Step S200: Based on the microbial abundance data, the transformed value of each sample is obtained through central log ratio transformation, and then the distance between any samples is obtained by Euclidean distance quantization.
[0085] It should be noted that the microbial abundance data includes the raw abundance value of each species in the sample. In some embodiments, such as... Figure 3 As shown, based on microbial abundance data, the transformed value of each sample is obtained through central logarithmic ratio transformation, which may include the following steps: S210, performing a cumulative multiplication operation on the original abundance values of all species in a single sample to obtain the cumulative multiplication result; S220, using the number of species as the square root, performing a square root operation on the cumulative multiplication result to obtain the geometric mean; S230, performing a logarithmic transformation on the ratio of the original abundance value to the geometric mean to obtain the transformed value of each species in each sample.
[0086] In some alternative implementations, data cleaning and quality control can also be performed: species that are zero in all samples (i.e., these species do not exist in the dataset) are removed. Species with low abundance or low variance are filtered based on certain criteria (such as frequency of occurrence using a preset threshold) to reduce noise.
[0087] For example, in some specific implementations, the central logarithmic ratio transformation (CLR) can be implemented as follows:
[0088] Since microbial abundance data is component data (i.e., the sum of abundance for each sample is 1 or a constant), CLR transformation is used to eliminate closure effects. For each sample's microbial abundance vector... (m is the number of species) Let x represent the original abundance value of the i-th species (i∈[1,m]) in the sample, and calculate the geometric mean of sample x. Then, a logarithmic transformation was performed on the abundance of each species: ,in This represents the value of the i-th species after CLR transformation.
[0089] In some alternative implementations, if the data contains zero values, pseudo-counting is required first, multiplying the minimum non-zero value of the sample by a coefficient. Replace the zero value with 0.5.
[0090] Zero-value handling (pseudo-counting method): ,in This represents the abundance of the j-th species in the i-th sample. This represents the pseudo-counting coefficient (taken as 0.5). The abundance of the j-th species in the i-th sample after zero-value processing, min( ) represents the minimum value of all non-zero abundance values in the sample.
[0091] It should be noted that in some embodiments, Euclidean distance quantization is used to obtain the distance value between any samples, including the following steps: calculating the Euclidean distance matrix based on the transformation value of each species in the sample to obtain the distance value between any samples;
[0092] For example, in some specific implementations, the Euclidean distance matrix can be calculated as follows:
[0093] After CLR transformation, the data is in Euclidean space and can be measured using Euclidean distance.
[0094] Calculate the distance matrix between samples ,in ,in The Euclidean distance between the i-th sample and the j-th sample This represents the abundance of the k-th species in the i-th sample. This represents the abundance of the k-th species in the j-th sample, where k represents the k-th species. This represents the CLR transformation value of the i-th sample and the k-th species.
[0095] Step S300: Adaptively set the minimum number of samples and neighborhood radius based on the number of samples and species corresponding to the microbial data;
[0096] It should be noted that in some embodiments, such as Figure 4As shown, step S300 may include the following steps: S310, when the number of species is less than or equal to the first threshold and the number of species at a preset multiple does not exceed a preset proportion of the number of samples, the minimum number of samples is set to the number of species at a preset multiple; otherwise, the minimum number of samples is set based on a preset fixed value; S320, the k value is determined based on the minimum number of samples; S330, a distance sorting line graph is drawn using the k-distance graph method based on the k value and the distance values between all samples; S340, the neighborhood radius is determined based on the inflection point of the distance sorting line graph.
[0097] For example, in some specific implementations, the intelligent initialization of the min_samples parameter can be achieved as follows: Set a default minimum value, min_samples = 3. If the number of species is <= 50 and 2 * number of species does not exceed 1 / 3 of the total sample size, then min_samples = 2 * number of species; otherwise, use a fixed value (min_samples = 3).
[0098] For example, in some specific implementations, determining the eps parameter (neighborhood radius) can be achieved as follows: using a k-distance graph method: for each sample, calculate its distance to all other samples and sort them in ascending order. Then, for each sample, take the distance of the k-th nearest neighbor (k is min_samples-1); sort the k-th distance values of all samples and plot them as a line graph, find the inflection point (i.e., the point where the curve suddenly rises), and the distance value corresponding to the inflection point is the recommended value of eps.
[0099] In some alternative implementations, parameter optimization can also be performed through parameter sensitivity analysis and stability verification, which can be achieved as follows:
[0100] After initial parameter setting through the aforementioned steps, parameter sensitivity analysis and stability verification are performed. This involves a grid search within a small range around the recommended parameters (eps, min_samples), systematically fine-tuning the parameters and observing changes in the clustering results. The silhouette coefficient is used as the core clustering quality evaluation index. This coefficient is calculated by determining the average distance between each sample and other samples in the same cluster (a(i), intra-cluster cohesion) and the average distance to the nearest other cluster sample (b(i), inter-cluster separation). The closer the average silhouette coefficient of all samples is to 1, the better the clustering effect. By analyzing the stability of the silhouette coefficient and the number of clusters under different parameter combinations, a set of optimal parameters that are robust and insensitive to small parameter changes is finally determined.
[0101] Silhouette coefficient: For each sample, calculate its average distance a(i) to other samples in the same cluster and its average distance b(i) to samples in other clusters. Then the silhouette coefficient s(i) = (b(i) - a(i)) / max(a(i), b(i)).
[0102] Step S400: Density clustering is performed on the samples in the microbial data to obtain clustering results. Based on the clustering results, the sample type of each sample is labeled by combining the minimum number of samples and the neighborhood radius.
[0103] It should be noted that the clustering results include multiple sample clusters, and the sample types include core samples, boundary samples, and noise samples. In some embodiments, such as... Figure 5 As shown, based on the clustering results, the sample type of each sample is labeled by combining the minimum number of samples and the neighborhood radius, which may include the following steps: S410, when there is at least one sample with the minimum number of samples in the neighborhood radius of the first sample, the first sample is determined to be a core sample; S420, when the second sample is in the neighborhood radius of the core sample and the second sample does not meet the conditions of the core sample, the second sample is determined to be a boundary sample; S430, when the third sample does not belong to any sample cluster, the third sample is determined to be a noise sample.
[0104] For example, in some specific implementations, after obtaining the optimized parameters, the DBSCAN clustering algorithm is executed, which divides the samples according to the natural density distribution of the data and automatically identifies three types of samples.
[0105] Core samples are points located in high-density regions within a cluster. Each core sample contains at least `min_samples` other samples within its `eps` neighborhood. Biologically, these core samples represent a stable, typical microbial ecotype; they constitute the "core" population of that ecotype, exhibiting highly similar and concentrated characteristics.
[0106] Boundary samples are also members of a cluster, but they do not meet the criteria for core samples (i.e., they have insufficient points in their neighborhood). They are included in a cluster only because they are located in the neighborhood of a core sample. These samples are located on the edge of the cluster and can be biologically interpreted as belonging to a certain ecotype, but their characteristics are not typical. They may be in a transitional or variant state of that ecotype, or have been affected by other factors.
[0107] Noise samples are sample points that do not belong to any cluster; they are neither core points nor reachable by any core point density. In the context of microbial ecology, these noise points may be genuine outliers, such as samples from special or extreme ecological niches; they may also be samples in transitional zones between different ecotypes that cannot be clearly classified into any known category; in addition, measurement or technical errors may also cause samples to be identified as noise.
[0108] Step S500: Perform multi-level analysis on the clustering results based on the sample type to obtain the analysis results of the clustering results;
[0109] It should be noted that the sample types include core samples, boundary samples, and noise samples. The clustering result includes multiple sample clusters, each containing both core and boundary samples while excluding noise samples. In some embodiments, such as... Figure 6 As shown, step S500 may include the following steps: S510, performing a core ecotype depth analysis on each sample cluster to obtain the key species driving ecotype differentiation; wherein, the core ecotype depth analysis includes differential abundance analysis methods and standard methods for determining significantly enriched species; S520, constructing a similarity matrix between the boundary sample and each sample cluster based on the Brectis distance between each boundary sample and the core sample of each sample cluster; S530, determining the distribution position of the boundary sample among the sample clusters through principal coordinate analysis based on the similarity matrix, thereby identifying transitional state samples; S540, performing transitional state verification on the transitional state samples based on environmental gradients to obtain the analysis results of the transitional state samples; S550, identifying technical errors based on the sequencing depth, sequence quality indicators, and DNA extraction quality of noise samples, thereby identifying low biomass samples after excluding technical anomalies; S560, performing ecological specificity analysis and biological significance assessment on noise samples to obtain the results of noise sample ecological significance mining.
[0110] For example, in some specific implementations, multi-level result parsing can be achieved as follows:
[0111] (1) In-depth analysis of core ecotypes: Differential abundance analysis was performed on each identified cluster (excluding noise points) to reveal the key species driving ecotype differentiation. The following statistical methods were used:
[0112] Differential abundance analysis methods:
[0113] Multiple group comparisons were performed using the nonparametric statistical test Kruskal-Wallis test, and pairwise comparisons were further performed using the Mann-Whitney U test for species with significant differences. Considering the sparsity and zero-expansion characteristics of microbial data, a statistical method specifically designed for microbiome data, the ALDEx2 method, was adopted to better handle component data and the zero-value problem.
[0114] Criteria for identifying significantly enriched species:
[0115] False positive rates were controlled using Benjamini-Hochberg FDR correction through multiple tests. Species with a corrected p-value less than 0.05 and an effect size (Fold Change greater than 2) were defined as significantly enriched marker species in the cluster. The LDA (Linear Discriminant Analysis Effect Size) value of each marker species was also calculated to assess its contribution to cluster differentiation.
[0116] (2) Analysis of transition state samples:
[0117] Boundary samples represent the transitional state of microbial community succession or environmental gradient changes, and have important ecological significance.
[0118] Boundary sample feature analysis methods:
[0119] Calculate the Bray-Curtis distance between each boundary sample and the centroid of each cluster's core sample to construct a similarity matrix between the boundary samples and each cluster. Visualize the distribution of boundary samples among clusters using principal coordinate analysis, and identify samples that maintain high similarity with multiple clusters as typical transitional state samples.
[0120] Transition state verification:
[0121] We examined whether these transitional samples were in the middle of the environmental gradient (methane concentration, depth, temperature, etc.) and confirmed through community composition analysis whether they actually contained the iconic species combination of the source cluster and the destination cluster, thereby verifying their possibility as an intermediate state of ecological succession.
[0122] (3) Mining the ecological significance of noise samples:
[0123] Noise samples may represent rare ecotypes, technological outliers, or special habitat samples, requiring systematic evaluation.
[0124] Technical error identification: First, check the sequencing depth, sequence quality indicators, and DNA extraction quality of the noise samples to rule out anomalies caused by technical factors. Calculate the community rarity index and phylogenetic diversity of each noise sample to identify truly low biomass samples.
[0125] Ecological specificity analysis: Integrating environmental metadata, this study analyzes whether noise samples cluster under specific environmental conditions (such as extreme temperatures, unique geological features, etc.). Through indicator species analysis, it identifies rare species that uniquely appear in the noise samples, suggesting that these species may be adapted to specific ecological niches.
[0126] Biological significance assessment: Focus on samples that are consistently identified as noise under different parameter settings, and examine whether these samples occupy a special position in the microbial co-occurrence network (such as module connection point or isolated node) through network analysis, so as to determine whether they represent bridge samples connecting different ecotypes or truly unique ecological niches.
[0127] Step S600: Based on the analysis results and methane leakage data, a correlation analysis is performed to obtain the microbial grouping analysis results of the target study area;
[0128] It should be noted that the clustering results include multiple sample clusters. In some embodiments, such as... Figure 7 As shown, step S600 may include the following steps: S610, spatially correlate the clustering results with geochemical data, and then use statistical methods to test the differences in geochemical parameters among different sample clusters; S620, based on the first average distance from the target sample to the same sample cluster and the second average distance from the target sample to other sample clusters, quantify the silhouette coefficient of the target sample as a clustering quality assessment index; S630, based on the clustering quality assessment index and the analysis results, compile a standardized analysis report as the microbial grouping analysis results of the target study area.
[0129] For example, in some specific implementations, the final microbial grouping analysis results can be obtained by grouping and analyzing the correlation strength between the clustering results (cluster labels) and methane leakage, and combining the analysis results:
[0130] Spatial correlation analysis was performed between the clustering results and geochemical data such as δ13C and methane concentration. A standardized analysis report was automatically generated, including: clustering quality assessment indicators, descriptions of microbial characteristics of each ecotype, and correlation strength analysis with methane leakage.
[0131] First, the clustering results (cluster labels) were correlated with geochemical data (δ13C, methane concentration). A statistical method (PERMANOVA) was used to test whether there were significant differences in geochemical parameters among different clusters. Then, based on the research background, a methane leakage sample was defined as one whose methane concentration exceeded a certain threshold (e.g., twice the background concentration). The proportion of methane leakage samples was then calculated for each cluster, and the proportions of methane leakage samples among different clusters were compared to see if there were significant differences.
[0132] Finally, a standardized analysis report is automatically generated. The report includes the silhouette coefficient (mean silhouette coefficient and silhouette coefficient of each cluster) as a cluster quality assessment indicator, the description of the microbial characteristics of each ecotype, the species (biomarkers) that are significantly enriched in each cluster (ecotype), obtained through differential abundance analysis (ALDEx2), including species name, corrected p-value, effect size (LDA score), geochemical data association analysis results: PERMANOVA results, methane leakage sample proportion and methane concentration descriptive statistics of each cluster, visualization of cluster results (PCA plot, colored with cluster labels), visualization of geochemical parameters and cluster results (box plot showing methane concentration of each cluster), and heat map of significantly enriched species.
[0133] Silhouette coefficient: For each sample, calculate its average distance to other samples in the same cluster, a(i), and its average distance to samples in other clusters, b(i). The silhouette coefficient s(i) = (b(i) - a(i)) / max(a(i), b(i)). The silhouette coefficient is used as the core indicator for evaluating clustering quality. This coefficient is calculated by calculating the average distance of each sample to other samples in the same cluster (a(i), intra-cluster cohesion) and the average distance to the nearest other cluster sample (b(i), inter-cluster separation). The closer the average silhouette coefficient of all samples is to 1, the better the clustering effect.
[0134] After the above steps, the correlation between microbial community clustering results and methane leakage can be comprehensively evaluated, thereby understanding the indicative role of different microbial ecotypes in the natural gas hydrate environment and realizing dynamic grouping of microbial communities during the development of natural gas hydrates.
[0135] To explain in detail the principle of the technical solution of the present invention, the overall process of the present invention will be described below with reference to some specific embodiments. It is easy to understand that the following is an explanation of the technical principle of the present invention and should not be regarded as a limitation of the present invention.
[0136] First, it should be noted that, addressing the core shortcomings of manual analysis methods in the clustering and grouping stage, the key technical problem this invention aims to solve is: how to establish an objective and automated method for natural grouping of microbial communities, capable of automatically identifying the natural grouping boundaries and quantities of microbial communities in a pre-defined natural gas hydrate survey area without human intervention; effectively discovering microbial ecotypes of methane leakage of arbitrary shapes and densities; clearly distinguishing between core ecotypes and transitional state samples, preserving complete ecological gradient information; and providing repeatable and standardized grouping results, reducing the influence of human subjectivity in traditional methods.
[0137] In some specific application scenarios, such as Figure 8 As shown, embodiments of the present invention can be implemented through the following process steps:
[0138] Microbial OTU abundance data were collected from a pre-defined natural gas hydrate survey area. A pseudo-counting method was used to handle zero values in the microbial abundance data, avoiding mathematical errors during logarithmic transformation and minimizing disturbance to the original data structure. A central logarithmic ratio transformation technique was applied to effectively address the inherent "closure effect" problem in microbial abundance data, converting the component data into coordinates in Euclidean space, laying the foundation for subsequent distance calculations. The transformed data was standardized to eliminate the impact of differences in abundance levels among different microbial species on the clustering results. The Euclidean distance matrix between samples was calculated, providing a suitable metric for density clustering.
[0139] Step 1: Collect and preprocess microbial OTU abundance data of the study area:
[0140] Input data: Microbial OTU / ASV abundance table, where rows represent samples, columns represent microbial species (OTU or ASV), and values are abundance (which can be absolute abundance or relative abundance).
[0141] Data cleaning and quality control:
[0142] Remove species that are zero in all samples (i.e., these species do not exist in the dataset). Filter species with low abundance or low variance based on certain criteria (such as frequency of occurrence) to reduce noise.
[0143] Central Logarithmic Ratio Transformation (CLR):
[0144] Since microbial abundance data is component data (i.e., the sum of abundance for each sample is 1 or a constant), CLR transformation is used to eliminate closure effects. For each sample's microbial abundance vector... (m is the number of species) Let x represent the original abundance value of the i-th species (i∈[1,m]) in the sample, and calculate the geometric mean of sample x. Then, a logarithmic transformation was performed on the abundance of each species: ,in This represents the value of the i-th species after CLR transformation.
[0145] If the data contains zero values, pseudo-counting is required first, which involves multiplying the minimum non-zero value of the sample by a coefficient. Replace the zero value with 0.5.
[0146] Zero-value handling (pseudo-counting method): ,in This represents the abundance of the j-th species in the i-th sample. This represents the pseudo-counting coefficient (taken as 0.5). The abundance of the j-th species in the i-th sample after zero-value processing, min( ) represents the minimum value of all non-zero abundance values in the sample.
[0147] Calculate the Euclidean distance matrix:
[0148] After CLR transformation, the data is in Euclidean space and can be measured using Euclidean distance.
[0149] Calculate the distance matrix between samples ,in ,in The Euclidean distance between the i-th sample and the j-th sample This represents the abundance of the k-th species in the i-th sample. This represents the abundance of the k-th species in the j-th sample, where k represents the k-th species. This represents the CLR transformation value of the i-th sample and the k-th species.
[0150] Step 2: Adaptive parameter optimization based on the distribution characteristics of microbial data
[0151] Based on the total number of samples and data types, the `min_samples` parameter is intelligently set, and the core parameter `eps` (neighborhood radius) of DBSCAN is automatically determined using the k-distance graph method. Parameter sensitivity analysis is implemented to ensure the stability of the clustering results.
[0152] The specific steps are as follows: initialize min_samples based on data characteristics → automatically identify the optimal eps through k-distance plot → run DBSCAN preliminary clustering with initial parameters → calculate quality indicators such as silhouette coefficient for quality assessment → iteratively optimize parameters based on assessment results → sensitivity analysis → stability verification.
[0153] (1) Intelligent initialization of the min_samples parameter:
[0154] Microbial data contains a large number of species (features) and a relatively small number of samples (e.g., 100 samples, 1000 species). Therefore, directly using twice the data dimension (number of species) may result in very large min_samples, which is unsuitable for high-dimensional data because points in high-dimensional space are usually sparse and difficult to form dense regions.
[0155] Therefore, the following strategy is adopted: If the data dimensionality (number of species) is not particularly large (e.g., less than 50), use twice the data dimensionality as min_samples. If the data dimensionality is large, instead of using twice the data dimensionality, use a smaller fixed value (min_samples=3), because distance calculations in high-dimensional spaces are affected by the curse of dimensionality, making it difficult to achieve the required point density. Furthermore, considering the total sample size, if the total sample size is small, such as only 10 species, then even with a small number of species (e.g., 10), 2*10=20 will still exceed the total sample size. Therefore, it is also necessary to ensure that min_samples does not exceed the total sample size.
[0156] Therefore, the intelligent setting process is as follows:
[0157] Set a default minimum value: min_samples = 3.
[0158] If the number of species is less than or equal to 50 and 2 * number of species does not exceed 1 / 3 of the total sample size, then min_samples = 2 * number of species.
[0159] Otherwise, use a fixed value (min_samples = 3).
[0160] (2) Determine the eps parameter (neighborhood radius):
[0161] Using the k-distance graph method: For each sample, calculate its distance to all other samples and sort them in ascending order. Then, for each sample, take the distance to its k-th nearest neighbor (k is min_samples-1).
[0162] Sort the k-th distance values of all samples and plot them as a line graph, then find the inflection point (i.e., the point where the curve suddenly rises). The distance value corresponding to the inflection point is the recommended value for eps.
[0163] (3) Parameter sensitivity analysis and stability verification:
[0164] After initial parameter setting, parameter sensitivity analysis and stability verification were performed. This involved a grid search within a small range around the recommended parameters (eps, min_samples), systematically fine-tuning the parameters and observing changes in the clustering results. The silhouette coefficient was used as the core clustering quality evaluation index. This coefficient is calculated by calculating the average distance between each sample and other samples in the same cluster (a(i), intra-cluster cohesion) and the average distance to the nearest other cluster sample (b(i), inter-cluster separation). The closer the average silhouette coefficient of all samples is to 1, the better the clustering effect. By analyzing the stability of the silhouette coefficient and the number of clusters under different parameter combinations, a set of optimal parameters that are robust and insensitive to small parameter changes was finally determined.
[0165] Silhouette coefficient: For each sample, calculate its average distance a(i) to other samples in the same cluster and its average distance b(i) to samples in other clusters. Then the silhouette coefficient s(i) = (b(i) - a(i)) / max(a(i), b(i)).
[0166] Step 3: DBSCAN density clustering execution:
[0167] After obtaining the optimized parameters, the DBSCAN clustering algorithm is executed. This algorithm divides the samples according to the natural density distribution of the data and automatically identifies three types of samples.
[0168] Core samples are points located in high-density regions within a cluster. Each core sample contains at least `min_samples` other samples within its `eps` neighborhood. Biologically, these core samples represent a stable, typical microbial ecotype; they constitute the "core" population of that ecotype, exhibiting highly similar and concentrated characteristics.
[0169] Boundary samples are also members of a cluster, but they do not meet the criteria for core samples (i.e., they have insufficient points in their neighborhood). They are included in a cluster only because they are located in the neighborhood of a core sample. These samples are located on the edge of the cluster and can be biologically interpreted as belonging to a certain ecotype, but their characteristics are not typical. They may be in a transitional or variant state of that ecotype, or have been affected by other factors.
[0170] Noise samples are sample points that do not belong to any cluster; they are neither core points nor reachable by any core point density. In the context of microbial ecology, these noise points may be genuine outliers, such as samples from special or extreme ecological niches; they may also be samples in transitional zones between different ecotypes that cannot be clearly classified into any known category; in addition, measurement or technical errors may also cause samples to be identified as noise.
[0171] Step 4: Multi-level result analysis:
[0172] (1) In-depth analysis of core ecotypes: Differential abundance analysis was performed on each identified cluster (excluding noise points) to reveal the key species driving ecotype differentiation. The following statistical methods were used:
[0173] Differential abundance analysis methods:
[0174] Multiple group comparisons were performed using the nonparametric statistical test Kruskal-Wallis test, and pairwise comparisons were further performed using the Mann-Whitney U test for species with significant differences. Considering the sparsity and zero-expansion characteristics of microbial data, a statistical method specifically designed for microbiome data, the ALDEx2 method, was adopted to better handle component data and the zero-value problem.
[0175] Criteria for identifying significantly enriched species:
[0176] False positive rates were controlled using Benjamini-Hochberg FDR correction through multiple tests. Species with a corrected p-value less than 0.05 and an effect size (Fold Change greater than 2) were defined as significantly enriched marker species in the cluster. The LDA (Linear Discriminant Analysis Effect Size) value of each marker species was also calculated to assess its contribution to cluster differentiation.
[0177] (2) Analysis of transition state samples:
[0178] Boundary samples represent the transitional state of microbial community succession or environmental gradient changes, and have important ecological significance.
[0179] Boundary sample feature analysis methods:
[0180] Calculate the Bray-Curtis distance between each boundary sample and the centroid of each cluster's core sample to construct a similarity matrix between the boundary samples and each cluster. Visualize the distribution of boundary samples among clusters using principal coordinate analysis, and identify samples that maintain high similarity with multiple clusters as typical transitional state samples.
[0181] Transition state verification:
[0182] We examined whether these transitional samples were in the middle of the environmental gradient (methane concentration, depth, temperature, etc.) and confirmed through community composition analysis whether they actually contained the iconic species combination of the source cluster and the destination cluster, thereby verifying their possibility as an intermediate state of ecological succession.
[0183] (3) Mining the ecological significance of noise samples:
[0184] Noise samples may represent rare ecotypes, technological outliers, or special habitat samples, requiring systematic evaluation.
[0185] Technical error identification: First, check the sequencing depth, sequence quality indicators, and DNA extraction quality of the noise samples to rule out anomalies caused by technical factors. Calculate the community rarity index and phylogenetic diversity of each noise sample to identify truly low biomass samples.
[0186] Ecological specificity analysis: Integrating environmental metadata, this study analyzes whether noise samples cluster under specific environmental conditions (such as extreme temperatures, unique geological features, etc.). Through indicator species analysis, it identifies rare species that uniquely appear in the noise samples, suggesting that these species may be adapted to specific ecological niches.
[0187] Biological significance assessment: Focus on samples that are consistently identified as noise under different parameter settings, and examine whether these samples occupy a special position in the microbial co-occurrence network (such as module connection point or isolated node) through network analysis, so as to determine whether they represent bridge samples connecting different ecotypes or truly unique ecological niches.
[0188] Step 5: Perform grouped analysis on the correlation strength between clustering results (cluster labels) and methane leakage:
[0189] Spatial correlation analysis was performed between the clustering results and geochemical data such as δ13C and methane concentration. A standardized analysis report was automatically generated, including: clustering quality assessment indicators, descriptions of microbial characteristics of each ecotype, and correlation strength analysis with methane leakage.
[0190] First, the clustering results (cluster labels) were correlated with geochemical data (δ13C, methane concentration). A statistical method (PERMANOVA) was used to test whether there were significant differences in geochemical parameters among different clusters. Then, based on the research background, a methane leakage sample was defined as one whose methane concentration exceeded a certain threshold (e.g., twice the background concentration). The proportion of methane leakage samples was then calculated for each cluster, and the proportions of methane leakage samples among different clusters were compared to see if there were significant differences.
[0191] Finally, a standardized analysis report is automatically generated. The report includes the silhouette coefficient (mean silhouette coefficient and silhouette coefficient of each cluster) as a cluster quality assessment indicator, the description of the microbial characteristics of each ecotype, the species (biomarkers) that are significantly enriched in each cluster (ecotype), obtained through differential abundance analysis (ALDEx2), including species name, corrected p-value, effect size (LDA score), geochemical data association analysis results: PERMANOVA results, methane leakage sample proportion and methane concentration descriptive statistics of each cluster, visualization of cluster results (PCA plot, colored with cluster labels), visualization of geochemical parameters and cluster results (box plot showing methane concentration of each cluster), and heat map of significantly enriched species.
[0192] Silhouette coefficient: For each sample, calculate its average distance to other samples in the same cluster, a(i), and its average distance to samples in other clusters, b(i). The silhouette coefficient s(i) = (b(i) - a(i)) / max(a(i), b(i)). The silhouette coefficient is used as the core indicator for evaluating clustering quality. This coefficient is calculated by calculating the average distance of each sample to other samples in the same cluster (a(i), intra-cluster cohesion) and the average distance to the nearest other cluster sample (b(i), inter-cluster separation). The closer the average silhouette coefficient of all samples is to 1, the better the clustering effect.
[0193] After the above steps, the correlation between microbial community clustering results and methane leakage can be comprehensively evaluated, thereby understanding the indicative role of different microbial ecotypes in the natural gas hydrate environment and realizing dynamic grouping of microbial communities during the development of natural gas hydrates.
[0194] In summary, this invention achieves fully data-driven identification of microbial ecotypes through a density clustering algorithm, solving the subjective bias problem introduced by human intervention in traditional methods. Unlike traditional hierarchical clustering, which requires researchers to subjectively determine the cutting positions of the dendrogram, the technical solution of this invention automatically determines the number and boundaries of ecotypes based on the inherent density distribution characteristics of microbial community composition, ensuring the objectivity and repeatability of the grouping results.
[0195] In applications involving pre-defined methane leakage zones, the technical solution of this invention demonstrates a superior ability to identify complex ecological patterns. Compared to traditional methods that are limited to identifying approximately spherical cluster structures, this invention can accurately discover microbial ecotypes of arbitrary shapes and different distribution densities, effectively capturing the nonlinear distribution patterns and ecological transition zones of microbial communities along the methane concentration gradient in this region, providing a more realistic characterization for understanding the microbial ecological processes in natural gas hydrate zones.
[0196] This invention enables refined analysis of microbial community states, clearly distinguishing between core ecotypes (representing stable community states), boundary samples (reflecting transitional states in community succession), and noise samples (potentially indicating specific niches or new technology signals). This classification system overcomes the limitations of traditional methods that forcibly categorize all samples, fully preserving the state diversity information of microbial communities and providing a richer data foundation for identifying characteristic microbial combinations under different methane leakage intensities.
[0197] At the technical implementation level, this invention establishes a fully automated system from parameter optimization to result analysis. By automatically determining algorithm parameters through objective methods such as k-distance maps and combining them with standardized result output formats, it ensures high consistency and repeatability in sample analysis at different stations and depths, greatly enhancing the practical application value of this method in natural gas hydrate exploration.
[0198] Most importantly, by analyzing the spatial distribution patterns of core ecotypes, boundary states, and noise samples, the technical solution of this invention can infer the aggregation mechanism and ecological succession trajectory of microbial communities, breaking through the limitation of traditional methods that can only provide static grouping results, and providing important technical support for predicting the response dynamics of microbial communities during natural gas hydrate development.
[0199] Compared with the prior art, the technical solution of the present invention achieves at least the following beneficial technical means:
[0200] DBSCAN parameter optimization technology for methane leakage characteristics: Based on the high heterogeneity and complexity of microbial community data in a pre-defined marine area, a dedicated parameter determination method and verification process were developed. By analyzing the specific distribution patterns of microbial communities at methane leakage points in the region, a method combining k-distance maps and silhouette coefficients was used to adaptively determine the optimal neighborhood radius (ε) and minimum sample size (min_samples). This ensures that the clustering results can accurately identify various microbial ecotypes from active leakage areas to background areas, providing reliable technical parameter support for natural gas hydrate exploration.
[0201] A multi-level clustering analysis framework for methane leakage environments is proposed, establishing a differentiated analysis strategy for core, boundary, and noise samples applicable to methane leakage areas. Core samples correspond to a stable symbiotic combination of anaerobic methane-oxidizing archaea (ANME) and sulfate-reducing bacteria (SRB), representing a typical leakage ecosystem. Boundary samples reflect the transitional states between different leakage intensities, revealing the environmental adaptive evolution of the microbial community. Noise samples may indicate specific micro-leakage environments or new functional groups. This framework fully explores the indicative significance of various sample types in natural gas hydrate exploration.
[0202] An automated clustering quality assessment system integrating geological significance: Developing internal validation indicators specifically for density clustering results, focusing on the geological rationality and exploration application value of the clustering results. By calculating indicators such as the intra-cluster microbial functional consistency index and the inter-cluster environmental gradient response difference, combined with the known distribution patterns of seepage points, the system ensures that the grouping results not only meet statistical requirements but also possess clear geochemical indicative significance and exploration practical value.
[0203] like Figure 9 As shown, this embodiment of the invention also provides a microbial grouping analysis device 900, which can implement the above-described method. This device may include:
[0204] The first module 910 is used to acquire microbial data of the target study area; the microbial data includes microbial abundance data and methane leakage data of multiple samples;
[0205] The second module 920 is used to obtain the transformed value of each sample based on the microbial abundance data through the central log ratio transformation, and then use Euclidean distance quantization to obtain the distance value between any samples.
[0206] The third module 930 is used to adaptively set the minimum number of samples and the neighborhood radius based on the number of samples and species corresponding to the microbial data;
[0207] The fourth module 940 is used to perform density clustering on samples in microbial data, obtain clustering results, and label the sample type of each sample based on the clustering results, combined with the minimum number of samples and the neighborhood radius.
[0208] The fifth module 950 is used to perform multi-level analysis of the clustering results based on sample type to obtain the analysis results of the clustering results;
[0209] Module 6, 960, is used to perform correlation analysis based on the analysis results and methane leakage data to obtain the microbial grouping analysis results of the target study area.
[0210] It is understood that the content of the above method embodiments is applicable to the present device embodiments. The specific functions implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0211] This invention also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described above. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0212] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0213] like Figure 10 As shown, Figure 10 The hardware structure of an electronic device 1000 according to another embodiment is illustrated. The electronic device 1000 includes:
[0214] The processor 1001 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (aSIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present invention.
[0215] The memory 1002 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RaM). The memory 1002 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1002 and is called and executed by the processor 1001.
[0216] Input / output interface 1003 is used to implement information input and output;
[0217] The communication interface 1004 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0218] Bus 1005 transmits information between various components of the device (e.g., processor 1001, memory 1002, input / output interface 1003, and communication interface 1004);
[0219] The processor 1001, memory 1002, input / output interface 1003 and communication interface 1004 are connected to each other within the device via bus 1005.
[0220] The electronic device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0221] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.
[0222] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0223] This invention also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0224] It is understood that the content of the above method embodiments is applicable to the embodiments of this program product. The specific functions implemented by the embodiments of this program product are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0225] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0226] The microbial grouping analysis method, apparatus, electronic device, storage medium, and program product provided in this invention acquire microbial data of a target study area. The microbial data includes microbial abundance data and methane leakage data from multiple samples. Based on the microbial abundance data, a central logarithmic ratio transformation is performed to obtain a transformed value for each sample, and then Euclidean distance quantization is used to obtain the distance between any two samples. A minimum sample size and neighborhood radius are adaptively set based on the number of samples and species corresponding to the microbial data. Density clustering is performed on the samples in the microbial data to obtain clustering results. Based on the clustering results, the sample type of each sample is labeled using the minimum sample size and neighborhood radius. Multi-level analysis is performed on the clustering results based on the sample type to obtain analytical results. Correlation analysis is performed based on the analytical results and the methane leakage data to obtain the microbial grouping analysis results for the target study area. This invention automatically groups preprocessed microbial abundance data using density clustering, eliminating the manual pre-setting grouping steps that rely on prior knowledge in traditional methods. This ensures that the grouping results are entirely determined by the distribution characteristics of the data itself, effectively eliminating subjective bias. Specifically, this invention uses clustering based on the density relationships between samples, effectively discovering natural microbial ecotypes of arbitrary shapes and densities within the data. It is particularly suitable for capturing nonlinear distribution patterns and ecological transition zones formed along environmental gradients (such as methane concentration). In detail, this invention standardizes and automates the analysis process from data preprocessing, adaptive parameter optimization, density clustering execution, to multi-level result analysis and geochemical correlation, significantly improving analysis efficiency, result consistency, and repeatability, laying the foundation for large-scale exploration applications.
[0227] The embodiments described in this invention are for the purpose of more clearly illustrating the technical solutions of the embodiments of this invention, and do not constitute a limitation on the technical solutions provided by the embodiments of this invention. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this invention are also applicable to similar technical problems.
[0228] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present invention, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0229] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0230] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0231] The preferred embodiments of the present invention have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and spirit of the present invention should be within the scope of the claims of the present invention.
Claims
1. A method of analyzing grouping of microorganisms, characterized by, The method includes the following steps: Acquire microbial data for the target study area; wherein, the microbial data includes microbial abundance data and methane leakage data from multiple samples; Based on the microbial abundance data, the transformed value of each sample is obtained through central log ratio transformation, and then the distance between any samples is obtained by Euclidean distance quantization. The microbial abundance data includes the original abundance value of each species in the sample. The process of obtaining the transformed value of each sample based on the microbial abundance data through central logarithmic ratio transformation includes the following steps: The original abundance values of all species in a single sample are multiplied together to obtain the cumulative result; Using the number of species as the root of the summation result, the geometric mean is obtained by performing a root operation on the summation result. Logarithmically transform the ratio of the original abundance value to the geometric mean to obtain the transformed value for each species in each sample; The minimum number of samples and the neighborhood radius are adaptively set based on the number of samples and species corresponding to the microbial data; The step of adaptively setting the minimum sample number and neighborhood radius based on the number of samples and species corresponding to the microbial data includes the following steps: If the number of species is less than or equal to a first threshold and the number of species at a preset multiple does not exceed a preset proportion of the number of samples, the minimum number of samples is set to the number of species at the preset multiple; otherwise, the minimum number of samples is set based on a preset fixed value. The value of k is determined based on the minimum number of samples. Based on the k value and the distance values between all the samples, a distance-ranked line chart is drawn using the k-distance graph method; The neighborhood radius is determined based on the inflection points of the distance sorted line graph; Density clustering is performed on the samples in the microbial data to obtain clustering results. Based on the clustering results, the sample type of each sample is labeled by combining the minimum number of samples and the neighborhood radius. The clustering results are analyzed at multiple levels based on the sample type to obtain the analysis results of the clustering results; Based on the analysis results and the methane leakage data, a correlation analysis was performed to obtain the microbial grouping analysis results of the target study area.
2. The method of claim 1, wherein, The method of obtaining the distance value between any two samples using Euclidean distance quantization includes the following steps: Calculate the Euclidean distance matrix based on the transformation value of each species in the sample to obtain the distance value between any of the samples; The expression for the distance value is: ; In the formula, represents the distance value of the ith sample and the jth sample; m represents the number of species; k represents the kth species; represents the conversion value of the kth species in the ith sample; represents the conversion value of the kth species in the jth sample.
3. The method of claim 1, wherein, The clustering result includes multiple sample clusters, and the sample types include core samples, boundary samples, and noise samples. The step of labeling the sample type of each sample based on the clustering result, combined with the minimum number of samples and the neighborhood radius, includes the following steps: If at least one sample of the minimum sample number exists within the neighborhood radius of the first sample, the first sample is determined to be the core sample. When the second sample is within the neighborhood radius of the core sample and the second sample does not meet the conditions of the core sample, the second sample is determined to be the boundary sample; When the third sample does not belong to any of the sample clusters, the third sample is determined as the noise sample.
4. The method of claim 1, wherein, The sample types include core samples, boundary samples and noise samples, the clustering results include a plurality of sample clusters, the sample clusters contain the core samples and the boundary samples and exclude the noise samples, the clustering results are analyzed based on the sample types to obtain analysis results of the clustering results, including the following steps: Performing core ecological depth analysis on each of the sample clusters to obtain key species driving ecological type differentiation; The core ecological depth analysis includes a differential abundance analysis method and a significant enrichment species determination standard method; Based on the Bray-Curtis distance between each of the boundary samples and the core samples of each of the sample clusters, a similarity matrix of the boundary samples and each of the sample clusters is constructed; Based on the similarity matrix, the distribution positions of the boundary samples among the sample clusters are determined by principal coordinate analysis, and then transition state samples are identified; Based on the environmental gradient, the transition state samples are verified to obtain transition state sample analysis results; Based on the sequencing depth, sequence quality indicators and DNA extraction quality of the noise samples, technical errors are identified, and then low biomass samples are identified after excluding technical anomalies; Performing ecological specificity analysis and biological significance evaluation on the noise samples to obtain noise sample ecological significance mining results.
5. The method according to any one of claims 1 to 4, characterized in that, The clustering results include a plurality of sample clusters, the correlation analysis is performed based on the analysis results and the methane leakage data to obtain microbial grouping analysis results of the target research area, including the following steps: Spatially correlating the clustering results with geochemical data, and then testing differences in geochemical parameters of different sample clusters by statistical methods; Based on the first average distance of the target sample to the same sample cluster and the second average distance of the target sample to other sample clusters, the profile coefficient of the target sample is quantified as a clustering quality evaluation index; Based on the clustering quality evaluation index and the analysis results, a standardized analysis report is prepared as the microbial grouping analysis results of the target research area.
6. A microorganism grouping analysis device characterized by comprising: The device comprises: A first module for obtaining microbial data of a target research area; wherein the microbial data includes microbial abundance data of a plurality of samples and methane leakage data; A second module for obtaining a conversion value of each of the samples by center logarithmic ratio conversion processing based on the microbial abundance data, and then quantifying the distance value between any of the samples by Euclidean distance; The microbial abundance data includes the original abundance value of each species in the sample, and the conversion value of each of the samples is obtained by center logarithmic ratio conversion processing based on the microbial abundance data, including the following steps: Performing multiplication operation on the original abundance values of all species in a single sample to obtain a multiplication result; Taking the number of species as the square root times, performing square root operation on the multiplication result to obtain a geometric mean value; The device comprises: A first module for obtaining microbial data of a target research area; wherein the microbial data includes microbial abundance data of a plurality of samples and methane leakage data; A second module for obtaining a conversion value of each of the samples by center logarithmic ratio conversion processing based on the microbial abundance data, and then quantifying the distance value between any of the samples by Euclidean distance; The microbial abundance data includes the original abundance value of each species in the sample, and the conversion value of each of the samples is obtained by center logarithmic ratio conversion processing based on the microbial abundance data, including the following steps: Performing multiplication operation on the original abundance values of all species in a single sample to obtain a multiplication result; Taking the number of species as the square root times, performing square root operation on the multiplication result to obtain a geometric mean value; Logarithmically transform the ratio of the original abundance value to the geometric mean to obtain the transformed value for each species in each sample; The third module is used to adaptively set the minimum number of samples and the neighborhood radius based on the number of samples and species corresponding to the microbial data; The step of adaptively setting the minimum sample number and neighborhood radius based on the number of samples and species corresponding to the microbial data includes the following steps: If the number of species is less than or equal to a first threshold and the number of species at a preset multiple does not exceed a preset proportion of the number of samples, the minimum number of samples is set to the number of species at the preset multiple; otherwise, the minimum number of samples is set based on a preset fixed value. The value of k is determined based on the minimum number of samples. Based on the k value and the distance values between all the samples, a distance-ranked line chart is drawn using the k-distance graph method; The neighborhood radius is determined based on the inflection point of the distance sorting line graph; the fourth module is used to perform density clustering on the samples in the microbial data to obtain clustering results, and based on the clustering results, the sample type of each sample is marked by combining the minimum number of samples and the neighborhood radius; The fifth module is used to perform multi-level analysis on the clustering results based on the sample type to obtain the analysis results of the clustering results; The sixth module is used to perform correlation analysis based on the analysis results and the methane leakage data to obtain the microbial grouping analysis results of the target study area.
7. An electronic device, comprising: The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method according to any one of claims 1 to 5.
8. A computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 5.
Citation Information
Patent Citations
Environmental microbial community state assessment method based on unsupervised clustering
CN119252350A