Method for mining methane leakage microbial markers based on multi-source data fusion and related equipment

By using a multi-source data fusion method, a multi-dimensional feature space is constructed. Random forest and regression modeling are used to quantify the importance, stability and universality of biomarker samples. This solves the problem of limited dimensionality and reliability of biomarker identification in existing technologies, and realizes efficient identification and cross-regional application of methane leakage microbial biomarkers.

CN121709088BActive Publication Date: 2026-07-14GUANGZHOU MARINE GEOLOGICAL SURVEY SANYA SOUTH CHINA SEA INST OF GEOLOGY +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGZHOU MARINE GEOLOGICAL SURVEY SANYA SOUTH CHINA SEA INST OF GEOLOGY
Filing Date
2025-12-22
Publication Date
2026-07-14

Smart Images

  • Figure CN121709088B_ABST
    Figure CN121709088B_ABST
Patent Text Reader

Abstract

The application discloses a methane seepage microbial marker mining method based on multi-source data fusion and related equipment, and the embodiment of the application fuses microbial gene data, spatial information and environmental geochemical data, constructs a multi-dimensional feature space, and improves the data dimension and reliability of marker mining; moreover, the embodiment of the application adopts a combination of random forest and regression modeling, realizes quantitative screening of the marker, and reduces manual intervention and subjective bias; specifically, the embodiment of the application grades the marker based on the comprehensive score of the three dimensions of importance, stability and universality, facilitating priority selection in subsequent practical applications; the embodiment of the application is suitable for methane seepage identification in different sea areas and different geological backgrounds, has good generalizability and practicality, and can be widely applied in the technical field of marker analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biomarker analysis technology, and in particular to a method and related equipment for mining microbial biomarkers of methane leakage based on multi-source data fusion. Background Technology

[0002] Currently, the identification of microbial biomarkers from methane leakage in natural gas hydrate mining areas mainly employs traditional bioinformatics analysis methods based on a single OTU / ASV abundance table. This approach first obtains microbial sequence data through 16S rRNA gene sequencing. After quality control, OTU clustering, and species annotation, a species abundance table is generated. Based on this, researchers manually divide samples into predefined groups such as "leaking group" and "non-leaking group" according to sampling location or geochemical parameter thresholds. Then, differential abundance analysis is used to identify species with statistically significant differences between groups as candidate biomarkers.

[0003] However, existing technologies rely solely on the single data type of OTU / ASV abundance tables, severely limiting the dimensionality and reliability of biomarker discovery. Traditional methods treat each sample as an independent observation point, completely ignoring the spatial autocorrelation and geographical distribution patterns of microbial communities. This fails to establish a correlation between biomarkers and spatial distribution patterns, significantly weakening their guiding value in regional exploration. Correlation and differential analysis methods based on linear statistical assumptions have limited ability to capture the ubiquitous nonlinear relationships, threshold effects, and complex interactions between microorganisms and the environment, making it difficult to discover indicative signals of multi-species synergistic combinations. Biomarker validation relies solely on resampling or simple group comparisons within the original dataset, lacking a systematic external independent validation framework across marine areas and geological backgrounds, and failing to distinguish between universally applicable core biomarkers and region-specific indicators. The entire process, from manually pre-defined grouping to differential species screening, depends on researcher experience, leaving considerable room for subjective judgment and potentially leading to the neglect or misinterpretation of important biological indicators. Summary of the Invention

[0004] The main objective of this invention is to propose a method, apparatus, electronic device, storage medium, and program product for mining microbial biomarkers of methane leakage based on multi-source data fusion, aiming to solve at least one problem of the prior art.

[0005] To achieve the above objectives, one aspect of this invention proposes a method for mining microbial biomarkers of methane leakage based on multi-source data fusion, the method comprising:

[0006] Obtain raw multi-source data of microbial genes, preprocess the raw multi-source data to obtain a multi-source dataset;

[0007] The multi-source dataset includes a standardized OTU abundance table, spatial information, and environmental information containing methane leakage signals for each biomarker sample.

[0008] Based on multi-source datasets, a multi-dimensional feature matrix is ​​constructed through feature engineering; the multi-dimensional feature matrix includes the multi-dimensional feature space corresponding to each marker sample.

[0009] Based on a multi-dimensional feature matrix, random forest and regression modeling are used to quantify the importance score of each biomarker sample, thereby determining the first candidate biomarker set.

[0010] Based on the multi-dimensional feature matrix, the stability score of each biomarker sample is obtained through regression modeling and regularization parameter optimization quantization, thereby determining the set of second candidate biomarkers.

[0011] The first candidate flag set and the second candidate flag set are synergistically integrated to obtain a comprehensive flag candidate set;

[0012] Based on the comprehensive biomarker candidate set, the universality score of each biomarker sample is obtained by migration verification quantization of biomarker positions;

[0013] The comprehensive score for each biomarker sample is obtained by converting the importance score, stability score, and universality score. Based on the comprehensive score, the biomarker samples are classified into biomarker levels to obtain the results of methane leakage microbial biomarkers.

[0014] In some embodiments, the raw multi-source data includes microbiome data from public databases, as well as raw sequence data, environmental geochemical data, and spatial coordinate data determined based on sampling at preset sampling points. Preprocessing the raw multi-source data to obtain a multi-source dataset includes the following steps:

[0015] Quality control is performed on the original sequence data by removing low-quality sequences to obtain high-quality sequence data, and then the original OTU abundance table is established.

[0016] Among them, the original sequence data was obtained by gene amplicon sequencing of microbial gene data collected at the sampling points, the environmental geochemical data included the concentration gradient curve established by the pore methane concentration corresponding to the sampling points and the methane leakage signal obtained by carbon isotope analysis based on foraminifera shells, and the spatial coordinate data was constructed by measuring water depth and sediment sampling depth using the latitude and longitude coordinates corresponding to the sampling points.

[0017] The original OTU abundance table was subjected to a first normalization process to obtain a normalized OTU abundance table. The first normalization process included center-log ratio transformation, threshold-based filtering of species with low abundance and low occurrence frequency, and sequencing depth normalization for gene amplicon sequencing.

[0018] Z-score standardization was performed on continuous environmental variables in environmental geochemical data, and one-hot encoding was performed on categorical variables in environmental geochemical data to obtain environmental information.

[0019] The spatial coordinate data is converted into a unified spatial reference system, and then spatial autocorrelation feature variables are constructed based on the spatial distance matrix between each sampling point to obtain spatial information.

[0020] The standardized OTU abundance table, environmental information, and spatial information are linked to the corresponding sampling points as measured biomarker samples.

[0021] Extracting public biomarker samples from microbiome data in public databases;

[0022] Based on public marker samples, a dataset for cross-regional comparative analysis was constructed by combining measured marker samples, resulting in a multi-source dataset.

[0023] In some embodiments, a multi-dimensional feature matrix is ​​constructed through feature engineering, including the following steps:

[0024] Based on the standardized OTU abundance table, the relative abundance characteristics of a single species are extracted and the phylogenetic diversity index is calculated based on the phylogenetic tree. Then, a microbial functional feature vector is constructed to obtain the microbial characteristics.

[0025] Based on spatial information, a spatial weight matrix is ​​constructed using the distance decay criterion through spatial feature engineering to quantify spatial dependencies, and then multi-scale spatial pattern features are extracted to obtain spatial features.

[0026] Based on environmental information, the polynomial characteristics of environmental factors are established using the principle of statistical significance, thus obtaining environmental characteristics.

[0027] Microbial features, spatial features, and environmental features are concatenated into a unified feature matrix to obtain a multi-dimensional feature space.

[0028] By using a multi-level matching strategy to align marker samples from different data sources, and then combining the multi-dimensional feature space of each marker sample, a multi-dimensional feature matrix is ​​obtained.

[0029] In some embodiments, based on a multi-dimensional feature matrix, the importance score of each biomarker sample is quantified using random forest and regression modeling to determine the first candidate biomarker set, including the following steps:

[0030] Based on the multi-dimensional feature space corresponding to all marker samples, a training set and a test set are obtained.

[0031] Based on the training set, a microbial biomarker mining model is constructed using random forest. The microbial biomarker mining model includes a preset number of decision trees, and the training set of each decision tree is constructed by sampling biomarker samples with replacement using the self-sampling method.

[0032] The Gini importance of each biomarker sample is determined based on the reduction in impurity corresponding to node splits in the decision tree, and then the first sub-score is obtained through normalization.

[0033] The baseline performance metrics of the microbial biomarker mining model were obtained based on the test set validation.

[0034] The feature values ​​of the target biomarker samples in the test set are randomly arranged to re-evaluate the model performance of the microbial biomarker mining model.

[0035] Based on the baseline performance metrics and model performance, record the degree of performance degradation, and return to the step of randomly arranging the feature values ​​in the test set until the performance degradation of the preset number of groups is obtained;

[0036] The second sub-score of the target biomarker sample is determined based on the average of all performance degradation levels;

[0037] Based on a multi-dimensional feature matrix, regression modeling is performed using Lasso regression, and then the regression coefficient of each biomarker sample is collected as the third sub-score.

[0038] The importance score of the corresponding marker sample is obtained by performing a first weighted summation on the first sub-score, second sub-score, and third sub-score of each marker sample.

[0039] Based on the importance score, the first candidate flag set is obtained by comparison and filtering using a preset importance threshold.

[0040] In some embodiments, based on a multi-dimensional feature matrix, a stability score for each biomarker sample is obtained through regression modeling and regularization parameter optimization quantization, thereby determining a second candidate biomarker set, including the following steps:

[0041] Based on a multi-dimensional feature matrix, multiple subsets of data are extracted with replacement using Bootstrap resampling.

[0042] Lasso regression was run independently on each subset of the dataset for regression modeling.

[0043] The stability of each biomarker sample across different subsets was evaluated using k-fold cross-validation, and the fourth sub-score was obtained by quantification.

[0044] The fifth sub-score is determined based on the reciprocal of the coefficient variability of each marker sample in each Bootstrap resampling.

[0045] Based on a multi-dimensional feature matrix, spatial cross-validation is used to determine the sixth sub-score by quantifying the consistency of the performance of biomarker samples in different spatial blocks.

[0046] The stability score of the corresponding biomarker sample is obtained by performing a second weighted summation on the fourth, fifth, and sixth sub-scores of each biomarker sample.

[0047] Statistical analysis is performed on all Bootstrap resampled subsets. When the proportion of a marker sample selected in all subsets exceeds a preset proportion, the corresponding marker sample is included in the second candidate marker set.

[0048] In some embodiments, based on a comprehensive biomarker candidate set, the universality score of each biomarker sample is obtained through biomarker migration verification quantization, including the following steps:

[0049] By applying the marker samples corresponding to the comprehensive marker candidate set to different sea area datasets, cross-sea area mobility is quantified to determine the seventh sub-score;

[0050] The eighth sub-score is determined based on the consistency of feature ranking of the marker samples corresponding to the comprehensive marker candidate set in different sea area datasets.

[0051] The universality score of the corresponding biomarker sample is obtained by performing a third weighted summation on the seventh and eighth sub-scores of each biomarker sample.

[0052] In some embodiments, a comprehensive score for each biomarker sample is obtained based on importance score, stability score, and universality score. Based on the comprehensive score, all biomarker samples are classified into biomarker rank groups, including the following steps:

[0053] Based on the preset weights, the importance score, stability score and universality score are summed in a fourth weighted manner to obtain the comprehensive score of each biomarker sample.

[0054] Based on the overall score, all marker samples are divided into multiple marker levels using a preset score range.

[0055] To achieve the above objectives, another aspect of the present invention proposes a device for mining microbial biomarkers of methane leakage based on multi-source data fusion, the device comprising:

[0056] The first module is used to acquire raw multi-source data of microbial genes, preprocess the raw multi-source data, and obtain a multi-source dataset.

[0057] The multi-source dataset includes a standardized OTU abundance table, spatial information, and environmental information containing methane leakage signals for each biomarker sample.

[0058] The second module is used to construct a multi-dimensional feature matrix based on multi-source datasets through feature engineering; the multi-dimensional feature matrix includes the multi-dimensional feature space corresponding to each marker sample;

[0059] The third module is used to determine the first candidate marker set by quantifying the importance score of each marker sample based on the multi-dimensional feature matrix using random forest and regression modeling.

[0060] The fourth module is used to determine the second candidate marker set by obtaining the stability score of each marker sample based on the multi-dimensional feature matrix through regression modeling and regularization parameter optimization quantization.

[0061] The fifth module is used to collaboratively integrate the first candidate flag set and the second candidate flag set to obtain a comprehensive candidate flag set;

[0062] The sixth module is used to obtain the universality score of each biomarker sample by using migration verification quantization of biomarker bits based on the comprehensive biomarker candidate set.

[0063] The seventh module is used to convert the importance score, stability score, and universality score into a comprehensive score for each biomarker sample. Based on the comprehensive score, all biomarker samples are classified into biomarker levels to obtain the results of methane leakage microbial biomarkers.

[0064] To achieve the above objectives, another aspect of the present invention provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the aforementioned method.

[0065] To achieve the above objectives, another aspect of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned method.

[0066] To achieve the above objectives, another aspect of the present invention provides a computer program product, including a computer program that, when executed by a processor, implements the aforementioned method.

[0067] The embodiments of the present invention include at least the following beneficial effects: The present invention provides a method, apparatus, electronic device, storage medium, and program product for mining methane leakage microbial biomarkers based on multi-source data fusion. This scheme obtains raw multi-source data of microbial genes, preprocesses the raw multi-source data to obtain a multi-source dataset; wherein, the multi-source dataset includes a standardized OTU abundance table, spatial information, and environmental information containing methane leakage signals for each biomarker sample; based on the multi-source dataset, a multi-dimensional feature matrix is ​​constructed through feature engineering; wherein, the multi-dimensional feature matrix includes a multi-dimensional feature space corresponding to each biomarker sample; based on the multi-dimensional feature matrix, random forest and regression modeling are used to quantify and obtain each The importance scores of each biomarker sample are used to determine the first candidate biomarker set. Based on the multi-dimensional feature matrix, the stability score of each biomarker sample is obtained through regression modeling and regularization parameter optimization, thus determining the second candidate biomarker set. The first and second candidate biomarker sets are synergistically integrated to obtain a comprehensive biomarker candidate set. Based on the comprehensive biomarker candidate set, the universality score of each biomarker sample is obtained through biomarker migration verification. The comprehensive score of each biomarker sample is obtained by transforming the importance score, stability score, and universality score. Based on the comprehensive score, all biomarker samples are classified into biomarker levels to obtain the methane leakage microbial biomarker results. This invention, through the fusion of microbial genetic data, spatial information, and environmental geochemical data, constructs a multi-dimensional feature space, thereby improving the data dimensionality and reliability of biomarker mining. Furthermore, this invention employs a combination of random forest and regression modeling to achieve quantitative screening of biomarkers, reducing human intervention and subjective bias. Specifically, this invention classifies biomarkers based on a comprehensive score across three dimensions: importance, stability, and universality, facilitating priority selection in subsequent practical applications. This invention is applicable to methane leakage identification in different sea areas and geological backgrounds, possessing good scalability and practicality. Attached Figure Description

[0068] Figure 1 This is a schematic diagram of an implementation environment for a method for mining microbial biomarkers of methane leakage based on multi-source data fusion, provided in an embodiment of the present invention.

[0069] Figure 2 This is a flowchart illustrating a method for mining microbial biomarkers of methane leakage based on multi-source data fusion, provided in an embodiment of the present invention.

[0070] Figure 3 This is a schematic diagram of the preprocessing flow for raw multi-source data provided in an embodiment of the present invention;

[0071] Figure 4This is a schematic diagram of the process of expanding a multi-dimensional feature matrix through feature engineering, as provided in an embodiment of the present invention.

[0072] Figure 5 This is a schematic diagram of the unfolding process of step S400 provided in an embodiment of the present invention;

[0073] Figure 6 This is a schematic diagram of the unfolding process of step S600 provided in an embodiment of the present invention;

[0074] Figure 7 This is a schematic diagram of the unfolding process of step S700 provided in the embodiment of the present invention;

[0075] Figure 8 This is a schematic diagram of the overall process of the method for mining microbial biomarkers of methane leakage based on multi-source data fusion provided in the embodiments of the present invention;

[0076] Figure 9 This is a schematic diagram of a methane leakage microbial marker mining device based on multi-source data fusion provided in an embodiment of the present invention;

[0077] Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0078] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of this invention; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this invention as detailed in the appended claims.

[0079] It is understood that the terms “first,” “second,” etc., used in this invention may be used herein to describe various concepts, but unless specifically stated otherwise, these concepts are not limited by these terms. These terms are used only to distinguish one concept from another. For example, first information may also be referred to as second information without departing from the scope of embodiments of the invention, and similarly, second information may also be referred to as first information. Depending on the context, the words “if,” “when,” or “in response to determination” as used herein may be interpreted as “when…” or “when…” or “in response to determination.”

[0080] The terms “at least one,” “multiple,” “each,” “any,” etc., used in this invention, “at least one” includes one, two, or more than two; “multiple” includes two or more than two; “each” refers to each of the corresponding multiple; and “any” refers to any one of the multiple.

[0081] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein is for the purpose of describing embodiments of the invention only and is not intended to limit the invention.

[0082] In related technologies, existing technologies rely solely on a single data type, the OTU / ASV abundance table, which severely limits the dimensionality and reliability of biomarker mining.

[0083] In view of this, this invention provides a method and related equipment for mining methane leakage microbial biomarkers based on multi-source data fusion. This method acquires raw multi-source data of microbial genes, preprocesses the raw multi-source data to obtain a multi-source dataset; wherein the multi-source dataset includes a standardized OTU abundance table, spatial information, and environmental information containing methane leakage signals for each biomarker sample; based on the multi-source dataset, a multi-dimensional feature matrix is ​​constructed through feature engineering; wherein the multi-dimensional feature matrix includes a multi-dimensional feature space corresponding to each biomarker sample; based on the multi-dimensional feature matrix, the importance of each biomarker sample is quantified using random forest and regression modeling. The scores are used to determine the first candidate biomarker set; based on the multi-dimensional feature matrix, the stability score of each biomarker sample is obtained through regression modeling and regularization parameter optimization, thereby determining the second candidate biomarker set; the first and second candidate biomarker sets are synergistically integrated to obtain a comprehensive biomarker candidate set; based on the comprehensive biomarker candidate set, the universality score of each biomarker sample is obtained through biomarker migration verification quantification; based on the importance score, stability score, and universality score, the comprehensive score of each biomarker sample is obtained; according to the comprehensive score, all biomarker samples are classified into biomarker levels to obtain the methane leakage microbial biomarker results. This invention, through the fusion of microbial genetic data, spatial information, and environmental geochemical data, constructs a multi-dimensional feature space, thereby improving the data dimensionality and reliability of biomarker mining. Furthermore, this invention employs a combination of random forest and regression modeling to achieve quantitative screening of biomarkers, reducing human intervention and subjective bias. Specifically, this invention classifies biomarkers based on a comprehensive score across three dimensions: importance, stability, and universality, facilitating priority selection in subsequent practical applications. This invention is applicable to methane leakage identification in different sea areas and geological backgrounds, possessing good scalability and practicality.

[0084] It is understood that the method for mining methane leakage microbial biomarkers based on multi-source data fusion provided by this invention can be applied to any computer device with data processing and computing capabilities, and this computer device can be various terminals or servers. When the computer device in the embodiments is a server, the server is an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. Optionally, the terminal can be a smartphone, tablet, laptop, or desktop computer, but it is not limited to these.

[0085] like Figure 1 The diagram shown is a schematic representation of an implementation environment provided by an embodiment of the present invention. (Refer to...) Figure 1 The implementation environment includes at least one terminal 102 and a server 101. The terminal 102 and the server 101 can be connected via a network, either wirelessly or via a wired connection, to complete data transmission and exchange.

[0086] Server 101 can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0087] Additionally, server 101 can also be a node server in a blockchain network. Blockchain is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms.

[0088] Terminal 102 can be a smartphone, tablet computer, laptop computer, desktop computer, smart speaker, smartwatch, etc., but is not limited to these. Terminal 102 and server 101 can be directly or indirectly connected via wired or wireless communication, and this embodiment of the invention does not impose any limitations.

[0089] For example, based on Figure 1The implementation environment shown in this embodiment of the invention provides a method for mining methane leakage microbial biomarkers based on multi-source data fusion. The following description uses the application of this method for mining methane leakage microbial biomarkers based on multi-source data fusion in server 101 as an example. It can be understood that this method for mining methane leakage microbial biomarkers based on multi-source data fusion can also be applied in terminal 102.

[0090] Reference Figure 2 , Figure 2 This is an optional flowchart of a method for mining microbial biomarkers of methane leakage based on multi-source data fusion provided in an embodiment of the present invention. The execution subject of this method for mining microbial biomarkers of methane leakage based on multi-source data fusion can be any of the aforementioned computer devices (including servers or terminals). Figure 2 The method may include, but is not limited to, steps S100 to S700.

[0091] Step S100: Obtain raw multi-source data of microbial genes, preprocess the raw multi-source data to obtain a multi-source dataset;

[0092] The multi-source dataset includes a standardized OTU abundance table, spatial information, and environmental information containing methane leakage signals for each biomarker sample.

[0093] It should be noted that the original multi-source data includes microbiome data from public databases, as well as raw sequence data, environmental geochemical data, and spatial coordinate data determined based on sampling at preset sampling points. In some embodiments, such as... Figure 3As shown, preprocessing the original multi-source data to obtain a multi-source dataset may include the following steps: S110, performing quality control by removing low-quality sequences from the original sequence data to obtain quality sequence data and then establishing an original OTU abundance table; wherein, the original sequence data is obtained by performing gene amplicon sequencing on microbial gene data collected from sampling points, the environmental geochemical data includes the concentration gradient curve established by the pore methane concentration corresponding to the sampling point and the methane leakage signal obtained based on the carbon isotope analysis of foraminifera shells, and the spatial coordinate data is constructed by measuring the water depth and sediment sampling depth using the latitude and longitude coordinates corresponding to the sampling points; S120, performing a first normalization process on the original OTU abundance table to obtain a normalized OTU abundance table; wherein, the first normalization process includes central logarithmic ratio transformation, threshold-based low abundance and low emission... The process includes: S130, S140, S150, S160, S170, S180, S180, S19 ...

[0094] For example, in some specific implementations, cold seep microbiome data can be downloaded from public databases such as NCBI; columnar sediment samples can be collected at a station in the South China Sea and 16S rRNA amplicon sequencing can be performed to obtain raw sequence data; the sequences can be quality controlled, denoised, and clustered to generate an OTU table; the OTU table can be transformed to a central logarithmic ratio to filter out low-abundance species; continuous variables such as methane concentration and sulfate can be standardized using Z-scores; the coordinates of the sampling points can be converted to the WGS84 coordinate system and the spatial distance matrix can be calculated; and the measured data can be aligned with public data through unique sample identifiers (such as hash IDs) to construct a cross-regional dataset.

[0095] Step S200: Based on the multi-source dataset, a multi-dimensional feature matrix is ​​constructed through feature engineering;

[0096] The multi-dimensional feature matrix includes the multi-dimensional feature space corresponding to each marker sample;

[0097] It should be noted that in some embodiments, such as Figure 4As shown, the multi-dimensional feature matrix constructed through feature engineering can include the following steps: S210, based on the standardized OTU abundance table, extract the relative abundance features of a single species and calculate the phylogenetic diversity index based on the phylogenetic tree, and then construct a microbial functional feature vector to obtain microbial features; S220, based on spatial information, construct a spatial weight matrix using the distance decay criterion through spatial feature engineering to quantify spatial dependencies, and then extract multi-scale spatial pattern features to obtain spatial features; S230, based on environmental information, establish multinomial features of environmental factors using the statistical significance principle to obtain environmental features; S240, concatenate the microbial features, spatial features, and environmental features into a unified feature matrix to obtain a multi-dimensional feature space; S250, perform sample alignment of biomarker samples from different data sources through a multi-level matching strategy, and then combine the multi-dimensional feature space of each biomarker sample to establish a multi-dimensional feature matrix.

[0098] For example, in some specific implementations, the abundance of a single species can be extracted from the OTU table as a feature; Faith's PD index (phylogenetic diversity index) can be calculated based on the phylogenetic tree; the abundance of KEGG functional pathways can be inferred using the PICRUSt2 software platform to construct a functional feature vector; a spatial weight matrix can be constructed based on the distance between sampling points to extract spatial lag variables; a second-order polynomial feature can be constructed for methane concentration; and the above three types of features can be concatenated into a unified feature matrix, with sample alignment ensured through a multi-level matching strategy.

[0099] Step S300: Based on the multi-dimensional feature matrix, use random forest and regression modeling to quantify the importance score of each marker sample and then determine the first candidate marker set;

[0100] It should be noted that in some embodiments, step S300 may include the following steps: dividing the multi-dimensional feature space corresponding to all biomarker samples into a training set and a test set; constructing a microbial biomarker mining model using a random forest based on the training set; wherein the microbial biomarker mining model includes a preset number of decision trees, and the training set of each decision tree is constructed by sampling biomarker samples with replacement using an autonomous sampling method; determining the Gini importance of each biomarker sample based on the reduction in impurity corresponding to node splits in the decision tree, and then obtaining the first sub-score through normalization; verifying the benchmark performance index of the microbial biomarker mining model based on the test set; randomly arranging the feature values ​​of the target biomarker samples in the test set, and then... The model performance of the microbial biomarker mining model is re-evaluated; the degree of performance degradation is recorded based on the baseline performance index and model performance, and the step of randomly arranging the feature values ​​in the test set is repeated until the performance degradation of the preset number of groups is obtained; the second sub-score of the target biomarker sample is determined based on the average of all performance degradation; regression modeling is performed using Lasso regression based on the multi-dimensional feature matrix, and the regression coefficient of each biomarker sample is collected as the third sub-score; the first sub-score, second sub-score and third sub-score of each biomarker sample are summed in a first weighted manner to obtain the importance score of the corresponding biomarker sample; based on the importance score, the first candidate biomarker set is obtained by comparison and screening using a preset importance threshold.

[0101] For example, in some specific implementations, the samples can be divided into training and test sets proportionally; a model can be trained using a random forest (500 trees); the Gini importance of each feature can be calculated based on the reduction in Gini impurity; each feature can be randomly permuted 50 times on the test set, and the permutation importance can be calculated; the importance score can be calculated by combining the Lasso regression coefficient; and an importance threshold (such as the 95th percentile) can be set to select the first candidate biomarker set.

[0102] Step S400: Based on the multi-dimensional feature matrix, the stability score of each marker sample is obtained through regression modeling and regularization parameter optimization quantization, thereby determining the second candidate marker set;

[0103] It should be noted that in some embodiments, such as Figure 5As shown, step S400 may include the following steps: S410, based on the multi-dimensional feature matrix, extract multiple subsets with replacement through Bootstrap resampling; S420, run Lasso regression independently on each subset for regression modeling; S430, evaluate the stability of each biomarker sample in different subsets through k-fold cross-validation, and quantify to obtain the fourth sub-score; S440, determine the fifth sub-score based on the inverse of the coefficient variation of each biomarker sample in each Bootstrap resampling; S450, based on the multi-dimensional feature matrix, quantify the consistency of biomarker sample performance in different spatial blocks through spatial cross-validation to determine the sixth sub-score; S460, perform a second weighted summation on the fourth, fifth, and sixth sub-scores of each biomarker sample to obtain the stability score of the corresponding biomarker sample; S470, perform statistics on all Bootstrap resampling subsets, and when the selection ratio of a biomarker sample in all subsets exceeds a preset ratio, include the corresponding biomarker sample in the second candidate biomarker set.

[0104] For example, in some specific implementations, 100 subsets can be generated by Bootstrap resampling; Lasso regression can be run on each subset to record the frequency of feature selection; the stability score of the feature in the subset can be calculated by k-fold cross-validation; the inverse of the coefficient variability can be calculated as the Bootstrap stability; spatial cross-validation can be performed based on spatial block partitioning to evaluate feature consistency; the stability scores can be weighted and summarized, and features that are stable in more than 70% of the subsets can be selected as the second candidate set.

[0105] Step S500: The first candidate flag set and the second candidate flag set are collaboratively integrated to obtain a comprehensive flag candidate set;

[0106] For example, in some implementations, the screening results of Lasso regression are cross-validated with the feature importance ranking of random forest, comparing biomarkers that are important in both models and calculating the ranking relevance of biomarkers in different models. Complementary integration preserves unique biomarkers that perform well in a single model, assesses the biological plausibility of these biomarkers, and constructs a comprehensive biomarker candidate set.

[0107] Step S600: Based on the comprehensive biomarker candidate set, the universality score of each biomarker sample is obtained through biomarker migration verification quantization.

[0108] It should be noted that in some embodiments, such as Figure 6As shown, step S600 may include the following steps: S610, applying the marker samples corresponding to the comprehensive marker candidate set to different sea area datasets, quantifying the cross-sea area mobility to determine the seventh sub-score; S620, determining the eighth sub-score based on the feature ranking consistency of the marker samples corresponding to the comprehensive marker candidate set in different sea area datasets; S630, performing a third weighted summation on the seventh and eighth sub-scores of each marker sample to obtain the universality score of the corresponding marker sample.

[0109] For example, in some specific implementations, the selected markers can be applied to other marine datasets; their predictive accuracy retention (transferability) on these datasets can be calculated; the consistency of their feature importance ranking across different marine areas can be evaluated; and a universality score can be calculated by weighting transferability and ranking consistency.

[0110] Step S700: Based on the importance score, stability score and universality score, the comprehensive score of each biomarker sample is obtained. According to the comprehensive score, the biomarker samples are classified into biomarker levels to obtain the results of methane leakage microbial biomarkers.

[0111] It should be noted that in some embodiments, such as Figure 7 As shown, step S700 may include the following steps: S710, based on preset weights, perform a fourth weighted summation on the importance score, stability score and universality score to obtain the comprehensive score of each marker sample; S720, according to the comprehensive score, use a preset score range to divide all marker samples into multiple marker level.

[0112] For example, in some specific implementations, the weights for importance, stability, and universality are set to 0.4, 0.35, and 0.25, respectively; a comprehensive score is calculated for each marker; the markers are divided into three levels based on their scores: ≥0.8 are core markers, 0.6–0.8 are auxiliary markers, and 0.4–0.6 are candidate markers; the classification results and corresponding confidence intervals are output.

[0113] To explain in detail the principle of the technical solution of the present invention, the overall process of the present invention will be described below with reference to some specific embodiments. It is easy to understand that the following is an explanation of the technical principle of the present invention and should not be regarded as a limitation of the present invention.

[0114] First, it should be noted that, addressing the shortcomings of the aforementioned existing technologies, the core technical problem this invention aims to solve is: how to overcome the analytical limitations of a single OTU abundance table and establish a multi-source data fusion-driven intelligent identification and verification method for methane leakage microbial biomarkers. This method can: achieve collaborative analysis of multi-source data, systematically integrate microbial composition data, spatial location information, environmental geochemical parameters, and cross-oceanic datasets to construct a multi-dimensional feature space, providing a richer and more comprehensive data foundation for biomarker mining; capture complex ecological relationships, employing machine learning algorithms capable of effectively handling nonlinear relationships and interactions to deeply mine multi-level bioindicator signals such as single-species biomarkers, multi-species synergistic combinations, and microbial network modules; integrate spatial geographic location as a key feature into the machine learning model to reveal the spatial distribution patterns and heterogeneity of microbial biomarkers, enhancing their practical value in regional exploration; and comprehensively evaluate the stability, spatial predictive ability, and cross-regional universality of biomarkers through a multi-level verification framework of internal cross-validation, spatial validation, and cross-oceanic external validation to ensure their reliability. Reduce human intervention and subjective judgment in the analysis process, establish a standardized and repeatable biomarker mining process, and improve the consistency and scalability of the method.

[0115] In some specific application scenarios, such as Figure 8 As shown, embodiments of the present invention can be implemented through the following process steps:

[0116] Step 1: Multi-source data acquisition and preprocessing:

[0117] Microbial genetic data collection included collecting columnar sediment samples from the seabed, performing 16S rRNA amplicon sequencing across different depth gradients, and covering archaea, bacteria, and fungal communities. Raw sequence data underwent quality control, removing low-quality sequences. Environmental geochemical data acquisition included measuring methane concentration in pore water, establishing concentration gradient curves, analyzing carbon isotopes (δ¹³C) in foraminifera shells, and identifying methane leakage signals. Key geochemical parameters such as sulfate concentration and sulfide content were determined. Spatial location information was recorded. The latitude and longitude coordinates of each sampling point were accurately recorded, and water depth and sediment sampling depth were measured to establish a three-dimensional spatial coordinate database. Microbiome data from other cold seep areas were collected from public databases, and relevant microbial community composition information from literature was compiled to establish a dataset for cross-regional comparative analysis. Data preprocessing included microbial data standardization, central logarithmic ratio transformation of the OTU abundance table, filtering low-abundance and low-frequency species, sequencing depth standardization, environmental data normalization, Z-score standardization for continuous environmental variables, and one-heat encoding for categorical variables. To handle missing and outlier values, spatial data is unified by converting geographic coordinates into a unified spatial reference system, calculating the spatial distance matrix between sampling points, and constructing spatial autocorrelation feature variables.

[0118] Step Two: Feature Engineering and Multi-Source Data Fusion

[0119] Feature construction process: Microbial feature extraction, extracting the relative abundance features of single species. Calculating the phylogenetic diversity index, constructing microbial functional feature vectors, spatial feature engineering, calculating spatial lag variables to reflect the influence of neighboring samples, constructing a spatial weight matrix, quantifying spatial dependencies, extracting multi-scale spatial pattern features, constructing environmental gradient features, establishing multinomial features of environmental factors, quantifying the environmental distance matrix, and constructing an environmental heterogeneity index. Features from different sources are concatenated into a unified feature matrix, and feature scaling and standardization are performed to construct a multi-dimensional feature space. Ensuring a one-to-one correspondence between samples from different data sources, handling sample missing and misalignment issues, and establishing a complete sample-feature matrix.

[0120] In microbial feature extraction, the phylogenetic diversity index is calculated using Faith's phylogenetic diversity index, which reflects the breadth of evolutionary history by calculating the total branch length of all extant species in the sample on the phylogenetic tree.

[0121] The formula for Faith's PD index (calculated based on phylogenetic tree) is:

[0122]

[0123] in, This represents Faith's PD index (i.e., the phylogenetic diversity index). This indicates traversing each branch in the phylogenetic tree (there are n branches in total). Let be the branch length of the i-th branch in the phylogenetic tree. It represents the English word "length". It is an indicator function (1 if the branch has at least one descendant species in the sample, and 0 otherwise). It represents the English word "Indicator".

[0124] Specifically, this index is equal to the sum of the branch lengths of the smallest subset required to connect all species within a sample in a phylogenetic tree, and can effectively characterize the level of evolutionary diversity of a community.

[0125] The construction of functional feature vectors mainly relies on the PICRUSt2 software platform. This tool infers metagenomic functional content based on 16S rRNA gene sequences, annotates the metabolic functional potential of microbial communities through the KEGG Orthology database, and finally generates feature vectors containing the abundance of multiple functional pathways.

[0126] The spatial weight matrix is ​​constructed primarily using the distance decay criterion to determine the degree of mutual influence between spatial units. Specifically, the spatial weights are calculated based on the spherical distance between sampling points. Sample pairs with distances within a set threshold are assigned non-zero weights, while sample pairs exceeding the threshold have zero weights. This construction method accurately quantifies the strength and scope of spatial dependencies.

[0127] The selection of the order of polynomial features follows the principle of statistical significance. The optimal order is determined by the forward selection method. Usually, the order is gradually increased starting from the second-order polynomial. After each increase, the significance of the new feature is evaluated using the F test until the increase in order no longer brings significant model improvement. This process ensures that overfitting is avoided while capturing nonlinear relationships.

[0128] The sample alignment process employs a multi-level matching strategy. First, a unified sample identification system is established, assigning a unique identifier to each physical sample. A unique sample ID is generated using a hash algorithm to ensure consistency across databases. This identifier includes core information such as the sampling area, station number, sampling depth, and timestamp. During the data integration phase, this identifier enables precise matching of samples from different data sources. For samples that cannot be directly matched, a spatial-temporal proximity principle is used for association, prioritizing the matching of sample pairs that are spatially closest and have the most similar sampling times. Simultaneously, a sample metadata verification mechanism is established, cross-validating multiple pieces of information, including sampling records, laboratory numbers, and spatial coordinates, to ensure the accuracy of sample alignment.

[0129] In some optional implementations, a tiered approach is used to handle missing samples. First, a random missing data mechanism is used to detect and distinguish different types of missing data. For samples with completely random missing data, simple deletion or mean imputation methods are employed. When the proportion of missing samples is high, a multiple imputation process is implemented to generate multiple complete imputation datasets, and the results are merged and analyzed using Rubin rules. For samples with non-random missing data, a missing data mechanism model is established, and the bias caused by missing data is corrected by introducing missing data indicator variables and selecting a model. Throughout the processing, the imputation quality is continuously monitored, and the imputation effect is evaluated by comparing the distribution characteristics of the imputed values ​​with the actual observed values ​​to ensure the reliability of data fusion.

[0130] Step 3: Random Forest Marker Mining

[0131] Model: Random Forest;

[0132] A microbial biomarker mining model was constructed using the random forest algorithm. First, the key parameters of the model were determined: the number of decision trees was set to 500, the maximum depth of each tree was dynamically determined through cross-validation, the minimum number of samples required for node splitting was set to 5, and the number of features considered at each split was the square root of the total number of features. During model training, a self-sampling method was used to draw samples with replacement from the original dataset to construct the training set for each decision tree, ensuring the model's diversity and robustness.

[0133] Feature importance calculation employs a dual-validation method combining Gini importance and permutation importance. Gini importance is calculated based on the reduction in impurity at each node split in the decision tree. It is obtained by accumulating the total Gini impurity reduction resulting from each split across all decision trees and then normalizing the sum to obtain a relative importance score. Specifically, for each feature, the sum of its impurity reduction at each split node across all trees is calculated, divided by the total number of trees, and finally, the importance scores of all features are normalized to a sum of 1. Permutation importance calculation uses a more rigorous method: first, the baseline performance index of the model is calculated on the test set; then, the values ​​of each feature are randomly permuted, breaking their relationship with the target variable, and the model performance is re-evaluated, with the degree of performance degradation recorded. This process is repeated 50 times, and the average of the performance degradation is taken as the permutation importance score for that feature. The advantage of this method is that it can identify features that have a genuine causal relationship with the target variable rather than being merely coincidental.

[0134] The importance threshold is determined using a multi-criteria decision framework, comprehensively considering statistical significance, effect size, and stability indices. First, a statistical significance threshold is set based on permutation tests, requiring that the feature's importance score be significantly higher than the score after random permutation. Specifically, the distribution of permutation importance scores is calculated, and 95% is taken as the statistical significance threshold. Second, the cumulative distribution curve of feature importance scores is analyzed using the elbow rule to find the point of maximum curvature as the initial threshold. Simultaneously, Cohen's d effect size is calculated, requiring the effect size d value of important features to be greater than 0.5, ensuring that the selected features have practical significance rather than just statistical significance. Stability assessment is achieved through repeated cross-validation, requiring important features to remain significant in at least 70% of the subsets. Finally, confidence intervals for feature importance are calculated using Bootstrap resampling, and features with a 95% confidence interval lower bound greater than 0 are selected as candidate biomarkers.

[0135] Biomarker Identification and Validation: In the biomarker identification stage, all microbial features are first ranked according to their comprehensive importance score. Then, the aforementioned multi-criteria thresholding is applied for screening to ensure that the final biomarkers are both statistically significant and biologically plausible. For the screened biomarkers, their stability under different environmental conditions and consistency across different spatial scales are further analyzed. Model validation employs a rigorous spatial hierarchical cross-validation strategy, dividing the study area into multiple spatial blocks to ensure that the training and test sets are spatially independent. This validation method effectively evaluates the model's generalization ability in unknown regions and avoids performance overestimation due to spatial autocorrelation.

[0136] Synergistic Effect Identification: To identify multi-species combinations exhibiting synergistic effects, conditional feature importance analysis was employed. By calculating the conditional importance of features given other relevant features, microorganisms that are not highly important when considered individually but have significant predictive value in specific combinations were identified. Simultaneously, by analyzing the splitting paths of the decision tree, frequently co-occurring feature combinations were discovered; these combinations often represent ecologically significant microbial functional groups.

[0137] Step 4: Lasso regression biomarker screening:

[0138] Model: Lasso regression regularization parameter optimization;

[0139] Model selection criteria: Prioritize predictive performance, selecting the model that performs best on the validation set, balancing model complexity and prediction accuracy. Ensure the model's generalization ability; consider interpretability, prioritizing models with easily interpretable coefficients, taking into account the biological significance of features, and ensuring the interpretability of the results.

[0140] Coefficient Stability Assessment: To ensure the reliability of the biomarker selection results, a rigorous coefficient stability assessment is implemented. Using Bootstrap resampling, multiple subsets of the original data are extracted with replacement, and Lasso regression is run independently on each subset. The frequency with which each feature is selected into the model is calculated to form a stability score. A feature is considered a reliable biomarker with stable predictive ability only when it is selected in more than 70% of the Bootstrap samples. This process effectively reduces the risk of misselection due to random data fluctuations.

[0141] Biomarker Screening and Interpretation: After Lasso regression screening, a concise set of biomarkers was obtained. These biomarkers underwent in-depth functional annotation and pathway analysis, mapping them to known microbial metabolic pathways, particularly key pathways related to the methane cycle. By analyzing the interactions between biomarkers, potential functional modules or synergistic groups were identified. Simultaneously, by combining the magnitude and direction of their regression coefficients, the contribution and direction of each biomarker to methane leakage prediction were quantified, providing statistical evidence for positive and negative correlations.

[0142] Step 5: Multi-model collaborative integration:

[0143] The selection results from Lasso regression were cross-validated with the feature importance ranking from random forest. Biomarkers important in both models were compared, and the ranking correlation of biomarkers across different models was calculated. Complementary integration was performed to retain unique biomarkers that performed well in a single model, assessing the biological plausibility of these biomarkers and constructing a comprehensive biomarker candidate set.

[0144] Biomarker classification system: Core biomarkers were identified, including highly important species that remained stable in both models, and key taxa with clearly defined ecological functions. Species appearing in different validation sets were also included.

[0145] Screening of auxiliary biomarkers: species that perform well in a model and are ecologically significant, secondary indicator species that provide supplementary information, and candidate biomarkers with potential application value.

[0146] Combination biomarker identification: Identify multi-species combinations with synergistic indicative effects, construct biomarkers for microbial functional modules, and determine the optimal combination indicative pattern.

[0147] Step Six: Implementation of a Multi-Level Verification System

[0148] Internal validation phase: Cross-validation is performed, hierarchical k-fold cross-validation is implemented, the mean and variance of model performance metrics are calculated, and the stability of the biomarker in different subsets is evaluated.

[0149] Bootstrap validation: Perform multiple Bootstrap resampling operations, calculate the confidence intervals of marker importance, and evaluate the robustness of marker selection.

[0150] Spatial validation phase: Spatial cross-validation design, which divides the training set and test set based on spatial location to ensure that the test set is spatially separated from the training set.

[0151] Evaluating the model's spatial predictive power: Spatial autocorrelation test. Calculate the spatial autocorrelation of the residuals. Verify whether the model adequately captures spatial patterns and ensures the correct handling of spatial dependencies.

[0152] External validation phase: Independent dataset testing, using completely independent datasets to validate the model.

[0153] Evaluate the generalization performance of the model: examine the external validity of the biomarker.

[0154] Cross-ocean applicability assessment: Applying the model to data from different sea areas to assess the geographical universality of the markers and identify region-specific adjustment needs.

[0155] Step 7: Construction of a High-Reliability Biomarker System:

[0156] Biomarker Quantitative Evaluation: A comprehensive score is calculated, combining importance, stability, and universality. A tiered evaluation system for biomarkers is established to determine final biomarker priorities. Uncertainty is quantified, assessing the selection uncertainty of each biomarker. Confidence intervals for importance scores are calculated, providing quantitative indicators of reliability.

[0157] A comprehensive scoring model based on three dimensions—importance, stability, and universality—is established. The importance dimension integrates Gini and permutation importance scores from random forests and the absolute values ​​of standardized coefficients from Lasso regression, with principal component analysis determining the weight allocation for each indicator. The stability dimension is quantified by calculating the selection frequency of the marker in cross-validation, the coefficient variability in bootstrap resampling, and the consistency of performance in spatial block validation. The universality dimension evaluates the transfer performance of the marker across different marine datasets, including the preservation of prediction accuracy and the ranking stability of feature importance.

[0158] (1) The formula for calculating the overall score is:

[0159] Overall score = ×Importance Score+ ×Stability Score+ × Universality score, where the weighting coefficients were determined through expert review and actual verification results, and are set as follows: =0.4, =0.35, =0.25, ensuring both strong predictive power and robustness of results.

[0160] Weighting basis:

[0161] 1.1 Importance Dimension (Weight 0.40): Includes three sub-indicators:

[0162] Gini importance (40%): Feature importance derived from random forest;

[0163] Ranking importance (40%): Contribution verified through feature ranking;

[0164] Regression coefficients (accounting for 20%): Absolute values ​​of standardized coefficients in Lasso regression;

[0165] Importance score = 0.4 × Gini importance + 0.4 × Ranking importance + 0.2 × Regression coefficient.

[0166] 1.2 Stability Dimension (Weight 0.35): Includes three sub-indicators:

[0167] Cross-validation stability (40%): the frequency with which it is selected in k-fold cross-validation;

[0168] Bootstrap stability (40%): the reciprocal of the coefficient variability in Bootstrap resampling;

[0169] Spatial stability (20%): Consistency in performance across different spatial blocks;

[0170] Stability score = 0.4 × CV stability + 0.4 × Bootstrap stability + 0.2 × Spatial stability.

[0171] 1.3 Universality Dimension (Weight 0.25): Includes two sub-indicators:

[0172] Cross-oceanic mobility (60%): Predictive performance retention across other oceanic datasets;

[0173] Feature ranking consistency (40%): The stability of feature importance ranking across different sea areas;

[0174] Universality score = 0.6 × transferability + 0.4 × ranking consistency.

[0175] All sub-indices are standardized to the [0,1] range using min-max (normalization) to ensure uniformity of dimensions.

[0176] (2) Construction of a graded evaluation system: Based on the comprehensive score, the markers are divided into three grades:

[0177] 2.1 Primary core biomarkers: A comprehensive score ≥ 0.8, with excellent performance in all three dimensions. These biomarkers are typically microbial groups that are prevalent in multiple sea areas, have a clear metabolic association with the methane cycle, and are statistically highly stable.

[0178] 2.2 Secondary auxiliary biomarkers: These biomarkers have a comprehensive score between 0.6 and 0.8, and demonstrate outstanding performance in one or two dimensions. While they may be region-specific or coupled with other environmental factors, they still possess significant indicative value.

[0179] 2.3 Level III candidate biomarkers: These biomarkers have a composite score between 0.4 and 0.6 and exhibit predictive ability under specific conditions. These biomarkers require further validation and research, but may reveal new ecological indication patterns.

[0180] (3) Uncertainty Quantification and Reliability Assessment: Bayesian statistical methods were used to quantify the selection uncertainty of each biomarker. Markov chain Monte Carlo sampling was used to obtain the posterior distribution of biomarker importance scores, and 95% confidence intervals were calculated. For biomarkers selected by Lasso regression, the probability of their inclusion in the model was estimated using a stability selection method, providing a probabilistic reliability index.

[0181] 3.1 Comprehensive Uncertainty Indicators:

[0182] Construct a comprehensive uncertainty score: Uncertainty score = 0.4 × confidence interval width + 0.3 × (1 - stability selection probability) + 0.3 × cross-model coefficient of variation;

[0183] The confidence interval width is standardized to the [0,1] interval;

[0184] Stability selection probability takes the complement;

[0185] Cross-model coefficient of variation is used to calculate the degree of variation in importance scores among different algorithms.

[0186] 3.2 Reliability Classification Criteria:

[0187] High reliability: Uncertainty score < 0.3, confidence interval width < 0.2;

[0188] Medium reliability: 0.3 ≤ Uncertainty score < 0.6, confidence interval width < 0.4;

[0189] Low reliability: Uncertainty score ≥ 0.6, or confidence interval width ≥ 0.4;

[0190] Finally, a dynamic optimization mechanism for the biomarker system is constructed. This involves continuously collecting new sample data, periodically re-evaluating biomarker performance, and establishing version-based management of the biomarker system. Biomarkers that perform poorly in practical applications are downgraded or removed, while a fast-track verification mechanism is established for newly discovered potential biomarkers.

[0191] In summary, the method for mining methane leakage microbial biomarkers based on multi-source data fusion in this invention mainly includes four core steps: multi-source data acquisition and fusion, multi-dimensional feature engineering, multi-model collaborative biomarker mining, and a multi-level verification system. This invention, by introducing a machine learning algorithm combining random forest and Lasso regression, solves the problems of strong subjectivity, incomplete feature selection, and imperfect verification systems in traditional microbial biomarker mining methods, providing an objective, systematic, and repeatable intelligent identification and verification technology system for microbial biomarkers in natural gas hydrate exploration.

[0192] Compared with the prior art, the embodiments of the present invention have at least the following beneficial effects:

[0193] Advantages of multi-source data fusion: It breaks through the limitations of a single OTU abundance table and provides a more comprehensive and multi-dimensional perspective for biomarker mining by integrating spatial, environmental and cross-oceanic data;

[0194] Spatial context awareness: By introducing a spatial explicit machine learning model, microbial biomarkers with clear spatial distribution patterns can be discovered, providing direct guidance for regional exploration;

[0195] Nonlinear Relationship Capture: Using machine learning algorithms such as random forest, the complex nonlinear relationships and interactions between microorganisms and methane leakage are effectively identified, overcoming the limitations of traditional linear statistical methods;

[0196] Automated marker screening: By ranking features by importance and regularizing feature selection, the objectivity and automation of marker mining are achieved, eliminating the subjective bias of manual pre-grouping;

[0197] Multi-level reliability verification: Establish a triple verification system of internal verification, spatial verification and external verification to ensure the reliability and robustness of the markers in different dimensions;

[0198] Feature engineering richness: Through professional feature engineering techniques, more biologically meaningful and predictive feature variables are extracted from raw data, improving the depth and accuracy of biomarker mining.

[0199] like Figure 9 As shown, this embodiment of the invention also provides a methane leakage microbial biomarker mining device 900 based on multi-source data fusion, which can implement the above-described method. This device may include:

[0200] The first module 910 is used to acquire raw multi-source data of microbial genes, preprocess the raw multi-source data, and obtain a multi-source dataset.

[0201] The multi-source dataset includes a standardized OTU abundance table, spatial information, and environmental information containing methane leakage signals for each biomarker sample.

[0202] The second module 920 is used to construct a multi-dimensional feature matrix based on a multi-source dataset through feature engineering; wherein, the multi-dimensional feature matrix includes the multi-dimensional feature space corresponding to each marker sample;

[0203] The third module 930 is used to determine the first candidate marker set by using random forest and regression modeling to quantify the importance score of each marker sample based on a multi-dimensional feature matrix.

[0204] The fourth module 940 is used to determine the second candidate marker set by obtaining the stability score of each marker sample based on the multi-dimensional feature matrix through regression modeling and regularization parameter optimization quantization.

[0205] The fifth module 950 is used to collaboratively integrate the first candidate flag set and the second candidate flag set to obtain a comprehensive flag candidate set;

[0206] The sixth module 960 is used to obtain the universality score of each marker sample by migration verification quantization of marker bits based on the comprehensive marker candidate set.

[0207] Module 7, 970, is used to convert the importance score, stability score, and universality score to obtain the comprehensive score of each biomarker sample. Based on the comprehensive score, all biomarker samples are classified into biomarker levels to obtain the results of methane leakage microbial biomarkers.

[0208] It is understood that the content of the above method embodiments is applicable to the present device embodiments. The specific functions implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0209] This invention also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described above. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.

[0210] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0211] like Figure 10 As shown, Figure 10The hardware structure of an electronic device 1000 according to another embodiment is illustrated. The electronic device 1000 includes:

[0212] The processor 1001 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (aSIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present invention.

[0213] The memory 1002 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RaM). The memory 1002 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1002 and is called and executed by the processor 1001.

[0214] Input / output interface 1003 is used to implement information input and output;

[0215] The communication interface 1004 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0216] Bus 1005 transmits information between various components of the device (e.g., processor 1001, memory 1002, input / output interface 1003, and communication interface 1004);

[0217] The processor 1001, memory 1002, input / output interface 1003 and communication interface 1004 are connected to each other within the device via bus 1005.

[0218] The electronic device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0219] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.

[0220] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.

[0221] This invention also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.

[0222] It is understood that the content of the above method embodiments is applicable to the embodiments of this program product. The specific functions implemented by the embodiments of this program product are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0223] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0224] The present invention provides a method, apparatus, electronic device, storage medium, and program product for mining methane leakage microbial biomarkers based on multi-source data fusion. It acquires raw multi-source data of microbial genes, preprocesses the raw multi-source data to obtain a multi-source dataset. The multi-source dataset includes a standardized OTU abundance table, spatial information, and environmental information containing methane leakage signals for each biomarker sample. Based on the multi-source dataset, a multi-dimensional feature matrix is ​​constructed through feature engineering. The multi-dimensional feature matrix includes a multi-dimensional feature space corresponding to each biomarker sample. Based on the multi-dimensional feature matrix, the weight of each biomarker sample is quantified using random forest and regression modeling. Importance scores are used to determine the first candidate biomarker set; based on the multi-dimensional feature matrix, the stability score of each biomarker sample is obtained through regression modeling and regularization parameter optimization quantification, thereby determining the second candidate biomarker set; the first and second candidate biomarker sets are synergistically integrated to obtain a comprehensive biomarker candidate set; based on the comprehensive biomarker candidate set, the universality score of each biomarker sample is obtained through biomarker migration verification quantification; based on the importance score, stability score, and universality score, the comprehensive score of each biomarker sample is obtained; according to the comprehensive score, all biomarker samples are classified into biomarker levels to obtain the methane leakage microbial biomarker results. This invention, through the fusion of microbial genetic data, spatial information, and environmental geochemical data, constructs a multi-dimensional feature space, thereby improving the data dimensionality and reliability of biomarker mining. Furthermore, this invention employs a combination of random forest and regression modeling to achieve quantitative screening of biomarkers, reducing human intervention and subjective bias. Specifically, this invention classifies biomarkers based on a comprehensive score across three dimensions: importance, stability, and universality, facilitating priority selection in subsequent practical applications. This invention is applicable to methane leakage identification in different sea areas and geological backgrounds, possessing good scalability and practicality.

[0225] The embodiments described in this invention are for the purpose of more clearly illustrating the technical solutions of the embodiments of this invention, and do not constitute a limitation on the technical solutions provided by the embodiments of this invention. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this invention are also applicable to similar technical problems.

[0226] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present invention, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0227] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0228] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0229] The preferred embodiments of the present invention have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and spirit of the present invention should be within the scope of the claims of the present invention.

Claims

1. A method for mining microbial biomarkers of methane leakage based on multi-source data fusion, characterized in that, The method includes the following steps: Obtain raw multi-source data of microbial genes, and preprocess the raw multi-source data to obtain a multi-source dataset; The multi-source dataset includes a standardized OTU abundance table, spatial information, and environmental information containing methane leakage signals for each biomarker sample. Based on the multi-source dataset, a multi-dimensional feature matrix is ​​constructed through feature engineering; wherein, the multi-dimensional feature matrix includes a multi-dimensional feature space corresponding to each of the marker samples; Based on the multi-dimensional feature matrix, the importance score of each marker sample is obtained by using random forest and regression modeling, thereby determining the first candidate marker set. Based on the multi-dimensional feature matrix, the stability score of each biomarker sample is obtained through regression modeling and regularization parameter optimization quantization, thereby determining the second candidate biomarker set; The first candidate flag set and the second candidate flag set are collaboratively integrated to obtain a comprehensive flag candidate set; Based on the comprehensive biomarker candidate set, the universality score of each biomarker sample is obtained through migration verification quantization of biomarker bits; Based on the importance score, stability score, and universality score, a comprehensive score is obtained for each biomarker sample. According to the comprehensive score, all biomarker samples are classified into biomarker levels to obtain the results of methane leakage microbial biomarkers.

2. The method according to claim 1, characterized in that, The raw multi-source data includes microbiome data from public databases, as well as raw sequence data, environmental geochemical data, and spatial coordinate data determined based on sampling at preset sampling points. The preprocessing of the raw multi-source data to obtain a multi-source dataset includes the following steps: The original sequence data is subjected to quality control by removing low-quality sequences to obtain high-quality sequence data, and then the original OTU abundance table is established. The original sequence data is obtained by performing gene amplicon sequencing on the microbial gene data collected at the sampling points. The environmental geochemical data includes a concentration gradient curve established by the pore methane concentration corresponding to the sampling points and a methane leakage signal obtained by carbon isotope analysis of foraminifera shells. The spatial coordinate data is constructed by measuring water depth and sediment sampling depth using the latitude and longitude coordinates corresponding to the sampling points. The original OTU abundance table is subjected to a first normalization process to obtain the normalized OTU abundance table; wherein, the first normalization process includes center-log ratio transformation, threshold-based filtering of species with low abundance and low occurrence frequency, and sequencing depth normalization for the gene amplicon sequencing. Z-score standardization is performed on continuous environmental variables in the environmental geochemical data, and one-hot encoding is performed on categorical variables in the environmental geochemical data to obtain the environmental information. The spatial coordinate data is converted into a unified spatial reference system, and then spatial autocorrelation feature variables are constructed based on the spatial distance matrix between each sampling point to obtain the spatial information. The standardized OTU abundance table, the environmental information, and the spatial information are associated with the corresponding sampling points as measured biomarker samples. Extract public biomarker samples from the microbiome data in the public database; Based on the public marker samples, a cross-regional comparative analysis dataset is constructed by combining the measured marker samples, thus obtaining the multi-source dataset.

3. The method according to claim 1, characterized in that, The process of constructing a multi-dimensional feature matrix through feature engineering includes the following steps: Based on the standardized OTU abundance table, the relative abundance features of a single species are extracted and the phylogenetic diversity index is calculated based on the phylogenetic tree. Then, a microbial functional feature vector is constructed to obtain the microbial features. Based on the spatial information, a spatial weight matrix is ​​constructed using the distance decay criterion through spatial feature engineering to quantify spatial dependencies, and then multi-scale spatial pattern features are extracted to obtain spatial features. Based on the aforementioned environmental information, the polynomial characteristics of environmental factors are established using the principle of statistical significance, thereby obtaining environmental characteristics. The microbial features, spatial features, and environmental features are concatenated into a unified feature matrix to obtain the multi-dimensional feature space. The marker samples from different data sources are aligned using a multi-level matching strategy, and then the multi-dimensional feature matrix is ​​established by combining the multi-dimensional feature space of each marker sample.

4. The method according to claim 1, characterized in that, The process of determining the first candidate marker set by using random forest and regression modeling to quantify the importance score of each marker sample based on the multi-dimensional feature matrix includes the following steps: Based on the multi-dimensional feature space corresponding to all the aforementioned marker samples, a training set and a test set are obtained. Based on the training set, a microbial biomarker mining model is constructed using random forest; wherein, the microbial biomarker mining model includes a preset number of decision trees, and the training set of each decision tree is constructed by sampling biomarker samples with replacement using the self-sampling method; The Gini importance of each biomarker sample is determined based on the amount of impurity reduction corresponding to node splits in the decision tree, and then the first sub-score is obtained through normalization. The baseline performance metrics of the microbial biomarker mining model were obtained based on the test set. The feature values ​​of the target biomarker samples in the test set are randomly arranged to re-evaluate the model performance of the microbial biomarker mining model. Based on the baseline performance index and the performance record of the model, the step of randomly arranging the feature values ​​in the test set is returned to until the performance degradation of the preset number of groups is obtained. The second sub-score of the target biomarker sample is determined based on the average of all the stated performance degradation levels; Based on the multi-dimensional feature matrix, regression modeling is performed using Lasso regression, and then the regression coefficient of each biomarker sample is collected as the third sub-score. The importance score of the corresponding marker sample is obtained by performing a first weighted summation on the first sub-score, the second sub-score, and the third sub-score of each marker sample; Based on the importance score, the first set of candidate flags is obtained by comparison and filtering using a preset importance threshold.

5. The method according to claim 1, characterized in that, The process of determining the second candidate marker set by obtaining the stability score of each marker sample based on the multi-dimensional feature matrix through regression modeling and regularization parameter optimization includes the following steps: Based on the multi-dimensional feature matrix, multiple subsets of data are extracted with replacement using Bootstrap resampling. The regression modeling is performed by running Lasso regression independently on each of the said subsets of datasets; The stability of each biomarker sample in different subsets of data was evaluated by k-fold cross-validation, and the fourth sub-score was obtained by quantification. The fifth sub-score is determined based on the reciprocal of the coefficient variability of each of the marker samples in each Bootstrap resampling. Based on the multi-dimensional feature matrix, spatial cross-validation is used to quantify the consistency of the performance of the marker samples in different spatial blocks to determine the sixth sub-score. The stability score of the corresponding marker sample is obtained by performing a second weighted summation on the fourth sub-score, the fifth sub-score, and the sixth sub-score of each marker sample. Statistical analysis is performed on all the Bootstrap resampled subsets. When the selection ratio of a marker sample in all the subsets exceeds a preset ratio, the corresponding marker sample is included in the second candidate marker set.

6. The method according to claim 1, characterized in that, The process of obtaining the universality score of each biomarker sample based on the comprehensive biomarker candidate set through migration verification quantization of biomarker bits includes the following steps: The marker samples corresponding to the comprehensive marker candidate set are applied to different sea area datasets to quantify cross-sea area mobility and determine the seventh sub-score. The eighth sub-score is determined based on the consistency of feature ranking of the marker samples corresponding to the comprehensive marker candidate set in different sea area datasets. The seventh and eighth sub-scores of each marker sample are summed using a third weighted average to obtain the universality score of the corresponding marker sample.

7. The method according to any one of claims 1 to 6, characterized in that, The process of converting the importance score, stability score, and universality score into a comprehensive score for each biomarker sample, and then classifying all biomarker samples into biomarker rank based on the comprehensive score, includes the following steps: Based on preset weights, the importance score, the stability score, and the universality score are summed using a fourth weighted average to obtain the comprehensive score for each biomarker sample. Based on the comprehensive score, all the marker samples are divided into multiple marker levels using a preset score range.

8. A device for mining microbial biomarkers from methane leakage based on multi-source data fusion, characterized in that, The device includes: The first module is used to acquire raw multi-source data of microbial genes, and to preprocess the raw multi-source data to obtain a multi-source dataset; The multi-source dataset includes a standardized OTU abundance table, spatial information, and environmental information containing methane leakage signals for each biomarker sample. The second module is used to construct a multi-dimensional feature matrix based on the multi-source dataset through feature engineering; wherein, the multi-dimensional feature matrix includes a multi-dimensional feature space corresponding to each of the marker samples; The third module is used to determine the first candidate marker set by using random forest and regression modeling to quantify the importance score of each marker sample based on the multi-dimensional feature matrix. The fourth module is used to determine the second candidate marker set by obtaining the stability score of each marker sample based on the multi-dimensional feature matrix through regression modeling and regularization parameter optimization quantization. The fifth module is used to collaboratively integrate the first candidate flag set and the second candidate flag set to obtain a comprehensive flag candidate set; The sixth module is used to obtain the universality score of each of the marker samples by means of migration verification quantization of the marker bits based on the comprehensive marker candidate set. The seventh module is used to convert the importance score, the stability score and the universality score to obtain a comprehensive score for each biomarker sample, and to classify all biomarker samples into biomarker levels according to the comprehensive score to obtain the methane leakage microbial biomarker results.

9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method according to any one of claims 1 to 7.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 7.

Citation Information

Patent Citations

  • CN115064219A

  • CN116935973A