Method for identifying leakage strength of cold spring based on dominant species of perforated worms
By collecting and analyzing sediment samples from areas with high cold seep seepage intensity, and combining them with machine learning models, the intensity of cold seep seepage was identified. This solved the problem of accuracy in assessing the variation of benthic foraminifera species assemblage and enabled precise evaluation of cold seep seepage intensity.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-18
- Publication Date
- 2026-03-24
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In existing technologies, the variation of benthic foraminifera genera and species combinations under different cold seep infiltration intensities lacks specificity, which may lead to differences in genera and species combinations in different regions under similar environmental conditions, increasing the difficulty of accurately identifying the intensity of cold seep infiltration.
Sediment samples were collected from areas with different cold seep seepage intensities. Through the analysis of benthic foraminifera species characteristics and the determination of geochemical parameters, a machine learning model was constructed. An integrated identification model was established by combining multiple classification algorithms to identify the intensity level of cold seep methane seepage.
It improves the accuracy of identifying changes in benthic foraminifera genus and species combinations under different seepage intensities, provides an objective and repeatable assessment of cold seep activity intensity, and solves the problem of misjudgment caused by differences in genus and species combinations under similar environmental conditions.
Smart Images

Figure CN121725909A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of paleoceanography research technology, specifically a method for identifying the intensity of cold seep leakage based on dominant foraminifera species. Background Technology
[0002] In paleoceanography and biogeochemistry research, the decomposition of natural gas hydrates releases methane, which can potentially impact global climate change. Cold seep systems are key areas for methane release. By analyzing changes in the structure of benthic foraminifera communities in seabed sediments, the intensity and duration of cold seep methane leakage activities in historical periods can be effectively inferred. In cold seep activity areas, differences in methane leakage intensity can lead to corresponding changes in the sediment geochemical environment, thereby affecting the survival and shell formation of foraminifera.
[0003] In existing technologies, the changes in the species composition of benthic foraminifera under different cold seep intensities are not specific; that is, the species composition may differ in different regions under similar environmental conditions, increasing the difficulty of accurate identification. Therefore, how to use machine learning algorithms to analyze a large amount of foraminifera sample data from different regions and under different intensities, identify species composition patterns, and improve the accuracy of identifying changes in the species composition of benthic foraminifera under different cold seep intensities is the problem that this invention aims to solve. To this end, a method for identifying cold seep intensities based on dominant foraminifera species is proposed. Summary of the Invention
[0004] To solve the above-mentioned technical problems, the present invention provides the following technical solution: a method for identifying the intensity of cold seep seepage based on dominant foraminifera species, comprising the following steps: S1. Collect sediment column samples from areas with different cold seep seepage intensities, and perform washing, sieving and drying to obtain benthic foraminifera samples. S2. Identify the genus and species of benthic foraminifera in the sediment column samples, and compile genus and species characteristics including abundance, diversity and the proportion of dominant species to form a genus and species characteristic set. S3. Simultaneously determine the geochemical indices of sediment pore water and the stable carbon isotope composition of benthic foraminifera shells. S4. Based on geochemical parameters, assign a corresponding cold seep methane leakage intensity level label to each sample; S5. Extract genus and species features from the genus and species feature set and construct a structured dataset for machine learning training. S6. Use multiple classification algorithms to train the structured dataset, establish initial recognition models respectively, rank each initial recognition model, extract the top three, and merge them to establish an integrated recognition model; S7. Deploy the integrated identification model and apply the constructed integrated identification model to identify the leakage intensity of unknown samples in order to identify the leakage intensity level of cold seep methane.
[0005] Preferably, S1 specifically includes: Traversing the different cold seep seepage intensities to be studied, undisturbed sediment columnar samples of different known seepage intensities were collected using box-type or gravity samplers to ensure that the samples fully represent different seepage environments and to obtain complete original seabed sediment information. The collected sediment column samples were placed in the laboratory, physically separated, and gently rinsed with deionized water to remove salt and fine clay particles, avoiding damage to the integrity of the foraminifera shell structure. The washed sediment column sample was wet-sieved through a standard sieve to collect the target particle size fraction, and the fraction below 50 μm was collected. o Drying at a constant temperature of C yields benthic foraminifera samples for identification.
[0006] Preferably, S2 specifically includes: The prepared benthic foraminifera samples for identification were placed under a microscope, and the benthic foraminifera were preliminarily classified according to their morphological characteristics. Different genera and species were distinguished and the preliminary observation results were recorded. Based on classification literature and illustrations, the genera and species of the preliminarily classified benthic foraminifera were identified, and the genera and species characteristics of each genera and species, including abundance, diversity and the proportion of dominant species, were statistically analyzed. The statistical data on genus and species characteristics are integrated into a unified database to form a structured set of genus and species characteristics that includes information on abundance, diversity, and the proportion of dominant species.
[0007] Preferably, S3 specifically includes: Low-temperature centrifugation was performed on the sediment column samples to separate and obtain pore water. At the same time, sufficient and clean benthic foraminifera shells (single dominant species) were selected from the sieved benthic foraminifera samples. The shells of benthic foraminifera of the target particle size were separated by wet sieving. After acid washing to remove carbonate impurities, the shells were extracted and purified using a vacuum condensation system to ensure that the samples for isotope analysis were pure and uncontaminated. The concentration of sulfate ions in pore water was determined using ion chromatography, and the methane content was analyzed using gas chromatography. Simultaneously, isotope mass spectrometry was used to test the shells of acid-etched benthic foraminifera to determine their stable carbon isotope composition, where the stable carbon isotope is δ¹⁸. 13 C.
[0008] Preferably, S4 specifically includes: Geochemical parameters, including methane content and sulfate concentration, were collected from each sediment column sample. Based on the correlation between geochemical parameters and cold seep methane leakage intensity, a leakage intensity classification standard was established, defining the threshold range and combination pattern of each geochemical parameter corresponding to different levels. The leakage intensity levels include weak leakage level, moderate leakage level, and strong leakage level. Based on the established grading standards, each sediment column sample was assigned a cold seep methane leakage intensity rating label, and cross-validation was used to ensure the accuracy and rationality of the label assignment.
[0009] Preferably, S5 specifically includes: From the established set of genera and species characteristics, extract genera and species characteristics including abundance, diversity and proportion of dominant species, and perform standardization and normalization to ensure uniform data scale. The standardized and normalized genus and species features are combined into feature vectors and organized into a table format according to machine learning requirements. This ensures that each feature vector corresponds to the assigned cold seep methane leakage intensity level label, forming a structured dataset with clearly defined feature columns and label columns, and dividing the dataset into training, validation, and test sets.
[0010] Preferably, S6 specifically includes: Multiple classification algorithms are trained in parallel, including random forest, gradient boosting decision tree, support vector machine, multilayer perceptron, XGBoost and K-nearest neighbor algorithm, and multiple initial recognition models are trained based on structured datasets and each classification algorithm; Based on the comprehensive performance of multiple initial recognition models on the validation set, the top three initial recognition models are extracted, and the stacked ensemble method is used to merge the initial recognition models to construct an integrated recognition model, which integrates the advantages of different algorithms and improves the comprehensive recognition capability. The integrated identification model is evaluated based on the test set to confirm its accuracy and stability, and then the integrated identification model is output as a reliable identification model for leakage intensity identification.
[0011] Preferably, the process of constructing the integrated recognition model is as follows: Based on the validation set data, a comprehensive performance evaluation was conducted on multiple initial recognition models. Multi-dimensional indicators such as accuracy, recall, F1 score, and AUC-ROC curve were used to quantify model performance. The initial recognition models were sorted in descending order according to their comprehensive scores, and the top three performing initial recognition models were selected as ensemble candidates to ensure that their classification ability and stability reached a high level. Among them, the weights of accuracy and recall were both 0.3, and the weights of F1 score and AUC-ROC curve were both 0.2, with a focus on balancing overall accuracy and class recall. The stacked ensemble method is adopted, which uses the top three initial recognition models as base learners, utilizes their prediction results on the validation set, i.e., the predicted category (One-Hot encoding), and outputs them to construct meta-features. The meta-features constructed from the prediction results of each base learner are concatenated into a new feature vector to form a meta-feature set. A logistic regression-based meta-classifier is trained on the meta-feature set to complete the secondary fusion of the prediction results of the base learners. By optimizing the hyperparameters of the meta-classifier, an ensemble recognition model is constructed to output the final predicted category.
[0012] Preferably, the expression for the predicted category is as follows: ; In the formula: The final predicted category of the integrated identification model is the cold spring seepage intensity level of the sample output by the model. The operator for taking the independent variable k that maximizes the function value indicates that the category with the highest probability is selected as the final output. This is a set of intensity levels for cold spring seepage. ; Represents the weight vector for category k The transpose of is used as a parameter to calculate the score for category k; Represents the weight vector for category j The transpose of is used to calculate the score for category j; j represents the index of the category, used for traversal. All categories; This represents the bias term for category k, a parameter used to adjust the score of category k; This represents the bias term for category j, a parameter used to adjust the score of category j; This indicates the prediction output of the first base learner for the current sample, which is the predicted category after One-Hot encoding; This indicates that the prediction output of the second base learner for the current sample is the predicted category after One-Hot encoding; This indicates the prediction output of the third base learner for the current sample, which is the predicted category after One-Hot encoding.
[0013] Preferably, S7 specifically includes: Serialize the trained ensemble recognition model (including base learners and meta-classifiers) into a loadable format (.pkl or .h5), deploy it to the production environment server, configure the input data preprocessing process, and complete the environment dependency installation and model initialization. The sediment column sample data collected from areas with unknown cold seep seepage intensity were standardized to unify data dimensions and format. Preprocessing operations consistent with those in the training phase were performed to generate feature vectors that meet the model input requirements, ensuring data quality and model compatibility. The features of sediment column samples from areas with unknown cold seep seepage intensity after preprocessing are input into the deployed integrated recognition model. Meta-features are generated sequentially through the base learner, and then the meta-classifier outputs the category label of the cold seep methane seepage intensity level. The prediction results and confidence levels are recorded, and a visualization report is generated simultaneously.
[0014] This invention provides a method for identifying the intensity of cold seep seepage based on dominant foraminifera species. It has the following beneficial effects: (i) The method for identifying the intensity of cold seep seepage based on dominant foraminifera species collects sediment samples from areas with different cold seep seepage intensities, analyzes the genus and species combination characteristics of benthic foraminifera, and constructs an identification model by combining machine learning algorithms. Compared with traditional methods that rely on the genus and species combination characteristics of a single or limited area, this method improves the accuracy of the model in identifying changes in the genus and species combination of benthic foraminifera under different seepage intensities by training with multi-regional data and ensemble learning technology, and solves the problem of misjudgment caused by differences in genus and species combination under similar environmental conditions.
[0015] (II) This method for identifying the intensity of cold seep leakage based on dominant foraminifera species establishes a quantitative grading standard for cold seep methane leakage intensity by simultaneously measuring the geochemical indicators of sediment pore water and the stable carbon isotope composition of foraminifera shells. Combined with genus and species characteristic data, it can assign a precise leakage intensity level label to each sample, avoiding the subjectivity of traditional qualitative analysis and providing an objective and repeatable scientific basis for assessing the intensity of cold seep activity. Attached Figure Description
[0016] Figure 1 This is a schematic diagram of the workflow of a method for identifying the intensity of cold seep leakage based on dominant foraminifera species according to the present invention. Figure 2 This is a flowchart of a method for identifying the intensity of cold spring seepage based on a dominant foraminifera species according to the present invention. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] Example 1, please refer to Figure 1 , Figure 2This invention provides a technical solution: a method for identifying the intensity of cold seep seepage based on dominant foraminifera species, comprising the following steps: S1. Collect sediment column samples from areas with different cold seep seep intensities, and perform washing, sieving, and drying to obtain benthic foraminifera samples. Traverse the different cold seep seep seep intensities under study, using box-type or gravity samplers to collect undisturbed sediment column samples from areas with different known seep intensities, ensuring that the samples fully represent different seep environments and completely obtain original seafloor sedimentary information. Place the collected sediment column samples in the laboratory for physical segmentation and gently rinse with deionized water to remove salt and fine clay particles, avoiding damage to the integrity of the foraminifera shell structure. Wet-sieve the washed sediment column samples through a standard sieve set to collect target particle size components, with those below 50 μm being collected. o Drying at a constant temperature of C yielded benthic foraminifera samples for identification. The specific work involves: within the multiple cold seep activity areas to be studied, based on existing geochemical survey data, selecting typical stations with varying known seepage intensities (high, medium, and low), and collecting undisturbed sediment column samples using box-type or gravity samplers. During sampling, the integrity of the sediment column is ensured to avoid stratification, fracturing, or human disturbance. A multi-point, multi-depth systematic sampling approach is employed to cover areas with varying seepage intensities (high, medium, and low). Simultaneously, water depth, temperature, and surrounding cold seep activity indicators (e.g., bubble escape and biological community distribution) at the sampling points are recorded. After transferring the collected sediment column samples to the laboratory, they are first physically segmented, cut into sub-samples at 5cm intervals. Community changes on the vertical profile are analyzed. The segmented samples are then gently rinsed with deionized water to remove surface-attached salts and fine clay particles, preventing salt crystallization or clay clogging of the sieves during subsequent analysis. The water flow intensity is controlled during rinsing to avoid damaging the porous structure. The complete structure of the foraminifera shells, especially the vitreous or ceramic shells, was ensured. After cleaning, the samples were wet-sieved through a standard sieve (63 μm to 1 mm particle size range) to separate the target particle size fraction rich in foraminifera. The wet sieve process was repeated 2-3 times to ensure separation efficiency. Finally, the residue was dried in a constant temperature oven below 50°C to prevent the decomposition of carbonates or degradation of organic matter in the shells due to high temperature. The dried samples were further processed to prepare specimens for identification. Residual particles attached to the shell surface were removed using an air blowing device. The samples were briefly soaked in a low concentration of hydrogen peroxide (3%) to dissolve organic matter. The time was controlled (<1 minute), and the reaction was immediately terminated by rinsing with deionized water. The processed samples were then placed on a microscope slide and sealed with glycerol or chloroform hydrate for long-term preservation and observation. Before identification, the samples were randomly sampled and re-examined to ensure that there was no shell damage or contamination. Finally, benthic foraminifera samples that could be used for identification were obtained. S2. Identify the species and genera of benthic foraminifera in the sediment column samples, and statistically analyze the species and genera characteristics, including abundance, diversity, and dominant species ratio, to form a species and genera characteristic set. Place the prepared benthic foraminifera samples for identification under a microscope, and perform preliminary classification of benthic foraminifera according to morphological characteristics. Distinguish between different genera and species and record the preliminary observation results. Based on classification literature and illustrations, identify the species and genera of the preliminarily classified benthic foraminifera, and statistically analyze the species and genera characteristics of each genera and species, including abundance, diversity, and dominant species ratio. Integrate the statistical species and genera characteristic data into a unified database to form a structured species and genera characteristic set containing information on abundance, diversity, and dominant species ratio. The specific work involves: placing the prepared benthic foraminifera samples under a microscope and using the microscope's high-resolution imaging function to observe the shell morphology and structural characteristics of the benthic foraminifera, including shell shape (round, elliptical, spiral, etc.), wall thickness (thin, medium, thick), and pore morphology (round, elliptical, slit-like, etc.). Based on morphological and structural characteristics, a preliminary classification of benthic foraminifera is performed, grouping morphologically similar individuals into one category and distinguishing different genera and species. During the classification process, preliminary observation results are recorded, including the approximate number of each genera and species, morphological characteristics, and significant differences from other genera and species. The final classification is based on authoritative taxonomic literature, standard atlases, and established regional foraminifera fossil atlases. The genus and species of the preliminarily classified benthic foraminifera were identified by comparing detailed descriptions and illustrations in literature and atlases, combined with microscopic observations. After identification, the specific genus and species characteristics of each genus and species were statistically analyzed, including abundance (the number of individuals of that genus and species per unit area or volume), heterogeneity (the number of different genera and species in the sample), and dominant species ratio (the proportion of individuals of one or more dominant genera and species to the total number of individuals). The statistically obtained genus and species characteristic data were systematically organized and integrated into a unified database. The data format was standardized to map the abundance, heterogeneity, and dominant species ratio to the corresponding genera, species, and samples, thus forming a structured set of genus and species characteristics. S3. Simultaneously determine the geochemical parameters of sediment pore water and the stable carbon isotope composition of benthic foraminifera shells. The sediment column samples were centrifuged at low temperature to separate pore water. Simultaneously, sufficient and clean benthic foraminifera shells (single dominant species) were selected from the sieved benthic foraminifera samples. Target-size benthic foraminifera shells were separated using a wet sieving method. After acid washing to remove carbonate impurities, the shells were purified using a vacuum condensation system to ensure the purity and lack of contamination in the isotope analysis samples. The sulfate ion concentration in the pore water was determined using ion chromatography, and the methane content was analyzed using gas chromatography. Simultaneously, isotope mass spectrometry was used to test the acid-etched benthic foraminifera shells to determine their stable carbon isotope composition, where the stable carbon isotope is δ¹⁸. 13 C; The specific work involved: performing low-temperature centrifugation on the collected sediment column samples. Under low-temperature conditions, pore water was separated and obtained through centrifugation. The low-temperature environment reduces changes in material properties and ensures the stability of the pore water composition. Simultaneously, sufficient and clean benthic foraminifera shells were selected from the sieved benthic foraminifera samples, prioritizing single dominant species. Wet sieving was used to separate benthic foraminifera shells of the target particle size, obtaining shells within the required size range to ensure sample consistency. Acid washing was then performed to effectively remove carbonate impurities from the shell surface, preventing interference from impurities in subsequent analyses. The shells were extracted and purified using a vacuum condensation system. Under vacuum, the condensation process further purified the shells, ensuring the purity and lack of contamination in the isotope analysis samples. Ion chromatography was used to determine the sulfate ion concentration in the separated pore water, which reflects the chemical environment of the cold seep area. Gas chromatography was used to analyze the methane content in the pore water. Methane is an important marker of cold seep activity, and its content indicates the intensity of cold seep activity. Simultaneously, isotope mass spectrometry was used to test the acid-etched benthic foraminifera shells, determining their stable carbon isotopes (δ¹²⁻¹). 13 C) Composition, which provides information about biogeochemical cycles and paleoenvironmental changes through stable carbon isotope composition; S4. Based on geochemical parameters, assign a corresponding cold seep methane leakage intensity level label to each sample. Collect geochemical parameters of each sediment column sample, including methane content and sulfate concentration. Based on the correlation between geochemical parameters and cold seep methane leakage intensity, establish a leakage intensity level classification standard, define the threshold range and combination pattern of each geochemical parameter corresponding to different levels. The leakage intensity level includes weak leakage level, moderate leakage level and strong leakage level. Based on the established classification standard, assign a cold seep methane leakage intensity level label to each sediment column sample, and ensure the accuracy and rationality of the label assignment through cross-validation. The specific work involves collecting geochemical parameters from various sediment columns, accurately determining the methane content and sulfate ion concentration in the pore water. Methane, as a direct product of cold seep seepage, is represented by its concentration, indicating seepage flux. Sulfate ions, as reactants in the anoxic oxidation process of methane, are consumed to directly reflect the intensity of this biogeochemical process. Based on the established causal relationship between geochemical parameters and cold seep methane seepage intensity, a grading standard for seepage intensity is established. Threshold ranges and specific combination patterns of geochemical parameters corresponding to different seepage intensity levels (divided into weak, moderate, and strong levels) are defined. The weak seepage level is characterized by high sulfate and low methane concentrations. The combination of strong leakage and high methane concentration is used to classify each sediment column sample as a cold seep methane leakage intensity level. Based on the established grading criteria, each sample is assigned a cold seep methane leakage intensity level label. For each sample, the measured methane content and sulfate ion concentration data are precisely compared with the threshold ranges and combination patterns defined in the grading criteria for each leakage intensity level, thus classifying it into the most appropriate leakage intensity level: weak leakage, moderate leakage, or strong leakage. The accuracy and reasonableness of the labeling process are cross-validated by comparing the consistency of chemical profiles at different depths of the same sample, and verifying the stable carbon isotopes (δ¹⁸O) of foraminifera shells. 13 C) Whether the data matches the leakage intensity level label determined based on pore water chemistry, ensuring that the final leakage intensity level label can truly reflect the deposition environment when the sample was formed; S5. Extract genus and species features from the genus and species feature set and construct a structured dataset for machine learning training. S6. Use multiple classification algorithms to train the structured dataset, establish initial recognition models respectively, rank each initial recognition model, extract the top three, and merge them to establish an integrated recognition model; S7. Deploy the integrated identification model and apply the constructed integrated identification model to identify the leakage intensity of unknown samples in order to identify the leakage intensity level of cold seep methane.
[0019] Example 2, as Figure 1 , Figure 2 As shown, based on Example 1, the present invention provides a technical solution: S5 specifically includes: extracting species features including abundance, diversity and dominant species ratio from the established species feature set, performing standardization and normalization processing to ensure data scale uniformity, combining the standardized and normalized species features into feature vectors, and organizing them into a table format according to machine learning requirements to ensure that each feature vector corresponds to the assigned cold seep methane leakage intensity level label, forming a structured dataset, clarifying feature columns and label columns, and dividing the training set, validation set and test set; The specific work involves: extracting species feature parameters covering abundance, diversity, and dominant species ratio from the established system-built genus and species feature set; preprocessing the extracted genus and species features using standardization (Z-score standardization, making the data mean 0 and standard deviation 1) and normalization (Min-Max normalization, scaling the data to the [0,1] interval) methods to ensure that all feature data are on a uniform scale and enhance the comparability between features; combining the preprocessed abundance, diversity, and dominant species ratio features in a preset order to form a multidimensional feature vector; organizing the feature vector in a table format according to the input requirements of the machine learning model, where each row represents a sediment column sample and each column corresponds to a feature parameter, ensuring that each feature vector accurately corresponds to the previously assigned cold seep methane leakage intensity level label; constructing a structured dataset containing feature columns and label columns; and dividing the dataset proportionally into a training set (for model parameter learning), a validation set (for model hyperparameter tuning), and a test set (for final model performance evaluation) using a random partitioning strategy to ensure that the data distribution of each subset is consistent. S6 specifically includes: parallel training of multiple classification algorithms, including random forest, gradient boosting decision tree, support vector machine, multilayer perceptron, XGBoost and K-nearest neighbor algorithm; training multiple initial recognition models based on structured datasets and each classification algorithm; ranking the multiple initial recognition models according to their comprehensive performance on the validation set; extracting the top three performance initial recognition models; using a stacked ensemble method to fuse the initial recognition models to construct an integrated recognition model; integrating the advantages of different algorithms to improve comprehensive recognition ability; performing a final performance evaluation on the fused integrated recognition model based on the test set to confirm its accuracy and stability; and finally outputting the integrated recognition model as a reliable recognition model for leakage intensity identification. The specific work involves: simultaneously training multiple classic and cutting-edge classification algorithms using a multi-algorithm parallel training strategy, including: Random Forest, which improves generalization ability by constructing multiple decision trees and combining the results; Gradient Boosting Decision Tree, which iteratively optimizes residuals to gradually enhance model performance; Support Vector Machine, which achieves classification by finding the optimal hyperplane; Multilayer Perceptron, which, as the basic structure of neural networks, can handle complex nonlinear relationships; XGBoost, an optimized implementation of the gradient boosting framework, which has efficient computation and regularization characteristics; and the K-Nearest Neighbors algorithm, which classifies based on neighboring sample voting. Based on the constructed structured dataset, multiple initial recognition models are trained using the characteristics of each classification algorithm. After the initial recognition models are trained, they are ranked according to their comprehensive performance on the validation set, and the evaluation metrics include accuracy and recall. The F1 score and AUC-ROC curve comprehensively reflect the classification ability and stability of the initial identification model. Through quantitative comparison, the top three initial identification models are extracted to ensure that the fusion model is based on the optimal candidate. Then, the stacked ensemble method is used to fuse the initial identification models. The top three initial identification models are used as base learners, and the predicted outputs of multiple base learners are used as new meta-features. On this basis, a meta-classifier is trained to perform the final decision, which is the ensemble identification model. After the ensemble identification model is built, the final performance is evaluated based on the test set. The accuracy of the ensemble identification model in the three-level classification of cold seep methane leakage intensity is analyzed. If the ensemble identification model achieves an accuracy of ≥90% on the test set and a balanced recall rate for each leakage intensity level, it is confirmed as a reliable identification model and output for actual leakage intensity identification tasks. The process of constructing the ensemble recognition model is as follows: Based on the validation set data, a comprehensive performance evaluation is performed on multiple initial recognition models. Multi-dimensional metrics such as accuracy, recall, F1 score, and AUC-ROC curve are used to quantify model performance. The initial recognition models are ranked in descending order based on their comprehensive scores, and the top three performing initial recognition models are selected as ensemble candidates to ensure that their classification ability and stability are both at a high level. Accuracy and recall are each weighted at 0.3, while F1 score and AUC-ROC curve are each weighted at 0.2. The focus is on balancing overall accuracy and class recall, using a stacking approach. The ensemble approach uses the top three initial recognition models as base learners, leveraging their predictions on the validation set (i.e., predicted categories, One-Hot encoding) to construct meta-features. The meta-features constructed from the predictions of each base learner are concatenated into a new feature vector, forming a meta-feature set. A logistic regression-based meta-classifier is trained on this meta-feature set to achieve a secondary fusion of the base learner predictions. By optimizing the hyperparameters of the meta-classifier, its ability to fit complex classification boundaries is improved. The ensemble recognition model then outputs the final predicted category, achieving a significant improvement in classification performance by integrating the advantages of the base learners. Furthermore, the expression for predicting the category is as follows: ; In the formula: The final predicted category of the integrated identification model is the cold spring seepage intensity level of the sample output by the model. The operator for taking the independent variable k that maximizes the function value indicates that the category with the highest probability is selected as the final output. This is a set of intensity levels for cold spring seepage. ; Represents the weight vector for category k The transpose of is used as a parameter to calculate the score for category k; Represents the weight vector for category j The transpose of is used to calculate the score for category j; j represents the index of the category, used for traversal. All categories; This represents the bias term for category k, a parameter used to adjust the score of category k; This represents the bias term for category j, a parameter used to adjust the score of category j; This indicates the prediction output of the first base learner for the current sample, which is the predicted category after One-Hot encoding; This indicates that the prediction output of the second base learner for the current sample is the predicted category after One-Hot encoding; This indicates that the prediction output of the third base learner for the current sample is the predicted category after One-Hot encoding. When multiple base learners make consistent predictions for a certain category k (i.e., the predicted category after One-Hot encoding has more 1s at position k), the signal of the corresponding category in the meta-features received by the meta-classifier is stronger. After weighted summation and Softmax, the probability of category k will be the highest, and thus it will be selected as the final result. S7 specifically includes: serializing the trained ensemble recognition model (including base learners and meta-classifiers) into a loadable format (.pkl or .h5), deploying it to the production environment server, configuring the input data preprocessing process, ensuring seamless integration between the model interface and the data pipeline, completing environment dependency installation and model initialization, standardizing the sediment column sample data collected from areas with unknown cold seep seepage intensity, unifying data dimensions and formats, performing preprocessing operations consistent with the training phase, generating feature vectors that meet the model input requirements, ensuring data quality and model compatibility, inputting the preprocessed sediment column sample features from areas with unknown cold seep seepage intensity into the deployed ensemble recognition model, generating meta-features sequentially through the base learners, and then having the meta-classifier comprehensively output the category label of the cold seep methane seepage intensity level, recording the prediction results and confidence levels, and simultaneously generating a visualization report; The specific tasks are as follows: Serialize the trained ensemble recognition model (including base learners and meta-classifiers) into a loadable format, preferably .pkl (Python object serialization) or .h5 (HDF5 format, supporting deep learning models). Use pickle or joblib libraries to save the model structure, parameters, and dependent meta-feature construction logic as a whole, ensuring the complete encapsulation of the prediction functions of the base learners and the weight information of the meta-classifiers. When deploying to the production environment server, configure a standardized runtime environment, including installing specified versions of Python, CUDA, Scikit-learn, NumPy, and other dependent libraries, and isolate environment conflicts using Docker containerization technology. During the initialization phase, load the serialized model, verify the compatibility of the input / output interfaces of the base learners and meta-classifiers, and establish a logging system to monitor the model loading status and initial prediction performance, ensuring the deployment process is reproducible and without compatibility issues. For sediment column sample data collected from areas with unknown cold seep seepage intensity, perform the same preprocessing procedure as the training phase. Data cleaning is performed to handle missing values, outliers, and unit standardization. Based on the feature engineering rules of the training set, the raw data is dimensionality reduced, normalized, or encoded to generate feature vectors that match the model input dimensions. Simultaneously, a data quality verification mechanism is established to automatically detect the dimension, type, and value range of the input data, ensuring that the preprocessed feature vectors are fully compatible with the data structure used during model training. The preprocessed feature vectors are then input into the deployed ensemble recognition model, where three base learners generate initial predicted categories and probability distributions in parallel, which are then concatenated into a meta-feature vector. Subsequently, the meta-classifier performs secondary fusion based on the meta-features, outputting the final category label (weak / moderate / strong) and confidence score for the intensity level of the cold spring methane leak. Simultaneously, the independent prediction results of each base learner, the meta-feature concatenation logic, and the decision basis of the meta-classifier are recorded, generating a structured log file. A visualization report module is also developed to display the predicted categories and confidence scores in chart form, and to annotate the comparison analysis between the prediction results and historical data. The output includes the original data identifier, prediction timestamp, model version number, and visualization report path.
[0020] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0021] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for identifying the intensity of cold seep seepage based on dominant foraminifera species, characterized in that, Includes the following steps: S1. Collect sediment column samples from areas with different cold seep seepage intensities and process them to obtain benthic foraminifera samples. S2. Identify the genus and species of benthic foraminifera in the sediment column samples, and compile genus and species characteristics including abundance, diversity and the proportion of dominant species to form a genus and species characteristic set. S3. Simultaneously determine the geochemical indices of sediment pore water and the stable carbon isotope composition of benthic foraminifera shells. S4. Based on geochemical parameters, assign a corresponding cold seep methane leakage intensity level label to each sample; S5. Extract genus and species features from the genus and species feature set and construct a structured dataset for machine learning training. S6. Use multiple classification algorithms to train the structured dataset, establish initial recognition models respectively, rank each initial recognition model, extract the top three, and merge them to establish an integrated recognition model; S7. Deploy the integrated identification model and apply the constructed integrated identification model to identify the leakage intensity of unknown samples in order to identify the leakage intensity level of cold seep methane.
2. The method for identifying cold seep seepage intensity based on dominant foraminifera species according to claim 1, characterized in that: S1 specifically includes: Traversing the regions with different cold seep seepage intensities to be studied, undisturbed sediment columnar samples were collected from regions with different known seepage intensities using box-type or gravity samplers. The collected sediment column samples were placed in the laboratory, physically separated, and gently rinsed with deionized water to remove salts and fine clay particles. The washed sediment column sample was wet-sieved through a standard sieve to collect the target particle size fraction, and the fraction below 50 μm was collected. o Drying at a constant temperature of C yields benthic foraminifera samples for identification.
3. The method for identifying the intensity of cold seep seepage based on the dominant species of foraminifera according to claim 2, characterized in that: S2 specifically includes: The prepared benthic foraminifera samples for identification were placed under a microscope, and the benthic foraminifera were preliminarily classified according to their morphological characteristics. Different genera and species were distinguished and the preliminary observation results were recorded. Based on classification literature and illustrations, the genera and species of the preliminarily classified benthic foraminifera were identified, and the genera and species characteristics of each genera and species, including abundance, diversity and the proportion of dominant species, were statistically analyzed. The statistical data on genus and species characteristics are integrated into a unified database to form a structured set of genus and species characteristics that includes information on abundance, diversity, and the proportion of dominant species.
4. The method for identifying the intensity of cold seep seepage based on dominant foraminifera species according to claim 1, characterized in that: S3 specifically includes: Low-temperature centrifugation was performed on the sediment column samples to separate and obtain pore water, and benthic foraminifera shells were simultaneously selected from the sieved benthic foraminifera samples. The shells of benthic foraminifera of the target particle size were separated by wet sieving. After acid washing to remove carbonate impurities, the shells were extracted and purified using a vacuum condensation system. The concentration of sulfate ions in pore water was determined using ion chromatography, and the methane content was analyzed using gas chromatography. Simultaneously, isotope mass spectrometry was used to test the shells of acid-etched benthic foraminifera to determine their stable carbon isotope composition, where the stable carbon isotope is δ¹⁸. 13 C.
5. The method for identifying the intensity of cold seep seepage based on dominant foraminifera species according to claim 1, characterized in that: S4 specifically includes: Geochemical parameters, including methane content and sulfate concentration, were collected from each sediment column sample. Based on the correlation between geochemical parameters and cold seep methane leakage intensity, a leakage intensity classification standard was established, defining the threshold range and combination pattern of each geochemical parameter corresponding to different levels. The leakage intensity levels include weak leakage level, moderate leakage level, and strong leakage level. Based on the established grading criteria, each sediment column sample was assigned a cold seep methane leakage intensity rating label.
6. The method for identifying the intensity of cold seep seepage based on dominant foraminifera species according to claim 1, characterized in that: S5 specifically includes: From the established set of genus and species characteristics, extract genus and species characteristics including abundance, diversity and proportion of dominant species, and perform standardization and normalization. The standardized and normalized genus and species features are combined into feature vectors and organized into a table format according to machine learning requirements. This ensures that each feature vector corresponds to the assigned cold seep methane leakage intensity level label, forming a structured dataset with clearly defined feature columns and label columns, and dividing the dataset into training, validation, and test sets.
7. The method for identifying the intensity of cold seep seepage based on dominant foraminifera species according to claim 1, characterized in that: S6 specifically includes: Multiple classification algorithms are trained in parallel, including random forest, gradient boosting decision tree, support vector machine, multilayer perceptron, XGBoost and K-nearest neighbor algorithm, and multiple initial recognition models are trained based on structured datasets and each classification algorithm; Based on the overall performance of multiple initial recognition models on the validation set, the top three initial recognition models are extracted, and the stacked ensemble method is used to fuse the initial recognition models to construct an integrated recognition model. The integrated identification model is then evaluated based on the test set, and the integrated identification model is output as a reliable identification model for leakage intensity identification.
8. The method for identifying the intensity of cold seep seepage based on dominant foraminifera species according to claim 7, characterized in that: The process of constructing the integrated recognition model is as follows: Based on the validation set data, a comprehensive performance evaluation was conducted on multiple initial recognition models. The model performance was quantified using multi-dimensional indicators such as accuracy, recall, F1 score, and AUC-ROC curve. The initial recognition models were sorted in descending order according to their comprehensive scores, and the top three performing initial recognition models were selected as integration candidates. The weights of accuracy and recall were both 0.3, and the weights of F1 score and AUC-ROC curve were both 0.
2. The stacked ensemble method is adopted, which uses the top three initial recognition models as base learners, uses their prediction results on the validation set to predict the category, and outputs meta-features. The meta-features constructed by the prediction results of each base learner are concatenated into a new feature vector to form a meta-feature set. A logistic regression-based meta-classifier is trained on the meta-feature set to complete the secondary fusion of the prediction results of the base learners. By optimizing the hyperparameters of the meta-classifier, an ensemble recognition model is constructed to output the final predicted category.
9. The method for identifying the intensity of cold seep seepage based on dominant foraminifera species according to claim 8, characterized in that: The expression for the predicted category is as follows: ; In the formula: The final predicted category of the integrated identification model is the cold spring seepage intensity level of the sample output by the model. The operator that takes the value of the function by maximizing the value of the independent variable k; This is a set of intensity levels for cold spring seepage. ; Represents the weight vector for category k transpose; Represents the weight vector for category j The transpose of; j represents the index of the category, used for traversal. All categories; This represents the bias term with respect to category k; This represents the bias term for category j; This represents the prediction output of the first base learner for the current sample; This represents the prediction output of the second base learner for the current sample; This represents the prediction output of the third base learner for the current sample.
10. The method for identifying the intensity of cold seep seepage based on dominant foraminifera species according to claim 1, characterized in that: Specifically, S7 includes: The trained ensemble recognition model is serialized into a loadable format, deployed to the production environment server, the input data preprocessing process is configured, and the environment dependency installation and model initialization are completed. The sediment column sample data collected from areas with unknown cold seep seepage intensity were standardized and preprocessed in the same way as during the training phase to generate feature vectors that meet the model input requirements. The features of sediment column samples from areas with unknown cold seep seepage intensity after preprocessing are input into the deployed integrated recognition model. Meta-features are generated sequentially through the base learner, and then the meta-classifier outputs the category label of the cold seep methane seepage intensity level. The prediction results and confidence levels are recorded, and a visualization report is generated simultaneously.