Seawater trace gas identification method based on multi-chromatographic column cooperation and machine learning
By combining multi-column collaboration and machine learning, along with multi-dimensional peak feature extraction and dual-model fusion, the problems of insufficient sensitivity and misjudgment in seawater trace gas identification were solved, achieving high-precision and rapid gas identification and detection.
Patent Information
- Application Number
- CN202610007254.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-06
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2046-01-06
AI Technical Summary
Existing seawater trace gas identification technologies suffer from problems such as insufficient sensitivity, difficulty in identification due to co-escape of multiple gas components, misjudgments caused by column temperature drift, weak anti-interference ability of machine learning models, and cumbersome detection procedures.
By employing a multi-column collaborative and machine learning approach, KNN and CNN model architectures are constructed through multi-dimensional peak feature extraction, dual-model fusion, and SHAP analysis. Combined with wavelet denoising, normalization, and outlier removal, high-precision identification of gas samples is achieved.
It effectively resists retention time drift and peak shape distortion, improves gas identification accuracy, shortens detection cycle, adapts to multi-component and multi-concentration gas identification, has interpretability and reliability, and meets industrial testing and regulatory requirements.
Smart Images

Figure CN121476498A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of gas chromatography analysis technology, and in particular to a method for identifying trace gases in seawater based on multi-column synergy and machine learning. Background Technology
[0002] Trace gases in seawater refer to gaseous components that dissolve in seawater at extremely low concentrations (typically less than 1 micromole per liter). Although they are present in minute quantities, they play a crucial role in marine chemistry, atmospheric science, global climate change, and marine ecosystem research.
[0003] Current methods for identifying trace gases in seawater primarily rely on traditional gas chromatography (GC) techniques. These methods classify gases through single-column separation and fixed threshold judgment, but they suffer from significant drawbacks: First, seawater samples have complex matrices and extremely low concentrations of trace gases (often below 1 ppm), making them susceptible to interference from salinity and organic matter. Traditional GC detectors lack sufficient sensitivity, leading to missed detections. Second, multi-component gases are prone to co-elution, making complete separation difficult with a single column, resulting in overlapping peaks that interfere with identification. Third, column temperature drift and column aging can cause retention time shifts, making traditional threshold judgments prone to misjudgment. Fourth, existing gas identification technologies incorporating machine learning do not fully utilize the synergistic separation characteristics of multiple columns, relying only on single-dimensional peak features, resulting in a lack of targeted feature selection and weak model anti-interference capabilities. Fifth, seawater samples require complex pretreatment; current technologies lack an integrated "pretreatment-separation-identification" solution, making the detection process cumbersome. Summary of the Invention
[0004] To overcome the aforementioned problems in the existing technology, this invention proposes a method for identifying trace gases in seawater based on multi-column synergy and machine learning.
[0005] The technical solution adopted by this invention to solve its technical problem is: a method for identifying trace gases in seawater based on multi-column synergy and machine learning, comprising the following steps: Step 1: Prepare the gas sample to be identified and perform gas chromatography detection on the gas sample; Step 2: Denoise, normalize, and remove outliers from the gas chromatogram; Step 3: Extract manual peak features and time-series peak features to form a multi-dimensional feature set; Step 4: Perform stratified filtering on the feature set obtained in Step 3, retaining the core features; Step 5: Construct a dual-model fusion architecture combining the KNN and CNN models; Step 6: Divide the features obtained in Step 4 into training set and validation set, train and validate the dual-model architecture, and optimize the dual-model architecture parameters. Step 7: Apply the SHAP analysis model to the decision-making process and generate a feature importance visualization report.
[0006] The above-mentioned method for identifying trace gases in seawater based on multi-column collaboration and machine learning includes, in step 1, a gas sample comprising multiple gases to be identified and interfering gases with different concentration gradients.
[0007] The above-mentioned method for identifying trace gases in seawater based on multi-column collaboration and machine learning, specifically step 4, involves dividing the features into physical features adapted to the KNN model and temporal features adapted to the CNN model, and filtering them through reverse verification and single-branch verification respectively, retaining the core features.
[0008] The above-mentioned method for identifying trace gases in seawater based on multi-column collaboration and machine learning, specifically step 5, involves: using manual peak features as input, calculating sample similarity using L1 distance, optimizing the k value according to the gas type, and achieving rapid matching and classification; using time-series peak features as input, constructing a lightweight 1D-CNN architecture, avoiding gradient vanishing through residual connections, and mining deep correlations between features; and using a voting mechanism for model fusion.
[0009] The voting mechanism of the above-mentioned seawater trace gas identification method based on multi-column collaboration and machine learning is as follows: when the prediction results of the KNN model and the CNN model are consistent, the result is directly output; when the results are inconsistent, the SHAP value is used for weighted judgment.
[0010] The above-mentioned method for identifying trace gases in seawater based on multi-column collaboration and machine learning, specifically step 7, involves: calculating the SHAP value of each peak feature, taking the absolute value of the SHAP value of each feature and then calculating the average value to measure the average impact on the model output.
[0011] The beneficial effects of this invention are that, through the fusion of multi-dimensional chromatographic peak features and the dual-model collaborative mechanism, it can effectively resist interference such as retention time drift and peak shape distortion. At the same time, it can be specifically optimized for the weak signal characteristics of low-concentration gases, solving the problems of misjudgment and missed judgment that are prone to occur in traditional methods, and significantly improving the accuracy of gas identification.
[0012] Improved efficiency and adaptability: No need to rely on precise standard sample calibration; rapid matching of unknown gases can be achieved simply by collecting feature samples in advance to build a model, significantly shortening the detection cycle. At the same time, the feature set and model parameters can be flexibly adjusted to accommodate the identification needs of multi-component and multi-concentration gases and adapt to the personalized detection requirements of different scenarios.
[0013] Enhanced interpretability and reliability: The introduction of SHAP technology clarifies the contribution of key identification features, breaking the "black box" limitation of machine learning models and enabling the detection results to have traceable and verifiable characteristics, which can meet the reliability requirements of identification results in industrial inspection, supervision and other scenarios. Attached Figure Description
[0014] Figure 1 This is a schematic diagram of the process of this invention; Figure 2 This is a schematic diagram of the dual-model fusion architecture of the present invention; Figure 3 This is a feature importance analysis diagram of the SHAP feature of this invention. Detailed Implementation
[0015] To enable those skilled in the art to better understand the technical solution of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0016] This invention achieves accurate gas identification through seven steps: sample collection and GC detection → data preprocessing → peak feature extraction → feature screening → dual model construction → validation and optimization → interpretability analysis. The core of this invention is the deep integration of multi-dimensional features of gas chromatographic peaks with machine learning algorithms. The specific process is as follows: Figure 1 As shown.
[0017] (I) Sample collection and GC detection 1. Sample preparation: Prepare gas samples to be identified (covering multiple different concentration gradients of the target gas and interfering gases). Collect at least 30 valid samples for each gas sample, and increase the number of low-concentration (≤5ppm) samples to 50 to ensure sample balance.
[0018] 2. GC Instrument Parameter Settings: A multi-column combination (5Å molecular sieve, HayesepD, GDX-502) is integrated into a single column oven. The column oven uses a programmed temperature ramp-up, reaching and holding at 60°C; the carrier gas is high-purity helium, with a constant flow rate of 30 mL / min; the detector is an EPD. The core of the system lies in achieving central cutting and separation of each component through precise control of valve switching time.
[0019] (ii) Data preprocessing 1. Denoising: Wavelet denoising algorithm is used to remove instrument noise and baseline drift interference while preserving the core peak features.
[0020] 2. Normalization: This is achieved through the formula... Peak correlation data are mapped to the [0,1] interval to eliminate dimensional differences.
[0021] 3. Outlier removal: Hotelling's T2 method is used to remove samples with peak distortion caused by abnormal injection or instrument malfunction, ensuring data reliability.
[0022] (III) Peak Feature Extraction Two types of core features are extracted to form a multi-dimensional feature set: 1. Manual peak features (physically interpretable features): including 7 types of features such as peak height, peak area, relative retention time (replacing absolute retention time to resist drift), peak rising slope (resisting peak distortion), peak start time, and peak end time, covering the key physical indicators for gas identification.
[0023] 2. Time-series peak features: Extract the original time-series data (approximately 500 sampling points) before and after the feature peak, retaining detailed information about the peak shape, to meet the feature mining needs of deep learning models.
[0024] (iv) Feature Filtering A combined strategy of "hierarchical screening + dual-model collaborative validation" is adopted: Layered filtering: Based on the dual-model characteristics, features are divided into KNN-adapted physical features and CNN-adapted temporal features, and filtered through reverse verification / single-branch verification respectively to retain core features.
[0025] Collaborative validation: The dual-model architecture with integrated feature inputs is validated with a fusion accuracy of ≥95% and a feature contribution of ≥20% as the standard, retaining core features and eliminating redundant information.
[0026] (V) Dual Model Construction and Fusion A dual-model fusion architecture of "KNN+CNN" is constructed to balance fast recognition and accurate classification. The dual-model fusion architecture is as follows: Figure 2 As shown: 1. KNN model: Using manual peak features as input, it calculates sample similarity using L1 distance and optimizes the k value (range of 30-60) according to the gas type to achieve fast matching and classification, and is suitable for emergency identification scenarios without standard samples.
[0027] 2. CNN Model: Using "temporal peak features" as input, a lightweight 1D-CNN architecture is constructed (3 convolutional layers, 16-64 filters, kernel size 3-5). Gradient vanishing is avoided through residual connections, and deep correlations between features are discovered.
[0028] 3. Model fusion: A voting mechanism is adopted. When the prediction results of the two models are consistent, the results are directly output; when the results are inconsistent, the SHAP value is used to weight the results to improve the reliability of classification.
[0029] (vi) Model Validation and Optimization 1. Dataset splitting: The training and test sets are split in a 7:3 ratio, and 5-fold cross-validation is used to evaluate model performance.
[0030] 2. Performance metrics: The core metrics are recognition accuracy, recall, and F1 score, with a target accuracy of ≥95%.
[0031] 3. Optimization strategy: Optimize model parameters (k value of KNN, learning rate of CNN) through grid search; use oversampling to enhance data for low-concentration samples; if features are redundant, reselect feature subsets.
[0032] (vii) Interpretability Analysis The SHAP (Shapley Additive exPlanations) technique is applied to analyze the model's decision-making process: calculate the SHAP value of each peak feature, quantify the contribution of features to the identification results, clarify key identification criteria, generate a visualization report on feature importance, and improve the credibility of the detection results.
[0033] This embodiment focuses on the identification of CH4 (covering a concentration gradient of 1-10 ppm, with simultaneous mixing of interfering gases such as C2H6), and primarily verifies its resistance to interference from "retention time drift and peak distortion". The specific process is as follows: I. CH4-specific sample and GC detection configuration Sample preparation: Prepare six concentration gradient samples of CH4 (1ppm, 2ppm, 4ppm, 6ppm, 8ppm, 10ppm), collect 30 basic samples for each concentration, and supplement with 20 additional samples for the low concentrations of 1ppm and 2ppm (total of 50 samples); simultaneously mix in equal concentrations of C2H6, CO, N2O, H2, and 50 times the amount of CO2 as interfering components to simulate a real complex gas environment.
[0034] GC anti-interference parameter settings: A combination of 5Å molecular sieve column (suitable for CH4 separation) + HayesepD column is used. CH4 is directionally cut into the 5Å molecular sieve column through a valve switching system. The column oven is programmed to be heated to 60℃ and held (to reduce retention time drift caused by temperature fluctuations). The carrier gas (high-purity helium) is kept at a constant flow rate of 30mL / min. The EPD detector focuses on the characteristic response range of CH4 to reduce the superposition of interference gas signals.
[0035] II. Preprocessing for CH4 Interference 1. Preprocessing to combat peak distortion: The CH4 peak is prone to distortion due to fluctuations in injection pressure, resulting in "increased peak width and additional small inflection points". A wavelet denoising algorithm was used to perform targeted filtering on the time interval (386~466s) where the CH4 peak is located (using the db4 wavelet basis and 3-level decomposition) to retain the core rising / falling edge slope characteristics of the CH4 peak. Then, the Hotelling's T2 method was used to remove distorted samples with "peak splitting and sudden drop in peak height" (a total of 8 distorted samples at a low concentration of 2ppm were removed).
[0036] 2. Normalization to reduce retention time drift: The retention time data of CH4 were normalized by "relative retention time" (using the stable CO2 peak as a reference) to convert the absolute retention time into a relative value, which initially offset the drift interference within ±2s.
[0037] III. Feature Extraction and Anti-interference Screening of CH4 1. Anti-interference feature extraction: (1) Manual feature extraction: extracting the peak height of CH4 , where y peak y_baseline is the baseline corresponding to the maximum value of y within the peak interval; peak area. Where s: the sampling point index corresponding to the start of the peak, e: the sampling point index corresponding to the end of the peak, t k y k : The time and response value of the kth sampling point; The total peak area is obtained by summing the areas of the trapezoids enclosed by two adjacent sampling points and the baseline; relative retention time. ,in The retention time of the CH4 peak; Retention time of the CO2 peak; slope of the rising edge of the peak. ,in The response value at the peak starting point. y represents the time corresponding to the peak start point. peak t peak : Peak response value and time; high area ratio Tail trail factor T=W 0.05h / 2d1, W 0.05h : The entire peak width at 5% peak height, 2d1: the distance between the peak apex and the peak front edge, totaling 7 types of physical characteristics; (2) Temporal features: Extract the original signal of 500 sampling points before and after the CH4 peak (including distortion details such as peak width and inflection point) and retain the temporal texture information after peak distortion.
[0038] 2. Anti-interference feature screening: During the stratified screening, the characteristic combination of "peak width, peak symmetry, and peak rise slope" (sensitive to chromatographic separation quality and providing time dynamic information) and "peak height" (sensitive to concentration information) is prioritized for retention, while the characteristics of "retention time, peak area, and peak height-to-area ratio" which are sensitive to interference are eliminated, and finally, four core characteristics are retained.
[0039] IV. Dual-model defense mechanism against CH4 interference KNN resists retention time drift: Using "peak width, peak symmetry, and peak rising slope" as inputs, the L1 distance is used to calculate sample similarity, and the k value is optimized to 10 to adapt to the concentration gradient distribution of CH4. Even if the CH4 retention time drifts by ±3s, the fluctuation range of the relative retention time is compressed, and the stable features of peak symmetry and peak rising slope will compensate for the deviation of retention time, ensuring that the matching accuracy of KNN for CH4 is not affected by drift.
[0040] CNNs combat peak distortion: Using "500-point temporal peak data" as input, a 1D-CNN is constructed with 3 convolutional layers, kernel sizes of 3 / 5 / 3, and filter numbers of 32 / 64 / 32. When the CH4 peak exhibits distortion such as "broadening and the addition of small inflection points", the convolutional layers can extract the distorted temporal texture features, such as the CH4-specific texture of "smooth rising edge + narrow falling edge". Deep features are preserved through residual connections to avoid feature loss caused by peak distortion.
[0041] Fusion verification of anti-interference effect: When CH4 is simultaneously subjected to a compound interference of "retention time drift of 2s + peak width increase of 1 time", the retention time feature of KNN shows a slight deviation, but CNN can identify CH4 through temporal peak texture. At this time, SHAP weighted judgment is used, and the SHAP contribution of CH4's "temporal peak data" reaches 40%, and the correct CH4 identification result is finally output.
[0042] SHAP feature importance analysis, such as Figure 3 As shown, the SHAP value of each feature in the model prediction process is calculated. The absolute value of the SHAP value of each feature is taken and the average value is calculated to measure the average influence of the feature on the model output, corresponding to the contribution ratio of the feature to CH4 gas. In Figure 3, the average SHAP value of peak width in the KNN model is higher (the bar is longer), which is the core feature with a greater impact on the model decision; peak symmetry and peak rise slope are next, while peak height has a relatively smaller average impact; in the CNN model, the 78th point of the time series peak (the steepest point of the rise edge) contributes the most, followed by the 66th point (acceleration steepening node) and the 59th point (initial rise edge), and the 101st point (peak transition region) contributes the least, accurately reflecting the CNN's ability to capture key nodes of the rise edge of the CH4 time series peak.
[0043] In this embodiment, the CH4 identification accuracy remained stable despite interference from "retention time drift ±3s and peak distortion (peak width increased by 1 time)", verifying the effective resistance of this method to the two types of interference.
[0044] The above embodiments are merely exemplary embodiments of the present invention and are not intended to limit the present invention. Those skilled in the art can make various modifications or equivalent substitutions to the present invention within its scope and spirit, and such modifications or equivalent substitutions should also be considered to fall within the scope of protection of the present invention.
Claims
1. A method for identifying trace gases in seawater based on multi-column synergy and machine learning, characterized in that, Includes the following steps: Step 1: Prepare the gas sample to be identified and perform gas chromatography detection on the gas sample; Step 2: Denoise, normalize, and remove outliers from the gas chromatogram; Step 3: Extract manual peak features and time-series peak features to form a multi-dimensional feature set; Step 4: Perform stratified filtering on the feature set obtained in Step 3, retaining the core features; Step 5: Construct a dual-model fusion architecture combining the KNN and CNN models; Step 6: Divide the features obtained in Step 4 into training set and validation set, train and validate the dual-model architecture, and optimize the dual-model architecture parameters. Step 7: Apply the SHAP analysis model to the decision-making process and generate a feature importance visualization report.
2. The method for identifying trace gases in seawater based on multi-column collaboration and machine learning according to claim 1, characterized in that, In step 1, the gas sample includes multiple gases to be identified and interfering gases with different concentration gradients.
3. The method for identifying trace gases in seawater based on multi-column collaboration and machine learning according to claim 1, characterized in that, Step 4 specifically involves dividing the features into physical features adapted to the KNN model and temporal features adapted to the CNN model, and filtering them through reverse verification and single-branch verification respectively, retaining the core features.
4. The method for identifying trace gases in seawater based on multi-column collaboration and machine learning according to claim 1, characterized in that, Step 5 specifically involves: the KNN model using manually generated peak features as input, calculating sample similarity using L1 distance, optimizing the k value based on gas type, and achieving fast matching and classification; the CNN model using time-series peak features as input, constructing a lightweight 1D-CNN architecture, avoiding gradient vanishing through residual connections, and mining deep correlations between features; and using a voting mechanism for model fusion.
5. The method for identifying trace gases in seawater based on multi-column collaboration and machine learning according to claim 4, characterized in that, The voting mechanism is as follows: when the prediction results of the KNN model and the CNN model are consistent, the result is output directly; when the results are inconsistent, the result is determined by weighting the SHAP value.
6. The method for identifying trace gases in seawater based on multi-column collaboration and machine learning according to claim 1, characterized in that, Step 7 specifically involves: calculating the SHAP value of each peak feature, taking the absolute value of the SHAP value of each feature, and then calculating the average value to measure the average impact on the model output.
Citation Information
Patent Citations
Method for prior oxidation of co in hydrogen-riched gas
CN101024490A
Special detection device for analyzing three-channel mine drainage gas
CN102590422A
Analysis device and method for content of trace impurities in hydrogen isotope gas and / or helium gas
CN108020612A
Multi-label image completion method based on deep convolution features and semantic neighbors
CN111080551A
Electronic nose gas sensitive-chromatographic information fusion and flavor substance on-site detection and analysis method
CN111443161A