A method for detecting hydrogen sulfide gas using an electronic nose

By combining an electronic nose with CNN and XGBoost models, the problem of high cost and inability to establish linear correlation in gas chromatography is solved, enabling low-cost, portable, rapid and accurate detection of hydrogen sulfide gas concentration.

CN116935981BActive Publication Date: 2025-10-17SOUTHWEST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310879706.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-18
Publication Date
2025-10-17
Estimated Expiration
2043-07-18

AI Technical Summary

Technical Problem

Existing gas chromatography methods are expensive to detect hydrogen sulfide gas and cannot establish a linear correlation between substances and olfactory stimuli. Humans cannot effectively identify high concentrations of hydrogen sulfide, making detection difficult.

Method used

An electronic nose detection method is adopted, which combines a convolutional neural network (CNN) and a gradient boosting regression model (XGBoost). By collecting odor data of hydrogen sulfide gas, preprocessing and feature extraction are performed, and CNN and XGBoost models are constructed to achieve real-time prediction of hydrogen sulfide gas concentration.

Benefits of technology

It enables accurate prediction of hydrogen sulfide gas concentration on low-cost and portable devices, establishes a linear correlation between substances and olfactory stimuli, and improves the accuracy and speed of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116935981B_ABST
    Figure CN116935981B_ABST
Patent Text Reader

Abstract

A kind of electronic nose detection method for hydrogen sulfide gas, comprising the following steps: 1) collecting the odor data of hydrogen sulfide gas with different concentrations;2) pre-processing and feature extraction are carried out on the odor data to obtain feature point data;3) a CNN model is constructed;4) the CNN model is trained using the feature point data and the odor data to obtain a feature extraction model;5) an XGBoost model is constructed;6) the XGBoost model is trained using the feature point data and the concentration of hydrogen sulfide gas to obtain a regression prediction model;7) real-time collection of odor data of the environment to be monitored and input into the feature extraction model to obtain feature point data;8) input the feature point data into the regression prediction model to obtain the predicted value of the concentration of hydrogen sulfide gas in the environment to be monitored.The combination of the CNN model and the XGBoost model realizes the prediction of hydrogen sulfide gas and provides valuable guidance for further research and application.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of gas concentration detection, and in particular to an electronic nose detection method for hydrogen sulfide gas. BACKGROUND

[0002] Air pollution refers to the accumulation of certain biological molecules, particles and gases (such as NH3 and H2S) in the Earth's atmosphere at excessive or harmful concentrations. Chemical pollutants generally come from human activities, such as livestock farming, automobile exhaust, petroleum products, sewage treatment facilities, landfills and industry. They are also produced by natural processes, such as forest fires and volcanic eruptions. Some of these pollutants are harmful to humans, animals and the environment. They can cause serious health problems, including allergies, asthma, fatigue, headaches, coughing, and can have carcinogenic properties, especially at high concentrations, and can cause premature death. In addition, they can have an odor that can cause air quality to deteriorate. H2S is one of the most toxic pollutants, detectable at low concentration levels (0.13 ppm) and has a characteristic rotten egg smell. H2S is released most during the decomposition of organic matter in wastewater. H2S is harmful even at low concentrations. Therefore, it is important to monitor it, especially in industrial areas where they are present at high concentrations.

[0003] Hydrogen sulfide (H2S) is a highly toxic colorless gas with a characteristic rotten egg smell. In most well-documented studies, the effects of exposure to hydrogen sulfide on health have been negative, especially when the level of hydrogen sulfide in the air is higher than 1 ppm. The human nervous system and respiratory tract are still the most sensitive to exposure to hydrogen sulfide. Although the odor of hydrogen sulfide is detected by the human nose at 0.0005 ppm, at higher concentrations (100 ppm), the sense of smell disappears after 2-15 minutes of exposure. Therefore, the odor of the gas is still an ineffective warning of the presence and detection of the respective gas. It is therefore important to detect high concentrations of hydrogen sulfide within the range where the human body cannot recognize it.

[0004] Gas chromatography has been widely used to analyze odorized air samples, allowing the identification of specific odor components. However, this method is expensive and does not provide information on human perception, so it cannot establish a linear correlation between the quantification of substances and olfactory stimulation. SUMMARY

[0005] The purpose of the present application is to propose an electronic nose detection method for hydrogen sulfide gas, comprising the following steps:

[0006] 1) Collecting odor data of hydrogen sulfide gas at different concentrations.

[0007] 2) Preprocessing and feature extraction of odor data to obtain feature point data, and constructing a CNN training data set with feature point data and odor data.

[0008] 3) Construct a CNN model.

[0009] 4) Train the CNN model using a CNN training dataset to obtain a feature extraction model.

[0010] 5) Construct an XGBoost model.

[0011] 6) Construct an XGBoost model training dataset using feature point data and hydrogen sulfide gas concentration.

[0012] 7) Train the XGBoost model using the XGBoost model training dataset to obtain a regression prediction model.

[0013] 8) Collect odor data in real time from the environment to be monitored and input it into the feature extraction model to obtain feature point data.

[0014] 9) Input the feature point data into the regression prediction model to obtain the predicted value of the hydrogen sulfide gas concentration in the environment to be monitored.

[0015] Further, the device for collecting odor data includes an electronic nose arranged in the environment to be monitored.

[0016] The odor data is the resistance value monitored by the electronic nose.

[0017] Further, the preprocessing of the odor data includes noise filtering and data correction.

[0018] The feature extraction method for the odor data includes statistical feature extraction, frequency domain feature extraction, time domain feature extraction, and feature selection method.

[0019] The feature values extracted by the statistical feature extraction include mean, variance, and peak value.

[0020] The frequency domain feature extraction method includes Fourier transform and wavelet transform.

[0021] The time domain feature extraction method includes autocorrelation function and sliding window statistics.

[0022] The feature selection method includes principal component analysis and linear discriminant analysis.

[0023] Further, the CNN model includes a 1D CNN model.

[0024] Further, the CNN model includes a convolutional layer, a pooling layer, and a fully connected layer.

[0025] The convolutional layer extracts local features in the time series of odor data by using several convolution kernels to perform convolution operations on the time series of odor data.

[0026] The pooling layer is used to reduce the spatial dimension of the output feature map.

[0027] The fully connected layer converts the feature map extracted by the convolution layer and the pooling layer into a final classification result, as follows:

[0028] Z = W·X + b (1)

[0029] In the formula, Z represents the output result of the fully connected layer, W is a weight matrix, X is a vector output by the pooling layer, and b is a bias vector.

[0030] Further, after obtaining the feature extraction model and the regression prediction model, the regression performance of the feature extraction model and the regression prediction model is evaluated by an evaluation index, and if the evaluation is not suitable, the data set is updated and the feature extraction model and the regression prediction model are retrained.

[0031] Further, the evaluation index includes a determination coefficient, a mean square error, and a mean absolute error.

[0032] Further, the determination coefficient is as follows:

[0033]

[0034] In the formula, R 2 is the determination coefficient, which represents an estimated value of the percentage of variance in the response variable explained by the relationship between the response variable and the explanatory variable. i represents the ordinal number of the explanatory variable, i = 1, 2, …, n, and n represents the total number of explanatory variables. y i represents the i-th explanatory variable. represents the response variable. represents the average value of the explanatory variable.

[0035] Further, the mean square error is as follows:

[0036]

[0037] In the formula, MSE is the mean square error. i represents the ordinal number of the explanatory variable, i = 1, 2, …, n, and n represents the total number of explanatory variables. y i represents the i-th explanatory variable. represents the response variable.

[0038] Further, the mean absolute error is as follows:

[0039]

[0040] In the formula, MAE is the mean absolute error. i represents the ordinal number of the explanatory variable, i = 1, 2, …, n, and n represents the total number of explanatory variables. y i represents the i-th explanatory variable. represents a response variable.

[0041] The technical effect of the present application is self-evident, and the present application realizes the prediction of hydrogen sulfide gas by combining the CNN model and the XGBoost model, and provides valuable guidance for further research and application.

[0042] The electronic nose detection method for hydrogen sulfide gas provided by the present application can establish a linear correlation between the quantified substance and the olfactory stimulus, realize a lower prediction error, and more accurately predict the numerical value of the hydrogen sulfide gas concentration. Compared with traditional detection instruments, it is more inexpensive, portable, and has a faster reaction and simple operation.

[0043] Further, 1DCNN is a kind of CNN model, which only performs one-dimensional convolution, so that its structure is simpler and has fewer parameters. Therefore, 1DCNN can save computing resources and time. Generally speaking, 1DCNN with proper structure can mine enough features and flexible forms from input data.

[0044] Further, it is found through experiments that the XGBoost model achieves satisfactory performance in high-concentration hydrogen sulfide regression prediction. The model has high prediction accuracy and robustness, which can help us better understand and predict the trend of high-concentration hydrogen sulfide. This has important practical significance for small sample gas concentration regression prediction. BRIEF DESCRIPTION OF DRAWINGS

[0045] Figure 1 is a 1DCNN_XGB framework diagram;

[0046] Figure 2 is a complete data acquisition device diagram;

[0047] Figure 3 is a response diagram of the data set;

[0048] Figure 4 is a part of the extracted feature data diagram;

[0049] Figure 5 is an SVM principle diagram;

[0050] Figure 6 is a data label diagram;

[0051] Figure 7 is a scatter plot of true value and predicted value;

[0052] Figure 8 is a prediction line chart. DETAILED DESCRIPTION

[0053] The application will be further described in conjunction with the examples below, but should not be understood as limiting the above-mentioned subject matter of the application to the following examples. Various substitutions and modifications can be made according to ordinary technical knowledge and conventional means in the art without departing from the technical idea of the application, and all such substitutions and modifications shall be included in the protection scope of the application.

[0054] Example 1

[0055] Referring to Figures 1 to 8 A method for detecting hydrogen sulfide gas by using an electronic nose, comprising the following steps:

[0056] 1) Collecting odor data of hydrogen sulfide gas with different concentrations.

[0057] 2) Preprocessing and feature extraction of the odor data to obtain feature point data, and constructing a CNN training data set with the feature point data and the odor data.

[0058] 3) Constructing a CNN model.

[0059] 4) Training the CNN model using the CNN training data set to obtain a feature extraction model.

[0060] 5) Constructing an XGBoost model.

[0061] 6) Constructing an XGBoost model training data set with the feature point data and the concentration of hydrogen sulfide gas.

[0062] 7) Training the XGBoost model using the XGBoost model training data set to obtain a regression prediction model.

[0063] 8) Collecting odor data of the environment to be monitored in real time and inputting it into the feature extraction model to obtain feature point data.

[0064] 9) Inputting the feature point data into the regression prediction model to obtain the predicted value of the concentration of hydrogen sulfide gas in the environment to be monitored.

[0065] Example 2

[0066] A method for detecting hydrogen sulfide gas by using an electronic nose, the main technical content is shown in Example 1, further, the device for collecting odor data comprises an electronic nose arranged in the environment to be monitored.

[0067] The odor data is the resistance value monitored by the electronic nose.

[0068] Example 3

[0069] A method for detecting hydrogen sulfide gas by using an electronic nose, the main technical content is shown in any one of Examples 1 to 2, further, the preprocessing of the odor data comprises noise filtering and data correction.

[0070] The method for feature extraction of the odor data includes statistical feature extraction, frequency domain feature extraction, time domain feature extraction, and feature selection method.

[0071] The feature values extracted by the statistical feature extraction include mean, variance, and peak value.

[0072] The frequency domain feature extraction method includes Fourier transform and wavelet transform.

[0073] The time domain feature extraction method includes autocorrelation function and sliding window statistics.

[0074] The feature selection method includes principal component analysis and linear discriminant analysis.

[0075] Embodiment 4:

[0076] A method for detecting hydrogen sulfide gas by using an electronic nose, the main technical content of which is shown in any one of embodiments 1 to 3, further, the CNN model includes a 1D CNN model.

[0077] Embodiment 5:

[0078] A method for detecting hydrogen sulfide gas by using an electronic nose, the main technical content of which is shown in any one of embodiments 1 to 4, further, the CNN model includes a convolution layer, a pooling layer, and a full connection layer.

[0079] The convolution layer extracts local features in the time sequence of the odor data by using a plurality of convolution kernels to perform convolution operation on the time sequence of the odor data, as follows:

[0080] C(i,j) = ∑∑(A(x,y)·B(i-x,j-y)) (1)

[0081] In the formula, C(i,j) represents the output feature of the convolution layer. B represents the convolution kernel. A(x,y) represents the input feature of the convolution layer.

[0082] The input of the convolution layer is a three-dimensional tensor with a shape of (sample number, time step length, feature number). In the present application, (105, 60, 10) is taken as an example. Specifically, each sample is a time sequence representing the odor data in a period of time. The time step length refers to the length of each time sequence, i.e. the number of observation points of the odor data. The feature number represents the number of features of each observation point, which is 10 in this case.

[0083] The pooling layer is used to reduce the spatial dimension of the output feature map.

[0084] The full connection layer converts the feature map extracted by the convolution layer and the pooling layer into the final classification result, as follows:

[0085] Z = W·X + b (2)

[0086] In the formula, Z represents the output result of the full connection layer, W is a weight matrix, X is a vector output by the pooling layer, and b is a bias vector.

[0087] Embodiment 6:

[0088] The electronic nose detection method for hydrogen sulfide gas mainly includes the technical content of any one of embodiments 1 to 5, further, after obtaining the feature extraction model and the regression prediction model, the regression performance of the feature extraction model and the regression prediction model is evaluated through the evaluation index, if the evaluation is not suitable, the data set is updated, and the feature extraction model and the regression prediction model are retrained.

[0089] Embodiment 7:

[0090] The electronic nose detection method for hydrogen sulfide gas mainly includes the technical content of any one of embodiments 1 to 6, further, the evaluation index includes a determination coefficient, a mean square error and a mean absolute error.

[0091] Embodiment 8:

[0092] The electronic nose detection method for hydrogen sulfide gas mainly includes the technical content of any one of embodiments 1 to 7, further, the determination coefficient is as follows:

[0093]

[0094] In the formula, R 2 is the determination coefficient, which represents an estimated value of the percentage of variance in the response variable explained by the relationship between the response variable and the explanatory variable. i represents the ordinal number of the explanatory variable, i = 1, 2, …, n, and n represents the total number of explanatory variables. y i represents the i th explanatory variable. represents the response variable. represents the average value of the explanatory variable.

[0095] Embodiment 9:

[0096] The electronic nose detection method for hydrogen sulfide gas mainly includes the technical content of any one of embodiments 1 to 8, further, the mean square error is as follows:

[0097]

[0098] In the formula, MSE is the mean square error. i represents the ordinal number of the explanatory variable, i = 1, 2, …, n, and n represents the total number of explanatory variables. y i represents the i th explanatory variable. represents the response variable.

[0099] Embodiment 10:

[0100] A method for detecting hydrogen sulfide gas by an electronic nose, the main technical content of which is shown in any one of Examples 1 to 9. Furthermore, the mean absolute error is as follows:

[0101]

[0102] Where MAE is the mean absolute error. i represents the ordinal number of the explanatory variable, i = 1, 2, ..., n, and n represents the total number of explanatory variables. i represents the i-th explanatory variable. represents the response variable.

[0103] Example 11:

[0104] See also Figures 1 to 8 , a method for detecting hydrogen sulfide gas by an electronic nose, comprising the following steps:

[0105] 1) Collect odor data of hydrogen sulfide gas at different concentrations.

[0106] The gas collection equipment used is the JF02F produced by Guizhou Guiyan Jinfeng Company. The entire test system consists of a computer, a test chassis, a gas distribution chassis, a detection gas chamber (detection module), etc. The JF02F software can be used to open and close the gas circuit and control the flow. The JF02F's dedicated software can achieve precise gas circuit opening and closing and flow control, ensuring the accuracy and repeatability of data collection. The equipment uses advanced sensor technology and high-precision flow control modules, which can monitor and adjust parameters such as pressure, temperature and flow in real time during the gas collection process, thereby providing high-quality data samples. Through the collaborative operation of this technical system, the data set collection process is efficient, stable and controllable, providing a reliable foundation for subsequent data analysis and application.

[0107] like Figure 2 As shown:

[0108] (1) During the control phase: Open the gas valve and simultaneously introduce a background gas (70% nitrogen and 30% oxygen mixture (synthetic air)) and a test gas (500 ppm hydrogen sulfide) into the gas mixing control box through the gas line. The total gas flow rate is 300 ml / min, and the gas concentration is controlled by adjusting the flow ratio of the background gas to the test gas.

[0109] (2) Software parameter selection: Since only one gas is tested, pulse ventilation response is set to two gas inlets. The test voltage is 8V and the heating voltage is 5V.

[0110] (3) Selection of experimental parameters: The sensor was preheated for 48 hours before the experiment.

[0111] (4) Gradient setting: initial cleaning time 120S, aeration time 60S, emptying time 120S. The software controls the test gas concentration by regulating the flow rate of the background gas. A total of 15 gradients are set, each with an interval of 20 ppm. The test hydrogen sulfide concentration is from 0-300 ppm.

[0112] (5) Sensor selection:

[0113] The sensor selection is shown in the following table:

[0114] Table 1. Sensor name and parameters

[0115] Sensor name Main test gas type MQ-7B Carbon monoxide TGS816 Methane TGS822 Alcohol TGS826 Ammonia MQ136 Hydrogen sulfide

[0116] (6) After setting the conditions above, the experiment is performed.

[0117] 2) Preprocess and feature extraction of odor data, get feature point data, and construct CNN training data set with feature point data and odor data.

[0118] 3) Construct CNN model.

[0119] Convolutional Neural Network (CNN) is a deep learning algorithm widely used in image processing and pattern recognition. CNN performs well in processing grid-structured data such as images, and it extracts and learns features in images by simulating the human visual system. The core idea of CNN is to build a deep network structure through convolutional layers, pooling layers and fully connected layers. In the convolutional layer, multiple convolution kernels are used to convolve the input image to extract local features in the image. The input image is constructed with the number of samples in the time series of odor data as the x-axis, the time step as the y-axis, and the number of features as the z-axis. The number of samples in the time series of odor data represents the odor data within the time series. The time step represents the length of the time series, i.e. the number of observation points of odor data. The number of features represents the number of features of each observation point. The pooling layer is used to reduce the spatial dimension of the output feature map. These convolution kernels move across the entire image, capturing features at different locations through local perception. The pooling layer is used to reduce the spatial dimension of the feature map, reducing the amount of calculation and extracting more robust features. Common pooling operations include max pooling and average pooling, which sample local regions to retain the most significant features. Finally, the fully connected layer maps the features extracted by the convolution and pooling layers to the final classification results. Through the backpropagation algorithm in the training process, CNN can learn the weight and bias parameters that adapt to specific tasks.

[0120] 4) Train the CNN model using the CNN training data set to obtain the feature extraction model.

[0121] Extracting features that contain essential information about the gas sensor response is important for easy interpretation of the dataset. Feature extraction methods usually involve signal processing and pattern recognition techniques.

[0122] Common feature extraction methods include statistical feature extraction (such as mean, variance, peak, etc.), frequency domain feature extraction (such as Fourier transform, wavelet transform), time domain feature extraction (such as autocorrelation function, sliding window statistics), and feature selection methods (such as principal component analysis, linear discriminant analysis). These methods can convert odor data into feature vectors with representativeness and discrimination.

[0123] 5) Build XGBoost model.

[0124] XGBoost (EXtreme Gradient Boosting) is an ensemble learning model based on gradient boosting decision trees, which performs well in prediction tasks and has achieved excellent results in various machine learning competitions. XGBoost combines gradient boosting algorithms and regularization techniques, with high efficiency and accuracy. The core idea of XGBoost is to iteratively train weak learners (decision trees) and optimize the model prediction of the current round based on the results of the previous round. It is based on gradient descent, which updates model parameters by minimizing the loss function, so that each round of training can better fit the target variable.

[0125] XGBoost uses a series of techniques to improve the performance of the model, including regularization, learning rate scaling, column sampling, row sampling, etc. These techniques help to reduce overfitting, increase the generalization ability of the model, and can handle large-scale datasets. XGBoost has good robustness and scalability, suitable for various types of data and tasks, including classification, regression and ranking, etc. It also has certain robustness in feature engineering, capable of handling missing values and different types of features. It provides efficient and accurate prediction ability through gradient boosting decision trees and regularization techniques, widely used in data mining and prediction modeling tasks.

[0126] 6) Build XGBoost model training dataset with feature point data and hydrogen sulfide gas concentration.

[0127] 7) Train the XGBoost model using the XGBoost model training dataset to obtain the regression prediction model.

[0128] 8) Real-time acquisition of odor data of the environment to be monitored, and input into the feature extraction model to obtain feature point data.

[0129] 9) inputting the feature point data into a regression prediction model to obtain a predicted value of the hydrogen sulfide gas concentration in the environment to be monitored.

[0130] Embodiment 12:

[0131] A method for detecting hydrogen sulfide gas by using an electronic nose, the main technical content of which is shown in Embodiment 11, further, the device for collecting odor data comprises an electronic nose arranged in the environment to be monitored.

[0132] The odor data is the resistance value monitored by the electronic nose.

[0133] For sensor data, each data point itself can be a feature. These data points can contain the measured value of a certain physical quantity. In this article, the resistance corresponding to each data point is the physical measurement value.

[0134] Embodiment 13:

[0135] A method for detecting hydrogen sulfide gas by using an electronic nose, the main technical content of which is shown in any one of Embodiments 11 to 12, further, the preprocessing of the odor data comprises noise filtering and data correction.

[0136] The method for extracting features from the odor data comprises statistical feature extraction, frequency domain feature extraction, time domain feature extraction, and feature selection method.

[0137] The feature values extracted by the statistical feature extraction comprise mean, variance, and peak value.

[0138] The frequency domain feature extraction method comprises Fourier transform and wavelet transform.

[0139] The time domain feature extraction method comprises autocorrelation function and sliding window statistics.

[0140] The feature selection method comprises principal component analysis and linear discriminant analysis.

[0141] Embodiment 14:

[0142] A method for detecting hydrogen sulfide gas by using an electronic nose, the main technical content of which is shown in any one of Embodiments 11 to 13, further, the CNN model comprises a 1D CNN model.

[0143] Although the complexity of XGBR (XGBoost Regressor) is low, the possibility of overfitting is low, but its ability to extract a sufficient number of features is poor, which can limit its performance. 1DCNN is a kind of CNN model, which only performs one-dimensional convolution, making its structure simpler and fewer parameters. Therefore, 1DCNN can save computing resources and time. In general, 1DCNN with proper structure can mine enough features and flexible forms from input data. However, 1DCNN requires more samples for training than traditional statistical models, and the number of samples required may make the electronic nose application unrealistic, which may lead to overfitting. Therefore, the 1DCNN-RFR framework is developed by combining the advantages of DCNN and XGBR to predict the concentration of hydrogen sulfide gas.

[0144] Embodiment 15:

[0145] An electronic nose detection method for hydrogen sulfide gas, the main technical content is any one of embodiments 11 to 14, further, the CNN model comprises a convolutional layer, a pooling layer and a fully connected layer.

[0146] The convolutional layer extracts local features in the time series of odor data by performing convolution operations on the time series of odor data using several convolution kernels, as follows:

[0147] C(i,j) = ∑∑(A(x,y)·B(i-x,j-y)) (1)

[0148] In the formula, C(i,j) represents the value of position (i,j) in the output feature map. B represents the convolution kernel. A represents the input feature map. (x,y) represents the position in the input feature map.

[0149] The input feature map is constructed with the number of samples of the time series of odor data as the x-axis, the time step as the y-axis, and the number of features as the z-axis.

[0150] The number of samples of the time series of odor data represents the odor data within the time series.

[0151] The time step represents the length of the time series, i.e. the number of observation points of odor data.

[0152] The number of features represents the number of features of each observation point.

[0153] The pooling layer is used to reduce the spatial dimension of the output feature map.

[0154] The input of the CNN model is a three-dimensional tensor with a shape of (sample number, time step, feature number). In the present application, it is specifically (105, 60, 10). Each sample is a time series, representing the smell data in a period of time. The time step refers to the length of each time series, i.e. the number of observation points of the smell data. The feature number represents the number of features of each observation point, which is 10 features here.

[0155] The fully connected layer converts the feature map extracted by the convolutional layer and the pooling layer into the final classification result, as follows:

[0156] Z = W·X + b (2)

[0157] In the formula, Z represents the output result of the fully connected layer, W is a weight matrix, X is a vector output by the pooling layer, and b is a bias vector.

[0158] Embodiment 16:

[0159] A method for detecting hydrogen sulfide gas by an electronic nose, the main technical content of which is any one of embodiments 11 to 15, further, after obtaining the feature extraction model and the regression prediction model, the regression performance of the feature extraction model and the regression prediction model is evaluated by an evaluation index, if the evaluation is not suitable, the data set is updated, and the feature extraction model and the regression prediction model are retrained.

[0160] Embodiment 17:

[0161] A method for detecting hydrogen sulfide gas by an electronic nose, the main technical content of which is any one of embodiments 11 to 16, further, the evaluation index includes a determination coefficient, a mean square error and a mean absolute error.

[0162] Embodiment 18:

[0163] A method for detecting hydrogen sulfide gas by an electronic nose, the main technical content of which is any one of embodiments 11 to 17, further, the determination coefficient is as follows:

[0164]

[0165] In the formula, R 2 is the determination coefficient, which represents the estimated value of the percentage of the variance of the response variable explained by the relationship between the response variable and the explanatory variable. i represents the ordinal number of the explanatory variable, i = 1, 2, …, n, and n represents the total number of explanatory variables. y i represents the i-th explanatory variable. represents the response variable. represents the average value of the explanatory variable.

[0166] Embodiment 19:

[0167] A method for detecting hydrogen sulfide gas by electronic nose, the main technical content is shown in any one of embodiments 11 to 18, further, the mean square error is as follows:

[0168]

[0169] In the formula, MSE is the mean square error. i represents the ordinal number of the explanatory variable, i = 1, 2, …, n, and n represents the total number of explanatory variables. y i represents the i-th explanatory variable. represents the response variable.

[0170] Example 20:

[0171] A method for detecting hydrogen sulfide gas by electronic nose, the main technical content is shown in any one of embodiments 11 to 19, further, the mean absolute error is as follows:

[0172]

[0173] In the formula, MAE is the mean absolute error. i represents the ordinal number of the explanatory variable, i = 1, 2, …, n, and n represents the total number of explanatory variables. y i represents the i-th explanatory variable. represents the response variable.

[0174] Example 21:

[0175] Referring to Figures 1 to 8 , a method for detecting hydrogen sulfide gas by electronic nose, the content is as follows:

[0176] Data set acquisition: The gas collection equipment selected is JF02F generated by Guizhou Guyan Jin Feng Company for collection, and the whole test system is composed of computer, test case, gas distribution case, detection gas chamber (detection module) and etc. The opening and closing of the gas circuit and the flow control can be realized through the software of JF02F. Through the special software of JF02F, accurate gas circuit opening and closing and flow control can be realized, ensuring the accuracy and repeatability of data acquisition. The equipment adopts advanced sensor technology and high-precision flow control module, which can monitor and adjust the pressure, temperature and flow parameters in the gas collection process in real time, so as to provide high-quality data samples. Through the collaborative operation of the technical system, the data set acquisition process has the characteristics of high efficiency, stability and controllability, providing a reliable basis for subsequent data analysis and application.

[0177] (1) In the control stage: open the gas valve to pass through the gas circuit to mix the background gas by 70% nitrogen and 30% oxygen (synthetic air), and the test gas (500 ppm of hydrogen sulfide) into the gas mixing control box at the same time. The total flow of gas is 300 ml / min, and the gas concentration change is controlled by adjusting the flow ratio of background gas and measured gas.

[0178] (2) Selection of software parameters: Since only one gas was tested, the two-way gas inlet setting for pulse ventilation response was performed. The test voltage was 8 V, and the heating voltage was 5 V.

[0179] (3) Selection of experimental parameters: The sensor was tested after preheating for 48 h.

[0180] (4) Gradient setting: The initial cleaning time was 120 s, the ventilation time was 60 s, and the emptying time was 120 s. The software controlled the test gas concentration by adjusting the flow rate of the background gas. A total of 15 gradients were set, with each gradient interval being 20 ppm. The test hydrogen sulfide concentration was from 0-300 ppm.

[0181] (5) Selection of sensors:

[0182] Table 1. Sensor name and parameters

[0183] Sensor name Main test gas type MQ-7B Carbon monoxide TGS816 Methane TGS822 Alcohol TGS826 Ammonia MQ136 Hydrogen sulfide

[0184] (6) After setting the above conditions, the experiment was performed.

[0185] Feature extraction data preprocessing

[0186] Extracting some features containing basic information of gas sensor response is important for easy interpretation of the data set. Feature extraction methods usually involve signal processing and pattern recognition techniques. The odor data obtained by the sensor array need to be preprocessed, such as noise filtering and data correction. Then, common feature extraction methods include statistical feature extraction (such as mean, variance, peak value, etc.), frequency domain feature extraction (such as Fourier transform, wavelet transform), time domain feature extraction (such as autocorrelation function, sliding window statistics), and feature selection methods (such as principal component analysis, linear discriminant analysis). These methods can convert odor data into feature vectors with representativeness and discrimination. For sensor data, each data point itself can be a feature. These data points may contain the measurement value of a certain physical quantity. In this paper, the resistance corresponding to each data point is the physical measurement value. Since the data at the beginning and end of the ventilation stage are not the most stable, in this paper, the method of selecting features is to select the data of the middle 60 s in each ventilation response stage for extraction. Each data corresponds to the resistance response value of 10 sensors. A total of 63,000 feature points were extracted from 15*10*60*7.

[0187] According to Figure 3As shown, a gas data set containing 7 cycles was obtained, which covered data collected from 10 sensors. This data set provided information on gas concentration, sensor resistance, and other key parameters. Each cycle recorded the sensor's measurements at different time points, forming a series of time series data. The data from these sensors revealed trends and fluctuations in gas concentration. By analyzing the data set, which contained 2819 measurement points collected over 7 cycles, the data set covered 105 sensor response gradients for measuring hydrogen sulfide concentrations ranging from 0 to 300 ppm. The data set feature point graph is shown in Figure 4 As shown. Each time point's data recorded the sample time (SampleTime) and the hydrogen sulfide concentration set value (GasConcentrationSet). In addition, the data set also contained the corresponding sensor response values of 10 channels (Channel1 to Channel10). The data of these channels provided information on different aspects of hydrogen sulfide concentration. By analyzing this data set, the relationship between hydrogen sulfide concentration and sensor response can be understood. For each sensor, the linear or nonlinear relationship between its response value and hydrogen sulfide concentration can be found, and the performance of the sensor at different concentration ranges can be understood.

[0188] Support Vector Machine (SVM) is a widely used supervised learning algorithm in the field of machine learning. It is a non-probabilistic binary classification model, but can also be extended to multi-classification and regression tasks. In two-dimensional space, two classes of points are completely separated by a straight line, which is called linearly separable. In two-dimensional coordinates, find a straight line in the sample space to separate samples of different classes. The straight line that separates the data set is called the separating hyperplane.

[0189] That is:

[0190] w T x+b = 0 (1)

[0191] The core idea of SVM is to achieve classification by finding an optimal hyperplane that maximizes the margin between different classes. This hyperplane is chosen to best separate the samples of the two classes and has good generalization ability for unknown samples. In SVM, support vectors are the sample points in the training set that are closest to the hyperplane, and they play a key role in determining the decision boundary. The goal of SVM is to find the solution of an optimization problem that aims to maximize the margin of support vectors to the hyperplane while also limiting the classification error of training samples. SVM can be applied to linearly separable and linearly inseparable cases. In the case of linear inseparability, SVM introduces the concept of kernel functions to map samples to higher-dimensional feature spaces, thereby achieving non-linear classification in space. When training an SVM model, appropriate kernel functions and related parameters, such as the regularization parameter C and the parameters of the kernel function, need to be selected. The selection of these parameters can be optimized through techniques such as cross-validation. SVM has some advantages, such as good adaptability to high-dimensional data and small sample size, better generalization performance, and anti-overfitting ability. However, for large-scale data sets, the computational complexity of SVM is high, and it takes a long time to train. Support vector machine is a powerful classification algorithm that achieves data classification by finding the optimal hyperplane and performs well in handling linearly separable and linearly inseparable problems.

[0192] XGBoost, an ensemble learning model based on gradient boosting tree algorithm, is a machine learning algorithm based on gradient boosting decision trees, which performs well in prediction tasks and achieves excellent results in various machine learning competitions. XGBoost combines gradient boosting algorithm and regularization techniques, with high efficiency and accuracy. The core idea of XGBoost is to iteratively train weak learners (decision trees) and optimize the model prediction of the current round based on the results of the previous round. It is based on gradient descent, which updates model parameters by minimizing the loss function, so that each round of training can better fit the target variable.

[0193] XGBoost adopts a series of techniques to improve the performance of the model, including regularization, learning rate scaling, column sampling, row sampling, etc. These techniques help to reduce overfitting, increase the generalization ability of the model, and can handle large-scale data sets. XGBoost has good robustness and scalability, and is suitable for various types of data and tasks, including classification, regression, and ranking. It also has certain robustness in feature engineering, and can handle missing values and different types of features. It provides efficient and accurate prediction ability through gradient boosting decision trees and regularization techniques, and is widely used in data mining and prediction modeling tasks.

[0194] Convolutional Neural Network (CNN) is a deep learning algorithm widely used in image processing and pattern recognition. CNN excels in handling data with grid structure, such as images, by mimicking the human visual system to extract and learn features in images. The core idea of CNN is to build a deep network structure through convolutional layers, pooling layers, and fully connected layers. In the convolutional layer, multiple convolution kernels are used to convolve the input image to extract local features. These convolution kernels move across the entire image, capturing features at different positions through local perception. Pooling layers are used to reduce the spatial dimension of feature maps, reducing computational complexity and extracting more robust features. Common pooling operations include max pooling and average pooling, which sample local regions to retain the most significant features. Finally, the fully connected layer maps the features extracted by convolution and pooling layers to the final classification results. Through the backpropagation algorithm during training, CNN can learn the weight and bias parameters that adapt to specific tasks.

[0195] The advantage of CNN is its ability to automatically extract features from images and have certain translation, scale, and rotation invariance. In addition, by stacking multiple convolutional and fully connected layers, CNN can learn more complex feature representations, improving the model's expressive power and classification accuracy.

[0196] C(i,j) = ∑∑p(A(x,y)·B(i-x,j-y)) (2)

[0197] Forward propagation process of convolutional layer. In the convolution operation, the convolution kernel B slides over the input feature map A and performs element-wise multiplication and accumulation operations on the local region at each position. C(i,j) in the formula represents the value of position (i,j) in the output feature map, which is obtained by element-wise multiplication and summation of input feature map A and convolution kernel B.

[0198] Z = W·X + b (3)

[0199] The calculation process of the fully connected layer is shown in the formula: where Z represents the output result of the fully connected layer, W is the weight matrix, X is the input vector, and b is the bias vector. The multiplication operation in the formula represents the matrix multiplication operation of the input vector X and the weight matrix W, and then adds the bias vector b to obtain the linear combination result Z.

[0200] In this embodiment, three evaluation indicators, coefficient of determination (R2), mean square error (MSE), and mean absolute error (MAE), are used to evaluate the regression performance of the four models and the proposed framework. R 2Generally denoted as the estimated value of the percentage of variance within the response variable explained by the (linear) relationship of the response variable with the explanatory variables. MSE represents the difference between the predicted value of the sample and the observed value. MAE is defined as the average absolute difference between the predicted value and the observed value of the sample. The evaluation metrics are defined in equations (4)-(6).

[0201]

[0202]

[0203]

[0204] Although XGBR (XGBoost Regressor) has a lower complexity and a lower possibility of overfitting, its ability to extract a sufficient number of features can limit its performance. 1DCNN is a type of CNN model that only performs one-dimensional convolution, making its structure simpler and having fewer parameters. Therefore, 1DCNN can save computational resources and time. In general, 1DCNN with a proper structure can mine sufficient features and flexible forms from input data. However, 1DCNN requires more samples for training than traditional statistical models, and the required number of samples can make the electronic nose application unrealistic, which can lead to overfitting. Therefore, the 1DCNN-RFR framework is developed by combining the advantages of 1DCNN and XGBR to predict the concentration of hydrogen sulfide gas.

[0205] Example 22:

[0206] Referring to Figures 1 to 8 , an electronic nose detection method for hydrogen sulfide gas, the main technical content is shown in Example 21, and further, the data label graph is shown in Figure 6 : According to the given data sample, the concentration of hydrogen sulfide (H2S) gas is predicted. In this regression task, the concentration range is set to 0 to 300 ppm, and labeled with an interval of 20 ppm. When the concentration of hydrogen sulfide gas is 0 ppm, it means that there is almost no H2S gas in the environment. When the concentration is 20 ppm, it means that the concentration of H2S gas has slightly increased, but is still at a very low level. As the concentration increases by 20 ppm, the concentration of H2S gradually increases until it reaches a maximum concentration of 300 ppm. Within the range of 20 ppm to 300 ppm, a label is marked every 20 ppm. This means that 15 labels will be generated, corresponding to 20 ppm-300 ppm, respectively.

[0207] According to the comparison of the experimental results in Table 2, it can be concluded that:

[0208] In the feature extraction and regression prediction tasks, the performance of different models was evaluated, and the results showed that the 1D CNN model exhibited significant advantages in both tasks. The R2 score of this model reached 99.45%, indicating that it could explain about 99.45% of the variance in the target variable. Additionally, the 1D CNN model achieved satisfactory results in terms of MSE and MAE, with values of 48.68 and 3.700, respectively. This indicates that the model can achieve lower prediction errors and more accurately predict the numerical values of the target variable.

[0209] By adding a random forest as a regression forest to improve the performance of the 1D_CNN model, we obtained the 1DCNN_RFR model, which further improved its performance. The R2 score of this model reached 99.85%, with an MSE of 2.869 and an MAE of 2.783, meaning that the model can more accurately explain the variance of the target variable and achieve more precise predictions. However, this is not the optimal solution.

[0210] In addition, we also explored the combination of 1D CNN and XGBoost models (1DCNN_XGBR). The results showed that this combined model achieved the best performance in feature extraction and regression prediction tasks. Its R2 score reached 99.98%, with an MSE of only 1.002 and an MAE of 0.312. This indicates that the model can explain the variance of the target variable with extremely high accuracy and achieve precise predictions.

[0211] In summary, the 1D_CNN model exhibited significant advantages in feature extraction and regression prediction tasks, and by adding regularization terms and combining with the XGBoost model, the performance of the model was further improved. These findings are of great significance for a deeper understanding of the target variable and accurate prediction, and provide valuable guidance for further research and application.

[0212] Table 2. Comparison of experimental results

[0213] Model R2 MSE MAE XGBoost 98.85% 73.84 6.02 1DCNN 99.45% 48.68 3.700 1DCNN_RFR 99.85% 12.89 2.783 1DCNN_XGBR 99.97% 1.520 0.298

[0214] The scatter plot is shown in Figure 7 : it shows the relationship between the true values and the predicted values, with an MSE of 1.520. In the scatter plot, most of the points are distributed in a relatively tight area, indicating that the predicted values are generally close to the true values. However, some outliers or points scattered in the far area can also be observed, indicating that there are some large differences between some predicted results and the true values. Although there are some deviations with an MSE of 1.520, overall, it can be said that the prediction model is relatively accurate in predicting the true values.

[0215] In this example, the content of high-concentration hydrogen sulfide is predicted by establishing a regression model. By comparing different regression models, it is found that the XGBoost model performs well in the regression prediction of high-concentration hydrogen sulfide. This model shows high prediction accuracy and reliability.

[0216] The results show that the XGBoost model achieves satisfactory performance in the regression prediction of high-concentration hydrogen sulfide. The R2 score of this model is 99.97%, and in addition, the MSE of the model is observed to be 1.520 and the MAE is 0.298, which indicates that the prediction error of the model is relatively small and can provide reliable prediction results. Based on these results, it can be concluded that the XGBoost model is an effective tool for the regression prediction of high-concentration hydrogen sulfide content. This model has high prediction accuracy and robustness, which can help better understand and predict the trend of high-concentration hydrogen sulfide. This has important practical significance for small sample gas concentration regression prediction. However, it needs to be noted that the prediction performance of the model is still affected by the quality of the data and the selection of features. Therefore, when further applying the model, it is necessary to ensure the accuracy and reliability of the data and select appropriate features for modeling. In addition, the generalization ability of the model also needs to be further verified and evaluated to ensure its prediction performance in different scenarios and data sets.

[0217] This example demonstrates the potential of the XGBoost model in the regression prediction of high-concentration hydrogen sulfide. This provides strong support for decision-making in related fields and also provides new directions and inspirations for future research and application.

Claims

1. A method for detecting hydrogen sulfide gas using an electronic nose, characterized in that: The following steps are involved: 1) Collect odor data of hydrogen sulfide gas at different concentrations; The equipment for collecting odor data includes an electronic nose placed in the environment to be monitored; The odor data is the resistance value detected by the electronic nose; 2) Preprocessing and feature extraction of odor data to obtain feature point data, and constructing a CNN training dataset based on the feature point data and odor data; The preprocessing of the odor data includes noise filtering and data correction; The method for extracting features from odor data includes statistical feature extraction, frequency domain feature extraction, time domain feature extraction, and feature selection method; The characteristic values ​​extracted by the statistical feature extraction include mean, variance, and peak value; The frequency domain feature extraction method includes Fourier transform and wavelet transform; The time domain feature extraction method includes autocorrelation function and sliding window statistics; The feature selection method includes principal component analysis and linear discriminant analysis; 3) Build a CNN model; The CNN model includes a 1DCNN model; The CNN model includes a convolutional layer, a pooling layer and a fully connected layer; The convolution layer performs convolution operations on the time series of the odor data using a number of convolution kernels to extract local features in the time series of the odor data; The pooling layer is used to reduce the spatial dimension of the output feature map; The fully connected layer converts the feature maps extracted by the convolutional layer and the pooling layer into the final classification result, as shown below: Z=W·X+b (1) Where Z represents the output of the fully connected layer, W is the weight matrix, X is the vector output by the pooling layer, and b is the bias vector; 4) Use the CNN training data set to train the CNN model to obtain a feature extraction model; 5) Build the XGBoost model; 6) Construct an XGBoost model training dataset based on feature point data and hydrogen sulfide gas concentration; 7) Use the XGBoost model training data set to train the XGBoost model to obtain a regression prediction model; 8) Collect odor data of the monitored environment in real time and input it into the feature extraction model to obtain feature point data; 9) Input the characteristic point data into the regression prediction model to obtain the predicted value of the hydrogen sulfide gas concentration in the monitored environment.

2. The electronic nose detection method for hydrogen sulfide gas according to claim 1, characterized in that: After obtaining the feature extraction model and regression prediction model, the regression performance of the feature extraction model and regression prediction model is evaluated through evaluation indicators. If the evaluation is inappropriate, the data set is updated and the feature extraction model and regression prediction model are retrained.

3. The electronic nose detection method for hydrogen sulfide gas according to claim 2, characterized in that: The evaluation indicators include determination coefficient, mean square error and mean absolute error.

4. The electronic nose detection method for hydrogen sulfide gas according to claim 3, characterized in that: The coefficient of determination is as follows: Where R 2 is the coefficient of determination, which represents an estimate of the percentage of variance in the response variable explained by the relationship between the response variable and the explanatory variables; i represents the ordinal number of the explanatory variable, i=1,2,…,n, n represents the total number of explanatory variables; y i represents the i-th explanatory variable; represents the response variable; represents the mean value of the explanatory variable.

5. The electronic nose detection method for hydrogen sulfide gas according to claim 3, characterized in that: The mean square error is as follows: Where MSE is the mean square error; i represents the ordinal number of the explanatory variable, i = 1, 2, ..., n, n represents the total number of explanatory variables; y i represents the i-th explanatory variable; represents the response variable.

6. The electronic nose detection method for hydrogen sulfide gas according to claim 3, characterized in that: The mean absolute error is as follows: Where MAE is the mean absolute error; i represents the ordinal number of the explanatory variable, i = 1, 2, ..., n, n represents the total number of explanatory variables; y i represents the i-th explanatory variable; represents the response variable.