A method for extracting response fingerprint features of a gas sensor based on parameter fitting
By calculating the fitting parameters and information entropy contribution rate of the gas sensor response data, compact and interpretable fingerprint features are selected, which solves the problem of insufficient interpretability in gas type and concentration identification in traditional methods and is suitable for small sample set identification tasks.
Patent Information
- Application Number
- CN202511235674.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-01
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-09-01
AI Technical Summary
Existing gas sensor response data processing methods lack interpretability, traditional feature extraction methods cannot effectively distinguish gas types and concentrations, and deep learning methods are not friendly to small sample sets.
By extracting the fitting parameters from the gas sensor response data, calculating the inter-class and intra-class divergence ratios and information entropy contribution rates, fingerprint features with high compactness and interpretability are selected.
It achieves efficient dimensionality reduction and feature extraction of gas sensor response data, and obtains fingerprint features with high information entropy contribution rate and interpretability, which are suitable for small sample set recognition tasks.
Smart Images

Figure CN120724116B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of gas fingerprint feature extraction, and particularly relates to a gas sensor response fingerprint feature extraction method based on parameter fitting. BACKGROUND
[0002] Toxic and harmful gas monitoring is crucial for environmental protection, industrial safety and health risk assessment. Gas sensors are widely used in gas detection industry due to their low power consumption, high sensitivity and compact size. The response data of gas sensors is a kind of time series data, which contains rich gas component information, including gas type, gas concentration and environmental interference information. The response behavior in the static temperature working mode is related to the gas concentration, and the saturation value of the response increases with the increase of the gas concentration; the response fingerprint in the dynamic temperature modulation mode is related to both the gas type and the concentration, and the response fingerprint mode of different types of gases can be distinguished obviously, and the fingerprint mode of the same type of gas has good consistency and the fingerprint amplitude increases with the increase of the gas concentration. With the development of machine learning and deep learning algorithms, data processing and analysis of gas sensor response data can realize autonomous detection and rapid analysis. Sensor calibration driven by response fingerprint can accelerate the construction and deployment process of identification model, and improve the identification accuracy and robustness of gas sensor.
[0003] Feature extraction is crucial for pattern recognition. Traditional feature extraction methods based on machine learning algorithms include principal component analysis, linear discriminant analysis, factor analysis and independent component analysis. These methods map high-dimensional data to low-dimensional space based on statistical principles, thereby realizing data dimension reduction and feature extraction. However, the above methods lack interpretability, and cannot provide a clear optimization direction for pattern recognition work. In recent years, end-to-end feature extractors based on deep learning algorithms are widely used for data coding, including convolutional neural networks, recurrent neural networks and attention mechanisms. Although the above deep learning methods simplify the feature extraction work by not requiring manual extraction of original data, the large feature extractor requires a large amount of training data, which is not friendly to small sample learning tasks such as gas sensor calibration. SUMMARY
[0004] In view of the deficiencies of the prior art, the present application provides a gas sensor response fingerprint feature extraction method based on parameter fitting. For the small sample response data set of the gas sensor, the present application extracts the response curve fitting parameters, calculates the inter-class and intra-class scatter ratio of the features, thereby obtains the information entropy contribution rate, and performs feature screening according to the information entropy contribution rate and the cumulative contribution rate of the features, thereby obtaining fingerprint features with high compactness and interpretability.
[0005] A gas sensor response fingerprint feature extraction method based on parameter fitting, comprising the following steps:
[0006] Step 1, obtaining time sequence response data of a gas sensor;
[0007] The gas sensor includes a semiconductor gas sensor, an electrochemical gas sensor, a catalytic combustion gas sensor, a MEMS gas sensor, and a surface acoustic wave gas sensor;
[0008] The time sequence response data is a one-dimensional time sequence, including static temperature test response data and dynamic temperature modulation response data; wherein the static temperature test is a working mode with constant working temperature of the gas sensor, and the dynamic temperature modulation is a working mode with periodic change of working temperature of the gas sensor;
[0009] The time sequence response data needs to be labeled according to specific tasks, specifically: when performing a gas species identification task, the label is labeled as a gas species; when performing a gas concentration prediction task, the label is labeled as a concentration type;
[0010] The time sequence response data includes C data categories, and each category includes N c time sequence response data;
[0011] Step 2, based on curve parameter fitting, obtaining fitting parameters of the time sequence response data as a feature sequence, and constructing a feature sequence data set;
[0012] Step 2.1, based on curve parameter fitting, data fitting is performed on the time sequence response data of the gas sensor, and the formula is as follows:
[0013] ;
[0014] Wherein, f(t) is the time sequence response data, F(t) is the parameter fitting model, and R(t) is the residual term;
[0015] The parameter fitting model includes a linear function, an exponential function, a power function, a polynomial function, a Fourier series, a Gaussian function, a trigonometric function, or a sum of the above functions;
[0016] Step 2.2, calculating the fitting performance index according to f(t) and F(t), and comparing with the index requirement, if the fitting performance index meets the index requirement, executing step 2.3, if the fitting performance index does not meet the index requirement, adjusting the parameter fitting model F(t), and executing step 2.1;
[0017] The fitting performance index is one or more of the coefficient of determination R 2 , mean absolute error MAE, mean square error MSE, and root mean square error RMSE;
[0018] Step 2.3, obtain the model parameters of the parameter fitting model F(t), and aggregate all the model parameters into a sequence {x(i)} as the extracted feature sequence, where x(i) is each feature, i represents the feature index in the feature sequence, i satisfies i = 1, 2,..., n, and n represents the number of features in the feature sequence;
[0019] Step 2.4, perform steps 2.1-2.3 on each piece of time series response data to obtain a feature sequence dataset;
[0020] The feature sequence dataset contains C categories in total, and each category contains N c feature sequences;
[0021] The jth feature sequence in the cth category is represented as follows:
[0022] i = 1, 2,..., n;
[0023] Step 3, calculate the between-class and within-class scatter ratio R(i) of each feature in the feature sequence;
[0024] Step 3.1, calculate the mean value μ c (i) of each feature by category, as follows:
[0025] ;
[0026] where μ c (i) is the mean value of the ith feature in the cth category, and c satisfies 1 ≤ c ≤ C;
[0027] Step 3.2, calculate the between-class scatter S b (i) of each feature, as follows:
[0028] ;
[0029] where μ p (i) and μ q (i) are the mean values of the ith feature in the pth category and the qth category;
[0030] Step 3.3, calculate the within-class scatter S w (i) of each feature, as follows:
[0031] ;
[0032] Step 3.4, calculate the between-class and within-class scatter ratio R(i) of each feature, as follows:
[0033] ;
[0034] Step 4, calculate the information entropy contribution rate H(i) of each feature according to the inter-class intra-class dispersion ratio R(i);
[0035] Step 4.1, normalize the inter-class intra-class dispersion ratio R(i) to obtain the normalized inter-class intra-class dispersion ratio R'(i), the formula is as follows:
[0036] ;
[0037] Wherein, the R(i) min and R(i) max are the minimum and maximum values of R(i) respectively;
[0038] Step 4.2, perform Softmax calculation on the normalized inter-class intra-class dispersion ratio R'(i) to map R'(i) to probability distribution P(i), the formula is as follows:
[0039] ;
[0040] Wherein, P(i) satisfies the following conditions:
[0041] ;
[0042] Step 4.3, calculate the information entropy H of the entire feature sequence according to P(i), the formula is as follows:
[0043] ;
[0044] Step 4.4, calculate the information entropy contribution rate H(i) of each feature, the formula is as follows:
[0045] ;
[0046] Step 5, perform feature screening according to the information entropy contribution rate H(i);
[0047] Step 5.1, reorder H(i) according to the numerical value of H(i) from large to small to obtain H(i'), and calculate the permutation relationship σ of the element position of H(i') and H(i); Wherein the value of i' satisfies i'=1, 2,..., n;
[0048] The calculation process of the permutation relationship σ of the element position is as follows:
[0049] The sequence length of H(i) and H(i') is n, so the index set of the two is represented as [n] = {1, 2,...,n};
[0050] The element position index in the original sequence H(i) is i, and i ∈ [n]; the element position index in the sorted sequence H(i') is i', and i' ∈ [n];
[0051] The element position order adjustment is defined by a permutation function σ(i'): [n]→[n]; the permutation function σ(i') represents the original position corresponding to the adjusted position i', that is, the index of the element at the new position i' in the original sequence;
[0052] Step 5.2, calculate the cumulative information entropy contribution rate ∑H(i') of the first i' features in H(i'), the formula is as follows:
[0053] ;
[0054] Wherein, k = 1, 2,..., i';
[0055] Step 5.3, compare ∑H(i') with the cumulative information entropy contribution rate threshold δ, find the minimum i' value i' that satisfies ∑H(i') not less than δ min , the formula is as follows:
[0056] ;
[0057] Wherein, the δ is the cumulative information entropy contribution rate threshold;
[0058] Step 5.4, obtain the index corresponding to the first i' min Information entropy contribution rate in H(i') }, the formula is as follows:
[0059] ;
[0060] Step 5.5, according to the permutation relationship σ of the element position, calculate the position index } of the sorted element position index } in the original sequence, the formula is as follows:
[0061] ;
[0062] Step 5.6, according to the position index From the original feature sequence Get the filtered feature As the extracted fingerprint feature, the formula is as follows:
[0063] .
[0064] Preferably, the gas sensor in step 1 is a semiconductor gas sensor;
[0065] Preferably, the parameter fitting model in step 2.1 is a Fourier series;
[0066] Preferably, the fitting performance index in step 2.2 is a determination coefficient R 2 ;
[0067] Preferably, the determination coefficient R 2 in step 2.2 is required to be not less than 0.95;
[0068] Preferably, the base of the logarithmic function in step 4.3 is 2;
[0069] Preferably, the cumulative information entropy contribution rate threshold δ in step 5.3 is 60%.
[0070] The beneficial effects produced by the above technical solutions are:
[0071] The application provides a gas sensor response fingerprint feature extraction method based on parameter fitting, uses curve parameter fitting to obtain fitting parameters of gas sensor response data as an original feature sequence, calculates the information entropy contribution rate of each feature through the inter-class and intra-class scatter ratio of each feature, and performs further feature screening on the extracted feature sequence based on the information entropy contribution rate, so that high compactness and interpretable fingerprint features are obtained. The application can effectively reduce the dimension of semiconductor gas response data while retaining the original information beneficial to the identification task, and extract fingerprint features with high information entropy contribution rate and interpretability. The application provides favorable conditions for gas sensor response data dimension reduction, fingerprint feature extraction, sensor calibration, gas type and concentration identification, and lightweight identification model construction. Compared with the existing method, the application has the following advantages:
[0072] (1) Compared with traditional statistical-based feature extraction methods (such as principal component analysis, linear discriminant analysis, and factor analysis), the feature extraction method described in the application has high interpretability, the fitting parameters of the response curve can highly explain the mode and amplitude information of the response fingerprint, and the inter-class and intra-class scatter ratio and the information entropy contribution rate of the feature can highly explain the information gain of each feature.
[0073] (2) Compared with traditional deep learning-based end-to-end feature extraction methods (such as convolutional neural networks, recurrent neural networks, and attention mechanisms), the feature extraction method described in the application has smaller calculation and development costs, and has low requirements for sample size, and is suitable for both large sample set feature extraction and small sample set feature extraction.
[0074] (3) The fingerprint of the gas sensor response curve contains rich gas component information, and the mode and amplitude distribution of the response fingerprint can fully reflect the type and concentration information of the gas. Compared with the traditional feature extraction method based on statistics and deep learning, the fingerprint feature obtained by using the feature extraction method of the application has high interpretability and compactness, and highly concentrates the gas type and concentration information contained in the response fingerprint. BRIEF DESCRIPTION OF DRAWINGS
[0075] Figure 1 The figure is the overall flow chart of the gas sensor response fingerprint feature extraction method in the embodiment of the application.
[0076] Figure 2 The figure is the response data graph of the SnO2 semiconductor gas sensor to four kinds of 100 ppm toxic gases in embodiment 1 of the application.
[0077] Wherein (a) is acetone, (b) is butanone, (c) is ethanol, and (d) is diethyl ether.
[0078] Figure 3 The figure is the fitting curve graph of the response data of different types of toxic gases in embodiment 1 of the application.
[0079] Wherein (a) is acetone, (b) is butanone, (c) is ethanol, and (d) is diethyl ether.
[0080] Figure 4 The figure is the fingerprint feature distribution graph under different cumulative information entropy contribution rates in embodiment 1 of the application.
[0081] Wherein (a) is 22.2%, (b) is 39.8%, (c) is 62.4%, and (d) is 100%.
[0082] Figure 5 The figure is the response data graph of the SnO2 semiconductor gas sensor to different concentrations of ethanol gas in embodiment 2 of the application.
[0083] Figure 6 The figure is the fitting curve graph of the response data of different concentrations of ethanol gas in embodiment 2 of the application.
[0084] Wherein (a) is 100 ppm ethanol, (b) is 200 ppm ethanol, (c) is 300 ppm ethanol, (d) is 400 ppm ethanol, and (e) is 500 ppm ethanol.
[0085] Figure 7 The figure is the distribution of the extracted gas concentration fingerprint feature and its linear relationship graph in embodiment 2 of the application. DETAILED DESCRIPTION
[0086] The specific embodiments of the present application are described in further detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present application, but are not used to limit the scope of the present application.
[0087] Example 1
[0088] Taking a SnO2 semiconductor gas sensor as an example, the time sequence response data of 100 ppm of acetone, butanone, ethanol and diethyl ether is obtained in a dynamic temperature modulation mode, the time sequence response data is extracted by the method described in the present application, and the identification task faced by the extracted features is a gas type classification task, as shown in Figure 1 The specific process includes:
[0089] Step 1, obtaining time sequence response data of a gas sensor;
[0090] In this embodiment:
[0091] The gas sensor is a SnO2 semiconductor gas sensor;
[0092] The time sequence response data is dynamic temperature modulation response data, the working temperature range is 135-375 ℃, the temperature modulation period is 50 s, the data sampling frequency is 5 Hz, and each data contains 250 data points, as shown in Figure 2 , wherein (a), (b), (c), (d) respectively represent the response curves of acetone, butanone, ethanol and diethyl ether;
[0093] The time sequence response data is a one-dimensional time sequence;
[0094] The label of the time sequence response data is labeled as a gas type;
[0095] The time sequence response data contains 4 data categories, which are acetone, butanone, ethanol and diethyl ether;
[0096] The time sequence response data contains 10 time sequence response data under each category.
[0097] Step 2, based on curve parameter fitting, obtaining fitting parameters of the time sequence response data as a feature sequence, and constructing a feature sequence data set;
[0098] Step 2.1, based on curve parameter fitting, the time sequence response data of the gas sensor is fitted, and the formula is as follows:
[0099] ;
[0100] Wherein, f(t) is the time sequence response data, F(t) is the parameter fitting model, and R(t) is the residual term;
[0101] In this embodiment, the parameter fitting model F(t) is a Fourier series, and the formula is as follows:
[0102]
[0103] Step 2.2, calculate the fitting performance index according to f(t) and F(t), and compare it with the index requirement, if the fitting performance index meets the index requirement, execute step 2.3, if the fitting performance index does not meet the index requirement, adjust the parameter fitting model F(t), and execute step 2.1;
[0104] In this embodiment, the fitting curve of the time series response data is as shown in the following figure: Figure 3 , wherein (a), (b), (c), (d) respectively represent acetone, butanone, ethanol, and diethyl ether, the fitting performance index is the determination coefficient R 2 , the fitting performance index requirement is R 2 > 0.95, the average value of the fitting performance index (R 2 ) of different categories of response data is as shown in Table 1, the fitting performance index meets the index requirement, and step 2.3 is continued to be executed;
[0105] Gas species Response data tag Average R 2 ]] 100 ppm acetone 0 0.9968>0.95 100 ppm butanone 1 0.9982>0.95 100 ppm ethanol 2 0.9913>0.95 100 ppm diethyl ether 3 0.9987>0.95
[0106] Step 2.3, obtain the model parameters of the parameter fitting model F(t), and collect all the model parameters into a sequence {x(i)} as the extracted feature sequence;
[0107] In this embodiment,
[0108] The x(i) is each feature, including a0, a1-a8, b1-b8;
[0109] The i represents the feature index in the feature sequence, and the value of i satisfies i = 1, 2,..., n;
[0110] The n represents the number of features of the feature sequence, which is 17 here;
[0111] Step 2.4, execute the operations of steps 2.1-2.3 on each time series response data to obtain a feature sequence dataset;
[0112] In this embodiment, the feature sequence dataset contains 4 categories, and each category contains 10 feature sequences;
[0113] The jth feature sequence in the cth category is represented as follows: i = 1, 2,..., 17.
[0114] Step 3, calculate the inter-class and intra-class scatter ratio R(i) of each feature in the feature sequence;
[0115] Step 3.1, calculate the average value μ of each feature by category c (i), the formula is as follows: ;
[0116] Wherein, μ c (i) is the average value of the i-th feature under the c-th category, and c is 1-C;
[0117] Step 3.2, calculate the inter-class dispersion S of each feature b (i), the formula is as follows: ;
[0118] Wherein, μ p (i) and μ q (i) are the average values of the i-th feature under the p-th category and the q-th category;
[0119] Step 3.3, calculate the intra-class dispersion S w (i) of each feature, the formula is as follows: ;
[0120] Step 3.4, calculate the inter-class intra-class dispersion ratio R(i) of each feature, the formula is as follows:
[0121] ;
[0122] In this embodiment: the calculation results of μ c (i) in step 3 are shown in Table 2;
[0123] μ c (i)]]> 1 2 3 4 5 6 7 8 9 [CDATA[μ0(i)]] 2466.2 -341.8 1166.2 -478.4 378.1 29.8 476.7 5.6 283.3 [μ1(i)] 3101.7 -293.2 1061.3 -394.3 502.1 20.6 569.3 48.6 319.6 [CDATA[μ2(i)]] 1907.6 -29.1 534.6 -14.7 174.7 -122.9 249 -38.6 104.3 [Mu3(i)] 750.5 -593.1 707.7 -335.8 670.2 24 680.8 235.6 417.3 μ c (i)]]> 10 11 12 13 14 15 16 17 - [CDATA[μ0(i)]] 191.9 85.8 110.8 59.8 66.4 -52.7 54.3 0.2 - [μ1(i)] 189.6 129.9 95.5 48.5 59.3 -36.1 25.3 2.7 - [CDATA[μ2(i)]] -66. 124.2 -11. 18.4 -16.5 42.9 -0.8 -25 - [μ3(i)] 292.8 200.3 197.1 24.8 102.6 -24.4 20.8 -18.4 -
[0124] S b (i), S w (i) and R(i) are shown in Table 3;
[0125] i 1 2 3 4 5 6 S b (i)] 11951800.2 641126.3 1052579.1 493093.6 523126.2 65684.0 <![CDATA[S w (i)]]> 734.5 350.8 614.0 223.2 146.0 95.9 R(i) 16271.0 1827.6 1714.3 2208.9 3582.8 684.8 i 7 8 9 10 11 12 S b (i)]]> 403509.9 174518.5 205161.7 281700.2 27308.3 87588.6 S w (i)]]> 157.4 61.7 141.8 45.3 55.7 48.0 R(i) 2563.4 2827.4 1446.4 6220.1 490.6 1825.7 i 13 14 15 16 17 - S b (i)]]> 4593.5 30086.8 21160.6 6199.4 2246.2 - S w (i)]]> 27.3 9.5 9.9 18.0 4.0 - R(i) 168.4 3165.3 2132.3 343.6 567.5 -
[0126] Step 4, calculate the information entropy contribution rate H(i) of each feature according to the inter-class intra-class dispersion ratio R(i);
[0127] Step 4.1, normalize the inter-class intra-class dispersion ratio R(i) to obtain the normalized inter-class intra-class dispersion ratio R'(i), the formula is as follows:
[0128] ;
[0129] Wherein, the R(i) min and R(i) max are the minimum value and the maximum value of R(i) respectively;
[0130] Step 4.2, Softmax calculation is performed on the normalized inter-class intra-class scatter ratio R'(i) to map R'(i) to a probability distribution P(i), and the formula is as follows:
[0131] ;
[0132] Wherein, P(i) satisfies the following conditions:
[0133] ;
[0134] Step 4.3, according to P(i), the information entropy H of the whole feature sequence is calculated, and the formula is as follows:
[0135] ;
[0136] Step 4.4, the information entropy contribution rate H(i) of each feature is calculated, and the formula is as follows:
[0137] ;
[0138] In this embodiment: the calculated value of H in step 4 is 4.03;
[0139] The calculated values of R'(i), P(i) and H(i) are shown in Table 4.
[0140] i 1 2 3 4 5 6 7 8 9 R'(i) 1.00 0.10 0.10 0.13 0.21 0.03 0.15 0.17 0.08 P(i) 0.13 0.05 0.05 0.05 0.06 0.05 0.06 0.06 0.05 H(i) 0.095 0.056 0.056 0.057 0.060 0.054 0.058 0.058 0.055 i 10 11 12 13 14 15 16 17 - R'(i) 0.38 0.02 0.10 0.00 0.19 0.12 0.01 0.02 - P(i) 0.07 0.05 0.05 0.05 0.06 0.05 0.05 0.05 - H(i) 0.067 0.053 0.056 0.052 0.059 0.057 0.053 0.053 -
[0141] Step 5, feature screening is performed according to the information entropy contribution rate H(i);
[0142] Step 5.1, according to the numerical value of H(i), H(i) is reordered in descending order to obtain H(i'), and the permutation relationship σ of the element positions of H(i') and H(i) is calculated; wherein the value of i' satisfies i'=1, 2,..., n;
[0143] The calculation process of the permutation relationship σ of the element positions is as follows:
[0144] The sequence length of H(i) and H(i') is n, so the index set of the two is represented as [n] = {1, 2,..., n};
[0145] The element position index in the original sequence H(i) is i, and i∈[n]; the element position index in the sorted sequence H(i') is i', and i'∈[n];
[0146] The element position order adjustment is defined by a permutation function σ(i'): [n]→[n]; the permutation function σ(i') represents the original position corresponding to the adjusted position i', that is, the index of the element at the new position i' in the original sequence;
[0147] Step 5.2, calculate the cumulative information entropy contribution rate ∑H(i') of the first i' features in H(i'), the formula is as follows:
[0148] ;
[0149] Wherein, k = 1, 2,..., i';
[0150] Step 5.3, compare ∑H(i') with the cumulative information entropy contribution rate threshold δ, find the minimum i' value i' that satisfies ∑H(i') not less than δ min , the formula is as follows:
[0151] ;
[0152] Wherein, the δ is the cumulative information entropy contribution rate threshold;
[0153] Step 5.4, obtain the index corresponding to the first i' min Information entropy contribution rate in H(i') }, the formula is as follows:
[0154] ;
[0155] Step 5.5, according to the permutation relationship σ of the element position, calculate the position index } of the sorted element position index } in the original sequence, the formula is as follows:
[0156] ;
[0157] Step 5.6, according to the position index } from the original feature sequence Obtain the screened feature As the extracted fingerprint feature, the formula is as follows:
[0158] .
[0159] In this embodiment:
[0160] The value of i' in step 5.1 satisfies i' = 1, 2,..., 17; the calculation result of H(i') is shown in Table 5;
[0161] i 1 2 3 4 5 6 7 8 9 H(i') 0.095 0.067 0.060 0.059 0.058 0.058 0.057 0.057 0.056 i 10 11 12 13 14 15 16 17 - H(i') 0.056 0.056 0.055 0.054 0.053 0.053 0.053 0.052 -
[0162] The sequence length of H(i) and H(i') is 17, so the index set of both can be expressed as [n] = {1, 2,..., 17};
[0163] The element position order adjustment in step 5.1 is defined by a permutation function σ(i'): [n]→[n], which is expressed as follows:
[0164] σ: [1, 10, 5, 14, 7, 8, 4, 15, 2, 12, 3, 9, 6, 17, 11, 16, 13];
[0165] The calculation result of ∑H(i') in step 5.2 is shown in Table 6;
[0166] i 1 2 3 4 5 6 7 8 9 ∑H(i') 0.095 0.162 0.222 0.282 0.340 0.398 0.455 0.512 0.568 i 10 11 12 13 14 15 16 17 - ∑H(i') 0.624 0.680 0.735 0.788 0.842 0.895 0.948 1.000 -
[0167] The cumulative information entropy contribution rate threshold δ in step 5.3 is 60%, and the calculation value of i' is 10; min
[0168] The calculation value of { in step 5.4 is {1, 2, 3, 4, 5, 6, 7, 8, 9, 10};
[0169] The calculation value of { in step 5.5 is {1, 10, 5, 14, 7, 8, 4, 15, 2, 12};
[0170] The calculation value of in step 5.6 is { , , , , , , , , , , }.
[0171] The distribution of fingerprint features under different cumulative information entropy contribution rates ∑H(i') is shown in Figure 4 Fig. 6, where (a), (b), (c), and (d) represent the distribution of fingerprint features under ∑H(i') of 22.2%, 39.8%, 62.4%, and 100%, respectively, and the fingerprint feature sequence is reduced to two dimensions for distribution visualization by t-SNE. Figure 4 It can be seen that when ∑H(i') is less than 60%, the extracted fingerprint information is insufficient, resulting in unclear boundaries of the class distribution of the fingerprint features; when ∑H(i') is greater than 60%, the extracted fingerprint information is sufficient, and the boundaries of the class distribution of the fingerprint features are clear, which can significantly distinguish the gas categories. Considering the compactness and interpretability of the features, 10 fitting parameters are extracted from each original time series response data as the fingerprint features in this embodiment.
[0172] Embodiment 2:
[0173] Taking a SnO2 semiconductor gas sensor as an example, in a static temperature test mode, time series response data of 100-500 ppm of ethanol gas is obtained, and the time series response data is subjected to fingerprint feature extraction using the method described in the application, and the identification task faced by the extracted features is a gas concentration prediction task, as shown in Figure 1 , and the specific process includes:
[0174] Step 1, obtaining time series response data of a gas sensor;
[0175] In this embodiment:
[0176] The gas sensor is a SnO2 semiconductor gas sensor;
[0177] The time series response data is static temperature test response data, the working temperature is 350 ℃, the response time is 30 s, the data sampling frequency is 2 Hz, and each piece of data contains 60 data points, as shown in Figure 5 ;
[0178] The time series response data is a one-dimensional time series;
[0179] The label of the time series response data is labeled as a concentration category;
[0180] The time series response data contains 5 data categories, which are 100 ppm, 200 ppm, 300 ppm, 400 ppm and 500 ppm;
[0181] The time series response data contains 10 time series response data under each category.
[0182] Step 2, based on curve parameter fitting, obtaining fitting parameters of the time series response data as a feature sequence, and constructing a feature sequence dataset;
[0183] Step 2.1, based on curve parameter fitting, data fitting is performed on the time series response data of the gas sensor, and the formula is as follows:
[0184] ;
[0185] Wherein, f(t) is time series response data, F(t) is a parameter fitting model, and R(t) is a residual term;
[0186] In this embodiment:
[0187] The parameter fitting model F(t) is a polynomial function, and the formula is as follows:
[0188]
[0189] Step 2.2, calculate the fitting performance index according to f(t) and F(t), and compare it with the index requirement, if the fitting performance index meets the index requirement, execute step 2.3, if the fitting performance index does not meet the index requirement, adjust the parameter fitting model F(t), and execute step 2.1;
[0190] In this embodiment:
[0191] The fitting curve of the time series response data is as shown in Figure 6 , wherein (a), (b), (c), (d) and (e) respectively represent the fitting curves under 100 ppm ethanol, 200 ppm ethanol, 300 ppm ethanol, 400 ppm ethanol and 500 ppm ethanol;
[0192] The fitting performance index is the coefficient of determination R 2 , the fitting performance index requirement is R 2 > 0.95, the average value of the fitting performance index (R 2 ) of different categories of response data is as shown in Table 7, the fitting performance index meets the index requirement, and step 2.3 is continued to be executed;
[0193] Gas concentration Response data tag Average R 2 ]] 100 ppm ethanol 0 0.9700>0.95 200 ppm ethanol 1 0.9927>0.95 300 ppm ethanol 2 0.9757>0.95 400 ppm ethanol 3 0.9879>0.95 500 ppm ethanol 4 0.9909>0.95
[0194] Step 2.3, obtain the model parameters of the parameter fitting model F(t), and all the model parameters are summarized into a sequence {x(i)} as the extracted feature sequence;
[0195] In this embodiment:
[0196] The x(i) is each feature, including a0, a1 and a2;
[0197] The i represents the feature index in the feature sequence, and the value of i satisfies i = 1, 2,..., n;
[0198] The n represents the number of features of the feature sequence, which is 3 here;
[0199] Step 2.4, execute steps 2.1-2.3 for each time series response data to obtain a feature sequence data set;
[0200] In this embodiment:
[0201] The feature sequence dataset contains 5 categories, and each category contains 10 feature sequences.
[0202] The jth feature sequence in the cth category is represented as follows:
[0203] i = 1, 2, 3.
[0204] Step 3, calculate the inter-class intra-class scatter ratio R(i) of each feature in the feature sequence;
[0205] Step 3.1, calculate the mean value μ c (i) of each feature by category, as follows:
[0206] ;
[0207] Where μ c (i) is the mean value of the ith feature in the cth category, and c takes values from 1 to C.
[0208] Step 3.2, calculate the inter-class scatter S b (i) of each feature, as follows:
[0209] ;
[0210] Where μ p (i) and μ q (i) are the mean values of the ith feature in category p and category q.
[0211] Step 3.3, calculate the intra-class scatter S w (i) of each feature, as follows:
[0212] ;
[0213] Step 3.4, calculate the inter-class intra-class scatter ratio R(i) of each feature, as follows:
[0214] ;
[0215] In this embodiment:
[0216] The calculation values of μ c (i), S b (i), S w (i) and R(i) are shown in Table 8 and Table 9.
[0217] μ c (i)]]> 1 2 3 [CDATA[μ0(i)]] -0.0046 0.51 3.86 [μ1(i)] -0.008 0.88 1.92 [Mu2(i)] -0.008 0.92 0.34 [Mu3(i)] -0.0082 1.02 0.80 [μ4(i)] -0.0096 1.18 2.52
[0218] i 1 2 3 S b (i)]]> 0.00006824 1.232864 39.356216 S w (i)]]> 0.00000144 0.00656 0.177016 R(i) 47.39 187.94 222.33
[0219] Step 4, calculate the information entropy contribution rate H(i) of each feature according to the between-intra cluster scatter ratio R(i);
[0220] Step 4.1, normalize the between-intra cluster scatter ratio R(i) to obtain the normalized between-intra cluster scatter ratio R'(i), and the formula is as follows:
[0221] ;
[0222] Wherein, the R(i) min and R(i) max are the minimum and maximum values of R(i) respectively;
[0223] Step 4.2, perform Softmax calculation on the normalized between-intra cluster scatter ratio R'(i) to map R'(i) to a probability distribution P(i), and the formula is as follows:
[0224] ;
[0225] Wherein, P(i) satisfies the following conditions:
[0226] ;
[0227] Step 4.3, calculate the information entropy H of the entire feature sequence according to P(i), and the formula is as follows:
[0228] ;
[0229] Step 4.4, calculate the information entropy contribution rate H(i) of each feature, and the formula is as follows:
[0230] ;
[0231] In this embodiment: the calculated value of H in step 4 is 1.48;
[0232] The calculated values of R'(i), P(i) and H(i) are shown in Table 10.
[0233] i 1 2 3 R'(i) 0.00 0.80 1.00 P(i) 0.17 0.38 0.46 H(i) 0.292 0.359 0.349
[0234] Step 5, perform feature screening according to the information entropy contribution rate H(i);
[0235] Step 5.1, reorder H(i) according to the numerical value of H(i) from large to small to obtain H(i'), and calculate the permutation relationship σ of H(i') and H(i) element position; wherein the value of i' satisfies i'=1, 2,..., n;
[0236] The calculation process for the permutation relationship σ of the element positions is as follows:
[0237] The sequence lengths of H(i) and H(i') are both n, so their index set is represented as [n] = {1, 2, ...,n};
[0238] In the original sequence H(i), the element index is i, and i∈[n]; in the sorted sequence H(i'), the element index is i', and i'∈[n];
[0239] The element position order is adjusted by the permutation function σ(i'): [n]→[n]; the permutation function σ(i') represents the original position corresponding to the adjusted position i', that is, the index of the element at the new position i' in the original sequence;
[0240] Step 5.2: Calculate the cumulative information entropy contribution rate ∑H(i') of the first i' features in H(i'), as shown in the following formula:
[0241] ;
[0242] Where k = 1, 2, ..., i';
[0243] Step 5.3: Compare ∑H(i') with the cumulative information entropy contribution rate threshold δ, and find the smallest i' value i' that satisfies ∑H(i') not less than δ. min The formula is as follows: ;
[0244] Wherein, δ is the cumulative information entropy contribution rate threshold;
[0245] Step 5.4: Obtain the first i' in H(i') min The index corresponding to the information entropy contribution rate { The formula is as follows:
[0246] ;
[0247] Step 5.5: Calculate the index of the sorted elements based on the permutation relationship σ. } Position index in the original sequence { The formula is as follows:
[0248] ;
[0249] Step 5.6: Based on the position index { From the original feature sequence Extracting filtered features The extracted fingerprint features are represented by the following formula:
[0250] .
[0251] In this embodiment:
[0252] The values of i' mentioned in step 5.1 satisfy i'=1, 2, 3;
[0253] The calculated values of H(i') in step 5.1 are shown in Table 11;
[0254] i 1 2 3 H(i') 0.359 0.349 0.292
[0255] The sequence lengths of H(i) and H(i') mentioned in step 5.1 are both 3, so their index sets can be represented as [n] = {1, 2, 3};
[0256] The element position order adjustment in step 5.1 is defined by a permutation function σ(i'): [n]→[n], as follows:
[0257] σ: [2, 3, 1];
[0258] The calculated values of ∑H(i') in step 5.2 are shown in Table 12;
[0259] i 1 2 3 ∑H(i') 0.359 0.708 1.000
[0260] In step 5.3, the cumulative information entropy contribution rate threshold δ is set to 30%, and the i' min The calculated value is 1;
[0261] Step 5.4 describes { The calculated value of} is {1};
[0262] Step 5.5 describes { The calculated value of} is {2};
[0263] Step 5.6 The calculated value is { }
[0264] The distribution of fingerprint features and their linear relationship of gas sensor concentration response data extracted based on the method described in this invention are as follows: Figure 7 As shown. By Figure 7 As can be seen, in this embodiment, only the coefficients of the first-order terms are extracted from the fitting parameters as fingerprint features. The extracted fingerprint features have an ideal linear relationship with the gas concentration. The linear prediction model obtained with the fingerprint features can achieve continuous prediction of gas concentration.
[0265] The various embodiments in this application are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.
[0266] The scope of the protection made on the basis of the aspects covered by the claims should not be restricted to the examples described above, but rather those going beyond the present disclosure comprise every novel object that comes into the protection scope of the claims or that based on its properties and its citation in this document can be included in the categories of aspects covered by the claims when they are interpreted according to the principles of patent law.
Claims
1. A method for extracting response fingerprint features of a gas sensor based on parameter fitting, characterized in that, The method comprises the following steps: Step 1, obtaining time sequence response data of a gas sensor; Step 2, obtaining fitting parameters of the time sequence response data as a feature sequence based on curve parameter fitting, and constructing a feature sequence dataset; Step 3, calculating the between-class and within-class scatter ratio R(i) of each feature in the feature sequence; Step 2.1, performing data fitting on the time sequence response data of the gas sensor based on curve parameter fitting, and the formula is as follows: ; Wherein, f(t) is the time sequence response data, F(t) is the parameter fitting model, and R(t) is the residual term; Step 2.2, calculating the fitting performance index according to f(t) and F(t), and comparing it with the index requirement, if the fitting performance index meets the index requirement, executing step 2.3, if the fitting performance index does not meet the index requirement, adjusting the parameter fitting model F(t), and executing step 2.1; Step 2.3, obtaining the model parameters of the parameter fitting model F(t), and all the model parameters are summarized into a sequence {x(i)} as the extracted feature sequence, wherein x(i) is each feature, i represents the feature index in the feature sequence, and i satisfies i = 1, 2,..., n, and n represents the number of features in the feature sequence; Step 2.4, performing the operations of steps 2.1-2.3 on each time sequence response data to obtain the feature sequence dataset; The feature sequence dataset contains C categories in total, and each category contains N c characteristic sequences. The representation of the jth feature sequence under the cth class is as follows: i = 1, 2,..., n; Step 4, calculating the information entropy contribution rate H(i) of each feature according to the between-class and within-class scatter ratio R(i); Step 5, performing feature screening according to the information entropy contribution rate H(i) to extract the fingerprint feature.
2. The method according to claim 1, wherein, The gas sensor in step 1 includes a semiconductor gas sensor, an electrochemical gas sensor, a catalytic combustion gas sensor, a MEMS gas sensor, and a surface acoustic wave gas sensor; The time sequence response data is a one-dimensional time sequence, including static temperature test response data and dynamic temperature modulation response data; wherein the static temperature test is a working mode with constant working temperature of the gas sensor, and the dynamic temperature modulation is a working mode with periodic change of working temperature of the gas sensor; The time sequence response data needs to be labeled according to specific tasks, specifically: when performing a gas species identification task, the label is labeled as a gas species; when performing a gas concentration prediction task, the label is labeled as a concentration type; The timing response data includes C data categories in total, each category including N c bar timing response data.
3. The method of claim 1, wherein, The parameter fitting model in step 2.1 includes a linear function, an exponential function, a power function, a polynomial function, a Fourier series, a Gaussian function, a trigonometric function, or a sum of the above functions.
4. The method of claim 1, wherein, The fit performance indicator described in step 2.2 is one or several of the coefficient of determination R 2 , the mean absolute error MAE, the mean squared error MSE and the root mean squared error RMSE.
5. The method of claim 1, wherein, The step 3 specifically comprises the following steps: Step 3.
1. Calculate the mean value μ of each feature by class c (i) as follows: ; wherein μ c (i) is the average value of the i-th feature under the c-th category, and c takes values from 1 to C; Step 3.2, compute the between-class scatter S for each feature b (i) as follows: ; where μ p (i) and μ q (i) is the average value of the i-th feature in category p and category q, respectively. Step 3.3, compute the within-class scatter S for each feature w (i) as follows: ; Step 3.4, calculating the between-class and within-class scatter ratio R(i) of each feature, and the formula is as follows: 。 6. The method of claim 1, wherein, The step 4 specifically comprises the following steps: Step 4.1, performing normalization calculation on the between-class and within-class scatter ratio R(i) to obtain the normalized between-class and within-class scatter ratio R'(i), and the formula is as follows: ; wherein R(i) min and R(i) max are the minimum and maximum values of R(i), respectively; Step 4.2, performing Softmax calculation on the normalized between-class and within-class scatter ratio R'(i) to map R'(i) to a probability distribution P(i), and the formula is as follows: , Wherein, P(i) satisfies the following conditions: ; Step 4.3, calculate the information entropy H of the whole feature sequence according to P(i), and the formula is as follows: ; Step 4.4, calculate the information entropy contribution rate H(i) of each feature, and the formula is as follows: 。 7. The method of claim 1, wherein, The step 5 specifically includes the following steps: Step 5.1, reorder H(i) according to the numerical size of H(i) in descending order to obtain H(i'), and calculate the permutation relationship σ of H(i') and the element position of H(i); wherein the value of i' satisfies i'=1, 2,..., n; The calculation process of the permutation relationship σ of the element position is as follows: The sequence length of H(i) and H(i') is n, so the index set of the two is represented as [n] = {1, 2,..., n}; The element position index of the original sequence H(i) is i, and i∈[n]; the element position index of the sorted sequence H(i') is i', and i'∈[n]; The order of the element position is adjusted by the permutation function σ(i'): [n]→[n]; the permutation function σ(i') represents the original position corresponding to the adjusted position i', that is, the index of the element in the original sequence after the new position i'; Step 5.2, calculate the cumulative information entropy contribution rate ∑H(i') of the first i' features in H(i'), and the formula is as follows: ; Wherein, k = 1, 2,..., i'; Step 5.3, compare ∑H(i') with the accumulated information entropy contribution rate threshold δ, find the minimum i' value i' satisfying ∑H(i') not less than δ min The formula is as follows: , Wherein, the δ is the cumulative information entropy contribution rate threshold; Step 5.4, obtain the index {i'}* corresponding to the first i' information entropy contribution rate in H(i'), the formula is as follows: min Step 5.4, obtain the index {i'}* corresponding to the first i' information entropy contribution rate in H(i'), the formula is as follows: ; Step 5.5, according to the permutation relationship σ of the element position, calculate the position index {i*} of the sorted element position index {i'} in the original sequence, and the formula is as follows: ; Step 5.
6. Obtain filtered features from the original feature sequence according to the position index {i*} As the extracted fingerprint features, the formula is as follows: 。
Citation Information
Patent Citations
Method for detecting living body fingerprint based on thin plate spline deformation model
CN101226589A
Lightweight time sequence information fusion perception control method and system
CN120358211A