Gas Concentration Prediction Method Based on Improved DBSCAN Algorithm and Multi-Scale LSTM Neural Network
Through the improved DBSCAN clustering algorithm and multi-scale LSTM neural network, the problem of improper processing of abnormal data in gas concentration detection is solved, the prediction accuracy is improved, the cost and complexity is reduced, and the accurate monitoring and prediction of changes in atmospheric gas concentration is achieved.
Patent Information
- Application Number
- CN202311124140.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-01
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2043-09-01
AI Technical Summary
The prior art has improper processing of abnormal data in gas concentration detection, resulting in reduced measurement accuracy, high equipment cost and complex operation.
An improved DBSCAN clustering algorithm is used to screen and correct abnormal data points, and a gas concentration prediction model is constructed in combination with a multi-scale LSTM neural network, and the local correlation characteristics between variables are extracted through feature extraction networks.
It improves the accuracy of gas concentration prediction, reduces the interference of abnormal data on experiments, reduces equipment cost and operation complexity, and realizes accurate monitoring and prediction of changes in atmospheric gas concentration.
Smart Images

Figure CN117171641B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of gas concentration detection, and in particular to a gas concentration prediction method based on an improved DBSCAN algorithm and a multi-scale LSTM neural network. Background Technique
[0002] The atmosphere contains various gases. When the external environment changes, the concentration of the gases will also change accordingly. Therefore, by monitoring the change of the gas concentration in the atmosphere, the change trend of the environment can be observed. Many gases in the atmosphere are associated with pollutants. By monitoring the change of the gas concentration in the atmosphere, the pollution degree and air quality of the atmosphere can also be known.
[0003] Laser spectroscopy technology is an important method for detecting the gas concentration in the atmosphere. Among many laser spectroscopy measurement technologies, direct absorption spectroscopy technology is the simplest and most convenient, but it also needs to perform a series of processing on the spectral data to be able to invert the gas concentration at the current moment. Although wavelength modulation technology and cavity ring-down technology can accurately measure the gas concentration, the required instrument equipment is expensive and requires professional technical personnel to operate. In the process of these laser spectroscopy measurements, some abnormal data will inevitably appear, and the abnormal data will reduce the measurement accuracy. To solve the above problems, a method for inverting and predicting the gas concentration by constructing a prediction model through software is proposed. Summary of the Invention
[0004] The present invention is to solve the deficiencies of the above-mentioned existing technologies, and proposes a gas concentration prediction method based on an improved DBSCAN algorithm and a multi-scale LSTM neural network, in order to be able to screen and correct abnormal data, so as to improve the accuracy of gas concentration prediction, so as to achieve the purpose of monitoring and predicting the change of the gas concentration in the atmosphere.
[0005] The present invention adopts the following technical solutions to achieve the above invention purpose:
[0006] A gas concentration prediction method based on an improved DBSCAN algorithm and a multi-scale LSTM neural network of the present invention is characterized by including the following steps:
[0007] Step 1: Use a gas absorption spectrum experimental device to obtain n groups of gas absorption spectrum data in the atmosphere, and calculate the integral area of the absorption peak of each group of gases, denoted as A = {A1, A2,..., A i ,..., A n}, where A i represents the integral area of the absorption peak calculated from the i-th group of gas absorption spectrum data;
[0008] Obtain the temperature data set T = {T1, T2,..., Ti ,...,T n}, and the humidity dataset R = {R1, R2,..., R i ,..., R n}, where T i and R i respectively represent the temperature and humidity corresponding to the i-th group of gas absorption spectrum data;
[0009] Obtain the historical true concentration data of n groups of gases, denoted as C = {C1, C2,..., C i ,...C n}, where C i represents the historical true concentration corresponding to the i-th group of gas absorption spectrum data; n represents the number of data groups;
[0010] Step 2. Use the improved DBSCAN clustering algorithm to screen out abnormal outlier data points;
[0011] Step 2.1. Define the minimum number of samples in the neighborhood as P min , define and randomly initialize the radius ε of the neighborhood, and define and randomly initialize the threshold σ;
[0012] Step 2.2. Initialize i = 1;
[0013] Step 2.3. Take A i as the i-th sample point, and judge whether the number of sample points P i within the neighborhood with a radius of ε centered on A i is greater than or equal to P min . If so, it means that A i is a core point, and store A i in the core point set H, and form the i-th cluster Q i consisting of the core point A i and all sample points within its radius ε; otherwise, directly execute Step 2.4;
[0014] Step 2.4. After assigning i + 1 to i, judge whether i > n holds. If it holds, it means that several clusters and the final core point set H are obtained. Otherwise, return to Step 2.3 and execute sequentially;
[0015] Step 2.5. Judge whether there are the same sample points in several clusters. If there are, merge the clusters corresponding to the same sample points into one cluster. Otherwise, do not merge; thus forming a cluster set Q without intersection with each other; and form an outlier set I consisting of the sample points in A that are not included in the cluster set Q;
[0016] Step 2.6. Calculate the distances between each outlier in the outlier set I and each core point in the core point set H to obtain the distance matrix 1 ≤ i ≤ e, 1 ≤ j ≤ f, where e represents the number of outliers and f represents the number of core points, and l ij represents the distance between the i-th outlier and the j-th core point;
[0017] Step 2.7: Obtain the minimum value in each row vector of the distance matrix L and get the nearest distance matrix L min =[l 1,min ,..., l i,min ,…, l e,min , where l i,min represents the distance between the i-th outlier and the nearest core point;
[0018] Step 2.8: Judge whether the difference S i,min between l i and the radius ε is less than the threshold σ. If so, retain the corresponding difference; otherwise, set the corresponding difference to -1, so as to obtain the difference matrix S = [S1, S2, …, S i ,…, S e ;
[0019] Step 2.9: Judge whether there is a difference greater than 0 in S. If so, select the minimum value S min among the differences greater than 0 in S, and assign ε + S min to ε and σ - S min to σ, then return to Step 2.2 and execute sequentially; otherwise, it means that the final outlier set I is obtained;
[0020] Step 3: Correct the abnormal data in the measurement and construct the neural network dataset;
[0021] Step 3.1: Obtain e nearest core points according to L min , and replace the corresponding outliers with the average value of all sample points within the neighborhood range of each nearest core point, so as to obtain the corrected outlier set I'. After merging it with the cluster set Q, the corrected gas absorption spectrum data A' is obtained;
[0022] Step 3.2: Construct the dataset X = {C, A', T, R} of the LSTM neural network from C, A', T, and R, and perform normalization processing on the dataset X to obtain the normalized X';
[0023] Step 4: Use the variable feature extraction network to extract the local correlation features between variables in X';
[0024] Step 4.1: Combine the variables in X' pairwise to obtain a three-dimensional matrix X of dimension (a - 1)! × n × 2 cnn; where, (a - 1)! represents the number of combinations,! represents factorial, and a represents the number of variables in X';
[0025] Step 4.2: Input the three-dimensional matrix X cnn into the variable feature extraction network, and use a three-dimensional convolutional layer with a convolutional kernel size of (a - 1)! × q × 2 and a stride of 1 to extract the features between every two of the a variables, obtaining a three-dimensional feature matrix Q1 with a dimension of (a - 1)! × (n - q + 1) × 1; where, q represents the window length;
[0026] Step 4.3: After the variable feature extraction network reconstructs the dimension of the three-dimensional feature matrix Q1, a two-dimensional feature matrix Q2 with a dimension of (n - q + 1) × (a - 1)! is obtained;
[0027] Step 4.4: The variable feature extraction network uses a two-dimensional convolutional layer with a convolutional kernel size of 1 × (a - 1)! and a number of convolutional kernels of p and a stride of 1 to extract features from the two-dimensional feature matrix Q2, obtaining p one-dimensional feature matrices with a dimension of (n - q + 1) × 1. Finally, after splicing the p one-dimensional feature matrices, a two-dimensional feature matrix Q3 with a dimension of (n - q + 1) × p is obtained;
[0028] Step 5: Use a multi-scale LSTM neural network for prediction;
[0029] Step 5.1: Divide the two-dimensional feature matrix Q3 into a training set and a test set, and use the training set as the input of the LSTM neural network and C as the output of the LSTM neural network, so as to train the LSTM neural network to obtain a gas concentration prediction model at one scale;
[0030] Step 5.2: After changing the window length q r times, respectively return to Step 4.2 and execute sequentially, so as to obtain r gas concentration prediction models at different scales;
[0031] Step 5.3: Input the test set into the gas concentration prediction models at r scales respectively, and use Equation (1) to perform weighted summation on the gas concentrations predicted by the gas concentration prediction models at each scale, so as to obtain the predicted comprehensive value y at r scales;
[0032] y = w1y1 + w2y2 + … + w k y k + … + w r y r (1)
[0033] In Equation (1), y k represents the gas concentration predicted by the gas concentration prediction model at the kth scale, w k represents the kth weight, and there is:
[0034]
[0035] In formula (2), MSE k represents the predicted gas concentration y at the k-th scale k and the mean square error between the true concentration value C.
[0036] An electronic device of the present invention includes a memory and a processor, characterized in that the memory is used to store a program for supporting the processor to execute the gas concentration prediction method, and the processor is configured to execute the program stored in the memory.
[0037] A computer-readable storage medium of the present invention, characterized in that a computer program is stored on the computer-readable storage medium, and when the computer program is run by a processor, it executes the steps of the gas concentration prediction method.
[0038] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0039] 1. The present invention screens and corrects abnormal data in the data set by improving the DBSCAN clustering algorithm, eliminates the interference of abnormal data on the experiment, and thus improves the accuracy of gas concentration prediction.
[0040] 1. The present invention obtains the integral area of the gas absorption peak from the original gas spectrum data, combines relevant variables such as the temperature and humidity of the environment, uses a feature extraction network to extract the features between the variables, and designs an LSTM prediction model to achieve accurate prediction of the gas concentration at the next moment, providing a reference basis for atmospheric environmental gas monitoring.
[0041] 2. Compared with the traditional laser spectroscopy technology, the present invention does not require further processing of the original spectrum, ignores the influence of the pressure and optical path inside the absorption cell that basically remain unchanged during the experiment, and fully utilizes the characteristic data that has a greater impact on the gas concentration, thereby improving the accuracy rate of the gas concentration prediction result. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 is a schematic flow chart of the method of the present invention;
[0043] Figure 2 is a structural block diagram of the gas absorption spectrum experimental device of the present invention;
[0044] Figure 3 is a flow chart of the improved DBSCAN algorithm of the present invention;
[0045] Figure 4 is a process diagram of the feature extraction network of the present invention;
[0046] Figure 5This is the structural diagram of the multi-scale LSTM neural network model of the present invention;
[0047] Reference numerals in the figure: 1 is a signal generator; 2 is a laser controller; 3 is a QCL quantum cascade laser; 4 is a mirror; 5 is a long-path absorption cell; 6 is a three-way valve; 7 and 8 are vacuum valves; 9 is a vacuum pump; 10 is a pressure gauge; 11 is a photodetector; 12 is a data acquisition card; 13 is a terminal industrial control computer; 14 is a temperature and humidity meter. Specific implementation manner
[0048] In this embodiment, as Figure 1 shown, a gas concentration prediction method based on an improved DBSCAN algorithm and a multi-scale LSTM neural network takes into account that during the experiment, due to environmental changes and fluctuations generated during long-term use of the instrument, the measured data contains a small amount of abnormal data. An improved DBSCAN clustering algorithm is used to screen and correct the abnormal data, and a multi-scale LSTM neural network is constructed based on the corrected data set, thereby realizing the prediction of gas concentration values and providing a basis for monitoring gas concentration changes in the environment. Specifically, the method is carried out according to the following steps:
[0049] Step 1, as Figure 2 shown, a gas absorption spectroscopy experimental device using a tunable laser as a light source is used to collect methane CH4 absorption spectrum data in the atmosphere, obtain the original data set of the methane CH4 absorption spectrum, and use a temperature and humidity meter to obtain the temperature and humidity of the experimental environment;
[0050] Step 1.1, the QCL quantum cascade laser 3 sets the parameters of the laser through the laser controller 2 and the signal generator 1. The laser emitted by the QCL quantum cascade laser 3 enters the absorption cell through the light inlet hole of the long-path absorption cell 5 through the mirror 4, and the laser is reflected multiple times inside the absorption cell and then exits from the light outlet hole of the absorption cell;
[0051] Step 1.2, a vacuum pump 9 and a pressure gauge 10 for measuring the internal pressure of the absorption cell are provided at the gas outlet end of the long-path absorption cell 5. The gas flow rate inside the absorption cell is controlled by adjusting the vacuum valves 6 and 8 to keep the pressure inside the long-path absorption cell 5 constant;
[0052] Step 1.3, the laser enters the photodetector 11 through the light outlet hole of the long-path absorption cell 5. The photodetector 11 is connected to the data acquisition card 12 and transmits the gas absorption spectrum data to the terminal industrial control computer 13. A data acquisition program is written by itself using LabView software in the terminal industrial control computer 13 to complete the acquisition and storage of the laser spectrum data. In a specific embodiment, data is continuously measured for 12 hours, and the acquisition interval is 1 minute;
[0053] Step 2: Use a gas absorption spectroscopy experimental device to obtain n groups of gas absorption spectroscopy data in the atmosphere, and calculate the integrated area of the absorption peak for each group of gases, denoted as A = {A1, A2, …, A i , …, A n}, where A i represents the integrated area of the absorption peak calculated from the i-th group of gas absorption spectroscopy data;
[0054] Obtain the temperature data set T = {T1, T2, …, T i , …, T n} and the humidity data set R = {R1, R2, …, R i , …, R n} of the experimental environment, where T i and R i represent the temperature and humidity corresponding to the i-th group of gas absorption spectroscopy data, respectively;
[0055] Obtain the historical true concentration data of n groups of gases, denoted as C = {C1, C2, …, C i , …, C n}, where C i represents the historical true concentration corresponding to the i-th group of gas absorption spectroscopy data; n represents the number of data groups;
[0056] Step 3: As shown in Figure 3 , use the improved DBSCAN clustering algorithm to screen out abnormal outlier data points;
[0057] Step 3.1: Define the minimum number of samples in the neighborhood as P min , define and randomly initialize the radius ε of the neighborhood, and define and randomly initialize the threshold σ;
[0058] Step 3.2: Initialize i = 1;
[0059] Step 3.3: Take A i as the i-th sample point, and judge whether the number of sample points P i within the neighborhood with a radius of ε centered on A i is greater than or equal to P min . If so, it means that A i is a core point, and store A i in the core point set H, and form the i-th cluster Q i from the core point A i and all sample points within its radius ε; otherwise, directly execute Step 3.4;
[0060] Step 3.4: After assigning i + 1 to i, judge whether i > n holds. If it holds, it means that several clusters and the final core point set H are obtained; otherwise, return to Step 3.3 and execute sequentially;
[0061] Step 3.5: Determine whether there are the same sample points in several clusters. If there are, merge the clusters corresponding to the same sample points into one cluster; otherwise, do not merge, thereby forming a set of clusters Q that are mutually non - intersecting; and form an outlier set I consisting of the sample points in A that are not included in the set of clusters Q.
[0062] Step 3.6: Calculate the distances between each outlier in the outlier set I and each core point in the core point set H to obtain a distance matrix where e represents the number of outliers, f represents the number of core points, and l ij represents the distance between the i - th outlier and the j - th core point.
[0063] Step 3.7: Obtain the minimum value in each row vector of the distance matrix L and get the nearest - distance matrix L min = [l 1,min ,…,l i,min ,...,l e,min , where l i,min represents the distance between the i - th outlier and the nearest core point.
[0064] Step 3.8: Determine whether the difference between the set L min and the radius ε is less than the threshold σ. If so, retain the corresponding difference; otherwise, set the corresponding difference to - 1, thereby obtaining a difference matrix S = [S1, S2,..., S i ,..., S e , where S i represents the i - th difference.
[0065] Step 3.9: Determine whether there is a difference greater than 0 in S. If there is, select the minimum value S min greater than 0 from S, and assign ε + S min to ε and σ - S min to σ, then return to Step 3.2 and execute sequentially; otherwise, it means obtaining the final outlier set I.
[0066] Step 4: Correct the abnormal data in the measurement and construct a neural network dataset.
[0067] Step 4.1: Obtain e nearest core points according to L min and replace the corresponding outliers with the average values of all sample points within the neighborhood range of each nearest core point, thereby obtaining a corrected outlier set I’. After merging with the set of clusters Q, the corrected gas absorption spectrum data A’ is obtained.
[0068] Step 4.2: Construct the dataset X = {C, A', T, R} of the LSTM neural network from C, A', T, and R, and perform normalization processing on the dataset X to obtain the normalized X'.
[0069] In specific implementation, the maximum-minimum normalization is adopted for the dataset X to normalize the entire dataset X to the range of 0 to 1. The normalization formula is as shown in Equation (3):
[0070]
[0071] In Equation (3), max(X (i,j) ) and min(X (i,j) ) respectively represent the maximum and minimum values in the dataset X; X (i,j) represents the data in the i-th row and j-th column, and X′ (i,j) represents the value after normalizing X (i,j) .
[0072] Step 5: Use the variable feature extraction network to extract the local correlation features between variables in X'.
[0073] Step 5.1: Combine the variables in X' pairwise to obtain a three-dimensional matrix X cnn with a dimension of (a - 1)! × n × 2; where (a - 1)! represents the number of combinations,! represents factorial, and a represents the number of variables in X'; as Figure 4 shown, in a specific embodiment, 4 variables are selected, and there are 6 pairwise combination methods.
[0074] Step 5.2: Input the three-dimensional matrix X cnn into the variable feature extraction network, and use a three-dimensional convolutional layer with a convolutional kernel size of (a - 1)! × q × 2 and a stride of 1 to extract the features between each pair of a variables, obtaining a three-dimensional feature matrix Q1 with a dimension of (a - 1)! × (n - q + 1) × 1; where q represents the window length; in the gas concentration measurement experiment, the measured variables are correlated, so the first convolutional layer is used to extract the correlation features between the variables in the dataset X' respectively.
[0075] Step 5.3: After the variable feature extraction network performs dimension reconstruction on the three-dimensional feature matrix Q1, a two-dimensional feature matrix Q2 with a dimension of (n - q + 1) × (a - 1)! is obtained;
[0076] Step 5.4: The variable feature extraction network uses a two-dimensional convolutional layer with a convolutional kernel size of 1×(a - 1)! and p convolutional kernels and a stride of 1 to extract features from the two-dimensional feature matrix Q2, obtaining p one-dimensional feature matrices with dimensions of (n - q + 1)×1. Finally, after concatenating the p one-dimensional feature matrices, a two-dimensional feature matrix Q3 with dimensions of (n - q + 1)×p is obtained; the second convolutional layer uses multiple convolutional kernels to perform multiple convolutions on the two-dimensional feature matrix Q2 to expand the final feature matrix.
[0077] Step 6: Use a multi-scale LSTM neural network for prediction;
[0078] Step 6.1: Divide the two-dimensional feature matrix Q3 into a training set and a test set, and use the training set as the input of the LSTM neural network and C as the output of the LSTM neural network, so as to train the LSTM neural network to obtain a gas concentration prediction model at one scale;
[0079] The rows in the two-dimensional feature matrix Q3 represent the overall features of the variables in the window, and the columns represent the feature values of the variables in the window extracted using different convolutional kernels.
[0080] In the specific implementation, the network structure consists of two LSTM layers, two Dropout layers, and one fully connected layer. Since the gas concentration will not be negative, the Relu function is selected as the activation function of the LSTM layer to avoid the appearance of negative predicted concentration values. The output is the predicted gas concentration value, so the Linear function is selected as the activation function of the fully connected layer. To avoid the problem of overfitting in the neural network, the most commonly used method in deep learning to prevent neural network overfitting is used, that is, a Dropout layer is added after each LSTM layer, and the Dropout rate is set to 0.2. The adaptive optimizer Adam is selected, the initial learning rate is set to 0.001, the number of iterations is 100, the batch size is 50, and the loss function selects the mean squared error MSE, and the calculation formula is as shown in Equation (4):
[0081]
[0082] In Equation (4), y i represents the i-th predicted gas concentration value, and Y i represents the corresponding i-th true gas concentration value.
[0083] Step 6.2: After changing the window length q r times, return to Step 5.2 and execute sequentially, so as to obtain r gas concentration prediction models at different scales;
[0084] In the input layer of the LSTM neural network, the number of input data can be selected. For example, when the window size is 3, the first three groups of data are selected as the input of the network each time, and the output is the predicted value of the gas concentration at the next moment. In this network structure, after selecting data with different window sizes, the data in the window will first be subjected to feature extraction using a convolutional kernel with the same size as the window, and then the second convolutional layer is used to expand the data features. Finally, the input of the LSTM neural network each time will be 1×p feature data.
[0085] Step 6.3: Input the test set into the gas concentration prediction models at r scales respectively, and use Equation (5) to perform weighted summation on the gas concentrations predicted by the gas concentration prediction models at each scale, so as to obtain the comprehensive prediction value y at r scales;
[0086] y = w1y1 + w2y2 + … + w k y k + … + w r y r (5)
[0087] In Equation (5), y k represents the gas concentration predicted by the gas concentration prediction model at the kth scale, w k represents the kth weight, and there is:
[0088]
[0089] In Equation (6), MSE k represents the mean square error between the predicted gas concentration y k at the kth scale and the true concentration value C.
[0090] In a specific embodiment, the network model structure diagram is as Figure 5 shown. Select the LSTM networks at 3 scales, that is, three different window lengths are selected. The mean absolute error MAE, root mean square error RMSE, and Pearson correlation coefficient R are used to evaluate the prediction results of the model, and the calculation formulas are respectively:
[0091]
[0092]
[0093]
[0094] In Equations (7), (8), and (9), y i is the predicted value of the ith gas concentration, Y i is the true value of the ith gas concentration, n is the number of data in the test set, Cov is the covariance calculation, and Var is the variance calculation.
[0095] In this embodiment, an electronic device includes a memory and a processor. The memory is used to store a program that supports the processor to execute the above method, and the processor is configured to execute the program stored in the memory.
[0096] In this embodiment, a computer-readable storage medium stores a computer program, and when the computer program is run by a processor, it executes the steps of the above method.
Claims
1. A gas concentration prediction method based on an improved DBSCAN algorithm and a multi-scale LSTM neural network, characterized in that, Including the following steps: Step 1: Use a gas absorption spectroscopy experimental device to obtain n sets of gas absorption spectrum data in the atmosphere, and calculate the integrated area of the absorption peak of each group of gases, denoted as A = {A1, A2,..., A i ,...A n}, where A i represents the integrated area of the absorption peak calculated from the i-th set of gas absorption spectrum data; Obtain the temperature data set T = {T1, T2,..., T i ,..., T n} and the humidity data set R = {R1, R2,..., R i ,..., R n}, where T i and R i respectively represent the temperature and humidity corresponding to the i-th group of gas absorption spectrum data; Obtain the historical true concentration data of n groups of gases, denoted as C = {C1, C2,..., C i ,...C n}, where C i represents the historical true concentration corresponding to the absorption spectrum data of the i-th group of gases; n represents the number of data groups; Step 2: Use the improved DBSCAN clustering algorithm to screen out abnormal outlier data points; Step 2.1: Define the minimum number of samples in the neighborhood as P min , define and randomly initialize the radius ε of the neighborhood, and define and randomly initialize the threshold σ; Step 2.2: Initialize i = 1; Step 2.3: Take A i as the i-th sample point, and determine whether the number of sample points P i within the neighborhood with a radius of ε centered at A i is greater than or equal to P min . If so, it means that A i is a core point, and store A i into the core point set H, and form the i-th cluster Q i with all the sample points within the radius ε of the core point A i ; otherwise, directly execute Step 2.4; Step 2.4: After assigning i + 1 to i, determine whether i > n holds. If it holds, it means that several clusters and the final core point set H are obtained. Otherwise, return to Step 2.3 and execute sequentially; Step 2.5: Determine whether there are the same sample points in several clusters. If there are, merge the clusters corresponding to the same sample points into one cluster. Otherwise, do not merge; thus forming a set of clusters Q that are mutually disjoint; and form an outlier set I from the sample points in A that are not included in the set of clusters Q; Step 2.6: Calculate the distances between each outlier in the outlier set I and each core point in the core point set H to obtain a distance matrix where e represents the number of outliers, f represents the number of core points, and l ij represents the distance between the i-th outlier and the j-th core point; Step 2.7: Obtain the minimum value in each row vector of the distance matrix L, and obtain the nearest distance matrix L min = [l 1,min ,..., l i,min ,..., l e,min , where l i,min represents the distance between the i-th outlier and the nearest core point; Step 2.8, determine whether the difference S i,min between l i and the radius ε is less than the threshold σ. If so, retain the corresponding difference; otherwise, set the corresponding difference to -1, thereby obtaining the difference matrix S = [S1, S2,..., S i ,..., S e ; Step 2.
9. Determine whether there is a difference greater than 0 in S. If so, select the minimum value S among the differences greater than 0 in S min , and assign ε + S min to ε, and assign σ - S min to σ, then return to Step 2.2 and execute sequentially; otherwise, it means that the final outlier set I is obtained; Step 3: Correct the abnormal data in the measurement and construct a neural network data set; Step 3.1: According to L min e nearest core points are obtained, and the outlier corresponding to each nearest core point is replaced by the average value of all sample points within the neighborhood range of the nearest core point, so as to obtain the corrected outlier set I'. After merging with the cluster set Q, the corrected gas absorption spectrum data A' is obtained; Step 3.2: Construct a data set X = {C, A', T, R} of the LSTM neural network from C, A', T, and R, and perform normalization processing on the data set X to obtain the normalized X'; Step 4: Use a variable feature extraction network to extract the local correlation features between variables in X'; Step 4.1: Combine the variables in X' pairwise to obtain a three-dimensional matrix X with dimensions of (a - 1)! × n × 2 cnn ; where (a - 1)! represents the number of combinations,! represents factorial, and a represents the number of variables in X'. Step 4.2: Input the three-dimensional matrix X cnn into the variable feature extraction network, and use a three-dimensional convolutional layer with a convolutional kernel size of (a - 1)! × q × 2 and a stride of 1 to extract the features between every two of the a variables, obtaining a three-dimensional feature matrix Q1 with a dimension of (a - 1)! × (n - q + 1) × 1; where q represents the window length; Step 4.3: After the variable feature extraction network reconstructs the dimensions of the three-dimensional feature matrix Q1, a two-dimensional feature matrix Q2 with dimensions of (n - q + 1) × (a - 1)! is obtained; Step 4.4: The variable feature extraction network uses a two-dimensional convolutional layer with a convolutional kernel size of 1 × (a - 1)!, a convolutional kernel number of p, and a stride of 1 to extract features from the two-dimensional feature matrix Q2, obtaining p one-dimensional feature matrices with dimensions of (n - q + 1) × 1. Finally, after splicing the p one-dimensional feature matrices, a two-dimensional feature matrix Q3 with dimensions of (n - q + 1) × p is obtained; Step 5: Use a multi-scale LSTM neural network for prediction; Step 5.1: Divide the two-dimensional feature matrix Q3 into a training set and a test set, and use the training set as the input of the LSTM neural network and C as the output of the LSTM neural network to train the LSTM neural network to obtain a gas concentration prediction model at one scale; Step 5.2: After changing the window length q r times, return to Step 4.2 and execute sequentially to obtain r gas concentration prediction models at different scales; Step 5.3: Input the test set into the gas concentration prediction models at r scales respectively, and use Equation (1) to perform weighted summation on the gas concentrations predicted by the gas concentration prediction models at each scale to obtain a predicted comprehensive value y at r scales; y = w1y1 + w2y2 + … + w k y k + … + w r y r (1) In Equation (1), y k represents the gas concentration predicted by the gas concentration prediction model at the k-th scale, and w k represents the k-th weight, and there is: In formula (2), MSE k represents the mean square error between the predicted gas concentration y k at the k-th scale and the true concentration value C.
2. An electronic device, including a memory and a processor, characterized in that, The memory is used to store a program that supports the processor to execute the gas concentration prediction method described in Claim 1, and the processor is configured to execute the program stored in the memory.
3. A computer-readable storage medium, on which a computer program is stored, characterized in that, When the computer program is run by the processor, it executes the steps of the gas concentration prediction method described in Claim 1.
Citation Information
Patent Citations
Random vibration drive ring-down cavity calibration-free gas concentration measurement system and method
CN110672554A
Gas concentration calculation method, device and equipment and storage medium
CN113758890A