R-LSTM air quality missing data interpolation algorithm based on sample screening

By applying the correlation LSTM algorithm based on sample screening and Bi-LSTM model in the air quality monitoring data, the problem of missing values ​​in the air quality monitoring data is solved, and the accuracy and effect of interpolation are improved.

CN120067546APending Publication Date: 2025-05-30SHENYANG INST OF COMPUTING TECH CO LTD THE CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410978047.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-07-22
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The missing values ​​in the air quality monitoring data caused by transmission failure or hardware damage have affected subsequent pollutant analysis, trend research and policy formulation.

Method used

A correlation LSTM (R-LSTM) algorithm based on sample screening is proposed, and missing value interpolation of air quality time series data is combined with Bi-LSTM model.

Benefits of technology

The accuracy of missing value interpolation of air quality time series data is improved, and the interpolation effect is improved by utilizing the intrinsic correlation in air quality data and the two-way long and short-term memory network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067546A_ABST
    Figure CN120067546A_ABST
Patent Text Reader

Abstract

The invention relates to an R-LSTM air quality missing data interpolation algorithm based on sample screening. According to the method, firstly, relevance calculation is carried out on input air quality time sequence data learning samples, after several most relevant features are obtained, learning sample screening is carried out according to the missing condition of the features, and data input into a model is made to be complete data. And finally, carrying out interpolation on the missing data segment through the BI-LSTM network and outputting complete time sequence data. According to the method, related experiments prove that the method has adaptability and better accuracy in the aspect of air quality time sequence missing value processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of environmental monitoring and data imputation using deep learning, and specifically to an algorithm for improving the accuracy of imputing missing values in air quality time series data. Background Art

[0002] In the context of the increasing attention to the natural environment in current society, the issue of urban air quality has gradually become the focus of public concern. Continuously monitoring urban air quality, regularly publishing relevant data to the public, and archiving records have become a normal task in environmental management. Thanks to the rapid progress of Internet of Things sensing technology and intelligent terminal monitoring technology, the monitoring of environmental air pollutants has reached a new level of intelligence and scale. However, in the actual data collection process, monitoring equipment may encounter technical problems such as transmission failures and hardware damages, which may lead to a large number of missing air quality monitoring data. The incompleteness of data may have an adverse impact on subsequent air quality pollutant analysis, trend research, and even policy making. Therefore, ensuring the integrity of air quality data is of crucial significance for achieving effective environmental management and scientific decision-making.

[0003] Facing the challenge of data missing in air quality monitoring, the academic and industrial communities have actively promoted a series of studies aimed at developing appropriate data imputation techniques. Currently, data imputation methods are mainly divided into two categories: traditional statistical methods and deep learning-based methods. Statistical methods are usually applicable to scenarios with a relatively low data missing rate and relatively loose requirements for task accuracy. In contrast, in many practical applications, deep learning methods, especially recurrent neural networks (RNNs) and attention mechanism-based models, have shown advantages in imputing time series data and can provide more accurate results. The reason why deep learning methods are effective is that they can deeply explore and learn the internal patterns and relationships of time series data, and have more advantages in dealing with complex data scenarios with a high missing rate compared to statistical methods that only rely on the surface features of data. In particular, long short-term memory networks (LSTMs), as a derivative model of RNNs, have been widely used in the prediction and imputation tasks of air quality data due to their unique ability to process sequence data. At the same time, attention mechanism-based models can capture and utilize long-term context information in time series, providing an efficient solution for accurately imputing missing data. Summary of the Invention

[0004] Although there are already various methods for current missing data imputation, these methods generally belong to the general formatting data imputation methods and do not sufficiently incorporate the knowledge of domain experts in the field of air quality. They fail to fully utilize the professional internal correlations in the monitored variables. In the invention, we propose the Relevance LSTM (R-LSTM) to combine the internal correlations in air quality data and introduce a data screening mechanism into the model, which improves the accuracy of imputing missing sample data.

[0005] The present invention adopts the following technical solutions: An R-LSTM air quality missing data imputation algorithm based on sample screening, comprising the following steps:

[0006] Obtain the air quality time series data containing missing values, and use a masking variable to mark the data to obtain the masked time series data;

[0007] For the air quality time series data containing missing values, perform a correlation analysis on the data characteristics between samples, and select several features with the highest correlation as the correlation features;

[0008] Screen the masked time series data through the correlation features to obtain the screened data:

[0009] Input the screened data into the Bi-LSTM model to obtain the combined data of the forward and backward directions of the model, and update the positions of the data gaps in the original air quality time series data containing missing values to complete the imputation of the missing values in the time series.

[0010] The masking variable is as follows:

[0011]

[0012] Wherein, is the data x of the feature dimension d at the time step t.

[0013] For the air quality time series data, performing a correlation analysis on the data characteristics between samples and selecting several features with the highest correlation as the correlation features includes the following steps:

[0014] For the air quality time series data containing missing values, the features are labeled as d 1 ,d 2 ……d n ,where n is the dimension of the data, and the Pearson coefficient is used to calculate the correlation between features;

[0015] Take the absolute value of each Pearson coefficient, and sort according to the absolute value, and select several features with the highest correlation as the correlation features.

[0016] The screening of the masked time series data by the correlation features includes the following steps:

[0017] For the masked time series data, if the masked variables of several features with the highest correlation in the sample data at a certain time step t are 0, it is considered that this feature is missing, and this sample is skipped.

[0018] The input of the screened data into the Bi-LSTM model to obtain the combined data of the forward and backward directions of the model is as follows:

[0019] Through the calculation of the forward and backward hidden states in the Bi-LSTM model, the imputation sequence generated forward is where represents the i-th result of the forward imputation output, and the sequence generated backward is where represents the i-th result of the backward imputation output;

[0020] Fill the mean value of the forward and backward sequences into the data gaps in the original air quality time series data with missing values, that is, the sample positions where the masked variable is 0, to complete the process of imputing the missing values in the time series.

[0021] An R-LSTM air quality missing data imputation system based on sample screening includes:

[0022] A data marking module, used to obtain the air quality time series data with missing values, mark the data using masked variables to obtain the masked time series data;

[0023] A correlation analysis module, used to perform a correlation analysis on the data features between samples for the air quality time series data with missing values, and select several features with the highest correlation as the correlation features;

[0024] A screening module, used to screen the masked time series data through the correlation features to obtain the screened data:

[0025] A missing data imputation module, used to input the screened data into the Bi-LSTM model to obtain the combined data of the forward and backward directions of the model, update the positions of the data gaps in the original air quality time series data with missing values, and complete the imputation of the missing values in the time series.

[0026] An R-LSTM air quality missing data imputation device includes a memory and a processor; the memory is used to store a computer program; the processor is used to implement the R-LSTM air quality missing data imputation method based on sample screening when executing the computer program.

[0027] A computer-readable storage medium has a computer program stored thereon. When the computer program is executed by a processor, the described R-LSTM air quality missing data interpolation method based on sample screening is implemented.

[0028] The present invention has the following beneficial effects and advantages:

[0029] 1. The present invention combines the inherent characteristics of pollutants in air quality data for preprocessing, including data cleaning and missing value identification.

[0030] 2. The present invention conducts sample correlation analysis and uses the correlation between air quality indicators for sample screening, making the data input into the model more valuable for training and facilitating the improvement of the accuracy of the output data.

[0031] 3. The present invention proposes and constructs an R-LSTM model, uses a bidirectional long short-term memory network for data interpolation, and effectively combines the internal correlations in air quality monitoring, thus effectively improving the interpolation effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 is a flowchart of the method of the present invention.

[0033] Figure 2 is a schematic diagram of data feature correlation in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0034] The following further elaborates the present invention in detail in conjunction with the drawings and embodiments.

[0035] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following provides a detailed description of the specific embodiments of the present invention in conjunction with the drawings. Many specific details are set forth in the following description to facilitate a thorough understanding of the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the spirit of the invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0036] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this invention belongs. The terms used in the description of the invention herein are for the purpose of describing specific embodiments only and are not intended to limit the present invention.

[0037] Figure 1 For the overall steps of the invention, the R-LSTM air quality missing data interpolation algorithm based on sample screening of the present invention specifically includes the following steps:

[0038] Step 1: The user inputs time series data with missing values, where the input data format is a matrix of d feature dimensions * t time steps;

[0039] As Figure 2 shown, the time series data has a dimension of d 1, d 2 …d 5 There are a total of 5 feature dimensions, and d 1 refers to all the data in the first row (the first feature dimension), indicating the data at the position of dimension d i at time step t i in dimension d, where d and t are vectors, i.e., t = {t1, t2…tn}, d = {d 1 , d 2 ……d n}. Among them, t represents the collection time, and d represents the collected pollutants, including sulfur dioxide, ozone, etc.

[0040] Step 2: Discriminate the time series data input in Step 1 to identify which data are samples with missing values and which data are complete samples, and use a mask variable to mark the data. The definition of the mask variable is as follows. The meaning is the data x of feature dimension d at time step t:

[0041]

[0042] Step 3: Since there is often a correlation between pollutants, then conduct a correlation analysis on the data characteristics between samples. First, label the characteristics of the time series data input in Step 1 as d 1 , d 2 ……d n where n is the dimension of the data, and then use the Pearson coefficient to calculate the correlation between the characteristics.

[0043] Step 4: Take the absolute value of each Pearson coefficient calculated in Step 3 and sort them according to the absolute value, and select the 3 features with the highest correlation as the correlation features;

[0044] Step 5: Input the time series data in Step 1 into the R-LSTM model after adding the mask variable in Step 2. When the data is input into the model, filter according to the correlation values calculated in Steps 3 and 4. If there is data missing among the 3 features with the highest correlation in the sample data at a certain time step t (the basis for judging the missing is obtained from the mask variable in Step 2. If the mask variable of this feature is 0, it is considered missing), then skip it, so as to filter out the data with relatively serious missing of correlation features;

[0045] AsFigure 2 As shown, in the data queue, the features are in a certain row and the samples are in a certain column. For example, the sample at time t1 is the data in the first column. For example Figure 2 there is missing data at time step t4. This algorithm is used to impute this missing data, that is, to impute the red data in dimension d3 at time t4. First, an observation mask is applied to the samples with missing data to confirm that the feature data in d3 is missing. Calculate the correlation between d3 and other features and sort them, and select the top 3 features with the highest correlation, namely d2, d3, and d4. Since the samples need to be put into the model to let the model give the missing values for filling, it can be seen that there is missing data in the features of samples d2, d3, and d4 at time steps t1 and t6. At this time, this sample is directly skipped and the next sample is learned.

[0046] Step 6: The data screened in Step 5 enters the R-LSTM model. The internal calculation formula of the R-LSTM model is as follows. In the formula, Wxi, Wxf, Wxc, Wxo, Whi, Whf, Whc, and Who represent weight matrices, and bi, bf, bc, and bo represent bias terms, while xt represents the input data at time step t.

[0047] i t = σ(W xi x t + W hi h t-1 + b i )

[0048] f t = σ(W xf x t + W hf h t-1 + b f )

[0049]

[0050] o t = σ(W xo x t + W ho h t-1 + b o )

[0051]

[0052] where σ represents the sigmoid calculation.

[0053] Step 7: In the Bi-LSTM model, the transfer and update of the hidden state are carried out in the forward and backward directions respectively. The sets of these two hidden states are defined as follows:

[0054]

[0055] Among them, respectively represent the state of the nth hidden layer.

[0056] Step 8: The forward hidden state is updated through the following recurrence relation, where h t-1 is the forward hidden state at the previous time step:

[0057]

[0058] Among them, LSTM cell represents an LSTM cell.

[0059] Step 9: Similarly, the backward hidden state is updated through the following recurrence relation, where h t+1 is the backward hidden state at the next time step:

[0060]

[0061] Step 10: Finally, calculate the output result of the fully connected layer according to the following calculation formula. The calculation formula is as follows, and

[0062]

[0063] Among them, h i represents the output of the i-th hidden layer, T represents the state of the i-th hidden layer.

[0064] Step 11: By combining the results calculated from the forward and backward hidden states, Bi-LSTM can make more full use of the bidirectional time series information, which is more beneficial for imputing missing data. Since the imputation is bidirectional, it is assumed that the imputation sequence generated forward is Among them represents the i-th result of the forward imputation output, and the sequence generated backward is: Among them represents the i-th result of the backward imputation output. The final value is the mean of the forward and backward sequences, and the calculated mean is filled into the data gaps in the original time series data to complete the process of imputing missing values in the time series.

Claims

1. An R-LSTM air quality missing data interpolation algorithm based on sample screening, characterized in that: The following steps are involved: Get the air quality time series data with missing values, mark the data with mask variables, and get masked time series data; For air quality time series data with missing values, correlation analysis is performed on the data features between samples, and several features with the highest correlation are selected as correlation features; The masked time series data is filtered by correlation features to obtain the filtered data: The filtered data is input into the Bi-LSTM model to obtain the combined forward and backward data of the model, update the data gaps in the original air quality time series data containing missing values, and complete the missing value interpolation of the time series.

2. The R-LSTM air quality missing data interpolation algorithm based on sample screening according to claim 1 is characterized in that: The mask variables are as follows: in, is the data x of feature dimension d at time step t.

3. The R-LSTM air quality missing data interpolation algorithm based on sample screening according to claim 1 is characterized in that: For the air quality time series data, correlation analysis is performed on the data features between samples, and several features with the highest correlation are selected as correlation features, including the following steps: For air quality time series data with missing values, the feature annotation is d 1, d 2…… d n , where n is the dimension of the data, and the Pearson coefficient is used to calculate the correlation between features; Take the absolute value of each Pearson coefficient, sort them according to the absolute value, and select several features with the highest correlation as correlation features.

4. The R-LSTM air quality missing data interpolation algorithm based on sample screening according to claim 1 is characterized in that: The step of screening the masked time series data by correlation features comprises the following steps: For masked time series data, if the mask variables of the most correlated features in the sample data at a certain time step t are 0, then this feature is considered missing and the sample is skipped.

5. The R-LSTM air quality missing data interpolation algorithm based on sample screening according to claim 1 is characterized in that: The filtered data is input into the Bi-LSTM model to obtain the combined forward and backward data of the model, as follows: Through the forward and reverse hidden state calculations in the Bi-LSTM model, the interpolation sequence generated forward is obtained as in, It is represented as the i-th result of the forward interpolation output, and the sequence generated backward is in, Represented as the i-th result of the reverse interpolation output; The mean of the forward and backward sequences is used to fill in the data gaps in the original air quality time series data containing missing values, that is, the sample positions where the mask variable is 0, to complete the process of missing value interpolation of the time series.

6. An R-LSTM air quality missing data interpolation system based on sample screening, characterized in that: include: The data labeling module is used to obtain the air quality time series data with missing values, label the data using mask variables, and obtain masked time series data; The correlation analysis module is used to perform correlation analysis on the data features between samples for the air quality time series data with missing values, and select several features with the highest correlation as correlation features; The filtering module is used to filter the masked time series data by correlation features to obtain the filtered data: The missing data interpolation module is used to input the filtered data into the Bi-LSTM model, obtain the forward and backward combined data of the model, update the data gaps in the original air quality time series data containing missing values, and complete the missing value interpolation of the time series.

7. An R-LSTM air quality missing data interpolation device based on sample screening, characterized in that: It comprises a memory and a processor; the memory is used to store a computer program; the processor is used to implement an R-LSTM air quality missing data interpolation method based on sample screening as described in any one of claims 1 to 5 when executing the computer program.

8. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, an R-LSTM air quality missing data interpolation method based on sample screening as described in any one of claims 1 to 5 is implemented.