Improved auto-encoder and data cleaning method thereof

By introducing a multi-layer structure of convolutional layer, bidirectional long and short-term memory layer, self-attention layer and deconvolution layer into the autoencoder, the problem of identifying and processing abnormal points and missing values ​​when processing high-dimensional data is solved, and a more efficient and accurate data cleaning effect is achieved.

CN120012857APending Publication Date: 2025-05-16CHINA ORDNANCE IND EXPLOSIVES ENG & SAFETY TECH RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510003383.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

When processing high-dimensional data, it is difficult to effectively identify and process abnormal points and missing values, especially in the problems of non-Gaussian data distribution, high computational complexity and poor processing of high-dimensional data.

Method used

A data cleaning method based on an improved autoencoder is adopted, and the multi-layer structure of the data is captured by combining the convolutional layer, the bidirectional long and short-term memory layer, the self-attention layer and the deconvolution layer to capture the spatiotemporal characteristics and complex relationships, so as to eliminate the abnormal points and fill the missing values.

Benefits of technology

It significantly improves the accuracy and efficiency of abnormal point detection and missing value filling, can effectively extract the spatial and timing characteristics of the data, dynamically allocate data points, ensure the accuracy of data filling and elimination, and is suitable for large-scale and high-dimensional data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012857A_ABST
    Figure CN120012857A_ABST
Patent Text Reader

Abstract

The invention discloses a data cleaning method based on an improved auto-encoder, and the method comprises the following steps: firstly, carrying out the local feature extraction of a convolution operation, the processing of a long-term time sequence of a long-short-term memory network, and the effective feature extraction of a self-attention mechanism; constructing an improved auto-encoder by combining convolution operation, a bidirectional long-short-term memory network and a self-attention mechanism; then, normal and non-abnormal data is used for training the improved auto-encoder; then, inputting data needing to be judged into the trained improved auto-encoder to obtain an output result, calculating a residual error between the data needing to be judged and output data, and judging and eliminating abnormal points through a preset threshold value; and finally, further inputting the data of which the abnormal points are eliminated into the improved auto-encoder to complete filling of missing values. According to the method, the elimination of abnormal points in the data and the filling of missing values can be effectively realized, and the data cleaning is accurately and effectively realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data processing, and in particular, relates to a method based on an improved autoencoder and a data cleaning method thereof. Background Art

[0002] With the rapid development of big data technology, the identification and processing of outliers in data and the filling of missing values ​​have become hot research directions in the field of data analysis and machine learning. Although traditional methods such as statistical outlier detection (such as standard deviation method, box plot method, etc.) are simple and easy to use, they are too strict in assuming data distribution and are difficult to process non-Gaussian distributed data. Clustering methods (such as K-Means, DBSCAN, etc.) and distance-based methods (such as isolation forest, etc.) can identify outliers to a certain extent, but they have high computational complexity and poor processing effect on high-dimensional data. Similarly, traditional missing value filling methods (such as mean filling, median filling, linear interpolation, spline interpolation, K-nearest neighbor filling, regression filling, etc.) work well when processing simple data sets, but when facing high-dimensional and complex data sets, the effect is often unsatisfactory. For example, mean filling will introduce a lot of noise, affecting the accuracy of subsequent analysis; although K-nearest neighbor filling and regression filling can consider the correlation between data, they have high computational complexity and poor processing effect on high-dimensional data.

[0003] In recent years, outlier detection and missing value filling methods based on deep learning have gradually become a research hotspot. As an unsupervised learning model, autoencoders reconstruct input data by training a neural network, thereby achieving dimensionality reduction and feature extraction of data. Common autoencoder structures include fully connected autoencoders, convolutional autoencoders, and long short-term memory (LSTM) autoencoders. These models have high efficiency and accuracy when processing high-dimensional data and can capture complex relationships between data. However, existing autoencoder methods still have some shortcomings in dealing with outliers and missing values, such as insufficient consideration of the spatiotemporal dependence of data and limited generalization ability of the model.

[0004] In order to address the shortcomings of the existing methods, the present invention proposes a method for outlier removal and missing value filling based on an autoencoder. The method can more comprehensively capture the spatiotemporal characteristics and complex relationships of the data by combining a multi-layer structure such as a convolutional layer, a bidirectional long short-term memory layer, a self-attention layer, a bidirectional long short-term memory layer, and a deconvolution layer. The convolutional layer can effectively extract the spatial characteristics of the data and can capture the local structure and pattern in the data. The bidirectional long short-term memory layer can simultaneously consider the front-end and back-end dependencies of the data, thereby more accurately modeling the temporal characteristics of the data. The self-attention layer can dynamically assign the importance of different data points, effectively capture the long-range dependencies between data, and improve the generalization ability of the model. The deconvolution layer is used to restore the extracted features to the form of original data to achieve the removal of outliers and the filling of missing values.

[0005] The method of the present invention has higher efficiency and accuracy when processing high-dimensional data, and is suitable for large-scale, high-dimensional data sets. In addition, through the introduction of the self-attention mechanism, the present invention can more effectively capture the long-range dependencies between data and improve the generalization ability of the model. The method of the present invention can be widely used in various data analysis and machine learning tasks, including text data, time series data, and sensor data. In text data, it can be used to identify and remove abnormal words or sentences, and fill in missing words or sentences, thereby improving the readability and accuracy of the text; in time series data, it can be used to identify and remove abnormal time series points, and fill in missing time series points, thereby improving the continuity and accuracy of time series data; in sensor data, it can be used to identify and remove abnormal sensor readings, and fill in missing sensor readings, thereby improving the reliability and accuracy of sensor data.

[0006] Compared with the existing outlier detection and missing value filling methods based on statistics and traditional machine learning, the method of the present invention can detect outliers and fill missing values ​​more accurately, and has higher efficiency and stronger generalization ability. Therefore, the method proposed by the present invention is not only innovative in technology, but also has important value and prospects in practical applications. Summary of the invention

[0007] The technical problem to be solved by the present invention is to provide a data cleaning method based on an improved autoencoder, which can effectively eliminate outliers and fill missing values ​​in the data while ensuring the original structure and morphology of the data, thereby accurately and effectively cleaning the data.

[0008] To solve the above problems, the present invention provides a data cleaning method based on an improved autoencoder, which adopts the following technical solution.

[0009] A data cleaning method based on improved autoencoding, characterized in that it comprises the following steps:

[0010] Step 1), constructing the network structure of the improved autoencoder;

[0011] Step 2), using the data without outliers as input and output respectively to train the improved autoencoder to obtain a trained improved autoencoder;

[0012] Step 3), the improved autoencoder is used to analyze the data, remove outliers and fill in new data.

[0013] Preferably, in the above-mentioned data cleaning method, in step 1), the network structure of the improved autoencoder consists of a one-dimensional convolutional layer, a bidirectional long short-term memory layer, a self-attention layer, and a one-dimensional deconvolutional layer.

[0014] Preferably, in the above-mentioned data cleaning method, the improved autoencoder extracts features from the data through a one-dimensional convolutional layer, a bidirectional long short-term memory layer and a self-attention layer; and then reconstructs the features through a bidirectional long short-term memory layer and a one-dimensional deconvolution layer.

[0015] Preferably, in the above data cleaning method, in step 3, the step of removing outliers is as follows:

[0016] For the data of abnormal points, the abnormal point data is input into the improved autoencoder to obtain the output data, the residual between the input data and the output data is calculated, and the abnormal values ​​in the data are located and eliminated according to the preset threshold to obtain the data after the abnormal values ​​are eliminated. The data after the abnormal values ​​are eliminated is input into the trained improved autoencoder to obtain the final completed filled data.

[0017] Preferably, in the above data cleaning method, the residual is calculated according to the following formula:

[0018] For the outlier data X, input X into the trained improved autoencoder to get the corresponding output Compute the residual between the output and the input:

[0019]

[0020] In the formula, X is abnormal data;

[0021] To improve the data obtained after the autoencoder processes abnormal data;

[0022] ε is the residual.

[0023] On the other hand, the present invention also provides an improved autoencoder for executing the above-mentioned data cleaning method, which adopts the following technical solution.

[0024] Based on an improved autoencoder, a first layer, a second layer, a third layer, a fourth layer, a fifth layer, a sixth layer, a seventh layer, an eighth layer, a ninth layer, a tenth layer and an eleventh layer are arranged in sequence from the input layer to the output layer;

[0025] The first, second, and third layers are convolutional layers;

[0026] The fourth, fifth, seventh and eighth layers are bidirectional long short-term memory layers;

[0027] The sixth layer is the self-attention layer;

[0028] The ninth, tenth and eleventh layers are deconvolution layers.

[0029] Preferably, in the above-mentioned improved autoencoder, the convolution kernel size of the first, second and third convolution layers is 25, and the number of channels of the first convolution layer is 16, the number of channels of the second convolution layer is 32, and the number of channels of the third convolution layer is 64.

[0030] Preferably, in the above-mentioned improved autoencoder, the number of hidden unit nodes of the fourth bidirectional long short-term memory layer is 70; the number of hidden unit nodes of the fifth bidirectional long short-term memory layer is 30; the number of hidden unit nodes of the seventh bidirectional long short-term memory layer is 30, and is connected to a fully connected layer with 70 output nodes; the number of hidden unit nodes of the eighth bidirectional long short-term memory layer is 70.

[0031] Preferably, in the above-mentioned improved autoencoder, the number of attention heads of the sixth self-attention layer is 2, the number of channels is 4, and it is connected to a fully connected layer with 30 output nodes.

[0032] Preferably, in the above-mentioned improved autoencoder, the convolution kernel size of the ninth, tenth and eleventh deconvolution layers is 25, and the number of channels of the ninth deconvolution layer is 32, the number of channels of the tenth deconvolution layer is 16, and the number of channels of the eleventh deconvolution layer is 1.

[0033] Compared with the prior art, the present invention has the following beneficial effects:

[0034] The present invention proposes a data cleaning method based on an improved autoencoder, which significantly improves the accuracy and efficiency of outlier detection and missing value filling by combining a convolutional layer, a bidirectional long short-term memory layer, a self-attention layer, and a deconvolution layer. The method can effectively extract the spatial and temporal characteristics of the data, dynamically allocate the importance of data points, and ensure the accuracy of data filling and elimination. In addition, the present invention also provides an improved autoencoder for executing the above-mentioned data cleaning method, and the accuracy of data cleaning is guaranteed by executing the above-mentioned data cleaning method by the improved autoencoder. At the same time, the present invention has low computational complexity when processing high-dimensional data, is suitable for real-time processing of large-scale data sets, and has strong adaptability and flexibility, without complex preprocessing and parameter adjustment. The present invention performs well in multiple scenarios such as text, time series, and sensor data, significantly improves the quality and integrity of the data, and provides reliable support for subsequent data analysis and machine learning tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 It is a flowchart of data cleaning in the present invention;

[0036] Figure 2 This is a structural diagram of the improved autoencoder in the present invention;

[0037] Figure 3 This is the result diagram of Example 1;

[0038] in: Figure 3 (a) is the original data graph in Example 1;

[0039] Figure 3 (b) is the data graph after constructing the abnormal point in Example 1;

[0040] Figure 3 (c) is a diagram showing the residual calculation results in Example 1;

[0041] Figure 3 (d) is the result diagram after the abnormal points are removed in Example 1;

[0042] Figure 3 (e) is the result diagram after missing value filling in Example 1.

[0043] Figure 4 It is the result diagram in Example 2;

[0044] in: Figure 4 (a) is the original data graph in Example 2;

[0045] Figure 4 (b) is the data graph after constructing the abnormal point in Example 2;

[0046] Figure 4(c) is a diagram showing the residual calculation results in Example 2;

[0047] Figure 4 (d) is the result diagram after the abnormal points are removed in Example 2;

[0048] Figure 4 (e) is the result after missing value filling in Example 2 DETAILED DESCRIPTION

[0049] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings.

[0050] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0051] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not refer to the same embodiment, nor is it a separate or selective embodiment that is mutually exclusive with other embodiments. The present invention provides the following embodiments.

[0052] This embodiment provides a data cleaning method based on an improved autoencoder, which specifically includes the following steps.

[0053] First, construct an improved autoencoder network structure, such as Figure 2As shown in the figure, the input layer is a one-dimensional time series; the first layer is a convolution layer, whose convolution kernel size is 25, the number of channels is 16, and the Relu activation function is used; the second layer is a convolution layer, whose convolution kernel size is 25, the number of channels is 32, and the Relu activation function is used; the third layer is a convolution layer, whose convolution kernel size is 25, the number of channels is 64, and the Relu activation function is used; the fourth layer is a bidirectional long short-term memory layer, whose hidden unit nodes are 70; the fifth layer is a bidirectional long short-term memory layer, whose hidden unit nodes are 30; the sixth layer is a self-attention layer, whose number of attention heads is 2, and the number of channels is The number of nodes is 4, and it is connected to a fully connected layer with 30 output nodes; the seventh layer is a bidirectional long short-term memory layer, whose hidden unit nodes are 30, and it is connected to a fully connected layer with 70 output nodes; the eighth layer is a bidirectional long short-term memory layer, whose hidden unit nodes are 70; the ninth layer is a deconvolution layer, whose convolution kernel size is 25, the number of channels is 32, and the Relu activation function is used; the tenth layer is a deconvolution layer, whose convolution kernel size is 25, the number of channels is 16, and the Relu activation function is used; the eleventh layer is a deconvolution layer, whose convolution kernel size is 25, the number of channels is 1, and the regression layer is used as the output.

[0054] Secondly, the improved autoencoder is trained using the data without outliers as input and output respectively to obtain a trained improved autoencoder.

[0055] Finally, the improved autoencoder is used to analyze the data. For the data with abnormal points, the abnormal point data is input into the improved autoencoder to obtain the output data, and the residual between the input data and the output data is calculated. The abnormal values ​​in the data are located and eliminated according to the preset threshold value to obtain the data after eliminating the abnormal values. The data after eliminating the abnormal values ​​is input into the trained improved autoencoder to obtain the final completed filled data, thereby completing the data cleaning operation.

[0056] In this embodiment, the above-mentioned data cleaning method is used to significantly improve the accuracy of data anomaly detection through a convolutional layer, a bidirectional long-short term and a layer, a self-attention layer and a deconvolution layer, and the filled data is accurate, which provides an effective guarantee for subsequent data analysis. On the one hand, this embodiment provides the above-mentioned data cleaning method, and on the other hand, it also provides an improved autoencoder for running the above-mentioned data cleaning method. The encoder is respectively provided with a first layer, a second layer, a third layer, a fourth layer, a fifth layer, a sixth layer, a seventh layer, an eighth layer, a ninth layer, a tenth layer and an eleventh layer from the input layer to the output layer.

[0057] The first layer is a convolutional layer with a convolution kernel size of 25 and 16 channels.

[0058] The second layer is a convolutional layer with a convolution kernel size of 25 and a channel number of 32.

[0059] The third layer is the convolution layer, whose convolution kernel size is 25 and the number of channels is 64.

[0060] The fourth layer is a bidirectional long short-term memory layer with 70 hidden unit nodes.

[0061] The fifth layer is a bidirectional long short-term memory layer with 30 hidden unit nodes.

[0062] The sixth layer is the self-attention layer, which has 2 attention heads and 4 channels, and is connected to a fully connected layer with 30 output nodes.

[0063] The seventh layer is a bidirectional long short-term memory layer with 30 hidden unit nodes and a fully connected layer with 70 output nodes.

[0064] The eighth layer is a bidirectional long short-term memory layer with 70 hidden unit nodes.

[0065] The ninth layer is the deconvolution layer, whose convolution kernel size is 25 and the number of channels is 32.

[0066] The tenth layer is a deconvolution layer with a convolution kernel size of 25 and 16 channels.

[0067] The eleventh layer is the deconvolution layer, whose convolution kernel size is 25 and the number of channels is 1.

[0068] Through the above-mentioned improvements, the autoencoder can effectively remove abnormal data and fill in the data. The removed data and filled data are highly accurate, ensuring the accuracy of the subsequent data analysis process.

[0069] In order to verify the effectiveness of the proposed data cleaning method based on the improved autoencoder, the present invention verifies the effectiveness of the improved autoencoder through Example 1 and Example 2 respectively. Example 1 and Example 2 respectively use 900 sets of data to verify the effectiveness of the improved autoencoder, and the results are as follows.

[0070] Example 1

[0071] like Figure 3 As shown in the figure, anomalies are artificially constructed in the non-abnormal data outside the 900 sets of training data. The original data of the 900 sets of training data are as follows Figure 3 As shown in (a), the result of constructing the outlier point is as follows Figure 3 As shown in (b), the proposed method is used to remove outliers and fill in missing values.

[0072] For the outlier data X, input X into the trained improved autoencoder to get the corresponding output Calculate the residual between the output and the input, and the result is as follows Figure 3 (c) as shown:

[0073]

[0074] Where:

[0075] X is abnormal data;

[0076] To improve the data obtained after the autoencoder processes abnormal data;

[0077] ε is the residual.

[0078] like Figure 3 As shown in the figure, the outliers in the data are located and removed according to the preset threshold value, and the data after the outliers are removed is obtained. The results are as follows Figure 3 As shown in (d), the data after removing outliers is further Input into the trained improved autoencoder to get the final completed filled data, such as Figure 3 (e) By comparison Figure 3 It can be seen from the original data and the data after missing value filling that after processing based on the improved autoencoder data cleaning method, not only can the outliers in the data be effectively removed, but the results after data filling are basically consistent with the original data, which can effectively restore the structure and morphology of the original data.

[0079] Example 2

[0080] like Figure 4 As shown in the figure, anomalies are artificially constructed in the non-abnormal data outside the 900 sets of training data. The original data of the 900 sets of training data are as follows Figure 4 As shown in (a), the result of constructing the outlier point is as follows Figure 4 As shown in (b), the proposed method is used to remove outliers and fill in missing values.

[0081] For the outlier data X, input X into the trained improved autoencoder to get the corresponding output Calculate the residual between the output and the input, and the result is as follows Figure 4 (c) as shown:

[0082]

[0083] Where:

[0084] X is abnormal data;

[0085] To improve the data obtained after the autoencoder processes abnormal data;

[0086] ε is the residual.

[0087] like Figure 4As shown in the figure, the outliers in the data are located and removed according to the preset threshold value, and the data after the outliers are removed is obtained. The results are as follows Figure 4 As shown in (d), the data after removing outliers is further Input into the trained improved autoencoder to get the final completed filled data, such as Figure 4 (e) By comparison Figure 4 It can be seen from the original data and the data after missing value filling that after processing based on the improved autoencoder data cleaning method, not only can the outliers in the data be effectively removed, but the results after data filling are basically consistent with the original data, which can effectively restore the structure and morphology of the original data.

[0088] The above content is a further detailed description of the present invention in combination with specific implementation methods. It cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as belonging to the scope of protection determined by the claims submitted for the present invention.

Claims

1. A data cleaning method based on improved autoencoding, characterized in that: The steps include: Step 1), constructing the network structure of the improved autoencoder; Step 2), using the data without outliers as input and output respectively to train the improved autoencoder to obtain a trained improved autoencoder; Step 3), the improved autoencoder is used to analyze the data, remove outliers and fill in new data.

2. The data cleaning method based on improved autoencoding according to claim 1, characterized in that: In step 1), the network structure of the improved autoencoder consists of a one-dimensional convolutional layer, a bidirectional long short-term memory layer, a self-attention layer, and a one-dimensional deconvolutional layer.

3. The data cleaning method based on improved autoencoding according to claim 2 is characterized in that: The improved autoencoder extracts features from the data through a one-dimensional convolution layer, a bidirectional long short-term memory layer, and a self-attention layer; and then reconstructs the features through a bidirectional long short-term memory layer and a one-dimensional deconvolution layer.

4. The data cleaning method based on improved autoencoding according to claim 1, characterized in that: In step 3, the steps for removing outliers are as follows: For the data of abnormal points, the abnormal point data is input into the improved autoencoder to obtain the output data, the residual between the input data and the output data is calculated, and the abnormal values ​​in the data are located and eliminated according to the preset threshold to obtain the data after the abnormal values ​​are eliminated. The data after the abnormal values ​​are eliminated is input into the trained improved autoencoder to obtain the final completed filled data.

5. The data cleaning method based on improved autoencoding according to claim 4 is characterized in that: The residual is calculated according to the following formula: For the outlier data X, input X into the trained improved autoencoder to get the corresponding output Compute the residual between the output and the input: In the formula, X is abnormal data; To improve the data obtained after the autoencoder processes abnormal data; ε is the residual.

6. A method based on an improved autoencoder, characterized in that: From the input layer to the output layer, there are respectively arranged a first layer, a second layer, a third layer, a fourth layer, a fifth layer, a sixth layer, a seventh layer, an eighth layer, a ninth layer, a tenth layer and an eleventh layer; The first, second, and third layers are convolutional layers; The fourth, fifth, seventh and eighth layers are bidirectional long short-term memory layers; The sixth layer is the self-attention layer; The ninth, tenth and eleventh layers are deconvolution layers.

7. The improved autoencoder according to claim 6, characterized in that: The convolution kernel size of the first, second and third convolutional layers is 25, and the number of channels of the first convolutional layer is 16, the number of channels of the second convolutional layer is 32, and the number of channels of the third convolutional layer is 64.

8. The improved autoencoder according to claim 6, characterized in that: The fourth bidirectional long short-term memory layer has 70 hidden unit nodes; the fifth bidirectional long short-term memory layer has 30 hidden unit nodes; the seventh bidirectional long short-term memory layer has 30 hidden unit nodes and is connected to a fully connected layer with 70 output nodes; the eighth bidirectional long short-term memory layer has 70 hidden unit nodes.

9. The improved autoencoder according to claim 6, characterized in that: The sixth self-attention layer has 2 attention heads and 4 channels, and is connected to a fully connected layer with 30 output nodes.

10. The improved autoencoder according to claim 6, characterized in that: The convolution kernel size of the ninth, tenth and eleventh deconvolution layers is 25, and the number of channels of the ninth deconvolution layer is 32, the number of channels of the tenth deconvolution layer is 16, and the number of channels of the eleventh deconvolution layer is 1.