An unsupervised learning method for detecting abnormal fluctuations in battery production processes
The HFFRN network addresses inefficiencies in anomaly detection by integrating CNN-ConvLSTM and attention mechanisms to enhance feature extraction and adaptive thresholding, improving anomaly detection in battery production processes.
Patent Information
- Application Number
- CN202210375802.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-11
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-04-11
AI Technical Summary
The existing time series anomaly detection methods are not effective in battery production, difficult to effectively capture the correlation of time series and are sensitive to noise, resulting in high risk of misclassification.
The HFFRN network is reconstructed by hierarchical feature fusion, combined with the CNN-ConvLSTM-ATTENTENT network, and features are extracted through multi-layer convolution and attention mechanisms, and abnormal detection is performed using an adaptive threshold mechanism.
It improves the detection efficiency of abnormal fluctuations in battery production processes, reduces noise interference, enhances feature interaction and information perception, and achieves efficient abnormal point recognition.
Smart Images

Figure CN114648076B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of industrial production process fluctuation analysis, and particularly to an unsupervised learning method for detecting abnormal process fluctuations. Background Art
[0002] With the development of the Internet of Things and sensor technology, a large amount of high-dimensional sensor data is recorded in industrial production to monitor the operating status of equipment. As an important way to monitor the equipment status, the anomaly detection of time series data has received extensive attention in urban resource management, computer network intrusion detection, healthcare, and industrial production. However, most of the data collected by sensors is in a normal state, while abnormal data is less due to its difficulty in obtaining, which increases the burden of anomaly detection. How to complete the anomaly point detection through existing technologies, thereby improving the efficiency of industrial production process fluctuation analysis and reducing the operation and maintenance costs of production equipment has great application value.
[0003] Traditional anomaly point detection methods include distance-based methods and classification-based methods. Among them, a class of support vector machine algorithms is a representative algorithm of distance-based methods, which distinguishes normal points and abnormal points in data by modeling the training data. Although this method can effectively solve the data discrimination, it only focuses on the information of the data itself and does not capture the correlation of the time series, resulting in poor results. Another method is the classification-based method, such as isolation forest, which builds a binary tree by randomly selecting features to complete the anomaly point detection of the time series. This method requires a large amount of time for feature design and preprocessing, resulting in unstable detection effects.
[0004] In view of the shortcomings of traditional anomaly point detection methods, a large number of deep learning-based methods have emerged. Among them, the Long Short-Term Memory (LSTM) is one of the typical deep learning models for processing time series. The data anomaly point detection method based on the LSTM model mainly predicts the data within the next time step and calculates the distance between the predicted data and the real data, and then determines whether an anomaly occurs at this time point based on the size of the distance. The LSTM model can learn the non-linear relationship in short-term or long-term time series data, but it is very sensitive to noise, which increases the risk of misclassification in anomaly detection.
[0005] In summary, in dealing with the problem of process fluctuation analysis in the battery production process, finding an effective method for time series anomaly detection to replace the existing technology has become an urgent problem to be solved. Summary of the Invention
[0006] The present invention proposes an unsupervised learning-based method for detecting abnormal fluctuations in battery production processes, namely the Hierarchical Feature Fusion Reconstruction Network (HFFRN). The algorithm flow chart is as shown in Figure 1 , and this algorithm can realize the function of detecting abnormal process fluctuations during battery production.
[0007] The technical solution of the present invention includes the following steps:
[0008] 1) Data collection: Through sensors installed on battery production equipment, collect the time series of process parameters during production to establish a data set;
[0009] 2) Data processing: Divide the original data set into a training set X train , a validation set X val , and a test set X test , where X train only contains normal samples, and X val and X test contain both normal and abnormal samples; after processing each data set, generate a multi-channel feature matrix M;
[0010] 3) Build the HFFRN model and use the training set X train to train the model;
[0011] Build the HFFRN model, and the specific construction steps are as follows:
[0012] The HFFRN model contains a CNN-ConvLSTM-ATTENTION network. The first layer of the model consists of a 4-layer two-dimensional convolutional layer as the encoding layer. Each two-dimensional convolutional layer contains a two-dimensional convolution 2D-Conv, a batch normalization layer BN, and a leaky rectified linear unit LRelu activation function. By receiving a multi-channel feature matrix sequence related to the process, the weights of different convolutional kernels and the width of the window are selected. The 4-layer 2D-Conv extracts the spatial features of the feature matrix layer by layer, completing the encoding of the spatial features. The output of each layer is used as the input of the middle layer of the lower layer. The middle layer consists of four parallel ConvLSTM blocks, which are used to capture the features in both the time and space dimensions of the multi-channel feature matrix sequence. At the same time, an attention mechanism is introduced on this basis. The attention mechanism completes the weight allocation, allocating more attention weights to key features and reducing interference. The decoding layer consists of 4 two-dimensional deconvolution layers. Each layer contains a two-dimensional deconvolution 2D-Deconv, a batch normalization layer BN, and a leaky rectified linear unit LRelu activation function. The two-dimensional deconvolution 2D-Deconv, as the inverse process of 2D-Conv, can restore the feature map feature information by adjusting the convolution stride. The deconvolution operation decodes from the last layer to the first layer in reverse, extracting the feature matrix of each layer. For the last layer, the feature extraction matrix is directly obtained by performing deconvolution on the output of ConvLSTM. For the other three layers, the feature extraction matrix of this layer is obtained by concatenating the output of the ConvLSTM of this layer with the feature extraction matrix of the i+1 layer and then performing deconvolution. The feature extraction matrix is represented by the following formula:
[0013]
[0014] Among them, c1, c2, c3, and c4 are the outputs of each block of ConvLSTM. represents the deconvolution operation. represents the concatenation operation, W de i and b di i are the weights and biases of the deconvolution kernel in the i-th layer.
[0015] In the hierarchical feature fusion layer, the feature extraction matrix is converted into four feature matrices with the same dimension and size through 4 parallel two-dimensional convolutions 2D-Conv. The formula is as follows:
[0016] p i = f(W re i * r i + b re i ) i = 1, 2, 3, 4 (2)
[0017] Among them, pi are feature matrices with the same dimensions and sizes, and W re i is the weight of the convolution kernel, and b re i is the bias of the convolution kernel;
[0018] After that, these four feature matrices are input into the connection layer for splicing; finally, by using a 2×2 convolution kernel to perform information fusion on the eigenvalues in the feature matrix, a fused feature matrix is obtained, which is also the final required reconstructed feature matrix. The reconstructed feature matrix is expressed as:
[0019] O = f([p1;p2;p3;p4]*W + b) (3)
[0020] where O is the reconstructed feature matrix, W is the weight of the convolution kernel, b is the bias of the convolution kernel; [p1;p2;p3;p4] represents the splicing of p1, p2, p3, and p4;
[0021] 4) Use the validation set X val to verify the trained network model, and by using a fully adaptive optimal threshold determination mechanism, find a suitable threshold T;
[0022] Use a fully adaptive optimal threshold determination mechanism to find a suitable threshold T to divide the anomaly scores. The points with anomaly scores below this threshold are normal points, and vice versa; the specific steps are as follows:
[0023] Since the model is trained using normal samples, the model can achieve good reconstruction for normal samples. In the data of the validation set X val there are both normal data and abnormal data, and a large error will occur during reconstruction. The anomaly score is represented by the following formula:
[0024] ε = ||O - M||2 (4)
[0025] where || || represents the L2 norm;
[0026] The anomaly scores ε1, ε2,..., ε obtained in the validation set X val are sorted in ascending order to obtain new anomaly scores n Suppose the optimal threshold is T, such that when the anomaly score ε ≤ T, the event is classified as a normal point, and vice versa; at this time, T satisfies the following conditions: t ≤ T, the event is classified as a normal point, and vice versa; at this time, T satisfies the following conditions:
[0027]
[0028] where \(1\leq t\lt n\); at this time, the proportions of normal events and abnormal events are \(w_0\) and \(w_1\), and their mean square deviations are \(\delta_0\) and \(\delta_1\), which are expressed by the following formulas:
[0029]
[0030]
[0031] where \(x\) i represents the \(i\)-th anomaly score, and \(\mu\) represents the mean of the overall scores;
[0032] Assume that at the optimal threshold \(T\), the between-class variance of the maximized anomaly scores after reconstructing normal and abnormal samples is \(\delta\) 2 , then
[0033] \(\delta\) 2 = \(w_0(t)\cdot\delta_0(t)+w_1(t)\cdot\delta_1(t)\) (8)
[0034] For finding the optimal threshold \(T\), it only needs to traverse all times \(t\), find the \(t\) that maximizes the above formula, and take the corresponding as the optimal threshold, that is
[0035] 5) Use the test set \(X\) test to measure the model performance and generalization ability.
[0036] Preferably, in the step 1) data acquisition, a sensor is used to obtain the time series under each process in the battery production process, where the processes include mixing, coating, laminating, and formation; and it is established into the original data set \(X=(x\) 0 , \(x\) 1 ,…, \(x\) N ) T \(\in R\) N×T , where \(N\) is the dimension of the sequence and \(T\) is the length of each sequence; each sample contains \(T\) sample points, that is, \(x\) i =(x1 i , x2 i ,…, \(x\) T i ), where \(i\) represents the \(i\)-th sample in the original data set.
[0037] Preferably, in the step 2) data processing, the data set is divided into a training set \(X\) train a validation set \(X\) val and a test set \(X\) test in a ratio of 2:1:1; for each part of the data, it is segmented into equal-length segments by a sliding window, the length of the window is \(w\), and then the cross-correlation calculation is performed on any two time series in the window to generate a multi-channel feature matrix \(M\).
[0038] Preferably, the mean squared error is used as the loss function of the model during the training of the HFFRN model.
[0039] Preferably, the test set X is used test to measure the model performance and generalization ability. Specifically, for the model with training and parameter adjustment completed, use the test set X test to perform model testing on the data, obtain the anomaly scores of each sample in the test set, and evaluate the performance of the model through precision, recall rate, and F1 score.
[0040] The beneficial effects of the present invention are as follows:
[0041] The feature fusion reconstruction network (HFFRN) proposed by the present invention. In this network, the encoder network performs spatial feature extraction on the input multi-channel feature matrix through multi-layer convolution operations. The ConvLSTM network can extract the temporal features of the input multi-channel feature matrix sequence at different time steps, complete the feature capture of the data. At the same time, the added attention mechanism can complete the weight allocation, allocate more attention weights to the key features and reduce the interference of noise. Through the decoder network, the feature map obtained in the previous step can be decoded. At the same time, the asymmetric ability of the feature matrix information is used to construct the feature extraction matrix, thereby enhancing the feature reuse between layers. The hierarchical feature fusion model increases the feature interaction between layers, enabling the model to perceive more information in the feature dimension and time dimension. Finally, an adaptive threshold mechanism is proposed to effectively and adaptively divide the anomaly scores. Ultimately, the detection of abnormal fluctuations in the battery production process is realized. Description of the Drawings
[0042] Figure 1 is a flowchart of the method of the present invention;
[0043] Figure 2 is the overall structure diagram of the neural network of the present invention. Detailed Embodiments
[0044] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0045] As Figure 1 shown, the present invention includes the following steps.
[0046] 1) Data acquisition: Through the sensors installed on the battery production equipment, collect the time series of process parameters during the production process and establish a data set.
[0047] 2) Data processing: Divide the original data set into a training set X train , a validation set X val , and a test set X test , where Xtrain only contains normal samples, X val and X test contains both normal and abnormal samples.
[0048] 3) Build the HFFRN model and use the training set X train to train the model.
[0049] 4) Use the validation set X val to validate the trained network model and find a suitable threshold T by using a fully adaptive optimal threshold determination mechanism.
[0050] 5) Use the test set X test to measure the model performance and generalization ability.
[0051] In the said step 1) data acquisition, use sensors to obtain the time series under each process in the battery production process, where the processes include mixing, coating, laminating, and formation. Establish it into the original data set X = (x 0 , x 1 , …, x N ) ∈ R N×T (where N is the dimension of the sequence and T is the length of each sequence). Each sample contains T sample points, that is, x i = (x1 i , x2 i , …, x T i ), where i represents the i-th sample in the original data set.
[0052] In the said step 2) data processing, divide the data set into the training set X train validation set X val and test set X test according to the ratio of 2:1:1. For each part of the data, perform equal-length segmentation of the data through a sliding window (the length of the window is w), and then calculate the cross-correlation between any two time series in the window to generate a multi-channel feature matrix M.
[0053] The overall structure diagram of the HFFRN model in the said step 3) is as shown in Figure 2 shown, and the specific construction steps are as follows:
[0054] The HFFRN model contains a CNN-ConvLSTM-ATTENTION network. The first layer of the model consists of a 4-layer two-dimensional convolutional layer as the encoding layer. Each two-dimensional convolutional layer contains two-dimensional convolution (2D-Conv, Two-dimension convolution), batch normalization layer (BN, Batch Normalization), and leaky rectified linear unit (LRelu, Leaky rectified linear unit) activation function. By receiving a multi-channel feature matrix sequence related to the process, selecting the weights of different convolutional kernels and the width of the window, the 4-layer 2D-Conv extracts the spatial features of the feature matrix layer by layer, completing the encoding of the spatial features. The output of each layer is used as the input of the intermediate layer of the lower layer. The intermediate layer consists of four parallel ConvLSTM blocks, which are used to capture the features in both the time and space dimensions of the multi-channel feature matrix sequence. At the same time, an attention mechanism is introduced on this basis. The attention mechanism completes the weight allocation, allocating more attention weights to the key features and reducing interference. The decoding layer consists of 4 two-dimensional deconvolution layers, and each layer contains two-dimensional deconvolution (2D-Deconv, Two-dimension deconvolution), BN layer, and activation function LRelu. 2D-Deconv, as the inverse process of 2D-Conv, can restore the feature information of the feature map by adjusting the convolution stride. The deconvolution operation decodes from the last layer to the first layer in reverse, extracting the feature matrix of each layer. For the last layer (i = 4), the feature extraction matrix is directly obtained by performing deconvolution operation on the output of ConvLSTM. For the remaining three layers (0 < i < 4), the feature extraction matrix of this layer is obtained by concatenating the output of the ConvLSTM of this layer with the feature extraction matrix of the i + 1 layer and then performing deconvolution operation. The feature extraction matrix can be represented by the following formula:
[0055]
[0056] Among them, c1, c2, c3, and c4 are the outputs of each block of ConvLSTM, represents the deconvolution operation, represents the concatenation operation, W de i and b di i are respectively the weights and biases of the deconvolution kernel in the i-th layer.
[0057] In the hierarchical feature fusion model, the feature extraction matrix is converted into four feature matrices with the same dimension and size through 4 parallel 2D-Conv. The formula is as follows:
[0058] p i = f(W rei *r i +b re i ) i = 1, 2, 3, 4 (2)
[0059] Among them, p i is a feature matrix with the same dimension and size, and W re i is the weight of the convolution kernel, and b re i is the bias of the convolution kernel.
[0060] After that, these four feature matrices are input into the connection layer for splicing. Finally, information fusion is performed on the eigenvalues in the feature matrix by using a 2×2 convolution kernel to obtain a fused feature matrix, that is, the final required reconstructed feature matrix, and the reconstructed feature matrix is expressed as:
[0061] O = f([p1; p2; p3; p4] * W + b) (3)
[0062] Among them, O is the reconstructed feature matrix, W is the weight of the convolution kernel, and b is the bias of the convolution kernel. [p1; p2; p3; p4] represents the splicing of p1, p2, p3, and p4.
[0063] After the model is constructed, the training set X train data is used to train the HFFRN model, and it is trained repeatedly for many times to make the model effect reach the optimal. Among them, the mean square error is used as the loss function of the model.
[0064] In step 4), the validation set X val data is used to verify the trained network model, and a suitable threshold T is found by using a fully adaptive optimal threshold determination mechanism.
[0065] In this step, a fully adaptive optimal threshold determination mechanism is used to find a suitable threshold T to divide the anomaly scores. The anomaly scores below this threshold are normal points, and vice versa. The specific steps are as follows:
[0066] Since the model is trained with normal samples, the model can achieve good reconstruction for normal samples. The validation set X val data contains both normal data and abnormal data, and a large error will occur during reconstruction. The anomaly score can be expressed by the following formula:
[0067] ε = ||O - M||2 (4)
[0068] Among them, || || represents the L2 norm.
[0069] It will be in the validation set X valThe obtained anomaly scores ε1, ε2,..., ε n are sorted in ascending order to obtain the new anomaly scores as Assume that the optimal threshold is T, such that when the anomaly score ε t ≤T, the event is classified as a normal point, otherwise as an abnormal point. At this time, T satisfies the following conditions:
[0070]
[0071] where 1 ≤ t < n. At this time, the proportions of normal events and abnormal events are w0 and w1, and their mean square deviations are δ0 and δ1, which can be expressed by the following formulas:
[0072]
[0073]
[0074] Assume that at the optimal threshold T, the between-class variance of the maximum anomaly scores after reconstructing normal and abnormal samples is δ 2 , then
[0075] δ 2 = w0(t)·δ0(t) + w1(t)·δ1(t) (8)
[0076] For finding the optimal threshold T, only need to traverse all times t, find the t that maximizes the above formula, and take the corresponding as the optimal threshold, that is
[0077] In step 5), the model with training and adjustment parameters completed is tested with the data of the test set X test to obtain the anomaly scores of each sample in the test set, and the performance of the model is evaluated by precision, recall, and F1 score.
Claims
1. An unsupervised learning method for detecting abnormal fluctuations in battery production processes, characterized in that, The method specifically includes the following steps: 1) Data collection: Collect the time series of process parameters during the production process through sensors installed on battery production equipment to establish a data set; 2) Data processing: Divide the original dataset into a training set X train , a validation set X val , and a test set X test , where X train only contains normal samples, and X val and X test contain both normal and abnormal samples; After processing each dataset, generate a multi-channel feature matrix M; 3) Build the HFFRN model and use the training set X train Train the model; Build an HFFRN model, and the specific construction steps are as follows: The HFFRN model contains a CNN-ConvLSTM-ATTENTION network. The first layer of the model consists of 4 layers of two-dimensional convolutional layers to form an encoding layer. Each two-dimensional convolutional layer contains a two-dimensional convolution 2D-Conv, a batch normalization layer BN, and a leaky rectified linear unit LRelu activation function; by receiving a multi-channel feature matrix sequence related to the process, select the weights of different convolutional kernels and the width of the window. The 4 layers of 2D-Conv extract the spatial features of the feature matrix layer by layer to complete the encoding of the spatial features. The output of each layer is used as the input of the middle layer of the lower layer; the middle layer consists of four parallel ConvLSTM blocks, which are used to capture the features in both the time and space dimensions of the multi-channel feature matrix sequence, and at the same time introduce an attention mechanism on this basis; the attention mechanism completes the weight allocation, allocating more attention weights to key features and reducing interference; the decoding layer consists of 4 two-dimensional transposed convolutional layers, and each layer contains a two-dimensional transposed convolution 2D-Deconv, a batch normalization layer BN, and a leaky rectified linear unit LRelu activation function; the two-dimensional transposed convolution 2D-Deconv is the inverse process of 2D-Conv, and the feature map feature information can be restored by adjusting the convolution stride; the transposed convolution operation decodes from the last layer to the first layer in reverse to extract the feature matrix of each layer; for the last layer, the feature extraction matrix is directly obtained by performing a transposed convolution operation on the output of ConvLSTM; for the other three layers, after splicing the output of the ConvLSTM of this layer with the feature extraction matrix of the i+1 layer and then performing a transposed convolution operation, the feature extraction matrix of this layer is obtained; the feature extraction matrix is represented by the following formula: Among them, c1, c2, c3, and c4 are the outputs of each block of ConvLSTM. represents the transposed convolution operation. represents the concatenation operation, W de i and b de i are the weight and bias of the transposed convolution kernel in the i-th layer, respectively. In the hierarchical feature fusion layer, the feature extraction matrix is converted into four feature matrices with the same dimension and size through 4 parallel two-dimensional convolutions 2D-Conv. The formula is as follows: p i = f(W re i * r i + b re i ) i = 1, 2, 3, 4 (2) Among them, p i is a feature matrix with the same dimension and size, W re i is the weight of the convolution kernel, b re i is the bias of the convolution kernel; After that, these four feature matrices are input into the connection layer for splicing; finally, information fusion is performed on the eigenvalues in the feature matrix by using a 2×2 convolutional kernel to obtain a fused feature matrix, that is, the final required reconstructed feature matrix. The reconstructed feature matrix is represented as: O = f([p1;p2;p3;p4]*W + b)(3) Among them, O is the reconstructed feature matrix, W is the weight of the convolutional kernel, and b is the bias of the convolutional kernel; [p1;p2;p3;p4] represents the splicing of p1, p2, p3, and p4; 4) Use the validation set X val To validate the trained network model and find an appropriate threshold T by using a fully adaptive optimal threshold determination mechanism; Use a fully adaptive optimal threshold determination mechanism to find a suitable threshold T to divide the anomaly scores. The points with anomaly scores below this threshold are normal points, and vice versa are anomaly points; the specific steps are as follows: Since the model is trained using normal samples, it can achieve good reconstruction for normal samples. In the validation set X val the data contains both normal and abnormal data, and a large error will occur during reconstruction. The anomaly score is represented by the following formula: ε = ||O - M||2 (4) Among them, || || represents the L2 norm; Will be in the validation set X val The obtained anomaly scores ε1, ε2,..., ε n Sort them in ascending order to obtain the new anomaly scores as Assume that the optimal threshold is T, such that when the anomaly score ε t ≤T events are classified as normal points, otherwise as anomaly points; at this time, T satisfies the following conditions: where 1 ≤ t < n; at this time, the proportions of normal events and abnormal events are w0 and w1, and their mean square deviations are δ0 and δ1, which are expressed by the following formula: where x i represents the i-th anomaly score, and μ represents the mean of the overall scores; Assume that at the optimal threshold T, the between-class variance of the maximized anomaly score after reconstructing normal and abnormal samples is δ 2 , then δ 2 = w0(t)·δ0(t) + w1(t)·δ1(t) (8) For finding the optimal threshold T, it is only necessary to traverse all times t, find the t that maximizes the above formula, and use its corresponding as the optimal threshold, that is 5) Use the test set X test Measure the model performance and generalization ability.
2. The method for detecting abnormal fluctuations in a battery production process by unsupervised learning according to claim 1, characterized in that: In the step 1) of data acquisition, sensors are used to obtain time series under each process in the battery production process, where the processes include mixing, coating, stacking, and formation; Build it into the original dataset \(X=(x 0 ,x 1 ,…,x N ) T \in\mathbb{R} N×T , where \(N\) is the dimension of the sequence and \(T\) is the length of each sequence; each sample contains \(T\) sample points, that is, \(x i =(x_1 i ,x_2 i ,…,x T i ), where \(i\) represents the \(i\)-th sample in the original dataset.
3. The method for detecting abnormal fluctuations in a battery production process by unsupervised learning according to claim 1, characterized in that: In the data processing of step 2), the data set is divided into a training set X, a validation set X, and a test set X in a ratio of 2:1:
1. For each part of the data, equal-length segmentation of the data is performed through a sliding window. The length of the window is w. Then, cross-correlation calculation is performed on any two time series in the window to generate a multi-channel feature matrix M. train validation set X val and test set X test ; For each part of the data, equal-length segmentation of the data is performed through a sliding window. The length of the window is w. Then, cross-correlation calculation is performed on any two time series in the window to generate a multi-channel feature matrix M.
4. A method for detecting abnormal fluctuations in a battery production process through unsupervised learning according to claim 1, characterized in that: The mean square error is used as the loss function of the model in the training of the HFFRN model.
5. The method for detecting abnormal fluctuations in a battery production process based on unsupervised learning according to claim 1, characterized in that: The described use of the test set X test to measure the model performance and generalization ability, specifically: the model that has completed training and adjusting parameters is tested with the test set X test data to obtain the anomaly scores of each sample in the test set, and the performance of the model is evaluated through precision, recall, and F1 score.
Citation Information
Patent Citations
Anomaly detection method for continuous space-time refueling data
CN110232082A
Abnormal event detection method of time-space variational self-encoding network based on self-attention enhancement
CN113449660A