Steel production system fault diagnosis method based on ConvLSTM-AE

By constructing the ConvLSTM-AE model, combining the advantages of ConvLSTM and AE, spatiotemporal features are extracted and reconstructed, and using kernel density estimation (KDE) to calculate thresholds, the shortcomings of autoencoder and ConvLSTM in the existing technology are solved, and efficient fault diagnosis is achieved.

CN120011985APending Publication Date: 2025-05-16HANGZHOU ZETA TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411958639.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In the prior art, the autoencoder (AE) lacks the ability to model time series features, while the ConvLSTM lacks the feature expression when processing high-dimensional complex data, resulting in insufficient fault diagnosis accuracy.

Method used

A fault diagnosis method for steel production system based on ConvLSTM-AE is proposed. By constructing the ConvLSTM-AE model, combining the advantages of ConvLSTM and AE, spatiotemporal features are extracted and reconstructed, and the threshold is calculated using kernel density estimation (KDE) to realize fault monitoring.

Benefits of technology

It effectively improves fault diagnosis capabilities, especially when processing high-dimensional and multi-variable industrial data, it shows superior performance, can accurately detect and diagnose system failures, and improves the accuracy and reliability of fault diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011985A_ABST
    Figure CN120011985A_ABST
Patent Text Reader

Abstract

The invention relates to a steel production system fault diagnosis technology, and aims to provide a ConvLSTM-AE-based steel production system fault diagnosis method. The method comprises the following steps: collecting process variable data in an offline training stage, preprocessing, constructing a training set, inputting the training set into a ConvLSTM-AE model, and obtaining hidden features and residual errors; respectively establishing statistics T2 and SPE, and calculating a threshold value by using KDE; in the online monitoring stage, variable data in the operation process are collected in real time, the preprocessed data are input into the trained model, and hidden features and residual errors are output; calculating a statistical magnitude based on the output result, and comparing the statistical magnitude with an off-line training stage threshold value; if both are greater than the threshold value, the system is considered to have a fault and an alarm signal is sent out; otherwise, considering that no fault occurs and continuously monitoring. According to the method, under the clamping of the convolutional neural network and the recurrent neural network, the unicity of variables extracted by a traditional auto-encoder can be made up, the features of the spatial dimension and the features of the time dimension are effectively fused, and more comprehensive process monitoring is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to process diagnosis technology, and specifically to a fault diagnosis method for a steel production system based on ConvLSTM-AE. Background Art

[0002] If a key process or equipment of the system fails during the steel production process, it may lead to production losses, reduced energy efficiency, energy waste and additional operating costs. Through system fault diagnosis, the root cause of the problem can be quickly determined so that appropriate repair measures can be taken to reduce downtime and maximize production recovery. Through system fault diagnosis, energy efficiency problems in the system can also be determined, and appropriate measures can be taken to improve the energy efficiency of the system, reduce energy consumption and operating costs. In particular, in most industrial data, the time dimension characteristics and spatial dimension characteristics coexist. If only the characteristics of a single dimension are used for fault diagnosis, the key characteristics of other dimensions will be lost, resulting in missed faults and false alarms, which may miss the critical period for repairing the fault. The data-driven fault diagnosis method does not require the establishment of an accurate mechanism model of the target system, but only requires a large amount of historical data to achieve its purpose. Therefore, it has been favored for a long time in industrial processes with multiple loops, multiple devices, and difficult mechanism modeling.

[0003] Among the existing data-driven fault diagnosis methods, autoencoder (AE) and convolutional long short-term memory network (ConvLSTM) are two commonly used neural network models.

[0004] An autoencoder (AE) is an unsupervised neural network model that can learn the implicit features of input data (this is called encoding) and reconstruct the original input data with the learned new features (this is called decoding). Autoencoders can automatically learn useful feature representations from data without manually designing features, which makes it very effective in data preprocessing and feature extraction. Autoencoders can compress high-dimensional data into low-dimensional encodings, thereby achieving dimensionality reduction and compression of data, which is conducive to saving storage space and computing costs. Autoencoders can perform denoising and reconstruction through training data, that is, restore the original data from damaged or noisy data, which helps to recover and reconstruct data. Autoencoders can achieve nonlinear modeling by stacking multiple hidden layers, making them more flexible in dealing with complex data and nonlinear relationships. However, autoencoders are prone to overfitting during training, especially when the encoding dimension is high or the amount of data is small.

[0005] ConvLSTM is a neural network structure that combines a convolutional neural network (CNN) and a long short-term memory network (LSTM). It is an extension of a recurrent neural network (RNN) and is used to process data with temporal characteristics, such as time series, video sequences, or sequence data. ConvLSTM combines the advantages of CNN and LSTM, aiming to solve the gradient vanishing and gradient exploding problems of traditional RNN in long time series tasks. By using convolution operations to capture local features of the input sequence and using the gating mechanism of LSTM to model long-term dependencies between sequences, ConvLSTM has better results when processing sequence data.

[0006] When used alone, these two neural network models have the following defects: Autoencoder (AE) extracts implicit features from data in an unsupervised manner and reconstructs the input data. However, AE has limited capabilities in processing time series features and has difficulty capturing the temporal dependencies of data. Due to AE's lack of modeling capabilities for data temporal features, it may lead to incomplete feature extraction of complex industrial data. ConvLSTM combines the advantages of convolutional neural networks (CNN) and long short-term memory networks (LSTM) to capture local spatial features and temporal characteristics of data. However, ConvLSTM may have problems with insufficient feature expression when processing high-dimensional complex data, and lacks an effective reconstruction mechanism, which may lead to insufficient fault diagnosis accuracy.

[0007] Therefore, based on the above-mentioned drawbacks in the processing of industrial data in the generation process, existing research results are limited to using a single method to achieve fault diagnosis, and it is difficult to find a suitable method to organically combine AE and ConvLSTM to comprehensively reflect the technical advantages of both.

[0008] The present invention proposes a new solution to solve the above problems. Summary of the invention

[0009] The technical problem to be solved by the present invention is to overcome the deficiencies in the prior art and provide a steel production system fault diagnosis method based on ConvLSTM-AE.

[0010] To solve the technical problem, the solution of the present invention is:

[0011] A steel production system fault diagnosis method based on ConvLSTM-AE is provided, which specifically includes an offline training stage and an online monitoring stage; wherein:

[0012] The offline training phase includes:

[0013] (1.1) Collect sufficient process variable data for the blast furnace ironmaking system under normal operation and fault conditions; preprocess the data and construct a training set;

[0014] (1.2) Build ConvLSTM-AE model;

[0015] (1.3) Input the training set into the ConvLSTM-AE model to obtain hidden features and residuals;

[0016] (1.4) Using hidden features to establish statistics T 2 , and use KDE to calculate the threshold; use the residual to establish the statistic SPE, and use KDE to calculate the threshold;

[0017] The online monitoring phase includes:

[0018] (2.1) Real-time collection of variable data of the blast furnace ironmaking system during operation and pre-processing;

[0019] (2.2) Input the processed data into the trained ConvLSTM-AE model and output the hidden features and residuals;

[0020] (2.3) Based on the output results, calculate and obtain the statistic T 2 and the statistic SPE, and compare them with the thresholds obtained in the offline training phase respectively; if both are greater than the thresholds, the system is considered to have failed and an alarm signal is issued; otherwise, it is considered that no failure has occurred and monitoring continues;

[0021] Among them, the ConvLSTM-AE model constructed in step (1.2) adopts the encoder (Encoder) and decoder (Decoder) of the AE model (Auto-Encoder) as the construction body, and uses the ConvLSTM network as the encoder of the AE model; among them, the encoder is used to extract the spatiotemporal features of the input data and generate hidden features, and the decoder is used to reconstruct the latent variables and generate outputs similar to the input.

[0022] The present invention also provides a computer device, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor executes the aforementioned steel production system fault diagnosis method based on ConvLSTM-AE.

[0023] The present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the aforementioned steel production system fault diagnosis method based on ConvLSTM-AE.

[0024] Description of the invention principle:

[0025] Considering that the autoencoder (AE) lacks the ability to model the temporal characteristics of data, which may lead to incomplete feature extraction of industrial complex data, and the ConvLSTM network has insufficient feature expression capabilities for high-dimensional complex data and lacks an effective reconstruction mechanism, which may lead to insufficient fault diagnosis accuracy. This application innovatively proposes to organically combine the two to achieve fault monitoring in steel production systems.

[0026] To this end, the present invention proposes an innovatively constructed ConvLSTM-AE model, in which the ConvLSTM network is used as the encoder part of the AE model to extract spatiotemporal features; the decoder reversely reconstructs the data through the DeConvLSTM layer. Based on the organic combination of the two, the hidden features can be extracted from the original process data, and the corresponding statistics can be constructed using the extracted hidden variables and the residuals between input and output; and the Kernel Density Estimation (KDE) method is used to determine its threshold, thereby realizing process monitoring.

[0027] The present invention innovatively improves the fault diagnosis capability through the ConvLSTM-AE model processing method, especially showing superior performance when processing high-dimensional, multivariate industrial data. Through experimental verification, this method can effectively predict and diagnose potential faults in steel production systems, providing strong technical support for the intelligent upgrading of the manufacturing industry.

[0028] The following technical defects or difficulties are common in the existing technology: (1) Autoencoders (AE) lack the ability to model time series features. Traditional AE mainly focuses on extracting features from static data and performing dimensionality reduction or denoising. It is not good at processing dynamic changes and dependencies in time series data. (2) ConvLSTM is not good at expressing the features of high-dimensional complex data: Although ConvLSTM can capture spatial and temporal features at the same time, it may not be able to fully express high-dimensional and complex industrial process data, especially in terms of feature reconstruction.

[0029] When combining the two, the existing technology has the following difficulties: (1) Increased model complexity: Combining AE with ConvLSTM makes the model more complex, which not only increases the difficulty of training, but also makes hyperparameter adjustment more difficult. (2) Increased computing resource requirements: The combined model requires more computing resources for training and reasoning, which places higher requirements on hardware. (3) Overfitting risk: Due to the introduction of more network layers and parameters, the combined model is more prone to overfitting, especially when the number of training samples is limited.

[0030] In view of the above problems, the present invention focuses on making innovative and specific improvements from the following points: (1) Spatiotemporal feature fusion: The present invention adopts ConvLSTM to replace the traditional fully connected layer as part of the encoder, so that spatiotemporal features can be effectively extracted from the input data. This solves the limitation that AE is difficult to capture time series characteristics. By using the DeConvLSTM layer to reversely reconstruct data for decoding, not only the data compression and reconstruction capabilities of AE are retained, but also the learning of multi-dimensional features is enhanced. (2) Two-dimensional feature processing: The present invention realizes the simultaneous processing of spatial dimension features and temporal dimension features of data, making up for the deficiency that traditional AE can only process a single data dimension, and providing more comprehensive process monitoring. (3) Combination of feature extraction and reconstruction: While maintaining the powerful spatiotemporal feature extraction capability of ConvLSTM, the reconstruction mechanism of the autoencoder is added, so that the model can not only learn the intrinsic characteristics of the data, but also verify the effectiveness of the model through the reconstructed output. (4) Real-time monitoring and statistical analysis: By real-time monitoring of the implicit features and residuals extracted by the model and combining statistical analysis methods, the present invention can more accurately detect and diagnose system faults, improving the accuracy and reliability of fault diagnosis. (5) Reduce the risk of overfitting: Through carefully designed loss functions and regularization techniques, the present invention effectively reduces the risk of overfitting and ensures the generalization performance of the model on new data.

[0031] Compared with the prior art, the present invention has the following beneficial effects:

[0032] 1. The ConvLSTM-AE method proposed in the present invention can make up for the singleness of the variables extracted by the traditional autoencoder with the support of convolutional neural network and recurrent neural network, effectively integrate the features of spatial dimension and time dimension, and realize more comprehensive process monitoring.

[0033] 2. The present invention can establish temporal relationships like LSTM and describe local spatial features like CNN. It establishes statistics for the hidden variables extracted by the ConvLSTM-AE model and the residuals between input and output, and uses KDE to calculate the threshold and model construction: The ConvLSTM-AE model first captures the temporal and spatial characteristics of the input data through the ConvLSTM layer, and then compresses and reconstructs the data through the autoencoder structure. In the encoding stage, convolution operations are used to capture local features; in the decoding stage, deconvolution operations are used to reconstruct data.

[0034] 3. The present invention realizes dual-dimensional feature fusion: by combining ConvLSTM and AE, it can simultaneously process and learn the features of the spatial and temporal dimensions of data. This combination effectively makes up for the deficiency that traditional AE can only process a single data dimension.

[0035] 4. The present invention adopts an improved network structure: the structure of the autoencoder is introduced into the traditional ConvLSTM, so that the network can not only extract features, but also reconstruct the input data through the decoding process, so as to better learn the intrinsic characteristics of the data.

[0036] 5. The present invention achieves enhanced fault detection capability: by real-time monitoring of implicit features and residuals extracted by the model, combined with statistical analysis methods (such as KDE), system faults can be detected and diagnosed more accurately. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 It is a brief flow chart of the fault diagnosis method in the present invention.

[0038] Figure 2 It is a flow chart of the specific implementation process of the present invention. DETAILED DESCRIPTION

[0039] The specific implementation modes of the present invention are described in detail below with reference to the accompanying drawings.

[0040] like Figure 1 As shown, the steel production system fault diagnosis method based on ConvLSTM-AE of the present invention is to establish a ConvLSTM-AE model for process variables to obtain spatiotemporal characteristics, use the latent variables extracted by the model and the residuals between input and output to establish statistics, and use KDE to calculate the threshold and realize process monitoring.

[0041] like Figure 2 As shown, the method specifically includes an offline training phase and an online monitoring phase in practical application.

[0042] (I) Offline training phase includes:

[0043] 1. Collect sufficient process variable data for the blast furnace ironmaking system in normal operation and fault conditions; pre-process the data and construct a training set;

[0044] The process variable data here refers to the key signal variables in the steel production process, including temperature, pressure, flow rate, oxygen content and raw material ratio, etc.; the preprocessing includes standardization and noise reduction. In the training set, the matrix X = [x1 x2…x N ]∈R m×N .

[0045] 2. Build the ConvLSTM-AE model;

[0046] The traditional ConvLSTM model is only applicable to time series prediction or classification, and fails to directly deal with the problem of anomaly detection. The present invention abandons its limitations in specific tasks, and retains the modeling ability of the traditional ConvLSTM model for spatiotemporal sequences in the ConvLSTM-AE model, which can effectively capture the dynamic characteristics in the time dimension and the local characteristics in the spatial dimension. Specifically, the model uses the encoder (Encoder) and decoder (Decoder) of the AE model (Auto-Encoder) as the main structure, and uses the ConvLSTM network as the encoder of the AE model; wherein the encoder is used to extract the spatiotemporal features of the input data and generate hidden features, and the decoder is used to reconstruct the hidden variables and generate outputs similar to the input.

[0047] The encoder includes a convolution layer Conv for extracting spatial features and an LSTM unit for extracting temporal features; the formula of the LSTM unit is as follows:

[0048] H t =σ(W*X t +U+H t-1 +b)

[0049] Among them, H t represents the hidden state; σ represents the Sigmoid function in the activation function; * represents the convolution operation; X t Indicates input; H t-1 represents the hidden state at the previous moment; b represents the bias term; W and U represent the weights of each gate;

[0050] The decoder includes a DeConvLSTM unit for decoding hidden features into the original data space and reconstructing data, and a linear layer for outputting the final reconstruction result.

[0051] 3. Input the training set into the ConvLSTM-AE model to obtain hidden features and residuals; specifically:

[0052] (a) The ConvLSTM network is used as the encoder of the AE model, and the convolutional layer Conv is used instead of the fully connected layer to extract data features; the specific construction formula is as follows:

[0053] E=U⊙t t-1 +W⊙x t

[0054] Where E represents the output of the current time step; ⊙ represents the convolution operation; h t-1 Represents the output of the previous ConvLSTM unit, x t represents the input data at time t, W and U represent the weights of each gate;

[0055] (b) Use the LSTM unit in the ConvLSTM network to extract the time feature. The specific construction formula is as follows:

[0056] i t =sigmoid(W i ·[h t-1 ,x t ]+b i )

[0057] f t =sigmoid(W f ·[h t-1 ,x t ]+b f )

[0058] o t =sigmoid(W o ·[h t-1 ,x t ]+b o )

[0059]

[0060] h t =o t ×tanh(c t )

[0061] Among them, i t 、f t , o t , c t 、h t Represents the corresponding input gate, forget gate, output gate, input node, current memory cell state and output state; W i , W f , W o , W crepresents the recursive connection weights of the corresponding thresholds, and represents the weight matrices of the input gate, forget gate, output gate, and candidate memory unit states; sigmoid() and tanh() are two activation functions; b i 、b f 、b o 、b c 、c t-1 Represent the bias items of the input gate, forget gate, output gate and candidate memory unit states respectively;

[0062] (c) The final result of the ConvLSTM network is determined by the following formula:

[0063] o t =sigmoid(W o ·[h t-1 ,x t ]+b o )

[0064]

[0065] h t =o t ×tanh(c t )

[0066] The final result is the hidden features extracted by the encoder.

[0067] (d) The decoder is composed of DeConvLSTM layers, which receives the hidden features extracted by the encoder; t Reconstructed as the input of ConvLSTM-AE;

[0068] The DeConvLSTM layer is composed of several DeConvLSTM units, and uses deconvolution operation to extract features; the specific construction formula is as follows:

[0069]

[0070] Where D represents the output of the current time step; Represents the weight of the DeConvLSTM gate; represents the deconvolution operation; represents the output of the encoder; Represents the output of DeConvLSTM at the current moment;

[0071] (e) The gate calculation of DeConvLSTM is performed in the same way as the encoder, and the output result is expressed as:

[0072]

[0073] in, represents the output of the output gate; Indicates the current memory unit status; Indicates the current hidden state.

[0074] (f) The decoder reconstructs the information with its last linear layer, and the final output is expressed as:

[0075]

[0076] in, represents the final output result, the output of the nth sample corresponding to the time step t; σ represents the Sigmod activation function operation; W y Represents the weight matrix, which is used to linearly transform the hidden state of the current time step; represents the hidden state of the current time step; b y Represents the bias term.

[0077] (g) Minimizing the residual between the output and the output is used as the training objective of the ConvLSTM-AE model;

[0078] The residual is used to reflect the difference between the input and the reconstructed output, reflecting the degree of adaptation of the model to the input data; its specific construction formula is as follows:

[0079]

[0080] Among them, R is the residual; X is the input data, is the reconstructed output of the model;

[0081] If the residual exceeds the set threshold, it is considered that an abnormal operating condition exists.

[0082] 4. Use hidden features to establish statistics T 2 , use the residual to establish the statistic SPE, and use KDE to calculate the threshold.

[0083] (h) Calculate the statistic T according to the following formula 2 And the statistic SPE:

[0084]

[0085] SPE=R 2

[0086] Among them, H T represents the transpose of the matrix H; Representative Matrix The inverse matrix of ; H represents the hidden features; R is the residual.

[0087] (i) Using the kernel density estimation method, calculate the statistic T 2The threshold value is denoted as J(T 2 ):

[0088] The statistic T 2 The probability density function is denoted as p(T 2 ), as follows:

[0089]

[0090] in, Representation Statistics The probability density function of ; N represents the number of samples; K is the kernel function, μ is its kernel width; T 2 (j) is the T of the sample under normal working conditions 2 value, j = 1, 2, ..., N; T 2 is a statistic that measures the difference between the data and the model; T is calculated under normal conditions. 2 The value of .

[0091] The statistics The threshold value is denoted as J(T 2 ), set a confidence level as α, and obtain J(T 2 ) value:

[0092]

[0093] (j) Calculate the threshold of the statistic SPE, denoted as J(SPE);

[0094] The probability density function of the statistic SPE is recorded as p(SPE), as shown below:

[0095]

[0096] Where p(SPE) represents the probability density function of the statistic SPE; N represents the number of samples; K is the kernel function, μ is its kernel width; SPE(j) is the SPE value of the jth sample under normal conditions, j = 1, 2, ..., N;

[0097] The threshold of the statistic SPE is recorded as J(SPE), and a confidence level is set as α. The value of J(SPE) is obtained by solving the following equation:

[0098]

[0099] The kernel function K described above is defined as follows:

[0100]

[0101] Among them, K(g) represents the value of kernel function K; g represents the input variable.

[0102] (II) The online monitoring stage includes:

[0103] (1) Real-time collection of variable data of the blast furnace ironmaking system during operation and pre-processing;

[0104] (2) Input the processed data into the trained ConvLSTM-AE model and output the hidden features and residuals;

[0105] (3) Based on the output results, calculate and obtain the statistic T 2 and the statistic SPE, and compare them with the thresholds obtained in the offline training phase respectively; if both are greater than the thresholds, the system is considered to have failed and an alarm signal is issued; otherwise, it is considered that no failure has occurred and monitoring continues;

[0106] (III) To implement the above method, the present invention provides an applicable computer device and a computer-readable storage medium.

[0107] The computer device includes: at least one processor, and a memory connected to the at least one processor in communication, wherein the memory stores instructions executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the steel production system fault diagnosis method based on ConvLSTM-AE. The computer readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the steel production system fault diagnosis method based on ConvLSTM-AE.

[0108] Compared with traditional technologies, this invention breaks through the limitations of a single model in anomaly detection tasks, and innovatively proposes a threshold judgment method based on statistics and KDE by combining the ConvLSTM-AE model with spatiotemporal modeling and reconstruction capabilities. By inputting process variables into the ConvLSTM-AE model to obtain spatiotemporal features, statistics are established using the extracted latent variables and the residuals between input and output, and KDE is used to calculate thresholds and realize process monitoring. The traditional ConvLSTM model's modeling capabilities for spatiotemporal sequences are retained, and it can effectively capture dynamic characteristics in the time dimension and local features in the spatial dimension.

Claims

1. A steel production system fault diagnosis method based on ConvLSTM-AE, characterized in that: The method specifically includes an offline training phase and an online monitoring phase; wherein, The offline training phase includes: (1.1) Collect sufficient process variable data for the blast furnace ironmaking system under normal operation and fault conditions; preprocess the data and construct a training set; (1.2) Build ConvLSTM-AE model; (1.3) Input the training set into the ConvLSTM-AE model to obtain hidden features and residuals; (1.4) Using hidden features to establish statistics T 2 , and use KDE to calculate the threshold; use the residual to establish the statistic SPE, and use KDE to calculate the threshold; The online monitoring phase includes: (2.1) Real-time collection of variable data of the blast furnace ironmaking system during operation and pre-processing; (2.2) Input the processed data into the trained ConvLSTM-AE model and output hidden features and residuals; (2.3) Based on the output results, calculate and obtain the statistic T 2 and the statistic SPE, and compare them with the thresholds obtained in the offline training phase respectively; if both are greater than the thresholds, the system is considered to have failed and an alarm signal is issued; otherwise, it is considered that no failure has occurred and monitoring continues; Among them, the ConvLSTM-AE model constructed in step (1.2) adopts the encoder (Encoder) and decoder (Decoder) of the AE model (Auto-Encoder) as the construction body, and uses the ConvLSTM network as the encoder of the AE model; among them, the encoder is used to extract the spatiotemporal features of the input data and generate hidden features, and the decoder is used to reconstruct the latent variables and generate outputs similar to the input.

2. The method according to claim 1, characterized in that In the step (1.1), the process variable data refers to key signal variables in the steel production process, including temperature, pressure, flow rate, oxygen content and raw material ratio; the preprocessing includes standardization and noise reduction processing.

3. The method according to claim 1, characterized in that In the step (1.1), the training set includes a matrix X=[x1 x2…x N ]∈R m×N .

4. The method according to claim 1, characterized in that: In the ConvLSTM-AE model: The encoder includes a convolution layer Conv for extracting spatial features and an LSTM unit for extracting temporal features; the formula of the LSTM unit is as follows: H t =σ(W*X t +U+H t-1 +b) Among them, H t represents the hidden state; σ represents the Sigmoid function in the activation function; * represents the convolution operation; X t Indicates input; H t-1 represents the hidden state at the previous moment; b represents the bias term; W and U represent the weights of each gate; The decoder includes a DeConvLSTM unit for decoding hidden features into the original data space and reconstructing data, and a linear layer for outputting the final reconstruction result.

5. The method according to claim 1, characterized in that In the step (1.3), the encoder of the ConvLSTM-AE model obtains hidden features in the following way: (a) The ConvLSTM network is used as the encoder of the AE model, and the convolutional layer Conv is used instead of the fully connected layer to extract data features; the specific construction formula is as follows: E=U⊙h t-1 +W⊙x t Where E represents the output of the current time step; ⊙ represents the convolution operation; h t-1 Represents the output of the previous ConvLSTM unit, x t represents the input data at time t, W and U represent the weights of each gate; (b) Use the LSTM unit in the ConvLSTM network to extract the time feature. The specific construction formula is as follows: i t =sigmoid(W i ·[h t-1 ,x t ]+b i ) f t =sigmoid(W f ·[h t-1 ,x t ]+b f ) o t =sigmoid(W o ·[h t-1 ,x t ]+b o ) h t =o t ×tanh(c t ) Among them, i t 、f t , o t , c t 、h t Represents the corresponding input gate, forget gate, output gate, input node, current memory cell state and output state; W i , W f , W o , W c represents the recursive connection weights of the corresponding thresholds, and represents the weight matrices of the input gate, forget gate, output gate, and candidate memory unit states; sigmoid() and tanh() are two activation functions; b i , b f , b o , b c 、c t-1 Represent the bias items of the input gate, forget gate, output gate and candidate memory unit states respectively; (c) The final result of the ConvLSTM network is determined by the following formula: o t =sigmoid(W o ·[h t-1 ,x t ]+b o ) h t =o t ×tanh(c t ) The final result is the hidden features extracted by the encoder.

6. The method according to claim 1, characterized in that In the step (1.3), the decoder of the ConvLSTM-AE model obtains the output result in the following manner: (d) The decoder is composed of DeConvLSTM layers and receives the hidden features extracted by the encoder; h t Reconstructed as the input of ConvLSTM-AE; The DeConvLSTM layer is composed of several DeConvLSTM units, and uses deconvolution operation to extract features; the specific construction formula is as follows: Where D represents the output of the current time step; Represents the weight of the DeConvLSTM gate; represents the deconvolution operation; represents the output of the encoder; Represents the output of DeConvLSTM at the current moment; (e) The gate calculation of DeConvLSTM is performed in the same way as the encoder, and the output result is expressed as: in, represents the output of the output gate; Indicates the current memory unit status; Indicates the current hidden state; (f) The decoder reconstructs the information with its last linear layer, and the final output is expressed as: in, represents the final output result, the output of the nth sample corresponding to the time step t; σ represents the Sigmod activation function operation; W y Represents the weight matrix, which is used to linearly transform the hidden state of the current time step; represents the hidden state of the current time step; b y Represents the bias term.

7. The method according to claim 1, characterized in that In the step (1.3), minimizing the residual between the output and the output is used as the training objective of the ConvLSTM-AE model; The residual is used to reflect the difference between the input and the reconstructed output, reflecting the degree of adaptation of the model to the input data; its specific construction formula is as follows: Among them, R is the residual; X is the input data, is the reconstructed output of the model; If the residual exceeds the set threshold, it is considered that an abnormal operating condition exists.

8. The method according to claim 1, wherein in step (1.4), the statistic T is calculated according to the following formula: 2 And the statistic SPE: SPE=R 2 in, H T represents the transpose of the matrix H; Representative Matrix The inverse matrix of ; H represents hidden features; R is the residual.

9. The method according to claim 1, characterized in that: In the step (1.4), the kernel density estimation method is used to calculate the statistic T 2 And the threshold of the statistic SPE: (i) Calculate the statistic T 2 Threshold: The statistic T 2 The probability density function is denoted as p(T 2 ), as follows: in, Representation Statistics The probability density function of ; N represents the number of samples; K is the kernel function, μ is its kernel width; T 2 (j) is the T of the sample under normal working conditions 2 value, j = 1, 2, ..., N; T 2 is a statistic that measures the difference between the data and the model; T is calculated under normal conditions. 2 The value of The statistics The threshold value is denoted as J(T 2 ), set a confidence level as α, and obtain J(T 2 ) value: (j) Calculate the threshold of the statistic SPE, denoted as J(SPE); The probability density function of the statistic SPE is recorded as p(SPE), as shown below: Where p(SPE) represents the probability density function of the statistic SPE; N represents the number of samples; K is the kernel function, μ is its kernel width; SPE(j) is the SPE value of the jth sample under normal conditions, j = 1, 2, ..., N; The threshold of the statistic SPE is recorded as J(SPE), and a confidence level is set as α. The value of J(SPE) is obtained by solving the following equation: The kernel function K is defined as follows: Among them, K(g) represents the value of kernel function K; g represents the input variable.

10. A computer device, characterized in that: include: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor executes the steel production system fault diagnosis method based on ConvLSTM-AE according to any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the steel production system fault diagnosis method based on ConvLSTM-AE according to any one of claims 1 to 8.

Citation Information

Cited By

  • Industrial process fault detection method based on space-time causal graph auto-encoder

    CN120974245A