Complex equipment anomaly detection method based on dual-time convolutional network
By introducing a dual-time convolutional network and an adaptive threshold mechanism in complex equipment anomaly detection, the problem of low accuracy in predicting anomaly detection in the prior art is solved, and higher detection accuracy and system reliability are achieved.
Patent Information
- Application Number
- CN202510112676.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-01-24
AI Technical Summary
In the prior art, the accuracy of predicted anomaly detection of complex equipment is low, mainly due to insufficient static threshold settings based on global prediction deviations, and the problem of uneven distribution of training data cannot be effectively handled.
A complex equipment anomaly detection method based on dual-time convolution network (TCN) is adopted, and the time series prediction network (P network) and prediction deviation estimation network (R network) are combined, and the basic threshold is set using extreme value theory (EVT), and adaptive adjustment is performed through the R network, and the threshold is dynamically adjusted to adapt to the difference in prediction deviations of different inputs.
It improves the real-time and accuracy of abnormal detection of complex equipment, can timely identify early potential failures of equipment, enhances the reliability and operating efficiency of the system, overcomes the limitations of static thresholds.
Smart Images

Figure CN120046071A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of equipment health management, and particularly relates to a complex equipment anomaly detection method based on a dual temporal convolutional network. Background Art
[0002] Industrial complex equipment integrates mechanical, hydraulic, electrical, magnetic, and thermal components, and processes a large amount of multi-dimensional operation data, which has the characteristics of high volume, complexity, and interdependence. Anomaly detection plays a crucial role in monitoring the operating state of equipment, and it is achieved by identifying data deviations of potential faults in signals. Fast and accurate anomaly detection can perform timely maintenance to ensure the reliability and efficiency of the system. Data-driven technologies based on machine learning and deep learning have effectively replaced the traditional threshold method based on expert experience. Among them, machine learning technologies such as clustering and classification often have difficulty dealing with high-dimensional and complex data, which limits their applicability; in contrast, deep learning methods are good at mining complex patterns in large-scale high-dimensional data. Currently, typically, the anomaly detection method based on prediction realizes anomaly detection by evaluating the deviation between the predicted data points and the actual data points.
[0003] In current similar anomaly detection methods based on prediction, usually, the threshold of the prediction deviation is set according to the overall prediction error of the prediction model in the known normal training data. Existing methods usually assume that the tolerance of all input prediction deviations is the same, that is, a fixed threshold is set for the prediction deviation. However, the training objective of the prediction model is to minimize the global error. Then, for data patterns that appear more frequently in the training data, the learning effect will be better and the prediction will be more accurate. Therefore, due to the uneven distribution of the training data itself, the trained model performs differently on different inputs. Thus, setting the prediction deviation threshold for all samples based on the global prediction deviation of the training data has deficiencies. Summary of the Invention
[0004] In view of the above problems, the present invention provides a complex equipment anomaly detection method based on a dual temporal convolutional network, which solves the technical problem of low accuracy of complex equipment prediction anomaly detection in the prior art.
[0005] The present invention provides a complex equipment anomaly detection method based on a dual temporal convolutional network, including the following steps:
[0006] Step S1: Install sensors on the complex equipment, and collect the historical operation data of the complex equipment through the sensors as training data. The training data includes a plurality of historical samples arranged in chronological order, and the historical samples include multiple operation parameter values of the complex equipment;
[0007] Step S2: Determine the network structure of the time series prediction network, and train the time series prediction network based on the training data to obtain a trained time series prediction network;
[0008] Step S3: Process the training data by the trained time series prediction network to obtain prediction deviation data, and determine a prediction deviation base threshold based on the prediction deviation data;
[0009] Step S4: Determine the network structure of the prediction deviation estimation network, and train the prediction deviation estimation network based on the training data and the prediction deviation data to obtain a trained prediction deviation estimation network;
[0010] Step S5: Collect online operation data of the complex equipment through sensors; the online operation data includes online samples at the current moment and online samples at multiple moments before the current moment, and the online samples include multiple operation parameter values of the complex equipment;
[0011] Process the online operation data by using the trained time series prediction network and the trained prediction deviation estimation network to obtain the actual prediction deviation at a future moment, and determine an adaptive threshold based on the prediction deviation base threshold;
[0012] Compare the actual prediction deviation at the future moment with the adaptive threshold to determine whether the operation state of the complex equipment at the future moment is abnormal.
[0013] Preferably, the step S2 specifically includes:
[0014] Step S2-1: Determine the network structure of the time series prediction network as follows: Use multiple historical samples arranged in chronological order in the training data as multi-dimensional time series data as the input layer, successively pass through 3 stacked TCN blocks and a fully connected layer, and finally pass through the output layer to obtain the predicted operation parameter value output by the time series prediction network;
[0015] Step S2-2: Use the operation parameters of the past 50 historical samples in the training data as the network input, and the operation parameters of the next historical sample as the supervision target to determine the loss function loss of the time series prediction network P , and minimize loss P as the optimization target to complete the training of the time series prediction network to obtain a trained time series prediction network.
[0016] Preferably, the expression of the loss function loss of the time series prediction network P is:
[0017]
[0018] Among them, y i , are respectively the supervised target value and the predicted value of the operating parameters of the i-th historical sample, and N is the number of training samples.
[0019] Preferably, the step S3 specifically includes:
[0020] Step S3-1: Calculate the prediction deviation of each historical sample in the training data by using the trained time series prediction network as the prediction deviation data, and the calculation expression is:
[0021]
[0022] Among them, r i represents the prediction deviation of the i-th historical sample in the training data.
[0023] Step S3-2: According to the prediction deviation data of the training data, use the extreme value theory method to set the basic threshold for the prediction deviation, and the expression is:
[0024]
[0025] Among them, z q is the basic threshold of the prediction deviation, t represents the preset peak value, q is the target probability parameter, N is the number of training samples, and N t is the number of samples in the sample that exceed the peak value t, and γ and σ are respectively the shape parameter and scale parameter of the Pareto distribution, are respectively the estimated values of γ and σ.
[0026] Preferably, the step S4 specifically includes:
[0027] Step S4-1: Determine the network structure of the prediction deviation estimation network as follows: Use multiple historical samples arranged in chronological order in the training data as multi-dimensional time series data as the input layer, and pass through 3 stacked TCN blocks and a fully connected layer in sequence, and finally obtain the predicted deviation estimation value output by the prediction deviation estimation network through the output layer;
[0028] Step S4-2: Use the operating parameters of the past 50 historical samples in the training data as the network input, and use the prediction deviation corresponding to the next historical sample in the prediction deviation data as the supervised target to determine the loss function loss R of the prediction deviation estimation network, and use minimizing loss R as the optimization goal to complete the training of the time series prediction network, and obtain the trained time series prediction network.
[0029] Preferably, the loss function loss of the prediction deviation estimation networkR The expression is:
[0030]
[0031] where r i represents the prediction deviation of the i-th historical sample as the supervision target, represents the estimated value of the prediction deviation of the i-th historical sample.
[0032] Preferably, the step S5 specifically includes:
[0033] Step S5-1: Collect the online operation data of the complex equipment through sensors. The online operation data includes the online sample at the current moment and the online samples at multiple moments before the current moment. The online sample includes multiple operation parameter values of the complex equipment;
[0034] Step S5-2: Process the online operation data by using the trained time series prediction network to obtain the predicted values of the operation parameters at future moments; process the online operation data by using the trained prediction deviation estimation network to obtain the estimated values of the prediction deviations at future moments;
[0035] Step S5-3-1: Obtain the actual prediction deviation at future moments based on the predicted values of the operation parameters at future moments;
[0036] Step S5-3-2: Determine the adaptive threshold from the estimated value of the prediction deviation at future moments, the actual prediction deviation at future moments, and the basic threshold of the prediction deviation;
[0037] Step S5-3-3: Compare the actual prediction deviation at future moments with the adaptive threshold to determine whether the operation state of the complex equipment at future moments is abnormal.
[0038] Preferably, the step S5-3-1 specifically includes:
[0039] Calculate the absolute value of the difference between the predicted value of the operation parameter at future moments and the true value of the operation parameter at future moments as the actual prediction deviation at future moments;
[0040] The adaptive threshold th ad + in the step S5-3-2 is calculated by the following expression:
[0041]
[0042] where is the estimated value of the prediction deviation at future moments, The prediction deviation of the i-th historical sample in the training data, N is the number of training samples, and r represents the actual prediction deviation at a future time.
[0043] The step S5-3-3 specifically includes: if the actual prediction deviation exceeds the adaptive threshold, it is detected as an abnormal state; otherwise, it is normal.
[0044] Preferably, the TCN block includes a causal convolutional layer, a dilated convolutional layer, and a residual connection output layer connected in sequence, where
[0045] The calculation expression of the causal convolutional layer is:
[0046]
[0047] where x T represents the T-th value of the input time series of the causal convolutional layer, F(x T ) represents the causal convolution value corresponding to x T , f k represents the k-th value in the convolutional kernel, K represents the length of the convolutional kernel, and x T-K+k represents the (T-K+k)-th value of the input time series of the causal convolutional layer;
[0048] The calculation expression of the dilated convolutional layer is:
[0049]
[0050] where Fd(x T ) represents the dilated convolution value corresponding to x T , x T-(K-k)d represents the (T-(K-k)d)-th value of the input time series of the causal convolutional layer, d represents the dilation coefficient, and c is the number of network layers of the dilated convolutional layer;
[0051] The calculation expression of the residual connection output layer is:
[0052] F(x T ) = F d (x T ) + x T
[0053] where F(x T ) represents the residual connection output value corresponding to x T .
[0054] Compared with the prior art, the present invention has at least the following beneficial effects:
[0055] (1) By introducing the Dual Temporal Convolutional Network (TCN), the present invention effectively improves the real-time performance and accuracy of anomaly detection for complex equipment. First of all, compared with other deep learning algorithms, TCN has higher computational efficiency and can process a large amount of online operation data in real time, ensuring that during the operation of complex equipment, anomalies can be detected and responded to in a timely manner, and the safe and stable operation of the system is guaranteed.
[0056] (2) The present invention introduces an adaptive threshold mechanism, which overcomes the limitations of static thresholds. By analyzing the differences in prediction deviations between different inputs, the adaptive threshold can be dynamically adjusted to effectively manage the tolerance of prediction deviations caused by uneven distribution of training data, improving the accuracy of anomaly detection and enabling the system to maintain high detection performance in the face of a changing operating environment.
[0057] (3) By combining the P network for time series prediction and the R network for deviation estimation, and using the Extreme Value Theory (EVT) to set the basic threshold and perform adaptive adjustment through the R network, the present invention achieves high-precision anomaly detection. It can identify early potential faults of complex equipment in a timely manner, and also provides strong guarantees for the safety and stability of the equipment, improving the reliability and operating efficiency of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] The drawings are only for the purpose of illustrating specific embodiments and are not considered to be a limitation of the present invention.
[0059] Figure 1 It is a flowchart of the method for anomaly detection of complex equipment based on the Dual Temporal Convolutional Network provided by the present invention;
[0060] Figure 2 It is a schematic structural diagram of the method for anomaly detection of complex equipment based on the Dual Temporal Convolutional Network provided by the present invention.
[0061] Figure 3 It is a schematic diagram of the calculation of the Temporal Convolutional Network.
[0062] Figure 4 It is a structural diagram of the sequence prediction network and the prediction deviation estimation network provided by the present invention.
[0063] Figure 5 It is a schematic diagram of the construction of training data for the sequence prediction network and the prediction deviation estimation network provided by the present invention.
[0064] Figure 6 It is a schematic diagram of the basic threshold based on EVT of the present invention.
[0065] Figure 7 It is a schematic diagram of the anomaly detection threshold of a certain equipment provided by the present invention.
[0066] Figure 8Schematic diagram of the abnormal detection results of a certain equipment provided by the present invention. Detailed implementation manners
[0067] In order to more clearly understand the above objects, features and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners. It should be noted that, without conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other. In addition, the present invention may also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited by the specific embodiments disclosed below.
[0068] In order to illustrate the effectiveness of the method proposed by the present invention, the above technical solutions of the present invention will be described in detail through a specific embodiment as follows. Figure 1 、 Figure 2 As shown in the figure, a complex equipment abnormal detection method based on a dual-time convolutional network is disclosed, and the specific implementation steps are as follows:
[0069] Step S1: Install sensors on the complex equipment, and collect the historical operation data of the complex equipment through the sensors as training data. The training data includes a plurality of historical samples arranged in chronological order, and the historical samples include a plurality of operation parameter values of the complex equipment.
[0070] The present invention accumulates a large amount of normal state data in the historical operation process of the complex equipment as the training data of the neural network. Specifically, different types of sensors can be installed on the complex equipment, and the sensors can continuously collect various operation parameters in the historical operation process of the equipment during normal operation. The operation parameters may include physical quantities such as temperature, pressure, vibration, current, and rotational speed. The sampling frequency for collecting the operation parameters can be set according to actual working conditions requirements, such as sampling once per second or once per minute, etc.
[0071] In some embodiments, the total time span of collecting the operation parameters of the complex equipment can be not less than 3 months to ensure that enough operation condition data is collected. After the collected data is preprocessed, including operations such as outlier removal and data normalization, the historical data in the normal operation state is saved in the database as the training data for subsequent neural network training.
[0072] In some embodiments, the training data of the present invention is stored and processed in the form of time series data. Specifically, the training data includes a plurality of historical samples arranged in chronological order. Each historical sample includes a timestamp and measurement values of corresponding different sensors. The data structure can be: the first column is the timestamp, and the subsequent columns are the measurement values of different sensors, all represented by floating-point numbers. For example, for the case of n sensors, the format of each historical sample record is: [timestamp, value_1, value_2,..., value_n]. Among them, value_1 to value_n respectively represent the measurement values of multiple different operating parameters of different sensors at this timestamp.
[0073] Step S2: Determine the network structure of the time series prediction network, and train the time series prediction network based on the training data to obtain a trained time series prediction network.
[0074] The time series prediction network of the present invention is established based on the Temporal Convolutional Network (TCN). As Figure 3 shown in the calculation schematic diagram of the time convolutional network. The parameters that need to be preset in the TCN network include the total number of network layers L and the convolutional kernel size. The calculation of each layer includes three components: causal convolution, dilated convolution, and residual connection, specifically as follows:
[0075] (1) Causal convolution
[0076] As Figure 3 shown, for the input of each layer, the time series data X = [x 1 , x 2 , L, x T with a length of T, and the convolutional kernel F = [f 1 , f 2 , L, f K with a length of K, for the data at each moment t, only the data at moment t and its previous moments are used for convolution operation. The calculation expression of the causal convolution value is:
[0077]
[0078] Among them, x T represents the T-th value of the input time series of the causal convolution layer, F c (x T ) represents the causal convolution value corresponding to x T , f k represents the k-th value in the convolutional kernel, K represents the length of the convolutional kernel, and x T-K+k represents the (T - K + k)-th value of the input time series of the causal convolution layer.
[0079] (2) Dilated Convolution
[0080] The TCN network is improved based on the above causal network, such as Figure 2 shown, for the time series data input at each layer, dilated convolution with a dilation coefficient d = 2 c is used for processing, where c is the number of network layers. Specifically, the dilated convolution of the c-th layer inserts d - 1 holes between adjacent inputs, enabling the convolutional kernel to cover a larger range of historical data. For example, when c = 1 and d = 2, the convolutional kernel samples every other data point in the input sequence; when c = 2 and d = 4, it samples every three data points. This processing method significantly increases the receptive field of the network, enabling it to capture longer-term temporal dependencies.
[0081] Therefore, the expression for the dilated causal convolution of the c-th layer is:
[0082]
[0083] d = 2 c
[0084] where x T represents the T-th value of the input time series of the dilated convolutional layer, F d (x T ) represents the dilated convolution value corresponding to x T , x T-(K-k)d represents the (T - (K - k)d)-th value of the input time series of the causal convolutional layer, d represents the dilation coefficient, and c is the number of network layers of the dilated convolutional layer.
[0085] (3) Residual Connection
[0086] The residual connection of TCN directly passes the original input to the output end through a shortcut connection and performs an element-wise addition operation with the features after a series of transformations. This processing method effectively alleviates the problem of gradient disappearance during the training of deep networks and improves the training effect of the network. Specifically, the input time series values of the causal convolutional layer and the features after dilated convolution processing are added to obtain the final output, and the expression is:
[0087] F(x T ) = F d (x T ) + x T
[0088] where F(x T ) represents the output value of the residual connection corresponding to x T .
[0089] The network composed of the above three parts of causal convolution, dilated convolution, and residual connection is also called a TCN block. The network structure of the time series prediction network (also called the P network) of the present invention is as Figure 4 shown. Multiple historical samples arranged in chronological order in the training data are used as multi-dimensional time series data as the input layer, and sequentially pass through 3 stacked TCN blocks and a fully connected layer, and finally obtain the predicted value of the operating parameters output by the time series prediction network.
[0090] The steps of training the time series prediction network include: using the operating parameters of the past 50 historical samples in the training data as the network input, and the operating parameters of the next historical sample as the supervision target. Determine the loss function for training the time series prediction network. The loss function loss of the P network P is the root mean square error of the predicted value, and the expression is:
[0091]
[0092] where y i , are respectively the supervision target value and the predicted value of the operating parameters of the i-th historical sample, and N is the number of training samples.
[0093] In some embodiments, an error backpropagation algorithm such as the gradient descent method can be used to minimize the loss function loss of the P network P as the optimization target to complete the training of the time series prediction network, and obtain the trained time series prediction network. There are relatively mature training optimization methods for various neural networks in the prior art, and the present invention does not limit the specific optimization method adopted for neural network training.
[0094] Step S3: Process the training data by the trained time series prediction network to obtain prediction deviation data, and determine the prediction deviation base threshold based on the prediction deviation data.
[0095] In this step, as Figure 5 shown, use the trained time series prediction network to calculate the prediction deviation of each historical sample in the training data. The prediction deviation r of the i-th historical sample in the training data i has the following expression:
[0096]
[0097] As Figure 6As shown in the figure, based on the prediction deviation data of the training data, the extreme value theory method is used to set the basic threshold for the prediction deviation. The extreme value theory (EVT) believes that the extreme events of different things satisfy the same distribution, that is, the part by which the extreme value exceeds a threshold satisfies the generalized Pareto distribution (GPD), which can be expressed as follows:
[0098]
[0099] Among them, Z represents a certain prediction deviation value, z represents the prediction deviation random variable, and t represents the preset threshold. represents the probability distribution of the extreme value of the prediction deviation exceeding t. represents the conditional probability that the excess amount Z - t is greater than the prediction deviation random variable z under the condition that Z is greater than t. τ represents the peak value. represents that when t is set as a peak value, it follows the Pareto distribution with parameters γ and σ. γ is the shape parameter and σ is the scale parameter.
[0100] Estimated values of parameters γ and σ can be obtained by using the method of moments (MOM). The basic threshold z of the prediction deviation q is calculated as follows:
[0101]
[0102] Among them, q is the target probability parameter, and the value of q is set artificially, usually taking 0.01%. N is the number of training samples, and N t is the number of samples in the sample that exceed the peak value t. are the estimated values of γ and σ respectively.
[0103] Step S4: Determine the network structure of the prediction deviation estimation network, and train the prediction deviation estimation network based on the training data and the prediction deviation data to obtain a trained prediction deviation estimation network.
[0104] As Figure 4 shown, the network structure of the prediction deviation estimation network (also called the R network) of the present invention is the same as that of the time series prediction network (also called the P network).
[0105] Taking multiple historical samples arranged in chronological order in the training data as multi-dimensional time series data as input, passing through the input layer, 3 stacked TCN blocks, and the fully connected layer in sequence, and finally obtaining the prediction deviation estimation value output by the time series prediction network.
[0106] The steps of training the prediction deviation estimation network include: using the operation parameter data of the past 50 historical samples in the training data as the network input, and using the prediction deviation corresponding to the next historical sample in the prediction deviation data as the supervision target. Determine the loss function for training the prediction deviation estimation network, and the loss function of the R network is loss R is the root mean square error of the prediction deviation, and the expression is:
[0107]
[0108] where r i represents the prediction deviation of the i-th historical sample as the supervision target, represents the estimated value of the prediction deviation of the i-th historical sample.
[0109] In some embodiments, similarly, error backpropagation algorithms such as the gradient descent method can be used to minimize the loss function loss of the R network R as the optimization target to complete the training of the prediction deviation estimation network and obtain the trained prediction deviation estimation network. There are relatively mature training optimization methods for various neural networks in the prior art, and the specific optimization method adopted by the present invention for neural network training is not limited.
[0110] Step S5-1: Collect the online operation data of the complex equipment through sensors. The online operation data includes the online sample at the current moment and the online samples at multiple moments before the current moment. The online sample includes multiple operation parameter values of the complex equipment.
[0111] After the P network and the R network are trained, the online operation data of the complex equipment can be detected. The "online" here specifically refers to the real-time nature of the data, that is, the acquisition, transmission, and processing of the data are all carried out in real time during the operation of the equipment, rather than offline analysis afterwards.
[0112] The acquisition method of the online operation data of the present invention is similar to that of the historical operation data, and various operation parameters during the operation of the equipment are also continuously collected through sensors. Specifically, the operation parameter values at the current moment are saved, and the operation parameter values at multiple moments before the current moment are also saved. In some embodiments, the operation parameter values of the latest 50 online samples before the current moment can be saved.
[0113] Similar to the historical samples of the present invention, the online samples of the present invention also include a time stamp and corresponding measurement values of different sensors. The data structure can be: the first column is the time stamp, and the subsequent columns are the measurement values of different sensors.
[0114] Step S5-2: Process the online operation data using the trained time series prediction network to obtain the predicted values of the operation parameters at future times;
[0115] Process the online operation data using the trained prediction deviation estimation network to obtain the predicted deviation estimation values at future times.
[0116] In some embodiments, the latest 50 online samples can be input into the trained time series prediction network to obtain the predicted values of the operation parameters at future times; the latest 50 online samples can be input into the trained prediction deviation estimation network to obtain the predicted deviation estimation values at future times.
[0117] Step S5-3-1: Obtain the actual predicted deviation at future times based on the predicted values of the operation parameters at future times;
[0118] Step S5-3-2: Determine the adaptive threshold from the predicted deviation estimation value at future times, the actual predicted deviation at future times, and the basic threshold of the predicted deviation;
[0119] Step S5-3-3: Compare the actual predicted deviation at future times with the adaptive threshold to determine whether the operation state of the complex equipment at future times is abnormal.
[0120] The adaptive threshold of the present invention is dynamically and adaptively adjusted based on the predicted deviation estimation value output by the R network on the basis of the basic threshold of the predicted deviation. The unoptimized adaptive threshold th ad The expression is:
[0121]
[0122] Where is the estimated value of the predicted deviation at future times, is the predicted deviation of the i-th historical sample in the training data, and N is the total number of training samples.
[0123] In this step, the actual predicted deviation at future times is obtained in a manner similar to the processing process of historical samples. The calculation method of the actual predicted deviation at future times is: calculate the absolute value of the difference between the predicted value of the operation parameter at future times obtained by the trained time series prediction network (P network) based on the online operation data and the true value of the operation parameter at future times as the actual predicted deviation at future times.
[0124] In addition to dynamic threshold adjustment, predicted deviations significantly exceeding the basic threshold can be directly classified as abnormal. To optimize the calculation of the adaptive threshold, when the actual predicted deviation r satisfies r > 2*z qIf the prediction deviation is considered to be significantly greater than the basic threshold, then the final adaptive threshold th ad + The expression is:
[0125]
[0126] where r represents the actual prediction deviation at a future time.
[0127] As Figure 7 shown is a schematic diagram of the prediction deviation and the deviation threshold during the abnormal detection of a certain equipment, including the actual prediction deviation, the estimated value of the prediction deviation output by the R network, the basic threshold based on EVT, and the adaptive dynamic threshold of the present invention.
[0128] Finally, compare the actual prediction deviation with the adaptive threshold. If the actual prediction deviation exceeds the adaptive threshold, it is detected as an abnormal state; otherwise, it is normal.
[0129] As Figure 8 shown is a schematic diagram of the abnormal detection result of a certain equipment, and the tiny abnormalities in the data are accurately detected.
[0130] Through the complex equipment abnormal detection method based on the dual temporal convolutional network provided by the present invention, it is possible to effectively make up for the deficiencies of the existing deep learning-based abnormal detection methods and further improve the accuracy of abnormal detection. The TCN network with higher computational efficiency is used to establish the P network for time series prediction and the R network for deviation estimation to ensure the real-time performance of abnormal detection; the present invention introduces an adaptive threshold mechanism that considers the differences in prediction deviations between different inputs. By combining the basic threshold based on EVT with the adaptive adjustment of the R network, the accuracy of abnormal detection is improved, and the safety and stability of the operation of complex equipment are enhanced.
[0131] Although the specific embodiments of the present invention depict various actions or steps in a particular order, it should be understood that such actions or steps are required to be performed in the specific order shown or in a sequential order, or that all of the illustrated actions or steps should be performed to achieve the desired result. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented in combination in a single implementation. Conversely, the various features described in the context of a single implementation may also be implemented separately or in any suitable sub-combination in multiple implementations. As described above, only the preferred specific embodiments of the present invention are provided, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention.
[0132] As described above, only the preferred specific embodiments of the present invention are provided, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention.
Claims
1. A complex equipment anomaly detection method based on dual-time convolutional network, characterized in that: The following steps are involved: Step S1, installing a sensor on the complex equipment, and collecting historical operation data of the complex equipment as training data through the sensor, wherein the training data includes a plurality of historical samples arranged in chronological order, and the historical samples include a plurality of operation parameter values of the complex equipment; Step S2, determining the network structure of the time series prediction network, and training the time series prediction network based on the training data to obtain a trained time series prediction network; Step S3, the trained time series prediction network processes the training data to obtain prediction deviation data, and determines a prediction deviation basic threshold based on the prediction deviation data; Step S4, determining the network structure of the prediction deviation estimation network, and training the prediction deviation estimation network based on the training data and the prediction deviation data to obtain a trained prediction deviation estimation network; Step S5, collecting online operation data of the complex equipment through sensors; the online operation data includes online samples at the current moment and online samples at multiple moments before the current moment, and the online samples include multiple operation parameter values of the complex equipment; The online operation data is processed using the trained time series prediction network and the trained prediction deviation estimation network to obtain an actual prediction deviation at a future moment, and an adaptive threshold is determined by the prediction deviation basic threshold; The actual prediction deviation at the future moment is compared with the adaptive threshold to determine whether the operating state of the complex equipment at the future moment is abnormal.
2. The complex equipment anomaly detection method based on dual-time convolutional network according to claim 1 is characterized in that: The step S2 specifically includes: Step S2-1, determining the network structure of the time series prediction network as follows: taking multiple historical samples arranged in chronological order in the training data as multidimensional time series data as the input layer, sequentially passing through 3 layers of stacked TCN blocks and a fully connected layer, and finally passing through the output layer to obtain the operating parameter prediction value output by the time series prediction network; Step S2-2: Using the operating parameters of the past 50 historical samples in the training data as network input and the operating parameters of the next historical sample as the supervision target, determine the loss function of the time series prediction network. P , to minimize the loss P The training of the time series prediction network is completed as an optimization target to obtain a trained time series prediction network.
3. The complex equipment anomaly detection method based on dual-time convolutional network according to claim 2 is characterized in that: The loss function of the time series prediction network is loss P The expression is: Among them, y i , are the supervised target value and predicted value of the operating parameters of the i-th historical sample respectively, and N is the number of training samples.
4. The complex equipment anomaly detection method based on dual-time convolutional network according to claim 3 is characterized in that: The step S3 specifically includes: Step S3-1, using the trained time series prediction network to calculate the prediction deviation of each historical sample in the training data as the prediction deviation data, the calculation expression is: Among them, r i Represents the prediction deviation of the i-th historical sample in the training data; Step S3-2: According to the prediction deviation data of the training data, a basic threshold value is set for the prediction deviation using the extreme value theory method, and the expression is: Among them, z q is the prediction deviation basic threshold, t represents the preset peak value, q is the target probability parameter, N is the number of training samples, and N t is the number of samples that exceed the peak value t in the sample, γ and σ are the shape parameter and scale parameter of the Pareto distribution respectively, are the estimated values of γ and σ respectively.
5. The complex equipment anomaly detection method based on dual-time convolutional network according to claim 4 is characterized in that: The step S4 specifically includes: Step S4-1, determining the network structure of the prediction deviation estimation network as follows: taking multiple historical samples arranged in chronological order in the training data as multidimensional time series data as the input layer, sequentially passing through 3 layers of stacked TCN blocks and a fully connected layer, and finally passing through the output layer to obtain the prediction deviation estimation value output by the prediction deviation estimation network; Step S4-2: Using the operating parameters of the past 50 historical samples in the training data as network input, using the prediction deviation corresponding to the next historical sample in the prediction deviation data as the supervision target, and determining the loss function loss of the prediction deviation estimation network R , to minimize the loss R The training of the time series prediction network is completed as an optimization target to obtain a trained time series prediction network.
6. The complex equipment anomaly detection method based on dual-time convolutional network according to claim 5 is characterized in that: The prediction deviation estimates the loss function of the network loss R The expression is: Among them, r i represents the prediction deviation of the i-th historical sample as the supervision target, Represents the estimated value of the forecast deviation of the i-th historical sample.
7. The complex equipment anomaly detection method based on dual-time convolutional network according to claim 6 is characterized in that: The step S5 specifically includes: Step S5-1, collecting online operation data of complex equipment through sensors, wherein the online operation data includes online samples at the current moment and online samples at multiple moments before the current moment, and the online samples include multiple operation parameter values of the complex equipment; Step S5-2, using the trained time series prediction network to process the online operation data to obtain the predicted value of the operation parameter at the future time; using the trained prediction deviation estimation network to process the online operation data to obtain the predicted deviation estimated value at the future time; Step S5-3-1, obtaining the actual prediction deviation at the future moment based on the predicted value of the operating parameter at the future moment; Step S5-3-2, determining an adaptive threshold value based on the prediction deviation estimate value at the future moment, the actual prediction deviation at the future moment, and the prediction deviation basic threshold value; Step S5-3-3: compare the actual predicted deviation at the future moment with the adaptive threshold to determine whether the operating status of the complex equipment at the future moment is abnormal.
8. The complex equipment anomaly detection method based on dual-time convolutional network according to claim 7 is characterized in that: The step S5-3-1 specifically includes: Calculating the absolute value of the difference between the predicted value of the operating parameter at the future moment and the true value of the operating parameter at the future moment as the actual prediction deviation at the future moment; The adaptive threshold th in step S5-3-2 ad + The calculation expression is: in, is the estimated value of the forecast deviation at the future moment, is the prediction deviation of the i-th historical sample in the training data, N is the number of training samples, and r represents the actual prediction deviation at the future moment; The step S5-3-3 specifically includes: if the actual prediction deviation exceeds the adaptive threshold, it is detected as an abnormal state, otherwise it is normal.
9. The complex equipment anomaly detection method based on dual-time convolutional network according to claim 8 is characterized in that: The TCN block includes a causal convolutional layer, a dilated convolutional layer, and a residual connection output layer connected in sequence, wherein: The computational expression of the causal convolutional layer is: Among them, x T represents the Tth value of the input time series of the causal convolutional layer, F(x T ) represents x T The corresponding causal convolution value, f k represents the kth value in the convolution kernel, K represents the length of the convolution kernel, x T-K+k Represents the T-K+kth value of the input time series of the causal convolutional layer; The calculation expression of the dilated convolutional layer is: d=2 c Among them, Fd(x T ) represents x T The corresponding dilated convolution value, x T-(K-k)d represents the T-(Kk)dth value of the input time series of the causal convolutional layer, d represents the dilation coefficient, and c is the number of network layers of the dilated convolutional layer; The calculation expression of the residual connection output layer is: F(x T )=F d (x T )+x T Among them, F(x T ) represents x T The corresponding residual connection output value.
Citation Information
Patent Citations
Transformer DGA data prediction method based on multi-dimensional time sequence frame convolution LSTM
CN110674604A
Fault diagnosis method based on AM-TCN
CN114297921A
Industrial control system-oriented anomaly detection system and method
CN115484102A
Aircraft telemetry data anomaly detection method and system based on time convolutional network
CN115510950A
Industrial internet time series data anomaly detection method and device
CN116522265A