Industrial time series data anomaly detection method based on Skip-Range attention
The Skip-Range Transformer method converts multi-attribute time series data into a two-dimensional waveform for enhanced feature extraction, improving anomaly detection accuracy by capturing complex interdependencies and robustness to noise.
Patent Information
- Application Number
- CN202510445871.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-15
AI Technical Summary
The prior art is difficult to effectively explore the complex dependencies between different attributes in industrial timing data, resulting in reduced efficiency of the model in multi-attribute data analysis.
Using the Skip-Range attention method, the data sequence is reconstructed and the error threshold is judged to detect abnormalities by converting multi-attribute timing data into two-dimensional waveform graphs and extracting image features.
It improves the modeling ability of complex timing relationships, enhances the accuracy of anomaly detection, reduces the computational complexity, and enhances the focus ability of key segments, and has fault tolerance and timing coherence.
Smart Images

Figure CN120316680A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of industrial time-series data anomaly detection, and in particular to an industrial time-series data anomaly detection method based on Skip-Range attention. Background Art
[0002] Industrial time-series data contains multiple attributes, such as temperature, pressure, vibration, current, voltage, etc. Each attribute exhibits different numerical changes at different time points, and complex interdependent relationships exist among these features. For example, an increase in temperature may cause a corresponding increase in pressure, and abnormal vibration patterns are usually related to irregular current fluctuations.
[0003] For example, Chinese Patent CN119167244A discloses an industrial time-series data anomaly detection method and system based on data augmentation. The method includes: collecting industrial time-series data in the industrial production process; combining data interpolation methods and data preprocessing methods to perform data augmentation on the industrial time-series data; using a 1D single-layer convolutional neural network model with a time attention mechanism to perform anomaly detection on the data-augmented industrial time-series data.
[0004] The above method enhances the original collected sequence by interpolation and uses the enhanced data for anomaly detection, unable to uncover the complex interdependent relationships between data of different attributes.
[0005] The richness of features and the complexity of their relationships in multi-attribute data pose significant challenges to algorithms in dealing with complex interdependent relationships. The algorithm needs to accurately capture and understand the interactions between these features to achieve effective analysis and prediction of industrial time-series data.
[0006] In the field of multi-attribute data analysis, complex interdependent relationships pose great challenges to models in precisely learning this complexity. Relying solely on checking a single attribute or independent analysis of multiple attributes is usually insufficient to accurately detect anomalies, while comprehensive analysis of multiple factors can more effectively identify potential anomalies in the system. Existing methods usually cannot uncover the complex interdependent relationships between different attributes, limiting the model's ability to discover the interactions between multi-attributes and resulting in reduced model efficiency. Summary of the Invention
[0007] The purpose of the present invention is to provide an industrial time-series data anomaly detection method based on Skip-Range attention.
[0008] The purpose of the present invention can be achieved through the following technical solutions:
[0009] An industrial time-series data anomaly detection method based on Skip-Range attention, comprising:
[0010] Step S1: Obtain multi-attribute time series data, and perform normalization on each attribute respectively to obtain the first multi-attribute time series data sequence;
[0011] Step S2: Convert the first multi-attribute time series data sequence into a two-dimensional waveform diagram, where the horizontal axis of the two-dimensional waveform diagram represents time, the vertical axis represents values, the connection lines of data points corresponding to different attributes at the same moment are parallel to the vertical axis, and the vertical axis values increase gradually outward from the origin;
[0012] Step S3: Extract image features based on the two-dimensional waveform diagram;
[0013] Step S4: Use the sequence after one or more Skip-Range transformations of the first multi-attribute time series data sequence as the time series feature;
[0014] Step S5: Input the image feature and the time series feature into the trained fusion decoder to obtain the second multi-attribute time series data sequence reconstructed by the fusion decoder;
[0015] Step S6: Calculate the squared error between the first multi-attribute time series data sequence and the second multi-attribute time series data sequence, and determine whether the squared error is greater than the pre-configured error threshold. If so, it is determined as abnormal.
[0016] In the single time of Step S4, it includes:
[0017] Step S4-1: Adopt the method of interval sampling to obtain multiple first segments by interval sampling of the sequence before transformation;
[0018] Step S4-2: Apply self-attention transformation to all the first segments to obtain the corresponding second segments;
[0019] Step S4-3: Re-organize all the second segments according to the sampling order of their corresponding first segments to obtain the first intermediate sequence;
[0020] Step S4-4: Sequentially split the first intermediate sequence to obtain multiple third segments;
[0021] Step S4-5: Apply self-attention transformation to all the third segments to obtain the corresponding fourth segments;
[0022] Step S4-6: Re-organize all the fourth segments according to the splitting order of their corresponding third segments to obtain the transformed sequence.
[0023] The number of the third segments is the same as the number of the first segments.
[0024] The position of an element at any position in any second segment in the first intermediate sequence is the same as the position of the element at the same position in the first segment corresponding to the second segment in the sequence before transformation;
[0025] The position of an element at any position in any fourth fragment in the transformed sequence is the same as the position of the element at the same position in the third fragment corresponding to this fourth fragment in the first intermediate sequence.
[0026] The multi-attribute time series data at least includes temperature time series data, pressure time series data, and vibration time series data.
[0027] The step S5 includes:
[0028] Step S5-1: Concatenate the image feature and the time series feature to obtain a feature vector, and input the feature vector into the trained fusion decoder;
[0029] Step S5-2: Obtain the second multi-attribute time series data sequence reconstructed by the fusion decoder:
[0030] Z = Norm(F + FFN(Norm(F + MultiHead(F))))
[0031] Where: F is the feature vector, Norm is the normalization function, FFN is the feed-forward neural network, MultiHead is the attention mechanism, and Z is the second multi-attribute time series data sequence reconstructed.
[0032] The loss function in the training process of the fusion decoder is:
[0033]
[0034] Where, is the loss function, Mean is the mean function, and X is the first multi-attribute time series data sequence.
[0035] The step S3 is implemented by using the image encoder ViT.
[0036] An industrial time series data anomaly device based on Skip-Range attention includes a memory, a processor, and a program stored in the memory. When the processor executes the program, the above-mentioned method is implemented.
[0037] A storage medium stores a program, and when the program is executed, the above-mentioned method is implemented.
[0038] Compared with the prior art, the present invention has the following beneficial effects:
[0039] 1. By introducing the DGR method, the normalized multi-attribute time-series data is converted into a two-dimensional waveform diagram, so that data with different attributes are interleaved and arranged in the image area of the same range. Subsequently, image features are obtained by extracting this two-dimensional waveform diagram, enhancing the model's understanding of the interrelationships between multi-attributes from a visual perspective. In addition, by using the Skip-Range method to extract time-series features, both global features are learned through the skip mechanism and local features are learned through the range mechanism, improving the modeling ability for complex time-series relationships, thus effectively improving the accuracy of anomaly detection for industrial time-series data.
[0040] 2. First perform Skip and then Range, so as to balance long-term dependencies and short-term mutations, and it is more suitable for multi-modal data with different frequency characteristics such as temperature, pressure, and vibration. Moreover, in the Skip process, attention is calculated at a coarse-grained interval first, and then refined learning is carried out using Range, which can reduce the computational complexity while enhancing the model's focusing ability on key segments. In addition, the Skip process uses a coarse-grained interval to smooth local noise, while the fine-grained interval analysis of Range can retain sensitive features during key periods, which is more fault-tolerant for sensor noise or data loss and can avoid weight allocation deviation caused by noise when directly performing attention on the original sequence.
[0041] 3. The number of the third segments is the same as the number of the first segments. Thus, through quantity alignment, the weight balance between global and local features is ensured, avoiding a certain scale from dominating the model learning. At the same time, the coherence of the time-series context is also guaranteed, that is, both the noise reduction ability at the coarse-grained level is retained, and the time-series dependence relationship can be restored locally.
[0042] 4. The position of any element at any position in any second segment in the first intermediate sequence is the same as the position of the element at the same position in the first segment corresponding to the second segment in the sequence before transformation. The position of any element at any position in any fourth segment in the transformed sequence is the same as the position of the element at the same position in the third segment corresponding to the fourth segment in the first intermediate sequence. Thus, it is ensured that for the second segment after the first-layer self-attention transformation, the timestamp information of its elements in the original sequence is not lost, and the final sequence after the second-layer transformation can still accurately reflect the physical time order of the original time series.
[0043] 5. The adopted fusion decoder integrates the time dependence relationship in the time-series data and the local morphological features in the image data, enhancing the feature expression ability through multi-modal information complementarity. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 It is a schematic diagram of the main step flow of the method of the present invention;
[0045] Figure 2 It is a schematic diagram of the technical route of this application. Detailed implementation mode
[0046] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and gives detailed implementation manners and specific operation processes, but the protection scope of the present invention is not limited to the following embodiments.
[0047] An industrial time series data anomaly detection method based on Skip-Range attention can be applied to anomaly detection in industrial time series data. This method uses DGR to integrate all attribute data onto a waveform diagram, providing an additional visual perspective for feature extraction. In order to enhance the model's ability to learn time series data, Skip-Range Transformer is proposed. Among them, the skip mechanism is used to enhance the model's ability to learn global features, and the range mechanism is used to learn local features of time series data. By integrating the features extracted from DGR and SRT, the model's understanding of complex multi-attribute time series data is enriched.
[0048] The experimental hardware environment is an Ubuntu20.04 server equipped with an AMD Ryzen 9 5900X CPU and an NVIDIA RTX 3090 GPU. The software environment includes CUDA version 11.3 and cuDNN version 8.2.1. The training and inference of the model code are based on python version 3.7 and PyTorch version 1.11.0.
[0049] As Figure 1 and Figure 2 shown, it includes:
[0050] Step S1: Obtain multi-attribute time series data, and normalize each attribute to obtain a first multi-attribute time series data sequence;
[0051] In this embodiment, the multi-attribute time series data at least includes temperature time series data, pressure time series data, and vibration time series data, which can comprehensively reflect the operating state of the device or system, thereby avoiding the limitations of single-sensor monitoring. And through multi-dimensional data correlation analysis (such as increased vibration accompanied by temperature anomalies), fault modes can be identified more accurately to reduce the false alarm rate. Of course, in other embodiments, other types can also be used.
[0052] Among them, the multi-attribute time series data is collected by corresponding sensors of different types and generally needs to be preprocessed and normalized. The preprocessing process includes: removing the obviously unreasonable parts in the data, such as missing values, abnormal magnitudes, etc.; then calculating the mean and standard deviation of each attribute of the multi-attribute time series data, subtracting the mean from the original data, and then dividing by the standard deviation to obtain the normalized data. Since preprocessing and normalization are conventional processes in this field, they will not be elaborated in this embodiment. However, it should be particularly noted that in this application, data of different types of attributes should be normalized independently to ensure that the curves of the data of any attribute can be located within the same image range in the two-dimensional waveform diagram obtained later. In this way, the image features extracted subsequently can enhance the model's understanding of the mutual relationship between multi-attributes from a visual perspective.
[0053] Step S2: Convert the first multi-attribute time series data sequence into a two-dimensional waveform diagram, where the horizontal axis of the two-dimensional waveform diagram represents time, the vertical axis represents the value, the connection lines of the data points of different attributes corresponding to the same moment are parallel to the vertical axis, and the vertical axis values increase gradually outward from the origin;
[0054] Step S3: Extract image features based on the two-dimensional waveform diagram;
[0055] In this embodiment, the image encoder ViT is specifically used to extract meaningful features from the DGR. The image encoder ViT adopts the Patch Embedding method to decompose the input graph into a series of structured and non-overlapping image patches. The input image is represented as Graph ∈ R H×W×C , where H and W respectively represent the height and width of the graph, and C represents the number of channels. For example, an RGB image has 3 channels. The Patch Embedding process divides the graph into a grid composed of smaller image patches, and the size of each image patch is P×P. The number of such image patches is denoted as m, and its calculation formula is m = (H * W) / P 2 . This division reduces the spatial dimension while retaining the local structure of the graph, making it suitable for subsequent processing. Subsequently, each image patch generates a fixed-size feature vector through convolution and expansion. This transformation is achieved through a learnable linear projection, which maps each image patch from its original pixel space to a higher-dimensional embedding space, thereby forming a sequence, denoted as:
[0056] V = {v0, v1, v2,..., v m-1}
[0057] where: V is the image feature, and each v i ∈R D, where D represents the dimension of the feature vector. This feature sequence is then input into the Transformer layer, where each layer contains a multi-head self-attention module and a feed-forward network. The output of each layer is passed through residual connections and layer normalization to stabilize training and improve gradient flow. Finally, the image features are obtained after encoding by the Transformer layer.
[0058] Step S4: Use the sequence obtained by performing one or more Skip-Range transformations on the first multi-attribute time-series data sequence as the time-series feature, where a single time includes:
[0059] Step S4-1: In an interval sampling manner, perform interval sampling on the sequence before transformation to obtain multiple first segments:
[0060] K i = {x i , x i+k , x i+2k ,...}, i ∈ {0, 1, 2,..., k - 1}
[0061] where: K i is the i-th first segment, x i is the i-th element of the sequence before transformation and also the first element in the i-th first segment, x i+2k is the (i + 2k)-th element of the sequence before transformation and also the third element in the i-th first segment.
[0062] In this way, k first segments can be obtained.
[0063] Step S4-2: Apply self-attention transformation to all the first segments to obtain the corresponding second segments;
[0064] In this process, the model parameters for self-attention transformation can be determined by learning, and K i can be obtained as K i ′
[0065] Step S4-3: Reorganize all the second segments according to the sampling order of their corresponding first segments to obtain the first intermediate sequence;
[0066] In this embodiment, the position of an element at any position in any second segment in the first intermediate sequence is the same as the position of the element at the same position in the corresponding first segment in the sequence before transformation;
[0067] Step S4-4: Sequentially split the first intermediate sequence to obtain multiple third segments:
[0068] R j = {x j , x j+1 , xj+2 ,…, x j+r-1}, j ∈ {0, r, 2r, ..., (k - 1)r}
[0069] Where: R j is the j-th third segment.
[0070] This method divides the data into k segments, each segment with a length of
[0071] Step S4-5: Apply self-attention transformation to all the third segments to obtain the corresponding fourth segments. Similarly, R j is obtained through self-attention transformation as R j ′
[0072] Step S4-6: Reorganize all the fourth segments according to the segmentation sorting of their corresponding third segments to obtain the transformed sequence.
[0073] In this embodiment, the position of an element at any position in any fourth segment in the transformed sequence is the same as the position of the element at the same position in the third segment corresponding to this fourth segment in the first intermediate sequence.
[0074] After multiple Skip-Range transformations, the final time-series data is obtained.
[0075] Generally, in this embodiment, the number of third segments is the same as the number of first segments.
[0076] Of course, in other embodiments, other designs can also be adopted.
[0077] Step S5: Input the image feature and the time-series feature into the trained fusion decoder to obtain the second multi-attribute time-series data sequence reconstructed by the fusion decoder, including:
[0078] Step S5-1: Concatenate the image feature and the time-series feature to obtain a feature vector, and input the feature vector into the trained fusion decoder:
[0079] S = {x0, x1, x2, ..., x n-1}
[0080] V = {v0, v1, v2, ..., v m-1}
[0081] F = {x0, x1, x2, ..., x n-1 , v0, v1, v2, ..., v m-1}
[0082] Where S is the time-series feature, V is the image feature, and F is the feature vector.
[0083] Step S5-2: Obtain the second multi-attribute time-series data sequence reconstructed by the fusion decoder:
[0084] Z = Norm(F + FFN(Norm(F + MultiHead(F))))
[0085] Where: F is the feature vector, Norm is the normalization function, FFN is the feed-forward neural network, MultiHead is the attention mechanism, and Z is the reconstructed second multi-attribute time-series data sequence.
[0086] The loss function in the training process of the fusion decoder is:
[0087]
[0088] Where, is the loss function, Mean is the mean function, and X is the first multi-attribute time-series data sequence.
[0089] Step S6: Calculate the mean square error between the first multi-attribute time-series data sequence and the second multi-attribute time-series data sequence, and determine whether the mean square error is greater than a pre-configured error threshold. If so, it is determined as an anomaly.
[0090] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs, etc., which can store program codes.
Claims
1. An industrial time series data anomaly detection method based on Skip-Range attention, characterized in that, Including: Step S1: Obtain multi-attribute time series data, and perform normalization on each attribute respectively to obtain the first multi-attribute time series data sequence; Step S2: Convert the first multi-attribute time series data sequence into a two-dimensional waveform diagram, where the horizontal axis of the two-dimensional waveform diagram represents time, the vertical axis represents a value, the connection lines of data points corresponding to different attributes at the same moment are parallel to the vertical axis, and the vertical axis value increases gradually outward from the origin; Step S3: Extract image features based on the two-dimensional waveform diagram; Step S4: Use the sequence after one or more Skip-Range transformations of the first multi-attribute time series data sequence as time series features; Step S5: Input the image features and time series features into the trained fusion decoder to obtain the second multi-attribute time series data sequence reconstructed by the fusion decoder; Step S6: Calculate the squared error between the first multi-attribute time series data sequence and the second multi-attribute time series data sequence, and determine whether the squared error is greater than a pre-configured error threshold. If so, it is determined as abnormal.
2. The industrial time series data anomaly detection method based on Skip-Range attention according to claim 1, wherein In the single time of the said step S4, it includes: Step S4-1: Adopt an interval sampling method to perform interval sampling on the sequence before transformation to obtain multiple first segments; Step S4-2: Use self-attention transformation for all the first segments to obtain the corresponding second segments; Step S4-3: Re-organize all the second segments according to the sampling order of their corresponding first segments to obtain the first intermediate sequence; Step S4-4: Sequentially split the first intermediate sequence to obtain multiple third segments; Step S4-5: Use self-attention transformation for all the third segments to obtain the corresponding fourth segments; Step S4-6: Re-organize all the fourth segments according to the splitting order of their corresponding third segments to obtain the transformed sequence.
3. The industrial time series data anomaly detection method based on Skip-Range attention according to claim 2, wherein, The number of the said third segments is the same as that of the first segments.
4. The industrial time series data anomaly detection method based on Skip-Range attention according to claim 2, wherein The position of an element at any position in any second segment in the first intermediate sequence is the same as the position of the element at the same position in the first segment corresponding to this second segment in the sequence before transformation; The position of an element at any position in any fourth segment in the transformed sequence is the same as the position of the element at the same position in the third segment corresponding to this fourth segment in the first intermediate sequence.
5. The industrial time series data anomaly detection method based on Skip-Range attention according to claim 1, characterized in that, The said multi-attribute time series data includes at least temperature time series data, pressure time series data, and vibration time series data.
6. The industrial time series data anomaly detection method based on Skip-Range attention according to claim 1, characterized in that, The said step S5 includes: Step S5-1: Concatenate the image features and time series features to obtain a feature vector, and input the feature vector into the trained fusion decoder; Step S5-2: Obtain the second multi-attribute time series data sequence reconstructed by the fusion decoder: Z = Norm(F + FFN(Norm(F + MultiHead(F)))) Where: F is the feature vector, Norm is the normalization function, FFN is the feed-forward neural network, MultiHead is the attention mechanism, and Z is the second multi-attribute time series data sequence reconstructed.
7. The anomaly detection method for industrial time-series data based on Skip-Range attention according to claim 6, characterized in that, The loss function in the training process of the said fusion decoder is: Among them, is the loss function, Mean is the mean function, and X is the first multi-attribute time series data sequence.
8. A method for abnormal detection of industrial time-series data based on Skip-Range attention according to claim 1, characterized in that, The said step S3 is implemented by the image encoder ViT.
9. An industrial time-series data anomaly device based on Skip-Range attention, comprising a memory, a processor, and a program stored in the memory, characterized in that, When the said processor executes the said program, it implements the method as described in any one of claims 1-8.
10. A storage medium having a program stored thereon, characterized in that, When the described program is executed, it implements the method described in any one of claims 1-8.
Citation Information
Patent Citations
Industrial time series data anomaly detection method and system based on data enhancement
CN119167244A