Industrial time series data anomaly detection method based on self-supervised contrast learning
Through the self-supervised comparison learning method, the residual network is trained using anchor points, positive and negative samples and neighbor sets, which solves the problem of abnormal detection in industrial sensor data, and achieves more efficient abnormal recognition and accurate abnormal detection effects.
Patent Information
- Application Number
- CN202510332481.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-25
AI Technical Summary
The existing time series anomaly detection methods are difficult to effectively distinguish between normal and abnormal behavior due to the lack of labeled data and scarcity of abnormal samples in industrial sensor data. The existing comparison learning methods assume that the time window after data enhancement may be converted into negative samples, which reduces the detection performance.
The self-supervised comparison learning method is adopted, the current time window is used as an anchor point, the adjacent time window is obtained as a positive sample, and negative samples are generated through exception injection. The residual network is trained using the triple loss function and the self-supervised classification loss function to learn to distinguish between normal and abnormal features, and self-supervised classification training is performed through the nearest neighbor and the farthest neighbor mechanism.
It improves the accuracy and performance of industrial sensor data abnormality detection, effectively reduces the false alarm rate, can identify abnormalities more accurately and provide a rich data foundation, and improves the model's ability to identify abnormalities.
Smart Images

Figure CN120372494A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of anomaly detection of time series data, and particularly relates to an anomaly detection method for industrial time series data based on self-supervised contrast learning. Background Art
[0002] Anomaly detection of time series is a core technology to ensure the reliability of scenarios such as industrial Internet of Things device monitoring, server operation security, and medical health data analysis. Its goal is to identify anomaly points or anomaly segments deviating from the normal pattern from continuous time series data. In recent years, deep learning technology has also been widely applied to the field of time series anomaly detection.
[0003] In industrial scenarios, the fault diagnosis of equipment depends on time series data such as temperature, pressure, and amplitude collected by sensors. Taking the temperature sensor as an example, the data collected by a temperature sensor on a production line under normal working conditions is 25.1, 25.2, 24.8, 25.1, 25.3, 25.2 (unit: °C), and the fluctuation range is within ±0.5 °C; but when the equipment bearing is worn, the temperature data at a certain time point collected is 30 °C. Therefore, through the abnormal data on industrial sensors, it is possible to effectively detect whether the equipment is operating normally. However, industrial sensor data often has volatility and complexity, resulting in many challenges for existing methods in anomaly detection.
[0004] Most existing time series anomaly detection models rely on unsupervised learning, that is, learning normal behavior from an unlabeled dataset and treating samples deviating from normal behavior as anomalies.
[0005] However, in industrial scenarios, the lack of labeled industrial sensor data makes it difficult for the model to understand the difference between normal and abnormal behaviors. Due to the lack of clear labels, the model often defines the normal boundary too tightly, resulting in minor deviations being misjudged as anomalies, thus generating a large number of false alarms (such as regarding a minor fluctuation of ±1 °C as a fault).
[0006] Another method is to use the self-supervised contrast learning method to perform time series anomaly detection on industrial sensor data. Contrast learning trains the model to distinguish similar and dissimilar sample pairs in the dataset, thereby learning feature representations that can distinguish different samples. For example, in the bearing fault detection task, contrast learning will make the normal vibration data and the vibration data during crack faults form an obvious boundary in the feature space. Through contrast learning, the difference between normal and abnormal data becomes clearer and more obvious, which helps the model to identify more accurate boundaries.
[0007] However, due to the scarcity of abnormal samples in industrial sensor data, existing contrastive learning methods for time series anomaly detection usually assume that the time windows after data augmentation are positive samples, while the time windows far from the current time window are negative samples. This assumption is risky because the time window after data augmentation may be converted into a negative sample (for example, scaling the amplitude of vibration time series data by 10% may be converted into abnormal data), while the time window far away may actually represent normal samples. This may cause the model to be unable to effectively distinguish normal and abnormal behaviors, thus reducing the performance of anomaly detection. Summary of the Invention
[0008] To solve the deficiencies of the prior art and achieve the purpose of improving the accuracy and performance of industrial sensor data anomaly detection, the present invention adopts the following technical solutions:
[0009] An industrial time series data anomaly detection method based on self-supervised contrastive learning, comprising the following steps:
[0010] Step 101, divide the time series data of the industrial sensor into multiple time windows;
[0011] Step 102, take the current time window as an anchor point, obtain the adjacent time windows as positive samples, and take the current time window injected with industrial equipment abnormal data as a negative sample; the positive samples retain the continuity and context information of the time series, providing a rich data basis for the model to learn the normal mode and abnormal mode of the industrial sensor time segment, and effectively improving the model's ability to identify anomalies.
[0012] Step 103, extract features from the anchor point and its corresponding positive and negative samples, and map the extracted features to the feature space;
[0013] Step 104, based on the feature distributions of the anchor point, positive samples, and negative samples in the feature space, determine the nearest neighbor set and the farthest neighbor set for each time window, and obtain an anchor point neighbor set composed of each anchor point and its neighbors;
[0014] Step 105, perform self-supervised classification training based on the anchor point neighbor set to classify the normal and abnormal data collected by the industrial sensor;
[0015] Step 106, generate an anomaly score according to the classification result, and judge whether the data of the current time window collected by the industrial sensor is abnormal.
[0016] Further, in the step 102, if the current time window belongs to a multi-dimensional time window, randomly select multiple dimensions to inject anomalies.
[0017] Further, in step 102, for each injected anomaly, randomly select the start position of the anomaly and randomly select the duration of the anomaly, and this duration is not greater than 90% of the current time window length.
[0018] Further, in step 102, the injected anomalies include but are not limited to global anomalies, context anomalies, seasonal anomalies, shape anomalies, and trend anomalies;
[0019] For the injection operation of global anomalies: randomly select at least one data point within the window and adjust its value to u ω ±g*σ ω where u ω and σ ω respectively represent the mean and standard deviation of the current time window, and g represents a preset multiple coefficient;
[0020] For the injection operation of context anomalies: randomly select a subsequence [s, e] within the window and adjust its value to u se ±x*σ se where u se and σ se represent the mean and standard deviation of the subsequence, and x represents a preset multiple coefficient;
[0021] For the injection operation of seasonal anomalies: randomly adjust the frequency coefficient f of the subsequence and generate a non-seasonal pattern through interpolation or compression;
[0022] For the injection operation of shape anomalies: randomly select a subsequence [s, e] within the window and limit the value of each data point in this subsequence within the same anomaly interval;
[0023] For the injection operation of trend anomalies: randomly select a subsequence [s, e] within the window. For each time point t within the subsequence, adjust its value to where ω(t) represents the value of the original time point t, b represents the trend coefficient, σ ω represents the standard deviation of the current time window, represents the normalized linear factor, s represents the start time of the sequence, and e represents the end time of the sequence.
[0024] Further, in step 103, during the training phase, some anchor points and positive and negative samples are obtained and input into the initial residual network. The initial residual network extracts the feature representations of each sample, calculates the loss according to the constructed triplet loss function, then uses the backpropagation algorithm to calculate the gradients, updates the parameters of the initial residual network, and repeats this step multiple times to minimize the triplet loss function, obtaining a once-trained residual network. The triplet loss function can enable the residual network to reduce the distance between the anchor point and its corresponding positive sample, while increasing the distance between the anchor point and the negative sample, encouraging the initial residual network to learn to distinguish the representations of normal windows and abnormal windows; the anchor point and its corresponding positive and negative samples are input into the once-trained residual network to extract features, the extracted features are dimensionally reduced by a multi-layer perceptron, and the dimensionally reduced features are mapped into the feature space to obtain the corresponding feature distribution.
[0025] Further, the loss function in step 103 is the triplet loss function The formula is as follows:
[0026]
[0027] Among them, φ P represents the residual network for training, that is, the initial residual network; T represents the set containing all triplets; α represents the margin parameter, controlling the minimum distance between the positive sample and the negative sample; |T| represents the size of the set T, that is, the number of triplets; a represents the anchor point; p represents the positive sample; n represents the negative sample; represents the squared Euclidean distance, which is the standard for controlling the minimum distance between the positive and negative samples; φ P (a) represents the anchor point feature obtained after passing through the residual network φ P ; φ P (p) represents the positive sample feature obtained after passing through the residual network φ P ; φ P (n) represents the negative sample feature obtained after passing through the residual network φ P .
[0028] Further, in step 104, for each anchor point, according to the feature distribution in the feature space, calculate the Euclidean distance between the anchor point feature and the positive sample feature, and the Euclidean distance between the anchor point feature and the negative sample feature respectively. Determine the set of positive sample features with the closest Euclidean distance to the anchor point feature as the nearest neighbor set, and find the set of negative sample features with the farthest Euclidean distance from the anchor point as the farthest neighbor set. Based on the anchor point feature, the farthest neighbor set and the nearest neighbor set, construct an anchor neighbor record. Each anchor point has a corresponding anchor neighbor record. The combination of all anchor neighbor records constitutes the anchor neighbor set, where each element represents an anchor neighbor record. The anchor neighbor set includes all anchor points, as well as the corresponding nearest neighbor set and farthest neighbor set. Step 104 constructs a data set containing high-quality training data, providing a priori knowledge for the model. This a priori knowledge uses the nearest and farthest neighbor sets of anchor points to capture the semantic similarity and semantic dissimilarity between window representations, providing a basis for subsequent self-supervised classification training.
[0029] Further, in step 105, construct a self-supervised classification loss function. For each anchor neighbor record, input the anchor point, its nearest neighbor set, and farthest neighbor set into the first residual network for feature extraction to generate classification features. Based on the classification features, obtain classification results for different categories, calculate the value of the self-supervised classification loss function according to the classification results, and use the gradient descent algorithm to reduce the loss and update the parameters until the value of the self-supervised classification loss function converges, obtaining the majority class and the second residual network. By step 105, maximize the similarity between the anchor point and the nearest neighbor set, and minimize the similarity between the anchor point and the farthest neighbor set, thus further distinguishing between the normal mode and the abnormal mode, and improving the model's ability to identify anomalies in industrial sensor time series.
[0030] Further, the self-supervised classification loss function in step 105 is as follows:
[0031]
[0032] where represents the consistency loss function; represents the inconsistency loss function; represents the entropy loss function; β represents the weight parameter used to adjust the contribution intensity of the entropy loss; φ s represents the first residual network; B represents the anchor neighbor set; N represents the set composed of all nearest neighbor sets, that is, each element in this set is the nearest neighbor set of a certain anchor point; F represents the set composed of all farthest neighbor sets, that is, each element in this set is the farthest neighbor set of a certain anchor point; C represents the category set. The goal of the self-supervised classification loss function is to learn highly discriminative representations.
[0033] The consistency loss function calculates the logarithm of the similarity between each anchor point and all elements in the nearest neighbor set and sums them up. Its function is to increase the similarity between the anchor point and its nearest neighbor set. The formula is as follows:
[0034]
[0035] Where |B| represents the size of set B; ω represents the anchor point; ω n represents the elements in the nearest neighbor set; Nω represents the nearest neighbor set for a single anchor point; similarity(φ s ,ω,ω n ) represents the anchor points ω and ω n The similarity between .
[0036] The inconsistency loss function calculates the logarithm of the similarity between each anchor point and all elements in its farthest neighbor set and sums them up. Its function is to reduce the similarity between the anchor point and its farthest neighbor set. The formula is as follows:
[0037]
[0038] Among them, |B| represents the size of set B; ω represents the anchor point, ω n represents the element in the farthest neighbor set; Fω represents the farthest neighbor set for a single anchor point; similarity(φ s ,ω,ω n ) represents the anchor point ω and its ω n The similarity between .
[0039] The entropy loss function Probability Take the logarithm and multiply by Finally, all categories are summed to prevent the model from being overly dependent on specific patterns in the training data, improve the generalization ability of the model, and thus improve the accuracy of anomaly detection. The formula is as follows:
[0040]
[0041] Among them, c represents the element in the category set, that is, the category; It represents the average probability of being predicted as category c, that is, for each category c, the average probability of being predicted as c in all samples is calculated to obtain the average probability of category c
[0042] Further, in step 106, after the second residual network extracts features, classification is performed through normalization. The normalization process calculates the probability of being classified into the majority class. For the classification result being the majority class, it is determined that the current time window is normal. Otherwise, the anomaly score is calculated, that is, the probability of the minority class. The smaller the probability of being assigned to the majority class, the higher the anomaly score and the greater the probability of anomaly. The obtained anomaly score is compared with a set threshold. If the value of the anomaly score is not greater than the threshold, it is determined that the current time window is still normal. If it is greater than the threshold, it is determined that an anomaly has occurred in the current time window.
[0043] The advantages and beneficial effects of the present invention are as follows:
[0044] The present invention constructs negative samples by means of anomaly injection and performs contrastive learning with normal time series windows, thereby effectively learning the feature representations for distinguishing normal and abnormal in industrial scenarios. At the same time, the present invention uses the nearest and farthest neighbor mechanisms for self-supervised classification training, thereby solving the deficiencies of the anomaly detection method based on contrastive learning in the current industrial scenario and further improving the performance of anomaly detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 is the flowchart of the method according to the embodiment of the present invention.
[0046] Figure 2 is a schematic diagram of time window division in the embodiment of the present invention.
[0047] Figure 3 is a schematic diagram of sample acquisition in the embodiment of the present invention.
[0048] Figure 4 is a schematic diagram of contrastive learning in the embodiment of the present invention.
[0049] Figure 5 is a schematic diagram of the set of anchor neighbors in the embodiment of the present invention.
[0050] Figure 6 is a schematic diagram of classification training in the embodiment of the present invention.
[0051] Figure 7 is a schematic diagram of category division in the embodiment of the present invention.
[0052] Figure 8 is a schematic diagram of anomaly judgment in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0053] The following further details the specific embodiments of the present invention with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only for illustrating and explaining the present invention and are not used to limit the present invention.
[0054] As Figure 1 shown, the industrial time series data anomaly detection method based on self-supervised contrastive learning includes the following steps:
[0055] Step 101: Divide the input time series data into multiple time windows.
[0056] For the given time series data, it can be divided into multiple time windows through the sliding window technique. The sliding window technique is a method of splitting time series data into multiple time windows, and each time window contains a fixed number of consecutive time points. The sliding window technique divides the time series into multiple time windows by sliding a window of a set size along the time series data and extracting the data points within each window.
[0057] In one embodiment, as Figure 2 shown, it is part of the temperature data recorded by a temperature sensor every 1 hour, and its sequence is [24.9, 25.3, 25.0, 25.0, 24.8, 25.1, 24.9, 24.8, 25.1, 25.2]. Assuming the size of the sliding window is 4, then the four numbers [24.9, 25.3, 25.0, 25.0] form a time window, and the four numbers [24.8, 25.1, 24.9, 24.8] form another time window, and so on. Slide the window from left to right, and each movement can form a time window until the right side of the sliding window slides to the last number.
[0058] In the present invention, the size of the sliding window is set to 200, and thus the size of each time window is also 200.
[0059] Thus, through Step 101, the input industrial sensor time series data is divided into multiple time windows.
[0060] Step 102: For each time window, use itself as an anchor point, obtain the adjacent time windows as positive samples, and at the same time perform anomaly injection on the current time window to obtain the negatively sampled time window after anomaly injection.
[0061] In one embodiment, as Figure 3 shown, for one time window ω i , use itself as an anchor point, and randomly select one from the y time windows adjacent to the current time window ω i as a positive sample, because the adjacent time windows are usually in the same pattern as the current time window. Next, perform a copy operation on the time window ω i to obtain a copy of the time window ω i , and this copy is exactly the same as the time window ω i .
[0062] The types of anomalies include five types: global anomaly, context anomaly, seasonal anomaly, shape anomaly, and trend anomaly.
[0063] For the injection operation of global anomalies: Randomly select at least one data point within the window and adjust its value to u ω ±g*σ ω , where u ω and σ ω represent the mean and standard deviation of the current time window respectively, and g represents a preset multiple coefficient;
[0064] For the injection operation of context anomalies: Randomly select a subsequence [s, e] within the window and adjust its value to u se ±x*σ se , where u se and σ se represent the mean and standard deviation of the subsequence respectively, and x represents a preset multiple coefficient;
[0065] For the injection operation of seasonal anomalies: Randomly adjust the frequency coefficient f of the subsequence and generate a non-seasonal pattern through interpolation or compression;
[0066] For the injection operation of shape anomalies: Randomly select a subsequence [s, e] within the window and limit the value of each data point in the subsequence within the same anomaly interval;
[0067] For the injection operation of trend anomalies: Randomly select a subsequence [s, e] within the window. For each time point t within the subsequence, adjust its value to where ω(t) represents the value of the original time point t, b represents the trend coefficient, and σ ω represents the standard deviation of the current time window, represents the normalized linear factor.
[0068] Randomly select one of these five anomalies and inject it into the copy to obtain a negative sample. Figure 3 The negative sample shown is obtained by selecting the global anomaly to inject into the copy. For a time window, there will be a corresponding anchor point, a positive sample, and a negative sample.
[0069] For each injected anomaly, randomly select the start position of the anomaly and randomly select the duration of the anomaly, which is between 1 data point and 90% of the length of the current time window.
[0070] If the current time window belongs to a multi-dimensional time window, randomly select d dimensions to inject anomalies, where Dim is the total number of dimensions of the current time window.
[0071] So far, through step 102, positive samples that preserve the continuity and context information of the time series are obtained. At the same time, the negative samples generated by anomaly injection contain various types of anomalies. The construction of positive and negative samples provides a rich data basis for the model to learn the normal and abnormal patterns of industrial sensor time segments, effectively improving the model's ability to identify anomalies.
[0072] Step 103: Input the anchor points, positive samples, and negative samples into the residual network to construct a triplet loss function. The residual network outputs features according to the loss function, and then through the processing of a multi-layer perceptron, the corresponding feature distribution is obtained.
[0073] In one embodiment, as Figure 4 shown, first construct a triplet loss function in the form of a triplet (a, p, n), and its formula is:
[0074]
[0075] where denotes the triplet loss function, φ P denotes the residual network for training, that is, the initial residual network; T denotes the set containing all triplets; α denotes the margin parameter, which controls the minimum distance between the positive sample and the negative sample; |T| denotes the size of the set T, that is, the number of triplets; a denotes the anchor point; p denotes the positive sample; n denotes the negative sample; denotes the squared Euclidean distance, which is the standard for controlling the minimum distance between the positive and negative samples; φ P (a) denotes the anchor point feature obtained after passing through the residual network φ P ; φ P (p) denotes the positive sample feature obtained after passing through the residual network φ P ; φ P (n) denotes the negative sample feature obtained after passing through the residual network φ P .
[0076] First is the training phase. Take some anchor points and positive and negative samples and input them into the initial residual network. The initial residual network extracts the feature representations of each sample, calculates the loss according to the triplet loss function, then uses the backpropagation algorithm to calculate the gradients, and updates the parameters of the initial residual network. Repeat this step multiple times to minimize the triplet loss function. The triplet loss function enables the residual network to reduce the distance between the anchor point and its corresponding positive sample, while increasing the distance between the anchor point and the negative sample, encouraging the initial residual network to learn to distinguish the representations of normal windows and abnormal windows. After the initial residual network is trained, a once-trained residual network is obtained. Then enter the feature mapping phase. Input the anchor points, the positive samples corresponding to the anchor points, and the negative samples corresponding to the anchor points into the once-trained residual network. The once-trained residual network extracts their features, and then outputs the processed anchor point features, positive sample features, and negative sample features. These features are then input into a multi-layer perceptron. The multi-layer perceptron performs dimensionality reduction on the input features, obtains their dimensionality-reduced features, and maps the dimensionality-reduced features to a feature space.
[0077] Repeat the above steps until all anchor points and their corresponding positive and negative samples are input. Finally, the multi-layer perceptron will map all anchor points and positive and negative samples into the same feature space to form a feature distribution.
[0078] So far, through step 103, the residual network can learn the features of positive and negative samples. At the same time, it encourages the residual network to shorten the distance between the anchor point and the positive sample in the feature space, widen the distance between the anchor point and the negative sample in the feature space, distinguish the feature representations of positive and negative samples, and lay a foundation for the subsequent nearest and farthest neighbor retrieval steps. At the same time, through dimensionality reduction, the multi-layer perceptron can effectively reduce the complexity of the output features.
[0079] Step 104, based on the feature distributions of anchor points, positive samples, and negative samples, determine the nearest neighbor set and the farthest neighbor set for each time window, and obtain an anchor neighbor set composed of each anchor point and its neighbors.
[0080] Based on the feature distributions of all anchor point features, positive sample features, and negative sample features in the feature space, for each time window, that is, each anchor point, calculate the Euclidean distances between the anchor point feature and the positive sample feature, and between the anchor point feature and the negative sample feature. Find the 5 positive sample features with the closest Euclidean distance to the anchor point as the nearest neighbor set, and find the 5 negative sample features with the farthest Euclidean distance to the anchor point as the farthest neighbor set. Concatenate the anchor point feature, the farthest neighbor set, and the nearest neighbor set to form an anchor neighbor record. Each anchor point has a corresponding anchor neighbor record, and the combination of all anchor neighbor records constitutes the anchor neighbor set.
[0081] In an embodiment, as Figure 5As shown, it is a schematic diagram of the anchor neighbor set. Among them, Anchor 1 has 5 nearest neighbors, namely neighbor N11, neighbor N12, neighbor N13, neighbor N14, and neighbor N15. These 5 nearest neighbors form the nearest neighbor set of Anchor 1. There are 5 farthest neighbors, namely far neighbor F11, far neighbor F12, far neighbor F13, far neighbor F14, and far neighbor F15. These 5 farthest neighbors form the farthest neighbor set of Anchor 1. Anchor 1 and its nearest neighbor set and farthest neighbor set form an anchor neighbor record. The same is true for Anchor 2, Anchor 3 up to Anchor X. All the anchor neighbor records form the anchor neighbor set.
[0082] So far, through Step 104, a data set containing high-quality training data is constructed, providing a prior knowledge for the model. This prior knowledge uses the nearest and farthest neighbor sets of the anchors to capture the semantic similarity and semantic dissimilarity between window representations, providing a basis for subsequent self-supervised classification training.
[0083] Step 105, design a self-supervised classification loss function, input the anchor neighbor set into the residual network, and the residual network is trained according to the self-supervised classification loss function to optimize its own classification ability.
[0084] In one embodiment, as Figure 6 shown, first construct a consistency loss function, and its formula is:
[0085]
[0086] Among them, represents the consistency loss function; φ s represents training the residual network once; B represents the anchor neighbor set; N represents the set composed of all nearest neighbor sets, that is, each element in this set is the nearest neighbor set of a certain anchor; |B| represents the size of set B; ω represents the anchor; ω n represents the element in the nearest neighbor set; Nω represents the nearest neighbor set for a single anchor; (similarity(φ s , ω, ω n ) represents the similarity between anchor ω and ω n . The consistency loss function calculates the logarithm of the similarity between each anchor and all elements in the nearest neighbor set and sums them. Its role is to increase the similarity between the anchor and its nearest neighbor set.
[0087] Next, construct an inconsistency loss function, and its formula is:
[0088]
[0089] Among them, represents the inconsistency loss function; φs Let A denote the one - time training residual network; B denote the set of anchor neighbors; F denote the set composed of all the farthest neighbor sets, that is, each element in this set is the farthest neighbor set of a certain anchor; |B| denote the size of set B; ω denote the anchor, ω n denote the elements in the farthest neighbor set; Fω denote the farthest neighbor set for a single anchor; (similarity(φ s , ω, ω n ) denote the similarity between anchor ω and its ω n Among them, the dissimilarity loss function calculates the logarithm of the similarity between each anchor and all elements in its farthest neighbor set and sums them up. Its role is to reduce the similarity between the anchor and its farthest neighbor set.
[0090] Finally, construct the entropy loss function, and its formula is:
[0091]
[0092] where denote the entropy loss function; φ s denote the one - time training residual network; B denote the set of anchor neighbors; C denote the set of categories; c denote the elements in the set of categories, that is, the categories; denote the average probability of being predicted as category c; The calculation process of this formula is: for each category c, calculate the average value of the probabilities of being predicted as c in all samples to obtain the average probability of category c Then calculate the entropy loss, that is, take the logarithm of the probability and then multiply by Finally, sum over all categories. The role of the entropy loss function is to prevent the model from over - relying on specific patterns in the training data, improve the generalization ability of the model, and thus improve the accuracy of anomaly detection.
[0093] After having the consistency loss function, the dissimilarity loss function and the entropy loss function, perform a linear transformation on these three to obtain the self - supervised classification loss function. The formula of the self - supervised classification loss function is:
[0094]
[0095] where denote the self - supervised classification loss function; denote the consistency loss function; denote the dissimilarity loss function; denote the entropy loss function; β denote the weight parameter used to adjust the contribution intensity of the entropy loss; φ sLet A represent the one-time trained residual network; B represent the set of anchor neighbors; N represent the set composed of all nearest neighbor sets, that is, each element in this set is the nearest neighbor set of a certain anchor; F represent the set composed of all farthest neighbor sets, that is, each element in this set is the farthest neighbor set of a certain anchor; C represent the set of categories; the objective of the self-supervised classification loss function is to learn highly discriminative representations.
[0096] In one embodiment, as Figure 7 shown, under the training of the self-supervised classification loss function, the residual network can assign the elements in the nearest neighbor set to the same class as the anchor, while the elements in the farthest neighbor set are assigned to different classes from the anchor. By combining the semantically meaningful nearest and farthest neighbors, the anomaly recognition becomes more accurate.
[0097] To complete the classification task, the number of categories needs to be set. In the present invention, the number of categories is set to 10, that is, there are 10 categories.
[0098] After establishing the self-supervised classification loss function and the number of categories, the classification training phase begins. For each anchor neighbor record, the anchor, its nearest neighbor set, and farthest neighbor set are all input into the one-time trained residual network. The one-time trained residual network extracts the features of each element and then outputs the classification features of each element. The output classification features are input into the Softmax layer. After the Softmax layer performs normalization processing, the probabilities of 10 categories are obtained, and the category with the highest probability is selected as the classification result. According to the classification result, the value of the self-supervised classification loss function is calculated, and then the gradient descent algorithm is used to reduce the loss and update the model parameters until the value of the self-supervised classification loss function converges. At this time, the classification training ends, and a majority class and a two-time trained residual network are obtained.
[0099] So far, through step 105, the model maximizes the similarity between the anchor and the nearest neighbor set and minimizes the similarity between the anchor and the farthest neighbor set, thereby further distinguishing the normal mode and the abnormal mode and enhancing the model's ability to identify anomalies in industrial sensor time series.
[0100] Step 106, in the test stage, according to the category probabilities output by the residual network, an anomaly score is generated to determine whether the current time window is abnormal.
[0101] In one embodiment, as Figure 7 shown, in the test stage, given a time window, the two-time trained residual network extracts its features and then passes through the Softmax layer for classification. The Softmax layer will calculate the probability of being classified into the majority class If for the other 9 classes among the set of 10 classes set, there are all It means that the probability of being classified into the majority class is higher than the probability of being classified into any other class. That is, if the classification result is the majority class, it is determined that the time window is normal, where c represents any one of the other nine classes. represents the probability of being classified into class c; otherwise, the anomaly score is calculated. The anomaly score: That is, 1 minus the probability of being classified into the majority class. The smaller the probability of being classified into the majority class, the higher the anomaly score and the greater the probability of anomaly. The obtained anomaly score is compared with the set threshold. If the value of the anomaly score is not greater than the threshold, it is determined that the time window is normal. If it is greater than the threshold, it is determined that an anomaly has occurred in the time window.
[0102] After relevant tests, the present invention performs excellently in the time series anomaly detection task and can outperform existing self-supervised, semi-supervised, and unsupervised methods on multiple benchmark datasets. In industrial sensor data, the present invention can more accurately identify anomalies and effectively reduce the false alarm rate, bringing new breakthroughs to related fields.
[0103] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An industrial time series data anomaly detection method based on self-supervised contrastive learning, characterized in that The method includes the following steps: Step 101: Divide the time series data of the industrial sensor into multiple time windows; Step 102: Take the current time window as an anchor point, obtain adjacent time windows as positive samples, and take the current time window injected with abnormal data of the industrial device as a negative sample; Step 103: Extract features from the anchor point and its corresponding positive and negative samples, and map the extracted features to the feature space; Step 104: Based on the feature distributions of the anchor point, positive samples, and negative samples in the feature space, determine the nearest neighbor set and the farthest neighbor set for the time window, and obtain an anchor neighbor set composed of each anchor point and its neighbors; Step 105: Perform self-supervised classification training based on the anchor neighbor set to classify the normal and abnormal data collected by the industrial sensor; Step 106: Generate an anomaly score according to the classification result, and determine whether the data of the current time window collected by the industrial sensor is abnormal.
2. The industrial time series data anomaly detection method based on self-supervised contrastive learning according to claim 1, wherein: In step 102, if the current time window belongs to a multi-dimensional time window, randomly select multiple dimensions to inject anomalies.
3. The industrial time series data anomaly detection method based on self-supervised contrastive learning according to claim 1, characterized in that: In step 102, for each injected anomaly, randomly select the start position of the anomaly and randomly select the duration of the anomaly, and the duration is not greater than 90% of the length of the current time window.
4. The industrial time series data anomaly detection method based on self-supervised contrastive learning according to claim 1, wherein: In step 102, the injected anomalies include but are not limited to global anomalies, context anomalies, seasonal anomalies, shape anomalies, and trend anomalies; For the injection operation of global exceptions: randomly select at least one data point within the window and adjust its value to u ω ±g*σ ω , where u ω and σ ω represent the mean and standard deviation of the current time window respectively, and g represents the preset multiple coefficient; For the injection operation with abnormal context: randomly select the subsequence [s, e] within the window and adjust its value to u se ±x*σ se , where u se and σ se represent the mean and standard deviation of the subsequence, and x represents the preset multiple coefficient; For the injection operation of seasonal anomalies: randomly adjust the frequency coefficient f of the subsequence, and generate a non-seasonal pattern through interpolation or compression; For the injection operation of shape anomalies: randomly select a subsequence [s, e] within the window, and limit the value of each data point in the subsequence within the same anomaly interval; For the injection operation with abnormal trends: randomly select the subsequence [s, e] within the window, and for each time point t within the subsequence, adjust its value to where ω(t) represents the value at the original time point t, b represents the trend coefficient, and σ ω represents the standard deviation of the current time window, represents the normalized linear factor, s represents the start time of the sequence, and e represents the end time of the sequence.
5. The industrial time series data anomaly detection method based on self-supervised contrastive learning according to claim 1, characterized in that: In step 103, during the training phase, input the anchor point and positive and negative samples into the initial residual network. The initial residual network extracts the feature representations of each sample, calculates the loss according to the constructed loss function, then uses the backpropagation algorithm to calculate the gradient, and updates the parameters of the initial residual network to minimize the loss function, obtaining a once-trained residual network; input the anchor point and its corresponding positive and negative samples into the once-trained residual network to extract features, reduce the dimensions of the extracted features through a multi-layer perceptron, and map the reduced-dimensional features to the feature space to obtain the corresponding feature distributions.
6. The industrial time series data anomaly detection method based on self-supervised contrastive learning according to claim 5, wherein: The loss function in the step 103 is a triplet loss function The formula is as follows: Among them, φ P represents the residual network for training, that is, the initial residual network; T represents the set containing all triples; α represents the margin parameter, controlling the minimum distance between positive and negative samples; |T| represents the size of the set T, that is, the number of triples; a represents the anchor; p represents the positive sample; n represents the negative sample; represents the squared Euclidean distance, which is the criterion for controlling the minimum distance between positive and negative samples; φ P (a) represents the anchor feature obtained after passing through the residual network φ P ; φ P (p) represents the positive sample feature obtained after passing through the residual network φ P ; φ P (n) represents the negative sample feature obtained after passing through the residual network φ P .
7. The industrial time series data anomaly detection method based on self-supervised contrastive learning according to claim 1, characterized in that: In step 104, for each anchor point, according to the feature distribution in the feature space, calculate the Euclidean distance between the anchor point feature and the positive sample feature, and the Euclidean distance between the anchor point feature and the negative sample feature respectively. Determine a group of positive sample features with the closest Euclidean distance to the anchor point feature as the nearest neighbor set, and find a group of negative sample features with the farthest Euclidean distance from the anchor point as the farthest neighbor set; based on the anchor point feature, the farthest neighbor set, and the nearest neighbor set, construct an anchor neighbor record. Each anchor point has a corresponding anchor neighbor record, and the combination of all anchor neighbor records constitutes the anchor neighbor set.
8. The industrial time series data anomaly detection method based on self-supervised contrastive learning according to claim 7, characterized in that: In step 105, a self-supervised classification loss function is constructed. For the anchor neighbor records, the anchor and its sets of nearest neighbors and farthest neighbors are all input into the first residual network for feature extraction to generate classification features. Based on the classification features, classification results for different classes are obtained. The value of the self-supervised classification loss function is calculated according to the classification results. The gradient descent algorithm is used to reduce the loss and update the parameters until the value of the self-supervised classification loss function converges, obtaining the majority class and the second residual network.
9. The industrial time series data anomaly detection method based on self-supervised contrastive learning according to claim 8, characterized in that: The self-supervised classification loss function in step 105 has the following formula: Among them, represents the consistency loss function; represents the inconsistency loss function; represents the entropy loss function; β represents the weight parameter used to adjust the contribution intensity of the entropy loss; φ s represents the first residual network; B represents the set of anchor neighbors; N represents the set composed of all nearest neighbor sets; F represents the set composed of all farthest neighbor sets; C represents the category set; The consistency loss function calculates the logarithm of the similarity between each anchor and all elements in the set of nearest neighbors and sums them up. The formula is as follows: where |B| represents the size of set B; ω represents an anchor point; ω n represents an element in the nearest neighbor set; Nω represents the nearest neighbor set for a single anchor point; similarity(φ s , ω, ω n ) represents the similarity between anchor points ω and ω n ; The inconsistency loss function calculates the logarithm of the similarity between each anchor and all elements in its set of farthest neighbors and sums them up. The formula is as follows: where |B| represents the size of set B; ω represents an anchor point, ω n represents an element in the farthest neighbor set; Fω represents the farthest neighbor set for a single anchor point; similarity(φ s , ω, ω n ) represents the similarity between the anchor point ω and its ω n therebetween; The entropy loss function takes the logarithm of the probability and then multiplies it by Finally, sum over all classes, and the formula is as follows: Among them, c represents an element in the category set, that is, the category; represents the average probability of being predicted as category c. That is, for each category c, calculate the average value of the probabilities of all samples being predicted as c to obtain the average probability of category c 10. The industrial time series data anomaly detection method based on self-supervised contrastive learning according to claim 8, characterized in that: In step 106, after the second residual network extracts features, classification is performed through normalization. The normalization process calculates the probability of being classified as the majority class. For the classification result being the majority class, it is determined that the current time window is normal. Otherwise, the anomaly score, that is, the probability of the minority class, is calculated. The smaller the probability of being classified into the majority class, the higher the anomaly score and the greater the probability of anomaly. The obtained anomaly score is compared with a set threshold. If the value of the anomaly score is not greater than the threshold, it is determined that the current time window is still normal. If it is greater than the threshold, it is determined that an anomaly has occurred in the current time window.
Citation Information
Patent Citations
Image anomaly detection method and system based on comparative learning
CN118115450A
Industrial internet time series data anomaly detection method and system
CN118898045A
Method for detecting abnormal defect on steel surface based on semi-supervised contrastive learning
US20240210329A1