A Time Series Mutation Point Detection Method Based on Adversarial Alternating Sliding Window
By introducing an adversarial alternating adaptive sliding window in the detection of time series mutation points, dynamically adjusting the window size, the window size selection problem in the prior art is solved, and the accuracy and real-time detection are improved.
Patent Information
- Application Number
- CN202210967432.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-12
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-08-12
AI Technical Summary
In the prior art, the size of the sliding window has a great impact on the accuracy of detection of time series mutation points. Too large or too small window size will lead to a reduction in detection accuracy, making it difficult to adaptively select a suitable window size.
A time series mutation point detection method based on an adversarial adaptive sliding window is proposed. Data preprocessing and abnormal sequence positioning are performed through adversarial sliding windows, mutation point detection is performed in combination with alternating sliding windows, and window size is dynamically adjusted according to local detection results.
It improves the real-time and accuracy of time series mutation point detection, effectively solves the problem of window size selection, and achieves more efficient mutation point detection.
Smart Images

Figure CN115293274B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for detecting mutation points in time series based on an adversarial alternating sliding window, and belongs to the technical field of artificial intelligence data mining. Background Art
[0002] With the rapid development of science and technology and modern industry, the mechatronic equipment in industrial production is developing towards the direction of intelligence, high speed, complexity, digitization and high power. A large amount of production data and equipment operation data are generated every day. Among them, time series data is the most widely existing data type. It is a sequence formed by arranging the values of the same statistical indicator in the order of their occurrence time, and can explain the characteristics information closely related to production activities such as environmental changes and equipment operation states. In actual production, these industrial process variable data are collected through multi-sensor technology, and the internal correlation and difference of the data are deeply analyzed and mined, so as to obtain the current operation state of the equipment, which is of great significance for promoting the high-quality development of industrial production.
[0003] Time series mutation point detection is an important research branch of data mining technology, which refers to finding a change point in the time series, and the distributions of the two segments of data before and after this change point are not the same distribution. In the classification problem, the data before and after the mutation point can be divided into two categories. When the data distribution of the time series changes, how to quickly and accurately find the position of the mutation point is the basis for further fault prediction, anomaly location and root cause analysis of faults. For building a safe and reliable industrial production system, it is of great significance to minimize the occurrence of faults in a timely manner. Therefore, improving the detection accuracy and efficiency of time series mutation points is the problem that needs to be focused on solving at present.
[0004] The sliding window model is one of the key technologies for mutation point detection. Through the sliding window technology, the data to be detected can be sliced into several subsequences, and each subsequence is detected according to the window order to achieve the rapid detection of multiple mutation points in time series data, while meeting the real-time requirements of online detection. However, in the sliding window model, the window size has a great influence on the detection accuracy of mutation points. If the window is too large, the data fluctuations within the window will be masked, resulting in a decrease in the detection accuracy of mutation points. If the window is too small, the amount of information data carried is small, which will also lead to a decrease in the detection accuracy of mutation points. Therefore, how to select an appropriate window size and detection method is the key to improving the mutation point detection technology. How to adaptively select the window length has become the key to improving the detection accuracy. Summary of the Invention
[0005] To solve the above technical problems, the present invention proposes a method for detecting mutation points in time series based on an adversarial alternating adaptive sliding window.
[0006] The specific technical solution of the present invention is as follows:
[0007] S1: Preprocess the time series data of the process variable to obtain a subsequence set;
[0008] S2: Adaptively adjust the size of the adversarial sliding window for intercepting subsequences;
[0009] S3: Use the adversarial sliding window to locate abnormal subsequences;
[0010] S4: Use the alternating sliding window to detect mutation points.
[0011] Further, the preprocessing of the time series data in step S1 is specifically as follows:
[0012] (1) A time series of process variables with a length of T collected in real time from an industrial process is represented as: X = {x 1 ,..., x t ,..., x T}, where x t is the value of the sequence X at time t, and t ∈ {1, 2,..., T};
[0013] (2) For the time series X, after being segmented by the adversarial adaptive sliding window, it is intercepted into a subsequence set and represented as X = {S 1 ,..., S n ,..., S N}, where S n represents the nth subsequence of X; for the data x n in the subsequence S t , it is mapped to the [0, 1] interval based on the max-min window normalization method in formula (1)
[0014]
[0015] In the formula, max(S n ) is the maximum value in the subsequence S n , min(S n ) is the minimum value in the subsequence S n , is the normalized data; then, the normalized time series can be denoted as and there is
[0016] Further, the window size adaptive adjustment strategy based on the local detection result in step S2 is specifically as follows:
[0017] (1) For the subsequence Divide it into 4 sub-sequences of the same length from left to right in chronological order, denoted as The mutation point detection result of is defined as a local detection result, denoted as Denote the sub-sequence The i-th sub-sequence in The mutation point detection result, initially Indicates that no mutation point is detected. If a mutation point is detected after performing the mutation point determination in Then
[0018] (2) Define the sub-sequence The i-th sub-sequence in The detection weight of is The sub-sequence The cumulative weight of is Calculated by formula (2):
[0019]
[0020] S2-3: Set the initial size of the adversarial window for the intercepted sub-sequence S 1 To be w 1 , As the value of n increases, the size w of the adversarial window for the intercepted sub-sequence S n Can be adaptively adjusted according to formula (3): n
[0021]
[0022] Furthermore, the step S3 is based on the abnormal sequence positioning of the adversarial sliding window, specifically as follows:
[0023] (1) For the sub-sequence Construct a positive and negative adversarial training sample set with two adjacent sub-sequences Where the superscript "+" means assigning a "positive" label to all the data in the sub-sequence, and the superscript "-" means assigning a "negative" label to all the data in the sub-sequence;
[0024] (2) For The positive and negative adversarial sample sets in Construct a binary classification model based on the support vector machine (SVM) as the abnormal distribution detection model;
[0025] (3) According to the training sample set Construct a binary classification model based on the support vector machine (SVM) as the abnormal distribution detection model;
[0026] (4) The area under the ROC curve (AUC) value is used as the evaluation index of the SVM binary classification model, as shown in formula (4).
[0027]
[0028] Wherein, M and N respectively represent the numbers of positive samples and negative samples, and Tol represents the number of cases where the predicted probability of a positive sample is greater than that of a negative sample among M×N pairs of samples; The value range is [0,1], and the larger this value, the higher the classification accuracy of the SVM model; If Then it is determined that the adjacent two sub-sequences and belong to the same distribution and there is no need to detect mutation points; Otherwise, go to step S4 to detect mutation points using an alternating sliding window;
[0029] Furthermore, the mutation point detection based on the alternating sliding window in step S4 is as follows:
[0030] (1) For any data point in, use an alternating window with a length of to intercept the data set about as Define the anomaly determination threshold of d t as g(d t )
[0031] g(d t ) = min(max(d t ) - mean(d t ), mean(d t ) - min(d t )) (5)
[0032] Wherein, max(d t ), min(d t ) and mean(d t ) respectively represent the maximum value, minimum value and mean value in d t ;
[0033] (2) Determine mutation points according to formulas (6) and (7)
[0034]
[0035] Wherein, std(d T ) represents the standard deviation of d t , and α is an adjustment parameter;
[0036]
[0037] In the formula, g 0 represents the initial abnormal quantity threshold, and the symbol represents rounding down; represents the number of data points determined to be abnormal points among the past 2d data points;
[0038] When simultaneously satisfies formulas (6) and (7), then it is determined that is a mutation point, and mark the LTR of the sub-sequence where it is located is 1, and store it in the global mutation point result set Test;
[0039] (3) If is determined to be a mutation point in step (2), then perform smoothing processing on it according to formula (8)
[0040]
[0041] In the formula, k represents the mutation point smoothing coefficient;
[0042] (4) After performing mutation point detection on using steps S3 to S4 , use the sub-sequence as the next round of adversarial sample set, and repeat steps S3 to S4 for the next round of mutation point detection; until the sub-sequence the detection is completed, and according to its local detection result use step S2 to obtain the width w of the next round of adversarial window n+1 , adaptively intercept the sub-sequence S n+1 , and after preprocessing in step S1, repeat steps S3 to S4 for a new round of mutation point detection.
[0043] The beneficial effects of the present invention are as follows: The time series anomaly detection method of the present invention is a data-driven method based on a sliding window model. It uses the time series data in two adjacent sliding windows as training samples to construct a binary classification model for adversarial distribution testing to achieve anomaly sequence positioning. Then, based on an alternating window, a dynamic decision threshold is set for mutation point detection, and the size of the sliding window is dynamically adjusted according to the local detection result, effectively improving the real-time performance and accuracy of mutation point detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. By referring to the drawings, the features and advantages of the present invention will be more clearly understood. The drawings are schematic and should not be construed as limiting the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. Among them:
[0045] Figure 1 is the flow chart of the method of the present invention;
[0046] Figure 2 is the turntable after adding counterweight screws in the present invention;
[0047] Figs. 3(a)-(c) are time series diagrams of a rotating mechanical device in different states collected by a sensor in the method embodiment of the present invention;
[0048] Figure 4 is the flow chart of the window adjustment strategy;
[0049] Figure 5 is the flow chart for performing adversarial distribution test in the present invention. Detailed implementation manners
[0050] In order to more clearly understand the above objects, features and advantages of the present invention, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0051] In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited by the specific embodiments disclosed below.
[0052] As Figure 1 shown, the present invention is based on collecting vibration acceleration signals from the normal state of the motor rotor and four rotor imbalance fault states, obtaining time series from the normal state to any one of the fault states; intercepting subsequences according to the adversarial window, dividing them into 4 equal-length sub-sequences and performing windowed normalization processing; then using two adjacent sub-sequences as positive and negative training sample sets to construct an SVM binary classification model for adversarial data distribution test; transferring the sub-sequences that may have mutation points to an alternating sliding window for detection according to the model AUC detection threshold; performing mutation point confidence test through the sliding mean and sliding variance of the alternating window where the time series data point to be detected is located, combined with the empirical threshold; determining mutation points for the data points exceeding the confidence upper and lower bounds in an iterative manner; finally storing the mutation points in the global mutation point result set, and smoothing the mutation points. When the detection of the sub-sequence is completed, according to the local detection result, adaptively update the length of the adversarial window in the next round, intercept a new sub-sequence for the next round of mutation point detection.
[0053] In order to facilitate the understanding of the above technical solutions of the present invention, the above technical solutions of the present invention will be described in detail below through specific embodiments.
[0054] Embodiment
[0055] This embodiment includes the following steps:
[0056] S1: Taking the process of the motor rotor in the rotating machinery fault evolving from the normal state to four rotor imbalance fault levels as an example; the occurrence of rotor faults usually causes abnormal vibrations of the motor rotor, manifested as an increase or decrease in the amplitude of specific frequency components by the abnormal vibration information source (sensor); in this embodiment of the present invention, 4 types of fault states are set: normal (D 1 ), rotor imbalance fault (D 2 ), rotor imbalance fault (D 3 ), rotor imbalance fault (D 4 ), rotor imbalance fault (D 5 ). The larger the recorded signal value, the higher the degree of rotor imbalance fault.
[0057] As Figure 2 shown, for the turntable installed on the rotor test bench, 1 - 4 counterweight screws are installed in the same direction respectively to simulate the four levels of imbalance in a certain direction. The more screws are installed, the more obvious the imbalance effect; in this example, the experimental equipment is the motor flexible rotor ZHS - 2 multi - function experimental platform. The rotational speed of the motor rotor is set to 1500 r / min, the sampling frequency is 1280 Hz, and the fundamental frequency 1X = 25 H; according to the above experimental conditions, three vibration acceleration sensors are placed equidistantly in the horizontal direction of the middle base of the rotor test bench, and four fault conditions with different degrees of imbalance are alternately run on the experimental platform. The HG8902 data collector is used for data acquisition, and the time series of the process variables monitored by the three sensors as shown in Figures 3(a) - (c) is obtained: X = {x 1 , x 2 ,..., x T}, x t = {x t,1 , x t,2 , x t,3} T .
[0058] For the time series X = {x 1 , x 2 ,..., x T}, after adaptive adversarial sliding window segmentation, it is intercepted into a subsequence set and represented as X = {S 1 , S 2 ,..., S N}, S n represents the nth subsequence of X, and the length of each subsequence is determined by the corresponding adversarial window size; for the subsequence S n, normalize the time series within the window according to formula (1) using the maximum and minimum values of the sequence within the window, map the original data to the interval [0, 1], and solve the problem of amplitude differences in sequences under different measurement conditions when the time span is large; denote the normalized time series as
[0059]
[0060] In the formula, max(S n ) is the maximum value in the subsequence S n , min(S n ) is the minimum value in the subsequence S n , is the normalized data.
[0061] S2: In this embodiment, the initial adversarial window size w 1 is selected based on the empirical parameter Table 1, and the subsequence obtained by intercepting according to the adversarial window is The initial sub-sequence detection weight is The local detection result is Initial If a mutation point is detected in the sub-sequence , then
[0062] It is stipulated that the detection result of the sub-sequence closer to the current moment has a higher weight and can better reflect the mutation point detection situation at the current moment; define the detection weight of the i-th sub-sequence in the subsequence as The subsequence The cumulative weight of is Calculated by formula (2):
[0063]
[0064] During the sliding process of the adversarial window with an initial size of w 1 , it is adaptively adjusted according to formula (3), and the window adjustment strategy flow chart is as Figure 4 shown:
[0065]
[0066] When , it means that the probability of not detecting a mutation point in the subsequence is relatively large, and the current data distribution is relatively flat. At this time, the window size should be enlarged to improve the mutation point detection efficiency; when , it means that the probability of detecting a mutation point is close to the probability of not detecting a mutation point, and the window size w remains unchanged; when When it indicates that the probability of not detecting a mutation point is relatively large and the current data distribution is relatively oscillating, the window size is reduced to prevent missed detection caused by an overly large window.
[0067] S3: For the subsequence Construct a positive and negative adversarial training sample set with two adjacent sub-sequences The adversarial test process is as Figure 5 shown;
[0068] Taking as an example, using the SVM model trained with it and then using as the test set for classification; if when, the classification performance of the model is poor and it cannot distinguish between positive and negative samples, indicating that and belong to the same data distribution and do not contain mutation points; if or then it indicates that the model has the ability to separate positive and negative samples, and there is probably a mutation point in two adjacent sub-sequences. At this time, switch to an alternating window for mutation point detection;
[0069] S4: According to the classification results of the SVM binary classification model, find the mutation boundary between positive and negative classes, and use this as a benchmark to iteratively detect mutation points in the time series data points before and after;
[0070] (1) Taking as the mutation boundary, use an alternating window with a length of to intercept the data set about as d t The anomaly determination threshold of t ) is g(d
[0071] g(d t ) = min(max(d t ) - mean(d t ), mean(d t ) - min(d t )) (5)
[0072] In the formula, max(d t ), min(d t ) and mean(d t ) respectively represent the maximum value, minimum value and mean value in d t ;
[0073] (2) Determine mutation points according to formulas (6) and (7)
[0074]
[0075] In the formula, std(dT ) represents the standard deviation of d t , and α is an adjustment parameter;
[0076]
[0077] In the formula, g 0 represents the initial abnormal quantity threshold, and the symbol represents rounding down; represents the number of data points determined to be abnormal points among the past 2d data points;
[0078] When simultaneously satisfies formulas (6) and (7), then it is determined that is a mutation point, mark the LTR of the sub-sequence where it is located is 1, and store it in the global mutation point result set Test;
[0079] (3) Set the outlier smoothing coefficient k, and smooth the mutation point x T according to formula (8);
[0080]
[0081] (4) After using steps S3 to S4 to detect mutation points for the data points in , use the sub-sequence as the next round of adversarial sample set, and repeat steps S3 to S4 for the next round of mutation point detection; until the sub-sequence detection is completed, according to its local detection result use step S2 to obtain the width w n+1 of the next round of adversarial window, adaptively intercept the sub-sequence S n+1 , and after the preprocessing of step S1, repeat steps S3-1 to S4-3 for a new round of mutation point detection.
[0082] Finally, when all sub-sequences of the time series X = {x 1 , x 2 ,..., x T} are detected, conduct an evaluation; if there are n actual mutation points in the time series X, and the present invention detects c mutation points, including TP actual mutation points and FP misdetected pseudo-mutation points; among the remaining undetected data, there are FN undetected actual mutation points and TN actual non-mutation points; thus, the following equations can be obtained:
[0083]
[0084] Based on the above variable assumptions, evaluate according to the evaluation indexes of formulas (10) and (11),
[0085]
[0086]
[0087] In the formula, Hit is the hit rate, representing the ratio of the number of correctly identified mutation points to the number of actual mutation points; MaE is the mean absolute error, representing the ratio of the sum of the distances between the detected mutation points and the actual mutation points to the number of mutation points; Test represents the set of detected mutation points, and Actual represents the set of actual mutation points;
[0088] In this embodiment, the values corresponding to the parameters w 1 , d, α, and g 0 in Table 1 are randomly combined 200 times. According to the combined parameter set {w 1 , d, α, g 0}, the mutation points of the time series in Figures 3(a)-3(c) are respectively detected. Then, according to the positions of the actual mutation points marked in the figure and the positions of the mutation points detected by the present invention, the algorithm is evaluated in combination with Formulas (9)-(11). The evaluation results are shown in Table 2.
[0089] Table 1 Empirical parameter table involved in the method of the present invention
[0090]
[0091] Table 2 Statistical table of 200 detection results of mutation points under the method of the present invention
[0092]
Claims
1. A method for detecting mutation points in time series based on an adversarial alternating sliding window, characterized in that, it includes the following steps: S1: Preprocess the time series data of rotating mechanical equipment in different states collected by sensors to obtain a subsequence set; S2: Adaptively adjust the size of the adversarial sliding window for intercepting subsequences; S3: Locate abnormal subsequences using the adversarial sliding window; S4: Detect mutation points using the alternating sliding window; The specific steps of step S2 are as follows: S2-1: For the subsequence Divide it into four sub-sequences of the same length from left to right in chronological order, denoted as The mutation point detection result of is defined as a local detection result, denoted as Denote the subsequence The i-th sub-sequence in The mutation point detection result, (i = 1, 2, 3, 4, initial Indicates that no mutation point is detected. If a mutation point is detected after the mutation point determination in then S2-2: Define subsequence The detection weight of the i-th sub-sequence is subsequence The cumulative weight of is calculated by formula (2): S2-3: Set the intercepted subsequence S 1 The initial size of the adversarial window is w 1 , As the value of n increases, the size w of the adversarial window of the intercepted subsequence S n is adaptively adjusted according to formula (3): n The specific steps of step S4 are as follows: S4-1: For any one of the data points in use an alternating window of length d to intercept the data set regarding as Define the anomaly determination threshold for d t as g(d t ): g(d t ) = min(max(d t ) - mean(d t ), mean(d t ) - min(d t )) (5) where max(d t ), min(d t ), and mean(d t ) represent the maximum value, minimum value, and mean value in d t respectively. The superscript "+" means that the data in the bisection sequence are all assigned the "positive" label, and the superscript "-" means that the data in the bisection sequence are all assigned the "negative" label; S4-2: Determine mutation points according to formulas (6) and (7) where std(d T ) represents the standard deviation of d t , and α is an adjustment parameter; where, g 0 represents the initial abnormal quantity threshold, and the symbol represents rounding down; represents the quantity of data points determined to be abnormal among the past 2d data points; When When both formula (6) and (7) are satisfied, it is determined that is a mutation point, mark the LTR of the sub-sequence where it is located = 1, and store it in the global mutation point result set Test; If, in step S4-2, is determined as a mutation point, it is smoothed according to formula (8). In the formula, k represents the mutation point smoothing coefficient; S4-4 pair After performing mutation point detection, use the sub-sequence as the next round of adversarial sample set and repeat the next round of mutation point detection; Until the subsequence The detection ends. According to its local detection result Use formula (3) in step S2 to obtain the width w of the next round of adversarial window n+1 and adaptively intercept the subsequence S n+1 After preprocessing in step S1, repeat a new round of mutation point detection.
2. The method for detecting mutation points in time series based on an adversarial alternating sliding window according to claim 1, characterized in that, the specific steps of step S1 are as follows: S1-1 A process variable time series of length T collected in real time from an industrial process, denoted as: X = {x 1 ,..., x t ,..., x T}, where x t is the value of the process variable time series X at time t, and t ∈ {1, 2,..., T}; S1-2: For the time series X, after adversarial adaptive sliding window segmentation, it is intercepted into a subsequence set, which is denoted as X = {S 1 ,..., S n ,..., S N}, where S n represents the nth subsequence of X; For subsequence S n and data x t in it, map it to the interval [0, 1] based on the max-min window normalization method in formula (1) where max(S n ) is the maximum value in the subsequence S n , min(S n ) is the minimum value in the subsequence S n , is the normalized data; then, the normalized time series can be denoted as and there is 3. The method for detecting mutation points in time series based on an adversarial alternating sliding window according to claim 1, characterized in that, the specific steps of step S3 are as follows: S3-1: For the subsequence Construct a positive and negative adversarial training sample set with two adjacent sub-sequences S3-2: For the positive and negative adversarial sample sets build a binary classification model based on support vector machine as the outlier distribution detection model; S3-3: Use the area value under the ROC curve as the evaluation index of the SVM binary classification model, as shown in formula (4) Wherein, M and N respectively represent the numbers of positive samples and negative samples, and Tol represents the number of cases where the predicted probability of a positive sample is greater than that of a negative sample among M×N pairs of samples; The value range is [0, 1]. The larger this value is, the higher the classification accuracy of the SVM model; if then it is determined that the adjacent two sub-sequences and belong to the same distribution and there is no need to detect mutation points; otherwise, go to step S4 to detect mutation points using an alternating sliding window.