A method for labeling anomalies in time series data
By dividing the time series into four categories and performing processed exception marking methods, the problems of high complexity and poor scalability of large-scale time series exception marking are solved, efficient exception marking of time series is achieved, and data analysis capabilities are improved.
Patent Information
- Application Number
- CN202210090576.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-25
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-01-25
AI Technical Summary
Large-scale time series exception marking faces the problems of high algorithm complexity, poor scalability and difficulty in manual labeling.
By dividing the time series into four categories and abnormal marking is performed according to the characteristics of each category of data, process-based operations are used to mark out the outliers of the time series. Specific steps include time series data preprocessing, classification and exception marking.
It realizes accurate and efficient labeling of outliers, context outliers and jump values of time series, improves data integrity and analysis capabilities, and is suitable for large-scale time series data analysis.
Smart Images

Figure CN114490156B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data analysis, and in particular relates to a method for marking anomalies in big data time series. Background Art
[0002] Time series is a sequence of data arranged continuously in chronological order. It is one of the most common data types in the fields of production, business activities, etc. In the process of data collection, transmission and storage, software and hardware failures often lead to the generation of erroneous data; abnormalities in the data generator will also lead to the generation of abnormal data (for example, an abnormal temperature sensor records inaccurate temperature information). Correctly marking abnormal / erroneous data in time series is a key step in correcting erroneous data, improving data integrity, and analyzing abnormal events.
[0003] With the rapid increase in the scale of various business data, time series anomaly labeling is also facing more and more challenges. At present, large-scale time series anomaly labeling mainly faces problems such as high algorithm complexity, poor scalability and difficulty in manual labeling. In order to solve the above problems, a simple and efficient anomaly labeling algorithm suitable for the early data processing of large-scale time series is developed here, which can accurately and efficiently label the outliers of time series with process-based operations. Summary of the invention
[0004] In order to meet the needs of big data analysis, the present invention discloses a simple and efficient time series anomaly marking method applicable to massive data. The method divides the time series into four categories according to the data characteristics of the sequence, and then marks the anomalies of the time series according to the characteristics of each data type.
[0005] The time series data anomaly marking algorithm disclosed in the present invention specifically includes the following steps:
[0006] Step 1: Preprocessing of time series data;
[0007] Step 2: Time series classification;
[0008] Step 3: Mark abnormal time series data.
[0009] Preferably, step 1 is as follows:
[0010] Step 1.1, normalize time series data.
[0011] Let the original time series X = {x1, x2, ..., x N}, where N is the length of the sequence, x tis the specific value of the sequence X at a certain moment, and t∈{1, 2, …, N}. The original time series X is normalized using the Min-Max method, and the original data points are mapped to the interval [0, 1]. The normalized time series is recorded as The calculation formula is:
[0012]
[0013] Where max(X) is the maximum value of the original time series, min(X) is the minimum value of the original time series, For time series The specific data after normalization at time t.
[0014] Step 1.2: Supplement missing dates in the time series and fill in missing values.
[0015] During the data collection process, some data and dates may be missing due to various reasons. The missing dates of the series are filled according to the data collection rules, and the series are interpolated using the linear interpolation method. The missing values are filled to obtain a continuous time series
[0016] Step 1.3, smooth the time series.
[0017] The Exponentially Weighted Moving Average (EWMA) algorithm is used to analyze the time series Smoothing is performed and the processed time series is recorded as The formula for this algorithm is:
[0018]
[0019] in For sequence The value at time t, For sequence The smoothed data value at time t, the coefficient β represents the rate at which the moving weight decreases.
[0020] Preferably, step 2 is as follows:
[0021] Step 2.1: Preprocess the time series A Fast Fourier Transform (FFT) is performed to determine candidate periodic sequences.
[0022] will sequence The sequence after fast Fourier transform is denoted as F, and the frequency components with a period of one day and one week in sequence F are denoted as f d and f wIf in sequence F, in addition to the DC component, f d or w is the maximum value in sequence F, then the sequence is identified as a candidate periodic time series.
[0023] Step 2.2, determine whether the candidate periodic sequence is a true periodic sequence based on the autocorrelation coefficient and determine the period size.
[0024] The autocorrelation coefficient reflects the degree of correlation of the time series at different times. When the autocorrelation coefficient value is in the interval (0.6, 0.8], it is a strong positive correlation, and when the autocorrelation coefficient value is in the interval (0.8, 1.0], it is an extremely strong positive correlation. If the time series is a periodic time series, when the autocorrelation data delay phase is equal to the period, the value of the autocorrelation coefficient will be very large. According to this principle, the Pearson correlation coefficient is used here to calculate the autocorrelation coefficient values of different phase differences. The principle formula is:
[0025]
[0026] in Representing time series The data delay with phase l, cov(·) and std(·) represent the covariance and standard deviation respectively. represents the autocorrelation coefficient value with a delay of l. Let the period lengths of the time series with a period of one day and one week be d and w respectively, and the autocorrelation coefficient with a lag phase of one day d be recorded as The autocorrelation coefficient of the lag phase of one week is recorded as At the same time, let the autocorrelation coefficient threshold be ρ ts ,pass and ρ ts The comparison determines whether the candidate periodic sequence is a true periodic sequence. The comparison formula is:
[0027]
[0028] The above formula can be used to divide time series into periodic time series and non-periodic time series. In order to determine the period size of periodic series, it is necessary to compare the periodic series. and The size of the comparison formula is:
[0029]
[0030] The period size of the periodic time series is determined by the above formula. After this step, the time series can be divided into a time series with a period of one day, a time series with a period of one week, and a non-periodic time series.
[0031] Step 2.3, classify the non-periodic time series.
[0032] The non-periodic time series is divided into multi-zero non-periodic series and non-multi-zero non-periodic series. The two time series have different data characteristics. The ratio of the number of zero values in the time series is denoted as n, and the threshold of the number of zero values is denoted as n ts , according to the following formula, the non-periodic time series can be divided into multi-zero non-periodic series and non-multi-zero non-periodic series.
[0033]
[0034] Preferably, step 3 is as follows:
[0035] Four different anomaly marking methods are designed for four types of time series. The time series needs to match the corresponding marking method to the normalized time series according to the classification results. Mark outliers.
[0036] If the time series is determined to be a time series with a period of one day, the abnormal points are marked according to the process of method 1. The specific process of method 1 is:
[0037] Process (1.1): Take the values of the time series at the same time every day, perform anomaly detection on all the values at each time, and mark the outliers of the time series and some contextual anomalies. For example, select the time series All data at 0 o'clock Then, the N-sigma criterion is used to detect abnormal points in the sequence. The N value in the N-sigma criterion is an adjustable coefficient. The criterion determines the normal value interval according to the size of the N value with a specific probability, and marks the values outside the interval as abnormal data. Its expression is:
[0038]
[0039] where μ and σ are the mean and standard deviation of the corresponding sequence, respectively, and x i is the value corresponding to the time series at time i.
[0040] Process (1.2): Perform differential operation on the time series data and detect the contextual anomalies of the original time series through differential data. The sequence after difference is recorded as The difference formula is:
[0041]
[0042] in Represents the difference sequence The tth data of Representing time series To determine whether the value of the differential data at a certain moment is an abnormal value, it is necessary to determine the range of the value at the current moment through the statistical information of a certain length window before the moment. Here, let the length of the sliding window be s0, and the data of the historical window at this time is Then use the boxplot method to determine that the normal data range at the current moment is (Q1-k*IQR, Q3+k*IQR), where Q3 and Q1 are data The upper and lower quartiles of the data, IQR is Q3-Q1, and k is an adjustable coefficient. When the current value exceeds the normal data range, the value can be marked as an outlier. Some periodic data will jump at the same time every day, but from the overall data point of view, it is not an abnormality, and the boxplot algorithm can also detect such data. In order to avoid such misjudgment, the N-sigma criterion at the same time is also used for the differential data to eliminate the above misjudgment, and the abnormal points detected by the boxplot and the N-sigma algorithm at the same time are used as the final abnormal points of the differential data.
[0043] Process (1.3): Correct the differential data labels. After process (1.2), the differential sequence can be obtained. The tag sequence make The tag sequence The value at time t, and The purpose of differential analysis is to find points that do not conform to the changing trend. After differential analysis of this type of data, there will be one positive and one negative abnormal differential data point. Will be and There is a positive and a negative abnormal differential data in the data. If only one of the abnormal differential data is detected, the label sequence is corrected according to the positive and negative relationship before and after this point. To determine the exact location of the abnormal differential data corresponding to the original data.
[0044] Process (1.4): Resample the time series and detect the abnormal points of the resampled time series. The upper quartile of each day, to get the quantile time series Then the quantile time series Use sliding window combined with boxplot method to detect anomalies. If a certain moment is a normal point, the label is recorded as 0, and if it is an abnormal point, the label is recorded as 1. At the same time, the value of the abnormal point is removed from the sequence. In order to prevent the initial window from containing jump points, a window of the same size is also needed to slide from the end of the sequence to the beginning of the sequence to complete the reverse anomaly detection and obtain the forward label sequence. and reverse marker sequence
[0045] Process (1.5): Determine the time series jump point position based on the marked 0 and 1 sequence. The length of the longest abnormal sequence in is denoted as l f and l b , the specified abnormal sequence length threshold is recorded as l ts First, l f , l b Following the threshold l ts For comparison, if both are less than the threshold, it means that there is no long-term abnormal value jump in the time series, and the marked sequence and Add them together to form a comprehensive label, change labels greater than 0 to 1, and then differentiate the comprehensive labels to locate the abnormal jump point through the position of 1 and -1 values. f , l b At least one of them is greater than the threshold l ts , it means that there is a long jump in the original time series. The marker sequence with a longer jump time is taken as the main marker sequence. The sequence is differentiated, and the jump point is located by 1 and -1. The jump time when the longest jump occurs is recorded as t b , and then the difference data t for another label sequence b The values near the time are reset to zero to prevent the conflict between the two marking results and cause misjudgment.
[0046] If the time series is determined to be a time series with a cycle of one week, the abnormal points need to be marked according to the process of method 2. The specific process of method 2 is:
[0047] Process (2.1): Splice the time series with a cycle of one week into a weekday time series and a weekend time series. Since the data characteristics of weekdays and weekends are different, it is necessary to splice all the weekday data of the time series together to form a weekday series. All weekend data are spliced together to form a weekend time series Then, the sequences and Mark the exceptions.
[0048] Process (2.2-2.4): Sequence and sequence Perform the same operations as in the process (1.1-1.3) to obtain the time series outlier data and contextual anomaly data.
[0049] Process (2.5): Detect the jump point of the time series with a cycle of one week. Resample the weekday time series and weekend time series, and take the upper quartile of each day to obtain the weekday quantile time series and weekend quantile time series Then, the boxplot algorithm is used to detect anomalies in the two quantile time series as a whole. The anomaly point is the point where the data jumps.
[0050] If the time series is determined to be a non-multi-zero non-periodic time series, the abnormal points are marked according to the process of method 3. The specific process of method 3 is:
[0051] Process (3.1): Detect sequence outliers.
[0052] For the sequence If you want to judge a point Whether it is an outlier, we need to use the window data with a length of s1 before and after the point Perform statistical judgment. Use the N-sigma criterion to detect anomalies on the window data. If If it is an outlier, mark it as 1 and replace it with the mean of all values at the same time of day in the window data. After using this method to mark once, some time series still have some outliers. In order to mark the outliers more perfectly, the series Use this method to do the second marking.
[0053] Process (3.2): Resample the time series and detect outliers in the resampled time series.
[0054] Take time series The median of each day, to get the median time series Then for the median time series Use sliding window combined with boxplot method to detect anomalies. If a certain moment is a normal point, the label is recorded as 0, and if it is an abnormal point, the label is recorded as 1. At the same time, the value of the abnormal point is removed from the sequence. In order to prevent the initial window from containing jump points, a window of the same size is also needed to slide from the end of the sequence to the beginning of the sequence to complete the reverse anomaly detection and obtain the forward label sequence. and reverse marker sequence
[0055] Process (3.3): Determine the time series jump point position based on the marked 0 and 1 sequences. This process is the same as the operation of process (1.5) in method 1.
[0056] If the time series is determined to be a multi-zero non-periodic time series, the abnormal points are marked according to the process of method 4. The specific process of method 4 is:
[0057] Process (4.1): Detect outliers in the sequence. The operation is the same as that of process (3.1).
[0058] Process (4.2): Detect sequence jump value.
[0059] Resample the time series after removing outliers to obtain the median time series for each day and daily mean time series First judge Is the sequence all zero? If all zero means that the original time series has no jump point, no jump point detection is required. Otherwise, The sequence uses a sliding window combined with a boxplot method for anomaly detection. When a point in the mean time series is detected as an anomaly, it is necessary to return to the day of the mean point and use the box plot algorithm to remove the outliers of the day and then calculate the mean of the day. Then, the new mean is compared with the threshold of the mean time series. If it is still greater than the threshold, the point is determined to be an outlier and is removed from the mean sequence. It is also necessary to perform two anomaly detections on the mean sequence to obtain a positive label sequence. and reverse marker sequence
[0060] Process (4.3): Determine the time series jump point position according to the marked 0 and 1 sequence. This process is exactly the same as the operation of process (1.5) in method 1, and will not be repeated here.
[0061] The beneficial effects of the present invention are as follows:
[0062] The present invention provides a method for anomaly marking of large-scale time series, which can accurately and efficiently mark outliers, contextual anomalies, and jump values of time series. While providing a set of process-oriented marking algorithms, the present invention also proposes a method for locating abnormal jump points through 0 and 1 marking sequences. This method is superior to other algorithms in terms of comprehensive performance of efficiency and accuracy, and can achieve a satisfactory effect in large-scale time series. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 is a flow chart of an embodiment of the present invention;
[0064] Figure 2 This is a flow chart of a time series preprocessing algorithm according to an embodiment of the present invention;
[0065] Figure 3 This is a flow chart of a time series classification algorithm according to an embodiment of the present invention;
[0066] Figure 4 Flowchart of the algorithm for labeling a time series with a period of one day. DETAILED DESCRIPTION
[0067] The method disclosed in the present invention can automatically classify time series, and then quickly and accurately and efficiently mark abnormal data for each type of time series. In particular, the method has good detection performance for abnormal jump points of time series, and is suitable for data analysis and abnormal marking of large-scale time series.
[0068] The data set of this algorithm embodiment is base station data of certain areas collected by a telecom operator for a period of time. The data is collected every hour and includes key performance indicator data (KPI) and data (counter) used by the communication system to collect statistics on the use of equipment and system resources.
[0069] The technical solution provided by the present invention is further described below in conjunction with the accompanying drawings and embodiments:
[0070] The time series data anomaly marking algorithm disclosed in the present invention specifically includes the following steps:
[0071] Step 1: Preprocessing of time series data;
[0072] Step 1.1, normalize time series data.
[0073] Different business data have different data ranges. The purpose of normalization is to eliminate the dimensional differences between different data and facilitate data comparison and joint processing. Let the original time series X = {x1, x2, ..., x N}, where N is the length of the sequence, x t is the specific value of the sequence X at a certain moment, and t∈{1, 2, …, N}. The original time series X is normalized using the Min-Max method, and the original data points are mapped to the interval [0, 1]. The normalized time series is recorded as The calculation formula is:
[0074]
[0075] Where max(X) is the maximum value of the original time series, min(X) is the minimum value of the original time series, For time series The specific data after normalization at time t.
[0076] Step 1.2: Supplement missing dates in the time series and fill in missing values.
[0077] During the data collection process, some data and dates may be missing due to various reasons. The missing dates of the series are filled according to the data collection rules, and the series are interpolated using the linear interpolation method. Fill the missing large values to obtain a continuous time series
[0078] Step 1.3, smooth the time series.
[0079] In order to reduce the impact of data spike fluctuations and outliers when performing step 2-time series classification, the Exponentially Weighted Moving Average (EWMA) algorithm is used to classify the time series. Smoothing is performed and the processed time series is recorded as The algorithm formula is:
[0080]
[0081] in For sequence The value at time t, For sequence The smoothed data value at time t, the coefficient β represents the rate at which the moving weight decreases.
[0082] Step 2: Time series classification;
[0083] Step 2.1: Preprocess the time series A Fast Fourier Transform (FFT) is performed to determine candidate periodic sequences.
[0084] In actual data, many business data are time series with a cycle of one day or one week. For a periodic time series, after fast Fourier transform, in addition to the DC component, there is generally a large frequency component in the frequency domain. By analyzing this frequency component, it can be determined whether the time series is a candidate periodic series. The sequence after fast Fourier transform is denoted as F, and the frequency components with a period of one day and one week in sequence F are denoted as f d and f w If in sequence F, in addition to the DC component, f d or w is the maximum value in sequence F, then the sequence is identified as a candidate periodic time series.
[0085] Step 2.2, determine whether the candidate periodic sequence is a true periodic sequence based on the autocorrelation coefficient and determine the period size.
[0086] The autocorrelation coefficient reflects the degree of correlation of the time series at different times. When the autocorrelation coefficient value is in the interval (0.6, 0.8], it is a strong positive correlation, and when the autocorrelation coefficient value is in the interval (0.8, 1.0], it is an extremely strong positive correlation. If the time series is a periodic time series, when the autocorrelation data delay phase is equal to the period, the value of the autocorrelation coefficient will be very large. According to this principle, the Pearson correlation coefficient is used here to calculate the autocorrelation coefficient values of different phase differences. The principle formula is:
[0087]
[0088] in Representing time series The data delay with phase l, cov(·) and std(·) represent the covariance and standard deviation respectively. represents the autocorrelation coefficient value with a delay of l. Let the period lengths of the time series with a period of one day and one week be d and w respectively, and the autocorrelation coefficient with a lag phase of one day d be recorded as The autocorrelation coefficient of the lag phase of one week is recorded as At the same time, let the autocorrelation coefficient threshold be ρ ts ,pass and ρ ts The comparison determines whether the candidate periodic sequence is a true periodic sequence. The comparison formula is:
[0089]
[0090] The above formula can be used to divide time series into periodic time series and non-periodic time series. In order to determine the period size of periodic series, it is necessary to compare the periodic series. and The comparison formula is:
[0091]
[0092] The period size of the periodic time series is determined by the above formula. After this step, the time series can be divided into a time series with a period of one day, a time series with a period of one week, and a non-periodic time series.
[0093] Step 2.3, classify the non-periodic time series.
[0094] The non-periodic time series is divided into multi-zero non-periodic series and non-multi-zero non-periodic series. The two time series have different data characteristics. The ratio of the number of zero values in the time series is denoted as n, and the threshold of the number of zero values is denoted as n ts , according to the following formula, the non-periodic time series can be divided into multi-zero non-periodic series and non-multi-zero non-periodic series.
[0095]
[0096] Step 3: Mark abnormal time series data.
[0097] Four different anomaly marking methods are designed for four types of time series. The time series needs to match the corresponding marking method to the normalized time series according to the classification results. Mark outliers.
[0098] If the time series is determined to be a time series with a period of one day, the abnormal points are marked according to the process of method 1. The specific process of method 1 is:
[0099] Process (1.1): Take the values of the time series at the same time every day, perform anomaly detection on all the values at each time, and mark the outliers of the time series and some contextual anomalies. For example, select the time series All data at 0 o'clock Then, the N-sigma criterion is used to detect abnormal points in the sequence. The N value in the N-sigma criterion is an adjustable coefficient. The criterion determines the normal value interval according to the size of the N value with a specific probability, and marks the values outside the interval as abnormal data. Its expression is:
[0100]
[0101] where μ and σ are the mean and standard deviation of the corresponding sequence, respectively, and x i is the value corresponding to the time series at time i.
[0102] Process (1.2): Perform differential operation on the time series data and detect the contextual anomalies of the original time series through differential data. The sequence after difference is recorded as The difference formula is:
[0103]
[0104] in Represents the difference sequence The tth data of Representing time series To determine whether the value of the differential data at a certain moment is an abnormal value, it is necessary to determine the range of the value at the current moment through the statistical information of a certain length window before the moment. Here, let the length of the sliding window be s0, and the data of the historical window at this time is Then use the boxplot method to determine that the normal data range at the current moment is (Q1-k*IQR, Q3+k*IQR), where Q3 and Q1 are data The upper and lower quartiles of the data, IQR is Q3-Q1, and k is an adjustable coefficient. When the current value exceeds the normal data range, the value can be marked as an outlier. Some periodic data will jump at the same time every day, but from the overall data point of view, it is not an abnormality, and the boxplot algorithm can also detect such data. In order to avoid such misjudgment, the N-sigma criterion at the same time is also used for the differential data to eliminate the above misjudgment, and the abnormal points detected by the boxplot and the N-sigma algorithm at the same time are used as the final abnormal points of the differential data.
[0105] Process (1.3): Correct the differential data labels. After process (1.2), the differential sequence can be obtained. The tag sequence make The tag sequence The value at time t, and The purpose of differential analysis is to find points that do not conform to the changing trend. After differential analysis of this type of data, there will be one positive and one negative abnormal differential data point. Will be and There is a positive and a negative abnormal differential data in the data. If only one of the abnormal differential data is detected, the label sequence is corrected according to the positive and negative relationship before and after this point. To determine the exact location of the abnormal differential data corresponding to the original data.
[0106] Process (1.4): Resample the time series and detect the abnormal points of the resampled time series. The upper quartile of each day, to get the quantile time series Then the quantile time series Use sliding window combined with boxplot method to detect anomalies. If a certain moment is a normal point, the label is recorded as 0, and if it is an abnormal point, the label is recorded as 1. At the same time, the value of the abnormal point is removed from the sequence. In order to prevent the initial window from containing jump points, a window of the same size is also needed to slide from the end of the sequence to the beginning of the sequence to complete the reverse anomaly detection and obtain the forward label sequence. and reverse marker sequence
[0107] Process (1.5): Determine the time series jump point position based on the marked 0 and 1 sequence. The length of the longest abnormal sequence in is denoted as l f and l b , the specified abnormal sequence length threshold is recorded as l ts First, lf , l b Following the threshold l ts For comparison, if both are less than the threshold, it means that there is no long-term abnormal value jump in the time series, and the marked sequence and Add them together to form a comprehensive label, change labels greater than 0 to 1, and then differentiate the comprehensive labels to locate the abnormal jump point through the position of 1 and -1 values. f , l b At least one of them is greater than the threshold l ts , it means that there is a long jump in the original time series. The marker sequence with a longer jump time is taken as the main marker sequence. The sequence is differentiated, and the jump point is located by 1 and -1. The jump time when the longest jump occurs is recorded as t b , and then the difference data t for another label sequence b The values near the time are reset to zero to prevent the conflict between the two marking results and cause misjudgment.
[0108] If the time series is determined to be a time series with a cycle of one week, the abnormal points need to be marked according to the process of method 2. The specific process of method 2 is:
[0109] Process (2.1): Splice the time series with a cycle of one week into a weekday time series and a weekend time series. Since the data characteristics of weekdays and weekends are different, it is necessary to splice all the weekday data of the time series together to form a weekday series. All weekend data are spliced together to form a weekend time series Then, the sequences and Mark the exceptions.
[0110] Process (2.2-2.4): Sequence and sequence Perform the same operations as in the process (1.1-1.3) to obtain the time series outlier data and contextual anomaly data.
[0111] Process (2.5): Detect the jump point of the time series with a cycle of one week. Resample the weekday time series and weekend time series, and take the upper quartile of each day to obtain the weekday quantile time series and weekend quantile time series Then, the boxplot algorithm is used to detect anomalies in the two quantile time series as a whole. The anomaly point is the point where the data jumps.
[0112] If the time series is determined to be a non-multi-zero non-periodic time series, the abnormal points are marked according to the process of method 3. The specific process of method 3 is:
[0113] Process (3.1): Detect sequence outliers.
[0114] For the sequence If you want to judge a point Whether it is an outlier, we need to use the window data with a length of s1 before and after the point Perform statistical judgment. Use the N-sigma criterion to detect anomalies on the window data. If If it is an outlier, mark it as 1 and replace it with the mean of all values at the same time of day in the window data. After using this method to mark once, some time series still have some outliers. This is because the window data for detecting these outliers contains many abnormal values, which affects the judgment of the N-sigma criterion. In order to mark the outliers more perfectly, the sequence This method is used for the second labeling to mark as many outliers as possible for this type of time series.
[0115] Process (3.2): Resample the time series and detect outliers in the resampled time series.
[0116] Take time series The median of each day, to get the median time series Then for the median time series Use sliding window combined with boxplot method to detect anomalies. If a certain moment is a normal point, the label is recorded as 0, and if it is an abnormal point, the label is recorded as 1. At the same time, the value of the abnormal point is removed from the sequence. In order to prevent the initial window from containing jump points, a window of the same size is also needed to slide from the end of the sequence to the beginning of the sequence to complete the reverse anomaly detection and obtain the forward label sequence. and reverse marker sequence
[0117] Process (3.3): Determine the time series jump point position based on the marked 0 and 1 sequences.
[0118] Determine the time series jump point position based on the marked 0 and 1 sequence. The length of the longest abnormal sequence in is denoted as l f and l b , the specified abnormal sequence length threshold is recorded as l ts First, l f , l b Following the threshold lts For comparison, if both are less than the threshold, it means that there is no long-term abnormal value jump in the time series, and the marked sequence and Add them together to form a comprehensive label, change labels greater than 0 to 1, and then differentiate the comprehensive labels to locate the abnormal jump point through the position of 1 and -1 values. f , l b At least one of them is greater than the threshold l ts , it means that there is a long jump in the original time series. The marker sequence with a longer jump time is taken as the main marker sequence. The sequence is differentiated, and the jump point is located by 1 and -1. The jump time when the longest jump occurs is recorded as t b , and then the difference data t for another label sequence b The values near the time are reset to zero to prevent the conflict between the two marking results and cause misjudgment.
[0119] If the time series is determined to be a multi-zero non-periodic time series, the abnormal points are marked according to the process of method 4. The specific process of method 4 is:
[0120] Process (4.1): Detect outliers in the sequence.
[0121] For the sequence If you want to judge a point Whether it is an outlier, we need to use the window data with a length of s1 before and after the point Perform statistical judgment. Use the N-sigma criterion to detect anomalies on the window data. If If it is an outlier, mark it as 1 and replace it with the mean of all values at the same time of day in the window data. After using this method to mark once, some time series still have some outliers. This is because the window data for detecting these outliers contains many abnormal values, which affects the judgment of the N-sigma criterion. In order to mark the outliers more perfectly, the sequence This method is used for the second labeling to mark as many outliers as possible for this type of time series.
[0122] Process (4.2): Detect sequence jump value.
[0123] Resample the time series after removing outliers to obtain the median time series for each day and daily mean time series First judge Is the sequence all zero? If it is all zero, it means that the original time series has no jump point, and no jump point detection is required. Otherwise, The sequence uses a sliding window combined with a boxplot method for anomaly detection. When a point in the mean time series is detected as an anomaly, it is necessary to return to the day of the mean point and use the box plot algorithm to remove the outliers of the day and then calculate the mean of the day. Then, the new mean is compared with the threshold of the mean time series. If it is still greater than the threshold, the point is determined to be an outlier and is removed from the mean sequence. It is also necessary to perform two anomaly detections on the mean sequence to obtain a positive label sequence. and reverse marker sequence
[0124] Process (4.3): Determine the time series jump point position based on the marked 0 and 1 sequences.
[0125] Determine the time series jump point position based on the marked 0 and 1 sequence. The length of the longest abnormal sequence in is denoted as l f and l b , the specified abnormal sequence length threshold is recorded as l ts First, l f , l b Following the threshold l ts For comparison, if both are less than the threshold, it means that there is no long-term abnormal value jump in the time series, and the marked sequence and Add them together to form a comprehensive label, change labels greater than 0 to 1, and then differentiate the comprehensive labels to locate the abnormal jump point through the position of 1 and -1 values. f , l b At least one of them is greater than the threshold l ts , it means that there is a long jump in the original time series. The marker sequence with a longer jump time is taken as the main marker sequence. The sequence is differentiated, and the jump point is located by 1 and -1. The jump time when the longest jump occurs is recorded as t b , and then the difference data t for another label sequence b The values near the time are reset to zero to prevent the conflict between the two marking results and cause misjudgment.
[0126] Figure 1 The algorithm flow chart provided by the embodiment of the present invention. Figure 1 As shown, the base station data set is first preprocessed to obtain preprocessed time series data; then the preprocessed time series are classified to obtain periodic time series with a period of one day or one week, and non-multi-zero or multi-zero non-periodic time series; finally, the four types of time series are anomaly marked to obtain anomaly labels of all time series.
[0127] Figure 2Schematic diagram of the time series preprocessing process provided by this application. Specifically, execute steps 1.1-1.3, and execute the base station time series X according to formula (1) to obtain the normalized time series Then, according to the rule of collecting base station data every hour, the missing dates in the base station data set are filled, and the missing values are linearly interpolated to obtain a continuous time series. Finally, according to formula (2), the EWMA algorithm is used to analyze the sequence Smoothing is performed to obtain a smoothed sequence In this embodiment, the weighted decrease rate coefficient β of the EWMA algorithm is set to 0.25.
[0128] Figure 3 This is a flowchart of the time series classification algorithm of this application. Visual analysis of the base station data in the embodiment shows that the period size of the periodic series in the KPI data and counter data is one day or one week, that is, d = 24, w = 168. Steps 2.1-2.3 are performed on the preprocessed time series. First, the smoothed series Perform fast Fourier transform and use the frequency domain component f d and f w The size of the candidate period time series is determined. Then the autocorrelation coefficient of the candidate period time series is calculated using formula (3). In this embodiment, the autocorrelation coefficient threshold ρ ts Set to 0.6, according to formula (4) and threshold ρ ts Compare the alternative periodic sequence to determine whether it is a true periodic sequence, and then compare the correlation coefficients of the alternative sequences and The size of , thereby determining the period value of the periodic sequence. Finally, the zero value ratio threshold n in step 2.3 is ts It is set to 0.2. When the zero value ratio n in the base station non-periodic time series data is greater than 0.2, the sequence is judged as a multi-zero non-periodic sequence, otherwise it is judged as a non-multi-zero non-periodic time series.
[0129] Figure 4 This is a flow chart of the marking algorithm for a time series with a period of one day in this application. If the base station time series is determined to be a time series with a period of one day, the time series will execute the process of method 1 in step 3 to perform abnormal marking.
[0130] Specifically: Execute process (1.1) to process (1.3) to mark the outliers and contextual anomalies of the time series with a period of one day. In this embodiment, the N value of the N-sigma algorithm in process (1.1) and process (1.2) is set to 3.76, and the k value of the boxplot algorithm in process (1.2) is set to 2.0. Then execute process (1.4) and process (1.5) to mark the abnormal jump points of the time series, set the sliding window length s0 in process (1.4) to 35, and the k value of the boxplot algorithm to 2.66. After two anomaly detections with the sliding window, the marked sequence can be obtained. and The l of process (1.5) ts Set to 14, through l f , l b With threshold l ts By comparing the two marker sequences, the jumping point of the time series can be located using the difference sequence of the two marker sequences.
[0131] If the base station time series is determined to be a time series with a period of one week, the time series will execute the process of method 2 in step 3 to mark the abnormality.
[0132] Specifically: Execute process (2.1) to divide the base station time series into working day time series and weekend time series Then execute process (2.2-2.4) to find outliers and context anomalies in the time series with a cycle of one week. This process is the same as process (1.1-1.3) and the parameter settings are also the same. Finally, set the k value of the boxplot algorithm in process (2.5) to 2.52 to complete the marking of the jump points of this type of time series.
[0133] If the base station time series is determined to be a non-multi-zero non-periodic time series, the time series will execute the process of method 3 in step 3 to mark the abnormality.
[0134] Specifically: Execute process (3.1) to mark the outliers of the non-multi-zero non-periodic time series, set the size of the sliding window s1 of the process to 168, and the N value to 3.66. Execute process (3.2), set the sliding window length to 21, and the k value of the boxplot algorithm to 5.66. Finally, execute process (3.3) to set l ts Set it to 14 to complete the marking of the jump points of this type of time series.
[0135] If the base station time series is determined to be a multi-zero non-periodic time series, the time series will execute the process of method 4 in step 3 to mark the abnormality.
[0136] Specifically: Execute process (4.1) to mark the outliers of the non-periodic time series with multiple zeros. This process is the same as process (3.1), and the parameter settings are also the same. Then execute process (4.2) to detect anomalies on the resampled time series. The length of the sliding window of this method is set to 21, the box plot coefficient k value is set to 7.66, and the boxplot algorithm coefficient on the day of the mean time series anomaly is set to 2.0. Finally, execute process (4.3) to complete the location of the abnormal jump point. This process is the same as process (1.5), and the parameter settings are also the same.
[0137] The present invention has a good performance in marking large-scale time series jump point. It can quickly and efficiently complete the marking while also achieving a high accuracy rate. Table 1 is a summary of the data of each KPI jump point selected and counted in this embodiment, where the precision and recall of the sequence marking are represented by P and R respectively, and the evaluation formulas are:
[0138]
[0139]
[0140] TP is a true positive example, FP is a false positive example, and FN is a false negative example. It can be seen from the table that the algorithm has a good effect on marking the jump points of base station KPI data, with a comprehensive P value of 90.3% and a comprehensive R value of 92.9%.
[0141] Table 1
[0142] KPI TP FN FP P R KPI1 637 48 64 90.9% 93.0% KPI2 345 20 35 90.8% 94.5% KPI3 386 32 44 89.8% 92.3% KPI4 403 36 48 89.4% 91.8%
[0143] In summary, the present invention provides a method suitable for large-scale time series anomaly marking, which can accurately and efficiently mark outliers, contextual anomalies, and jump values of time series. While providing a set of process-oriented marking algorithms, the present invention also proposes a method for locating abnormal jump points through 0, 1 marking sequences. This method is superior to other algorithms in terms of comprehensive performance of efficiency and accuracy, and achieves a relatively satisfactory effect in the embodiment.
Claims
1. A method for marking anomalies in time series data, characterized in that: The following steps are involved: Step 1: Preprocessing of time series data; Step 2: Time series classification; Step 3: Mark abnormal time series data; Four different anomaly marking methods are designed for four types of time series. The time series needs to match the corresponding marking method to the normalized time series according to the classification results. Marking of outliers; If the time series is determined to be a time series with a period of one day, the abnormal points are marked according to the process of method 1; the specific process of method 1 is: Process (1.1): Take the values of the time series at the same time every day, perform anomaly detection on all values at each time, mark the outliers of the time series and some contextual anomalies; select the time series All data at 0 o'clock Then the N-sigma criterion is used to detect the outliers of the sequence; The N value in the N-sigma criterion is an adjustable coefficient. The criterion determines the normal value interval according to the set probability based on the N value, and marks the values outside the interval as abnormal data. Its expression is: where μ and σ are the mean and standard deviation of the corresponding sequence, respectively, and x i is the value corresponding to the time series at time i; Process (1.2): Perform differential operation on time series data, detect contextual anomalies of the original time series through differential data; convert the time series The sequence after difference is recorded as The difference formula is: in Represents the difference sequence The tth data of Representing time series To determine whether the value of the differential data at a certain moment is an abnormal value, it is necessary to determine the range of the value at the current moment through the statistical information of a certain length window before the moment; here, the length of the sliding window is s0, and the data of the historical window at this time is Then use the boxplot method to determine that the normal data range at the current moment is (Q1-k*IQR, Q3+k*IQR), where Q3 and Q1 are data The upper and lower quartiles, IQR is Q3-Q1, k is an adjustable coefficient; when the current moment value exceeds the normal data range, the value can be marked as an outlier; some periodic data will jump at the same moment every day, but from the overall data point of view, it is not an abnormality, and the boxplot algorithm can also detect such data. In order to avoid such misjudgment, the N-sigma criterion at the same moment is also used for the differential data to eliminate the above misjudgment, and the abnormal points detected by the boxplot and the N-sigma algorithm at the same moment are used as the final abnormal points of the differential data; Process (1.3): Correct the differential data labels; after process (1.2), the differential sequence can be obtained The tag sequence make The tag sequence The value at time t, and The purpose of differencing is to find points that do not conform to the changing trend. After differencing this type of data, one positive and one negative abnormal differential data point will appear; For a contextually abnormal data Will be and There is a positive and a negative abnormal differential data in the data. If only one of the abnormal differential data is detected, the label sequence is corrected according to the positive and negative relationship before and after this point. To determine the exact location of the abnormal differential data corresponding to the original data; Process (1.4): Resample the time series and detect the abnormal points of the resampled time series; take the time series The upper quartile of each day, to get the quantile time series Then the quantile time series Use sliding window combined with boxplot method to detect anomalies. If a certain moment is a normal point, the label is recorded as 0, and if it is an abnormal point, the label is recorded as 1. At the same time, the value of the abnormal point is removed from the sequence. To avoid contaminating the data of the next sliding window, we need to use a window of the same size to slide from the end of the sequence to the beginning of the sequence to complete the reverse anomaly detection and obtain the forward label sequence. and reverse marker sequence Process (1.5): Determine the time series jump point position based on the marked 0 and 1 sequence; The length of the longest abnormal subsequence in is denoted as l f and l b , the specified abnormal subsequence length threshold is recorded as l ts ; First, l f , l b Following the threshold l ts For comparison, if both are less than the threshold, it means that there is no long-term abnormal value jump in the time series, and the marked sequence and Add them together to form a comprehensive label, change labels greater than 0 to 1, and then differentiate the comprehensive labels to locate the abnormal jump point through the position of 1 and -1 values; if l f , l b At least one of them is greater than the threshold l ts , it means that there is a long jump in the original time series. The marker sequence with a longer jump time is taken as the main marker sequence, and the sequence is differentiated. The jump point is located by 1 and -1, and the jump time with the longest jump time is recorded as t b , and then the difference data t for another label sequence b The values near the time are reset to zero to prevent the conflict between the two marking results and cause misjudgment.
2. A time series data anomaly marking method according to claim 1, characterized in that: Step 1 is as follows: Step 1.1, normalize time series data; Let the original time series X = {x1, x2, ..., x N }, where N is the length of the sequence, x t is the specific value of the sequence X at a certain moment, and t∈{1, 2, …, N}; the original time series X is normalized using the Min-Max method, and the original data points are mapped to the interval [0, 1]; the normalized time series is recorded as The calculation formula is: Where max(X) is the maximum value of the original time series, min(X) is the minimum value of the original time series, For time series The specific data after normalization at time t; Step 1.2, supplement the missing dates of the time series and fill in the missing values; During the data collection process, some data and dates may be missing due to various reasons. The missing dates of the series are filled according to the data collection rules, and the series are interpolated using the linear interpolation method. Fill the missing values to obtain a continuous time series Step 1.3, smooth the time series; The moving weighted average algorithm is used to analyze the time series Smoothing is performed and the processed time series is recorded as The formula for this algorithm is: in For sequence The value at time t, For sequence The smoothed data value at time t, the coefficient β represents the rate at which the moving weight decreases.
3. A time series data anomaly marking method according to claim 2, characterized in that: Step 2 is as follows: Step 2.1: Preprocess the time series performing a fast Fourier transform (FFT) to determine a candidate periodic sequence; will sequence The sequence after fast Fourier transform is denoted as F, and the frequency components with a period of one day and one week in sequence F are denoted as f d and f w ; If in sequence F, in addition to the DC component, f d or w is the maximum value in sequence F, then the sequence is identified as a candidate periodic time series; Step 2.2, judging whether the candidate periodic sequence is a true periodic sequence based on the autocorrelation coefficient and determining the periodicity; The autocorrelation coefficient reflects the degree of correlation of the time series at different times. When the autocorrelation coefficient value is in the interval (0.6, 0.8], it is a strong positive correlation, and when the autocorrelation coefficient value is in the interval (0.8, 1.0], it is an extremely strong positive correlation. If the time series is a periodic time series, when the autocorrelation data delay phase is equal to the period, the value of the autocorrelation coefficient will be very large. Based on this principle, the Pearson correlation coefficient is used here to calculate the autocorrelation coefficient values of different phase differences. The principle formula is: in Representing time series The data delay with phase l, cov(·) and std(·) represent the covariance and standard deviation respectively. represents the autocorrelation coefficient value with a delay of l; let the period lengths of the time series with a period of one day and one week be d and w respectively, and record the autocorrelation coefficient with a lag phase of one day as The autocorrelation coefficient of the lag phase of one week is recorded as At the same time, let the autocorrelation coefficient threshold be ρ ts ,pass and ρ ts The comparison determines whether the candidate periodic sequence is a true periodic sequence. The comparison formula is: The above formula can be used to divide time series into periodic time series and non-periodic time series. In order to determine the period size of periodic series, it is necessary to compare periodic series. and The size of the comparison formula is: The period size of the periodic time series is determined by the above formula; after this step, the time series can be divided into a time series with a period of one day, a time series with a period of one week, and a non-periodic time series; Step 2.3, classify the non-periodic time series; The non-periodic time series is divided into multi-zero non-periodic series and non-multi-zero non-periodic series. The two time series have different data characteristics. The ratio of zero values in the time series is recorded as n, and the threshold of zero value ratio is recorded as n ts , the non-periodic time series can be divided into multi-zero non-periodic series and non-multi-zero non-periodic series according to the following formula; 4. A time series data anomaly marking method according to claim 3, characterized in that: Step 3 specifically also includes: If the time series is determined to be a time series with a cycle of one week, the abnormal points need to be marked according to the process of method 2. The specific process of method 2 is: Process (2.1): Splice the time series with a cycle of one week into weekday time series and weekend time series; for time series Since the data characteristics of weekdays and weekends are different, it is necessary to splice all the weekday data of the time series together to form a weekday series. All weekend data are spliced together to form a weekend time series Then, the sequences and Mark abnormalities; Process (2.2-2.4): Sequence and sequence Perform the same operations as in the process (1.1-1.3) to obtain the time series outlier data and contextual anomaly data; Process (2.5): Detect the jump points of the time series with a cycle of one week; resample the weekday time series and weekend time series, and take the upper quartile of each day to obtain the weekday quantile time series and weekend quantile time series Then, the boxplot algorithm is used to detect anomalies in the two quantile time series as a whole. The anomaly point is the point where the data jumps. If the time series is determined to be a non-multi-zero non-periodic time series, the abnormal points are marked according to the process of method 3. The specific process of method 3 is: Process (3.1): Detect sequence outliers; For the sequence If you want to judge a point Whether it is an outlier, we need to use the window data with a length of s1 before and after the point Perform statistical judgment; use the N-sigma criterion to detect anomalies on the window data. If If it is an outlier, mark the value as 1 and replace it with the mean of all values at the same time of day in the window data; After using this method to mark once, some time series still have some outliers. In order to mark the outliers more perfectly, the series Use this method to do a second markup; Process (3.2): Resample the time series and detect outliers in the resampled time series; Take time series The median of each day, to get the median time series Then for the median time series Use sliding window combined with boxplot method to detect anomalies. If a certain moment is a normal point, the label is recorded as 0, and if it is an abnormal point, the label is recorded as 1. At the same time, the value of the abnormal point is removed from the sequence. To avoid contaminating the data of the next sliding window, we need to use a window of the same size to slide from the end of the sequence to the beginning of the sequence to complete the reverse anomaly detection and obtain the forward label sequence. and reverse marker sequence Process (3.3): Determine the time series jump point position based on the marked 0 and 1 sequences; this process is the same as the operation of process (1.5) in method 1; If the time series is determined to be a multi-zero non-periodic time series, the abnormal points are marked according to the process of method 4. The specific process of method 4 is: Process (4.1): Detect outliers in the sequence; the operation is the same as process (3.1); Process (4.2): Detect sequence jump value; Resample the time series after removing outliers to obtain the median time series for each day and daily mean time series First judge Is the sequence all zero? If it is all zero, it means that the original time series has no jump point, and no jump point detection is required. Otherwise, The sequence uses a sliding window combined with a boxplot method for anomaly detection. When a point in the mean time series is detected as an anomaly, it is necessary to return to the day of the mean point and use the box plot algorithm to remove the outliers of the day and then calculate the mean of the day. Then, the new mean is compared with the threshold of the mean time series. If it is still greater than the threshold, the point is determined to be an outlier and is removed from the mean sequence. It is also necessary to perform two anomaly detections on the mean sequence to obtain a positive label sequence. and reverse marker sequence Process (4.3): Determine the time series jump point position based on the marked 0 and 1 sequences; this process is exactly the same as the operation of process (1.5) in method 1, and will not be repeated here.
Citation Information
Patent Citations
Power consumer electricity consumption anomaly detection method based on machine learning
CN111695639A
Method and device for processing and detecting data acquired by industrial equipment
CN111967509A