Time sequence anomaly detection method based on parallel learning framework

Through the time sequence anomaly detection method based on the parallel learning framework, combined with the characteristics of real-time data and historical data, the limitations of traditional methods in dealing with complex anomaly and multi-dimensional correlations are solved, and energy data anomaly detection with high precision and low computing cost is achieved.

CN119961848AInactive Publication Date: 2025-05-09ZHEJIANG UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510444756.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-05-09
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional energy data anomaly detection methods are difficult to effectively deal with complex anomaly types and multidimensional correlations, and deep neural network methods have problems of long training time and dynamic processing challenges when dealing with spatiotemporal dependencies and real-time data.

Method used

The timing anomaly detection method based on the parallel learning framework is adopted. The model consists of a historical data detection module, a real-time data detection module, a cascade module and an evaluation module. Real-time streaming data is processed through preprocessing, exponential weighted moving average method, two-step smoothing method and automatic threshold setting based on the moment quantity method, and historical data features are extracted in combination with the classification regression tree method to identify abnormal patterns.

Benefits of technology

It significantly improves the accuracy of abnormal detection of multi-dimensional energy data, effectively reduces noise and highlights abnormal points, comprehensively explores the dependence between historical mode and real-time dynamics, and realizes abnormal detection with high accuracy and low computing cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961848A_ABST
    Figure CN119961848A_ABST
Patent Text Reader

Abstract

The invention discloses a time sequence anomaly detection method based on a parallel learning framework, which belongs to the technical field of time sequence anomaly detection methods, and comprises the following steps: firstly, a real-time data detection module processes real-time streaming data through preprocessing, an exponential weighted moving average method, a two-step smoothing method and automatic threshold setting based on a moment method; therefore, noise reduction and dynamic anomaly detection are realized. Next, after being cached, real-time streaming data is aligned with original historical data in a cascade module according to a corresponding time window and transmitted to a historical data detection model, and features are extracted from the historical data by adopting a classification regression tree method; according to the multi-dimensional energy consumption data abnormal detection method and system, the historical data is extracted, the abnormal mode in the historical data is recognized by combining path length calculation, abnormal score generation and threshold setting, and finally, the detection result is verified through the evaluation module, so that the accuracy of multi-dimensional energy consumption data abnormal detection is effectively improved, and the production safety and operation stability of enterprises are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a time series anomaly detection method, in particular to a time series anomaly detection method based on a parallel learning framework, and belongs to the technical field of time series anomaly detection methods. Background Art

[0002] With the continuous growth of industrial energy demand and the increasing complexity of energy systems, energy data anomaly detection has become an important part of energy management. Accurate and timely anomaly detection enables energy companies to quickly discover potential problems, optimize the allocation and configuration of energy resources, and effectively reduce energy losses. Many traditional anomaly detection methods, such as the Z-score method based on statistical rules, the Grubbs test, and the K-means clustering algorithm based on clustering, mainly focus on identifying significant outliers in the data.

[0003] However, these traditional methods show limitations in dealing with the complex anomaly types and multi-dimensional correlations inherent in energy data, and it is difficult to meet the needs of modern energy management for efficient and accurate anomaly detection. In order to explore the complex patterns in energy data anomaly detection, some deep neural network-based methods have gradually emerged, such as using autoencoders, generative adversarial networks, long short-term memory networks, and graph convolutional networks for anomaly detection.

[0004] These methods can usually only process a single data type and cannot fully capture the spatiotemporal dependencies and information flows between historical data and real-time streaming data. In addition, due to their complex network architecture, these methods usually require a long training time, making real-time updates and dynamic streaming data processing challenging. Energy data is also affected by many external factors such as seasonal changes, production plans, and market fluctuations, making the anomaly detection task of historical data and real-time data more complicated. Therefore, a time series anomaly detection method based on a parallel learning framework is proposed to solve the above problems. Summary of the invention

[0005] The main purpose of the present invention is to provide a time series anomaly detection method based on a parallel learning framework. The entire model consists of four modules, namely a historical data detection module, a real-time data detection module, a cascade module and an evaluation module. First, the real-time data detection module processes real-time streaming data through preprocessing, exponentially weighted moving average method, two-step smoothing method and automatic threshold setting based on moment method, thereby achieving noise reduction and dynamic anomaly detection. Then, after the real-time streaming data is cached, it is passed to the historical data detection model in the cascade module according to the corresponding time window alignment with the original historical data, and the classification regression tree method is used to extract features from the historical data, and the path length calculation, anomaly score generation and threshold setting are combined to identify abnormal patterns in the historical data. Finally, the detection results are verified by the evaluation module. Specifically comprising the following steps: Step 1, build a real-time data detection module, and use the streaming time series of real-time data as the input of the real-time data detection module. After data preprocessing, fluctuation extraction, two-step smoothing and automatic threshold setting based on the moment method, the abnormal conditions of all energy sources at different times are detected.

[0006] Step 2: The real-time streaming data after cache processing is aligned with the original historical data through the cascade module, and the length of the time window is dynamically adjusted according to the data sampling frequency and anomaly detection requirements (for the set detection time window, for example, 24 hours, if the real-time detection data is 1 hour, the historical data is the remaining 23 hours; the 1-hour data processed by the real-time detection module is aligned with the undetected 23-hour historical data to form a complete 24-hour data for the historical data detection model input; similarly, if the detection window is 2 hours and the real-time detection data is 10 minutes, the historical data is 110 minutes, and the 10-minute data processed by the real-time detection module is aligned with the undetected 110-minute historical data to form a complete 2-hour data). Subsequently, these data are concatenated in columns to form a unified feature vector, combining the long-term pattern of historical data with the short-term dynamic changes of real-time data, and passing it as an input matrix to the historical data detection model.

[0007] Step 3: The historical data detection model uses the classification and regression tree method to extract features from historical data, and combines path length calculation, anomaly score generation, and threshold setting to identify anomalies in historical data.

[0008] Step 4: In step 3, the anomaly score is generated and the threshold is set. In step 4, the anomaly score of each point is compared with the threshold. If it exceeds the threshold, it is considered an anomaly. If it does not exceed the threshold, it is not considered an anomaly. Then, the ratio of the true anomaly points in the obtained anomaly points to the total anomaly points in the input data is calculated.

[0009] Furthermore, in step 1, in order to reduce noise while retaining the normal mode as much as possible, a segmented filling method based on the length of the missing segment is proposed in the data preprocessing stage. The method includes two cases: (1) For short missing segments (less than five observation points), first-order linear interpolation is used for filling; (2) For long missing segments (greater than or equal to five observation points): use the value of the same time point in the previous cycle for estimation, and add an offset term. The offset is defined as: ; in, It represents the offset between two adjacent time periods, which is used to measure the difference in the average values ​​of the two time periods; is the current time period ( arrive time point), Indicates the length of a time period. It is the time series data data points; This is the previous time period ( arrive time point), Index variables, which represent specific moments in the time series; The length of a time segment, which is the window size used for segment calculations; At a point in time in a time series Observed value of Normalization factor for the mean calculation, used to ensure that the contributions of all points within a time period are averaged; The index of the current time point, indicating the reference time for calculating the offset.

[0010] Furthermore, in time series data, abnormal points usually show significant deviations from the trend of historical data. In order to quantitatively describe the degree of this deviation, we introduce the concept of "fluctuation eigenvalue" as an important feature to capture abnormal fluctuations in data. The fluctuation eigenvalue can be obtained by calculating the difference between the current point and the expected value of its historical data, and the data fluctuation eigenvalue is calculated using the exponentially weighted moving average. , specifically expressed as: ; ; in, Indicates the current time point The fluctuation value indicates the deviation between the current point and the expected value, which directly characterizes the fluctuation characteristics.

[0011] Represents a time series at a point in time The actual observed value of Indicates the expected value at the current time point; Represents the smoothing coefficient. The larger the value, the higher the weight of recent data and the stronger the dependence on recent data; the smaller the value, the greater the influence of historical data.

[0012] Indicates the length of the time window. Defines the length of historical data used to calculate the smoothed expected value (for example, the data of the last 5 or 10 points).

[0013] Represents a time series at a point in time historical observations.

[0014] Furthermore, a two-step smoothing mechanism is designed to make the outliers more prominent, while retaining the abnormal fluctuations and reducing the normal fluctuation residuals to near zero. The first step of smoothing extracts the fluctuation values ​​by continuous processing. To eliminate local noise, it can be expressed as: ; ; Indicates the increment of the current window standard deviation, indicating that when a new data point is introduced After that, the change in the sliding window standard deviation.

[0015] Indicates the current time point Fluctuation value of Indicates the moment before the current time point. Indicates the starting point of the window, and the window length is ; Current time point The processing result.

[0016] If the current point If the addition of leads to a significant increase in the window standard deviation, it is considered as a potential anomaly; otherwise, the fluctuation value is set to zero; The second step of smoothing removes periodic noise through period-based processing. Before operating on the data, a simple method is used to deal with the data drift problem: ; ; ; in, Indicates the current time point The corresponding maximum fluctuation value.

[0017] Indicates that the time series is The fluctuation value of (i.e. the fluctuation value two cycles ago).

[0018] Indicates the current time point The fluctuation value of .

[0019] It is expressed as the length of a time period; Indicates the current time point The fluctuation difference of Represents the maximum fluctuation value sequence within multiple periods; represents the number of cycles. If The current point is considered normal and its value is set to 0.

[0020] Furthermore, the smoothed fluctuation characteristics obtained in the double-step smoothing In , the moment estimation method is used to estimate the parameters, and the generalized Pareto distribution is used to set the dynamic threshold, which is specifically expressed as: ; in, is the initial threshold, and are the shape parameter and scale parameter of the generalized Pareto distribution, is the cumulative distribution function of the generalized Pareto distribution, which is used to describe the probability of the tail distribution. is the current fluctuation value Relative to threshold The excess part, that is , and The moment estimate of is derived as follows: ; in, and Represent the mean and variance of the samples respectively. The dynamic threshold is iteratively updated in the following way: ; in, is a dynamically updated threshold value, which adjusts the initial threshold value based on the distribution parameters calculated in real time. is the initial threshold, used as a baseline for the start of detection. Represents the risk factor, which is used to control the sensitivity of anomaly detection; Indicates that the initial threshold is exceeded points.

[0021] Furthermore, in step 2, after the real-time streaming data is cached, it is aligned with the original historical data according to the corresponding time window. Then the data is concatenated column by column to form a unified feature vector. The feature vector is passed to the historical data detection module as an input matrix.

[0022] Furthermore, in step 3, the historical data detection module uses a classification regression tree method to extract features from the historical data, and performs the following steps: (1) Feature selection: randomly select a feature ,in is the total number of features, Represents data points No. Features.

[0023] (2) Split point selection: For the selected features . Select a split point , so that its value is located in the feature The split point is between the minimum and maximum values ​​of Usually selected randomly and evenly distributed.

[0024] Recursive partitioning: For each node, the data set is divided into two subsets, and the partitioning is continued recursively for each subset until the stopping condition is met, the number of samples is 1, or the maximum tree depth is reached.

[0025] Each data point in the tree has a path length , represents the distance from the root node to the leaf node where the data point is located. , path length It is calculated based on the distance from the root node to the leaf node in each tree. exist The average path length in a tree is calculated as: ; in, represent In the The path length in the tree.

[0026] Anomaly score The average path length It is concluded that according to the path length and the size of the data set Normalize the relationship between: ; in, It is a data point The average path length among all trees in the forest.

[0027] is the normalization factor, defined as: ; constant According to the data set The path length calculation is resized to ensure that the anomaly scores are consistent across datasets of different sizes.

[0028] Once the anomaly score is calculated , you can classify data points as normal or abnormal by setting a threshold.

[0029] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: In an embodiment of the present application, it is characterized in that the entire model consists of four modules, namely a historical data detection module, a real-time data detection module, a cascade module and an evaluation module. First, the real-time data detection module processes real-time streaming data through preprocessing, exponentially weighted moving average method, double-step smoothing method and automatic threshold setting based on moment method, thereby realizing noise reduction and dynamic anomaly detection. Then, after the real-time streaming data is cached, it is passed to the historical data detection model in the cascade module according to the corresponding time window alignment with the original historical data, and the classification regression tree method is used to extract features from the historical data, and the abnormal patterns in the historical data are identified by combining path length calculation, abnormal score generation and threshold setting. Finally, the detection results are verified by the evaluation module. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 A schematic diagram of a framework of a time series anomaly detection method based on a parallel learning framework provided in an embodiment of the present invention.

[0031] Figure 2 Result graph of the hyperparameter grid search for the real-time data detection module.

[0032] Figure 3 Schematic diagram of the two-step smoothing mechanism.

[0033] Figure 4 This is the result diagram of anomaly detection of energy consumption time series data. DETAILED DESCRIPTION

[0034] In order to make the technical solution of the present invention more clear and specific to those skilled in the art, the present invention is further described in detail below in conjunction with embodiments and drawings, but the implementation manner of the present invention is not limited thereto.

[0035] The data uses energy consumption data of 500 industrial enterprises in a city in Zhejiang Province for a total of 570 days from 23 to 24 years. Each data contains six different types of energy data: natural gas, water, steam, coal, diesel and gasoline. The flow chart of this detection method is as follows Figure 1 As shown, the specific implementation of this method includes the following steps: Step 1, build a real-time data detection module, and use the streaming time series of real-time data as the input of the real-time data detection module. After data preprocessing, fluctuation extraction, two-step smoothing and automatic threshold setting based on the moment method, the abnormal conditions of all energy sources at different times are detected.

[0036] Step 2: The real-time streaming data after cache processing is aligned with the original historical data through the cascade module, and the length of the time window is dynamically adjusted according to the data sampling frequency and anomaly detection requirements (for the set detection time window, for example, 24 hours, if the real-time detection data is 1 hour, the historical data is the remaining 23 hours; the 1-hour data processed by the real-time detection module is aligned with the undetected 23-hour historical data to form a complete 24-hour data for the historical data detection model input; similarly, if the detection window is 2 hours and the real-time detection data is 10 minutes, the historical data is 110 minutes, and the 10-minute data processed by the real-time detection module is aligned with the undetected 110-minute historical data to form a complete 2-hour data). Subsequently, these data are concatenated in columns to form a unified feature vector, combining the long-term pattern of historical data with the short-term dynamic changes of real-time data, and passing it as an input matrix to the historical data detection model.

[0037] Step 3: The historical data detection model uses the classification and regression tree method to extract features from historical data, and combines path length calculation, anomaly score generation, and threshold setting to identify anomalies in historical data.

[0038] Step 4: Obtain the final energy data anomaly detection result through the evaluation module.

[0039] Furthermore, in step 1, the time series is defined as an ordered set ,in Indicates the length of the time series, each value are recorded at specific equally spaced timestamps, i.e., multidimensional data points belong .

[0040] In order to reduce noise while preserving the normal pattern as much as possible, a segmented filling method based on the length of the missing segment is proposed in the data preprocessing stage. The method includes two cases: (1) For short missing segments (less than five observation points), the first-order linear interpolation method is used for filling; (2) For long missing segments (greater than or equal to five observation points): the value at the same time point in the previous period is used for estimation, and an offset term is added. The offset is defined as: ; in, and Represents the average value of adjacent time periods.

[0041] Furthermore, the exponentially weighted moving average is used to calculate the data fluctuation characteristic value , specifically expressed as: ; ; in, represents the smoothing coefficient, Indicates the window length. Fluctuation value It is used to indicate the deviation between the current point and the expected value, directly characterizing the fluctuation characteristics.

[0042] A hyperparameter selection process is designed to determine the parameter configuration that achieves the best performance, where performance is measured in terms of accuracy. This paper uses a traditional grid optimization method to fully explore the hyperparameter space through an exhaustive search strategy to find the parameter combination with the best performance. and The grid search strategy is used for the setting. Within the parameter range, search with a step size of 0.01; The search was performed with a step size of 10, thus achieving a total of 270 different hyperparameter combinations for comprehensive search and testing. Figure 2 As shown, two hyper parameters of the real-time data detection module are determined is 0.08 and When it is 50, the model reaches the highest accuracy BestAccuracy.

[0043] Further, such as Figure 3 As shown in Figure 1, a two-step smoothing mechanism is designed to make the outliers more prominent, while retaining the abnormal fluctuations and reducing the normal fluctuation residuals to near zero. The first step of smoothing extracts the fluctuation values ​​by continuous processing. To eliminate local noise, it can be expressed as: ; ; If the current point If the addition of causes a significant increase in the window standard deviation, it is considered as a potential anomaly; otherwise, the fluctuation value is set to zero.

[0044] The second step of smoothing removes periodic noise through period-based processing. Before operating on the data, a simple method is used to deal with the data drift problem: ; ; ; in, Expressed as the length of a time period, Indicates the number of cycles, Indicates the size of the data drift window. If , the current point is considered normal and its value is set to 0.

[0045] Furthermore, the smoothed fluctuation characteristics obtained in the double-step smoothing In , the moment estimation method is used to estimate the parameters, and the generalized Pareto distribution is used to set the dynamic threshold, which is specifically expressed as: ; in, is the initial threshold, and are the shape parameter and scale parameter of the generalized Pareto distribution, and The moment estimate of is derived as follows: ; in, and Represent the mean and variance of the samples respectively. The dynamic threshold is iteratively updated in the following way: ; in, Represents the risk factor, which is used to control the sensitivity of anomaly detection; Indicates that the initial threshold is exceeded points.

[0046] Furthermore, in step 2, after the real-time streaming data is cached, it is aligned with the original historical data according to the corresponding time window. Then the data is concatenated column by column to form a unified feature vector. The feature vector is passed to the historical data detection module as an input matrix.

[0047] Furthermore, in step 3, the historical data detection module uses a classification regression tree method to extract features from the historical data, and performs the following steps: (1) Feature selection: randomly select a feature ,in is the total number of features, Represents data points No. Features.

[0048] (2) Split point selection: For the selected features . Select a split point , so that its value is located in the feature The split point is between the minimum and maximum values ​​of Usually selected randomly and evenly distributed.

[0049] Recursive partitioning: For each node, the data set is divided into two subsets, and the partitioning is continued recursively for each subset until the stopping condition is met, the number of samples is 1, or the maximum tree depth is reached.

[0050] Each data point in the tree has a path length , represents the distance from the root node to the leaf node where the data point is located. , path length It is calculated based on the distance from the root node to the leaf node in each tree. exist The average path length in a tree is calculated as: ; in, represent In the The path length in the tree.

[0051] Anomaly score The average path length It is concluded that according to the path length and the size of the data set Normalize the relationship between: ; in, It is a data point The average path length among all trees in the forest.

[0052] is the normalization factor, defined as: ; constant According to the data set The path length calculation is resized to ensure that the anomaly scores are consistent across datasets of different sizes.

[0053] Once the anomaly score is calculated , you can classify data points as normal or abnormal by setting a threshold.

[0054] The trained model was used to perform multi-dimensional energy consumption anomaly detection on 7,257,600 data points over 14 days. The output results were event types, which were divided into two categories: no anomaly and anomaly. The detection results at three time scales are shown in Table 1. The recognition rate of abnormal events reached 93.78%. Figure 4 As shown in the figure, natural gas, water, steam, coal, diesel, and gasoline represent the consumption of natural gas, water, steam, coal, diesel, and gasoline respectively, and Anomalies represent abnormal values. This model can be used for anomaly detection of energy time series data.

[0055]

[0056] Compared with the prior art, the present invention has the following beneficial effects: the present invention significantly improves the accuracy of anomaly detection of multidimensional energy data by introducing a parallel learning framework and combining the characteristics of real-time streaming data and historical data in time series anomaly detection. First, a two-step smoothing mechanism and a dynamic threshold setting method based on the method of moments are adopted to effectively reduce noise and highlight abnormal points, thereby overcoming the limitations of traditional methods in dealing with complex multidimensional correlations; secondly, the historical data features are extracted by a classification regression tree model and anomaly scores are generated in combination with path length calculations, which comprehensively mines the dependency between historical patterns and real-time dynamics. Finally, hyperparameters are optimized through grid search to improve the consistency and universality of model performance. The present invention can complete anomaly detection of energy data with high precision and low computational cost, which is of great significance to the production safety and operational stability of energy enterprises.

[0057] The above description is only a further embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical solutions and concepts of the present invention within the scope disclosed by the present invention, which belong to the protection scope of the present invention.

Claims

1. A time series anomaly detection method based on a parallel learning framework, characterized by: The steps include: Step 1: Build a real-time data detection module, take the streaming time series of real-time data as the input of the real-time data detection module, and detect the abnormal conditions of all energy sources at different times through data preprocessing, fluctuation extraction, two-step smoothing and automatic threshold setting based on moment method; Step 2: Use a cascade module to align the data processed by the real-time data detection module with the original historical data according to the corresponding time window, concatenate the data by column to form a unified feature vector, integrate the long-term pattern of historical data with the instantaneous change of real-time data, and pass the feature vector as the input matrix to the historical data detection model; Step 3: The historical data detection model uses the classification and regression tree method to extract features from historical data, and combines path length calculation, anomaly score generation and threshold setting to identify anomalies in historical data; Step 4: Compare the anomaly score of each point with the threshold based on step 3. If it exceeds the threshold, it is considered an anomaly. If it does not exceed the threshold, it is not considered an anomaly. Calculate the ratio of the true anomaly points in the anomaly points to the total anomaly points in the input data.

2. The method for detecting anomalies in a time series based on a parallel learning framework according to claim 1, characterized in that: In step 1, the data preprocessing adopts a segment filling method based on the length of the missing segment, which includes using a first-order linear interpolation method to fill in short missing segments or using the value of the same time point in the previous cycle to estimate long missing segments, and adding an offset term; The offset is defined as: ; in, It represents the offset between two adjacent time periods, which is used to measure the difference in the average values ​​of the two time periods; For the current time period arrive The average value of the time points; For time series data data points; For the previous time period arrive The average value of the time points; is an index variable, indicating a specific moment in the time series; is the length of a time period, which is the window size used for segment calculation; For a time series at a point in time Observed value of The normalization factor calculated for the mean, used to ensure that the contributions of all points within a time period are averaged; The index of the current time point, indicating the reference time for calculating the offset.

3. The method for detecting time series anomalies based on a parallel learning framework according to claim 2, characterized in that: In time series data, abnormal points are manifested as significant deviations from the trend of historical data. In order to quantitatively describe the degree of this deviation, the fluctuation characteristic value is introduced as an important feature to capture abnormal fluctuations in data. The fluctuation characteristic value is obtained by calculating the difference between the current point and the expected value of its historical data. , specifically expressed as: ; ; in, For the current time point The fluctuation value indicates the deviation between the current point and the expected value, which directly represents the fluctuation characteristics; For a time series at a point in time The actual observed value of is the expected value at the current time point; is the smoothing coefficient , the larger the value, the higher the weight of recent data, and the stronger the dependence on recent data; The smaller the value, the greater the influence of historical data; is the time window length, which defines the length of historical data used to calculate the smoothed expected value; For a time series at a point in time historical observations.

4. The method for detecting time series anomalies based on a parallel learning framework according to claim 3 is characterized in that: The two-step smoothing mechanism used in step 1 includes the following steps: Smoothing out the fluctuating values ​​extracted through continuous processing To eliminate local noise, it can be expressed as: ; ; in, is the increment of the current window standard deviation, indicating that when a new data point is introduced After that, the change in the sliding window standard deviation; For the current time point Fluctuation value of is the moment before the current time point, Indicates the starting point of the window, and the window length is ; For time point The processing result; Current point If the addition of leads to a significant increase in the window standard deviation, it is considered as a potential anomaly; otherwise, the fluctuation value is set to zero; Smoothing eliminates periodic noise through period-based processing. Before operating on the data, the following methods are used to deal with data drift: ; ; ; in, For the current time point The corresponding maximum fluctuation value; For time series Fluctuation value of For the current time point Fluctuation value of is the length of a time period; For the current time point The fluctuation difference of is the maximum fluctuation value sequence within multiple periods; is the number of cycles.

5. The method for detecting time series anomalies based on a parallel learning framework according to claim 4 is characterized in that: Smoothed fluctuation characteristics obtained in double-step smoothing In , the moment estimation method is used to estimate the parameters, and the generalized Pareto distribution is used to set the dynamic threshold, which is specifically expressed as: ; in, is the initial threshold, and are the shape parameter and scale parameter of the generalized Pareto distribution, respectively; is the cumulative distribution function of the generalized Pareto distribution, which is used to describe the probability of the tail distribution; Current fluctuation value Relative to threshold The excess part, that is , and The moment estimate of is derived as follows: ; in, and Respectively represent the mean and variance of the samples, and the dynamic threshold is iteratively updated in the following way: ; in, For dynamically updated thresholds, the initial threshold is adjusted based on the distribution parameters calculated in real time; is the initial threshold, used as the benchmark at the beginning of detection; is the risk factor, which is used to control the sensitivity of anomaly detection; To exceed the initial threshold 's points.

6. The method for detecting time series anomalies based on a parallel learning framework according to claim 5, characterized in that: In step 3, the historical data detection module uses the classification regression tree method to extract features from the historical data and performs the following steps: Feature selection: randomly select a feature ,in is the total number of features, Represents data points No. Features Split point selection: For the selected features Select a split point , so that its value is located in the feature between the minimum and maximum values ​​of ; Split Point Usually selected randomly and evenly distributed; Recursive partitioning: For each node, the data set is divided into two subsets, and each subset is recursively partitioned until the stopping condition is met, the number of samples is 1, or the maximum tree depth is reached; Each data point in the tree has a path length , represents the distance from the root node to the leaf node where the data point is located. For the data point , path length It is calculated based on the distance from the root node to the leaf node in each tree. exist The average path length in a tree is calculated as: ; in, represent In the The path length on the tree; Anomaly score The average path length It is concluded that according to the path length and the size of the data set Normalize the relationship between: ; in, It is a data point The average path length among all trees in the forest, is the normalization factor, defined as: ; constant According to the data set Resize the path length calculation to ensure that the anomaly score remains consistent across datasets of different sizes; Once the anomaly score is calculated , data points are classified as normal or abnormal by setting a threshold.

Citation Information

Patent Citations

  • Data processing method and device, electronic equipment and storage medium

    CN112712549A

  • Abnormality detection method and device, electronic equipment and computer program product

    CN117112339A

  • Abnormality detection method and device and computer readable storage medium

    CN117370849A

  • Industrial internet time series data anomaly detection method and system

    CN118898045A

  • Power data anomaly detection and analysis method and system based on big data

    CN119150189A