Anomaly Detection Method, Device, and Storage Medium for Time-Series Data
Through the calculation of multi-dimensional index system and global standard score values, the problem of insufficient accuracy of time series data abnormality detection in the prior art is solved, efficient identification and accurate detection of small anomalies are achieved, and abnormal detection is suitable for spacecraft and other key areas.
Patent Information
- Application Number
- CN202510398165.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-01
AI Technical Summary
The existing standard score outlier detection method has insufficient accuracy in time series data, especially in small window timing data, which is unstable, making it difficult to meet the needs of high-precision anomalies detection. Especially in applications such as spacecraft, it is difficult to identify early slow-modification anomalies and high leakage detection rates in high noise environments.
By obtaining multiple dimension reference index value sequences of target time series data, the standard score values of the points to be detected are calculated, and indicators such as amplitude difference, cumulative amplitude, over-average rate, wave peak value, and trough value are analyzed under different time window scales to form a multi-dimensional index system, and abnormality is judged based on global standard score values, and preset thresholds and abnormality detection models are used to improve detection accuracy.
It improves the accuracy and reliability of time series data anomaly detection, can effectively identify small anomalies, reduce the misjudgment rate and misjudgment rate, and ensure the safe operation of key areas such as spacecraft.
Smart Images

Figure CN119917980B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the technical field of data processing and analysis, and particularly to an anomaly detection method, device, and storage medium for time series data. Background Art
[0002] In today's big data era, time series data, as an important type of data, widely exists in various fields, such as finance, healthcare, industry, transportation, etc. Time series data refers to data collected in chronological order, where each data point is associated with a specific time point, and these data points are usually measured and recorded at uniform time intervals. For example: daily stock prices, monthly sales amounts, annual total population, etc. With the development of technology and the progress of data acquisition technology, the magnitude and complexity of time series data are constantly increasing, and the demand for data analysis and mining is also becoming increasingly urgent.
[0003] Among them, anomaly detection, as one of the mainstream applications in time series data analysis, aims to identify abnormal events or behaviors from normal time series. Accurate anomaly detection technology has wide application value in the real world, and it can involve multiple fields such as quantitative trading, network security detection, autonomous driving vehicles, and daily maintenance of large industrial equipment. Taking an on-orbit spacecraft as an example, due to the high cost and complex system of the spacecraft, failure to detect hazards may lead to serious or even irreparable damage, and anomalies may develop into serious faults at any time. Therefore, accurate and timely anomaly detection is crucial for reminding aerospace engineers to take measures as early as possible.
[0004] Currently, the commonly used method for anomaly detection of time series data is the standard score outlier detection method, which is an outlier detection technology based on statistical principles. This method determines whether a data point is an outlier by calculating the standardized distance between the data point and the average value of the data set. However, this method has certain limitations. In time series data, the average value and standard deviation are greatly affected by outliers, especially when the calculation time window is small, the standard deviation will be more unstable, which may lead to a decrease in the accuracy of anomaly detection and cannot meet the requirements of high-precision anomaly detection in practical applications. Summary of the Invention
[0005] This specification provides an anomaly detection method, device, and storage medium for time series data to partially solve the above problems existing in the prior art.
[0006] This specification adopts the following technical solutions:
[0007] This specification provides an anomaly detection method for time series data, including:
[0008] Obtain the reference index value sequences of multiple dimensions corresponding to the points to be detected in the target time series data; wherein, the reference index value sequences include the index values of at least one reference point; the time interval between the reference point and the point to be detected is within a specified range; the index value of the reference point is used to reflect the fluctuation characteristics or statistical characteristics of the target time series data at the reference point;
[0009] For each dimension, determine the standard score value of the point to be detected according to the reference index value sequence of this dimension and the index value of the point to be detected; the standard score value is used to reflect the relative deviation degree of the index value of the point to be detected in the overall data distribution composed of the index value of the point to be detected and the index values in the reference index value sequence;
[0010] Determine the global standard score value according to the standard score values corresponding to the points to be detected, and determine whether there is an abnormality in the points to be detected according to the global standard score value.
[0011] Optionally, the index value of each dimension includes at least one of the amplitude difference calculated at different time window scales, the cumulative amplitude value calculated at different time window scales, the over-mean rate calculated at different time window scales, the peak value calculated at different time window scales, the valley value calculated at different time window scales, and the peak-valley difference calculated at different time window scales.
[0012] Optionally, determining the standard score value of the point to be detected according to the reference index value sequence of this dimension and the index value of the point to be detected specifically includes:
[0013] Filter the reference index value sequence of this dimension according to the standard score value of each reference point in the reference index value sequence of this dimension to obtain a filtered index value sequence;
[0014] Determine the standard score value of the point to be detected according to the filtered index value sequence and the index value of the point to be detected.
[0015] Optionally, filtering the reference index value sequence of this dimension according to the standard score value of each reference point in the reference index value sequence of this dimension to obtain a filtered index value sequence specifically includes:
[0016] For each reference point in the reference index value sequence of this dimension, judge whether the standard score value corresponding to this reference point for this dimension exceeds a preset screening threshold;
[0017] If not, retain the index value of this reference point under this dimension in the filtered index value sequence.
[0018] Optionally, according to the filtered sequence of index values and the index value of the point to be detected, determine the standard score value of the point to be detected, specifically including:
[0019] If it is determined that the number of index values included in the filtered sequence of index values is less than a preset number threshold, then determine that the standard score value of the point to be detected corresponding to this dimension is the preset maximum standard score value;
[0020] Otherwise, according to the statistical characteristic values of each index value included in the filtered sequence of index values and the index value of the point to be detected under this dimension, determine the standard score value of the point to be detected corresponding to this dimension.
[0021] Optionally, according to the standard score values corresponding to the point to be detected, determine the global standard score value, specifically including:
[0022] Take the maximum value among the standard score values corresponding to the point to be detected as the global standard score value.
[0023] Optionally, the method further includes:
[0024] Input the standard score values corresponding to the point to be detected into a preset anomaly detection model to obtain the anomaly detection result of the point to be detected, and the anomaly detection result is used to characterize whether there is an anomaly in the point to be detected.
[0025] Optionally, according to the global standard score value, determine whether there is an anomaly in the point to be detected, specifically including:
[0026] Judge whether the global standard score value is higher than a preset reference threshold;
[0027] If so, determine the time-related reference points of the point to be detected. If the global standard score values of the determined time-related reference points are all higher than the reference threshold, then determine that there is an anomaly in the point to be detected; the time interval between the time-related reference point and the point to be detected is within a preset associated time range.
[0028] This specification provides an anomaly detection device for time series data, including:
[0029] An acquisition module, configured to acquire a sequence of reference index values of multiple dimensions corresponding to a point to be detected in target time series data; wherein, the sequence of reference index values includes index values of at least one reference point; the time interval between the reference point and the point to be detected is within a specified range; the index value of the reference point is used to reflect the fluctuation characteristic or statistical characteristic of the target time series data at the reference point;
[0030] A determination module, configured to determine a standard score value of the to-be-detected point for each dimension according to the reference index value sequence of the dimension and the index value of the to-be-detected point; the standard score value is used to reflect the relative deviation degree of the index value of the to-be-detected point in the overall data distribution composed of the index value of the to-be-detected point and the index values in the reference index value sequence;
[0031] A detection module, configured to determine a global standard score value according to the standard score values corresponding to the to-be-detected point, and determine whether there is an abnormality in the to-be-detected point according to the global standard score value.
[0032] This specification provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the above-mentioned anomaly detection method for time series data is implemented.
[0033] This specification provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the above-mentioned anomaly detection method for time series data is implemented.
[0034] At least one of the above technical solutions adopted in this specification can achieve the following beneficial effects:
[0035] In the anomaly detection method for time series data provided in this specification, first, a reference index value sequence of multiple dimensions corresponding to the to-be-detected point in the target time series data is obtained. Among them, the reference index value sequence includes the index values of at least one reference point, and the time interval between the reference point and the to-be-detected point is within a specified range. The index value of the reference point is used to reflect the fluctuation characteristics or statistical characteristics of the target time series data at the reference point. For each dimension, according to the reference index value sequence of the dimension and the index value of the to-be-detected point, the standard score value of the to-be-detected point is determined. The standard score value is used to reflect the relative deviation degree of the index value of the to-be-detected point in the overall data distribution composed of the index value of the to-be-detected point and the index values in the reference index value sequence. According to the standard score values corresponding to the to-be-detected point, the global standard score value is determined, and whether there is an abnormality in the to-be-detected point is determined according to the global standard score value.
[0036] As can be seen from the above method, in this method, not only can the reference index value sequences of the to-be-detected point in multiple different dimensions be extracted for accurate detection of the data points included in the tiny abnormal intervals. A standard score value can also be introduced for each to-be-detected point to determine the global standard score value of the to-be-detected point after comprehensively considering the data characteristics reflected by the reference index value sequence of each dimension and the index value of the to-be-detected point. Therefore, when performing anomaly detection on the target time series data, the deviation can be effectively removed, ensuring the accuracy and reliability of the anomaly detection result. Description of the Drawings
[0037] The drawings described herein are provided to further understand the present specification and form a part of the present specification. The illustrative embodiments of the present specification and their descriptions are used to explain the present specification and do not constitute an improper limitation of the present specification. In the drawings:
[0038] Figure 1 It is a schematic flowchart of an anomaly detection method for time series data provided in the present specification;
[0039] Figure 2A It is a schematic diagram of the first and second index value sequences provided in the present specification;
[0040] Figure 2B It is a schematic diagram of the third and fourth index value sequences provided in the present specification;
[0041] Figure 2C It is a schematic diagram of the fifth and sixth index value sequences provided in the present specification;
[0042] Figure 3 It is a schematic diagram of the global standard score sequence provided in the present specification;
[0043] Figure 4 It is a schematic diagram of the anomaly detection process provided in the present specification;
[0044] Figure 5 It is a schematic diagram of an anomaly detection device for time series data provided in the present specification;
[0045] Figure 6 It is provided in the present specification corresponding to Figure 1 schematic diagram of an electronic device. Detailed Embodiments
[0046] To make the objectives, technical solutions, and advantages of the present specification clearer, the technical solutions of the present specification will be clearly and completely described below in conjunction with the specific embodiments of the present specification and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present specification, rather than all the embodiments. Based on the embodiments in the present specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present specification.
[0047] Currently, time series data, as a type of structured data with time-dependent characteristics, has the observed values of its data points strictly corresponding to specific time points and presents continuous observation characteristics in equally-spaced sampling scenarios (such as minute-level sensor readings, daily financial indicators, monthly economic statistics, etc.). Such data constitutes the core analysis object in key fields such as industrial monitoring, financial transactions, and spacecraft operation and maintenance. Among them, anomaly detection technology, as one of the core tasks of time series analysis, aims to identify data points or event sequences that deviate from the normal pattern through algorithmic models, thereby realizing risk warning and decision support.
[0048] Taking the health management of on-orbit spacecraft as an example, the time series generated by its sensor network (such as attitude angle, propellant pressure, thermal control temperature) needs to monitor micro-anomalies in real time to prevent catastrophic failures. Although the existing standard score method can detect some obvious anomalies (such as abrupt outliers caused by sensor failures), it has insufficient recognition sensitivity for early slow-changing anomalies (such as trend offsets caused by the progressive degradation of components), and is prone to an increase in the missed detection rate due to the distortion of statistics in a high-noise background. In addition, the operating conditions of the spacecraft change frequently during on-orbit missions, and fixed-parameterized detection models are difficult to accurately identify the anomalies existing in the time series data, further restricting the reliability of the detection results and posing a major operation and maintenance risk to high-value space assets.
[0049] The following will, in conjunction with the accompanying drawings, elaborate on the technical solutions provided in each embodiment of this specification in detail.
[0050] Figure 1 The following is a schematic flowchart of an anomaly detection method for time series data provided in this specification, including the following steps:
[0051] S101: Obtain a sequence of reference index values for multiple dimensions corresponding to the point to be detected in the target time series data; wherein, the sequence of reference index values includes the index values of at least one reference point; the time interval between the reference point and the collection of the point to be detected is within a specified range; the index value of the reference point is used to reflect the fluctuation characteristics or statistical characteristics of the target time series data at the reference point.
[0052] In this specification, the business platform can perform anomaly detection on the points to be detected included in the collected target time series data, and send an alarm message when it is determined that there is an anomaly in the point to be detected, so that the business personnel can perform anomaly handling tasks based on the alarm message.
[0053] Specifically, the above-mentioned target time-series data can be the complete time-series data collected in advance. At this time, the business platform can use each data point included in the target time-series data as a point to be detected, and then can obtain the reference index value sequences of multiple dimensions corresponding to the point to be detected, so as to determine whether there is an abnormality in the target time-series data at the point to be detected based on the reference index value sequences of multiple dimensions corresponding to the point to be detected.
[0054] Among them, the above-mentioned reference index value sequence includes the index values of at least one reference point, and the time interval between the reference point and the point to be detected is within a specified range. In other words, the interval between the time point of collecting the reference point and the time point of collecting the point to be detected is within a specified range.
[0055] For example: If the above-mentioned target time-series data consists of 10 data points (i.e., the data points collected by the sensor from the 1st second to the 10th second), and the point to be detected is the data point collected at the 6th second in the target time-series data. At this time, the business platform can, according to actual needs, use the data points whose time intervals from the time point corresponding to the point to be detected are within a specified range (for example, when the specified range is 2 seconds, the data points within the specified range are the data points collected at the 4th second, 5th second, 7th second, and 8th second in the target data) as the reference points of the point to be detected. These reference points together form the reference time-series data sequence corresponding to the point to be detected, and the index values corresponding to each reference point among these reference points together form the reference index value sequence corresponding to the point to be detected.
[0056] Furthermore, in order to be able to detect the tiny abnormal intervals included in the target time-series data, the index value of each data point included in the above-mentioned target time-series data can be the index values of different dimensions. The index values of each dimension here can include: amplitude difference (i.e., the difference between the maximum amplitude value and the minimum amplitude value within a preset time window), cumulative amplitude value (i.e., the cumulative sum of the absolute values of the differences between the amplitude values of every two adjacent data points within a preset time window), over-mean rate (i.e., the ratio of the number of times that the values of two adjacent data points among the data points included in a preset time window are respectively on different sides of the average value of the entire target time-series data, that is, one data point value is higher than the average value and the other data point value is lower than the average value, to the maximum possible number of over-mean times of the target time-series data), peak value (i.e., the maximum peak value included in a preset time window), valley value (i.e., the minimum valley value included in a preset time window), peak-valley difference (i.e., the cumulative sum of the differences between all adjacent peak values and valley values included in a preset time window).
[0057] It should be noted that when the time window scales of the above-mentioned time windows are different, the sensitivities for detecting tiny abnormal intervals contained in the target time series data are also different. Therefore, the index values of each dimension mentioned above can also include at least one of the amplitude differences calculated under different time window scales, the cumulative amplitude values calculated under different time window scales, the over-mean rates calculated under different time window scales, the peak values calculated under different time window scales, the trough values calculated under different time window scales, and the peak-trough differences calculated under different time window scales.
[0058] For example: The amplitude difference of the data points determined by the time window with a time window scale of 10 and the amplitude difference of the data points determined by the time window with a time window scale of 20 are two-dimensional index values.
[0059] Another example: The amplitude difference of the data points determined by the time window with a time window scale of 10 and the cumulative amplitude value of the data points determined by the time window with a time window scale of 10 can also be two-dimensional index values.
[0060] Based on this, for each dimension, the index values of each reference point contained in the target time series data under this dimension together form a reference index value sequence corresponding to the point to be detected. It can be understood that there can be multiple reference index value sequences of the point to be detected obtained by the business platform, and different reference index value sequences contain index values of different dimensions, specifically as Figure 2A 、 Figure 2B 、 Figure 2C shown.
[0061] Figure 2A are the schematic diagrams of the first and second index value sequences provided in this specification.
[0062] In Figure 2A , 100 is the target time series data, 101 is the index value sequence composed of the amplitude differences calculated by each data point contained in the target time series data on one time window scale, and 102 is the index value sequence composed of the amplitude differences calculated by each data point contained in the target time series data on another time window scale.
[0063] Figure 2B are the schematic diagrams of the third and fourth index value sequences provided in this specification.
[0064] In Figure 2BAmong them, 201 is a sequence of metric values composed of cumulative amplitude values calculated for each data point included in the target time-series data on one time window scale, and 202 is a sequence of metric values composed of cumulative amplitude values calculated for each data point included in the target time-series data on another time window scale.
[0065] Figure 2C These are schematic diagrams of the fifth and sixth sequences of metric values provided in this specification.
[0066] In Figure 2B Among them, 301 is a sequence of metric values composed of peak-valley differences calculated for each data point included in the target time-series data on one time window scale, and 302 is a sequence of metric values composed of peak-valley differences calculated for each data point included in the target time-series data on another time window scale. By analogy, the service platform can obtain sequences of reference metric values for multiple dimensions corresponding to the points to be detected in the target time-series data according to actual needs, and this specification will not list them one by one here.
[0067] It should be noted that in application scenarios such as in-orbit spacecraft health management, it is often necessary to perform real-time monitoring of micro-anomalies on time-series data (such as attitude angles, propellant pressures, thermal control temperatures) generated by sensor networks to prevent catastrophic failures.
[0068] At this time, the service platform can also use the time-series data that needs to be collected through the sensor network as the target time-series data. Furthermore, during the process of collecting the target time-series data through the sensor network, each newly collected data point can be used as the point to be detected. Then, at least some of the data points that make up the target time-series data and have been collected before the point to be detected can be used as reference points for the point to be detected, and real-time anomaly detection can be performed on the point to be detected based on the reference points of the point to be detected, so as to timely discover possible anomalies at the point to be detected and send warning messages, thereby avoiding losses.
[0069] For example: If the above point to be detected is the data point at the 6th second in the target time-series data. At this time, the service platform can use the data points that make up the target time-series data and have been collected before the point to be detected and whose time intervals from the time point corresponding to the point to be detected are within the specified range as the reference points for the point to be detected. For example, when the specified range is within 3 seconds, the reference points for the point to be detected can be the data points at the 3rd second, 4th second, and 5th second in the target time-series data that have been collected.
[0070] In this specification, the entity that implements the anomaly detection method for time series data can refer to a designated device such as a server set up on a business platform, or can also refer to a terminal device such as a desktop computer or a laptop. For the sake of convenience in description, hereinafter, only the case where the server is the entity will be taken as an example to illustrate the anomaly detection method for time series data provided in this specification.
[0071] S102: For each dimension, determine the standard score value of the point to be detected according to the reference index value sequence of this dimension and the index value of the point to be detected; the standard score value is used to reflect the relative deviation degree of the index value of the point to be detected in the overall data distribution composed of the index value of the point to be detected and the index values in the reference index value sequence.
[0072] It can be seen from the above content that when the server collects each data point in the target time series data, it can use this data point as the point to be detected, and calculate the index values of different dimensions of the point to be detected according to different time window scales. And according to the index values of each reference point of the point to be detected in different dimensions, multiple reference index value sequences corresponding to the point to be detected can be obtained. Furthermore, anomaly detection can be performed on the point to be detected according to the index values of different dimensions of the point to be detected and the multiple reference index value sequences corresponding to multiple dimensions of the point to be detected.
[0073] Among them, for the index value of each dimension, the server can determine the standard score value corresponding to the dimension of the point to be detected according to the reference index value sequence of this dimension and the index value of the point to be detected in this dimension (here, the standard score value is used to reflect the relative deviation degree of the index value of the point to be detected in the overall data distribution composed of the index value of the point to be detected and the index values in the reference index value sequence), for use when performing anomaly detection on the point to be detected.
[0074] Specifically, for each dimension, the server can filter the reference index value sequence of this dimension according to the standard score value of each reference point in the reference index value sequence of this dimension, and obtain the filtered index value sequence. Furthermore, according to the filtered index value sequence and the index value of the point to be detected in this dimension, the standard score value corresponding to the dimension of the point to be detected can be determined.
[0075] Among them, the method for the server to filter the reference index value sequence of this dimension to obtain the filtered index value sequence can be to judge whether the standard score value corresponding to each reference point in the reference index value sequence of this dimension exceeds a preset screening threshold (here, the screening threshold can be a prior parameter set in advance according to actual requirements). If not, the index value of this reference point in this dimension is retained in the filtered index value sequence.
[0076] It should be noted that in fields such as the health detection of on-orbit spacecraft, when the collected spacecraft telemetry data is used as the target time-series data, since the spacecraft telemetry data usually does not strictly follow a normal distribution and may exhibit characteristics such as skewed distribution and heavy-tailed distribution, there may be many outliers and noises in the spacecraft telemetry data, which may lead to abnormal data being easily mixed into each reference point in the above reference index value sequence. In order to improve the accuracy of screening abnormal data from each reference point in the above reference index value sequence, the above screening threshold can be set to 4 to reduce the false positive rate and false negative rate of anomaly detection and ensure the reliable operation of the spacecraft.
[0077] In the actual application scenario, there may be an abnormal situation where the filtered index value sequence is empty. At this time, when the server determines that the number of index values included in the filtered index value sequence is less than the preset quantity threshold, it can determine that the standard score value of the point to be detected corresponding to this dimension is the preset maximum standard score value (the maximum standard score value here can be a priori parameter set in advance).
[0078] Otherwise, the server can determine the standard score value of the point to be detected corresponding to this dimension according to the statistical characteristic values of the index values included in the filtered index value sequence and the index value of the point to be detected under this dimension. Specifically, it can refer to the following formula:
[0079]
[0080]
[0081]
[0082] In the above formula, is the average value of the index values included in the filtered index value sequence, is the standard deviation of the index values included in the filtered index value sequence, is the nth index value in the filtered index value sequence, is the standard score value of the point to be detected corresponding to this dimension.
[0083] It can be seen from the above formula that for each reference index value sequence corresponding to the point to be detected, the server can calculate the standard score value of the point to be detected corresponding to this reference index value sequence based on this reference index value sequence through the above method.
[0084] It should be noted that Figure 2A 103 in is the standard score value sequence composed of the standard score values of each data point included in the target time-series data calculated according to a reference index value sequence 101 of the target time-series data, Figure 2AThe 104 in it is the standard score value sequence composed of the standard score values of each data point included in the target time series data calculated based on a reference index value sequence 102 of the target time series data. By analogy, it can be known that Figure 2B 、 Figure 2C The 203, 204, 303, and 304 in it are the standard score value sequences corresponding to the target time series data.
[0085] S103: Determine the global standard score value according to the standard score values corresponding to the to-be-detected point, and determine whether there is an abnormality in the to-be-detected point according to the global standard score value.
[0086] Furthermore, after the server determines the standard score values corresponding to the to-be-detected point, it can determine the global standard score value according to the standard score values corresponding to the to-be-detected point, and perform abnormality detection on the to-be-detected point according to the global standard score value to obtain the abnormality detection result of the to-be-detected point. Here, the abnormality detection result is used to reflect whether there is an abnormality in the to-be-detected point. Specifically, as Figure 3 shown.
[0087] Figure 3 This is a schematic diagram of the global standard score sequence provided in this specification.
[0088] In Figure 3 it, 401 is the global standard score value sequence composed of the global standard score values corresponding to each data point included in the target time series data. Here, the length of the global standard score value sequence is the same as the length of the target time series data.
[0089] Among them, there are two methods for the server to determine the global standard score value according to the standard score values corresponding to the to-be-detected point. The following will explain these two methods in detail respectively.
[0090] The first method can be that the server can determine the global standard score value according to the standard score values corresponding to the to-be-detected point according to a preset calculation rule. The above calculation rule can be set according to actual needs.
[0091] For example: Take the maximum value among the standard score values corresponding to the to-be-detected point as the global standard score value.
[0092] For another example, if it is determined that the standard score values corresponding to the point to be detected conform to a preset rule (the rule here can be: when it is determined that there is 1 standard score value greater than 5 among the standard score values, it can be regarded as conforming to the preset rule; or when it is determined that there are at least 2 standard score values greater than 4 among the standard score values, it can be regarded as conforming to the preset rule; or when it is determined that there are at least 3 standard score values greater than 3 among the standard score values, it can be regarded as conforming to the preset rule), then the preset specified global standard score value can be determined as the global standard score value corresponding to the point to be detected.
[0093] The second method can be to input the standard score values corresponding to the point to be detected into a preset anomaly detection model to obtain the anomaly detection result of the point to be detected.
[0094] It should be noted that the above anomaly detection model needs to be trained before it can be deployed to the server. Among them, the method for training the above anomaly detection model can be to obtain the sample standard score values corresponding to the sample detection points and the actual anomaly detection results corresponding to the sample detection points, input the sample standard score values corresponding to the sample detection points into the preset anomaly detection model to obtain the reference anomaly detection result, and then the target loss value can be determined according to the deviation between the reference anomaly detection result and the actual anomaly detection result (the deviation between the reference anomaly detection result and the actual anomaly detection result is positively correlated with the target loss value), and the anomaly detection model is trained with the goal of minimizing the target loss value to obtain the trained anomaly detection model.
[0095] Furthermore, the method for the server to perform anomaly detection on the point to be detected based on the global standard score value to obtain the anomaly detection result of the point to be detected can be that the server can determine whether the global standard score value is higher than a preset reference threshold. If so, the server can determine the time-related reference points of the point to be detected. When it is determined that the global standard score values of the time-related reference points are all higher than the reference threshold, it is determined that there is an anomaly in the point to be detected. Here, the time interval between the time-related reference point and the point to be detected is within a preset associated time range, and the associated time range and the reference threshold can be prior parameters determined in advance according to actual needs. Specifically, as Figure 4 shown.
[0096] Figure 4 This is a schematic diagram of the anomaly detection process provided in this specification.
[0097] Combined with Figure 4It can be seen that when the server determines that the global standard score values corresponding to each reference point before the point to be detected are continuously greater than the reference threshold within the preset correlation time range, it can be determined that the point to be detected and the global standard score values greater than the preset reference threshold within the correlation time range of the point to be detected jointly form an abnormal interval in the target time series data.
[0098] For example: It can be seen from Figure 4 that 501, 502, 503, 504, and 505 in Figure 4 are each alternative abnormal interval in the global standard score value sequence corresponding to the target time series data where the global standard score value is greater than the reference threshold. Among them, since the time range corresponding to each data point forming the alternative abnormal interval 501 is less than the above-mentioned correlation actual range, the data points forming the alternative abnormal interval 501 do not belong to abnormal data points, while the data points forming the alternative abnormal intervals 502, 503, 504, and 505 belong to abnormal data points.
[0099] It can be seen from the above method that the server not only considers the amplitude difference of the time series data within a certain time period (i.e., the difference between the maximum and minimum values of the data within the time period), but also comprehensively analyzes the cumulative amplitude sum, the rate of crossing the mean, the peak value, and the trough value. These indicators reflect the fluctuation characteristics and statistical characteristics of the time series data from different perspectives, forming a comprehensive multi-dimensional index system. In addition, the server can also calculate the above indicators on different time scales, enabling the server to capture the change trends and abnormal characteristics of the time series data within different time windows. For example, a short-term window can capture sudden anomalies, while a long-term window can detect trend-like anomalies.
[0100] On this basis, in order to further improve the detection accuracy, the server can also introduce a standard score sequence for each dimension. By calculating the standard score of each data point, the influence of abnormal values can be effectively removed when calculating the mean and standard deviation, making the statistical analysis more robust and reliable.
[0101] The above is an abnormal detection method for one or more implementation time series data in this specification. Based on the same idea, this specification also provides a corresponding abnormal detection device for time series data, such as Figure 5 shown.
[0102] Figure 5 is a schematic diagram of an abnormal detection device for time series data provided in this specification, including:
[0103] An acquisition module 501, configured to acquire a sequence of reference metric values for multiple dimensions corresponding to a point to be detected in target time series data; wherein, the sequence of reference metric values includes metric values of at least one reference point; a time interval between the reference point and the point to be detected is within a specified range; the metric value of the reference point is used to reflect a fluctuation feature or a statistical feature of the target time series data at the reference point;
[0104] A determination module 502, configured to, for each dimension, determine a standard score value of the point to be detected according to the sequence of reference metric values of the dimension and the metric value of the point to be detected; the standard score value is used to reflect a relative deviation degree of the metric value of the point to be detected in an overall data distribution formed by the metric value of the point to be detected and the metric values in the sequence of reference metric values;
[0105] A detection module 503, configured to determine a global standard score value according to the standard score values corresponding to the point to be detected, and determine whether the point to be detected is abnormal according to the global standard score value.
[0106] Optionally, the metric value of each dimension includes at least one of an amplitude difference calculated at different time window scales, an accumulated amplitude value calculated at different time window scales, a rate of crossing the mean calculated at different time window scales, a peak value calculated at different time window scales, a valley value calculated at different time window scales, and a peak-valley difference calculated at different time window scales.
[0107] Optionally, the determination module 502 is specifically configured to filter the sequence of reference metric values of the dimension according to the standard score value of each reference point in the sequence of reference metric values of the dimension, to obtain a filtered sequence of metric values; and determine the standard score value of the point to be detected according to the filtered sequence of metric values and the metric value of the point to be detected.
[0108] Optionally, the determination module 502 is specifically configured to, for each reference point in the sequence of reference metric values of the dimension, determine whether the standard score value corresponding to the reference point for the dimension exceeds a preset screening threshold; if not, retain the metric value of the reference point under the dimension in the filtered sequence of metric values.
[0109] Optionally, the determination module 502 is specifically configured to, if it is determined that the number of metric values included in the filtered sequence of metric values is less than a preset quantity threshold, determine that the standard score value of the point to be detected corresponding to the dimension is a preset maximum standard score value; otherwise, determine the standard score value of the point to be detected corresponding to the dimension according to the statistical feature value of each metric value included in the filtered sequence of metric values and the metric value of the point to be detected under the dimension.
[0110] Optionally, the detection module 503 is specifically configured to use the maximum value among the standard score values corresponding to the point to be detected as the global standard score value.
[0111] Optionally, the detection module 503 is specifically configured to input the standard score values corresponding to the point to be detected into a preset anomaly detection model to obtain an anomaly detection result for the point to be detected, where the anomaly detection result is used to indicate whether there is an anomaly at the point to be detected.
[0112] Optionally, the detection module 503 is specifically configured to determine whether the global standard score value is higher than a preset reference threshold; if so, determine the time-related reference points for the point to be detected, and if it is determined that the global standard score values of the time-related reference points are all higher than the reference threshold, then determine that there is an anomaly at the point to be detected; the time interval between the time-related reference point and the point to be detected is within a preset associated time range.
[0113] This specification also provides a computer-readable storage medium storing a computer program that can be used to execute the above Figure 1 provided method for anomaly detection of time series data.
[0114] This specification also provides Figure 6 a schematic structural diagram of an electronic device corresponding to Figure 1 as shown. As Figure 6 described, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory, and of course may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above Figure 1 described method for anomaly detection of time series data. Of course, in addition to the software implementation, this specification does not exclude other implementation methods, such as logical devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logical unit, and can also be hardware or a logical device.
[0115] For an improvement in a technology, it can be clearly distinguished whether it is a hardware improvement (e.g., improvement in circuit structures such as diodes, transistors, switches, etc.) or a software improvement (improvement in method processes). However, with the development of technology, many improvements in method processes today can be regarded as direct improvements in hardware circuit structures. Almost all designers obtain the corresponding hardware circuit structure by programming the improved method process into the hardware circuit. Therefore, it cannot be said that an improvement in a method process cannot be implemented with a hardware entity module. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. The designer can program by himself to "integrate" a digital system on a piece of PLD, without having to ask the chip manufacturer to design and manufacture a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL), and there is not only one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply making a little logical programming of the method process with the above-mentioned several hardware description languages and programming it into the integrated circuit, it is easy to obtain the hardware circuit that implements the logical method process.
[0116] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that, in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to make the controller implement the same function in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or structures within the hardware component.
[0117] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.
[0118] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0119] Those skilled in the art should understand that the embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.
[0120] This specification is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the specification. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general purpose computers, special purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices generate means for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.
[0121] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.
[0122] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.
[0123] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0124] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.
[0125] A computer-readable medium includes both permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information accessible by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media such as modulated data signals and carrier waves.
[0126] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or also includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.
[0127] Those skilled in the art should understand that the embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.
[0128] This specification can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0129] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and for the relevant parts, reference can be made to the corresponding description in the method embodiment.
[0130] The above is only the embodiment of this specification and is not intended to limit this specification. For those skilled in the art, various changes and modifications can be made to this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the scope of the claims of this specification.
Claims
1. An anomaly detection method for time series data, characterized in that, Including: Obtain a sequence of reference index values for multiple dimensions corresponding to the point to be detected in the target time series data; wherein, the sequence of reference index values includes the index values of at least one reference point; the time interval between the reference point and the point to be detected is within a specified range; the index value of the reference point is used to reflect the fluctuation characteristics or statistical characteristics of the target time series data at the reference point; the target time series data includes one of the attitude angle data of the spacecraft, the propellant pressure data of the spacecraft, and the thermal control temperature data of the spacecraft; For each dimension, determine the standard score value of the point to be detected according to the sequence of reference index values of this dimension and the index value of the point to be detected; the standard score value is used to reflect the relative deviation degree of the index value of the point to be detected in the overall data distribution composed of the index value of the point to be detected and the index values in the sequence of reference index values; Determine the global standard score value according to the standard score values corresponding to the point to be detected, and determine whether there is an abnormality at the point to be detected according to the global standard score value.
2. The method according to claim 1, wherein The index value of each dimension includes at least one of the amplitude difference calculated under different time window scales, the cumulative amplitude value calculated under different time window scales, the over-mean rate calculated under different time window scales, the peak value calculated under different time window scales, the trough value calculated under different time window scales, and the peak-to-valley difference calculated under different time window scales.
3. The method according to claim 2, wherein Determine the standard score value of the point to be detected according to the sequence of reference index values of this dimension and the index value of the point to be detected, specifically including: Filter the sequence of reference index values of this dimension according to the standard score value of each reference point in the sequence of reference index values of this dimension to obtain a filtered sequence of index values; Determine the standard score value of the point to be detected according to the filtered sequence of index values and the index value of the point to be detected.
4. The method according to claim 3, wherein Filter the sequence of reference index values of this dimension according to the standard score value of each reference point in the sequence of reference index values of this dimension to obtain a filtered sequence of index values, specifically including: For each reference point in the sequence of reference index values of this dimension, determine whether the standard score value corresponding to this reference point for this dimension exceeds a preset screening threshold; If not, retain the index value of this reference point under this dimension in the filtered sequence of index values.
5. The method according to claim 3, wherein Determine the standard score value of the point to be detected according to the filtered sequence of index values and the index value of the point to be detected, specifically including: If it is determined that the number of index values included in the filtered sequence of index values is less than a preset quantity threshold, determine that the standard score value corresponding to the point to be detected for this dimension is the preset maximum standard score value; Otherwise, determine the standard score value corresponding to the point to be detected for this dimension according to the statistical characteristic values of the index values included in the filtered sequence of index values and the index value of the point to be detected under this dimension.
6. The method according to claim 1, wherein Determine the global standard score value according to the standard score values corresponding to the point to be detected, specifically including: Take the maximum value among the standard score values corresponding to the point to be detected as the global standard score value.
7. The method according to claim 1, wherein The method further includes: Input the standard score values corresponding to the point to be detected into a preset anomaly detection model to obtain the anomaly detection result of the point to be detected, where the anomaly detection result is used to characterize whether there is an anomaly at the point to be detected.
8. The method according to claim 1, wherein Determine whether there is an anomaly at the point to be detected according to the global standard score value, specifically including: Judge whether the global standard score value is higher than a preset reference threshold; If so, determine the time-related reference points of the point to be detected. When it is determined that the global standard score values of all the time-related reference points are higher than the reference threshold, it is determined that there is an anomaly at the point to be detected; the time interval between the time-related reference point and the point to be detected is within a preset associated time range.
9. An abnormal detection device for time series data, characterized in that, Including: An acquisition module, configured to acquire a sequence of reference index values of multiple dimensions corresponding to the point to be detected in the target time series data; wherein, the sequence of reference index values includes the index values of at least one reference point; the time interval between the reference point and the point to be detected is within a specified range; the index value of the reference point is used to reflect the fluctuation feature or statistical feature of the target time series data at the reference point; the target time series data includes one of the attitude angle data of the spacecraft, the propellant pressure data of the spacecraft, and the thermal control temperature data of the spacecraft; A determination module, configured to determine the standard score value of the point to be detected according to the sequence of reference index values of each dimension and the index value of the point to be detected; the standard score value is used to reflect the relative deviation degree of the index value of the point to be detected in the overall data distribution composed of the index value of the point to be detected and the index values in the sequence of reference index values; A detection module, configured to determine the global standard score value according to the standard score values corresponding to the point to be detected, and determine whether there is an anomaly at the point to be detected according to the global standard score value.
10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and when the computer program is executed by a processor, the method described in any one of claims 1 to 8 above is implemented.
Citation Information
Patent Citations
Fault detection method and device
CN109861857A
Anomaly detection method and device, electronic equipment and computer readable storage medium
CN113518011A