Time series data anomaly detection method and device, electronic equipment and storage medium

By training and correcting parameter prediction on the time series data annotation sample set, the target anomaly detection model is optimized, and the problem of poor interpretability of machine learning models is solved, and the accuracy and reliability of time series data anomaly detection is improved.

CN120561791APending Publication Date: 2025-08-29SHENZHEN CONSYS SCI&TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510500913.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

Among the existing timing data anomaly detection methods, the machine learning model has poor interpretability, resulting in inaccurate abnormality detection results.

Method used

By obtaining time series data and labeling sample sets, the pre-constructed initial anomaly detection model is trained, and the normal data fluctuation interval and detection result correction parameters are used to optimize the target anomaly detection model, perform probability and interval correction, and finally perform information nesting and merging to improve detection accuracy.

Benefits of technology

It improves the accuracy and reliability of timing data abnormality detection, ensures that the target abnormality detection model can learn accurate normal and abnormal data characteristics, and provides reliable detection benchmarks and optimization methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120561791A_ABST
    Figure CN120561791A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a time series data anomaly detection method and device, electronic equipment and a storage medium, and belongs to the technical field of artificial intelligence. The method comprises the following steps: acquiring a time sequence data annotation sample set; performing model training on the initial anomaly detection model based on the time series data labeling sample set to obtain a target anomaly detection model; performing fluctuation interval prediction on the time series data annotation sample set to obtain a normal data fluctuation interval; performing correction parameter prediction based on the time series data annotation sample set and the normal data fluctuation interval to obtain a detection result correction parameter; performing anomaly detection on the target time series data based on the target anomaly detection model to obtain a data anomaly detection preliminary result; and based on the normal data fluctuation interval, the detection result correction parameter and the target time sequence data, performing correction processing on the data anomaly detection preliminary result to obtain time sequence data anomaly detection information. According to the embodiment of the invention, the accuracy of time series data anomaly detection can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method and device for detecting anomalies in time series data, an electronic device, and a storage medium. Background Art

[0002] Time series data anomaly detection refers to the process of identifying data points or intervals in time series data that deviate significantly from the normal pattern.

[0003] Currently, common time series data anomaly detection usually uses a trained machine learning model to perform anomaly detection on the input time series data, and uses the output of the machine learning model as the anomaly detection result of the time series data. However, the interpretability of the machine learning model is poor, resulting in inaccurate output anomaly detection results. Therefore, how to improve the accuracy of time series data anomaly detection has become a technical problem that needs to be solved urgently. Summary of the Invention

[0004] The main purpose of the embodiments of the present application is to propose a time series data anomaly detection method and device, an electronic device and a storage medium, in order to improve the accuracy of time series data anomaly detection.

[0005] To achieve the above objectives, a first aspect of an embodiment of the present application provides a method for detecting anomalies in time series data, the method comprising:

[0006] Obtain a time series data annotated sample set, wherein the time series data annotated sample set includes time series sample data, positive and negative data labels, and abnormal interval labels, wherein the data labels are used to indicate whether the time series sample data is normal or abnormal, and the abnormal interval labels are used to indicate abnormal data intervals in the time series sample data;

[0007] Based on the time series data labeled sample set, a pre-built initial anomaly detection model is trained to obtain a target anomaly detection model;

[0008] Predicting the fluctuation range of the time series data annotated sample set to obtain a normal data fluctuation range;

[0009] Based on the time series data annotated sample set and the normal data fluctuation range, correction parameter prediction is performed to obtain the detection result correction parameter;

[0010] Based on the target anomaly detection model, anomaly detection is performed on the preset target time series data to obtain a data detection probability pair and a data detection anomaly interval, wherein the data prediction probability pair is used to characterize the probability of normality and anomaly of the target time series data;

[0011] Based on the normal data fluctuation interval, the detection result correction parameter and the target time series data, probability correction is performed on the data detection probability pair to obtain a data correction probability pair;

[0012] Based on the normal data fluctuation interval, performing interval correction on the data detection abnormal interval to obtain a data correction abnormal interval;

[0013] Information nesting and merging is performed on the data correction probability pair and the data correction abnormal interval to obtain time series data abnormality detection information.

[0014] In some embodiments, the training of a pre-built initial anomaly detection model based on the time series data labeled sample set to obtain a target anomaly detection model includes:

[0015] Based on the initial anomaly detection model, abnormal data prediction is performed on the time series data labeled sample set to obtain predicted positive and negative probability pairs and predicted data anomaly intervals;

[0016] Based on the positive and negative data labels, performing loss calculation on the predicted positive and negative probability pairs to obtain probability pair loss data;

[0017] Based on the abnormal interval label, performing loss calculation on the abnormal interval of the predicted data to obtain interval loss data;

[0018] Based on the probability pair loss data and the interval loss data, parameters of the initial anomaly detection model are adjusted to obtain the target anomaly detection model.

[0019] In some embodiments, performing fluctuation range prediction on the time series data labeled sample set to obtain a normal data fluctuation range includes:

[0020] Based on the positive and negative data labels, the time series sample data is screened to obtain normal sample data;

[0021] Performing mean calculation on the normal sample data to obtain normal mean data;

[0022] Performing heteroscedasticity calculation on the normal mean data to obtain data heteroscedasticity;

[0023] Based on the data heteroskedasticity, the normal data fluctuation range is determined.

[0024] In some embodiments, the correction parameter prediction based on the time series data labeled sample set and the normal data fluctuation interval to obtain the detection result correction parameter includes:

[0025] Based on the positive and abnormal data labels, the time series sample data is classified to obtain the normal sample data and the abnormal sample data;

[0026] Based on the normal data fluctuation interval, the normal sample data is recorded as out-of-interval data to obtain the number of data outside the positive sample interval;

[0027] Based on the normal data fluctuation interval, the abnormal sample data is recorded as out-of-interval data to obtain the number of data outside the abnormal sample interval;

[0028] Based on the number of data outside the positive sample interval and the number of data outside the outlier sample interval, a correction parameter is calculated to obtain the detection result correction parameter.

[0029] In some embodiments, the calculation of the correction parameter based on the number of data outside the positive sample interval and the number of data outside the outlier sample interval to obtain the detection result correction parameter includes:

[0030] Performing maximum screening on the number of data outside the positive sample interval to obtain the target number of positive samples;

[0031] Perform minimum screening on the number of data outside the outlier sample interval to obtain the target number of outlier samples;

[0032] Calculating the quantity difference between the number of outlier targets and the number of positive sample targets to obtain a quantity difference parameter;

[0033] Performing median calculation on the target number of the heterogeneous samples to obtain a median parameter of the number;

[0034] The quantity difference parameter and the quantity median parameter are subjected to data nesting and merging to obtain the detection result correction parameter.

[0035] In some embodiments, the performing probability correction on the data detection probability pair based on the normal data fluctuation interval, the detection result correction parameter, and the target time series data to obtain the data correction probability pair includes:

[0036] Based on the normal data fluctuation interval, the target time series data is recorded as out-of-interval data to obtain the number of data outside the target interval;

[0037] Based on the amount of data outside the target interval and the detection result correction parameter, performing normal probability correction on the data detection probability pair to obtain a data normal correction probability;

[0038] Determining a data abnormality correction probability based on the data normal correction probability;

[0039] The normal data correction probability and the abnormal data correction probability are nested and merged to obtain the data correction probability pair.

[0040] In some embodiments, the data detection abnormal interval includes an abnormal interval start point and an abnormal interval end point. Based on the normal data fluctuation interval, performing interval correction on the data detection abnormal interval to obtain a data correction abnormal interval includes:

[0041] Based on the data detection abnormal interval, the target time series data is interval-screened to obtain a target detection interval, wherein the target detection interval includes a detection interval start point and a detection interval end point, the detection interval start point is less than or equal to the abnormal interval start point, and the detection interval end point is greater than or equal to the abnormal interval end point;

[0042] Recording data for the target detection interval to obtain the number of interval data;

[0043] Based on the normal data fluctuation interval, recording the data outside the target detection interval to obtain the number of data outside the detection interval;

[0044] Based on the number of data in the interval and the number of data outside the detection interval, the abnormal data ratio is calculated to obtain abnormal proportion data;

[0045] Based on the abnormal ratio data, the target detection interval is screened to obtain the data correction abnormal interval, wherein the abnormal ratio data of the data correction abnormal interval is the largest and the interval range of the data correction abnormal interval is the largest.

[0046] To achieve the above-mentioned purpose, a second aspect of an embodiment of the present application provides a time series data anomaly detection device, the device comprising:

[0047] A sample data acquisition module is used to acquire a time series data annotated sample set, wherein the time series data annotated sample set includes time series sample data, positive and negative data labels, and abnormal interval labels. The data labels are used to indicate whether the time series sample data is normal or abnormal, and the abnormal interval labels are used to indicate abnormal data intervals in the time series sample data.

[0048] An initial model training module is used to train a pre-built initial anomaly detection model based on the time series data labeled sample set to obtain a target anomaly detection model;

[0049] A fluctuation interval prediction module is used to predict the fluctuation interval of the time series data annotated sample set to obtain a normal data fluctuation interval;

[0050] A correction parameter prediction module is used to predict the correction parameters based on the time series data labeled sample set and the normal data fluctuation range to obtain the correction parameters of the detection result;

[0051] A data anomaly detection module is used to perform anomaly detection on the preset target time series data based on the target anomaly detection model to obtain a data detection probability pair and a data detection anomaly interval, wherein the data prediction probability pair is used to represent the probability of normality and anomaly of the target time series data;

[0052] a probability pair correction module, configured to perform probability correction on the data detection probability pair based on the normal data fluctuation interval, the detection result correction parameter, and the target time series data to obtain a data correction probability pair;

[0053] An abnormal interval correction module is used to correct the data detection abnormal interval based on the normal data fluctuation interval to obtain a data correction abnormal interval;

[0054] The correction data merging module is used to perform information nesting and merging on the data correction probability pairs and the data correction anomaly intervals to obtain time series data anomaly detection information.

[0055] To achieve the above-mentioned purpose, the third aspect of an embodiment of the present application proposes an electronic device, which includes a memory and a processor, the memory stores a computer program, and the processor implements the method of the above-mentioned first aspect when executing the computer program.

[0056] To achieve the above-mentioned purpose, the fourth aspect of the embodiments of the present application proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the method of the above-mentioned first aspect.

[0057] The time series data anomaly detection method and device, electronic device and storage medium proposed in the present application provide a reliable data basis for time series data anomaly detection by obtaining a time series data annotated sample set, and then perform model training on the pre-built initial anomaly detection model based on the time series data annotated sample set to obtain a target anomaly detection model, thereby ensuring that the target anomaly detection model can learn accurate normal data features and abnormal data features. Furthermore, the fluctuation interval of the time series data annotated sample set is predicted to obtain the normal data fluctuation interval, thereby providing a detection benchmark for time series data anomaly detection. Then, based on the time series data annotated sample set and the normal data fluctuation interval, correction parameter prediction is performed to obtain the detection result correction parameter, thereby facilitating the input of the target anomaly detection model. The optimization is carried out to improve the accuracy of time series data anomaly detection. Secondly, based on the target anomaly detection model, anomaly detection is performed on the preset target time series data to obtain data detection probability pairs and data detection anomaly intervals, which provides an optimizable preliminary detection result for time series data anomaly detection. Finally, according to the normal data fluctuation interval, the detection result correction parameter and the target time series data, the data detection probability pairs are probability corrected to obtain data correction probability pairs. According to the normal data fluctuation interval, the data detection anomaly interval is interval corrected to obtain data correction anomaly interval. The data correction probability pairs and the data correction anomaly interval are nested and merged to obtain time series data anomaly detection information, which can improve the accuracy and reliability of time series data anomaly detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 This is a flowchart of a method for detecting anomalies in time series data provided by an embodiment of the present application;

[0059] Figure 2 yes Figure 1 Flowchart of step S102 in FIG.

[0060] Figure 3 yes Figure 2 Flowchart of step S103 in FIG.

[0061] Figure 4 yes Figure 1 Flowchart of step S104 in FIG.

[0062] Figure 5 yes Figure 1 Flowchart of step S404 in FIG.

[0063] Figure 6 yes Figure 5 Flowchart of step S106 in FIG.

[0064] Figure 7 yes Figure 5 Flowchart of step S107 in FIG.

[0065] Figure 8 Schematic diagram of the structure of the time series data anomaly detection device provided in an embodiment of the present application;

[0066] Figure 9 This is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0067] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0068] It should be noted that although the device schematics illustrate functional module divisions and the flowcharts illustrate logical sequences, in certain circumstances, the steps shown or described may be performed in a sequence that differs from the module divisions in the device or the sequence in the flowcharts. The terms "first," "second," and so on, in the specification, claims, and drawings, are used to distinguish similar items and are not necessarily used to describe a specific sequence or precedence.

[0069] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0070] First, let’s analyze some of the terms used in this application:

[0071] Generalized Autoregressive Conditional Heteroskedasticity Model (GACHM): GACHM is a statistical model used in time series analysis, primarily to describe and predict data volatility. It is particularly widely used in the financial sector. By incorporating an autoregressive structure into the conditional variance, it considers not only the influence of past residuals but also the influence of past conditional variance, effectively capturing volatility persistence with fewer parameters. GACHM can capture dynamic characteristics of financial markets, such as volatility clustering, where periods of high and low volatility alternate, providing important volatility inputs for risk management and derivative pricing.

[0072] Time series data: A collection of data points arranged in chronological order, time series data records the observed values ​​of a specific variable or multiple variables at different points in time. It is time-dependent, with a sequential order between data points. It often exhibits characteristics such as trends, seasonality, and random fluctuations. It is widely used in fields such as finance, meteorology, and healthcare to analyze and predict dynamic behavior over time.

[0073] Artificial intelligence (AI) is a new technical discipline that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. A branch of computer science, AI seeks to understand the essence of intelligence and create new intelligent machines that can respond in a manner similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thinking. It also encompasses theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results.

[0074] Time series data anomaly detection refers to the process of identifying data points or intervals in time series data that deviate significantly from the normal pattern.

[0075] Currently, common time series data anomaly detection usually uses a trained machine learning model to perform anomaly detection on the input time series data, and uses the output of the machine learning model as the anomaly detection result of the time series data. However, the interpretability of the machine learning model is poor, resulting in inaccurate output anomaly detection results. Therefore, how to improve the accuracy of time series data anomaly detection has become a technical problem that needs to be solved urgently.

[0076] Based on this, the embodiments of the present application provide a time series data anomaly detection method and device, an electronic device, and a storage medium, aiming to improve the accuracy of time series data anomaly detection.

[0077] The time series data anomaly detection method and device, electronic device, and storage medium provided in the embodiments of the present application are specifically illustrated through the following embodiments. First, the time series data anomaly detection method in the embodiments of the present application is described.

[0078] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results.

[0079] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0080] The time series data anomaly detection method provided in the embodiment of the present application relates to the field of artificial intelligence technology. The time series data anomaly detection method provided in the embodiment of the present application can be applied to a terminal, can be applied to a server side, and can also be software running in a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, or can be configured as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the time series data anomaly detection method, etc., but is not limited to the above forms.

[0081] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.

[0082] Figure 1 This is an optional flowchart of the time series data anomaly detection method provided in an embodiment of the present application. Figure 1 The method may include but is not limited to steps S101 to S108.

[0083] Step S101: Acquire a time series data annotated sample set, wherein the time series data annotated sample set includes time series data, positive and negative data labels, and abnormal interval labels. The data labels are used to indicate whether the time series data is normal or abnormal, and the abnormal interval labels are used to indicate abnormal data intervals in the time series data.

[0084] Step S102: Based on the time series data labeled sample set, the pre-built initial anomaly detection model is trained to obtain a target anomaly detection model;

[0085] Step S103: performing fluctuation range prediction on the time series data annotated sample set to obtain a normal data fluctuation range;

[0086] Step S104: Based on the time series data labeled sample set and the normal data fluctuation range, correction parameter prediction is performed to obtain the detection result correction parameter;

[0087] Step S105: Based on the target anomaly detection model, perform anomaly detection on the preset target time series data to obtain a data detection probability pair and a data detection anomaly interval, wherein the data prediction probability pair is used to characterize the probability of the target time series data being normal and abnormal;

[0088] Step S106, based on the normal data fluctuation range, the detection result correction parameter and the target time series data, the data detection probability pair is probability corrected to obtain a data correction probability pair;

[0089] Step S107, based on the normal data fluctuation interval, performing interval correction on the data detection abnormal interval to obtain a data correction abnormal interval;

[0090] Step S108 , nesting and merging the data correction probability pairs and the data correction abnormal intervals to obtain time series data abnormality detection information.

[0091] Steps S101 to S108 shown in the embodiment of the present application provide a reliable data basis for time series data anomaly detection by obtaining a time series data annotated sample set, and then perform model training on the pre-built initial anomaly detection model based on the time series data annotated sample set to obtain a target anomaly detection model, thereby ensuring that the target anomaly detection model can learn accurate normal data features and abnormal data features. Furthermore, a fluctuation interval prediction is performed on the time series data annotated sample set to obtain a normal data fluctuation interval, thereby providing a detection benchmark for time series data anomaly detection. Then, a correction parameter prediction is performed based on the time series data annotated sample set and the normal data fluctuation interval to obtain a detection result correction parameter, thereby facilitating the optimization of the output of the target anomaly detection model. ization, thereby improving the accuracy of time series data anomaly detection. Secondly, based on the target anomaly detection model, anomaly detection is performed on the preset target time series data to obtain data detection probability pairs and data detection anomaly intervals, providing optimizable preliminary detection results for time series data anomaly detection. Finally, according to the normal data fluctuation interval, the detection result correction parameter and the target time series data, the data detection probability pairs are probability corrected to obtain data correction probability pairs. According to the normal data fluctuation interval, the data detection anomaly interval is interval corrected to obtain data correction anomaly interval. The data correction probability pairs and the data correction anomaly interval are nested and merged to obtain time series data anomaly detection information, which can improve the accuracy and reliability of time series data anomaly detection.

[0092] In step S101 of some embodiments, the time series data annotation sample set refers to a collection of multiple time series sample data, and the time series sample data refers to data points sorted in chronological order, such as daily closing prices of stocks, temperature records, heart rate monitoring data, etc. It should be noted that each time series sample data contains positive and negative data labels and abnormal interval labels, wherein the positive and negative data labels include normal data labels and abnormal data labels. The normal data label indicates that the time series sample data is normal data, and the abnormal data label indicates that the time series sample data is abnormal data. The abnormal interval label refers to the time series sample data. The abnormal data in this data is located in the interval of the time series data. For example, for the time series sample data {100, 101, 102, 100, 103, 85, 90, 87, 99, 102} with the abnormal data {85, 90, 87}, the positive data label can be the abnormal data label, and the abnormal interval label can be (6, 8), where 6 indicates that the starting point of the interval where the abnormal data in the time series sample data is located is the sixth data point in the time series sample data, and 8 indicates that the end point of the interval where the abnormal data in the time series sample data is located is the eighth data point in the time series sample data.

[0093] The embodiments of the present application can collect initial time series data through various methods such as database query, file reading or network crawling, and then align the initial time series data to obtain time series sample data to ensure that the data length of the time series sample data remains consistent. Furthermore, the time series sample data can be labeled with positive and negative data labels and abnormal interval labels based on the labeling information attached when the time series sample data is collected.

[0094] In an embodiment of the present application, by obtaining time series sample data, a reliable data basis is provided for the training of the initial anomaly detection model. In addition, by labeling the time series sample data with positive and negative data labels and abnormal interval labels, the accuracy of the initial anomaly detection model in identifying normal data and abnormal data in the time series data can be improved.

[0095] In step S102 of some embodiments, the initial anomaly detection model refers to a machine learning model capable of detecting anomaly probabilities in time series data. Furthermore, the initial anomaly detection model can also detect intervals within the time series data containing anomalies. The target anomaly detection model refers to an anomaly detection model applicable to the type of time series sample data. For example, when the time series sample data represents the mean pulmonary artery pressure of adult males over a specific period of time, the target anomaly detection model is applicable to anomaly detection in the field of mean pulmonary artery pressure of adult males.

[0096] The embodiments of the present application can train a pre-built initial anomaly detection model through supervised learning, thereby improving the detection accuracy of the trained target anomaly detection model.

[0097] For details, see Figure 2 In some embodiments, step S102 may include but is not limited to steps S201 to S204:

[0098] Step S201: Based on the initial anomaly detection model, perform anomaly data prediction on the time series data labeled sample set to obtain predicted positive and negative probability pairs and predicted data anomaly intervals;

[0099] Step S202: Based on the positive and negative data labels, the loss of the predicted positive and negative probability pairs is calculated to obtain probability pair loss data;

[0100] Step S203: Based on the abnormal interval label, calculate the loss of the abnormal interval of the predicted data to obtain interval loss data;

[0101] Step S204 , based on the probability pair loss data and interval loss data, adjust the parameters of the initial anomaly detection model to obtain a target anomaly detection model.

[0102] In step S201 of some embodiments, the predicted positive-differential probability pair includes a predicted normal probability and a predicted abnormal probability, wherein the predicted normal probability represents the probability that the time series sample data detected by the initial anomaly detection model is normal data, and the predicted abnormal probability represents the probability that the time series sample data detected by the initial anomaly detection model is abnormal data. The predicted data abnormal interval includes a predicted interval start point and a predicted interval end point, wherein the predicted interval start point represents the starting position of the abnormal data in the time series sample data detected by the initial anomaly detection model, and the predicted interval end point represents the ending position of the abnormal data in the time series sample data detected by the initial anomaly detection model.

[0103] In an embodiment of the present application, the initial anomaly detection model can perform hidden state update processing on the data points in the time series sample data in sequence according to the input order of the time series sample data, and after processing all the data points and obtaining the target hidden state vector, substitute the target hidden state vector into a pre-set activation function, which can generate the probability of normal data and abnormal data for each time series sample data in the time series data annotation sample set, as well as the interval position of the abnormal data in the time series sample data.

[0104] In steps S202 and S203 of some embodiments, the probability pair loss data characterizes the gap between the predicted positive and negative probability pairs and the actual positive and negative probability of the time series sample data, wherein the actual normal probability and abnormal probability of the time series sample data are related to the positive and negative data labels of the time series sample data. For example, the positive and negative data labels of the time series sample data {100, 101, 102, 100, 103, 85, 90, 87, 99, 102} are abnormal data labels, then the actual positive and negative probability of the time series sample data is a normal probability of 0% and an abnormal probability of 100%. The interval loss data represents the gap between the predicted data anomaly interval and the actual data anomaly interval of the time series sample data, wherein the actual data anomaly interval of the time series sample data is related to the anomaly interval label of the time series sample data. For example, the anomaly interval label of the time series sample data {100, 101, 102, 100, 103, 85, 90, 87, 99, 102} is (6, 8), then the actual data anomaly interval of the time series sample data is also (6, 8).

[0105] In an embodiment of the present application, the initial anomaly detection model can determine the actual positive deviation probability of each time series sample data in the time series data annotation sample set based on the positive deviation data label, and then use the pre-constructed probability pair loss calculation function to calculate the gap between the predicted positive deviation probability pair and the actual positive deviation probability. Similarly, the initial anomaly detection model can also determine the actual data anomaly interval of each time series sample data in the time series data annotation sample set based on the anomaly interval label, and then use the pre-constructed interval loss calculation function to calculate the gap between the predicted data anomaly interval pair and the actual data anomaly interval.

[0106] It is important to know that the above probability loss calculation function is as follows:

[0107] l BCE (s2,r)=-[rlogs2+(1-r)log(1-s2)]

[0108] Among them, l BCE Represents the probability of loss data, s2 represents the predicted abnormal probability in the predicted positive and negative probability, and r represents the positive and negative data label.

[0109] It is also important to know that the above interval loss calculation function is as follows:

[0110]

[0111] Among them, l MSE represents the interval loss data, t represents the abnormal interval of the predicted data, t * Represents the abnormal interval label, and i represents the i-th data abnormal interval.

[0112] In step S204 of some embodiments, after obtaining the probability pair loss data and the interval loss data, the initial anomaly detection model is back-propagated based on the probability pair loss data and the interval loss data, so that the parameters of the initial anomaly detection model are adjusted, thereby obtaining a target anomaly detection model suitable for the field of time series sample data.

[0113] In steps S201 to S204 shown in the embodiment of the present application, the initial anomaly detection model is used to predict abnormal data for the time series data labeled sample set to obtain predicted positive and negative probability pairs and predicted data abnormal intervals, which provide basic data support for the optimization of the initial anomaly detection model. Furthermore, based on the positive and negative data labels, the loss of the predicted positive and negative probability pairs is calculated to obtain probability pair loss data. Based on the abnormal interval labels, the loss of the predicted data abnormal interval is calculated to obtain interval loss data. Based on the probability pair loss data and the interval loss data, the parameters of the initial anomaly detection model are adjusted to obtain the target anomaly detection model, which ensures the accuracy of the target anomaly detection model in classifying normal and abnormal time series sample data, enhances the accuracy of the target anomaly detection model in locating abnormal intervals of time series sample data, and ensures the practicality of the target anomaly detection model.

[0114] In step S103 of some embodiments, the normal data fluctuation interval refers to the data fluctuation range of the normal time series sample data.

[0115] The embodiment of the present application calculates the data heteroscedasticity of normal data in the time series sample data, and then determines the normal data fluctuation range of the time series data labeled sample set according to a pre-set confidence level.

[0116] For details, see Figure 3 In some embodiments, step S103 may include but is not limited to steps S301 to S304:

[0117] Step S301: Based on the positive and negative data labels, the time series sample data is screened to obtain normal sample data;

[0118] Step S302, performing mean calculation on normal sample data to obtain normal mean data;

[0119] Step S303, performing heteroscedasticity calculation on the normal mean data to obtain data heteroscedasticity;

[0120] Step S304: determining a normal data fluctuation range based on data heteroskedasticity.

[0121] In step S301 of some embodiments, normal sample data refers to time series data in which the positive and negative data labels are normal data labels in the time series sample data.

[0122] The embodiment of the present application can filter out the time series sample data that meets the filtering conditions by setting the filtering conditions, thereby leaving the required normal sample data. It should be noted that the filtering conditions can be used to filter out the time series sample data whose positive and negative data labels are abnormal data labels, and retain the time series sample data whose positive and negative data labels are normal data labels.

[0123] In step S302 of some embodiments, normal mean data refers to the mean data of each normal sample data. For example, normal sample data A is {20, 21, 21, 19, 18}, and normal sample data B is {18, 19, 17, 21, 18}, then the normal mean data can be {19, 20, 19, 20, 18}.

[0124] The embodiment of the present application calculates the mean of the data at the same time point in each normal sample data to obtain the mean data of each time point, and then sorts the mean data of each time point according to the order of the time points in the normal sample data to obtain normal mean data.

[0125] In step S303 of some embodiments, data heteroskedasticity refers to heteroskedastic data of normal mean data.

[0126] In the embodiment of the present application, a pre-trained generalized autoregressive conditional heteroscedasticity model can be used to calculate the heteroscedasticity data of the above-mentioned normal mean data, and the data heteroscedasticity can be obtained.

[0127] In step S304 of some embodiments, after obtaining the data heteroscedasticity of the time series sample data, the corresponding confidence interval can be constructed using the data heteroscedasticity according to a preset confidence level. For example, when the confidence level is set to 95.44% and the heteroscedasticity is σ 2 When , the confidence interval corresponding to the confidence level can be [-2σ, 2σ].

[0128] In steps S301 to S304 shown in the embodiment of the present application, the time series sample data is screened according to the positive and negative data labels to obtain normal sample data, which provides a reliable data basis for the prediction of the normal data fluctuation range. Furthermore, the mean of the normal sample data is calculated to obtain normal mean data, which provides the central trend of the data for the prediction of the normal data fluctuation range. Secondly, the heteroscedasticity of the normal mean data is calculated to obtain data heteroscedasticity, which realizes the calculation of the volatility characteristics of the normal sample data. Finally, the normal data fluctuation range is determined based on the data heteroscedasticity, which improves the reliability of time series data anomaly detection and ensures adaptability and practicality in complex time series data environments.

[0129] In step S104 of some embodiments, the detection result correction parameter refers to a parameter used to correct the abnormality detection result of the time series data.

[0130] After obtaining the normal data fluctuation range, the embodiment of the present application can obtain the number of data outside the normal data fluctuation range by recording the number of data in the time series sample data that exceeds the normal data fluctuation range, and then determine the detection result correction parameter based on the number of data outside the normal data fluctuation range and the pre-set correction parameter calculation formula.

[0131] For details, see Figure 4 In some embodiments, step S104 may include but is not limited to steps S401 to S404:

[0132] Step S401: classify the time series sample data based on the positive and abnormal data labels to obtain normal sample data and abnormal sample data;

[0133] Step S402: Based on the normal data fluctuation interval, record the out-of-interval data of the normal sample data to obtain the number of data outside the positive sample interval;

[0134] Step S403: Based on the normal data fluctuation interval, record the abnormal sample data outside the interval to obtain the number of data outside the abnormal sample interval;

[0135] Step S404 : Based on the number of data outside the positive sample interval and the number of data outside the outlier sample interval, correction parameters are calculated to obtain detection result correction parameters.

[0136] In step S401 of some embodiments, abnormal sample data refers to time series data in which the positive data label is an abnormal data label in the time series sample data.

[0137] The embodiments of the present application can classify the time series sample data according to the positive and negative data labels of the time series sample data. Specifically, the time series sample data whose positive and negative data labels are normal data labels are classified as normal sample data, and the time series sample data whose positive and negative data labels are abnormal data labels are classified as abnormal sample data.

[0138] In steps S402 and S403 of some embodiments, the number of data outside the positive sample interval refers to the set of data numbers in the normal sample data that exceed the normal data fluctuation interval. For example, there are normal sample data a{99, 100, 100}, normal sample data b{100, 100, 101}, normal sample data c{98, 101, 102}, and the normal data fluctuation interval is {99.5~100.5, 100.5~101, 101~102}, then the amount of data outside the normal interval can be {3, 2, 1}, where 3 indicates that the number of data in the normal sample data a that exceeds the normal data fluctuation interval is 3, 2 indicates that the number of data in the normal sample data b that exceeds the normal data fluctuation interval is 2, and 1 indicates that the number of data in the normal sample data c that exceeds the normal data fluctuation interval is 1. The number of data outside the negative sample interval refers to the set of data numbers in the abnormal sample data that exceed the normal data fluctuation range. For example, there are abnormal sample data d{85, 87, 101}, abnormal sample data e{90, 110, 115}, and abnormal sample data f{80, 101, 102}, and the normal data fluctuation range is {99.5~100.5, 100.5~101, 101~102}, then the amount of data outside the normal interval can be {2, 3, 1}, where 2 means that the number of data in the abnormal sample data d that exceeds the normal data fluctuation range is 2, 3 means that the number of data in the abnormal sample data e that exceeds the normal data fluctuation range is 3, and 1 means that the number of data in the abnormal sample data f that exceeds the normal data fluctuation range is 1.

[0139] The embodiment of the present application can traverse the normal sample data and the normal data fluctuation interval, and find out the positive sample out-of-interval data that exceeds the normal data fluctuation interval from the normal sample data. Further, the number of positive sample out-of-interval data is recorded to obtain the number of data outside the positive sample interval. Similarly, by traversing the abnormal sample data and the normal data fluctuation interval, and finding out the abnormal sample out-of-interval data that exceeds the normal data fluctuation interval from the abnormal sample data, and then recording the number of abnormal sample out-of-interval data, the number of abnormal sample out-interval data can be obtained.

[0140] In step S404 of some embodiments, after obtaining the number of data outside the positive sample interval and the number of data outside the abnormal sample interval, the number of data outside the positive sample interval and the number of data outside the abnormal sample interval can be calculated to determine the data difference between the normal sample data and the abnormal sample data that exceeds the normal data fluctuation range, and then the number of data outside the abnormal sample interval is median-processed to obtain the data median of the abnormal sample data that exceeds the normal data fluctuation range. Finally, the above data representing the data difference and the data median are merged to obtain the detection result correction parameter.

[0141] For details, see Figure 5 In some embodiments, step S404 may include but is not limited to steps S501 to S505:

[0142] Step S501, performing maximum screening on the number of data outside the positive sample interval to obtain the target number of positive samples;

[0143] Step S502, performing minimum screening on the number of data outside the outlier interval to obtain the target number of outlier samples;

[0144] Step S503, calculating the quantity difference between the number of outlier targets and the number of positive targets to obtain a quantity difference parameter;

[0145] Step S504, performing median calculation on the target number of different samples to obtain a median parameter of the number;

[0146] Step S505 , performing data nesting and merging on the quantity difference parameter and the quantity median parameter to obtain a detection result correction parameter.

[0147] In some embodiments, in steps S501 to S503, the number of positive target samples refers to the data with the largest value among the data outside the positive sample interval. For example, when the number of data outside the positive sample interval is {1, 1, 2, 2, 3}, the number of positive target samples is 3. The number of outlier targets refers to the data with the smallest value among the data outside the outlier interval. For example, when the number of data outside the outlier interval is {10, 11, 12, 9, 10}, the number of outlier targets is 9. The quantity difference parameter refers to the quantity data obtained by subtracting the number of positive target samples from the number of outlier targets. For example, when the number of positive target samples is 3 and the number of outlier targets is 9, the quantity difference parameter may be 6.

[0148] The embodiment of the present application can find the data with the largest value from the number of data outside the positive sample interval, that is, the number of positive sample targets, and find the data with the smallest value from the number of data outside the abnormal sample interval, that is, the number of abnormal sample targets, by traversing the number of data outside the positive sample interval and the number of data outside the abnormal sample interval. Furthermore, the quantity difference parameter can be obtained by subtracting the above-mentioned positive sample target number from the above-mentioned number of abnormal sample targets.

[0149] In step S504 of some embodiments, a pre-established formula for calculating the median number of different samples may be used to calculate the median number parameter of the target number of different samples. Specifically, the formula for calculating the median number of different samples is as follows:

[0150]

[0151] Among them, m represents the median parameter, j represents the number of abnormal sample data, G(ano j) represents the number of data outside the heterogeneous sample interval, minG j (ano j ) represents the number of different sample targets.

[0152] In step S505 of some embodiments, after obtaining the quantity difference parameter and the quantity median parameter, the quantity difference parameter and the quantity median parameter are embedded in a pre-constructed empty set to obtain the detection result correction parameter.

[0153] In steps S501 to S505 shown in the embodiment of the present application, the number of positive sample targets is obtained by performing maximum value screening on the number of data outside the positive sample interval, and the number of outlier targets is obtained by performing minimum value screening on the number of data outside the positive sample interval. Then, the number of outlier targets and the number of positive sample targets are calculated for quantity difference to obtain quantity difference parameters, which can improve the fault tolerance of the parameters obtained by calculating the correction parameters. Secondly, the number of outlier targets is median calculated to obtain the quantity median parameter, which stabilizes the statistical characteristics of the number of outlier targets. Finally, the quantity difference parameter and the quantity median parameter are nested and merged to obtain the detection result correction parameter, which improves the accuracy and reliability of the detection result correction parameter.

[0154] In steps S401 to S404 shown in the embodiment of the present application, the time series sample data is classified according to the positive and abnormal data labels to obtain normal sample data and abnormal sample data, which provides an accurate data basis for the calculation of the correction parameters. Furthermore, the normal sample data is recorded as out-of-interval data to obtain the number of data outside the positive sample interval, and the abnormal sample data is recorded as out-of-interval data to obtain the number of data outside the abnormal sample interval, which facilitates the understanding of the distribution characteristics of the normal sample data and the abnormal sample data, and provides a basis for the calculation of the correction parameters. Finally, the correction parameters are calculated based on the number of data outside the positive sample interval and the number of data outside the abnormal sample interval to obtain the correction parameters of the detection results, which can improve the accuracy and reliability of the time series data anomaly detection results.

[0155] In step S105 of some embodiments, the target time series data refers to time series data whose data is normal or abnormal. For example, the time series data {100, 101, 102, 100, 103, 85, 90, 87, 99, 102} without positive or negative data labels and abnormal interval labels can be the target time series data. The data detection probability pair includes a normal probability and an abnormal probability, wherein the normal probability refers to the probability that the target time series data detected by the target anomaly detection model is normal data, and the abnormal probability refers to the probability that the target time series data detected by the target anomaly detection model is abnormal data. For example, for the target time series data {100, 101, 102, 100, 103, 85, 90, 87, 99, 102}, the data detection probability pair detected by the target anomaly detection model can be (0.3, 0.7), wherein 0.3 is the normal probability, indicating that there is a 30% probability that the target time series data is normal data, and 0.7 is the abnormal probability, indicating that there is a 70% probability that the target time series data is abnormal data. The data detection anomaly interval includes the anomaly interval starting point and the anomaly interval end point, wherein the anomaly interval starting point refers to the starting position of the anomaly data in the target time series data detected by the target anomaly detection model, and the anomaly interval end point refers to the end position of the anomaly data in the target time series data detected by the target anomaly detection model. For example, for the target time series data {100, 101, 102, 100, 103, 85, 90, 87, 99, 102}, the data detection anomaly interval detected by the target anomaly detection model can be (7, 8), wherein 7 is the anomaly interval starting point, indicating that the starting point of the data detection anomaly interval is the seventh data point in the target time series data, and 8 is the anomaly interval end point, indicating that the end point of the data detection anomaly interval is the eighth data point in the target time series data.

[0156] By inputting the pre-acquired target time series data into the trained target anomaly detection model, the embodiment of the present application can obtain the data detection probability pair and data detection anomaly interval of the target time series data, realize the preliminary anomaly detection of the target time series data, and provide a data basis for the detection data correction of the target time series data.

[0157] In step S106 of some embodiments, the data correction probability pair refers to the corrected data detection probability pair. For example, when the target time series data is {100, 101, 102, 100, 103, 85, 90, 87, 99, 102} and the data detection probability pair is (0.3, 0.7), the data correction probability pair can be (0.15, 0.85), where 0.15 indicates that there is a 15% probability that the target time series data is normal data, and 0.85 indicates that there is an 85% probability that the target time series data is abnormal data.

[0158] The embodiment of the present application records the data that exceeds the normal data fluctuation range in the target time series data, and thus can obtain the number of out-of-interval data of the target time series data. Furthermore, according to the number of out-of-interval data of the target time series data and the above-mentioned detection result correction parameters, the data detection probability pair is modified to obtain the data correction probability pair.

[0159] For details, see Figure 6 In some embodiments, step S106 may include but is not limited to steps S601 to S604:

[0160] Step S601: Based on the normal data fluctuation interval, record the out-of-interval data of the target time series data to obtain the number of data outside the target interval;

[0161] Step S602, based on the number of data outside the target interval and the detection result correction parameter, the data detection probability pair is corrected to normal probability to obtain the data normal correction probability;

[0162] Step S603, determining the probability of data abnormality correction based on the probability of data normality correction;

[0163] Step S604 , performing data nesting and merging on the data normal correction probability and the data abnormal correction probability to obtain a data correction probability pair.

[0164] In step S601 of some embodiments, the amount of data outside the target interval refers to the amount of data in the target time series data that exceeds the normal data fluctuation interval.

[0165] The embodiment of the present application can find out-of-range data that exceeds the normal data fluctuation range from the target time series data by traversing the target time series data and the normal data fluctuation range, and record the number of out-of-range data to obtain the number of data outside the target range.

[0166] In step S602 of some embodiments, the data normal correction probability refers to the probability data obtained by correcting the normal probability in the data detection probability.

[0167] In the embodiment of the present application, a normal probability correction formula can be pre-constructed to calculate the normal probability of the data after the normal probability of the data detection probability is corrected. The normal probability correction formula is as follows:

[0168]

[0169] Among them, x represents the target time series data, s ′ (x)1 represents the probability of normal correction of data, s(x)1 represents the normal probability of the data detection probability pair, G(x) represents the number of data outside the target interval, m represents the median parameter of the number, and l represents the number difference parameter.

[0170] In step S603 of some embodiments, the data anomaly correction probability refers to the probability data obtained by correcting the anomaly probability in the data detection probability.

[0171] After obtaining the probability of normal data correction, the embodiment of the present application can calculate the probability of abnormal data correction based on the reason why the sum of the probability pairs is 1. For example, when the probability of normal data correction is 10%, the probability of abnormal data correction is 100%-10%=90%.

[0172] In step S604 of some embodiments, the data detection probability pair can be converted into a data correction probability pair by replacing the normal probability in the data detection probability pair with the above-mentioned corrected data normal probability and replacing the abnormal probability in the data detection probability pair with the corrected data abnormal probability.

[0173] In steps S601 to S604 shown in the embodiment of the present application, out-of-interval data is recorded for the target time series data according to the normal data fluctuation interval, and the number of data outside the target interval is obtained, which provides basic data support for the correction of the data detection probability pair. Furthermore, based on the number of data outside the target interval and the detection result correction parameters, the data detection probability pair is corrected for normal probability, and based on the data normal correction probability, the data abnormal correction probability is determined, and then the data normal correction probability and the data abnormal correction probability are nested and merged to obtain the data correction probability pair, thereby improving the accuracy and reliability of the conclusion of time series data anomaly detection.

[0174] In step S107 of some embodiments, the data correction anomaly interval refers to the corrected data detection anomaly interval. For example, when the target time series data is {100, 101, 102, 100, 103, 85, 90, 87, 99, 102} and the data detection anomaly interval is (7, 8), the data correction anomaly interval can be (6, 8), where 6 indicates that the starting point of the data detection anomaly interval is the sixth data point in the target time series data, and 8 indicates that the end point of the data detection anomaly interval is the eighth data point in the target time series data.

[0175] The embodiment of the present application constructs a detection interval for detecting the position of abnormal data intervals in the target time series data, so as to obtain the number of data exceeding the normal data fluctuation interval in each detection interval, and based on the number of data exceeding the normal data fluctuation interval in each detection interval, screen out the intervals that meet the pre-set screening conditions from the detection interval as the data correction abnormality interval.

[0176] For details, see Figure 7 In some embodiments, step S107 may include but is not limited to steps S701 to S705:

[0177] Step S701: Based on the data detection abnormal interval, the target time series data is interval-screened to obtain a target detection interval, wherein the target detection interval includes a detection interval start point and a detection interval end point, the detection interval start point is less than or equal to the abnormal interval start point, and the detection interval end point is greater than or equal to the abnormal interval end point;

[0178] Step S702, record data of the target detection interval to obtain the number of interval data;

[0179] Step S703: Based on the normal data fluctuation interval, record the data outside the target detection interval to obtain the number of data outside the detection interval;

[0180] Step S704: Calculate the abnormal data ratio based on the number of interval data and the number of data outside the detection interval to obtain abnormal proportion data;

[0181] Step S705 : Based on the abnormal ratio data, the target detection interval is screened to obtain a data correction abnormal interval, wherein the abnormal ratio data of the data correction abnormal interval is the largest and the interval range of the data correction abnormal interval is the largest.

[0182] In step S701 of some embodiments, the target detection interval refers to the interval between the start point of the detection interval and the end point of the detection interval, wherein the start point of the detection interval represents any position in the target time series data that is less than or equal to the start point of the above-mentioned abnormal interval, and the end point of the detection interval represents any position in the target time series data that is greater than or equal to the end point of the above-mentioned abnormal interval. For example, when the target time series data includes ten data, and the start point of the abnormal interval is 6 and the end point of the abnormal interval is 8, the start point of the detection interval can be any position from 1 to 6, and the end point of the detection interval can be any position from 8 to 10.

[0183] The embodiment of the present application detects abnormal intervals based on data and can determine the starting point and end point of the abnormal interval. On this basis, an interval position whose range is greater than or equal to the abnormal interval, the starting point of the detection interval is less than or equal to the starting point of the abnormal interval, and the end point of the detection interval is greater than or equal to the end point of the abnormal interval is selected to obtain the target detection interval.

[0184] In steps S702 to S704 of some embodiments, the number of interval data refers to the number of data within the target detection interval. The number of data outside the detection interval refers to the number of data within the target detection interval that exceeds the normal data fluctuation range. The abnormality ratio data refers to the ratio of the number of data outside the detection interval to the number of interval data. It should be noted that the larger the abnormality ratio data, the greater the probability that the target detection interval is an interval containing abnormal data, and the smaller the abnormality ratio data, the lower the probability that the target detection interval is an interval containing abnormal data.

[0185] The embodiment of the present application can determine the number of interval data by traversing the data in the target detection interval, and then compare the data in the target detection interval with the data interval in the normal data fluctuation interval, filter out the data that exceeds the normal data fluctuation interval from the data in the target detection interval, and record the number, so as to obtain the number of data outside the detection interval. Finally, the ratio of the number of data outside the detection interval to the number of interval data is calculated to obtain abnormal proportion data.

[0186] In step S705 of some embodiments, a pre-built interval correction formula may be used to filter out data correction abnormal intervals from the target detection interval, wherein the interval correction formula is as follows:

[0187]

[0188] Among them, A represents the target detection interval set that meets the formula conditions, t ′ (x)1 represents the starting point of the detection interval, t ′ (x)2 represents the end point of the detection interval, t(x)1 represents the starting point of the abnormal interval, t(x)2 represents the end point of the abnormal interval, R represents the abnormal ratio data, and B represents the data correction abnormal interval.

[0189] In steps S701 to S705 shown in the embodiment of the present application, the target time series data is interval-screened according to the data detection anomaly interval to obtain the target detection interval, which provides a clear interval range for the interval correction of the data detection anomaly interval. Secondly, data is recorded for the target detection interval to obtain the number of interval data. Based on the normal data fluctuation interval, the out-of-interval data of the target detection interval is recorded to obtain the number of data outside the detection interval. The abnormal data ratio is calculated for the number of interval data and the number of data outside the detection interval to obtain abnormal proportion data, which provides a quantitative basis for the interval correction of the data detection anomaly interval. Finally, the target detection interval is screened based on the abnormal proportion data to obtain the data correction abnormal interval, thereby improving the accuracy and reliability of the abnormal interval of time series data anomaly detection.

[0190] In step S108 of some embodiments, the time series data anomaly detection information refers to the information obtained after the target time series data is detected for anomalies. It should be noted that the time series data anomaly detection information includes information such as the normal probability, abnormal probability and interval position of the abnormal data of the target time series data.

[0191] The embodiment of the present application replaces the data detection probability pair with the data correction probability and replaces the data detection anomaly interval with the data correction anomaly interval, thereby nesting and merging the information of the data correction probability pair and the data correction anomaly interval, thereby obtaining time series data anomaly detection information and improving the accuracy of time series data anomaly detection.

[0192] This application provides a reliable data basis for time series data anomaly detection by obtaining a time series data annotated sample set, and then trains the pre-built initial anomaly detection model based on the time series data annotated sample set to obtain a target anomaly detection model, ensuring that the target anomaly detection model can learn accurate normal data features and abnormal data features. Furthermore, the fluctuation interval of the time series data annotated sample set is predicted to obtain the normal data fluctuation interval, which provides a detection benchmark for time series data anomaly detection. Then, based on the time series data annotated sample set and the normal data fluctuation interval, the correction parameter prediction is performed to obtain the detection result correction parameter, which is convenient for optimizing the output of the target anomaly detection model, thereby improving the time series data. The accuracy of anomaly detection. Secondly, based on the target anomaly detection model, anomaly detection is performed on the preset target time series data to obtain data detection probability pairs and data detection anomaly intervals, providing optimizable preliminary detection results for time series data anomaly detection. Finally, according to the normal data fluctuation interval, the detection result correction parameters and the target time series data, the data detection probability pairs are probability corrected to obtain data correction probability pairs. According to the normal data fluctuation interval, the data detection anomaly interval is interval corrected to obtain data correction anomaly intervals. The data correction probability pairs and the data correction anomaly intervals are information nested and merged to obtain time series data anomaly detection information, which can improve the accuracy and reliability of time series data anomaly detection.

[0193] See also Figure 8 The present application also provides a device for detecting anomalies in time series data, which can implement the above-mentioned method for detecting anomalies in time series data. The device includes:

[0194] The sample data acquisition module 801 is used to acquire a time series data annotated sample set, wherein the time series data annotated sample set includes time series sample data, positive and negative data labels, and abnormal interval labels. The data labels are used to indicate whether the time series sample data is normal or abnormal, and the abnormal interval labels are used to indicate abnormal data intervals in the time series sample data.

[0195] The initial model training module 802 is used to train the pre-built initial anomaly detection model based on the time series data labeled sample set to obtain a target anomaly detection model;

[0196] The fluctuation interval prediction module 803 is used to predict the fluctuation interval of the time series data annotated sample set to obtain the normal data fluctuation interval;

[0197] The correction parameter prediction module 804 is used to predict the correction parameters based on the time series data labeled sample set and the normal data fluctuation range to obtain the correction parameters of the detection results;

[0198] The data anomaly detection module 805 is used to perform anomaly detection on the preset target time series data based on the target anomaly detection model to obtain a data detection probability pair and a data detection anomaly interval, wherein the data prediction probability pair is used to represent the probability of normality and anomaly of the target time series data;

[0199] The probability pair correction module 806 is used to perform probability correction on the data detection probability pair based on the normal data fluctuation range, the detection result correction parameter and the target time series data to obtain a data correction probability pair;

[0200] The abnormal interval correction module 807 is used to correct the data detection abnormal interval based on the normal data fluctuation interval to obtain the data correction abnormal interval;

[0201] The correction data merging module 808 is used to perform information nesting and merging on the data correction probability pairs and the data correction anomaly intervals to obtain time series data anomaly detection information.

[0202] The specific implementation of the time series data anomaly detection device is basically the same as the specific embodiment of the above-mentioned time series data anomaly detection method, and will not be repeated here.

[0203] The present application also provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the above-described time series data anomaly detection method when executing the computer program. The electronic device can be any smart terminal, such as a tablet computer or an in-vehicle computer.

[0204] See also Figure 9 , Figure 9 The hardware structure of an electronic device according to another embodiment is shown. The electronic device includes:

[0205] The processor 901 can be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;

[0206] The memory 902 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 902 and is called by the processor 901 to execute the time series data anomaly detection method of the embodiment of the present application;

[0207] Input / output interface 903, used to implement information input and output;

[0208] Communication interface 904, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);

[0209] Bus 905 , which transmits information between various components of the device (e.g., processor 901 , memory 902 , input / output interface 903 , and communication interface 904 );

[0210] The processor 901 , the memory 902 , the input / output interface 903 and the communication interface 904 are connected to each other in communication within the device via a bus 905 .

[0211] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned time series data anomaly detection method is implemented.

[0212] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0213] The embodiments of the present application provide a time series data anomaly detection method, a time series data anomaly detection device, an electronic device, and a storage medium. The method obtains a time series data annotated sample set, wherein the time series data annotated sample set includes time series sample data, positive and negative data labels, and abnormal interval labels. The data labels are used to characterize whether the time series sample data is normal or abnormal, and the abnormal interval labels are used to characterize the abnormal data interval in the time series sample data. Based on the time series data annotated sample set, a pre-built initial anomaly detection model is trained to obtain a target anomaly detection model. Furthermore, a fluctuation interval prediction is performed on the time series data annotated sample set to obtain a normal data fluctuation interval. Based on the time series data annotated sample set and the normal data fluctuation interval, a pre-built initial anomaly detection model is trained to obtain a target anomaly detection model. Interval, correction parameter prediction is performed to obtain the correction parameter of the detection result. Secondly, based on the target anomaly detection model, anomaly detection is performed on the preset target time series data to obtain data detection probability pairs and data detection anomaly intervals, wherein the data prediction probability pairs are used to characterize the probabilities of normal and abnormal target time series data. Finally, based on the normal data fluctuation interval, the detection result correction parameter and the target time series data, the data detection probability pairs are probability corrected to obtain data correction probability pairs. Based on the normal data fluctuation interval, the data detection anomaly interval is interval corrected to obtain data correction anomaly interval. The data correction probability pairs and the data correction anomaly interval are information nested and merged to obtain time series data anomaly detection information.

[0214] The embodiments described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0215] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.

[0216] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.

[0217] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.

[0218] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0219] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0220] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. The mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0221] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0222] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0223] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes multiple instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: various media that can store programs, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0224] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.

Claims

1. A time series data anomaly detection method, characterized in that: The method comprises: Obtain a time series data annotated sample set, wherein the time series data annotated sample set includes time series sample data, positive and negative data labels, and abnormal interval labels, wherein the data labels are used to indicate whether the time series sample data is normal or abnormal, and the abnormal interval labels are used to indicate abnormal data intervals in the time series sample data; Based on the time series data labeled sample set, a pre-built initial anomaly detection model is trained to obtain a target anomaly detection model; Predicting the fluctuation range of the time series data annotated sample set to obtain a normal data fluctuation range; Based on the time series data annotated sample set and the normal data fluctuation range, correction parameter prediction is performed to obtain the detection result correction parameter; Based on the target anomaly detection model, anomaly detection is performed on the preset target time series data to obtain a data detection probability pair and a data detection anomaly interval, wherein the data prediction probability pair is used to characterize the probability of normality and anomaly of the target time series data; Based on the normal data fluctuation interval, the detection result correction parameter and the target time series data, probability correction is performed on the data detection probability pair to obtain a data correction probability pair; Based on the normal data fluctuation interval, performing interval correction on the data detection abnormal interval to obtain a data correction abnormal interval; Information nesting and merging is performed on the data correction probability pair and the data correction abnormal interval to obtain time series data abnormality detection information.

2. The method according to claim 1, characterized in that The method of training a pre-built initial anomaly detection model based on the time series data labeled sample set to obtain a target anomaly detection model includes: Based on the initial anomaly detection model, abnormal data prediction is performed on the time series data labeled sample set to obtain predicted positive and negative probability pairs and predicted data anomaly intervals; Based on the positive and negative data labels, performing loss calculation on the predicted positive and negative probability pairs to obtain probability pair loss data; Based on the abnormal interval label, performing loss calculation on the abnormal interval of the predicted data to obtain interval loss data; Based on the probability pair loss data and the interval loss data, parameters of the initial anomaly detection model are adjusted to obtain the target anomaly detection model.

3. The method according to claim 1, characterized in that The performing of fluctuation range prediction on the time series data labeled sample set to obtain a normal data fluctuation range includes: Based on the positive and negative data labels, the time series sample data is screened to obtain normal sample data; Performing mean calculation on the normal sample data to obtain normal mean data; Performing heteroscedasticity calculation on the normal mean data to obtain data heteroscedasticity; Based on the data heteroskedasticity, the normal data fluctuation range is determined.

4. The method according to claim 3, characterized in that The method of predicting correction parameters based on the time series data labeled sample set and the normal data fluctuation interval to obtain the detection result correction parameters includes: Based on the positive and abnormal data labels, the time series sample data is classified to obtain the normal sample data and the abnormal sample data; Based on the normal data fluctuation interval, the normal sample data is recorded as out-of-interval data to obtain the number of data outside the positive sample interval; Based on the normal data fluctuation interval, the abnormal sample data is recorded as out-of-interval data to obtain the number of data outside the abnormal sample interval; Based on the number of data outside the positive sample interval and the number of data outside the outlier sample interval, a correction parameter is calculated to obtain the detection result correction parameter.

5. The method according to claim 4, characterized in that The calculation of the correction parameter based on the number of data outside the positive sample interval and the number of data outside the outlier sample interval to obtain the detection result correction parameter includes: Performing maximum screening on the number of data outside the positive sample interval to obtain the target number of positive samples; Perform minimum screening on the number of data outside the outlier sample interval to obtain the target number of outlier samples; Calculating the quantity difference between the number of outlier targets and the number of positive sample targets to obtain a quantity difference parameter; Performing median calculation on the target number of the heterogeneous samples to obtain a median parameter of the number; The quantity difference parameter and the quantity median parameter are subjected to data nesting and merging to obtain the detection result correction parameter.

6. The method according to any one of claims 1 to 5, characterized in that The method of performing probability correction on the data detection probability pair based on the normal data fluctuation interval, the detection result correction parameter, and the target time series data to obtain a data correction probability pair includes: Based on the normal data fluctuation interval, the target time series data is recorded as out-of-interval data to obtain the number of data outside the target interval; Based on the amount of data outside the target interval and the detection result correction parameter, performing normal probability correction on the data detection probability pair to obtain a data normal correction probability; Determining a data abnormality correction probability based on the data normal correction probability; The normal data correction probability and the abnormal data correction probability are nested and merged to obtain the data correction probability pair.

7. The method according to any one of claims 1 to 5, characterized in that The data detection abnormal interval includes an abnormal interval start point and an abnormal interval end point. Based on the normal data fluctuation interval, the data detection abnormal interval is corrected to obtain a data correction abnormal interval, including: Based on the data detection abnormal interval, the target time series data is interval-screened to obtain a target detection interval, wherein the target detection interval includes a detection interval start point and a detection interval end point, the detection interval start point is less than or equal to the abnormal interval start point, and the detection interval end point is greater than or equal to the abnormal interval end point; Recording data for the target detection interval to obtain the number of interval data; Based on the normal data fluctuation interval, recording the data outside the target detection interval to obtain the number of data outside the detection interval; Based on the number of data in the interval and the number of data outside the detection interval, the abnormal data ratio is calculated to obtain abnormal proportion data; Based on the abnormal ratio data, the target detection interval is screened to obtain the data correction abnormal interval, wherein the abnormal ratio data of the data correction abnormal interval is the largest and the interval range of the data correction abnormal interval is the largest.

8. A time series data anomaly detection device, characterized in that: The device comprises: A sample data acquisition module is used to acquire a time series data annotated sample set, wherein the time series data annotated sample set includes time series sample data, positive and negative data labels, and abnormal interval labels. The data labels are used to indicate whether the time series sample data is normal or abnormal, and the abnormal interval labels are used to indicate abnormal data intervals in the time series sample data. An initial model training module is used to train a pre-built initial anomaly detection model based on the time series data labeled sample set to obtain a target anomaly detection model; A fluctuation interval prediction module is used to predict the fluctuation interval of the time series data annotated sample set to obtain a normal data fluctuation interval; A correction parameter prediction module is used to predict the correction parameters based on the time series data labeled sample set and the normal data fluctuation range to obtain the correction parameters of the detection result; A data anomaly detection module is used to perform anomaly detection on the preset target time series data based on the target anomaly detection model to obtain a data detection probability pair and a data detection anomaly interval, wherein the data prediction probability pair is used to represent the probability of normality and anomaly of the target time series data; a probability pair correction module, configured to perform probability correction on the data detection probability pair based on the normal data fluctuation interval, the detection result correction parameter, and the target time series data to obtain a data correction probability pair; An abnormal interval correction module is used to correct the data detection abnormal interval based on the normal data fluctuation interval to obtain a data correction abnormal interval; The correction data merging module is used to perform information nesting and merging on the data correction probability pairs and the data correction anomaly intervals to obtain time series data anomaly detection information.

9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the data table partitioning method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the data table partitioning method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Method and device for detecting abnormal points in time sequence, electronic device and storage medium

    CN109784042A

  • Electrocardiogram abnormity labeling method and device

    CN110693486A

  • Abnormal sample detection method and device, electronic equipment and storage medium

    CN115204278A

  • Method and device for detecting abnormality of time series data of database

    CN116541743A