Data anomaly detection method and device, equipment, storage medium and program product
By calculating the differential statistics of time-series data, the data fluctuation range is determined, which solves the problem of low accuracy in anomaly detection in existing technologies and achieves more efficient anomaly identification.
Patent Information
- Application Number
- CN202410598242.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-14
- Publication Date
- 2025-11-14
AI Technical Summary
Existing technologies have low accuracy in detecting abnormal data in large-scale network environments, and are prone to false alarms, especially in non-stationary data recovery scenarios.
By calculating differential statistics on the time series data to be detected, the data fluctuation range is determined, and abnormal data is identified based on the differential statistics. The accuracy of detection is improved by utilizing the differential statistics, which include the sub-differential statistics of each sub-data.
It improves the accuracy of anomaly detection, reduces false alarms, and enables more precise identification of anomaly data.
Smart Images

Figure CN120956435A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a data anomaly detection method, apparatus, device, storage medium, and program product. Background Technology
[0002] In large-scale network environments, anomaly detection has a wide range of applications. Related technologies for anomaly detection are typically model-based, studying the characteristics of data distribution and then matching them with corresponding probability distributions for detection. When anomalies occur, such as network latency or server failures, abnormal fluctuations in system business-related data can occur. When related technologies transform non-stationary data, they often misjudge the time-series data corresponding to the detected anomaly, resulting in low accuracy in anomaly detection.
[0003] Currently, there is no good method to improve the accuracy of abnormal data detection in related technologies. Summary of the Invention
[0004] This application provides a data anomaly detection method, apparatus, electronic device, computer-readable storage medium, and computer program product, which can improve the accuracy of anomaly data detection.
[0005] The technical solution of this application embodiment is implemented as follows:
[0006] This application provides a data anomaly detection method, the method comprising:
[0007] Acquire the time series data to be detected, wherein the time series data to be detected includes multiple sub-data with different time sequences;
[0008] For each of the sub-data, the following processing is performed: the sub-data is taken as the first sub-data, and based on the first sub-data and a plurality of second sub-data corresponding to the first sub-data, the sub-difference statistics of the first sub-data are determined, wherein the timing of the second sub-data is before the timing of the first sub-data.
[0009] The data fluctuation range corresponding to the time series data to be detected is determined based on the differential statistics, wherein the differential statistics include: the sub-differential statistics value of each sub-data;
[0010] Based on the differential statistics and the data fluctuation range, abnormal sub-data in the time series data to be detected are determined.
[0011] This embodiment provides a data anomaly detection device, including:
[0012] The data acquisition module is used to acquire the time series data to be detected, wherein the time series data to be detected includes multiple sub-data with different time sequences;
[0013] The data processing module is configured to perform the following processing for each of the sub-data: taking the sub-data as first sub-data, determining the sub-difference statistics of the first sub-data based on the first sub-data and multiple second sub-data corresponding to the first sub-data, wherein the time sequence of the second sub-data is earlier than the time sequence of the first sub-data; determining the data fluctuation range corresponding to the time sequence data to be detected based on the difference statistics, wherein the difference statistics include: the sub-difference statistics of each of the sub-data;
[0014] An anomaly detection module is used to determine abnormal sub-data in the time series data to be detected based on the differential statistics and the data fluctuation range.
[0015] This application provides an electronic device, including:
[0016] Memory is used to store executable instructions for a computer;
[0017] The processor, when executing computer-executable instructions stored in the memory, implements the data anomaly detection method provided in the embodiments of this application.
[0018] This application provides a computer-readable storage medium storing a computer program or computer-executable instructions, which, when executed by a processor, implements the data anomaly detection method provided in this application.
[0019] This application provides a computer program product, including a computer program or computer executable instructions. When the computer program or computer executable instructions are executed by a processor, they implement the data anomaly detection method provided in this application.
[0020] The embodiments of this application have the following beneficial effects:
[0021] By calculating the differential statistics of sub-data points in the time series data to be detected, corresponding differential statistics are obtained. Based on these differential statistics, the data fluctuation range corresponding to the time series data is determined. Anomalies are then identified based on these fluctuation ranges, enabling more accurate detection of anomalies. Compared to related technologies that determine differential statistics based on a single preceding data point, the differential statistics for each sub-data point are determined based on multiple preceding sub-data points, improving the accuracy of determining the differential statistics and consequently improving the accuracy of determining the fluctuation range. This allows for more accurate detection of anomalies based on the fluctuation range. Attached Figure Description
[0022] Figure 1 This is a schematic diagram illustrating the application mode of the data anomaly detection method provided in the embodiments of this application;
[0023] Figure 2 This is a schematic diagram of the server structure provided in an embodiment of this application;
[0024] Figure 3A This is a first flowchart illustrating the data anomaly detection method provided in this application embodiment;
[0025] Figure 3B This is a schematic diagram of the second process of the data anomaly detection method provided in the embodiments of this application;
[0026] Figure 3C This is a schematic diagram of the third process of the data anomaly detection method provided in the embodiments of this application;
[0027] Figure 3D This is a schematic diagram of the fourth process of the data anomaly detection method provided in the embodiments of this application;
[0028] Figure 3E This is a schematic diagram of the fifth process of the data anomaly detection method provided in the embodiments of this application;
[0029] Figure 3F This is a schematic diagram of the sixth process of the data anomaly detection method provided in the embodiments of this application;
[0030] Figure 4 This is a schematic diagram of the seventh process of the data anomaly detection method provided in the embodiments of this application;
[0031] Figure 5 This is a schematic diagram illustrating the principle of non-stationary sequence transformation provided in the embodiments of this application;
[0032] Figure 6 This is a comparison diagram of the non-stationary sequence and the converted stationary sequence provided in the embodiments of this application;
[0033] Figure 7 This is a schematic diagram of anomaly detection based on first-order difference calculation;
[0034] Figure 8 This is a first schematic diagram of anomaly detection based on improved differential provided in an embodiment of this application;
[0035] Figure 9 This is a second schematic diagram of anomaly detection based on improved differential provided in an embodiment of this application. Detailed Implementation
[0036] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0037] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0038] In the following description, the terms "first, second, third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0039] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0040] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in the embodiments of this application is for the purpose of describing the embodiments of this application only and is not intended to limit this application.
[0041] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.
[0042] 1) Time Series Data: refers to data recorded in chronological order, such as the daily click count of an advertisement, the processor usage of a laptop, and the indoor temperature at different times of the day.
[0043] 2) Data characteristics: The attributes or features used to distinguish different data items in a dataset, and the statistical measures of certain characteristics of indicator data, such as the mean, variance, first difference, slope, and divergence of indicator data.
[0044] 3) The Three-Standard-Degree Rule (3-sigma): Used in quality management and process control, this rule is used to determine whether data is normal and to detect outliers. It states that under a normal distribution, approximately 68.27% of data values will be within one standard deviation of the mean, 95.45% will be within two standard deviations, and approximately 99.73% will be within three standard deviations.
[0045] 4) First-order difference: is the difference between two consecutive adjacent terms in a discrete function.
[0046] 5) Stationary series: refers to a series whose statistical characteristics over time, including the mean and variance, do not change over time.
[0047] 6) Non-stationary series: refers to a series whose statistical characteristics over time, including the mean and variance, change over time.
[0048] 7) Granularity: refers to the frequency of data collection, i.e. the detection frequency. For example, daily granularity means that one data point represents the statistical value of the indicator in the past day, and it is detected once a day.
[0049] 8) Statistics: Values obtained by performing numerical calculations on the original indicator data.
[0050] 9) False Positive Rate (FPR): This refers to the ratio of normal data to abnormal data, that is, the frequency at which the system incorrectly identifies normal advertising data or data points as abnormal data or points.
[0051] 10) Advertising data: This is a textual record, a textual description and record of advertising data, which typically includes: ad display data, user click data, conversion data, advertising cost data, advertising channel data, etc., and is used as the basis for advertising data analysis and mining.
[0052] 11) Time series model: a statistical model used to describe the trend, seasonality and other characteristics of a set of random variables observed at consecutive time points, and used to predict or analyze time series data.
[0053] 12) Cross-Entropy Loss: A loss function that measures the difference between the predicted probability distribution and the true distribution.
[0054] Related technologies preprocess non-stationary sequences to convert them into stationary sequences before using an appropriate probability distribution function for anomaly detection. However, in scenarios where data is recovered after anomalies, these technologies are prone to generating multiple alarms, resulting in false alarms.
[0055] This application provides a data anomaly detection method, a data anomaly detection device, an electronic device, a computer-readable storage medium, and a computer program product, which can improve the accuracy of anomaly data detection. The following describes exemplary applications of the electronic device provided in this application. The device provided in this application can be implemented as various types of terminals such as laptops, tablets, desktop computers, set-top boxes, smartphones, smart speakers, smartwatches, smart TVs, and vehicle terminals, or it can be implemented as a server. The following will describe exemplary applications when the device is implemented as a terminal or server.
[0056] See Figure 1 , Figure 1 This is a schematic diagram illustrating the application mode of the data anomaly detection method provided in this application embodiment. To support a data anomaly detection application, an example is provided. Figure 1 The system involves server 200, network 300, and terminal device 400. Terminal device 400 is connected to server 200 and database 500 through network 300. Network 300 can be a wide area network, a local area network, or a combination of both.
[0057] In some embodiments, the user is a tester performing anomaly detection who can issue operation instructions to perform anomaly detection; the server 200 is a server that processes the data to be detected; the terminal device 400 is a computer or mobile phone that can be used by the tester; and the database 500 stores the data to be detected and differential statistics data.
[0058] For example, server 200 is used to collect statistics on the data to be tested, and terminal device 400 is used to display the data to be tested pushed to the user. Terminal device 400 sends a data anomaly detection request to server 200 via network 300. Server 200 performs differential statistics calculation and anomaly detection processing based on the data to be tested, and sends the anomaly detection result to terminal device 400. Terminal device 400 provides feedback on the anomaly detection result to the tester, which can be notified to the tester in the form of an alarm message. Testers can identify abnormal data and analyze and determine the cause of the anomaly based on the anomaly detection result.
[0059] In some embodiments, the data anomaly detection method of this application can also be applied in the following application scenarios:
[0060] (1) In the advertising network system, during the actual delivery process, ad clicks and conversions are generated. The ad data production process may have abnormal movement speed, which may lead to data abnormalities. For example, the production task may be abnormally interrupted or the data source of the production task may be abnormal, resulting in incomplete data. The data abnormality detection method provided in this application embodiment can detect the data quality and feed back the abnormal data to the data source for alarm prompts, reminding the data source to conduct self-checks.
[0061] (2) In the data center system, there are a large amount of key performance indicator data in the network of the data center, such as the retransmission rate of the transmission protocol of network traffic, the uplink bit rate and the download speed. The data anomaly detection method provided in this application embodiment can detect key performance indicator data, detect faults in time and issue alarms, and can also help operation and maintenance personnel to accurately locate fault data and support rapid fault recovery.
[0062] It is understood that, in the embodiments of this application, the collection and processing of relevant data (e.g., display data of advertising data, user click data, conversion data, etc.) should be strictly in accordance with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.
[0063] This application embodiment can be implemented using database technology. A database, simply put, can be viewed as an electronic filing cabinet storing electronic files, where users can perform operations such as adding, querying, updating, and deleting data. A "database" is a collection of data stored together in a certain way, capable of being shared by multiple users, having minimal redundancy, and being independent of application programs.
[0064] A Database Management System (DBMS) is a computer software system designed to manage databases, generally possessing basic functions such as storage, retrieval, security, and backup. DBMSs can be classified according to the database model they support, such as relational or XML (Extensible Markup Language); or according to the type of computer they support, such as server clusters or mobile devices; or according to the query language used, such as Structured Query Language (SQL) or XQuery; or according to performance priorities, such as maximum scale or maximum operating speed; or other classification methods. Regardless of the classification method used, some DBMSs can cross categories, for example, simultaneously supporting multiple query languages.
[0065] This application embodiment can also be implemented through artificial intelligence (AI). AI is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new type of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess perception, reasoning, and decision-making capabilities.
[0066] This application embodiment can also be implemented using cloud technology. Cloud technology is a general term for network technology, information technology, integration technology, management platform technology, and application technology based on cloud computing business models. It can form a resource pool, available on demand, offering flexibility and convenience. Cloud computing technology will become a crucial support. Backend services of technical network systems require substantial computing and storage resources, such as video websites, image websites, and many portal websites. With the rapid development and application of the internet industry, and driven by demands for search services, social networks, mobile commerce, and open collaboration, every item may eventually possess its own hash-coded identification mark, requiring transmission to a backend system for logical processing. Data at different levels will be processed separately, and various industry data will require robust system support, which can only be achieved through cloud computing.
[0067] In some embodiments, server 200 may be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Terminals and servers can be connected directly or indirectly via wired or wireless communication, which is not limited in this embodiment.
[0068] See Figure 2 , Figure 2 This is a schematic diagram of the server structure provided in an embodiment of this application. Figure 2The server 200 shown includes at least one processor 410, memory 450, and at least one network interface 420. The various components of server 200 are coupled together via a bus system 440. It is understood that the bus system 440 is used to implement communication between these components. In addition to a data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for clarity, ... Figure 2 The general labeled all buses as Bus System 440.
[0069] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.
[0070] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 450 may optionally include one or more storage devices physically located away from the processor 410.
[0071] The memory 450 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 450 described in this application embodiment is intended to include any suitable type of memory.
[0072] In some embodiments, memory 450 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.
[0073] Operating system 451 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic business functions and handling hardware-based tasks;
[0074] The network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420, exemplary network interfaces 420 including: Bluetooth, WiFi, and Universal Serial Bus (USB), etc.
[0075] In some embodiments, the apparatus provided in this application can be implemented in software. Figure 2 A data anomaly detection device 455 stored in memory 450 is shown. This device can be software in the form of programs and plug-ins, and includes the following software modules: a data acquisition module 4551, a data processing module 4552, and an anomaly detection module 4553. These modules are logically linked and can therefore be arbitrarily combined or further separated according to their implemented functions. Figure 2 For ease of explanation, all the above modules are shown at once, but this should not be interpreted as excluding the implementation of the data anomaly detection device 455 which may only include the data acquisition module 4551. The functions of each module will be explained below.
[0076] In some embodiments, the terminal or server can implement the data anomaly detection method provided in this application by running various computer-executable instructions or computer programs. For example, computer-executable instructions can be microprogram-level commands, machine instructions, or software instructions. Computer programs can be native programs or software modules in an operating system; they can be native applications (APPs), i.e., programs that need to be installed in the operating system to run, such as live streaming APPs or instant messaging APPs; or they can be applets that can be embedded in any APP, i.e., programs that only need to be downloaded to a browser environment to run. In summary, the aforementioned computer-executable instructions can be any form of instruction, and the aforementioned computer programs can be any form of application, module, or plugin.
[0077] The data anomaly detection method provided in this application will be described in conjunction with exemplary applications and implementations of the electronic devices provided in the embodiments of this application.
[0078] The data anomaly detection method provided in the embodiments of this application will be described below. As mentioned above, the electronic device implementing the data anomaly detection method in the embodiments of this application can be a terminal or a server, or a combination of both. Therefore, the executing entity of each step will not be described again below.
[0079] See Figure 3A , Figure 3A This is a first flowchart of the data anomaly detection method provided in this application embodiment, which will be described in conjunction with the steps shown in Figure 3.
[0080] In step 301, the timing data to be detected is obtained.
[0081] Here, the time series data to be detected includes multiple sub-data with different time series.
[0082] For example, to perform anomaly detection on data, it is necessary to obtain time-series data over a period of time from the object to be detected, and use this as the time-series data to be detected.
[0083] Time-series data is data recorded in chronological order. Taking an advertising data quality inspection scenario as an example, in the process of acquiring the time-series data to be inspected in an advertising data quality inspection scenario, the quantity of statistical data is divided by time. The advertising data is divided according to daily granularity to obtain the number of statistical data per day. Granularity refers to the frequency of data collection. Daily granularity represents the statistical values within the past day. Inspection is performed once a day, and the number of advertising data collected each day is recorded as num. The data indicators and formats after being divided by time are stored in the time-series data (time-series data to be inspected).
[0084] Historical statistical metrics are retrieved from time-series data using time as an index. Taking the detection of advertising data metrics within 30 days as an example, where i is the number of days corresponding to the time index, the historical statistical metrics of advertising data within 30 days retrieved from the storage system are: [num i-29 ,num i-28 , ...,num i The time series data to be tested consists of multiple sub-data points divided into days. Each sub-data point corresponds to a different time series, representing advertising data metrics for different number of days.
[0085] In some embodiments, see Figure 3B , Figure 3B This is a schematic diagram of the second process of the data anomaly detection method provided in this application embodiment; to illustrate the acquisition process of the time series data to be detected in more detail, during the execution Figure 3A Before step 301, execute Figure 3B Steps 3041 to 3042 are explained in detail below.
[0086] In step 3041, sample time series data is obtained.
[0087] Here, the sample time series data includes: multiple sample data with different time series of sample objects, and the actual types corresponding to the multiple sample data, including abnormal sample data and normal sample data.
[0088] For example, the sample data is multi-category data, and the sample object can be an advertisement. Each sample data has a different time series, and the time series data includes normal sample data and abnormal sample data. For instance, in an advertisement data quality inspection scenario, the advertisement data includes multiple sample categories such as the number of advertisements and the number of clicks. The data for each sample category of advertisements contains data corresponding to different time series. When an anomaly occurs in a certain time series, abnormal data is generated accordingly. In this case, the actual sample type obtained includes not only normal sample data but also abnormal sample data.
[0089] In step 3042, a time series model is trained based on the sample time series data to obtain the trained time series model.
[0090] Here, the trained time series model is used to identify anomalous subdata in the time series data to be detected.
[0091] For example, the processing of time series data is based on time series models. A time series dataset with a time order is constructed to train the time series model. The time series model is used to analyze and predict series data. The model parameters are constructed by training on the time series dataset. The trained model can detect abnormal data in the time series data.
[0092] In some embodiments, see Figure 3C , Figure 3C This is a schematic diagram of the third process of the data anomaly detection method provided in the embodiments of this application; Figure 3B Step 3042 in the process can be achieved through Figure 3C Steps 30421 to 30424 are implemented, and the details are explained below.
[0093] In step 30421, a time series model is called based on the sample time series data to perform feature extraction processing, thereby obtaining the sequence features of the sample time series data.
[0094] Here, sequence features are used to characterize difference statistics and data fluctuation ranges.
[0095] For example, a time series model is used to fit the sample time series data. After training the sample time series dataset, the parameters of the time series model are determined. After the time series model is completed, the parameters of the model and the residual sequence are used to extract features from the sample time series data, such as the difference statistics of the time series data, the mean and standard deviation of the difference statistics.
[0096] In step 30422, a time series model is invoked based on the sequence features to perform prediction processing, thereby obtaining the prediction type for each sample data.
[0097] For example, prediction processing is performed based on extracted sequence features. The predicted type of sample data can be normal sample data or abnormal sample data. The predicted type obtained by the time series model may differ from the actual type of the sample data, that is, the time series model has prediction errors. The accuracy of the time series model in determining the actual type of the sample can be improved through training and parameter tuning.
[0098] In step 30423, the cross-entropy loss of the time series model is determined based on the predicted type and the actual type of each sample data.
[0099] For example, the loss function used in the training process of a time series model can be cross-entropy loss. The purpose of cross-entropy loss is to measure the difference between the model's predicted distribution and the true distribution. The cross-entropy loss determined based on the predicted type and the actual type of each sample data can be used to optimize the parameters of the time series model.
[0100] In step 30424, the time series model is backpropagated based on cross-entropy loss to obtain the trained time series model.
[0101] For example, the deterministic cross-entropy loss characterizes the difference between the predicted type and the actual type. The gradient of the time series model is calculated based on backpropagation. The parameters of the time series model are updated according to the gradient information, and the value of the loss function is gradually reduced to optimize the trained time series model.
[0102] In this embodiment, a time series model is trained using sample time series data, and the parameters of the time series model are optimized based on the cross-entropy loss function to improve the accuracy and stability of the time series model in predicting anomalous sub-data.
[0103] In some embodiments, see Figure 3D , Figure 3D This is a schematic diagram of the fourth process of the data anomaly detection method provided in this application embodiment; step 301 can be implemented through steps 3011 to 3013, as described in detail below.
[0104] In step 3011, statistical data of the object to be detected within the pre-configured time period is obtained.
[0105] For example, when performing anomaly detection on data, it is necessary to pre-configure the duration of the time-series data to be detected. Based on the pre-configured duration, statistical data within the corresponding duration is obtained from the time-series data. The pre-configured duration is the time from the period when anomaly detection is required to the current time, which is set by the tester in advance. The pre-configured duration can be set according to the needs of the tester in the actual application scenario. For example, if the tester needs to perform anomaly detection on the data of an advertising platform within one month, then the pre-configured duration is 30 days (one month).
[0106] For example, in the above scenario of advertising data quality detection, it is necessary to detect anomalies in the data within the last 30 days. The pre-configured duration is 30 days, the object to be detected is the advertising data metrics, and the statistical data obtained from the storage system is the statistical data of the corresponding advertising data metrics within the last 30 days.
[0107] In step 3012, the statistical data is divided according to the time index of the statistical data to obtain multiple sub-data with different sampling times.
[0108] For example, a time index refers to an index used to represent a time interval or point in time in time series data. After obtaining statistical data of a pre-configured duration from the time series data, the statistical data is divided according to the time interval to obtain multiple sub-data corresponding to different sampling times.
[0109] In step 3013, multiple sub-data are sorted according to the chronological order of each time index to obtain the time series data to be detected.
[0110] For example, the time index has a chronological order. The multiple sub-data after division are sorted according to the corresponding time index order. The sorted sub-data are used to obtain a data sequence with a time sequence, which is the time series data to be detected.
[0111] Continue to refer to Figure 3A In step 302, the following processing is performed for each sub-data: the sub-data is taken as the first sub-data, and the sub-difference statistics of the first sub-data are determined based on the first sub-data and the multiple second sub-data corresponding to the first sub-data.
[0112] Here, the first sub-data is any single sub-data, and the timing of the second sub-data precedes that of the first sub-data.
[0113] For example, multiple sub-data points, divided and sorted according to time index, form a time series sequence. Based on the multiple sub-data points preceding each sub-sequence, the sub-difference statistics corresponding to each sub-data point are calculated. Any one sub-data point is taken as the first sub-data point. For example, if the first sub-data point is the sub-data point of time series i, the sub-data points of time series i-1, i-2, and i-3 preceding time series i are taken as the second sub-data points, representing the sub-data points one day, two days, and three days before time series i. Based on the sub-data points of time series i-1, i-2, and i-3, the sub-difference statistics of the first sub-data point are determined.
[0114] In some embodiments, see Figure 3E , Figure 3E This is a schematic diagram of the fifth process of the data anomaly detection method provided in this application embodiment; step 302 can be implemented by the following steps 3021 to 3023, which are described in detail below.
[0115] In step 3021, a first difference between the first sub-data and multiple second sub-data is obtained.
[0116] For example, the first sub-data is any single sub-data, and the first sub-data is the sub-data of time series i, denoted as x. i The second sub-data is the sub-data of time series i-1, i-2, and i-3. i-1 x represents the first second sub-data of the time series preceding time series i. i-2 x represents the second sub-data point of the two time series preceding time series i. i-3 This represents the second sub-data point of the third time series before time series i, and obtains the difference x between the first sub-data point and the sub-data point of time series i-1. i -x i-1 The difference x between the i-2 time series sub-data i -x i-2 The difference x between the i-3 time series sub-data i -x i-3 , which serves as the first difference between the first sub-data and multiple second sub-data.
[0117] In step 3022, the product between each first difference is obtained.
[0118] For example, the first difference between the first sub-data and multiple second sub-data is x. i -x i-1 x i -x i-2 and x i -x i-3 Multiply the multiple first differences to obtain the product (x) between each first difference. i -x i-1 (x) i -x i-2 (x)i -x i-3 ).
[0119] In step 3023, the Nth root of the product is used as the sub-difference statistic of the first sub-data.
[0120] Here, N is the number of multiple second sub-data.
[0121] For example, multiply the products (x) between each first difference. i -x i-1 (x) i -x i-2 (x) i -x i-3 The Nth root of ) is used as the sub-difference statistic δ of the first sub-data xi.
[0122] In some embodiments, step 3023 can be characterized by the following formula (1):
[0123]
[0124] Where, x i This represents the first sub-data point in time series i, where i is the detection time point, and x... i-1 x represents the first second sub-data of the time series preceding time series i. i-2 x represents the second sub-data point of the two time series preceding time series i. i-3 δ represents the third second sub-data point of the first three time series of time series i, and δ is the sub-difference statistic of the first sub-data point of time series i. The mean and standard deviation of the corresponding sub-difference statistic are calculated based on the sub-difference statistic δ.
[0125] Formula (1) is applicable to cases where the number of days in the time series data to be detected exceeds three days. As can be seen from Formula (1), the sub-difference statistics of the first sub-data are calculated based on the first sub-data of time series i and multiple second sub-data earlier than time series i. In this embodiment, three days is taken as an example. If the amount of data in the time series cannot meet the requirement of more than three days, this formula is not applicable to the calculation of the sub-difference statistics.
[0126] Continuing with the example of advertising data quality detection, we will detect advertising data from the past 30 days. The data will be divided into segments based on a daily granularity, with the time index of the segmented data being the day. The sub-data segments will correspond to advertising data from different days. By arranging the sub-data segments according to the page order of the sampled days, we can obtain a time-series sequence of advertising data. For example, to detect the number of advertisements in the last 30 days, since the sub-difference statistic requires the calculation of sub-data based on the previous three time series of sub-data, it is necessary to retrieve the number of advertisements for the last 32 days from the storage system. Otherwise, it is impossible to calculate the sub-data for the 30 days from the current time. After dividing and sorting by daily granularity, the time series data to be detected for the number of advertisements in the last 30 days (one month) is: [25196153, 25279188, 25328248, 25369165, 25409647, 25448752, 25493504, 25527187, 255 [78003, 25626621, 25675075, 25712293, 25719726, 25794092, 25838934, 25874082, 25912147, 25945568, 25977842, 26015397, 26051339, 26081039, 26114203, 26150438, 26185062, 26220729, 26269075, 26280907, 26357199, 26366491, 9854014], where each data point represents the number of advertisements pushed to the user within one day.
[0127] For example, in the scenario of detecting the number of advertisements in the last 30 days, it was found that the number of data entries decreased significantly on the 30th day. According to the formula (1) in step 302, the sub-difference statistic value corresponding to the 30th day was calculated to be -16480807.017739916.
[0128] In this embodiment, the differential statistics corresponding to the first sub-data are calculated using the first sub-data and multiple second sub-data that are earlier in time than the first sub-data. The calculated differential statistics have time-series characteristics and are more accurate.
[0129] Continue to refer to Figure 3A In step 303, the data fluctuation range corresponding to the time series data to be detected is determined based on the differential statistics.
[0130] Here, difference statistics include: sub-difference statistics of sub-data.
[0131] For example, to determine whether the data to be detected is abnormal, it is necessary to first determine the data fluctuation range of the time series data to be detected. Based on the difference statistics, the upper limit and lower limit of the data fluctuation range are calculated. The data between the upper limit and lower limit is considered normal data fluctuation and is used as the data fluctuation range.
[0132] In some embodiments, see Figure 3F , Figure 3F This is a schematic diagram of the sixth process of the data anomaly detection method provided in this application embodiment; step 303 can be implemented through steps 3031 to 3033, as described in detail below.
[0133] In step 3031, the mean and standard deviation corresponding to the difference statistic are determined.
[0134] In some embodiments, step 3031 can be implemented by the following method: determining a first sum between each of the sub-difference statistics, dividing the first sum by the total number of the sub-data to obtain the mean corresponding to the difference statistic; determining the square of the difference between each of the sub-difference statistics and the mean, and determining a second sum between each of the squares; and taking the square root of the ratio between the second sum and the total number of the sub-data as the standard deviation corresponding to the difference statistic.
[0135] The mean of the difference statistic can be represented by the following formula (2):
[0136] mean=∑y i / N (2)
[0137] Among them, y i This represents the sub-difference statistics of the sub-data corresponding to time series i, where N is the number of statistics in the time series data to be detected, and ∑y i The first summation among the sub-difference statistics represents the accumulation of the sub-difference statistics of all time series corresponding to time series i within the statistical time period. The mean is the ratio between the accumulated value and the total number of sub-data N.
[0138] The standard deviation (std) of the difference statistic can be represented by the following formula (3):
[0139]
[0140] Specifically, the difference between the sub-difference statistic of each sub-data point and the mean is calculated, and the squares of these differences are summed to form the second sum ∑(y) between each set of squares. i -mean) 2 The square root of the ratio between the sum of the sub-data and the total number of sub-data N is used as the standard deviation of the difference statistic.
[0141] In step 3032, the upper limit and lower limit of the interval are determined based on the mean and standard deviation, respectively.
[0142] In some embodiments, step 3032 can be implemented by: taking the third sum between the mean and the standard deviation of the pre-configured multiple as the upper limit of the interval, wherein the pre-configured multiple is the number of second sub-data used to determine the sub-difference statistics of the first sub-data; and taking the difference between the mean and the standard deviation of the pre-configured multiple as the lower limit of the interval.
[0143] For example, after calculating the mean and standard deviation (std) of the difference statistic, the upper and lower limits of the interval are calculated using the n-sigma algorithm. Here, n is a pre-configured multiple, which is the number of second sub-data points used to determine the sub-difference statistic of the first sub-data point. The upper limit of the interval is the sum of the mean and the standard deviation of the pre-configured multiple, which is the difference between the mean and the standard deviation of the pre-configured multiple, which is the lower limit of the interval, which is the difference between the mean and the standard deviation of the pre-configured multiple ... standard deviation of the pre-configured multiple, which is the difference between the mean and the standard deviation of the standard deviation of the standard deviation of the standard deviation of the standard deviation of the standard deviation of the standard deviation of the standard deviation of the standard deviation of the standard deviation of the standard deviation of the standard deviation of the standard deviation of the standard deviation of the standard deviation of the standard deviation of the standard deviation of the standard deviation of the standard deviation of the standard deviation of the standard deviation of the standard deviation of the standard deviation of the standard deviation of the standard deviation of the standard deviation of the standard deviation of the standard deviation of the standard deviation of the standard deviation of the standard
[0144] In step 3033, the range between the upper limit and the lower limit of the interval is taken as the data fluctuation range.
[0145] For example, the range between the upper and lower limits of the interval is taken as the data fluctuation range. The data fluctuation range is used as the criterion for data anomaly detection. That is, the data fluctuation range is: [mean + n*std, mean - n*std], where n*std is the standard deviation of the pre-configured multiple, and mean is the mean.
[0146] For example, when the pre-configuration multiplier is 5, that is, n is 5, in the scenario of detecting the number of advertisements in the last 30 days, the reasonable fluctuation range is determined based on the determined mean and standard deviation as: [-42797.26067859252, 123510.22619583391].
[0147] In this embodiment, a reasonable fluctuation range is determined by summing and differing the mean of the difference statistic with the standard deviation of the pre-configured multiple, which can improve the accuracy of identifying abnormal sub-data and increase the accuracy of abnormal data detection.
[0148] Continue to refer to Figure 3A In step 304, based on the differential statistics and the data fluctuation range, abnormal sub-data in the time series data to be detected is determined.
[0149] For example, after determining the data fluctuation range, the difference statistics corresponding to the sub-data in the time series data to be tested are compared with the data fluctuation range. Based on whether the difference statistics corresponding to the sub-data are within the data fluctuation range, it is determined whether the sub-data in the time series data to be tested is abnormal.
[0150] In some embodiments, step 304 can be implemented by the following method: in response to any sub-difference statistic in the difference statistics being outside the data fluctuation range, determining that the sub-data corresponding to the sub-difference statistic outside the data fluctuation range is abnormal.
[0151] For example, in the scenario described above, where the number of advertisements within the last 30 days is being analyzed, the difference statistic for the sub-data 9854014 on day 30 is -16480807.017739916, which is below the lower limit of the reasonable fluctuation range. Therefore, this data is considered abnormal. If the sub-data for day 31 is 26366492, the calculated difference statistic for day 31 is 5353.72504490245, which is within the reasonable fluctuation range. Therefore, the data for day 31 is considered normal.
[0152] In this embodiment of the application, by determining a reasonable data fluctuation range, while maintaining the detected abnormal data points, it is possible to accurately detect the data after it has returned to normal, thereby reducing the number of false alarms after data recovery.
[0153] In some embodiments, step 304 can also be implemented by the following method: based on the difference statistics and the data fluctuation range, calling the trained time series model for prediction processing to obtain the abnormal sub-data in the time series data to be detected.
[0154] For example, the differential statistics and data fluctuation range of the data to be detected are input into the trained time series model for prediction processing to obtain sub-data of type abnormal sub-data.
[0155] See Figure 9 , Figure 9 This is a second schematic diagram of anomaly detection based on improved differential provided in the embodiments of this application, which will be described in detail below. Figure 9 The improved differential anomaly data diagram uses the number of days as the horizontal axis and the number of data entries as the vertical axis. The vertical axis 1E+7 indicates that the unit of the vertical axis data value is 10 to the power of 7. The broken line 901 shows the trend of the number of data entries in the original time series with the number of days. The broken line 902 shows the trend of the differential statistics data after the original time series is transformed by the improved differential algorithm with the number of days. The dashed line 903 shows the lower limit of the reasonable fluctuation range, and the dashed line 904 shows the upper limit of the reasonable fluctuation range. The dashed line 904 is above the dashed line 903.
[0156] For example, Figure 9 The raw data in the table represents the number of advertisements, with each data point representing the amount of data per day. The reasonable fluctuation range determined above is: [-42797.26067859252, 123510.22619583391]. Figure 9 The data shows a significant decrease on day 30, with the difference statistic for day 30 falling below the lower limit of the interval. In this case, the first sub-data point corresponding to day 30 is an anomaly, and an alert should be issued to the data source in response to the anomaly.
[0157] In some embodiments, after determining the anomalous sub-data in the time series data to be detected based on the differential statistics and the data fluctuation range in step 304, the method further includes: issuing an anomalous data alarm in response to the anomalous sub-data in the time series data to be detected meeting a pre-configured condition, wherein the pre-configured condition includes at least one of the following:
[0158] Condition 1: The amount of abnormal sub-data reaches the pre-configured quantity.
[0159] For example, in the scenario of advertising data quality detection, a value is pre-configured for the number of ads based on the past advertising data quality. The pre-configured value can be the value corresponding to the software crash on the user side. When the sub-data reaches the value pre-set based on historical experience, the sub-data is abnormal data. The time series corresponding to the sub-data may have experienced a software crash. It is necessary to issue an alert to the data source to indicate that there is an anomaly.
[0160] Condition 2: The difference between the abnormal sub-data and the upper or lower limit of the data fluctuation range reaches a threshold.
[0161] For example, after identifying abnormal sub-data, an abnormal data alarm is generated for the abnormal sub-data that meets the pre-configured conditions.
[0162] For example, if the data fluctuation trend of the data source remains stable for a long time, when the difference between the upper and lower limits of the pre-configured data fluctuation range reaches a threshold, it indicates that there is unstable sub-data. When the sub-data reaches the threshold, it is confirmed that the data is abnormal data, and it is necessary to issue an alert to the data source to indicate that there is unstable abnormal sub-data.
[0163] In this embodiment, a sub-difference statistical value is calculated based on the first sub-data and multiple second sub-data whose time sequence is earlier than the first sub-data to determine the data fluctuation range. Abnormal sub-data is determined based on the difference statistics of the time sequence data to be detected and the data fluctuation range. During the anomaly detection process, the identification of anomaly points is retained, which improves the detection accuracy of the recovered data, avoids false alarms caused during data recovery, and reduces the number of false alarms.
[0164] The following will describe an exemplary application of the data anomaly detection method of this application in a real-world application scenario.
[0165] In large-scale network environments, anomaly detection has a wide range of applications. Related technologies typically involve studying the distribution characteristics of the data and creating a probability distribution function that matches the data. If the observed data does not fit the model well, it is considered an outlier. When anomalies occur, such as network latency or server failures, fluctuations in system business-related data can occur. For example, if software pushing advertising data to users suddenly crashes, causing many users to be unable to use the software for an extended period and thus unable to receive the pushed advertising data, the advertising business data obtained by the data source for that day will drop sharply. In this case, the observed data does not fit well with the probability distribution function matching the data, and is therefore an outlier. For non-stationary sequences, due to the instability of variance and mean, preprocessing is generally used to convert the non-stationary sequence into a stationary sequence before employing an appropriate probability distribution function for data anomaly detection.
[0166] In related technologies, first-order difference algorithms are usually used to convert non-stationary sequences to stationary sequences. However, detection algorithms based on first-order difference statistics are prone to misjudging normal data as abnormal data in scenarios where data is recovering from anomalies, resulting in multiple alarm messages and false alarms.
[0167] This application embodiment calculates differential statistics on sub-data in the time series data to be detected, and obtains the corresponding differential statistics based on the improved differential statistics method. It determines the reasonable data fluctuation range corresponding to the time series data, which can more accurately detect abnormal data in the fluctuation range, reduce the false alarm rate, and improve the accuracy of data anomaly detection.
[0168] The following explanation is in conjunction with the accompanying drawings. Figure 4 , Figure 4 This is a schematic diagram of the seventh process of the data anomaly detection method provided in this application embodiment. The executing entity can be a terminal device, a server, or a combination of both. This application embodiment takes a server as the executing entity as an example, and will combine... Figure 4 The steps shown are explained in detail.
[0169] In step 401, the data is sorted according to time order to obtain the time sequence data to be detected.
[0170] For example, in real-world applications, when performing anomaly detection, time-series data over a past period is obtained based on the object to be detected. Time-series data refers to data recorded in chronological order.
[0171] For example, based on daily granular statistical data and divided by time, the daily statistical data is used as an indicator, denoted as `num`. This data, along with the indicator format, is stored in the storage system. Before data anomaly detection, historical statistical indicators are retrieved from the storage system using time as the index. Granularity refers to the frequency of data collection; daily granularity means that one data point represents the statistical value of that indicator within the past day, and detection is performed once per day. Taking a 30-day data indicator as an example, where `i` represents the number of days corresponding to the time index, the 30-day time-series data to be detected retrieved from the storage system is: [num...] i-29 ,num i-28 , ...,num i ].
[0172] To more clearly illustrate the data anomaly detection method provided in this application embodiment, the following uses an application in the quality detection scenario of advertising data as an example to describe the anomaly detection process in detail.
[0173] The advertising data system refers to the actual network system in use, that is, an advertising network system that has been built and is available to users. It typically includes an advertising delivery platform and an advertising data collection system, used to achieve advertising delivery, monitoring, and optimization. The generated advertising data refers to the data produced during the actual delivery process; specifically, it is a textual record, a written description and record of advertising data, such as ad display data, user click data, conversion data, and advertising cost data. This information helps advertisers and advertising platforms better understand the advertising performance and effectiveness, serving as the foundation for advertising data analysis and mining, and supporting subsequent data analysis and applications.
[0174] Anomalies in the advertising data production process can lead to data corruption, such as interruptions in production tasks or abnormalities in the data source, resulting in incomplete data. Directly feeding this abnormal data into the recommendation model for training will negatively impact the training results; therefore, data quality control is necessary. Data quality control refers to the process of determining whether the data is normal and meets expectations for validity. This is primarily achieved by analyzing advertising data metrics, including the number of data entries, data size, and numerical indicators within the data, to determine if they fall within a reasonable fluctuation range.
[0175] Because advertising data and content are constantly updated every day, the statistical advertising data indicators obtained are non-stationary series. Non-stationary series refer to series whose statistical characteristics over time, including the mean and variance, change over time.
[0176] See Figure 5 , Figure 5 This is a schematic diagram illustrating the principle of non-stationary sequence transformation provided in the embodiments of this application, which will be explained in detail below.
[0177] In step 501, the time series is obtained.
[0178] For example, based on the object to be detected, obtain the corresponding time series data over a past period.
[0179] In step 502, it is determined whether the sequence is stationary.
[0180] For example, a stationary sequence is a sequence whose statistical characteristics over time, including the mean and variance, do not change with time. If step 502 determines that the obtained time series is not a stationary sequence, the difference calculation in step 503 is performed. Based on the difference algorithm, the non-stationary sequence can be converted into a stationary sequence.
[0181] See Figure 6 , Figure 6 This is a comparison diagram of the non-stationary sequence and the converted stationary sequence provided in the embodiments of this application, which will be explained in detail below. Figure 6 With time as the horizontal axis and data volume as the vertical axis, the solid line 601 represents the trend of data volume change in the original non-stationary sequence, and the dashed line 602 represents the trend of data volume change after transformation. It can be seen that before the transformation, the sequence data from the first day onwards showed a continuous upward trend, while the transformed sequence data tended to be stable after the first day, indicating that the data sequence became stable after transformation.
[0182] See also Figure 5 The transformed stationary sequence is then processed in step 504 to improve the calculation of difference statistics.
[0183] For example, after obtaining the transformed stationary sequence through first-order differencing, the difference statistic is calculated on the stationary data. Here, the first-order difference value is taken as the difference statistic, and the mean and standard deviation of the difference statistic are calculated. Based on the detection algorithm, the reasonable fluctuation range of the difference statistic is obtained. See also Figure 7 , Figure 7 This is a schematic diagram of anomaly detection based on first-order difference calculation, which is explained in detail below. Figure 7 The time series comparison chart uses time as the horizontal axis and data volume as the vertical axis. Line 701 represents the data volume change trend of the original time series over time, line 702 represents the data volume change trend of the sequence after the original time series is transformed by the first-order difference algorithm, dashed line 703 represents the lower limit of the reasonable fluctuation range, and dashed line 704 represents the upper limit of the reasonable fluctuation range.
[0184] For example, at time t4, the curve data volume decreases due to a fault, and at time t5, the curve data volume returns to normal. Line 701 represents the trend of data volume change over time in the original time series, indicating that the original time series is non-stationary. A one-stage differencing transformation is performed on this set of time series data, resulting in line 702. Using the first-order difference value as the difference statistic δ, the mean and standard deviation (std) of the difference statistic δ are calculated. Based on the three-standard-deviation law detection algorithm, the reasonable fluctuation range of the statistic δ is calculated as [mean + 3*std, mean – 3*std]. Dashed line 703 represents the lower limit of the reasonable fluctuation range, i.e., mean – 3*std, and dashed line 704 represents the upper limit of the reasonable fluctuation range, i.e., mean + 3*std. The three-standard-deviation law describes that under a normal distribution, approximately 99.73% of the data values will fall within three standard deviations.
[0185] If a data anomaly occurs at time t4, and the difference statistic δ is outside the reasonable fluctuation range at both times t4 and t5, meaning the data will be judged as abnormal at both times t4 and t5, generating two alarms. However, the data has actually returned to normal at time t5, making the alarm at time t5 a false alarm. Therefore, detection algorithms based on first-order difference statistics are prone to generating multiple alarms and causing false alarms in scenarios where data recovers after anomalies.
[0186] See Figure 8 , Figure 8 This is a first schematic diagram of anomaly detection based on improved differential provided in an embodiment of this application, which will be described in detail below. Figure 8 The time series comparison chart uses time as the horizontal axis and data volume as the vertical axis. Line 801 represents the trend of data volume change of the original time series over time, line 802 represents the trend of difference statistics after the original time series is transformed by the improved difference algorithm over time, dashed line 803 represents the lower limit of the reasonable fluctuation range, and dashed line 804 represents the upper limit of the reasonable fluctuation range.
[0187] If a data anomaly occurs at time t4, by improving the calculation of the differential statistics, only the data at time t4 is considered abnormal. The data at time t5 has recovered and will not be considered abnormal, remaining within a reasonable fluctuation range. By improving the differential algorithm for statistical calculation, the data is no longer considered abnormal after recovery, reducing the number of error alarms.
[0188] See also Figure 5 Based on the calculated improved difference statistics, step 505, anomaly detection, is performed.
[0189] For example, after calculating the statistic using the improved difference algorithm, the mean and standard deviation (std) of the difference statistic δ are calculated. Based on the three-standard-deviation law detection algorithm, the upper and lower limits of the reasonable fluctuation range of the statistic δ are calculated. Data within the reasonable fluctuation range is judged as normal data, and data outside the reasonable fluctuation range is judged as abnormal data.
[0190] See also Figure 4 In step 402, the differential statistics of the time series data to be detected are calculated based on the current sub-data and the sub-data preceding the current sub-data in time series.
[0191] For example, the difference statistic is calculated for the time series sequence to be detected obtained in step 401. The formula (1) for calculating the statistic is as follows:
[0192]
[0193] Where, x i The time series data representing time series i is the detection time point, and δ is the corresponding difference statistic. Then, the mean and standard deviation std are calculated based on the difference statistic δ. Formula (1) is applicable when the number of days of the time series data to be detected exceeds three days. As can be seen from Formula (1), the improved difference statistic is calculated based on time series i and data earlier than time series i. If the time series data cannot meet the requirements, this formula is not applicable.
[0194] For example, the mean is the ratio of the sum of the difference statistics to the number of days of testing, and the standard deviation is calculated using the following formula (3):
[0195]
[0196] Among them, y i This represents the value of the difference statistic, where mean is the average value and N is the number of days monitored.
[0197] In step 403, a reasonable fluctuation range for the statistic to be detected is determined.
[0198] For example, based on the mean and standard deviation (std) of the difference statistic, the reasonable fluctuation range of the statistic to be detected can be determined using the n-sigma algorithm. The upper limit of the reasonable fluctuation range is mean + n * std, and the lower limit of the reasonable fluctuation range is mean - n * std.
[0199] See Figure 9 , Figure 9 This is a second schematic diagram of anomaly detection based on improved differential provided in the embodiments of this application, which will be described in detail below. Figure 9The improved differential anomaly data diagram uses the number of days as the horizontal axis and the number of data entries as the vertical axis. The unit of the data value on the vertical axis is 10 to the power of 7 (1E+7). Line 901 represents the trend of the number of data entries in the original time series with the number of days. Line 902 represents the trend of the differential statistics data after the original time series is transformed by the improved differential algorithm with the number of days. Dashed line 903 represents the lower limit of the reasonable fluctuation range, and dashed line 904 represents the upper limit of the reasonable fluctuation range.
[0200] For example, Figure 9 The raw data in the dataset consists of the number of ad displays, with each data point representing the amount of data per day, i.e., the number of ad displays. Data from 30 days is used as the calculation data for a reasonable fluctuation range. The number of ad displays can be the total number of various types of ads shown and pushed to users multiple times within a day on a specific application platform. When the application's server crashes, or when users are not interested in the ads and have blocked them, causing a sudden drop in the number of displayed ads, abnormal data will appear.
[0201] The time series sequence retrieved over the past 30 days is: [25196153, 25279188, 25328248, 25369165, 25409647, 25448752, 25493504, 25527187, 25578003, 25626621, 25675075, 25712293, 25719726, 25794092, 25838934, 2587408]. 2, 25912147, 25945568, 25977842, 26015397, 26051339, 26081039, 26114203, 26150438, 26185062, 26220729, 26269075, 26280907, 26357199, 26366491, 9854014], where each data point represents the number of advertisements pushed to the user within one day.
[0202] When the number of data entries decreases significantly on day 30, the difference statistic for day 30 is calculated as -16480807.017739916 based on the improved difference method. Using the n-sigma anomaly detection algorithm, where n is 5, the reasonable fluctuation range is determined to be: [-42797.26067859252, 123510.22619583391].
[0203] See also Figure 4 In step 404, abnormal data in the time series to be detected are determined based on the reasonable fluctuation range of the statistic to be detected.
[0204] For example, after obtaining the reasonable fluctuation range of the statistic to be detected, the data in the time series to be detected is judged. When the difference statistic corresponding to the data is within the reasonable fluctuation range, the data is judged to be normal data; when the difference statistic corresponding to the data is outside the reasonable fluctuation range, the data is judged to be abnormal data. At this time, an alarm is issued for the detected abnormal data.
[0205] See also Figure 9 The determined reasonable fluctuation range is [-42797.26067859252, 123510.22619583391]. The difference statistic for day 30 is -16480807.017739916, which is lower than the lower limit of the determined reasonable fluctuation range and is therefore judged as abnormal data. If the data for day 31 is 26366492, the calculated difference statistic for day 31 is 5353.72504490245, which is within the determined reasonable fluctuation range and is therefore judged as normal data.
[0206] For example, when abnormal data is detected, an alarm message is generated and fed back to the data source to remind the data source to make modifications or conduct further checks. The data source is usually an advertising platform that provides the time-series data to be checked. The alarm message can help the technical staff on the advertising platform side to discover the anomaly and analyze the cause of the anomaly.
[0207] The audio detection method provided in this application has the following beneficial effects:
[0208] By calculating the improved differential statistics of the time series to be detected, the reasonable fluctuation range of the data is determined, and abnormal data outside the reasonable fluctuation range is more accurately identified. While retaining the identification of abnormal points during the anomaly detection process, false detections caused by data recovery are avoided, thereby reducing the false alarm rate of generated alarm information and improving the accuracy of anomaly data detection.
[0209] The following description continues to illustrate the exemplary structure of the data anomaly detection device 455 provided in the embodiments of this application as a software module. In some embodiments, such as... Figure 2As shown, the software modules in the data anomaly detection device 455 stored in the memory 450 may include: a data acquisition module 4551, used to acquire time-series data to be detected, wherein the time-series data to be detected includes multiple sub-data with different time sequences; a data processing module 4552, used to perform the following processing for each sub-data: taking the sub-data as first sub-data, determining the sub-difference statistical value of the first sub-data based on the first sub-data and multiple second sub-data corresponding to the first sub-data, wherein the time sequence of the second sub-data is earlier than the time sequence of the first sub-data; determining the data fluctuation range corresponding to the time-series data to be detected based on the difference statistics, wherein the difference statistics include: the sub-difference statistical value of each sub-data; and an anomaly detection module 4553, used to determine the abnormal sub-data in the time-series data to be detected based on the difference statistics and the data fluctuation range.
[0210] In some embodiments, the data acquisition module 4551 is further configured to acquire statistical data of the object to be detected within a pre-configured time period; divide the statistical data according to the time index of the statistical data to obtain a plurality of sub-data with different sampling times; and sort the plurality of sub-data according to the order of each time index to obtain the time-series data to be detected.
[0211] In some embodiments, the data processing module 4552 is further configured to obtain a first difference between the first sub-data and the plurality of second sub-data; obtain a product between each of the first differences; and use the Nth root of the product as a sub-difference statistical value of the first sub-data, wherein N is the number of the plurality of second sub-data.
[0212] In some embodiments, the data processing module 4552 is further configured to determine the mean and standard deviation corresponding to the difference statistic; determine the upper limit and lower limit of the interval based on the mean and the standard deviation; and take the range between the upper limit and the lower limit of the interval as the data fluctuation range.
[0213] In some embodiments, the data processing module 4552 is further configured to determine a first sum between each of the sub-difference statistics, divide the first sum by the total number of the sub-data to obtain the mean corresponding to the difference statistic; determine the square of the difference between each of the sub-difference statistics and the mean, and determine a second sum between each of the squares; and take the square root of the ratio between the second sum and the total number of the sub-data as the standard deviation corresponding to the difference statistic.
[0214] In some embodiments, the data processing module 4552 is further configured to use a third sum between the mean and the standard deviation of a pre-configured multiple as the upper limit of the interval, wherein the pre-configured multiple is the number of second sub-data used to determine the sub-difference statistics of the first sub-data; and to use the difference between the mean and the standard deviation of the pre-configured multiple as the upper limit of the interval.
[0215] In some embodiments, the anomaly detection module 4553 is further configured to determine that the sub-data corresponding to the sub-difference statistical value outside the data fluctuation range is abnormal in response to any sub-difference statistical value in the difference statistics being outside the data fluctuation range.
[0216] In some embodiments, the anomaly detection module 4553 is further configured to, after determining the abnormal sub-data in the time series data to be detected based on the differential statistics and the data fluctuation range, issue an abnormal data alarm in response to the abnormal sub-data in the time series data to be detected meeting a pre-configured condition, wherein the pre-configured condition includes at least one of the following: the amount of the abnormal sub-data reaches a pre-configured quantity; the difference between the abnormal sub-data and the upper limit or lower limit of the data fluctuation range reaches a threshold.
[0217] In some embodiments, the data acquisition module 4551 is further configured to acquire sample time series data before acquiring the time series data to be detected, wherein the sample time series data includes: multiple sample data with different time series of sample objects, and actual types corresponding to the multiple sample data respectively, wherein the actual types include abnormal sample data and normal sample data; and to train a time series model based on the sample time series data to obtain the trained time series model, wherein the trained time series model is used to determine abnormal sub-data in the time series data to be detected.
[0218] In some embodiments, the data acquisition module 4551 is further configured to: call the time series model based on the sample time series data to perform feature extraction processing to obtain the sequence features of the sample time series data, wherein the sequence features are used to characterize the difference statistics and the data fluctuation range; call the time series model based on the sequence features to perform prediction processing to obtain the prediction type of each sample data; determine the cross-entropy loss of the time series model based on the prediction type and the actual type of each sample data; and perform backpropagation processing on the time series model based on the cross-entropy loss to obtain the trained time series model.
[0219] In some embodiments, the anomaly detection module 4553 is further configured to, based on the differential statistics and the data fluctuation range, call the trained time series model to perform prediction processing to obtain abnormal sub-data in the time series data to be detected.
[0220] This application provides a computer program product, which includes a computer program or computer-executable instructions stored in a computer-readable storage medium. The processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the data anomaly detection method described in this application.
[0221] This application provides a computer-readable storage medium storing computer-executable instructions or a computer program. When the computer-executable instructions or the computer program are executed by a processor, the processor will execute the data anomaly detection method provided in this application. For example, ... Figure 3A The data anomaly detection method is shown.
[0222] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.
[0223] In some embodiments, computer-executable instructions may take the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as stand-alone programs or as modules, components, subroutines, or other units suitable for use in a computing environment.
[0224] As an example, computer-executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple co-located files (e.g., files that store one or more modules, subroutines, or code sections).
[0225] As an example, computer-executable instructions can be deployed to execute on a single electronic device, or on multiple electronic devices located at one location, or on multiple electronic devices distributed across multiple locations and interconnected via a communication network.
[0226] In summary, the embodiments of this application can perform differential statistical value calculation on sub-data in the time series data to be detected, determine a reasonable data fluctuation range, accurately identify abnormal data in the time series data to be detected based on the data fluctuation range, improve the accuracy of determining differential statistical values and data fluctuation range, and improve the accuracy of data anomaly detection.
[0227] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.
Claims
1. A method for detecting data anomalies, characterized in that, The method includes: Acquire the time series data to be detected, wherein the time series data to be detected includes multiple sub-data with different time sequences; For each of the sub-data, the following processing is performed: the sub-data is taken as the first sub-data, and based on the first sub-data and a plurality of second sub-data corresponding to the first sub-data, the sub-difference statistics of the first sub-data are determined, wherein the timing of the second sub-data is before the timing of the first sub-data. The data fluctuation range corresponding to the time series data to be detected is determined based on the differential statistics, wherein the differential statistics include: the sub-differential statistics value of each sub-data; Based on the differential statistics and the data fluctuation range, abnormal sub-data in the time series data to be detected are determined.
2. The method according to claim 1, characterized in that, The acquisition of the time series data to be detected includes: Obtain statistical data of the object to be detected within a pre-configured time period; The statistical data is divided according to the time index of the statistical data to obtain multiple sub-data with different sampling times; The multiple sub-data are sorted according to the chronological order of each time index to obtain the time series data to be detected.
3. The method according to claim 1, characterized in that, The step of taking the sub-data as the first sub-data, and determining the sub-difference statistics of the first sub-data based on the first sub-data and multiple second sub-data corresponding to the first sub-data, includes: Obtain the first difference between the first sub-data and the plurality of second sub-data; Obtain the product between each of the first differences; The Nth root of the product is used as the sub-difference statistic of the first sub-data, where N is the number of the plurality of second sub-data.
4. The method according to claim 1, characterized in that, The step of determining the data fluctuation range corresponding to the time series data to be detected based on differential statistics includes: Determine the mean and standard deviation corresponding to the difference statistic; Based on the mean and the standard deviation, the upper limit and lower limit of the interval are determined respectively. The range between the upper limit and the lower limit of the interval is defined as the data fluctuation range.
5. The method according to claim 4, characterized in that, Determining the mean and standard deviation corresponding to the difference statistic includes: Determine a first sum between each of the sub-difference statistics, and divide the first sum by the total number of the sub-data to obtain the mean value corresponding to the difference statistics; Determine the square of the difference between each of the sub-difference statistics and the mean, and determine the second summation between each of the squares; The square root of the ratio between the second sum and the total number of the sub-data is taken as the standard deviation of the difference statistic.
6. The method according to claim 4, characterized in that, The step of determining the upper limit and lower limit of the interval based on the mean and the standard deviation includes: The third sum between the mean and the standard deviation of the pre-configured multiple is used as the upper limit of the interval, wherein the pre-configured multiple is the number of second sub-data points used to determine the sub-difference statistics of the first sub-data points; The difference between the mean and the standard deviation of the pre-configured multiple is used as the lower limit of the interval.
7. The method according to claim 1, characterized in that, The step of determining abnormal sub-data in the time series data to be detected based on the differential statistics and the data fluctuation range includes: In response to any sub-difference statistic in the difference statistics being outside the data fluctuation range, it is determined that the sub-data corresponding to the sub-difference statistic outside the data fluctuation range is abnormal.
8. The method according to any one of claims 1 to 7, characterized in that, After determining the anomalous sub-data in the time series data to be detected based on the differential statistics and the data fluctuation range, the method further includes: In response to the abnormal sub-data in the time series data to be detected meeting the pre-configured conditions, an abnormal data alarm is issued, wherein the pre-configured conditions include at least one of the following: the amount of abnormal sub-data reaches the pre-configured quantity; the difference between the abnormal sub-data and the upper or lower limit of the data fluctuation range reaches a threshold.
9. The method according to any one of claims 1 to 7, characterized in that, Before acquiring the time series data to be detected, the method further includes: Acquire sample time series data, wherein the sample time series data includes: multiple sample data with different time series of sample objects, and the actual types corresponding to the multiple sample data respectively, wherein the actual types include abnormal sample data and normal sample data; A time series model is trained based on the sample time series data to obtain the trained time series model, wherein the trained time series model is used to determine abnormal sub-data in the time series data to be detected.
10. The method according to claim 9, characterized in that, The step of training a time series model based on the sample time series data to obtain the trained time series model includes: Based on the sample time series data, the time series model is called to perform feature extraction processing to obtain the sequence features of the sample time series data, wherein the sequence features are used to characterize the difference statistics and the data fluctuation range; Based on the sequence features, the time series model is invoked for prediction processing to obtain the prediction type for each sample data; The cross-entropy loss of the time series model is determined based on the predicted type and the actual type of each sample data. The time series model is backpropagated based on the cross-entropy loss to obtain the trained time series model.
11. The method according to claim 9, characterized in that, The step of determining abnormal sub-data in the time series data to be detected based on the differential statistics and the data fluctuation range includes: Based on the difference statistics and the data fluctuation range, the trained time series model is invoked for prediction processing to obtain abnormal sub-data in the time series data to be detected.
12. A data anomaly detection device, characterized in that, The device includes: The data acquisition module is used to acquire the time series data to be detected, wherein the time series data to be detected includes multiple sub-data with different time sequences; A data processing model is used to perform the following processing for each of the sub-data: taking the sub-data as first sub-data, determining the sub-difference statistics of the first sub-data based on the first sub-data and multiple second sub-data corresponding to the first sub-data, wherein the time sequence of the second sub-data is earlier than the time sequence of the first sub-data; determining the data fluctuation range corresponding to the time sequence data to be detected based on the difference statistics, wherein the difference statistics include: the sub-difference statistics of each of the sub-data; An anomaly detection module is used to determine abnormal sub-data in the time series data to be detected based on the differential statistics and the data fluctuation range.
13. A device, characterized in that, The device includes: Memory is used to store executable instructions for a computer; A processor, when executing computer-executable instructions or computer programs stored in the memory, implements the data anomaly detection method according to any one of claims 1 to 11.
14. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by a processor, the data anomaly detection method according to any one of claims 1 to 11 is implemented.
15. A computer program product comprising computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by a processor, the data anomaly detection method according to any one of claims 1 to 11 is implemented.