Streaming data anomaly detection method and device, electronic equipment and storage medium
Through the combination of sliding window and local outlier factor algorithm, the problems of high memory consumption and low efficiency in streaming data abnormality detection are solved, and efficient abnormality detection under finite memory conditions is achieved, adapting to dynamic changes in data flow, improving detection accuracy and real-timeness.
Patent Information
- Application Number
- CN202510517816.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-04
AI Technical Summary
The problems of high memory consumption and low detection efficiency in streaming data abnormality detection methods are difficult to meet the requirements of high efficiency and low latency in real-time data services.
The sliding window mechanism and local outlier factor algorithm are used to cache streaming data points by preset sliding windows, detect abnormal data points based on local outlier factors, and mark and delete them. The data points are sampled and merged at the same time to optimize memory usage and calculation overhead.
While reducing memory usage, it improves the efficiency of streaming data abnormal detection, dynamically adapts to data flow changes, reduces calculation overhead, and improves detection accuracy and real-time performance.
Smart Images

Figure CN120263690A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of communication technology, and in particular to a method, device, electronic device and storage medium for detecting anomalies in streaming data. Background Art
[0002] Streaming data refers to data that arrives in the form of a continuous, unbounded data stream. Unlike traditional batch data processing, streaming data is generated and processed in real time, and usually needs to be analyzed and responded to at the same time as the data is generated. This data processing method is particularly suitable for application scenarios that require instant response, such as real-time monitoring, event-driven applications, and online transaction processing. With the development of real-time data services, the processing and analysis of streaming data has gradually become an important part of data development and decision support. In the field of big data, the demand for real-time data processing continues to rise, and the timely detection and processing of abnormal data in the process of streaming data processing has become a key issue to ensure system stability and data quality.
[0003] In the process of synchronizing and processing streaming data, problems such as data loss, duplication or anomalies often occur, which not only interrupts task operation, but also may cause deviations in downstream data calculation results, thus affecting the stability and accuracy of real-time business. Traditional data synchronization technology usually lacks an efficient data quality detection mechanism, making it difficult to timely detect and process abnormal data during data transmission, and it is difficult to effectively ensure the integrity and consistency of real-time data.
[0004] Anomaly detection of streaming data is a crucial task in the streaming data processing process. Traditional anomaly detection methods are mainly designed for offline data and usually require storing the entire data set and the distance relationship between its data points. This detection method is suitable for small-scale data sets. However, in streaming data scenarios, with the continuous growth of data volume, the demand for memory resources continues to increase, resulting in high memory consumption and poor scalability. At the same time, streaming data is continuous and dynamic. Traditional anomaly detection methods need to recalculate relevant indicators (such as the distance between data points) for the entire data set when the data increment changes. The computational overhead increases significantly, making it difficult to meet the requirements of real-time data services for high efficiency and low latency.
[0005] Therefore, how to improve the efficiency of streaming data anomaly detection under the condition of limited memory is one of the technical problems that need to be urgently solved in the existing technology. Summary of the invention
[0006] In order to solve the problems of high memory consumption and low detection efficiency in streaming data anomaly detection methods, embodiments of the present application provide a streaming data anomaly detection method, device, electronic device, and storage medium.
[0007] In a first aspect, an embodiment of the present application provides a method for detecting anomalies in streaming data, including:
[0008] In response to a data anomaly detection request sent by a client, obtain a data stream to be detected from a data source, where the data stream to be detected includes data points generated in chronological order;
[0009] Store each obtained data point into a preset sliding window in chronological order;
[0010] When the amount of data in the sliding window reaches a first preset ratio, determine the local outlier factor of each data point according to each data point in the current sliding window and its k-nearest neighbor data points respectively;
[0011] Determine the data points with local outlier factors greater than the outlier factor threshold as anomaly data points, mark the anomaly data points, return the marked anomaly data points to the client, and then delete them from the sliding window;
[0012] Sample data points with a second preset ratio from the remaining data points according to the local outlier factors of the remaining data points in the sliding window, and merge the sampled data points with the data points newly arriving at the sliding window, where the second preset ratio is less than one-half of the sliding window;
[0013] When the amount of data in the sliding window reaches the capacity limit, re-determine the local outlier factor of each data point in the sliding window that has slid out, and perform the steps of marking anomaly data points, sampling data points, and merging data points.
[0014] In one implementation, determining the local outlier factor of each data point according to each data point in the current sliding window and its k-nearest neighbor data points respectively specifically includes:
[0015] Calculate the reachable distance between each data point in the current sliding window and its k-nearest neighbor data points respectively;
[0016] Determine the local reachability density of each data point according to the reachable distance between each data point and its k-nearest neighbor data points;
[0017] Determine the local outlier factor of each data point according to the local reachability density of each data point and its k-nearest neighbor data points respectively.
[0018] In one implementation, sampling data points with a second preset ratio from the remaining data points according to the local outlier factors of the remaining data points in the sliding window specifically includes:
[0019] Arrange the remaining data points in the sliding window in descending order of local outlier factors;
[0020] Divide the remaining data points in the rearranged sliding window into a first preset number of intervals;
[0021] Sample data points at a third preset ratio from the first second preset number of intervals and sample data points at a fourth preset ratio from the last third preset number of intervals, where the third preset ratio is greater than the fourth preset ratio.
[0022] In one implementation, calculate the reachability distance between each data point in the current sliding window and its k-nearest neighbor data points, specifically including:
[0023] Calculate the reachability distance between each data point in the current sliding window and its k-nearest neighbor data points through the following formula:
[0024] reach-dist(p i ,p j ) = max(dist(p i ,p j ), k-distance(p j ))
[0025] where reach-dist(p i ,p j ) represents the reachability distance between the i-th data point p i and its j-th nearest neighbor data point p j in the current sliding window, p j ∈N k (p i ), N k (p i ) represents the set of k nearest neighbor data points of the i-th data point p i , p j represents the j-th nearest neighbor data point of the i-th data point p i ;
[0026] dist(p i ,p j ) represents the distance between the i-th data point p i and its j-th nearest neighbor data point p j ;
[0027] k-distance(p j ) represents the distance from the data point p j to its k-th nearest neighbor data point.
[0028] In one implementation, determine the local reachability density of each data point according to the reachability distance between each data point and its k-nearest neighbor data points, specifically including:
[0029] The local reachability density of each data point is determined by calculating the reachability distance between each data point and its k nearest neighbor data points through the following formula:
[0030]
[0031] where LRD(p i ) represents the local reachability density of the i-th data point p in the current sliding window i .
[0032] In one embodiment, the local outlier factor of each data point is determined according to the local reachability density of each data point and its respective k nearest neighbor data points, specifically including:
[0033] The local outlier factor of each data point is calculated through the following formula:
[0034]
[0035] where LOF(p i ) represents the local outlier factor of the i-th data point p in the current sliding window i ;
[0036] LRD(p i ) represents the local reachability density of the i-th data point p i ;
[0037] LRD(p j ) represents the local reachability density of the j-th nearest neighbor data point p i of the i-th data point p j ;
[0038] |N k (p i )| represents the modulus of the set of k nearest neighbor data points of the i-th data point p i .
[0039] In a second aspect, an embodiment of the present application provides a streaming data anomaly detection device, including:
[0040] An acquisition module, configured to obtain a data stream to be detected from a data source in response to a data anomaly detection request sent by a client, where the data stream to be detected includes data points generated in chronological order;
[0041] A caching module, configured to sequentially store each obtained data point into a preset sliding window in chronological order;
[0042] A determination module, configured to, when the data volume in the sliding window reaches a first preset ratio, determine the local outlier factor of each data point according to each data point in the current sliding window and its k nearest neighbor data points respectively;
[0043] An anomaly recognition module, configured to determine the data points with local outlier factors greater than the outlier factor threshold as anomaly data points, mark the anomaly data points, return the marked anomaly data points to the client, and then delete them from the sliding window;
[0044] A sampling and merging module, configured to sample data points with a second preset ratio from the remaining data points according to the local outlier factors of the remaining data points in the sliding window, and merge the sampled data points with the data points newly arriving at the sliding window, where the second preset ratio is less than one-half of the sliding window;
[0045] A processing module, configured to, when the data volume in the sliding window reaches the capacity limit, re-determine the local outlier factors of each data point in the sliding-out window, and perform the steps of marking anomaly data points, sampling and merging data points.
[0046] In one implementation, the determination module is specifically configured to calculate the reachable distance between each data point in the current sliding window and its k nearest neighbor data points respectively; determine the local reachability density of each data point according to the reachable distance between each data point and its k nearest neighbor data points; and determine the local outlier factor of each data point according to the local reachability density of each data point and its k nearest neighbor data points respectively.
[0047] In one implementation, the sampling and merging module is specifically configured to arrange the remaining data points in the sliding window in descending order of the local outlier factor; divide the remaining data points in the re-arranged sliding window into a first preset number of intervals; sample data points with a third preset ratio from the first second preset number of intervals, and sample data points with a fourth preset ratio from the last third preset number of intervals, where the third preset ratio is greater than the fourth preset ratio.
[0048] In one implementation, the determination module is specifically configured to calculate the reachable distance between each data point in the current sliding window and its k nearest neighbor data points through the following formula:
[0049] reach-dist(p i ,p j )=max(dist(p i ,p j ),k-distance(p j ))
[0050] where, reach-dist(p i , p j ) represents the reachability distance between the i-th data point p i and its j-th nearest neighbor data point p j in the current sliding window. p j ∈ N k (p i ) where N k (p i ) represents the set of k nearest neighbor data points of the i-th data point p i , and p j represents the j-th nearest neighbor data point of the i-th data point p i ;
[0051] dist(p i , p j ) represents the distance between the i-th data point p i and its j-th nearest neighbor data point p j ;
[0052] k-distance(p j ) represents the distance from the data point p j to its k-th nearest neighbor data point.
[0053] In one embodiment, the determining module is specifically configured to calculate the reachability distance between each data point and its k nearest neighbor data points through the following formula to determine the local reachability density of each data point:
[0054]
[0055] where, LRD(p i ) represents the local reachability density of the i-th data point p i in the current sliding window.
[0056] In one embodiment, the determining module is specifically configured to calculate the local outlier factor of each data point through the following formula:
[0057]
[0058] where, LOF(p i ) represents the local outlier factor of the i-th data point p i in the current sliding window;
[0059] LRD(p i ) represents the local reachability density of the i-th data point p i ;
[0060] LRD(pj ) represents the j-th nearest neighbor data point p i of the i-th data point p j and its local reachability density;
[0061] |N k (p i )| represents the modulus of the set of k nearest neighbor data points of the i-th data point p i .
[0062] In a third aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the streaming data anomaly detection method described in the present application is implemented.
[0063] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps in the streaming data anomaly detection method described in the present application are implemented.
[0064] The beneficial effects of the present application are as follows:
[0065] The streaming data anomaly detection method, device, electronic device, and storage medium provided by the embodiments of this application respond to a data anomaly detection request sent by a client, obtain a data stream to be detected from a data source, where the data stream to be detected includes data points generated in chronological order, and store each obtained data point into a preset sliding window in chronological order. When the amount of data in the sliding window reaches a first preset ratio, determine the Local Outlier Factor (LOF) of each data point according to each data point in the current sliding window and its k nearest neighbor data points respectively, determine the data points with local outlier factors greater than the outlier factor threshold as anomaly data points, mark the anomaly data points, return the marked anomaly data points to the client and then delete them from the sliding window, sample data points with a second preset ratio from the remaining data points according to the local outlier factors of the remaining data points in the sliding window, and merge the sampled data points with the data points newly arriving at the sliding window, where the second preset ratio is less than one-half of the sliding window. When the amount of data in the sliding window reaches the capacity limit, re-determine the local outlier factors of each data point sliding out of the window, and perform the steps of marking anomaly data points, sampling data points, and merging. In the embodiments of this application, a sliding window with a fixed size is preset to cache streaming data. When the amount of data in the sliding window reaches a first preset ratio, the local outlier factor algorithm is combined to detect the anomaly data points, that is, outlier points, within the current sliding window, mark the anomaly data points and feedback them to the client and then delete them from the sliding window, and perform data sampling and merging based on the local outlier factor values of the data points. Thus, while maintaining the distribution characteristics of the data, the memory usage is optimized, full-scale calculation is avoided through the local update strategy, and the calculation overhead is significantly reduced. Therefore, the streaming data anomaly detection efficiency is improved while reducing the memory usage rate.
[0066] Other features and advantages of this application will be described in the following specification, and, in part, will become apparent from the specification, or will be understood by implementing this application. The objectives and other advantages of this application can be achieved and obtained through the structures specifically pointed out in the written specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] The drawings described herein are used to provide a further understanding of this application, and constitute a part of this application. The illustrative embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application. In the drawings:
[0068] Figure 1 It is a schematic diagram of the application scenario of the streaming data anomaly detection method provided by the embodiments of this application;
[0069] Figure 2Schematic flowchart of the streaming data anomaly detection method provided by the embodiment of the present application;
[0070] Figure 3 Schematic flowchart of determining the local outlier factor of each data point in the current sliding window provided by the embodiment of the present application;
[0071] Figure 4 An example diagram for calculating the local outlier factor of a data point provided by the embodiment of the present application;
[0072] Figure 5 Schematic flowchart of sampling the remaining data points in the current sliding window provided by the embodiment of the present application;
[0073] Figure 6 An example diagram of data point sampling provided by the embodiment of the present application;
[0074] Figure 7 An example diagram of data point sampling and merging provided by the embodiment of the present application;
[0075] Figure 8 Schematic structural diagram of the streaming data anomaly detection device provided by the embodiment of the present application;
[0076] Figure 9 Schematic structural diagram of the electronic device provided by the embodiment of the present application. Detailed implementation manners
[0077] To solve the problems of high memory consumption and low detection efficiency of the streaming data anomaly detection method, the embodiment of the present application provides a streaming data anomaly detection method, device, electronic device and storage medium.
[0078] The following describes the preferred embodiments of the present application with reference to the accompanying drawings of the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application. And without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.
[0079] In this article, it should be understood that among the technical terms involved in the present application:
[0080] 1. Streaming data: Also known as time-series data, it refers to data generated in chronological order, that is, data continuously generated over time. Streaming data consists of data points generated in chronological order. The generation times of different data points are different. That is, one data can be generated at one time point, and this one data can be called a data point. The data points generated at multiple adjacent time points in sequence constitute the streaming data. Streaming data has characteristics such as large data volume and continuous arrival of data points.
[0081] 2. Local Reachability Density (LRD): A density measurement method mainly used to measure the density of a data point relative to its neighborhood. LRD calculates the density by considering the reachability distance of a data point to its neighborhood, aiming to eliminate the possible noise effects in traditional density measurement methods. LRD calculates the reachability distance between a data point and its k nearest neighbors as the density of this point, reflecting the density change in the local area.
[0082] 3. Local Outlier Factor: A density-based outlier detection algorithm used to identify outliers in a dataset. It evaluates whether a point is an outlier by comparing the density difference between the data point and other points in its neighborhood. The Local Outlier Factor calculates the local density of a point and compares it with the density of its neighborhood points. It is suitable for discovering outliers with large local density differences and can be used in fields such as network security, financial fraud, and medical diagnosis. This method has strong local outlier detection capabilities and is suitable for uneven density distributions.
[0083] First, refer to Figure 1, which is a schematic diagram of an application scenario of the streaming data anomaly detection method provided by the embodiments of the present application. It may include a client 101, a monitoring device 102, and a data source 103. The client 101 is connected to the monitoring device 102 through a network, and the monitoring device 102 is connected to the data source 103 through a network. The data source 103 contains real-time generated business data. The business data is streaming data, and the streaming data includes data points generated in chronological order. Each data point carries its own generation timestamp. The business data can be, but is not limited to, the following real-time business data: monitoring data (such as video monitoring data, etc.), sensor data (such as temperature sensor data, pressure sensor data, etc.), online financial transaction data, etc. Correspondingly, the data source 103 may include: video monitoring devices (such as cameras), temperature sensors, pressure sensors, online financial transaction monitoring devices, etc., and can also be any other device that can generate streaming data. The embodiments of the present application do not limit this. In response to a data anomaly detection request sent by the client 101, the monitoring device 102 obtains the data stream to be detected from the data source 103. The data stream to be detected includes data points generated in chronological order. The monitoring device 102 stores each obtained data point into a preset sliding window in chronological order. When the data volume in the sliding window reaches the first preset ratio, the local outlier factor of each data point is determined respectively according to each data point in the current sliding window and its k nearest neighbor data points. The data points with local outlier factors greater than the outlier factor threshold are determined as abnormal data points, and the abnormal data points are marked. After the marked abnormal data points are returned to the client 101, they are deleted from the sliding window. The second preset ratio of data points is sampled from the remaining data points according to the local outlier factors of the remaining data points in the sliding window, and the sampled data points are merged with the data points newly arriving at the sliding window, where the second preset ratio is less than half of the sliding window. When the data volume in the sliding window reaches the capacity limit, the local outlier factor of each data point sliding out of the window is re-determined, and the steps of abnormal data point marking, data point sampling, and merging are executed. Thus, while reducing the memory usage rate, the efficiency of streaming data anomaly detection is improved.
[0084] In the present application, a streaming data processing platform, such as a Kafka cluster, may also be included. First, the Kafka cluster can use real-time stream processing tools such as Flink or Spark to obtain streaming data from the data source 103 and store the streaming data into the Kafka message queue in chronological order. When the monitoring device 102 receives a data anomaly detection request sent by the client 101, the monitoring device 102 can obtain the data stream to be detected from the Kafka message queue. The embodiments of the present application do not limit this.
[0085] The monitoring device can be any server, cluster, device or equipment with computing capabilities. The server can be an independent physical server or a cloud server that provides basic cloud computing services such as cloud servers, cloud databases, and cloud storage. The embodiments of the present application do not limit this.
[0086] Based on the above application scenarios, the exemplary embodiments of the present application will be described in more detail below with reference to the attached Figures 2 to 7 It should be noted that the above application scenarios are only shown for the convenience of understanding the spirit and principle of the present application, and the embodiments of the present application are not restricted by any of them. On the contrary, the embodiments of the present application can be applied to any applicable scenario.
[0087] As Figure 2 shown, it is a schematic diagram of the implementation process of the streaming data anomaly detection method provided by the embodiments of the present application. This streaming data anomaly detection method can be applied to the above-mentioned monitoring device 102 and specifically includes the following steps:
[0088] S21. In response to a data anomaly detection request sent by the client, the monitoring device obtains a data stream to be detected from the data source.
[0089] In specific implementation, when the monitoring device receives a data anomaly detection request sent by the client, it uses real-time stream processing tools such as Flink or Spark to obtain a real-time data stream to be detected from the data source 103. The data stream to be detected is a streaming data, which includes data points generated in chronological order, and each data point carries its own generation timestamp. That is, a data point is a piece of data carrying its generation timestamp.
[0090] S22. Store each obtained data point into a preset sliding window in chronological order.
[0091] In specific implementation, the monitoring device pre-sets a preset sliding window. This sliding window is a time sliding window with a fixed size of W, where W is a preset time range and can be set according to actual needs, such as it can be set to 1 minute. The embodiments of the present application do not limit this. This sliding window serves as a data buffer for receiving and storing the continuous streaming data obtained by the monitoring device from the data source. This sliding window divides the data stream into non-overlapping continuous time blocks according to timestamps. Whenever a new data point arrives, all the data points in the sliding window are updated according to the timestamps to ensure that the data points in the sliding window remain available within the time range determined by the window size.
[0092] In this step, the monitoring device stores each obtained data point into the preset sliding window in chronological order.
[0093] S23. When the amount of data in the sliding window reaches the first preset ratio, determine the local outlier factor of each data point according to each data point in the current sliding window and its k nearest neighbor data points respectively.
[0094] In specific implementation, when the amount of data in the sliding window reaches the first preset ratio, the monitoring device determines the local outlier factor value of each data point according to each data point that arrives in the current sliding window and its k nearest neighbor data points respectively. The first preset ratio can be set by itself according to requirements. For example, it can be set to 50% but not limited to this. That is, when the capacity W of the sliding window is 1 minute, the amount of data in the sliding window reaching 50% means the data points in the first half minute that arrive at the sliding window.
[0095] Specifically, the local outlier factor of each data point in the current sliding window can be determined according to the steps as Figure 3 shown, including the following steps:
[0096] S31. Calculate the reachable distance between each data point in the current sliding window and its k nearest neighbor data points respectively.
[0097] In specific implementation, the value of k is set in advance, and this application embodiment does not limit this. For each data point, the monitoring device calculates the Euclidean distance between this data point and each other data point, and determines the k nearest neighbor data points of this data point as the k data points with the smallest Euclidean distance.
[0098] Then, the reachable distance between each data point in the current sliding window and its k nearest neighbor data points can be calculated through the following formula:
[0099] reach-dist(p i ,p j ) = max(dist(p i ,p j ), k-distance(p j ))
[0100] where, reach-dist(p i ,p j ) represents the reachable distance between the i-th data point p i and its j-th nearest neighbor data point p j in the current sliding window. p j ∈N k (p i ), N k (p i ) represents the set of k nearest neighbor data points of the i-th data point p i . j = 1, 2,..., k. p j represents the i-th data point pi The j-th nearest neighbor data point of
[0101] dist(p i , p j ) represents the distance between the i-th data point p i and its j-th nearest neighbor data point p j .
[0102] k-distance(p j ) represents the distance from the data point p j to its k-th nearest neighbor data point.
[0103] During implementation, the dist(p i , p j ) value can be obtained by calculating the Euclidean distance between the i-th data point p i and its j-th nearest neighbor data point p j . Similarly, the k-distance(p j ) value can be obtained by calculating the Euclidean distance from the data point p j to its k-th nearest neighbor data point p j .
[0104] Specifically, for each data point (i.e., each piece of data), the monitoring device first inputs the data point into a text embedding model (Embedding) for embedding to obtain a vector of the data point. The text embedding model is used to generate a vector representation of the text. The text embedding model can be, but is not limited to, the following models: the text-embedding-large-3 model of OpenAI, the BCEmbedding model of Youdao, etc. Any other model that can generate text vectors can also be used. This application embodiment does not make any limitations in this regard.
[0105] Then, the monitoring device obtains the dist(p i , p j ) value by calculating the Euclidean distance between the vector of the i-th data point p i and the vector of its j-th nearest neighbor data point p j . The k-distance(p j ) value is obtained by calculating the Euclidean distance from the vector of the data point p j to the vector of its k-th nearest neighbor data point p j .
[0106] S32. Determine the local reachability density of each data point according to the reachable distance between each data point and its k-nearest neighbor data points.
[0107] In specific implementation, after the monitoring device calculates the reachable distance between each data point in the current sliding window and its k nearest neighbor data points, the local reachability density of each data point can be calculated through the following formula:
[0108]
[0109] where LRD(p i ) represents the local reachability density of the i-th data point p i in the current sliding window;
[0110] reach-dist(p i ,p j ) represents the reachable distance between the i-th data point p i and its j-th nearest neighbor data point p j .
[0111] The local reachability density reflects the density characteristics of data points in their local neighborhoods. When the local reachability density value of a data point is relatively high, it indicates that the data distribution around this data point is relatively dense. When the local reachability density value of a data point is relatively low, it indicates that this data point is in a sparse area of the data distribution.
[0112] S33. Determine the local outlier factor of each data point according to the local reachability density of each data point and its respective k nearest neighbor data points.
[0113] In specific implementation, the local outlier factor of each data point can be calculated through the following formula:
[0114]
[0115] where LOF(p i ) represents the local outlier factor of the i-th data point p i in the current sliding window;
[0116] LRD(p i ) represents the local reachability density of the i-th data point p i ;
[0117] LRD(p j ) represents the local reachability density of the j-th nearest neighbor data point p i of the i-th data point p j ;
[0118] |N k (p i )| represents the modulus of the set of k nearest neighbor data points of the i-th data point p i .
[0119] When the local outlier factor value of a data point is significantly greater than 1, it indicates that the density of this data point in the current distribution is significantly lower than that of other data points, and this data point is an isolated outlier data point, that is: an outlier.
[0120] As Figure 4 shown, it is an example diagram for calculating the local outlier factor of a data point, and k is set to 3.
[0121] S24. Determine the data points with local outlier factors greater than the outlier factor threshold as outlier data points, mark the outlier data points, and after returning the marked outlier data points to the client, delete them from the sliding window.
[0122] In specific implementation, the outlier factor threshold Q can be preset according to empirical values. For example, the outlier factor threshold Q can be set to 10, or it can be set to other values. This application embodiment does not make a limitation in this regard. After the monitoring device calculates the local outlier factor of each data point in the current sliding window, it compares the local outlier factor value of each data point with the outlier factor threshold, determines the data points with local outlier factor values greater than the outlier factor threshold as outlier data points, and can mark the outlier data points through the specified fields of the data points. For example, this data can be marked as outlier data by setting the reserved field of this data to "01", and it can also be marked with other characters. This application embodiment does not make a limitation in this regard. Furthermore, the monitoring device returns the marked outlier data points to the client for the operation and maintenance personnel to process through the client. Furthermore, the monitoring device deletes the outlier data points from the sliding window. In this way, while real-time identifying outlier data, the memory of the monitoring device can be saved.
[0123] S25. Sample data points with a second preset ratio from the remaining data points according to the local outlier factors of the remaining data points in the sliding window, and merge the sampled data points with the data points newly arriving at the sliding window.
[0124] In this step, the second preset ratio is less than the first preset ratio, and the second preset ratio is less than half of the sliding window, and it can be set by itself during implementation. For example, the second preset ratio can be set to 25% of the sliding window, or it can also be set to any ratio less than 50%. This application embodiment does not make a limitation in this regard.
[0125] During the implementation process, in order to maintain the accuracy of outlier detection, the sampled data samples should meet the following requirements:
[0126] (1) It can maintain the distribution density of the original data.
[0127] (2) Retain the diversity of the original data as much as possible.
[0128] To meet the above requirements, it can be carried out asFigure 5 The process shown samples the remaining data points in the current sliding window, including the following steps:
[0129] S41. Arrange the remaining data points in the sliding window in descending order of the local outlier factor.
[0130] Specifically in implementation, the monitoring device arranges the remaining data points in the current sliding window after deleting the abnormal data points in descending order of their respective local outlier factor values.
[0131] S42. Divide the remaining data points in the rearranged sliding window into a first preset number of intervals.
[0132] Specifically in implementation, the first preset number m can be set by itself, m > 3, for example, it can be set to 10, and the embodiments of the present application do not limit this. The remaining data points in the rearranged sliding window are evenly divided into m intervals:
[0133] D = {D1, D2, ……, D m}
[0134] S43. Sample data points at a third preset ratio from the first second preset number of intervals from the front, and sample data points at a fourth preset ratio from the third preset number of intervals from the back.
[0135] Among them, the sum of the second preset number and the third preset number is the first preset number, and the third preset ratio is greater than the fourth preset ratio. The second preset number, the third preset number, the third preset ratio, and the fourth preset ratio can be set by itself. Assuming that the first preset number is set to 10, the second preset number is 3, the third preset number is 7, the third preset ratio is set to 80%, and the fourth preset ratio is set to 20%, then the monitoring device can randomly sample 80% of the data points from the first 3 intervals, denoted as S1, and sample 20% of the data points from the last 7 intervals, denoted as S2. The proportion of the total sampled data points in the sliding window is 25%, that is: 25%W: D sample = S1 ∪ S2, as Figure 6 shown. In this way, a higher proportion of data points are collected from the intervals where the local outlier factor values are high, which can ensure that the sampled data can retain the distribution characteristics of the dense area of the data set as much as possible. A lower proportion of data points are collected from the intervals where the local outlier factor values are low, which can ensure the diversity of data sampling and avoid over-concentration in the data-dense area, thereby improving the perception ability of the sparse data area.
[0136] Furthermore, the monitoring device combines the sampled data points with the newly arrived data points in the sliding window, deletes the other unsampled data from the sliding window, and the 25%W of the sampled data points D sampleMerge with the newly arrived 50% W data points D recent (i.e., the most recently arrived 50% of the data points) to obtain the merged data set D final :
[0137] D final = D sample ∪D recent
[0138] After the above data point sampling and merging, the capacity of the sliding window will be compressed to 75%, as Figure 7 shown, to continuously receive new data points.
[0139] S26. When the amount of data in the sliding window reaches the capacity limit, re-determine the local outlier factor of each data point in the sliding window, and perform the steps of abnormal data point marking, data point sampling and merging.
[0140] Specifically, when the amount of data in the sliding window reaches the capacity limit W, re-determine the local outlier factor of each data point in the sliding window according to step S23. This is because when new data points arrive, the data distribution in the sliding window changes, and the arrival of new data points may affect the nearest neighbor relationship of existing data points, thus changing their local reachability density. To ensure the accuracy of the local reachability density of data points and improve the abnormal data point detection accuracy, it is necessary to re-determine the local outlier factor of each data point according to the new data distribution and then make a judgment.
[0141] Furthermore, perform the steps of abnormal data point determination, abnormal data point marking, feedback of the marked data points to the client and then deletion of the abnormal data points according to step S24, and perform the steps of data point sampling and merging according to step S25 to ensure continuous incremental update and abnormal detection in the sliding window. Through the data sampling scheme based on interval local reachability density, each new data point of the data stream is added to the end of the sliding window. When the amount of data in the sliding window reaches the maximum capacity W, the sampling and merging operations are automatically performed. Thus, it is ensured that the amount of data in the sliding window always remains within W, and at the same time, some data with little impact is discarded, enabling real-time abnormal detection of streaming data under the condition of a fixed-size memory (the size of the sliding window W), effectively solving the problems of large memory consumption and low detection efficiency faced by traditional streaming data abnormal detection methods.
[0142] This application can dynamically adapt to the incremental changes in the data stream. Through data sampling and merging based on the interval local reachability density, it is not necessary to recalculate the local reachability density for all data points, significantly reducing memory consumption and computational overhead. At the same time, it overcomes the deficiencies of traditional streaming data anomaly detection methods in terms of real-time performance and scalability, significantly improving the efficiency and accuracy of large-scale streaming data in scenarios such as real-time monitoring, fault diagnosis, and alarm response, providing efficient and reliable technical support for streaming data processing tasks.
[0143] In the streaming data anomaly detection method provided by the embodiments of this application, the monitoring device responds to a data anomaly detection request sent by the client, obtains the data stream to be detected from the data source. The data stream to be detected includes data points generated in chronological order, and each obtained data point is sequentially stored in a preset sliding window in chronological order. When the data volume in the sliding window reaches the first preset ratio, the local outlier factor of each data point is determined respectively according to each data point in the current sliding window and its k-nearest neighbor data points. The data points with the local outlier factor greater than the outlier factor threshold are determined as abnormal data points, and the abnormal data points are marked. After the marked abnormal data points are returned to the client, they are deleted from the sliding window. A second preset ratio of data points is sampled from the remaining data points according to the local outlier factors of the remaining data points in the sliding window, and the sampled data points are merged with the data points newly arriving at the sliding window, where the second preset ratio is less than half of the sliding window. When the data volume in the sliding window reaches the capacity limit, the local outlier factor of each data point sliding out of the window is re-determined, and the steps of marking abnormal data points, data point sampling, and merging are performed. In the embodiments of this application, a sliding window with a fixed size is preset to cache the streaming data. When the data volume in the sliding window reaches the first preset ratio, the local outlier factor algorithm is combined to detect the abnormal data points, that is, the outlier points, in the current sliding window. The abnormal data points are marked and fed back to the client and then deleted from the sliding window, and data sampling and merging are performed based on the local outlier factor values of the data points. Thus, while maintaining the distribution characteristics of the data, the memory usage is optimized, and full-scale calculation is avoided through the local update strategy, significantly reducing the computational overhead. Therefore, the efficiency of streaming data anomaly detection is improved while reducing the memory usage rate.
[0144] Based on the same inventive concept, the embodiments of this application also provide a streaming data anomaly detection device. Since the principle of the above streaming data anomaly detection device for solving problems is similar to that of the above streaming data anomaly detection method, the implementation of the above device can refer to the implementation of the method, and the repeated parts will not be elaborated.
[0145] As Figure 8 shown, it is a schematic structural diagram of the streaming data anomaly detection device provided by the embodiments of this application, which can be applied to such as Figure 1In the monitoring device shown, the streaming data anomaly detection device may include:
[0146] An acquisition module 51, configured to obtain a data stream to be detected from a data source in response to a data anomaly detection request sent by a client, where the data stream to be detected includes data points generated in chronological order;
[0147] A cache module 52, configured to sequentially store each obtained data point into a preset sliding window in chronological order;
[0148] A determination module 53, configured to, when the amount of data in the sliding window reaches a first preset ratio, determine the local outlier factor of each data point according to each data point in the current sliding window and its k-nearest neighbor data points respectively;
[0149] An anomaly recognition module 54, configured to determine data points with local outlier factors greater than an outlier factor threshold as anomaly data points, mark the anomaly data points, return the marked anomaly data points to the client, and then delete them from the sliding window;
[0150] A sampling and merging module 55, configured to sample data points with a second preset ratio from the remaining data points according to the local outlier factors of the remaining data points in the sliding window, and merge the sampled data points with the data points newly arriving at the sliding window, where the second preset ratio is less than one-half of the sliding window;
[0151] A processing module 56, configured to, when the amount of data in the sliding window reaches the capacity limit, re-determine the local outlier factors of each data point in the sliding-out window, and perform the steps of marking anomaly data points, sampling and merging data points.
[0152] In one implementation, the determination module 53 is specifically configured to calculate the reachable distance between each data point in the current sliding window and its k-nearest neighbor data points respectively; determine the local reachability density of each data point according to the reachable distance between each data point and its k-nearest neighbor data points; and determine the local outlier factor of each data point according to the local reachability density of each data point and its k-nearest neighbor data points respectively.
[0153] In one implementation, the sampling and merging module 55 is specifically configured to arrange the remaining data points in the sliding window in descending order of local outlier factors; divide the remaining data points in the re-arranged sliding window into a first preset number of intervals; sample data points with a third preset ratio from the first second preset number of intervals, and sample data points with a fourth preset ratio from the last third preset number of intervals, where the third preset ratio is greater than the fourth preset ratio.
[0154] In one embodiment, the determining module 53 is specifically configured to calculate the reachability distance between each data point in the current sliding window and its k nearest neighbor data points through the following formula:
[0155] reach-dist(p i ,p j )=max(dist(p i ,p j ),k-distance(p j ))
[0156] where reach-dist(p i ,p j ) represents the reachability distance between the i-th data point p i and its j-th nearest neighbor data point p j in the current sliding window, p j ∈N k (p i ), N k (p i ) represents the set of k nearest neighbor data points of the i-th data point p i , and p j represents the j-th nearest neighbor data point of the i-th data point p i ;
[0157] dist(p i ,p j ) represents the distance between the i-th data point p i and its j-th nearest neighbor data point p j ;
[0158] k-distance(p j ) represents the distance from the data point p j to its k-th nearest neighbor data point.
[0159] In one embodiment, the determining module 53 is specifically configured to calculate the local reachability density of each data point by calculating the reachability distance between each data point and its k nearest neighbor data points through the following formula:
[0160]
[0161] where LRD(p i ) represents the local reachability density of the i-th data point p i in the current sliding window.
[0162] In one embodiment, the determining module 53 is specifically configured to calculate the local outlier factor of each data point through the following formula:
[0163]
[0164] Among them, LOF(p i ) represents the local outlier factor of the i-th data point p i in the currently described sliding window;
[0165] LRD(p i ) represents the local reachability density of the i-th data point p i ;
[0166] LRD(p j ) represents the local reachability density of the j-th nearest neighbor data point p i of the i-th data point p j ;
[0167] |N k (p i )| represents the modulus of the set of k nearest neighbor data points of the i-th data point p i .
[0168] Based on the same inventive concept, an embodiment of the present application further provides an electronic device 600. Referring to Figure 9 shown, the electronic device 600 is used to implement the streaming data anomaly detection method described in the above method embodiment. The electronic device 600 in this embodiment may include: a memory 601, a processor 602, and a computer program stored in the memory and executable on the processor, such as a streaming data anomaly detection program. When the processor executes the computer program, the steps in the above various streaming data anomaly detection method embodiments are implemented.
[0169] In the embodiment of the present application, the specific connection medium between the above-mentioned memory 601 and the processor 602 is not limited. In the embodiment of the present application Figure 9 , the memory 601 and the processor 602 are connected through a bus 603. The bus 603 is represented by a thick line in Figure 9 . The connection manners between other components are only for illustrative purposes and are not to be construed as limiting. The bus 603 may be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 9 only one thick line is used to represent it in
[0170] , but it does not mean that there is only one bus or one type of bus.The memory 601 can be a volatile memory, such as a random-access memory (RAM); the memory 601 can also be a non-volatile memory, such as a read-only memory, a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD), or the memory 601 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 601 can be a combination of the above memories.
[0171] The processor 602 is used to implement the streaming data anomaly detection method provided by the embodiments of the present application.
[0172] The embodiments of the present application also provide a computer-readable storage medium, storing computer-executable instructions required to be executed by the above processor, which includes a program required to be executed by the above processor.
[0173] In some possible implementation manners, various aspects of the streaming data anomaly detection method provided by the present application can also be implemented in the form of a program product, which includes program code. When the program product runs on an electronic device, the program code is used to cause the electronic device to execute the steps in the streaming data anomaly detection method according to various exemplary embodiments of the present application described above in this specification.
[0174] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a device, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.
[0175] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (apparatus), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate for implementation in the process Figure 1one or more processes and / or blocks Figure 1 means for the functions specified in one or more blocks
[0176] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including an instruction means that implements the functions in the process Figure 1 one or more processes and / or blocks Figure 1 specified in one or more blocks
[0177] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions in the process Figure 1 one or more processes and / or blocks Figure 1 specified in one or more blocks
[0178] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present application
[0179] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations
Claims
1. A method for detecting anomalies in streaming data, characterized in that, including: In response to a data anomaly detection request sent by a client, obtain a data stream to be detected from a data source, where the data stream to be detected includes data points generated in chronological order; Store each obtained data point into a preset sliding window in chronological order; When the data volume in the sliding window reaches a first preset ratio, determine the local outlier factor of each data point according to each data point in the current sliding window and its respective k-nearest neighbor data points; Determine the data points with local outlier factors greater than the outlier factor threshold as abnormal data points, mark the abnormal data points, return the marked abnormal data points to the client, and then delete them from the sliding window; Sample data points with a second preset ratio from the remaining data points according to the local outlier factors of the remaining data points in the sliding window, and merge the sampled data points with the data points newly arriving at the sliding window, where the second preset ratio is less than one-half of the sliding window; When the data volume in the sliding window reaches the capacity limit, re-determine the local outlier factor of each data point in the sliding window that has slid out, and perform the steps of marking abnormal data points, sampling data points, and merging data points.
2. The method according to claim 1, wherein Determine the local outlier factor of each data point according to each data point in the current sliding window and its respective k-nearest neighbor data points, specifically including: Calculate the reachability distance between each data point in the current sliding window and its k-nearest neighbor data points respectively; Determine the local reachability density of each data point according to the reachability distance between each data point and its k-nearest neighbor data points; Determine the local outlier factor of each data point according to the local reachability density of each data point and its respective k-nearest neighbor data points respectively.
3. The method according to claim 1, characterized in that, Sample data points with a second preset ratio from the remaining data points according to the local outlier factors of the remaining data points in the sliding window, specifically including: Arrange the remaining data points in the sliding window in descending order of local outlier factor; Divide the remaining data points in the rearranged sliding window into a first preset number of intervals; Sample data points with a third preset ratio from the first second preset number of intervals, and sample data points with a fourth preset ratio from the last third preset number of intervals, where the third preset ratio is greater than the fourth preset ratio.
4. The method according to claim 2, wherein Calculate the reachability distance between each data point in the current sliding window and its k-nearest neighbor data points respectively, specifically including: Calculate the reachability distance between each data point in the current sliding window and its k-nearest neighbor data points through the following formula: reach-dist(p i ,p j ) = max(dist(p i ,p j ), k-distance(p j )) Among them, reach-dist(p i , p j ) represents the reachable distance between the $i$-th data point $p$ i in the current sliding window and its $j$-th nearest neighbor data point $p$ j . $p$ j ∈N k (p i ), where N k (p i ) represents the set of $k$ nearest neighbor data points of the $i$-th data point $p$ i . $p$ j represents the $j$-th nearest neighbor data point of the $i$-th data point $p$ i ; dist(p i , p j ) represents the distance between the i-th data point p i and its j-th nearest neighbor data point p j ; k-distance(p j ) represents the distance from the data point p j to its k-th nearest neighbor data point.
5. The method according to claim 4, characterized in that, Determine the local reachability density of each data point according to the reachability distance between each data point and its k-nearest neighbor data points, specifically including: Calculate the local reachability density of each data point by calculating the reachability distance between each data point and its k-nearest neighbor data points through the following formula: where LRD(p i ) represents the local reachability density of the i-th data point p i in the currently described sliding window.
6. The method according to claim 5, wherein Determine the local outlier factor of each data point according to the local reachability density of each data point and its respective k-nearest neighbor data points respectively, specifically including: Calculate the local outlier factor of each data point through the following formula: where LOF(p i ) represents the local outlier factor of the i-th data point p i in the currently described sliding window; LRD(p i ) represents the local reachability density of the i-th data point p i ; LRD(p j ) represents the local reachability density of the j-th nearest neighbor data point p i of the i-th data point p j ; |N k (p i )| represents the modulus of the set of the k nearest neighbor data points of the i-th data point p i .
7. A streaming data anomaly detection device, characterized in that, including: An acquisition module, configured to obtain a data stream to be detected from a data source in response to a data anomaly detection request sent by a client, where the data stream to be detected includes data points generated in chronological order; A caching module, configured to sequentially store each obtained data point into a preset sliding window in chronological order; A determination module, configured to, when the amount of data in the sliding window reaches a first preset ratio, respectively determine the local outlier factor of each data point according to each data point in the current sliding window and its k-nearest neighbor data points; An anomaly recognition module, configured to determine the data points with local outlier factors greater than the outlier factor threshold as anomaly data points, mark the anomaly data points, return the marked anomaly data points to the client, and then delete them from the sliding window; A sampling and merging module, configured to sample data points at a second preset ratio from the remaining data points according to the local outlier factors of the remaining data points in the sliding window, and merge the sampled data points with the data points newly arriving at the sliding window, where the second preset ratio is less than one half of the sliding window; A processing module, configured to, when the amount of data in the sliding window reaches the capacity limit, re-determine the local outlier factor of each data point in the sliding window that has slid out, and perform the steps of marking anomaly data points, sampling and merging data points; 8. The apparatus according to claim 7, wherein the determination module is specifically configured to calculate the reachable distance between each data point in the current sliding window and its k-nearest neighbor data points respectively; determine the local reachability density of each data point according to the reachable distance between each data point and its k-nearest neighbor data points; respectively determine the local outlier factor of each data point according to the local reachability density of each data point and its k-nearest neighbor data points.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein When the processor executes the program, it implements the streaming data anomaly detection method according to any one of claims 1 to 6.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps in the streaming data anomaly detection method according to any one of claims 1 to 6.