Edge service anomaly detection method based on concept drift
By using the reservoir sampling method based on random pairing and the distribution change index DCM calculated by the degree of difference in edge service anomaly detection, the uniformity and representativeness of the sampling data are solved, the sensitivity and accuracy of the detection are improved, and more accurate concept drift anomaly detection is achieved.
Patent Information
- Application Number
- CN202510150250.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-11
AI Technical Summary
The existing edge service concept drift anomaly detection method is difficult to meet uniformity and representativeness in the sampling data, and at the same time, there is room for improvement in detection sensitivity and accuracy of the measurement method.
The reservoir sampling method based on random pairing is used, and the distribution change index DCM is calculated in combination with the degree of difference, which is used to extract representative historical sample library features from historical data to perform more accurate detection of drift anomalies in edge service concepts.
It realizes the extraction of representative historical sample library features under storage cost constraints, improves detection sensitivity and accuracy, and solves the uniformity and representative problems of sampling data.
Smart Images

Figure CN119989231A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of edge services, and in particular relates to an edge service anomaly detection method based on concept drift. Background Art
[0002] In order to deal with concept drift, researchers have proposed a variety of detection methods, including statistics-based drift detection methods, prediction-based drift detection methods, sliding window-based drift detection methods, and clustering-based drift detection methods.
[0003] Statistical drift detection methods usually use statistical indicators to monitor changes in data streams, such as mean, variance, or more complex statistical tests. Baidari et al. developed a novel concept drift detection mechanism that relies on Bhattacharyya distance, supplemented by the calculation of mean and standard deviation to identify gradual or sudden changes in data distribution. The essence of this method is that it evaluates the degree of change in data distribution by measuring the center position and dispersion of consecutive data blocks and comparing them with the data of historical reference windows. When the observed changes in mean and standard deviation exceed the established limits, the system will diagnose the occurrence of concept drift. Statistical drift detection methods usually rely on simple statistics such as mean and variance, so they are not sensitive enough to changes in data distribution, especially when there are nonlinear or complex patterns in the data. In addition, such methods may be very sensitive to noise, resulting in erroneous drift detection. Prediction-based drift detection methods focus on training prediction models using historical data and monitoring changes in model prediction accuracy. When the prediction performance decreases, this indicates that concept drift has occurred in the data stream. Yang designed an integrated prediction framework, called Performance Weighted Probability Average Ensemble (PWPAE), to predict and identify concept drift. The framework integrates multiple anomaly prediction models and achieves adaptive adjustment to drift by considering the performance weights of each model and its prediction probability. This method relies on the performance of the prediction model. If the model itself is overfitted or underfitted (such as insufficient training data, improper feature selection, etc.), it may lead to inaccurate drift detection. In addition, such methods usually require a long period of historical data, so when the data flow is large, there may be lags, affecting timeliness. The drift detection method based on sliding windows captures the latest data trends by using sliding windows on the data stream, and compares the statistical characteristics of the data in the window with the statistical characteristics of the historical data to find the shift in the distribution. Liu developed a novel concept drift detection framework, Nearest Neighbor Density Change Identification (NN-DVI), which is applied to the drift detection task under the sliding window mechanism. NN-DVI uses a k-nearest neighbor-based spatial partitioning pattern (NNPS) to transform discrete and difficult-to-measure data instances into a set of shared subspaces, then accumulates density differences in these subspaces and quantifies overall differences to determine the concept drift confidence interval. The sliding window method needs to define the window size when processing data streams. If the window is too large, it may not be able to capture rapidly changing drifts in time. If the window is too small, it may lead to insufficient information and increase the inaccuracy of detection. In addition, the sliding window method usually assumes that the data distribution is locally stable, which may not be true in practical applications. Clustering-based drift detection methods analyze the intrinsic structural changes of data by clustering data points. When the clustering structure changes significantly, it indicates that concept drift has occurred. Jain et al. proposed a concept drift detection method that combines error rate and data distribution information.Their method analyzes drift based on captured data of sliding windows and is combined with the K-Means clustering algorithm to reduce the data size and update the training dataset. In addition, they also adopted a support vector machine (SVM) classifier for anomaly detection and triggered the retraining of the model after confirming that drift has occurred. Clustering methods rely on the selection of clustering algorithms and parameter adjustment, such as the selection of K value in K-Means. Different clustering methods may have large differences in adaptability to data and are prone to unstable results. Especially for high-dimensional data, the clustering effect may not be ideal, resulting in inaccurate detected drift information. In addition, the computational complexity of clustering algorithms is high, and performance bottlenecks may be faced when processing large-scale data.
[0004] In edge computing environments, the problem of concept drift is particularly prominent. Although research in this area is relatively limited, some researchers have begun to explore this challenge. Wang et al. developed an advanced detection mechanism called A-Detection, which integrates reservoir sampling and singular value decomposition (SVD) techniques for efficient sampling and feature extraction of large-scale data streams. They used Jensen-Shannon (JS) divergence and created a new metric method, FDC (Fractional Distribution Change), to measure the inconsistency of data stream distribution. FDC provides a new tool for real-time monitoring of edge service operation anomalies by quantifying the distribution changes of edge service reliability data streams. Although A-Detection has shown its advancedness in detecting concept drift, it does not consider the representativeness of samples in the design of the sampling algorithm, which to some extent affects the accuracy of the detection mechanism, and the FDC metric cannot accurately detect anomalies under specific drift patterns. Wang et al. further introduced two methods, B-Detection and L-Detection. B-Detection uses LSTM autoencoders to detect anomalies in the reliability data of MEC services. It identifies data points that are significantly different from the normal pattern by learning the distribution characteristics of normal data and comparing the reconstruction errors. B-Detection also incorporates weighted reservoir sampling and boosting strategies to adapt to the dynamic changes in data flow distribution in MEC environments. L-Detection combines sliding window technology with a sampling method based on locality sensitive hashing (LSH) to extract features from the latest and historical data, and uses JS divergence to evaluate the reliability anomaly of edge services during runtime. Both B-Detection and L-Detection methods take into account the representativeness of the sampling algorithm, but their improvements are mainly achieved through weight allocation. Although this method improves the representativeness of the sampled data to a certain extent, the weight allocation is very sensitive to the choice of initial weights, which is difficult to achieve without sufficient prior knowledge. Summary of the invention
[0005] In view of the above-mentioned deficiencies in the prior art, the present invention provides an edge service anomaly detection method based on concept drift, which solves the problem that the sampling data of the existing edge service concept drift anomaly detection method cannot meet the uniformity and representativeness, and the current measurement method also has room for improvement in detection sensitivity and accuracy. It realizes the extraction of representative historical sample library features from historical data under the constraint of storage cost, and performs more accurate edge service concept drift anomaly detection.
[0006] In order to achieve the above-mentioned object of the invention, the technical solution adopted by the present invention is: a method for detecting anomalies of edge services based on concept drift, comprising the following steps:
[0007] S1. Calculate the reliability score of the data stream and obtain the data stream sequence;
[0008] S2, sampling the data stream sequence by a reservoir sampling method based on random pairing, and storing the sampling results in a sample library;
[0009] S3, calculate the distribution change index of historical data in the sample library and current data in the sliding window;
[0010] S4. Perform concept drift anomaly detection based on distribution change indicators.
[0011] Further: In S1, the reliability score of the data stream is specifically:
[0012] The probability that the edge server will respond incorrectly, partially fail to respond, or time out when processing a user request r Δt , its specific expression is:
[0013] r Δt =e - λ ×len(Δt)
[0014] In the formula, Δt is a fixed time period, len(.) is the length of the fixed time period interval Δt, and λ is the failure rate of the edge service within the maximum response time. The specific expression is:
[0015] λ=1-[1-f(Δt)] n
[0016] Where n is the number of call requests, and f(Δt) is the probability of failure response of the edge server within a limited time. The specific expression is:
[0017] f(Δt)=L0 / n
[0018] Where L0 is the number of failed calls in n call requests.
[0019] Further: S2 comprises the following sub-steps:
[0020] S21, initializing a deletion counter, a prohibition counter and a total exclusion counter;
[0021] S22, sampling the elements in the data stream sequence in sequence through a sliding window, and calculating the difference between the sampled elements and the first sequence;
[0022] S23, judging whether the calculated difference is greater than the prohibition threshold, if so, allowing the element to enter the sample library, and entering S24; if not, prohibiting the element from entering the sample library, and entering S26;
[0023] S24, determine whether the total exclusion counter is 0, if so, execute the ordinary reservoir sampling method and enter S27; if not, enter S25;
[0024] S25. Generate a random decimal by random pairing, calculate the probability of new data being added to the sample library according to the values of the deletion counter and the prohibition counter, and determine whether the probability is greater than the random decimal; if so, add the element to the sample library and reduce the value of the deletion counter by 1; if not, prohibit the element from entering the sample library and reduce the value of the prohibition counter by 1;
[0025] Enter S27;
[0026] S26, prohibit the sampled element from entering the sample library, increase the value of the prohibition counter by 1, traverse the sample library to calculate the Euclidean distance between the sampled element and each element in the sample library, if there is an element in the sample library whose Euclidean distance is less than the deletion threshold, then delete the element whose Euclidean distance is less than the deletion threshold from the library, and increase the value of the deletion counter by 1, and enter S27;
[0027] S27, determine whether all elements in the data stream sequence have been sampled, if so, obtain the final sample library, if not, return to S21 to sample the next element in the data stream sequence.
[0028] Further: in said S21, the deletion counter records the number of deleted elements in the sample, the prohibition counter records the number of prohibited sampling elements in the window, and the total exclusion counter records the total exclusion number, whose value is the sum of the deletion counter and the prohibition counter values.
[0029] Further: in S22, for the element D[i] at position i in the sampled data stream sequence D, the expression of the difference z between it and the first sequence D[i+1:i-1+W] is specifically:
[0030] z=Z(D[i],D[i+1:i-1+W])
[0031] Where W is the sliding window size and Z(.) is the Z-Score formula.
[0032] Further: In S25, the expression of the probability P1 of new data being added to the sample library is specifically:
[0033] P1=C1 / (C1+C2)
[0034] Where C1 is the value of the delete counter and C2 is the value of the disable counter.
[0035] Further: In S3, the expression of the distribution change index DCM is specifically:
[0036]
[0037] Where D history is the historical data in the sample library, D current is the current data in the sliding window, P current is the probability distribution of the current data, P history is the probability distribution of historical data, γ is the set of all possible joint distributions, Π(P history ,P current ) for all history and P current The set of joint distributions γ conditioned on marginal distributions, (x, y) represents sample pairs drawn from the joint distribution γ, where x comes from the probability distribution of historical data and y comes from the probability distribution of current data, ||xy|| represents the distance between sample pairs (x, y), and Δ(·) is the difference measure between the two distributions.
[0038] Further: S4 is specifically:
[0039] Determine whether the distribution change index is greater than the abnormal threshold. If so, an abnormality occurs; if not, no abnormality occurs.
[0040] The beneficial effects of the present invention are as follows: the present invention provides an edge service anomaly detection method based on concept drift, which combines the difference and the distribution change index of the reservoir sampling algorithm based on random pairing, solves the technical problem that the sampling data of the existing edge service concept drift anomaly detection method cannot meet the uniformity and representativeness, and also solves the problem that the current measurement method has room for improvement in detection sensitivity and accuracy, and realizes the extraction of representative historical sample library features from historical data under the constraint of storage cost, so as to perform more accurate edge service concept drift anomaly detection, and has the following effects compared with the existing technology:
[0041] (1) The present invention effectively utilizes the sliding window and improves the traditional reservoir sampling algorithm by introducing the difference degree, thereby realizing the extraction of representative historical sample library features from historical data under the constraint of storage cost.
[0042] (2) The present invention introduces the concept of random pairing to achieve the extraction of uniform historical sample library features from historical data.
[0043] (3) In terms of measurement value, the present invention innovatively proposes the data flow distribution change index DCM as an indicator for judging anomalies. By calculating the minimum cost required for the probability distribution conversion between the current reliability data flow and the historical reliability data flow in the service environment, more accurate concept drift anomaly detection is achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 This is a flow chart of an edge service anomaly detection method based on concept drift of the present invention. DETAILED DESCRIPTION
[0045] The specific implementation modes of the present invention are described below so that those skilled in the art can understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific implementation modes. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the attached claims, these changes are obvious, and all inventions and creations utilizing the concept of the present invention are protected.
[0046] like Figure 1 As shown, in one embodiment of the present invention, a method for detecting anomalies of edge services based on concept drift includes the following steps:
[0047] S1. Calculate the reliability score of the data stream and obtain the data stream sequence;
[0048] S2, sampling the data stream sequence by a reservoir sampling method based on random pairing, and storing the sampling results in a sample library;
[0049] S3, calculate the distribution change index of historical data in the sample library and current data in the sliding window;
[0050] S4. Perform concept drift anomaly detection based on distribution change indicators.
[0051] In this embodiment, the present invention adds the concept of difference based on the reservoir sampling algorithm, and uses the difference to determine whether an element is allowed to be sampled into the sample library, so as to ensure that representative data can be stored in a fixed-size sample library. In order to ensure that the uniformity of the sampled data can be met after adding the difference to the sampling method, the concept of random pairing is introduced. After obtaining more uniform and more representative sampled data, the appropriate distribution change index DCM is selected based on the bulldozer distance idea to achieve more accurate concept drift anomaly detection.
[0052] The present invention calculates the reliability score of the data flow to quantify the reliability of the data flow. The present invention defines it as the probability that the edge server successfully processes the user request and avoids the occurrence of error response, partial response failure or timeout under specific operating conditions and time constraints. Δt express.
[0053] In S1, the reliability score of the data flow is specifically:
[0054] The probability that the edge server will respond incorrectly, partially fail to respond, or time out when processing a user request r Δt , its specific expression is:
[0055] r Δt =e - λ ×len(Δt)
[0056] In the formula, Δt is a fixed time period, len(.) is the length of the fixed time period interval Δt, and λ is the failure rate of the edge service within the maximum response time. The specific expression is:
[0057] λ=1-[1-f(Δt)] n
[0058] Where n is the number of call requests, and f(Δt) is the probability of failure response of the edge server within a limited time. The specific expression is:
[0059] f(Δt)=L0 / n
[0060] Where L0 is the number of failed calls in n call requests.
[0061] In this embodiment, within a fixed time period Δt, a call request is sent to the edge server at fixed intervals, and a total of n call requests are executed. Each call request is regarded as a Bernoulli test, in which only when the request is within the longest time T max The test is considered successful only when a complete response is obtained within n consecutive Bernoulli trials; otherwise, it is considered a failure. Through n consecutive Bernoulli trials, a rate value can be obtained to estimate the ability of the edge server to perform the required function within a fixed time period Δt. The rate value can be expressed as [1-f(Δt)] n .
[0062] The S2 comprises the following sub-steps:
[0063] S21, initializing a deletion counter, a prohibition counter and a total exclusion counter;
[0064] S22, sampling the elements in the data stream sequence in sequence through a sliding window, and calculating the difference between the sampled elements and the first sequence;
[0065] S23, judging whether the calculated difference is greater than the prohibition threshold, if so, allowing the element to enter the sample library, and entering S24; if not, prohibiting the element from entering the sample library, and entering S26;
[0066] S24, determine whether the total exclusion counter is 0, if so, execute the ordinary reservoir sampling method and enter S27; if not, enter S25;
[0067] In this embodiment, the common reservoir sampling algorithm is specifically as follows: a sample library of size k is initialized during the sliding window sliding process. The first k data in the sliding window are directly stored in the sample library. For the i-th (i>k) data, the probability k / i is used to determine whether to replace the data in the reservoir. If replaced, a data in the reservoir is randomly selected for replacement. This ensures that the probability of each data point being selected is equal.
[0068] S25. Generate a random decimal by random pairing, calculate the probability of new data being added to the sample library according to the values of the deletion counter and the prohibition counter, and determine whether the probability is greater than the random decimal; if so, add the element to the sample library and reduce the value of the deletion counter by 1; if not, prohibit the element from entering the sample library and reduce the value of the prohibition counter by 1;
[0069] Enter S27;
[0070] S26, prohibit the sampled element from entering the sample library, increase the value of the prohibition counter by 1, traverse the sample library to calculate the Euclidean distance between the sampled element and each element in the sample library, if there is an element in the sample library whose Euclidean distance is less than the deletion threshold, then delete the element whose Euclidean distance is less than the deletion threshold from the library, and increase the value of the deletion counter by 1, and enter S27;
[0071] S27, determine whether all elements in the data stream sequence have been sampled, if so, obtain the final sample library, if not, return to S21 to sample the next element in the data stream sequence.
[0072] In S21, the deletion counter records the number of deleted elements in the sample, the prohibition counter records the number of prohibited sampling elements in the window, and the total exclusion counter records the total exclusion number, whose value is the sum of the deletion counter and the prohibition counter.
[0073] In S22, for the element D[i] at position i in the sampled data stream sequence D, the expression of the difference z between it and the first sequence D[i+1:i-1+W] is specifically:
[0074] z=Z(D[i],D[i+1:i-1+W])
[0075] Where W is the sliding window size and Z(.) is the Z-Score formula.
[0076] In this embodiment, the sliding window slides on the data stream sequence, the window size is W, and the sliding step is 1.
[0077] In S25, the probability P1 of new data being added to the sample library is specifically expressed as:
[0078] P1=C1 / (C1+C2)
[0079] Where C1 is the value of the delete counter, C2 is the value of the disable counter, and the generated random decimal value range is [0,1).
[0080] In this embodiment, the present invention introduces the concept of random pairing. Pairing means that element A and element B form a "partnership". Operations on element A need to rely on certain operations of element B, that is, certain operations of element B are prerequisites for certain operations of element A. Random pairing means that when the current data D[i] is allowed to be sampled into the sample library, data D[i] randomly selects an element D[j] that is prohibited from sampling and has not been paired from the elements in the window D[0] to D[i-1] to form a pair. Element D[j] is called a pairing item. If an element in the sample library is removed when the pairing item D[j] is prohibited from sampling, the current data is added to the sample library; otherwise, the current data D[i] will not be added to the sample library even if it is allowed to be sampled.
[0081] In S3, the expression of the distribution change index DCM is specifically:
[0082]
[0083] Where D history is the historical data in the sample library, D current is the current data in the sliding window, P current is the probability distribution of the current data, P history is the probability distribution of historical data, γ is the set of all possible joint distributions, which describe the history To P current A transformation of Π(P history ,P current ) for all history and P current The set of joint distributions γ with marginal distribution conditions, (x, y) represents the sample pairs drawn from the joint distribution γ, where x comes from the probability distribution of historical data, y comes from the probability distribution of current data, ||xy|| represents the distance between the sample pairs (x, y), and the present invention adopts the Euclidean distance representation, Δ(·) is the difference measure between the two distributions, ∫||xy||dγ(x, y) represents the weighted sum of the distances between the sample pairs (x, y) under all possible joint distributions γ, which can be regarded as the weighted sum of the distances between the sample pairs (x, y) from P history Transform to P current The required “cost” is to convert P history Transformed to P currentThe minimum value of this weighted sum in all possible joint distributions is the minimum energy consumption under the optimal path planning, that is, the DCM value.
[0084] In this embodiment, the present invention uses the idea of Wasserstein distance to define the Distribution Change Metric (DCM). history and the current data D in the sliding window current The distribution difference is expressed as Δ(D history ,D current ), Δ represents the difference measure between the two distributions.
[0085] The S4 is specifically:
[0086] Determine whether the distribution change index is greater than the abnormal threshold. If so, an abnormality occurs; if not, no abnormality occurs.
[0087] The beneficial effects of the present invention are as follows: the present invention provides an edge service anomaly detection method based on concept drift, which combines the difference and the distribution change index of the reservoir sampling algorithm based on random pairing, solves the technical problem that the sampling data of the existing edge service concept drift anomaly detection method cannot meet the uniformity and representativeness, and also solves the problem that the current measurement method has room for improvement in detection sensitivity and accuracy, and realizes the extraction of representative historical sample library features from historical data under the constraint of storage cost, so as to perform more accurate edge service concept drift anomaly detection, and has the following effects compared with the existing technology:
[0088] (1) The present invention effectively utilizes the sliding window and improves the traditional reservoir sampling algorithm by introducing the difference degree, thereby realizing the extraction of representative historical sample library features from historical data under the constraint of storage cost.
[0089] (2) The present invention introduces the concept of random pairing to achieve the extraction of uniform historical sample library features from historical data.
[0090] (3) In terms of measurement value, the present invention innovatively proposes the data flow distribution change index DCM as an indicator for judging anomalies. By calculating the minimum cost required for the probability distribution conversion between the current reliability data flow and the historical reliability data flow in the service environment, more accurate concept drift anomaly detection is achieved.
[0091] In the description of the present invention, it is necessary to understand that the orientation or positional relationship indicated by the terms "center", "thickness", "upper", "lower", "horizontal", "top", "bottom", "inner", "outer", "radial", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention. In addition, the terms "first", "second", and "third" are used only for descriptive purposes, and cannot be understood as indicating or implying the relative importance or the number of implicitly specified technical features. Therefore, the features defined by "first", "second", and "third" may explicitly or implicitly include one or more of the features.
Claims
1. A method for detecting anomalies in edge services based on concept drift, characterized in that: The following steps are involved: S1. Calculate the reliability score of the data stream and obtain the data stream sequence; S2, sampling the data stream sequence by a reservoir sampling method based on random pairing, and storing the sampling results in a sample library; S3, calculate the distribution change index of historical data in the sample library and current data in the sliding window; S4. Perform concept drift anomaly detection based on distribution change indicators.
2. The edge service anomaly detection method based on concept drift according to claim 1 is characterized in that: In S1, the reliability score of the data flow is specifically: The probability that the edge server will respond incorrectly, partially fail to respond, or time out when processing a user request r Δt , its specific expression is: r Δt =e -λ×len(Δt) In the formula, Δt is a fixed time period, len(.) is the length of the fixed time period interval Δt, and λ is the failure rate of the edge service within the maximum response time. The specific expression is: λ=1-[1-f(Δt)] n Where n is the number of call requests, and f(Δt) is the probability of failure response of the edge server within a limited time. The specific expression is: f(Δt)=L0 / n Where L0 is the number of failed calls in n call requests.
3. The edge service anomaly detection method based on concept drift according to claim 1 is characterized in that: The S2 comprises the following sub-steps: S21, initializing a deletion counter, a prohibition counter and a total exclusion counter; S22, sampling the elements in the data stream sequence in sequence through a sliding window, and calculating the difference between the sampled elements and the first sequence; S23, judging whether the calculated difference is greater than the prohibition threshold, if so, allowing the element to enter the sample library, and entering S24; If not, the element is prohibited from entering the sample library and the process goes to S26; S24, determine whether the total exclusion counter is 0, if so, execute the ordinary reservoir sampling method and enter S27; if not, enter S25; S25, generate a random decimal by random pairing, calculate the probability of new data being added to the sample library according to the values of the deletion counter and the prohibition counter, and determine whether the probability is greater than the random decimal; if so, add the element to the sample library and reduce the value of the deletion counter by 1; If not, the element is prohibited from entering the sample library, and the value of the prohibition counter is reduced by 1; Enter S27; S26, prohibit the sampled element from entering the sample library, increase the value of the prohibition counter by 1, traverse the sample library to calculate the Euclidean distance between the sampled element and each element in the sample library, if there is an element in the sample library whose Euclidean distance is less than the deletion threshold, then delete the element whose Euclidean distance is less than the deletion threshold from the library, and increase the value of the deletion counter by 1, and enter S27; S27, determine whether all elements in the data stream sequence have been sampled, if so, obtain the final sample library, if not, return to S21 to sample the next element in the data stream sequence.
4. The edge service anomaly detection method based on concept drift according to claim 3 is characterized in that: In S21, the deletion counter records the number of deleted elements in the sample, the prohibition counter records the number of prohibited sampling elements in the window, and the total exclusion counter records the total exclusion number, whose value is the sum of the deletion counter and the prohibition counter.
5. The edge service anomaly detection method based on concept drift according to claim 3 is characterized in that: In S22, for the element D[i] at position i in the sampled data stream sequence D, the expression of the difference z between it and the first sequence D[i+1:i-1+W] is specifically: z=Z(D[i],D[i+1:i-1+W]) Where W is the sliding window size and Z(.) is the Z-Score formula.
6. The edge service anomaly detection method based on concept drift according to claim 3 is characterized in that: In S25, the probability P1 of new data being added to the sample library is specifically expressed as: P1=C1 / (C1+C2) Where C1 is the value of the delete counter and C2 is the value of the disable counter.
7. The edge service anomaly detection method based on concept drift according to claim 1 is characterized in that: In S3, the expression of the distribution change index DCM is specifically: Where D history is the historical data in the sample library, D current is the current data in the sliding window, P current is the probability distribution of the current data, P history is the probability distribution of historical data, γ is the set of all possible joint distributions, Π(P history ,P current ) for all history and P current The set of joint distributions γ conditioned on marginal distributions, (x, y) represents sample pairs drawn from the joint distribution γ, where x comes from the probability distribution of historical data and y comes from the probability distribution of current data, ||xy|| represents the distance between sample pairs (x, y), and Δ(·) is the difference measure between the two distributions.
8. The edge service anomaly detection method based on concept drift according to claim 1 is characterized in that: The S4 is specifically: Determine whether the distribution change index is greater than the abnormal threshold. If so, an abnormality occurs; if not, no abnormality occurs.
Citation Information
Patent Citations
Unbalanced data flow mining method based on active drift detection
CN112000705A
Concept drift detection method and system based on weighted sampling and electronic equipment
CN113033643A
Conceptual drift-oriented adaptive interpretable industrial control system anomaly detection method
CN116991137A
Cardinality estimation method and system based on dynamic sample recommendation
CN117909363A
Automated dataset drift detection
US20230139718A1