Novel abnormal data detection method
Bilateral anomaly detection is performed in high-dimensional datasets through the BikNN method, combined with Mahalanobis and weighted Minkowski evaluation, which solves the shortcomings of multi-cluster datasets and local outlier detection in existing technologies, realizes efficient anomaly detection and visualization, and improves detection accuracy.
Patent Information
- Application Number
- CN202510764258.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-09-19
AI Technical Summary
Existing outlier detection methods have shortcomings when dealing with high-dimensional datasets, especially multi-cluster datasets and local outlier detection, and it is difficult to provide effective interpretability and efficient detection performance.
The BikNN method is used to perform bilateral anomaly detection in the density domain and spatial domain through the k-nearest neighbor algorithm. The anomaly coordinate system is constructed in the two-dimensional coordinate system by combining the Mahalanobis anomaly evaluation and the weighted Minkowski anomaly evaluation. The ECOD and DIF methods are used for integrated anomaly detection and visualization.
We achieve high-performance anomaly detection on synthetic and real datasets, can classify and visualize anomalies, provide useful guidance on anomaly characteristics, and improve detection accuracy and robustness.
Smart Images

Figure CN120671044A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of data science and machine learning, and in particular to a method for detecting outliers in high-dimensional data. Background Art
[0002] Generally speaking, extreme values (outliers), or outliers, are small data points that exhibit characteristics that differ from those of normal instances. The existence of outliers is very common in data science and machine learning. Outliers often provide important information in various applications, such as detecting unauthorized network access and credit card fraud. Outlier detection is also used in medicine to monitor patients' vital functions. It is also used to detect faults in complex systems. Outliers can also significantly affect the performance of statistical models. Therefore, it is crucial to identify outliers to reveal new insights or to remove them from normal instances to maintain the performance of machine learning models.
[0003] Some typical existing anomaly detection methods, such as the minimum covariance determinant anomaly detector, provide robust estimators of the mean and covariance matrix of the observations. It seeks to minimize the impact of outliers by finding the subset of observations whose covariance matrix has the minimum determinant. However, as a linear model, it is not suitable for datasets with multiple clusters. Histogram-based outlier detectors, for example, assume that each dimension is independent and divide each dimension into a certain number of intervals. The anomaly score is estimated by aggregating the density of each interval. It focuses on effectively detecting global outliers but may perform poorly in detecting local outliers. Other methods include distance-based methods, methods based on binary space partitioning, probabilistic methods, ensemble-based methods, and neural network-based methods. These methods identify outliers in their own way, each with its own advantages and limitations. Summary of the Invention
[0004] To address the above issues, the present invention provides a novel abnormal data detection method, which aims to provide simple anomaly explainability. It can also achieve better results on certain data sets by integrating methods such as ECOD (Unsupervised Outlier Detection Using Empirical Cumulative Distribution Functions) and DIF (Deep Isolation Forest for Anomaly Detection), and even further classify and visualize anomalies.
[0005] The technical solution provided by the present invention includes:
[0006] Step a: Preprocess the given high-dimensional data set, estimate two unilateral anomalies in the density domain and the spatial domain using the k-nearest neighbor algorithm, and standardize them into a two-dimensional coordinate system.
[0007] First, the data points are converted from the original space to the ECDF space through the projection function to obtain the ECDF value corresponding to the data point in the ECDF space. The ECDF value of each data point is calculated through the Euclidean distance to obtain the density domain distance, and the spatial domain distance of each data point in the original space is calculated through the Euclidean distance.
[0008] Then, the distance K from the k nearest neighbor points in the density domain and the spatial domain is obtained respectively by the k-nearest neighbor algorithm. e and K p , and standardize it.
[0009] Step b: Use the two-dimensional space to establish a two-dimensional coordinate system, perform Mahalanobis anomaly evaluation and weighted Minkowski anomaly evaluation on the data points in the two-dimensional coordinate system, and obtain the estimated outlier values of the BikNN (ANOMALY ESTIMATION IN BILATERAL DOMAINSWITH K-NEAREST NEIGHBORS) method.
[0010] Two-sided outlier K across all data points e and K p , construct a global anomaly coordinate system that maps all data points into a two-dimensional space, where the horizontal axis represents density domain anomalies and the vertical axis represents spatial domain anomalies; in this space, each point x i It is represented by two ordered sets of anomalies: (K e (x i ),K p (x i )), and the two coordinate components will be combined to evaluate the anomaly score for each point.
[0011] Based on the constructed anomaly coordinate system, the anomaly score of each data point is estimated, which includes the Mahalanobis anomaly estimate and the weighted Minkowski anomaly estimate.
[0012] The Mahalanobis anomaly estimate M(x i ): For a given data point x i , estimate its projected outlier point v(x i ) to the Mahalanobis distance of the center of the dense point in the two-dimensional anomaly space.
[0013] The weighted Minkowski anomaly estimate W(x i): Introduce parameters [w1,w2] to control the importance of the two anomalies.
[0014] Estimated outliers using the BikNN method It is estimated by combining the Mahalanobis anomaly and the weighted Minkowski anomaly and is expressed as follows:
[0015]
[0016] The parameter μ∈[0,1] is used to balance the two anomalies.
[0017] Step c: Based on the estimated outlier value of the BikNN method, ECOD or DIF anomaly detection methods can be integrated on different data sets to obtain the final anomaly score. The anomaly detection results are obtained according to the anomaly threshold and visualized.
[0018] Beneficial effects of the present invention:
[0019] Experiments on synthetic and real-world datasets show that the proposed BikNN model, integrated with other models, performs well and achieves the highest average performance. It can also classify and visualize anomalies in data points on a two-dimensional plane. This visualization tool can provide useful guidance for further investigation into which features make certain points potential outliers. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the present invention, a brief description of the technology in the present invention will be given below using the accompanying drawings, which will give people in this field a clearer understanding of the performance of this technology.
[0021] Figure 1 This is a flow chart of a novel anomaly detection method technology in the present invention;
[0022] FIG2( a ) is a distribution diagram of the original data in the present invention;
[0023] FIG2( b ) is a graph showing the distribution of synthetic data in ECDF space in the present invention;
[0024] FIG2( c ) is a distribution diagram of the synthetic data in the abnormal space in the present invention;
[0025] FIG3( a ) is a Mahalanobis (Mahalanobis distance) anomaly graph in the present invention;
[0026] FIG3( b ) is a weighted Minkowski anomaly map of the present invention;
[0027] FIG3( c ) is a combined anomaly graph of Mahalanobis and weighted Minkowski in the present invention;
[0028] Figure 4(a) shows the bilateral classification of outliers in the outlier space;
[0029] Figure 4(b) is a scatter plot in the original space;
[0030] Figure 4(c) shows a synthetic data point graph with three clusters;
[0031] Figure 4(d) shows the bilateral classification map of outliers in the outlier space. DETAILED DESCRIPTION
[0032] In order to further explain the present invention in detail, specific examples are described below with reference to the accompanying drawings.
[0033] A new abnormal data detection method, such as Figure 1 As shown, the following process is included:
[0034] First, each data point in the original graph dataset is represented as a density vector. The density vector calculation method can be selected based on the specific application scenario and data characteristics, such as kernel density estimation. By converting data points into density vectors, the present invention can better capture the inherent structure and distribution of the data.
[0035] k-nearest neighbor search:
[0036] For each data point, an appropriate distance metric (such as Euclidean distance) is used to search for the k nearest neighbors in its neighborhood to measure its sparsity. Points with high sparsity values can be considered very different from the surrounding points and can therefore be considered outliers. The value of k can be adjusted according to the specific application to balance computational complexity and accuracy of outlier detection.
[0037] Density domain and space domain distance calculation:
[0038] Anomalies in data points are often related to the data distribution, so it's important to consider the data distribution when estimating the distance between data points. On the one hand, closer points in Euclidean space are more likely to be similar. On the other hand, the greater the density between two points, the less similar they are. A greater density indicates a greater number of other points between the two points. Therefore, we introduce the density domain distance between points for anomaly estimation. The density domain distance concept is defined as:
[0039] d(x1,x2)=F X (x2)-F X (x1) = P(x1 < X ≤ x2)
[0040] Among them F X(x) = P(X≤x) is the cumulative distribution function (CDF) of X, and P(X≤x) represents the probability that the random variable X takes a value less than or equal to x.
[0041] In order to define the density domain distance of a dataset, the empirical cumulative distribution function ECDF of a multivariate dataset is calculated, which is an estimate of the underlying CDF of the generated samples. is a d-dimensional dataset with n observations, is the i-th observation in the j-th dimension. The j-th ECDF is given by It is defined as:
[0042]
[0043] where I(·) is an indicator function with two possible values: if x i 1 if ≤x is true, 0 otherwise. The result is a step function that increases by 1 / n at each data point.
[0044] Assume that the features of the data are independent of each other. Given two data points and The density domain distance between them is specifically defined as:
[0045]
[0046] where, for j = 1, 2, ..., d, the vector The jth entry of p is the p norm of the vector. First, the data points are transformed from the original space to the ECDF space through the projection function.
[0047] P(x i )= <P1(X≤x i,1 ),P2(X≤x i,2 ),...,P d (X≤x i,d )
[0048] The density domain distance is then calculated in the ECDF space in the same manner as the Euclidean distance, and the spatial domain distance is calculated in the original space. The original space and ECDF space of the synthetic dataset are shown in Figures 2(a) and 2(b). As can be seen from the figures, the two clusters with different densities in the original space are transformed into clusters with similar densities in the ECDF space. Furthermore, the two points are clearly isolated in the ECDF space, which facilitates outlier detection.
[0049] Bilateral outlier extraction:
[0050] For each data point, the bilateral outlier value between it and its k-nearest neighbor set is calculated. The calculation method of the bilateral outlier value can adopt existing technologies, such as distance-based measurement methods or density-based measurement methods. In the present invention, a distance-based measurement method is used to calculate the distance between each data point and its k-nearest neighbor in the spatial domain and the density domain, and combine them to obtain the bilateral outlier value. First, calculate x i The distance K to its kth nearest neighbor e (x i ),
[0051]
[0052] where N k (i) represents x i The k nearest neighbor points of D1(x i ,x j ) is the Minkowski distance between two points in the original space. By default, the Euclidean distance is used. The function max is optional, and other functions such as the mean or median of the k nearest neighbor points can also be used. This distance can be used alone to estimate the anomaly score, which is called kNN (K-NEAREST NEIGHBORS) spatial anomaly. Then from x i To the k nearest neighbor points N k (i) Maximum distance K p (x i ) is defined as,
[0053]
[0054] Here, P(·) is the projection function that transforms the data point from the original space to the ECDF space, and D2(·,·) is the Minkowski distance between two points in the ECDF space. The neighborhood of the point in the original space is used instead of the k nearest neighbors in the ECDF space. This distance can also be used alone to estimate anomaly scores, which is called kNN probability density anomaly.
[0055] Construct a two-dimensional coordinate system:
[0056] By collecting the bilateral outliers of all data points as described above, a global outlier coordinate system can be constructed. This coordinate system maps all data points into a two-dimensional space, where the horizontal axis represents the kNN probability density anomaly and the vertical axis represents the anomaly in the kNN spatial domain. In this space, each point is represented by an ordered set of two anomalies: (K e (x i ),K p (x i), and the two coordinate components are combined to evaluate the anomaly score for each point. This allows for intuitive observation of the distribution and degree of anomaly of the data points. This is shown in Figure 2(c). As can be seen from the figure, the two coordinate components are not completely on the same diagonal line, indicating some inconsistency between the two unilateral anomalies. However, with the exception of a few points at the periphery, most points are very close. The correlation between the two anomalies also needs to be considered.
[0057] Calculate the anomaly score:
[0058] Based on the constructed two-dimensional coordinate system, this paper proposes two schemes to jointly estimate the anomaly score of each data point in the two-dimensional anomaly space. The score reflects the degree of anomaly of the point in the bilateral domain. One scheme is Mahalanobis anomaly estimation. Mahalanobis distance is a common metric that can capture the non-isotropic characteristics of the feature space. For a given point x i , estimate its projected outlier point v(x i ) to the Mahalanobis distance of the center of the dense point in the two-dimensional anomaly space.
[0059]
[0060] where v(x i )=[K e (x i ),K p (x i )] T , is the center of the outlier, and ∑ is the covariance matrix estimated based on the outlier. The Mahalanobis anomaly plot of the experimental dataset is shown in Figure 3(a). As shown in the figure, in this case, the anomaly score increases faster along the density axis. This does not mean that density anomalies are more important than spatial anomalies, but rather that the principal axis depends on the specific data. Another solution is weighted Minkowski anomaly estimation. The weighted Minkowski anomaly estimation is as follows:
[0061]
[0062] in p is the p-norm of the vector. Unlike the Mahalanobis anomaly, this scheme ignores the distribution of anomalies. The weighted Minkowski anomaly plot of the synthetic data is shown in Figure 3(b). Different functions or norms can be used to define the anomaly score. In the current implementation, this scheme is used for two main reasons. On the one hand, points in the upper right region of the anomaly space tend to have larger anomalies, which the Mahalanobis scheme does not take into account. The introduction of this scheme compensates for points in this region. On the other hand, the parameters [w1, w2] are introduced to control the importance of the two anomalies and make the method include the traditional kNN method.
[0063] Example x i Anomaly score It is estimated by combining the Mahalanobis anomaly and the weighted Minkowski anomaly and is expressed as follows,
[0064]
[0065] The parameter μ∈[0,1] is used to balance these two anomalies. The main parameters that need to be adjusted are w1, w2, and μ. When the parameters w1=1, w2=0, and μ=1 are set, the combination becomes the traditional kNN method. The combined anomaly of the synthetic data is shown in Figure 3(c).
[0066] The anomaly scores obtained by the above BikNN method can be integrated with different anomaly detection methods on different data sets, such as the anomaly scores obtained by the ECOD method or the DIF method. i The abnormal score S(x i ) by combining BikNN anomaly scores and anomaly scores of different methods To estimate, it is expressed as follows,
[0067]
[0068] Where a1 and a2 are obtained through Bayesian optimization.
[0069] Visualization and classification:
[0070] Finally, the calculated anomaly scores are visualized, and the data points are classified as normal or abnormal according to the threshold. In the present invention, the anomaly points can be visualized by drawing ordered anomaly points on a two-dimensional plane, and then the anomaly points are manually marked through simple interaction. Specifically, the outliers are first extracted according to the threshold in the spatial anomaly coordinates and the density anomaly coordinates, respectively, and the two vertical lines divide the plane into four regions. In addition to the area where the normal data points are located, the points in the remaining three areas can be simply divided into three types. Type I points are points with high spatial anomalies and high density anomalies, Type II points are points with low density anomalies and high spatial anomalies, and Type III points are points with high density anomalies and low spatial anomalies. Figure 4(a) to Figure 4(d) As shown, Figure 4(a) is a bilateral classification diagram of outliers in the anomaly space; Figure 4(b) is a scatter plot in the original space; Figure 4(c) is a synthetic data point diagram with three clusters; Figure 4(d) is a bilateral classification diagram of outliers in the anomaly space.
[0071] This threshold can be adjusted according to the actual application to balance the false positive rate and the false negative rate. Through visualization and classification operations, we can better understand the anomalies in the dataset and provide a basis for subsequent processing.
[0072] In summary, the technical features of this invention are: First, it detects anomalies in both spatial and density domains and establishes a two-dimensional anomaly coordinate system. Experiments on synthetic and real datasets demonstrate that this method performs well and achieves the highest average performance. Second, this method can also classify and visualize anomalies in data points on a two-dimensional plane. This visualization tool can provide useful information for further studying the characteristics that make certain points potential outliers.
[0073] Table 1 Method performance (anomaly detection accuracy)
[0074]
[0075] Table 2 Method performance (ROC-AUC score)
[0076]
[0077] Table 1 compares the anomaly detection accuracy of the ECOD, DIF, and BikNN ensembles. Table 2 compares the anomaly detection ROC-AUC (Receiver Operating Characteristic-Area Under the Curve) scores of the ECOD, DIF, and BikNN ensembles, which measure the model's ability to distinguish between positive and negative samples at different thresholds. BikNN performs best on the letter dataset, BikNN-ECOD ensemble performs best on the arrhythmia and cardio datasets, and BikNN-DIF ensemble performs best on the satimage-2 and wbc datasets. BikNN alone is unstable, performing well on some datasets, such as the letter dataset, but underperforming the ensemble on others. On complex datasets, such as satimage-2, performance improvements exceed 200%. If the baseline ECOD or DIF method performs well, BikNN can be integrated with that method to achieve better performance. If performance is poor, BikNN alone can be used, such as on the letter dataset. In addition, when the data dimension is low and the anomaly ratio is moderate, the single BikNN is the best, and it is necessary to rely on integration in high-dimensional data.
Claims
1. A new abnormal data detection method, characterized in that: The following steps are involved: Step 1: Preprocess the data set, estimate two unilateral anomalies in the density domain and spatial domain using the k-nearest neighbor algorithm, and standardize them into a two-dimensional space; Step 2: Use the two-dimensional space to establish a two-dimensional coordinate system, perform Mahalanobis anomaly evaluation and weighted Minkowski anomaly evaluation on the data points in the two-dimensional coordinate system, and obtain the estimated outlier value of the BikNN method; Step 3: Based on the estimated outlier value of the BikNN method, the ECOD anomaly detection method is integrated to obtain the final anomaly score. The anomaly detection result is obtained according to the anomaly threshold and visualized.
2. A novel abnormal data detection method according to claim 1, characterized in that: The specific implementation process of step 1 is: First, the data points are converted from the original space to the ECDF space through the projection function to obtain the ECDF value corresponding to the data point in the ECDF space. The ECDF value of each data point is calculated by the Euclidean distance to obtain the density domain distance, and the spatial domain distance of each data point in the original space is calculated by the Euclidean distance. Then, the distance K from the k nearest neighbor points in the density domain and the spatial domain is obtained respectively by the k-nearest neighbor algorithm. e and K p , and standardize it.
3. A novel abnormal data detection method according to claim 2, characterized in that: The specific process of establishing the two-dimensional coordinate system is as follows: Two-sided outlier K across all data points e and K p ,construct a global anomaly coordinate system that maps all data points into a two-dimensional space, where the horizontal axis represents density domain anomalies and the vertical axis represents space domain anomalies; In this space, every point x i It is represented by two ordered sets of anomalies: (K e (x i ),K p (x i ), and the two coordinate components will be combined to evaluate the anomaly score for each point.
4. A novel abnormal data detection method according to claim 3, characterized in that: The specific implementation process of obtaining the estimated outlier value of the BikNN method is as follows: Based on the constructed anomaly coordinate system, the anomaly score of each data point is estimated. The anomaly score includes the Mahalanobis anomaly estimate and the weighted Minkowski anomaly estimate. The Mahalanobis anomaly estimate M(x i ): For a given data point x i , estimate its projected outlier point v(x i ) to the Mahalanobis distance of the center of the dense point in the two-dimensional anomaly space; The weighted Minkowski anomaly estimate W(x i ): Introduce parameters [w1,w2] to control the importance of the two anomalies; Estimated outliers using the BikNN method It is estimated by combining the Mahalanobis anomaly and the weighted Minkowski anomaly and is expressed as follows: The parameter μ∈[0,1] is used to balance the two anomalies.
5. A novel abnormal data detection method according to claim 4, characterized in that: The integrated ECOD anomaly detection method is specifically implemented as follows: By integrating the ECOD method on the anomaly score obtained by the above BikNN method, the final anomaly score is obtained; Example x i The abnormal score S(x i ) by combining BikNN anomaly scores and ECOD anomaly score To estimate, it is expressed as follows: Where a1 and a2 are obtained through Bayesian optimization.
6. A novel abnormal data detection method according to claim 1, characterized in that: The step 3 further includes obtaining a final anomaly score by integrating the DIF anomaly detection method.