Abnormity detection method and device based on deep isolation forest
The data set is mapped and featured by neural network model, combined with the deep isolation forest algorithm, and the abnormal score is calculated using the deviation between the average path length, node characteristic value and split threshold and local density, to solve the problem of robustness and low accuracy of abnormal detection in high-dimensional data scenarios in the existing technology, and achieve more efficient abnormal data recognition.
Patent Information
- Application Number
- CN202510313632.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-08
AI Technical Summary
The existing anomaly detection methods are relatively robust and accurate when facing high-dimensional data and diversified data scenarios, making it difficult to effectively identify abnormal data.
The data set is mapped and featured by neural network model, mapped the data into low-dimensional latent representations, built an isolated forest, and combined the average path length, the deviation between node characteristic values and split thresholds, and local density, anomaly scores are calculated to filter out abnormal data points.
It improves the accuracy and robustness of abnormal data identification under data types and data scales in different application scenarios, and improves the efficiency and accuracy of abnormal detection.
Smart Images

Figure CN120277561A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of anomaly detection, and particularly relates to an anomaly detection method and device based on a deep isolation forest. Background Art
[0002] When data in industrial and medical scenarios is abnormal, it may lead to problems such as production interruption, equipment damage, or patient safety risks. Therefore, effective anomaly detection methods are of great significance for ensuring the stability and reliability of the system. Anomaly detection algorithms are widely used in fields such as industry, network security, and healthcare. By real-time monitoring data streams and automatically identifying abnormal patterns, problems can be detected early, improving the stability and security of the system. Traditional anomaly detection algorithms may be limited by difficulties in obtaining labels, false alarm rates, and high-dimensional data, resulting in low efficiency and accuracy of anomaly detection. To solve the above problems, a Chinese patent document with the authorization announcement number CN117647367B proposes a method and system for locating leak points in an aircraft fuel tank based on machine learning. It uses a Deep Isolation Forest (DIF) to identify abnormal points. First, a CNN model is used to extract features and divide the processed fuel tank data set. Then, the isolated forest is trained using the divided feature data. By calculating the path length of each data point on each tree, abnormal points in the data are extracted, and then the leak point position is accurately located. This method can achieve non-linear isolation, thereby improving the accuracy of non-linear data anomaly detection. However, with the exponential growth of data scale and dimension, in fields such as industry, healthcare, and finance, the problems faced by anomaly detection are becoming increasingly large, and the application scenarios are also more and more, the data types are more and more, and the data scales are also different. As a result, when using the path length of data points in the deep isolation forest algorithm to identify abnormal points, the robustness and accuracy are relatively low, resulting in errors in the output anomaly detection accuracy. Summary of the Invention
[0003] The purpose of the present invention is to provide an anomaly detection method and device based on a deep isolation forest to solve the problem of low robustness and accuracy when existing anomaly detection methods detect abnormal data.
[0004] The present invention provides an anomaly detection method based on a deep isolation forest to solve the above technical problems, including:
[0005] Step 1): Input the data set to be detected into a neural network model for data mapping and non-linear division to obtain a divided data representation set, and construct an isolation forest using the divided data representation set;
[0006] Step 2): Calculate the average path length of data points in the entire isolation forest and the deviation density measure for each data point, where the deviation density measure is calculated based on the local density and the deviation between the node eigenvalue and the splitting threshold;
[0007] Step 3): Calculate the anomaly score for each data point based on the average path length and the deviation density measure, compare the anomaly score with a set threshold, and screen out the anomaly data points.
[0008] Furthermore, in Step 2), the deviation density measure is calculated according to the following formula:
[0009]
[0010] where, G(x y |t i ) is the deviation density measure, N(x i ) is the neighbor set of the data point xi, d(x i ) is the position of xi, d(x j ) is the position of the neighbor x i of x j , x j is the j-th node, is the node path of the traversal of x y on the tree, x y is the representation of the data object on a certain isolation tree, is the left or right branch that x y enters according to the threshold j k judgment, j k is the threshold for judging the split, and k is the k-th node.
[0011] Furthermore, in Step 3), the anomaly score for each data point is calculated according to the anomaly score formula, and the anomaly score formula is:
[0012]
[0013] where, S(d|T) is the anomaly score function, T is the total number of isolation trees, t i is the i-th isolation tree, is the traversal path length of x y in the isolation tree t i , and Z(T) is the normalization factor.
[0014] Furthermore, the deviation density measure used when calculating the anomaly score for each data point in Step 3) is obtained by adjusting the deviation density measure calculated in Step 2) according to the size of the data set and the structure of the isolation tree.
[0015] Further, adjust the deviation density metric by means of experimental verification, data cross-verification, or an adaptive strategy.
[0016] Further, there are N neural network models, and each neural network model outputs a representation of data. The N neural network models constitute a deep representation integration model; N ≥ 2.
[0017] The beneficial effects of the above technical solutions are as follows: The present invention is an improved invention. When using the deep isolation forest algorithm for anomaly detection, first, the neural network model is used to map and feature partition the dataset to be detected, and the original high-dimensional data is mapped into a low-dimensional latent representation, realizing a highly flexible data space partition for high-dimensional non-linear data. Then, an isolation forest is constructed using the obtained data representation set after partitioning. When performing anomaly scoring on the data points of the isolation forest, the average path length, the deviation between the node feature value and the splitting threshold, and the local density are fused, thereby improving the accuracy and robustness in identifying anomaly data under different application scenarios and data scales.
[0018] To solve the above technical problems, the present invention also provides an anomaly detection device based on a deep isolation forest, including a processor, and the processor is used to implement an anomaly detection method based on a deep isolation forest, including:
[0019] Step 1): Input the dataset to be detected into the neural network model for data mapping and non-linear partitioning to obtain a partitioned data representation set, and construct an isolation forest using the partitioned data representation set;
[0020] Step 2): Calculate the average path length of the data points in the entire isolation forest and the deviation density metric of each data point, where the deviation density metric is calculated based on the local density and the deviation between the node feature value and the splitting threshold;
[0021] Step 3): Calculate the anomaly score of each data point based on the average path length and the deviation density metric, compare the anomaly score with a set threshold, and screen out the anomaly data points.
[0022] Further, in step 2), the deviation density metric is calculated according to the following formula:
[0023]
[0024] where G(x y |t i ) is the deviation density metric, N(x i ) is the neighbor set of the data point xi, d(x i ) is the position of xi, d(x j ) is the neighbor x i of x jPosition, x j Is the j-th node, Is x y The node path of the traversal on the tree, x y Is the representation of the data object on a certain isolation tree, Is x y According to the threshold j k Judge the left and right branches to enter, j k Is the threshold for judging splitting, and k is the k-th node.
[0025] Further, in step 3), calculate the anomaly score of each data point according to the anomaly score formula. The anomaly score formula is:
[0026]
[0027] Among them, S(d|T) is the anomaly score function, T is the total number of isolation trees, and t i Is the i-th isolation tree, Is x y The traversal path length in the isolation tree t i Z(T) is the normalization factor.
[0028] Further, the deviation density metric used when calculating the anomaly score of each data point in step 3) is obtained by adjusting the deviation density metric calculated in step 2) according to the size of the data set and the structure of the isolation tree.
[0029] The beneficial effects of the above technical solutions are as follows: The present invention is an improved invention. When using the deep isolation forest algorithm for anomaly detection, first, the data set to be detected is mapped and feature partitioned through a neural network model, and the original high-dimensional data is mapped into a low-dimensional latent representation to achieve a highly flexible data space partition for high-dimensional non-linear data. Then, an isolation forest is constructed using the data representation set obtained after partitioning. When calculating the anomaly score of the data points in the isolation forest, the average path length, the deviation between the node feature value and the splitting threshold, and the local density are fused, thereby improving the accuracy and robustness in identifying anomaly data under different application scenarios and data scales. Description of the Drawings
[0030] Figure 1 Is the flow chart of anomaly detection based on deep isolation forest in the method embodiment of the present invention. Detailed Embodiment
[0031] In order to make the purpose, technical solutions and advantages of the present invention clearer and more understandable, the specific embodiments of the present invention will be further described below with reference to the drawings.
[0032] When the present invention uses the deep isolation forest algorithm for anomaly detection, it first maps and divides the features of the dataset to be detected through a neural network model, maps the original high-dimensional data into a low-dimensional latent representation, and achieves a highly flexible data space division for high-dimensional non-linear data. Then, it constructs an isolation forest using the obtained data representation set after division. When performing anomaly scoring on the data points of the isolation forest, it combines the average path length, the deviation between the node feature value and the splitting threshold, and the local density, thereby improving the accuracy and robustness in identifying anomaly data under different application scenarios and data scales.
[0033] Method Embodiment
[0034] An anomaly detection method based on the deep isolation forest of the present invention is used in the data cleaning process after machine vision assistance collection, applicable to large-scale and high-dimensional complex datasets, achieving efficient division and anomaly detection of data in industrial and other scenarios, improving the accuracy and efficiency of anomaly detection, reducing algorithm bias, maintaining the scalability of the isolation forest algorithm, and having good performance and adaptability. As Figure 1 shown, it includes the following steps:
[0035] 1. Input the dataset to be detected into the neural network model to obtain a divided data representation set.
[0036] For example, in an industrial scenario, an industrial robot collects data such as the dimensions, shapes, and surface defects of parts through sensors such as machine vision on an automobile manufacturing assembly line; in a power plant or substation, intelligent monitoring devices collect the operating state data of power equipment (such as generators, transformers, etc.), including data such as the appearance, temperature, and vibration of the equipment, through machine vision and various sensors (such as temperature sensors, vibration sensors, etc.); in a food processing factory, automated detection equipment uses machine vision and other relevant sensors to collect relevant data on food raw materials (such as the appearance quality and size of fruits and vegetables) and processed foods (such as whether the packaging is intact and whether there are defects in the appearance of the food); and sends the dataset to be detected collected through sensors such as machine vision to the processor, and the processor uses the neural network model to perform data mapping and random division of the dataset to be detected. After data mapping of the dataset to be detected through a randomly initialized neural network model, diverse data representations are created in the new space, making it easier to isolate the anomaly points in the data, and then a divided data representation set is obtained through axis-parallel division. Preferably, the neural network model uses a deep neural network model.
[0037] In a preferred embodiment, there are N neural network models, and each neural network model outputs a representation of the data. The N neural network models constitute a deep representation ensemble model; N ≥ 2. The deep representation ensemble method is an ensemble learning method based on deep learning, aiming to improve the performance of the model by integrating the representations (i.e., feature representations) of multiple deep learning models. In this method, each deep learning model is responsible for learning different representations of the data, and the ensemble method combines these representations to obtain a more comprehensive and robust feature representation, thereby improving the generalization ability of the model and its adaptability to complex data patterns, effectively enhancing the accuracy and efficiency of anomaly detection. The time complexity of the ensemble process is similar to the feed-forward process of a single neural network because, in a given mini-batch, all ensemble members can be computed simultaneously, improving the partitioning efficiency of the neural network for the dataset, reducing memory occupancy and computation time. The process of the deep representation ensemble model for partitioning data is as follows:
[0038] 1) Initialize the neural network: Use multiple deep neural networks, and the weights of each deep neural network are randomly initialized without optimization or training.
[0039] 2) Parallel feed-forward process: These deep neural networks perform feed-forward calculations simultaneously on the given mini-batch data to generate the data representations corresponding to each deep neural network. This method utilizes the advantage of parallel computing to improve the computing efficiency.
[0040] 3) Representations of ensemble members: The data representations generated by each deep neural network are members of the ensemble, and the set of all members constitutes a random representation ensemble. This ensemble captures the diversity of the original data in different representation spaces.
[0041] 4) Partition the ensemble: In these randomly generated representation spaces, apply axis-parallel partitioning to the random representation ensemble for data partitioning to obtain a set of partitioned data representations. These partitions are equivalent to non-linear partitions in the original data space in the new space, thereby improving the ability to isolate outliers.
[0042] 2. Construct an isolation forest using the set of partitioned data representations.
[0043] In a deep isolation forest, an isolation forest is constructed using a set of data representations obtained through axis-parallel partitioning. The construction method is as follows: multiple isolation trees are assigned to each data representation, and these isolation trees together form the isolation forest to achieve non-linear segmentation of the original data. Each isolation tree is jointly constructed by multiple ensemble members to ensure diversity, and the isolation tree parameters are set to the conventional isolation forest parameters. An isolation tree is a random binary tree, and its construction process includes: randomly selecting an attribute and an attribute value, and then dividing the data set into two parts according to this attribute value, and recursively performing the same operation on these two parts until the stopping condition is met, such as only one record or all records are the same in the data set, or the preset tree height limit is reached. The isolation forest consists of multiple isolation trees, and each tree is constructed independently. In the isolation forest, an outlier is defined as a point that is easily isolated, and the path length of these points in the tree is usually short.
[0044] 3. Calculate the average path length of data points in the entire isolation forest and the deviation density measure of each data point.
[0045] According to the generated isolation forest structure, calculate the average path length of each data point in the isolation tree. Each representation data in the data representation set is obtained by mapping the data points in the original data set, and they are used for isolation operations in the new space to help detect anomalies in the data points of the original data set. In the isolation forest algorithm, to calculate the average path length of a node, it is first necessary to recursively randomly select features and partitioning thresholds in each isolation tree to isolate data points until reaching the leaf node. Then record the path length from each leaf node to the root node, and calculate the average of these lengths to obtain the average path length of a single tree. If there are multiple isolation trees in the isolation forest, then average the average path lengths of each tree to obtain the average path length of the entire isolation forest, which is used to evaluate the anomaly degree of data points. At the same time, calculate the deviation between the node feature value and the splitting threshold on the isolation tree path of each data point and the local density around this point.
[0046] Local density is an indicator used to describe the degree of tightness of data distribution around a data point. In the anomaly detection method based on the deep isolation forest, it plays an important role in measuring whether a data point is anomalous. Intuitively, a high local density means that there are more "neighbor" data points around the data point, and the data distribution is relatively dense; while a low local density indicates that the data around the data point is relatively sparse. For example, in a two-dimensional data space, if there are many other data points clustered around a certain data point, then the local density of this data point is relatively high; on the contrary, if a data point is alone in a relatively empty area, its local density is low. In the context of anomaly detection, anomaly points usually have a low local density. Because anomaly points are often data points that deviate from the normal data pattern, their distribution in the data space is relatively isolated, so their local density will be different from that of normal data points, which makes local density a valuable feature for identifying anomaly points.
[0047] Assume a dataset D, for each data point x in the dataset i , define its local density D(x i ) as follows:
[0048]
[0049] where N(xi) is the neighbor set of the data point xi, d(xi) is the position of xi, d(x j ) is the position of the neighbor x j of xi, in this invention, the Euclidean distance is used as the standard for calculating the average density, |d(x i ) - d(x j )| is the average density of each neighbor x i of x j , and x j is the j-th node.
[0050] And the calculation method for the deviation between the node feature value of a certain point and the splitting threshold is as follows: assume x y is the representation of a data object on a certain isolation tree, is the node path of the traversal of x y on the tree, and the deviation between the node feature value and the splitting threshold is defined as:
[0051] Based on the local density and the deviation between the node feature value and the splitting threshold, the deviation density metric is calculated according to the deviation density metric formula, which is used to quantify the degree of deviation of the data point in the newly created data space. The deviation density metric formula is:
[0052]
[0053] Among them, G(x y |t i ) is the deviation density metric, for x y judge to enter the left and right branches according to the threshold j k where j k is the threshold for judging splitting, and k is the k-th node.
[0054] 4. Calculate the anomaly score for each data point based on the average path length and the deviation density metric, compare the anomaly score with the set threshold, filter out the anomaly data points, mark the anomaly points, and return the anomaly detection result to support further processing and response in subsequent industrial scenarios.
[0055] Calculate the anomaly score for each data point according to the anomaly score formula. The anomaly score formula is:
[0056]
[0057] Among them, S(d|T) is the anomaly score function, T is the total number of isolation trees, and t i is the i-th isolation tree, for x y the traversal path length in the isolation tree t i and Z(T) is the normalization factor used to normalize the anomaly score.
[0058] Furthermore, dynamically adjust the deviation density metric based on the size of the data set and the structure of the isolation tree. Dynamically adjusting the deviation density metric generally involves adapting the algorithm to different data sets and scenarios according to the scale of the data set and the structure of the isolation tree to optimize the performance of the isolation forest algorithm. This may include adjusting the number of trees, the maximum number of samples per tree, the contamination ratio, the number of feature selections, the maximum depth of the tree, and in some variants, adjusting the deviation density metric itself. The deviation density metric can be adjusted through experimental verification, data cross-validation, or adaptive strategies to ensure that the algorithm can achieve the best anomaly detection effect on different data sets. Certain parameters in the anomaly score formula, such as weights, exponents, etc. in the anomaly score formula, can also be adjusted according to the size of the data set or specific features to adapt to the characteristics of different industrial scenarios. Introduce a deviation density adjustment factor in the anomaly score to dynamically adjust the deviation between the node feature value and the splitting threshold, the local density of the node, and the contribution of the average path length to the final score, improving the anomaly detection efficiency and accuracy, with higher flexibility, and better coping with different types of complex data through this method.
[0059] Device Embodiment
[0060] An anomaly detection device based on deep isolation forest of the present invention includes a processor, which is used to implement an anomaly detection method based on deep isolation forest introduced in the method embodiment of the present invention. The specific implementation process of this method has been described in detail in the method embodiment and will not be elaborated here.
Claims
1. An anomaly detection method based on deep isolation forest, characterized in that, Including: Step 1): Input the dataset to be detected into the neural network model for data mapping and non-linear partitioning, obtain the partitioned data representation set, and construct an isolation forest using the partitioned data representation set; Step 2): Calculate the average path length of data points in the entire isolation forest and the deviation density metric for each data point, where the deviation density metric is calculated based on the local density and the deviation between the node feature value and the splitting threshold; Step 3): Calculate the anomaly score for each data point based on the average path length and the deviation density metric, compare the anomaly score with the set threshold, and filter out the anomaly data points.
2. The anomaly detection method based on deep isolation forest according to claim 1, characterized in that In Step 2), the deviation density metric is calculated according to the following formula: where G(x y |t i ) is the deviation density metric, N(x i ) is the set of neighbors of the data point xi, d(x i ) is the position of xi, d(x j ) is the position of the neighbor x i of x j , x j is the j-th node, is the node path of the traversal of x y on the tree, x y is the representation of the data object on a certain isolation tree, is the left and right branches entered by x y according to the threshold j k judgment, j k is the threshold for judging splitting, and k is the k-th node.
3. The anomaly detection method based on deep isolation forest according to claim 2, characterized in that, In Step 3), calculate the anomaly score for each data point according to the anomaly score formula, and the anomaly score formula is: Among them, S(d|T) is the anomaly scoring function, T is the total number of isolation trees, and t i is the i-th isolation tree, is x y the traversal path length in the isolation tree t i and Z(T) is the normalization factor.
4. The anomaly detection method based on deep isolation forest according to claim 1, characterized in that The deviation density metric used to calculate the anomaly score for each data point in Step 3) is obtained by adjusting the deviation density metric calculated in Step 2) according to the size of the dataset and the structure of the isolation tree.
5. The anomaly detection method based on deep isolation forest according to claim 4, characterized in that Adjust the deviation density metric through experimental verification, data cross-validation, or an adaptive strategy.
6. The anomaly detection method based on deep isolation forest according to claim 1, characterized in that There are N neural network models, and each neural network model outputs a representation of the data. The N neural network models constitute a deep representation integration model; N ≥ 2.
7. An anomaly detection device based on deep isolation forest, comprising a processor, characterized in that, The processor is used to implement an anomaly detection method based on a deep isolation forest, including: Step 1): Input the dataset to be detected into the neural network model for data mapping and non-linear partitioning, obtain the partitioned data representation set, and construct an isolation forest using the partitioned data representation set; Step 2): Calculate the average path length of data points in the entire isolation forest and the deviation density metric for each data point, where the deviation density metric is calculated based on the local density and the deviation between the node feature value and the splitting threshold; Step 3): Calculate the anomaly score for each data point based on the average path length and the deviation density metric, compare the anomaly score with the set threshold, and filter out the anomaly data points.
8. The anomaly detection device based on the deep isolation forest according to claim 7, characterized in that In Step 2), the deviation density metric is calculated according to the following formula: where G(x y |t i ) is the deviation density measure, N(x i ) is the set of neighbors of the data point xi, d(x i ) is the position of xi, d(x j ) is the position of the neighbor x i of x j , x j is the j-th node, is the node path of the traversal of x y on the tree, x y is the representation of the data object on a certain isolation tree, is the left and right branches entered by x y according to the threshold j k , j k is the threshold for judging splitting, and k is the k-th node.
9. The anomaly detection device based on deep isolation forest according to claim 8, characterized in that, In Step 3), calculate the anomaly score for each data point according to the anomaly score formula, and the anomaly score formula is: Among them, S(d|T) is the anomaly scoring function, T is the total number of isolation trees, and t i is the i-th isolation tree, is x y the traversal path length in the isolation tree t i and Z(T) is the normalization factor.
10. The anomaly detection device based on the deep isolation forest according to claim 7, characterized in that, The deviation density metric used to calculate the anomaly score for each data point in Step 3) is obtained by adjusting the deviation density metric calculated in Step 2) according to the size of the dataset and the structure of the isolation tree.
Citation Information
Patent Citations
A method and system for locating aircraft fuel tank leaks based on machine learning
CN117647367B