Distance-based double-center deviation anomaly detection method, detection system and application
By adopting a distance-based dual-center biased anomaly detection method in abnormality detection, using the dual-center mechanism and anchor a to distinguish target anomaly from other samples, the problems of high false positive rates and incomplete label data in the existing methods are solved, and efficient target anomaly recognition and false positive rates are achieved.
Patent Information
- Application Number
- CN202311611084.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-29
- Publication Date
- 2025-05-30
AI Technical Summary
Existing anomaly detection methods are difficult to effectively identify high-risk target anomalies and ignore low-risk non-target anomalies, resulting in high false positive rates and difficulty in obtaining non-target anomaly tags.
Using a distance-based dual-center biased anomaly detection method, the target anomaly sample is separated from the supersphere by mapping the normal sample from the non-target anomaly in hidden space and away from the supersphere, using the dual-center mechanism and anchor a to distinguish target anomaly from other samples.
It significantly reduces the false positive rate of abnormality detection, improves the recognition rate of target abnormalities, can effectively respond to the challenge of incomplete label data, and is suitable for biased abnormality detection and standard abnormality detection.
Smart Images

Figure CN120067838A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of anomaly detection, and particularly relates to a distance-based dual-center biased anomaly detection method, a detection system and an application in the case of coexistence of multiple levels of anomalies. Background Art
[0002] In recent years, with the vigorous development of information technology, data in various fields has grown rapidly, and some illegal behaviors have accompanied it. Anomaly detection is a task aimed at identifying objects that deviate significantly from most data, and has important applications in various fields, such as fraud detection in the financial field, risk management in the banking industry, intrusion detection in network security, and disease prediction in healthcare. Since it is very challenging to obtain a sufficient number of accurate labeled samples in the anomaly detection scenario, unsupervised anomaly detection methods that do not use any labeled data have dominated this field for decades. However, in practical applications, a small number of labeled anomalies can be obtained. Therefore, many semi-supervised anomaly detection algorithms have emerged in recent years, which effectively improve the performance of anomaly detection algorithms by using a small number of labeled anomaly samples as prior knowledge.
[0003] However, in the practical application of anomaly detection, anomalies usually have the characteristic of different priorities. Due to practical scenarios and human efficiency considerations, business owners often only focus on one or several types of anomalies, and have a higher tolerance for other types of anomalies. For example, in the risk control scenario of an aggregated payment platform, high-risk anomalies, such as gambling and money laundering, pose a major threat to the platform's security, while low-risk anomalies such as cash-out and false orders are relatively less harmful. Considering the importance of risk management and the limited human resources, it is crucial to accurately identify high-risk anomalies, while low-risk behaviors are tolerable for business owners. Simply classifying all samples that significantly deviate from normal samples as anomalies is not sufficient to meet the actual needs, and existing methods often ignore this data characteristic and requirement. Therefore, it is of great significance to introduce biased anomaly detection (BAD), that is, to focus on identifying target anomalies of interest while deliberately ignoring those non-target anomalies with lower risks.
[0004] Solving biased anomaly detection needs to address two main challenges: First, there is widespread class overlap among target anomalies, non-target anomalies, and marginal normal samples. Effectively distinguishing them to minimize the false positive rate is the key to solving the biased anomaly detection problem. Second, it is difficult to obtain all types of non-target anomaly class labels because they are often not objects of interest to enterprises. In many cases, only the labels of the target anomalies of interest can be obtained. Therefore, maximizing the utilization rate of labeled data is crucial. Existing methods cannot fully address the above challenges: Unsupervised methods do not utilize any labeled anomaly data and often treat all possible anomalies equally. Therefore, non-target anomalies will result in a significant false positive rate. Although most semi-supervised methods can identify target anomalies with a high recall rate, their performance will significantly degrade when dealing with a large number of non-target anomalies. Due to the often severe class overlap between non-target anomalies and target anomalies, classical semi-supervised methods will misclassify most non-target anomalies as target anomalies when solving the biased anomaly detection problem, resulting in an increase in the false positive rate. Another class of newer semi-supervised methods intentionally uses some labeled anomalies to guide the model to identify all types of potential anomalies, which is different from the goal of the biased anomaly detection problem. Summary of the Invention
[0005] To address the deficiencies of the existing technology, the objective of the present invention is to provide a distance-based dual-center biased anomaly detection method that can effectively and accurately separate the target anomalies of interest from other samples, and ensure a high recognition rate and a low false alarm rate.
[0006] Based on the position and distance of sample points in the latent space, the present invention maps normal sample points and non-target sample points into a tight hypersphere in the hyperspace, while moving the target anomaly sample points away from this hypersphere. Different from other existing methods, the present invention adopts a dual-center mechanism and uses a set of auxiliary samples to distinguish target anomaly sample points from other sample points. There are two optional sources for this set of auxiliary samples: provided as prior information, or selected from the unlabeled dataset. The key advantages of the present invention are: The proposed models BiasedAD and the model variant BiasedAD^M in the present invention improve the AUPRC metric by an average of 17.96% and 16.23% respectively compared to the existing methods, indicating a significant reduction in the false positive rate of anomaly detection. Its dual-center mechanism enables the model to effectively cope with the challenge of incomplete labeled data. In addition, the proposed variant model BiasedAD^M can be applied to both biased anomaly detection and standard anomaly detection tasks simultaneously.
[0007] In the present invention, the dual-center mechanism refers to the simultaneous use of the center point c of normal samples and the anchor point a to distinguish and expand the distance between target abnormal samples and non-target hyperspheres. Among them, the center point of normal samples represents the average position of unlabeled data in the latent space, while the anchor point represents the average position of non-target abnormal samples or marginal samples in the latent space.
[0008] The specific technical solution for achieving the object of the present invention is: a distance-based dual-center biased anomaly detection method, which includes five steps: identifying the object to be recognized, constructing a data set, pre-training the model, optimizing and iterating the model, and calculating the anomaly score. The method includes the following specific steps:
[0009] Step 1: Identify the object to be recognized. Different from the traditional anomaly detection methods that only include normal samples and abnormal samples, the present invention needs to clarify four sample concepts: normal samples, marginal samples, non-target abnormal samples, and target abnormal samples. Their specific meanings are as follows:
[0010] Normal samples: Samples that occupy most of the data and do not exhibit any abnormal behaviors.
[0011] Marginal samples: Samples that exist on the edge of normal samples and may include noise or normal samples that are difficult to correctly classify from unlabeled data.
[0012] Non-target abnormal samples: Samples that exhibit characteristics different from normal samples and belong to a type of anomaly, but are not the target of the anomaly detection model.
[0013] Target abnormal samples: The main recognition target of the biased anomaly detection model, and it is necessary to separate the target abnormal samples from the other two types of samples (marginal samples, non-target abnormal samples).
[0014] The definition of the semi-supervised biased anomaly detection problem is as follows: Given a large unlabeled data set, which mainly contains normal samples, but also mixed with some target abnormal samples and non-target abnormal samples, and a very small labeled target abnormal sample set (labeled target abnormal data set), the task of semi-supervised biased anomaly detection is to develop a model that can accurately identify target abnormal samples and distinguish the target abnormal samples from the large unlabeled data set.
[0015] Step 2: Construct a data set. In the present invention, in addition to the unlabeled data set D u , the labeled target abnormal data set D t , it is also necessary to use a set of auxiliary data sets D aTo improve the learning of representations for targeted and biased identification of target abnormal samples. The uniqueness of the present invention lies in that two different semi-supervised abnormal detection data settings can be introduced to deal with data scenarios where non-target abnormal label samples are available and unavailable respectively. In the first case, if labeled non-target abnormal samples are available, D a is composed of the labeled non-target abnormal samples, corresponding to the model BiasedAD; in the second case, that is, when the labeled non-target abnormal samples are unavailable, the marginal samples selected from the unlabeled dataset form D a , corresponding to the model BiasedAD^M. It is worth noting that BiasedAD^M is also applicable to the classical abnormal detection setting, without emphasizing the priority level in abnormalities. All possible abnormal types are target abnormalities, the same as the currently widely used standard abnormal detection setting. At this time, D a is composed of the marginal samples selected from the unlabeled dataset.
[0016] Step 3: Model pre-training and initialization. First, pre-train using an autoencoder on the unlabeled dataset D u . The autoencoder used for pre-training consists of two components: an encoder E and a decoder D. Among them, the encoder E can map high-dimensional raw data into low-dimensional latent variables, while the decoder D maps the low-dimensional latent variables into vectors with the same dimension as the original high-dimensional data. The entire autoencoder is trained through the reconstruction error between the vectors before and after encoding. After the autoencoder converges, use the network weights of the encoder E to initialize the neural network Φ(·; W) with the same structure as the encoder E. The neural network Φ is the model for subsequent biased abnormal detection and is used for the next model optimization iteration.
[0017] Step 4: Model optimization iteration. The models used in the distance-based dual-center biased abnormal detection method proposed by the present invention are named BiasedAD and its variant BiasedAD^M respectively. The difference between the two lies in the difference in the auxiliary dataset D a , as well as the difference in sampling during the target optimization process described in step 4-1-2. The goal of the distance-based dual-center biased abnormal detection method is to train a neural network Φ to effectively compress normal samples and auxiliary samples into a compact non-target hypersphere in the latent space, while pushing the target abnormal samples away from the non-target hypersphere.
[0018] The overall objective functions of the models BiasedAD and its variant BiasedAD^M are defined as follows:
[0019]
[0020] Among them, L compactGuide Φ to minimize the volume of the hypersphere diffused by normal samples and auxiliary samples in the latent space, while L target Guide Φ to increase the distance between the target abnormal samples and this hypersphere. Specifically, it includes the following sub-steps:
[0021] Step 4-1: Use L compact to concentrate a large number of normal samples and non-target abnormal samples in a compact non-target hypersphere. Acting on D u and D a The L compact formula is as follows:
[0022]
[0023] Among them, D u represents the unlabeled dataset, D a represents the auxiliary dataset, Φ is the output of the neural network, x i represents the i-th sample, W represents the weights of the neural network, c represents the center point of the normal samples, and η 0 is the weight parameter.
[0024] Step 4-1-1: For the unlabeled dataset D u , since D u mainly contains normal samples, punish the distance from the learned representations of all samples in D u to c in the latent space, forcing Φ to map normal samples to the vicinity of c in the latent space. Among them, c is obtained by averaging the output of the first forward pass of Φ for D u , and this output is initialized using a pre-trained autoencoder:
[0025]
[0026] Among them, D u represents the unlabeled dataset, Φ is the output of the neural network, x i represents the i-th sample, W represents the weights of the neural network.
[0027] Step 4-1-2: Make the auxiliary dataset D a close to the center point c of the normal samples.
[0028] The auxiliary dataset D a has two possible compositions: If labeled non-target abnormal samples can be obtained, then D a is composed of non-target abnormal samples marked by labels, corresponding to the model BiasedAD; otherwise D a is composed of D uIt consists of the selected edge samples. The edge samples are the samples that are farthest from the center point c of the normal samples in the latent space, corresponding to the variant BiasedAD^M of the model BiasedAD, and are sampled in real time during the iteration process. The specific update method is shown in Step 4-2-1. In both cases, the auxiliary data exhibits characteristics different from most normal samples and tends to be located at a certain distance from the center point c of the normal samples in the latent space.
[0029] L compact The second term in changes the position of the auxiliary data by guiding Φ, which is similar to the term acting on the unlabeled dataset D u In addition, the hyperparameter η 0 is introduced to control the balance of the influence of the unlabeled samples and the auxiliary samples. By increasing the value of η 0 (for example, η 0 takes 1 under standard circumstances, and can take η 0 as 10 when the number of auxiliary samples is extremely small), the neural network Φ can emphasize the learning of the auxiliary samples, thus obtaining a more compact hypersphere.
[0030] Step 4-2: Use the auxiliary samples to enhance the learning of the target abnormal representation to reduce the false positive rate. It specifically includes the following sub-steps:
[0031] Step 4-2-1: Definition and initialization of the anchor point a. The auxiliary samples are usually located at the edge of the hypersphere centered at c. Other anomaly detection methods often misclassify the auxiliary samples as target abnormal samples, which will lead to a significant increase in the false positive rate. To effectively solve this problem, the present invention introduces the anchor point a:
[0032]
[0033] where D a represents the auxiliary dataset, Φ is the output of the neural network, x i represents the i-th sample, and W represents the weights of the neural network. The anchor point a is represented by the average vector of the auxiliary samples in the latent space, representing the average position of the auxiliary samples, and is initially set as the average output of the first forward pass of Φ for D a .
[0034] Step 4-2-2: Use the anchor point a to separate the target abnormal samples from other samples. To effectively detect the target anomalies, while ensuring a certain distance between the target abnormal samples and the center c, it is very important to keep a sufficient distance from the auxiliary samples. Based on this goal, the present invention designs a new loss function, denoted as L target . L targetIt can be the sum of the reciprocal of the distance from the latent space representation of the Φ-penalized target outlier samples to the center point c and the reciprocal of the distance from them to the anchor point a. However, it is noted that when both the target sample center point c and the anchor point a are considered, it may have conflicting effects on the target outlier samples located inside the hypersphere, confusing the optimization direction of the model. To solve this problem, the present invention proposes a refined loss considering the respective positions of each target outlier sample, and divides the labeled target outlier dataset D t into a difficult-to-identify set D hard and an easy-to-identify set D easy in real time according to their positions in the latent space. Specifically, assume a hypersphere centered at c with a radius equal to the distance between c and a. The target outlier samples located inside the hypersphere are severely mixed with normal samples and non-target outlier samples, which are more challenging for the model to identify. Therefore, this part of the target outlier samples is called the difficult-to-identify set D hard . On the contrary, the target outlier samples located outside the hypersphere are more separated from normal samples and non-target outlier samples, so they are easy to be identified by the model. Therefore, this part of the target outlier samples is called the easy-to-identify set D easy , that is:
[0035]
[0036]
[0037]
[0038] where D easy represents the set of easy-to-identify target outlier samples, D hard represents the set of difficult-to-identify target outlier samples, Φ is the output of the neural network, x represents the sample, W represents the weights of the neural network, c represents the center point of the normal samples, a represents the anchor point, dist(·, ·) represents the distance between two samples in the latent space, and D t represents the labeled target outlier dataset.
[0039] Therefore, the target outlier set D t is divided into D easy and D hard according to the above formula according to their positions in the latent space, and is updated in real time according to the above formula during the model optimization iteration process. For the target samples belonging to D easy , Φ imposes penalties on the reciprocals of the distances from them to the center point c of the normal samples and the anchor point a; while for the target samples belonging to D hardFor the target samples, Φ reduces the penalty and only refines the penalty on the reciprocal of the distance to the center point c of the normal samples, that is, only penalizes the reciprocal of the distance to the center point c of the normal samples to prevent the anchor points from exerting conflicting influences and confusing the optimization of the model. The loss function L acting on the target abnormal samples target is specifically defined as follows:
[0040]
[0041] where D easy represents the set of easily recognizable target abnormal samples, and D hard represents the set of difficult-to-recognize target abnormal samples, Φ is the output of the neural network, x k represents the k-th sample, W represents the weights of the neural network, c represents the center point of the normal samples, a represents the anchor point, η 1 and η 2 are used to adjust the weights acting on D easy and D hard in the formula.
[0042] Step 4-3: Update the position of the anchor point a
[0043] During the entire training process, the center point c of the normal samples is only initialized in Step 4-1-1 and remains unchanged, acting as a fixed center point of the normal samples. The anchor point a is dynamically updated according to Step 4-2-1 in each round of training to ensure that the anchor point a can always accurately reflect the current position of the auxiliary samples, enabling the model to better distinguish target abnormal samples from other samples. As the model iterates, the difficult-to-recognize set D hard and the easily recognizable set D easy will change continuously with the update of a: all samples in the labeled target abnormal set D t will gradually move away from c and gradually transfer from the difficult-to-recognize set D hard to the easily recognizable set D easy .
[0044] Continuously repeat Step 4-1 to Step 4-3 until the overall objective function of the model BiasedAD and its variant BiasedAD^M no longer decreases and the model converges.
[0045] Step 5: Calculate the anomaly score. After the model converges, calculate the anomaly score of the samples in the test set according to the following formula:
[0046]
[0047] Among them, s(x) represents the distance from the sample x after being mapped to the latent space by the neural network Φ to the center point c of the normal samples; Φ is the output of the neural network, x represents the sample, W represents the weights of the neural network, c represents the center point of the normal samples, and ||·|| 2 represents the L2 distance in the latent space. The higher the score of s(x), the greater the probability that the sample x is the target abnormal sample, and the lower the score of s(x), the smaller the probability that the sample x is the target abnormal sample.
[0048] The present invention provides a detection system for implementing the above-mentioned anomaly detection method. The detection system includes: a data pre-training module, a bias target optimization module, and an anomaly score calculation module;
[0049] The model pre-training module is used for the mutual mapping between high-dimensional and low-dimensional data, and the initialization of the bias anomaly detection model after the convergence of the autoencoder in the pre-training. The autoencoder used in the pre-training can map the high-dimensional original data into low-dimensional latent variables, and then map the low-dimensional latent variables into vectors with the same dimension as the original high-dimensional data, and train through the reconstruction error between the vectors before and after the reconstruction coding. After convergence, the encoder network weights are used to initialize the neural network Φ(·; W) for the next model optimization iteration;
[0050] The bias target optimization module is used to separate the target abnormal samples and other samples in the latent space. The present invention proposes a distance-based dual-center bias anomaly detection method, and the models used are named BiasedAD and its variant BiasedAD^M respectively. The goal of the bias target optimization module is to train a neural network Φ to effectively compress the normal samples and auxiliary samples into a compact non-target hypersphere in the latent space, while pushing the target abnormal samples away from the non-target hypersphere;
[0051] The anomaly score calculation module is used to calculate the anomaly score after the convergence of the bias anomaly detection model to complete the bias anomaly detection. The anomaly score is defined as the distance from the sample x after being mapped to the latent space by the neural network Φ to the center point c of the normal samples. The higher the score, the greater the probability that the sample x is the target abnormal sample. The anomaly score calculation module will assign a significantly higher anomaly score to the target abnormal samples than to other samples.
[0052] The present invention also provides applications of the above-mentioned anomaly detection method or detection system in anomaly detection of financial data (such as anomaly detection in the aggregated payment platform), image anomaly detection, medical data anomaly detection, network intrusion tabular data anomaly detection, etc. For example, in the risk control scenario of the aggregated payment platform, high-risk anomalies such as gambling and money laundering pose a significant threat, while low-risk anomalies such as cash withdrawal and false orders are relatively less harmful. Considering the importance of risk management and limited human resources, it is crucial to accurately identify high-risk behaviors as target anomalies, while low-risk behaviors as non-target anomalies do not need to be identified. The present invention is applicable to the above scenarios and is also applicable to anomaly detection of image data and network intrusion tabular data, which can significantly reduce the false positive rate of non-target anomalies.
[0053] Compared with the prior art, the present invention provides a new and general method for the actual needs in the real anomaly detection scenario, with good applicability. The present invention can be directly applied to various types of tasks, with a simple method and high efficiency, including the following beneficial technical effects:
[0054] (1) The present invention first defines the biased anomaly detection task and first proposes that in the anomaly detection task, all possible anomalies cannot be treated equally, emphasizes the concept of priority among anomalies, and focuses on accurately identifying the target anomalies of interest.
[0055] (2) Compared with the existing anomaly detection methods, the present invention has strong generality. The model proposed by the present invention does not depend on the data type and is applicable to tabular data and image data at the same time, and is applicable to fields such as network intrusion, financial fraud detection, and image anomaly detection. For different data types, only the model structure of the pre-training module needs to be changed, and after mapping to the latent space, the general biased target optimization module and anomaly score calculation module can be used for processing.
[0056] (3) The present invention proposes a simple and effective anomaly detection model specifically designed to solve the biased anomaly detection problem. Through the dual-center mechanism and innovative loss function in step 4, the model effectively reduces the number of false positives caused by the class overlap between normal samples, non-target and target anomalies. The proposed models BiasedAD and BiasedAD^M respectively improve the AUPRC index by 17.96% and 16.23% on average compared with the existing methods, reflecting a significant reduction in the false positive rate of anomaly detection. Its dual-center mechanism enables the model to effectively cope with the challenge of incomplete label data.
[0057] (4) The present invention effectively solves the biased anomaly detection tasks in two scenarios, namely the scenarios where labeled non-target anomalies are available and where labeled non-target anomalies are not available. The BiasedAD model proposed by the present invention only uses a very small number of labeled target anomaly samples (accounting for 0.31%-4.05% of all data according to different test problems), demonstrating the characteristics of high flexibility and high labeled data utilization efficiency. Even in the case where non-target anomaly samples are not available, the BiasedAD^M model proposed by the present invention still improves the key indicator AUPRC by an average of 16.23% and significantly reduces the false positive rate.
[0058] (5) The BiasedAD^M model proposed by the present invention is also applicable to the standard anomaly detection scenario. In the standard anomaly detection scenario (i.e., all types of anomalies need to be detected), the dual-center mechanism and the innovative loss function design can be used to significantly reduce the false positive rate of anomaly recognition. Under the test, the key indicator AUPRC is increased by an average of 40.21%, significantly reducing the cost of manual review.
[0059] (6) The loss function designed by the present invention makes full use of the labeled data. On the premise of achieving comparable performance with the best comparison method, it reduces the utilization amount of labeled anomalies by 76.67% and 99% respectively, showing extremely high utilization rate of labeled target anomalies, effectively reducing the manually labeled samples required for anomaly detection, and can significantly reduce the manual labeling cost.
[0060] (7) The innovative loss function design used by the present invention has strong robustness to the noise in the unlabeled data, demonstrating a strong resistance to noise interference. This characteristic relaxes the requirements of the present invention for unlabeled data and is more widely applicable in real scenarios.
[0061] The present invention is applied to datasets in three different categories and fields, all showing significant performance improvement: implemented on the merchant anomaly detection data of the aggregated payment platform in the financial field, it improves the key indicator AUPRC by at least 4.38% compared with the existing methods; implemented on the UNSW_NB15 dataset in the network intrusion field, it improves the key indicator AUPRC by at least 6.92% compared with the existing methods; implemented on the image data in the fashion field, it improves the key indicator AUPRC by 2.71%-11.71% compared with the existing methods in different settings.
[0062] The present invention is applied to the UNSW_NB15 dataset in the field of network intrusion, and shows significantly better utilization of labeled data compared to other existing semi-supervised anomaly detection methods. Compared with the existing best-performing methods, the models BiasedAD and BiasedAD^M can reduce the utilization of labeled data by 76.67% and 99% respectively while achieving the same good detection effect. Description of the Drawings
[0063] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following-described drawings are only some embodiments of the present invention. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0064] Figure 1 It is the overall flowchart of the present invention.
[0065] Figure 2 It is the optimization objective of the model proposed by the present invention.
[0066] Figure 3 It is the specific statistical information of the relevant datasets in the specific implementation manner of the present invention.
[0067] Figure 4 It is the comparison results of the present invention applied to a total of five datasets in different fields, data types, and settings in the specific implementation manner. Detailed Description of the Invention
[0068] Combined with the following specific embodiments and drawings, the present invention will be further described in detail. The processes, conditions, experimental methods, etc. for implementing the present invention, except for the specifically mentioned content below, are all common knowledge and well-known common sense in the art, and the present invention has no particularly restricted content.
[0069] The present invention provides a distance-based dual-center biased anomaly detection method. Existing anomaly detection methods usually ignore the different priorities in anomalies and uniformly identify all types of anomaly samples. The present invention proposes the concept of "biased anomaly detection", whose goal is to accurately identify target anomalies while deliberately ignoring non-target anomalies that are not of interest. Existing methods are difficult to solve the problem of high false positive rate caused by the class overlap between target anomalies, non-target anomalies, and normal samples, and it is very difficult to obtain a sufficient number of labeled non-target samples. The present invention proposes a simple and effective semi-supervised biased anomaly detection model, which uses a dual-center mechanism to expand the boundary between target anomalies and other samples, significantly reducing the false positive rate and effectively coping with the challenge of limited labeled data. It has good effectiveness, robustness, and label utilization rate, and at the same time proposes a brand-new solution for the standard anomaly detection problem.
[0070] Specifically, the present invention provides a distance-based dual-center bias anomaly detection method. The overall process of the anomaly detection method is as Figure 1 shown, including the following steps:
[0071] Step 1: Identify the objects to be recognized, and divide the objects to be distinguished and detected into normal samples, marginal samples, non-target anomaly samples, and target anomaly samples;
[0072] Step 2: Construct a data set, which includes an unlabeled data set D u , a labeled target anomaly data set D t , and an auxiliary data set D a ;
[0073] Step 3: Model pre-training and initialization. Use an autoencoder to perform pre-training on the unlabeled data set D u , and initialize the bias anomaly detection model with the network weights after the convergence of the autoencoder;
[0074] Step 4: Optimize and iterate the bias anomaly detection model, and train the neural network Φ to effectively compress the normal samples and the auxiliary samples in a non-target hypersphere in the latent space, and at the same time push the target anomaly samples away from the non-target hypersphere;
[0075] Step 5: Calculate the anomaly score to obtain the anomaly test result.
[0076] In the present invention, the anomalies are classified. For the anomaly detection problem, most traditional methods classify all anomalies into one category without considering the differences within the anomalies, and only uniformly perform binary classification. This overly rough classification method will lead to a decrease in recognition accuracy and a high false alarm rate. Other methods notice the differences and diversities within the anomalies, but the purpose of this type of method is to use some known anomaly types to identify all possible anomaly types in a perceptive way. From the perspective of the recognition effect, there is still no targeted distinction for the anomalies, and it will still lead to a high false alarm rate.
[0077] The present invention proposes a brand-new bias anomaly detection method that can accurately identify target anomalies, emphasizing the significant priority difference between abnormal samples. The proposed method can accurately separate the target anomalies of concern from other samples without causing confusion of other non-concerned anomalies and normal marginal samples, significantly reducing the false alarm rate and improving the recognition accuracy. The present invention can also handle the problem of standard anomaly detection and significantly reduce the false alarm rate in the standard scenario.
[0078] The present invention is the first known method that emphasizes the priority in anomalies. The proposed method preferentially classifies target anomaly samples, effectively reducing the interference of non-target anomalies, ensuring a high recognition rate of target anomalies, and very effectively reducing the labor cost in practical applications, with strong practical value.
[0079] In the present invention, a dual-center mechanism is proposed. An anchor point a is introduced to effectively solve the problems of biased and standard anomaly detection in multiple scenarios. Based on the position and distance of sample points in the latent space, normal samples and non-target anomaly sample points are mapped into a tight hypersphere in the hyperspace, while the target anomaly samples are kept away from this hypersphere, and the difference between the non-target hypersphere and the target anomaly samples is further enlarged with the anchor point a. The present invention adopts the dual-center mechanism and uses a set of auxiliary samples to distinguish target anomaly samples from other samples.
[0080] The anchor point a is represented by the average vector of the auxiliary samples in the latent space, representing the average position of the auxiliary samples, and is initially set to the average output of the first forward pass of Φ to D a and is updated in real time with the model iteration later. The anchor point a has two functions: one is to keep the target anomaly samples away from the normal sample center point c and away from the anchor point a at the same time, ensuring a certain distance between the target anomaly samples and the normal sample center point c while maintaining a sufficient distance from the auxiliary samples to separate them from other samples; the other is to assist unsupervised samples in learning a more compact normal hypersphere. Under the control of the hyperparameter η 0 , the model learns a more compact normal sample representation, making the target anomaly samples more significantly distinguishable from the normal samples. Under this mechanism, the key advantage of the present invention is that the proposed model significantly reduces the false positive rate of anomaly detection.
[0081] In the present invention, there are two optional sources for the auxiliary samples used to determine the position of the anchor point a: provided as prior information or selected from an unlabeled dataset.
[0082] When dealing with a biased anomaly detection task (i.e., distinguishing priorities for anomalies, only identifying target anomalies and not non-target anomalies), if non-target anomaly samples are available, the auxiliary dataset D a consists of a small number of non-target anomaly samples with accurate labels; if non-target anomaly samples are not available, the auxiliary dataset D a consists of D uIt consists of the marginal samples selected from D, where the marginal samples are the part of samples that are farthest from the center point c in the latent space and are sampled in real time during the iteration process. In both cases, the auxiliary data exhibits characteristics different from those of most normal samples and tends to be located at a certain distance from the center c in the latent space. Taking this part of the auxiliary data as the anchor point and updating the model according to the biased optimization objective proposed by the present invention can effectively act on the unlabeled samples and the target abnormal samples respectively to obtain a more obvious discrimination effect.
[0083] When dealing with the standard anomaly detection task (i.e., without distinguishing the priorities of anomalies and uniformly identifying all possible anomalies), the auxiliary data set D a is composed of the marginal samples selected from D u and is updated in real time as the model iterates. Using this part of the marginal samples can effectively expand the normal hypersphere and the abnormal samples, widen the boundary between the positive and negative samples, and reduce the false alarms caused by class overlap.
[0084] The key advantage of the present invention is that it is applicable to both the biased anomaly detection problem and the standard anomaly detection problem, and shows a significant improvement in AUPRC compared with other state-of-the-art methods, and effectively addresses the challenge of incomplete labeled data.
[0085] In the present invention, using the auxiliary data set D a to obtain the anchor point and acting on the training objective can effectively solve the recognition difficulty caused by class overlap.
[0086] In the biased anomaly detection problem, the target abnormal samples and other samples often have a certain overlap in the latent space, making the biased anomaly detection more challenging. The dual-center mechanism adopted by the present invention can effectively use the auxiliary data set D a to obtain the anchor point and accurately distinguish the overlapping region. When the overlap degree between the target abnormal and normal samples is more serious, the labeled non-target abnormal can make the target abnormal and normal samples more distinguishable. As the class overlap degree deepens, the detection difficulty increases, and the strategy of introducing the auxiliary data set D a and the anchor point a becomes more effective.
[0087] In the standard anomaly detection scenario, when there is a significant overlap between the positive and negative samples, the strategy of using the marginal samples to increase the gap between the normal and abnormal samples is very effective. In this case, the existence of the marginal samples helps the model better distinguish between normal and abnormal samples, and finally improves the overall detection performance. This shows the robustness and adaptability of the method proposed by the present invention when dealing with complex and challenging anomaly detection scenarios.
[0088] In the present invention, it has the highest label sample utilization efficiency compared with other semi-supervised methods. In actual anomaly detection scenarios, label samples are often difficult to obtain or the quantity is very small, which brings great difficulties to improving the effect of anomaly detection. The method proposed in the present invention has a higher utilization rate of labeled data compared with other methods: only using a very small number of target anomaly label samples can achieve a recognition effect comparable to that of the best competitor. The dual-center mechanism including anchor points can exert two aspects of influence simultaneously: the anchor point a pushes the target anomaly samples away from the normal samples and at the same time compresses the normal samples. When the number of labeled non-target anomaly samples increases, the position of the anchor point becomes more stable and accurate, thereby improving the detection performance. Even if there are only a small number of non-target anomaly samples, a rough position of the anchor point can be obtained to increase the gap between the normal samples and the target anomaly samples.
[0089] In the present invention, the detection method has strong anti-pollution ability. In the problem of anomaly detection, pure normal samples are often unavailable, and there are often a large number of noise data of target / non-target anomalies in the obtained unlabeled data. When the noise ratio is relatively high, it will cause insufficient learning of the pattern features of the normal category by the anomaly detection method, resulting in a decrease in the anomaly detection accuracy and an increase in the false alarm rate. The dual-center mechanism and the design of the learning objective used in the present invention can make the model very robust to the noise in the unsupervised samples. Therefore, even when the pollution ratio in the data increases, the distance-based dual-center bias anomaly detection method proposed in the present invention can maintain a stable and superior AUPRC. Therefore, the key advantage of the present invention is that it is insensitive to anomaly pollution, has low requirements for data quality, has stable anomaly detection effects, and is applicable to various scenarios containing data pollution.
[0090] In the present invention, it is insensitive to the changes of various model hyperparameters. Models sensitive to hyperparameters need to spend a lot of effort and data for hyperparameter tuning, and as the time span changes, the model needs to be retrained and tuned repeatedly, resulting in a waste of a large amount of human and material resources. The method proposed in the present invention has a wide range of hyperparameter changes and has relatively stable performance for the changes of different hyperparameters. Therefore, it has strong robustness and high practicality.
[0091] Specifically, referring to Appendix Figure 1 、 2 ,the distance-based dual-center bias anomaly detection of the present invention is carried out according to the following steps:
[0092] Step 1: Identify the objects to be recognized. Different from the traditional anomaly detection methods that only include normal samples and abnormal samples, four sample concepts need to be defined in the present invention: normal samples, marginal samples, non-target abnormal samples, and target abnormal samples. Except for normal samples, the anomalies that are of great concern and need to be accurately identified with extremely high accuracy are defined as target anomalies, which are the main objects to be recognized by the model. The anomalies that are different from normal samples but are not concerned due to low risk, etc. are defined as non-target abnormal samples, and the model will not recognize this type of samples.
[0093] Step 2: Construct the dataset.
[0094] Step 2-1: Use the mixed data with the vast majority of samples being normal as the unlabeled dataset Du. u 。
[0095] Step 2-2: Collect a small amount of target abnormal data with exact labels to construct the labeled target abnormal dataset Dt. t 。
[0096] Step 2-3: In the standard anomaly detection scenario, that is, when all anomalies need to be detected and recognized by the model, no additional data is required. After pre-training, a small number of marginal samples will be selected from the unlabeled dataset Du to form Dm. In the biased anomaly detection scenario, if labeled non-target abnormal samples are available, Dm will be composed of this small group of labeled non-target abnormal samples; otherwise, that is, when labeled non-target abnormal samples are not available, Dm will be composed of the marginal samples selected from the unlabeled dataset. a 。In the biased anomaly detection scenario, if labeled non-target abnormal samples are available, Dm a will be composed of this small group of labeled non-target abnormal samples; otherwise, that is, when labeled non-target abnormal samples are not available, Dm will be composed of the marginal samples selected from the unlabeled dataset. a 。
[0097] Step 2-4: Preprocess the dataset. For tabular data, convert categorical features to numerical features, use one-hot encoding, and replace missing values with column means. For image data and tabular data, all features are mapped to the range [0, 1] through min-max normalization.
[0098] Step 3: Model pre-training and initialization. Use the unlabeled dataset Du uPut it into the autoencoder for pre-training. The structure of the autoencoder is determined according to the type and category of the input data. The autoencoder used for pre-training consists of two components: an encoder E and a decoder D. The encoder E can map the high-dimensional raw data into low-dimensional latent variables, while the decoder D maps the low-dimensional latent variables into a vector with the same dimension as the original high-dimensional data. Taking the network intrusion dataset UNSW_NB15 as an example, the present invention uses a neural network with 3 hidden layers, where each layer has 168, 64, and 32 nodes respectively, that is, the data is mapped from 168 dimensions after preprocessing to 32 dimensions in the latent space, and the decoder is designed with a structure opposite to that of the encoder. For the image dataset, the present invention uses a variant of LeNet as the autoencoder. The entire autoencoder is trained by the reconstruction error between the vectors before and after encoding. After the autoencoder converges, the network weights of the encoder E are used to initialize the neural network Φ(·;W) with the same structure as the encoder E for the next model optimization iteration.
[0099] Step 4: Model optimization iteration. During the model optimization iteration, the Leaky ReLU function g(z)=max(0,z)+0.01*min(0,z) is used for gradient propagation, and L2 regularization is applied to each hidden layer (if applicable) to mitigate overfitting.
[0100] The models used in the distance-based dual-center bias anomaly detection method proposed by the present invention are named BiasedAD and BiasedAD^M respectively, and its overall objective function is:
[0101]
[0102] Specifically, it includes the following sub-steps:
[0103] Step 4-1: Use L compact to concentrate a large number of normal samples and non-target anomalies in a compact hypersphere. Acting on D u and D a The L compact formula is as follows:
[0104]
[0105] Step 4-1-1: For the constructed unlabeled dataset D u , since D u mainly contains normal samples, the distance between the learned representations of all samples in D u to c in the latent space is penalized, forcing Φ to map normal samples to the vicinity of c in the latent space. First, use the pre-trained autoencoder E to initialize Φ, and the target sample center point c obtained by averaging the output of the first forward pass of Φ is used as the center point for representation learning.
[0106] Input all the sample points in D u into the neural network Φ, map them from the original data dimension to the hidden layer, obtain the low-dimensional vector representation of the hidden layer, and calculate the distances from all the sample hidden vectors in D u to the center point c as the losses corresponding to the samples in Du, and add them to the total optimization objective. Since the number of samples in D u is very large and most of them are normal, punishing this part of the losses can make Φ map the normal sample data to a position closer to the center c, so as to obtain a more compact normal hypersphere (non-target hypersphere).
[0107] Step 4-1-2: Make the auxiliary data set D a close to the center point c.
[0108] The auxiliary data set D a may consist of labeled non-target abnormal samples, or may consist of marginal samples selected from D u . The marginal samples are the samples that are farthest from the center c in the hidden space and are sampled in real time during the iteration process. The specific update method is shown in Step 4-2-1.
[0109] Input all the sample points in D a into the neural network Φ, map them from the original data dimension to the hidden layer, obtain the low-dimensional vector representation of the hidden layer, and calculate the distances from all the sample hidden vectors in D a to the center point c as the losses corresponding to the samples in D u , and add them to the total optimization objective. Since the samples in D a are not the target abnormal samples of concern and are not the objects that the model needs to identify, generally speaking, it is hoped that the samples in D a can also approach c in the hidden space. In addition, regardless of whether D a is composed of non-target abnormal samples or marginal samples, the samples in D a always show characteristics different from most normal samples, tend to be located at a certain distance from the center point c of the target samples in the hidden space, and are distributed on the edge of the normal hypersphere. Therefore, appropriately increasing the hyperparameter η 0 can make the neural network Φ emphasize the learning of the auxiliary data, so as to emphasize the compactness of the edge of the normal hypersphere, and thus obtain a more compact normal hypersphere.
[0110] Step 4-2: Use the auxiliary samples to enhance the learning of the target abnormal representation, so as to use the prior knowledge to reduce the false positive rate. Specifically, it includes the following sub-steps:
[0111] Step 4-2-1: Definition and initialization of anchor point a. Auxiliary samples are usually located at the edge of the hypersphere centered at c. Other anomaly detection methods often misclassify auxiliary samples as target anomaly samples, which will lead to a significant increase in the false positive rate. To effectively solve this problem, the present invention introduces anchor point a:
[0112]
[0113] Anchor point a is represented by the average vector of auxiliary samples in the latent space, representing the average position of the auxiliary samples, and is initially set to the average output of the first forward pass of Φ for D a . All samples in D a are input into the neural network Φ, and after mapping from high dimension to low dimension latent vectors, the average vector is obtained using the above formula as the auxiliary anchor point a.
[0114] Step 4-2-2: Using anchor point a to separate target anomaly samples from other samples. To effectively detect target anomaly samples, while ensuring a certain distance between the target anomaly samples and the center c, it is very important to keep a sufficient distance from the auxiliary samples. Based on this goal, the loss function L target is used to make Φ simultaneously penalize the reciprocal distances between the latent space representations of target anomalies and c and a. When both center points are considered, in order to consider the refined loss of each target anomaly's respective position, D t needs to be divided into a difficult-to-recognize set D hard and an easy-to-recognize set D easy in real time according to their positions in the latent space.
[0115] First, all samples in the labeled target anomaly set D t are input into the neural network Φ to obtain the low-dimensional vector representation in the latent space. Then, assume a hypersphere centered at c with a radius equal to the distance between c and a. Calculate the distances from all samples in the target anomaly set D t located inside the hypersphere to the center c, and compare their distances with c and a. If the distance is less than this distance, then this part of the target anomaly samples is divided into the difficult-to-recognize set D hard . On the contrary, if the distance is greater than the distance between c and a, then the sample is located outside the hypersphere. This part of the target anomaly samples is more clearly separated from normal samples and non-target anomaly samples, so it is easy to be recognized by the model. This part of the target anomaly samples is called the easy-to-recognize set D easy , and the specific implementation is based on the following formula:
[0116]
[0117]
[0118]
[0119] Here, dist(·, ·) represents the distance between two samples in the latent space.
[0120] Divide the target anomaly set D t into D easy and D haard in two parts according to their positions in the latent space according to the above formula, and update them in real time according to the above formula during the model optimization iteration process.
[0121] For the target samples belonging to D easy , calculate the reciprocal of the distance from their output from Φ to the center c and the anchor point a to impose a penalty; for the target samples belonging to D hard , after output from Φ, only calculate the reciprocal of the distance from their output from Φ to the center c for refined penalty to prevent the anchor point from imposing a conflicting influence on it and confusing the optimization of the model. The loss function L target acting on the target anomaly is specifically defined as follows:
[0122]
[0123] where η 1 and η 2 are used to adjust the weights of the two parts acting on D easy and D hard in the formula. Increasing the corresponding hyperparameter can emphasize the learning of the corresponding part.
[0124] Step 4-3: Update the position of the anchor point a
[0125] During the entire training process, the center point c of the target samples is only initialized in Step 4-1-1 and remains unchanged, acting as a fixed center point of normal samples. While the anchor point a is dynamically updated in each epoch according to Step 4-2-1 to ensure that the anchor point a can always accurately reflect the current position of the auxiliary samples, enabling the model to better distinguish target anomalies from other samples. That is, before each round of model training, first calculate the anchor point position of all samples in D a using formula 4-2-1 for use in the loss function calculation of that round.
[0126] As the model is continuously iterated, the difficult-to-identify set D hard and the easy-to-identify set D easy will change continuously with the update of a: all samples in the target anomaly set D t will gradually move away from c and gradually transfer from D hard to the easy-to-identify set D easyTherefore, after calculating the position of the current round's anchor point a, immediately re-partition the difficult-to-recognize set D of the current round according to the formula in step 4-2-2 hard and the easy-to-recognize set D easy , and then calculate the corresponding loss function according to the corresponding data set to update the model.
[0127] Keep repeating step 4 until the overall objective function no longer decreases and the model converges.
[0128] Step 5: Calculate the anomaly score. After the model converges, calculate the anomaly score of the samples in the test set according to the following formula:
[0129]
[0130] where s(x) represents the distance from the sample x after being mapped to the latent space by the neural network Φ to the normal center point c. The higher the score, the greater the probability that x is the target anomaly.
[0131] Input all the samples in the test set into the neural network Φ to obtain their latent variable representations, then calculate the distance from each sample's latent variable to the center point c and sort them. The closer the distance, the smaller the probability that x is the target anomaly, and the farther the distance, the greater the probability that x is the target anomaly.
[0132] The present invention is implemented on data sets in three different fields: The merchant anomaly detection data of the aggregated payment platform in the financial field is a tabular data set with a dimension of 182 and a total of 317,281 data; The data set UNSW_NB15 in the network intrusion field is tabular data, and after processing, its dimension is 196 and there are 103,220 data; The image data in the fashion field is image data with an original dimension of 28*28 and contains 7,700 data, as Figure 3 shown, Figure 3 are the specific statistical information of the relevant data sets in the specific implementation manner of the present invention.
[0133] Refer to Appendix Figure 4 , From the results of target anomaly detection on five data scenarios of the three data sets, whether it is a biased anomaly detection task with multiple anomaly types (UNSWNB 15, SQB, FMNIST 1 , FMNIST 2 ) or a standard anomaly detection task that regards all anomaly types as outliers (FMNIST 3 ), using the models (BiasedAD, BiasedAD M ) proposed by the present invention, the test results on the key anomaly detection index AUPRC are the highest, indicating that the proposed model has the best effect, the highest recognition accuracy and the lowest false alarm rate. Figure 4This is the comparison result of the present invention applied to a total of five datasets in different fields, data types, and settings.
[0134] Example 1
[0135] BiasedAD is implemented on the network intrusion dataset UNSW_NB15. The tested AUPRC: 79.16 ± 0.22; the tested AUROC: 97.37 ± 0.06;
[0136] BiasedAD^M is implemented on the network intrusion dataset UNSW_NB15. The tested AUPRC: 75.91 ± 2.79; the tested AUROC: 97.35 ± 0.1;
[0137] Example 2
[0138] BiasedAD is implemented on the fashion domain image dataset FMNIST2. The tested AUPRC: 76.29 ± 1.02; the tested AUROC: 92.54 ± 0.74;
[0139] BiasedAD^M is implemented on the fashion domain image dataset FMNIST2. The tested AUPRC: 74.43 ± 1.62; the tested AUROC: 91.97 ± 0.84;
[0140] Comparative Example 1
[0141] The best-performing existing model Deep SAD is implemented on the network intrusion dataset UNSW_NB15. The tested AUPRC: 72.24 ± 1.04; the tested AUROC: 96.35 ± 0.11;
[0142] Comparative Example 2
[0143] The best-performing existing model Deep SAD is implemented on the fashion domain image dataset FMNIST2. The tested AUPRC: 64.58 ± 4.68; the tested AUROC: 89.0 ± 1.69;
[0144] The protection scope of the present invention is not limited to the above embodiments. Without departing from the spirit and scope of the inventive concept, changes and advantages that can be conceived by those skilled in the art are included in the present invention, and the scope of protection is defined by the appended claims.
Claims
1. A distance-based dual-center bias anomaly detection method, characterized in that, it includes the following steps: Step 1: Identify the objects to be recognized, and divide the objects to be distinguished and detected into normal samples, marginal samples, non-target anomaly samples, and target anomaly samples; Step 2: Construct a data set, where the data set contains an unlabeled data set D u , a labeled target anomaly data set D t , and an auxiliary data set D a ; Step 3: Model pre-training and initialization. Use an autoencoder to perform pre-training on the unlabeled dataset D u and initialize the bias anomaly detection model with the network weights after the autoencoder converges; Step 4: Optimize and iterate the bias anomaly detection model, and train the neural network Φ to effectively compress the normal samples and auxiliary samples in a non-target hypersphere in the latent space, and at the same time push the target anomaly samples away from the non-target hypersphere; Step 5: Calculate the anomaly score to obtain the anomaly test result.
2. The anomaly detection method according to claim 1, characterized in that, in Step 1, the normal sample refers to a sample that does not show any abnormal behavior; the marginal sample refers to a sample existing on the edge of the normal sample, including noise or normal samples that are difficult to correctly classify from unlabeled data; the non-target anomaly sample refers to an anomaly sample that does not belong to the detection target of the bias anomaly detection model; the target anomaly sample refers to an anomaly sample that belongs to the detection target of the bias anomaly detection model.
3. The anomaly detection method according to claim 1, characterized in that, In step 2, the auxiliary dataset D a is set according to two different semi-supervised data scenarios; if labeled non-target abnormal samples can be obtained, then the labeled non-target abnormal samples are used as the auxiliary dataset D a ; if labeled non-target abnormal samples cannot be obtained, then the edge samples selected from the unlabeled dataset D u are used as the auxiliary dataset D a .
4. The anomaly detection method according to claim 1, characterized in that, in Step 3, the autoencoder includes an encoder E and a decoder D; the encoder E is used to map the high-dimensional original data into low-dimensional latent variables, and the decoder D maps the low-dimensional latent variables into a vector with the same dimension as the original high-dimensional data; the autoencoder is trained by the reconstruction error between the vectors before and after encoding; after the autoencoder converges, use the network weights of the encoder E to initialize the neural network Φ(·; W) with the same structure as the encoder E as the bias anomaly detection model; The biased anomaly detection model includes the model BiasedAD and the model variant BiasedAD^M; the BiasedAD model is used for scenarios where labeled non-target anomaly samples can be obtained, and the labeled non-target anomaly samples are used as the auxiliary dataset D a ; BiasedAD^M is used for scenarios where labeled non-target anomaly samples cannot be obtained, and the marginal samples selected from the unlabeled dataset D u are used as the auxiliary dataset D a .
5. The anomaly detection method according to claim 1, characterized in that, in Step 4, the bias anomaly detection model includes BiasedAD and BiasedAD^M; the objective functions of the models BiasedAD and BiasedAD^M are as follows: Among them, L compact guides Φ to minimize the volume of the hypersphere diffused by normal samples and auxiliary samples in the latent space, and L target guides Φ to increase the distance between the target abnormal samples and this hypersphere; Step 4 further includes the following steps: Step 4-1: Use L compact to concentrate the normal samples and non-target abnormal samples in the non-target hypersphere; Step 4-2: Use auxiliary samples to enhance the learning of the representation of the target anomaly samples; Step 4-3: Dynamically update the anchor point a, so that all samples in the labeled target anomaly set D t gradually move away from c, and gradually transfer from the difficult-to-recognize set D hard to the easy-to-recognize set D easy ; Step 4-4: Repeat Step 4-1 to Step 4-3 until the overall objective function of the bias anomaly detection model no longer decreases and the model converges.
6. The anomaly detection method according to claim 5, characterized in that, In step 4-1, the L compact is expressed as follows: Among them, D u represents the unlabeled data set, D a represents the auxiliary data set, Φ is the output of the neural network, x i represents the i-th sample, W represents the weights of the neural network, c represents the center point of the normal samples, η 0 is the weight parameter; Step 4-1 further includes the following steps: Step 4-1-1: Penalize the distance from the learned representation of all samples in D u to c in the latent space, so that Φ maps normal samples to the vicinity of c in the latent space; c is obtained by averaging the output of the first forward pass of Φ on D u as shown in the following formula: Among them, D u represents the unlabeled dataset, Φ is the output of the neural network, and x i represents the i-th sample, and W represents the weights of the neural network; Step 4-1-2: By introducing η 0 Adjust and control the influence balance between the unlabeled samples and the auxiliary samples, so that Φ emphasizes the learning of the auxiliary samples.
7. The anomaly detection method according to claim 5, characterized in that, Step 4-2 further includes the following steps: Step 4-2-1: Define and initialize the anchor point a, and the anchor point a is represented as follows: Among them, D u represents the unlabeled data set, D a represents the auxiliary data set, Φ is the output of the neural network, x i represents the i-th sample, W represents the weights of the neural network; the anchor point a represents the average position of the auxiliary samples in the latent space, and is initially set to the average output of the first forward pass of Φ on D a ; Step 4-2-2: Design the loss function L target , and use the anchor point a to separate the target abnormal samples from other samples; The loss function L target is expressed by the following formula: Among them, D easy represents the easily recognizable target abnormal sample set, and D hard represents the difficult-to-recognize target abnormal sample set. Φ is the output of the neural network, and x k represents the k-th sample, W represents the weights of the neural network, c represents the center point of the normal samples, a represents the anchor point, η 1 and η 2 are used to adjust the weights acting on D easy and D hard in two parts; The said D easy and D hard represent two types of sets that are divided in real time according to their positions in the latent space. The said D t refers to the set of difficult-to-recognize samples inside the hypersphere centered at c with a radius equal to the distance between c and a. The said D hard refers to the set of easy-to-recognize samples outside the hypersphere centered at c with a radius equal to the distance between c and a; The said D easy refers to the set of easy-to-recognize samples outside the hypersphere centered at c with a radius equal to the distance between c and a; The said D easy and D hard are respectively expressed as follows: Among them, D easy represents the set of easily recognizable target abnormal samples, and D hard represents the set of difficult-to-recognize target abnormal samples, Φ is the output of the neural network, x represents the sample, W represents the weights of the neural network, c represents the center point of the normal samples, a represents the anchor point, dist(·,·) represents the distance between two samples in the latent space, D t represents the labeled target abnormal data set; For the target samples belonging to D easy , Φ imposes penalties on the reciprocals of their distances to the center point c of the normal samples and the anchor point a; for the target samples belonging to D hard , Φ only imposes a penalty on the reciprocal of the distance to the center c.
8. The anomaly detection method according to claim 1, characterized in that, in Step 5, the anomaly score is calculated by the following formula: Among them, s(x) represents the distance from the sample x after being mapped by the neural network Φ to the center point c of the normal samples in the latent space; Φ is the output of the neural network, x represents the sample, W represents the weights of the neural network, c represents the center point of the normal samples, and ||·|| 2 represents the L2 distance in the latent space, and D t represents the labeled target abnormal data set.
9. A detection system for implementing the anomaly detection method according to any one of claims 1-8, characterized in that, the detection system includes: a data pre-training module, a bias target optimization module, and an anomaly score calculation module; The model pre-training module is used for the mutual mapping between high-dimensional and low-dimensional data, and the initialization of the bias anomaly detection model after the convergence of the auto-encoder in pre-training; The bias target optimization module is used to separate target anomaly samples from other samples in the latent space; The anomaly score calculation module is used to calculate the anomaly score after the bias anomaly detection model converges, and complete the bias anomaly detection.
10. The application of the anomaly detection method according to any one of claims 1-8, or the detection system according to claim 9 in financial data anomaly detection, medical data anomaly detection, image data anomaly detection, and network intrusion tabular data anomaly detection.