Anomaly Detection Method, System and Computer Storage Medium Based on Contrastive Learning

Through comparative learning, designing positive and negative sample pairs and comparative loss training models, the problem of insufficient extraction of abnormal detection features in the prior art is solved, and more efficient exception data distinction and detection are achieved.

CN114330572BActive Publication Date: 2025-07-08HARBIN INST OF TECH SHENZHEN GRADUATE SCHOOL
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111666302.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-31
Publication Date
2025-07-08
Estimated Expiration
2041-12-31

AI Technical Summary

Technical Problem

The prior art is prone to losing important information when abnormal detection is unsupervised in unsupervised scenarios, and deep learning-based methods are difficult to effectively distinguish features when normal samples are close to abnormal samples, ignoring the guidance and feedback of known abnormal data.

Method used

Using contrast learning method, positive and negative sample pairs are designed and the anomaly detection model is trained through contrast loss, features with high distinction are extracted, and abnormal data are distinguished in abstract feature space using contrast learning.

Benefits of technology

Differentiated features are extracted in the feature space, and the abnormality scores output by the network are highly distinguished, which improves the data set detection effect in real life.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114330572B_ABST
    Figure CN114330572B_ABST
Patent Text Reader

Abstract

The present invention proposes an anomaly detection method, system and computer storage medium based on contrastive learning. The method includes an anomaly detection model training stage and an anomaly detection stage. In the anomaly detection model training stage, the feature vectors of the input samples are extracted, and the feature vectors are discriminated. According to the discrimination results, the contrastive loss of the anomaly detection model is calculated, and the anomaly detection model is trained using the contrastive loss. In the anomaly detection stage, the samples in the sample set to be detected are input into the trained anomaly detection model, and the discrimination results output are calculated to obtain anomaly scores. The anomaly scores of all samples are normalized to obtain normalized anomaly scores. By setting the normalized anomaly score threshold, it is determined whether the sample is abnormal. The present invention extracts discriminative features in the feature space, and the discriminatively output anomaly scores have high discrimination. There is a significant improvement compared with other methods in the anomaly detection of datasets in real life.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of anomaly detection, and particularly to an anomaly detection method, system and computer storage medium based on contrastive learning. Background Art

[0002] Anomaly occurs when a system or entity behaves abnormally. Therefore, anomalies contain a lot of useful information to reflect the abnormal characteristics of the system or entity. By observing anomalies, we can study the system from another perspective. Therefore, anomaly detection is of great significance and has wide applications in real life. For example:

[0003] (1) Intrusion detection system: In many computer systems, different types of data are generated by the operating system, network operation or user operation. When a malicious intrusion occurs, these data will become abnormal. By using anomaly detection algorithms to find the anomalies in the data, this illegal intrusion can be detected, providing a new solution for detecting network intrusions.

[0004] (2) Bank card fraud: Due to the theft of some sensitive information such as bank card numbers, more and more bank card fraud incidents are occurring now. Illegal use of others' bank cards will show some different information, such as consumption at special locations or large - amount consumption. These information can be used for anomaly detection in bank card transaction data to prevent further losses.

[0005] (3) Sensor anomaly: Abnormal data generated by factory sensors may indicate that a part of the system has failed. Timely discovery and maintenance of the failure can reduce the losses of the factory. In real life, it is also very meaningful to study the anomaly of sensor data for detecting the environment.

[0006] (4) Medical diagnosis: In many medical applications, data comes from various devices, such as various data from magnetic resonance imaging scans, positron emission tomography scans or electrocardiograms (ECGs). Unusual patterns in these data usually reflect disease conditions. Detecting these abnormal data helps doctors make disease judgments.

[0007] (5) Earth science: A large amount of data related to weather changes, climate transitions or land cover can be collected through devices such as satellites or remote sensing. Anomalies in these data provide important insights into possible causes of human activities or environmental trends that may have caused the change. Experts may obtain some unexpected results by studying these data.

[0008] Therefore, it is of great significance to detect abnormal data contained in the data in a timely manner. Anomalies are "few and different". Compared with normal data, the proportion of abnormal data is very low, and anomalies vary in different scenarios and are complex. Although traditional machine learning anomaly detection algorithms can detect anomalies in unsupervised scenarios, they are prone to losing important information in the time series data representation stage and have high complexity when calculating the distance between sequences. In recent years, a large number of anomaly detection algorithms based on deep learning have emerged. Although they can better utilize neural networks to extract more complex deep features, they also rely on a large number of normal data samples for model training. Most deep learning-based anomaly detection algorithms only focus on extracting deep features of normal data during training. When normal samples are similar to abnormal samples, the extracted features cannot effectively distinguish between normal and abnormal. In actual application scenarios, in fact, some abnormal data is known, but most existing anomaly detection algorithms do not consider using this prior knowledge and ignore the guidance and feedback of abnormal data on the model. Summary of the Invention

[0009] In view of the above problems, the present invention provides an anomaly detection method, system and computer storage medium based on contrast learning. Taking contrast learning as the starting point and integrating abnormal data into the training of the anomaly detection model, an anomaly detection method based on contrast learning is proposed. The research is carried out from two aspects: positive and negative sample pairs and contrast loss. Among them, the design of positive and negative sample pairs will make more full use of scenario-related abnormal data; in the design of contrast loss, the direct goal is to generate highly discriminative anomaly scores, guiding the anomaly detection model to extract more discriminative features, ensuring that when the anomaly detection model performs anomaly detection in the abstract feature space, it can more easily distinguish abnormal data.

[0010] In the first aspect of the present invention, an anomaly detection method based on contrast learning is provided. The method includes the following steps:

[0011] Anomaly detection model training stage: The anomaly detection model includes a feature extraction encoder module and a discriminator module. Among them, the feature extraction encoder module is used to extract the feature vector of the input sample, and the discriminator module is used to discriminate the feature vector output by the feature extraction module and output a discrimination result; calculate the contrast loss of the anomaly detection model according to the discrimination result, and use the contrast loss to train the anomaly detection model;

[0012] Anomaly detection stage: Input the samples in the sample set to be detected into the trained anomaly detection model, input the feature vector obtained by the feature extraction encoder module into the discriminator module, calculate the discrimination result output by the discriminator module to obtain an anomaly score; normalize the anomaly scores of all samples in the sample set to be detected to obtain a normalized anomaly score, and determine whether the sample is abnormal by setting a normalized anomaly score threshold.

[0013] Furthermore, the feature extraction encoder module is used to extract the feature vector of the input sample, specifically including:

[0014] The feature extraction encoder module includes two linear fully connected networks and an encoder. After the input sample is output by the encoder, the two linear fully connected networks calculate the mean and variance of the encoder output respectively, and add random noise to the mean and variance to obtain the feature vector.

[0015] Furthermore, the discriminator module includes three parts: discriminant network D xz , discriminant network D zz , discriminant network D zf , where the discriminant network D xz discriminates the sample pair composed of (x, z); the discriminant network D zz discriminates the sample pair composed of (z, z); the discriminant network D zf discriminates the sample pair composed of (z, f). x represents the input sample, z represents the feature vector output by the feature extraction encoder module, and f represents the encoder output vector in the feature extraction encoder module.

[0016] Furthermore, in the discriminator module, the normal sample and the feature vector corresponding to the normal sample are composed into a positive sample pair, and the normal sample and the feature vector corresponding to the abnormal sample are composed into a negative sample pair. The output probability corresponding to the positive sample pair (x xz , z nor , nor ) in the discriminant network D nor is set to 1, and the output probability corresponding to the negative sample pair (x ano , z zz ) is set to 0; the output probability corresponding to the positive sample pair (z nor , z nor ) in the discriminant network D nor is set to 1, and the output probability corresponding to the negative sample pair (z ano , z zf ) is set to 0; the output probability corresponding to the positive sample pair (z nor , f nor ) in the discriminant network D nor is set to 1, and the output probability corresponding to the negative sample pair (z ano , f nor ) is set to 0; where x nor represents the normal sample, z ano represents the feature vector of the normal sample, z nor represents the encoder output vector corresponding to the normal sample, and f ano represents the encoder output vector corresponding to the abnormal sample. ​

[0017] Furthermore, the specific expression of the anomaly score in the anomaly detection stage is as follows:

[0018] Anoscore i =-(D xz (x i ,z i )+D zz (z i ,z i )+D zf (z i ,f i ))), where D xz (x i ,z i ), D zz (z i ,z i ), D zf (z i ,f i )

[0019] are the outputs of discriminant network D xz , discriminant network D zz , and discriminant network D zf for sample x i respectively.

[0020] Furthermore, the normalized anomaly score Anoscore i 's specific expression is:

[0021] where min(Anoscore) is the minimum anomaly score in the sample set to be detected, and max(Anoscore) is the maximum anomaly score in the sample set to be detected.

[0022] In the second aspect of the present invention, an anomaly detection system based on contrast learning is provided. The system includes:

[0023] An anomaly detection model training unit for training the anomaly detection model. The anomaly detection model includes a feature extraction encoder module and a discriminator module. Among them, the feature extraction encoder module is used to extract the feature vector of the input sample, and the discriminator is used to discriminate the feature vector output by the feature extraction module and output a discrimination result; calculate the contrast loss of the anomaly detection model according to the discrimination result, and use the contrast loss to train the anomaly detection model;

[0024] Anomaly detection unit, which is used to perform anomaly detection on samples in the sample set to be detected. Input the samples in the sample set to be detected into the trained anomaly detection model, input the feature vectors obtained by the feature extraction encoder module into the discriminator module, calculate the discrimination results output by the discriminator module to obtain anomaly scores; normalize the anomaly scores of all samples in the sample set to be detected to obtain normalized anomaly scores, and determine whether the samples are abnormal by setting the normalized anomaly score threshold.

[0025] Further, the feature extraction encoder module includes two linear fully connected networks and an encoder. After the input sample is output by the encoder, the two linear fully connected networks calculate the mean and variance of the encoder output respectively, and add random noise to the mean and variance to obtain feature vectors.

[0026] In the third aspect of the present invention, there is provided an anomaly detection system based on contrast learning, including: a processor; and a memory, wherein, computer-executable programs are stored in the memory, and when the computer-executable programs are executed by the processor, the above-mentioned anomaly detection method based on contrast learning is executed.

[0027] In the fourth aspect of the present invention, there is provided a computer-readable storage medium, on which instructions are stored, and when the instructions are executed by a processor, the processor executes the above-mentioned anomaly detection method based on contrast learning.

[0028] An anomaly detection method, system and computer storage medium based on contrast learning provided by the present invention focus on distinguishing data in an abstract feature space, and the model optimization is simpler. This algorithm makes full use of scenario-related abnormal data when designing positive and negative sample pairs; in the model design, a contrast loss is directly designed with the goal of generating highly discriminative anomaly scores to guide the feature extraction network in the model to extract discriminative features, and at the same time ensure that the discriminant network can well distinguish the feature vectors corresponding to normal samples and the feature vectors corresponding to abnormal samples, so as to generate highly discriminative anomaly scores. The finally achieved beneficial effects are: compared with existing anomaly detection methods, the anomaly detection method, system and computer storage medium based on contrast learning provided by the present invention can extract discriminative features in the feature space, and the anomaly scores output by the discriminant network have high discriminability, and there is a large improvement compared with other methods on the data sets in real life, and it has great practical value. Description of the Drawings

[0029] Figure 1 It is a flowchart for training an anomaly detection model in the anomaly detection method based on contrast learning according to an embodiment of the present invention;

[0030] Figure 2 It is a schematic structural diagram of the anomaly detection system based on contrast learning in an embodiment of the present invention;

[0031] Figure 3 is the architecture of the computer device in the embodiments of the present invention. Detailed implementation manners

[0032] To further elaborate on the technical solution of the present invention in detail, this embodiment is implemented on the premise of the technical solution of the present invention, and detailed implementation manners and specific steps are given.

[0033] Embodiments of the present invention are directed to an anomaly detection method, system, and computer storage medium based on contrastive learning. The following embodiments are provided:

[0034] Based on Embodiment 1 of the present invention

[0035] This embodiment is used to illustrate the principle of how the present invention solves technical problems and the training steps of the anomaly detection model. Under the guidance of the contrastive learning idea, how to enable the anomaly detection model to learn more representative features for normal samples, while abnormal samples cannot extract such features. Therefore, the feature representations between the two will have "high distinguishability", making it easier for the anomaly detection model to detect abnormal samples. Good features should be able to represent the "unique" information of the sample. Mutual information can be used to measure whether the extracted information is unique to this type of sample: Mutual information can measure the mutual dependence between two random variables. The greater the mutual information, the higher the correlation between the two.

[0036] Assume that X is the set of original data, x ∈ X represents a certain data sample, Z is the set of feature vectors extracted by the model, and z ∈ Z represents a certain feature vector; represents the distribution of x, and p(z|x) represents the distribution of the feature vector corresponding to x; represents the distribution of the entire Z given p(z|x), then the mutual information I(X, Z) between X and Z is as follows;

[0037]

[0038] I(X, Z) represents the relative entropy of the product of the joint distribution and the independent distribution of the corresponding distributions of x and z. Maximizing the mutual information is to maximize and The KL divergence between them. In addition, it is necessary to constrain the coding space where the feature vectors are located, because a scattered coding space often reflects overfitting of the model, which learns the training samples too well and may have poor actual test results; while a regular and continuous coding space is more likely to generalize to unknown new samples and improve the model's processing ability for unknown new samples. Therefore, the present invention introduces a prior distribution p(z) of a standard normal distribution and constrains z to follow this prior distribution by minimizing the KL divergence between p(z) and q(z). Considering the above two parts of constraints comprehensively, the loss function is shown in the following formula, where λ, α, and β are all weights that can be set independently, where λ ∈ (0, 1), α ∈ (0, 1), β ∈ (0, 1), and α + β = 1.

[0039]

[0040] As Figure 1 shown, it is the flowchart of abnormal detection model training in the abnormal detection method based on contrast learning according to the embodiment of the present invention. The specific steps include the abnormal detection model training stage and the abnormal detection stage, where the abnormal detection model training stage:

[0041] The abnormal detection model includes a feature extraction encoder module and a discriminator module. Among them, the feature extraction encoder module is used to extract the feature vectors of the input samples, and the discriminator module is used to discriminate the feature vectors output by the feature extraction module and output the discrimination results; calculate the contrast loss of the abnormal detection model according to the discrimination results, and use the contrast loss to train the abnormal detection model.

[0042] Furthermore, the feature extraction encoder module is used to extract the feature vectors of the input samples. The feature extraction encoder module maps the input samples to the latent space, and the low-dimensional vectors in the latent space are the extracted feature vectors. Specifically, it includes: the feature extraction encoder module includes two linear fully connected networks and an encoder. After the input samples are output by the encoder, the two linear fully connected networks calculate the mean and variance of the encoder output respectively, and add random noise to the mean and variance to obtain the feature vectors.

[0043] In the specific implementation process, the feature extraction encoder module maps the input sample x to a latent vector z containing important information, providing data for the subsequent design of positive and negative sample pairs. At the same time, the encoder should ensure that the coding space where the feature vectors are located is regular and continuous. Drawing on the reparameterization technique in the variational autoencoder VAE, the inventor adds two linear fully connected networks to calculate the mean μ and variance σ of the output f of the encoder respectively, and then adds random noise to it to convert f into the final feature vector z. Therefore, the corresponding KL divergence constraint can be converted into the following formula loss KL , where M is the dimension of the feature vector z corresponding to the normal samples nor_i and ∈ is the random noise, μnor_i and σ nor_i are the mean and variance corresponding to the feature vector z nor_i .

[0044] z = μ + σ·∈, (∈ ∼ N(0, 1))

[0045]

[0046] Furthermore, the discriminator module includes three parts: the discriminant network D xz , the discriminant network D zz , the discriminant network D zf , where the discriminant network D xz discriminates the sample pair composed of (x, z); the discriminant network D zz discriminates the sample pair composed of (z, z); the discriminant network D zf discriminates the sample pair composed of (z, f), x represents the input sample, z represents the feature vector output by the feature extraction encoder module, and f represents the encoder output vector in the feature extraction encoder module.

[0047] Furthermore, since good positive and negative sample pairs can help the anomaly detection model extract the key features of the samples more effectively during training. Because contrastive learning can learn the common features between similar instances and distinguish non-similar instances, and for anomaly detection, normal samples and normal samples can be divided into similar instances, and normal samples and abnormal samples can be divided into non-similar. At the same time, it is also hoped that the correlation between the extracted feature vector and the sample is relatively high, and it is easier to distinguish between normal samples and abnormal samples under this feature representation. Therefore, in the discriminator module, the feature vectors corresponding to normal samples and normal samples are composed into positive sample pairs, and the feature vectors corresponding to normal samples and abnormal samples are composed into negative sample pairs. The output probability of the positive sample pair (x xz , z nor ) in the discriminant network D nor is set to 1, and the output probability of the negative sample pair (x nor , z ano ) is set to 0; the output probability of the positive sample pair (z zz , z nor ) in the discriminant network D nor is set to 1, and the output probability of the negative sample pair (z nor , z ano ) is set to 0; the output probability of the positive sample pair (z zf , f nor ) in the discriminant network D nor is set to 1, and the output probability of the negative sample pair (z nor , f ano ) is set to 0; where, xnor Denote the normal sample as z nor Denote the feature vector of the normal sample as z ano Denote the feature vector of the abnormal sample as f nor Denote the encoder output vector corresponding to the normal sample as f ano Denote the encoder output vector corresponding to the abnormal sample.

[0048] In the specific implementation process, the discriminator module is used to discriminate the feature vector and output the corresponding result. For the positive sample pair, the discriminator module should output a higher value, while for the negative sample pair, it outputs a lower value. According to this result, the contrast loss can be calculated for updating the parameters of the anomaly detection model. Among them, the discriminant network D xz Determine the sample pair composed of (x, z) and output the corresponding probability score. Among them, for the positive sample pair (x nor , z nor ), the output probability D xz (x nor , z nor ) is set to 1, and for the negative sample pair (x nor , z ano ), the output probability D xz (x nor , z ano ) is set to 0. The corresponding constraint is as follows:

[0049] loss Dxz = E[logD xz (x nor , z nor )] + E[log(1 - D xz (x nor , z ano ))]

[0050] The discriminant network D zz Determine the sample pair composed of (z, z) and output the corresponding probability score. Among them, for the positive sample pair (z nor , z nor ), the output probability D zz (z nor , z nor ) is set to 1; for the negative sample pair (z nor , z ano ), the output probability D zz (z nor , z ano ) is set to 0. The corresponding constraint is as follows:

[0051] loss Dzz = E[logD zz (z nor , znor )] + E[log(1 - D zz (z nor , z ano ))]

[0052] Discriminative network D zf Determine the sample pairs composed of (z, f), and output the corresponding probability scores. Among them, for the positive sample pairs (z nor , f nor ), the corresponding output probability D zf (z nor , f nor ) is set to 1, and for the negative sample pairs (z nor , f ano ), the corresponding output probability D zf (z nor , f ano ) is set to 0. The corresponding constraints are as follows:

[0053] loss Dzf = E[logD zf (z nor , f nor )] + E[log(1 - D zf (z nor , f ano ))]

[0054] Furthermore, according to the discriminators corresponding to the positive and negative sample pairs input into the discriminative network, calculate the corresponding probabilities. Finally, according to the contrast loss loss CLAD calculate the loss corresponding to the entire model. With the goal of minimizing the contrast loss, update and optimize the parameters of the anomaly detection model. In the contrast loss loss CLAD formula, both α and β are weights that can be set independently. In the preferred embodiment, α and β evenly divide the weight 1 and are 0.5 respectively. CLAD In the formula loss

[0055] loss CLAD = -α · [loss Dxz + loss Dzz + loss Dzf + β · loss KL

[0056] Based on Embodiment 2 of the present invention

[0057] This embodiment provides the specific implementation steps in the anomaly detection phase based on the anomaly detection model trained in Embodiment 1. The samples in the sample set to be detected are input into the trained anomaly detection model, the feature vectors obtained by the feature extraction encoder module are input into the discriminator module, and the discrimination results output by the discriminator module are calculated to obtain an anomaly score. The anomaly scores of all samples in the sample set to be detected are normalized to obtain a normalized anomaly score, and whether the sample is abnormal is determined by setting a normalized anomaly score threshold.

[0058] In the specific implementation process, the sample x in the sample set to be detected i is input into the trained anomaly detection model, and the corresponding feature vector z is obtained through the feature extraction encoder module i , and then three discriminant networks, discriminant network D xz , discriminant network D zz , and discriminant network D zf will output the corresponding probabilities. The larger the probability, the more the extracted features tend to be the features unique to normal samples. Then the probabilities are added and negated to obtain the corresponding anomaly score. The higher the anomaly score, the more likely the sample is an abnormal sample. The specific expression of the anomaly score in the anomaly detection phase is:

[0059] Anoscore i = -(D xz (x i , z i ) + D zz (z i , z i ) + D zf (z i , f i ))), where D xz (x i , z i ), D zz (z i , z i ), D zf (z i , f i )

[0060] are the outputs of discriminant network D xz , discriminant network D zz , and discriminant network D zf for the sample x i .

[0061] Furthermore, the anomaly scores of all samples in the sample set to be detected are normalized to the interval [0, 1]. The more abnormal the data, the closer the anomaly score is to 1. The specific expression of the normalized anomaly score Anoscore i ' is:

[0062] Where min(Anoscore) is the minimum anomaly score in the sample set to be detected, and max(Anoscore) is the maximum anomaly score in the sample set to be detected.

[0063] By setting a threshold As long as it is considered that the sample is abnormal. Preferably,

[0064] Based on Embodiment 3 of the present invention

[0065] Hereinafter, with reference to Figure 2 The system corresponding to the methods according to Embodiments 1 and 2 of the present disclosure will be described. An anomaly detection system based on contrast learning, System 100 includes: an anomaly detection model training unit 101 for training an anomaly detection model. The anomaly detection model includes a feature extraction encoder module and a discriminator module. Among them, the feature extraction encoder module is used to extract the feature vector of the input sample, and the discriminator is used to discriminate the feature vector output by the feature extraction module and output a discrimination result; calculate the contrast loss of the anomaly detection model according to the discrimination result, and use the contrast loss to train the anomaly detection model; an anomaly detection unit 102 for performing anomaly detection on the samples in the sample set to be detected, inputting the samples in the sample set to be detected into the trained anomaly detection model, inputting the feature vector obtained by the feature extraction encoder module into the discriminator module, calculating the discrimination result output by the discriminator module to obtain an anomaly score; normalizing the anomaly scores of all samples in the sample set to be measured to obtain a normalized anomaly score, and determining whether the sample is abnormal by setting a normalized anomaly score threshold. In addition to the above two units, System 100 may further include other components. However, since these components are not related to the content of the embodiments of the present disclosure, their illustrations and descriptions are omitted here.

[0066] Furthermore, the feature extraction encoder module includes two linear fully connected networks and an encoder. After the input sample is output by the encoder, the two linear fully connected networks calculate the mean and variance of the encoder output respectively, and add random noise to the mean and variance to obtain the feature vector.

[0067] The specific working process of an anomaly detection system 100 based on contrast learning refers to the descriptions of Embodiments 1 and 2 of the above-mentioned anomaly detection method based on contrast learning, and will not be elaborated here.

[0068] Based on Embodiment 4 of the present invention

[0069] The device according to the embodiments of the present invention can also be implemented by means of Figure 3 the architecture of the computing device shown. Figure 3shows the architecture of the computing device. As Figure 3 shown, computer system 201, system bus 203, one or more CPUs 204, input / output 202, memory 205, etc. Memory 205 can store various data or files used for computer processing and / or communication, as well as program instructions executed by the CPU, including the methods of Embodiment 1 - Embodiment 2. Figure 3 The architecture shown is only exemplary. When implementing different devices, one or more components in Figure 3 are adjusted according to actual needs.

[0070] Based on Embodiment 5 of the present invention

[0071] Embodiments of the present invention can also be implemented as a computer-readable storage medium. Computer-readable instructions are stored on the computer-readable storage medium according to Embodiment 5. When the computer-readable instructions are run by a processor, the contrastive learning-based anomaly detection method according to Embodiments 1 and 2 of the present invention described with reference to the above figures can be executed.

[0072] Embodiments of the present invention are directed to the above-described embodiments of the contrastive learning-based anomaly detection method, system embodiments, and computer storage medium embodiments. The results of the above 5 embodiments are compared with the performance of the current optimal anomaly detection algorithms. The embodiments are carried out on data from two UCR public datasets, the MIT-BIH public dataset, and the BIDMC public dataset. The anomalies in the dataset include one or more. Each set of electrocardiogram data records a person's heartbeat activity and includes one or more types of abnormal heartbeats. Sensor data comes from different sensor devices, such as radio frequency (RF) instruments and accelerometers. Action data is data collected from different activities, such as walking and gun-holding actions. Image data is sequential data extracted from images.

[0073] A comparison of the performance of the anomaly detection method of the present invention and other methods is shown in Table 1. It can be seen that the contrastive learning-based anomaly detection method proposed by the present invention performs excellently overall: 1) It obtains the highest AUC-ROC on 13 / 15 datasets and is the second highest in terms of AUC-ROC score on the remaining 2 datasets; 2) In terms of the average AUC-ROC, the method proposed by the present invention has the highest average score, and there is a performance improvement of nearly 8.5% compared to the highest average value among other comparison algorithms.

[0074] Table 1 AUC-ROC scores of the method Ours proposed by the present invention and the comparison methods (bold for the first, underlined for the second)

[0075]

[0076] Combining the above-provided anomaly detection method, system, and computer storage medium based on contrastive learning, it is possible to extract discriminative features, and the anomaly scores output by the discriminant network have high discriminability, showing a significant improvement compared to other algorithms on real-life datasets and having great practical value.

[0077] In this article, the terms "including", "comprising", or any other variant thereof are intended to cover non-exclusive inclusion, such that a step or method including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such a step or method.

[0078] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as falling within the protection scope of the present invention.

Claims

1. An anomaly detection method based on contrastive learning, characterized in that The method includes the following steps: Abnormal detection model training stage: The abnormal detection model includes a feature extraction encoder module and a discriminator module. Among them, the feature extraction encoder module is used to extract the feature vector of the input sample, and the discriminator module is used to discriminate the feature vector output by the feature extraction module and output the discrimination result; calculate the contrast loss of the abnormal detection model according to the discrimination result, and use the contrast loss to train the abnormal detection model; Abnormal detection stage: Input the samples in the sample set to be detected into the trained abnormal detection model, input the feature vector obtained by the feature extraction encoder module into the discriminator module, calculate the discrimination result output by the discriminator module to obtain the abnormal score; normalize the abnormal scores of all samples in the sample set to be detected to obtain the normalized abnormal score, and determine whether the sample is abnormal by setting the normalized abnormal score threshold; The feature extraction encoder module is used to extract the feature vector of the input sample, specifically including: The feature extraction encoder module includes two linear fully connected networks and an encoder. After the input sample is output by the encoder, the two linear fully connected networks calculate the mean and variance of the encoder output respectively, and add random noise to the mean and variance to obtain the feature vector; The discriminator module consists of three parts: discriminant network D xz , discriminant network D zz , discriminant network D zf , where discriminant network D xz discriminates the sample pairs composed of (x, z); discriminant network D zz discriminates the sample pairs composed of (z, z); discriminant network D zf discriminates the sample pairs composed of (z, f), where x represents the input sample, z represents the feature vector output by the feature extraction encoder module, and f represents the encoder output vector in the feature extraction encoder module; The input sample and the sample to be detected are text data or image data.

2. The anomaly detection method according to claim 1, wherein In the discriminator module, positive sample pairs are formed by normal samples and their corresponding feature vectors, and negative sample pairs are formed by normal samples and the feature vectors of abnormal samples. For the discriminant network D xz the output probability corresponding to the positive sample pair (x nor , z nor ) is set to 1, and the output probability corresponding to the negative sample pair (x nor , z ano ) is set to 0; for the discriminant network D zz the output probability corresponding to the positive sample pair (z nor , z nor ) is set to 1, and the output probability corresponding to the negative sample pair (z nor , z ano ) is set to 0; for the discriminant network D zf the output probability corresponding to the positive sample pair (z nor , f nor ) is set to 1, and the output probability corresponding to the negative sample pair (z nor , f ano ) is set to 0; where x nor represents a normal sample, z nor represents the feature vector of a normal sample, z ano represents the feature vector of an abnormal sample, f nor represents the encoder output vector corresponding to a normal sample, and f ano represents the encoder output vector corresponding to an abnormal sample.

3. The anomaly detection method according to claim 2, wherein The specific expression for the anomaly score in the anomaly detection stage is: Anoscore i = -(D xz (x i , z i ) + D zz (z i , z i ) + D zf (z i , f i ))), where D xz (x i , z i ), D zz (z i , z i ), and D zf (z i , f i ) are the outputs of discriminative network D xz , discriminative network D zz , and discriminative network D zf for sample x i .

4. The anomaly detection method according to claim 3, wherein Normalization processing yields the normalized anomaly score Anoscore i The specific expression is as follows: where min(Anoscore) is the minimum anomaly score in the sample set to be detected, and max(Anoscore) is the maximum anomaly score in the sample set to be detected.

5. An anomaly detection system based on contrastive learning, characterized in that, The system includes: An abnormal detection model training unit, which is used to train the abnormal detection model. The abnormal detection model includes a feature extraction encoder module and a discriminator module. Among them, the feature extraction encoder module is used to extract the feature vector of the input sample, and the discriminator is used to discriminate the feature vector output by the feature extraction module and output the discrimination result; calculate the contrast loss of the abnormal detection model according to the discrimination result, and use the contrast loss to train the abnormal detection model; An abnormal detection unit, which is used to perform abnormal detection on the samples in the sample set to be detected. Input the samples in the sample set to be detected into the trained abnormal detection model, input the feature vector obtained by the feature extraction encoder module into the discriminator module, calculate the discrimination result output by the discriminator module to obtain the abnormal score; normalize the abnormal scores of all samples in the sample set to be detected to obtain the normalized abnormal score, and determine whether the sample is abnormal by setting the normalized abnormal score threshold; The feature extraction encoder module is used to extract the feature vector of the input sample, specifically including: The feature extraction encoder module includes two linear fully connected networks and an encoder. After the input sample is output by the encoder, the two linear fully connected networks calculate the mean and variance of the encoder output respectively, and add random noise to the mean and variance to obtain the feature vector; The discriminator module consists of three parts: discriminant network D xz , discriminant network D zz , discriminant network D zf , where discriminant network D xz discriminates the sample pairs composed of (x, z); discriminant network D zz discriminates the sample pairs composed of (z, z); discriminant network D zf discriminates the sample pairs composed of (z, f), where x represents the input sample, z represents the feature vector output by the feature extraction encoder module, and f represents the encoder output vector in the feature extraction encoder module; The input sample and the sample to be detected are text data or image data.

6. An anomaly detection system based on contrastive learning, characterized in that, It includes: A processor; And a memory, wherein the memory stores a computer executable program, and when the processor executes the computer executable program, it executes the contrast learning-based abnormal detection method according to any one of claims 1-4.

7. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the contrast learning-based abnormal detection method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Traffic data anomaly detection method and device and storage medium

    CN112702329A

  • Sound anomaly detection method and device, computer equipment and storage medium

    CN113470695A

  • Anomaly detection method and device based on unsupervised learning and storage medium

    CN113643292A