Unsupervised anomaly detection method and device based on comparative learning
By adopting a dual-task learning model based on contrast learning in unsupervised anomaly detection, the problem of difficult to distinguish samples deviating from feature patterns in the prior art is solved, and a more efficient anomaly detection effect is achieved.
Patent Information
- Application Number
- CN202510011437.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-06
AI Technical Summary
During actual detection, it is difficult to distinguish whether samples deviate from known feature patterns are normal or abnormal, resulting in a large number of mispredictions.
Unsupervised anomaly detection method based on contrast learning is adopted, and multi-level features and semantic knowledge of training samples are captured in the training stage through the dual-task learning model, and model parameters are updated through contrast learning in the testing stage to improve the ability to recognize deviation patterns.
Improved the effectiveness of anomaly detection, allowing the dual-task learning model to better capture distinctions that deviate from the learned normal mode, thereby reducing mispredictions.
Smart Images

Figure CN119939464A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of anomaly detection, and in particular to an unsupervised anomaly detection method and device based on contrastive learning. Background Art
[0002] Anomaly detection is used to identify data points in the data set to be tested that are significantly deviated from the majority of instance points. It plays an important role in medical diagnosis, network intrusion detection, financial fraud detection, industrial fault detection, etc. Since it is very difficult to obtain annotation information for abnormal data in most practical scenarios, unsupervised anomaly detection that does not rely on abnormal labels for model training has become one of the most practical anomaly detection methods.
[0003] Existing unsupervised anomaly detection methods usually focus on learning feature patterns from normal samples during the training phase. However, it is impossible for limited training samples to contain all normal patterns. When the normal pattern in actual detection is significantly different from the learned normal pattern, it is difficult for the model to distinguish whether the samples that deviate from the known feature pattern are normal or abnormal, resulting in a large number of incorrect predictions.
[0004] Therefore, there is an urgent need to provide a new unsupervised anomaly detection method to improve the effectiveness of anomaly detection. Summary of the invention
[0005] The present invention provides an unsupervised anomaly detection method and device based on contrastive learning. The technical solution is as follows:
[0006] In one aspect, a method for unsupervised anomaly detection based on contrastive learning is provided, the method comprising:
[0007] Based on multiple training samples in the training sample set, an unsupervised training method is used to obtain a dual-task learning model that has been preliminarily trained; the training samples are tabular data;
[0008] Using the test samples in the test sample set, using the current dual-task learning model to calculate the loss of each test sample, selecting high-confidence normal samples and high-confidence abnormal samples from the test sample set according to the loss of the test samples, and comparing the high-confidence normal samples and high-confidence abnormal samples with the training samples, so as to update the model parameters of the current dual-task learning model through comparative learning, using the updated dual-task learning model as the current dual-task learning model, deleting the high-confidence normal samples and high-confidence abnormal samples selected this time from the test sample set, and continuing to perform this step using the remaining test sample sets until no high-confidence normal samples and high-confidence abnormal samples can be selected from the test sample set, and obtaining a trained dual-task learning model; the test samples are tabular data;
[0009] The trained dual-task learning model is used to perform anomaly detection on the input tabular data.
[0010] On the other hand, an unsupervised anomaly detection device based on contrastive learning is provided, the device comprising:
[0011] A preliminary training unit, used to obtain a dual-task learning model that has completed preliminary training by adopting an unsupervised training method based on multiple training samples in a training sample set; the training samples are tabular data;
[0012] A parameter updating unit is used to use the test samples in the test sample set and the current dual-task learning model to calculate the loss of each test sample, select high-confidence normal samples and high-confidence abnormal samples from the test sample set according to the loss of the test samples, and compare the high-confidence normal samples and high-confidence abnormal samples with the training samples to update the model parameters of the current dual-task learning model through comparative learning, use the updated dual-task learning model as the current dual-task learning model, delete the high-confidence normal samples and high-confidence abnormal samples selected this time from the test sample set, and continue to perform this step using the remaining test sample sets until no high-confidence normal samples and high-confidence abnormal samples can be selected from the test sample set, thereby obtaining a trained dual-task learning model; the test samples are tabular data;
[0013] The anomaly detection unit is used to perform anomaly detection on input table data using the trained dual-task learning model.
[0014] On the other hand, a computer device is provided, comprising a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory to implement the steps of the unsupervised anomaly detection method based on contrastive learning described above.
[0015] On the other hand, a computer-readable storage medium is provided, wherein a computer program is stored in the storage medium, and when the computer program is executed by a processor, the steps of the unsupervised anomaly detection method based on contrastive learning are implemented.
[0016] On the other hand, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the steps of the unsupervised anomaly detection method based on contrastive learning are implemented.
[0017] The technical solution provided by the present invention can at least bring the following beneficial effects:
[0018] In the training phase, the dual-task learning model is initially trained using a training sample set that contains all normal samples, so that the dual-task learning model can capture the multi-level features and semantic knowledge of the training samples; then, a test sample set that is different from the normal samples in the training sample set is used, and the dual-task learning model selects high-confidence normal samples and high-confidence abnormal samples from the test sample set based on the loss of the test samples, and compares the selected high-confidence normal samples and high-confidence abnormal samples with the training samples, so as to update the model parameters in the dual-task learning model through comparative learning, so that the dual-task learning model can further capture the distinction that deviates from the learned normal pattern, thereby improving the effectiveness of anomaly detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0020] Figure 1 is a flow chart of an unsupervised anomaly detection method based on contrastive learning provided by an embodiment of the present invention;
[0021] Figure 2 It is a schematic diagram of the architecture of the dual-task learning model and its input processing process provided by an embodiment of the present invention;
[0022] Figure 3 It is a schematic diagram of the processing process of the model parameter updating stage provided by an embodiment of the present invention;
[0023] Figure 4 is a structural diagram of an unsupervised anomaly detection device based on contrastive learning provided by an embodiment of the present invention;
[0024] Figure 5 It is a hardware architecture diagram of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0025] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0026] Please refer to Figure 1, an unsupervised anomaly detection method based on contrastive learning provided by an embodiment of the present invention, the method comprising:
[0027] Step 100, based on multiple training samples in the training sample set, an unsupervised training method is used to obtain a dual-task learning model that has been preliminarily trained; the training samples are tabular data;
[0028] Step 102, using the test samples in the test sample set, using the current dual-task learning model to calculate the loss of each test sample, selecting high-confidence normal samples and high-confidence abnormal samples from the test sample set according to the loss of the test samples, and comparing the high-confidence normal samples and high-confidence abnormal samples with the training samples, so as to update the model parameters of the current dual-task learning model through comparative learning, and using the updated dual-task learning model as the current dual-task learning model, deleting the high-confidence normal samples and high-confidence abnormal samples selected this time from the test sample set, and continuing to perform this step using the remaining test sample sets, until no high-confidence normal samples and high-confidence abnormal samples can be selected from the test sample set, and obtaining a trained dual-task learning model; the test samples are tabular data;
[0029] Step 104: Use the trained dual-task learning model to perform anomaly detection on the input table data.
[0030] In an embodiment of the present invention, a dual-task learning model is preliminarily trained using a training sample set that contains all normal samples during a training phase, so that the dual-task learning model can capture multi-level features and semantic knowledge of the training samples; then, a test sample set that has a pattern different from that of the normal samples in the training sample set is used, and the dual-task learning model selects high-confidence normal samples and high-confidence abnormal samples from the test sample set based on the loss of the test samples, and the selected high-confidence normal samples and high-confidence abnormal samples are compared with the training samples, so as to update the model parameters in the dual-task learning model through comparative learning, so that the dual-task learning model can further capture the distinction that deviates from the learned normal pattern, thereby improving the effectiveness of anomaly detection.
[0031] Described below Figure 1 How the various steps are performed.
[0032] First, for step 100, based on multiple training samples in the training sample set, an unsupervised training method is used to obtain a dual-task learning model that has been preliminarily trained.
[0033] In the embodiment of the present invention, the dual-task learning model includes two stages, one stage is a preliminary training stage using a training sample set, and the other stage is a stage for updating model parameters in the dual-task learning model using a comparative learning method using a test sample set and a training sample set.
[0034] In the embodiment of the present invention, since the dual-task learning model is trained by unsupervised training in the initial training stage, the training samples in the training sample set are all normal samples, so that the dual-task learning model learns the normal mode of normal samples.
[0035] The training sample is tabular data, which is row and column data. In one embodiment, each column may correspond to different parameters, and each row represents the full parameter value of a time node. The training sample may be tabular data in different fields, which may be medical diagnosis, network intrusion detection, financial fraud detection, industrial fault detection, etc. However, it should be noted that tabular data in different fields correspond to different training sample sets, and training samples in the same training sample set correspond to tabular data in the same field.
[0036] The execution subject of the present invention may be a terminal device, and the dual-task learning model is obtained by the terminal device.
[0037] In one embodiment of the present invention, a dual-task learning model is obtained by jointly training the main task and the auxiliary task. The main task reconstructs the original input sample, so that the dual-task learning model can learn the low-level feature representation of the input sample; the auxiliary task reconstructs the embedded representation of the input sample, providing high-level semantic information as supplementary guidance for the dual-task learning model.
[0038] Please refer to Figure 2 , is a schematic diagram of the architecture of the dual-task learning model and its input processing process in an embodiment of the present invention. The dual-task learning model includes: a mask encoder E and a reconstructor D and a multi-layer perceptron MLP respectively connected to the mask encoder; the mask encoder includes a mask generator and an encoder;
[0039] When the dual-task learning model that has been preliminarily trained is obtained, the dual-task learning model processes the input training samples including:
[0040] Using the mask generator to generate multiple mask matrices for the training sample X, and calculating the element-by-element product of the multiple mask matrices and the original input training sample; using an encoder to convert the element-by-element product into an encoding vector;
[0041] Using the reconstructor based on the encoding vector e mask Reconstruct the original input training samples so that the model can learn the feature representation of the original input training samples Complete low-level feature capture of the main task;
[0042] Using the multilayer perceptron based on the encoding vector e mask Reconstruct the encoding vector of the unmasked input Capturing high-level semantic knowledge to accomplish auxiliary tasks.
[0043] It can be seen that in the embodiment of the present invention, on the one hand, the main task uses the codec structure to reconstruct the training samples of the original input, so that the dual-task learning model can learn the feature representation of the original input; on the other hand, the auxiliary task is based on the embedding space representation obtained by the encoder, and learns how to restore the embedding representation of the unmasked input from the masked embedding representation, thereby capturing the high-level semantic knowledge of the input.
[0044] In an embodiment of the present invention, the loss function of the dual-task learning model in the preliminary training stage is the sum of the first reconstruction loss of the reconstructor, the second reconstruction loss of the multilayer perceptron, and the mask diversity loss of the mask encoder;
[0045] The first reconstruction loss for:
[0046]
[0047] The second reconstruction loss for:
[0048]
[0049] Where K is the number of masks, X is the training sample of the original input, and e is the embedded representation of the original input. is the feature representation corresponding to the i-th mask of the original input training sample reconstructed by the reconstructor, The unmasked coding vector corresponding to the i-th mask reconstructed by the multilayer perceptron based on the coding vector;
[0050] In order to allow the mask generator to generate more diverse masks and capture information from normal samples, the mask generator in the embodiment of the present invention is trained according to the admission mask diversity loss:
[0051] The mask diversity loss for:
[0052]
[0053] Among them, < > represents the inner product operation; represents the i-th and j-th mask matrices; I i≠j represents the indicator function, whose value is 1 when i≠j and 0 otherwise; τ represents the temperature parameter, and scale represents the scaling factor for adjusting the range of diversity loss.
[0054] Therefore, the loss function of the dual-task learning model in the initial training stage is for:
[0055]
[0056] Among them, λ and γ are hyperparameters.
[0057] As mentioned above, the dual-task learning model can be obtained by preliminary training using the training sample set.
[0058] Then, for step 102, the test samples in the test sample set are used to calculate the loss of each test sample using the current dual-task learning model, and high-confidence normal samples and high-confidence abnormal samples are selected from the test sample set according to the loss of the test samples, and the high-confidence normal samples and high-confidence abnormal samples are compared with the training samples to update the model parameters of the current dual-task learning model through comparative learning. The updated dual-task learning model is used as the current dual-task learning model, and the high-confidence normal samples and high-confidence abnormal samples selected this time are deleted from the test sample set, and this step is continued using the remaining test sample sets until high-confidence normal samples and high-confidence abnormal samples can no longer be selected from the test sample set, thereby obtaining a trained dual-task learning model.
[0059] In an embodiment of the present invention, the test samples in the test sample set are also tabular data, which are in the same field as the tabular data of the training samples. Unlike the training sample set which is entirely normal samples, the test sample set includes normal samples and abnormal samples. Therefore, in the model parameter updating stage, the dual-task learning model can be trained based on the comparative learning method, so that the dual-task learning model outputs high-confidence normal samples and high-confidence abnormal samples for the test samples, and then dynamically adjusts the model parameters to improve the model's ability to distinguish between normal samples and abnormal samples in the model parameter updating stage.
[0060] In an embodiment of the present invention, each test sample is used as an input of the dual-task learning model, so as to calculate the corresponding loss for each test sample using the dual-task learning model. Specifically, selecting a high-confidence normal sample and a high-confidence abnormal sample from the test sample set according to the loss of the test sample may include:
[0061] Normalize the loss of the test sample to the range of [0,1];
[0062] The number of high-confidence normal samples and high-confidence abnormal samples to be selected is determined according to the contamination rate of the test sample set, and high-confidence normal samples are selected from the test samples whose losses are closer to 0, and high-confidence abnormal samples are selected from the test samples whose losses are closer to 1 according to the selection number.
[0063] The contamination rate is the ratio of abnormal samples in the test sample set to the total number of test samples. Based on this ratio, the ratio of normal samples to abnormal samples in the test sample set can be determined. The number of high-confidence normal samples and high-confidence abnormal samples is the ratio. For example, in the test sample set, the ratio of normal samples to abnormal samples is 9:1, so the number of high-confidence normal samples to high-confidence abnormal samples is 9:1.
[0064] It should be noted that after the proportional relationship of the selection quantity is determined, the specific number to be selected needs to be determined based on the total number of test samples in the test sample set. For example, if the number of test samples in the test sample set is 1,000, then each time you select, you can select 9 high-confidence normal samples and 1 high-confidence abnormal sample, or you can select 18 high-confidence normal samples and 2 high-confidence abnormal samples. This selection quantity needs to ensure that the test samples in the test sample set can go through multiple rounds of selection, thereby achieving multiple rounds of iterative updates of the model parameters in the dual-task learning model.
[0065] The comparative learning method in the embodiment of the present invention is to compare and optimize the selected high-confidence samples (high-confidence normal samples and high-confidence abnormal samples) with the embedded representations of the training samples in the training sample set in the embedding space. The comparative optimization includes two parts: one part is test sample adaptation, and the other part is comparison between normal samples and abnormal samples.
[0066] Specifically, updating the model parameters of the current dual-task learning model by contrastive learning may include:
[0067] For each high-confidence normal sample selected this time, the following operations are performed: using the dual-task learning model to calculate the corresponding first reconstruction loss, second reconstruction loss and mask diversity loss for the high-confidence normal sample, and calculating the total loss of the high-confidence normal sample based on the weights of different losses;
[0068] The mean of the total losses of all high-confidence normal samples selected this time is taken as the first contrast loss;
[0069] For each high confidence sample selected this time, the following steps are performed: selecting k training samples that are closest to the high confidence sample from the training sample set, and calculating the mean distance between the high confidence sample and the k training samples;
[0070] The second contrast loss is obtained by subtracting the sum of the distance means of all high-confidence normal samples selected this time from the sum of the distance means of all high-confidence abnormal samples selected this time;
[0071] The sum of the first contrast loss and the second contrast loss is used as the final loss for updating the model parameters of the current dual-task learning model.
[0072] In the embodiment of the present invention, in the first part of the contrast optimization, the test sample adaptation part uses the main task and auxiliary task in the training phase to learn from the selected high-confidence normal samples, so that the model can better adapt to the normal samples in the test distribution. for:
[0073]
[0074] In an embodiment of the present invention, in the second part of the contrast optimization, the contrast between normal samples and abnormal samples is mainly used to reduce the embedding distance between normal samples in the test sample set and training samples in the training sample set, while increasing the embedding distance between abnormal samples in the test sample set and training samples in the training sample set. However, it is impractical to require that the embedding of each high-confidence test sample is close to or far away from the embedding of all training samples. On the one hand, a given sample embedding will naturally be far away from the embedding of some training samples, which makes it challenging to continuously increase the distance from all embeddings. On the other hand, calculating the distance between the embedding of the test sample and the embedding of all training samples will incur a very high computational cost. Therefore, in order to alleviate the above problems, k-nearest neighbor contrast learning is adopted in an embodiment of the present invention, which can dynamically improve the learning of the embedding representation of normal samples and abnormal samples by the dual-task learning model without incurring excessively high computational costs.
[0075] Based on this, the second contrast loss of the comparison between normal samples and abnormal samples for:
[0076]
[0077] in, represents the selected x-th high-confidence normal sample and y-th high-confidence abnormal sample; C 1 , C 2 Indicates the number of selected high-confidence normal samples and high-confidence abnormal samples; represents the training sample set, k is the k training samples selected from the training sample set with the smallest distance to the selected high confidence samples.
[0078] In summary, the final loss of updating the model parameters of the current dual-task learning model is for:
[0079]
[0080] Among them, δ is a hyperparameter.
[0081] Please refer to Figure 3 , which is a schematic diagram of the processing process of the model parameter updating phase (that is, the testing phase).
[0082] Furthermore, in order to improve the learning feature diversity of the dual-task learning model in the model parameter updating stage, the high-confidence normal samples selected each time can be added to the training sample set to enrich the normal patterns of normal samples in the training sample set, so that each subsequent selection of the k nearest distance k training samples can be selected from the updated training sample set.
[0083] Finally, for step 104, the trained dual-task learning model is used to perform anomaly detection on the input table data.
[0084] In the embodiment of the present invention, the terminal device can use the trained dual-task learning model to perform anomaly detection on the input table data and identify the anomaly probability in the table data.
[0085] In the embodiment of the present invention, the two core stages of model training based on collaborative dual-task learning and model updating based on contrastive learning can enable the model to capture the multi-level features of samples, enhance the model's ability to discriminate test samples, and effectively improve the model's anomaly detection performance on multiple data sets.
[0086] Compared with 10 baseline models, the dual-task learning model trained by the embodiment of the present invention achieved the best performance in the three indicators of AUC-ROC, AUC-PR, and F1. In terms of the F1 score, which is the most commonly used in anomaly detection tasks, the dual-task learning model trained by the embodiment of the present invention achieved the best performance in 10 of the 15 datasets and achieved suboptimal performance in 3 datasets. For AUC-ROC, the dual-task learning model trained by the embodiment of the present invention achieved the best results on 8 datasets and the second best results on 3 datasets. For AUC-PR, the dual-task learning model trained by the embodiment of the present invention achieved the best results on 7 datasets and achieved suboptimal results on 4 datasets.
[0087] Please refer to Figure 4 The embodiment of the present invention provides an unsupervised anomaly detection device based on contrastive learning, the device comprising:
[0088] A preliminary training unit 400 is used to obtain a dual-task learning model that has been preliminarily trained by adopting an unsupervised training method based on multiple training samples in a training sample set; the training samples are tabular data;
[0089] A parameter updating unit 402 is used to use the test samples in the test sample set and the current dual-task learning model to calculate the loss of each test sample, select high-confidence normal samples and high-confidence abnormal samples from the test sample set according to the loss of the test samples, and compare the high-confidence normal samples and high-confidence abnormal samples with the training samples to update the model parameters of the current dual-task learning model through comparative learning, use the updated dual-task learning model as the current dual-task learning model, delete the high-confidence normal samples and high-confidence abnormal samples selected this time from the test sample set, and continue to perform this step using the remaining test sample sets until no high-confidence normal samples and high-confidence abnormal samples can be selected from the test sample set, thereby obtaining a trained dual-task learning model; the test samples are tabular data;
[0090] The anomaly detection unit 404 is used to perform anomaly detection on the input table data using the trained dual-task learning model.
[0091] In one embodiment of the present invention, the dual-task learning model includes: a mask encoder and a reconstructor and a multi-layer perceptron respectively connected to the mask encoder; the mask encoder includes a mask generator and an encoder;
[0092] When the dual-task learning model that has been preliminarily trained is obtained, the dual-task learning model processes the input training samples including:
[0093] Using the mask generator to generate multiple mask matrices for the training samples, and calculating the element-by-element product of the multiple mask matrices and the original input training samples; using an encoder to convert the element-by-element product into a coding vector;
[0094] Reconstructing the original input training samples based on the encoding vector using the reconstructor so that the model learns the feature representation of the original input training samples and completes the low-level feature capture of the main task;
[0095] The multilayer perceptron is used to reconstruct the encoding vector that is not masked input based on the encoding vector, thereby completing the capture of high-level semantic knowledge of the auxiliary task.
[0096] In one embodiment of the present invention, the loss function of the dual-task learning model in the preliminary training stage is the sum of the first reconstruction loss of the reconstructor, the second reconstruction loss of the multilayer perceptron, and the mask diversity loss of the mask encoder;
[0097] The first reconstruction loss for:
[0098]
[0099] The second reconstruction loss for:
[0100]
[0101] Where K is the number of masks, X is the training sample of the original input, and e is the embedded representation of the original input. is the feature representation corresponding to the i-th mask of the original input training sample reconstructed by the reconstructor, The unmasked coding vector corresponding to the i-th mask reconstructed by the multilayer perceptron based on the coding vector;
[0102] The mask diversity loss for:
[0103]
[0104] Among them, < > represents the inner product operation; represents the i-th and j-th mask matrices; I i≠j represents the indicator function, whose value is 1 when i≠j and 0 otherwise; τ represents the temperature parameter, and scale represents the scaling factor for adjusting the range of diversity loss.
[0105] In one embodiment of the present invention, when the parameter updating unit performs the step of selecting a high-confidence normal sample and a high-confidence abnormal sample from the test sample set according to the loss of the test sample, the step specifically includes:
[0106] Normalize the loss of the test sample to the range of [0,1];
[0107] The number of high-confidence normal samples and high-confidence abnormal samples to be selected is determined according to the contamination rate of the test sample set, and high-confidence normal samples are selected from the test samples whose losses are closer to 0, and high-confidence abnormal samples are selected from the test samples whose losses are closer to 1 according to the selection number.
[0108] In one embodiment of the present invention, when the parameter updating unit performs the updating of the model parameters of the current dual-task learning model by the comparative learning method, the following steps are specifically performed:
[0109] For each high-confidence normal sample selected this time, the following operations are performed: using the dual-task learning model to calculate the corresponding first reconstruction loss, second reconstruction loss and mask diversity loss for the high-confidence normal sample, and calculating the total loss of the high-confidence normal sample based on the weights of different losses;
[0110] The mean of the total losses of all high-confidence normal samples selected this time is taken as the first contrast loss;
[0111] For each high confidence sample selected this time, the following steps are performed: selecting k training samples that are closest to the high confidence sample from the training sample set, and calculating the mean distance between the high confidence sample and the k training samples;
[0112] The second contrast loss is obtained by subtracting the sum of the distance means of all high-confidence normal samples selected this time from the sum of the distance means of all high-confidence abnormal samples selected this time;
[0113] The sum of the first contrast loss and the second contrast loss is used as the final loss for updating the model parameters of the current dual-task learning model.
[0114] In one embodiment of the present invention, the first contrast loss for:
[0115]
[0116] The second contrast loss for:
[0117]
[0118] in, represents the selected x-th high-confidence normal sample and y-th high-confidence abnormal sample; C 1 , C 2 Indicates the number of selected high-confidence normal samples and high-confidence abnormal samples; represents the training sample set, k is the k training samples selected from the training sample set with the smallest distance to the selected high confidence samples.
[0119] It should be noted that the unsupervised anomaly detection device based on contrastive learning provided in the above embodiment is only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the unsupervised anomaly detection device based on contrastive learning provided in the above embodiment and the unsupervised anomaly detection method embodiment based on contrastive learning belong to the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0120] The embodiment of the present application also provides a computer device, please refer to Figure 5 The computer device includes a processor and a memory, wherein at least one instruction, at least one program, a code set or an instruction set is stored in the memory, and the at least one instruction, at least one program, a code set or an instruction set is loaded and executed by the processor to implement the unsupervised anomaly detection method based on contrastive learning provided in the above-mentioned method embodiments.
[0121] An embodiment of the present application also provides a computer-readable storage medium, on which is stored at least one instruction, at least one program, code set or instruction set, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by a processor to implement the unsupervised anomaly detection method based on contrastive learning provided in the above-mentioned method embodiments.
[0122] An embodiment of the present application also provides a computer program product, which includes a computer program. A processor of a computer device reads the computer program from a computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the unsupervised anomaly detection method based on contrastive learning described in any of the above embodiments.
[0123] For the convenience of description, the above system or device is described by dividing it into various modules or units according to its functions. Of course, when implementing the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0124] It can be known from the description of the above implementation methods that those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application can be essentially or partly contributed to the prior art in the form of a software product, which can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present application or certain parts of the embodiments.
[0125] Finally, it should be noted that, in this article, relational terms such as first, second, third and fourth are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the statement "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0126] The above is only a preferred implementation of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. An unsupervised anomaly detection method based on contrastive learning, characterized in that: The method comprises: Based on multiple training samples in the training sample set, an unsupervised training method is used to obtain a dual-task learning model that has been preliminarily trained; the training samples are tabular data; Using the test samples in the test sample set, using the current dual-task learning model to calculate the loss of each test sample, selecting high-confidence normal samples and high-confidence abnormal samples from the test sample set according to the loss of the test samples, and comparing the high-confidence normal samples and high-confidence abnormal samples with the training samples, so as to update the model parameters of the current dual-task learning model through comparative learning, using the updated dual-task learning model as the current dual-task learning model, deleting the high-confidence normal samples and high-confidence abnormal samples selected this time from the test sample set, and continuing to perform this step using the remaining test sample sets until no high-confidence normal samples and high-confidence abnormal samples can be selected from the test sample set, and obtaining a trained dual-task learning model; the test samples are tabular data; The trained dual-task learning model is used to perform anomaly detection on the input tabular data.
2. The method according to claim 1, characterized in that The dual-task learning model includes: a mask encoder and a reconstructor and a multi-layer perceptron respectively connected to the mask encoder; the mask encoder includes a mask generator and an encoder; When the dual-task learning model that has been preliminarily trained is obtained, the dual-task learning model processes the input training samples including: Using the mask generator to generate multiple mask matrices for the training samples, and calculating the element-by-element product of the multiple mask matrices and the original input training samples; using an encoder to convert the element-by-element product into a coding vector; Reconstructing the original input training samples based on the encoding vector using the reconstructor so that the model learns the feature representation of the original input training samples and completes the low-level feature capture of the main task; The multilayer perceptron is used to reconstruct the encoding vector that is not masked input based on the encoding vector, thereby completing the capture of high-level semantic knowledge of the auxiliary task.
3. The method according to claim 2, characterized in that The loss function of the dual-task learning model in the preliminary training stage is the sum of the first reconstruction loss of the reconstructor, the second reconstruction loss of the multi-layer perceptron, and the mask diversity loss of the mask encoder; The first reconstruction loss for: The second reconstruction loss for: Where K is the number of masks, X is the training sample of the original input, and e is the embedded representation of the original input. is the feature representation corresponding to the i-th mask of the original input training sample reconstructed by the reconstructor, The unmasked coding vector corresponding to the i-th mask reconstructed by the multilayer perceptron based on the coding vector; The mask diversity loss for: Among them, <> represents the inner product operation; represents the i-th and j-th mask matrices; i i≠j represents the indicator function, whose value is 1 when i≠j and 0 otherwise; τ represents the temperature parameter, and scale represents the scaling factor for adjusting the range of diversity loss.
4. The method according to claim 1, characterized in that: The step of selecting a high-confidence normal sample and a high-confidence abnormal sample from the test sample set according to the loss of the test sample includes: Normalize the loss of the test sample to the range of [0,1]; The number of high-confidence normal samples and high-confidence abnormal samples to be selected is determined according to the contamination rate of the test sample set, and high-confidence normal samples are selected from the test samples whose losses are closer to 0, and high-confidence abnormal samples are selected from the test samples whose losses are closer to 1 according to the selection number.
5. The method according to claim 3, characterized in that: The updating of model parameters of the current dual-task learning model by contrastive learning includes: For each high-confidence normal sample selected this time, the following operations are performed: using the dual-task learning model to calculate the corresponding first reconstruction loss, second reconstruction loss and mask diversity loss for the high-confidence normal sample, and calculating the total loss of the high-confidence normal sample based on the weights of different losses; The mean of the total losses of all high-confidence normal samples selected this time is taken as the first contrast loss; For each high confidence sample selected this time, the following steps are performed: selecting k training samples that are closest to the high confidence sample from the training sample set, and calculating the mean distance between the high confidence sample and the k training samples; The second contrast loss is obtained by subtracting the sum of the distance means of all high-confidence normal samples selected this time from the sum of the distance means of all high-confidence abnormal samples selected this time; The sum of the first contrast loss and the second contrast loss is used as the final loss for updating the model parameters of the current dual-task learning model.
6. The method according to claim 5, characterized in that The first contrast loss for: The second contrast loss for: in, represents the xth high-confidence normal sample and the yth high-confidence abnormal sample selected; C1 and C2 represent the number of high-confidence normal samples and high-confidence abnormal samples selected; represents the training sample set, k is the k training samples selected from the training sample set with the smallest distance to the selected high confidence samples.
7. An unsupervised anomaly detection device based on contrastive learning, characterized in that: The device comprises: A preliminary training unit, used to obtain a dual-task learning model that has completed preliminary training by adopting an unsupervised training method based on multiple training samples in a training sample set; the training samples are tabular data; A parameter updating unit is used to use the test samples in the test sample set and the current dual-task learning model to calculate the loss of each test sample, select high-confidence normal samples and high-confidence abnormal samples from the test sample set according to the loss of the test samples, and compare the high-confidence normal samples and high-confidence abnormal samples with the training samples to update the model parameters of the current dual-task learning model through comparative learning, use the updated dual-task learning model as the current dual-task learning model, delete the high-confidence normal samples and high-confidence abnormal samples selected this time from the test sample set, and continue to perform this step using the remaining test sample sets until no high-confidence normal samples and high-confidence abnormal samples can be selected from the test sample set, thereby obtaining a trained dual-task learning model; the test samples are tabular data; The anomaly detection unit is used to perform anomaly detection on input table data using the trained dual-task learning model.
8. A computer device, characterized in that: The computer device includes a memory and a processor, the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory to implement the steps of any one of the methods described in claims 1-6.
9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 6 are implemented.
10. A computer program product, characterized in that The method comprises a computer program, wherein when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.