Weakly Supervised Video Anomaly Detection Method Based on Learning with Dynamically Selected Comparison Examples

Through the method of dynamic selection and comparison instance learning, instance feature learning, dynamic instance selection and instance feature domain adaptive modules are constructed, which solves the problems of high manual annotation cost and impacts of scene similarity in video anomaly detection, and realizes high-precision weakly supervised video anomaly detection.

CN120107868BActive Publication Date: 2025-07-29MINJIANG UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510591902.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-07-29
Estimated Expiration
2045-05-09

AI Technical Summary

Technical Problem

The existing video anomaly detection technology has problems such as high manual labeling costs, insufficient model understanding of abnormal patterns, ignoring the relationship between normal and abnormal instances within abnormal videos, and scene similarity affecting the detection effect.

Method used

Using a method based on dynamic selection and comparison instance learning, the instance feature learning module, dynamic instance selection module and instance feature domain adaptive module are constructed to weaken the influence of irrelevant information, enhance the separability of normal and abnormal instances, and optimize model performance using a multi-instance learning framework.

Benefits of technology

It significantly improves the performance of weakly supervised video anomaly detection, reduces manual labeling costs, improves the accuracy of the model for abnormal detection, and overcomes the negative impact of scene similarity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107868B_ABST
    Figure CN120107868B_ABST
Patent Text Reader

Abstract

The present invention relates to a weakly supervised video anomaly detection method based on dynamically selecting comparison instances for learning, belonging to the field of video anomaly detection. The method first constructs an instance feature learning module based on contrastive learning to encourage the aggregation of normal instances while separating abnormal instances, thereby promoting the distinction between normal instances and abnormal instances; then, a dynamic instance selection module is constructed, which identifies the most likely normal instance candidates with the smallest feature quantity in abnormal videos; finally, a feature domain adaptation module is constructed to further enhance the separability between normal and abnormal instances through domain adaptation learning. The method of the present invention enables the model to learn the features of normal segments in abnormal videos, thereby more precisely distinguishing normal from abnormal; and the method of the present invention fully weakens the negative impact of scene similarity on the model through contrastive learning and domain adaptation learning, thereby improving the anomaly detection performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of video anomaly detection, and particularly relates to a weakly supervised video anomaly detection method based on dynamically selecting contrastive instances for learning. Background Art

[0002] Video anomaly detection refers to the detection of abnormal events in videos, and its application scope is very wide, including road traffic monitoring, crowd monitoring, etc. In recent years, video anomaly detection technology has made remarkable progress. However, due to the unboundedness of abnormal events in practical applications and the difficulty of collecting large-scale annotated data, video anomaly detection still faces major challenges.

[0003] Among the existing video anomaly detection technologies, they can be divided into three categories according to the annotation of the required training videos: supervised video anomaly detection, unsupervised video anomaly detection, and weakly supervised video anomaly detection. Supervised video anomaly detection uses video data with frame-level labels for training. Unsupervised video anomaly detection usually assumes that only normal video data is used for training, while weakly supervised video anomaly detection uses video data with video-level labels to train the model. In practical scenarios, the manual annotation cost of the supervised method is extremely expensive and time-consuming. The unsupervised method trains the model through methods based on frame prediction and methods based on frame reconstruction. However, since this method can only use normal data for training, it has an obvious drawback, that is, there are only normal data samples in the training data, and the model lacks awareness of abnormal patterns during the training process. Therefore, the anomaly boundary of the unsupervised video anomaly detection method is not clear. Although the weakly supervised video anomaly detection method balances the detection accuracy and the annotation cost, and has a much lower cost than the supervised method while having a relatively high detection accuracy, the weakly supervised video anomaly detection method often based on the multi-instance learning strategy, and this type of method currently has two drawbacks: 1. They only focus on the instances with the largest feature quantities in normal and abnormal videos, while ignoring the potentially valuable information in other instances, which reduces the overall performance of the model. 2. They ignore the relationship between normal and abnormal instances within abnormal videos, and the visual and semantic similarities of these instances themselves make it challenging to distinguish between abnormal and normal. Summary of the Invention

[0004] The purpose of the present invention is to overcome the defects of the prior art and provide a weakly supervised video anomaly detection method based on dynamically selecting contrastive instances for learning. First, feature extraction and feature projection are performed on the input video. Secondly, an instance feature learning module based on contrastive loss is proposed to weaken the negative impact of information unrelated to anomalies on the model. Then, a MIL (multi-instance learning) module for dynamic instance selection is proposed based on the multi-instance learning framework to improve the model performance by learning normal instances in abnormal samples. Finally, an instance feature domain adaptation module is introduced to weaken the differences caused by different background information in the video.

[0005] To achieve the above object, the technical solution of the present invention is: a weakly supervised video anomaly detection method based on dynamically selecting comparison examples for learning, including:

[0006] Construct an instance feature learning module based on contrastive learning to aggregate normal instances while separating abnormal instances;

[0007] Construct a dynamic instance selection module to identify the normal instance candidates with the smallest feature amount in the abnormal video;

[0008] Construct an instance feature domain adaptation module to enhance the separability between normal instances and abnormal instances through domain adaptation learning.

[0009] Furthermore, before the input video passes through the instance feature learning module based on contrastive learning, feature extraction and feature projection need to be performed on the input video.

[0010] Furthermore, the specific implementation method of performing feature extraction and feature projection on the input video is as follows:

[0011] (1) Feature extraction:

[0012] Collect a set of abnormal videos and a set of normal videos with video-level labels;

[0013] First, divide a single video C into several segments, each segment consisting of 16 consecutive frames, and each segment is called an instance, denoted as c; at this time, a single video is represented as , where M is the number of segments owned by a single video C, represents the m-th instance;

[0014] Subsequently, use the I3D network as the backbone network for feature extraction: First, perform 10-crop data augmentation on each instance in a single video C, that is: intercept five sub-images of the upper left, upper right, lower left, lower right, and middle from each original video frame, and intercept five sub-images of the upper left, upper right, lower left, lower right, and middle again after mirroring and flipping the original image; then, input the augmented video into the pre-trained I3D network to obtain the video feature , where represents the m-th instance after feature extraction;

[0015] Finally, after a set of abnormal videos and a set of normal videos are subjected to feature extraction, an abnormal training set and a normal training set are formed, represents the i-th abnormal video feature, represents the i-th normal video feature;

[0016] (2) Feature projection:

[0017] First, use the normal training set to offline learn a static sparse dictionary D. The static sparse dictionary D consists of R sparse centers d with the same shape as the instances, denoted as , representing the r-th sparse center; the learning process of the sparse dictionary is as follows:

[0018]

[0019] where is the coefficient vector constrained by the sparsity prior, is the regularization coefficient, which is set to 0.1 in the present invention; the learning process of the sparse dictionary is to encourage each sparse center in the static sparse dictionary D to be as similar as possible to the features of each normal instance, so as to learn a sparse dictionary containing normal semantic information;

[0020] Then, use the static sparse dictionary D and the attention mechanism to perform feature projection on the video; the specific calculation process is:

[0021]

[0022] where respectively represent the features of the query, key, and value obtained through the linear function, is from or input features; is the projection of on D, and its feature dimension is the same as .

[0023] Furthermore, the instance feature learning module based on contrastive learning is specifically implemented as follows:

[0024] Given a pair of video features and respectively from and input to the model, after feature projection, and are obtained respectively; and are input into the multi-scale temporal network MTN together with their respective projection inputs to learn the long-term and short-term spatio-temporal dependencies of the video information, denoted as:

[0025]

[0026] where represents MTN, and are respectively the , after passing through MTN, and embedding features; introducing channel attention, taking and M instances in as the channel dimension, and taking the instance features corresponding to each instance as the spatio-temporal dimension, calculating the difference between the channel weights of and as the attention weight: where

[0027]

[0028] is the global average pooling layer, E is the fully connected network, is the Sigmoid function; similarly, the attention weights of and and are obtained ; multiplying and by their attention weights and respectively, to obtain the video features after normal information is suppressed and as:

[0029]

[0030] where "·" represents matrix multiplication;

[0031] Define the instances in and as and respectively, that is , , and represent and the m-th instance in respectively; design an instance feature contrast loss function, for , encourage the instance features in it to be far away from each other; for , encourage the instance features in it to be close to each other; in contrastive learning, define a non-linear projection head H, and the process of obtaining the embedding features through the non-linear projection head is as follows:

[0032]

[0033] where represents and ;

[0034] The instance feature contrast loss function is defined as:

[0035]

[0036] Among them, is the weight coefficient, is the natural exponential function, is the indicator function, indicating that the function value is 1 when and 0 otherwise.

[0037] Furthermore, the dynamic instance selection module is a multi-instance learning (MIL) module for dynamic instance selection.

[0038] Furthermore, the dynamic instance selection module is specifically implemented as follows:

[0039] Encourage the top K largest instance feature sizes in the abnormal video samples to increase, and at the same time encourage the top K largest instance feature sizes in the normal video samples to decrease; specifically as follows,

[0040] First, the feature size is defined as:

[0041]

[0042] Among them, is or the instance feature in is the symbol for calculating the two-norm of the feature vector;

[0043] Then, to select the top K largest instance features in the video samples, the following function is introduced:

[0044]

[0045] Among them, represents the top K instance features with the largest feature sizes from represents the average feature size of the top K largest instance features in the video sample

[0046] To learn the possible normal instance features in the abnormal video samples, the average value of the feature sizes of all the instance features in all the normal video samples in the small batch with the current batch size of I will be calculated and used as the threshold to dynamically screen the instance features in the abnormal video sample in each round of training; The calculation process is as follows:

[0047]

[0048] Then, for each abnormal video sample in the small batch There are the following operations: Define an empty set , for each in , if , then store it in the set ;

[0049] Next, select the first smallest instance features in the abnormal video samples, and propose the following function:

[0050]

[0051] where represents the number of elements in the set . When , represents the smallest instance features in terms of feature size, represents the average feature size of the first smallest instance features in the video sample ; when , represents the smallest instance features in terms of feature size, represents the average feature size of the first smallest instance features in the video sample ;

[0052] Next, define the feature size loss function as:

[0053]

[0054] where b is a predefined boundary, is a weight hyperparameter, and "*" is ordinary multiplication;

[0055] Finally, define a discriminator Z, which consists of a multi-layer perceptron network L and a Sigmoid function . Input into the discriminator to obtain the anomaly score :

[0056]

[0057] At the same time, input into the discriminator to obtain the anomaly score :

[0058]

[0059] Define the feature classification loss function as:

[0060]

[0061] Among them, BCE is the binary cross-entropy loss function.

[0062] Furthermore, the instance feature domain adaptation module is specifically implemented as follows:

[0063] The video features after the abnormal information is suppressed and are:

[0064]

[0065] Then, perform global average pooling on and respectively to obtain the average values representing normal instances for each:

[0066]

[0067] Next, input and into the discriminator with a gradient reversal layer GRL:

[0068]

[0069] Among them, is a multi-layer perceptron network; GRL acts as an identity function during forward propagation:

[0070]

[0071] Among them, is the input of GRL;

[0072] During backpropagation, GRL reverses the current gradient by multiplying by the reversal coefficient so that the optimization objective of is opposite to the optimization objective of the task discriminator;

[0073] Finally, the reversed loss function is defined as:

[0074]

[0075] Among them, represents or When represents at this time, ; when represents at this time, .

[0076] Compared with the prior art, the present invention has the following beneficial effects:

[0077] 1. The present invention optimizes the video anomaly detection model by using the method of weakly supervised multi-instance learning. The data used for training the model are normal and abnormal data with video-level labels. Compared with the supervised method, the manual annotation cost of the video-level labels used in the proposed model is much lower than that of the supervised method, and it has more practical value. Compared with the unsupervised video anomaly detection method, the weak supervision information can greatly improve the model performance, making the anomaly detection performance of the model far exceed that of the unsupervised method.

[0078] 2. Existing models only focus on several instances with the largest feature sizes in normal and abnormal videos, while ignoring the information contained in other instance features, especially the beneficial information contained in the instances with the smallest functional sizes in abnormal videos, which limits the performance of the models. The method of the present invention can enable the model to learn the features of normal segments in abnormal videos, so as to more precisely distinguish between normal and abnormal.

[0079] 3. Existing models ignore the fact that the instances in the same video have a high similarity to each other due to the same scene, and this similarity still exists between normal and abnormal instances, which is not conducive to MIL to distinguish abnormal and normal instances. The method of the present invention fully weakens the negative impact of scene similarity on the model through contrastive learning and domain adaptation learning, thereby improving the anomaly detection performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0080] Figure 1 It is a diagram of the overall execution process of the method of the present invention.

[0081] Figure 2 It is a sample of video data frames: (a) normal video frame; (b) abnormal video frame.

[0082] Figure 3 It is the feature extraction and feature projection module.

[0083] Figure 4 It is the instance feature learning module based on contrastive learning.

[0084] Figure 5 It is the MIL module for dynamic instance selection.

[0085] Figure 6 It is the instance feature domain adaptation module.

[0086] Figure 7 It is the ShanghaiTech loss curve.

[0087] Figure 8 It is the UCF-Crime loss curve.

[0088] Figure 9 For visualizing the detection results.

[0089] Figure 10 The AUC-ROC curves of the detection results on ShanghaiTech and UCF-Crime; among them, the left figure is the AUC-ROC curve of the detection results on ShanghaiTech, and the right figure is the AUC-ROC curve of the detection results on UCF-Crime.

[0090] Figure 11 The flow chart for implementing the method of the present invention. Specific implementation manners

[0091] The technical solution of the present invention will be specifically described below in conjunction with the accompanying drawings.

[0092] The present invention provides a weakly supervised video anomaly detection method based on dynamically selecting contrastive instance learning, including:

[0093] Constructing an instance feature learning module based on contrastive learning to aggregate normal instances and separate abnormal instances at the same time;

[0094] Constructing a dynamic instance selection module to identify the normal instance candidates with the smallest feature amount in the abnormal video;

[0095] Constructing an instance feature domain adaptation module to enhance the separability between normal instances and abnormal instances through domain adaptation learning.

[0096] The following is the specific implementation process of the present invention.

[0097] As Figure 1 shown, a weakly supervised video anomaly detection method based on dynamically selecting contrastive instance learning of the present invention first performs feature extraction and feature projection on the input video; secondly, an instance feature learning module based on contrastive loss is proposed to weaken the negative impact of information unrelated to anomalies on the model. Then, a MIL (multi-instance learning) module of dynamic instance selection is proposed on the basis of the multi-instance learning framework to improve the model performance by learning normal instances in abnormal samples. Finally, an instance feature domain adaptation module is introduced to weaken the differences caused by different background information in the video. The specific implementation steps are as follows:

[0098] Step 1: Input of video data

[0099] The data used in this invention is sourced from the publicly available benchmark video datasets ShanghaiTech and UCF-Crime. Among them, ShanghaiTech is a challenging multi-scene dataset consisting of 13 campus scenes with different lighting conditions and camera angles. This dataset contains a total of 437 videos, including 238 training videos and 199 test videos. UCF-Crime is a large-scale anomaly detection dataset composed of 1900 untrimmed videos captured from real-world streets and indoor surveillance cameras. Compared with ShanghaiTech, the background of UCF-Crime is more complex and diverse. This dataset includes 1610 training videos and 290 test videos. Both datasets contain different video surveillance scenarios, including abnormal human behaviors and abnormal events caused by vehicles or devices. The division of the two datasets is shown in Table 1, and sample video data frames in the datasets are as shown in Figure 2 shown in Figure 2 (a) in which is a normal video frame; Figure 2 (b) in which is an abnormal video frame).

[0100]

[0101] Step 2: Feature Extraction and Feature Projection Module

[0102] The feature extraction and feature projection process is as shown in Figure 3 Collect a set of abnormal videos and a set of normal videos with video-level labels to train our model.

[0103] Feature Extraction: First, divide a single video C into several segments, each segment consisting of 16 consecutive frames, and each segment is called an instance, denoted as c; at this time, a single video is represented as , where M is the number of segments owned by a single video C, represents the m-th instance. Subsequently, use the I3D network as the backbone network for feature extraction. First, perform 10-crop data augmentation on each instance in a single video C, that is, intercept five sub-images of the upper left, upper right, lower left, lower right, and middle from each original video frame, and then mirror and flip the original image and intercept five sub-images again in the same way. Then, input the augmented video into the pre-trained I3D network to obtain the video feature , where represents the m-th instance after feature extraction. Finally, perform the above operations on all abnormal videos and normal videos respectively to form the abnormal training set and the normal training set .

[0104] Feature Projection: Use the normal training set Learn a static sparse dictionary D offline. The static sparse dictionary D consists of R sparse centers d with the same shape as the instances, denoted as , denoting the r-th sparse center; the learning process of the sparse dictionary is as follows:

[0105]

[0106] where is the coefficient vector constrained by the sparsity prior, is the regularization coefficient, which is set to 0.1 in the present invention; the learning process of the sparse dictionary is to encourage each sparse center in the static sparse dictionary D to be as similar as possible to the features of each normal instance, so as to learn a sparse dictionary containing normal semantic information;

[0107] Then, use the static sparse dictionary D and the attention mechanism to perform feature projection on the video; the specific calculation process is:

[0108]

[0109] where respectively represent the features of the query, key, and value obtained through the linear function, is from or input feature; is the projection of on D. Since D is the sparse dictionary of normal samples, so is the normal feature obtained by projecting on D, and its feature dimension is the same as .

[0110] Step 3: Instance Feature Learning Module Based on Contrastive Learning

[0111] The instance feature learning module based on contrastive learning aims to encourage a large feature distance between instances in abnormal video samples and a small feature distance between instances in normal video samples, so as to weaken the correlation between each instance in abnormal samples and reduce the negative impact of irrelevant information in instance features on anomaly detection. Its structure is as Figure 4 shown. The specific method is as follows. Given a pair of video features and input to the model, which are from and respectively, and obtain and respectively through feature projection; then, input and into their respective projection input multi-scale temporal network MTN to learn the long-term and short-term spatio-temporal dependencies of video information. This process is expressed as:

[0112]

[0113] Among them, represents MTN, and are respectively the embedding features of , , and after MTN; introducing channel attention to suppress normal event information in the video embedding features, taking the M instances in and as the channel dimension, taking the instance features corresponding to each instance as the spatio-temporal dimension, and calculating the difference between the channel weights of and as the attention weight:

[0114]

[0115] Among them, is the global average pooling layer, E is the fully connected network FC Layer, is the Sigmoid function; similarly, the attention weights and of are obtained; multiplying and by their respective attention weights and respectively, the video features with normal information suppressed are obtained and as:

[0116]

[0117] Among them, "·" represents matrix multiplication;

[0118] Define and the instances in and respectively as , , and respectively represent and the m-th instance in , which may contain some normal instances and some abnormal instances. However, since these instances all come from the same video, the feature information of irrelevant abnormal events in the instances (such as the picture background information) is very similar, especially when the proportion of the abnormal event main body in the picture is small. For the video anomaly detection task, the model needs to pay as much attention as possible to the information related to the abnormal event main body, and at the same time avoid being interfered by the information of irrelevant abnormal events as much as possible. Therefore, an instance feature contrast loss function is designed. For , each instance feature in it is encouraged to be far away from the other instance features, so as to break the problem that normal and abnormal instances are similar due to irrelevant information such as the background; while for , the instance features in it are encouraged to be close to each other, so as to prevent the model from over-focusing on the feature differences between normal instances. Using a non-linear transformation projection head in contrastive learning can form and retain more information in the encoder. For this purpose, a non-linear projection head H is defined, and the process of obtaining the embedded features through the non-linear projection head is as follows:

[0119]

[0120] Among them, represents and ;

[0121] The instance feature contrast loss function is defined as:

[0122]

[0123] Among them, is the weight coefficient, is the natural exponential function, is the indicator function, indicating that the function value is 1 when , otherwise the function value is 0.

[0124] Step 4: The MIL module for dynamic instance selection

[0125] In weakly supervised video anomaly detection, each instance in the normal video sample is normal, while in the abnormal video sample, often a part of the instances are abnormal and another part of the instances are normal. Existing weakly supervised video anomaly detection methods adopt the traditional MIL strategy, by encouraging the top K instances with the largest feature sizes in the abnormal video sample to be as large as possible, and vice versa, the top K instances with the largest feature sizes in the normal video sample to be as small as possible, so as to distinguish between normal and abnormal. However, the existing methods ignore the learning of the feature of the possible normal instances in the abnormal video sample, resulting in limited detection performance. Therefore, we propose the MIL module for dynamic instance selection, as shown in Figure 5 .

[0126] Based on the prior information that abnormal instances in MIL often have larger feature sizes than normal instances, we encourage the top K largest instance feature sizes in abnormal video samples to increase, while encouraging the top K largest instance feature sizes in normal video samples to decrease. Specifically, first, the feature size is defined as:

[0127]

[0128] where is or the instance feature in ; is the symbol for calculating the two-norm of the feature vector;

[0129] Then, to select the top K largest instance features in the video sample, the following function is introduced:

[0130]

[0131] where represents the K instance features with the largest feature sizes from ; represents the average feature size of the top K largest instance features in the video sample ; in the video sample;

[0132] To learn the normal instance features that may exist in the abnormal video sample, the average value of the feature sizes of the instance features in all normal video samples in the minibatch with the current batch-size of I will be calculated , and will be used as the threshold in each round of training to dynamically screen the instance features in the abnormal video sample . The calculation process is as follows:

[0133]

[0134] Then, for each abnormal video sample in the minibatch the following operations are performed: Define an empty set , and for each in , if , then it is stored in the set

[0135] Next, select the top smallest instance features in the abnormal video sample, and the following function is proposed:

[0136]

[0137] Among them, represents the number of elements in the set . When , represents the instance features with the smallest feature size from ; represents the average feature size of the first smallest instance features in the video sample; when , represents the instance features with the smallest feature size from ; represents the average feature size of the first smallest instance features in the video sample;

[0138] Next, the feature size loss function is defined as:

[0139]

[0140] where b is a predefined boundary, is a weight hyperparameter, and "*" is ordinary multiplication;

[0141] Finally, a discriminator Z is defined. Z consists of a multi-layer perceptron network L and a Sigmoid function . Input into the discriminator to obtain an anomaly score :

[0142]

[0143] Meanwhile, input into the discriminator to obtain an anomaly score :

[0144]

[0145] The feature classification loss function is defined as:

[0146]

[0147] where BCE is the binary cross-entropy loss function.

[0148] Step 5, Instance Feature Domain Adaptation Module

[0149] In the MIL-based method, only the top-K instance features with the largest feature sizes in abnormal / normal video samples are encouraged to be amplified / attenuated, so as to separate abnormal instances and normal instances. Existing MIL-based methods are based on the prior knowledge that the features of normal instances in abnormal video samples are similar to those of normal instances in normal video samples. Therefore, encouraging the attenuation of the top-K instance features with the largest feature sizes in normal video samples indirectly encourages the attenuation of the features of normal instances in abnormal video samples. In the method of the backup, it is also hoped that the features of normal instances in abnormal video samples have a high similarity with those of normal instances in normal video samples.

[0150] Inspired by domain adaptation, an instance feature domain adaptation module is proposed, as Figure 6 shown. First, in step 3, channel attention is introduced to suppress the normal event information in the video embedding features, and the attention weights and that focus on abnormal event information are obtained, while and are the attention weights that focus on normal event information. Therefore, the video features and after suppressing abnormal information are:

[0151]

[0152] Then, global average pooling is performed on and respectively to obtain the respective averages representing normal instances:

[0153]

[0154] Among them, is the global average pooling layer, which targets the channel dimension, that is, an average operation is performed on the original M instance features. Therefore, and have the same dimensional shape as a single instance . Then, and are input into a discriminator with a gradient reversal layer (GRL):

[0155]

[0156] Among them, is a multi-layer perceptron network; GRL acts as an identity function during forward propagation:

[0157]

[0158] Among them, Is the input for GRL; during backpropagation, GRL multiplies by the inversion coefficient to invert the current gradient ( ), such that the optimization objective of the MTN feature extractor is opposite to that of the task discriminator, that is, it encourages to ignore the differences between the normal instance features in abnormal video samples and the normal instance features in normal video samples when extracting features, making the two have high similarity.

[0159] Finally, the inversion loss function is defined as:

[0160]

[0161] where represents or , when represents , ; when represents , .

[0162] Step 6, Abnormal Frame Judgment

[0163] During the model inference process, input a video feature extracted by the pre-trained I3D network , and use the output of the discriminator Z as the anomaly score . The overall inference process is shown in Table 2.

[0164]

[0165] Step 7, Model Training and Evaluation

[0166] Load the unsupervised video anomaly detection model based on the multi-subcluster memory prototype; input the training set video data of ShanghaiTech and UCF-Crime into the model for training. The method is trained using the Adam optimizer, with the learning rate set to 0.001, the weight decay to 0.005, and the Batch-size set to 32. The training is divided into two stages. The first 1000 epochs are the first stage, at this time , and the last 1000 epochs are the second stage, at this time .

[0167] During the iteration process, the model loss curve reflects the model's learning degree of video features and the ability to distinguish abnormal and normal instances. The overall loss curve on the ShanghaiTech dataset is as Figure 7 shown, while the overall loss curve on the UCF-Crime dataset is asFigure 8 As shown, after the 1000th iteration period, there is a small upward trend in the total loss curves on both datasets. The reason for this phenomenon is that at this time, the parameters , and the dynamic instance selection module begins to take effect, affecting the network parameters, thus causing fluctuations. Subsequently, the curves continue to show a downward trend.

[0168] The unsupervised video anomaly detection model based on multi-subcluster memory prototypes is used to detect the videos in the test set. Randomly selected from ShanghaiTech (“ Figure 9 (a) in Figure 9 (b) in Figure 9 (c) in Figure 9 (d) in Figure 9 (e) in Figure 9 (f) in Figure 9 ), and the anomaly scores (Anomaly) and ground truth labels of UCF-Crime are visualized. The detection results are as

[0169] The evaluation metric adopted by this method is the area under the frame-level ROC curve for video anomaly detection: AUC. The AUC-ROC curves of the detection results of the proposed model on ShanghaiTech and UCF-Crime are as Figure 10 shown:

[0170] The test set is input into the model for video anomaly detection. As can be seen from the AUC-ROC curve Figure 10 shown, the detection effect can reach a relatively good result. According to the AUC-ROC curve, there are still a small number of false positives (misjudging normal samples as abnormal samples) and false negatives (misjudging abnormal samples as normal samples) during the detection process of the model. The reason is that there are some difficult samples in the test set videos, such as minor abnormal events or rare but normal events.

[0171] The present invention defines a baseline model and conducts ablation experiments on three key modules of the proposed model. The experimental results are shown in Table 3, where A represents the instance feature learning module based on contrastive learning, B represents the MIL module for dynamic instance selection, and C represents the instance feature domain adaptation module.

[0172]

[0173] As can be seen from Table 3, when any one of the three proposed modules is added alone, the model performance is improved; when any two modules are added simultaneously, the performance will be further enhanced; and when the three modules work together, the proposed model achieves the best frame-level AUC of 97.57% on the ShanghaiTech dataset and 85.65% on the UCF-Crime dataset. The results of the ablation experiments prove that the addition of each module proposed in the present invention can improve the model performance.

[0174] Table 4 presents the frame-level AUC results of the proposed method on the ShanghaiTech and UCF-Crime datasets and the comparison with the current SOTA (State-of-the-art) methods. Among the comparison methods, RTFM is a robust temporal feature amplitude learning model that trains a feature amplitude learning function to effectively identify positive samples; Wang et al is a video anomaly detection model based on a multi-instance ranking algorithm; BE-WVAD is the BE-WSVAD model that uses a binary network enhancement strategy during training, significantly improving the detection accuracy; DAR is a decoupled and solved model for video anomaly detection, which consists of two modules, a temporal proposal producer and an online anomaly locator, to jointly complete the detection task; Park et al is a normal-directed multi-instance learning model that encodes multiple normal patterns from normal videos to construct a similarity-based anomaly classifier; Tsiktsiris et al is an intelligent surveillance model based on a spatio-temporal autoencoder architecture that uses a center-weighted loss function to learn the features of normal videos; HSN is a human-scene-based model that learns discriminative representations by dissociating to capture subtle and significant features. It can be seen from the figure that the proposed method outperforms the existing SOTA unsupervised learning methods and obtains the best frame-level AUC results of 97.57% and 85.65% on the ShanghaiTech dataset and the UCF-Crime dataset, respectively.

[0175]

[0176] Step 8, Model invocation

[0177] Read the trained video anomaly detection model stored in the cloud or locally to perform video anomaly detection. Specifically, obtain the video stream from the surveillance video and input it into the model. Compare the value of the anomaly score with a pre-set threshold. If the anomaly score is greater than the threshold, it is determined that an anomaly has occurred and an alarm is generated; if the anomaly score is less than or equal to the threshold, it is determined to be normal.

[0178] The above are the preferred embodiments of the present invention. Any changes made to the technical solution of the present invention that do not exceed the scope of the technical solution of the present invention in terms of the functions and effects produced shall fall within the protection scope of the present invention.

Claims

1. A weakly supervised video anomaly detection method based on learning with dynamically selected contrastive examples, characterized in that Including: Obtain a video stream from a surveillance video and input it into a video anomaly detection model. Compare the anomaly score output by the video anomaly detection model with a preset threshold. If the anomaly score is greater than the threshold, determine that the input video is an abnormal video and generate an alarm. If the anomaly score is less than or equal to the threshold, determine that the input video is a normal video; The video anomaly detection model is constructed in the following way: Construct an instance feature learning module based on contrastive learning, which is used to aggregate instances in normal videos and separate instances in abnormal videos according to the abnormal training set, normal training set, and video features after feature projection of the abnormal training set and normal training set; Construct a dynamic instance selection module, which is used to identify the first several instance features sorted from smallest to largest in terms of feature size in abnormal videos according to the instance features of the aggregated instances in normal videos and the separated instances in abnormal videos, and calculate the anomaly score; Construct an instance feature domain adaptation module, which is used to enhance the separability of separating instances in abnormal videos in the instance feature learning module based on contrastive learning; The instance feature domain adaptation module is specifically implemented as follows: Video features after abnormal information is suppressed and are as follows: Then, for and perform global average pooling respectively to obtain the average values representing normal instances for each: Next, and are input into a discriminator with a gradient reversal layer GRL: Among them, is a multi-layer perceptron network; GRL acts as an identity function during forward propagation; during backpropagation, GRL reverses the current gradient by multiplying by the inversion coefficient to reverse the current gradient, , such that the optimization objective of is opposite to the optimization objective of the task discriminator; Finally, the inversion loss function is defined as: Among them, represents or When represents then ; When represents then .

2. The weakly supervised video anomaly detection method based on learning with dynamically selected comparison examples according to claim 1, wherein Before the input video passes through the instance feature learning module based on contrastive learning, feature extraction and feature projection need to be performed on the input video.

3. The weakly supervised video anomaly detection method based on learning with dynamically selected comparison examples according to claim 2, wherein The specific implementation method of performing feature extraction and feature projection on the input video is as follows: S1, Feature extraction: Collect a set of abnormal videos and a set of normal videos with video-level labels; First, divide a single video C into several segments, each segment consists of 16 consecutive frames, and each segment is called an instance, denoted as c; At this time, a single video is represented as , where M is the number of segments owned by a single video C, represents the m-th instance; Subsequently, the I3D network is used as the backbone network for feature extraction: Each instance in a single video C is subjected to 10-crop data augmentation, and the 10-crop data augmentation is as follows: For each original video frame, five sub-images are intercepted from the upper left, upper right, lower left, lower right, and middle parts, and after the original image is mirrored and flipped, five sub-images are intercepted again from the upper left, upper right, lower left, lower right, and middle parts; The augmented video is input into the pre-trained I3D network to obtain video features , where represents the m-th instance after feature extraction; Finally, after feature extraction is performed on a set of abnormal videos and a set of normal videos, an abnormal training set is formed and a normal training set , represents the feature of the i-th abnormal video, represents the feature of the i-th normal video; S2, Feature projection: First, use a normal training set to offline learn a static sparse dictionary D, which consists of R sparse centers with the same shape as the instances, denoted as , representing the r-th sparse center; The learning process of the sparse dictionary is as follows: Among them, is the coefficient vector constrained by the sparsity prior, is the regularization coefficient; Then, use the static sparse dictionary D and the attention mechanism to perform feature projection on normal videos and abnormal videos; the specific calculation process is: Among them, respectively represent the features of the query, key, and value obtained through a linear function, are the input features from or , and softmax represents the activation function; is the projection on D, and its feature dimension is the same as .

4. The weakly supervised video anomaly detection method based on learning with dynamically selected comparison examples according to claim 3, wherein The instance feature learning module based on contrastive learning is specifically implemented as follows: A pair of video features respectively from and for the given input model and are respectively obtained through feature projection to get and ; , , and are input into the multi-scale temporal network MTN to learn the long-term and short-term spatio-temporal dependencies of video information, expressed as: Among them, represents MTN, and are the embedded features of , , and after passing through the multi-scale time network MTN; Channel attention is introduced, taking the M instances in and as the channel dimension, taking the instance features corresponding to each instance as the spatio-temporal dimension, calculating the difference in the channel weights of and and using it as the attention weight: Among them, is and 's attention weight, is the global average pooling layer, E is the fully connected network, is the Sigmoid function; Similarly calculate and 's attention weight ; Multiply and respectively with and to obtain the video features after the normal information is suppressed and are: Among them, "·" represents matrix multiplication; Definition and The instances in are respectively and , that is , , and respectively represent and the m-th instance in; design an instance feature contrast loss function, for , make the instance features in it move away from each other; for , make the instance features in it move closer to each other; in contrastive learning, define a non-linear projection head H, and the process of obtaining the embedded features through the non-linear projection head is as follows: Among them, represents and ; The instance feature contrast loss function is defined as: Among them, is the weight coefficient, is the natural exponential function, is the indicator function, indicating that when the function value is 1, otherwise the function value is 0.

5. The weakly-supervised video anomaly detection method based on learning with dynamically selected comparison examples according to claim 4, wherein The dynamic instance selection module is a multi-instance learning MIL module for dynamic instance selection.

6. The weakly supervised video anomaly detection method based on learning with dynamically selected comparison examples according to claim 4, characterized in that, The dynamic instance selection module is specifically implemented as follows: Increase the size of the first K largest instance features in the abnormal video sample and decrease the size of the first K largest instance features in the normal video sample; specifically as follows, First, the feature size is defined as: Among them, is or an instance feature in the calculation symbol of the second norm of the feature vector; Then, to select the first K largest instance features in the video sample, introduce the following function: Among them, represents the top K instance features with the largest feature sizes from and denotes the average feature size of the top K instance features in To learn the features of normal instances that may exist in abnormal video samples, the average value of the feature sizes of the instances in all normal video samples in the current mini-batch with a size of I will be calculated , and in each round of training, will be used as the threshold to dynamically screen the instance features in the abnormal video samples; The calculation process is as follows: Then, the following operations are performed on each abnormal video sample in the small batch: Define an empty set , for each in make a judgment. If , then store it in the set ; Next, select the smallest instance features in the abnormal video samples and propose the following function: Among them, represents the number of elements in the set When then represents the set of instance features with the smallest first feature sizes from and represents the average feature size of the first smallest instance features in When then represents the instance features with the smallest first feature sizes from and represents the average feature size of the first smallest instance features in Next, define the feature size loss function as: where b is a predefined boundary, is a weight hyperparameter, and "*" is ordinary multiplication; Finally, define a discriminator Z, which consists of a multi-layer perceptron network L and a Sigmoid function and input it into the discriminator to obtain the first anomaly score : Meanwhile, input it into the discriminator to obtain the second anomaly score : Define the feature classification loss function as: Among them, BCE is the binary cross-entropy loss function.

Citation Information

Patent Citations

  • New crown diagnosis system based on deep convolutional neural network and multi-instance learning

    CN112150442A

  • Equipment fault diagnosis method based on adversarial transfer learning and class balance loss

    CN117312950A