Weak supervision video anomaly detection method based on dynamic selection comparison instance learning
By introducing dynamic selection of comparative instance learning methods in video anomaly detection, the problems of high labeling costs, insufficient model understanding of abnormal patterns, and neglecting instance relationships in the prior art are solved, and more efficient video anomaly detection performance is achieved.
Patent Information
- Application Number
- CN202510591902.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-09
AI Technical Summary
Existing video anomaly detection technology faces problems such as high manual labeling costs, insufficient awareness of abnormal patterns of unsupervised methods, and weak supervision methods ignore instance relationships, resulting in limited detection accuracy and performance.
A weakly supervised video anomaly detection method based on dynamic selection contrast instance learning is proposed. By constructing an instance feature learning module, a dynamic instance selection module and an instance feature domain adaptive module based on contrast learning, the video anomaly detection model is optimized and the model performance is improved.
This method significantly reduces the cost of manual labeling, improves the abnormal detection performance of the model, can distinguish between normal and abnormal more accurately, and weakens the negative impact of scene similarity on the model.
Smart Images

Figure CN120107868A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of video anomaly detection, and in particular relates to a weakly supervised video anomaly detection method based on dynamic selection contrast instance learning. Background Art
[0002] Video anomaly detection refers to the detection of abnormal events in videos, and its application range is very wide, including road traffic monitoring, crowd monitoring, etc. In recent years, video anomaly detection technology has made significant progress. However, due to the unbounded nature of abnormal events in practical applications and the difficulty in collecting large-scale annotated data, video anomaly detection still faces major challenges.
[0003] Existing video anomaly detection technologies can be divided into three categories according to the annotation of the required training videos: supervised video anomaly detection, unsupervised video anomaly detection, and weakly supervised video anomaly detection. Supervised video anomaly detection uses video data with frame-level labels for training, unsupervised video anomaly detection usually assumes that only normal video data is used for training, and weakly supervised video anomaly detection uses video data with video-level labels to train the model. In actual scenarios, the manual annotation cost of supervised methods is very expensive and time-consuming. Unsupervised methods train models through frame prediction-based methods and frame reconstruction-based methods, but since this method can only use normal data for training, this method has an obvious disadvantage, that is, there are only normal data samples in the training data, and the model lacks cognition of abnormal patterns during the training process, so the abnormal boundaries of unsupervised video anomaly detection methods are not clear. Although weakly supervised video anomaly detection methods balance detection accuracy and annotation costs, and have higher detection accuracy while having a much lower cost than supervised methods, weakly supervised video anomaly detection methods are often based on multi-instance learning strategies. Such methods currently have two shortcomings: 1. They only focus on the instances with the largest feature volume in normal and abnormal videos, while ignoring potentially valuable information in other instances, which reduces the overall performance of the model. 2. They ignore the relationship between normal and abnormal instances within abnormal videos. The visual and semantic similarities of these instances themselves make it challenging to distinguish between abnormal and normal. Summary of the invention
[0004] The purpose of the present invention is to overcome the defects of the prior art and provide a weakly supervised video anomaly detection method based on dynamic selection contrastive instance learning. First, the input video is subjected to feature extraction and feature projection. Secondly, an instance feature learning module based on contrast loss is proposed to weaken the negative impact of information irrelevant to the anomaly on the model. Then, a MIL (multiple instance learning) module with dynamic instance selection is proposed based on the multi-instance learning framework to improve the model performance by learning normal instances in abnormal samples. Finally, an instance feature domain adaptation module is introduced to reduce the differences caused by different background information in the video.
[0005] To achieve the above object, the technical solution of the present invention is: a weakly supervised video anomaly detection method based on dynamic selection contrast instance learning, comprising: Construct an instance feature learning module based on contrastive learning to aggregate normal instances and separate abnormal instances; Construct a dynamic instance selection module to identify normal instance candidates with the smallest feature value in abnormal videos; An instance feature domain adaptation module is constructed to enhance the separability between normal instances and abnormal instances through domain adaptation learning.
[0006] Furthermore, before the input video passes through the instance feature learning module based on contrastive learning, feature extraction and feature projection are required for the input video.
[0007] Furthermore, the specific implementation method of feature extraction and feature projection for the input video is as follows: (1) Feature extraction: Collect a set of abnormal videos and a set of normal videos with video-level labels; First, a single video C is divided into several segments, each of which consists of 16 consecutive frames. Each segment is called an instance, denoted by c. At this time, a single video is represented as , where M is the number of segments of a single video C, represents the mth instance; Subsequently, the I3D network is used as the backbone network for feature extraction: First, each instance in a single video C is subjected to 10-corp data augmentation, that is, five sub-images of the upper left, upper right, lower left, lower right, and middle are intercepted for each original video frame, and the original image is mirrored and then five sub-images of the upper left, upper right, lower left, lower right, and middle are intercepted again; then, the enhanced video is input into the pre-trained I3D network to obtain the video features ,in represents the mth instance after feature extraction; Finally, after feature extraction, a group of abnormal videos and a group of normal videos form an abnormal training set. and the normal training set , represents the i-th abnormal video feature, represents the i-th normal video feature; (2) Feature projection: First, use the normal training set Learn a static sparse dictionary D offline. The static sparse dictionary D consists of R sparse centers d with the same shape as the instance, expressed as , Represents the rth sparse center; the learning process of the sparse dictionary is as follows: in, is a coefficient vector subject to a sparsity prior constraint, is the regularization coefficient, which is set to 0.1 in the present invention; the learning process of the sparse dictionary is to encourage each sparse center in the static sparse dictionary D to be as similar as possible to the features of each normal instance, thereby learning a sparse dictionary containing normal semantic information; Then, the static sparse dictionary D and the attention mechanism are used to perform feature projection on the video. The specific calculation process is: in, Respectively represent the features of query, key and value obtained by linear function, For or Input features of yes The projection on D has the same characteristic dimension as same.
[0008] Furthermore, the instance feature learning module based on contrastive learning is specifically implemented as follows: Given a pair of input models from and Video Features and , respectively obtained through feature projection and ;Will and With their respective projections, they are input into the multi-scale temporal network MTN to learn the long-term and short-term spatiotemporal dependencies of video information, expressed as: in, Indicates MTN, and After MTN , , and The embedding features of and The M instances in are used as the channel dimension, and each instance The corresponding instance features are used as the spatiotemporal dimensions to calculate and The difference of the channel weights is used as the attention weight: in, is the global average pooling layer, E is the fully connected network, is the Sigmoid function; similarly, and The attention weight ;Will and and their attention weights respectively and Multiply them together to get the video features after the normal information is suppressed. and for: Where “·” represents matrix multiplication; definition and The examples in are and ,Right now , , and Respectively and The mth instance in ; design an instance feature contrast loss function for , encouraging the features of each instance to stay away from each other; , encouraging the features of each instance to be close to each other; in contrastive learning, a nonlinear projection head H is defined, and the process of obtaining embedded features through the nonlinear projection head is as follows: in, represent and ; The instance feature contrast loss function is defined as: in, is the weight coefficient, is the natural exponential function, is the indicator function, which means when The function value is 1 when , otherwise the function value is 0.
[0009] Furthermore, the dynamic instance selection module is a multiple instance learning MIL module for dynamic instance selection.
[0010] Furthermore, the dynamic instance selection module is specifically implemented as follows: The feature sizes of the top K instances in abnormal video samples are encouraged to increase, while the feature sizes of the top K instances in normal video samples are encouraged to decrease; specifically, First, the feature size is defined as: in, for or Instance features in Compute the sign of the two-norm of the eigenvector; Then, in order to select the top K instance features in the video sample, the following function is introduced: in, Representatives from The K instance features with the largest feature size in , Represents a video sample The average feature size of the top K largest instance features; In order to learn the normal instance features that may exist in abnormal video samples, all normal video samples in the current batch size I are calculated The average feature size of the instance features in , and in each round of training As a threshold for abnormal video samples Dynamically filter the instance features in The calculation process is as follows: Then, for each abnormal video sample in the small batch There are the following operations: define an empty collection ,for Each of Make a judgment, if , then store it in the collection middle; Next, select the front For small instance features, the following function is proposed: in, Representing a collection The number of elements in , when hour, Representatives from The smallest feature size Instance features, Represents a video sample Center front The average feature size of small instance features; when hour, Representatives from The smallest feature size Instance features, Represents a video sample Center front The average feature size of small instance features; Next, the feature size loss function is defined as: Where b is a predefined boundary, is the weight hyperparameter, “*” is ordinary multiplication; Finally, a discriminator Z is defined, which consists of a multi-layer perceptron network L and a Sigmoid function Composition, will Input the discriminator to get the anomaly score : At the same time, Input the discriminator to get the anomaly score : The feature classification loss function is defined as: Among them, BCE is the binary cross entropy loss function.
[0011] Furthermore, the instance feature domain adaptation module is specifically implemented as follows: Video features after abnormal information is suppressed and for: Then, yes and Perform global average pooling to obtain the average value of each representative normal instance: Next, and Input to the discriminator with gradient reversal layer GRL: in, is a multi-layer perceptron network; GRL acts as the identity function during forward propagation: in, is the input of GRL; During backpropagation, GRL is multiplied by the inversion coefficient to reverse the current gradient, , so that The optimization goal of is opposite to that of the task discriminator; Finally, invert the loss function Defined as: in, represent or ,when represent hour, ;when represent hour, .
[0012] Compared with the prior art, the present invention has the following beneficial effects: 1. The present invention optimizes the video anomaly detection model by using weakly supervised multi-instance learning. The data used to train the model are normal and abnormal data with video-level labels. Compared with the supervised method, the manual labeling cost of the video-level labels used by the proposed model is much lower than that of the supervised method, and it is more valuable. Compared with the unsupervised video anomaly detection method, weakly supervised information can greatly improve the model performance, making the model's anomaly detection performance far exceed the unsupervised method.
[0013] 2. The existing model only focuses on several instances with the largest feature size in normal and abnormal videos, while ignoring the information contained in the features of other instances, especially the useful information contained in the instances with the smallest feature size in abnormal videos, which limits the performance of the model. The method of the present invention enables the model to learn the features of normal segments in abnormal videos, thereby distinguishing normal from abnormal more finely.
[0014] 3. Existing models ignore the fact that instances in the same video have high similarity due to the same scene, and this similarity still exists between normal and abnormal instances, which is not conducive to MIL distinguishing abnormal and normal instances. The method of the present invention fully weakens the negative impact of scene similarity on the model through contrastive learning and domain adaptive learning, thereby improving the anomaly detection performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 The figure is a diagram of the overall execution process of the method of the present invention.
[0016] Figure 2 Examples of video data screens: (a) normal video screen; (b) abnormal video screen.
[0017] Figure 3 It is the feature extraction and feature projection module.
[0018] Figure 4 It is an instance feature learning module based on contrastive learning.
[0019] Figure 5 The MIL module selected for the dynamic instance.
[0020] Figure 6 It is an instance feature domain adaptation module.
[0021] Figure 7 This is the loss curve of ShanghaiTech.
[0022] Figure 8 is the UCF-Crime loss curve.
[0023] Fig. 9 Visualize the detection results.
[0024] Fig.10 The left figure is the AUC-ROC curve of the detection results on ShanghaiTech and UCF-Crime, and the right figure is the AUC-ROC curve of the detection results on UCF-Crime.
[0025] Fig.11 The present invention is a flowchart for implementing the method. DETAILED DESCRIPTION
[0026] The technical solution of the present invention is described in detail below in conjunction with the accompanying drawings.
[0027] The present invention provides a weakly supervised video anomaly detection method based on dynamic selection contrast instance learning, comprising: Construct an instance feature learning module based on contrastive learning to aggregate normal instances and separate abnormal instances; Construct a dynamic instance selection module to identify normal instance candidates with the smallest feature value in abnormal videos; An instance feature domain adaptation module is constructed to enhance the separability between normal instances and abnormal instances through domain adaptation learning.
[0028] The following is the specific implementation process of the present invention.
[0029] like Figure 1 As shown, the present invention is a weakly supervised video anomaly detection method based on dynamic selection contrast instance learning. First, the input video is subjected to feature extraction and feature projection. Secondly, an instance feature learning module based on contrast loss is proposed to weaken the negative impact of information irrelevant to the anomaly on the model. Then, based on the multi-instance learning framework, a MIL (multiple instance learning) module with dynamic instance selection is proposed to improve the model performance by learning normal instances in abnormal samples. Finally, an instance feature domain adaptation module is introduced to weaken the differences caused by different background information in the video. The specific implementation steps are as follows: Step 1: Video data input The data used in this invention comes from the public benchmark video datasets ShanghaiTech and UCF-Crime. Among them, ShanghaiTech is a challenging multi-scene dataset consisting of 13 campus scenes with different lighting conditions and camera angles. The dataset contains a total of 437 videos, including 238 training videos and 199 test videos. UCF-Crime is a large-scale anomaly detection dataset consisting of 1,900 untrimmed videos captured from real-world street and indoor surveillance cameras. Compared with ShanghaiTech, the background of UCF-Crime is more complex and diverse. The dataset includes 1,610 training videos and 290 test videos. Both datasets contain different video surveillance scenes, including abnormal human behavior and abnormal events caused by vehicles or equipment. The division of the two datasets is shown in Table 1. The video data screen samples in the dataset are as follows. Figure 2 As shown ( Figure 2 (a) is a normal video screen; Figure 2 (b) is an abnormal video image).
[0030] Step 2: Feature extraction and feature projection module The feature extraction and feature projection process is as follows Figure 3 As shown in Figure 2, we collect a set of abnormal videos and a set of normal videos with video-level labels to train our model.
[0031] Feature extraction: First, a single video C is divided into several segments, each of which consists of 16 consecutive frames. Each segment is called an instance, denoted by c; at this time, a single video is represented as , where M is the number of segments of a single video C, represents the mth instance. Then, the I3D network is used as the backbone network for feature extraction. First, each instance in a single video C is enhanced by 10-corp data, that is, five sub-images are captured from each original video frame, namely, the upper left, upper right, lower left, lower right, and middle. Then, the original image is flipped and five sub-images are captured again in the same way. Then, the enhanced video is input into the pre-trained I3D network to obtain the video features. ,in represents the mth instance after feature extraction. Finally, perform the above operations on all abnormal videos and normal videos to form an abnormal training set and the normal training set .
[0032] Feature projection: using normal training set Learn a static sparse dictionary D offline. The static sparse dictionary D consists of R sparse centers d with the same shape as the instance, expressed as , Represents the rth sparse center; the learning process of the sparse dictionary is as follows: in, is a coefficient vector subject to a sparsity prior constraint, is the regularization coefficient, which is set to 0.1 in the present invention; the learning process of the sparse dictionary is to encourage each sparse center in the static sparse dictionary D to be as similar as possible to the features of each normal instance, thereby learning a sparse dictionary containing normal semantic information; Then, the static sparse dictionary D and the attention mechanism are used to perform feature projection on the video. The specific calculation process is: in, Respectively represent the features of query, key and value obtained by linear function, For or Input features of yes Projection on D. Since D is a sparse dictionary of normal samples, for The normal features obtained by projection on D have the same feature dimension as same.
[0033] Step 3: Instance feature learning module based on contrastive learning The instance feature learning module based on contrastive learning aims to encourage instances in abnormal video samples to have larger feature distances and instances in normal video samples to have smaller feature distances, thereby weakening the correlation between each instance in the abnormal samples and reducing the negative impact of irrelevant information in instance features on anomaly detection. Its structure is as follows Figure 4 The specific method is as follows: given a pair of input models from and Video Features and , respectively obtained through feature projection and ; Then, and The projections are then fed into a multi-scale temporal network MTN to learn the long-term and short-term spatiotemporal dependencies of video information. The process is expressed as: in, Indicates MTN, and After MTN , , and The channel attention is introduced to suppress the normal event information in the video embedding feature. and The M instances in are used as the channel dimension, and each instance The corresponding instance features are used as the spatiotemporal dimensions to calculate and The difference of the channel weights is used as the attention weight: in, is the global average pooling layer, E is the fully connected network FC Layer, is the Sigmoid function; similarly, and The attention weight ;Will and and their attention weights respectively and Multiply them together to get the video features after the normal information is suppressed. and for: Where “·” represents matrix multiplication; definition and The examples in are and ,Right now , , and Respectively and For the mth instance in the video level label , which may contain some normal instances and some abnormal instances, but since these instances are from the same video, the feature information of the instances that are not related to abnormal events (such as the background information of the picture) is very similar, especially when the abnormal event subject occupies a small proportion of the picture. For the video anomaly detection task, the model needs to pay as much attention to the information related to the abnormal event subject as possible, while avoiding interference from the information of irrelevant abnormal events as much as possible. Therefore, an instance feature comparison loss function is designed. For , encouraging each instance feature to stay away from the other instance features, thereby breaking the problem of normal and abnormal instances being similar due to irrelevant information such as background; , the instance features are encouraged to be close to each other, thus avoiding the model from focusing too much on the feature differences between normal instances. In contrastive learning, the use of nonlinear transformation projection heads can form and maintain more information in the encoder. To this end, a nonlinear projection head H is defined. The process of obtaining embedded features through the nonlinear projection head is as follows: in, represent and ; The instance feature contrast loss function is defined as: in, is the weight coefficient, is the natural exponential function, is the indicator function, which means when The function value is 1 when , otherwise the function value is 0.
[0034] Step 4: MIL module for dynamic instance selection In weakly supervised video anomaly detection, every instance in a normal video sample is normal, while in an abnormal video sample, some instances are abnormal and others are normal. Existing weakly supervised video anomaly detection methods use the traditional MIL strategy to distinguish between normal and abnormal by encouraging the top K instances with the largest feature size in abnormal video samples to be as large as possible, and vice versa, the top K instances with the largest feature size in normal video samples to be as small as possible. However, existing methods ignore the learning of normal instance features that may exist in abnormal video samples, resulting in limited detection performance. Therefore, we propose a MIL module for dynamic instance selection, such as Figure 5 shown.
[0035] Based on the prior information that abnormal instances in MIL tend to have larger feature sizes than normal instances, we encourage the feature sizes of the top K largest instances in abnormal video samples to increase, while encouraging the feature sizes of the top K largest instances in normal video samples to decrease. Specifically, the feature size is defined as: in, for or Instance features in Compute the sign of the two-norm of the eigenvector; Then, in order to select the top K instance features in the video sample, the following function is introduced: in, Representatives from The K instance features with the largest feature size in , Represents a video sample The average feature size of the top K largest instance features; In order to learn the normal instance features that may exist in abnormal video samples, all normal video samples in the minibatch of the current batch-size I are calculated. The average feature size of the instance features in , and in each round of training As a threshold for abnormal video samples Dynamically filter the instance features in . The calculation process is as follows: Then, for each abnormal video sample in the minibatch There are the following operations: define an empty collection ,for Each of Make a judgment, if , then store it in the collection middle.
[0036] Next, select the front For small instance features, the following function is proposed: in, Representing a collection The number of elements in , when hour, Representatives from The smallest feature size Instance features, Represents a video sample Center front The average feature size of small instance features; when hour, Representatives from The smallest feature size Instance features, Represents a video sample Center front The average feature size of small instance features; Next, the feature size loss function is defined as: Where b is a predefined boundary, is the weight hyperparameter, “*” is ordinary multiplication; Finally, a discriminator Z is defined, which consists of a multi-layer perceptron network L and a Sigmoid function Composition, will Input the discriminator to get the anomaly score : At the same time, Input the discriminator to get the anomaly score : The feature classification loss function is defined as: Among them, BCE is the binary cross entropy loss function.
[0037] Step 5: Instance feature domain adaptation module In the MIL-based method, only the instance features with the largest feature sizes in the abnormal / normal video samples are encouraged to be enlarged / reduced, so as to separate abnormal instances from normal instances. The existing MIL-based method is based on the prior knowledge that the features of normal instances in abnormal video samples are similar to those of normal instances in normal video samples. Therefore, encouraging the reduction of the features of the instances with the largest feature sizes in the normal video samples indirectly encourages the reduction of the features of normal instances in abnormal video samples. In the backup method, it is also hoped that the features of normal instances in abnormal video samples have a high similarity with the features of normal instances in normal video samples.
[0038] Inspired by domain adaptation, an instance feature domain adaptation module is proposed, such as Figure 6 First, in step 3, the channel attention is introduced to suppress the normal event information in the video embedding feature, and the attention weight for focusing on the abnormal event information is obtained. and ,and and is the attention weight for normal event information. Therefore, the video features after abnormal information is suppressed are and for: Then, yes and Perform global average pooling to obtain the average value of each representative normal instance: in, is the global average pooling layer, which targets the channel dimension, that is, it performs an average operation on the original M instance features, so and The shape and single instance The dimensions and shapes of and Input to the discriminator with gradient reversal layer (GRL): in, is a multi-layer perceptron network; GRL acts as the identity function during forward propagation: in, is the input of GRL; during back propagation, GRL multiplies the inversion coefficient by To reverse the current gradient ( ), so that the MTN feature extractor The optimization goal of is opposite to that of the task discriminator, that is, to encourage When extracting features, the difference between the normal instance features in the abnormal video samples and the normal instance features in the normal video samples is ignored, so that the two have a high similarity.
[0039] Finally, invert the loss function Defined as: in, represent or ,when represent hour, ;when represent hour, .
[0040] Step 6: Abnormal frame determination During the model inference process, input a video feature extracted by the pre-trained I3D network , the output of the discriminator Z is used as the anomaly score ,The overall reasoning process is shown in Table 2.
[0041] Step 7: Model training and evaluation Load the unsupervised video anomaly detection model based on multi-subcluster memory prototype; input the training set video data of ShanghaiTech and UCF-Crime into the model for training. The method uses the Adam optimizer for training, the learning rate is set to 0.001, the weight decay is 0.005, and the batch size is set to 32. The training is divided into two stages. The first 1000 epochs are the first stage. , the next 1000 epochs are the second stage. .
[0042] During the iteration process, the model loss curve reflects the model's learning degree of video features and its ability to distinguish abnormal and normal instances. The total loss curve on the ShanghaiTech dataset is as follows: Figure 7 As shown, the total loss curve on the UCF-Crime dataset is as follows Figure 8 As shown in the figure, the total loss curves on both datasets show a small upward trend after the 1000th iteration. This phenomenon is caused by the fact that the parameter ,The dynamic instance selection module begins to play a role and has an impact on the network parameters, thus causing fluctuations, and then the curve continues to show a downward trend.
[0043] The unsupervised video anomaly detection model based on multi-subcluster memory prototypes is used to detect the videos in the test set. ShanghaiTech (" Fig. 9 (a)", Fig. 9 (b)”, Fig. 9 (c)") and UCF-Crime (" Fig. 9 in (d),” Fig. 9 in (e)"," Fig. 9 (f)” to visualize the anomaly score and the true value label. The detection results are shown in Fig. 9 As shown, the curve in the figure represents the abnormal score curve, and the dark background area represents the abnormal events manually marked.
[0044] The evaluation index used in this method is the area under the frame-level ROC curve of video anomaly detection: AUC. The AUC-ROC curves of the proposed model's detection results on ShanghaiTech and UCF-Crime are shown in Figure 2. Fig.10 As shown: The test set is input into the model for video anomaly detection. Fig.10 The AUC-ROC curve shows that the detection effect can achieve relatively impressive results. According to the AUC-ROC curve, the model still has a small number of false positives (normal samples are misclassified as abnormal samples) and false negatives (abnormal samples are misclassified as normal samples) during the detection process. The reason is that there are some difficult samples in the test set video, such as minor abnormal events or rare but normal events.
[0045] The present invention defines a baseline model and conducts ablation experiments on the three key modules of the proposed model. The experimental results are shown in Table 3, where A represents the instance feature learning module based on contrastive learning, B represents the MIL module for dynamic instance selection, and C represents the instance feature domain adaptation module.
[0046] As can be seen from Table 3, when any one of the three modules proposed is added separately, the model performance is improved; when any two modules are added at the same time, the performance will be further improved; and when the three modules work together, the proposed model achieves the best frame-level AUC of 97.57% on the ShanghaiTech dataset and 85.65% on the UCF-Crime dataset. The ablation experiment results prove that the addition of each module proposed by the present invention can improve the model performance.
[0047] Table 4 gives the frame-level AUC results of the proposed method on the ShanghaiTech and UCF-Crime datasets and the comparison with the current SOTA (State-of-the-art) method. Among the comparative methods, RTFM is a robust temporal feature amplitude learning model, which trains feature amplitude learning functions to effectively identify positive samples; Wang et al. is a video anomaly detection model based on a multi-instance ranking algorithm; BE-WVAD is a BE-WSVAD model, which uses a binary network enhancement strategy during training to significantly improve the detection accuracy; DAR is a decoupling and resolution model for video anomaly detection, which consists of two modules: a temporal proposal producer and an online anomaly locator to collaboratively complete the detection task; Park et al. is a normal-guided multi-instance learning model, which encodes multiple normal patterns from normal videos to build a similarity-based anomaly classifier; Tsiktsiris et al. is an intelligent surveillance model based on a spatiotemporal autoencoder architecture and uses a center-weighted loss function to learn the features of normal videos; HSN is a human scene-based model that learns discriminative representations by capturing subtle and significant features in a dissociated manner. It can be seen from the figure that the proposed method outperforms the existing SOTA unsupervised learning method and achieves the best frame-level AUC results of 97.57% and 85.65% on the ShanghaiTech dataset and UCF-Crime dataset, respectively.
[0048] Step 8: Model call Read the trained video anomaly detection model stored in the cloud or locally to perform video anomaly detection. Specifically, obtain the video stream from the surveillance video and input it into the model. The value is compared with a preset threshold. If the anomaly score is greater than the threshold, it is determined to be abnormal and an alarm is generated. If the anomaly score is less than or equal to the threshold, it is determined to be normal.
[0049] The above are preferred embodiments of the present invention. Any changes made according to the technical solution of the present invention, as long as the resulting functions do not exceed the scope of the technical solution of the present invention, belong to the protection scope of the present invention.
Claims
1. A weakly supervised video anomaly detection method based on dynamic selection contrastive instance learning, characterized in that: include: Construct an instance feature learning module based on contrastive learning to aggregate normal instances and separate abnormal instances; Construct a dynamic instance selection module to identify normal instance candidates with the smallest feature value in abnormal videos; An instance feature domain adaptation module is constructed to enhance the separability between normal instances and abnormal instances through domain adaptation learning.
2. The weakly supervised video anomaly detection method based on dynamic selection contrastive instance learning according to claim 1 is characterized in that: Before the input video passes through the instance feature learning module based on contrastive learning, feature extraction and feature projection are required for the input video.
3. The weakly supervised video anomaly detection method based on dynamic selection contrastive instance learning according to claim 2 is characterized in that: The specific implementation of feature extraction and feature projection for input video is as follows: (1) Feature extraction: Collect a set of abnormal videos and a set of normal videos with video-level labels; First, a single video C is divided into several segments, each segment consists of 16 consecutive frames, each segment is called an instance, denoted by c; At this point, a single video is represented as , where M is the number of segments of a single video C, represents the mth instance; Subsequently, the I3D network is used as the backbone network for feature extraction: First, each instance in a single video C is subjected to 10-corp data augmentation, that is, five sub-images of the upper left, upper right, lower left, lower right, and middle are intercepted for each original video frame, and the original image is mirrored and then five sub-images of the upper left, upper right, lower left, lower right, and middle are intercepted again; then, the enhanced video is input into the pre-trained I3D network to obtain the video features ,in represents the mth instance after feature extraction; Finally, after feature extraction, a group of abnormal videos and a group of normal videos form an abnormal training set. and the normal training set , represents the i-th abnormal video feature, represents the i-th normal video feature; (2) Feature projection: First, use the normal training set Learn a static sparse dictionary D offline. The static sparse dictionary D consists of R sparse centers d with the same shape as the instance, expressed as , represents the rth sparse center; The learning process of the sparse dictionary is as follows: in, is a coefficient vector subject to a sparsity prior constraint, is the regularization coefficient; the learning process of the sparse dictionary is to encourage each sparse center in the static sparse dictionary D to be as similar as possible to the features of each normal instance, so as to learn a sparse dictionary containing normal semantic information; Then, the static sparse dictionary D and the attention mechanism are used to perform feature projection on the video. The specific calculation process is: in, Respectively represent the features of query, key and value obtained by linear function, For or Input features, softmax represents the activation function; yes The projection on D has the same characteristic dimension as same.
4. The weakly supervised video anomaly detection method based on dynamic selection contrastive instance learning according to claim 3 is characterized in that: The instance feature learning module based on contrastive learning is specifically implemented as follows: Given a pair of input models from and Video Features and , respectively obtained through feature projection and ;Will and With their respective projections, they are input into the multi-scale temporal network MTN to learn the long-term and short-term spatiotemporal dependencies of video information, expressed as: in, Indicates MTN, and They are respectively after the multi-scale time network MTN , , and The embedding features of and The M instances in are used as the channel dimension, and each instance The corresponding instance features are used as the spatiotemporal dimensions to calculate and The difference of the channel weights is used as the attention weight: in, is the global average pooling layer, E is the fully connected network, is the Sigmoid function; similarly calculated and The attention weight ;Will and Respectively and Multiply them together to get the video features after the normal information is suppressed. and for: Where "·" represents matrix multiplication; definition and The examples in are and ,Right now , , and Respectively and The mth instance in ; design an instance feature contrast loss function for , encouraging the features of each instance to stay away from each other; , encouraging the features of each instance to be close to each other; in contrastive learning, a nonlinear projection head H is defined, and the process of obtaining embedded features through the nonlinear projection head is as follows: in, represent and ; The instance feature contrast loss function is defined as: in, is the weight coefficient, is the natural exponential function, is the indicator function, which means when The function value is 1 when , otherwise the function value is 0.
5. The weakly supervised video anomaly detection method based on dynamic selection contrastive instance learning according to claim 4 is characterized in that: The dynamic instance selection module is a multiple instance learning MIL module for dynamic instance selection.
6. The weakly supervised video anomaly detection method based on dynamic selection contrastive instance learning according to claim 4 is characterized in that: Dynamic instance selection module, specifically implemented as follows: The feature sizes of the top K instances in abnormal video samples are encouraged to increase, while the feature sizes of the top K instances in normal video samples are encouraged to decrease; specifically, First, the feature size is defined as: in, for or Instance features in Compute the sign of the two-norm of the eigenvector; Then, in order to select the top K instance features in the video sample, the following function is introduced: in, Representatives from The K instance features with the largest feature size in , Represents a video sample The average feature size of the top K largest instance features; In order to learn the normal instance features that may exist in abnormal video samples, all normal video samples in the current batch size I are calculated The average feature size of the instance features in , and in each round of training As a threshold for abnormal video samples Dynamically filter the instance features in The calculation process is as follows: Then, for each abnormal video sample in the small batch There are the following operations: define an empty collection ,for Each of Make a judgment, if , then store it in the collection middle; Next, select the front For small instance features, the following function is proposed: in, Representing a collection The number of elements in , when hour, Representatives from The smallest feature size Instance features, Represents a video sample Center front The average feature size of small instance features; when hour, Representatives from The smallest feature size Instance features, Represents a video sample Center front The average feature size of small instance features; Next, the feature size loss function is defined as: Where b is a predefined boundary, is the weight hyperparameter, "*" is ordinary multiplication; Finally, a discriminator Z is defined, which consists of a multi-layer perceptron network L and a Sigmoid function Composition, will Input the discriminator to get the anomaly score : At the same time, Input the discriminator to get the anomaly score : The feature classification loss function is defined as: Among them, BCE is the binary cross entropy loss function.
7. The weakly supervised video anomaly detection method based on dynamic selection contrastive instance learning according to claim 6 is characterized in that: The instance feature domain adaptation module is specifically implemented as follows: Video features after abnormal information is suppressed and for: Then, yes and Perform global average pooling to obtain the average value of each representative normal instance: Next, and Input to the discriminator with gradient reversal layer GRL: in, is a multi-layer perceptron network; GRL acts as the identity function during forward propagation: in, is the input of GRL; During backpropagation, GRL is multiplied by the inversion coefficient to reverse the current gradient, , so that The optimization goal of is opposite to that of the task discriminator; Finally, invert the loss function Defined as: in, represent or ,when represent hour, ;when represent hour, .
Citation Information
Patent Citations
New crown diagnosis system based on deep convolutional neural network and multi-instance learning
CN112150442A
Equipment fault diagnosis method based on adversarial transfer learning and class balance loss
CN117312950A
KR20240175437A