Video anomaly detection method based on two-stream dilated 3D convolutional network and new loss function

By using a dual-stream expanded 3D convolutional network and a new loss function method in video anomaly detection, the features of video instances are extracted and the neural network is trained, which solves the problems of insufficient data and high labeling costs in the prior art, and realizes high-precision abnormal event detection.

CN115731418BActive Publication Date: 2025-06-06NANJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211475696.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-23
Publication Date
2025-06-06
Estimated Expiration
2042-11-23

AI Technical Summary

Technical Problem

The existing video anomaly detection methods face the problems of insufficient data and high labeling costs during training and detection, resulting in low detection accuracy.

Method used

The video anomaly detection method based on a dual-stream expanded 3D convolutional network and a new loss function is adopted. The visual features and optical flow characteristics of video instances are extracted through the I3D network, and the three-layer fully connected neural network is trained using the new loss function to build a network model that can effectively detect abnormal events.

Benefits of technology

Through weak supervision learning, this method reduces the cost of manual labeling, improves the accuracy of abnormal event detection, can effectively distinguish abnormal and normal events, and significantly improves the detection effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115731418B_ABST
    Figure CN115731418B_ABST
Patent Text Reader

Abstract

The video anomaly detection method based on the dual-stream dilated 3D convolutional network and the new loss function first defines each video as a video package, divides the positive package and the negative package, and then divides it into several video clips according to the same number of frames, and each clip is regarded as a video instance; then it is input into the dual-stream dilated 3D convolutional network I3D to extract visual features; at the same time, the optical flow image of the video instance is obtained through the optical flow extraction network, and then input into the dual-stream dilated 3D convolutional network to extract optical flow features; then the image features and optical flow features are spliced ​​to obtain the I3D features of the video instance; finally, the I3D features are input into the fully connected classification network, and the fully connected classification network is trained through the new loss function, and finally a classification network that can effectively detect abnormal events is obtained. This method uses video-level labels instead of frame-level labels during training, which effectively reduces the cost of manual annotation and can significantly improve the accuracy of the model in detecting abnormal events.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of video anomaly detection, and specifically is a video anomaly detection method based on a dual-stream dilated 3D convolutional network and a new loss function. Background Art

[0002] Video anomaly detection is not only an important and active research topic in the field of computer vision, but also a very challenging task. The task of video anomaly detection is to find events that are obviously inconsistent with normal events in a video, such as fire, traffic accidents, or crowd trampling. However, due to the wide variety of abnormal events and their low frequency of occurrence, it is impossible to collect all and enough abnormal events to train the model. At the same time, labeling a large number of video frames will also consume a lot of manpower and time.

[0003] Traditional video anomaly detection methods only use normal events for training, and are mainly divided into two categories: video frame reconstruction and video frame prediction. Video frame reconstruction is to pass the current frame through an autoencoder to reconstruct a reconstructed frame, and calculate the reconstruction error between the current frame and the reconstructed frame to detect abnormal events, such as the AE method (Mahmudul Hasan, Jonghyun Choi, Jan Neumann, Amit K Roy-Chowdhury, and Larry S Davis. Learning temporal regularity in video sequences. In CVPR, 2016.), but due to the strong generalization ability of the autoencoder, abnormal events can also be well reconstructed, resulting in low accuracy. Video frame prediction is to predict the future frames by predicting the previous frames, and then input the future frames into the discriminator to determine whether they are abnormal events, thereby realizing the detection of abnormal events, such as (Wen Liu, Weixin Luo, Dongze Lian, Shenghua Gao. Future Frame Prediction for Anomaly Detection--A New Baseline. In CVPR, 2018.), but some normal events such as people squatting, opening doors, etc. cannot be well predicted, resulting in low detection accuracy. Summary of the invention

[0004] To solve the above problems, the present invention provides a video anomaly detection method based on a dual-stream inflated 3D convolutional network and a new loss function. The method uses a dual-stream inflated 3D convolutional network (I3D) as the backbone network to extract I3D features of video instances, and then inputs the I3D features into a three-layer fully connected layer network guided by a new loss function to train the network, thereby constructing a network model that can effectively detect abnormal events.

[0005] The video anomaly detection method based on the dual-stream dilated 3D convolutional network and the new loss function specifically includes the following steps:

[0006] S1, collect video data, divide the video data into training set and test set, define each video as a video package, in the training set, define the video package containing one or more abnormal events as a positive package, and define the video package containing only normal events as a negative package;

[0007] S2, divide each video packet of the training set and the test set in S1 into several video segments according to the same number of frames, and regard each video segment as a video instance;

[0008] S3, input the training set video instance obtained in S2 into the I3D network, so as to extract the visual features of each video instance;

[0009] S4, inputting the training set video instance obtained in S2 into the optical flow extraction network, thereby extracting the optical flow image of each video instance, and then inputting the optical flow image into the I3D network to obtain the optical flow feature of each video instance;

[0010] S5, concatenates the visual features and optical flow features obtained in S3 and S4 to form the I3D features required for the final classification;

[0011] S6, inputting the I3D features obtained in S5 into a three-layer fully connected neural network, and training the neural network through a new loss function, and finally training a network model that can distinguish abnormal events from normal events;

[0012] S7, sequentially performing steps S3, S4 and S5 on the test set video instance obtained in S2 to obtain the I3D features of the test set video instance;

[0013] S8, input the I3D features of the test set video instances obtained in S7 into the video anomaly detection model trained in S6, and then obtain the anomaly score of the test set video instances, thereby achieving effective detection of abnormal events.

[0014] The beneficial effects of the present invention are:

[0015] (1) The present invention combines the weakly supervised learning method and uses video-level labels instead of frame-level labels. While not requiring a large amount of manual labeling costs, it can greatly improve the accuracy of the model in detecting abnormal events.

[0016] (2) The present invention introduces an I3D network as a feature extraction backbone network, which can well capture the visual features and optical flow features contained in the video instance, that is, the spatiotemporal features, thereby facilitating subsequent classification and recognition work.

[0017] (3) The present invention concatenates the visual features and optical flow features of the video instance as the final classification features, which significantly increases the amount of information contained in the features and is beneficial to subsequent classification and recognition work.

[0018] (4) The present invention constructs a new loss function, which is proposed based on the sparsity of abnormal events and the similarity of normal events. With the help of this loss function, a video anomaly detection network model with significantly improved accuracy can be trained. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 4 is a flow chart of an anomaly detection method in an embodiment of the present invention.

[0020] Figure 2 4 is a structural diagram of an I3D network in an embodiment of the present invention.

[0021] Figure 3 4 is a network model structure diagram of the anomaly detection method in an embodiment of the present invention. DETAILED DESCRIPTION

[0022] The technical solution of the present invention is further described in detail below in conjunction with the accompanying drawings.

[0023] The present invention is a video anomaly detection method based on a dual-stream dilated 3D convolutional network and a new loss function, and the anomaly detection method specifically comprises the following steps:

[0024] S1, collect video data, divide the video data into training set and test set, define each video as a video package, and define the video package containing one or more abnormal events as positive package B in the training set. p , the video packet containing only normal events is defined as negative packet B n .

[0025] S2, each video packet of the training set and the test set in S1 is divided into several video segments according to 16 consecutive frames, and each video segment is regarded as a video instance. The positive packet video instances are represented by Indicates that the negative bag video instances are used in turn Indicates that L pand L n Represents the number of video instances of positive and negative packets respectively.

[0026] S3, inputs the training set video instance obtained in S2 into a two-stream inflated 3D convolutional network (I3D) with two parallel 3D convolution modules, so as to extract the visual features of each video instance. Figure 2 ,The I3D network consists of two parallel 3D convolution modules, which extract the ,visual features and optical flow features of the input video instance ,respectively, and finally concatenate the two features to obtain the ,I3D features of the video instance.

[0027] S4, input the training set video instance obtained in S2 into the commonly used flownet2 network, so as to extract the optical flow image of each video instance, and then input the optical flow image into the I3D network to obtain the optical flow features of each video instance.

[0028] S5, the visual features and optical flow features obtained in S3 and S4 are concatenated in the channel dimension to form the I3D features required for the final classification. The concatenation operation is a simple dimensional concatenation, such as the concatenation of the features of (1, a) and (1, b) dimensions to form the features of (1, a+b).

[0029] S6, input the I3D features obtained in S5 into a three-layer fully connected neural network (structure is 2048, 512, 1, as Figure 3 ), and the neural network is trained by the proposed new loss function, and finally a network model that can effectively distinguish abnormal events from normal events is obtained. The construction process of the new loss function is as follows:

[0030] S6-1, for a given positive packet video instance p i and negative bag video instance n i , expect p i The anomaly score is greater than n i , so the following loss function can be obtained:

[0031] s(p i )>s(n i )

[0032] Among them, s(p i ) and s(n i ) represent the video instance p i and n i The anomaly score.

[0033] S6-2, in order to make the anomaly score of the positive video instance greater than that of the negative video instance, the loss function in S6-1 is further improved to obtain a new loss function as follows:

[0034]

[0035] The mean method means taking the average of all instance anomaly scores in the video package.

[0036] S6-3, since only some instances in the positive packet are abnormal events, the rest are still normal events, and the first L of the positive and negative packets are taken respectively. P / K and L n / The K largest instance scores are sorted and compared, and the improved loss function is as follows:

[0037]

[0038] Among them, k_mean means taking the first L P / K and L n / K video instance scores with the largest anomaly scores, and calculate their average; K is a hyperparameter, that is, the ratio of the number of selected instance scores to the total number of video instances; so that while ensuring the separation of positive and negative instance scores, it is beneficial to separate the normal instance and abnormal instance scores in the negative package, thereby improving the detection accuracy. K is a hyperparameter, that is, the ratio of the number of selected instance scores to the total number of video instances. For different data sets, the value of K is adjusted accordingly.

[0039] S6-4, appropriately adjust the loss function obtained in S6-3 to obtain the new loss function as follows:

[0040]

[0041] S6-5, for negative packets, since they only contain normal instances, the scores of all negative packet video instances should be close, and then a smoothness constraint is proposed for the scores of negative packet video instances. The loss function is as follows:

[0042]

[0043] The var method takes the variance of the scores of all video instances in the negative bag.

[0044] S6-6, the final constructed loss function is as follows:

[0045] loss = k_loss + λs_loss

[0046] Among them, λ is a hyperparameter, that is, the weight coefficient corresponding to s_loss.

[0047] S7, sequentially pass the test set video instance obtained in S2 through steps S3, S4 and S5 to obtain the I3D features of the test set video instance.

[0048] S8, input the I3D features of the test set video instances obtained in S7 into the video anomaly detection model trained in S6, and then obtain the anomaly score of the test set video instances, thereby achieving effective detection of abnormal events.

[0049] The present invention adopts a weakly supervised learning method, which only needs to mark the video-level labels. This approach can avoid the cost of a large number of manual annotations and can significantly improve the detection accuracy. The present invention uses an I3D network as a feature extraction network, which can effectively extract the visual features and optical flow features of video instances, that is, spatiotemporal features (visual features are spatial features, and optical flow features are temporal features), combined with the proposed new loss function, taking into account the high scores of abnormal instances and the similar scores of normal instances, which can effectively improve the model's detection ability for abnormal events.

[0050] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments.

[0051] A. Experimental conditions

[0052] 1. Experimental Database

[0053] The training and testing are performed on the shanghaiTech dataset. The dataset is divided into a training set and a test set during anomaly detection. See Table 1 for details.

[0054] Table 1 Detailed introduction of the dataset

[0055]

[0056] 2. Experimental parameter settings

[0057] The fixed parameter settings of the model are shown in Table 2 below:

[0058] Table 2 Model fixed parameters

[0059] K λ 4 20

[0060] B. Experimental results evaluation criteria

[0061] This model is designed for anomaly detection settings, and uses AUC to measure the detection effect. The higher the AUC value, the better the model effect. The calculation formula of AUC is as follows:

[0062]

[0063] Among them, M is the number of abnormal event samples, N is the number of normal event samples, and predpos Indicates the abnormal event sample score, pred neg represents the score of normal event samples, and the numerator means the total number of combinations in which the score of abnormal event samples is greater than the score of normal event samples.

[0064] C. Comparative test plan

[0065] This example is compared with other cutting-edge anomaly detection methods on the shanghaiTech dataset.

[0066] Table 3. Performance comparison of anomaly detection methods

[0067] Model ShanghaiTech Dataset Binaryclassifier 50.0 IBL 82.50 AR-net 91.24 MIST 93.13 This method 93.76

[0068] The comparison results in Table 3 with the current cutting-edge anomaly detection methods show that the anomaly detection effect of this method exceeds that of other compared methods. AR-net is the basic model of this method. After using the loss function proposed by this method, the model effect has been improved by more than 2%, which shows the effectiveness of this method.

[0069] The above description is only a preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiment. Any equivalent modifications or changes made by ordinary technicians in this field based on the contents disclosed by the present invention should be included in the protection scope recorded in the claims.

Claims

1. Video anomaly detection method based on two-stream dilated 3D convolutional network and new loss function, Features: The method specifically comprises the following steps: S1, collect video data, divide the video data into training set and test set, define each video as a video package, in the training set, define the video package containing one or more abnormal events as a positive package, and define the video package containing only normal events as a negative package; S2, divide each video packet of the training set and the test set in S1 into several video segments according to the same number of frames, and regard each video segment as a video instance; S3, input the training set video instance obtained in S2 into the I3D network, so as to extract the visual features of each video instance; S4, inputting the training set video instance obtained in S2 into the optical flow extraction network, thereby extracting the optical flow image of each video instance, and then inputting the optical flow image into the I3D network to obtain the optical flow feature of each video instance; S5, concatenates the visual features and optical flow features obtained in S3 and S4 to form the I3D features required for the final classification; S6, inputs the I3D features obtained in S5 into a three-layer fully connected neural network, and trains the neural network through a new loss function, and finally trains to obtain a network model that can distinguish abnormal events from normal events; the specific steps of constructing the new loss function used in the model training proposed in S6 are as follows: S6-1, for a given positive packet video instance p i and negative bag video instance n i , expect p i The anomaly score is greater than n i , so the following loss function is obtained: s(p i )>s(n i ) Among them, s(p i ) and s(n i ) represent the video instance p i and n i The abnormal score of S6-2, in order to make the anomaly score of the positive video instance greater than that of the negative video instance, the loss function in S6-1 is further improved to obtain a new loss function as follows: Among them, the mean method means taking the average of all instance anomaly scores in the video package; S6-3, take the first L for positive and negative packages respectively P / K and L n / The K largest instance scores are sorted and compared, and the improved loss function is as follows: Among them, k_mean means taking the first L in the positive bag and the negative bag respectively. P / K and L n / K video instance scores with the largest anomaly scores, and calculate their average; K is a hyperparameter, that is, the ratio of the number of selected instance scores to the total number of video instances; S6-4, appropriately adjust the loss function obtained in S6-3 to obtain the new loss function as follows: S6-5, for negative bags, a smoothness constraint is proposed, and the loss function is as follows: Among them, the var method means taking the variance of the scores of all video instances in the negative bag; S6-6, the final constructed loss function is as follows: loss = k_loss + λs_loss Among them, λ is a hyperparameter, that is, the weight coefficient corresponding to s_loss; S7, sequentially performing steps S3, S4 and S5 on the test set video instance obtained in S2 to obtain the I3D features of the test set video instance; S8, input the I3D features of the test set video instances obtained in S7 into the video anomaly detection model trained in S6, and then obtain the anomaly score of the test set video instances, thereby achieving effective detection of abnormal events.

2. According to claim 1, the video anomaly detection method based on dual-stream dilated 3D convolutional network and new loss function, Features: The positive video packets in S1 are represented by B p Indicates that negative video packets are represented by B n express.

3. The video anomaly detection method based on a dual-stream dilated 3D convolutional network and a new loss function according to claim 1, Features: The same number of frames in S2 is 16 frames, that is, every 16 consecutive frames are regarded as a video instance; the positive packet video instances are sequentially represented by Indicates that the negative bag video instances are used in turn Indicates that L p and L n Represents the number of video instances of positive and negative packets respectively.

4. The video anomaly detection method based on a dual-stream dilated 3D convolutional network and a new loss function according to claim 1, Features: The I3D network in S3 is composed of two parallel 3D convolution modules, which respectively extract the visual features and optical flow features of the input video instance, and finally concatenate the two features to obtain the I3D features of the video instance.

5. The video anomaly detection method based on a dual-stream dilated 3D convolutional network and a new loss function according to claim 1, Features: The optical flow extraction network in S4 adopts the flownet2 network.

Citation Information

Patent Citations

  • Abnormal behavior detection method based on improved pseudo three-dimensional residual neural network

    CN110263728A

  • Real anomaly detection method in monitoring video

    CN113312968A