Video anomaly detection method based on bidirectional skip frame prediction network

Through the bidirectional jump frame prediction network, the differential channel and context space attention module are used to solve the problem of indistinguishable boundaries between normal and abnormal features in video abnormality detection, and achieve a more efficient abnormal detection effect.

CN120339718APending Publication Date: 2025-07-18XIAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510507907.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the existing video anomaly detection methods, the characteristic boundaries between normal and abnormal events are difficult to distinguish, resulting in a degradation of model performance.

Method used

Using a method based on a bidirectional jump frame prediction network, video frame prediction is performed by constructing forward and backward autoencoders, combining differential channel attention and context space attention modules, and abnormality detection is performed by calculating the errors of predicted frames and real frames.

Benefits of technology

The fuzzy boundary and difference between normal and abnormal features is expanded, and the performance of video anomaly detection is improved, especially in abnormal events with insufficient boundaries, which significantly improves detection accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339718A_ABST
    Figure CN120339718A_ABST
Patent Text Reader

Abstract

The invention discloses a video anomaly detection method based on a bidirectional skip frame prediction network, which comprises the following steps: acquiring a video data set only containing a normal sample as a training set, and acquiring a video data set containing the normal sample and an abnormal sample as a test set; preprocessing videos in the training set and the test set; constructing a bidirectional jump prediction network; constructing a loss function; iteratively training the bidirectional jump prediction network to obtain a video frame prediction model; and preprocessing a to-be-detected video, inputting the preprocessed to-be-detected video into the video frame prediction model to obtain a prediction frame, and performing video anomaly detection by calculating an error between the prediction frame and a real frame. According to the video anomaly detection method based on the bidirectional jump frame prediction network, the problem that in an existing video anomaly detection method, normal event feature boundaries and abnormal event feature boundaries are difficult to distinguish, so that the model performance is reduced is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image or video recognition, and particularly relates to a video anomaly detection method based on a bidirectional skip-frame prediction network. Background Art

[0002] With the continuous development of computer vision, anomaly detection has been widely studied in many fields and applications such as medical images, industrial product quality inspection, and traffic monitoring. Especially with the wide popularization and application of surveillance video devices in different fields of society, it plays an important role in many aspects to ensure social security and stability. At the same time, the amount of surveillance video data has shown an explosive growth. The traditional manual screening method cannot meet the timely and accurate detection requirements for massive video information. Therefore, related intelligent technologies are gradually replacing manual detection. Video anomaly detection (VAD) has been widely studied by researchers with the development of computer vision.

[0003] Video anomaly detection can be divided into supervised learning, weakly supervised learning, and unsupervised learning. Supervised video anomaly detection needs to use fine-grained labels or image-level label data to complete the discrimination between normal and abnormal. However, since abnormal events are usually discrete and low-probability events in the entire video sequence, it is time-consuming and inefficient. In contrast, weakly supervised video anomaly detection only uses image-level labels to make a simple anomaly judgment on video information. The training data set includes normal and abnormal events with labels, and thus distinguishes normal and abnormal through multi-instance learning ranking or classification models. Unsupervised anomaly detection does not require any labels, and only needs to complete the classification preprocessing of normal and abnormal events when constructing the data set, that is, only model and feature learning are performed on normal events in the training stage, and anomaly detection is completed by measuring the reconstruction deviation of abnormal events in the test stage, and the specific time frame and spatial position of the abnormal event can be located.

[0004] All unsupervised VAD methods follow the general assumption of the anomaly detection task, that is, all events that do not occur in the training set are regarded as abnormal. Therefore, the discrimination boundary between normal and abnormal events is particularly important. However, there is a severe challenge in this assumption, that is, when the boundary is difficult to distinguish, the anomaly detection model will be greatly affected, resulting in performance degradation. Summary of the Invention

[0005] The purpose of the present invention is to provide a video anomaly detection method based on a bidirectional skip-frame prediction network, which solves the problem that the feature boundary between normal and abnormal events in the existing video anomaly detection method is difficult to distinguish, resulting in performance degradation of the model.

[0006] The technical solution adopted by the present invention is a video anomaly detection method based on a bidirectional skip-frame prediction network, which specifically includes the following steps:

[0007] Step 1: Obtain a video dataset containing only normal samples as the training set, and a video dataset containing both normal samples and abnormal samples as the test set; preprocess the videos in the training set and the test set;

[0008] Step 2: Construct a bidirectional jump prediction network;

[0009] Step 3: Construct a loss function;

[0010] Step 4: Iteratively train the bidirectional jump prediction network to obtain a video frame prediction model;

[0011] Step 5: Preprocess the video to be detected and input it into the video frame prediction model to obtain a predicted frame, and perform video anomaly detection by calculating the error between the predicted frame and the real frame.

[0012] The features of the present invention also lie in:

[0013] In Step 1, the preprocessing of the videos in the training set and the test set is as follows:

[0014] For each video in the training set, first divide it into n consecutive image frames at a frame rate of T f During the training process, use the forward frames I of a video i+1,i+3,i+5 to predict the forward predicted frames Use the backward frames I of a video i+6,i+4,i+2 to predict the backward predicted frames where i = 0, 1, …, n - 6;

[0015] For each video in the test set, first divide it into m consecutive image frames at a frame rate of T f During the test process, use the forward frames I of a video j+1:j+3 and the backward frames I j+7:j+5 to jointly predict the intermediate frames where j = 0, 1, …, m - 7.

[0016] The bidirectional jump prediction network constructed in Step 2 includes a forward autoencoder and a backward autoencoder;

[0017] The forward autoencoder includes an encoder, a decoder, and a differential channel attention module with skip connections between the encoder and the decoder; a context spatial attention module is embedded after upsampling in each decoding stage of the encoder;

[0018] The encoder is repeatedly composed of two convolutional layers and a max pooling layer; the decoder is repeatedly composed of a transposed convolution module and two convolutional layers;

[0019] The differential channel attention module includes parallel variance attention maps Mvar and the channel attention map M c ; the context spatial attention module includes a context feature extraction module con(·) and a spatial attention map M s and a 1×1 two-dimensional convolution;

[0020] The encoder successively extracts features through convolution and downsamples the input multi-frame images of 256×256, and finally obtains the depth-compressed features of 32×32; in the skip connection operation, the features of different scales from the encoder are used for differential channel attention module feature extraction and are correspondingly input into the subsequent decoder to complete feature splicing; in the decoder, the depth-compressed features of 32×32 are upsampled through successive convolution feature extraction and deconvolution to obtain the final predicted image frames of 256×256. Among them, after each deconvolution, by inputting the fused features of the spliced skip connection features into the context spatial attention module, the fusion of different features is realized;

[0021] The backward autoencoder has exactly the same structure as the forward autoencoder.

[0022] The differential channel attention module includes parallel variance attention map M var and the channel attention map M c , and the processing process for the input feature F is as follows:

[0023] Multiply the input feature F element-wise with the result of the variance attention map M var to obtain the variance feature F var . At the same time, multiply the feature F element-wise with the result of the channel attention map M c to obtain the channel feature F c . Finally, add the variance feature F var , the channel feature F c and the input feature F element-wise to obtain the variance channel feature F vc , which is expressed as follows:

[0024] F vc =F + F var + F c (1)

[0025]

[0026] In the formula, F a ∈R b×c×1 represents the modeled global feature, and b, c, and D represent the batch size, channel dimension, and spatial dimension respectively; represents element-wise multiplication, represents the mean square error;

[0027]

[0028] In the formula, δ and σ respectively represent the Sigmoid and ReLU activation functions, W1 represents a two-dimensional convolution with a convolution kernel of 1×1, and P avg represents the global average pooling operation.

[0029] The context spatial attention includes a context feature extraction module con(·) and a spatial attention map M s and a 1×1 two-dimensional convolution; for the input feature F u the processing process is as follows:

[0030] First, the context feature extraction module con(·) uses two-dimensional convolutions with four different convolution kernel sizes and dilation rates to extract features respectively and splices them into the context feature F t , which is expressed as follows:

[0031] F t = con(F u )(4)

[0032] Then, the context feature F t is multiplied element-wise with the result of the context feature F t passing through the spatial attention map M s to obtain the spatial feature map F s , which is expressed as follows:

[0033]

[0034] In the formula, W s represents a two-dimensional convolution with a convolution kernel of 3×3;

[0035] Finally, the spatial feature map F s is input into the two-dimensional convolution W1 with the same number of channels as the feature F u to finally obtain the feature F ts , which is expressed as follows:

[0036] F ts = W1(F s )(6)

[0037] In the formula, W1 represents a two-dimensional convolution with a convolution kernel of 1×1.

[0038] The context feature extraction module con(·) uses two-dimensional convolutions with convolution kernels of 1, 3, 3, and 3 respectively, and the corresponding dilation rates are 0, 3, 5, and 7.

[0039] The loss function constructed in step 3 is expressed as follows:

[0040] L = L fp + L bp + L con (7)

[0041] In the formula, L represents the total loss function, L fp represents the forward prediction loss, L bp represents the backward prediction loss, L con represents the consistency loss;

[0042] The forward prediction loss L fp and the backward prediction loss L bp , are expressed as follows:

[0043]

[0044] In the formula, represents the forward prediction frame, represents the backward prediction frame, I f represents the forward true frame, I b represents the backward true frame, represents calculating the mean square error;

[0045] The consistency loss L con , is expressed as follows:

[0046]

[0047] Step 5 is specifically as follows:

[0048] After processing the video to be detected according to the preprocessing method of the videos in the test set, input it into the video frame prediction model to obtain two identical prediction frames, and fuse the forward error and the backward error are expressed as follows:

[0049]

[0050] In the formula, ω f and ω b are the weights of e f and e b respectively, and ω f +ω b = 1;

[0051] Based on the fused error, use a new anomaly evaluation method to calculate the anomaly score, and through the multi-scale pyramid error, achieve a more comprehensive detection of anomalies of different sizes, which is expressed as follows:

[0052]

[0053] In the formula, v i is the maximum prediction error based on patches in size i obtained through mean pooling;

[0054] Normalize the PSNR to [0, 1], that is, after smoothing S(I t ) to obtain the final anomaly score, which is expressed as follows:

[0055]

[0056] According to the obtained anomaly scores, by setting different thresholds, the final anomaly detection results are obtained.

[0057] The beneficial effects of the present invention are:

[0058] The video anomaly detection method based on the bidirectional jump frame prediction network of the present invention, firstly, realizes different video frame input strategies through the jump frame mechanism, initially expanding the essential differences between different features; secondly, the proposed context spatial attention expands the scale and position features of different targets in the video frame; then, the proposed variance channel attention pays more attention to and expands the motion differences between different features in the video frame; finally, the present invention expands the fuzzy boundary and differences between normal and abnormal features from the perspectives of data preprocessing, motion features, and different targets, and uses the fusion error scores of bidirectional prediction to improve the performance of video anomaly detection. Description of the Drawings

[0059] Figure 1 is the block diagram of the bidirectional jump prediction network in the video anomaly detection method based on the bidirectional jump frame prediction network of the present invention;

[0060] Figure 2 is the schematic diagram of video frame preprocessing in the training set and the test set in the video anomaly detection method based on the bidirectional jump frame prediction network of the present invention;

[0061] Figure 3 is the structure diagram of the forward autoencoder in the video anomaly detection method based on the bidirectional jump frame prediction network of the present invention;

[0062] Figure 4 is the structure diagram of the differential channel attention module in the video anomaly detection method based on the bidirectional jump frame prediction network of the present invention;

[0063] Figure 5 is the structure diagram of the context spatial attention module in the video anomaly detection method based on the bidirectional jump frame prediction network of the present invention;

[0064] Figure 6 are the predicted qualitative results of four datasets;

[0065] Figure 7 is an example of the blurred boundary situation in the ShanghaiTech dataset. Detailed Embodiments

[0066] The present invention will be described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0067] The video anomaly detection method based on a bidirectional skip frame prediction network of the present invention can be used in fields such as security monitoring and traffic monitoring. As Figure 1 shown, a video is divided into forward skip frames I fwd and backward skip frames I bwd , and then the two skip frames are respectively input into the forward autoencoder P f ={E f , D f} and the backward autoencoder P b ={E b , D b}, and the forward prediction frame and the backward prediction frame

[0068] Specifically, it includes the following steps:

[0069] Step 1: Obtain a video dataset containing only normal samples as the training set, and a video dataset containing both normal samples and abnormal samples as the test set; preprocess the videos in the training set and the test set.

[0070] Preprocess the videos in the training set and the test set. As Figure 2 shown, the specific process is as follows:

[0071] For each video in the training set, first divide it into consecutive n image frames at a frame rate of T f . During the training process, use the forward frames I i+1,i+3,i+5 of a video to predict the forward prediction frame Use the backward frames I i+6,i+4,i+2 of a video to predict the backward prediction frame where i = 0, 1,..., n - 6;

[0072] For each video in the test set, first divide it into consecutive m image frames at a frame rate of T f . During the test process, use the forward frames I j+1:j+3 and the backward frames I j+7:j+5 of a video to jointly predict the intermediate frame where j = 0, 1,..., m - 7.

[0073] Step 2: Construct a bidirectional skip prediction network.

[0074] The constructed bidirectional skip prediction network includes a forward autoencoder P f ={E f , D f} and a backward autoencoder P b ={Eb , D b}。The structures of the forward autoencoder and the backward autoencoder are exactly the same, except for the different inputs. The forward autoencoder takes the forward frame as input, and the backward encoder takes the backward frame as input.

[0075] Taking the structure of the forward autoencoder as an example, as Figure 3 shown, the forward autoencoder includes an encoder, a decoder, and a differential channel attention module with skip connections between the encoder and the decoder; a context spatial attention module is embedded after the upsampling in each decoding stage in the encoder.

[0076] The encoder is composed of two convolutional layers and one max-pooling layer repeated, specifically including a first convolutional module, a second convolutional module, a third convolutional module, and a fourth convolutional module connected in sequence. The first convolutional module includes two convolutional layers with 32 output channels, the second convolutional module includes a max-pooling layer and two convolutional layers with 64 output channels, the third convolutional module includes a max-pooling layer and two convolutional layers with 128 output channels, the fourth convolutional module includes a max-pooling layer and two convolutional layers with 256 output channels. The convolutional layer includes a 3×3 convolution, a normalization layer, and a ReLU activation function;

[0077] The decoder is composed of a transposed convolution module and two convolutional layers repeated, specifically including a fifth convolutional module, a sixth convolutional module, a seventh convolutional module, and an eighth convolutional module connected in sequence. The fifth convolutional module includes a context spatial attention module and two convolutional layers with 256 output channels, the sixth convolutional module includes a transposed convolution layer, a context spatial attention module, and two convolutional layers with 128 output channels, the seventh convolutional module includes a transposed convolution layer, a context spatial attention module, and two convolutional layers with 64 output channels, the eighth convolutional module includes a transposed convolution layer, a context spatial attention module, and two convolutional layers with 32 output channels, as well as a 3×3 convolution and a Tanh activation function. The convolutional layer includes a 3×3 convolution, a normalization layer, and a ReLU activation function, and the transposed convolution layer includes a 3×3 transposed convolution, a normalization layer, and a ReLU activation function.

[0078] The differential channel attention module includes parallel variance attention map M var and channel attention map M c , and the specific structure is as Figure 4 shown, where the variance attention map M var includes a 1×1 convolution and a softmax activation function, which are used to perform global feature modeling on the input feature F to obtain F a , and then calculate the mean square error through Variance, and then pass through the softmax activation function; the channel attention map M cIt includes first performing a global average pooling operation on the feature F, followed by a 1×1 convolution, a ReLU activation function, a 1×1 convolution, and a Sigmoid activation function. The specific calculation process for the input feature F is as follows:

[0079] F vc = F + F var + F c (1)

[0080]

[0081] In the formula, F a ∈R b×c×1 represents the modeled global feature, b, c, and D represent the batch size, channel dimension, and spatial dimension respectively; represents element-wise multiplication, represents calculating the mean squared error;

[0082]

[0083] In the formula, δ and σ represent the Sigmoid and ReLU activation functions respectively, W1 represents a two-dimensional convolution with a 1×1 convolution kernel, and P avg represents the global average pooling operation.

[0084] Through the differential channel attention module, the extracted temporal features are enhanced in terms of attention and are connected to the decoder through skip connections to increase the gap between normal times and abnormal events.

[0085] The context spatial attention module includes a context feature extraction module con(·) and a spatial attention map M s and a 1×1 two-dimensional convolution; the specific structure is as Figure 5 shown, where the context feature extraction module con(·) includes four parallel two-dimensional convolutions with different convolution kernel sizes and dilation rates. The convolution kernels of the two-dimensional convolutions are 1, 3, 3, 3 respectively, and the corresponding dilation rates are 0, 3, 5, 7. The spatial attention map M s includes an average pooling operation, a 3×3 convolution, and a Sigmoid activation function. For the input feature F u the processing process is as follows:

[0086] First, the context feature extraction module con(·) extracts features using four two-dimensional convolutions with different convolution kernel sizes and dilation rates respectively and concatenates them into the context feature F t , which is expressed as follows:

[0087] F t = con(F u ) (4)

[0088] Then, the context feature Ft With the context feature F t Multiply element - by - element with the result of the spatial attention map M s to obtain the spatial feature map F s , which is expressed as follows:

[0089]

[0090] In the formula, W s represents a two - dimensional convolution with a 3×3 convolution kernel;

[0091] Finally, input the spatial feature map F s into a two - dimensional convolution W1 with the same number of channels as the feature F u to finally obtain the feature F ts , which is expressed as follows:

[0092] F ts = W1(F s ) (6)

[0093] In the formula, W1 represents a two - dimensional convolution with a 1×1 convolution kernel.

[0094] Introduce a context - spatial attention module and embed it into the decoder to reduce the differences in different target scales and features.

[0095] Based on the above - mentioned structure of the forward auto - encoder, its processing process for the input image frames is summarized as follows: The encoder extracts features through successive convolutions and downsamples the input multi - image frames of 256×256, and finally obtains 32×32 depth - compressed features; in the skip - connection operation, features from different scales in the encoder are used for differential channel attention module feature extraction and are correspondingly input into the subsequent decoder to complete feature splicing; in the decoder, the 32×32 depth - compressed features are upsampled through successive convolution feature extraction and de - convolution to obtain the final predicted image frames of 256×256. Among them, after each de - convolution, the fused features obtained by splicing the skip - connection features are input into the context - spatial attention module to achieve the fusion of different features.

[0096] Step 3: Construct the loss function.

[0097] The minimum objective function for network optimization consists of the forward prediction loss, the backward prediction loss, and the consistency loss. Therefore, the constructed loss function is expressed as follows:

[0098] L = L fp + L bp + L con (7)

[0099] In the formula, L represents the total loss function, L fp represents the forward prediction loss, Lbp Denote the backward prediction loss as L con Denote the consistency loss;

[0100] To make the forward prediction frame and the backward prediction frame similar to the image pixels of the ground truth frame, calculate the mean squared error between the forward and backward prediction frames and their corresponding ground truth value I f and I b to obtain the forward prediction loss L fp and the backward prediction loss L bp , which are expressed as follows:

[0101]

[0102] In the formula, denotes the forward prediction frame, denotes the backward prediction frame, I f denotes the forward ground truth frame, I b denotes the backward ground truth frame, denotes calculating the mean squared error;

[0103] Since the prediction targets in the training and testing phases are different, to ensure the structural consistency of the two final prediction frames in the testing phase, this paper uses the Structural Similarity Index measure (SSIM) to construct the consistency loss L con of the two prediction frames, which is expressed as follows:

[0104]

[0105] Step 4: Iteratively train the bidirectional skip prediction network to obtain the video frame prediction model.

[0106] In the training phase, use the forward frame I i+1,i+3,i+5 of the video in the training set to predict the forward prediction frame Use the backward frame I i+6,i+4,i+2 of a video to predict the backward prediction frame where i = 0, 1, …, n - 6; stop after reaching the number of iterations, and use the forward frame I j+1:j+3 and the backward frame I j+7:j+5 of the video in the test set to jointly predict the intermediate frame where j = 0, 1, …, m - 7, and test the model performance. If the performance index required by the model is not reached, repeat the training until the model with the optimal performance, i.e., the video frame prediction model, is obtained.

[0107] Step 5: Preprocess the video to be detected and input it into the video frame prediction model to obtain the prediction frame, and perform video anomaly detection by calculating the error between the prediction frame and the ground truth frame.

[0108] After preprocessing the video to be detected in the same way as the videos in the test set, input it into the video frame prediction model to obtain two identical prediction frames, and fuse the pre-error and the post-error as follows:

[0109]

[0110] In the formula, ω f and ω b are the weights of e f and e b respectively; because the two identical frames predicted are complementary, therefore, ω f +ω b = 1;

[0111] Anomaly detection obtains the anomaly score S(t) by calculating the prediction quality of the prediction frame and makes a decision based on the anomaly score. Most methods use the peak signal-to-noise ratio (PSNR) to measure S(t), where PSNR is calculated through the mean squared error between the real frame I t and the prediction frame .

[0112] Based on the fusion error, a new anomaly evaluation method is used to calculate the anomaly score. Through the multi-scale pyramid error, a more comprehensive detection of anomalies of different sizes is achieved, as follows:

[0113]

[0114] In the formula, v i is the maximum prediction error based on blocks in size i obtained through mean pooling; N = 3 indicates that the error pyramid consists of three different scales, and the maximum prediction error for each size is determined by the highest value among different blocks. For a single frame, the final prediction error is calculated by summing the maximum prediction errors in each size.

[0115] Normalize PSNR to [0,1], that is, smooth S(I t ) through a Gaussian filter to obtain the final anomaly score, as follows:

[0116]

[0117] According to the obtained anomaly score, by setting different thresholds, the final anomaly detection result is obtained.

[0118] Embodiment 1

[0119] This embodiment provides a video anomaly detection method based on a bidirectional skip frame prediction network, which specifically includes the following steps:

[0120] Step 1: Obtain a video dataset containing only normal samples as the training set, and a video dataset containing both normal samples and abnormal samples as the test set; preprocess the videos in the training set and the test set;

[0121] Step 2: Construct a bidirectional jump prediction network;

[0122] Step 3: Construct a loss function;

[0123] Step 4: Iteratively train the bidirectional jump prediction network to obtain a video frame prediction model;

[0124] Step 5: Preprocess the video to be detected and input it into the video frame prediction model to obtain a predicted frame, and perform video anomaly detection by calculating the error between the predicted frame and the real frame.

[0125] Embodiment 2

[0126] Based on Embodiment 1, preprocess the videos in the training set and the test set, and the specific process is as follows:

[0127] For each video in the training set, first segment it into n consecutive image frames at a frame rate of T f During the training process, use the forward frames I of a video i+1,i+3,i+5 to predict the forward predicted frames Use the backward frames I of a video i+6,i+4,i+2 to predict the backward predicted frames where i = 0, 1,..., n - 6;

[0128] For each video in the test set, first segment it into m consecutive image frames at a frame rate of T f During the test process, use the forward frames I of a video j+1:j+3 and the backward frames I j+7:j+5 to jointly predict the intermediate frames where j = 0, 1,..., m - 7.

[0129] Embodiment 3

[0130] Based on Embodiment 2, the bidirectional jump prediction network constructed in Step 2 includes a forward autoencoder and a backward autoencoder;

[0131] The forward autoencoder includes an encoder, a decoder, and a differential channel attention module with skip connections between the encoder and the decoder; a context spatial attention module is embedded after upsampling in each decoding stage of the encoder;

[0132] The encoder is repeatedly composed of two convolutional layers and one max-pooling layer; the decoder is repeatedly composed of a transposed convolution module and two convolutional layers;

[0133] The differential channel attention module includes parallel variance attention map M var and channel attention map M c ; the context spatial attention module includes a context feature extraction module con(·), a spatial attention map M s and a 1×1 two-dimensional convolution;

[0134] The encoder successively extracts convolutional features and downsamples multiple input image frames of 256×256, and finally obtains 32×32 depth-compressed features; in the skip connection operation, differential channel attention module feature extraction is performed on features from different scales in the encoder, and they are correspondingly input into the subsequent decoder to complete feature splicing; in the decoder, the 32×32 depth-compressed features are successively extracted through convolutional features and upsampled through deconvolution to obtain the final predicted image frames of 256×256. Among them, after each deconvolution, by inputting the fused features of the spliced skip connection features into the context spatial attention module, the fusion of different features is realized;

[0135] The backward autoencoder has exactly the same structure as the forward autoencoder.

[0136] Example 4

[0137] Based on Example 3, the differential channel attention module includes parallel variance attention map M var and channel attention map M c , and the processing process for the input feature F is as follows:

[0138] Multiply the input feature F element-wise with the result of the variance attention map M var to obtain the variance feature F var . At the same time, multiply the feature F element-wise with the result of the channel attention map M c to obtain the channel feature F c . Finally, add the variance feature F var , the channel feature F c and the input feature F element-wise to obtain the variance-channel feature F vc , which is expressed as follows:

[0139] F vc =F + F var + F c (1)

[0140]

[0141] In the formula, F a ∈R b×c×1 represents the modeled global feature, and b, c, and D respectively represent the batch size, channel dimension, and spatial dimension; represents element-wise multiplication, Indicates the calculation of the mean square error;

[0142]

[0143] In the formula, δ and σ respectively represent the Sigmoid and ReLU activation functions, W1 represents a two-dimensional convolution with a convolution kernel of 1×1, and P avg represents the global average pooling operation.

[0144] Context spatial attention, including a context feature extraction module con(·) and a spatial attention map M s and a 1×1 two-dimensional convolution; for the input feature F u The processing process is as follows:

[0145] First, the context feature extraction module con(·) uses two-dimensional convolutions with four different convolution kernel sizes and dilation rates to extract features respectively and splices them into the context feature F t , which is expressed as follows:

[0146] F t = con(F u ) (4)

[0147] The context feature extraction module con(·) uses two-dimensional convolutions with convolution kernels of 1, 3, 3, and 3 respectively, and the corresponding dilation rates are 0, 3, 5, and 7;

[0148] Then, the context feature F t is multiplied element-wise with the result of the context feature F t passing through the spatial attention map M s to obtain the spatial feature map F s , which is expressed as follows:

[0149]

[0150] In the formula, W s represents a two-dimensional convolution with a convolution kernel of 3×3;

[0151] Finally, the spatial feature map F s is input into the two-dimensional convolution W1 with the same number of channels as the feature F u to finally obtain the feature F ts , which is expressed as follows:

[0152] F ts = W1(F s ) (6)

[0153] In the formula, W1 represents a two-dimensional convolution with a convolution kernel of 1×1.

[0154] Example 5

[0155] Based on Example 4, the loss function constructed in step 3 is expressed as follows:

[0156] L = L fp + L bp + L con (7)

[0157] In the formula, L represents the total loss function, L fp represents the forward prediction loss, L bp represents the backward prediction loss, L con represents the consistency loss;

[0158] The forward prediction loss L fp and the backward prediction loss L bp are expressed as follows:

[0159]

[0160] In the formula, represents the forward prediction frame, represents the backward prediction frame, I f represents the forward true frame, I b represents the backward true frame, represents calculating the mean square error;

[0161] The consistency loss L con is expressed as follows:

[0162]

[0163] Example 6

[0164] Based on Example 5, step 5 is specifically as follows:

[0165] After processing the video to be detected in the same preprocessing manner as the videos in the test set, input it into the video frame prediction model to obtain two identical prediction frames, and fuse the forward error and the backward error as follows:

[0166]

[0167] In the formula, ω f and ω b are the weights of e f and e b respectively, and ω f + ω b = 1;

[0168] Based on the fused error, use a new anomaly evaluation method to calculate the anomaly score, and through the multi-scale pyramid error, achieve a more comprehensive detection of anomalies of different sizes, as follows:

[0169]

[0170] where v i is the maximum prediction error based on patch blocks in size i obtained through mean pooling;

[0171] Normalize PSNR to [0,1], that is, after smoothing S(I t ) with a Gaussian filter to obtain the final anomaly score, which is expressed as follows:

[0172]

[0173] According to the obtained anomaly scores, by setting different thresholds, the final anomaly detection results are obtained.

[0174] Simulation experiments

[0175] All simulation experiments were conducted on a single RTX 4090 GPU. Based on the PyTorch deep learning framework, the experimental results were evaluated using the Area Under the Curve (AUC) metric. Experiments were carried out on four unsupervised video anomaly detection datasets: UCSD Ped1 & Ped2, CUHK Avenue, and ShanghaiTech. All datasets contain a training set and a test set. The training set only contains normal events, and the test set contains both normal and abnormal events.

[0176] This simulation experiment compared the detection method of the present invention with several current advanced video anomaly detection methods, and the results are shown in Table 1. The experimental results show that the present invention can significantly improve the detection accuracy compared with other methods. Especially in Ped2 and ShanghaiTech, the method of the present invention significantly expands the boundary between normal and abnormal, thus greatly improving the detection performance.

[0177] Table 1 Comparison results of AUC with other methods

[0178] Ped1 Ped2 Avenue Shanghaitech MemAE - 94.1 83.3 71.2 MNAD - 97.0 88.5 70.5 DAST-Net 85.4 97.9 89.8 73.7 SSAGAN 84.2 96.9 88.8 74.3 MGAN-CL - 96.5 87.1 73.6 The method of the present invention 86.3 98.6 89.5 76.4

[0179] Figure 6 Shows the visualization results of the present invention in different datasets and marks the exact positions of anomalies. Obviously, due to bidirectional prediction, abnormal events show an overlapping and blurred state, and this effect is particularly obvious in abnormal events with obvious boundaries. The brightness of the prediction error map is significantly higher, while the error of normal events is negligible.

[0180] Figure 7Further show examples of fuzzy boundaries. Taking the Shanghaitech dataset as an example, it includes four abnormal events, namely jumping, throwing a bag, running, and wandering. It can be seen that these fuzzy abnormal events are only distinguished by human behavior. In the error graph, the present invention can effectively distinguish Figure 7 the fuzzy abnormalities therein, which proves that the present invention realizes the required discrimination ability of the model by expanding the internal difference between normal and abnormal.

Claims

1. A video anomaly detection method based on a bidirectional jump frame prediction network, characterized in that Specifically, it includes the following steps: Step 1: Obtain a video dataset containing only normal samples as the training set, and a video dataset containing normal samples and abnormal samples as the test set; preprocess the videos in the training set and the test set; Step 2: Construct a bidirectional jump prediction network; Step 3: Construct a loss function; Step 4: Iteratively train the bidirectional jump prediction network to obtain a video frame prediction model; Step 5: Preprocess the video to be detected and input it into the video frame prediction model to obtain a predicted frame, and perform video anomaly detection by calculating the error between the predicted frame and the real frame.

2. The video anomaly detection method based on a bidirectional skip frame prediction network according to claim 1, wherein In Step 1, the preprocessing of the videos in the training set and the test set is as follows: For each video in the training set, first at a frame rate of T f it is segmented into n consecutive image frames. During the training process, the forward frames I of a video are used respectively i+1,i+3,i+5 to predict the forward prediction frames The backward frames I of a video are used i+6,i+4,i+2 to predict the backward prediction frames where i = 0, 1, …, n - 6; For each video in the test set, first, at a frame rate of T f it is divided into m consecutive image frames. During the testing process, the forward frames I j+1:j+3 and the backward frames I j+7:j+5 of a video are jointly used to predict the intermediate frames where j = 0, 1, …, m - 7.

3. The video anomaly detection method based on a bidirectional skip-frame prediction network according to claim 1, wherein The bidirectional jump prediction network constructed in Step 2 includes a forward autoencoder and a backward autoencoder; The forward autoencoder includes an encoder, a decoder, and a differential channel attention module with a skip connection between the encoder and the decoder; A context spatial attention module is embedded after the upsampling in each decoding stage in the encoder; The encoder is repeatedly composed of two convolutional layers and a max pooling layer; the decoder is repeatedly composed of a deconvolution module and two convolutional layers; The differential channel attention module includes a parallel variance attention map M var and a channel attention map M c ; The context spatial attention module includes a context feature extraction module con(·) and a spatial attention map M s and a 1×1 two-dimensional convolution; The encoder extracts features from multiple input image frames of 256×256 through successive convolutional feature extractions and downsamplings, and finally obtains 32×32 deep compressed features; In the skip connection operation, features of different scales from the encoder are used for feature extraction by the differential channel attention module and are correspondingly input into the subsequent decoder for feature splicing; in the decoder, the 32×32 deep compressed features are upsampled through successive convolutional feature extractions and deconvolutions to obtain the final predicted image frames of 256×256. Among them, after each deconvolution, the fused features obtained by splicing the skip connection features are input into the context spatial attention module to achieve the fusion of different features; The structure of the backward autoencoder is exactly the same as that of the forward autoencoder.

4. The video anomaly detection method based on a bidirectional skip frame prediction network according to claim 3, wherein The differential channel attention module includes a parallel variance attention map M var and a channel attention map M c , and the processing process for the input feature F is as follows: Multiply the input feature F element-wise with the variance attention map M var to obtain the variance feature F var . Meanwhile, multiply the feature F element-wise with the channel attention map M c to obtain the channel feature F c . Finally, add the variance feature F var , the channel feature F c and the input feature F element-wise to obtain the variance-channel feature F vc , which is expressed as follows: F vc = F + F var + F c (1) where, F a ∈R b×c×1 represents the global feature of the model, and b, c, and D represent the batch size, the channel dimension, and the spatial dimension, respectively; represents element-wise multiplication, represents the mean squared error; where δ and σ represent the Sigmoid and ReLU activation functions respectively, W1 represents a two-dimensional convolution with a convolution kernel of 1×1, and P avg represents the global average pooling operation.

5. The video anomaly detection method based on a bidirectional skip-frame prediction network according to claim 3, wherein The context spatial attention includes a context feature extraction module con(·) and a spatial attention map M s and a 1×1 two-dimensional convolution; for the input feature F u the processing process is as follows: First, the context feature extraction module con(·) extracts features using 2D convolutions with four different convolutional kernel sizes and dilation rates respectively and concatenates them into the context feature F t , which is expressed as follows: F t = con(F u ) (4) Then, the context feature F t and the context feature F t are multiplied element-wise by the spatial attention map M s to obtain the spatial feature map F s , which is expressed as follows: In the formula, W s represents a two-dimensional convolution with a 3×3 convolution kernel; Finally, the spatial feature map F s is then input into the two-dimensional convolution W1 with the same number of channels as the feature F u to finally obtain the feature F ts , expressed as follows: F ts = W1(F s ) (6) In the formula, W1 represents a two-dimensional convolution with a convolution kernel of 1×1.

6. The video anomaly detection method based on a bidirectional skip frame prediction network according to claim 5, wherein For the context feature extraction module con(·), the convolution kernels of the four two-dimensional convolutions are 1, 3, 3, and 3 respectively, and the corresponding dilation rates are 0, 3, 5, and 7.

7. The video anomaly detection method based on a bidirectional skip-frame prediction network according to claim 1, wherein The loss function constructed in Step 3 is expressed as follows: L = L fp + L bp + L con In equation (7), L represents the total loss function, L fp represents the forward prediction loss, L bp represents the backward prediction loss, L con represents the consistency loss; Forward prediction loss L fp and backward prediction loss L bp , are expressed as follows: In the formula, represents a forward prediction frame, represents a backward prediction frame, I f represents a forward real frame, I b represents a backward real frame, represents calculating the mean square error; Consistency loss L con , is expressed as follows:

8. The video anomaly detection method based on a bidirectional skip frame prediction network according to claim 2, wherein Step 5 is specifically: After processing the video to be detected in the same way as the preprocessing of the videos in the test set, it is input into the video frame prediction model to obtain two identical predicted frames, and the pre-error and the post-error are expressed as follows: where ω f and ω b are the weights of e f and e b respectively, and ω f + ω b = 1; Based on the fusion error, a new anomaly evaluation method is used to calculate the anomaly score, and through the multi-scale pyramid error, a more comprehensive detection of anomalies of different sizes is realized, which is expressed as follows: where, v i is the maximum prediction error based on patch blocks in dimension i obtained through mean pooling; Normalize the PSNR to [0, 1], that is, after smoothing S(I t ) to obtain the final anomaly score, which is expressed as follows: According to the obtained anomaly score, different thresholds are set to obtain the final anomaly detection result.