A video anomaly detection method, system, medium and device

By integrating time shifting and dynamic large kernel convolution, the method enhances sensitivity to time and spatial information, improving the capture of local features and overall accuracy in video anomaly detection.

CN119516437BActive Publication Date: 2025-07-15SHANDONG JIANZHU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411576841.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-06
Publication Date
2025-07-15
Estimated Expiration
2044-11-06

AI Technical Summary

Technical Problem

The existing video anomaly detection methods are insensitive to time and spatial information, and the local feature capture rate is low during feature extraction, resulting in low detection accuracy.

Method used

Combining time shift and dynamic large kernel convolution, the distribution of input features over different time periods is changed through time shift, and spatial features are accumulated through recursive aggregation method to capture multi-scale context information, and enhance the model's understanding of local information.

Benefits of technology

It improves sensitivity to temporal and spatial information, enhances the capture rate of local features, achieves a larger receptive field, and improves the accuracy of video anomaly detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119516437B_ABST
    Figure CN119516437B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of computer vision technology. The present invention discloses a video anomaly detection method, system, medium and device, including: obtaining a continuous video frame sequence; based on the continuous video frame sequence, detecting an anomaly event through a video anomaly detection model; wherein, the video anomaly detection model performs a time shift operation between two consecutive feature maps, replaces a part of the channels of each feature map with the same channel part from the previous feature map to extract temporal features; and after performing the time shift operation, accumulates the spatial features through a recursive aggregation method to capture multi-scale context information. The sensitivity to temporal information and spatial information is improved, and the video anomaly detection accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology. Specifically, it relates to a video anomaly detection method, system, medium and device. Background Art

[0002] The statements in this part only provide background technical information related to the present invention and do not necessarily constitute prior art.

[0003] Video anomaly detection has great application potential in fields such as social security, medical monitoring, and disaster warning.

[0004] However, the current video anomaly detection methods have the following problems: insensitivity to time information and spatial information, relatively single information scale, and low capture rate of local features during feature extraction, resulting in low accuracy of video anomaly detection. Summary of the Invention

[0005] To solve the above problems, the present invention provides a video anomaly detection method, system, medium and device, which combines time shift and dynamic large kernel convolution to improve the sensitivity to time information and spatial information. Time shift is used to change the distribution of input features in different time periods, thereby improving the sensitivity to time information; the dynamic large kernel convolution operation is introduced, and the spatial features are accumulated through a recursive aggregation method to effectively capture multi-scale context information, enhance the model's understanding ability of local information, improve the capture rate of local features, and achieve a larger receptive field.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] The first aspect of the present invention provides a video anomaly detection method, which includes:

[0008] Obtain a continuous video frame sequence;

[0009] Based on the continuous video frame sequence, detect an abnormal event through a video anomaly detection model;

[0010] Wherein, the video anomaly detection model performs a time shift operation between two consecutive feature maps, replaces a part of the channels of each feature map with the same channel part from the previous feature map to extract time features; and after performing the time shift operation, accumulates the spatial features through a recursive aggregation method to capture multi-scale context information.

[0011] Further, the video anomaly detection model includes an encoder, and the operations of the encoder include: after integrating a continuous video frame sequence into a tensor, through linear embedding, and using a sliding window for feature extraction, a number of feature maps are obtained. The feature maps sequentially pass through two layers of ST-3D blocks, a time shift operation, two layers of dynamic large kernel convolution operations, downsampling, six layers of ST-3D blocks, one layer of dynamic large kernel convolution layer, two layers of ST-3D blocks, and one layer of dynamic large kernel convolution operation to obtain the encoder output.

[0012] Further, a first skip connection is established in the second dynamic large kernel convolution operation of the two layers of dynamic large kernel convolution operations.

[0013] Further, a second skip connection is established in the one layer of dynamic large kernel convolution layer between the six layers of ST-3D blocks and the two layers of ST-3D blocks.

[0014] Further, the video anomaly detection model includes a decoder, and the operations of the decoder include: after cascading the encoder output and the second skip connection, sequentially passing through an upsampling operation, six layers of ST-3D blocks, one layer of dynamic large kernel convolution operation, cascading with the first skip connection, upsampling, and two layers of ST-3D blocks to obtain the decoder output.

[0015] The second aspect of the present invention provides a video anomaly detection system, which includes:

[0016] A data acquisition module, which is configured to: acquire a continuous video frame sequence;

[0017] An anomaly detection module, which is configured to: based on the continuous video frame sequence, detect an anomaly event through the video anomaly detection model;

[0018] Wherein, the video anomaly detection model performs a time shift operation between two consecutive feature maps, replaces a part of the channels of each feature map with the same channel part from the previous feature map to extract time features; and after performing the time shift operation, accumulates the spatial features through a recursive aggregation method to capture multi-scale context information.

[0019] Further, the video anomaly detection model includes an encoder, and the operations of the encoder include: after integrating a continuous video frame sequence into a tensor, through linear embedding, and using a sliding window for feature extraction, a number of feature maps are obtained. The feature maps sequentially pass through two layers of ST-3D blocks, a time shift operation, two layers of dynamic large kernel convolution operations, downsampling, six layers of ST-3D blocks, one layer of dynamic large kernel convolution layer, two layers of ST-3D blocks, and one layer of dynamic large kernel convolution operation to obtain the encoder output.

[0020] Further, the video anomaly detection model includes a decoder, and the operations of the decoder include: after cascading the output of the encoder and the second skip connection, successively performing an upsampling operation, six layers of ST-3D blocks, one layer of dynamic large kernel convolution operation, cascading with the first skip connection, upsampling, and two layers of ST-3D blocks to obtain the decoder output.

[0021] The third aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps in a video anomaly detection method as described above are implemented.

[0022] The fourth aspect of the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and running on the processor. When the processor executes the program, the steps in a video anomaly detection method as described above are implemented.

[0023] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0024] The present invention combines the time shift technology with the dynamic large kernel convolution, improves the sensitivity to time information and space information, and improves the video anomaly detection accuracy.

[0025] The present invention uses time shift to change the distribution of input features in different time periods, thereby improving the sensitivity to time information.

[0026] The present invention introduces the dynamic large kernel convolution, accumulates the spatial features through a recursive aggregation method, effectively captures the multi-scale context information, enhances the model's understanding ability of local information, improves the capture rate of local features, and realizes a larger receptive field. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The accompanying drawings constituting a part of this specification are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute a limitation to the present invention.

[0028] Figure 1 It is a flowchart of a video anomaly detection method according to Embodiment 1 of the present invention;

[0029] Figure 2 It is a schematic diagram of the feature map undergoing time shift and dynamic large kernel convolution processing before the second downsampling in the encoder according to Embodiment 1 of the present invention;

[0030] Figure 3 It is a diagram of the experimental results according to Embodiment 1 of the present invention;

[0031] Figure 4 It is a diagram of the test results according to Embodiment 1 of the present invention. Detailed implementation manners

[0032] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0033] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0034] Without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other. The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0035] Embodiment 1

[0036] The purpose of this Embodiment 1 is to provide a method for video anomaly detection.

[0037] Aiming at the insensitivity to time information and space information in the previous video anomaly detection methods, as well as the problem of extracting local features during feature extraction, a method for video anomaly detection provided in this embodiment uses the combination of time shift and DLK (Dynamic Large Kernel Convolution Layer), which improves the sensitivity to time information and space information, and at the same time improves the capture rate of local features and realizes a larger receptive field.

[0038] A method for video anomaly detection provided in this embodiment includes the following steps:

[0039] Step 1: Select a data set and evaluation metrics.

[0040] Three data sets are used: Avenue (a video data set shot on the campus avenue of the Chinese University of Hong Kong), Ped2 (a video data set of a sidewalk shot by a fixed camera), and Shanghai Tech (a data set on the campus of the Shanghai University of Science and Technology).

[0041] Avenue (a video data set shot on the campus avenue of the Chinese University of Hong Kong): The Chinese University of Hong Kong released 37 videos in 2013. These videos are further classified into 16 normal videos for training purposes and 21 abnormal videos for testing purposes. This database contains RGB (Red, Green, Blue) single-scene data, covering a total of 47 abnormal events. Among them, the abnormal events can be divided into three categories: abnormal behaviors, wrong directions, and abnormal objects.

[0042] Ped2 (Sidewalk Video Dataset Captured by Fixed Cameras): The UCSD Anomaly Detection Dataset collected by the University of California, San Diego (UCSD) contains surveillance video clips captured by fixed cameras located at advantageous high positions overlooking sidewalks. The pedestrian densities on these sidewalks vary from sparse to dense. In standard scenarios, the video sequences are only characterized by pedestrians. The Peds2 subset includes 16 training video samples and 12 test video samples.

[0043] ShanghaiTech University Dataset: Released in 2016, it is divided into two parts, Part A (Part_A_final) and Part B (Part_B_final). Part A includes 300 training images and 182 test images, while Part B includes 400 training images and 316 test images.

[0044] In this embodiment, the performance evaluation metric of the model is PSNR (Peak Signal-to-Noise Ratio). PSNR is calculated by comparing the differences between the original image and the reconstructed image. The higher the PSNR value, generally the smaller the difference between the reconstructed image and the original image, indicating higher reconstruction quality.

[0045] For deep learning-based video anomaly detection techniques, various neural networks are used to identify abnormal behavior patterns or object movements in videos. With the success of the Transformer architecture in natural language processing, the Transformer (VIT) was first applied to the image classification task of 16×16 images in the field of computer vision, demonstrating powerful performance. Subsequently, various networks based on the Transformer architecture emerged. To capture more complex details, the Transformer in Transformer (a structure that nests another layer of Transformer inside the Transformer) model further subdivides each image patch (visual sentence) into smaller patches, enabling the model to simultaneously grasp local details and global context. Similarly, the Swin Transformer (a Transformer model based on sliding windows) based on the Transformer framework adopts a hierarchical structure, dividing the image into smaller patches, enabling it to capture local feature information while significantly reducing computational complexity. In addition, a sliding window mechanism has been introduced to increase the receptive field of the model.

[0046] Step 2: Data preprocessing.

[0047] For the three datasets, all images were uniformly cropped into the format of 224×224×3 as the input. Then, the images were normalized by scaling the pixel values from [0, 255] to [0, 1], and the dimensions were rearranged from [height, width, number of channels] to [number of channels, height, width] for subsequent processing.

[0048] Step 3: Hyperparameter setting.

[0049] The input video frames were divided into blocks of size (4, 4, 4) and fed into the 3D Swin Transformer (ST-3D) blocks under a three-dimensional scale that includes temporal information, where the local attention window size is (8, 7, 7). The number of layers in the three-dimensional Swin Transformer is [3, 6]. The ratio of the hidden dimension size of the MLP (multi-layer perceptron) layer to the input dimension size is 4. The number of attention heads in each 3D Swin Transformer block is [6, 12]. During training, the optimizer used is the AdamX optimization algorithm, and the initial learning rate for the three datasets is set to 0.001.

[0050] Step 4: Model improvement.

[0051] In this embodiment, improvements were made in temporal feature extraction and multi-scale context, and ablation experiments were conducted to evaluate its performance in the experimental results.

[0052] In this embodiment, as Figure 1 shown, the backbone network of the video anomaly detection model is based on the 3D-SwinTransformer model, and the reconstruction task of the input image is completed through two downsampling and two upsampling operations. The feature extractor in Swin Transformer is combined with a 3D encoder to learn the normal behavior patterns from the anomaly-free video frames, and a 3D decoder is used to predict future frames.

[0053] The video anomaly detection model in this embodiment adopts an unsupervised network structure decoder based on Swin Transformer (a Transformer model based on a sliding window) and 3D (three-dimensional), and uses skip connections to connect the feature maps from the encoder and the decoder. The ability of the network to capture multi-scale context information is also enhanced by using large and deep convolutional kernels of different sizes and dilation rates to process the features; using temporal shift to change the distribution of the input features at different time periods, thereby improving the sensitivity to temporal information.

[0054] A video anomaly detection method provided in this embodiment integrates time shift processing to process channel information across consecutive time dimensions, replaces some consecutive feature mapping channels, and uses convolution sizes of different scales to adaptively merge multi-scale local feature maps.

[0055] Test the shifted window masking of Swin Transformer in the ST-3D model, conduct experiments to remove the masking, and attempt to retain more feature information.

[0056] A video anomaly detection method provided in this embodiment mainly involves a 3D encoder and decoder based on Swin Transformer, channel shift and dynamic large kernel convolution (DLK) along the channel, as Figure 1 shown.

[0057] (1) Encoder.

[0058] (101) First downsampling.

[0059] The encoder inputs a sequence of consecutive video image frames T = [t1, t2... t n where t is the image frame sorted by time, and the size of each video frame is [224, 224, 3]. These consecutive time frames are integrated into a tensor of shape [T, 224, 224, 3]. After linear embedding, the high-dimensional data is mapped to a low-dimensional space to obtain T / 2 feature maps, and the periodic 3D sliding window method is used for feature extraction. The input is converted from [T, 224, 224, 3] to [T / 2, 96, 56, 56].

[0060] (102) After passing through two ST-3D blocks, channel shift (i.e., time domain translation technology) is applied between two consecutive feature maps, and the first one-eighth channels of each feature map are replaced with the same channel parts from the previous feature map. Among them, the time domain translation technology has extensive applications in video analysis and understanding. To better capture the time dimension of video data, in this embodiment, some channels in the time dimension are transferred to the next frame to capture the dynamic changes over time, which helps to improve feature learning in time series data. For each frame, the first 1 / 8 channel features of the video frame are extracted from each residual block, and the feature maps are cached in the memory. For the next frame, this embodiment replaces the first 1 / 8 of the current feature map with the cached feature map. This embodiment uses a combination of 7 / 8 of the current feature map and 1 / 8 of the old feature map to generate the feature map for the input of the next layer and repeats this process.

[0061] (103) Perform two - layer convolutional DLK operations on the feature map. The first layer uses a 5×5×5 kernel; while the second layer uses a 7×7×7 kernel with dilation rate to increase the receptive field through spatial dilation and group convolution, so as to enhance the model's attention mechanism and establish the first skip connection.

[0062] (104) After the feature map undergoes two - layer dynamic large - kernel convolution operations, through Patch Merging, perform the second downsampling, and the input is transformed from [T / 2, 96, 56, 56] to [T / 2, 192, 28, 28].

[0063] (105) After six - layer ST - 3D blocks and another dynamic large - kernel convolution layer with the established second skip connection, the encoder operation is completed after two - layer ST - 3D blocks and using dynamic large - kernel convolution.

[0064] This part mainly involves the 3D encoder and the decoder based on Swin Transformer, channel shifting, and dynamic large - kernel convolution modules.

[0065] (2) Decoder.

[0066] Similar to the encoder, the decoder also uses ST - 3D modules to process 3D image patches. However, it is different in that it gradually upsamples the feature map and processes them through convolution and self - attention, reduces the number of channels of the feature map to obtain a higher - resolution feature map, and performs two upsampling operations, thus halving the dimension.

[0067] For the data processing process of the encoder: starting from the concatenation using the second skip connection, then an upsampling operation using a convolutional kernel of size (1, 2, 2) and a stride that doubles the length and width of the feature map, converting it from [T / 2, 192, 28, 28] to [T / 2, 96, 56, 56]; then, through six - layer ST - 3D blocks and DLK to enhance the understanding of spatio - temporal dynamics in the feature map; then, the output of the dynamic large - kernel convolution is concatenated with the first skip connection of the encoder output from the same dimension; subsequently, upsampling and two - layer ST blocks are used to change the tensor dimension from [T / 2, 96, 56, 56] to [3, 224, 224], generating the final image output.

[0068] To better process context information, this embodiment uses a dynamic large - kernel convolution (DLK) module to enhance the existing Swin - Transformer module, which uses multiple large - depth convolutional kernels to extract multi - scale features. After performing the time - shift operation, applying DLK to the feature map allows capturing finer information and achieving a larger receptive field, which helps understand video behavior. As Figure 2As shown, in this embodiment, two convolution kernels with sizes of 5×5×5 and 7×7×7 respectively are used to enable interaction between these features among different spatial descriptors. Before the second downsampling in the encoder, the feature map undergoes time shift and DLK processing. First, the feature map is shifted along the channel dimension to extract temporal features, which is effective for analyzing behavior patterns over long time spans. Then, it enters the DLK model, and the spatial features are accumulated through a recursive aggregation method to effectively capture multi-scale context information and enhance the model's ability to understand local information.

[0069] Step 5: Use the training set to train the improved video anomaly detection model in Step 4.

[0070] Step 6: The trained video anomaly detection model was verified on three public video datasets, namely ShanghaiTech, Avenue, and Ped2, and achieved high accuracy rates.

[0071] Step 7: Perform video anomaly detection using the trained video anomaly detection model. This includes: obtaining a sequence of consecutive video frames; based on the sequence of consecutive video frames, detecting abnormal events through the video anomaly detection model. Among them, the video anomaly detection model performs a time shift operation between two consecutive feature maps, replaces a part of the channels of each feature map with the same channel part from the previous feature map to extract temporal features; and after performing the time shift operation, accumulates the spatial features through a recursive aggregation method to capture multi-scale context information.

[0072] Among them, the video anomaly detection model includes an encoder, and the operations of the encoder include: after integrating the sequence of consecutive video frames into a tensor, mapping multiple consecutive video frames to the embedding dimension through sliding window processing to obtain several feature maps. The feature maps pass through two layers of ST-3D blocks, a time shift operation, two layers of dynamic large kernel convolution operations, downsampling, six layers of ST-3D blocks, one layer of dynamic large kernel convolution layer, two layers of ST-3D blocks, and one layer of dynamic large kernel convolution operation in sequence to obtain the encoder output. Among them, the second convolution operation in the two layers of dynamic large kernel convolution operations establishes a first skip connection. Among them, one layer of dynamic large kernel convolution layer between the six layers of ST-3D blocks and the two layers of ST-3D blocks establishes a second skip connection.

[0073] Among them, the video anomaly detection model includes a decoder, and the operations of the decoder include: after cascading the encoder output with the second skip connection, passing through an upsampling operation, six layers of ST-3D blocks, one layer of dynamic large kernel convolution operation, cascading with the first skip connection, upsampling, and two layers of ST-3D blocks in sequence to obtain the decoder output.

[0074] Table 1: Results of ablation experiments

[0075] Time displacement Dynamic large kernel convolution Avenue dataset Ped2 dataset ShanghaiTech dataset × × 83.1% 95.3% 71.6% × √ 85.5% 96.1% 72.7% √ × 84.6% 96.8% 72.3% √ √ 86.5% 97.1% 73.7%

[0076] As shown in the second row of Table 1, if time offset is not used to extract time dimension information from the video and the dynamic large kernel convolution module is not used to understand context information, the performance of the three datasets is the worst. In the third and fourth rows, the dynamic large kernel convolution and time offset are combined separately, thus improving the model performance. In addition, the role of the mask in Swin Transformer when using a sliding window was tested, aiming to remove the mask to retain more image information. However, the experimental results show that the mask did not achieve the expected effect. The anomaly detection accuracies on the three datasets of Avenue, Ped2, and Shanghai Tech are 86.5%, 97.7%, and 73.7% respectively. It is proved that the model of this embodiment can improve the network reconstruction accuracy and improve the model performance.

[0077] Since the ShanghaiTech dataset has a large number of video frame data and contains various complex scene information, the ShanghaiTech dataset is more challenging for video anomaly detection.

[0078] As Figure 3 shown, the reconstruction error heatmaps of the Avenue dataset, ped2 dataset, and Shanghai Tech dataset are presented. The method of this embodiment is used to reconstruct the abnormal frames and display the reconstruction error heatmaps. The first row consists of the original images, the second row contains the reconstructed images, and the third row shows the reconstruction error heatmaps between the two. Among them, the behaviors of running, riding a motorcycle, and riding a bicycle are all considered abnormal behavior events.

[0079] As Figure 4 shown, an anomaly detection is performed on the first video dataset of Shanghai Tech using a pre-trained model. The first test video has a total of 265 frames. In the normal video frames, the road is relatively empty and there are few pedestrians. When pedestrians appear, the peak signal-to-noise ratio starts to decrease. In Figure 4 , the peak signal-to-noise ratio has been normalized. In fact, during the test process, a higher peak signal-to-noise ratio indicates normal, while a decrease indicates abnormal. Around the 100th frame, pedestrians start to appear and the value increases. Starting from the 150th frame, a pedestrian riding a bicycle appears, and around the 230th frame, the person riding the bicycle disappears and the value returns to normal.

[0080] As Figure 4 shown, after testing using the network of this embodiment, for a continuous video sequence with anomalies, when an abnormal frame appears, the anomaly score will increase.

[0081] Embodiment 2

[0082] The purpose of this Embodiment 2 is to provide a video anomaly detection system.

[0083] A data acquisition module, which is configured to: acquire a sequence of consecutive video frames;

[0084] An anomaly detection module, which is configured to: based on the sequence of consecutive video frames, detect an anomaly event through a video anomaly detection model;

[0085] Wherein, the video anomaly detection model performs a time shift operation between two consecutive feature maps, replaces a partial channel of each feature map with the same channel part from the previous feature map to extract time features; and after performing the time shift operation, accumulates the spatial features through a recursive aggregation method to capture multi-scale context information.

[0086] Wherein, the video anomaly detection model includes an encoder, and the operations of the encoder include: after integrating the sequence of consecutive video frames into a tensor, through linear embedding, and after feature extraction using a sliding window, obtaining a number of feature maps, and the feature maps pass through two layers of ST-3D blocks, a time shift operation, two layers of dynamic large kernel convolution operations, downsampling, six layers of ST-3D blocks, one layer of dynamic large kernel convolution layer, two layers of ST-3D blocks, and one layer of dynamic large kernel convolution operation in sequence to obtain the encoder output.

[0087] Wherein, the video anomaly detection model includes a decoder, and the operations of the decoder include: after cascading the encoder output with a second skip connection, successively passing through an upsampling operation, six layers of ST-3D blocks, one layer of dynamic large kernel convolution operation, cascading with a first skip connection, upsampling, and two layers of ST-3D blocks to obtain the decoder output.

[0088] It should be noted here that each module in this embodiment corresponds one by one to each step in Embodiment 1, and the specific implementation process is the same, so it will not be repeated here.

[0089] Embodiment 3

[0090] This embodiment provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps in a video anomaly detection method as described in Embodiment 1 above are implemented.

[0091] Embodiment 4

[0092] This embodiment provides a computer device, including a memory, a processor, and a computer program stored on the memory and running on the processor, and when the processor executes the program, the steps in a video anomaly detection method as described in Embodiment 1 above are implemented.

[0093] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

[0094] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, it is not a limitation to the protection scope of the present invention. Those skilled in the art should understand that based on the technical solutions of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present invention.

Claims

1. A video anomaly detection method, characterized in that, Including: Obtain a sequence of consecutive video frames; Based on the sequence of consecutive video frames, detect abnormal events through a video anomaly detection model; The backbone network of the video anomaly detection model is based on the 3D-Swin Transformer model, and completes the reconstruction task of the input image through two downsampling and two upsampling operations; the feature extractor in Swin Transformer is combined with a 3D encoder to learn normal behavior patterns from anomaly-free video frames, and a 3D decoder is used to predict future frames, and skip connections are used to connect feature maps from the encoder and the decoder; The video anomaly detection model includes an encoder and a decoder. The operations of the encoder include: the first downsampling. After integrating the sequence of consecutive video frames into a tensor, through linear embedding, and using a sliding window for feature extraction, several feature maps are obtained. The feature maps sequentially pass through two layers of ST-3D blocks, a time shift operation, and two layers of dynamic large kernel convolution operations; after the feature maps go through two layers of dynamic large kernel convolution operations, through block merging, the second downsampling is performed. After passing through six layers of ST-3D blocks, one layer of dynamic large kernel convolution layer, two layers of ST-3D blocks and one layer of dynamic large kernel convolution operation, the encoder output is obtained; The operations of the decoder include: after cascading the encoder output with the second skip connection, sequentially passing through an upsampling operation, six layers of ST-3D blocks, one layer of dynamic large kernel convolution operation, cascading with the first skip connection, upsampling and two layers of ST-3D blocks to obtain the decoder output; Among them, the video anomaly detection model performs a time shift operation between two consecutive feature maps, replacing a part of the channels of each feature map with the same channel part from the previous feature map to extract time features; extracting time features includes: extracting the first 1 / n channel features of the video frame, extracting feature maps from each residual block and caching them in memory. For the next frame, replacing the first 1 / n of the current feature map with the cached feature map, and using the combination of (n - 1) / n of the current feature map and 1 / n of the old feature map to generate the feature map for the input of the next layer, where n represents a positive integer; Then, enter two layers of dynamic large kernel convolution operations, accumulate spatial features through a recursive aggregation method, capture multi-scale context information, and extract multi-scale features.

2. The video anomaly detection method according to claim 1, characterized in that The second layer of dynamic large kernel convolution operation in the two layers of dynamic large kernel convolution operations establishes a first skip connection.

3. The video anomaly detection method according to claim 1, wherein One layer of dynamic large kernel convolution between the six layers of ST-3D blocks and the two layers of ST-3D blocks establishes a second skip connection.

4. A video anomaly detection system, characterized in that, Including: A data acquisition module configured to: obtain a sequence of consecutive video frames; An anomaly detection module configured to: based on the sequence of consecutive video frames, detect abnormal events through a video anomaly detection model; The video anomaly detection model adopts a Transformer model based on a sliding window of Swin Transformer and an unsupervised network structure decoder in 3D, and uses skip connections to connect feature maps from the encoder and the decoder; The video anomaly detection model includes an encoder and a decoder. The operations of the encoder include: the first downsampling. After integrating a sequence of consecutive video frames into a tensor, through linear embedding, and using a sliding window for feature extraction, a number of feature maps are obtained. The feature maps are successively passed through two layers of ST-3D blocks, a time shift operation, and two layers of dynamic large kernel convolution operations. After the feature maps undergo two layers of dynamic large kernel convolution operations, through block merging, the second downsampling is performed. After passing through six layers of ST-3D blocks, one layer of dynamic large kernel convolution layer, two layers of ST-3D blocks, and one layer of dynamic large kernel convolution operation, the encoder output is obtained. The operations of the decoder include: after cascading the encoder output with the second skip connection, successively passing through an upsampling operation, six layers of ST-3D blocks, one layer of dynamic large kernel convolution operation, cascading with the first skip connection, upsampling, and two layers of ST-3D blocks to obtain the decoder output. Among them, the video anomaly detection model performs a time shift operation between two consecutive feature maps, replacing a part of the channels of each feature map with the same channel part from the previous feature map to extract time features. Extracting time features includes: extracting the first 1 / n channel features of the video frames, extracting feature maps from each residual block and caching them in the memory. For the next frame, replacing the first 1 / n of the current feature map with the cached feature maps, and using a combination of n - 1 / n of the current feature map and 1 / n of the old feature map to generate the feature map for the input of the next layer, where n represents a positive integer. Then, it enters two layers of dynamic large kernel convolution operations, accumulating spatial features through a recursive aggregation method, capturing multi-scale context information, and extracting multi-scale features.

5. A computer-readable storage medium having a computer program stored thereon, the program being executed by a processor, characterized in that, When the program is executed by a processor, it implements the steps in a video anomaly detection method as described in any one of claims 1-3.

6. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in a video anomaly detection method as described in any one of claims 1-3.

Citation Information

Patent Citations

  • Video anomaly detection method based on clustering guide learning

    CN117746291A

  • KR20230095845A