A Video Abnormal Event Detection Method Based on Perceived Local Response

By adopting a method of perceived local response in video abnormal event detection, using the technology of video frame prediction and multi-module alternating use, the problem of insensitive to local area details in the prior art is solved, and high accuracy detection and prediction of parking lot abnormal events are achieved.

CN116863404BActive Publication Date: 2025-06-13QINGDAO SONLI SOFTWARE INFORMATION TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310847795.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-12
Publication Date
2025-06-13
Estimated Expiration
2043-07-12

AI Technical Summary

Technical Problem

Existing video anomaly event detection methods are insensitive to local area details, especially under the powerful fitting ability of deep learning networks, it is impossible to effectively perceive subtle differences in anomaly events.

Method used

The video abnormal event detection method that perceives local response is adopted, parking lot video is collected through monitoring equipment, abnormal video frame prediction method is used to detect abnormal video frames, and through modules such as space-time block cutting, local detail constraint enhancement, dynamic information perception and cross-scale information gated filtering, the receptive field is gradually expanded to realize the detection and prediction of abnormal events.

Benefits of technology

This method can effectively perceive the subtle information of abnormal events in parking lots, solve the problem of insensitive detection of local abnormal areas in the prior art, and improve the accuracy and intelligence of abnormal event detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116863404B_ABST
    Figure CN116863404B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of parking lot anomaly detection, and relates to a video anomaly event detection method for perceiving local response. First, a data set is constructed, features based on image patches are extracted, and then window merging is performed. By using a local detail constraint enhancement module, the local detail constraints between different image patches are learned to obtain window features after local detail constraint enhancement. Then, a dynamic information perception module is used to learn the motion information between image patches. Next, scale-level perception is performed, and a cross-scale information gating filtering module is used to remove redundant feature information. After obtaining the information required to restore the predicted frame of the video segment, it is converted into an output image, thereby predicting the next frame of the current video segment. Finally, the video frame prediction network is trained and the detection result is obtained through inference, which can not only be used for anomaly perception in anomaly event detection, but also be used in related fields such as motion detection and video object detection in complex scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of parking lot anomaly detection, and relates to a video anomaly event detection method for perceiving local response. Background Art

[0002] With the development of society and the progress of economy, cars have become very popular, which has led to an increasing number of parking lots. Therefore, the intelligent transformation of parking lots has become very important. With the popularization of intelligent devices, artificial intelligence and computer vision have become essential technologies for intelligent transformation. Whether it is the intelligent management of parking lots or the autonomous driving of vehicles, intelligence is indispensable. Along with the intelligent parking lot, as the information input end of intelligent parking, the intelligence of monitoring devices is particularly important. The intelligence of monitoring devices can monitor abnormal events occurring in the parking lot in real time, including events such as vehicles being scratched and vehicles being deliberately damaged by pedestrians, so as to discover and alarm vehicle abnormal events in the parking lot in a timely manner, thereby greatly improving the safety of the parking lot.

[0003] At present, the video data of parking lots is gradually increasing. Through artificial intelligence algorithms, it has become possible to analyze and predict relevant tasks for video data, especially the methods for analyzing video anomaly event detection are gradually increasing. However, the current video anomaly event detection methods are not sensitive to the details of local areas. Especially due to the powerful fitting ability of deep learning networks, when detecting abnormal events with slight differences, they cannot perceive the slight differences. Therefore, a video anomaly event detection method that can perceive slight abnormal information is needed to provide intelligent improvement for parking lot anomaly event detection. Summary of the Invention

[0004] The purpose of the present invention is to overcome the shortcomings of the existing technology and design and provide a video anomaly event detection method for perceiving local response. The monitoring device is used to collect parking lot videos, and the method of video frame prediction is adopted to detect abnormal video frames, so as to realize the detection and prediction of abnormal events.

[0005] To achieve the above purpose, the specific process of the present invention for realizing video anomaly event detection based on perceiving local response is as follows:

[0006] (1) Collect image data from the abnormal detection video dataset to construct a dataset, and divide the constructed dataset into a training set, a validation set, and a test set after binary labeling according to normal images and abnormal images;

[0007] (2) Use the spatio-temporal block cutting module to obtain features based on image blocks by means of 2D convolution;

[0008] (3) The image patch features obtained in step (2) are merged into windows by dividing the image window, and the local detail constraint enhanced module uses the multi-head attention mechanism to learn the local detail constraints between different image patches to obtain the window features after local detail constraint enhancement;

[0009] (4) The dynamic information perception module uses 3D convolution on the features after local detail constraint enhancement in step (3) to learn the motion information between image patches;

[0010] (5) The receptive field of the network is gradually expanded to global information in a cascaded manner, that is, the spatio-temporal block cutting module, the dynamic information perception module, and the local detail constraint enhanced module are alternately used for scale-level perception;

[0011] (6) The cross-scale information gating and filtering module is used to remove redundant feature information to obtain the information required to restore the predicted frame of the video segment;

[0012] (7) The information required for the predicted frame is converted into an output image, thereby predicting the next frame of the current video segment;

[0013] (8) The video frame prediction network is trained to obtain a trained video frame prediction network;

[0014] (9) The SSIM value between the original image and the predicted image output by the video frame prediction network is selected as the criterion for judging whether the video frame is an abnormal frame, and the threshold parameter of the SSIM value is calculated. If the SSIM value is less than the threshold, the video frame is considered an abnormal video frame. If the SSIM value is greater than the threshold, the predicted video image is considered to be very similar to the real video image and there is no abnormality.

[0015] As a further technical solution of the present invention, the image patch-based features obtained in step (2) are: , where represents the image patch-based features obtained after normalizing the input image , t represents the number of input images, represents the normalization process, represents the 2D convolution operation.

[0016] As a further technical solution of the present invention, the features after merging the image patch features into windows in step (3) are: , and the window features after local detail constraint enhancement are: , where represents the window features after local detail constraint enhancement of the i-th frame image, represents the position bias, which is used to determine the position relationship between image patches, represents the softmax function, It represents the feature after window merging of the image block features generated in the i-th frame. It represents the image block merging operation.

[0017] As a further technical solution of the present invention, the process of step (4) is as follows: , where It represents the window feature after the i-th frame image is perceived by dynamic information, It represents using 3D convolution operation to perceive motion information.

[0018] As a further technical solution of the present invention, the process of step (5) is as follows:

[0019] ,

[0020] In the of the dynamic information perception module, there are two operations including spatial-level convolution and temporal-level convolution. Among them, the temporal-level convolution is used to perceive motion information, and the spatial-level convolution is used to perceive object scale information.

[0021] As a further technical solution of the present invention, the process of step (6) is as follows:

[0022] ,

[0023] It represents that the reverse decoding network contains the feature after the superposition of the output features of the intermediate layer of the forward decoding network, It represents the feature superposition module, It represents the gated switch filtering module, and the gated switch filtering module is implemented by ConGRU; among them, the forward encoding network is:

[0024] ,

[0025] The reverse decoding network is:

[0026] .

[0027] As a further technical solution of the present invention, the image output in step (7) is:

[0028] ,

[0029] Among them, It represents the output image of the current prediction network, and the last is not connected to the in the .

[0030] As a further technical solution of the present invention, the specific process of step (8) is as follows: Select an image training video frame prediction network with images labeled as normal images. The energy loss function adopted by the network structure is the L2 loss. The continuous images in the training set are fed into the video frame prediction network. According to the number of images B required for each training, they are sequentially input into the video frame prediction network and then the next frame prediction image of the current video segment is output . After 429 full training set training iterations, the model parameters with the highest accuracy on the validation set are saved as the parameters of the finally trained model and saved to the local folder to obtain the trained video frame prediction network.

[0031] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0032] (1) The present invention provides a novel abnormal event detection method, which can perceive the occurrence of abnormal events in the parking lot, solves the problem of insensitive detection of details in local abnormal areas in abnormal detection, especially the problem that the predicted video frames cannot predict abnormalities caused by the strong fitting ability of the network, and even the problem of inaccurate thresholds.

[0033] (2) The method provided by the present invention can not only be used for abnormal perception in abnormal event detection, but also be used in related fields such as motion detection and video object detection in complex scenarios. Description of the Drawings

[0034] Figure 1 It is a work flow block diagram for the present invention to realize video abnormal event detection.

[0035] Figure 2 It is a network structure diagram for the present invention to realize video abnormal event detection. Detailed Embodiments

[0036] The present invention will be further described below with reference to the drawings and through embodiments, but the scope of the present invention is not limited in any way.

[0037] Embodiment:

[0038] As Figure 1 and Figure 2 shown, the present embodiment provides a video abnormal event detection method based on perceptual local response, which specifically includes the following steps:

[0039] (1) Dataset construction: To construct a video abnormal detection dataset, in this embodiment, one image data is collected every 5 frames from the abnormal detection video dataset to construct the dataset. If the video frame rate is 25, a total of 5 frames are collected, and the constructed image dataset is binarized (1 / 0) labeled according to normal images and abnormal images, and the dataset is divided into a training set, a validation set and a test set;

[0040] (2) Spatiotemporal block cutting: Since the pixel-based method brings huge computations, in order to capture the global strong correlation between abnormal regions, in this embodiment, the method of dividing into image blocks is adopted. By means of image blocks, the network can perceive the correlation between the global parts. At the same time, calculating the correlation between image blocks can be faster than calculating the pixel-level correlation. To implement the above-mentioned block cutting operation, the spatiotemporal block cutting module uses 2D convolution to obtain the features based on image blocks:

[0041] , where represents the features based on image blocks obtained after the input image is normalized. t represents the number of input images, represents the normalization process, represents the 2D convolution operation. In this embodiment, 3D convolution is not adopted because 3D convolution will lose the time dimension information. Through 2D convolution, the image can be divided into local image blocks without losing the timing information ( );

[0042] (3) Enhancement of local detail constraints: Through step (2), the features based on image blocks are obtained. To speed up the execution speed of the network, in this embodiment, the method of dividing image windows is adopted to merge the image blocks into windows ( ). At the same time, to increase the perceptual diversity of abnormal features in the video, the local detail constraint enhancement module adopts the multi-head attention mechanism ( ) , That is, the features of the abnormal region are passed through the self-attention mechanism to learn the local detail constraints between different image blocks to obtain the window features after local detail constraint enhancement:

[0043] ,

[0044] ,

[0045] where represents the window features after local detail constraint enhancement of the i-th frame image, represents the position bias, which is used to determine the position relationship between image blocks, represents the softmax function, represents the features after window merging of the image block features generated in the i-th frame, represents the image block merging operation. This process does not involve the network parameter learning process and mainly uses data dimension conversion and feature splitting;

[0046] (4) Dynamic information perception: Since motion information usually contains abnormal regions, in order to perceive motion information, existing methods often adopt the method of optical flow calculation. However, the optical flow method has many limitations, such as low optical flow quality and slow optical flow calculation speed, etc. To ensure the timeliness and quality of motion perception, this embodiment adopts the method of 3D convolution to perceive motion information, that is, the dynamic information perception module learns the motion information between image blocks through 3D convolution of the features enhanced by local detail constraint in step (3):

[0047] ,

[0048] Among them, represents the window feature after the i-th frame image passes through dynamic information perception, represents the use of 3D convolution operation to perceive motion information;

[0049] (5) Scale-level perception: The window features enhanced by local detail constraint obtained through step (3) and the window features after dynamic information perception obtained through step (4) can cover abnormal event detection in different scenarios. However, the above methods can only perceive local detail information. In order to integrate global constraints and scale-level constraint information into the network, this embodiment adopts a cascaded method to gradually expand the receptive field of the network to global information, that is, the spatio-temporal block cutting module, the dynamic information perception module, and the local detail constraint enhancement module are used alternately. Specifically:

[0050] ,

[0051] In there are two operations: spatial-level convolution and temporal-level convolution. Among them, the temporal-level convolution is used to perceive motion information, and the spatial-level convolution can be used to perceive object scale information. Therefore, in the scale-level operation process, in adopts (H / 2; W / 2; T / 2, where H represents height, W represents width, and T represents temporal length), which contains spatial-level scale information;

[0052] (6) Cross-scale information gating and filtering: Through the above steps, the abnormal information of the video can be encoded to obtain high-dimensional semantic information. To predict the next frame of the current video segment, it is necessary to decode the information (motion information, local detail information) encoded by the forward encoding network, so as to output the features required to predict the next frame in the current video segment.

[0053] The forward encoding network is:

[0054] ,

[0055] The reverse decoding network is:

[0056] ,

[0057] However, in the process of cross-scale information fusion, if the short-link method is directly adopted, feature redundancy will occur. When the network predicts the next frame of the image, it needs to transmit the information required for prediction. Therefore, in this embodiment, the cross-scale information gating filter method is adopted to remove redundant information.

[0058] ,

[0059] represents the features after the superposition of the output features of the intermediate layer of the forward decoding network in the reverse decoding network. represents the feature superposition module. represents the gating switch filter module, which is mainly implemented by ConGRU and is used to remove redundant information in the features.

[0060] (7) Prediction image output: Obtain the information required to restore the predicted frames of the video segment through step (6), and convert the information required for the predicted frames into the output image, so as to predict the next frame of the current video segment:

[0061] ,

[0062] Among them, represents the output image of the current prediction network. In this embodiment, in the last is not connected to in , because the final output will not need to restore the first t frames of the image, but only need to output the predicted image of the (t + 1)-th frame. At the same time, if the forward propagation features of the first t frames of the image are introduced, the difference of the predicted (t + 1)-th frame image will be too small;

[0063] (8) Video frame prediction network training: Select images labeled as normal images to train the video frame prediction network. The energy loss function adopted by the network structure is the L2 loss. The continuous images in the training set are sent into the video frame prediction network. According to the number of images B required for each training, they are sequentially input into the video frame prediction network and then the next frame prediction image of the current video segment is output . After 429 full training set training iterations, save the model parameters with the highest accuracy on the validation set as the parameters of the finally trained model, and save them to the local folder to obtain the trained video frame prediction network;

[0064] (9) Obtain the result through inference: Select the original image and the predicted image output by the prediction network The SSIM value between them is used as the criterion for judging whether a video frame is an abnormal frame, and the threshold parameter of the SSIM value is calculated. If the SSIM value is less than the threshold, the video frame is considered an abnormal video frame. If the SSIM value is greater than the threshold, it is considered that the predicted video image is very similar to the real video image and there is no abnormality.

[0065] The algorithms and calculation processes not detailed in this article are all common techniques in this field.

[0066] It should be noted that the purpose of publishing the embodiments is to help further understand the present invention. However, those skilled in the art can understand that various substitutions and modifications are possible without departing from the spirit and scope of the present invention and the appended claims. Therefore, the present invention should not be limited to the content disclosed in the embodiments, and the scope of protection claimed by the present invention is subject to the scope defined by the claims.

Claims

1. A method for detecting video abnormal events by perceiving local responses, characterized in that, the specific process is as follows: (1) Collect image data from the abnormal detection video dataset to construct a dataset, and divide the constructed dataset into a training set, a validation set, and a test set after binary annotation according to normal images and abnormal images; (2) Use a 2D convolution method through a spatio-temporal block cutting module to obtain features based on image blocks; (3) Use the method of dividing image windows to merge the image block features obtained in step (2), and use a multi-head attention mechanism through a local detail constraint enhancement module to learn the local detail constraints between different image blocks to obtain window features after local detail constraint enhancement; (4) The dynamic information perception module learns the motion information between image blocks from the features after local detail constraint enhancement in step (3) through 3D convolution; (5) Use a cascaded method to gradually expand the receptive field of the network to global information, that is, the spatio-temporal block cutting module, the dynamic information perception module, and the local detail constraint enhancement module are alternately used for scale-level perception; (6) Use a cross-scale information gating filtering module to remove redundant feature information to obtain the information required to restore the predicted frame of the video segment; (7) Convert the information required for the predicted frame into an output image, so as to predict the next frame of the current video segment; (8) Train the video frame prediction network to obtain a trained video frame prediction network; (9) Select the SSIM value between the original image and the predicted image output by the video frame prediction network as the judgment criterion for whether the video frame is an abnormal frame, and calculate the threshold parameter of the SSIM value. If the SSIM value is less than the threshold, it is considered that the video frame is an abnormal video frame. If the SSIM value is greater than the threshold, it is considered that the video predicted image is very similar to the real video image and there is no abnormality.

2. The method for detecting video abnormal events by perceiving local responses according to claim 1, characterized in that, The feature based on image patches obtained in step (2) is: , where represents the feature based on image patches obtained after normalizing the input image , t represents the number of input images, represents the normalization process, represents the 2D convolution operation.

3. The method for detecting video abnormal events by perceiving local responses according to claim 2, characterized in that, The feature after window merging of the image patch features in step (3) is: , and the window feature after enhanced local detail constraint is: , where represents the window feature after enhanced local detail constraint of the i-th frame image, represents the position bias, which is used to determine the positional relationship between image patches, represents the softmax function, represents the feature after window merging of the image patch features generated in the i-th frame, represents the image patch merging operation.

4. The method for detecting video abnormal events by perceiving local responses according to claim 3, characterized in that, The process of step (4) is as follows: , where represents the window feature after the dynamic information perception of the i-th frame image, represents the use of 3D convolution operation to perceive motion information.

5. The method for detecting video abnormal events by perceiving local responses according to claim 4, characterized in that, The process of step (5) is as follows: , In the dynamic information perception module, there are two operations: spatial-level convolution and temporal-level convolution. The temporal-level convolution is used to perceive motion information, and the spatial-level convolution is used to perceive object scale information.

6. The method for detecting video abnormal events by perceiving local responses according to claim 5, characterized in that, The process of step (6) is as follows: , The representative reverse decoding network includes the features after superimposing the output features of the intermediate layers of the forward decoding network. The representative feature superimposing module The representative gating switch filtering module is implemented by ConGRU. Where the forward encoding network is: , The reverse decoding network is: 。 7. The method for detecting video abnormal events by perceiving local responses according to claim 6, characterized in that, The image output in step (7) is: , Among them, represents the output image of the current prediction network, and the last is not connected in in .

8. The method for detecting video abnormal events by perceiving local responses according to claim 7, characterized in that, The specific process of step (8) is as follows: Select the images marked as normal images to train the video frame prediction network. The energy loss function adopted by the network structure is the L2 loss. The consecutive images in the training set are fed into the video frame prediction network. According to the number of images B required for each training, they are sequentially input into the video frame prediction network and then the predicted image of the next frame of the current video segment is output . After 429 training iterations of the complete training set, the model parameters with the highest accuracy on the validation set are saved as the parameters of the finally trained model and saved to the local folder to obtain the trained video frame prediction network.

Citation Information

Patent Citations

  • Stereoscopic video quality evaluation method based on motion significance

    CN106875389A

  • Parking lot abnormal event detection method

    CN115082870A