Abnormal behavior detection method based on multi-scale space-time prediction network
Through the combination of HDC_net and DB-ConvLSTM_net, multi-scale spatial features and timing information are extracted, and the problems of video frame timing relationship and spatial information loss in the prior art are solved, thereby improving the accuracy of abnormal behavior detection.
Patent Information
- Application Number
- CN202510206194.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-07-18
AI Technical Summary
Existing video prediction algorithms ignore the timing relationship between video frames, resulting in the loss of time information, and multiple downsampling operations lead to the lack of spatial hierarchical information, affecting the abnormal behavior detection effect.
HDC_net is used to extract multi-scale spatial features, combine DB-ConvLSTM_net to memorize the timing information between continuous video frames, and fuse multi-scale apparent features and dynamic information of continuous video frames, and strengthen the representation performance of the model through HDC_net and DB-ConvLSTM_net structures.
It improves the accuracy of the detection of abnormal behavior by the video prediction model, effectively retains the time and space information of the video frame, and improves the performance of abnormal detection.
Smart Images

Figure CN120339896A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of video prediction, and particularly relates to a method for abnormal behavior detection based on a multi-scale spatio-temporal prediction network. Background Art
[0002] In recent years, intelligent video surveillance such as intelligent vehicle monitoring systems and smart homes has developed greatly. As the main target in video surveillance, pedestrians have rich attribute tags and have a wide range of application scenarios in real life. Pedestrian attribute recognition technology can predict the visual attributes of a given pedestrian image. When retrieving a large amount of data, it can search for human targets according to the given attributes. In addition to pedestrian search for a large amount of data, pedestrian attributes are also commonly used in pedestrian re-identification tasks in video surveillance. In addition, pedestrian attribute tags are also commonly used in personnel identification tasks, and the facial attributes of pedestrians can be used in face recognition tasks.
[0003] Video prediction algorithms can reasonably model dynamic scenes, speculate on future changes, learn the internal visual representation of data in an unsupervised manner, assist the computer in making better decisions, and improve the abnormal detection performance of the algorithm. For example, using U-Net as a neural network to predict future frames of a video. By adopting skip connections in the network structure of the same depth, the low-level surface features are fused with the high-level semantic information, enriching the feature information and solving problems such as gradient disappearance and uneven hierarchical information. However, this network ignores the temporal relationship between video frames, resulting in the loss of time information in the video frame sequence. At the same time, the U-Net structure performs multiple downsampling and upsampling operations, which will cause the loss of some internal data structures and spatial hierarchical information, weaken the representation performance of the network, and thus affect the detection effect of the model on abnormal behaviors. Summary of the Invention
[0004] To solve the problems mentioned in the above background art; the purpose of the present invention is to provide a method for abnormal behavior detection based on a multi-scale spatio-temporal prediction network.
[0005] A method for abnormal behavior detection based on a multi-scale spatio-temporal prediction network of the present invention has the following detection method:
[0006] Step 1: Use HDC_net to extract the multi-scale spatial features of the target object and learn the change information at different scales;
[0007] Step 2: Use DB-ConvLSTM_net to extract the complex dynamic patterns between video frames by memorizing the temporal information between consecutive video frames.
[0008] Preferably, in the first step, based on the multi-scale spatial feature extraction module of HDC_net, first, multiple feature maps extracted in the early stage of the model are sequentially input into three different branches of the HDC_net structure for feature processing; different dilation rates are adopted for these branch structures, and target objects in the video are automatically extracted through corresponding different receptive fields; the feature maps extracted by each branch structure are concatenated with the initially input feature map and then input into the next layer of the network to perform related operations, further capturing the spatial context information of the target objects in the video image and effectively integrating multi-scale spatial features.
[0009] Preferably, in the second step, based on the time information extraction module of DB-ConvLSTM_net, the DB-ConvLSTM_net structure consists of a shallow forward layer and a deep backward layer.
[0010] Compared with the prior art, the beneficial effects of the present invention are as follows: the multi-scale apparent feature extraction module of the target object is fused with the time series information extraction module to fully extract the multi-scale apparent features of the target and the dynamic information of consecutive video frames. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] For ease of explanation, the present invention is described in detail by the following specific embodiments and the accompanying drawings.
[0012] Figure 1 is the overall framework diagram of the anomaly detection method;
[0013] Figure 2 is the structural diagram of the multi-scale spatio-temporal prediction network;
[0014] Figure 3 is the schematic diagram of the hybrid dilated convolution network structure;
[0015] Figure 4 is the schematic diagram of the DB-ConvLSTM_net structure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0016] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention is described below through specific embodiments shown in the accompanying drawings. However, it should be understood that these descriptions are only exemplary and do not intend to limit the scope of the present invention. The structures, ratios, sizes, etc. shown in the drawings of this specification are only used to cooperate with the content disclosed in the specification for those skilled in this technology to understand and read, and are not used to limit the limiting conditions under which the present invention can be implemented. Therefore, they do not have a technical essence. Any modification of the structure, change of the proportional relationship or adjustment of the size, without affecting the effects that the present invention can produce and the objectives that can be achieved, should still fall within the scope covered by the technical content disclosed in the present invention. In addition, in the following description, the descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present invention.
[0017] Here, it should also be noted that in order to avoid obscuring the present invention due to unnecessary details, only the structures and / or processing steps closely related to the solution according to the present invention are shown in the drawings, while other details less related to the present invention are omitted.
[0018] As Figure 1 shown, the following technical solutions are adopted in this specific embodiment: (1) HDC_net can be used to extract multi-scale spatial features of the target object and learn the change information at different scales; (2) DB-ConvLSTM_net can extract complex dynamic patterns between video frames by memorizing the temporal information between consecutive video frames. The algorithm predicts the future frames of the video from the spatial dimension and the time dimension, effectively improving the accuracy of the prediction model. The algorithm framework is as Figure 1 shown.
[0019] In this embodiment, the overall framework is mainly divided into two parts: the video prediction part and the anomaly detection part. The goal of the first part is to train a G network to predict the future frames of the video. In order to improve the image quality of the predicted frames, a GAN network and some common loss functions are used to optimize the network model. First, MSTP-Net is used as the G network, and the video frame sequence (I1, I2, I3, …, It) before the current frame It+1 is used as the input tensor. The predicted frame I*t+1 is the output tensor. Regarding the D network, PatchGAN is adopted, and the structure further discriminates the difference between the real frame and the predicted frame, prompting the G network to generate an image as consistent with the real frame as possible. Then, the total objective optimization function is used to minimize the gap between the predicted frame and the real frame. By continuously optimizing the model parameters, the predicted frame I*t+1 is made more similar to the real frame It+1. The second part of the overall framework is the anomaly detection part. The test video data is input into the pre-trained network model, and the anomaly degree of the test samples is judged by calculating the regular score of each frame.
[0020] I. Video Prediction:
[0021] The MSTP-Net model architecture is mainly designed based on the U-Net structure, and the specific implementation details are as Figure 2 shown. By adding HDC_net to extract the multi-scale spatial features of the training samples, and then inserting DB-ConvLSTM_net to process the temporal information between consecutive T video frames in a non-linear manner. The overall network architecture of MSTP-Net includes an encoding part and a decoding part. Among them, the input and output sizes of the network are both 256×256×3. The kernel sizes of all convolutions and deconvolutions are set to 3×3, and the convolution kernel size of the max pooling layer is set to 2×2.
[0022] 1.1. Multi-scale Spatial Feature Extraction Module Based on HDC_net:
[0023] To improve computational efficiency and avoid overfitting, U-Net introduces downsampling operations within its own structure. Performing multiple downsamplings will reduce the image quality and lose some spatial detail information. Due to the different positions and angles of the cameras, the styles and sizes of the objects in the video will also be different. When building a new model based on the U-Net structure to detect abnormal behaviors in the video, not only should we consider extracting the multi-scale spatial features of the target object, but also we should make up for the loss of some image detail information caused by the downsampling operation. To solve the problem of image spatial detail loss caused by the downsampling operation of the U-Net structure, it is necessary to retain as much feature information as possible before the downsampling operation. Therefore, starting from the second downsampling, HDC_net is placed in the previous convolutional layer of each downsampling layer to retain more detail information. The reason for not using HDC_net before the first downsampling is that several convolutional operations before the first downsampling will not cause serious loss of image information. The schematic diagram of the internal structure of HDC_net is as Figure 3 shown.
[0024] First, multiple feature maps extracted in the early stage of the model are sequentially input into three different branches of the HDC_net structure for feature processing. These branch structures adopt different dilation rates and automatically extract the target objects in the video through corresponding different receptive fields. After splicing the feature maps extracted by each branch structure with the initially input feature maps, they are input into the next layer of the network to perform related operations, further capturing the spatial context information of the target objects in the video image, effectively fusing the multi-scale spatial features, and strengthening the representation performance of the model.
[0025] 1.2. Temporal Information Extraction Module Based on DB-ConvLSTM_net:
[0026] Design a bidirectional ConvLSTM structure, considering both the forward and backward feature information of the video frame sequence, so as to more comprehensively capture the spatio-temporal information of the surveillance video data, which has important research significance for predicting future frames of the video. It is proposed to input T video frames into the encoder network one by one to sequentially generate corresponding feature maps, avoiding the interaction and mixing of temporal information between multiple video frames, and effectively retaining the temporal features of the video frame sequence.
[0027] As Figure 4As shown, the DB-ConvLSTM_net structure consists of a shallow forward layer and a deep backward layer. Specifically, {Ht^f} represents the forward-order feature map corresponding to the output after being processed by the ConvLSTM units included in the forward layer. After receiving the forward-order output feature map {Ht^f}, the deep backward layer processes it one by one through the ConvLSTM units included in the backward layer to generate the corresponding backward-order feature map {Ht^b}. Since the feature information of the input image is processed by the corresponding ConvLSTM units in the forward and backward layers, these feature information flows can be exchanged with each other, thereby extracting more detailed spatio-temporal information.
[0028] II. Anomaly Detection:
[0029] After training the prediction model with the video frame sequence of normal events, the difference between the predicted frame and its corresponding real frame is used for anomaly detection. In the test stage, the PSNR evaluation criterion is adopted to estimate the anomaly degree score of the current frame. The larger the PSNR value, the more similar the predicted frame is to the real frame, and the greater the probability that the frame is judged to be normal, and vice versa. For the convenience of comparison, each test video frame is normalized to the range of [0,1] one by one to obtain its regular score.
[0030]
[0031] In the formula, mintPSNR and maxtPSNR respectively correspond to the minimum and maximum PSNR values of all frames in each test video.
[0032] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention.
[0033] In addition, it should be understood that although this specification is described according to the embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A method for detecting abnormal behaviors based on a multi-scale spatio-temporal prediction network, characterized in that: Its detection method is as follows: Step 1: Use HDC_net to extract multi-scale spatial features of the target object and learn the change information at different scales; Step 2: Use DB-ConvLSTM_net to extract complex dynamic patterns between video frames by memorizing the temporal information between consecutive video frames.
2. The method for detecting abnormal behavior based on a multi-scale spatio-temporal prediction network according to claim 1, wherein: The first step is based on the multi-scale spatial feature extraction module of HDC_net. First, multiple feature maps extracted in the early stage of the model are sequentially input into three different branches of the HDC_net structure for feature processing; different dilation rates are used in these branch structures to automatically extract the target object in the video through corresponding different receptive fields; the feature maps extracted by each branch structure are concatenated with the initially input feature map and then input into the next layer of the network to perform related operations, further capturing the spatial context information of the target object in the video image and effectively fusing multi-scale spatial features.
3. The method for detecting abnormal behavior based on a multi-scale spatio-temporal prediction network according to claim 1, wherein: The second step is based on the time information extraction module of DB-ConvLSTM_net. The DB-ConvLSTM_net structure consists of a shallow forward layer and a deep backward layer.