A traffic abnormal behavior detection method and system
Through background modeling and twin cross-correlation combined with P3D-Attention, the state of traffic abnormal static vehicles and the start time of abnormal behavior is estimated, which solves the problems of high calculation costs and low detection efficiency in the prior art, and achieves efficient and accurate traffic abnormal behavior detection.
Patent Information
- Application Number
- CN202210457144.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-28
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-04-28
AI Technical Summary
The existing traffic abnormal behavior detection methods require a large amount of calculation costs, memory requirements and time costs, making it difficult to efficiently detect the status of traffic abnormalities and accurately estimate the start time of abnormal behavior.
Background modeling is used to remove the normally moving vehicles in the traffic surveillance video, detect abnormal target vehicles through perspective cropping, and use twin cross-correlation to combine P3D-Attention with network model to perform abnormal vehicle status detection and target matching, and estimate the start time of abnormal behavior.
The speed and performance of traffic abnormal static vehicle state detection is improved, the abnormal behavior start time is accurately estimated, the calculation cost and time cost are reduced, and more efficient traffic abnormal behavior detection is achieved.
Smart Images

Figure CN114821421B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent video analysis, and particularly relates to a traffic abnormal behavior detection method and system. Background Art
[0002] Computer vision has developed into a key technology for many applications. Compared with other information sources such as manual monitoring, GPS, and radar signals, visual data contains rich information. Therefore, visual data also plays a crucial role in detecting / predicting traffic congestion, accidents, and other abnormal phenomena.
[0003] Abnormal behavior detection is a subfield of behavior understanding in surveillance scenarios. An abnormality is usually expressed as a deviation of a scene entity from normal behavior, which is a data pattern that does not conform to the defined concept of good normal behavior.
[0004] Typical abnormal behavior detection methods learn normal behavior through training, use normal training samples to model the abnormal concept, and identify deviations from the normal pattern as abnormal phenomena. Any behavior that significantly deviates from normal behavior can be called an abnormality. For example, a vehicle stalling on the road, bypassing a traffic signal at a traffic intersection, or a vehicle making a U-turn at a red light are all abnormal phenomena. Some studies use a series of statistical patterns to learn abnormalities; such as sparse reconstruction, which calculates errors based on reconstruction in a semi-supervised manner to measure abnormalities, and regards behaviors that deviate from normal behavior as abnormal behaviors. And with the development of deep learning technology, deep autoencoders with reconstruction loss can be used to solve the abnormal prediction task. Although these methods have made great progress in abnormal datasets such as CUHKAvenue, they are not suitable for detecting road traffic with more complex and unknown scenarios and local abnormal behaviors.
[0005] For the classification-based abnormal behavior detection method, abnormal data with label information is used to model the abnormal abstraction, and a classifier is used to distinguish the normal class and the abnormal class in a given feature space. This requires a large number of labeled instances containing normal classes and abnormal classes. One method is to design a graph convolutional network to correct noisy labels, so that the network can provide clean data for the action classifier for supervised learning, and the frame-level AUC score obtained on UCF-Crime is 82.12%. Use weakly labeled training videos and regard video segments as instances in multi-instance learning (MIL) for training. The trained deep abnormal ranking model can obtain a high abnormal score on the Real-world dataset.
[0006] In basic computer vision problems such as image classification, object detection, and tracking, great success has been achieved with the rapid development of deep learning, and its classification accuracy even exceeds that of humans. Usually, in tracking-assisted anomaly detection, a large-scale labeled dataset is not required, and it can be easily transferred to other unknown scenarios. For road scenes, unsupervised anomaly detection methods relying on object detection and tracking have attracted quite a number of researchers. However, the assumptions used in anomaly detection cannot be generally applied to different traffic scenarios. Therefore, some methods systematically divide the road traffic analysis into four levels, namely image acquisition, static and dynamic feature extraction, behavior understanding, and its subsequent services.
[0007] In the context of traffic visual surveillance, local anomalies are also known as point anomalies. A vehicle stationary on a normal road can be called a point anomaly. Some methods use background modeling to remove all moving objects while keeping the stationary vehicles in the background for analyzing potential static vehicles. For example, trajectory estimation is performed by learning the dynamic patterns of vehicles in the vehicle trajectories to find abnormal trajectories for anomaly detection. "Zhao J, Yi Z, Pan S, et al. Unsupervised Traffic Anomaly Detection Using Trajectories[C] / / IEEE Computer Society Conference on Computer Vision and Pattern Recognition Workshops, 2019." designs a multi-object tracking method combining single-object tracking results for background image sequences to solve the unsupervised anomaly detection problem. In the method disclosed in "Bai S, He Z, Lei Y, et al. Traffic Anomaly Detection via Perspective Map Based on Spatial-Temporal Information Matrix[C] / / 2019 IEEE / CVF Conference on Computer Vision and Pattern Recongnition, 2019: 117–124.", the detection accuracy is improved through average background modeling method, unified perspective map cropping and re-detection by FPN-DCN network, and the spatio-temporal information matrix is constructed using the trajectory detection results to estimate the anomaly start time. In the method disclosed in "Wang G, Yuan X, Zhang A, et al. Anomaly Candidate Identification and Starting Time Estimation of Vehicles from Traffic Videos[C] / / IEEE Computer Society Conference on Computer Vision and Pattern Recognition Workshops, 2019: 382–390.", background modeling and YOLOv3 method are used for abnormal object detection, and TNT (TrackletNet Tracker) is trained to extract the trajectories of abnormal candidate objects to estimate the anomaly start time.In the method disclosed in 《Li Y, Wu J, Bai X, et al. Multi-Granularity Tracking with Modularlized Components for Unsupervised Vehicles Anomaly Detection[C] / / IEEE Computer Society Conference on Computer Vision and Pattern Recognition Workshops, 2020, 2020-June: 2501–2510.》, a detection network is constructed using MOG2 (Mixture of Gaussians) background modeling and Faster R-CNN, and a multi-granularity tracking algorithm is proposed. Box-level tracking and pixel-level tracking are respectively used to predict the start time of anomalies.
[0008] However, all these methods require using object tracking to obtain high-level trajectory features, which consume a large amount of computational cost, memory requirements, and time cost. Therefore, there is an urgent need and necessity to study traffic anomaly behavior detection methods and systems that can reduce computational cost, improve prediction efficiency, and shorten detection time. Summary of the Invention
[0009] The object of the present invention is: aiming at the deficiencies of the prior art, to provide a traffic anomaly behavior detection method and system, which can improve the speed and performance of detecting the state of abnormally stationary vehicles in traffic, can solve the problem of detecting traffic anomaly time, and thus accurately estimate the start time of abnormal behavior.
[0010] Specifically, the present invention is implemented by the following technical solutions.
[0011] On the one hand, the present invention provides a traffic anomaly behavior detection method, including;
[0012] Through background modeling, the normally moving vehicles in each frame of the traffic surveillance video are removed from the frame, so that the abnormally stationary target vehicles are retained in the background;
[0013] Perspective cropping is performed on each frame of the background extracted through background modeling, and a cropping frame for cropping the traffic surveillance video is obtained according to the vehicle size; Anomaly target detection is performed on the traffic surveillance video every first number of frames to detect the abnormally target vehicles in the video frames of the traffic surveillance video; After detecting the abnormally target vehicles, a cropped picture of the abnormally target vehicles is cropped out, and the traffic surveillance video is cropped forward or backward with the position of the abnormally target vehicle as the center to obtain a cropped video segment with a spatial capacity of several times the vehicle size and a time of the second number of frames for subsequent anomaly time estimation;
[0014] Estimate the abnormal start time based on the abnormal target vehicle detection result, including abnormal vehicle state detection and abnormal target matching;
[0015] For the abnormal vehicle state detection, input the cropped image of the abnormal target vehicle and the cropped video segment into the combined network model of siamese cross-correlation and P3D-Attention to detect whether the abnormal target in the cropped video segment is in a stationary state or a driving state. According to the abnormal vehicle state detection result, classify the cropped video segment and label it as abnormal, driving or normal respectively, and determine the video frame when the abnormal target vehicle has an abnormal behavior;
[0016] For the abnormal target matching, input the vehicle image to be matched and the cropped image of the abnormal target vehicle into the combined network of siamese cross-correlation and P3D-Attention to determine whether the vehicle to be matched is an abnormal target vehicle. Combine the video frame when the abnormal target vehicle has an abnormal behavior determined in the abnormal target vehicle detection result to determine the start time and end time of the traffic abnormal behavior.
[0017] Furthermore, the background modeling adopts the MOG2 algorithm.
[0018] Furthermore, the abnormal target detection adopts the YOLOv3 target detection method.
[0019] Furthermore, the inputs for the abnormal vehicle state detection are, one is the cropped image of the abnormal target vehicle, and the other is a video segment with a spatial capacity 2 times the size of the vehicle and a set number of frames as the second number of frames cropped centered on the abnormal target vehicle;
[0020] The abnormal vehicle state detection includes: respectively extracting the feature maps of the input cropped image of the abnormal target vehicle and the feature maps of each frame of the input video segment through the P3D-Attention module, and improving the correlation of the important channel features through P3D-Attention; then selecting the respectively extracted feature maps to be fused through the siamese cross-correlation operation in three layers with different receptive field sizes; the results of the first two siamese cross-correlation operations use the multiply method to fuse the feature maps extracted from the input image and the input video through the siamese cross-correlation operation with the feature maps of each frame of the input video segment extracted through the P3D-Attention module before cross-correlation; use GAP for pooling; finally, directly classify the input cropped image of the abnormal target vehicle and the input video segment through the softmax layer, classify the input video segment according to the abnormal vehicle state detection result and label it as abnormal, driving or normal respectively, and determine the video frame where the abnormal time point occurs.
[0021] Further, the inputs matched with the abnormal target vehicle are, firstly, the cropped images of the abnormal target vehicle, and secondly, the positive and negative class images for the abnormal target matching model;
[0022] The matching of the abnormal target vehicle includes: respectively extracting the feature maps of the two input images through the P3D-Attention module, and performing a siamese cross-correlation operation using the extracted feature maps of the two inputs; selecting the respectively extracted feature maps to be fused through the siamese cross-correlation operation at three different receptive field sizes; obtaining three feature maps of different sizes, and performing a concatenate merge operation; finally, directly classifying the cropped image of the input abnormal target vehicle and the positive and negative class images for the abnormal target matching model through the softmax layer, and the classification result is matching or non-matching.
[0023] Further, the P3D-Attention module simulates a 3×3×3 convolution in the spatial domain and the temporal domain by using 1×3×3 convolution kernels and 3×1×1 convolution kernels respectively, and decouples the 3×3×3 convolution in terms of time and space; the P3D-Attention module also includes a dual-channel attention module and a spatial attention module for improving the relevance of important features.
[0024] Further, the dual-channel attention module combines a 1×3×3 convolution kernel in the spatial domain and a 3×1×1 convolution kernel in the temporal domain into a 3×3×3 convolution to learn the weights of the frame attention module to express the attention to frames, and learn the weights of the channel attention module
[0025] to express the attention to channels; where F represents the number of frames of the feature map, C represents the number of channels of the feature map, H represents the height of the feature map, and W represents the width of the feature map.
[0026] Further, the spatial attention module learns the weight matrix of the single-channel feature map
[0027] through a 2D convolution kernel to determine the importance and relevance of each position in the video feature map; where F represents the number of frames of the
[0028] feature map, C represents the number of channels of the feature map, H represents the height of the feature map, and W represents the width of the feature map. On the other hand, the present invention also provides a traffic abnormal behavior detection system for implementing the above traffic abnormal behavior detection method, including:
[0027] A video acquisition and clipping module, which acquires real-time traffic video streams, is used to provide continuous traffic video stream information, display it on the display module, and transmit it as input information to the background modeling module;
[0028] The background modeling module reserves an interface. By inputting video stream data with the correct format and through background modeling, the normally moving vehicles in each frame of the traffic surveillance video are removed from the frame, so that the abnormally stationary target vehicles are retained in the background;
[0029] The abnormal target detection module performs perspective cropping on each frame of the background extracted through background modeling, and obtains a cropping frame for cropping the traffic surveillance video according to the vehicle size; performs abnormal target detection on the traffic surveillance video every first number of frames to detect abnormal target vehicles in the video frames of the traffic surveillance video; after detecting an abnormal target vehicle, crops a cropping picture of the abnormal target vehicle, and crops a cropping video segment with a spatial capacity several times the vehicle size and a time of the second number of frames forward or backward with the position of the abnormal target vehicle as the center in the traffic surveillance video for subsequent abnormal time estimation;
[0030] The abnormal time estimation module includes abnormal vehicle state detection and abnormal target matching; for the abnormal vehicle state detection, inputs the cropping picture of the abnormal target vehicle and the cropping video segment into a combined network model of siamese cross-correlation and P3D-Attention to detect whether the abnormal target in the cropping video segment is in a stationary state or a driving state, assigns classification labels to the cropping video segment according to the abnormal vehicle state detection results, marks them as abnormal, driving or normal respectively, and determines the video frame when the abnormal target vehicle has an abnormal behavior; for the abnormal target matching, inputs the vehicle picture to be matched and the cropping picture of the abnormal target vehicle into the combined network of siamese cross-correlation and P3D-Attention to determine whether the vehicle to be matched is an abnormal target vehicle, and combines the video frame when the abnormal target vehicle has an abnormal behavior determined in the abnormal target vehicle detection result to determine the start time and end time of the traffic abnormal behavior;
[0031] The display module displays the input video, image information and the output abnormal behavior detection position, abnormal time information and warning information.
[0032] The beneficial effects of the traffic abnormal behavior detection method and system of the present invention are as follows:
[0033] The traffic abnormal behavior detection method of the present invention is a two-stage road traffic abnormal event detection method based on the mechanism of combining siamese cross-correlation and P3D-Attention network. Siamese cross-correlation is mainly used to strengthen the attention to abnormal target vehicles, fully integrate spatio-temporal features and abnormal target features, focus on specific targets (such as abnormally stationary vehicles, as abnormal target vehicles), learn the abnormal target vehicle state detection method and image comparison method, so as to improve the speed and performance of detecting the state of abnormal target vehicles, solve the problem of traffic abnormal behavior time detection, and thus accurately estimate the start time of traffic abnormal behavior. The P3D-Attention module is based on decoupling spatial and temporal convolutions in the P3D module, and is respectively fused with an adaptive spatial attention module and a dual-channel attention module to fully integrate spatio-temporal features, improve the correlation of important channel features, increase the global correlation of feature maps, and improve the prediction performance of abnormal behavior vehicles. On the relevant data sets, the abnormalities in the same video are fused to obtain the final abnormal results. Experimental results show that the traffic abnormal behavior detection method of the present invention is effective in various traffic video scenarios, the F1 score of abnormal target detection is 0.9705, the root mean square error of abnormal time estimation is 9.22, and the average detection time for a 12-frame video is 12.2 ms, which consumes less compared to 342 ms using the tracking trajectory method and is more superior in terms of time efficiency. Description of the Drawings
[0034] Figure 1 It is a flowchart of Embodiment 1 of the present invention.
[0035] Figure 2 It is a schematic diagram of jitter perspective cropped video of the present invention. (a) is the video under the camera scene, (b) is the video under different jitter camera scenes randomly cropped from (a), and (c) is the target box generated according to (b).
[0036] Figure 3 It is a schematic diagram of the siamese cross-correlation and P3D-Attention combined network model of the present invention.
[0037] Figure 4 It is a schematic diagram of the depth cross-correlation operation of the present invention.
[0038] Figure 5 It is a schematic diagram of the principle of the depth cross-correlation mechanism of the present invention.
[0039] Figure 6 It is the adjusted dual-channel attention mechanism of the present invention. Detailed Embodiments
[0040] The present invention will be further described in detail below in conjunction with embodiments and with reference to the accompanying drawings.
[0041] Example 1:
[0042] An embodiment of the present invention is a traffic anomaly detection method, which is based on the mechanism of combining twin cross - correlation and P3D - Attention network and can be used to accurately estimate the start time and end time of road traffic anomaly behavior.
[0043] As Figure 1 shown, the input of the traffic anomaly detection method in this embodiment is a traffic surveillance video captured by a camera. The detection method includes:
[0044] I. Retain the abnormal target vehicle in the background through background modeling.
[0045] The abnormal target vehicle is the abnormally stationary target vehicle. Through background modeling, the normally moving vehicles in each frame of the traffic surveillance video are removed from the frame, so that the abnormally stationary target vehicle remains in the background.
[0046] The methods of background modeling generally include the moving average background modeling method and the MOG2 algorithm. The MOG2 algorithm is an adaptive algorithm based on Gaussian mixture probability density. The recursive equation is used to continuously update the parameters, and at the same time, an appropriate GMM component is selected for each pixel. In the background of small - target congestion and slow - speed roads, compared with the moving average background modeling method, the MOG2 algorithm has a worse effect. Moreover, in the congested traffic flow, the moving average background modeling method will retain more moving vehicle information in the modeled background, resulting in misdetection of stationary vehicles when subsequent abnormal target vehicle state detection methods are carried out. And in the scenario of camera shaking, the MOG2 algorithm is more stable than the moving average background modeling method.
[0047] Therefore, the present invention adopts the MOG2 algorithm to retain the abnormal target vehicle in the background. The MOG2 algorithm is an adaptive algorithm based on Gaussian mixture probability density. The recursive equation is used to continuously update the parameters, and at the same time, an appropriate GMM component is selected for each pixel. The GMM component is gradually updated over a long period of time and has better adaptability to different scenarios.
[0048] In this embodiment, in the surveillance video, the number of frames transmitted per second of the video is 30 frames. The update interval of the MOG2 algorithm is set to 120 frames, and the time required to detect the state of the abnormal target vehicle for these 120 - frame videos is 4 s. At this time, all normally moving vehicles are removed from the frame, and the static vehicles still remain in the background.
[0049] Through background modeling, the normally moving vehicles in each frame of the traffic surveillance video are removed from the frame, so that the abnormally stationary target vehicles are retained in the background. However, background modeling can cause problems with delayed vehicle appearance modeling. For example, forward background modeling can cause stationary vehicles to appear in the background some time after abnormal stillness occurs, resulting in an abnormal delay start time. To obtain an accurate abnormal start time, only the abnormal delay start time can be obtained through background modeling first, and then a more accurate time localization can be obtained by backtracking and analyzing the original images. For this purpose, the present invention also needs to perform abnormal target detection on the background to obtain the cropped pictures and cropped videos of the abnormal target vehicles; input them into the Siamese cross-correlation and P3D-Attention network model for abnormal time estimation to obtain accurate abnormal start and end times.
[0050] Second, perform abnormal target detection on the background to obtain the cropped pictures and cropped videos of the abnormal target vehicles.
[0051] Perform perspective cropping on each frame of the background extracted from the traffic surveillance video through background modeling, and obtain the cropping frame for cropping the surveillance video according to the vehicle size; for each frame of the traffic surveillance video, detect the background once every specified first number of frames (for example, 30 frames) using the abnormal target detection method to detect the abnormal target vehicles in the video frames; after detecting the abnormal target vehicles, crop a video segment with a spatial capacity of 2 times the vehicle size and a time of a set number of frames (for example, 12 frames) forward (or backward) centered on the position of the abnormal target vehicle for subsequent abnormal time estimation.
[0052] The abnormal target detection methods that can be adopted include general target detection methods such as YOLOv3 and Faster-RCNN. Preferably, the present invention adopts the YOLOv3 target detection method.
[0053] For videos, in addition to the problem of difficult detection of small targets, there is also the problem of camera jitter. In order to be able to well distinguish the abnormally stationary abnormal target vehicles from the starting moving vehicles, the present invention generates videos in different jitter camera scenarios in the form of random cropping as a simulated jitter dataset. The random cropping is carried out in two ways. One is to crop according to the magnitude of the jitter, and the other is to crop according to whether the jitter is random or back-and-forth tremor jitter. The magnitudes of the jitter are set to be half of the vehicle capacity and 1 vehicle capacity, etc., as Figure 2 shown. Preferably, considering that the data of the abnormal target vehicles cropped from the video is too small, in order to prevent overfitting caused by the learning bias of the combined network model of Siamese cross-correlation and P3D-Attention, a number of data samples (for example, more than 4000) of different vehicles are made through the target tracking algorithm.
[0054] III. Input the cropped images and cropped videos of the abnormal target vehicle into the Siamese cross-correlation and P3D-Attention network model for abnormal time estimation to obtain the accurate start time and end time of the abnormality.
[0055] For abnormal time estimation, the constructed Siamese cross-correlation and P3D-Attention combined network model is used to detect the abnormal vehicle state and match the abnormal target for the input images and input videos.
[0056] The abnormal vehicle state detection task mainly detects whether the abnormal target in the input video segment is in a stationary state or a driving state. According to the detection results of the abnormal vehicle state, classification labels are assigned to the video segments, marked as anomaly, driving, or normal respectively, and the video frame where the abnormal time point occurs is determined. The abnormal time point is the instant when the two states of driving and stationary change.
[0057] 1) Anomaly: The abnormal target vehicle stagnates on the road surface while the non-abnormal target vehicles are driving normally, and it is calibrated as abnormal at this time.
[0058] 2) Driving: The abnormal target vehicle is in a normal driving state.
[0059] 3) Normal: There is no abnormal target vehicle stagnating on the road surface, but it includes the state where other normal cars drive by and the road surface is empty.
[0060] The abnormal target vehicle matching task is used to assist the abnormal vehicle state detection. By image matching, it is determined whether the vehicle to be matched is an abnormal target vehicle, and combined with the video frame where the abnormal target vehicle has abnormal behavior determined in the detection results of the abnormal target vehicle, the start time and end time of the traffic abnormal behavior are thus determined.
[0061] Corresponding to the model combined with the Siamese cross-correlation and P3D-Attention network proposed in the present invention, there are three inputs: one is the cropped image of the abnormal target vehicle, the second is the video segment with a spatial capacity twice the size of the vehicle and a set first number of frames (such as 12 frames) cropped centered on the abnormal target vehicle, and the third is the positive and negative class images for abnormal target matching.
[0062] In a neural network, the convolution operation is essentially a cross-correlation operation, which has the characteristic of weight sharing, that is, all nodes in the same layer based on the convolutional neural network share the same connection weights.
[0063] The twin cross - correlation and P3D - Attention combined network model of the present invention is based on the P3D - Attention network structure as the main body, and the cross - correlation mechanism is integrated into the P3D - Attention network structure, such as Figure 3 shown. In the twin cross - correlation structure, features of pictures and videos are extracted by performing depth cross - correlation operations; the feature map is the feature extracted after the convolutional kernel operation of the twin cross - correlation in the twin cross - correlation and P3D - Attention combined network on the input picture or input video. In this embodiment, the inputs of the twin cross - correlation and P3D - Attention combined network model are two pictures and a video, and all three extract features through the P3D - Attention backbone network, that is, the backbone networks of the three inputs share weights.
[0064] For the input of the abnormal vehicle state detection task, one is the cropped picture of the abnormal target vehicle (such as ), and the other is a video clip with a spatial capacity 2 times the size of the vehicle centered on the abnormal target vehicle and a time of the set second number of frames (such as 12 frames, ). The abnormal vehicle state detection is performed on the video containing the abnormal target in the abnormal target detection result every specified first number of frames (such as 30 frames). First, the P3D - Attention module extracts the feature map of the input template picture and the feature map of each frame of the input video respectively, and the P3D - Attention is used to improve the correlation of important channel features; then the respectively extracted feature maps are selected to perform the twin cross - correlation operation at three different receptive field sizes (such as Figure 3The mutual correlations 1, 2, and 3 are fused. Preferably, the twin mutual correlation operation adopts the mutual correlation operation of the SiamFC model. The results of the first two twin mutual correlation operations use the multiply method to fuse the feature maps extracted from the input image and the input video after the twin mutual correlation operation with the video feature maps extracted by the P3D-Attention module before the mutual correlation; and use GAP (Global Average Pooling) to replace the fully connected layer for pooling; finally, through the softmax layer, directly classify the cropped images of the input abnormal target vehicles and the input video segments. Using the multiply method to fuse the feature maps extracted from the input image and the input video after the twin mutual correlation operation with the video feature maps extracted by the P3D-Attention module before the mutual correlation is similar to the spatial attention mechanism. The difference is that the spatial attention mechanism learns the parts that need to be focused on through gradient descent fitting of the network, while the input of the part to be focused on in the twin mutual correlation mechanism is the feature map extracted by the P3D-Attention module from the pre-set cropped images of abnormal target vehicles. The extracted feature map is used as the convolution kernel input for the depth mutual correlation operation to strengthen the features that match the cropped images of abnormal target vehicles with the input video. The important positions are directly strengthened through the combination network of twin mutual correlation and P3D-Attention, such as Figure 5 as shown. GAP is simpler and more natural in converting between feature maps and final classification. Unlike the fully connected layer that requires a large number of parameters for training and tuning, reducing the spatial parameters makes the model more robust and has a better anti-overfitting effect. The network using GAP to replace the fully connected layer usually still has good prediction performance, can also prevent the model from overfitting, and can greatly reduce the computational amount of model training.
[0065] The inputs for the abnormal target vehicle matching task are, one is the cropped image of the abnormal target vehicle (for example ), and the other is the positive and negative class images for the abnormal target matching model (for example ). First, the P3D-Attention module is used to extract the feature maps of the two input images respectively, and the twin mutual correlation operation is performed using the two extracted feature maps. The feature maps extracted separately are still selected to perform the twin mutual correlation operation at three different receptive field sizes (such as Figure 3 the mutual correlations 4, 5, and 6 in) for fusion. The twin mutual correlation operation uses the VALID type, that is, no padding is added for the mutual correlation operation to obtain , and Feature maps of a certain size are obtained and concatenated. The concatenation operation stitches tensors along a specified dimension. Finally, through the softmax layer, the cropped images of the abnormal target vehicles and the positive and negative class images for the abnormal target matching model are directly classified, and the classification result is matching or non-matching.
[0066] 1) Matching: Each pair of matching images is a cropped image of the same vehicle at different times.
[0067] 2) Non-matching: The template image is a cropped image of the abnormal target vehicle, and the non-matching images can be another vehicle in the positive and negative class images for the abnormal target matching model or any image that can appear in the traffic road scene. Considering that this task is to clarify the vehicle's stationary time, if the IOU of the same vehicle is less than 0.7, it will be regarded as non-matching.
[0068] In the second task, instead of using a distance function to calculate the similarity between images, the depthwise cross-correlation operation identical to SiamFC is used. The depthwise cross-correlation operation is performed on feature maps of the same shape and size, and the connection (concatenate) method is used to aggregate features of different scales, and then classification is carried out.
[0069] After processing all inputs to obtain the start time and end time of all abnormalities, anomaly fusion is performed. That is, abnormalities with intersecting abnormal times are regarded as the same abnormality; for example, abnormalities with an abnormal time interval not exceeding 5 seconds and an IOU greater than 0.5 are regarded as the same abnormality.
[0070] Each convolution kernel of the conventional cross-correlation operation operates on each channel of the input image simultaneously. The present invention chooses to use the depthwise cross-correlation operation. One convolution kernel of the depthwise cross-correlation operation is only responsible for one channel, and one channel only performs cross-correlation operation with the corresponding convolution kernel to obtain an output with a non-1 feature channel number to retain more features. The operation description is as Figure 4 shown, where is the depthwise cross-correlation operation.
[0071] Construct the Siamese cross-correlation and P3D-Attention combined network model as shown in Figure 3 specifically including the following steps.
[0072] Step 3.1, adjustment of the P3D-Attention network for two types of inputs.
[0073] The backbone network for feature extraction in this invention is the P3D-Attention network. When implementing the fusion of the siamese cross-correlation and the P3D-Attention network, the 2D convolution operation needs to be replaced with a 3D convolution operation. The P3D-Attention network first simulates the 3×3×3 convolution in the spatial domain and the temporal domain respectively using 1×3×3 and 3×1×1 convolution kernels, decoupling the 3×3×3 convolution in time and space. Secondly, the P3D-Attention network also includes a dual-channel attention mechanism and a spatial attention mechanism to improve the correlation of important features.
[0074] The two inputs of the combined network model of siamese cross-correlation and P3D-Attention are images and videos respectively. The images include the cropped images of abnormal target vehicles and the positive and negative class images for the abnormal target matching model, such as ( , that is ) images; the video is a video clip with a spatial capacity 2 times the size of the vehicle and a time of the second frame number cropped centered on the abnormal target vehicle, such as ( , that is ) video. Both inputs extract features through the P3D-Attention backbone network, that is, the backbone networks of the two inputs for abnormal vehicle state detection and the two inputs for abnormal target matching share weights. When the two inputs pass through the same network simultaneously, not only are the image sizes different (80:40), but the input dimensions are also different (4:3), and the network structure needs to be adjusted when combining the attention mechanisms.
[0075] First, it is necessary to adjust the image (such as image) to data with an added time dimension (such as ) to adapt to the convolution operation on the video dimension; however, there is still a situation where the sizes are different in the time dimension (12:1). For this reason, the implementation of the combined network model structure of siamese cross-correlation and P3D-Attention needs to be adjusted in the P3D module and the Attention module respectively.
[0076] For the P3D module, the difference in image size is not a problem, and both inputs can perform convolution operations on the 1×3×3 convolution kernel; for the difference in time dimension size, when performing the image input , the convolution operation of the 3×1×1 convolution kernel needs to be ignored because for an image with a time dimension of only 1, a convolution size of 3 is meaningless.
[0077] For the Attention module, the dual-channel attention module and the spatial attention module are respectively used to apply attention on the time frame and the spatial frame to strengthen the key frames. Different attention weights are automatically assigned to different joints according to the feature map, focusing on the positions mentioned in the prior knowledge, and removing the interference of the background and noise on the recognition. For the specific steps, please refer to Step 3.2. The attention mechanism is expressed as:
[0078]
[0079] In the formula, M represents the attention module, and F represents the feature map. represents the element-wise multiplication of matrices.
[0080] Step 3.2: Adjustment of the dual-channel attention module and the spatial attention module.
[0081] The P3D-Attention network, which combines the pseudo 3D convolutional neural network and the attention mechanism, uses the P3D-Attention module to apply attention to the channels and the feature map by using the attention mechanism, including the dual-channel attention module and the spatial attention module.
[0082] The dual-channel attention module applies attention between video frames and on the channels of each frame. The feature map , where F in R represents the frame, C represents the number of channels in each frame, H and W represent the features in different channels, and learn the weights to determine the importance of each channel, and transpose the feature map to , and embed it into the 2D spatial attention module to learn and the weights to express the attention to the frame and the channel respectively.
[0083] The spatial attention module is implemented by using the CBAM (Convolutional Block Attention Module). It focuses on the spatial features by learning the weight map in the spatial dimension. Taking the feature map of each frame of the video as an example, where F represents the frame, C represents the number of channels of the feature map, H represents the height of the feature map, and W represents the width of the feature map. The spatial attention module mainly learns the position information in the single-channel feature map weight matrix to determine the importance and relevance of each position in the video feature map. The size of the convolutional kernel is not related to the size of the feature map, so both video features and image features can be directly cascaded with the spatial attention module to improve the relevance of important positions.
[0084] Based on the decoupling of 3D convolution into spatial and temporal convolutions in the P3D module, the attention module is divided into the following three different P3D-Attention modules to implement the network model.
[0085] P3D-Attention-A:
[0086] 1D temporal convolution kernel Cascaded to the 2D spatial convolution kernel , by cascading a spatial attention module after and cascading a channel attention module after , the P3D-Attention-A structure is implemented. The 1D temporal convolution kernel
[0087]
[0088] wherein, represents the input feature map, represents the output after applying the attention mechanism, and have the same feature dimension.
[0089] P3D-Attention-B:
[0090] The original P3D-B uses the indirect influence between two convolution kernels, enabling the two convolution kernels to process convolution features in parallel; on the basis of removing the residual unit, after cascading a spatial attention module ( ) and then cascading a channel attention module ( ) at the
[0091]
[0092] P3D-Attention-C:
[0093] The P3D-C module is a compromise between P3D-a and P3D-B, and at the same time establishes the direct influence between , and the final output; in order to achieve the direct connection between and the final output based on the cascaded P3D-A architecture, the P3D-Attention-C is constructed by adding an attention module, expressed as:
[0094]
[0095] To adapt to the integration of the Siamese network and the attention mechanism, corresponding adjustments need to be made to different branches of the input. The adjusted dual-channel attention mechanism is as follows Figure 6 shown. Taking the feature map as an example, where F represents the number of frames, C represents the number of channels of the feature map, H represents the height of the feature map, and W represents the width of the feature map. In order to adapt to 3D convolution, the present invention constructs a dual-channel attention module (dual-channel attention model), combines a 1×3×3 convolution kernel in the spatial domain and a 3×1×1 convolution kernel in the temporal domain into a 3×3×3 convolution, and learns the frame attention module weights to express the attention to frames, and learns the channel attention module weights to express the attention to channels.
[0096] Regarding the attention to frames, that is, learning , it is necessary to transpose the feature map of to ( ), and eliminate the spatial dimensions H and W through max pooling and average pooling. Max pooling can extract feature textures and reduce the influence of useless information. Average pooling can retain background information. Then the dimensions of the two inputs (pictures and videos) after pooling will be different, which are and respectively. Therefore, when executing the frame attention mechanism for picture input, the present invention ignores the frame attention operation.
[0097] Regarding the attention to the channels of each frame, that is, learning , the F in the temporal dimension can be directly pooled and eliminated. The dimensions of the two inputs of pictures and videos after pooling are the same (C:C). Therefore, when executing the picture channel attention mechanism for picture input, the present invention ignores the picture channel attention operation.
[0098] For the traffic anomaly behavior detection method of the present invention, the RMSE (root mean square error) of the anomaly time estimation is 9.22, and the average detection time for 12-frame videos is 12 ms, which has great superiority in terms of time efficiency.
[0099] The traffic abnormal behavior detection method of the present invention estimates the abnormal time by combining abnormal vehicle state detection and abnormal target matching, and adopts multi-task training. Two input images and a video are respectively input into one branch of the combined network of siamese cross-correlation and P3D-Attention. The three branches have the same structure and share the same parameters. Two Softmax layers are used at the end of the network as the results of two classification tasks. The combined network model of siamese cross-correlation and P3D-Attention performs multi-task learning on the abnormal vehicle state detection method and the image matching method based on the siamese network simultaneously. By predicting the position state of the vehicle and using the abnormal target matching method to further accurately determine the frame where the abnormal start time is located, it can well solve the problem of estimating the start time of traffic anomalies and accurately estimate the start time of abnormal behaviors.
[0100] Furthermore, the entire combined network of siamese cross-correlation and P3D-Attention is a convolutional neural network architecture. In order to reduce the misjudgment of the model during training, the present invention uses F1-Score instead of accuracy to evaluate the performance of the model. F1-Score, also known as balanced F Score, is defined as the harmonic mean of precision P and recall R.
[0101] Among them, P (Precision) refers to precision.
[0102]
[0103] In the formula, TP represents the positive sample, that is, the classifier determines it as a positive example and the determination is correct; FP represents that the classifier determines it as a positive example, but the determination is wrong.
[0104] R (Recall) refers to recall.
[0105]
[0106] In the formula, FN represents that the classifier determines it as a negative example, but the determination is wrong.
[0107] The F1 score is the harmonic mean of precision and recall.
[0108]
[0109] Embodiment 2:
[0110] Another embodiment of the present invention is a traffic abnormal behavior detection system, including:
[0111] Video acquisition and clipping module, which acquires real-time traffic video streams to provide continuous traffic video stream information, displays it on the display module, and transmits it as input information to the background modeling module;
[0112] The background modeling module reserves an interface. By inputting video stream data with the correct format and through background modeling, it removes the normally moving vehicles in each frame of the traffic surveillance video from the frame, so that the abnormally stationary target vehicles remain in the background;
[0113] The abnormal target detection module performs perspective clipping on each frame of the background extracted through background modeling, and obtains a clipping frame for clipping the traffic surveillance video according to the vehicle size; it performs abnormal target detection on the traffic surveillance video every first number of frames, and detects abnormal target vehicles in the video frames of the traffic surveillance video; after detecting an abnormal target vehicle, it clips out a clipping picture of the abnormal target vehicle, and takes the position of the abnormal target vehicle as the center of the traffic surveillance video, and clips out a clipping video segment with a spatial capacity of 2 times the vehicle size and a time of the second number of frames forward or backward for subsequent abnormal time estimation;
[0114] The abnormal time estimation module includes abnormal vehicle state detection and abnormal target matching; for the abnormal vehicle state detection, it inputs the clipping picture of the abnormal target vehicle and the clipping video segment into the combined network model of siamese cross-correlation and P3D-Attention to detect whether the abnormal target in the clipping video segment is in a stationary state or a driving state, and respectively assigns classification labels to the clipping video segment according to the abnormal vehicle state detection results, and marks them as abnormal, driving or normal respectively, and determines the video frame when the abnormal target vehicle has an abnormal behavior; for the abnormal target matching, it inputs the vehicle picture to be matched and the clipping picture of the abnormal target vehicle into the combined network of siamese cross-correlation and P3D-Attention to determine whether the vehicle to be matched is an abnormal target vehicle, and combines the video frame when the abnormal target vehicle has an abnormal behavior determined in the abnormal target vehicle detection result to determine the start time and end time of the traffic abnormal behavior;
[0115] The display module displays the input video, image information, and the output abnormal behavior detection position, abnormal time information, and warning information.
[0116] In some embodiments, certain aspects of the above technologies may be implemented by one or more processors of a processing system executing software. The software includes one or more sets of executable instructions stored or otherwise tangibly embodied on a non-transitory computer-readable storage medium. The software may include instructions and certain data that, when executed by one or more processors, manipulate the one or more processors to perform one or more aspects of the above technologies. The non-transitory computer-readable storage medium may include, for example, magnetic or optical storage devices, solid-state storage devices such as flash memory, cache, random access memory (RAM), or other non-volatile memory devices. The executable instructions stored on the non-transitory computer-readable storage medium may be in source code, assembly language code, object code, or other instruction formats interpreted or otherwise executed by one or more processors.
[0117] A computer-readable storage medium may include any storage medium or combination of storage media that can be accessed by a computer system during use to provide instructions and / or data to the computer system. Such storage media may include, but are not limited to, optical media (e.g., compact discs (CDs), digital versatile discs (DVDs), Blu-ray discs), magnetic media (e.g., floppy disks, magnetic tapes, or magnetic hard disk drives), volatile memory (e.g., random access memory (RAM) or cache), non-volatile memory (e.g., read-only memory (ROM) or flash memory), or microelectromechanical systems (MEMS)-based storage media. The computer-readable storage medium may be embedded in the computing system (e.g., system RAM or ROM), fixedly attached to the computing system (e.g., magnetic hard disk drive), removably attached to the computing system (e.g., optical disc or universal serial bus (USB)-based flash memory), or coupled to the computing system via a wired or wireless network (e.g., network-attached storage (NAS)).
[0118] Note that not all activities or elements in the above general description are required, some of a particular activity or device may not be required, and one or more further activities or elements may be performed or included in addition to those described. Further, the order in which the activities are listed need not be the order in which they are performed. Also, these concepts have been described with reference to specific embodiments. However, those of ordinary skill in the art recognize that various modifications and changes can be made without departing from the scope of the disclosure set forth in the claims. Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive, and all such modifications are included within the scope of the disclosure.
[0119] Benefits, other advantages, and solutions to problems have been described above with respect to specific embodiments. However, no benefit, advantage, solution to problems, or any feature that may cause any benefit, advantage, or solution to occur or become more pronounced should be construed as a critical, required, or essential feature of any or all claims in any respect. Furthermore, the specific embodiments disclosed above are merely illustrative, as the disclosed subject matter may be modified and implemented in different but equivalent manners obvious to those skilled in the art that benefit from the teachings herein. Except as described in the following claims, there is no intention to limit the details of the construction or design shown herein. It is, therefore, evident that the specific embodiments disclosed above may be altered or modified, and all such variations are considered to be within the scope of the disclosed subject matter.
Claims
1. A traffic abnormal behavior detection method, characterized in that including; By background modeling, the vehicles moving normally in each frame of the traffic surveillance video are removed from the frame, so that the abnormally stationary target vehicle remains in the background; Perform perspective cropping on each frame of the background extracted by background modeling, and obtain a cropping frame for cropping the traffic surveillance video according to the vehicle size; Perform abnormal target detection on the traffic surveillance video every first number of frames to detect abnormal target vehicles in the video frames of the traffic surveillance video; After detecting an abnormal target vehicle, crop a cropped picture of the abnormal target vehicle, and crop a cropped video segment with a spatial capacity of several times the vehicle size and a time of the second number of frames forward or backward centered on the position of the abnormal target vehicle in the traffic surveillance video for subsequent abnormal time estimation; Estimate the abnormal start time according to the abnormal target vehicle detection result, including abnormal vehicle state detection and abnormal target matching; The abnormal vehicle state detection inputs the cropped picture of the abnormal target vehicle and the cropped video segment into a combined network model of siamese cross-correlation and P3D-Attention to detect whether the abnormal target in the cropped video segment is in a stationary state or a driving state. According to the abnormal vehicle state detection result, classify the cropped video segment and label it as abnormal, driving or normal respectively, and determine the video frame when the abnormal target vehicle has an abnormal behavior; The inputs of the abnormal vehicle state detection are, one is the cropped picture of the abnormal target vehicle, and the other is a video segment with a spatial capacity of 2 times the vehicle size and a time of the set second number of frames cropped centered on the abnormal target vehicle; The abnormal vehicle state detection includes: respectively extracting the feature maps of the input cropped picture of the abnormal target vehicle and the feature maps of each frame of the input video segment through the P3D-Attention module, and improving the correlation of the important channel features through P3D-Attention; then select the respectively extracted feature maps to be fused through siamese cross-correlation operations in three layers with different receptive field sizes; the results of the first two siamese cross-correlation operations use the multiply method to fuse the feature maps extracted from the input picture and the input video through the siamese cross-correlation operation with the feature maps of each frame of the input video segment extracted by the P3D-Attention module before cross-correlation; use GAP for pooling; finally, directly classify the input cropped picture of the abnormal target vehicle and the input video segment through the softmax layer, and classify the input video segment according to the abnormal vehicle state detection result and label it as abnormal, driving or normal respectively, and determine the video frame where the abnormal time point occurs; For the abnormal target matching, the cropped image of the vehicle to be matched and the cropped image of the abnormal target vehicle are input into the combined network of twin cross-correlation and P3D-Attention to determine whether the vehicle to be matched is an abnormal target vehicle. Combining the video frames when the abnormal target vehicle has abnormal behaviors determined in the abnormal target vehicle detection result, the start time and end time of the traffic abnormal behavior are determined; The inputs for the abnormal target vehicle matching are, one is the cropped image of the abnormal target vehicle, and the other is the positive and negative class images for the abnormal target matching model; The abnormal target vehicle matching includes: respectively extracting the feature maps of the two input images through the P3D-Attention module, and performing twin cross-correlation operations using the extracted feature maps of the two inputs; Selecting the respectively extracted feature maps to be fused through twin cross-correlation operations at three different receptive field sizes; Obtaining three feature maps of different sizes and performing concatenate merge operations; Finally, through the softmax layer, directly classifying the cropped image of the input abnormal target vehicle and the positive and negative class images for the abnormal target matching model, and the classification result is matching or not matching; The P3D-Attention module simulates 3×3×3 convolution in the spatial domain and temporal domain by 1×3×3 convolution kernels and 3×1×1 convolution kernels respectively, and decouples 3×3×3 convolution in time and space.
2. The traffic anomaly behavior detection method according to claim 1, wherein The background modeling adopts the MOG2 algorithm.
3. The traffic abnormal behavior detection method according to claim 1, characterized in that The abnormal target detection adopts the YOLOv3 target detection method.
4. The traffic abnormal behavior detection method according to claim 1, characterized in that The P3D-Attention module also includes a dual-channel attention module and a spatial attention module for improving the correlation of important features.
5. The traffic abnormal behavior detection method according to claim 4, wherein The dual-channel attention module combines a 1×3×3 convolutional kernel in the spatial domain and a 3×1×1 convolutional kernel in the temporal domain into a 3×3×3 convolution to learn the weights of the frame attention module to express attention to frames, and learn the weights of the channel attention module to express attention to channels; where F represents the number of frames of the feature map , C represents the number of channels of the feature map, H represents the height of the feature map, and W represents the width of the feature map.
6. The traffic anomaly behavior detection method according to claim 5, wherein The spatial attention module learns the weight matrix of the feature map of a single channel through a 2D convolutional kernel to determine the importance and relevance of each position in the video feature map; where F represents the number of frames of the feature map , C represents the number of channels of the feature map, H represents the height of the feature map, and W represents the width of the feature map.
7. A traffic abnormal behavior detection system, which implements the traffic abnormal behavior detection method according to any one of claims 1 to 6, characterized in that, It includes: A video acquisition and shear module, which acquires the real-time traffic video stream, is used to provide continuous traffic video stream information, display it on the display module, and transmit it as input information to the background modeling module; The background modeling module reserves an interface. By inputting video stream data with the correct format, through background modeling, the normally moving vehicles in each frame of the traffic monitoring video are removed from the frame, so that the abnormally stationary target vehicles are retained in the background; An abnormal target detection module, which performs perspective cropping on each frame of the background extracted through background modeling, and obtains a cropping frame for cropping the traffic monitoring video according to the vehicle size; Perform abnormal target detection on the traffic monitoring video every first number of frames to detect abnormal target vehicles in the video frames of the traffic monitoring video; After detecting an abnormal target vehicle, crop out the cropped image of the abnormal target vehicle, and crop out a cropped video segment with a spatial capacity of several times the vehicle size and a time of the second number of frames forward or backward centered on the position of the abnormal target vehicle in the traffic monitoring video for subsequent abnormal time estimation; Anomaly time estimation module, including anomaly vehicle state detection and anomaly target matching; for the anomaly vehicle state detection, the cropped image of the anomaly target vehicle and the cropped video segment are input into the combined network model of siamese cross-correlation and P3D-Attention to detect whether the anomaly target in the cropped video segment is in a stationary state or a driving state. According to the anomaly vehicle state detection results, classification labels are respectively assigned to the cropped video segment, marked as anomaly, driving or normal, and the video frame when the anomaly target vehicle has an abnormal behavior is determined; for the anomaly target matching, the vehicle image to be matched and the cropped image of the anomaly target vehicle are input into the combined network of siamese cross-correlation and P3D-Attention to determine whether the vehicle to be matched is an anomaly target vehicle. Combining the video frame when the anomaly target vehicle has an abnormal behavior determined in the anomaly target vehicle detection results, the start time and end time of the traffic anomaly behavior are determined. Display module, which displays the input video, image information, the detected position of the abnormal behavior, the abnormal time information and the warning information.