A crowd abnormal behavior detection method based on improved SSD

By improving the SSD network, using the lightweight network MobileNet v2, deformable convolution module and coordinate attention mechanism, the high computational complexity and occlusion problems in the detection of population abnormal behavior are solved, and the rapid and accurate detection effect is achieved.

CN115273234BActive Publication Date: 2025-07-18HARBIN LIHAI JIAYUAN TECH DEV CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210886980.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-26
Publication Date
2025-07-18
Estimated Expiration
2042-07-26

AI Technical Summary

Technical Problem

Existing deep learning methods have high computational complexity in the detection of abnormal behavior of populations, resulting in slow operation speed and low detection accuracy in complex scenarios such as occlusion conditions.

Method used

The lightweight network MobileNet v2 is used to replace the feature extraction network of the SSD model, and feature extraction is enhanced through the deformable convolution module and coordinate attention mechanism, to capture the remote dependence between spatial locations and improve the occlusion problem.

Benefits of technology

It improves the accuracy and speed of abnormal behavior detection for crowds, can accurately detect abnormal behavior in complex scenarios, and supports public safety monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115273234B_ABST
    Figure CN115273234B_ABST
Patent Text Reader

Abstract

A method for detecting abnormal human behaviors based on improved SSD, which trains and evaluates the SSD network model on a preprocessed dataset of abnormal human behaviors. It improves the problems existing in the SSD network model, including poor real-time performance of the model due to a large number of parameters and low detection accuracy caused by the inability to detect abnormal behaviors with partial occlusion. The lightweight network MobileNetv2 is used as the feature extraction network of the SSD model, and a deformable convolution module is embedded to construct the convolutional layer to enhance the receptive field. On this basis, the output feature map is enhanced through the coordinate attention mechanism. By learning the context relationship, it can capture the long-range dependence between spatial positions, and predict the occluded part based on the unoccluded part to effectively improve the occlusion problem. The abnormal behavior datasets targeted by the present invention all occur in a natural state, can accurately detect the categories and positions of abnormal human behaviors, and at the same time, the detection of the model for partial occlusion situations existing in the dataset scene is improved, having good application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of crowd abnormal behavior detection, and particularly relates to a crowd abnormal behavior detection method based on improved SSD. Background Art

[0002] Abnormal behavior detection, as a popular research direction in the fields of machine vision and image processing, has attracted much attention from researchers. The frequent occurrence of public security incidents has posed a serious threat to the personal safety of the people. As the most intuitive form of expression of the objective world, video resources play an important role in maintaining social public security. Traditional detection methods generally first segment the target to be detected from the video sequence, then extract features and compare the extracted crowd behavior features with the abnormal behavior samples in the standard library, and finally hand them over to the classifier to judge whether there is an abnormal behavior. However, if the data volume is large, this method shows problems such as insufficient computing power and inability to express deep features, making the monitoring device unable to alarm the abnormal behavior of the scene in time. If an intelligent monitoring system is used to monitor abnormal situations in real time and alarm for abnormal situations, this can reduce the impact of public security incidents on society.

[0003] With the vigorous development of research related to deep learning, researchers have begun to explore crowd abnormal behavior detection based on deep learning, and the deep learning method can solve problems more efficiently. Hu Xuemin et al. proposed an algorithm for detecting group abnormal events based on a deep spatio-temporal convolutional neural network, which uses the spatial features of each frame of video and the temporal features of the front and back frames, extends the two-dimensional convolutional operation to three-dimensional space, divides the video region into several sub-regions to obtain their spatial features, and finally inputs the spatial features into the deep spatio-temporal convolutional network for training and classification. Almazroey et al. proposed a deep learning-based algorithm to detect the abnormal behavior of people in surveillance videos. This algorithm uses the size, direction, and speed features of the optical flow of the key frames extracted from the video to generate multiple 2D model features, and finally inputs the 2D model features into the pre-trained AlexNet model for judgment. Mu Yonglin et al. proposed an algorithm for detecting crowd abnormal events based on Generative Adversarial Networks (GANs). This algorithm uses normal event samples to train a pair of generative adversarial networks, takes one of the generative adversarial networks as the input and generates the corresponding optical flow features, then inputs the optical flow features into the other generative adversarial network and generates the corresponding frames, and finally analyzes the difference between the generated frame images and the real frames to detect and locate abnormal events.

[0004] In some scenarios, the abnormal behavior characteristics of the crowd contained in the monitoring screen are usually affected by the complex background environment, crowding, occlusion, etc. These factors will greatly reduce the accuracy and detection speed of the abnormal behavior detection algorithm. Although the currently used deep learning methods have made good progress, most of the algorithms are highly complex and will consume a lot of computing resources in real scenarios, reducing the network operation speed, and cannot ensure the accuracy of abnormal behavior detection in complex scenarios (such as occlusion). Summary of the invention

[0005] In order to overcome the defects of the above-mentioned prior art, the purpose of the present invention is to provide a crowd abnormal behavior detection method based on improved SSD. Based on the standard SSD network, the method replaces the original feature extraction network VGG-16 with a lightweight network MobileNet v2, and constructs a convolution layer through a deformable convolution module to enhance the receptive field. Then, feature enhancement is performed by integrating position information into channel attention, which can capture the long-range dependency between spatial positions, thereby improving the technical problem of overlapping occlusion and has broad application prospects.

[0006] In order to achieve the above object, the technical solution adopted by the present invention is:

[0007] A method for detecting abnormal crowd behavior based on improved SSD specifically includes the following steps:

[0008] Step 1: Data preprocessing: including video frame serialization, image annotation, and data set division

[0009] 1) Video frame serialization: Through the multimedia operation open source program integrated in Python, the video can be cropped and intercepted;

[0010] 2) Image annotation: Use image annotation tools to annotate the abnormal behavior detection dataset, frame the abnormal behavior in each image with a rectangular frame, and indicate the category to which it belongs;

[0011] 3) Divide the data set: Divide the data set expanded in step 2) into a training set and a test set, and compress the original image to the default size of the SSD network;

[0012] Step 2: Train and evaluate the SSD network model

[0013] The image processed in step 3) of step 1 is used as the input image of the SSD network model, and the operating parameters of the SSD network model are set. The SSD network model is trained on the experimental operation platform, and then the effect of the trained SSD network model is evaluated using the evaluation indicators commonly used in the field of abnormal behavior detection;

[0014] Step 3. Improve the SSD network for the evaluation in Step 2

[0015] 1) Replace the feature extraction network

[0016] Replace the feature extraction network of the SSD network with the lightweight network MobileNet v2 network to reduce the scale of network model parameters;

[0017] 2) Embed the deformable convolution module

[0018] By adding direction vectors, change the fixed size and shape of the traditional convolution kernel so that it can adaptively adjust its own shape to cope with target objects of different scales and deformations, thereby better extracting input features;

[0019] 3) Design the attention mechanism module

[0020] The SSD network model detects target objects by extracting six different-scale feature maps default in the SSD network. The feature maps contain feature channels and position information. The content in the map contributes differently to the result of the target detection task, that is, the saliency is different. Using the coordinate attention module, through learning, suppress the insignificant features, enhance the expression ability of the features in the network, obtain the context relationship through learning, can capture the long-range dependence relationship between spatial positions, and predict the abnormal behavior of the occluded part according to the abnormal behavior of the unoccluded part in the crowd, thereby improving the target detection effect;

[0021] Step 4. Train the improved SSD network

[0022] Use the images processed in step 3) of step 1 as the input images of the improved SSD network in step 3, set the running parameters of the network model, train the improved SSD network model on the experimental operation platform, and use the evaluation indicators commonly used in the field of abnormal behavior detection to evaluate the detection results of the improved SSD, and output the final abnormal behavior detection results.

[0023] The specific method of step 2) in step 3 is as follows:

[0024] 2.1) The input image passes through a common convolution to obtain the input feature map;

[0025] 2.2) Convolve the upper half of the input feature map to obtain the offset ΔP n ;

[0026] 2.3) Use bilinear interpolation to represent the offset position;

[0027] 2.4) Add the obtained offset to the input feature map to get a new sampling position, and then use the convolution kernel to extract features to obtain the output feature map;

[0028] Definition of the convolutional kernel:

[0029] R = {(-1, -1), (-1, 0),...(0, 1), (1, 1)} (1)

[0030] Among them, R defines the size and dilation of the receptive field;

[0031] The output of ordinary convolution is:

[0032]

[0033] Among them, X is the input, Y is the output, W is the weight matrix, P0 is each point on the feature map, and P n is n points in the grid;

[0034] The output of deformable convolution is:

[0035]

[0036] Among them, ΔP n is the coordinate offset.

[0037] The specific method of step 3) in step three is as follows:

[0038] The coordinate attention module decomposes the channel attention into two parallel one-dimensional feature encoding processes, and effectively integrates the spatial coordinate information into the generated attention map; that is, average pooling is performed on the X horizontal and Y vertical directions to obtain two one-dimensional vectors. Next, concatenation in the spatial dimension and 1×1 convolution are used to compress the channels, and then the spatial information in the two directions is encoded through the BN layer and ReLU and split. Then, convolutions are respectively performed to obtain the same number of channels as the input feature map, normalized and weighted, and finally multiplied by the original feature map to adaptively adjust the features;

[0039] The coordinate attention module is placed after the convolutional layer and before the BN layer, and feature enhancement is performed by integrating the position information into the channel attention, and the output feature maps extracted by MobileNet v2 are enhanced by the coordinate attention module respectively.

[0040] The running parameters of the network model in step two and step four include the initial learning rate, learning momentum, weight decay rate, and the network parameters are updated using the stochastic gradient descent method SGD.

[0041] The commonly used evaluation metrics in the field of abnormal behavior detection in Step 2 and Step 4 are AUC (Area Under Curve); the Receiver Operating Characteristic is (Receiver Operating Characteristic, ROC); the horizontal and vertical coordinates of the ROC curve represent the False Positive Rate (FPR) and the True Positive Rate (TPR) respectively; where: TPR represents the ratio of samples that are correctly judged as positive among all true positive samples; FPR represents the ratio of samples that are wrongly judged as positive among all true negative samples;

[0042] Taking FPR as the abscissa and TPR as the ordinate, its calculation formula is as follows:

[0043]

[0044]

[0045] In the formula, represents being judged as positive and actually being positive; represents being judged as positive and actually being negative; represents being judged as negative and actually being negative; represents being judged as negative and actually being positive.

[0046] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0047] 1. Aiming at the deficiencies of the existing model, the present invention first uses the lightweight network MobileNet v2 as the feature extraction network of the SSD model, reduces the number of model parameters, and improves the model operation speed;

[0048] 2. Aiming at the problem of low detection accuracy in the task of crowd abnormal behavior detection, the present invention constructs a convolutional layer through a deformable convolution module to enhance the receptive field and can better extract input features;

[0049] 3. The present invention uses a coordinate attention module to enhance the output feature map extracted by MobileNet v2. By learning the context relationship and predicting the occluded part according to the unoccluded part, the occlusion problem can be effectively improved.

[0050] In summary, based on the standard SSD model, the present invention changes the model feature extraction network, embeds a deformable convolution module and adds a coordinate attention mechanism. Compared with the existing crowd abnormal behavior detection algorithms, the present invention improves the crowd abnormal behavior detection accuracy, improves the detection speed, has a good detection effect, proves that the model in this article can achieve fast and accurate detection of crowd abnormal behaviors, and can provide support for public safety monitoring. Description of the Drawings

[0051] Figure 1This is a diagram of the abnormal behavior detection model of a crowd based on improved SSD of the present invention.

[0052] Figure 2 It is a LabelImg annotation diagram in an embodiment of the present invention.

[0053] Figure 3 This is a diagram of the deformable convolution module of the present invention.

[0054] Figure 4 This is the coordinate attention module diagram of the present invention.

[0055] Figure 5 It is a structural diagram of the feature enhancement of the present invention. DETAILED DESCRIPTION

[0056] The present invention is further described in detail below with reference to the accompanying drawings and embodiments:

[0057] See also Figure 1 , a crowd abnormal behavior detection method based on improved SSD, specifically comprising the following steps:

[0058] Step 1: Data preprocessing: video frame serialization, image annotation, and data set division

[0059] 1) Video frame serialization: Through the multimedia operation open source program integrated in Python, the video can be cropped, intercepted, and other functions can be performed. Figure 2 ;

[0060] 2) Image annotation: Use the LabelImg image annotation tool to annotate the abnormal behavior detection dataset, frame the abnormal behavior in each image with a rectangular frame, and indicate the category to which it belongs;

[0061] 3) Divide the dataset: Divide the dataset obtained in step 2) into a training set and a test set, where 80% of the dataset is used as a training set and 20% as a test set. Compress the original image into an image of 300 pixels × 300 pixels as the input of the training model.

[0062] Step 2: Train the SSD network:

[0063] The image processed in step 3) of step 1 is used as the input image of the SSD network model, and the operating parameters of the SSD network model are set. The SSD network model is trained on the experimental operation platform, and then the training effect of the SSD network model is evaluated using the evaluation indicators commonly used in the field of abnormal behavior detection. The problems existing in the SSD network model include: the large number of parameters leads to poor real-time performance of the model, and the inability to detect abnormal behaviors with partial occlusion, resulting in low detection accuracy;

[0064] Step 3. Improve the SSD network for the evaluation in Step 2

[0065] 1) Replace the feature extraction network:

[0066] Replace the feature extraction network of the standard SSD network with the lightweight network MobileNet v2 network to reduce the scale of network model parameters;

[0067] 2) Embed the deformable convolution module:

[0068] See Figure 3 , for an input feature map, assuming the original convolution operation is 3×3, in order to learn the offset, another 3×3 convolution layer is defined, and the output dimension is the same as the original feature map, and the number of channels is equal to 2N; Figure 3 The deformable convolution in the lower part can be considered as an interpolation operation based on the offset generated in the upper part, and then an ordinary convolution is performed. By changing the feature extraction method, the network can learn more fully. Although it increases a small amount of computational complexity, it can greatly improve the network performance; the specific method is as follows:

[0069] 2.1) The input image passes through an ordinary convolution to obtain the input feature map;

[0070] 2.2) See the input feature map Figure 3 The upper part of the convolution obtains the offset ΔP n ;

[0071] 2.3) The offset may be a floating point number, while the image positions are all integers, so bilinear interpolation is used to represent the offset position, and secondly, it is also convenient for gradient backpropagation;

[0072] 2.4) Add the obtained offset to the input feature map to get the new sampling position, and then use a 3×3 convolution kernel to extract features to obtain the output feature map;

[0073] Definition of the convolution kernel:

[0074] R = {(-1, -1), (-1, 0),...(0, 1), (1, 1)} (1)

[0075] Among them, R defines the size and dilation of the receptive field.

[0076] The output of the ordinary convolution is:

[0077]

[0078] Among them, X is the input, Y is the output, W is the weight matrix, P0 is each point on the feature map, and P n is n points in the grid;

[0079] The output of the deformable convolution is:

[0080]

[0081] where ΔP n is the coordinate offset;

[0082] 3) Design the attention mechanism module

[0083] See Figure 4 and Figure 5 , the SSD network model detects target objects by extracting six feature maps of different scales default in the SSD network. Based on this, the coordinate attention module can be regarded as a computing unit for enhancing the feature representation ability. The coordinate attention module is used to suppress insignificant features, enhance the expression ability of features in the network, and capture the long-range dependence relationship between spatial positions by learning the context relationship, predict the occluded part according to the unoccluded part, and then improve the accuracy of abnormal behavior detection. The specific method is as follows:

[0084] The coordinate attention module decomposes the channel attention into two parallel one-dimensional feature encoding processes, and effectively integrates the spatial coordinate information into the generated attention map. The specific operation is to perform average pooling on X and Y (i.e., the horizontal and vertical directions) to obtain two one-dimensional vectors, then splice and 1×1 convolution in the spatial dimension to compress the channels, and then encode the spatial information in two directions through the BN layer and ReLU and split, then each obtains the same number of channels as the input feature map through convolution, and then normalizes and weights. Finally, it is multiplied by the original feature map to adaptively adjust the features.

[0085] Place the coordinate attention module after the convolutional layer and before the BN layer, and perform feature enhancement by integrating the position information into the channel attention. Use the coordinate attention module to enhance 6 output feature maps extracted by MobileNet v2 respectively.

[0086] Step 4: Train the improved SSD network:

[0087] Use the image processed in Step 1 as the input image of the improved SSD network model in Step 3, and set the running parameters of the network model. Train the improved SSD model on the experimental operation platform, and use the evaluation metrics commonly used in the field of abnormal behavior detection to evaluate the detection results of the improved SSD. Finally, the AUC values are increased by 20.20% and 39.36% respectively in the two scenarios compared with the standard SSD model, and the detection speeds are increased by 21.08% and 24.80% respectively. The experiment verifies the effectiveness of this method.

[0088] The running parameters of the network model in Step 2 and Step 4 are as follows: the initial learning rate is 0.01, the Stochastic Gradient Descent (SGD) method is used to update the network parameters, the learning momentum is 0.9, and the weight decay rate is 0.0005.

[0089] The experimental operation platform in Step 2 and Step 4 is the Windows 10 system, using the Pytorch deep learning framework. The CPU model is Intel(R) Xeon(R) E5-2678 v3 @ 2.50 GHz, and the graphics card (GPU) model is NVIDIA GeForce RTX 2080 Ti with a graphics card memory of 11 GB. The PyCharm compilation environment is used.

[0090] The commonly used evaluation metric in the field of abnormal behavior detection in Step 2 and Step 4 is AUC (Area Under Curve); AUC is the area enclosed by the Receiver Operating Characteristic (ROC) curve, which takes values between 0 and 1. Its meaning is the probability that positive examples are ranked in front of negative examples. If the AUC value of a certain detection algorithm is relatively high, it can be considered that the algorithm has good performance. The horizontal and vertical coordinates of the ROC curve represent the False Positive Rate (FPR) and the True Positive Rate (TPR), respectively. Among them: TPR represents the ratio of samples that are correctly judged as positive among all true positive samples; FPR represents the ratio of samples that are wrongly judged as positive among all true negative samples;

[0091] Taking FPR as the abscissa and TPR as the ordinate, its calculation formula is as follows:

[0092]

[0093]

[0094] In the formula, represents being judged as positive and actually being positive; represents being judged as positive and actually being negative; represents being judged as negative and actually being negative; represents being judged as negative and actually being positive.

[0095] The proposed method for crowd abnormal behavior detection based on the improved SSD in the present invention is compared and analyzed with the standard SSD model, the SocialForce model, and the method proposed by Pang et al., and is verified in the crowd abnormal behavior detection dataset. The results show that the present invention has varying degrees of improvement in detection speed and accuracy compared with other methods, and can effectively improve the occlusion problem existing in the crowd. It is proved that the model of the present invention can achieve fast and accurate detection of crowd abnormal behavior and can provide support for public security monitoring.

[0096] The experimental results are shown in the following table.

[0097]

Claims

1. An abnormal behavior detection method for crowd based on improved SSD, characterized in that: The specific steps include: Step 1: Data preprocessing: including video frame serialization, image annotation, and data set division 1) Video frame serialization: Through the multimedia operation open source program integrated in Python, the video can be cropped and intercepted; 2) Image annotation: Use image annotation tools to annotate the abnormal behavior detection dataset, frame the abnormal behavior in each image with a rectangular frame, and indicate the category to which it belongs; 3) Divide the data set: Divide the data set expanded in step 2) into a training set and a test set, and compress the original image to the default size of the SSD network; Step 2: Train and evaluate the SSD network model The image processed in step 3) of step 1 is used as the input image of the SSD network model, and the operating parameters of the SSD network model are set. The SSD network model is trained on the experimental operation platform, and then the effect of the trained SSD network model is evaluated using the evaluation indicators commonly used in the field of abnormal behavior detection; Step 3: Improve the SSD network based on the evaluation in step 2 1) Replace the feature extraction network Replace the feature extraction network of the SSD network with a lightweight MobileNet v2 network to reduce the size of the network model parameters; 2) Embedding Deformable Convolutional Module By adding direction vectors to change the fixed size of the traditional convolution kernel, it can adaptively adjust its shape to cope with target subjects of different scales and deformations, thereby better extracting input features. 3) Design attention mechanism module The SSD network model detects the target object by extracting the default six feature maps of different scales of the SSD network. The feature map contains feature channels and position information. The content in the map contributes differently to the result of the target detection task, that is, the significance is different. The coordinate attention module is used to suppress insignificant features through learning, enhance the expressiveness of features in the network, and obtain contextual relationships through learning. It can capture the long-range dependency between spatial positions, predict the abnormal behavior of the occluded part according to the abnormal behavior of the unoccluded part of the crowd, and thus improve the target detection effect; Step 4: Train and improve the SSD network The image processed in step 3) of step 1 is used as the input image of the improved SSD network in step 3, and the operating parameters of the network model are set. The improved SSD network model in step 3 is trained on the experimental operation platform, and the detection results of the improved SSD are evaluated using the evaluation indicators commonly used in the field of abnormal behavior detection, and the final abnormal behavior detection results are output.

2. The method for detecting abnormal crowd behavior based on improved SSD according to claim 1, characterized in that: The specific method of step 2) of step 3 is as follows: 2.1) The input image undergoes a normal convolution to obtain the input feature map; 2.2) Convolve the upper half of the input feature map to obtain the offset ΔP n ; 2.3) Use bilinear interpolation to represent the offset position; 2.4) Add the obtained offset to the input feature map to obtain the new sampling position, and then use the convolution kernel to extract features to obtain the output feature map; Definition of convolution kernel: R={(-1,-1),(-1,0),...(0,1),(1,1)} (1) Among them, R defines the size and expansion of the receptive field; The output of a normal convolution is: Where X is the input, Y is the output, W is the weight matrix, P0 is each point on the feature map, and P n are n points in the grid; The output of the deformable convolution is: where, ΔP n is the coordinate offset.

3. The method for detecting abnormal behaviors of a crowd based on an improved SSD according to claim 1, characterized in that: The specific method of step 3) is as follows: The coordinate attention module decomposes channel attention into two parallel one-dimensional feature encoding processes, effectively integrating spatial coordinate information into the generated attention map. That is, average pooling is performed in the X horizontal and Y vertical directions to obtain two one-dimensional vectors. Next, concatenation in the spatial dimension and 1×1 convolution are used to compress the channels. Then, the spatial information in both directions is encoded through a BN layer and ReLU and split. Subsequently, convolutions are respectively performed to obtain the same number of channels as the input feature map, followed by normalization and weighting. Finally, the result is multiplied by the original feature map to adaptively adjust the features. The coordinate attention module is placed after the convolutional layer and before the BN layer. Feature enhancement is performed by integrating position information into channel attention, and the output feature maps extracted by MobileNet v2 are enhanced using the coordinate attention module respectively.

4. The method for detecting abnormal human behaviors based on the improved SSD according to claim 1, wherein: The operating parameters of the network model in Step 2 and Step 4 include the initial learning rate, learning momentum, and weight decay rate. The network parameters are updated using the stochastic gradient descent method (SGD).

5. The method for detecting abnormal crowd behavior based on improved SSD according to claim 1, wherein: The commonly used evaluation metrics in the field of abnormal behavior detection in Step 2 and Step 4 are AUC; the receiver operating characteristic is ROC; the horizontal and vertical coordinates of the ROC curve represent the false positive rate (FPR) and the true positive rate (TPR) respectively. Among them: TPR represents the ratio of samples that are correctly judged as positive among all true positive samples; FPR represents the ratio of samples that are wrongly judged as positive among all true negative samples. Using FPR as the abscissa and TPR as the ordinate, its calculation formula is as follows: In the formula, represents being judged as positive and actually being positive; represents being judged as positive and actually being negative; represents being judged as negative and actually being negative; represents being judged as negative and actually being positive.