Lightweight Dual Frame Rate Network-Based Abnormal Behavior Recognition Method, Device, and System
By adopting a lightweight dual-frame rate network in the recognition of abnormal behavior of human bodies, combining the feature fusion of low frame rate and high frame rate branch networks, the problems of computing resource density and feature extraction independence in the prior art are solved, and the effect of efficiently identifying abnormal behavior on edge AI platforms is achieved.
Patent Information
- Application Number
- CN202210314736.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-28
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-03-28
AI Technical Summary
The prior art has the problem of high computing resources in the recognition of abnormal behaviors of humans and it is difficult to operate efficiently on edge AI platforms. The spatial and temporal feature extraction of dual-stream networks is independent, ignoring the intrinsic connection of features.
The lightweight dual-frame rate network is adopted to capture the spatial semantic information of video data through the low-frame rate branch network, and the high-frame rate branch network captures motion information, and the fusion feature is connected to the classifier for abnormal behavior recognition.
It realizes efficient identification of abnormal behavior on edge AI platforms, reduces the demand for computing resources, and better captures the spatio-temporal information of video data through feature fusion.
Smart Images

Figure CN114663816B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning video behavior recognition, and in particular to an abnormal behavior recognition method and device based on a lightweight dual-frame rate network. Background Art
[0002] In the past, AI had to rely on powerful cloud computing capabilities for data analysis and algorithm operation. With the maturity of technology and the emergence of new applications, chip capabilities have continued to improve and edge computing platforms have matured. The main reason for edge computing is the lack of cloud computing services. Cloud computing mostly uses centralized management methods, which enables cloud services to create higher economic benefits. In the context of the Internet of Everything, application services require low latency, high reliability and data security, and traditional cloud computing cannot meet these requirements.
[0003] With the development of edge AI, there is a clear trend for machine learning prediction to move down to embedded hardware that is closer to users, does not require network connections, and can solve complex problems (such as abnormal behavior detection) in real time. Then designing an abnormal behavior recognition method that can run on an edge AI platform is also a problem that the present invention needs to solve.
[0004] Abnormal behavior recognition refers to identifying human behaviors that are beyond the normal range in high-level semantic understanding from the monitored video stream. For example, pedestrians walking normally on the street, standing and answering phone calls, etc., these behaviors are all within the normal range; if two or more people on the street engage in physical fights or other fights, or damage public property such as pushing down trash cans on the street, or steal cars and motorcycles on the street, then the above behaviors are considered to be abnormal behaviors that are beyond the normal range. At this time, an alarm needs to be issued so that the area management personnel can stop the illegal and irregular behaviors in time; if the pedestrians walk normally, then the behavior is considered to be normal behavior that does not exceed the normal range, and no alarm is required.
[0005] At present, deep learning convolutional neural networks have achieved good results in the field of abnormal human behavior recognition, such as recurrent neural networks, two-stream networks and 3D convolutional neural networks.
[0006] The method based on recurrent neural network is suitable for simple sequence data such as skeletons. Some people have studied the method of enhancing the spatial expression ability of the human skeleton. The three-dimensional coordinates of the original skeleton joints are transferred to the human coordinate system and then scaled. Then they are input into the improved residual independent recurrent neural network to recognize the human skeleton sequence behavior. It performs well on the NTU RGB+D dataset, but the proposed network lacks spatial modeling capabilities.
[0007] The method based on the two-stream network calculates the dense optical flow for every two frames in the video sequence to obtain a sequence of dense optical flows; then, CNN models are trained for the video images and the dense optical flows respectively, and the two branches of the network judge the categories of actions respectively; finally, the training results of the two networks are directly fused to obtain the final classification result. Some people have proposed a method for detecting abnormal behaviors in videos based on a two-stream convolutional neural network. This method uses the RGB images and the optical flow information between video frames as the inputs of the two network branches respectively to learn spatial information and temporal information, and uses a long short-term neural network to model the long-term dependencies between video frames, so as to obtain the final behavior classification result, and has achieved good recognition effects on the Shanghai Tech, UCSD Ped1, and Pedestrian 2 datasets. However, the spatio-temporal features extracted by the model are independent and it is easy to ignore their internal connections.
[0008] The method based on the 3D convolutional neural network can directly extract spatial and temporal features from the original video, significantly improving the performance in the field of action recognition and detection. Some people improved the Inception structure based on the original I3D, and replaced the convolutional kernels in the original Inception structure with two-level convolutional kernels using the principle of convolutional kernel replacement. And drawing on the idea of channel mixing, the channel mixing strategy is integrated into the improved I3D neural network, and a new 3D network model of I3D-shufflenet is proposed. Although the model is fast, the model parameters increase exponentially. Summary of the Invention
[0009] Aiming at the defects in the prior art, the purpose of the present invention is to provide a method, device, and system for abnormal behavior recognition based on a lightweight dual-frame rate network.
[0010] According to the first aspect of the present invention, a method for abnormal behavior recognition based on a lightweight dual-frame rate network is provided, including:
[0011] Input the original video data into a lightweight dual-frame rate convolutional neural network;
[0012] The low-frame rate branch network of the lightweight dual-frame rate convolutional neural network captures the spatial semantic information of the original video data;
[0013] The high-frame rate branch network of the lightweight dual-frame rate convolutional neural network captures the motion information of the original video data;
[0014] Use lateral connections to perform feature fusion on each stage of the low-frame rate branch network and the high-frame rate branch network;
[0015] Merge the feature vectors output by the low-frame rate branch network and the high-frame rate branch network respectively to obtain a merged feature;
[0016] Input the combined features into a classifier to obtain the classification result for abnormal behavior recognition.
[0017] Preferably, the low-frame-rate branch network samples video frame images at a low frame rate with an interval of 16 frames;
[0018] The high-frame-rate branch network samples video frame images at a high frame rate with an interval of 2 frames.
[0019] Preferably, the low-frame-rate branch network includes a five-level cascaded generalized convolutional layer, namely: the first generalized convolutional layer A1, the second generalized convolutional layer A2, the third generalized convolutional layer A3, the fourth generalized convolutional layer A4, and the fifth generalized convolutional layer A5;
[0020] Wherein:
[0021] The first generalized convolutional layer A1 includes: 1 convolutional layer of 3×3×3 and 1 max-pooling layer of 3×3×3;
[0022] The second generalized convolutional layer A2 successively includes: 1 branch 1 and 3 branch 2s;
[0023] The third generalized convolutional layer A3 successively includes: 1 branch 1 and 7 branch 2s;
[0024] The fourth generalized convolutional layer A4 successively includes: 1 branch 1 and 3 branch 2s;
[0025] The fifth generalized convolutional layer A5 successively includes: 1 convolutional layer of 1×1×1 and 1 average-pooling layer of 8×1×1.
[0026] Preferably, branch 1 is divided into two paths, wherein:
[0027] The first path includes: 1 depth convolutional layer of 3×3×3 with a stride of 2, 2 batch normalizations, 1 convolutional layer of 1×1×1, and 1 activation layer;
[0028] The second path includes: 2 convolutional layers of 1×1×1, 3 batch normalizations, 2 activation layers, 1 depth convolutional layer of 3×3×3 with a stride of 2, and 1 squeeze-and-excitation module;
[0029] The squeeze-and-excitation module includes: 1 adaptive average-pooling layer, 2 fully-connected layers, and 2 activation layers;
[0030] Concatenate the feature vectors output by the two paths in the channel dimension, and output after channel shuffle operation;
[0031] The branch 2 is divided into two paths through a channel splitting operation, that is, the feature vector of the input branch 2 is evenly divided into two parts according to the number of channels, one part is used as the input of the first path, and the other part is used as the input of the second path, where:
[0032] The first path includes: a shortcut connection of an identity mapping;
[0033] The second path includes: 2 1×1×1 convolutional layers, 3 batch normalizations, 2 activation layers, 1 3×3×3 depth convolutional layer with a stride of 1, and 1 squeeze-and-excitation module;
[0034] The squeeze-and-excitation module includes: 1 adaptive average pooling layer, 2 fully connected layers, and 2 activation layers;
[0035] The feature vectors output by the two paths are concatenated in the channel dimension and output after a channel shuffle operation.
[0036] Preferably, the high frame rate branch network includes five cascaded generalized convolutional layers, namely the first generalized convolutional layer B1, the second generalized convolutional layer B2, the third generalized convolutional layer B3, the fourth generalized convolutional layer B4, and the fifth generalized convolutional layer B5;
[0037] Where:
[0038] The first generalized convolutional layer B1 sequentially includes: 1 3×3×3 convolutional layer and 1 3×3×3 max pooling layer;
[0039] The second generalized convolutional layer B2 sequentially includes: 1 branch 1 and 3 branch 2s;
[0040] The third generalized convolutional layer B3 sequentially includes: 1 branch 1 and 7 branch 2s;
[0041] The fourth generalized convolutional layer B4 sequentially includes: 1 branch 1 and 3 branch 2s;
[0042] The fifth generalized convolutional layer B5 sequentially includes: 1 1×1×1 convolutional layer and 1 8×1×1 average pooling layer.
[0043] Preferably, the branch 1 is divided into two paths, where:
[0044] The first path includes: 1 3×3×3 depth convolutional layer with a stride of 2, 2 batch normalizations, 1 1×1×1 convolutional layer, and 1 activation layer;
[0045] The second path includes: 2 1×1×1 convolutional layers, 3 batch normalizations, 2 activation layers, 1 3×3×3 depth convolutional layer with a stride of 2, and 1 squeeze-and-excitation module;
[0046] The compression excitation module includes: 1 adaptive average pooling layer, 2 fully connected layers, and 2 activation layers;
[0047] The feature vectors output by the two paths are concatenated in the channel dimension and output after a channel shuffle operation;
[0048] Branch 2 is divided into two paths through a channel splitting operation, that is, the feature vector input to Branch 2 is evenly divided into two parts according to the number of channels, one part is used as the input of the first path, and the other part is used as the input of the second path, where:
[0049] The first path includes: an identity mapping shortcut connection;
[0050] The second path includes: 2 1×1×1 convolutional layers, 3 batch normalizations, 2 activation layers, 1 3×3×3 depth convolutional layer with a stride of 1, and 1 compression excitation module;
[0051] The compression excitation module includes: 1 adaptive average pooling layer, 2 fully connected layers, and 2 activation layers;
[0052] The feature vectors output by the two paths are concatenated in the channel dimension and output after a channel shuffle operation.
[0053] Preferably, the lateral connection includes four convolutional layers, namely the first convolutional layer C1, the second convolutional layer C2, the third convolutional layer C3, and the fourth convolutional layer C4, where:
[0054] The first convolutional layer C1 includes: 1 5×1×1 convolutional layer with a stride of 4;
[0055] The second convolutional layer C2 includes: 1 5×1×1 convolutional layer with a stride of 4;
[0056] The third convolutional layer C3 includes: 1 5×1×1 convolutional layer with a stride of 4;
[0057] The fourth convolutional layer C4 includes: 1 5×1×1 convolutional layer with a stride of 4;
[0058] Matching the feature sizes before fusion includes:
[0059] The feature size of the low frame rate branch network is {T, S 2 , C},
[0060] The feature size of the high frame rate branch network is {αT, S 2 , βC},
[0061] T represents the time length, S 2Let \(h\) and \(w\) represent the height and width of the feature map, \(\alpha\) represent the ratio of the sampling density of the high frame rate branch network to that of the low frame rate branch network, \(\beta\) represent the ratio of the number of channels of the high frame rate branch network to that of the low frame rate branch network, and \(C\) represent the number of channels;
[0062] Perform 3D convolution on the features of the high frame rate branch network, with the output number of channels being \(2\beta C\) and the stride being \(\alpha\);
[0063] The output result of the high frame rate branch network is fused into the low frame rate branch network through splicing, including:
[0064] The output of the first general convolution layer B1 is used as the input of the first convolution layer C1, and the output of the first convolution layer C1 and the output of the first general convolution layer A1 are fused and then used as the input of the second general convolution layer A2;
[0065] The output of the second general convolution layer B2 is used as the input of the second convolution layer C2, and the output of the second convolution layer C2 and the output of the second general convolution layer A2 are fused and then used as the input of the third general convolution layer A3;
[0066] The output of the third general convolution layer B3 is used as the input of the third convolution layer C3, and the output of the third convolution layer C3 and the output of the third general convolution layer A3 are fused and then used as the input of the fourth general convolution layer A4;
[0067] The output of the fourth general convolution layer B4 is used as the input of the fourth convolution layer C4, and the output of the fourth convolution layer C4 and the output of the fourth general convolution layer A4 are fused and then used as the input of the fifth general convolution layer A5.
[0068] Preferably, merging the feature vectors finally output by the two branch networks as the input of the classifier to obtain the abnormal behavior recognition classification result includes:
[0069] After the two branch networks perform convolution operations, the vectors containing feature parameters finally output are concatenated and then input into the fully connected layer;
[0070] The fully connected layer inputs the calculated feature vector into the Sigmoid regression layer for regression calculation to obtain the classification result.
[0071] According to the second aspect of the present invention, there is provided an abnormal behavior recognition device based on a lightweight dual frame rate network, including:
[0072] A video extraction device, which acquires a video;
[0073] A processor, which performs loading processing on the video according to the video to obtain an abnormal behavior recognition result;
[0074] A memory that stores the lightweight dual-frame rate network model executed by the processor, and when the machine-readable instructions in the memory are executed by the processor, the method for identifying abnormal behaviors based on the lightweight dual-frame rate network is executed.
[0075] According to the third aspect of the present invention, there is provided an abnormal behavior recognition system based on a lightweight dual-frame rate network, including:
[0076] A low-frame rate branch network that captures the spatial semantic information of the original video data;
[0077] A high-frame rate branch network that captures the motion information of the original video data;
[0078] A lateral connection that performs feature fusion on the low-frame rate branch network and the high-frame rate branch network;
[0079] A merging network that merges the feature vectors respectively output by the low-frame rate branch network and the high-frame rate branch network to obtain a merged feature;
[0080] A classification network that inputs the merged feature to obtain an abnormal behavior recognition classification result.
[0081] Compared with the prior art, the present invention has the following beneficial effects:
[0082] An abnormal behavior recognition method based on a lightweight dual-frame rate network provided by an embodiment of the present invention, wherein the lightweight dual-frame rate network recognition model includes a low-frame rate branch network, a high-frame rate branch network, and a lateral connection. The low-frame rate branch network uses a relatively large time span, that is, the number of frames extracted per second. For example, image frames are extracted at a rate of every 16 frames at intervals, aiming to capture the spatial semantic information provided by sparse frames. The high-frame rate branch network uses a relatively small time span. For example, image frames are extracted at a rate of every 2 frames at intervals, aiming to capture the time semantic information of rapid changes. Through the lateral connection, the features of the low-frame rate branch network and the high-frame rate branch network are fused;
[0083] In addition, both branch networks use lightweight networks as the basic networks, which greatly improves the detection efficiency while maintaining the ability of the low-frame rate branch network to extract category spatial semantics and the high-frame rate branch network to extract time semantic information. BRIEF DESCRIPTION OF THE DRAWINGS
[0084] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objectives, and advantages of the present invention will become more obvious:
[0085] The following further describes the embodiments of the present invention with reference to the drawings:
[0086] Figure 1 The structural diagram of the lightweight dual-frame rate network provided by an embodiment of the present invention;
[0087] Figure 2 The schematic flowchart of the abnormal behavior recognition method provided by an embodiment of the present invention;
[0088] Figure 3 The schematic diagram of the channel shuffle operation provided by an embodiment of the present invention;
[0089] Figure 4 The structural diagram of branch 1 of the lightweight dual-frame rate network provided by an embodiment of the present invention;
[0090] Figure 5 The structural diagram of branch 2 of the lightweight dual-frame rate network provided by an embodiment of the present invention;
[0091] Figure 6 The process diagram of the abnormal behavior recognition method based on the lightweight dual-frame rate network for processing videos provided by an embodiment of the present invention;
[0092] Figure 7 The structural diagram of the electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0093] The present invention will be described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made. These all belong to the protection scope of the present invention.
[0094] Before introducing the human abnormal behavior recognition method provided by an embodiment of the present application, the application scenarios applicable to the human abnormal behavior recognition method are first introduced. The application scenarios include but are not limited to: identifying whether there are abnormal behaviors in a crowd through videos captured by cameras in public places.
[0095] As Figure 1 shown, it is the structural diagram of the lightweight dual-frame rate network provided by an embodiment of the present application. The main idea of this abnormal behavior recognition method includes:
[0096] Sampling image frames at a sampling rate of 16 frames at intervals from the video stream as the input of the low-frame rate branch network, sampling image frames at a sampling rate of 2 frames at intervals as the input of the high-frame rate branch network, fusing the features generated at each stage of the high-frame rate branch network with the features generated at each stage of the low-frame rate branch network, concatenating the final output of the low-frame rate branch network and the final output of the high-frame rate branch network to predict the classification result, and judging whether an alarm is needed according to the classification result.
[0097] In the above embodiments, by using a two-stream-like network structure, different from the feature that the spatio-temporal features extracted by the two-stream network are independent, the spatio-temporal features extracted by it will be fused during the extraction process, overcoming the shortcomings of the two-stream network. A lightweight 3D convolutional network is used as the basic network, starting from two aspects: FLOPs and the time cost of memory access, to reduce the overall computational amount of the network, overcoming the shortcoming of the large computational amount of the 3D network model, so that it can be used on the edge AI platform.
[0098] As Figure 2 shown, it is a flowchart of an abnormal behavior recognition method further optimized based on the above main idea provided by the present invention, including:
[0099] S11: Obtain a monitoring video sample data set, label the video sample data set, and perform preprocessing to obtain a preprocessed video data set;
[0100] S12: Divide the preprocessed video data in S11 into a training data set and a validation data set;
[0101] S13: Construct a lightweight dual-frame rate convolutional neural network recognition model;
[0102] S14: Use the training data set constructed in S12 to train the lightweight dual-frame rate convolutional neural network recognition model;
[0103] S15: Obtain monitoring video data, input it into the trained lightweight dual-frame rate convolutional neural network recognition model, and obtain the abnormal behavior recognition classification result.
[0104] In a preferred embodiment provided by the present invention, there are many implementation manners for S11 to obtain the monitoring video sample data set, including but not limited to:
[0105] The first acquisition method: Use acquisition devices such as video recorders, cameras, or color cameras to shoot the target object to obtain a video sample data set; then the acquisition device sends the video sample data set, and the electronic device receives the video sample data set sent by the acquisition device;
[0106] The second acquisition method: Obtain the video sample data set from the file system, database, or removable storage device of the video server.
[0107] In another preferred embodiment provided by the present invention, S11 and S12 include: selecting the currently publicly available crime video dataset UCF-Crime, selecting videos of three types (fighting, stealing, and vandalism) from it, performing frame preprocessing on the video data at 1 frame per second to obtain a training video frame labeled image set, and classifying the human behaviors in each training video frame image in the training video frame labeled image set and assigning label information, where there are 8132 labels in the training set and 2155 labels in the validation set;
[0108] Performing frame preprocessing on the video data at 30 frames per second results in a total of 388357 pictures, where there are 296350 pictures in the training set and 92007 pictures in the validation set.
[0109] In order to overcome the disadvantage of easily ignoring the internal connections of features in the two-stream network and the disadvantage of large computational complexity of using a 3D network model, the present invention provides a preferred embodiment. The construction of the lightweight dual-frame rate convolutional neural network recognition model in S13, whose network structure is as Figure 1 shown, includes a low-frame rate branch network, a high-frame rate branch network, and a lateral connection. Among them, the high-frame rate branch network is a convolutional network similar to the low-frame rate branch network, but has a ratio of β (β < 1) channels of the low-frame rate branch network. In the embodiment, the typical value is β = 1 / 8. The calculation (floating-point operations or FLOPs) of the common layer is usually the square of the channel scaling ratio, so the high-frame rate branch network is more computationally efficient than the low-frame rate branch network. In the embodiment, the high-frame rate branch network usually accounts for 20% of the total computational amount.
[0110] As Figure 1 shown, the low-frame rate branch network includes cascaded first generalized convolutional layer A1, second generalized convolutional layer A2, third generalized convolutional layer A3, fourth generalized convolutional layer A4, and fifth generalized convolutional layer A5, where:
[0111] The first generalized convolutional layer A1 includes a 3×3×3 convolutional layer, a max-pooling layer with a stride of 1×2×2 and 24 + 3×3×3 convolutional kernels, and a stride of 1×2×2;
[0112] The second generalized convolutional layer A2 includes 4 parts (1 branch 1 and 3 branch 2), and the specific structure is as follows:
[0113] The first part (i.e., 1 branch 1, the schematic diagram of the branch 1 structure is as Figure 4 shown) contains two paths, the output features of the two paths are concatenated by channel number, and after concatenation, they are output after a channel shuffle operation. The schematic diagram of the channel shuffle operation is as Figure 3As shown, the output feature 1 after grouped convolution is "shuffled evenly" through channel shuffle operation to ensure information exchange between feature maps of different groups after grouped convolution without increasing the computational complexity, so as to fully fuse the channels. Among them:
[0114] The first path includes: a 3×3×3 convolutional layer with a stride of 1×2×2, 30 convolutional kernels, group = 30 + batch normalization + a 1×1×1 convolutional layer with a stride of 1×1×1, 88 convolutional kernels + batch normalization + ReLU activation function;
[0115] The ReLU activation function is:
[0116]
[0117] The second path includes: a 1×1×1 convolutional layer with a stride of 1×1×1, 88 convolutional kernels + batch normalization + ReLU activation function + a 3×3×3 convolutional layer with a stride of 1×2×2, 88 convolutional kernels, group = 88 + batch normalization + a 1×1×1 adaptive average pooling layer + a fully connected layer + ReLU activation function + a fully connected layer + Sigmoid activation function + a 1×1×1 convolutional layer with a stride of 1×1×1, 88 convolutional kernels + batch normalization + ReLU activation function;
[0118] The Sigmoid activation function is:
[0119]
[0120] The second part, the third part and the fourth part (i.e., 3 branches 2, the schematic diagram of the branch 2 structure is as Figure 5 shown) are repeated and parallel structures, and the specific structures are as follows:
[0121] Each part contains two paths, the output features of the two paths are concatenated by the number of channels, and after concatenation, they are output after channel shuffle operation, where:
[0122] The first path includes: an identity mapping shortcut connection;
[0123] The second path includes: a 1×1×1 convolutional layer with a stride of 1×1×1, 88 convolutional kernels + batch normalization + Relu activation function + a 3×3×3 convolutional layer with a stride of 1×1×1, 88 convolutional kernels, group = 88 + batch normalization + a 1×1×1 adaptive average pooling layer + a fully connected layer + Relu activation function + a fully connected layer + Sigmoid activation function + a 1×1×1 convolutional layer with a stride of 1×1×1, 88 convolutional kernels + batch normalization + Relu activation function;
[0124] The third generalized convolutional layer A3 consists of 8 parts (1 branch 1 and 7 branch 2s), and the specific structure is as follows:
[0125] The first part (i.e., 1 branch 1) contains two paths, and the output features of the two paths are concatenated by the number of channels. After concatenation, it is output after a channel shuffle operation, where:
[0126] The first path includes: a convolutional layer of 3×3×3, a stride of 1×2×2, 220 convolutional kernels, group of 220 + batch normalization + a convolutional layer of 1×1×1, a stride of 1×1×1, 176 convolutional kernels + batch normalization + Relu activation function;
[0127] The second path includes: a convolutional layer of 1×1×1, a stride of 1×1×1, 176 convolutional kernels + batch normalization + Relu activation function + a convolutional layer of 3×3×3, a stride of 1×2×2, 176 convolutional kernels + batch normalization + a 1×1×1 adaptive average pooling layer + a fully connected layer + Relu activation function + a fully connected layer + Sigmoid activation function + a convolutional layer of 1×1×1, a stride of 1×1×1, 176 convolutional kernels + batch normalization + Relu activation function;
[0128] The second part to the eighth part (i.e., 7 branch 2s) are repeated and parallel structures, and the specific structure is as follows:
[0129] Each part contains two paths, and the output features of the two paths are concatenated by the number of channels. After concatenation, it is output after a channel shuffle operation, where:
[0130] The first path includes: a shortcut connection of identity mapping;
[0131] The second path includes: a convolutional layer of 1×1×1, a stride of 1×1×1, 176 convolutional kernels + batch normalization + Relu activation function + a convolutional layer of 3×3×3, a stride of 1×1×1, 176 convolutional kernels, group of 176 + batch normalization + a 1×1×1 adaptive average pooling layer + a fully connected layer + Relu activation function + a fully connected layer + Sigmoid activation function + a convolutional layer of 1×1×1, a stride of 1×1×1, 176 convolutional kernels + batch normalization + Relu activation function;
[0132] The fourth generalized convolutional layer A4 consists of 4 parts (1 branch 1 and 3 branch 2s), and the specific structure is as follows:
[0133] The first part (i.e., 1 branch 1) contains two paths, and the output features of the two paths are concatenated by the number of channels. After concatenation, it is output after a channel shuffle operation, where:
[0134] The first path includes: a 3×3×3 convolutional layer with a stride of 1×2×2, 440 convolutional kernels, a group of 440 + batch normalization + a 1×1×1 convolutional layer with a stride of 1×1×1, 352 convolutional kernels + batch normalization + Relu activation function;
[0135] The second path includes: a 1×1×1 convolutional layer with a stride of 1×1×1, 352 convolutional kernels + batch normalization + Relu activation function + a 3×3×3 convolutional layer with a stride of 1×2×2, 352 convolutional kernels, a group of 352 + batch normalization + a 1×1×1 adaptive average pooling layer + a fully connected layer + Relu activation function + a fully connected layer + Sigmoid activation function + a 1×1×1 convolutional layer with a stride of 1×1×1, 352 convolutional kernels + batch normalization + Relu activation function;
[0136] The second part, the third part, and the fourth part (i.e., 3 branches 2) are repetitive and parallel structures, and the specific structures are as follows:
[0137] Each part contains two paths, and the output features of the two paths are concatenated by the number of channels and then output after a channel shuffle operation, where:
[0138] The first path includes: a shortcut connection of an identity mapping;
[0139] The second path includes: a 1×1×1 convolutional layer with a stride of 1×1×1, 352 convolutional kernels + batch normalization + Relu activation function + a 3×3×3 convolutional layer with a stride of 1×1×1, 352 convolutional kernels, a group of 352 + batch normalization + a 1×1×1 adaptive average pooling layer + a fully connected layer + Relu activation function + a fully connected layer + Sigmoid activation function + a 1×1×1 convolutional layer with a stride of 1×1×1, 352 convolutional kernels + batch normalization + Relu activation function;
[0140] The fifth general convolutional layer A5 includes: a 1×1×1 convolutional layer with a stride of 1×1×1, 1024 convolutional kernels + an 8×1×1 average pooling layer.
[0141] As Figure 5 shown, the high frame rate branch network includes cascaded first general convolutional layer B1, second general convolutional layer B2, third general convolutional layer B3, fourth general convolutional layer B4, and fifth general convolutional layer B5, where:
[0142] The first generalized convolutional layer B1 includes a 3×3×3 convolutional layer, a max pooling layer with a stride of 1×2×2 and 3 + 3×3×3 convolutional kernels, and a stride of 1×2×2;
[0143] The second generalized convolutional layer B2 includes 4 parts (1 branch 1 and 3 branches 2), and the specific structure is as follows:
[0144] The first part (i.e., 1 branch 1) contains two paths, and the output features of the two paths are concatenated by the number of channels and then output after a channel shuffle operation, where:
[0145] The first path includes: a 3×3×3 convolutional layer, a stride of 1×2×2, 3 convolutional kernels, group is 3 + batch normalization + a 1×1×1 convolutional layer, a stride of 1×1×1, 11 convolutional kernels + batch normalization + Relu activation function;
[0146] The second path includes: a 1×1×1 convolutional layer, a stride of 1×1×1, 11 convolutional kernels + batch normalization + Relu activation function + a 3×3×3 convolutional layer, a stride of 1×2×2, 11 convolutional kernels, group is 11 + batch normalization + a 1×1×1 adaptive average pooling layer + fully connected layer + Relu activation function + fully connected layer + Sigmoid activation function + a 1×1×1 convolutional layer, a stride of 1×1×1, 11 convolutional kernels + batch normalization + Relu activation function;
[0147] The second part, the third part, and the fourth part (i.e., 3 branches 2) are repeated and parallel structures, and the specific structure is as follows:
[0148] Each part contains two paths, and the output features of the two paths are concatenated by the number of channels and then output after a channel shuffle operation, where:
[0149] The first path includes: an identity mapping shortcut connection;
[0150] The second path includes: a 1×1×1 convolutional layer, a stride of 1×1×1, 11 convolutional kernels + batch normalization + Relu activation function + a 3×3×3 convolutional layer, a stride of 1×1×1, 11 convolutional kernels, group is 11 + batch normalization + a 1×1×1 adaptive average pooling layer + fully connected layer + Relu activation function + fully connected layer + Sigmoid activation function + a 1×1×1 convolutional layer, a stride of 1×1×1, 11 convolutional kernels + batch normalization + Relu activation function;
[0151] The third generalized convolutional layer B3 includes 8 parts (1 branch 1 and 7 branches 2), and the specific structure is as follows:
[0152] The first part (i.e., 1 branch 1) contains two paths, and the output features of the two paths are concatenated by the number of channels and then output after a channel shuffle operation, where:
[0153] The first path includes: a 3×3×3 convolutional layer with a stride of 1×2×2, 22 convolutional kernels, group of 22 + batch normalization + a 1×1×1 convolutional layer with a stride of 1×1×1, 22 convolutional kernels + batch normalization + Relu activation function;
[0154] The second path includes: a 1×1×1 convolutional layer with a stride of 1×1×1, 22 convolutional kernels + batch normalization + Relu activation function + a 3×3×3 convolutional layer with a stride of 1×2×2, 22 convolutional kernels + batch normalization + a 1×1×1 adaptive average pooling layer + a fully connected layer + Relu activation function + a fully connected layer + Sigmoid activation function + a 1×1×1 convolutional layer with a stride of 1×1×1, 22 convolutional kernels + batch normalization + Relu activation function;
[0155] The second part to the eighth part (i.e., 7 branches 2) are repeated and parallel structures, and the specific structures are as follows:
[0156] Each part contains two paths, and the output features of the two paths are concatenated by the number of channels and then output after a channel shuffle operation, where:
[0157] The first path includes: an identity mapping shortcut connection;
[0158] The second path includes: a 1×1×1 convolutional layer with a stride of 1×1×1, 22 convolutional kernels + batch normalization + Relu activation function + a 3×3×3 convolutional layer with a stride of 1×1×1, 22 convolutional kernels, group of 22 + batch normalization + a 1×1×1 adaptive average pooling layer + a fully connected layer + Relu activation function + a fully connected layer + Sigmoid activation function + a 1×1×1 convolutional layer with a stride of 1×1×1, 22 convolutional kernels + batch normalization + Relu activation function;
[0159] The fourth generalized convolutional layer B4 includes 4 parts (1 branch 1 and 3 branches 2), and the specific structures are as follows:
[0160] The first part (i.e., 1 branch 1) contains two paths, and the output features of the two paths are concatenated by the number of channels and then output after a channel shuffle operation, where:
[0161] The first path includes: a 3×3×3 convolutional layer with a stride of 1×2×2, 44 convolutional kernels, a group of 44 + batch normalization + a 1×1×1 convolutional layer with a stride of 1×1×1, 44 convolutional kernels + batch normalization + Relu activation function;
[0162] The second path includes: a 1×1×1 convolutional layer with a stride of 1×1×1, 44 convolutional kernels + batch normalization + Relu activation function + a 3×3×3 convolutional layer with a stride of 1×2×2, 44 convolutional kernels, a group of 44 + batch normalization + a 1×1×1 adaptive average pooling layer + a fully connected layer + Relu activation function + a fully connected layer + Sigmoid activation function + a 1×1×1 convolutional layer with a stride of 1×1×1, 44 convolutional kernels + batch normalization + Relu activation function;
[0163] The second part, the third part, and the fourth part (i.e., 3 branches 2) are repetitive and parallel structures, and the specific structures are as follows:
[0164] Each part contains two paths, and the output features of the two paths are concatenated by the number of channels and then output after a channel shuffle operation, where:
[0165] The first path includes: an identity mapping shortcut connection;
[0166] The second path includes: a 1×1×1 convolutional layer with a stride of 1×1×1, 44 convolutional kernels + batch normalization + Relu activation function + a 3×3×3 convolutional layer with a stride of 1×1×1, 44 convolutional kernels, a group of 44 + batch normalization + a 1×1×1 adaptive average pooling layer + a fully connected layer + Relu activation function + a fully connected layer + Sigmoid activation function + a 1×1×1 convolutional layer with a stride of 1×1×1, 11 convolutional kernels + batch normalization + Relu activation function;
[0167] The fifth general convolutional layer B5 includes: a 1×1×1 convolutional layer with a stride of 1×1×1, 128 convolutional kernels + an 8×1×1 average pooling layer.
[0168] As Figure 1 shown, the lateral connections include: the first convolutional layer C1, the second convolutional layer C2, the third convolutional layer C3, the fourth convolutional layer C4, and the specific structures are as follows:
[0169] The first convolutional layer C1 includes a 5×1×1 convolutional layer with a stride of 4×1×1, 6 convolutional kernels: fusing the output features of the first general convolutional layer B1 and the output features of the first general convolutional layer A1 as the input of the second general convolutional layer A2;
[0170] The second convolutional layer C2 includes a convolutional layer of 5×1×1, a stride of 4×1×1, and 44 convolutional kernels: fusing the output features of the second generalized convolutional layer B2 and the output features of the second generalized convolutional layer A2 as the input of the third generalized convolutional layer A3;
[0171] The third convolutional layer C3 includes a convolutional layer of 5×1×1, a stride of 4×1×1, and 88 convolutional kernels: fusing the output features of the third generalized convolutional layer B3 and the output features of the third generalized convolutional layer A3 as the input of the fourth generalized convolutional layer A4;
[0172] The fourth convolutional layer C4 includes a convolutional layer of 5×1×1, a stride of 4×1×1, and 176 convolutional kernels: fusing the output features of the fourth generalized convolutional layer B4 and the output features of the fourth generalized convolutional layer A4 as the input of the fifth generalized convolutional layer A5.
[0173] As Figure 1 shown, the final output of the low frame rate branch network and the final output of the high frame rate branch network are concatenated according to the number of channels as the input of the fully connected layer, and the fully connected layer inputs the calculated feature vector into the Sigmoid regression layer for regression calculation to obtain the classification result.
[0174] To better perform classification and recognition, the present invention provides a preferred embodiment. In this embodiment, S14 uses the training data set constructed in S12 to train the lightweight dual frame rate convolutional neural network recognition model;
[0175] S14 further includes: As Figure 6Schematic diagram of the video processing process provided by an embodiment of the present application. In a specific practice process, a lightweight dual-frame rate convolutional neural network recognition model is deployed on NVIDIA Jetson AGX Xavier (i.e., the video processing module in the figure). The first step is to download and install the SDK Manager on the host. The second step is to connect the host and Jetson AGX Xavier with a USB data cable. The third step is to enter the recovery state after the host and Jetson AGX Xavier are connected and then install and start. The fourth step is to enter the installation interface after logging in to the nvidia account. First, select the download path, then burn the OS image and install the SDK components. After the burning is completed, Jetson AGX xavier will automatically power on and enter the Ubuntu system settings interface. After the settings are completed, Jetson AGX xavier will enter the ubuntu system. At this time, change the apt-get source of this system to a domestic source. At this time, basic software toolkits, including CUDA, cudnn, OpenCV, etc., have been deployed on Jetson AGX xavier, providing a basic environment for the deployment of the network model. Next, build the environment required for the network model on JetsonAGX xavier: Pytorch framework, Numpy, Detestron2, etc.
[0176] After the deployment is completed, the video frames collected by the video capture device can be input into the video processing module (i.e., NVIDIA Jetson AGX Xavier). After being processed by the lightweight dual-frame rate convolutional neural network recognition model deployed inside it, the final classification result is sent to the local (for example, abnormal situations such as fighting, theft, and vandalism).
[0177] S15 further includes: As Figure 6 shown, in the embodiment, video frames can be collected through an external camera, i.e., the video capture module. 64 image frames (the length of the video can also be changed according to actual requirements. In this embodiment, 64 frames are used) are used as a video segment and input into the trained lightweight dual-frame rate convolutional neural network recognition model in the video processing module for processing. Among them, the low-frame rate branch network samples 4 frames as input, and the high-frame rate branch network samples 32 frames as input. After processing, the prediction result of this video segment is obtained, and the result is uploaded to the local to determine whether there is abnormal behavior in this video segment. Next, 32 frames from the previous video segment and the next 32 image frames are combined into a video segment of 64 frames in total as the input of the network model. After processing, the processing result of this video segment is obtained. During prediction, the above operations are repeated to process the video segment.
[0178] Based on the same inventive concept, in other embodiments of the present invention, an electronic device for abnormal behavior recognition based on a lightweight dual frame rate network is provided. In one embodiment, its structure is as follows Figure 7 shown, including: a video acquisition device S11, a processor S12, and a memory S13. The memory S13 stores a network model executable by the processor S12. When the machine-readable instructions are executed by the processor S12, the above method is executed. The acquired video is input into the processor S12 through the device S11, and the abnormal behavior recognition result is obtained through model loading and processing in the processor S12.
[0179] Based on the same inventive concept of the present invention, in other embodiments of the present invention, an abnormal behavior recognition system based on a lightweight dual frame rate network is provided, including: a low frame rate branch network, a high frame rate branch network, a lateral connection, a merging network, and a classification network; the low frame rate branch network captures the spatial semantic information of the original video data; the high frame rate branch network captures the motion information of the original video data; the lateral connection performs feature fusion on each stage of the low frame rate branch network and the high frame rate branch network; the merging network merges the feature vectors output by the low frame rate branch network and the high frame rate branch network respectively to obtain a merged feature; the merged feature is input into the classification network to obtain an abnormal behavior recognition classification result.
[0180] Those of ordinary skill in the art can realize that each module in the abnormal behavior recognition device based on the lightweight dual frame rate network described in combination with the embodiments disclosed herein can be implemented in whole or in part by software, hardware, and their combination. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0181] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various deformations or modifications within the scope of the claims, which does not affect the essence of the present invention. The above preferred features can be combined arbitrarily without conflict.
Claims
1. An abnormal behavior recognition method based on a lightweight dual-frame rate network, characterized in that, it includes: Input the original video data into a lightweight dual-frame rate convolutional neural network. Among them, the low-frame rate branch network of the lightweight dual-frame rate convolutional neural network captures the spatial semantic information of the original video data; the high-frame rate branch network of the lightweight dual-frame rate convolutional neural network captures the motion information of the original video data; Use horizontal connections to perform feature fusion on each stage of the low-frame rate branch network and the high-frame rate branch network; Merge the feature vectors output by the low-frame rate branch network and the high-frame rate branch network respectively to obtain a merged feature; Input the merged feature into a classifier to obtain an abnormal behavior recognition classification result; The low-frame rate branch network includes five cascaded generalized convolutional layers, namely: the first generalized convolutional layer A1, the second generalized convolutional layer A2, the third generalized convolutional layer A3, the fourth generalized convolutional layer A4, and the fifth generalized convolutional layer A5; Among them: The first generalized convolutional layer A1 includes: 1 3×3×3 convolutional layer and 1 3×3×3 max pooling layer; The second generalized convolutional layer A2 sequentially includes: 1 branch 1 and 3 branch 2s; The third generalized convolutional layer A3 sequentially includes: 1 branch 1 and 7 branch 2s; The fourth generalized convolutional layer A4 sequentially includes: 1 branch 1 and 3 branch 2s; The fifth generalized convolutional layer A5 sequentially includes: 1 1×1×1 convolutional layer and 1 8×1×1 average pooling layer; The branch 1 is divided into two paths, among which: The first path includes: 1 3×3×3 depth convolutional layer with a stride of 2, 2 batch normalizations, 1 1×1×1 convolutional layer, and 1 activation layer; The second path includes: 2 1×1×1 convolutional layers, 3 batch normalizations, 2 activation layers, 1 3×3×3 depth convolutional layer with a stride of 2, and 1 squeeze-and-excitation module; The squeeze-and-excitation module includes: 1 adaptive average pooling layer, 2 fully connected layers, and 2 activation layers; Concatenate the feature vectors output by the two paths in the channel dimension, and output after channel shuffle operation; The branch 2 is divided into two paths through channel splitting operation, that is, the feature vector input to the branch 2 is evenly divided into two parts according to the number of channels, one part is used as the input of the first path, and the other part is used as the input of the second path, among which: The first path includes: an identity mapping shortcut connection; The second path includes: 2 1×1×1 convolutional layers, 3 batch normalizations, 2 activation layers, 1 3×3×3 depth convolutional layer with a stride of 1, and 1 squeeze-and-excitation module; The squeeze-and-excitation module includes: 1 adaptive average pooling layer, 2 fully connected layers, and 2 activation layers; Concatenate the feature vectors output by the two paths in the channel dimension, and output after channel shuffle operation.
2. The abnormal behavior recognition method based on a lightweight dual-frame rate network according to claim 1, characterized in that, The low-frame-rate branch network samples video frame images at a low frame rate with an interval of 16 frames; the high-frame-rate branch network samples video frame images at a high frame rate with an interval of 2 frames.
3. The abnormal behavior recognition method based on a lightweight dual-frame rate network according to claim 1, characterized in that the high-frame-rate branch network includes five cascaded generalized convolutional layers, namely the first generalized convolutional layer B1, the second generalized convolutional layer B2, the third generalized convolutional layer B3, the fourth generalized convolutional layer B4, and the fifth generalized convolutional layer B5; wherein: The first generalized convolutional layer B1 sequentially includes: 1 convolutional layer of 3×3×3 and 1 max-pooling layer of 3×3×3; The second generalized convolutional layer B2 sequentially includes: 1 branch 1 and 3 branch 2s; The third generalized convolutional layer B3 sequentially includes: 1 branch 1 and 7 branch 2s; The fourth generalized convolutional layer B4 sequentially includes: 1 branch 1 and 3 branch 2s; The fifth generalized convolutional layer B5 sequentially includes: 1 convolutional layer of 1×1×1 and 1 average-pooling layer of 8×1×1.
4. The abnormal behavior recognition method based on a lightweight dual-frame rate network according to claim 3, characterized in that the branch 1 is divided into two paths, wherein: The first path includes: 1 depth convolutional layer of 3×3×3 with a stride of 2, 2 batch normalizations, 1 convolutional layer of 1×1×1, and 1 activation layer; The second path includes: 2 convolutional layers of 1×1×1, 3 batch normalizations, 2 activation layers, 1 depth convolutional layer of 3×3×3 with a stride of 2, and 1 squeeze-and-excitation module; The squeeze-and-excitation module includes: 1 adaptive average-pooling layer, 2 fully connected layers, and 2 activation layers; The feature vectors output by the two paths are concatenated in the channel dimension, and after concatenation, they are output through a channel shuffle operation; The branch 2 is divided into two paths through a channel splitting operation, that is, the feature vector input to the branch 2 is evenly divided into two parts according to the number of channels, one part is used as the input of the first path, and the other part is used as the input of the second path, wherein: The first path includes: an identity mapping shortcut connection; The second path includes: 2 convolutional layers of 1×1×1, 3 batch normalizations, 2 activation layers, 1 depth convolutional layer of 3×3×3 with a stride of 1, and 1 squeeze-and-excitation module; The squeeze-and-excitation module includes: 1 adaptive average-pooling layer, 2 fully connected layers, and 2 activation layers; The feature vectors output by the two paths are concatenated in the channel dimension, and after concatenation, they are output through a channel shuffle operation.
5. An abnormal behavior recognition method based on a lightweight dual-frame rate network according to claim 1, characterized in that the lateral connection includes four convolutional layers, namely the first convolutional layer C1, the second convolutional layer C2, the third convolutional layer C3, and the fourth convolutional layer C4, wherein: The first convolutional layer C1 includes: 1 convolutional layer of 5×1×1 with a stride of 4; The second convolutional layer C2 includes: 1 convolutional layer of 5×1×1 with a stride of 4; The third convolutional layer C3 includes: 1 convolutional layer of 5×1×1 with a stride of 4; The fourth convolutional layer C4 includes: a 5×1×1 convolutional layer with a stride of 4; Match the feature sizes before fusion, including: The feature size of the low frame rate branch network is {T, S 2 , C} The feature size of the high frame rate branch network is {αT,S 2 ,βC} T represents the time length, S 2 represents the height and width of the feature map, α represents the ratio of the sampling density of the high frame rate branch network to the sampling density of the low frame rate branch network, β represents the ratio of the number of channels of the high frame rate branch network to the number of channels of the low frame rate branch network, and C represents the number of channels; Perform 3D convolution on the features of the high frame rate branch network, with the number of output channels being 2βC and the stride being α; The output result of the high frame rate branch network is fused into the low frame rate branch network through splicing, including: The output of the first generalized convolutional layer B1 serves as the input of the first convolutional layer C1, and the output of the first convolutional layer C1 and the output of the first generalized convolutional layer A1 are fused and then serve as the input of the second generalized convolutional layer A2; The output of the second generalized convolutional layer B2 serves as the input of the second convolutional layer C2, and the output of the second convolutional layer C2 and the output of the second generalized convolutional layer A2 are fused and then serve as the input of the third generalized convolutional layer A3; The output of the third generalized convolutional layer B3 serves as the input of the third convolutional layer C3, and the output of the third convolutional layer C3 and the output of the third generalized convolutional layer A3 are fused and then serve as the input of the fourth generalized convolutional layer A4; The output of the fourth generalized convolutional layer B4 serves as the input of the fourth convolutional layer C4, and the output of the fourth convolutional layer C4 and the output of the fourth generalized convolutional layer A4 are fused and then serve as the input of the fifth generalized convolutional layer A5.
6. A method for abnormal behavior recognition based on a lightweight dual frame rate network according to claim 1, characterized in that Merge the feature vectors finally output by the two branch networks as the input of the classifier to obtain the classification result of abnormal behavior recognition, including: After the two branch networks undergo convolutional operations, the vectors containing feature parameters finally output are concatenated and input into the fully connected layer; The fully connected layer inputs the calculated feature vectors into the Sigmoid regression layer for regression calculation to obtain the classification result.
7. An apparatus for abnormal behavior recognition based on a lightweight dual frame rate network, characterized in that it includes: An extraction video device that acquires a video; A processor that performs loading processing on the video according to the video to obtain an abnormal behavior recognition result; A memory that stores the lightweight dual frame rate network model executed by the processor, and the machine-readable instructions of the memory are executed by the processor to execute the method according to any one of claims 1-6.
8. A system for abnormal behavior recognition based on a lightweight dual frame rate network, characterized in that it includes: A low frame rate branch network that captures the spatial semantic information of the original video data; A high frame rate branch network that captures the motion information of the original video data; A horizontal connection that fuses the features of each stage of the low frame rate branch network and the high frame rate branch network; A merging network that merges the feature vectors respectively output by the low frame rate branch network and the high frame rate branch network to obtain merged features; A classification network that inputs the merged features into the classification network to obtain the classification result of abnormal behavior recognition; The low frame rate branch network includes five cascaded generalized convolutional layers, namely: the first generalized convolutional layer A1, the second generalized convolutional layer A2, the third generalized convolutional layer A3, the fourth generalized convolutional layer A4, and the fifth generalized convolutional layer A5; wherein: The first generalized convolutional layer A1 includes: a 3×3×3 convolutional layer and a 3×3×3 max pooling layer; The second generalized convolutional layer A2 sequentially includes: one branch 1 and three branch 2s; The third generalized convolutional layer A3 sequentially includes: one branch 1 and seven branch 2s; The fourth generalized convolutional layer A4 sequentially includes: one branch 1 and three branch 2s; The fifth generalized convolutional layer A5 sequentially includes: a 1×1×1 convolutional layer and an 8×1×1 average pooling layer; The branch 1 is divided into two paths, where: The first path includes: a 3×3×3 depth convolutional layer with a stride of 2, two batch normalizations, a 1×1×1 convolutional layer, and an activation layer; The second path includes: two 1×1×1 convolutional layers, three batch normalizations, two activation layers, a 3×3×3 depth convolutional layer with a stride of 2, and a squeeze-and-excitation module; The squeeze-and-excitation module includes: an adaptive average pooling layer, two fully connected layers, and two activation layers; The feature vectors output from the two paths are concatenated in the channel dimension, and after concatenation, they are output through a channel shuffle operation; The branch 2 is divided into two paths through a channel splitting operation, that is, the feature vector input to branch 2 is evenly divided into two parts according to the number of channels, one part is used as the input of the first path, and the other part is used as the input of the second path, where: The first path includes: an identity mapping shortcut connection; The second path includes: two 1×1×1 convolutional layers, three batch normalizations, two activation layers, a 3×3×3 depth convolutional layer with a stride of 1, and a squeeze-and-excitation module; The squeeze-and-excitation module includes: an adaptive average pooling layer, two fully connected layers, and two activation layers; The feature vectors output from the two paths are concatenated in the channel dimension, and after concatenation, they are output through a channel shuffle operation.
Citation Information
Patent Citations
Real-time behavior identification method based on time attention mechanism and double-flow network
CN113283298A
Real-time intelligent video monitoring abnormal behavior analysis method based on slowfast double-frame rate
CN113743306A