Behavior recognition method based on lightweight model
By using a lightweight model based on MobileNetV2 in the computer vision behavior recognition model and inserting a multi-scale one-way spatio-temporal migration-feature fusion module at each bottleneck layer, the existing model has many parameters, slow computing and poor performance of the time migration module, and efficient behavior recognition is achieved.
Patent Information
- Application Number
- CN202510067880.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-06-06
AI Technical Summary
The existing computer vision behavior recognition model has many parameters, large size, and complex algorithms, which leads to slow computing speed and is not suitable for deployment on edge devices. The time migration module dilutes the connection between adjacent time step characteristics during long migration, and performs poorly.
The lightweight behavior recognition model based on MobileNetV2 is adopted. By inserting a multi-scale one-way spatio-temporal migration-feature fusion module after 1×1 convolution of each bottleneck layer, including the improved TSM module for two-level cascade structure, the model's ability to extract timing information is enhanced.
It realizes small parameters and fast computing speed, enhances the ability to extract local features of feature maps, and improves the accuracy of video object behavior recognition.
Smart Images

Figure CN120108032A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a behavior recognition method based on a lightweight model. Background Art
[0002] A visual behavior recognition task applied to computers aims to enable intelligent monitoring equipment to judge the behavior of people in the scene based on real-time monitoring images. Deep learning models are often used to achieve this. However, on a macro level, most of the existing models have many parameters and large volumes. The complex algorithms lead to slow calculation speeds, which are not conducive to deployment on edge devices. On a micro level, the time migration modules designed for these models using time migration modules will dilute the connection between features between adjacent time steps when facing long-term migration, resulting in poor performance. Specifically, for example, the traditional TSM-related model has the defect that it cannot migrate long time steps, the connection in the time dimension is weak, and the number of parameters and volume are relatively large. Summary of the invention
[0003] The purpose of the present invention is to provide a behavior recognition method based on a lightweight model, which not only has a small number of parameters and a fast calculation speed, but also enhances the ability to extract local features of feature maps, so that the behavior recognition accuracy of video objects is high.
[0004] The present invention is achieved through the following technical solutions:
[0005] A behavior recognition method based on a lightweight model includes the following steps:
[0006] S1: First, convert the video into video frames;
[0007] S2: Input the video frames into a lightweight behavior recognition model that uses MobileNetV2 as the basic model and includes a multi-scale unidirectional spatiotemporal migration-feature fusion module;
[0008] S3: The input video frame is inferred by the lightweight behavior recognition model and finally the output result of the recognized behavior type is obtained.
[0009] Preferably, the multi-scale unidirectional spatiotemporal migration-feature fusion module in step S2 is composed of a two-stage cascade of improved TSM modules, which sequentially include a multi-scale unidirectional time migration module and a space migration module.
[0010] Furthermore, the operation process of the TSM module is divided into two steps: time migration and convolution. After the TSM module is cascaded in two stages, the model's ability to extract time series information can be enhanced without changing the original convolution kernel.
[0011] Furthermore, compared with the computational amount K×K×H×W×C×N generated by the standard convolution edge extraction feature edge combination, the lightweight behavior recognition model reduces the computational amount by:
[0012]
[0013] Where K×K is the size of the convolution kernel, C is the number of channels, H×W is the input video frame size, and N is the number of input video frames.
[0014] Preferably, the step S2 specifically includes the following steps:
[0015] (1) The input video frame first passes through point convolution to expand the number of channels in the bottleneck layer;
[0016] (2) After point convolution, the multi-scale unidirectional time migration module is used to perform time migration and feature fusion;
[0017] (3) After passing through the multi-scale unidirectional time migration module, it passes through the spatial migration module to perform spatial migration and feature fusion;
[0018] (4) The spatial migration module and deep convolution are used to independently apply convolution operations to each input channel to effectively extract local features;
[0019] (5) The local features after the deep convolution are projected back to the low-dimensional representation through a linear 1×1 convolution again;
[0020] (6) According to the above steps, the video frame passes through a total of 6 bottleneck layers and then passes through a convolutional layer, a pooling layer, and a linear layer in sequence to output the behavior inference result.
[0021] The present invention has the following beneficial effects:
[0022] (1) The lightweight behavior recognition model based on MobileNetV2 and including a multi-scale unidirectional spatiotemporal migration-feature fusion module in the present invention can accurately judge the behavior pattern of video objects. Compared with the prior art, the overall model parameters are reduced and the operation is faster, which solves the problem of sparse temporal feature correlation after the prior art model is migrated in the time dimension, and enhances the ability to extract local features of the feature map, thereby improving recognition accuracy;
[0023] (2) The multi-scale unidirectional spatiotemporal migration-feature fusion module can independently migrate temporal information. The cascaded modules will effectively modify the migrated tensors and effectively fuse the spatiotemporal information before the secondary migration, thus avoiding the one-sidedness and fragmentation of temporal migration. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1It is a schematic diagram of the local structure of the lightweight behavior recognition model in the present invention;
[0025] Figure 2 It is the working principle diagram of the TSM module in the present invention;
[0026] Figure 3 Schematic diagram of depth-wise separable convolution in the present invention;
[0027] Figure 4 This is a working principle diagram of the multi-scale unidirectional time migration module in the present invention;
[0028] Figure 5 This is a working principle diagram of the space migration module in the present invention. DETAILED DESCRIPTION
[0029] The present invention is further described in detail below in conjunction with specific embodiments, which are intended to explain the present invention rather than to limit it.
[0030] The present invention provides a behavior recognition method based on a lightweight model, comprising the following steps:
[0031] S1: First, convert the video into video frames, the mode of the video frames is "RGB";
[0032] S2: Using MobileNetV2 as the basic model, a multi-scale unidirectional spatiotemporal migration-feature fusion module is inserted after the 1×1 convolution of each bottleneck layer of MobileNetV2 to build a lightweight behavior recognition model;
[0033] S3: Input the video frames into the lightweight action recognition model;
[0034] S4: The input video frame first passes through the point convolution to expand the number of channels in the bottleneck layer;
[0035] S5: After point convolution, it passes through the multi-scale unidirectional time migration module to perform time migration and feature fusion;
[0036] S6: After passing through the multi-scale unidirectional time migration module, it passes through the spatial migration module for spatial migration and feature fusion;
[0037] S7: The spatial migration module and then the deep convolution are used to independently apply convolution operations to each input channel to effectively extract local features;
[0038] S8: The local features after the depthwise convolution are projected back to the low-dimensional representation through a linear 1×1 convolution again, which helps to maintain low computational cost and memory usage;
[0039] S9: The video frame passes through a total of 6 bottleneck layers according to steps S4 to S8, and then passes through a convolutional layer, a pooling layer, and a linear layer in sequence to output the behavior inference result. The output result is the recognized behavior type.
[0040] The lightweight behavior recognition model is constructed with MobileNetV2 as the basic model. A multi-scale unidirectional spatiotemporal migration-feature fusion module is inserted after the 1×1 convolution of each bottleneck layer of MobileNetV2. Its structure is as follows: Figure 1 shown.
[0041] The multi-scale unidirectional spatiotemporal migration-feature fusion module is composed of two-level cascades of TSM modules. The principle of the TSM module is as follows: Figure 2 As shown, Figure 2 (a) represents a video stream without time migration. C is the channel dimension, representing different channels of the video stream. T is the time dimension. Each layer of color represents a time step in the time series. Figure 2 (b) represents the video stream after time migration, in which some channels are moved up or down by one time step in the time dimension, while the rest are not moved; the features of the behavior in any three consecutive time steps in the original time dimension are extracted using one convolution, that is, only 2D convolution is needed to realize the spatiotemporal feature extraction of the action in one time step. Therefore, the operation process of the TSM module can be divided into two steps: time migration and convolution.
[0042] Assume there is a 1×3 convolution kernel W = (ω 1 ,ω 2 ,ω 3 ) and a 3-channel infinite video stream X k =(x i-k,j ,x i,j ,x i+k,j ) n×n , (n=∞), where k represents the number of time migrations. After one time migration, the video stream becomes X 1 =(x i-1,1 ,x i,2 ,x i+1,3 ) n×3 , (n=∞), and then perform convolution calculation, where y i Represents the convolution operation process, and Y represents the calculation result:
[0043] y i =ω 1 x i-1,1 +ω 2 x i,2 +ω 3 x i+1,3
[0044]
[0045] From the calculation process, it can be seen that the first step of the TSM module operation does not incur any computational cost, while the second step will generate a certain amount of computation. However, compared with the reduction in computational cost caused by the convolution method from 3D to 2D, the slight increase in computational cost is acceptable.
[0046] Then, the TSM module is simply cascaded in two stages to enhance the model's ability to extract time series information and achieve effective fusion of spatiotemporal information. Assume that there is an infinitely long one-dimensional vector X and a convolution kernel W of size 1×3 = (w 1 w 2 w 3 ), assuming that the vector Z has undergone one migration, where α, β, and γ are weight factors:
[0047] Z=αX -1 +βX 0 +γX +1
[0048] Then after two cascades, the convolution result Y is:
[0049] Y=w 1 Z -1 +w 2 Z 0 +w 3 Z +1
[0050] Y=w a X -2 +w b X -1 +w c X 0 +w d X +1 +w e X +2
[0051] in:
[0052] w a =αw 1
[0053] w b =βw 1 +αw 2
[0054] w c =γw 1 +βw 2 +αw 3
[0055] w d =γw 2 +βw3
[0056] w e =γw 3
[0057] Then, through inverse decoupling, the following conclusions can be drawn:
[0058] y i =w a x i-3 +w b x i-1 +w c x i +w d x i+1 +w e x i+2
[0059] This realizes the infinite length of the one-dimensional vector X and the new convolution kernel W′=(w a w b w c w d w e ), that is, without changing the original convolution kernel, through a simple secondary cascade, a convolution kernel of size 1×3 can achieve a convolution of size 1×5.
[0060] The above-mentioned twice-cascaded TSM module is added to the 1×1 convolution of each bottleneck layer in MobileNetV2 to form a multi-scale unidirectional spatiotemporal migration-feature fusion module. MobileNetV2 retains the core idea of MobileNets, that is, using depth-wise separable convolution to replace traditional standard convolution, thereby reducing the amount of calculation and parameters.
[0061] like Figure 3 As shown in Figure 2, the depth-separable convolution is to convert the standard convolution into point convolution as Figure 3 (a) and depthwise convolution as Figure 3 (b) The main function of deep convolution is to use a single filter to extract features from each input channel, while point convolution applies 1×1 convolution to combine the outputs of deep convolution; the amount of computation generated by this method is K×K×C×H×W+C×N×H×W, where K×K is the size of the convolution kernel, C is the number of channels, H×W is the input video frame size, and N is the number of input video frames. Compared with the amount of computation generated by standard convolution while extracting features and combining them, K×K×H×W×C×N, the amount of computation reduced is:
[0062]
[0063] The specific working principle of the multi-scale unidirectional spatiotemporal migration-feature fusion module is as follows: Figure 4 and Figure 5 As shown, Figure 4 This is a working diagram of the first-level TSM module, i.e., the multi-scale unidirectional time migration module. Figure 4 (a) is the initial input video stream, each column is a channel, each line is a frame, Figure 4 It is assumed that there are 4 frames and 6 channels; Figure 4 (b) To perform multi-scale unidirectional time migration, one channel migrates 1 time step and the other channel migrates 2 time steps; Figure 4 (c) After a time migration, the original input video stream and the migrated video stream are used for feature fusion.
[0064] Figure 5 This is a working diagram of the second-level TSM module, namely the space migration module. Figure 5 (a) is a feature image segment processed by time migration. Each frame of the feature image has 3 channels. Figure 5 (b) To perform spatial migration, each channel of a feature map migrates one time step in any direction of the four directions of up, down, left, and right; Figure (c) shows that after the spatial migration, the originally input video stream and the migrated feature map sequence are used for feature fusion. Finally, the corresponding behavior conclusion is output after the process of convolution layer, pooling layer, and linear layer. The above content is a further detailed description of the present invention in combination with specific embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these inventions. For ordinary technicians in the technical field of the present invention, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should be regarded as belonging to the scope of protection of the present invention.
Claims
1. A behavior recognition method based on a lightweight model, characterized in that: The following steps are involved: S1: First, convert the video into video frames; S2: Input the video frames into a lightweight behavior recognition model that uses MobileNetV2 as the basic model and includes a multi-scale unidirectional spatiotemporal migration-feature fusion module; S3: The input video frame is inferred by the lightweight behavior recognition model and finally the output result of the recognized behavior type is obtained.
2. According to the lightweight model-based behavior recognition method of claim 1, it is characterized in that: The multi-scale unidirectional spatiotemporal migration-feature fusion module in step S2 is composed of a two-stage cascade of improved TSM modules, which sequentially include a multi-scale unidirectional time migration module and a space migration module.
3. The behavior recognition method based on a lightweight model according to claim 2 is characterized in that: The operation process of the TSM module is divided into two steps: time migration and convolution. After the TSM module is cascaded in two stages, the model's ability to extract time series information can be enhanced without changing the original convolution kernel.
4. The behavior recognition method based on a lightweight model according to claim 2 is characterized in that: Compared with the computational amount K×K×H×W×C×N generated by the standard convolution edge extraction feature edge combination, the lightweight behavior recognition model reduces the computational amount by: Where K×K is the size of the convolution kernel, C is the number of channels, H×W is the input video frame size, and N is the number of input video frames.
5. The behavior recognition method based on a lightweight model according to claim 1, characterized in that: The step S2 specifically includes the following steps: (1) The input video frame first passes through point convolution to expand the number of channels in the bottleneck layer; (2) After point convolution, the multi-scale unidirectional time migration module is used to perform time migration and feature fusion; (3) After passing through the multi-scale unidirectional time migration module, it passes through the spatial migration module to perform spatial migration and feature fusion; (4) The spatial migration module and deep convolution are used to independently apply convolution operations to each input channel to effectively extract local features; (5) The local features after the deep convolution are projected back to the low-dimensional representation through a linear 1×1 convolution again; (6) According to the above steps, the video frame passes through a total of 6 bottleneck layers and then passes through a convolutional layer, a pooling layer, and a linear layer in sequence to output the behavior inference result.