Lightweight behavior recognition method and system based on global frequency domain pooling algorithm
Patent Information
- Application Number
- CN202311268267.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-28
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2043-09-28
AI Technical Summary
这种做法虽然实现了较大程度的信息压缩,减少了后续全连接层的参数量和浮点运算量,但当不同通道的特征图均值相同时,原本表示不同特征信息的特征图就会表达出相同的语义,使得压缩后的均值特征缺乏多样性,从而产生信息损失和信息冗余的问题
[0036] 1. A novel R3D behavior recognition network framework composed of residual blocks from universities was used, which lightened the network model by reducing the number of parameters and computation, while effectively realizing the fusion of short, medium and long time series of time information.
Smart Images

Figure CN117475349B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer vision and relates to the design and application of human behavior recognition methods. Background Technology
[0002] With the widespread adoption of smartphones and portable devices, and the booming development of short video apps, everyone can become a short video producer. The content of these videos encompasses all aspects of people's lives. Among these, the technology that focuses on people and aims to analyze the types of actions performed when people interact with each other and with objects in videos is called Human Action Recognition (HAR). HAR is an important research direction for video understanding using pattern recognition technology, and it has significant applications in fields such as intelligent security, human-computer interaction, and smart education.
[0003] In recent years, numerous deep learning-based action recognition methods have emerged, with mainstream methods broadly categorized into three types: Two-Stream based action recognition methods, RNN-based action recognition methods, and 3D-ConvNet-based action recognition methods. Currently, some research still uses optical flow to describe motion information in videos, but this places high demands on computation and storage, hindering large-scale training and deployment. Therefore, 3D-ConvNet has become an important tool for modeling temporal information in videos. In 3D-ConvNet-based action recognition methods, many algorithms such as Res3D, I3D, R(2+1)D, and TSM are largely based on ResNet as their framework. When transitioning from convolutional layers to fully connected layers, ResNet employs Global Average Pooling (GAP) to compress the feature map output from the final convolutional layer, retaining only a mean value to represent the high-level semantic information contained in that channel's feature map. While this approach achieves a high degree of information compression and reduces the number of parameters and floating-point operations in subsequent fully connected layers, when the mean of the feature maps in different channels is the same, the feature maps that originally represented different feature information will express the same semantics, resulting in a lack of diversity in the compressed mean features, thus causing information loss and information redundancy.
[0004] Therefore, existing methods cannot guarantee that, while reducing information loss and redundancy during average pooling, the risk of model overfitting can be reduced while keeping the model lightweight. Summary of the Invention
[0005] Purpose of the invention: To address the problems of information loss, information redundancy and network overfitting in global average pooling in 3D-ConvNet, a behavior recognition algorithm model based on global frequency domain pooling is proposed.
[0006] Technical solution: A lightweight behavior recognition method based on a global frequency domain pooling algorithm, including:
[0007] Obtain the video segments that need to be performed for action recognition, and perform preprocessing operations on the selected video segments to obtain images;
[0008] The preprocessed image is input into a human behavior recognition network model based on global frequency domain pooling for convolutional neural network training, and the corresponding human action classification results are output.
[0009] The human behavior recognition network model based on global frequency domain pooling includes an input layer, a 3D convolutional layer, a continuously stacked high-efficiency residual block (ERB), a global frequency domain pooling layer (GFDP), a fully connected layer, and a softmax output layer.
[0010] Furthermore, the preprocessing operation includes:
[0011] The video is kept in its original structure, and the video segment is decomposed frame by frame into more than n images, then consecutive images are randomly selected. Frame images are used as video frames input from the network.
[0012] The image is scaled up to a set ratio and then randomly cropped to the set size as input.
[0013] Furthermore, the efficient residual block (ERB) consists of two densely connected three-dimensional bottleneck convolutional blocks. The three-dimensional bottleneck convolutional block is composed of three convolutional layers with kernels of 3*1*1, 3*3*3, and 3*1*1 respectively. The first 3*1*1 convolution fits short-range temporal information and reduces the output channel dimension to 1 / 4 of the original. The middle 3*3*3 convolutional kernel extracts spatiotemporal information. Due to the shortening of its input and output dimensions, the number of parameters is greatly reduced. The last 3*1*1 convolutional layer increases the channel dimension to the original size and fits the temporal information again.
[0014] Furthermore, the benchmarks for lightweight network models include the number of parameters and computational cost. The formula for calculating the number of parameters in the current layer is as follows:
[0015] Parameters = k t ×k w ×k h ×c i ×c0+c0 (1)
[0016] The formula for calculating the computational cost (FLOPs) of the current layer is as follows:
[0017] FLOPs = 2 × k t ×k w ×k h ×t×w×h×c i ×c0 (2)
[0018] In the formula c i c0 is the number of input feature maps, and k is the number of output feature maps. t k w k h Let t represent the size of the convolution kernel in the three dimensions of time, width, and height, and let t, w, and h represent the time length, width, and height of the input feature map. Substitute the efficient residual block into the formula to calculate and verify its lightweight model effect.
[0019] Furthermore, the construction process of the Global Frequency Domain Pooling Layer (GFDP) is as follows: the input is the feature map output by the last convolutional layer of the CNN, and then the spectrum of the specified N frequency components is obtained through DCT, which is the basic feature information in the feature map. Finally, the different spectra are fused to obtain the compressed feature map frequency domain features.
[0020] Furthermore, the specific process of the Global Frequency Domain Pooling Layer (GFDP) is as follows:
[0021] Based on the introduced N frequency components, the expression for a single-channel GFDP is:
[0022]
[0023] In the formula n u ,n v It is the frequency value corresponding to the nth frequency component. It is the input single-channel feature map, DCT n express The nth spectral component; c(·) is a compensation coefficient, expressed as: N is the dimension of the input data, i represents the horizontal coordinate of the spatial location of the frequency component, j represents the vertical coordinate of the spatial location of the frequency component, and W and H represent the width and height of the input feature map.
[0024] Secondly, because matrix operations in computers are more time-efficient than traversal operations, in order to perform matrix operations, therefore... Since it is independent of the variable n, the input x of equation (3) is... 2d Decoupled from the DCT basis functions, the expression for the DCT basis functions is:
[0025]
[0026] Therefore, the decoupled single-channel GFDP expression is:
[0027]
[0028] Finally, due to the DCT basis functions Only with frequency component nu ,n v The GFDP of the entire feature map is related to the spatial location i,j, but not to the channel. Therefore, when calculating the GFDP of the entire feature map, we can use the matrix multiplication concept to calculate the weight matrix only once, and then multiply it with the feature map. The matrix form of the GFDP of the entire feature map is as follows:
[0029]
[0030] In the formula X 2d , They represent and The matrix form, as shown in formula (6), is obtained by adding the cardinal matrices of different frequencies to get the final DCT weight matrix, which is then combined with X. 2d Dot product, to get input X 2d After weighting, the multi-spectral feature matrix is finally compressed by channel summation to output the final feature information, thus completing the GFDP operation of the feature map.
[0031] A lightweight behavior recognition system based on global frequency domain pooling algorithm includes: a data acquisition and preprocessing module, and a human behavior recognition network module based on global frequency domain pooling.
[0032] The data acquisition and preprocessing module acquires the video segments that need to be performed for action recognition, and performs preprocessing operations on the selected video segments to obtain images.
[0033] The human behavior recognition network module based on global frequency domain pooling inputs the preprocessed image into the human behavior recognition network model based on global frequency domain pooling for convolutional neural network training, and outputs the corresponding human action classification results.
[0034] The human behavior recognition network module based on global frequency domain pooling includes an input layer, a 3D convolutional layer, a continuously stacked high-efficiency residual block (ERB), a global frequency domain pooling layer (GFDP), a fully connected layer, and a softmax output layer.
[0035] Beneficial effects:
[0036] 1. A novel R3D behavior recognition network framework composed of residual blocks from universities was used, which lightened the network model by reducing the number of parameters and computation, while effectively realizing the fusion of short, medium and long time series of time information.
[0037] 2. By extending the GAP operation to the frequency domain using two-dimensional DCT transformation, a novel global frequency domain pooling is proposed, thereby introducing more frequency components to increase the specificity between feature channels and reduce information redundancy after information compression; a batch normalization strategy of convolutional layers is introduced to extend the fully connected layer of the behavior recognition model to optimize data distribution and further improve recognition accuracy. Attached Figure Description
[0038] Figure 1 It is a detailed construction process of multiple efficient residual blocks (ERBs);
[0039] Figure 2 This is a detailed structure diagram of Global Frequency Domain Pooling (GFDP);
[0040] Figure 3 This is a diagram showing the combined effect of Global Average Pooling (GAP) and Global Frequency Domain Pooling (GFDP).
[0041] Figure 4 This is a diagram of a lightweight behavior recognition network model based on global frequency domain pooling. Detailed Implementation
[0042] The technical solution of the present invention will be further described below with reference to the accompanying drawings.
[0043] This invention provides a technical solution: a lightweight behavior recognition method based on a global frequency domain pooling algorithm, the specific technical solution of which is as follows:
[0044] Obtain the video clips that need to be performed on motion recognition;
[0045] Perform preprocessing operations on the selected video segments;
[0046] The preprocessing operations for the video segments include:
[0047] The video is kept in its original structure and the video segment is decomposed frame by frame into more than 16 images. Then, 8 consecutive frames are randomly selected as the video frames for network input.
[0048] The image is scaled proportionally to 171*128 and then randomly cropped to a size of 112*112 as input.
[0049] The preprocessed eight consecutive images are input into a human behavior recognition network based on global frequency domain pooling for training of a convolutional neural network, and the corresponding human action classification results are output.
[0050] The human behavior recognition network model based on global frequency domain pooling mainly consists of an input layer, a 3D convolutional layer, a continuously stacked efficient residual block (ERB), a global frequency domain pooling layer (GFDP), a fully connected layer, and a softmax output layer.
[0051] The construction process of the human behavior recognition network model based on global frequency domain pooling includes:
[0052] Construct the efficient residual block (ERB) and global frequency domain pooling structure (GFDP);
[0053] The efficient residual block (ERB) consists of two densely connected 3D bottleneck convolutional blocks, each composed of three convolutional layers with kernels of 3*1*1, 3*3*3, and 3*1*1 respectively. The first 3*1*1 convolution fits short-range temporal information and reduces the output channel dimension to one-quarter of its original size. The middle 3*3*3 convolutional kernel is used to extract spatiotemporal information; due to its shortened input and output dimensions, the number of parameters is significantly reduced. The final 3*1*1 convolutional layer increases the channel dimension back to its original size and fits the temporal information again.
[0054] The construction process of the high-efficiency residual block is as follows: Figure 1 As shown, the main function of this module is to fully integrate short, medium, and long-term time-series information based on a lightweight human behavior recognition network model. The main reference indicators for the lightweight network model are the number of parameters and computational cost. The calculation formula for the current layer's parameters is as follows:
[0055] Parameters = k t ×k w ×k h ×c i ×c0+c0 (1)
[0056] The formula for calculating the computational cost (FLOPs) of the current layer is as follows:
[0057] FLOPs = 2 × k t ×k w ×k h ×t×w×h×c i ×c0 (2)
[0058] In the formula c i c0 is the number of input feature maps, and k is the number of output feature maps. t k w k h Let be the size of the convolutional kernel in the three dimensions of time, width, and height, and let t, w, and h be the time length, width, and height of the input feature map. After substituting the efficient residual block g into the formula, it can be seen that the number of parameters and computational cost of the replaced convolutional block are reduced by about 8.5 times, effectively achieving the goal of lightweight model.
[0059] The construction process of the Global Frequency Domain Pooling (GFDP) structure is as follows:
[0060] The feature map is transformed to the frequency domain using DCT. While preserving the lowest frequency component (GAP), other frequency components are introduced to enrich the feature information, increase the specificity between channels, and reduce the redundancy of feature information. The structure of GFDP is as follows: Figure 2 As shown, the input is the feature map output by the last convolutional layer of the CNN. Then, the spectrum of the specified N frequency components is obtained through DCT, which is the basic feature information in the feature map. Finally, the different spectra are fused to obtain the compressed feature map frequency domain features.
[0061] GFDP employs an additive fusion method. Compared to the concatenated information fusion method, this approach does not increase the dimensionality of the compressed feature information, ensuring that the number of parameters and computational cost in subsequent fully connected layers are consistent with GAP. Regarding the computation method of GFDP, due to the multi-dimensional nature of the feature map, this paper reduces the redundant computation of DCT components through matrix dot multiplication. The specific process is as follows:
[0062] First, based on the introduced N frequency components, the expression for a single-channel GFDP is:
[0063]
[0064] In the formula n u ,n v It is the frequency value corresponding to the nth frequency component. It is the input single-channel feature map, DCT n express The nth spectral component.
[0065] Secondly, because matrix operations in computers are more time-efficient than traversal operations, in order to perform matrix operations, because... Since the input is independent of the variable n, this paper decouples the input in equation (3) from the DCT basis functions. The expression for the basis functions is:
[0066]
[0067] Therefore, the decoupled single-channel GFDP expression is:
[0068]
[0069] Finally, due to the basis functions Only with frequency component n u ,n vThe weight matrix is related to the spatial location i,j but not to the channel. Therefore, when calculating the GFDP of the entire feature map, we can use the matrix dot product approach, calculating the weight matrix only once and then multiplying it with the feature map. The matrix form of the GFDP of the entire feature map is as follows:
[0070]
[0071] In the formula X 2d , They represent and The matrix form. As shown in formula (6), the final DCT weight matrix is obtained by adding the cardinal matrices of different frequencies, and then... 2d Dot product, to get input X 2d After weighting, the multi-spectral feature matrix is finally compressed by channel summation to output the final feature information, thus completing the GFDP operation of the feature map.
[0072] As can be seen from the above calculation process, the difference between GAP and GFDP lies in their weight matrices. GAP contains only (0,0) components, making each position in the weight matrix a constant 1, thus easily leading to identical semantics after channel compression. GFDP, in addition to the (0,0) components, introduces other frequency components, altering the weight distribution across the weight matrix, significantly reducing the probability of identical semantics after channel compression. The compression effects of GAP and GFDP are shown below. Figure 3 As shown, taking two-channel feature maps as an example, after GAP, the two channels with different feature information express the same semantic information, resulting in information redundancy; while after GFDP introduces four frequency components, the two channels still retain their unique semantics, effectively suppressing the impact of information redundancy.
[0073] Input eight consecutive frames of images into this human behavior recognition network, such as... Figure 4 As shown, after several efficient residual blocks, global frequency domain pooling, and fully connected layers, the final behavior category is output by the softmax function.
[0074] Experimental section:
[0075] First, this paper introduces the public datasets and basic theoretical knowledge involved:
[0076] UCF101 dataset
[0077] This invention uses the UCF101 behavior recognition dataset for experimental verification. The videos in this dataset are mainly obtained from YouTube. The behavior categories involve five aspects: human-object interaction, simple body movements, human-human interaction, playing musical instruments, and sports, containing a total of 101 behavior categories. Each behavior category is divided into 25 groups, and each group contains 4 to 7 videos of one behavior, totaling 13,320 videos with a total duration of approximately 27 hours.
[0078] Regarding the partitioning of the UCF101 dataset, the official website provides three strategies for splitting the training and test sets. This invention selects the split01 method for partitioning, where the training set contains 9537 video sequences, accounting for approximately 70% of the total dataset, and the test set contains 3783 video sequences, accounting for approximately 30% of the total dataset.
[0079] ERB-Res3D network for behavior recognition
[0080] ERB-Res3D is a fusion of the Res3D network model and the ERB structure. The Res3D action recognition network is a three-dimensional extension of ResNet2D, transforming all 3×3 convolutional kernels in the residual structure into a 3×3×3 convolutional deep learning network model. Compared to Res3D, all cascaded 3D convolutions are replaced with ERB structures. Except for the first downsampling layer with a stride of 1×2×2, all other downsampling convolutional layers have a stride of 2×2×2. The network is divided into two types, ERB-Res3D-18 and ERB-Res3D-34, depending on the number of ERB structures.
[0081] A lightweight human recognition method based on global frequency domain pooling (ERB-Res3D-GDFP) based on ERB-Res3D includes the following steps:
[0082] 1. Obtain short videos of the behaviors to be identified;
[0083] You can use a camera or webcam to capture short, real-time videos of the behavior, or use an existing video dataset that includes the behavior: for example, the UCF101 dataset used in this case.
[0084] 2. Preprocess the acquired video to be identified;
[0085] The preprocessing process includes: decomposing the video frame by frame, randomly selecting 8 consecutive frames to form a new image sequence, and then performing image enhancement operations such as random flipping, random cropping, and noise reduction on the image sequence. In this implementation, the resolution of the input image sequence is controlled at 112×112.
[0086] 3. Image sequences were input into the behavior recognition network for ablation experiments, which included two phases: training and testing.
[0087] Training phase: The Xavier method is used to initialize network parameters, and training is performed using mini-batch data. The batch size is set to 12 based on the GPU performance. The learning rate is adjusted using a piecewise constant decay strategy, with a set value of 0.001 for the first 50 epochs, and then decayed to 1 / 10 of the original value every 20 epochs until the network converges. During backpropagation, the cross-entropy loss function is used to measure the network loss, and the Adam optimization algorithm is used to update the model parameters. L2 regularization and Batch Normalization (BN) are used to prevent overfitting.
[0088] Testing phase: First, 8 frames of images from a video in the test set are extracted at equal intervals. The extracted image sequence is processed by center cropping and used as network input. Then, 101 behavior classification scores are output through forward propagation. Finally, the category with the highest score is the prediction result.
[0089] Results analysis:
[0090] Table 1 compares the performance of the method used in this invention with other recognition methods on the UCF101 dataset. As can be seen from the table, our method demonstrates better recognition performance compared to classic networks and more recent networks. Both Efficient Residual Block (ERB) and Global Frequency Pooling (GFDP) can individually improve network performance. After applying the method used in this invention to the ERB-Res3D network structure, the computational cost of the network model is 3.5 GFlops, the number of parameters is 7.4M, and the final recognition accuracy is improved by 3.9% compared to the ERB-Res3D model and by 17.4% compared to the original Res3D model, achieving more accurate behavior recognition results efficiently.
[0091]
[0092] This invention, based on the Discrete Cosine Transform (DCT), recognizes GAP as a special case of eigenvalue decomposition in the frequency domain. It then introduces more frequency components to increase the specificity between feature channels and reduce information redundancy after compression. Furthermore, it introduces a batch normalization strategy for convolutional layers to reduce the risk of overfitting in the network model. This invention innovates upon traditional 3D convolution and global average pooling methods in behavior recognition, achieving both lightweight network models and efficient improvement in recognition accuracy.
Claims
1. A lightweight behavior recognition method based on a global frequency domain pooling algorithm, characterized in that, include: Obtain the video segments that need to be performed for action recognition, and perform preprocessing operations on the selected video segments to obtain images; The preprocessed image is input into a human behavior recognition network model based on global frequency domain pooling for convolutional neural network training, and the corresponding human action classification results are output. The human behavior recognition network model based on global frequency domain pooling includes an input layer, a 3D convolutional layer, a continuously stacked high-efficiency residual block (ERB), a global frequency domain pooling layer (GFDP), a fully connected layer, and a softmax output layer. The high-efficiency residual block (ERB) consists of two densely connected 3D bottleneck convolutional blocks, which are composed of three convolutional layers with kernels of 3*1*1, 3*3*3, and 3*1*1 respectively. The first 3*1*1 convolution fits short-range temporal information and reduces the output channel dimension to 1 / 4 of the original. The middle 3*3*3 convolutional kernel extracts spatiotemporal information. The last 3*1*1 convolutional layer increases the channel dimension to the original size and fits the temporal information again. The construction process of the global frequency domain pooling layer (GFDP) is as follows: the input is the feature map output by the last convolutional layer of the CNN, and then the spectrum of the specified N frequency components is obtained through DCT, which is the basic feature information in the feature map. Finally, the different spectra are fused to obtain the compressed feature map frequency domain features. The specific process of the Global Frequency Domain Pooling Layer (GFDP) is as follows: Based on the introduced N frequency components, the expression for a single-channel GFDP is: (3) In the formula It is the first The frequency value corresponding to each frequency component. It is the input single-channel feature map. express The One spectral component; It is a compensation coefficient, expressed as: N is the dimension of the input data. The horizontal axis represents the spatial location of the frequency component. The vertical coordinate represents the spatial location of the frequency component, and W and H represent the width and height of the input feature map. The input in equation (3) Decoupled from the DCT basis functions, the expression for the DCT basis functions is: (4) The decoupled single-channel GFDP expression is: (5) When calculating the GFDP of the entire feature map, the weight matrix is calculated only once, and then multiplied by the feature map. The GFDP matrix form of the entire feature map is as follows: (6) In the formula , They represent and The matrix form, as shown in formula (6), is obtained by adding the cardinal matrices of different frequencies to get the final DCT weight matrix, which is then combined with... Dot product, to get the input After weighting, the multi-spectral feature matrix is finally compressed by channel summation to output the final feature information, thus completing the GFDP operation of the feature map.
2. The lightweight behavior recognition method based on global frequency domain pooling algorithm according to claim 1, characterized in that, The preprocessing operations include: The video is kept in its original structure, and the video segment is decomposed frame by frame into n images, then consecutive images are randomly selected. n frames of images are used as video frames input to the network; The image is scaled up to a set ratio and then randomly cropped to the set size as input.
3. The lightweight behavior recognition method based on global frequency domain pooling algorithm according to claim 1, characterized in that, The key metrics for lightweight network models include the number of parameters and computational cost. The formula for calculating the number of parameters in the current layer is as follows: (1) The formula for calculating the computational cost (FLOPs) of the current layer is as follows: (2) In the formula The number of input feature maps. The number of output feature maps. , , The size of the convolution kernel in the three dimensions of time, width, and height. , , The time length, width, and height of the input feature map are used as inputs. The efficient residual block (ERB) is substituted into the formula for calculation to verify its lightweight model effect.
4. A lightweight behavior recognition system based on a global frequency domain pooling algorithm, characterized in that, include: The data acquisition and preprocessing module is based on a human behavior recognition network module with global frequency domain pooling. The data acquisition and preprocessing module acquires the video segments that need to be performed for action recognition, and performs preprocessing operations on the selected video segments to obtain images. The human behavior recognition network module based on global frequency domain pooling inputs the preprocessed image into the human behavior recognition network model based on global frequency domain pooling for convolutional neural network training, and outputs the corresponding human action classification results. The human behavior recognition network module based on global frequency domain pooling includes an input layer, a 3D convolutional layer, a continuously stacked high-efficiency residual block ERB, a global frequency domain pooling layer GFDP, a fully connected layer, and a softmax output layer. The efficient residual block (ERB) consists of two densely connected 3D bottleneck convolutional blocks. Each 3D bottleneck convolutional block is composed of three convolutional layers with kernels of 3*1*1, 3*3*3, and 3*1*1 respectively. The first 3*1*1 convolution fits short-range temporal information and reduces the output channel dimension to 1 / 4 of its original size. The middle 3*3*3 convolutional kernel extracts spatiotemporal information. The last 3*1*1 convolutional layer increases the channel dimension to its original size and fits the temporal information again. The construction process of the Global Frequency Domain Pooling Layer (GFDP) is as follows: the input is the feature map output by the last convolutional layer of the CNN, and then the spectrum of the specified N frequency components is obtained through DCT, which is the basic feature information in the feature map. Finally, the different spectra are fused to obtain the compressed feature map frequency domain features. The specific process of the Global Frequency Domain Pooling Layer (GFDP) is as follows: Based on the introduced N frequency components, the expression for a single-channel GFDP is: (3) In the formula It is the first The frequency value corresponding to each frequency component. It is the input single-channel feature map. express The One spectral component; It is a compensation coefficient, expressed as: N is the dimension of the input data. The horizontal axis represents the spatial location of the frequency component. The vertical coordinate represents the spatial location of the frequency component, and W and H represent the width and height of the input feature map. The input in equation (3) Decoupled from the DCT basis functions, the expression for the DCT basis functions is: (4) The decoupled single-channel GFDP expression is: (5) When calculating the GFDP of the entire feature map, the weight matrix is calculated only once, and then multiplied by the feature map. The GFDP matrix form of the entire feature map is as follows: (6) In the formula , They represent and The matrix form, as shown in formula (6), is obtained by adding the cardinal matrices of different frequencies to get the final DCT weight matrix, which is then combined with... Dot product, to get the input After weighting, the multi-spectral feature matrix is finally compressed by channel summation to output the final feature information, thus completing the GFDP operation of the feature map.