Lightweight grouping shuffling convolution personnel activity detection algorithm based on improved YOWOv3

By introducing lightweight grouped shuffled convolution (GSConv) optimization method into the video action detection model, the problem of difficulty in real-time deployment of existing models in complex indoor scenarios is solved, and the calculation complexity is reduced and feature interaction is improved.

CN120198965APending Publication Date: 2025-06-24CHANGCHUN UNIV OF TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510330634.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing video action detection model is difficult to deploy in real-time in complex indoor scenarios, with high computational complexity, making it difficult to achieve lightweight and efficient feature interaction.

Method used

A lightweight packet shuffling convolution (GSConv) optimization method is proposed, which replaces the standard 3×3 convolution through dynamic grouping strategy, combines channel shuffling operations to optimize the backbone network and feature pyramid module to reduce the computational complexity.

Benefits of technology

It significantly reduces the complexity of model calculation, improves the feature expression ability and real-time performance in complex indoor scenarios, and is suitable for the real-time detection requirements of edge devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The invention belongs to the technical field of image processing and computer vision, provides a lightweight grouping shuffling convolution personnel activity detection algorithm based on improved YOWOv3, and is applied to the field of real-time video behavior detection. According to the scheme, an original CSP residual module is reconstructed, and standard convolution is replaced by a dynamic grouping convolution and channel shuffling mechanism, so that the calculation complexity is remarkably reduced while the precision is ensured. The core innovation comprises the following steps: dividing an input channel into a plurality of subgroups for parallel processing, and reducing a parameter quantity to 1 / 4 of a standard convolution; feature interaction is realized through cross-group channel rearrangement, and the model expression ability is improved; convolution kernel parameters are multiplexed in the same network stage, memory occupation is further reduced, and efficient spatial-temporal feature fusion is realized in combination with a 3D convolution spatial-temporal detection head. According to the method, an edge calculation scene is adapted through lightweight design, real-time performance and detection precision are effectively balanced in a video analysis task, and an efficient solution is provided for dynamic behavior recognition in a low-resource environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision and deep learning, and specifically relates to a lightweight model optimization method for detecting human activity movements. By improving the convolution structure and feature interaction mechanism, the model calculation complexity is significantly reduced, and the method is suitable for edge computing scenarios such as smart homes and elderly care. Background Art

[0002] In recent years, video action detection technology based on deep learning has made significant progress, but real-time deployment in complex indoor scenes still faces severe challenges. Existing methods mainly rely on 3D convolutional neural networks (such as C3D, I3D) or two-stream networks (such as TSN, SlowFast) to model spatiotemporal features. Although they have achieved high accuracy on public datasets, they have the following inherent defects.

[0003] Traditional 3D convolution needs to process both spatial and temporal dimensions, resulting in an exponential increase in the number of parameters and computation. Taking the typical model YOWOv3 as an example, its backbone network uses 3D-ResNet34, which requires about 120 GFLOPs for a single inference and can only reach 15 FPS on edge devices (such as Jetson Nano), which cannot meet the needs of real-time monitoring. In addition, complex temporal modeling modules (such as LSTM and Transformer) further increase the computational burden and limit their application in low-power scenarios.

[0004] Existing methods generally use standard convolution (Conv2d / Conv3d) to extract spatial features. Although it can capture local texture information, the ability of cross-channel feature interaction is limited. Although grouped convolution (GroupConv) reduces the amount of calculation by channel grouping, the isolation of information between groups leads to a significant decrease in feature expression ability. For example, ShuffleNet alleviates this problem by shuffling channels, but its fixed grouping strategy cannot adapt to the dynamic changes of features at different levels, and it is easy to lose key motion clues in indoor scenes with frequent occlusion.

[0005] Compared with open environments, human activities (such as falling, gesture interaction) have the characteristics of small movement amplitude and more background interference (furniture occlusion, reflective surface). The feature extractor optimized by the existing model on simple scenes (Kinetics dataset) is difficult to effectively distinguish similar actions (such as "sitting down" and "bending down to pick up objects"), and is sensitive to changes in camera perspective, which restricts its practical application value. In summary, how to achieve lightweight and efficient feature interaction of indoor action detection models while ensuring detection accuracy is still a technical problem that needs to be solved urgently. Summary of the invention

[0006] Based on the YOWOv3 video action detection framework as the core foundation, aiming at the problems of computational redundancy and insufficient cross-channel feature interaction in its 2D feature extraction module, a lightweight grouped shuffle convolution (GSConv) optimization method is proposed. By reconstructing the backbone network, feature pyramid and residual module of YOWOv3, while retaining the spatio-temporal action detection ability of the original model, the computational complexity is significantly reduced, and the feature expression ability and real-time performance of the model in indoor complex scenarios (such as occlusion, multi-person interaction) are improved. Perform dynamic reconstruction of the backbone network, replace the computationally intensive standard 3×3 convolution in the 2D backbone network (such as DarkNet) of YOWOv3 with grouped shuffle convolution (GSConv), adapt different levels of features through a dynamic grouping strategy, and then embed GSConv in the bidirectional feature pyramid (BiFPN) of YOWOv3 to optimize the cross-scale feature fusion efficiency. Reconstruct the cross-stage partial network (CSP) residual module of YOWOv3, introduce grouped convolution and channel shuffle to solve the problem of inter-group information isolation. Based on the end-to-end characteristics of YOWOv3, design a progressive training strategy and a hardware acceleration deployment plan to achieve the balance of model efficiency and accuracy. To achieve the above objectives, the technical solutions of the present invention are as follows:

[0007] A person activity detection algorithm based on lightweight grouped shuffle convolution for improved YOWOv3, comprising the following steps: Step 1: Obtain a person activity action video dataset, modify the person activity action dataset according to the annotation format of the AVA dataset, and successfully apply it in the YOWOv3 network; Step 2: For the 2D backbone network of YOWOv3 (such as the Stage1 / 3 / 5 layers of DarkNet), identify its standard 3×3 convolutional layer, replace the original 3×3 standard convolution with a GSConv module, retain the number of input and output channels, but reduce the computational amount through grouped convolution, and dynamically select the maximum legal number of groups according to the number of input channels of the current layer to ensure uniform channel division (for example, when the input channels are 64, the number of groups can be selected as 1 / 2 / 4 / 8 / 16 / 32 / 64); Step 3: In the channel alignment layer (such as P5→P4, P4→P3) of the bidirectional feature pyramid (BiFPN) of YOWOv3, use GSConv to replace the original 1×1 standard convolution; insert a channel shuffle operation in the feature upsampling / downsampling path to enhance cross-level feature interaction; Step 4: Replace the second standard convolution of the residual module in the cross-stage partial network (CSP) of YOWOv3 with GSConv, share the grouped convolution weights in the residual path to avoid repeated calculations, and finally output the detection results of person activity actions.

[0008] The specific situation in the above Step 1 is as follows: (1)Using the publicly available Charades_v1_480 human activity dataset, the input video is temporally segmented, and the video stream is converted into a sequence of consecutive frames at a sampling rate of 30 frames per second; (2)Deploy the YOLOv5 object detection model for frame-by-frame object recognition. Screen out the main detection targets through human feature matching, clear the detection boxes of other irrelevant persons, and generate the bounding box coordinate data containing only the target persons; (3)Use the VIA annotation tool to perform fine-grained action annotation on the target persons, and record the action categories and corresponding spatial coordinates of the target persons in each frame. Construct a semantic system containing 12 basic actions (Stand, Lie-down, Sit, Tidy, Cook, Play-phone, Call, Eat-Drink, WatchTV, Open-Close-door, HandShakes-Hugs, Fall-Down), and encode the expert annotation results into a standardized label format suitable for neural network training; (4)Jointly input the preprocessed video frame sequence and the detection box coordinates output by YOLOv5 into the DeepSort multi-object tracking algorithm, and realize the cross-frame identity ID association of the target persons through appearance feature matching and trajectory prediction, and establish the individual behavior trajectories in the continuous time series.

[0009] The specific situation in step three is as follows: (1)Replace the 1×1 standard convolution (such as the P5→P4, P4→P3 paths) used to adjust the number of channels in BiFPN with the GSConv module, compress the computational amount through group convolution, and enhance cross-level feature interaction by using channel shuffle; (2)Insert channel shuffle operations in the up / down sampling paths of BiFPN to force the fusion of feature information at different scales (such as high-level semantics and low-level details), and improve the detection sensitivity of occluded and small target actions.

[0010] The specific situation in step four is as follows: Replace the second standard convolution of the CSP residual module with a lightweight GSConv module, compress the number of parameters through group convolution, and introduce channel shuffle to enhance feature interaction in occluded scenarios. Share the group convolution weights in the residual path to reduce redundant calculations, force the network to learn common features (such as human pose changes), and improve the detection generalization of complex indoor actions. The output of the optimized module is fused and sent to the detection head to accurately locate and classify actions, achieving a balance between lightweight and scene robustness. Compared with the prior art, the beneficial effects of the technical solution of the present invention are: (1) Replace the standard 3×3 convolution with dynamic group convolution (GSConv), combined with channel shuffle operation, to reduce the number of parameters while retaining the cross-group feature interaction ability, and adapt to the real-time detection requirements of edge devices; (2) Through residual path weight sharing and cross-level channel shuffle of BiFPN, force the model to learn common motion features and enhance multi-scale feature fusion, significantly reduce the missed detection rate of occlusions and small target actions, and improve the adaptability to illumination changes and multi-person interaction scenarios.

[0011] Figure 1 This is the overall flowchart of the present invention.

[0012] Figure 2 This is the improved YOWOv3 network model used in the present invention.

[0013] Figure 3 This is the improved lightweight grouped shuffle convolution module used in the present invention. Detailed implementation manners

[0014] For those skilled in the art, some well-known structures and their descriptions in the drawings can be omitted. The technical solutions of the present invention will be further described below with reference to the drawings and embodiments. In the 2D backbone network of YOWOv3 of the present invention, the 3×3 standard convolution in Stage1 / 3 / 5 is replaced with a GSConv module, and the number of input and output channels is retained. The number of groups is selected through a dynamic grouping strategy (the largest number of groups that the number of input channels can be divided by). For example, 8 groups are used when the input is 64 channels, reducing the computational amount while retaining the feature interaction ability.

[0015] Figure 1 This is the method flowchart of the present invention. First, the annotation format of the Charades_v1_480 dataset is converted to generate standardized input data. Subsequently, an improved YOWOv3 network model is constructed: in the 2D backbone network, the key 3×3 convolution is replaced with a GSConv module by a dynamic grouping strategy, the cross-level channel shuffle of the BiFPN feature pyramid is optimized, and the grouped convolution weights are shared in the CSP residual module to compress the parameters.

[0016] Figure 2 The YOWOv3 network model of the present invention includes two main branches: a 2D CNN branch and a 3D CNN branch. Among them, the 2D CNN branch introduces an improved GSConv convolution module, reducing the computational amount while retaining the feature interaction ability.

[0017] Figure 3The lightweight module of the present invention consists of a three - stage cascaded structure of dynamic group convolution, channel shuffle, and 1×1 convolution. The input features are divided by a dynamic grouping strategy (for example, 64 channels are divided into 8 groups), each group independently performs 3×3 convolution, and the output channels are halved; the cross - group information is rearranged through channel shuffle to enhance feature interaction in occluded scenarios; 1×1 convolution restores the target number of channels. The specific implementation steps are as follows:

[0018] Step1.1 Based on the publicly available Charades_v1_480 dataset of human activities, the video is divided into continuous segments of 30 frames to retain the time - continuity features; Step1.2 Use YOLOv5 to detect human targets in the video frame by frame, filter out the detection frames of non - target people, and only retain the coordinate information of the targets to be analyzed; Step1.3 Use the VIA annotation tool to manually annotate the actions of the target people. The annotation content includes the action categories frame by frame (such as "fall", "sit and stand") and the spatial positions (bounding box coordinates). Combine the expert annotations to define 12 basic actions: stand, lie down, sit, tidy up, cook, play with mobile phone, make a call, eat and drink, watch TV, open and close the door, shake hands and hug, fall; Step1.4 Input the video frames and the YOLOv5 detection coordinates into the DeepSort algorithm to associate the temporal ID of the target people and generate a continuous action sequence with identity labels for the model to learn spatio - temporal behavior patterns; Step2.1 In the 2D backbone network DarkNet of YOWOv3, identify the computationally intensive 3×3 standard convolution layers (concentrated in Stage 1 / 3 / 5), divide the input channels according to the dynamic number of groups, each group independently performs 3×3 convolution, and the output channels are halved; Step2.2 Generate a list of legal grouping candidate values according to the number of input channels (for example, when the input is 64 channels, 1 / 2 / 4 / 8 groups can be selected), and select the maximum number of groups (gopt = 8) to ensure uniform channel division; Generate all grouping candidate values that can be divided by C in and select the maximum value g opt, as shown in Equation 1: (1) Step3.1 In the channel alignment layer of BiFPN, locate the 1×1 standard convolution layer used to adjust the number of channels, replace the original 1×1 standard convolution with a GSConv module, and retain the number of input and output channels; Step3.2 In the up - sampling path (such as P5→P4) and down - sampling path (such as P3→P4) of BiFPN, insert a channel shuffle layer after the feature concatenation operation (Concat); After replacing the 1×1 convolution in BiFPN with GSConv, the number of parameters is reduced to 1 / g opt, as shown in Equation 2: (2) Step4.1 In the cross-stage partial network (CSP) residual module of YOWOv3, replace the second 3×3 standard convolution with a lightweight GSConv module, and retain the first convolutional layer to ensure the basic feature extraction ability; The original residual module contains two 3×3 standard convolutions. After improvement, the second convolution is replaced with GSConv (number of groups g opt), as shown in Equation 3: (3) Step4.2 Share the weight parameters of the GSConv grouped convolutional layer among all residual modules in the same CSP stage (for example, the 4 residual modules in Stage 3 reuse the same set of convolutional kernels), and define the GSConv layer as a globally reusable module through the shared parameter interface of the framework.

Claims

1. Lightweight group shuffle convolutional personnel activity detection algorithm based on improved YOWOv3, Its characteristics include the following steps: Step 1: Obtain a video dataset of human activity movements, modify the dataset according to the AVA dataset annotation format, and successfully apply it in the YOWOv3 network; Step 2: For the 2D backbone network of YOWOv3 (such as Stage 1 / 3 / 5 layers of DarkNet), identify its standard 3×3 convolution layer, replace the original 3×3 standard convolution with the GSConv module, retain the number of input and output channels, but reduce the amount of calculation through grouped convolution, and dynamically select the maximum legal number of groups according to the number of input channels of the current layer to ensure that the channels are evenly divided (for example, when the input channels are 64, the number of groups can be 1 / 2 / 4 / 8 / 16 / 32 / 64); Step 3: In the channel alignment layer (such as P5→P4, P4→P3) of the bidirectional feature pyramid (BiFPN) of YOWOv3, GSConv is used to replace the original 1×1 standard convolution; channel shuffle operations are inserted in the feature upsampling / downsampling path to enhance cross-level feature interaction; Step 4: Replace the second standard convolution of the residual module in the YOWOv3 cross-stage local network (CSP) with GSConv, share the grouped convolution weights in the residual path to avoid repeated calculations, and finally output the detection results of the personnel activity actions.

2. A lightweight group shuffle convolution based on improved YOWOv3 is used for human activity detection, which is characterized by The specific process in Step 1 is as follows: Step 1.1 Based on the public human activity dataset Charades_v1_480, the video is divided into 30-frame continuous segments to retain the temporal continuity feature; Step 1.2 Use YOLOv5 to detect the human targets in the video frame by frame, filter the detection frames of non-target people, and only retain the coordinate information of the target to be analyzed; Step 1.3 Use VIA annotation tools to manually annotate the target person's actions, including frame-by-frame action categories (such as "falling" and "sitting") and spatial positions (bounding box coordinates). Combined with expert annotations, 12 basic action categories are defined; Step 1.4 Input the video frames and YOLOv5 detection coordinates into the DeepSort algorithm, associate the temporal ID of the target person, and generate a continuous action sequence with identity labels for the model to learn spatiotemporal behavior patterns.

3. The personnel activity detection algorithm based on improved YOWOv3 lightweight group shuffle convolution according to claim 1 is characterized in that The specific process in Step 2 is as follows: Step 2.1 In the 2D backbone network DarkNet of YOWOv3, identify the computationally intensive 3×3 standard convolutional layers (concentrated on Stage 1 / 3 / 5), divide the input channels into dynamic groups, perform 3×3 convolutions independently for each group, and reduce the output channels by half; Step 2.2 Generate a list of legal grouping candidates based on the number of input channels (e.g., 1 / 2 / 4 / 8 groups can be selected when 64 channels are input), and select the maximum number of groups (gopt = 8) to ensure that the channels are evenly divided.

4. The personnel activity detection algorithm based on improved YOWOv3 lightweight group shuffle convolution according to claim 1 is characterized in that The specific process in Step 3 is as follows: Step 3.1 In the channel alignment layer of BiFPN, locate the 1×1 standard convolution layer used to adjust the number of channels, replace the original 1×1 standard convolution with the GSConv module, and retain the number of input and output channels; Step 3.2 In the upsampling path (such as P5→P4) and downsampling path (such as P3→P4) of BiFPN, insert a channel shuffle layer after the feature concatenation operation (Concat).

5. The personnel activity detection algorithm based on improved YOWOv3 lightweight group shuffle convolution according to claim 1 is characterized in that The specific process in Step 4 is as follows: Step 4.1 In the cross-stage local network (CSP) residual module of YOWOv3, the second 3×3 standard convolution is replaced with a lightweight GSConv module, and the first convolution layer is retained to ensure the basic feature extraction capability; Step 4.2 All residual modules in the same CSP stage share the weight parameters of the GSConv grouped convolution layer (for example, the four residual modules of Stage 3 reuse the same set of convolution kernels). Through the framework's shared parameter interface, the GSConv layer is defined as a globally reusable module.

Citation Information

Cited By

  • Weak supervision space-time action detection method and device based on reconstruction double-flow label matching

    CN120451881A

  • Weakly supervised spatiotemporal action detection method and device based on reconstructed dual-stream label matching

    CN120451881B