Lightweight live-action pedestrian recognition method based on attention mechanism and shared convolution
By introducing attention mechanism and shared convolution technology into the pedestrian detection model, the lightweight C2F-FasterNet module and redesigned detection head are designed, and the balance between computing complexity and detection accuracy of existing models is solved, achieving efficient real-time detection and resource consumption reduction.
Patent Information
- Application Number
- CN202510349406.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-13
AI Technical Summary
Existing pedestrian detection models are difficult to balance between computational complexity and detection accuracy, especially in terms of real-time and resource consumption.
Using a lightweight real-life pedestrian recognition method based on attention mechanism and shared convolution, the redundant feature calculation is reduced and the calculation efficiency is improved through the C2F-FasterNet module, the Context Anchor Attention module, the GELU activation function and the redesigned detection head.
Significantly reduces computing resource consumption, realizes efficient real-time detection, maintains detection accuracy while reducing the amount of model parameters, and is suitable for mobile or edge computing devices.
Smart Images

Figure CN120148071A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of computer vision and deep learning, and particularly relates to a lightweight real-scene pedestrian recognition method based on an attention mechanism and shared convolution, which is applicable to real-time target detection tasks in intelligent transportation scenarios. Background Art
[0002] The research on pedestrian detection technology has both significant engineering application value and frontier academic significance in the field of computer vision. In the intelligent transportation application scenario, as the core component of the environmental perception module of the autonomous driving system, this technology provides key input information for vehicle path planning and dynamic obstacle avoidance decision-making, and its detection accuracy is directly related to the functional safety level of ADAS (Advanced Driver Assistance System). In the field of urban security, through pedestrian detection algorithms that integrate spatio-temporal context, abnormal behavior pattern recognition and crowd density heat map generation are realized, providing technical support for the governance of smart cities. In addition, in the service robot SLAM (Simultaneous Localization and Mapping) system and AR / MR (Augmented / Mixed Reality) human-computer interaction interface, real-time pedestrian detection algorithms effectively improve the system's environmental understanding ability and user experience.
[0003] From the dimension of computer vision theory research, pedestrian detection has become a benchmark test platform for verifying new object detection paradigms due to its unique technical challenges. Compared with conventional object detection tasks, its technical difficulties are reflected in: multi-scale feature fusion (processing height changes in the range of 0.5 - 3m), dynamic occlusion modeling (dealing with more than 70% of local occlusion scenarios), dense crowd instance segmentation (solving the high overlap situation where the human body spacing is less than 15 pixels), and cross-modal feature alignment (fusing multi-source sensor data such as visible light and infrared). These technical bottlenecks have promoted the development of new architectures such as deformable convolutional networks and graph neural networks.
[0004] The current research paradigm is dominated by deep learning and mainly evolves along two technical routes: two-stage detectors based on region proposals (such as Faster R-CNN and its variants) and end-to-end one-stage detectors (such as YOLO series and RetinaNet). However, existing models generally face the trade-off dilemma between computational complexity and detection accuracy - the number of parameters of the ResNet-152 backbone network is as high as 60M, which is difficult to meet the real-time requirements of in-vehicle embedded platforms (such as NVIDIA Xavier). It is worth noting that lightweight technologies such as knowledge distillation (such as Logit Mimicking), neural architecture search (such as EfficientDet), and model quantization (INT8 precision) are breaking through the technical bottlenecks in this field.
[0005] These breakthroughs provide innovative solutions to break the inherent contradiction between model lightweight and detection accuracy. Summary of the Invention
[0006] The purpose of the present invention is to provide a lightweight real-scene pedestrian recognition method based on an attention mechanism and shared convolution, aiming to reduce the redundancy of the network structure, improve the computational efficiency, and reduce the consumption of computing resources while ensuring the detection accuracy.
[0007] The present invention is implemented as follows. A lightweight real-scene pedestrian recognition method based on an attention mechanism and shared convolution includes the following steps:
[0008] Step 1: Use a drone or install a fixed camera on a vehicle to collect real vehicle and pedestrian picture data, and perform picture annotation and preprocessing on it;
[0009] Step 2: Obtain a network model for the real-scene pedestrian target recognition task in a dynamic traffic scene, that is, the basic model of YOLOv8;
[0011] Step 3: Improve the model. For the real-scene pedestrian target recognition task in a dynamic traffic scene, design a dedicated detection model;
[0011] Step 4: Obtain the final model, and embed this model into a platform with a function of receiving photos to perform the pedestrian target recognition task.
[0012] Further, step 3 is specifically as follows:
[0013] The backbone network is composed of multiple sequentially connected C2F-FasterNet modules, and the C2F-FasterNet module compresses redundant features through a cross-stage partial connection strategy; a Context Anchor Attention module (CAA) is introduced before the SPPF module;
[0014] The output end of the neck network is connected to the GELU activation function, aiming to enhance the expression ability of non-linear features and is used to enhance the expression ability of non-linear features; GNConv and DEGNConv modules are designed to redesign the detection head to streamline the parameters of the obtained model and improve the computational efficiency.
[0015] Step 3-1: All C2f modules in the backbone part are replaced with C2f-FasterNet modules. The C2F-FasterNet divides the input feature map into two branches. The main branch is sequentially connected to n partial convolutions (PConv), and the PConv layer is followed by two PWConv layers or a 1x1 conventional convolution layer to form an inverted residual block. The auxiliary branch retains the original features; the outputs of the main branch and the auxiliary branch are channel-concatenated through a cross-stage partial connection strategy;
[0016] And adopt adaptive spatial feature fusion (ASFF) to adjust the fusion weights;
[0017] Step 3-2: Add a CAA module before the SPPF module of the backbone. This module first obtains local features through a global average pooling, then these features pass through a 1×1 convolutional layer, and then two depthwise separable strip convolutions (Strip Conv) are adopted, which are performed in the horizontal and vertical directions respectively. These designs are used to simulate the effect of large kernel convolution. The relevant formula is: where represents the pooled features, k b = 11 + 2×l, represents the features in the horizontal direction, represents the features in the vertical direction;
[0018] The obtained result will pass through another 1×1 Conv and Sigmoid function to generate a weighted attention map, enabling different regions in the feature map to be weighted according to their global context information.
[0019] Step 3-3: Replace the activation function used in the Head part with the GELU activation function. The relevant formula is:
[0020] GELU(x) = x·Φ(x),
[0021] where Φ(x) is the cumulative distribution function of the standard Gaussian distribution, and its approximate calculation formula is:
[0022]
[0023] Step 3-4: Replace the original detection head of the YOLO model with the redesigned detection head. The implementation logic of the redesigned detection head is as follows:
[0024] First, design a Group Norm Conv convolution (GNConv), which enables the convolution module to divide related feature channels into a group;
[0025] And calculate the mean and variance of each group in the channel direction, and perform normalization within each group
[0026] The pixel set for calculating the mean and variance is S i ,
[0027] where G is the number of groups, C / G is the number of channels in each group, N is the batch axis, indicates that k and i are in the channels of the same group;
[0028] In addition, based on the Group Norm convolution, the Detail Enhanced Conv convolution (DEConv) is added, and the Detail Enhanced Group Norm Conv (DEGNConv) is designed by combination;
[0029] Specifically, DEConv integrates prior information into the ordinary convolution layer, and then, by using the reparameterization technique, DEConv is equivalently converted into an ordinary convolution without additional parameters and computational costs.
[0030] The calculation formula of DEConv is: where K cut represents combining parallel convolutions together;
[0031] The redesigned detection head receives the features of P3, P4, and P5, and then performs 1*1 GNGonv operations respectively. Further, the processed features are input into two consecutive 3*3 DEGNConv for further information aggregation; finally, the extracted information is input into the classification and regression heads, and it performs feature scaling through the Scale layer.
[0032] Further, step 4 is specifically:
[0033] The model needs to be embedded in a platform with the function of receiving photos. The input module of the model supports the fusion of RGB and infrared bimodal data, and aligns heterogeneous features through the cascaded cross-modal attention mechanism (CMA). After the photos are detected in the model, the specific values of relevant evaluation indicators such as P, R, mAP50, mAP50-90, etc. and the situation pictures of pedestrian detection can be obtained.
[0034] The beneficial effects of the present invention are as follows: (1) Significantly reduce the consumption of computing resources, achieve efficient real-time detection. By introducing the C2F-FasterNet module, and in addition, combined with the attention mechanism and shared convolution technology, further reduce redundant feature calculations, making the model more suitable for deployment on mobile or edge computing devices. (2) The model has a wide range of applicable scenarios. The system is applicable to traffic monitoring scenarios and can handle pedestrian targets with occlusion, scale changes, and lighting interference. (3) The system of the present invention has strong practicability and deployment flexibility, supports dual-modal input expansion, can adapt to a variety of sensor data (such as visible light and infrared cameras), and broadens the scope of application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 It is a schematic diagram of the overall process provided by an embodiment of the present invention;
[0036] Figure 2 The structural diagram of the YOLOv8 base model provided by the embodiments of the present invention;
[0037] Figure 3 The structural diagram of the C2f-FasterNet module provided by the embodiments of the present invention;
[0038] Figure 4 The schematic diagram of the redesigned detection head provided by the embodiments of the present invention; Detailed implementation manners
[0039] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0040] Figure 1 The flow of a lightweight real-scene pedestrian recognition method based on the attention mechanism and shared convolution of the present invention is shown. As shown in the figure, the present invention is implemented as follows. A lightweight real-scene pedestrian recognition method based on the attention mechanism and shared convolution includes the following steps:
[0041] S1. Use a drone or install a fixed camera on a vehicle to collect real vehicle and pedestrian picture data, and perform picture annotation and preprocessing on it;
[0042] S2. Obtain a network model for the real-scene pedestrian target recognition task in a dynamic traffic scene, that is, the base model of YOLOv8, as Figure 2 shown;
[0043] S3. Improve the model, and design a dedicated detection model for the real-scene pedestrian target recognition task in a dynamic traffic scene;
[0044] S4. Obtain the final model, and embed the model into a platform with a function of receiving photos to perform the pedestrian target recognition task.
[0045] Furthermore, in step S3, when improving the model and designing a dedicated detection model for the pedestrian target recognition task, specifically:
[0046] The backbone network is composed of multiple sequentially connected C2F-FasterNet modules. A Context Anchor Attention module (CAA) is introduced before the SPPF module that compresses redundant features through a cross-stage partial connection strategy in the C2F-FasterNet module; the output end of the neck network is connected to the GELU activation function, aiming to enhance the expression ability of non-linear features and used to enhance the non-linear feature expression ability; the GNConv and DEGNConv modules are designed to re-design the detection head to streamline the parameters of the obtained model and improve the computational efficiency.
[0047] S31. All C2f modules in the backbone part are replaced with C2f-FasterNet modules. Figure 3 The structure diagram of the C2f-FasterNet module is shown. C2F-FasterNet divides the input feature map into two branches. The main branch is sequentially connected to n partial convolution (PConv) layers, and the PConv layer is followed by two PWConv layers or a 1x1 conventional convolution layer to form an inverted residual block. The auxiliary branch retains the original features; the outputs of the main branch and the auxiliary branch are concatenated in channels through a cross-stage partial connection strategy, and adaptive spatial feature fusion (ASFF) is used to adjust the fusion weights.
[0048] S32. Add a CAA module before the SPPF module in the backbone. This module first obtains local features through a global average pooling; then these features pass through a 1*1 convolution layer, and then two depthwise separable strip convolutions (Strip Conv) are used, respectively in the horizontal and vertical directions. These designs are used to simulate the effect of large kernel convolution. The relevant formula is: where represents the pooled features, k b = 11 + 2×l, represents the features in the horizontal direction, represents the features in the vertical direction;
[0049] The obtained result will then pass through a 1*1 Conv and a Sigmoid function to generate a weighted attention map, enabling different regions in the feature map to be weighted according to their global context information.
[0050] S33. The activation function used in the Head part is replaced with the GELU activation function. The relevant formula is:
[0051] GELU(x) = x·Φ(x);
[0052] where Φ(x) is the cumulative distribution function of the standard Gaussian distribution, and its approximate calculation formula is:
[0053]
[0054] S33: Feed the features obtained by multiplying the two branches into the maxpool layer for pooling to obtain the output features.
[0055] S34. Replace the original detection head of the YOLO model with the redesigned detection head. The implementation logic of the redesigned detection head is as follows: First, Group Norm Conv (GNConv) is designed. This convolution enables the convolution module to divide related feature channels into groups, calculate the mean and variance of each group in the channel direction, and perform normalization within each group.
[0056] The pixel set for calculating the mean and variance is S i ,
[0057] where G is the number of groups, C / G is the number of channels in each group, N is the batch axis, indicates that k and i are in the channels of the same group;
[0058] In addition, on the basis of the Group Norm convolution, Detail Enhanced Conv (DEConv) is added, and Detail Enhanced Group Norm Conv (DEGNConv) is designed in combination.
[0059] Specifically, DEConv integrates prior information into the ordinary convolution layer. Then, by using the reparameterization technique, DEConv is equivalently converted into an ordinary convolution without additional parameters and computational costs.
[0060] The calculation formula of DEConv is: where K cut indicates combining parallel convolutions together;
[0061] Finally, the extracted information is input into the classification and regression heads, and it is feature-scaled through the Scale layer. Figure 4 Figure [ID] shows the structural diagram of the entire redesigned detection head.
[0062] Furthermore, in step S4, the final model is obtained. The model needs to be embedded in a platform with a function of receiving photos. The input module of the model supports the fusion of RGB and infrared bimodal data, and aligns heterogeneous features through the cascaded cross-modal attention mechanism (CMA).
[0063] After the photo is input into the model for detection, the specific values of relevant evaluation metrics such as P, R, mAP50, mAP50-90, etc. and the situation picture of pedestrian detection can be obtained. The calculation formulas involved are as follows:
[0064]
[0065]
[0066]
[0067] Since the pedestrian detection method based on deep learning is prone to problems such as structural redundancy and large consumption of computing resources under general structural combinations, there are still certain limitations in the detection model in balancing the lightweight and accuracy of the model. The pedestrian detection method designed by the present invention can reduce the number of model parameters, improve the computing efficiency, and reduce the consumption of computing resources while maintaining the detection accuracy.
[0068] Although the specific implementation manners of the present invention are described above in conjunction with the accompanying drawings, it is not a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications or deformations that can be made without creative labor on the basis of the technical solutions of the present invention are still within the protection scope of the present invention.
Claims
1. A lightweight real-scene pedestrian recognition method based on attention mechanism and shared convolution, characterized in that: The steps include: Step 1: Use a drone or install a fixed camera on a car to collect real vehicle and pedestrian image data, and perform image annotation and preprocessing on it; Step 2: Obtain a network model for real-life pedestrian target recognition tasks in dynamic traffic scenarios, i.e., the basic model of YOLOv8; Step 3: Improve the model and design a dedicated detection model for real-life pedestrian target recognition tasks in dynamic traffic scenes; Step 4: Get the final model and embed it into a platform with photo receiving function to perform pedestrian target recognition tasks.
2. According to claim 1, a lightweight real-scene pedestrian recognition method based on attention mechanism and shared convolution is characterized in that: The improvements included in step 3 are as follows: The backbone network consists of multiple C2F-FasterNet modules connected in sequence, and the C2F-FasterNet modules compress redundant features through a cross-stage partial connection strategy; The Context Anchor Attention module (CAA) is introduced before the SPPF module; the output of the neck network is connected to the GELU activation function to enhance the expression ability of nonlinear features; The GNConv and DEGNConv modules are designed to redesign the detection head to simplify the parameters of the obtained model and improve the computational efficiency.
3. According to claim 2, a lightweight real-scene pedestrian recognition method based on attention mechanism and shared convolution is characterized in that: The structure of the C2F-FasterNet module is: The input feature map is divided into two branches. The main branch is connected to n partial convolutions (PConv) in sequence. The PConv layer is connected to two PWConv layers or 1x1 regular convolution layers to form an inverted residual block. The auxiliary branch retains the original features. The outputs of the main branch and the auxiliary branch are channel-wise spliced through a cross-stage partial connection strategy, and the adaptive spatial feature fusion (ASFF) is used to adjust the fusion weights.
4. According to claim 2, a lightweight real-scene pedestrian recognition method based on attention mechanism and shared convolution is characterized in that: The CAA attention mechanism is specifically implemented as follows: The module first obtains local features through a global average pooling; These features are then passed through a 1*1 convolutional layer, followed by two depth-separable strip convolutions (Strip Conv), performed in the horizontal and vertical directions respectively. These designs are used to simulate the effect of large kernel convolution. The relevant formula is: in represents the feature after pooling, k b =11+2×l, Represents the horizontal characteristics, Indicates features in the vertical direction; The result is then passed through a 1*1 Conv and Sigmoid function to generate a weighted attention map, so that different areas in the feature map can be weighted according to their global context information.
5. According to claim 2, a lightweight real-scene pedestrian recognition method based on attention mechanism and shared convolution is characterized in that: The expression of the GELU activation function is: GELU(x)=x·Φ(x), where Φ(x) is the cumulative distribution function of the standard Gaussian distribution, and its approximate calculation formula is:
6. According to claim 2, a lightweight real-scene pedestrian recognition method based on attention mechanism and shared convolution is characterized in that: The implementation logic of the redesigned detection head is as follows: First, we designed the Group Norm Conv convolution (GNConv), which enables the convolution module to group related feature channels into a group, calculate the mean and variance of each group in the channel direction, and perform normalization within each group. The pixel set for calculating mean and variance is S i , Among them, G is the number of groups, C / G is the number of channels in each group, and N is the batch axis. indicates that k and i are in the same channel of the same group; In addition, on the basis of Group Norm convolution, Detail Enhanced Conv convolution (DEConv) is added to design Detail Enhanced Group Norm Conv (DEGNConv) in combination; specifically, DEConv integrates prior information into the ordinary convolution layer, and then, by using the reparameterization technology, DEConv is equivalently converted to ordinary convolution, without the need for additional parameters and computational costs. The calculation formula of DEConv is: Where K cut It means combining parallel convolutions together; the redesigned detection head receives the features of P3, P4, and P5, and then performs 1*1 GNGonv operations on each of them, and further inputs the processed features into two consecutive 3*3 DEGNConvs for further information aggregation; Finally, the extracted information is input into the classification and regression head, and it is scaled through the Scale layer.
7. The lightweight real-scene pedestrian recognition method based on attention mechanism and shared convolution according to claim 1 is characterized in that: The input module supports the fusion of RGB and infrared bimodal data and aligns heterogeneous features through a cascaded cross-modal attention mechanism (CMA).
8. The lightweight real-scene pedestrian recognition method based on attention mechanism and shared convolution according to claim 1, characterized in that: The system is suitable for traffic monitoring scenarios and can handle pedestrian targets with occlusion, scale changes and lighting interference, with a detection frame rate of ≥45FPS (1080p resolution, GPU platform).