Lightweight high-efficiency dense pedestrian detection method based on YOLO-CPEE

By improving the YOLO11 model structure, introducing the C2CGA module, shallow high-resolution feature extraction and EMA attention mechanism, and designing an efficient detection head, the problems of high false detection rate and high missed detection rate in pedestrian detection in dense scenes are solved, and efficient and accurate pedestrian detection is achieved.

CN120726564APending Publication Date: 2025-09-30TIANJIN POLYTECHNIC UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510829227.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

Existing pedestrian detection algorithms have difficulty effectively handling pedestrian overlap, multi-scale changes, and complex background interference in dense scenes, resulting in high false detection rates and high missed detection rates, especially poor results in detecting small targets.

Method used

The YOLO-CPEE algorithm is adopted to enhance the feature extraction and detection capabilities by improving the YOLO11 model structure, introducing the C2CGA module, shallow high-resolution feature extraction, EMA attention mechanism and efficient detection head.

Benefits of technology

It improves the accuracy and stability of dense pedestrian detection, reduces computational complexity, enhances the detection capability of small targets and complex backgrounds, and improves detection precision and recall rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120726564A_ABST
    Figure CN120726564A_ABST
Patent Text Reader

Abstract

The invention discloses a lightweight high-efficiency dense pedestrian detection model based on a YOLO-CPEE algorithm, belongs to the technical field of computer vision, and aims to solve the detection problem caused by high-density overlapping between pedestrians and a small proportion of distant pedestrians in an image. On the basis of a YOLO11 model, a cascade group attention mechanism is integrated into an original feature extraction module C2PSA to form a new C2CGA module, and good balance between model calculation efficiency and precision is achieved; a high-resolution P2 feature layer is introduced to better capture shallow pedestrian feature details and position information of small targets; cross-scale feature interaction and adaptive weight distribution are realized through an efficient multi-scale attention mechanism, and the multi-scale feature extraction capability is further improved; a novel lightweight high-efficiency detection head is designed, and the target detection capability is remarkably improved while model parameters are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision technology, and relates to target detection, pedestrian recognition, and image classification. It is specifically manifested in a lightweight, efficient and dense pedestrian detection method based on YOLO-CPEE. Background Art

[0002] Pedestrian detection technology is widely used to obtain pedestrian information. As the foundation of our society and economy, pedestrians carry enormous benefits and risks. Therefore, for the sake of public safety and convenience, research on pedestrian detection deserves high attention. With the acceleration of urbanization and the increasing demand for public safety, pedestrian detection in densely populated scenarios has become a valuable research direction. This technology plays an irreplaceable role in intelligent traffic management systems, large-scale video surveillance networks, and autonomous mobile robot platforms.

[0003] Pedestrian detection methods are divided into traditional methods and deep learning-based methods. Traditional detection methods mainly follow manually designed features and cascade classifier architectures. The manual extraction method has cumbersome extraction steps and requires years of experience to design and process. It has high computational costs and unsatisfactory real-time performance. Pedestrian detection methods based on deep learning have made significant progress in recent years. It simulates the visual perception system of the human brain and learns the differential features in the response data through a large amount of data. Among them, the deep learning model based on the convolutional neural network (CNN) plays a key role in solving the high-level semantic feature extraction of detection. As a typical single-stage target detection framework, the YOLO series is widely used in pedestrian detection tasks with high real-time requirements due to its efficient detection speed and better detection accuracy.

[0004] Although target detection technology has made significant progress through traditional and deep learning methods, existing algorithms are still difficult to effectively generalize to various complex real-world scenarios. In surveillance and autonomous driving, pedestrians are usually detected from captured images. In the collected images, pedestrians in dense areas often overlap with each other and have different movements, resulting in an excessively high false detection rate for pedestrians. Due to the uneven distribution of pedestrian positions in the image, the target sizes show significant diversity. Especially in long-distance scenes, small-sized targets are easily overlooked, resulting in the missed detection of some pedestrians. In addition, the background in the image is generally more complex, especially in low-light environments. Objects are easily misidentified as pedestrians because of their similar shape, appearance, color and texture, which greatly interferes with the accuracy of detection. In response to this series of problems, there are still huge challenges in the research of pedestrian detection in dense scenes. Summary of the Invention

[0005] In response to the shortcomings of existing methods, this paper analyzes the robustness of pedestrian detection models in complex scenes. To address the high false detection rate caused by multi-scale image changes and dense crowd occlusion, this paper proposes a new lightweight, efficient, and dense pedestrian detection algorithm, YOLO-CPEE, based on the improved YOLO11. The algorithm includes the following steps:

[0006] S1. Acquisition of dense pedestrian dataset;

[0007] Through web crawlers, public dataset platforms, and field collection, we acquire pedestrian image data from densely populated areas such as urban business districts, subway stations, squares, and parks, building a dataset suitable for dense pedestrian detection. We use annotation tools to accurately label pedestrian targets, ensuring data quality and improving the adaptability of the detection model.

[0008] S2, image preprocessing;

[0009] The acquired pedestrian images are preprocessed by unifying their size, adjusting their brightness and contrast, enhancing noise, and cleaning data. Data enhancement methods, including random flipping, mosaic stitching, and scaling, are used to expand the training sample and improve the model's robustness and generalization capabilities in complex environments.

[0010] S3. Configure the required training environment;

[0011] We built a training environment for the object detection model based on the deep learning framework PyTorch 2.1.0. The configuration includes GPU acceleration hardware RTX 3080Ti, operating system Ubuntu 22.04, CUDA 12.1, and Python 3.10.14. We also set training hyperparameters, including learning rate, batch size, and optimizer type.

[0012] S4. Improve the YOLO11 algorithm model structure to obtain the YOLO-CPEE target detection model;

[0013] First, we propose a C2CGA module, which uses CGA block instead of Bottleneck. We also replace the C2PSA module in the YOLO11 backbone with the C2CGA module. Secondly, we design a shallow high-resolution feature extraction P2 layer, and add an F2 feature fusion layer and a D2 detection head to the neck network, specifically for feature enhancement and prediction of small targets. Then, we introduce the EMA attention mechanism at multiple stages of the neck network to enhance multi-scale feature extraction. Finally, we design an efficient detection head, replacing the original ordinary detection head with a lightweight dynamic detection head, which performs processing through a multi-scale feature fusion mechanism and detail-enhancing shared convolution.

[0014] S5. Train the YOLO-CPEE network model;

[0015] The training set obtained by preprocessing according to step S1 is input into the YOLO-CPEE network model for training, and the validation set is input during the training process. The training effect is verified by the test set, thereby obtaining a trained YOLO-CPEE target detection model;

[0016] S6. Input the dense pedestrian image into the trained YOLO-CPEE network model to obtain category, location, and confidence information.

[0017] Furthermore, the YOLO-CPEE network model structure in step S4 is as follows:

[0018] The YOLO-CPEE network model structure is mainly composed of three parts in series: backbone network, neck network and detection head;

[0019] First is the backbone part, where the input goes through eleven layers in sequence: convolution, convolution, C3K2, convolution, C3K2, convolution, C3K2, convolution, C3K2, SPPF, and C2CGA;

[0020] Then comes the neck part. The output of the trunk part goes through twenty-three layers in the neck, namely upsampling, splicing, C3K2, upsampling, splicing, C3K2, EMA, upsampling, splicing, C3K2, EMA, convolution, splicing, C3K2, EMA, convolution, splicing, C3K2, EMA, convolution, splicing, C3K2, EMA, convolution, splicing, C3K2, EMA, convolution, splicing, C3K2, EMA, totaling twenty-three layers. There are many feature map fusions at different stages in the neck part. The C3K2 layer of the third layer of the trunk part is used as an additional input of the ninth splicing layer of the neck, the C3K2 layer of the fifth layer of the trunk part is used as an additional input of the fifth splicing layer of the neck, the C3K2 layer of the seventh layer of the trunk part is used as an additional input of the second splicing layer of the neck, the C2CGA layer of the eleventh layer of the trunk part is used as an additional input of the first and twenty-first splicing layers of the neck, and the third and sixth C3K2 layers of the neck are used as additional inputs of the thirteenth and seventeenth splicing layers of the neck respectively.

[0021] Finally, the eleventh, fifteenth, and nineteenth layers of the neck are input into the efficient detection head respectively, forming the head of YOLO-CPEE.

[0022] Furthermore, the C2CGA structure in step S4 is specifically as follows:

[0023] The overall structure of the C2CGA module includes key stages: input convolution, feature partitioning, stacking multiple CGABlocks, feature fusion, and output convolution. Specifically, the input feature map first undergoes a preliminary transformation in the channel dimension through a convolutional layer. A split operation then divides the feature map into multiple subgroups along the channel dimension, which are then fed into multiple parallel CGABlock modules. Each CGABlock consists of two main submodules: a cascaded group attention mechanism module and a feedforward network module. The cascaded group attention mechanism module adopts a multi-head attention structure. In each split attention head, a linear transformation of the query, key, and value is first performed. A token interaction mechanism is then used to perform fine-grained interaction between tokens within the group, significantly improving feature representation while maintaining low computational complexity. The outputs of the multiple split heads are concatenated and linearly fused after the self-attention mechanism to produce the final group attention feature. The FFN part uses two serially connected 1×1 convolution operations to further transform the channel direction and nonlinearly enhance the features at each position. Finally, the output features within each CGABlock module are integrated through a concatenation operation and fused through a convolutional layer to produce the output of the current C2CGA module.

[0024] Furthermore, the EMA structure in step S4 is as follows:

[0025] The size of the input feature map is C×H×W, where C is the number of channels, H and W are the height and width of the feature map respectively. First, the input features are grouped by channels and divided into multiple sub-feature maps according to the number of groups g. The number of channels of each sub-feature map is C / g; each group of sub-feature maps enters two main branches respectively: the first branch is the local attention enhancement branch: average pooling is performed on the sub-feature maps in the X-axis direction and the Y-axis direction respectively to extract local context information; after splicing the pooling results in the two directions, they are integrated through 1×1 convolution, and then two Sigmoid activation functions are input respectively to generate two weight maps; the weight maps are used to enhance the original The sub-feature maps are multiplied element-by-element to achieve the suppression and enhancement of spatial attention; then group normalization is used and the results are fused with the results of the second branch; the second branch is the global relationship modeling branch: the sub-feature map first undergoes a 3×3 convolution to expand the context receptive field, and then average pooling is applied; then it enters the Softmax calculation module to generate spatial attention weights; matrix multiplication is used to interact with the down-sampled features to achieve long-range dependency modeling; the global weighted result is multiplied element-wise with the output of the first branch to form an enhanced feature; finally, the outputs of the two branches are fused element-by-element and output as the final result of the EMA module.

[0026] Furthermore, the structure of the efficient detection head in step S4 is as follows:

[0027] The efficient detection head structure adopts a multi-level feature input method, receiving multiple scale feature maps from the feature pyramid, including P2, P3, P4, and P5; each layer of feature map first passes through a lightweight convolution module CGS for channel compression and nonlinear transformation to enhance feature representation capabilities; then, each layer of features is input into a shared convolution module, which contains two consecutive detail enhancement convolutions, which enhances the spatial information of features while maintaining parameter efficiency; this design can improve the sensitivity of the detection head to targets of different scales, especially small targets; after the features are processed by shared convolution, they are input into the classification branch and the bounding box regression branch respectively. The former is used to predict the probability of the target category, and the latter is used to predict the probability of the target category. The bounding box parameters used to predict the target include the center point coordinates, width and height; both branches use independent convolution structures, and a scale coefficient Scale module is attached at the end. This module normalizes and fine-tunes the prediction results according to the distribution of the feature map to improve the prediction accuracy and stability; the detail enhancement convolution is composed of four operations: CDC module, i.e. center differential convolution, ADC module, i.e. angular differential convolution, HDC, i.e. horizontal differential convolution, and VDC, i.e. vertical differential convolution, which are used to model the edge and position information of the target in different directions; this module is connected in parallel with the VC module, i.e. ordinary convolution, in the main path of the feature map, and is fused through residual addition to further enhance the target boundary perception capability.

[0028] The beneficial effects of the present invention are: introducing a cascade group attention module to generate a new C2CGA module, which improves the feature representation extraction capability of the target area through the attention mechanism of the local window, while reducing unnecessary computational redundancy, making the backbone network more efficient in extracting key features; adding a P2 feature layer, fusing the shallow high-resolution features of the P2 layer with the deep features of P3, P4, and P5, so that the network can more effectively utilize shallow small target features and position information, reduce the loss of shallow features caused by the increase in model depth, and make the network more accurate and stable when detecting small targets in dense crowds; by adding an efficient multi-scale attention module at the neck, feature enhancement is performed on different feature layers, dynamically focusing on the information interaction between features of different scales, improving the multi-scale feature extraction capability and accuracy of the network, and maximizing the extraction of effective feature information with less computing power; designing a lightweight, dynamic and efficient detection head, using group convolution for feature extraction, and using a parameter sharing strategy to compress the model complexity, while enhancing the detection head's ability to adapt to geometric deformations of targets of different scales, while keeping the model lightweight and enhancing the model's detection capability of the target. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 This is the structure diagram of the YOLO-CPEE model in the present invention.

[0030] Figure 2 This is the structural diagram of the C2CGA module in the present invention.

[0031] Figure 3 This is a schematic diagram of the P2 structure in the present invention.

[0032] Figure 4 This is the structural diagram of the EMA module in the present invention.

[0033] Figure 5 This is a structural diagram of the high-efficiency detection head module in the present invention.

[0034] Figure 6 This is the detection effect diagram of the DensePed-LR dataset in this invention.

[0035] Figure 7 This is a comparison chart of the detection effects of different YOLO series models in the present invention. DETAILED DESCRIPTION

[0036] Exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings.

[0037] Example

[0038] The present invention provides a lightweight, efficient and dense pedestrian detection method based on YOLO-CPEE. The specific implementation steps are as follows:

[0039] S1. Acquisition of dense pedestrian dataset:

[0040] First, we acquired the dense pedestrian dataset, DensePed-LR. This dataset, used in this example, consists of pedestrian images captured by surveillance cameras in crowded urban areas, subway stations, and parks, images obtained through online crawlers, and a subset of publicly available dense pedestrian detection datasets from the Roboflow website. A total of 3,672 images were collected and, after data augmentation and preprocessing, the images encompassed a variety of complex scenarios, including dense crowds, occlusions, low lighting, and unusual postures. The final number of annotated instances reached 130,984.

[0041] S2. Image preprocessing:

[0042] To enhance the model's robustness in diverse environments, a noise reduction mechanism was introduced during data preprocessing to simulate rainy, foggy, or otherwise impaired conditions. Brightness and contrast were adjusted to enhance the model's detection capabilities at night or in low-light conditions. Furthermore, to improve data quality, comprehensive data cleaning was performed to minimize the impact of incorrect labeling and ensure the reliability of model evaluation. Ultimately, the 11,000 images obtained were screened and annotated, with 8,000 used for training, 2,000 for validation, and 1,000 for testing. All images were resized to a 640×640 resolution to ensure consistency and efficient training.

[0043] S3, experimental environment and parameter configuration:

[0044] After processing the data set, the data.yaml file of the data set is created next, and the paths of the training set and the validation set and all the category information in this data set are written under the data.yaml file. Then, the parameters including the number of training times and batch parameters under train.py are modified according to the situation required by the invention. The hardware used in the present invention is built on an AutoDL server, using RTX 3080Ti GPU, Intel (R) Xeon (R) Silver 4214R CPU, and the software environment runs on an ubuntu22.04 operating system, using PyTorch 2.1.0, Python 3.10.14 and CUDA 12.1. In order to train the model in the embodiment, a stochastic gradient descent optimizer is used, with momentum and weight decay of 0.937 and 0.0005 respectively, warmup_epochs of 3.0, and the learning rate is adjusted to 0.01. In order to find the best hyperparameter values, image size 640×640 and batch size 16 are selected, and the model runs for up to 300 cycles. To prevent model overfitting and improve generalization, data augmentation parameters were adjusted appropriately during training. Data augmentation can generate additional training samples by modifying existing training data or creating synthetic data. Classic methods were used in the examples, including image HSV enhancement, random scale sampling, random flipping, and mosaic enhancement.

[0045] S4. Improve the YOLO11 algorithm model structure to obtain the YOLO-CPEE target detection model. The YOLO-CPEE structure model is as follows:

[0046] The YOLO-CPEE framework is as follows Figure 1As shown in the figure, the main part first undergoes eleven layers: convolution, convolution, C3K2, convolution, C3K2, convolution, C3K2, convolution, C3K2, SPPF, and C2CGA. The neck part then undergoes twenty-three layers: upsampling, concatenation, C3K2, upsampling, concatenation, C3K2, EMA, upsampling, concatenation, C3K2, EMA, convolution, concatenation, C3K2, EMA, convolution, concatenation, C3K2, EMA, convolution, concatenation, C3K2, EMA, convolution, concatenation, C3K2, EMA, convolution, concatenation, C3K2, EMA. The neck part contains feature map fusion at many different stages. The C3K2 layer of the third backbone layer serves as an additional input to the ninth neck layer, the C3K2 layer of the fifth backbone layer serves as an additional input to the fifth neck layer, the C3K2 layer of the seventh backbone layer serves as an additional input to the second neck layer, the C2CGA layer of the eleventh backbone layer serves as an additional input to the first and twenty-first neck layers, and the C3K2 layers of the third and sixth neck layers serve as additional input to the thirteenth and seventeenth neck layers, respectively. Finally, the eleventh, fifteenth, and nineteenth neck layers are input into the efficient detection head, forming the YOLO-CPEE head.

[0047] Compared with the original YOLO11 model structure, the YOLO-CPEE network model of the present invention has the following improvements:

[0048] (1) Replace C2PSA with C2CGA: In dense pedestrian detection tasks, the original C2PSA module is replaced with a more efficient C2CGA module to achieve finer-grained calculations on feature maps and reduce redundant calculations. Although the C2PSA module plays a positive role in enhancing the spatial attention distribution of feature maps, its high computational and memory overheads, especially when dealing with large-scale dense scenes, limit the performance improvement. Specifically, C2PSA uses a global attention mechanism to directly process full-size feature maps, resulting in a significant increase in computational resource consumption under high-resolution input. In addition, the computational structure of C2PSA is relatively simple, lacking sufficient feature decomposition and refinement capabilities, making it difficult to effectively process complex features and high-dimensional information, ultimately restricting the feature discrimination capabilities of densely occluded targets and small-scale pedestrians.

[0049] In order to overcome these shortcomings, the C2CGA module is designed as Figure 2As shown. Define the CGABlock and C2CGA classes, integrate the cascade group attention mechanism into the existing C2PSA attention module system, and implement a local window-based cascade group attention mechanism, which can perform more fine-grained attention calculations on feature maps at different resolutions. The different situations of input feature map size and window resolution are handled to enhance the adaptability of the model. The cascade group attention mechanism gradually passes the output of each attention head to the input of the subsequent head through a cascade operation, allowing each attention head to calculate different partitions of the input features, greatly reducing redundant calculations and improving the efficiency of the model. Formally, it can be expressed as:

[0050]

[0051] X i+1 =Concat[X ij ] j=1:n W i P

[0052] where X i ∈R C×H×W is the complete input feature tensor of the i-th block, and the j-th head computes X ij Self-attention on X ij is the input feature X i The jth partition along the channel dimension, i.e. X i =Concat[X i1 ,X i2 ,...,X in ] and 1≤j≤n, n is the total number of attention heads. is a learnable parameter matrix that maps input features to different subspaces. Attn(·) is the self-attention calculation function, X ij is the input feature X i The attention output calculated by the jth head. The division features output by each head are spliced ​​along the channel dimension and passed through the linear projection layer W i P Project the concatenated features back to the same dimension as the input features. In short, W i P is a linear layer that adjusts the feature dimension. i(j-1) The result of all segmentation heads being concatenated and linearly mapped after the attention outputs. The cascade operation adds the output features of the jth head to the input of the subsequent (j+1)th head, and achieves progressive feature optimization through feature residual accumulation:

[0053] X′ ij =X ij +X i(j-1) ,1≤j≤n

[0054] where X′ ij is the input segmentation feature X of the j-th head ij And the output feature X of the (j-1)th head i(j-1) sum.

[0055] Furthermore, the C2CGA module introduces deep convolution after query projection. Assuming Q is the original query output, the query after deep convolution is represented as Q', using the formula: Q' = DepthwiseConv(Q). Since the original query already contains global context, deep convolution further extracts local spatial details, allowing the attention mechanism to more comprehensively understand the input features. This design enhances the module's ability to capture detailed features, improving target positioning accuracy in densely occluded and complex background scenes, thereby optimizing the robustness of dense pedestrian detection.

[0056] (2) Introducing a shallow high-resolution layer P2: Existing small target definition methods are mainly divided into the following two categories: definition based on relative scale and definition based on absolute scale. In the study of definition based on relative scale, the median ratio of the target bounding box area to the total image area of ​​0.08%-0.58% is used as the judgment standard for small targets. In terms of definition based on absolute scale, targets with a resolution lower than 32×32 pixels are classified as small targets. When extracting features in a convolutional neural network, due to the small pixel ratio of the target object, after several layers of downsampling, the feature information will gradually decrease and continue to be lost as the network layer deepens. The backbone network of YOLO11s adopts an FPN structure, which continuously extracts features through three downsampling layers and outputs downsampled feature maps P3, P4 and P5 with sizes of 1 / 8, 1 / 16 and 1 / 32 of the original image respectively. The 1 / 8 size version of the P3 feature map is used as the input of the neck network for feature fusion of small target detection. However, there are a large number of small targets in dense pedestrian datasets. If the target is smaller than 10×10 pixels, after 1 / 8 downsampling, the target features in the feature map become very weak or even disappear completely. Figure 3 As shown in the figure, a 1 / 4 downsampled high-resolution feature map P2 is introduced into the backbone network, and an F2 feature fusion layer and a D2 detection head are added to the neck network, which are specifically used for feature enhancement and prediction of small targets.

[0057] (3) Adding EMA attention mechanism: EMA is a new type of efficient multi-scale attention mechanism that can be run without dimensionality reduction. Its design idea is to effectively capture multi-scale features while maintaining high efficiency, thereby improving the accuracy of the model. Figure 4In the figure, "g" indicates the divided group, "X average pooling" indicates 1D horizontal global pooling, and "Y average pooling" indicates 1D vertical global pooling. The EMA attention mechanism not only encodes inter-channel information to adjust the importance of different channels, but also preserves the precise spatial structure information in the channels. After the convolution layer, EMA adopts a multi-scale branch responsible for extracting features of different scales. By introducing convolution kernels or pooling operations of different sizes in the network, feature representations of different scales can be obtained at different layers. These features of different scales are fed into their respective attention branches for processing. Each attention branch uses its own attention mechanism to dynamically adjust the importance of features based on specific contextual information. In this way, features of different scales can be assigned different weights according to their relative importance, thereby better capturing small objects in the image.

[0058] This algorithm has the following advantages: 1) It considers channel information and preserves spatial information. 2) It has few parameters and low computational complexity. 3) When faced with objects of varying scales, it can dynamically process information, improving the ability to recognize small objects in complex environments, thereby enhancing the performance of computer vision tasks. 4) It is highly flexible and can be easily inserted into the core modules of lightweight networks without retraining the entire model.

[0059] Introducing the EMA mechanism enhances the model's ability to learn pedestrian features, thereby improving detection accuracy and robustness. Because pedestrians in dense scenes may be subject to significant occlusion, varying postures, and complex backgrounds, the EMA mechanism helps the model more effectively extract key features from multi-scale images, especially when processing image regions with high levels of detail.

[0060] (4) Improve the original detection head to a high-efficiency detection head:

[0061] The detection accuracy and computational complexity of the model often show a mutually restrictive relationship. In order to avoid the computational redundancy of the model itself and the increase in computational cost and parameter amount caused by the optimization of the early feature extraction network, and because each detection head is input by a separate feature map, the lack of information interaction between each other can easily lead to the loss of some effective information, insufficient feature sharing, and thus affect the reasoning effect of the model. In the embodiment, the detection head part is optimized, and a lightweight multi-scale dynamic fusion detection head-efficient detection head is proposed. By introducing a multi-scale feature fusion mechanism, it is possible to capture contextual information of different scales at the same time, significantly improving the model's recognition ability for multi-scale targets. While ensuring high detection accuracy, the shared convolution optimization network structure and parameter configuration are introduced to effectively reduce the number of parameters and the amount of calculation, thereby achieving a balance between detection accuracy and reasoning efficiency.

[0062] Based on the parameter sharing concept, the YOLO11 detection head is redesigned, and a lightweight and efficient detection head based on detail enhancement convolution DEConv is proposed. Its network structure is as follows Figure 5 As shown in the figure. The efficient detection head first enables the feature maps of P2-P5 to exchange information across channels through group normalized standard convolutions (GN) with a convolution kernel of 1×1, enriching the target information and thus improving the performance of the detection head. Secondly, two detail enhancement convolutions with a convolution kernel of 3×3 are used to fuse the feature maps of three different scales for feature interaction and parameter sharing. The shared convolution with a convolution kernel of 3×3 can improve the receptive field, increase the probability of learning adjacent feature information, and achieve information fusion in various dimensions. Finally, the information extracted by the shared convolution is input into the classification and regression heads. In addition, to address the problem of inconsistent target sizes detected by each detection head, a Scale layer is introduced in the regression detection head to scale the features, thereby enhancing the model's ability to capture multi-scale features. The Scale layer also contains a learnable scaling factor X that can be multiplied by input features of any scale and continuously updated through gradient descent during the training process, thereby achieving dynamic scaling of the input features and improving the generalization ability of the model.

[0063] S5. Train with the modified model:

[0064] The proposed method divides the training set into training, validation, and test sets in a ratio of 8:2:1. The training rounds are set to 300, with 16 images input for each training session. The training data is observed in real time during the training process, and the trained weights are saved after the training is completed. The experimental results are shown in Table 1. The experimental results show that compared to the original model, the proposed YOLO-CPEE model improves accuracy by 0.7%, recall by 4.1%, mAP@0.5 by 2.6%, and mAP@0.5:0.95 by 3.2%. At the same time, the number of parameters is reduced by 3.5%, and the model size is reduced by 2.6%.

[0065] Table 1 Test results of YOLO11s and YOLO-CPEE models

[0066] Model P / % R / % mAP@0.5 / % mAP@0.5:0.95 / % Parameters / M Size / MB YOLOv11s 0.857 0.718 0.81 0.545 9.41 19.2 YOLO-CPEE 0.864 0.759 0.836 0.577 9.09 18.7

[0067] S6. Input the dense pedestrian image into the trained YOLO-CPEE network model to obtain category, location, and confidence information:

[0068] During the experiment, the YOLO-CPEE model was rigorously tested, and the results are as follows: Figure 6As shown in Figure 2, the experimental results verify the excellent performance of the model in the dense pedestrian target detection task. In particular, under challenging conditions such as some small targets, significant lighting changes, and complex backgrounds, the YOLO-CPEE model demonstrates excellent ability in maintaining high recognition rate and positioning accuracy.

[0069] S7. Analyze the experimental results:

[0070] In order to comprehensively evaluate the superiority and advancement of the model proposed in the embodiment in the target detection task, a detailed comparative experiment was conducted on the current leading target detection model, mainly using the YOLO series model for comparison. Based on precision (P), recall (R), mAP@0.5, mAP@0.5:0.95, Parameters, Size as key indicators for evaluation, the comparative experimental results are summarized in Table 2. The comparative evaluation shows that YOLO-CPEE's comprehensive performance in dense pedestrian detection tasks, including mAP@Size 0.5 and mAP@0.5:0.95, is significantly better than that of the mainstream YOLO series model, and the model parameters are also relatively lightweight. The performance improvement of YOLO-CPEE is significant, which proves that the detection method proposed in the present invention has superior performance.

[0071] Table 2 Comparative experimental results of different series of YOLO models

[0072] Model P / % R / % mAP@0.5 / % mAP@0.5:0.95 / % Parameters / M Size / MB YOLOv5s 0.854 0.712 0.803 0.539 7.81 16.0 YOLOv6s 0.852 0.707 0.795 0.543 15.98 32.2 YOLOv8s 0.856 0.71 0.807 0.543 11.13 22.5 YOLOv9s 0.844 0.723 0.803 0.551 6.19 13.3 YOLOv10s 0.828 0.723 0.803 0.53 7.22 16.5 YOLO11s 0.857 0.718 0.81 0.545 9.41 19.2 YOLO-CPEE 0.864 0.759 0.836 0.577 9.09 18.7

[0073] The test comparison results are as follows Figure 7 Figure 2 compares the performance of seven algorithms in a crowded street scene with densely packed pedestrians and significant overlap in the center. YOLO-CPEE significantly outperforms the other models in detection completeness, identifying nearly all visible pedestrians, while the other models inevitably miss some targets, particularly in highly occluded areas, as indicated by the red boxes in the figure. This demonstrates that YOLO-CPEE is more effective in handling overlapping and occluded scenes.

[0074] Of course, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Changes, modifications, additions or substitutions made by technicians in this technical field within the essential scope of the present invention should also fall within the scope of protection of the present invention.

Claims

1. A lightweight, efficient and dense pedestrian detection method based on YOLO-CPEE, characterized by: The steps include: S1. Acquisition of dense pedestrian dataset; Through web crawlers, public dataset platforms, and field collection, we acquire pedestrian image data from densely populated areas such as urban business districts, subway stations, squares, and parks, building a dataset suitable for dense pedestrian detection. We use annotation tools to accurately label pedestrian targets, ensuring data quality and improving the adaptability of the detection model. S2, image preprocessing; Perform size unification, brightness and contrast adjustment, noise enhancement, and data cleaning preprocessing operations on the acquired pedestrian images; Through data augmentation methods, including random flipping, mosaic splicing, and scale scaling, the training samples are expanded to improve the robustness and generalization ability of the model in complex environments; S3. Configure the required training environment; We built a training environment for the object detection model based on the deep learning framework PyTorch 2.1.

0. The configuration includes GPU acceleration hardware RTX 3080Ti, operating system Ubuntu 22.04, CUDA 12.1, and Python 3.10.

14. We also set training hyperparameters, including learning rate, batch size, and optimizer type. S4. Improve the YOLO11 algorithm model structure to obtain the YOLO-CPEE target detection model; First, we propose a C2CGA module. This module uses CGABlock instead of Bottleneck and replaces the C2PSA module in the YOLO11 backbone. Second, we design a shallow high-resolution feature extraction layer (P2), and add an F2 feature fusion layer and a D2 detection head to the neck network, specifically for feature enhancement and prediction of small targets. Then, the EMA attention mechanism is introduced at multiple stages of the neck network to enhance multi-scale feature extraction. Finally, an efficient detection head is designed, replacing the original ordinary detection head with a lightweight dynamic detection head, which is processed through a multi-scale feature fusion mechanism and detail-enhanced shared convolution. S5. Train the YOLO-CPEE network model; The training set obtained by preprocessing according to step S1 is input into the YOLO-CPEE network model for training, and the validation set is input during the training process. The training effect is verified by the test set, thereby obtaining a trained YOLO-CPEE target detection model; S6. Input the dense pedestrian image into the trained YOLO-CPEE network model to obtain category, location, and confidence information.

2. The lightweight, efficient and dense pedestrian detection method based on YOLO-CPEE according to claim 1 is characterized in that: The YOLO-CPEE network model structure is as follows: The YOLO-CPEE network model structure is mainly composed of three parts in series: backbone network, neck network and detection head; First is the backbone part, where the input goes through eleven layers in sequence: convolution, convolution, C3K2, convolution, C3K2, convolution, C3K2, convolution, C3K2, SPPF, and C2CGA; Then comes the neck part. The output of the trunk part goes through twenty-three layers in the neck, namely upsampling, splicing, C3K2, upsampling, splicing, C3K2, EMA, upsampling, splicing, C3K2, EMA, convolution, splicing, C3K2, EMA, convolution, splicing, C3K2, EMA, convolution, splicing, C3K2, EMA, convolution, splicing, C3K2, EMA, convolution, splicing, C3K2, EMA, totaling twenty-three layers. There are many feature map fusions at different stages in the neck part. The C3K2 layer of the third layer of the trunk part is used as an additional input of the ninth splicing layer of the neck, the C3K2 layer of the fifth layer of the trunk part is used as an additional input of the fifth splicing layer of the neck, the C3K2 layer of the seventh layer of the trunk part is used as an additional input of the second splicing layer of the neck, the C2CGA layer of the eleventh layer of the trunk part is used as an additional input of the first and twenty-first splicing layers of the neck, and the third and sixth C3K2 layers of the neck are used as additional inputs of the thirteenth and seventeenth splicing layers of the neck respectively. Finally, the eleventh, fifteenth, and nineteenth layers of the neck are input into the efficient detection head respectively, forming the head of YOLO-CPEE.

3. The lightweight, efficient and dense pedestrian detection method based on YOLO-CPEE according to claims 1 and 2 is characterized in that: The structure of the C2CGA is as follows: The overall structure of the C2CGA module includes key stages: input convolution, feature partitioning, stacking multiple CGABlocks, feature fusion, and output convolution. Specifically, the input feature map first undergoes a preliminary transformation in the channel dimension through a convolutional layer. A split operation then divides the feature map into multiple subgroups along the channel dimension, which are then fed into multiple parallel CGABlock modules. Each CGABlock consists of two main submodules: a cascaded group attention mechanism module and a feedforward network module. The cascaded group attention mechanism module adopts a multi-head attention structure. In each split attention head, a linear transformation of the query, key, and value is first performed. A token interaction mechanism is then used to perform fine-grained interaction between tokens within the group, significantly improving feature representation while maintaining low computational complexity. The outputs of the multiple split heads are concatenated and linearly fused after the self-attention mechanism to produce the final group attention feature. The FFN part uses two serially connected 1×1 convolution operations to further transform the channel direction and nonlinearly enhance the features at each position. Finally, the output features within each CGABlock module are integrated through a concatenation operation and fused through a convolutional layer to produce the output of the current C2CGA module.

4. The lightweight, efficient and dense pedestrian detection method based on YOLO-CPEE according to claims 1 and 2 is characterized in that: The structure of the EMA is as follows: The size of the input feature map is C×H×W, where C is the number of channels, H and W are the height and width of the feature map respectively. First, the input features are grouped by channels and divided into multiple sub-feature maps according to the number of groups g, and the number of channels of each sub-feature map is C / g; each group of sub-feature maps enters two main branches respectively: the first branch is the local attention enhancement branch: average pooling is performed on the sub-feature map in the X-axis direction and the Y-axis direction respectively to extract local context information; after splicing the pooling results in the two directions, they are integrated through 1×1 convolution, and then two Sigmoid activation functions are input respectively to generate two weight maps; the weight map is multiplied element-by-element with the original sub-feature map to achieve the suppression and enhancement of spatial attention; then group normalization is adopted and the results are fused with the results of the second branch; the second branch is the global relationship modeling branch: the sub-feature map first undergoes 3×3 convolution to expand the context receptive field, and then average pooling is applied; then it enters the Softmax calculation module to generate spatial attention weights; Matrix multiplication is used to interact with downsampled features to achieve long-range dependency modeling; the global weighted result is multiplied element-wise with the output of the first branch to form an enhanced feature; finally, the outputs of the two branches are fused element-by-element and output as the final result of the EMA module.

5. The lightweight, efficient and dense pedestrian detection method based on YOLO-CPEE according to claims 1 and 2 is characterized in that: The structure of the high-efficiency detection head is as follows: The efficient detection head structure adopts a multi-level feature input method, receiving multiple scale feature maps from the feature pyramid, including P2, P3, P4, and P5. Each layer of feature map first passes through a lightweight convolution module CGS for channel compression and nonlinear transformation to enhance feature representation capabilities. Subsequently, the features of each layer are input into a shared convolution module, which contains two consecutive detail enhancement convolutions, which enhance the spatial information of the features while maintaining parameter efficiency. This design can improve the sensitivity of the detection head to targets of different scales, especially small targets. After the features are processed by the shared convolution, they are input into the classification branch and the bounding box regression branch respectively. The former is used to predict the probability of the target category, and the latter is used to predict the bounding box parameters of the target including the center point coordinates, width and height. Both branches adopt an independent convolution structure and attach a scale coefficient Scale module at the end. This module normalizes and fine-tunes the prediction results according to the distribution of the feature map to improve the prediction accuracy and stability. The detail enhancement convolution is composed of four operations: the CDC module, i.e., center difference convolution, the ADC module, i.e., angular difference convolution, the HDC, i.e., horizontal difference convolution, and the VDC, i.e., vertical difference convolution, which are used to model the edge and position information of the target in different directions. This module is connected in parallel with the VC module, i.e., ordinary convolution, in the main path of the feature map, and is fused by residual addition to further enhance the target boundary perception ability.

Citation Information

Cited By

  • Forge piece surface defect detection method based on YOLO model

    CN121481935A