Dense pedestrian target detection method based on improved YOLOv11

By introducing GhostConv, WTConv and MLCA attention mechanisms into the backbone network of the YOLOv11 model, the problem of low accuracy of dense pedestrian detection in complex environments is solved, and higher robustness and detection accuracy are achieved.

CN119942598AActive Publication Date: 2025-05-06GUILIN UNIV OF ELECTRONIC TECH

Patent Information

Application Number
CN202510239308.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-05-06
Estimated Expiration
2045-03-03

AI Technical Summary

Technical Problem

The existing pedestrian detection algorithms are difficult to accurately identify pedestrian targets that are partially or completely blocked in complex environments where crowds are highly concentrated and mutually blocked, resulting in a significant reduction in detection accuracy.

Method used

By introducing the GhostConv module, WTConv module and MLCA attention mechanism into the backbone network of YOLOv11, an improved intensive pedestrian object detection model is built. The GhostConv module reduces the calculation cost of redundant information. The WTConv module captures a larger range of context information through wavelet transformation. The MLCA attention mechanism integrates multi-dimensional information to enhance pedestrian feature expression ability.

Benefits of technology

It significantly improves the robustness and detection accuracy of the model in dense scenarios, reduces the missed detection rate caused by occlusion, and enables the model to more effectively identify the obstructed pedestrian targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942598A_ABST
    Figure CN119942598A_ABST
Patent Text Reader

Abstract

The invention discloses a dense pedestrian target detection method based on improved YOLOv11, and the method achieves the efficient detection of dense pedestrians through the construction of an improved YOLOv11 model. The method comprises the following steps: firstly, collecting a dense pedestrian data set and carrying out image preprocessing, and then constructing an improved YOLOv11 model; the specific improvement steps are as follows: a GhostConv module is used to replace common convolution in a backbone network, redundant information in a feature map is reduced, the calculation cost is reduced, and model lightweight is realized; a C3K2 module is replaced by C3K2-WTConv, the convolution receptive field is expanded, low-frequency information in the image is effectively captured, and the dense target detection effect is remarkably improved; a lightweight MLCA attention mechanism is introduced into a C2PSA module to form a new C2PSA-MLCA module, and local and global information is combined to enhance the network expression ability. And after model construction is completed, training, testing and performance evaluation are carried out. According to the method, the model lightweight is realized, the target detection accuracy is improved, and the method is suitable for traffic monitoring and safety monitoring of public places.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of target detection, and in particular to a dense pedestrian target detection method based on improved YOLOv11. Background Art

[0002] As one of the research directions in the field of target detection, pedestrian detection has a wide range of application value in the field of intelligent transportation and security. This technology provides road environment perception capabilities for assisted driving systems by real-time identification and positioning of pedestrian targets, realizes accurate target tracking in the field of road monitoring, and plays an important role in early warning and protection systems. Especially in public places with dense crowds, such as commercial complexes and tourist attractions, pedestrian detection technology can provide reliable data support for passenger flow statistics, behavior analysis and security management, and effectively improve the overall efficiency of intelligent security systems. However, existing pedestrian detection algorithms still face significant technical bottlenecks when dealing with large-scale dense scenes. Especially in complex environments where people are highly concentrated and occluded, the algorithm often finds it difficult to accurately identify pedestrian targets that are partially or completely occluded, resulting in a significant decrease in detection accuracy. This technical difficulty has become a key factor restricting the actual application effect of pedestrian detection, and it is also a core challenge that academia and industry are currently concerned about. How to improve the robustness of the algorithm in crowded scenes and reduce the missed detection rate caused by occlusion is still a research focus that needs to be broken through.

[0003] In recent years, the rapid evolution of deep learning technology has promoted the innovation of target detection methods. In the framework of deep neural networks, detection algorithms have made breakthroughs in two main directions: one is the two-stage detection method represented by R-CNN and Fast R-CNN, which achieves target positioning through two steps: region proposal and feature extraction; the other is the single-stage detection method represented by the YOLO series and SSD series, which directly predicts the target position and category in an end-to-end manner. Compared with traditional detection methods, these new algorithms based on deep learning have shown significant advantages in detection accuracy and efficiency, and have become the mainstream technical route in the field of target detection.

[0004] In order to improve the target detection accuracy of the model and achieve lightweight, the present invention improves the YOLOv11 model. By introducing GhostConv convolution, the model is lightweight, so as to better adapt to the deployment requirements of terminal devices such as drones. In addition, combined with WTConv convolution and MLCA attention mechanism, the detection accuracy of the model is effectively improved without significantly increasing the number of parameters, and the problem of dense small target detection is also effectively solved. Summary of the invention

[0005] The purpose of the present invention is to propose a dense pedestrian target detection method based on improved YOLOv11 to solve the problems of low detection accuracy of dense small targets, large number of model parameters and high computational complexity.

[0006] To achieve the above purpose, the technical solution adopted by the present invention is as follows:

[0007] A dense pedestrian target detection method based on improved YOLOv11 includes the following steps:

[0008] Step 1: Collect dense pedestrian datasets and perform image expansion and label preprocessing on the datasets;

[0009] Step 2: Build a dense pedestrian target detection model based on improved YOLOv11;

[0010] Wherein, the step 2 specifically includes the following steps:

[0011] Step 2.1: In the backbone network of YOLOv11, replace the Conv ordinary convolution of the network with the GhostConv module;

[0012] Step 2.2: Introduce the WTConv module into the C3K2 module to form the C3K2-WTConv module, and replace the C3K2 module with the C3K2-WTConv module in the backbone network of YOLOv11;

[0013] Step 2.3: Introduce the MLCA attention mechanism into the C2PSA module to form the C2PSA-MLCA module, and replace the C2PSA module with the C2PSA-MLCA module in the backbone network of YOLOv11;

[0014] Step 3: Train the dense pedestrian target detection model based on improved YOLOv11;

[0015] Step 4: Use the trained dense pedestrian target detection model based on improved YOLOv11 to evaluate the performance of the dense pedestrian test set.

[0016] 2. The dense pedestrian target detection method based on improved YOLOv11 according to claim 1 is characterized in that: step 1 specifically comprises the following steps:

[0017] Step 1.1: Filter the collected dense pedestrian dataset and manually delete animated images and images with a small number of pedestrians;

[0018] Step 1.2: Write a corresponding classification program to remove the ignored area and crowd labels in the annotation file in the dataset, retain the pedestrian, rider, and partially visible person labels and classify them into the same label Person, and convert the annotation file into YOLO format;

[0019] Step 1.3: Divide the dataset into training set and test set;

[0020] 3. The dense pedestrian target detection method based on improved YOLOv11 according to claim 1 is characterized in that: step 2.1 specifically comprises the following steps:

[0021] Step 2.1.1: First, use 1×1 ordinary convolution, batch normalization layer and a ReLu activation function to compress the number of channels of the input image to generate an inherent feature map, and then apply a series of low-cost linear operations to the feature map Get more different feature maps and increase features. Among them, low-cost linear operations It consists of a depth-wise separable convolution, a batch normalization layer, and a ReLU activation function;

[0022] Step 2.1.2: Combine the feature map obtained through the low-cost linear operation and the feature map obtained through ordinary convolution, batch normalization and a ReLU activation function in step 2.1.1 through the Concat operation. The final feature map is the output.

[0023] 4. The dense pedestrian target detection method based on improved YOLOv11 according to claim 1, characterized in that step 2.2 specifically comprises the following steps:

[0024] Step 2.2.1: WTConv first uses the two-dimensional Haar wavelet transform to perform multi-level decomposition on the input image. The Haar wavelet transform uses four filters to decompose the image into four sub-bands: low-frequency components, horizontal high-frequency components, vertical high-frequency components, and diagonal high-frequency components. In each level of wavelet transform, the image is downsampled, so that the spatial resolution is halved, but the frequency information is more finely decomposed, thereby obtaining frequency components at different scales;

[0025] Step 2.2.2: Then perform group convolution operation, use a small-size depth convolution kernel on each frequency sub-band, use a 7×7 small convolution kernel, and perform convolution operation on each decomposed sub-band. The process is as follows: [X LL ,X LH ,X HL ,X HH ]=Conv([f LL ,f LH ,f HL,f HH ],X)

[0026] Where X is the input tensor, f LL is a low-pass filter, f LH ,f HL , f HH is a set of high-pass filters, X LL is the low frequency component, X LH , X HL , X HH They are the high frequency components of horizontal, vertical and diagonal lines respectively;

[0027] Step 2.2.3: The IWT operation is linear and can losslessly reconstruct the WT convolution results and the grouped convolution results into the original space. After the grouped convolution operation is completed, the IWT inverse wavelet transform is used to resynthesize the convolution results of each subband into a complete output. The process is as follows: Y = IWT(Conv(W, WT(X)))

[0028] Where X is the input tensor and W is the weight tensor of a k×k depthwise convolution kernel with four times the number of input channels as X.

[0029] 5. The dense pedestrian target detection method based on improved YOLOv11 according to claim 1, characterized in that step 2.3 specifically comprises the following steps:

[0030] Step 2.3.1: The input feature map is first processed by local average pooling and global average pooling. Secondly, the features after local pooling and the features after global pooling are transformed by a 1D convolution, so that the features are rearranged and adapted to subsequent operations;

[0031] Step 2.3.2: Use 1D convolution to rearrange the local pooled features. This process is equivalent to feature selection, which strengthens the focus on useful features;

[0032] Step 2.3.3: After 1D convolution, rearrangement and unpooling (UNAP) of the global pooled features, they are combined with the local pooled features through addition operations. This process incorporates global context information in the feature map;

[0033] Step 2.3.4: Finally, the feature maps processed by local and global attention are combined, and then go through the unpooling operation (UNAP) again, and then combined with the original input features through a multiplication operation to restore the original spatial dimension.

[0034] 6. The dense pedestrian target detection method based on improved YOLOv11 according to claim 1, characterized in that step 3 specifically comprises the following steps:

[0035] Step 3.1: Divide the preprocessed dense pedestrian dataset into a training set and a test set in a ratio of 7:3;

[0036] Step 3.2: Set the number of iterations to 200, the batch size to 16, and the optimizer to auto;

[0037] Step 3.3: Perform training according to the set parameters in the dense pedestrian target detection model based on the improved YOLOv11 to obtain a trained dense pedestrian target detection model based on the improved YOLOv11.

[0038] The beneficial effects of the present invention are as follows:

[0039] 1. The present invention replaces the Conv ordinary convolution with the GhostConv module in the backbone network of YOLOv11. By separating the generation process of the feature map, it significantly reduces the computational cost of redundant information while retaining the feature expression capability, so that the model can be deployed on a mobile terminal.

[0340] 2. In the backbone network of YOLOv11, the present invention replaces the C3K2 module with the C3K2-WTConv module. The wavelet transform enables the module to effectively capture a wider range of contextual information in the image without increasing too much computation, thereby improving the accuracy of target positioning. In addition, WTConv effectively alleviates the problem of fine-grained feature loss caused by excessive response of traditional convolution to high-frequency information by enhancing low-frequency feature extraction, thereby improving the accuracy of target detection.

[0041] 3. The present invention introduces the MLCA attention mechanism in C2PSA to form the C2PSA-MCLA module in the backbone network of YOLOv11, fuses multi-dimensional information through local and global pooling methods, enhances the ability to express pedestrian features, and through 1D convolution operations, significantly reduces parameter complexity and computational complexity while improving the accuracy of dense pedestrian target detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] The following will provide a detailed description through specific implementation methods and drawings.

[0043] Figure 1 This is a schematic diagram of the overall structure of the improved YOLOv11 network;

[0044] Figure 2 It is a schematic diagram of the structure of GhostConv;

[0045] Figure 3 It is a schematic diagram of the structure of WTConv;

[0046] Figure 4 Schematic diagram of the MLCA attention mechanism structure;

[0047] Figure 5 This is a flowchart of the dense pedestrian target detection method based on improved YOLOv11 proposed in the present invention. DETAILED DESCRIPTION

[0048] In order to make the implementation modes of the present invention and the prior art solutions clearer and easier to understand, the present invention is described in detail below with reference to the accompanying drawings.

[0049] Reference Figure 5 , the steps to obtain the dense pedestrian target detection model based on improved YOLOv11 are as follows:

[0050] Step 1: Use various large search engines to collect dense pedestrian datasets and pre-process the dataset images and labels;

[0051] Step 1.1: The images in the dataset collected for expansion contain many images that are not suitable for training, such as cartoon images, images with few pedestrians in the crowd, and images with high blur. Therefore, the collected dense pedestrian dataset is manually screened, and the images that are not suitable for training are deleted. Then, the pedestrians in each image are labeled using the rectangular annotation in the LabelImg image annotation tool, and the labels are uniformly set to the person class. After the annotation is completed, the corresponding txt label file is obtained to record the information of the bounding box;

[0052] Step 1.2: The image labels of some of the collected data sets include five types of labels: pedestrians, riders, partially visible persons, ignored areas, and crowds. Therefore, a corresponding classification program is written to remove the ignored areas and crowds in the annotation files in the data set, retain the pedestrians, riders, and partially visible persons, and classify them into the same annotation Person, and convert the annotation files into the YOLO format;

[0053] Step 1.3: Integrate the collected data sets, images and labels, and divide them into training and test sets in a 7:3 ratio;

[0054] Step 2: Build a dense pedestrian target detection model based on improved YOLOv11;

[0055] In order to solve the problem of dense pedestrian target detection, the YOLOv11 network is improved from the following three aspects: Figures 1 to 4 As shown:

[0056] Step 2.1: In order to reduce the amount of computation and complexity of the network, in the backbone network of YOLOv11, the Conv ordinary convolutions of the 0th, 1st, 3rd, 5th, and 7th layers of the network are replaced with GhostConv modules. First, a 1×1 convolution is performed on the input feature map to generate an intermediate feature map (intrinsic feature maps) with a smaller number of channels, and the information features between channels are aggregated; then a depthwise separable convolution is applied to the output feature map of the 1×1 convolution to generate redundant feature maps (Ghost feature maps) with the same number as the initial feature map; finally, the intermediate feature map and the redundant feature map are spliced ​​in the channel dimension to obtain the final feature map.

[0057] Step 2.2: Introduce the WTConv module into the C3K2 module to form the C3K2-WTConv module, and replace the 2nd, 4th, 6th, and 8th layer C3K2 modules of the network with C3K2-WTConv modules in the backbone network of YOLOv11;

[0058] Step 2.3: Introduce the MLCA attention mechanism into the C2PSA module to form the C2PSA-MLCA module, and replace the C2PSA module on the 10th layer of the YOLOv11 backbone network with the C2PSA-MLCA module;

[0059] Wherein, step 2.1 comprises the following steps:

[0060] Step 2.1.1: First, use 1×1 ordinary convolution, batch normalization layer and a ReLu activation function to compress the number of channels of the input image to generate an inherent feature map, and then apply a series of low-cost linear operations to the feature map Get more different feature maps and increase redundant features. If the original convolution needs to generate feature maps with C channels, this convolution will generate feature maps with C / 2 channels, and then generate new C / 2 new channels through 3×3 depthwise separable convolution. Among them, low-cost linear operations It consists of a depth-wise separable convolution, a batch normalization layer, and a ReLU activation function;

[0061] Step 2.1.2: Use the Concat operation to concatenate the feature map obtained through the low-cost linear operation and the feature map obtained through 1×1 ordinary convolution, batch normalization and a ReLU activation function in step 2.1.1. At this time, the number of channels of the output feature map is restored to the level of the original convolution, and the final feature map is the output.

[0062] Wherein, step 2.2 comprises the following steps:

[0063] Step 2.2.1: WTConv first uses the two-dimensional Haar wavelet transform to perform multi-level decomposition on the input image. The Haar wavelet transform uses four filters to decompose the image into four sub-bands: low-frequency components, horizontal high-frequency components, vertical high-frequency components, and diagonal high-frequency components. In each level of wavelet transform, the image is downsampled, so that the spatial resolution is halved, but the frequency information is more finely decomposed, thereby obtaining frequency components at different scales;

[0064] Step 2.2.2: Next, perform a convolution operation. Use a small-sized depth convolution kernel on each frequency subband. Use a 7×7 small convolution kernel to perform a convolution operation on each decomposed subband. The process is as follows: [X LL ,X LH ,X HL ,X HH ]=Conv([f LL ,f LH ,f HL ,f HH ],X)

[0065] Where X is the input tensor, f LL is a low-pass filter, f LH ,f HL , f HH is a set of high-pass filters, X LL is the low frequency component, X LH , X HL , X HH They are the high frequency components of horizontal, vertical and diagonal lines respectively;

[0066] Step 2.2.3: The IWT operation is linear and can losslessly reconstruct the convolution result to the original space. After the convolution operation is completed, the IWT inverse wavelet transform is used to re-synthesize the convolution results of each sub-band into a complete output. The process is as follows: Y = IWT(Conv(W, WT(X)))

[0067] Where X is the input tensor and W is the weight tensor of a k×k depthwise convolution kernel with four times the number of input channels as X.

[0068] Wherein, step 2.3 comprises the following steps:

[0069] Step 2.3.1: Input feature map X∈R B×C×H×W First, it is divided into multiple local area blocks. Adaptive average pooling (AdaptiveAvgPool2d) is usually used to compress the feature map to a fixed size to extract local spatial information, reduce the size of the feature map, and reduce the complexity of subsequent calculations. The process is as follows: X local =AdaptiveAvgPool2d(X,S)

[0070] Among them, B is the batch size, C is the number of input feature map channels, H and W are the height and width of the input feature map respectively;

[0071] Step 2.3.2: Next, perform local feature extraction: extract the local feature X after pooling local ∈R B×C×S×S Perform 1D convolution to calculate the local dependencies between channels, where B is the batch size, C is the number of input feature map channels, and S is the local block size;

[0072] Step 2.3.3: Further perform global average pooling (GAP) on the local features to generate the global channel feature X glocal ∈R B×C×1×1 , and then capture the global channel weights through 1D convolution. This process is equivalent to a feature selection, which strengthens the focus on useful features;

[0073] Step 2.3.4: Combine the global pooling feature weights with the local pooling feature weights through an addition operation. This process incorporates global context information in the feature map. The process is as follows: W final =λ·W local +(1-λ)·W glocal

[0074] Among them, W loca =Conv1d(X local ) and W glocal =Conv1d(X glocal ) are local weight and global weight respectively;

[0075] Step 2.3.5: Finally, after the feature maps processed by local and global attention are combined, they are again processed through the unpooling operation (UNAP), and then combined with the original input features through the multiplication operation to restore the original spatial dimension, enhance important features and suppress redundant information. The process is as follows:

[0076] Among them, the unpooling operation adjusts the weights back to the original input size H×W through adaptive upsampling (AdaptiveAvgPool2d) to ensure alignment with the input feature map, σ is the Sigmoid function, Represents element-wise multiplication.

[0077] Step 3: Training the dense pedestrian target detection model based on improved YOLOv11 includes the following steps;

[0078] Step 3.1: Divide the preprocessed dense pedestrian dataset into a training set and a test set in a ratio of 7:3, and input the training set into the improved YOLOv11;

[0079] Step 3.2: Design a training program. Before training, resize the images to 640×640 and perform data augmentation operations, including random cropping and random flipping, to increase the amount of data and improve the generalization ability of the model. Set the number of epochs to 300, the batch size to 16, and the optimizer to auto.

[0080] Step 3.3: Train the dense pedestrian target detection model based on the improved YOLOv11 according to the set parameters, calculate the total loss function and use this data to optimize the model parameters, and obtain the trained dense pedestrian target detection model based on the improved YOLOv11;

[0081] Among them, the loss function includes positioning loss (bounding box regression loss), confidence loss and classification loss. Positioning loss is used to measure the difference between the center coordinates (x, y) and size (w, h) of the predicted box and the real box. The calculation formula is as follows:

[0082] Among them, S 2 Indicates that the image is divided into S×S grids, B indicates the number of candidate boxes predicted for each grid, is an indicator function, which is 1 when the jth candidate box of the i-th grid is responsible for detecting the object, otherwise it is 0, and λ coord is the weight parameter used to enhance the weight of position loss, x i ,y i Indicates the offset of the center of the prediction box relative to the grid, is the true value;

[0083] Confidence loss is used to measure the confidence of whether the predicted box contains the object (i.e. the IoU between the predicted box and the real box). The calculation formula is as follows:

[0084] Among them, C i Indicates the confidence of the prediction box (the value range is 0 to 1), represents the true confidence (if the IoU between the candidate box and the real box is the largest, it is 1, otherwise it is 0), noobj Represents the weight parameter, which is used to reduce the loss weight of the area without objects. Represents an indicator function, which is 1 when the grid does not contain an object.

[0085] The classification loss is used to measure the accuracy of the category prediction of the prediction box, and the calculation formula is as follows:

[0086] in, Indicator function, which is 1 when the grid contains an object, p i (c) represents the predicted category probability, represents the true class probability.

[0087] In summary, the total loss of YOLOv11 is the weighted sum of the above three parts, and the calculation formula is as follows: L total =L coord +L conf +L class

[0088] Step 4: After training is completed, the model saved after every 10 epochs is obtained. Therefore, the total loss function and accuracy of every 10 epochs during the training process are compared, the model with the best performance is selected, and the performance of the model is evaluated on the dense pedestrian test set.

[0089] The performance evaluation of the present invention includes three key indicators: precision (p), recall (R), and mean average precision (mAP). The specific formula is as follows:

[0090]

[0091]

[0092]

[0093] Among them, TP (True Positives) is the number of correctly predicted positive samples, FN (False Negatives) is the number of incorrectly predicted negative samples, and FP (False Positives) is the number of incorrectly predicted positive samples.

[0094] The dense pedestrian target detection method based on improved YOLOv11 provided by the present invention fully pays attention to the redundant information in the dense pedestrian feature map, enhances the model's feature recognition ability for shape rather than texture, and enables the model to improve the accuracy of target detection while achieving lightweight, and is particularly suitable for public security fields such as subway station passenger flow analysis and urban intersection pedestrian detection.

[0095] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be within the protection scope of the present invention.

Claims

1. A dense pedestrian target detection method based on improved YOLOv11, characterized in that: The steps include: Step 1: Collect dense pedestrian dataset and perform image preprocessing on the dataset; Step 2: Build a dense pedestrian target detection model based on improved YOLOv11; Wherein, the step 2 specifically includes the following steps: Step 2.1: In the backbone network of YOLOv11, replace the ordinary convolution of the network with the GhostConv module; Step 2.2: Introduce the WTConv module into the C3K2 module to form the C3K2-WTConv module, and replace the C3K2 module with the C3K2-WTConv module in the backbone network of YOLOv11; Step 2.3: Introduce the MLCA attention mechanism into the C2PSA module to form the C2PSA-MLCA module, and replace the C2PSA module with the C2PSA-MLCA module in the backbone network of YOLOv11; Step 3: Train the dense pedestrian target detection model based on improved YOLOv11; Step 4: Use the trained dense pedestrian target detection model based on improved YOLOv11 to evaluate the performance of the dense pedestrian test set.

2. The dense pedestrian target detection method based on improved YOLOv11 according to claim 1, characterized in that: Step 1 specifically includes the following steps: Step 1.1: Filter the collected dense pedestrian dataset and manually delete animated images and images with a small number of pedestrians; Step 1.2: Write a corresponding classification program to remove the ignored area and crowd labels in the annotation file in the dataset, retain the pedestrian, rider, and partially visible person labels and classify them into the same label Person, and convert the annotation file into YOLO format; Step 1.3: Divide the dataset into training set and test set.

3. The dense pedestrian target detection method based on improved YOLOv11 according to claim 1, characterized in that: Step 2.1 specifically includes the following steps: Step 2.1.1: First, use 1×1 ordinary convolution, batch normalization layer and a ReLu activation function to compress the number of channels of the input image to generate an inherent feature map, and then apply a series of low-cost linear operations to the feature map Get more different feature maps and increase features, among which low-cost linear operations It consists of a depth-wise separable convolution, a batch normalization layer, and a ReLU activation function; Step 2.1.2: Combine the feature map obtained through the low-cost linear operation and the feature map obtained through ordinary convolution, batch normalization and a ReLU activation function in step 2.1.1 through the Concat operation. The final feature map is the output.

4. The dense pedestrian target detection method based on improved YOLOv11 according to claim 1, characterized in that: Step 2.2 specifically includes the following steps: Step 2.2.1: WTConv first uses the two-dimensional Haar wavelet transform to perform multi-level decomposition on the input image. The Haar wavelet transform uses four filters to decompose the image into four sub-bands: low-frequency component, horizontal high-frequency component, vertical high-frequency component, and diagonal high-frequency component. In each level of wavelet transform, the image is downsampled, so that the spatial resolution is halved, but the frequency information is more finely decomposed, thereby obtaining frequency components at different scales; Step 2.2.2: Next, perform a convolution operation. Use a small-sized depth convolution kernel on each frequency subband. Use a 7×7 small convolution kernel to perform a convolution operation on each decomposed subband. The process is as follows: [X LL ,X LH ,X HL ,X HH ]=Conv([f LL ,f LH ,f HL ,f HH ],X) Where X is the input tensor, f LL is a low-pass filter, f LH ,f HL , f HH is a set of high-pass filters, X LL is the low frequency component, X LH , X HL , X HH They are the high frequency components of horizontal, vertical and diagonal lines respectively; Step 2.2.3: The IWT operation is linear and can losslessly reconstruct the convolution result to the original space. After the convolution operation is completed, the IWT inverse wavelet transform is used to re-synthesize the convolution results of each sub-band into a complete output. The process is as follows: Y = IWT(Conv(W, WT(X))) Where X is the input tensor and W is the weight tensor of a k×k depthwise convolution kernel with four times the number of input channels as X.

5. The dense pedestrian target detection method based on improved YOLOv11 according to claim 1, characterized in that: Step 2.3 specifically includes the following steps: Step 2.3.1: The input feature map is first processed by local average pooling and global average pooling. Secondly, the features after local pooling and the features after global pooling are transformed by a 1D convolution, so that the features are rearranged and adapted to subsequent operations; Step 2.3.2: For the local pooled features, use 1D convolution and then rearrange them. This process is equivalent to a feature selection, which strengthens the focus on useful features. Step 2.3.3: After 1D convolution, rearrangement and unpooling operations, the global pooled features are combined with the local pooled features through addition operations. This process integrates the global context information in the feature map. Step 2.3.4: Finally, the feature maps processed by local and global attention are combined, and then unpooled again, and then combined with the original input features through multiplication operation to restore to the original spatial dimension.

6. The dense pedestrian target detection method based on improved YOLOv11 according to claim 1, characterized in that: Step 3 specifically includes the following steps: Step 3.1: Divide the preprocessed dense pedestrian dataset into a training set and a test set in a ratio of 7:3; Step 3.2: Set the number of iterations to 200, the batch size to 16, and the optimizer to auto; Step 3.3: Perform training according to the set parameters in the dense pedestrian target detection model based on the improved YOLOv11 to obtain a trained dense pedestrian target detection model based on the improved YOLOv11.

Citation Information

Patent Citations

  • Improved YOLOv5 lightweight community scene pedestrian detection method

    CN115862066A

  • Light-weight YOLO model method based on grouping fast spatial pyramid pooling

    CN116797910A

Cited By

  • Lightweight target detection method and device for remote sensing image, equipment and medium

    CN120279260A

  • Tomato image real-time detection method, tomato image real-time detection system, tomato picking method and tomato picking system

    CN120388367A

  • Multi-modal fusion-based ship course obstacle monitoring method and ship collision avoidance method

    CN120496001A

  • Bad driving behavior identification method and device based on YOLOv11 model

    CN120894766A