YOLOv8 Dense Pedestrian Detection Method Based on FocalNeXt Fold Focus Stacking Block
By introducing FocalNeXtFold focus stacking blocks, AFPN asymptotic feature pyramid network and VariFocal loss function in YOLOv8, the problem of low accuracy in small-target dense object detection is solved, and the performance and adaptability of dense pedestrian detection is significantly improved.
Patent Information
- Application Number
- CN202311098033.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-29
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2043-08-29
AI Technical Summary
The existing YOLOv8 has low accuracy in detecting dense objects in small targets, making it difficult to effectively detect scenes with dense pedestrians.
A YOLOv8 dense pedestrian detection method based on FocalNeXtFold focused stacking block was designed. By replacing the backbone network C2f block as a new focusing stacking block FocalNeXtFold, and using the AFPN asymptotic feature pyramid network in the head part, the loss function is changed to the VariFocal loss function.
The improved YOLOv8 backbone network has richer features for small target extraction, improving the performance of dense pedestrian detection, and improving detection accuracy and adaptability.
Smart Images

Figure CN117115736B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to dense pedestrian detection technology, and specifically to a YOLOv8 dense pedestrian detection method based on FocalNeXtFold focused stacking blocks. Background Art
[0002] Object detection, as one of the core problems of computer vision, aims to find the category and location of specific objects in an image and has now been widely applied in various fields such as autonomous driving, remote sensing images, video surveillance, and medical detection. Since the development of YOLO in 2016, continuous version updates have been made, and it has now reached v8. In 2016, a one-stage object detection method represented by YOLOv1 emerged. Looking at the development process of one-stage object detection methods, it can be found that from the first one-stage object detection method YOLOv1 proposed to YOLOv8 in 2023, the object detection methods of the YOLO series have developed along with the development of one-stage object detection and have become a typical representative of one-stage methods.
[0003] In recent years, there have been numerous improved versions of YOLOv7. In contrast, for YOLOv8, there are problems with low accuracy in the detection of small target dense objects in existing YOLOv8 applications.
[0004] In recent years, there have been numerous improved versions of YOLOv7, including modifications to the backbone network, feature fusion network, and loss function to improve the accuracy, recall rate, and mean average precision of object detection. Improvements to YOLOv8 have emerged successively since the end of July 2023, and the application scenarios are endless, such as photovoltaic cell defect detection, fish detection, etc. However, there are very few in small target dense detection. Summary of the Invention
[0005] The present invention designs a YOLOv8 dense pedestrian detection method based on FocalNeXtFold focused stacking blocks. By replacing the backbone C2f block with a new type of focused stacking block FocalNeXtFold and stacking several FocalNeXt blocks to form a FocalNeXtFold stacking block, the performance of dense pedestrian detection is improved.
[0006] The technical solution disclosed by the present invention is as follows: A YOLOv8 dense pedestrian detection method based on FocalNeXtFold focused stacking blocks, including the steps of:
[0007] Step S1: Obtain a dataset with a divided training set, test set, and validation set;
[0008] Step S2: Improve the YOLOv8 network by replacing all C2f modules in the Backbone main network with FocalNeXtFold Blocks, which are focused stacking blocks;
[0009] Step S3: Based on Step S2, replace the head part with a network that can enhance the direct interaction between non-adjacent layers;
[0010] Step S4: Based on Step S3, complete the improvement of the model by replacing the loss function of the YOLOv8 network;
[0011] Step S5: Use the divided dataset to train the improved YOLOv8 model in Step S4 to obtain a trained model;
[0012] Step S6: Use the trained model to detect the data to be detected and output the detection results.
[0013] On the basis of the above solution, preferably, the dataset in Step S1 is the WiderPerson outdoor pedestrian detection benchmark dataset.
[0014] On the basis of the above solution, preferably, the FocalNeXtFold Block includes a cv1 convolutional layer, n FocalNextBlock modules, and a cv2 convolutional layer, where n is at least 2.
[0015] On the basis of the above solution, preferably, each FocalNextBlock module includes dwconv, dwconv_3, norm, pwconv1, act, pwconv2, gamma, and drop_path.
[0016] On the basis of the above solution, preferably, the head part is replaced with an AFPN asymptotic feature pyramid network.
[0017] On the basis of the above solution, preferably, the loss function of the YOLOv8 network is changed to the VariFocal loss function.
[0018] Compared with the prior art, the present invention has the following beneficial effects:
[0019] By replacing the C2f block in the main network with a new type of focused stacking block, FocalNeXtFold, and forming the FocalNeXtFold stacking block by stacking several FocalNeXt block blocks, the improved YOLOv8 main network can extract richer features for small targets, thus improving the performance of dense pedestrian detection.
[0020] The head part of YOLOv8 is changed to the AFPN Asymptotic Feature Pyramid Network, which consists of multiple Conv convolutions, ASFF, and BasicBlock blocks to strengthen the direct interaction between non-adjacent layers.
[0021] To improve the accuracy of object detection in dense pedestrian scenarios, the original loss function CIoU is changed to the new VariFocal loss function, which is a variant of the binary cross-entropy loss function (nn.BCEWithLogitsLoss) in binary classification tasks, enhancing the detection accuracy and being more suitable for object detection in dense scenarios. Brief Description of the Drawings
[0022] Figure 1 It shows the position of the FocalNeXtFold Block in the YOLOv8 backbone network;
[0023] Figure 2 It shows the structure of the FocalNextFold Block;
[0024] Figure 3 It is the AFPN pyramid network structure. Detailed Implementation Manner
[0025] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the specific implementation manners of the present invention will be described below with reference to the accompanying drawings. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings and other implementation manners can be obtained.
[0026] As Figures 1-3 shown, the YOLOv8 dense pedestrian detection method based on the FocalNeXtFold focused stacking block includes the steps:
[0027] Step S1: The present invention selects the WiderPerson outdoor pedestrian detection benchmark dataset, in which the dataset has been divided and contains a training set, a test set, and a validation set.
[0028] Step S2: Based on the YOLOv8 network, all C2f modules in the Backbone backbone network are replaced with the designed new focused stacking block FocalNeXtFold Block to extract richer features for small targets.
[0029] FocalNextBlock Basic Structure: This is a class representing a neural network block. It includes some convolutional layers, normalization layers, linear layers, and GELU activation functions, etc., and implements the DropPath operation. This block includes depthwise convolution, layer normalization, linear transformation, activation function, optional scaling (gamma), and DropPath operation.
[0030] The FocalNextFold Block structure is composed of stacked FocalNextBlocks.
[0031] The model starts from the input data. First, it passes through a 1x1 convolutional layer (cv1) to increase the number of input channels from c1 to 2*c, and then enters n FocalNextBlocks. Each FocalNextBlock includes operations such as depthwise convolution, normalization, linear transformation, and activation function, as well as the DropPath operation. The model gradually extracts features through n such blocks. Finally, through another 1x1 convolutional layer (cv2), the number of channels is reduced from (2 + n)*c to c2 to obtain the final output.
[0032] Input: The model accepts the input data x, which is a feature map or image data.
[0033] cv1(1*1Conv): The input data x passes through a 1*1 convolutional layer, and the number of channels increases from c1 to 2*c, which helps to increase the feature dimension. By introducing more feature channels, it is possible to better capture different feature information in subsequent operations.
[0034] FocalNextBlock 1,2,...,n: The model contains n FocalNextBlocks, and each block operates as follows:
[0035] dwconv (Depthwise Convolution): Process the input using depthwise separable convolution. This helps to learn spatial information and enables the model to capture information at different scales.
[0036] dwconv_3 (Depthwise Convolution, dilation = 3): Similar to dwconv, but uses a larger dilated convolutional kernel to increase the receptive field and capture context information at a greater distance.
[0037] norm (Normalization): Normalize the convolutional features to have standard mean and variance. Helps to improve the training stability of the model.
[0038] pwconv1 (Pointwise Convolution 1): Perform a linear transformation on the normalized features to expand the feature dimension to 4*c to increase the feature representation ability.
[0039] act (GELU activation function): Apply the GELU activation function to introduce non-linear transformation, enhance the feature expression ability, and improve the non-linear fitting ability of the model.
[0040] pwconv2 (Linear transformation 2): Perform a linear transformation on the activated features to reduce the feature dimension back to c, compressing the feature representation.
[0041] gamma (Optional scaling): If layer_scale_init_value is set to be greater than 0, use the gamma parameter to scale the features for further adjustment of feature expression.
[0042] drop_path (DropPath operation): Apply the DropPath operation to enhance the robustness of the model, randomly discard a part of the features to avoid overfitting.
[0043] cv2 (1x1 Conv): After passing through n FocalNextBlocks, the features are further passed to a 1x1 convolutional layer to reduce the number of channels from (2 + n)*c to c2, preparing for the output stage.
[0044] Output: The final output of the model is the feature map after cv2, which has c2 channels and can be used for subsequent tasks such as classification and detection.
[0045] The FocalNextFold Block structure gradually transforms and extracts the input features through multiple FocalNextBlocks to generate more expressive feature representations. This helps to improve the performance and generalization ability of the model.
[0046] Step S3: Based on the YOLOv8 network, replace all the structures in the YOLOv8 head part with the AFPN (Asymptotic Feature Pyramid Network) that contains multiple Conv convolutions, ASFF, and BasicBlock blocks to strengthen the direct interaction between non-adjacent layers.
[0047] The AFPN (Adaptive Feature Pyramid Network) pyramid structure has the following functions. It integrates feature information at multiple scales to improve the performance of the model:
[0048] Multi-scale feature integration: AFPN integrates feature information at different scales, enabling the model to capture different scales and context information of the target from multiple levels. This helps to improve the model's detection and recognition ability for objects of different scales.
[0049] Adaptive Feature Fusion: Through the ASFF_2 and ASFF_3 modules, AFPN can adaptively fuse feature maps of different scales, and perform weighted combination of features at each scale according to the needs of the current task. This helps the model have better expressive ability in different scenarios.
[0050] Cascade Structure: In the AFPN pyramid structure, the ASFF_2 and ASFF_3 modules are cascaded step by step, and each module can process more scale feature information. This cascade structure helps the model gradually expand the receptive field, so as to better understand the context information in the image.
[0051] Channel Compression: Each ASFF module uses channel compression technology to reduce the number of channels of the input feature map, thereby reducing the computational amount and improving the efficiency of the model. This helps reduce resource consumption while maintaining performance.
[0052] Adaptive Weight Calculation: Each ASFF module adaptively calculates the fusion weights according to the importance of features at different scales. This enables the model to better focus on features at important scales, thereby improving the accuracy of object detection and segmentation.
[0053] Pyramid Pooling: Through the Upsample and Downsample modules, AFPN realizes the function of pyramid pooling and can perform feature aggregation at different scales. This helps the model obtain feature information from multiple levels and enhances the model's adaptability to scale changes.
[0054] The AFPN pyramid has significant functions in multi-scale feature integration, adaptive fusion, cascade structure, etc., and can effectively improve the performance of the model in object detection.
[0055] Step S4: Based on the YOLOv8 network, change the loss function of the YOLOv8 network to the latest VariFocal loss function to comprehensively improve the detection accuracy in dense scenes. Build an object detection model based on Steps S2 and S3.
[0056] VariFocal Loss, used for binary classification tasks, is mainly used to improve the cross-entropy loss function to better handle the problem of weight imbalance between positive and negative samples. The functions of VariFocal Loss are as follows:
[0057] Weight Adjustment: Different weight adjustments are made for positive and negative samples to cope with the problem of class imbalance. Usually, the number of negative samples is significantly more than that of positive samples, which will cause the model to be more prone to bias towards negative samples during the learning process and affect the performance of the model. VariFocal Loss improves the performance by adjusting the weights of positive and negative samples, making the model pay more attention to difficult-to-classify samples.
[0058] Adaptive weight: Adjust the weight adaptively according to the difficulty of the samples. Specifically, for samples that are easy to classify, their weights will pay more attention to the differences between samples to enhance the model's learning ability for these samples. For samples that are difficult to classify, their weights will pay more attention to the confidence of correct classification, thus reducing the impact of misclassification.
[0059] Convex loss curve: By introducing the gamma parameter, VariFocal Loss makes the loss curve convex when the confidence is high, thus paying more attention to samples with high confidence and reducing the contribution of samples with low confidence to the loss.
[0060] Alpha parameter: Through the alpha parameter, the weight of positive samples can be adjusted to further balance the influence of different classes. The alpha parameter controls the ratio of the weights of positive and negative samples, which can be adjusted for different tasks and data distributions to obtain better performance.
[0061] Step S5: Use the divided dataset to train the object detection model in step S4;
[0062] Step S6: Use the trained object detection model to perform object detection on the training set and output the detection results.
[0063] It should be noted that the above embodiments can be freely combined as needed. The above are only the preferred embodiments of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A YOLOv8 dense pedestrian detection method based on the FocalNeXt Fold focusing stack block, characterized in that Including the steps: Step S1: Obtain a dataset with the training set, test set, and validation set partitioned; Step S2: Improve the YOLOv8 network by replacing all C2f modules in the Backbone backbone network with FocalNeXtFold Blocks. The FocalNeXtFold Block includes a cv1 convolutional layer, n FocalNextBlock modules, and a cv2 convolutional layer, where n is at least 2. Each FocalNextBlock module includes dwconv, dwconv_3, norm, pwconv1, act, pwconv2, gamma, and drop_path; Step S3: Based on Step S2, replace the head part with a network that can enhance the direct interaction between non-adjacent layers; Step S4: Based on Step S3, replace the loss function of the YOLOv8 network to complete the improvement of the model; Step S5: Use the partitioned dataset to train the improved YOLOv8 model in Step S4 to obtain a trained model; Step S6: Use the trained model to detect the data to be detected and output the detection results.
2. The YOLOv8 dense pedestrian detection method based on the FocalNeXtFold focus stacking block according to claim 1, wherein, The dataset in Step S1 is the WiderPerson outdoor pedestrian detection benchmark dataset.
3. The YOLOv8 dense pedestrian detection method based on the FocalNeXtFold focus stacking block according to claim 1, wherein, The head part is replaced with an AFPN asymptotic feature pyramid network.
4. The YOLOv8 dense pedestrian detection method based on the FocalNeXtFold focus stacking block according to claim 1, wherein, The loss function of the YOLOv8 network is changed to the VariFocal loss function.
Citation Information
Patent Citations
Pedestrian occlusion detection method based on improved YOLOX algorithm
CN115082855A
Tunnel punching inclination angle correction method based on depth camera
CN116579989A