Multi-scale target detection method based on hierarchical feature aggregator (HFA)

The hierarchical feature aggregator (HFA) and dynamic weight fusion (DWFF) algorithm enhance multi-scale feature fusion to address the challenges of target detection in complex environments, improving accuracy and robustness, especially for small targets, while maintaining low computational overhead.

CN120318489APending Publication Date: 2025-07-15GUILIN UNIV OF ELECTRONIC TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510378842.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing target detection methods are difficult to effectively integrate multi-scale feature information in complex environments, especially under low illumination, rain and fog, which leads to insufficient detection accuracy of small targets and lack of effective weight adjustment mechanisms, affecting the detection effect.

Method used

Using a multi-scale object detection method based on hierarchical feature aggregator (HFA), a multi-scale feature map is extracted through the YOLOv8 backbone network, and a cross-scale weighted fusion is used to optimize the feature map quality with dynamic weighted feature fusion (DWFF) algorithm, and finally object detection is performed on the YOLOv8 head network.

Benefits of technology

It significantly improves the target detection accuracy in complex environments, especially the small target detection capability, improves the robustness and detection accuracy of the model, while maintaining a low computational complexity, and is suitable for real-time object detection tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318489A_ABST
    Figure CN120318489A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-scale target detection method based on a hierarchical feature aggregator (HFA), and belongs to the field of computer vision. In order to solve the problem that the existing target detection technology is poor in vehicle detection effect in a complex road environment (such as low illumination, rain and fog, glare and the like), the invention provides a novel multi-scale target detection method, and the vehicle detection precision in the complex road environment is improved. Firstly, an input image is preprocessed, and a YOLOv8 backbone network is utilized to extract bottom layer features; then, the HFA module designed in the invention carries out weighted fusion on features of different scales through cross-scale connection. Wherein the fusion algorithm comprises the steps of weight non-negatization, Swsh activation function, L2 regularization, weighted feature map superposition and Dropout operation. According to the method, the detection task of a multi-scale target can be efficiently processed in a complex scene, and the vehicle detection precision and robustness are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and particularly to a multi-scale object detection method based on a hierarchical feature aggregator (HFA), which is especially applicable to vehicle detection problems in complex road environments (such as low illumination, rain and fog, glare, etc.). Background Art

[0002] Object detection is a core task in the field of computer vision and is widely applied in multiple fields such as autonomous driving, security monitoring, and medical image analysis. With the progress of deep learning technology, object detection methods have achieved significant improvements in accuracy and speed. However, in complex environments, existing object detection methods still face great challenges, especially for the detection of small objects. Traditional methods are difficult to effectively extract multi-scale features, resulting in poor detection effects.

[0003] Existing object detection methods mainly rely on frameworks such as convolutional neural networks (CNNs) and region-based convolutional neural networks (R-CNNs), typically such as the YOLO series, SSD, etc. They extract features through multiple layers of convolution and perform object classification and localization. Although these methods have achieved good results in standard scenarios, their performance in complex environments such as low illumination, rain and fog is still insufficient. Especially in the detection of small objects, existing technologies cannot fully fuse feature information of different scales, resulting in a decrease in detection accuracy.

[0004] Currently, although methods such as the Feature Pyramid Network (FPN) and the Path Aggregation Network (PAN) have to some extent solved the multi-scale object detection problem, they still have certain limitations in complex backgrounds. Especially for the detection of small objects, traditional feature fusion methods fail to effectively extract fine-grained information, resulting in the loss of detailed information during the fusion process, thus affecting the detection accuracy. In addition, existing technologies lack an effective weight adjustment mechanism during feature fusion, resulting in insufficient fusion effects and further affecting the final detection results. Summary of the Invention

[0005] The purpose of the present invention is to provide a multi-scale object detection method based on a hierarchical feature aggregator (HFA) to solve the problem of insufficient multi-scale feature fusion in the prior art. Especially in complex environments, it can effectively improve the detection accuracy of small and large objects.

[0006] To achieve the above purpose, the present invention provides a multi-scale object detection method based on a hierarchical feature aggregator (HFA), including the following steps:

[0007] S1, obtain input image data, and preprocess the image to obtain an RGB three-channel image of 640×640;

[0008] S2. Use the YOLOv8 backbone network to extract multi-scale feature maps and obtain feature information at different scales;

[0009] S3. Perform cross-scale weighted fusion through a hierarchical feature aggregator (HFA) module to enhance the detection effects of small and large targets;

[0010] S4. Optimize the fused multi-scale features using the dynamic weighted feature fusion (DWFF) algorithm to improve the quality of the feature maps;

[0011] S5. Input the optimized feature maps into the YOLOv8 head network for object detection and output the position, category, and confidence of the objects.

[0012] A further technical solution is that the preprocessing step in S1 includes:

[0013] Adjust the size of the input image to 640×640 to meet the input requirements of the network;

[0014] Normalize the image and standardize the pixel values to the range of 0-1;

[0015] Perform data augmentation on the input image, including rotation, cropping, etc., to enhance the diversity of the data.

[0016] A further technical solution is that the YOLOv8 backbone network in S2 is used to extract multi-scale feature maps of the image, capture information at different scales, and provide a basis for subsequent feature fusion.

[0017] A further technical solution is that in S3, the hierarchical feature aggregator (HFA) module performs weighted fusion on features at different scales through cross-scale connections. The module includes:

[0018] Multiple convolutional layers for extracting features at different scales;

[0019] 1x1 convolution operation for adjusting the number of channels of the feature maps;

[0020] Upsampling operation for unifying the scales of the feature maps;

[0021] Concatenation operation to concatenate feature maps at different scales to form a multi-scale feature map.

[0022] A further technical solution is that in S4, the dynamic weighted feature fusion (DWFF) algorithm includes the following steps:

[0023] Initialize all elements of the weight vector to 1. The weight vector needs to perform gradient calculation and update during the training process, and use the ReLU activation function to ensure that the weight values are always non-negative;

[0024] Multiply the feature input value by the output of the Sigmoid function using the Swish activation function. The Swish activation function adds a very small constant to enable the model to learn more complex non-linear relationships;

[0025] Use the L2 regularization term to constrain the magnitude of the weights by calculating the sum of the squares of the weights, preventing overly large weights from causing overfitting;

[0026] Calculate the weighted values of each feature map and stack them. Stack all the weighted feature maps together and sum along the first dimension (i.e., the dimension of the number of feature maps) to represent the finally fused feature map;

[0027] Perform a Dropout operation on the fused feature map, randomly discarding a certain proportion of neurons to prevent overfitting.

[0028] A further technical solution lies in that in S5, the YOLOv8 head network includes a target localization, class prediction, and confidence calculation module. It uses the fused multi-scale features for target detection and removes redundant detection results through non-maximum suppression (NMS). BRIEF DESCRIPTION OF THE DRAWINGS

[0029] To more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following-described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0030] Figure 1 It is a step schematic diagram of a multi-scale target detection method based on a hierarchical feature aggregator (HFA) provided by the present invention.

[0031] Figure 2 It is a network structure diagram of YOLOv8 provided by the present invention.

[0032] Figure 3 It is a structure diagram of a hierarchical feature aggregator (HFA) provided by the present invention.

[0033] Figure 4 It is a flowchart of a feature weighted fusion algorithm provided by the present invention.

[0034] Figure 5It is a vehicle dataset sample provided by the present invention. (a) is a vehicle picture in a rainy day scene, (b) is a vehicle picture in a foggy day scene, (c) is a vehicle picture in a low-light scene, (d) is a vehicle picture in a strong-light scene, (e) is a vehicle image in a light interference scene, (f) is a vehicle image in a road surface reflection scene, (g) is a vehicle image in a daytime normal scene, (h) is a vehicle picture with Gaussian noise added to the picture in (g), and (i) is a vehicle picture with Gaussian noise added to the picture in (g).

[0035] Figure 6 It is the test visualization effect diagram provided by the present invention. Specific implementation manners

[0036] The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present invention, and should not be construed as a limitation to the present invention.

[0037] In the description of the present invention, "a plurality of" means two or more, unless otherwise specifically defined.

[0038] Please refer to Figure 1 、 Figure 2 、 Figure 3 , the present invention provides a multi-scale object detection method based on a hierarchical feature aggregator (HFA), including the following steps:

[0039] S101. Preprocess the input image to obtain an RGB three-channel image of 640×640, and input it into the network.

[0040] Specifically, first adjust the size of the input image to 640×640 by bilinear interpolation to meet the input requirements of the YOLOv8 network. Normalize the image to standardize the pixel values to the range of 0-1. Perform data augmentation on the input image, including rotation, cropping, etc., to enhance the diversity of the data.

[0041] S102. Use the YOLOv8 backbone network to extract multi-scale feature maps and obtain information at different scales.

[0042] Specifically, extract the multi-scale feature maps of the input image through the backbone part of the YOLOv8 network. This network extracts feature information at different levels through convolutional layers of different scales, respectively representing details at different scales. Obtain information at different scales from the extracted feature maps, and these features will be used for subsequent multi-scale feature fusion operations.

[0043] S103. Process the feature information of different scales through a hierarchical feature aggregator (HFA) module to achieve scale unification.

[0044] Specifically, use multiple convolutional layers to extract features of different scales, and adjust the number of channels of the feature map through 1x1 convolution so that features of different scales can be uniformly processed during fusion. Adopt upsampling to unify the sizes of feature maps of different scales.

[0045] S104. Through the dynamic weighted feature fusion (DWFF) algorithm, perform cross-scale weighted fusion on the feature maps of unified scales.

[0046] Specifically, in order to further improve the effect of feature fusion, the present invention adopts the dynamic weighted feature fusion (DWFF) algorithm. The DWFF algorithm weights the fused multi-scale features and balances the features to ensure that features of each scale can play a role in the final detection result. The algorithm flow of the DWFF module is as Figure 4 shown and includes the following steps:

[0047] Step (1), weight initialization and non-negativization:

[0048] First, initialize a weight vector w of length length:

[0049] w = nn.Parameter(torch.ones(length), requires_grad = True)

[0050] Among them, nn.Parameter: indicates that this weight vector is a trainable parameter, which will be automatically included in the gradient update of the model; torch.ones(length) means to create a vector of length length and initialize all its elements to 1; requires_grad = True means that this parameter needs to calculate gradients during the training process.

[0051] Perform non-negativization processing on the weight vector w to ensure that the weight values are always non-negative. This is achieved by using the ReLU activation function:

[0052] w relu = ReLU(w)

[0053] Among them, the definition of the ReLU activation function is ReLU(x) = max(0, x). It will change all negative values to 0 and keep positive values unchanged. The purpose of the activation function is to ensure that the weights are not negative, thus avoiding the gradient vanishing problem and improving the stability of the training process.

[0054] Step (2), weight normalization:

[0055] To avoid the weights in the network being too large or too small, which may affect the stability of the model, DWFF adopts the following normalization operation:

[0056]

[0057] where w relu represents the weight vector after ReLU activation; Swish(x) represents the Swish activation function, defined as Swish(x) = x·σ(x). The Swish activation function allows the model to learn more complex non-linear relationships by multiplying the input value with the output of the Sigmoid function. ∈ is a small constant (e.g., 10 -5 ), which is used to prevent the denominator from being zero and ensure the stability of the normalization operation.

[0058] where σ(x) is the Sigmoid function:

[0059]

[0060] The purpose of normalization is to stably scale the weights so that each weight remains relatively balanced during training, avoiding the training instability caused by some weights being too large or too small.

[0061] Step (3), weight L2 regularization:

[0062] To prevent overfitting of the model, DWFF adds an L2 regularization term to the loss function. The L2 regularization term constrains the magnitude of the weights by calculating the sum of the squares of the weights, preventing the overfitting caused by too large weights:

[0063]

[0064] where λ is the L2 regularization coefficient, which controls the strength of regularization. During training, a suitable λ usually needs to be selected through cross-validation; w i represents the i-th weight; is the sum of the squares of all weights, which is used to constrain the magnitude of the weights.

[0065] The role of L2 regularization is to avoid too large weights, prevent model overfitting, and improve the generalization ability of the model by constraining the magnitude of the weights.

[0066] Step (4), feature map weighting and superposition:

[0067] DWFF calculates the weighted values of each feature map and superimposes them. Each input feature map x i is weighted by the corresponding weight w i :

[0068]

[0069] Among them, x i represents the i-th input feature map; w i represents the i-th weight, which is the weighting coefficient corresponding to each feature map; represents the i-th weighted feature map.

[0070] Stack all the weighted feature maps together to form a new tensor, and sum along the first dimension (i.e., the dimension of the number of feature maps):

[0071]

[0072] Among them, F sum is the sum of the weighted feature maps, representing the finally fused feature map.

[0073] By weighting and stacking different feature maps, the fusion of multi-scale information is achieved, enabling the model to better utilize features at different levels.

[0074] Step (5), Dropout operation:

[0075] To enhance the generalization ability of the model, DWFF applies the Dropout operation to the output feature map. Dropout is a regularization technique that can randomly discard a certain proportion of neurons to prevent overfitting:

[0076] F dropout = Dropout(F sum )

[0077] Among them, Dropout(·) represents performing the Dropout operation on the input tensor during training, that is, randomly setting some elements to zero, reducing the model's dependence on certain neurons, thereby enhancing its robustness.

[0078] The role of Dropout is to avoid overfitting of the model during training and enhance the model's adaptability to unknown data.

[0079] Combining the above steps, the forward propagation process of DWFF can be expressed as:

[0080]

[0081] Among them, w[i] represents the i-th element in the weight vector; ReLU(w[i]) represents performing ReLU activation on the weight w[i]; Swish(w[j]) represents performing Swish activation on each element w[j] in the weight vector; x i represents the i-th input feature map; ∈ represents a very small constant used to prevent division-by-zero errors; F outRepresents the final output feature map, which has undergone weighted, normalized, weighted summation, and Dropout operations.

[0082] DWFF is a feature fusion module that combines weight non-negativization, Swish activation function, L2 regularization, weighted feature map superposition, and Dropout. Through these operations, DWFF can effectively extract and fuse features at different scales while ensuring computational efficiency, thereby improving the model's expressive ability and generalization performance.

[0083] S105: Input the fused multi-scale feature map into the YOLOv8 head network to finally output the detection results.

[0084] Specifically, target localization is performed through the YOLOv8 head network to determine the position of the target in the image, class prediction is performed on the detected target, and the confidence value is calculated. The non-maximum suppression algorithm is used to remove redundant detection results to ensure the accuracy of the output results.

[0085] The technical effect of this step is reflected in that the multi-scale feature map optimized by the HFA module and DWFF can significantly improve the network's detection accuracy for multi-scale targets in complex scenes, especially in terms of the robustness of small target detection and complex backgrounds, with significant advantages.

[0086] Through the above steps, the present invention combines the YOLOv8 framework and the hierarchical feature aggregator (HFA) module to propose a novel multi-scale object detection method. Through the hierarchical feature aggregation strategy of the HFA module and the dynamic weighted feature fusion (DWFF) optimization method, the object detection accuracy in complex environments is significantly improved, especially in the detection of small and large targets. This method can not only maintain high accuracy in a variety of complex scenes but also has a low computational burden, making it an effective and efficient object detection solution.

[0087] To perform vehicle detection in complex road environments, this paper collected video data under various different scenarios, covering environmental conditions such as normal weather, rainy and foggy weather, low light, and strong light. This paper used tools such as FFmpeg to extract one frame per second from the videos, ensuring wide coverage of the data in the time dimension, and extracted 6,000 images containing vehicles from them. To ensure the diversity of the dataset, this paper also selected 4,000 images containing vehicles from public datasets such as Cityscapes, Foggy Cityscapes, and UA-DETRAC. The dataset contains pictures of various resolutions. This paper used the LabelImg tool to perform detailed annotation on the vehicles in each image, marking the categories and bounding boxes of the vehicles. To enhance the robustness of the dataset and simulate the possible effects caused by dust accumulation in the camera, Gaussian noise and salt-and-pepper noise were randomly added to the dataset. Gaussian noise and salt-and-pepper noise simulated the degradation of image quality due to dust or blur, thereby further enriching the diversity and authenticity of the data. Samples of the vehicle dataset in complex scenarios are as Figure 5 shown.

[0088] To evaluate the performance of the trained model, this paper used evaluation metrics such as Precision, Recall, and mAP (mean average precision):

[0089]

[0090] Among them, TP is the number of true positives, FP is the number of false positives, FN is the number of false negatives, n represents the total number of categories, and AP i represents the average precision of the i-th category.

[0091] Precision represents the proportion of samples predicted as positive by the model that are actually positive. Recall represents the proportion of samples that are actually positive and are correctly predicted as positive. mAP is the mean of the average precisions of all categories. These metrics can comprehensively reflect the performance of the model. Especially in multi-class classification tasks, mAP is an important evaluation metric. In addition, this paper also used two metrics, mAP50 and mAP50-95, to more comprehensively evaluate the detection performance of the model:

[0092]

[0093] Among them, represents the average precision of the i-th category when the IoU threshold is 0.5, It represents the average precision when the IoU threshold of the $i$-th class varies from 0.5 to 0.95. mAP50 is the most commonly used metric, representing the average precision when the IoU threshold is 0.5; while mAP50-95 comprehensively considers the average precision under different IoU thresholds, more comprehensively reflecting the detection performance of the model. To understand the complexity and computational volume of the model in detail, this paper also counts the number of model parameters and the inference speed.

[0094] The model training server uses 2 NVIDIA L40 GPUs with 96G of video memory. The operating system is Ubuntu20.04 for all, the CUDA version is 12.0, the cuDNN version is 8.9, the programming language is Python 3.8, and the deep learning framework is Pytorch 2.2. The training adopts SGD optimization, no pre-trained model is loaded, and the learning rate is initialized to 0.01. The number of training epochs is set to 600, and the batch size is 16. If the model does not improve within 50 epochs, the training stops. This paper evaluates the model performance through metrics such as precision (Pre), recall (Rec), mAP@0.5 (mAP50), mAP@0.5:0.95 (mAP95), number of parameters (Par, unit is M, i.e., 10^6), computational volume (GFLOPs, unit is G), model occupied size (Size, unit is MB), and frame rate (FPS).

[0095] In the YOLO model, letters such as x, l, m, s, n, etc. represent variants of the model. These variant models are mainly achieved by reducing the stacking times of blocks and the number of channels of the convolution in the network bottleneck layer through layer accumulation. Reducing the depth reduces the repetition times of the key layers, and reducing the width reduces the number of feature channels of each layer of convolution. The input of the model in all experiments in this paper is a 640×640 RGB three-channel image. To verify the effectiveness of the HFA structure, this paper applies the HFA structure to the neck network of YOLOv8m. This paper verifies the effectiveness of HFA by setting the initial output channel number of the YOLOv8m network. As shown in Table 1, the baseline method is YOLOv8m. At different channel numbers, the network integrating HFA is superior to the baseline method in both model accuracy and model size. To further verify the HFA structure, YOLOv8m+HFA is compared with the RT-DETR model and other versions of the YOLO model. In this paper, the RT-DETR model used is the rtdetr-l model under the YOLO version. To verify the effectiveness of the HFA structure based on the YOLOv8 model. As shown in Table 2, the YOLOv8m network integrating HFA achieves excellent performance in mAP95 (0.766), Par (17.2), GFLOPs (64.4), and Size (34.8), higher than the performance achieved by other models.

[0096] Ablation experiments of HFA with different numbers of channels in Table 1

[0097]

[0098] HFA effectiveness experiment in Table 2

[0099]

[0100] Beneficial effects

[0101] The present invention proposes a multi-scale object detection method based on a hierarchical feature aggregator (HFA), which makes full use of the multi-scale feature fusion technology, significantly improves the accuracy and robustness of object detection, and the test results are as Figure 6 shown. By introducing the HFA module and the dynamic weighted feature fusion (DWFF) algorithm, the detection capabilities of small and large objects can be effectively improved in complex scenarios (such as low illumination, rain, fog, glare, etc.).

[0102] Compared with the existing object detection methods, the present invention improves the detection accuracy without significantly increasing the computational complexity. The HFA module effectively enhances the expression of objects at different scales through operations such as cross-scale connection and weighted fusion, especially for the detection of small objects and objects in complex backgrounds. The test results show that the present invention has achieved good detection results on multiple datasets, especially in complex scenarios, with a significant improvement compared to traditional methods. The innovation of the present invention lies in the efficient fusion of information at different scales through the hierarchical feature aggregator (HFA) module; the use of the dynamic weighted feature fusion (DWFF) method further optimizes the feature fusion effect, ensuring the stability and robustness of the network. This method not only effectively improves the accuracy of object detection but also maintains a low cost in terms of computational efficiency, making it suitable for real-time object detection tasks.

Claims

1. A multi-scale object detection method based on a hierarchical feature aggregator (HFA), characterized in that It includes the following steps: S1. Obtain the input image data, and preprocess the image to obtain an RGB three-channel image of 640×640; S2. Use the YOLOv8 backbone network to extract multi-scale feature maps and obtain feature information of different scales; S3. Perform cross-scale weighted fusion through the hierarchical feature aggregator (HFA) module to enhance the detection effects of small and large targets; S4. Adopt the dynamic weighted feature fusion (DWFF) algorithm to optimize the fused multi-scale features and improve the quality of the feature maps; S5. Input the optimized feature maps into the YOLOv8 head network for object detection, and output the position, category, and confidence of the object.

2. The multi-scale object detection method based on a hierarchical feature aggregator (HFA) according to claim 1, wherein, In S1, the preprocessing steps include: Adjust the size of the input image to 640×640 to meet the input requirements of the network; Normalize the image to standardize the pixel values to the range of 0-1; Perform data augmentation on the input image, including rotation, cropping, etc., to enhance the diversity of the data.

3. A multi-scale object detection method based on a hierarchical feature aggregator (HFA) according to claim 1, wherein In S2, the YOLOv8 backbone network is used to extract multi-scale feature maps of the image, capture information of different scales, and provide a basis for subsequent feature fusion.

4. A multi-scale object detection method based on a hierarchical feature aggregator (HFA) according to claim 1, characterized in that, In S3, the hierarchical feature aggregator (HFA) module performs weighted fusion on features of different scales through cross-scale connections. The module includes: Multiple convolutional layers for extracting features of different scales; 1x1 convolution operation for adjusting the number of channels of the feature maps; Upsampling operation for unifying the scales of the feature maps; Concatenation operation to concatenate feature maps of different scales together to form a multi-scale feature map.

5. A multi-scale object detection method based on a hierarchical feature aggregator (HFA) according to claim 4, characterized in that, In S4, the dynamic weighted feature fusion (DWFF) algorithm includes the following steps: Initialize all elements of the weight vector to 1. The weight vector needs to be calculated for gradients and updated during training. Use the ReLU activation function to ensure that the weight values are always non-negative; Multiply the feature input values by the output of the Sigmoid function using the Swish activation function. The Swish activation function adds a very small constant to enable the model to learn more complex non-linear relationships; Use the L2 regularization term to constrain the magnitude of the weights by calculating the sum of the squares of the weights to prevent overfitting caused by overly large weights; Calculate the weighted values of each feature map and stack them. Stack all the weighted feature maps together and sum along the first dimension (i.e., the dimension of the number of feature maps) to represent the final fused feature map; Perform Dropout operation on the fused feature map to randomly discard a certain proportion of neurons to prevent overfitting.

6. A multi-scale object detection method based on a hierarchical feature aggregator (HFA) according to claim 5, characterized in that, In S5, the YOLOv8 head network includes an object localization, category prediction, and confidence calculation module. It uses the fused multi-scale features for object detection and removes redundant detection results through non-maximum suppression (NMS).