Mixed feature and multi-scale fusion small target detection method and related equipment

By introducing HCTB, MDSKC and OKCSM modules into the object detection network, combining feature enhancement and multi-scale feature fusion technology, the problem of insufficient accuracy and robustness in small object detection in traditional methods is solved, and efficient small object detection under limited hardware conditions is achieved.

CN119942080APending Publication Date: 2025-05-06Chinese People's Liberation Army Cyberspace Force Information Engineering University
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510090507.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Under limited hardware conditions, traditional object detection methods are difficult to meet the accuracy and robustness requirements of small object detection when dealing with low resolution, high noise, target occlusion and background complexity.

Method used

A small object detection method for hybrid features and multi-scale fusion is proposed. By using the HybridConvTransBlock (HCTB) module in the backbone network for feature enhancement, and using the MultiDilatedSharedKernelConv (MDSKC) module for multi-scale feature extraction and fusion in the neck network, combined with the OKCSM module for feature fusion, optimize the FPN structure to improve the performance of small object detection.

Benefits of technology

This method significantly improves the accuracy rate, recall rate and mAP50 and mAP50:95 performance on the VisDrone2019 and TinyPerson datasets, and the model size and parameter quantity have also been optimized, which is suitable for devices with limited hardware resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942080A_ABST
    Figure CN119942080A_ABST
Patent Text Reader

Abstract

The invention provides a mixed feature and multi-scale fusion small target detection method and related equipment, and relates to the technical field of target detection. The method comprises the following steps: acquiring a to-be-detected image; inputting the to-be-detected remote sensing image into a preset target detection model, and outputting a detection result; the detection result comprises a target type and a corresponding confidence coefficient; the target detection model comprises a backbone network, a neck network and a detection head; wherein in the backbone network, an HCTB module is adopted to perform feature enhancement; the process of performing feature enhancement by the HCTB module comprises the following steps: splitting a channel of an input feature map into a first part and a second part, processing the first part through a Transformer and a CGSA module in sequence, processing the second part through a CNN and a Bottleneck module in sequence, and splicing the output of the CGSA module and the output of the Bottleneck module to obtain an enhanced feature map; wherein the CGSA module is used for carrying out feature extraction by combining a multi-head self-attention MHSA module and a convolution gated channel attention CGLU module.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of target detection technology, and in particular to a small target detection method and related equipment that integrates hybrid features and multi-scales. Background Art

[0002] Among many studies in the field of computer vision, target detection plays a key role. It aims to identify specific objects in images or videos and determine their locations. It is the basis of advanced visual tasks such as target tracking and behavior recognition. In recent years, with the rapid development of deep learning technology, the research on target detection has gradually deepened, especially the detection of small targets has attracted widespread attention. Small target detection has become an important and challenging research field. Small targets are widely present in various practical scenarios, such as pedestrians in the distance in surveillance video streams, tiny objects on the ground in drone aerial images, and traffic signs in autonomous driving systems. In practical application fields such as security monitoring, intelligent transportation systems, and drone military reconnaissance, the accurate detection of small targets is particularly critical, which is directly related to the effectiveness and safety of the system. Therefore, improving the accuracy and robustness of small target detection is one of the current researchers' concerns.

[0003] In recent years, although deep learning target detection technology has made significant progress, the detection task of small targets is still challenging because small targets occupy a small area in the image, have low contrast, insufficient information, and are easily disturbed by various factors such as background noise, image blur and lighting. To address this problem, some scholars have improved the detection accuracy of small targets by adding a small target detection head to the network structure and using large-scale feature maps to provide rich detail information, but this greatly increases the amount of calculation; or by integrating multiple different attention mechanisms to enhance the feature extraction ability of the network and design more complex neck FPN feature pyramid fusion structures, such as ASFF, NAS-FPN and BiFPN. However, a good balance has not been achieved in terms of detection accuracy and model scale, especially in devices with limited hardware resources, such as drones, car systems and devices with strict restrictions on onboard resources.

[0004] In general, the main challenges of small object detection in practical applications can be summarized into three points: insufficient feature representation, complex and cluttered background, and optimization of model size and accuracy under limited hardware conditions. Summary of the invention

[0005] In view of the problem that traditional detection methods are difficult to meet actual needs in terms of accuracy and robustness under limited hardware conditions due to factors such as low resolution, high noise, target occlusion and background complexity, the present invention proposes a small target detection method and related equipment that integrates hybrid features and multi-scale fusion.

[0006] In a first aspect, the present invention provides a small target detection method by hybrid feature and multi-scale fusion, comprising:

[0007] Acquire the image to be detected;

[0008] The remote sensing image to be tested is input into a preset target detection model, and the detection result is output; the detection result includes the target type and the corresponding confidence; the target detection model includes a backbone network, a neck network and a detection head; wherein, the HCTB module is used in the backbone network for feature enhancement; the process of feature enhancement by the HCTB module includes: splitting the channel of the input feature map into a first part and a second part, the first part is processed by the Transformer and CGSA modules in turn, the second part is processed by the CNN and Bottleneck modules in turn, and the output of the CGSA module and the output of the Bottleneck module are spliced ​​to obtain an enhanced feature map; wherein, the CGSA module is used to extract features by combining the multi-head self-attention MHSA module and the convolutional gated channel attention CGLU module.

[0009] Furthermore, the CGSA module is used to extract features by combining the multi-head self-attention MHSA module and the convolutional gated channel attention CGLU module, specifically including:

[0010] The input feature map is normalized, and the normalized features are passed to the MHSA module for processing, and the first intermediate feature map is output; the first intermediate feature map is regularized by DropPath, and the second intermediate feature map is output; the input feature map is added to the second intermediate feature map through residual connection, and the third intermediate feature map is output; the third intermediate feature map is normalized, and the normalized features are sent to the CGLU module for processing, and the fourth intermediate feature map is output; the fourth intermediate feature map is regularized by DropPath, and the fifth intermediate feature map is output; the third intermediate feature map is added to the fifth intermediate feature map through residual connection to obtain the output feature map.

[0011] Furthermore, it also includes: using the MDSKC module in the last layer of the backbone network to extract and fuse multi-scale features; the process of extracting and fusing multi-scale features by the MDSKC module includes:

[0012] The input feature map is processed by a 1×1 convolution layer to generate a sixth intermediate feature map; the sixth intermediate feature map is passed into a shared convolution kernel shared by multiple different expansion rates for processing, and multiple convolution results with different expansion rates are output, and the multiple convolution results with different expansion rates are spliced ​​with the sixth intermediate feature map, and the spliced ​​feature map is transformed by a 1×1 convolution layer to obtain an output feature map; wherein the processing process of the shared convolution kernel shared by multiple different expansion rates includes: setting the expansion rate of the shared convolution kernel, performing convolution processing on the input feature map at the current expansion rate and outputting the convolution result; updating the expansion rate of the shared convolution kernel, performing convolution processing on the previous convolution result at the current expansion rate and outputting the convolution result, and repeating the above process until the convolution processing is completed at all expansion rates.

[0013] Furthermore, in the neck network, the features of the P2 feature layer after being processed by SPDConv are fused with the P3 detection layer.

[0014] Furthermore, in the neck network, an OKCSM module is added before the P3 detection layer, and the OKCSM module is used to learn feature representation from global to local and perform multi-scale feature fusion; the process of the OKCSM module learning feature representation from global to local and performing multi-scale feature fusion includes:

[0015] The input feature map is processed by a 1×1 convolution layer to generate the seventh intermediate feature map; the channel of the seventh intermediate feature map is split into the third part and the fourth part, the third part is sent to the OmniKernel module for processing, and the eighth intermediate feature map is output; the eighth intermediate feature map is spliced ​​with the fourth part to obtain the output feature map.

[0016] Furthermore, the number of convolution kernels of the depthwise convolution in the large branch of the OmniKernel module is set to 31.

[0017] Furthermore, the target detection model uses the Yolov8 network as a basic framework.

[0018] In a second aspect, the present invention provides a small target detection device with hybrid features and multi-scale fusion, comprising:

[0019] An acquisition module, used for acquiring an image to be detected;

[0020] A detection module is used to input the remote sensing image to be tested into a preset target detection model and output a detection result; the detection result includes the target type and the corresponding confidence; the target detection model includes a backbone network, a neck network and a detection head; wherein, the HCTB module is used in the backbone network for feature enhancement; the process of feature enhancement by the HCTB module includes: splitting the channel of the input feature map into a first part and a second part, the first part is processed by the Transformer and CGSA modules in turn, the second part is processed by the CNN and Bottleneck modules in turn, and the output of the CGSA module and the output of the Bottleneck module are spliced ​​to obtain an enhanced feature map; wherein, the CGSA module is used to extract features by combining the multi-head self-attention MHSA module and the convolutional gated channel attention CGLU module.

[0021] In a third aspect, the present invention provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described in the first aspect when executing the program.

[0022] In a fourth aspect, the present invention provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method described in the first aspect.

[0023] The beneficial effects of the present invention are:

[0024] (1) The present invention designs a HybridConvTransBlock (HCTB) module that mixes CNN and Transformer modules in the backbone network to replace part of the C2f module or C3 module in the original backbone network, aiming to fuse local and global features and keep the computational cost controllable while ensuring efficient feature extraction, thereby optimizing the computational efficiency and feature extraction capability.

[0025] (2) Based on the idea of ​​dilated convolution, the present invention designs a multi-dilated shared convolution kernel module MultiDilatedSharedKernelConv (MDSKC) to replace the SPP module, SPPF module, their deformation modules or other multi-scale convolution operations in the original backbone network, so as to efficiently extract multi-scale features and retain important feature information while reducing the number of parameters.

[0026] (3) Based on the original PANFPN, the present invention fuses the P2 layer features rich in small target information with the P3 detection layer through the SPDConv operation, constructs a feature pyramid network for small targets, and constructs the OKCSM feature fusion module based on the OmniKernel and Cross Stage Partial ideas, which retains the information of small targets to a greater extent, effectively learns the feature representation from global to local, and improves the detection performance of small targets.

[0027] (4) The method of the present invention is applied to Yolov8n, and the improved model is compared with the mainstream algorithm on the VisDrone2019 and TinyPerson datasets. The experimental results show that the method of the present invention is superior to the original yolov8n in precision, recall, and mAP. 50 、mAP 50:95 The improvements are 1.3%, 3.1%, 3%, 1.9% and 3.6%, 1.3%, 2.1%, 0.7% respectively, and the model size and parameter amount are only 6.3MB and 3.14M, which are significantly better than other mainstream methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 A schematic diagram of a flow chart of a small target detection method using hybrid features and multi-scale fusion provided by an embodiment of the present invention;

[0029] Figure 2 A structural diagram of the HCTB module provided in an embodiment of the present invention;

[0030] Figure 3 A structural diagram of the MDSKC module provided in an embodiment of the present invention;

[0031] Figure 4 A structural diagram of the OKCSM module provided in an embodiment of the present invention;

[0032] Figure 5 A framework diagram of a target detection model formed by applying the method of the present invention in a Yolov8n network provided by an embodiment of the invention;

[0033] Figure 6 The heat map comparison results of the Yolov8n algorithm provided in the embodiment of the present invention and the method of the present invention in different scenarios;

[0034] Figure 7 The comparison algorithm provided in the embodiment of the present invention and the method of the present invention are compared with the experimental test results of the TinyPerson data set;

[0035] Figure 8The comparison algorithm provided in the embodiment of the present invention and the comparison experiment test results of the method of the present invention on the VisDrone2019 data set;

[0036] Fig. 9 A schematic diagram of the structure of a small target detection device with hybrid features and multi-scale fusion provided by an embodiment of the present invention;

[0037] Fig.10 A structural block diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0038] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution in the embodiment of the present invention will be clearly described below in conjunction with the drawings in the embodiment of the present invention. Obviously, the described embodiment is a part of the embodiment of the present invention, not all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0039] The present invention aims to design a high-precision small object detector in complex scenes, and improve the detection accuracy and robustness of the model while controlling the model size and parameters. The key to alleviating the problems of insufficient feature representation and background confusion lies in feature enhancement and fusion. Specifically, feature enhancement is performed in the backbone network, and the network's perception of small objects can be effectively enhanced by making full use of local and global context information. The present invention proposes a feature enhancement module HybridConvTransBlock (HCTB) and a multi-scale feature fusion module MultiDilatedSharedKernelConv (MDSKC). The HCTB module uses traditional convolutional CNN and transformer at the same time to extract global and local features and enrich local and global context features. MDSKC expands the receptive field of the backbone through dilated convolutions with different dilation rates, thereby obtaining rich feature information at different scales. Feature fusion is performed in the neck network, the FPN structure is optimized, and the OKCSM feature fusion module is constructed to fuse the small target branch extending from the P2 detection layer with the P3 detection layer to enhance the small target features, thereby achieving better performance.

[0040] In one embodiment, Figure 1 As shown, an embodiment of the present invention provides a small target detection method by hybrid feature and multi-scale fusion, comprising the following steps:

[0041] S101: Acquire an image to be detected;

[0042] S102: Input the remote sensing image to be tested into a preset target detection model and output the detection result; the detection result includes the target type and the corresponding confidence; the target detection model includes a backbone network, a neck network and a detection head; wherein, the HCTB module is used in the backbone network for feature enhancement; the process of feature enhancement by the HCTB module includes: splitting the channel of the input feature map into a first part and a second part, the first part is processed by the Transformer and CGSA modules in turn, the second part is processed by the CNN and Bottleneck modules in turn, and the output of the CGSA module and the output of the Bottleneck module are spliced ​​to obtain an enhanced feature map; wherein, the CGSA module is used to extract features by combining the multi-head self-attention MHSA module and the convolutional gated channel attention CGLU module.

[0043] Specifically, the main task of the backbone network is to extract rich features from the input image to be detected. Generally, it can output multiple feature maps of different scales, provide basic features for the neck network and detection head after the backbone network, and serve as the basis for subsequent target positioning and classification. The main task of the neck network is to fuse the feature maps of different scales output by the backbone network. Through feature fusion, feature information at different levels interacts to enhance the expression ability of features. The main task of the detection head is to output a bounding box based on the fused feature map, give the position of the predicted target through the bounding box, identify the category of the target and give the confidence of the bounding box. The Yolo series network generally includes a backbone network, a neck network and a detection head. Therefore, the target detection model in the small target detection method of hybrid feature and multi-scale fusion provided in the embodiment of the present invention can use the Yolo series network as the basic framework.

[0044] In computer vision tasks, many studies have shown that the receptive field size of CNN is small, which means that it can only extract local features. The Transformer structure has received widespread attention due to its powerful global feature extraction capabilities. However, due to the high computational complexity of the Transformer structure, directly applying it to all channels will result in significant computational overhead. In order to reduce the computational cost while ensuring efficient fusion of global and local features, an embodiment of the present invention designs a hybrid structure module called HybridConvTransBlock (HCTB), which combines the advantages of both convolutional neural network (CNN) and Transformer architectures. By including both CNN branches and Transformer branches in one module, it is possible to simultaneously utilize the advantages of CNN in local feature extraction and spatial structure maintenance, as well as the Transformer's capabilities in global information capture and long-distance dependency modeling. The design of this hybrid architecture enables the model to perform better in complex visual tasks.

[0045] The structure of HCTB is as follows Figure 2 As shown, the channel of the input feature is split into two parts, one part of which is sent to the Transformer branch, namely the ConvGLUwithSelfAttention (CGSA) module, and the other part is sent to the CNN branch, namely the Bottleneck module; then, the outputs of the two branches are concat spliced; finally, the spliced ​​features are subjected to 1×1 convolution for channel adjustment. Among them, in the CGSA module, the multi-head attention mechanism Mutil-Head-Self-Attention (MHSA) and the convolutional linear gated unit ConvolutionalGLU (CGLU) are combined. MHSA is responsible for extracting global features, and CGLU is used to enhance the expression ability of nonlinear features. Compared with the traditional FFN, CGLU has stronger performance and forms an efficient feature extraction and expression method. In some exemplary embodiments, the HCTB module can be used to replace some traditional C3 modules and C2f modules in the backbone network.

[0046] In one embodiment, the CGSA structure is as follows Figure 2 As shown, the feature extraction process of CGSA includes: normalizing the input feature map, passing the normalized features to the MHSA module for processing, and outputting the first intermediate feature map; the first intermediate feature map is regularized by DropPath, and the second intermediate feature map is output; the input feature map is added to the second intermediate feature map by residual connection, and the third intermediate feature map is output; the third intermediate feature map is normalized, and the normalized features are sent to the CGLU module for processing, and the fourth intermediate feature map is output; the fourth intermediate feature map is regularized by DropPath, and the fifth intermediate feature map is output; the third intermediate feature map is added to the fifth intermediate feature map by residual connection to obtain the output feature map.

[0047] Specifically, unlike the traditional self-attention mechanism, the MHSA module adopts a dynamic observation dimension allocation strategy when calculating queries, keys, and values, and achieves efficient integration of information in a multi-head setting by introducing the attention ratio (attn_ratio). This not only reduces the computational complexity, but also improves the flexibility and expressiveness of feature capture. The CGLU module strengthens the spatial information retention of features by using convolution operations instead of fully connected layers. This enables the model to better capture local features when processing high-dimensional inputs, thereby reducing the redundancy of feature dimensions while maintaining contextual information. Through the gating mechanism, the model can autonomously select important features and filter useless information, further improving the effectiveness of feature expression. At the same time, residual connections and variable DropPath strategies are used to enhance the fluidity of information, alleviate the gradient vanishing problem in deep network training, and ensure the robustness of the model at different training stages, thereby improving the overall performance.

[0048] In one embodiment, in order to further balance the computational cost and feature extraction capability of the target detection model, the MDSKC module is used in the last layer of the backbone network to perform multi-scale feature extraction and fusion; the process of the MDSKC module performing multi-scale feature extraction and fusion includes: processing the input feature map through a 1×1 convolution layer to generate a sixth intermediate feature map; passing the sixth intermediate feature map into a shared convolution kernel shared by multiple different expansion rates for processing, outputting multiple convolution results with different expansion rates, splicing the multiple convolution results with different expansion rates with the sixth intermediate feature map, and using a 1×1 convolution layer to transform the spliced ​​feature map to obtain an output feature map; wherein, the processing process of the shared convolution kernel shared by multiple different expansion rates includes: setting the expansion rate of the shared convolution kernel, performing convolution processing on the input feature map at the current expansion rate and outputting the convolution result; updating the expansion rate of the shared convolution kernel, performing convolution processing on the previous convolution result at the current expansion rate and outputting the convolution result, and repeating the above process until the convolution processing is completed at all expansion rates.

[0049] Specifically, in the traditional backbone network, the last layer uses the traditional SPP module, SPPF module, a deformed version of the two modules, or a multi-scale convolution operation; in these traditional modules or other multi-scale convolution operations, each scale usually has a separate convolution kernel for processing, resulting in parameter redundancy and increased computational cost. Unlike the traditional method, in the MultiDilatedSharedKernelConv (MDSKC) module of the embodiment of the present invention, all convolution operations of different scales share the same convolution kernel. This approach effectively reduces the number of parameters, and by sharing the convolution kernel, the model has a stronger feature reuse capability in multi-scale processing. Compared with using an independent convolution layer for each dilation rate, the shared convolution layer can reduce redundancy and improve model efficiency. At the same time, by introducing different dilation rates (dilations) on the shared convolution, the dilated convolution is used to increase the receptive field, thereby obtaining rich contextual information at different scales. Dilated convolution can usually expand the receptive field without increasing the amount of additional computation. Low dilation rates capture local details, and high dilation rates capture global context, which is particularly important for capturing multi-scale information, especially in visual tasks with objects of different scales.

[0050] In an exemplary embodiment, the overall structure of the MDSKC module is as follows: Figure 3 As shown in the figure, three different expansion rates are set for the shared convolution kernel; the process of multi-scale feature extraction and fusion of the MDSKC module includes: first, the number of channels is efficiently adjusted through the 1×1 convolution layer, the input features are processed, and a smaller feature map is generated. Subsequently, the feature information from different scales is effectively fused by splicing the convolution results of multiple different expansion rates. Finally, a 1×1 convolution is used to transform the spliced ​​feature map, and finally the fused features are output. This multi-scale feature fusion strategy helps the model better handle objects with different scales and shapes. In contrast, the pooling operation of SPPF may lose some detail information. The MDSK-Conv module has higher flexibility and expressiveness in feature extraction, and can better capture the details and complex patterns in the image. The above calculation process can be expressed as:

[0051]

[0052]

[0053] Where: represents a convolution operation with a convolution kernel size of 1×1. They represent dilation rates of 1, 3, and 5, respectively. Cat (Concat) represents the feature map concatenation operation. W1, W2, and W3 represent the feature maps of the three branches obtained after the input feature map undergoes conventional convolution and dilation convolution, respectively. F is the input feature map, and Y represents the final output feature map.

[0054] In one embodiment, in order to effectively extract small targets in the image to be detected, in the neck network, the features of the P2 feature layer after SPDConv processing are fused with the P3 detection layer.

[0055] Specifically, due to the downsampling operation in the target detection network, the features of small targets are lost, and small targets are slightly struggling on the normal P3, P4, and P5 detection layers. The P2 detection layer is usually located in the shallower layer of the network, and the feature map has a higher resolution, which means that small targets occupy more pixels on the feature map and have richer detail information. The traditional detection method is to add an additional P2 detection layer on the basis of the original P3, P4, and P5 detection layers to improve the detection ability of small targets. However, this approach will also bring a series of problems, such as excessive calculation after adding the P2 detection layer, and more time-consuming post-processing. To address this problem, this embodiment is based on the original PANFPN. By fusing the features of the P2 feature layer after SPDConv processing with the P3 detection layer, the detailed information of the small target on the feature map is retained without increasing the amount of calculation, forming an effective feature pyramid for small targets.

[0056] In one embodiment, in order to further improve the detection performance of small targets, an OKCSM module is added before the P3 detection layer, and the OKCSM module is used to learn feature representation from global to local and perform multi-scale feature fusion; the process of the OKCSM module learning feature representation from global to local and performing multi-scale feature fusion includes: processing the input feature map through a 1×1 convolution layer to generate a seventh intermediate feature map; splitting the channel of the seventh intermediate feature map into a third part and a fourth part, sending the third part to the OmniKernel module for processing, and outputting an eighth intermediate feature map; the eighth intermediate feature map is spliced ​​with the fourth part to obtain an output feature map.

[0057] Specifically, this embodiment is based on the OmniKernel and Cross Stage Partial ideas to obtain the OKCSM module for feature fusion, so as to effectively learn feature representation from global to local, and ultimately improve the detection performance of small targets.

[0058] In an exemplary embodiment, the OKCSM structure is as follows Figure 4As shown in the figure, the processing of this module includes: first, the input features are adjusted through 1×1 convolution, and the channels of the output features are split, one quarter of which are sent to the OmniKernel module, which inputs the features into three branches, namely the local branch, the large branch, and the global branch, to enhance the multi-scale representation. The results of the three branches are added and fused, and then modulated by another 1×1 convolution. Finally, the output of the OmniKernel module is spliced ​​and output with the other three quarters of the split channels.

[0059] In the above exemplary embodiment, the convolution kernel of the depthwise convolution in the large branch of the OmniKernel module is further adjusted from 63 to 31 to reduce the amount of calculation.

[0060] Yolov8 is an end-to-end object detection network, which is built on the basis of the historical version of the Yolo series, uses a new backbone network and anchor-free detection, and has the advantages of fast detection speed and good real-time performance. Therefore, in one embodiment, Yolov8n is used as the basic framework, and the method of the present invention is used to improve it. The overall framework is as follows Figure 5 shown.

[0061] Specifically, the embodiments of the present invention make improvements in the following three aspects: (1) Backbone network feature extraction. This embodiment adopts the constructed combined gated linear unit ConvGLU with Self Attention (CGSA) and multi-head self-attention module Mutil-Head-Self-Attention to replace part of the C2f module in the original backbone network, which can enrich local and global context features and realize the extraction of global and local features, thereby optimizing the computational efficiency and feature extraction capabilities, and can extract features more effectively while reducing the amount of computation. (2) Multi-scale feature extraction. Construct a multi-dilation rate shared convolution kernel module MultiDilatedSharedKernelConv (MDSKC) to replace the SPPF module, use a low dilation rate to capture local details, and a high dilation rate to capture global context, while reducing the number of parameters and retaining important feature information. (3) Small target feature enhancement. Due to downsampling in the network, the small target features are lost. In order to retain more small target information, the P2 feature layer is fused with the P3 feature layer to retain the information of the small target to a greater extent. At the same time, an OKCSM feature fusion module is constructed to effectively learn feature representations from global to local and perform multi-scale feature fusion to improve the detection performance of small objects.

[0062] In order to verify the effectiveness of the solution of the present invention, the present invention also provides the following comparative experiments.

[0063] 1.1 Dataset and Experimental Settings

[0064] This experiment uses the VisDrone2019 dataset and the TinyPerson dataset for training and performance evaluation. The VisDrone2019 dataset is taken by drones of different models in different cities, including different scenes, altitudes, weather and lighting conditions, covering 10 categories such as pedestrians and cars, with a total of 10,209 static images and about 2.6 million target instances. Compared with other datasets, VisDrone2019 has more obvious differences in the number and type scale of small targets, which has important research significance. The TinyPerson dataset is a dataset dedicated to small-scale target detection, containing 1,610 images and a total of 72,651 annotation boxes, including 1,498 normal label images, 78 dense label images, and 34 background images. The dataset has a high proportion of small targets, and the absolute scale of small targets is less than 36 pixels. On average, the absolute scale of small targets is about 18 pixels, and there are more than 200 targets in each image in dense scenes.

[0065] The experiment was conducted on Ubuntu 18.04 operating system, using NVIDIA GeForce RTX 4090 GPU, and configured with Python 3.8.10, PyTorch 2.0.0 and CUDA 11.8. Yolov8n was used as the basic framework, and the method of the present invention was used to improve it, and the improved target detection model was obtained as follows: Figure 5 As shown in the figure, the model training parameters are set as follows: batch size 8, number of training rounds 270, learning rate 0.01, weight decay coefficient 0.0005, and input image size 1024×1024. The pre-trained weights of yolov8n are used, and the other parameters remain default. The batch size of the TinyPerson dataset is 4, and the other parameters of the two datasets remain the same.

[0066] Performance evaluation indicators include precision P (Precision), recall R (Recall), and mean average precision mAP (Mean Average Precision). In target detection, TP (True Positive) is the number of correctly detected targets, FP (False Positive) is the number of incorrectly detected targets, and FN (False Negative) is the number of missed targets. The calculation formula for precision P is:

[0067]

[0068] The calculation formula of recall rate R is

[0069]

[0070] The area enclosed by the precision-recall (PR) curve represents AP, which is calculated as follows:

[0071]

[0072] mAP is the average of all APs of different categories. mAP@0.5 refers to the average detection accuracy of all target categories when the IoU threshold is 0.5. mAP@0.5:0.95 means the average of the detection accuracy calculated from all 10 IoU thresholds (from 0.5 to 0.95) with a step size of 0.05. The formula is as follows:

[0073]

[0074] Among them, N is the number of categories in the dataset, and a higher mAP value indicates that the model performs better in the object detection task. In addition, the model size and number of parameters are also used as evaluation indicators of the model.

[0075] (II) Ablation Experiment Results

[0076] In order to verify the effectiveness of the improved module of the method of the present invention, ablation experiments were carried out on each module on the VisDrone2019 dataset and the TinyPerson dataset. The evaluation results are shown in Table 1. It can be seen that after replacing the C2f module with the HCTB module, although the precision rate is slightly reduced, the recall rate, mAP50, and mAP50:95 accuracy are significantly improved, and the model size and parameter amount are also greatly reduced. Then, after adding the OKCSM module, the accuracy of P, R, mAP50, and mAP50:95 is further improved, while the model size and parameter amount are increased. The main reason is that compared with the original yolov8n model, the present invention extends an additional small branch from the P2 detection layer to enhance the small target features. Finally, after adding the MDSKC module, the accuracy of R, P, R, mAP50, and mAP50:95 reaches the highest, and the model size and parameter amount remain unchanged or decrease. Similarly, this experiment also verifies each module in the TinyPerson dataset. As shown in Table 2, after adding modules one by one, the accuracy of the experimental results shows an increasing trend, and finally reaches the optimal result.

[0077] Table 1 Evaluation results of detection accuracy of each algorithm module on the VisDrone2019 dataset

[0078]

[0079] Table 2 Evaluation results of detection accuracy of each algorithm module on the TinyPerson dataset

[0080]

[0081] In order to more intuitively illustrate the superiority of the method of the present invention in feature extraction and expression capabilities for small targets in complex scenes, Figure 6 As shown in the figure, this experiment visualizes the heat map of images in different scenes. It can be seen that compared with the original yolov8n algorithm, even in complex environments at night or in scenes with drastic changes in lighting, the method of the present invention can still more accurately identify the features of targets of different scales, especially the features of small targets. This shows that the method of the present invention has a stronger ability to express the features of small targets and can effectively extract the semantic feature information of small targets in complex scene images, thereby achieving better performance.

[0082] (III) Comparative experimental results

[0083] In order to further explore the detection performance of the method of the present invention, this experiment is compared with other classic target detection algorithms Yolov3-tiny, Yolv5s, Yolov6n, Yolov7-tiny, Yolov10n, and Yolo11n in terms of detection accuracy, model size, and parameter amount on the VisDrone2019 and TinyPerson datasets. The results are shown in Table 3. It can be clearly seen that the detection accuracy of the method of the present invention on the dataset is the best, and the model size and network parameters are relatively small. Compared with the latest Yolov11n, the model size and parameter amount are only increased by 0.8MB and 0.54M, and the various accuracy indicators are greatly improved, which has a strong competitive advantage. This shows that the method of the present invention can have better detection accuracy and detection speed while maintaining lightweight, meeting the needs of real-time detection.

[0084] Table 3 Performance comparison of the proposed algorithm and other algorithms on the VisDrone2019 dataset

[0085]

[0086] Note: Bold font indicates the best result, and underline “_” indicates the second best result.

[0087] Table 4 Performance comparison of the proposed algorithm and other algorithms on the TinyPerson dataset

[0088]

[0089] Note: Bold font indicates the best result, and underline “_” indicates the second best result.

[0090] like Figure 7 and Figure 8As shown in the figure, this experiment selects four representative scene images from the VisDrone2019 and TinyPerson datasets, and uses different algorithms to compare the detection results of the method of the present invention. It can be seen that each model has a good detection effect on relatively close medium and large targets, but for relatively dense small target areas, such as the red box marked area in the figure, due to the limited ability of the model to extract features for small targets, each model is difficult to effectively identify and locate, resulting in a large number of missed detections and false detections, even the latest yolov10n and yolo11n models are no exception. Compared with other algorithms, the improved algorithm of the present invention can accurately and effectively detect targets in dense areas and small targets at a distance, and the false detection rate and missed detection rate have been significantly improved, especially in the recognition of small targets at a distance, and the performance is better. At the same time, it can be seen from the heat map that the model can better focus on the target, and the model also generally improves the confidence of target detection. It shows that the method of the present invention has stronger robustness and practicality when using different scenarios and dealing with dense areas and small targets.

[0091] Based on the same inventive concept, Fig. 9 As shown, an embodiment of the present invention further provides a small target detection device integrating hybrid features and multi-scale fusion, including: an acquisition module and a detection module.

[0092] The acquisition module is used to acquire the image to be detected; the detection module is used to input the remote sensing image to be detected into a preset target detection model and output the detection result; the detection result includes the target type and the corresponding confidence; the target detection model includes a backbone network, a neck network and a detection head; wherein, the HCTB module is used in the backbone network for feature enhancement; the process of feature enhancement by the HCTB module includes: splitting the channel of the input feature map into a first part and a second part, the first part is processed by the Transformer and CGSA modules in turn, the second part is processed by the CNN and Bottleneck modules in turn, and the output of the CGSA module and the output of the Bottleneck module are spliced ​​to obtain an enhanced feature map; wherein, the CGSA module is used to extract features by combining the multi-head self-attention MHSA module and the convolutional gated channel attention CGLU module.

[0093] It should be noted that the small target detection device provided in the embodiment of the present invention is for implementing the above method. The specific functions thereof can be referred to the above method embodiments, which will not be described in detail here.

[0094] Fig.10 An example of a physical structure diagram of an electronic device is shown in FIG. Fig.10As shown, the electronic device may include: a processor (processor) 1001, a communication interface (Communications Interface) 1002, a memory (memory) 1003 and a communication bus 1004, wherein the processor 1001, the communication interface 1002, and the memory 1003 communicate with each other through the communication bus 1004. Processor 1001 can call logic instructions in memory 1003 to execute a small target detection method, which includes: obtaining an image to be detected; inputting the remote sensing image to be detected into a preset target detection model and outputting a detection result; the detection result includes a target type and a corresponding confidence level; the target detection model includes a backbone network, a neck network and a detection head; wherein, an HCTB module is used in the backbone network for feature enhancement; the process of feature enhancement by the HCTB module includes: splitting the channel of the input feature map into a first part and a second part, the first part is processed by the Transformer and CGSA modules in turn, the second part is processed by the CNN and Bottleneck modules in turn, and the output of the CGSA module and the output of the Bottleneck module are spliced ​​to obtain an enhanced feature map; wherein, the CGSA module is used to extract features by combining a multi-head self-attention MHSA module and a convolutional gated channel attention CGLU module.

[0095] In addition, when the logic instructions in the above-mentioned memory 1003 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0096] An embodiment of the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the small target detection method provided by the above-mentioned method embodiments.

[0097] An embodiment of the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the small target detection method provided by the above-mentioned method embodiments is implemented.

[0098] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0099] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A small target detection method based on hybrid features and multi-scale fusion, characterized in that: include: Acquire the image to be detected; Inputting the remote sensing image to be tested into a preset target detection model and outputting the detection result; The detection result includes the target type and the corresponding confidence; the target detection model includes a backbone network, a neck network and a detection head; wherein, the HCTB module is used in the backbone network for feature enhancement; the process of feature enhancement by the HCTB module includes: splitting the channel of the input feature map into a first part and a second part, the first part is processed by the Transformer and CGSA modules in turn, the second part is processed by the CNN and Bottleneck modules in turn, and the output of the CGSA module and the output of the Bottleneck module are spliced ​​to obtain an enhanced feature map; wherein, the CGSA module is used to extract features by combining the multi-head self-attention MHSA module and the convolutional gated channel attention CGLU module.

2. According to claim 1, a small target detection method combining hybrid features and multi-scale fusion is characterized in that: The CGSA module is used to extract features by combining the multi-head self-attention MHSA module and the convolutional gated channel attention CGLU module, specifically including: The input feature map is normalized, and the normalized features are passed to the MHSA module for processing, and the first intermediate feature map is output; the first intermediate feature map is regularized by DropPath, and the second intermediate feature map is output; the input feature map is added to the second intermediate feature map through residual connection, and the third intermediate feature map is output; the third intermediate feature map is normalized, and the normalized features are sent to the CGLU module for processing, and the fourth intermediate feature map is output; the fourth intermediate feature map is regularized by DropPath, and the fifth intermediate feature map is output; the third intermediate feature map is added to the fifth intermediate feature map through residual connection to obtain the output feature map.

3. The small target detection method of hybrid feature and multi-scale fusion according to claim 1 is characterized in that: Also includes: The MDSKC module is used in the last layer of the backbone network to extract and fuse multi-scale features; The process of multi-scale feature extraction and fusion performed by the MDSKC module includes: The input feature map is processed by a 1×1 convolution layer to generate a sixth intermediate feature map; the sixth intermediate feature map is passed into a shared convolution kernel shared by multiple different expansion rates for processing, and multiple convolution results with different expansion rates are output, and the multiple convolution results with different expansion rates are spliced ​​with the sixth intermediate feature map, and the spliced ​​feature map is transformed by a 1×1 convolution layer to obtain an output feature map; wherein the processing process of the shared convolution kernel shared by multiple different expansion rates includes: setting the expansion rate of the shared convolution kernel, performing convolution processing on the input feature map at the current expansion rate and outputting the convolution result; updating the expansion rate of the shared convolution kernel, performing convolution processing on the previous convolution result at the current expansion rate and outputting the convolution result, and repeating the above process until the convolution processing is completed at all expansion rates.

4. The small target detection method of hybrid feature and multi-scale fusion according to claim 1 is characterized in that: In the neck network, the features of the P2 feature layer after SPDConv processing are fused with the P3 detection layer.

5. The small target detection method of hybrid feature and multi-scale fusion according to claim 4 is characterized in that: In the neck network, an OKCSM module is added before the P3 detection layer, and the OKCSM module is used to learn feature representation from global to local and perform multi-scale feature fusion; The process of the OKCSM module learning feature representation from global to local and performing multi-scale feature fusion includes: The input feature map is processed by a 1×1 convolution layer to generate the seventh intermediate feature map; the channel of the seventh intermediate feature map is split into the third part and the fourth part, the third part is sent to the OmniKernel module for processing, and the eighth intermediate feature map is output; the eighth intermediate feature map is spliced ​​with the fourth part to obtain the output feature map.

6. The small target detection method of hybrid feature and multi-scale fusion according to claim 5 is characterized in that: The number of convolution kernels for the depthwise convolution in the large branch of the OmniKernel module is set to 31.

7. A small target detection method combining hybrid features and multi-scale fusion according to any one of claims 1 to 6, characterized in that: The target detection model uses the Yolov8 network as the basic framework.

8. A small target detection device with hybrid features and multi-scale fusion, characterized in that: include: An acquisition module, used for acquiring an image to be detected; A detection module, used to input the remote sensing image to be detected into a preset target detection model and output a detection result; The detection result includes the target type and the corresponding confidence; the target detection model includes a backbone network, a neck network and a detection head; wherein, the HCTB module is used in the backbone network for feature enhancement; the process of feature enhancement by the HCTB module includes: splitting the channel of the input feature map into a first part and a second part, the first part is processed by the Transformer and CGSA modules in turn, the second part is processed by the CNN and Bottleneck modules in turn, and the output of the CGSA module and the output of the Bottleneck module are spliced ​​to obtain an enhanced feature map; wherein, the CGSA module is used to extract features by combining the multi-head self-attention MHSA module and the convolutional gated channel attention CGLU module.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Method, device and equipment for evaluating water resistance of engine lubricating liquid and medium

    CN121049462A

  • Method and device for detecting target visual element in image and medium

    CN122023780A

  • Methods, equipment, and media for detecting target visual elements in images.

    CN122023780B