Small target detection method based on multi-scale feature fusion and feature enhancement
By adopting multi-scale feature fusion and feature enhancement methods in small object detection algorithm, the problems of insufficient feature information and background interference in small object detection are solved, and higher detection accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510191072.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-20
AI Technical Summary
Existing small object detection algorithms are difficult to achieve accurate detection effects, especially in the context of insufficient feature information and complex backgrounds, which makes small objects difficult to distinguish from backgrounds.
A small object detection method based on multi-scale feature fusion and feature enhancement is adopted to achieve accurate detection of small objects by building a detection model including input module, backbone network, feature fusion module, feature enhancement module and detector. The feature fusion module optimizes and fuses multi-scale features through the channel attention module and the spatial attention module. The feature enhancement module enhances the feature representation of small targets through global information extraction and local information extraction.
It improves the accuracy and robustness of small object detection, can more effectively distinguish small objects from complex backgrounds, reduces redundant information in the multi-scale feature fusion process, and gives full play to the advantages of high-resolution feature maps.
Smart Images

Figure CN120070865A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of object detection, and particularly to a small object detection method based on multi-scale feature fusion and feature enhancement. Background Art
[0002] Small object detection is an important research direction in the field of computer vision, aiming to classify and locate instances with limited areas in images with complex backgrounds. In the COCO dataset, objects with an area less than 32x32 pixels are usually classified as small objects. In recent years, object detection in remote sensing images has received extensive attention. Small objects are ubiquitous in remote sensing images, and detecting tiny objects in remote sensing images is of great significance for various application scenarios such as military reconnaissance, maritime rescue, and traffic management. In recent years, with the development of deep learning, many detection models based on deep learning have been proposed and have made important progress in object detection. Currently, small object detection algorithms are mainly divided into two-stage and single-stage methods:
[0003] The two-stage detection algorithm divides the detection process into two parts. The first stage generates candidate regions, and the second stage classifies and locates these regions. By training the two parts separately, such detectors can adapt to various object detection tasks, providing higher detection accuracy and precision, especially in large-scale object detection tasks. Common two-stage object detection algorithms include the R-CNN series, Faster R-CNN series, and Mask R-CNN series, etc. However, the two-stage detection algorithm usually includes two main stages, and each stage requires independent calculations, resulting in a relatively long overall inference time. Moreover, when the candidate regions are large, it will occupy a large amount of memory. Also, due to the long inference time, the two-stage detection algorithm is usually not suitable for real-time application scenarios such as autonomous driving and drone monitoring.
[0004] The single-stage detection algorithm has a relatively simple structure, is easy to implement and deploy. They directly predict the position and category of objects from the image without first generating candidate regions. With the end-to-end training method and high real-time detection performance, they are suitable for most object detection tasks, especially those with requirements for detection speed, and have great advantages. Common single-stage detection algorithms include the YOLO series, SSD, FCOS network, etc. However, compared with the two-stage algorithm, the single-stage algorithm is slightly inferior in terms of accuracy. Especially when dealing with small objects, it is easy to misdetect the noise in the background as objects, resulting in the single-stage detector being unable to accurately locate and classify small object objects.
[0005] However, whether it is a single-stage small object detection algorithm or a two-stage small object detection algorithm, due to the small pixel area of small objects in the image, their feature information is insufficient, and small objects are located in complex environmental backgrounds, such as urban streets or natural environments like forests. As a result, interference objects in the background may be misidentified as targets, or the real targets may be occluded by the background and difficult to recognize. These factors often lead to a decrease in the detection accuracy of the detection model when locating and classifying small objects, especially in remote sensing images. Therefore, enhancing the feature representation of small objects and the distinguishability between small objects and the background remains a challenging task.
[0006] Currently, the existing small object detection algorithms can be summarized as follows:
[0007] (1) Small object detection algorithms based on multi-scale representation learning: Multi-scale feature fusion has made remarkable progress in computer vision, especially in the field of object detection. However, compared with medium and large objects, small objects have fewer pixels, making feature extraction more difficult. Moreover, as the number of network layers increases, their feature and location information gradually gets lost, resulting in difficulty in being effectively detected. Although multi-scale feature fusion methods, such as FPN, can improve object localization and classification to a certain extent by combining shallow and deep features, these methods still have deficiencies. First, the features of small and large objects in existing multi-scale fusion methods may be confused at certain scales. Even at fine-grained scales, they may still be interfered by background noise or large objects. Second, effectively fusing features at different scales remains a challenge. If the fusion strategy is inappropriate, it may lead to information loss or redundancy. And existing multi-scale methods can capture objects of different sizes, but in some cases, smaller objects may still not be effectively distinguishable from the background, especially when the edges of the objects are blurred or the size of the objects themselves is too small and the information is insufficient, the algorithm may have difficulty accurately locating the objects. Although FPN and its improved versions (such as CE-FPN, PANet, NAS-FPN, BiFPN, and AugFPN) have improved multi-scale fusion to a certain extent, they still have not solved the problems of interference from noise or large objects and the difficulty of small objects to be distinguished from complex backgrounds in small object detection, which affects the detection accuracy and robustness.
[0008] (2) Small target detection algorithm based on feature enhancement: The application of feature enhancement in small target detection aims to strengthen the feature representation of small targets in images through specific algorithms or techniques, making these targets easier to be captured and recognized by detectors. The core idea of this method is to process the original input data or intermediate feature maps to improve the saliency of small targets relative to the background or other interference factors, and emphasize the use of context information to assist in classifying and locating small targets, thereby improving the detection accuracy. However, there are still certain limitations in some current feature enhancement methods. Some researchers have proposed a feature enhancement module specifically for small target detection in remote sensing images, which assigns weights to feature maps in the spatial and channel directions, enabling the model to not only effectively suppress the influence of noise but also highlight useful features, improving the target classification accuracy and localization precision. However, current feature enhancement methods often generate a large number of enhanced features, but not all of these features are necessarily helpful for small target detection. In some cases, redundant enhanced features may lead to information overload and even make it more difficult for the network to extract useful signals from them, thus affecting the detection accuracy. Secondly, some feature enhancement methods may strengthen background noise or irrelevant regions. Especially when the characteristics of complex backgrounds are not fully considered during the enhancement process, noise may be introduced to interfere with the detection of small targets, resulting in false detections or missed detections.
[0009] (3) Small target detection algorithm based on attention mechanism: The application of the attention mechanism in small target detection aims to simulate the characteristics of the human visual system, that is, it can automatically focus on the most relevant or important parts of the image. This mechanism allows the model to dynamically allocate weights to different regions when processing input data, thereby increasing the attention to key information, especially in the task of detecting small targets in complex backgrounds. The core of this method is to assign different weights to different parts of the input feature map, enabling the network to focus more on those regions that are most important for the current task. In recent years, to address the problem that feature maps of different scales have different importance, some researchers have proposed to dynamically adjust the importance of feature maps of different scales through the attention mechanism, thereby reducing redundant information in the feature layers.
[0010] In summary, most of the existing small object detection algorithms rely on cascaded methods. First, the important regions of features are roughly extracted, and then based on this, it is used to guide the more refined highlighting of the regions where small objects are located and the surrounding context information beneficial to small objects. However, this method requires multiple stages to gradually refine the detection results, and the networks in each stage usually need to be trained separately or jointly optimized. Due to the dependency relationship between different stages, it is difficult to achieve complete parallel processing. In addition, some current research methods start from the perspective of data imbalance of small objects and use label optimization strategies to make the proportion of small objects in the image equivalent to that of medium and large objects. However, due to the small size of small objects, even with label optimization strategies, it may not be possible to fully distinguish small objects from background noise. It can be seen from this that current small object detection faces the problems of insufficient feature information and being easily partially or completely occluded by other objects or the background, especially in dense scenes, making it difficult to separate small objects from the background. Although existing methods have made efforts to solve the above problems from various aspects, most methods ignore the advantages of high-resolution feature maps and the features of small objects are not fully represented, making it difficult to achieve accurate detection results. Summary of the Invention
[0011] The present invention provides a small object detection method based on multi-scale feature fusion and feature enhancement to solve the technical problem that it is difficult to achieve accurate detection results with existing small object detection methods.
[0012] To solve the above technical problems, the present invention provides the following technical solutions:
[0013] On the one hand, the present invention provides a small object detection method based on multi-scale feature fusion and feature enhancement, and the small object detection method based on multi-scale feature fusion and feature enhancement includes:
[0014] Construct a small object detection model based on multi-scale feature fusion and feature enhancement;
[0015] Train the constructed small object detection model;
[0016] Use the trained small object detection model to implement small object detection.
[0017] Further, the small object detection model includes: an input module, a backbone network, a feature fusion module, a feature enhancement module, a feature reconstruction module, and a detector;
[0018] Among them, the feature reconstruction module is only used in the model training stage and is discarded in the inference stage;
[0019] The process of the small object detection model implementing small object detection includes:
[0020] The image to be detected is input into the backbone network through the input module. After being processed by the backbone network, multi-scale feature maps are obtained. Then, the multi-scale feature maps are sent into the feature fusion module. In the feature fusion module, the multi-scale feature maps are first optimized in terms of features, and then the optimized multi-scale feature maps are fused from top to bottom to obtain the feature map after feature fusion. Subsequently, the feature map after feature fusion is sent into the feature enhancement module for processing. By capturing the local detail information of small targets and the global information captured by the interaction between pixels, the feature representation of small targets is enhanced to obtain the feature map after feature enhancement. Finally, the feature map after feature enhancement is sent into the detector to achieve the classification and localization of targets.
[0021] Further, the feature fusion module includes a channel attention module and a spatial attention module;
[0022] The feature fusion module first optimizes the multi-scale feature maps output by the backbone network through the channel attention module to highlight the important regions in the feature maps of different scales, and then fuses the optimized multi-scale feature maps through the spatial attention module to obtain the feature map after feature fusion.
[0023] Further, the optimization of the multi-scale feature maps output by the backbone network through the channel attention module to highlight the important regions in the feature maps of different scales includes:
[0024] The channel attention module first performs channel processing on the multi-scale feature maps output by the backbone network through adaptive max pooling. Then, the pooled feature vectors are fused and passed through the Sigmoid function to generate the final channel attention weights. Finally, the original feature maps are multiplied by the channel attention weights, thereby enhancing the features useful for small target detection and suppressing the irrelevant features. Finally, the features obtained by multiplying the original feature maps by the channel attention weights are subjected to point convolution so that the number of feature maps for each scale is 256, obtaining the optimized multi-scale feature maps to highlight the important regions in the feature maps of different scales.
[0025] Further, the fusion of the optimized multi-scale feature maps through the spatial attention module to obtain the feature map after feature fusion includes:
[0026] The low-level feature map and the high-level feature map in the feature map optimized by the channel attention module are respectively expanded into sequences. The sequence corresponding to the high-level feature map is used as the query, and the sequences corresponding to the low-level feature maps are used as the key and value. Self-Attention is performed between the high-level feature map and the low-level feature maps, and the result of Self-Attention is transformed into a two-dimensional image. Then, through upsampling, the dimension of the two-dimensional image is matched with that of the low-level feature map. Finally, the low-level feature map and the upsampled two-dimensional image are fused to obtain the feature map after feature fusion.
[0027] Further, the feature enhancement module includes a global information extraction module, a local information extraction module, and an inverted residual structure;
[0028] The global information extraction module captures global information by performing perception and interaction between pixels; the local information extraction module completes local information extraction through parallel calculation of three ordinary convolutions and three dilated convolutions; among them, the dilation rates of the three dilated convolutions are 2, 4, and 2 respectively; the features obtained by fusing the local information and the global information are fed into the inverted residual structure to obtain the feature map after feature enhancement.
[0029] Further, the global information extraction module first unfolds the two-dimensional image in units of pixels to make it into a spatially continuous sequence. Then, an average operation is performed on the channels, and then linear calculation is performed through FFN and the Softmax activation function is used to obtain the probability distribution of the features in their spatial dimensions, obtaining the weights between pixels on the feature map. Finally, it is reshaped into a two-dimensional image.
[0030] Further, the feature reconstruction module realizes feature reconstruction based on a masked autoencoder.
[0031] On the other hand, the present invention also provides an electronic device, which includes a processor and a memory; wherein, at least one instruction is stored in the memory, and the instruction is loaded and executed by the processor to implement the above method.
[0032] On another aspect, the present invention also provides a computer-readable storage medium, in which at least one instruction is stored, and the instruction is loaded and executed by the processor to implement the above method.
[0033] The beneficial effects brought by the technical solution provided by the present invention at least include:
[0034] 1. Compared with other small target detection methods based on multi-scale representation learning: The technical solution provided by the present invention fully considers the importance of feature maps of different scales and proposes a progressive and refined feature fusion strategy. This strategy first uses the channel attention module to screen the feature map, which can automatically learn important features for small target detection. Then, feature fusion is performed through a top-down feature fusion module, which helps the model's understanding of complex scenes and surrounding environments. It not only reduces the redundant information in the multi-scale feature fusion process, but also gives full play to the advantages of high-level feature maps and low-level feature maps, which helps solve the problem that small targets are difficult to distinguish from complex backgrounds.
[0035] 2. Compared with the small target detection method based on feature enhancement: small objects often need to rely on the surrounding environment information for more accurate classification and positioning. In this regard, the feature enhancement method proposed in the technical solution of the present invention focuses on both the local detail information and the global context information of the small target, which not only highlights the importance of the surrounding environment information, but also gives full play to the advantages of the local detail information, greatly improving the accuracy of small target detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0037] Figure 1 It is a schematic diagram of the execution flow of a small target detection method based on multi-scale feature fusion and feature enhancement provided by an embodiment of the present invention;
[0038] Figure 2 is an architecture diagram of a small target detection model provided by an embodiment of the present invention;
[0039] Figure 3 It is a SAFPN module architecture diagram provided by an embodiment of the present invention;
[0040] Figure 4 It is a system block diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0041] In order to make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.
[0042] First of all, it should be noted that in the embodiments of the present invention, words such as "exemplarily" and "for example" are used to give examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" in the present invention should not be construed as being more preferred or more advantageous than other embodiments or design solutions. Rather, the use of the word "exemplarily" is intended to present concepts in a specific manner. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one of the two.
[0043] The First Embodiment
[0044] This embodiment provides a small target detection method based on multi-scale feature fusion and feature enhancement, aiming to establish an effective and refined single-stage small target detection technology. This method enables the model to focus on the important regions of small targets through a feature fusion module SAFPN, solving the problem that small targets are difficult to distinguish from the background. High-resolution feature maps can often retain more detailed information of small targets, such as the boundaries, textures, and shapes of small targets, and this information is particularly important for small target detection. For example, the texture features of targets can effectively help the model distinguish adjacent complex backgrounds. Therefore, we introduce a new feature enhancement module LGFE to enhance the feature representation of small targets. This module obtains the context information and detailed information of small targets through the local information captured by convolution and the global information captured by the interaction between pixels and pixels. We also additionally introduce a feature reconstruction module FR, which can, as auxiliary information, force the model to learn richer and more detailed feature representations, alleviating the problem of information loss caused by downsampling of small targets in the deep neural network and the problem that small targets are difficult to distinguish from complex backgrounds.
[0045] This method can be implemented by an electronic device, and the execution process of this method is as Figure 1 shown, including the following steps:
[0046] S1, construct a small target detection model based on multi-scale feature fusion and feature enhancement;
[0047] S2, train the constructed small target detection model;
[0048] S3, use the trained small target detection model to achieve small target detection.
[0049] Specifically, as Figure 2 shown, the small target detection model of this embodiment is an improvement based on the FCOS network model, and it includes: an input module, a backbone network, a feature fusion SAFPN module, a feature enhancement LGFE module, a feature reconstruction FR module, and a detector; based on this, the working process of this small target detection model can be described as follows:
[0050] First, input data is processed by the backbone network (ResNet50) to obtain multi-scale feature maps. Then, the multi-scale feature maps are fed into the SAFPN module. In the SAFPN module, feature optimization is first performed, and then top-down feature fusion is carried out to sequentially obtain high-resolution to low-resolution feature maps, namely P2, P3, P4, P5, and P6. Subsequently, the high-resolution feature map P2 passes through the feature enhancement LGFE module, which enhances the feature representation of small targets by capturing the local detailed information of small targets and the global information captured by the interaction between pixels. Finally, the multi-scale feature maps are fed into the detector (FCOS network detector) for classification and localization.
[0051] Next, the core modules in this model (SAFPN module, LGFE module, FR module) will be introduced.
[0052] 1. Multi-scale Feature Fusion Module - SAFPN
[0053] This module is an effective and refined method for multi-scale feature fusion of small target features. By capturing the important features of small targets in multi-scale feature maps, it improves the distinguishability between small targets and complex backgrounds, prompting the model to focus more on the important regions related to small targets, thereby improving the detection performance of the model. First, it optimizes the multi-scale features passing through the backbone network through a channel attention module to highlight the important regions in feature maps of different scales, and then performs feature fusion through a spatial attention module. This module consists of two parts: feature optimization and feature fusion. As Figure 3 shown.
[0054] Feature Optimization: From the perspective of feature channels, first, the channel attention module first processes the input feature map through adaptive max pooling in the channel dimension. Then, the pooled feature vectors are fused, and the final channel attention weights are generated through the Sigmoid function. Finally, the original feature map is multiplied by the channel attention weights to enhance the features useful for small target detection and suppress irrelevant features. The obtained important features pass through PWConv so that the number of feature maps for each scale is 256, enabling feature map matching at different scales. Adaptive max pooling selects the most significant or important response values on each channel, emphasizing local and prominent feature points, which can better capture key but small targets or detailed information in the image, enhancing the model's feature selection ability in the channel dimension and helping to extract more accurate feature information from each channel.
[0055] Feature Fusion: In multi-scale feature maps, low-level feature maps come from shallow networks and are rich in spatial information. The feature resolution of spatial information is relatively high, which can better preserve the spatial structure information of the image and help capture local details in the image more effectively. However, due to only containing local pixel information, they lack the expression of high-level semantic content and have weak understanding ability of objects and context in complex scenes. High-level feature maps come from deep networks and have rich high-level semantic information. The feature resolution of semantic information is relatively low, such as scene context and semantic relationships, which helps to aggregate global information but loses the detailed information in the image. Especially when dealing with small targets, the positioning accuracy of small targets is not very accurate. The SAFF module in this paper gives full play to the advantages of low-level feature maps with accurate detailed information and high-level feature maps with rich semantic information, not only promoting the communication between features at different levels, enabling the model to dynamically adjust the importance of features according to the input content, but also helping the model better understand the relationship between the target and its surrounding environment by calculating the interaction between different positions of the feature map. As Figure 3 shown, the low-level feature map and the high-level feature map are respectively expanded into sequences. The sequence of the high-level feature map is Q, and the sequences of the low-level feature map are K and V. Perform Self-Attention between the low-level feature map and the high-level feature map, then turn it into a two-dimensional image, and make it match the dimension of the low-level feature map through upsampling, and finally fuse it with the two-dimensional image.
[0056] In summary, according to the characteristics of multi-scale features, this embodiment constructs a progressive and refined feature pyramid feature fusion module SAFPN. First, feature selection is performed on features of different scales in the channel dimension, and then based on the high-level feature map, it guides the feature fusion of the low-level feature map. As Figure 3 shown, CA in the figure represents the channel attention module, PWconv represents point convolution, Conv represents 3x3 convolution, and Q, K, and V respectively represent query, key, and value. The feature fusion SAFPN module enhances the discriminability of small target features, large target features, and complex backgrounds through the feature optimization module and the feature fusion module, making it easier for the model to distinguish small targets from complex backgrounds and making up for the problem of insufficient small target feature information.
[0057] 2. Feature Enhancement Module LGFE:
[0058] This embodiment designs a feature enhancement module LGFE, as Figure 2As shown, this module fuses local information and global information to improve the accuracy of small object detection. First, we introduce the Pixel Aware Attention (PAA) module, which is responsible for the perception and interaction between pixels to capture global information. The PAA module first unfolds the two-dimensional image in units of pixels to make it a spatially continuous sequence. Then, an average operation is performed on the channels. After that, linear calculations are carried out through a Feed-Forward Neural Network (FFN), and the Softmax activation function is used to obtain the probability distribution of the features in their spatial dimensions, obtaining the weights between pixels on the feature map. Finally, it is reshaped into a two-dimensional image. The capture of local information is completed by parallel calculations of three ordinary convolutions and three dilated convolutions, where the dilation rates of the dilated convolutions are 2, 4, and 2 respectively. The receptive field of an ordinary convolution is directly related to the size of the convolution kernel and can accurately capture local information. The dilated convolution expands the receptive field by inserting holes between the convolution kernel elements, and can effectively increase the receptive field without changing the size of the convolution kernel, which is very beneficial for maintaining the detailed information of the image. However, a large dilation rate may result in the captured long-distance information not being relevant. Fusing ordinary convolutions and dilated convolutions can mutually compensate for the deficiencies between the two. It has been proven in MobileNetv2 that the inverted residual structure not only helps the backpropagation of gradients, reduces the computational complexity, but also enables the model to better maintain the information integrity of the input features and ensures the effective transmission of information. Therefore, finally, we send the feature representation obtained by fusing local information and global information into the inverted residual structure to ensure the effective transmission of the enhanced small object feature representation. The fusion of local information and global information in the high-resolution feature map makes up for the detailed information lost in the global features and the lack of rich semantic information in the local features, reducing the information loss problem of small objects in the downsampling process, enabling the model to obtain a more abundant and accurate feature representation, and improving the model's ability to understand the surrounding environment of small objects.
[0059] 3. Feature Reconstruction Module FR:
[0060] The Masked Autoencoder (MAE) is a self-supervised learning method that reconstructs the entire image by masking a part of the input image and based on the unmasked part, which not only enhances the model's understanding of the image but also generates powerful feature representations. As Figure 2As shown in the figure, the feature reconstruction module proposed in this paper first generates a binary image object map with the same size as the original image according to the gt bboxes (the true detection boxes marked in the image), where 1 represents the area where small targets are located and 0 represents the background area. Then, through max pooling, a target map is obtained to match the size of the binary image object map with the high-resolution feature map P2. After that, we upsample the P4 feature map and the P6 feature map to get f4 and f6 respectively, so that they can match the size of the high-resolution feature map P2. MAE will mask out a large subset of random image patches, and it has been proven that a masking ratio of 75% has the best recognition effect in MAE. In the feature reconstruction module of this paper, a mask is generated according to the binary image object map. First, the binary image is unfolded into a set of spatially continuous patches with a patch_size of 8, and the sum of pixels in each patch is calculated. If the sum of pixels is greater than or equal to 1, it indicates that there are small targets in the patch, and masking is performed. Then, the ratio of the masked patches is counted. If it does not reach 75%, random masking is performed from the unmasked patches to ensure that the masking ratio of the patches is 75%. Then, according to the generated mask and the binary image target map, the original image is reconstructed for the high-resolution feature map P2, as well as f4 and f6 through the decoder of MAE. The mean square error loss is calculated by comparing the reconstructed original image with the binary image target map respectively. Finally, the losses of feature reconstruction are added up to obtain the final auxiliary loss. It should be noted that the feature reconstruction module proposed in this paper is only used in the training stage and discarded in the inference stage.
[0061] In summary, in this embodiment, a feature reconstruction module is constructed based on the advantages of the high-level feature map, and then the P2, P4, and P6 feature maps are respectively sent into the feature reconstruction FR module for feature reconstruction, forcing the model to learn richer and more detailed feature representations, as Figure 2 shown, where "DW Conv" represents depthwise convolution, which applies a convolutional kernel to each input channel separately, retaining the basic characteristics of traditional convolution while reducing the number of parameters. "DConv" represents dilated convolution. Among them, the high-resolution feature map contains more detailed information about small targets. Based on the high-resolution feature map, capturing both the local detailed information and the global context information of the high-resolution feature map can improve the model's understanding ability of complex backgrounds and surrounding environments.
[0062] Next, the effectiveness of the solution proposed by the present invention is verified.
[0063] Table 1 provides the comparison results of the performance of this method with other state-of-the-art methods on the AI-TOD-v2 dataset. As can be seen from Table 1, compared with other methods, the design of the feature fusion, feature enhancement, and feature reconstruction modules of this method enables the model to achieve an accuracy of 26.9%, exceeding other methods, and in the AP 50 , AP 75 , AP vt , AP t and AP s metrics, this method also reaches the highest, outperforming other methods. The experimental results verify the effectiveness of the method proposed in the present invention.
[0064] Table 1 Comparison results of the performance of this method with other state-of-the-art methods on the AI-TOD-v2 dataset
[0065]
[0066] The AI-TOD-v2 dataset contains 8 common small target objects, including: aircraft (AI), bridge (BR), storage tank (ST), ship (SH), swimming pool (SP), vehicle (SE), person (PE), and windmill (WM). As shown in Table 2, we conducted experimental evaluations on the AP of each category of the AI-TOD-v2 dataset respectively. It can be seen from the experimental data that this method performs better than previous other methods in each category of AI-TOD-v2. These results highlight the excellent performance of this method on this challenging small target dataset.
[0067] Table 2 AP evaluation results of each category of the AI-TOD-v2 dataset
[0068]
[0069]
[0070] As shown in Table 3, on the SODA-A dataset, compared with other methods based on rotated bounding boxes, this method reaches 33.7% in terms of APs, outperforming other methods. And in the comparison of each category, this method has better AP metrics than other methods in multiple categories. Generally speaking, the method provided by the present invention shows comparable results with existing methods on this challenging small target dataset, demonstrating its effectiveness on the SODA-A dataset.
[0071] Table 3 Comparison results of this method with other methods based on rotated bounding boxes on the SODA-A dataset
[0072]
[0073]
[0074] As can be seen from the above, extensive experiments conducted on the AI-TOD-v2 and SODA-A datasets show that the method proposed in the present invention is superior to other methods in terms of performance. The method proposed in the present invention provides a feasible solution to the problems of insufficient feature information of small targets and the difficulty of distinguishing small targets from complex backgrounds, and provides a new direction and idea for further research on small target detection algorithms.
[0075] In summary, the present embodiment provides a small target detection method based on multi-scale feature fusion and feature enhancement, and designs a progressive and refined feature fusion method. This method first performs feature selection on multi-scale features, and then, based on the advantages of high-level feature maps and low-level feature maps, conducts refined feature fusion. An LGFE (Local and Global Feature Enhancement) module for feature enhancement is introduced. Based on high-resolution feature maps, it captures both the local detail information and global context information of small targets, enabling the model to enhance the feature representation of small targets and reduce the influence of redundant noise. An additional FR (Feature Reconstruction) module is introduced. Through feature reconstruction, the model is forced to learn richer and more detailed feature representations, solving the problems of information loss of small targets during the downsampling process and the difficulty of distinguishing small targets from complex backgrounds. Thus, by using feature fusion and feature enhancement methods, and with the help of the feature reconstruction method, the problems of insufficient feature information of existing small target detection algorithms, and the easy occlusion of small targets and the difficulty of distinguishing them from complex backgrounds are solved, effectively improving the performance of small target detection.
[0076] The method of the present embodiment can be applied to multiple fields. For example, in the field of monitoring and security, the detection of small targets in surveillance cameras is one of the key tasks. In surveillance videos, it is necessary to detect small targets in the distance, such as intruders, suspicious items, etc., to ensure public safety. Accurately detecting their categories and positions is crucial for the normal operation of the urban security and traffic management systems; in the fields of aerospace and military applications, small target detection is a key issue in devices such as unmanned aerial vehicles (UAVs) and aircraft. For example, in military reconnaissance, it is necessary to detect hidden enemy targets, such as small targets like soldiers, vehicles, weapons, etc. Accurately detecting these targets plays an important role in military decision-making and flight safety; in the medical field, small target detection can be used to detect tiny lesion areas, such as tumor cells, vascular abnormalities, etc. For example, detecting micro-nodules in the early screening of lung cancer; detecting retinal lesions in the diagnosis of ophthalmic diseases. Accurately identifying tiny abnormal structures in medical images is crucial for the early diagnosis and treatment of diseases; in the field of autonomous driving, small target detection can be used to detect small targets such as traffic signs, pedestrians, bicycles, etc. in the distance to ensure driving safety and compliance with traffic rules. The detection accuracy of small targets directly affects the response time and safety of the system; in the field of industrial inspection, small target detection can be used to detect defects or small components in products to ensure product quality.
[0077] Second Embodiment
[0078] This embodiment provides an electronic device. As Figure 4 shown, the electronic device includes: a processor and a memory; wherein, the processor and the memory can be connected through a communication bus; at least one instruction is stored in the memory, and the instruction is loaded and executed by the processor to implement the method of the above first embodiment. In addition, the electronic device may further include a transceiver, and the processor and the transceiver can be connected through a communication bus, and the transceiver is used for communicating with other devices.
[0079] Next, in combination with Figure 4 each component of the electronic device will be specifically introduced:
[0080] Among them, the processor is the control center of the electronic device. The electronic device may include multiple processors, and each of these processors may be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). Here, the processor may be a single processor or a collective term for multiple processing elements. For example, the processor is one or more central processing units (central processing unit, CPU), or may be other general-purpose processors, application specific integrated circuits (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention. For example: one or more digital signal processors (DSP), or one or more field programmable gate arrays (field programmable gate array, FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The processor can execute various functions of the electronic device by running or executing software programs stored in the memory and calling data stored in the memory.
[0081] In a specific implementation, as an embodiment, the processor may include one or more CPUs, such as Figure 4 CPU0 and CPU1 shown in
[0082] The memory is used to store the software program for implementing the solution of the present invention and is controlled by the processor for execution. The specific implementation manner can refer to the above method embodiment and will not be elaborated here.
[0083] Optionally, the memory may be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, or may also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory may be integrated with the processor or exist independently and be coupled to the processor through the interface circuit ( Figure 4 not shown) of the electronic device. The embodiments of the present invention do not make specific limitations thereto.
[0084] The transceiver may include a receiver and a transmitter ( Figure 4 not shown separately). Among them, the receiver is used to implement the receiving function, and the transmitter is used to implement the transmitting function. The transceiver may be integrated with the processor or exist independently and be coupled to the processor through the interface circuit ( Figure 4 not shown) of the electronic device. The embodiments of the present invention do not make specific limitations thereto.
[0085] In addition, it should be noted that Figure 4 the structure of the electronic device shown in does not constitute a limitation to the device. The actual device may include more or fewer components than shown in the figure, or combine some components, or have a different component layout. In addition, the technical effects achieved by the electronic device when executing the method of the first embodiment above may refer to the technical effects described in the first embodiment above, so they will not be repeated here.
[0086] Third Embodiment
[0087] This embodiment provides a computer-readable storage medium storing at least one instruction, which is loaded and executed by a processor to implement the method of the first embodiment above. Among them, the computer-readable storage medium may be a ROM, a random access memory, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc. The instructions stored therein can be loaded and executed by the processor in the terminal to implement the above method.
[0088] In addition, it should be noted that the present invention can be provided as a method, apparatus, or computer program product. Therefore, the embodiments of the present invention can take the form of all or part of a hardware embodiment, all or part of a software embodiment, or an embodiment combining software and hardware aspects. Moreover, when implemented in software, the embodiments of the present invention can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center containing one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0089] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing terminal device generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0090] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1the functions specified in one or more boxes. These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in one or more processes and / or boxes Figure 1 one process or more processes and / or boxes Figure 1 the steps of the functions specified in one box or more boxes.
[0091] It should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising the element. In addition, the term "and / or" is only a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. Among them, A and B can be singular or plural. In addition, the character " / " in this article generally means that the objects before and after are in an "or" relationship, but it may also mean an "and / or" relationship, which can be understood specifically with reference to the context. "At least one" means one or more, and "a plurality" means two or more. "At least one of the following (items)" or similar expressions refer to any combination of these items, including any combination of single (item) or plural items (items). For example, at least one of a, b or c can mean: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, c can be single or multiple.
[0092] In addition, it can be understood that in various embodiments of the present invention, the magnitudes of the serial numbers of the above processes do not mean the sequence of execution. The execution sequence of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0093] Those of ordinary skill in the art will realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.
[0094] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of functional modules / units is only a logical functional division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be electrical, mechanical, or other forms. The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in each embodiment of the present invention, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0095] If the method is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The foregoing storage medium includes various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0096] Finally, it should be noted that the above are only the preferred embodiments of the present invention. It should be pointed out that although the preferred embodiments of the present invention have been described, for those of ordinary skill in the art, once the basic creative concept of the present invention is known, several improvements and refinements can be made without departing from the principle described in the present invention. These improvements and refinements should also be regarded as the protection scope of the present invention. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present invention.
Claims
1. A small target detection method based on multi-scale feature fusion and feature enhancement, characterized in that: include: Construct a small target detection model based on multi-scale feature fusion and feature enhancement; Train the constructed small target detection model; Small target detection is achieved using the trained small target detection model.
2. The small target detection method based on multi-scale feature fusion and feature enhancement as claimed in claim 1, characterized in that: The small target detection model includes: an input module, a backbone network, a feature fusion module, a feature enhancement module, a feature reconstruction module and a detector; The feature reconstruction module is only used in the model training phase and is discarded in the inference phase; The process of implementing small target detection by the small target detection model includes: The image to be detected is input into the backbone network through the input module, and is processed by the backbone network to obtain a multi-scale feature map. Then, the multi-scale feature map is sent to the feature fusion module. In the feature fusion module, the multi-scale feature map is first feature optimized, and then the optimized multi-scale feature map is feature fused from top to bottom to obtain a feature map after feature fusion. Subsequently, the feature map after feature fusion is sent to the feature enhancement module for processing. The feature representation of the small target is enhanced by capturing the local detail information of the small target and the global information captured by the interaction between pixels, to obtain a feature map after feature enhancement. Finally, the feature map after feature enhancement is sent to the detector to achieve target classification and positioning.
3. The small target detection method based on multi-scale feature fusion and feature enhancement as claimed in claim 2, characterized in that: The feature fusion module includes a channel attention module and a spatial attention module; The feature fusion module first optimizes the multi-scale feature map output by the backbone network through the channel attention module to highlight the important areas in the feature maps of different scales, and then performs feature fusion on the optimized multi-scale feature map through the spatial attention module to obtain the feature map after feature fusion.
4. The small target detection method based on multi-scale feature fusion and feature enhancement as claimed in claim 3, characterized in that: The multi-scale feature map output by the backbone network is optimized by the channel attention module to highlight the important areas in the feature maps of different scales, including: The channel attention module first performs channel processing on the multi-scale feature map output by the backbone network through adaptive maximum pooling, then fuses the pooled feature vectors, and generates the final channel attention weight through the Sigmoid function. Finally, the original feature map is multiplied by the channel attention weight to enhance the features useful for small target detection and suppress irrelevant features. Finally, the features obtained by multiplying the original feature map with the channel attention weight are point convolved so that the number of feature maps at each scale is 256, and the optimized multi-scale feature map is obtained to highlight the important areas in the feature maps of different scales.
5. The small target detection method based on multi-scale feature fusion and feature enhancement as claimed in claim 3, characterized in that: The step of performing feature fusion on the optimized multi-scale feature map through the spatial attention module to obtain a feature map after feature fusion includes: The low-level feature map and the high-level feature map in the feature map optimized by the channel attention module are expanded into sequences respectively, the sequence corresponding to the high-level feature map is used as the query, and the sequence corresponding to the low-level feature map is used as the key and value. Self-Attention is performed between the high-level feature map and the low-level feature map, and the Self-Attention result is converted into a two-dimensional image. Then, the dimension of the two-dimensional image is matched with the low-level feature map through upsampling. Finally, the low-level feature map is fused with the upsampled two-dimensional image to obtain a feature map after feature fusion.
6. The small target detection method based on multi-scale feature fusion and feature enhancement as claimed in claim 2, characterized in that: The feature enhancement module includes a global information extraction module, a local information extraction module and an inverted residual structure; The global information extraction module captures global information by perceiving and interacting between pixels; the local information extraction module completes local information extraction through parallel calculation of three ordinary convolutions and three extended convolutions; wherein the expansion rates of the three extended convolutions are 2, 4, and 2, respectively; the features obtained by fusing the local information and the global information are sent to the inverted residual structure to obtain a feature map after feature enhancement.
7. The small target detection method based on multi-scale feature fusion and feature enhancement as claimed in claim 6, characterized in that: The global information extraction module first expands the two-dimensional image in pixels to make it a spatially continuous sequence, then performs an average operation on the channel, and then uses FFN to perform linear calculations and use the Softmax activation function to obtain the probability distribution of the feature in its spatial dimension, obtain the weights between pixels on the feature map, and finally reshape it into a two-dimensional image.
8. The small target detection method based on multi-scale feature fusion and feature enhancement as claimed in claim 2, characterized in that: The feature reconstruction module realizes feature reconstruction based on a mask autoencoder.
Citation Information
Patent Citations
Multi-scale attention-fused traffic helmet small target detection system and method
CN116665156A
Small target detection system and method based on improved YOLOv5
CN117523267A
Non-reference frame image target detection method based on feature enhancement and multiple scales
CN117893868A
Remote sensing small target detection method based on multi-scale feature fusion and context information enhancement
CN117994658A
SAR small target detection method based on super-resolution pyramid network and sidelobe suppression
CN118366048A