A method and device for detecting small targets in complex environments based on images from a drone
Patent Information
- Application Number
- CN202411995372.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2044-12-31
AI Technical Summary
此类小目标检测方法大多基于单尺度或固定特征提取策略,无法充分应对目标的尺度变化和复杂背景问题
[0035]采用本发明能够基于无人机航拍图像实现复杂环境下的更准确的小目标检测。
Smart Images

Figure CN120032212B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of target detection and artificial intelligence technology, specifically, it relates to a method and apparatus for detecting small targets in complex environments based on UAV images. Background Technology
[0002] With the increasing application of drones in various fields, especially in scenarios such as military reconnaissance, traffic monitoring, and search and rescue operations, target detection technology based on drone platforms is developing rapidly.
[0003] However, due to the constantly changing altitude and angle required during drone aerial photography, the proportion of target objects in the acquired images often changes accordingly. Furthermore, when detecting small targets, the original small target has relatively few pixels in the image, and the image background can also cause varying degrees of interference during detection. These factors make accurately and effectively acquiring the features of small targets particularly difficult.
[0004] Currently, deep learning network methods used for detecting drone aerial images mainly fall into the following two categories:
[0005] A two-stage detection method based on candidate regions first generates corresponding candidate regions using drone aerial images, then extracts target features contained within the candidate regions, and finally classifies and regresses the targets based on the extracted feature information. This type of method has the advantage of superior detection accuracy and precision, but is relatively slow.
[0006] Single-level detection methods based on regression treat target detection in UAV aerial images as a regression problem. Instead of generating corresponding region proposals, they directly extract features and perform classification regression on the entire image, achieving end-to-end target detection. Representative networks include SSD, YOLO, and RetinaNet. Most of these small target detection methods rely on single-scale or fixed feature extraction strategies, which cannot adequately handle target scale variations and complex backgrounds. For example, while the YOLO series of algorithms can perform small target detection to some extent, their sensitivity to target scale is insufficient, especially in complex scenes where small targets are easily misidentified or missed. Summary of the Invention
[0007] To address the aforementioned problems, this invention discloses a method and apparatus for detecting small targets based on UAV images with complex backgrounds, thereby improving the accuracy of small target detection in complex environments.
[0008] A first aspect of the present invention provides a method for detecting small targets in complex environments based on UAV images, comprising:
[0009] Acquire images captured by the drone;
[0010] An adaptive feature extraction network is used to extract features from the image. The feature extraction network first extracts features through two extraction branch networks, and then adds the features extracted by the two extraction branch networks and outputs them. The first extraction branch network consists of a CBS module, a deformable convolutional module, and a coordinate attention module connected in sequence, and the second extraction branch network consists of a CBS module, a coordinate attention module, and a deformable convolutional module connected in sequence.
[0011] A multi-scale feature fusion network is used to fuse features extracted from different layers of the adaptive feature extraction network;
[0012] Target detection is performed using a target detection network; wherein, the target detection network includes multiple detection heads, one of which performs smaller target detection based on the fusion features of the high-resolution features of the shallow layer of the adaptive feature extraction network and the high-resolution features of the shallow layer of the multi-scale feature fusion network.
[0013] In other examples of the above method, the step of adding the features extracted by the two extraction branch networks and then outputting them includes: scaling the feature scales extracted by the two extraction branch networks to the same size, linearly adding the feature matrices, and then outputting them through the CBS module.
[0014] In other examples of the above method, the multi-scale feature fusion network includes an FPN network and a feature fusion network; the FPN network outputs three features at different scales; the feature fusion network includes three fusion branches, each of which is used to fuse features at one scale and their convolution results.
[0015] In other examples of the above method, the step of fusing features extracted from different layers of the adaptive feature extraction network using a multi-scale feature fusion network further includes: fusing the features output by the three fusion branches pairwise: F(i,j) = Resize[F(i)]*F(j)
[0016] Where F(i) and F(j) represent the features output by different fusion branches, Resize is used to adjust features at different scales to the same scale, and * indicates cross-correlation.
[0017] In other examples of the above method, the plurality of detection heads further includes:
[0018] The small target detection head receives high-resolution features from the shallow layer of the multi-scale feature fusion network to perform small target detection.
[0019] A medium-resolution target detection head receives medium-resolution features from the intermediate layer of the multi-scale feature fusion network for medium-resolution target detection.
[0020] A large target detection head receives low-resolution features from the deep layers of the multi-scale feature fusion network for large target detection.
[0021] A second aspect of the present invention provides a device for detecting small targets in complex environments based on UAV images, comprising:
[0022] The input module is configured to acquire images captured by the drone;
[0023] The feature extraction module is configured to extract features from the image using an adaptive feature extraction network. The feature extraction network first extracts features through two extraction branch networks, and then adds the features extracted by the two extraction branch networks together and outputs the sum. The first extraction branch network consists of a CBS module, a deformable convolution module, and a coordinate attention module connected in sequence, and the second extraction branch network consists of a CBS module, a coordinate attention module, and a deformable convolution module connected in sequence.
[0024] The feature fusion module is configured to fuse features extracted from different layers of the adaptive feature extraction network using a multi-scale feature fusion network.
[0025] The target detection module is configured to perform target detection using a target detection network; wherein the target detection network includes multiple detection heads, one of which performs smaller target detection based on the fusion features of the high-resolution features of the shallow layer of the adaptive feature extraction network and the high-resolution features of the shallow layer of the multi-scale feature fusion network.
[0026] In other examples of the above-mentioned device, the step of adding the features extracted by the two extraction branch networks and then outputting them includes: scaling the feature scales extracted by the two extraction branch networks to the same size, linearly adding the feature matrices, and then outputting them through the CBS module.
[0027] In other examples of the above-described device, the multi-scale feature fusion network includes an FPN network and a feature fusion network; the FPN network outputs three features at different scales; the feature fusion network includes three fusion branches, each of which is used to fuse features at one scale and their convolution results.
[0028] In other examples of the aforementioned device, the step of fusing the extracted features using a multi-scale feature fusion network to output features of different dimensions further includes: fusing the features output by the three fusion branches pairwise.
[0029] F(i, j) = Resize[F(i)] * F(j)
[0030] Where F(i) and F(j) represent the features output by different fusion branches, Resize is used to adjust features at different scales to the same scale, and * indicates cross-correlation.
[0031] In other examples of the aforementioned device, the plurality of detection heads further include:
[0032] The small target detection head receives high-resolution features from the shallow layer of the multi-scale feature fusion network to perform small target detection.
[0033] A medium-resolution target detection head receives medium-resolution features from the intermediate layer of the multi-scale feature fusion network for medium-resolution target detection.
[0034] A large target detection head receives low-resolution features from the deep layers of the multi-scale feature fusion network for large target detection.
[0035] The present invention enables more accurate small target detection in complex environments based on UAV aerial images. Attached Figure Description
[0036] Other features, objects, and advantages of the invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0037] Figure 1 This is a schematic diagram of an adaptive feature extraction network according to an embodiment of the present invention;
[0038] Figure 2 This is a schematic diagram of a multi-scale feature fusion network according to an embodiment of the present invention;
[0039] Figure 3 This is a schematic diagram of a target detection network according to an embodiment of the present invention;
[0040] Figure 4 This is a schematic diagram of the process for detecting small targets in complex environments based on UAV images according to an embodiment of the present invention;
[0041] Figure 5 This is a schematic diagram of a small target detection device in a complex environment based on UAV images according to an embodiment of the present invention. Detailed Implementation
[0042] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings.
[0043] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0044] When performing target detection tasks on UAV-based platforms, the complexity of the background, dynamic changes in the surrounding environment, lighting and weather conditions, and terrain undulations can severely affect the effectiveness and accuracy of target detection. Furthermore, changes in perspective caused by UAV rotation, tilting, and side-flying during flight alter the appearance and attitude of the target, resulting in significant variations in target size within the image. Small targets may blend into the background, while large targets may not appear completely within the field of view. To address these challenges, target detection algorithms need to possess strong robustness and good generalization ability, ensuring reliable operation even in highly variable environments.
[0045] According to an embodiment of the present invention, a target detection model is provided. This model, after training, can detect small targets in complex environments based on input UAV aerial images. The target detection model comprises three parts: an adaptive feature extraction network, a multi-scale feature fusion network, and a target detection network, which are described in detail below.
[0046] (I) Adaptive Feature Extraction Network
[0047] The system acquires complex environmental images captured by drones and performs feature extraction on these images to obtain features (feature maps) at different scales. The adaptive feature extraction network consists of a deformable convolution (DC) module and a coordinate attention (CA) module. It aims to dynamically adjust the receptive field to help the network focus on key features of small targets and eliminate redundant background interference.
[0048] Traditional convolutional operations sample a fixed local region, while deformable convolutions, by introducing learnable offsets, allow for dynamic adjustment of the sampling position. This enables the convolutional kernel to adapt to geometric transformations of the input features, thus selectively focusing on important feature regions. For small targets or irregularly shaped objects, convolutions with fixed receptive fields may not effectively capture their important features, while deformable convolutions adaptively adjust their receptive fields based on the input data, enhancing the network's ability to detect small targets.
[0049] Coordinate attention achieves efficient encoding of location information by decomposing spatial and channel information. It applies coordinated attention weights to the input feature map, enabling the network to better identify spatial structures and long-range dependencies. Through this attention mechanism, the network can highlight target regions in the feature map while suppressing redundant background information. This allows the network to focus more on key regions when detecting small targets, improving the resolution of local features.
[0050] This invention combines DC and CA. The feature extraction network can not only flexibly adjust the receptive field through DC to focus on important local features, but also further separate important and unimportant information in the global scope through CA. This reduces the interference of background noise, enabling the network to effectively identify small targets in complex scenes and improve the overall detection accuracy and efficiency.
[0051] like Figure 1 As shown, according to an embodiment of the present invention, a deformable convolutional module and a coordinate attention module are cross-fused. The input is a preprocessed image of a complex environment captured by a UAV. In this embodiment, the feature extraction network includes two branches: the first branch consists of a CBS module, a DC module, and a CA module connected in sequence; the second branch consists of a CBS module, a CA module, and a DC module connected in sequence. The CBS module consists of a Conv layer (convolutional layer), a BN layer (batch normalization layer), and a Silu layer (an activation function), a module commonly used in YOLO series object detection algorithms.
[0052] As can be seen, the difference between these two branches is that the deformable convolution module and the coordinate attention module are in different orders. This combination avoids feature loss during the branch processing, helps to extract more accurate feature information, and reduces interference from complex backgrounds.
[0053] Specifically, the different module order in the two branches provides diverse feature extraction capabilities. The first branch, using DC followed by CA, focuses on local geometric features earlier, then uses CA to concentrate on important regions. The second branch, using CA followed by DC, initially filters information across the entire feature map through an attention mechanism, then uses DC to achieve more precise feature alignment. This diversity helps the network capture different types of features from the input data, enhancing the model's generalization ability.
[0054] Furthermore, the design of applying DC followed by CA in the first branch allows the DC module to flexibly adjust its receptive field first, and then the CA module to strengthen the importance of local features. This order helps to build the global context from the details. The second branch, applying CA before DC, can identify globally important regions in advance through the attention mechanism. Subsequent DC can adjust its receptive field more accurately, ensuring in-depth analysis of the detailed structure of these regions.
[0055] Furthermore, the different processing paths of the two branches allow the model to integrate information through multi-path fusion, thereby reducing dependence on individual path anomalies. This can effectively cope with complex and varied input data, and is particularly robust when dealing with tasks that simultaneously contain local small targets and global contextual information. In particular, this design can resolve the contradiction between complex environments and the limited computing resources of UAVs, providing deep models with richer representation capabilities and adaptability.
[0056] Then, the feature scales extracted from these two branches are scaled to the same size, and the feature matrices are linearly added directly. Finally, the output is processed by the CBS module.
[0057] This invention utilizes an adaptive feature extraction network to automatically adjust convolutional operations to adapt to small targets of different sizes and shapes, achieving dynamic adjustment of the receptive field and adaptive allocation of spatial attention. This enhances the ability to focus on small targets and suppress background noise, thereby improving the accuracy and efficiency of detection and tracking.
[0058] (II) Multi-scale feature fusion network
[0059] According to an embodiment of the present invention, the following is employed: Figure 2 The multi-scale feature fusion network shown.
[0060] like Figure 2 As shown, the left half of the multi-scale feature fusion network uses a feature pyramid network (FPN).
[0061] FPN provides the detector with rich, semantically rich multi-scale features by constructing a top-down feature pyramid. The core idea of this architecture is to utilize the information from feature maps at different levels in a convolutional neural network (CNN) to form a feature pyramid from high to low levels, with each layer responsible for capturing targets at a specific scale.
[0062] FPN is typically based on classic CNN architectures (such as ResNet, VGG, etc.) as the backbone network. The backbone network extracts preliminary features from the input image and outputs feature maps at multiple levels.
[0063] In the backbone network, the higher the level, the lower the spatial resolution of the feature maps, but the richer their semantic information. FPN takes advantage of this characteristic and upsamples these high-level feature maps step by step (usually using bilinear interpolation) from top to bottom.
[0064] Furthermore, each cross-layer connection allows the upsampling results of higher-level feature maps to be added element-wise with the lower-level feature maps. These lateral connections help combine detailed low-level texture information with high-level semantic information, enhancing the expressive power of features at different scales.
[0065] By combining top-down paths and lateral connections, a set of feature maps with the same number of channels but different resolutions is generated. These feature maps together form a feature pyramid, with each map used to detect targets at a specific scale. For example, the lower-resolution top-level feature map is used to detect large targets, while the higher-resolution bottom-level feature map is used to detect small targets.
[0066] Therefore, FPN enhances the detection capability for targets of different sizes by combining feature maps of different resolutions, thus achieving efficient multi-scale detection. Vertical connections allow semantic information to be shared between feature maps, thereby improving the quality of high-level semantic information in the lower-level feature maps. Furthermore, compared to other methods that process different scales through multiple branches and multiple input paths, FPN is more lightweight and computationally efficient, significantly improving the model's detection performance and accuracy in complex scenes.
[0067] The FPN output represents three features from different layers. The right half of the multi-scale feature fusion network designs a feature fusion network, which includes three fusion branches. Each branch fuses features at one scale and their convolution results. During feature extraction, shallow and deep features are adaptively weighted and fused to ensure the network effectively captures details of small targets. The specific method is as follows:
[0068] First, the three fusion branches perform 1x1 convolutions on the three features respectively, and then fuse them with the original features (features output by FPN) to output the fused features. This process fuses the features between each channel of each feature map, which can be represented as:
[0069] F′1=Add(Conv 1×1 (F1), F1)
[0070] Then, the features output from the three fusion branches are fused pairwise. The fusion method is as follows:
[0071] F(i, j) = Resize[F(i)] * F(j)
[0072] The Resize method adjusts features at different scales to the same scale, resulting in three sets of cross-correlation feature maps F(1,2), F(1,3), and F(2,3).
[0073] In tasks like object detection that require multi-scale feature fusion, the resize operation is used to adjust feature maps from different scales to the same size to ensure spatial alignment during subsequent feature fusion (such as concat and add operations). Resize methods can be implemented using bilinear interpolation, bicubic interpolation, nearest neighbor interpolation, and pooling operations. The appropriate choice of resize method can significantly improve performance. e This method can effectively align the sizes of feature maps from different sources, laying the foundation for feature fusion or analysis in subsequent layers and ensuring that feature maps are fully utilized in deep learning models.
[0074] (III) Target Detection Network
[0075] The feature fusion process has yielded features at different scales, and the target detection network performs target detection based on these features at different scales.
[0076] In the detection scheme of the YOLO V5 network, three detection heads (Headl-3) are used to detect target information at different scales:
[0077] Head3, the small target detection head, receives high-resolution features from earlier layers of the network. Although the semantic information may not be as rich as that of deeper features, these features retain more detail and are used to detect small targets in the image, such as distant or small objects. In this branch, due to the high spatial resolution of the feature maps, the convolutional processing directly manages the small target detection operation.
[0078] The target detection head, Head2, receives feature maps from the intermediate layers. These features incorporate contextual information from the pyramid and are used to locate medium-sized targets in the image. Typically, the feature maps have medium spatial resolution, making them suitable for processing targets that are neither too small nor too large.
[0079] The large target detection head, Head1, receives feature maps from deeper layers of the network. These feature maps have low resolution but contain rich semantic information, and are used to process large targets in images. The lower spatial resolution allows this detection head to efficiently process large-scale targets with lower requirements for low-level details.
[0080] The three detection heads use different feature map resolutions to focus on target objects of different sizes. Their outputs include bounding box regression information, class probability, and confidence. Finally, they all undergo combination and post-processing steps (such as non-maximum suppression) to obtain the final detection results.
[0081] Although this multi-scale detection head design improves YOLO v5's performance in handling multi-scale target scenes, enabling accurate identification and localization of both distant and near targets, the gradual weakening of gradient feature information of small targets in UAV aerial images occurs as network depth increases. To prevent the loss of feature information from earlier layers and to enable the network to better detect the feature information of small targets in the image, according to embodiments of the present invention, such as... Figure 3 As shown, a smaller target head detection head, Head4, is added to the YOLO v5 network head detection.
[0082] exist Figure 3 In the target detection model shown, the feature extraction network on the left extracts features layer by layer. As sampling increases, the feature map resolution gradually decreases, while high-dimensional information gradually increases. The feature fusion network in the middle fuses the feature information and divides it into three dimensions: low, medium, and high. The rightmost detection head detects features of different dimensions, thus achieving multi-scale target detection. This invention extracts features from the original high-resolution features of the feature extraction network. These features contain rich information about small targets and are fused with the low-dimensional features in the feature fusion network. The fused features are then input into the added smaller target detection head, Head4, i.e., the fourth detection head, for target detection. This approach fully utilizes the feature information obtained from the previous layers and further improves the network's ability to perceive multiple small targets, enabling the network to more accurately detect small targets in images.
[0083] Based on the above target detection model, according to one embodiment of the present invention, a method for detecting small targets in complex environments based on UAV images is provided, such as... Figure 4 As shown, the method includes the following steps:
[0084] A first aspect of the present invention provides a method for detecting small targets in complex environments based on UAV images, comprising:
[0085] S401, acquire images gathered by the drone;
[0086] S402 employs an adaptive feature extraction network to extract features from the image. The feature extraction network first extracts features through two extraction branch networks, and then adds the features extracted by the two extraction branch networks together and outputs the result. The first extraction branch network consists of a CBS module, a deformable convolutional module, and a coordinate attention module connected in sequence, and the second extraction branch network consists of a CBS module, a coordinate attention module, and a deformable convolutional module connected in sequence.
[0087] S403 employs a multi-scale feature fusion network to fuse features extracted from different layers of the adaptive feature extraction network;
[0088] S404 employs an object detection network for object detection; wherein, the object detection network includes multiple detection heads, one of which performs smaller object detection based on the fusion features of the high-resolution features of the shallow layer of the adaptive feature extraction network and the high-resolution features of the shallow layer of the multi-scale feature fusion network.
[0089] In other examples of the above method, the features extracted by the two extraction branch networks are added together and then output, including: scaling the feature scales extracted by the two extraction branch networks to the same size, linearly adding the feature matrices and then outputting them through the CBS module.
[0090] In other examples of the above methods, the multi-scale feature fusion network includes an FPN network and a feature fusion network; the FPN network outputs three features at different scales; the feature fusion network includes three fusion branches, each of which is used to fuse features at one scale and their convolution results.
[0091] Other examples include: fusing the features output from the three fusion branches pairwise:
[0092] F(i, j) = Resize[F(i)] * F(j)
[0093] Where F(i) and F(j) represent the features output by different fusion branches, Resize is used to adjust features at different scales to the same scale, and * indicates cross-correlation.
[0094] In other examples of the above methods, multiple detection heads also include:
[0095] The small target detection head receives high-resolution features from the shallow layers of a multi-scale feature fusion network to perform small target detection.
[0096] The medium-resolution target detection head receives medium-resolution features from the intermediate layer of the multi-scale feature fusion network for medium-resolution target detection.
[0097] The large target detection head receives low-resolution features from deep layers of a multi-scale feature fusion network for large target detection.
[0098] According to another embodiment of the present invention, a device for detecting small targets in complex environments based on UAV images is provided, such as... Figure 5 As shown, the device includes:
[0099] A second aspect of the present invention provides a device for detecting small targets in complex environments based on UAV images, comprising:
[0100] Input module 501 is configured to acquire images collected by a drone;
[0101] The feature extraction module 502 is configured to use an adaptive feature extraction network to extract features from the image. The feature extraction network first extracts features through two extraction branch networks, and then adds the features extracted by the two extraction branch networks and outputs them. The first extraction branch network consists of a CBS module, a deformable convolution module, and a coordinate attention module connected in sequence, and the second extraction branch network consists of a CBS module, a coordinate attention module, and a deformable convolution module connected in sequence.
[0102] The feature fusion module 503 is configured to fuse features extracted from different layers of the adaptive feature extraction network using a multi-scale feature fusion network.
[0103] The target detection module 504 is configured to perform target detection using a target detection network; wherein the target detection network includes multiple detection heads, one of which performs smaller target detection based on the fusion features of the high-resolution features of the shallow layer of the adaptive feature extraction network and the high-resolution features of the shallow layer of the multi-scale feature fusion network.
[0104] In other examples of the above-mentioned device, the output of the sum of the features extracted by the two extraction branch networks includes: scaling the feature scales extracted by the two extraction branch networks to the same size, linearly adding the feature matrices, and then outputting them through the CBS module.
[0105] In other examples of the aforementioned devices, the multi-scale feature fusion network includes an FPN network and a feature fusion network; the FPN network outputs three features at different scales; the feature fusion network includes three fusion branches, each of which is used to fuse features at one scale and their convolution results.
[0106] Other examples include: fusing the features output from the three fusion branches pairwise:
[0107] F(i, j) = Resize[F(i)] * F(j)
[0108] Where F(i) and F(j) represent the features output by different fusion branches, Resize is used to adjust features at different scales to the same scale, and * indicates cross-correlation.
[0109] In other examples of the aforementioned devices, the multiple detection heads also include:
[0110] The small target detection head receives high-resolution features from the shallow layers of a multi-scale feature fusion network to perform small target detection.
[0111] The medium-resolution target detection head receives medium-resolution features from the intermediate layer of the multi-scale feature fusion network for medium-resolution target detection.
[0112] The large target detection head receives low-resolution features from deep layers of a multi-scale feature fusion network for large target detection.
[0113] This invention improves upon the YOLO v5 network:
[0114] By replacing the original CSP Darknet network with an adaptive feature extraction network for feature extraction, the receptive field can be dynamically adjusted, thereby improving the network's performance in small object detection tasks.
[0115] In addition, the feature fusion network replaces the original PAN network with three fusion branches, which can effectively enhance the network's ability to perceive targets of different scales and reduce the occurrence of missed detections and false detections of small targets.
[0116] Furthermore, a smaller target detection head is added to the original network's three detection heads. Small target detection is performed based on the fusion result of the original high-resolution features extracted by the feature extraction network and the high-resolution features of the feature fusion network, thereby enabling the network to detect smaller targets in the image more accurately.
[0117] Therefore, this invention enables more accurate small target detection in complex environments based on UAV aerial images.
[0118] Although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments. Modifications or equivalent substitutions to the technical solutions of the embodiments of the present invention without departing from the inventive concept of the present invention should not depart from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for detecting small targets in complex environments based on UAV images, characterized in that, include: Acquire images captured by the drone; An adaptive feature extraction network is used to extract features from the image; The feature extraction network first extracts features through two extraction branch networks, and then adds the features extracted by the two extraction branch networks and outputs them. The first extraction branch network consists of a CBS module, a deformable convolution module, and a coordinate attention module connected in sequence, and the second extraction branch network consists of a CBS module, a coordinate attention module, and a deformable convolution module connected in sequence. A multi-scale feature fusion network is used to fuse features extracted from different layers of the adaptive feature extraction network. The multi-scale feature fusion network includes an FPN network and a feature fusion network. The FPN network outputs three features at different scales. The feature fusion network includes three fusion branches, each of which is used to fuse features at one scale and their convolution results. Target detection is performed using a target detection network; wherein, the target detection network includes multiple detection heads, one of which performs smaller target detection based on the fusion features of the high-resolution features of the shallow layer of the adaptive feature extraction network and the high-resolution features of the shallow layer of the multi-scale feature fusion network.
2. The method for detecting small targets in complex environments according to claim 1, characterized in that, The step of adding the features extracted by the two extraction branch networks and then outputting them includes: scaling the feature scales extracted by the two extraction branch networks to the same size, linearly adding the feature matrices, and then outputting them through the CBS module.
3. The method for detecting small targets in complex environments according to claim 1, characterized in that, The step of fusing features extracted from different layers of the adaptive feature extraction network using a multi-scale feature fusion network further includes: fusing the features output from the three fusion branches pairwise. Where F(i) and F(j) represent the features output by different fusion branches, and Resize is used to adjust features at different scales to the same scale. Indicates mutual correlation.
4. The method for detecting small targets in complex environments according to claim 1, characterized in that, The plurality of detection heads also include: The small target detection head receives high-resolution features from the shallow layer of the multi-scale feature fusion network to perform small target detection. A medium-resolution target detection head receives medium-resolution features from the intermediate layer of the multi-scale feature fusion network for medium-resolution target detection. A large target detection head receives low-resolution features from the deep layers of the multi-scale feature fusion network for large target detection.
5. A device for detecting small targets in complex environments based on UAV imagery, characterized in that, include: The input module is configured to acquire images captured by the drone; The feature extraction module is configured to extract features from the image using an adaptive feature extraction network; The feature extraction network first extracts features through two extraction branch networks, and then adds the features extracted by the two extraction branch networks and outputs them. The first extraction branch network consists of a CBS module, a deformable convolution module, and a coordinate attention module connected in sequence, and the second extraction branch network consists of a CBS module, a coordinate attention module, and a deformable convolution module connected in sequence. The feature fusion module is configured to use a multi-scale feature fusion network to fuse features extracted from different layers of the adaptive feature extraction network to output features of different dimensions; the multi-scale feature fusion network includes an FPN network and a feature fusion network; the FPN network outputs three features at different scales; the feature fusion network includes three fusion branches, each fusion branch being used to fuse features at one scale and their convolution results; The target detection module is configured to perform target detection using a target detection network; wherein the target detection network includes multiple detection heads, one of which performs smaller target detection based on the fusion features of the high-resolution features of the shallow layer of the adaptive feature extraction network and the high-resolution features of the shallow layer of the multi-scale feature fusion network.
6. The complex environment small target detection device according to claim 5, characterized in that, The step of adding the features extracted by the two extraction branch networks and then outputting them includes: scaling the feature scales extracted by the two extraction branch networks to the same size, linearly adding the feature matrices, and then outputting them through the CBS module.
7. The complex environment small target detection device according to claim 5, characterized in that, The step of fusing features extracted from different layers of the adaptive feature extraction network using a multi-scale feature fusion network further includes: fusing the features output from the three fusion branches pairwise. Where F(i) and F(j) represent the features output by different fusion branches, and Resize is used to adjust features at different scales to the same scale. Indicates mutual correlation.
8. The complex environment small target detection device according to claim 5, characterized in that, The plurality of detection heads also include: The small target detection head receives high-resolution features from the shallow layer of the multi-scale feature fusion network to perform small target detection. A medium-resolution target detection head receives medium-resolution features from the intermediate layer of the multi-scale feature fusion network for medium-resolution target detection. A large target detection head receives low-resolution features from the deep layers of the multi-scale feature fusion network for large target detection.
Citation Information
Patent Citations
Unmanned aerial vehicle image detection method based on multi-scale feature fusion and context enhancement
CN117037004A